git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [ANNOUNCE] Cogito-0.12

From
Junio C Hamano <junkio@cox.net>
Date
Jul 7, 2005, 17:21 UTC
Message-ID
<7vk6k2sfa4.fsf@assigned-by-dhcp.cox.net>
In-Reply-To
<20050707144501.GG19781@pasky.ji.cz>
>>>>> "PB" == Petr Baudis <pasky@suse.cz> writes:

PB> It won't happen. Or rather, I hope the HTTP pulls become more efficient PB> soon. Actually, perhaps Linus has something done already, my workstation PB> is a bit derailed now so I couldn't pull from him in the last few days PB> (hopefully will sort that out today).

PB> Hmm, yes, I guess Linus won't be touching the HTTP backend at all. ;-) I PB> suggest you to check the last development in Linus' branch and sync with PB> Daniel Barkalow, who promised improving the pull tools as well.

If this weekend is not too late, I have been brewing what is called an "efficient pull from dumb servers" suite, which would hopefully fill this gap. I am still in the process of finishing the details, but basically it already seems to work.

Linus, please drop the patch I sent you earlier, privately by mistake not CCing the list, that implemented only the server end. I've changed some file formats already from that one.

The outline of how it works is like this.
 * I assume a dumb transport (read: static files only HTTP
   server) and no on-request server side processing.  All the
   smarts must go in the client.  The server side X.git being an
   ordinary GIT archive (no need for files in the work tree),
   plus:
   - X.git/objects/pack can have packed GIT archives.  I
     envision that this will be a series of 5 to 20 MB packs,
     occasionally adding a new incremental pack when
     X.git/objects/??/ directories accumulate enough standalone
     SHA1 files.  It is not necessary to have X.git/objects/??/
     files if an object is contained in one of the packs.
   - X.git/info/ has three extra files.
     - "inventory" lists all the branches stored in X.git/refs
       and looks like this (contents and path):
          ff83c8f3554ceb444b413beaeb49b4a781dae944 snap/0
          013e7c7ff498aae82d799f80da37fbd395545456 snap/10
          ff83c8f3554ceb444b413beaeb49b4a781dae944 heads/master
          dd7ba8b4949535c24e604a37709db0e3be9ccbbc heads/linus
       This is to facilitate discovery from a transport that is
       not so "ls" friendly, like HTTP.
     - "pack" lists available packs under X.git/objects/pack and
       looks like this (size and name):
          432495 pk-65fe69e9bc2e8a3e0881e008dde182522156ba7c.pack
       The file is there for discovery.  The size is used by the
       client to discover optimum set of packs to slurp.
     - "rev-cache" is a binary file that describes commit
       ancestry information in a dense format.  It lists all
       commits available from this repository along with who
       its parents are for each of the commit.  This file is
       produced append-only, so that the server side can use
       rsync based mirroring scheme.
   A new command "git-update-dumb-server" is used to prepare
   these three files.  There may need a helper script that uses
   git-pack-objects and friends to prepare packs partitioned to
   allow pulling a popular branch efficiently.
 * The client side is called "git-dumb-pull-script".  This
   downloads the above three files, and .idx files associated
   with packs described in "pack".  With the information in
   "inventory" about desired branch to pull from along with
   "rev-cache" ancestry information, it discovers the set of
   commits that is lacking from its local store.  By comparing
   that list with downloaded .idx files, along with size
   information for each pack, it comes up a list of packs to
   download to cover the most commits that it wants to obtain,
   and downloads them, verifies them and stores them in its
   .git/objects/pack/ directory.
   The above process of downloading packs would typically not
   cover all the things lacking, because some new commits may
   not be in any of the packs.  After this point, the usual
   commit-walking git-http-pull can be used to fill the rest,
   and it does not have to pull that many objects.  Dan's
   http-pull parallelism improvement would be very useful
   independently here.
Previous: Petr BaudisNext: Linus Torvalds
Message 4 of 66 in “[ANNOUNCE] Cogito-0.12”
  1. Petr BaudisJul 3, 2005
  2. Brian GerstJul 6, 2005
  3. Petr BaudisJul 7, 2005
  4. Junio C HamanoJul 7, 2005
  5. Linus TorvaldsJul 7, 2005
  6. Junio C HamanoJul 7, 2005
  7. Linus TorvaldsJul 7, 2005
  8. Junio C HamanoJul 7, 2005
  9. Junio C HamanoJul 7, 2005
  10. Eric W. BiedermanJul 7, 2005
  11. Linus TorvaldsJul 7, 2005
  12. Eric W. BiedermanJul 8, 2005
  13. Dumb servers (was: [ANNOUNCE] Cogito-0.12)Kevin Smith, Jul 8, 2005
  14. Linus TorvaldsJul 8, 2005
  15. Petr BaudisJul 7, 2005
  16. Linus TorvaldsJul 7, 2005
  17. Pull efficiently from a dumb git store.Junio C Hamano, Jul 7, 2005
  18. rev-list: add "--objects=self-sufficient" flag.Junio C Hamano, Jul 7, 2005
  19. Linus TorvaldsJul 7, 2005
  20. rev-list: add "--full-objects" flag.Junio C Hamano, Jul 8, 2005
  21. Linus TorvaldsJul 8, 2005
  22. Linus TorvaldsJul 8, 2005
  23. Junio C HamanoJul 8, 2005
  24. Linus TorvaldsJul 8, 2005
  25. Eric W. BiedermanJul 9, 2005
  26. Linus TorvaldsJul 10, 2005
  27. Junio C HamanoJul 10, 2005
  28. Sven VerdoolaegeJul 10, 2005
  29. Linus TorvaldsJul 10, 2005
  30. Eric W. BiedermanJul 11, 2005
  31. Linus TorvaldsJul 11, 2005
  32. Eric W. BiedermanJul 12, 2005
  33. Linus TorvaldsJul 12, 2005
  34. Eric W. BiedermanJul 12, 2005
  35. Linus TorvaldsJul 12, 2005
  36. Eric W. BiedermanJul 12, 2005
  37. Linus TorvaldsJul 12, 2005
  38. Linus TorvaldsJul 11, 2005
  39. Give --full-objects flag to rev-list when preparing a dumb server.Junio C Hamano, Jul 8, 2005
  40. Use --objects=self-sufficient flag to rev-list.Junio C Hamano, Jul 7, 2005
  41. Tony LuckJul 7, 2005
  42. Junio C HamanoJul 7, 2005
  43. Linus TorvaldsJul 7, 2005
  44. Tony LuckJul 8, 2005
  45. Linus TorvaldsJul 8, 2005
  46. Russell KingJul 9, 2005
  47. Russell KingJul 9, 2005
  48. Junio C HamanoJul 9, 2005
  49. Linus TorvaldsJul 10, 2005
  50. Linus TorvaldsJul 10, 2005
  51. Russell KingJul 10, 2005
  52. Junio C HamanoJul 10, 2005
  53. Russell KingJul 10, 2005
  54. Linus TorvaldsJul 10, 2005
  55. Russell KingJul 10, 2005
  56. Linus TorvaldsJul 10, 2005
  57. Russell KingJul 10, 2005
  58. Linus TorvaldsJul 10, 2005
  59. Russell KingJul 10, 2005
  60. Petr BaudisJul 10, 2005
  61. Chris WrightJul 11, 2005
  62. Linus TorvaldsJul 8, 2005
  63. Petr BaudisJul 8, 2005
  64. Daniel BarkalowJul 8, 2005
  65. Chris WrightJul 7, 2005
  66. Check packs and then files.Junio C Hamano, Jul 11, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.