git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC] Add --create-cache to repack

From
Nicolas Pitre <nico@fluxnic.net>
Date
Jan 28, 2011, 18:46 UTC
Message-ID
<alpine.LFD.2.00.1101281304270.8580@xanadu.home>
In-Reply-To
<AANLkTim+AUY9SdeAFfkny2_a3qQ9SCDLUHR3s9Q3M98u@mail.gmail.com>
On Fri, 28 Jan 2011, Shawn Pearce wrote:
Show 18 quoted lines
> This started because I was looking for a way to speed up clones coming
> from a JGit server.  Cloning the linux-2.6 repository is painful, it
> takes a long time to enumerate the 1.8 million objects.  So I tried
> adding a cached list of objects reachable from a given commit, which
> speeds up the enumeration phase, but JGit still needs to allocate all
> of the working set to track those objects, then go find them in packs
> and slice out each compressed form and reformat the headers on the
> wire.  Its a lot of redundant work when your kernel repository has
> 360MB of data that you know a client needs if they have asked for your
> master branch with no "have" set.
> 
> Later I realized, we can get rid of that cached list of objects and
> just use the pack itself.  Its far cleaner, as there is no redundant
> cache.  But either way (object list or pack) its a bit of a challenge
> to automatically identify the right starting points to use.  Linus
> Torvalds' linux-2.6 repository is the perfect case for the RFC I
> posted, its one branch with all of the history, and it never rewinds.
> But maybe Linus is just very unique in this world.  :-)

Playing my old record again... I know. But pack v4 should solve a big part of this enumeration cost.

I've changed the format slightly again in my WIP branch.  The idea is to:
1) Have a non compressed yet still really dense representation for tree 
   objects;
2) do the same thing for the first part of commit objects, and only 
   deflate the free form text part.
There is nothing new here.  However, it should be possible to:
3) replace all SHA1 references by an offset into the pack file directly, 
   just like we do for OFS_DELTA objects.  If the SHA1 is actually 
   needed then we can obtain it with a reverse lookup with given object offset 
   in the pack index file, but in practice that is not actually required that 
   often.

So walking the history graph and enumerating objects would require nothing more than simply following straight pointers in the pack data in 99% of the cases. No object decompression, no memory buffer allocation/deallocation to perform that decompression, no string parsing in the tree object case, etc. Only cross pack references would require a full SHA1 based lookup like we do now.

I still have to sit down and figure out the implications of this, especially with forward references, meaning that the offset might have to be an object index so to allow for variable length encoding, and also to make sure index-pack can reconstruct the pack index. But that would only be an indirect lookup which shouldn't be significantly costly.

So that's the idea. Keep the exact same functionality as we have now, without any need for cache management, but making the data structure in a form that should improve object enumeration by some magnitude.

Nicolas
Previous: Shawn PearceNext: Shawn Pearce
Message 8 of 26 in “[RFC] Add --create-cache to repack”
  1. Shawn O. PearceJan 28, 2011
  2. Johannes SixtJan 28, 2011
  3. Shawn PearceJan 28, 2011
  4. Johannes SixtJan 28, 2011
  5. Shawn PearceJan 28, 2011
  6. Jay SoffianJan 28, 2011
  7. Shawn PearceJan 28, 2011
  8. Nicolas PitreJan 28, 2011
  9. Shawn PearceJan 28, 2011
  10. Nicolas PitreJan 28, 2011
  11. Shawn PearceJan 29, 2011
  12. Shawn PearceJan 29, 2011
  13. Junio C HamanoJan 30, 2011
  14. Shawn PearceJan 30, 2011
  15. Junio C HamanoJan 30, 2011
  16. Shawn PearceJan 30, 2011
  17. Nicolas PitreJan 30, 2011
  18. Nicolas PitreJan 29, 2011
  19. Shawn PearceJan 29, 2011
  20. Junio C HamanoJan 30, 2011
  21. Nicolas PitreJan 30, 2011
  22. A Large Angry SCMJan 30, 2011
  23. Shawn PearceJan 30, 2011
  24. Shawn PearceJan 30, 2011
  25. Shawn PearceJan 31, 2011
  26. Nicolas PitreJan 31, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.