git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: CAREFUL! No more delta object support!

From
Linus Torvalds <torvalds@osdl.org>
Date
Jun 28, 2005, 16:45 UTC
Message-ID
<Pine.LNX.4.58.0506280921480.19755@ppc970.osdl.org>
In-Reply-To
<20050628103852.GB21533@64m.dyndns.org>
On Tue, 28 Jun 2005, Christopher Li wrote:
Show 5 quoted lines
>
> That is all nice improvement to address the space usage issue.
> 
> Should people just run repacking once a while or is it automaticly
> add new object to the pack file?

While adding a new object to a pack file is _possible_ (you add it to the end of the pack-file, and re-generate the index file), I would strongly suggest against it for several reasons:

 - It's a lot more complex and expensive than just writing a new file.  
   Much better to make the pack generation be an off-line thing, and make 
   new object creation really cheap.
 - it has serious locking issues, and if something goes wrong you are just 
   horribly screwed. This implies, for example, that to be safe you really 
   have to use fsync() etc at every point (and be careful about writing 
   the index), making the update even _more_ expensive. Over NFS you need 
   to be extremely careful to make sure that everybody got the right lock, 
   yadda yadda.
   Packing things off-line just means that _all_ of these problems go 
   away.
 - There are operations that want to remove objects (I do that all the 
   time: I do something stupid, and decide to undo it, or I just do a 
   "git-update-cache" and notice that I need to do more work so I edit it 
   some more and actually never commit the first version)
   If _adding_ to the file had some serious correctness issues, _removing_ 
   an object from a file is even worse. MUCH worse. Now you don't just 
   have to lock against other people creating new objects, now you have to 
   lock against updates (or totally re-write the whole big file and do an 
   atomic "rename").
 - it can actually generate worse packing. The current "offline" method 
   means that we can pack any version of a file against any other version 
   of a file, and we do. We pick the closest version we can find, and we 
   try to always pack against the bigger one (deletes are smaller deltas, 
   and the biggest one tends to be the latest version, so this not only
   means that the delta is denser, it also means that the latest version -
   which is likely to be the biggest and most often used - tends to be
   non-delta).
   In contrast, updating the pack file means that you always write the 
   latest version as a delta, which means that you're doing things 
   _exactly_ the wrong way around both for performance and size.
 - Finally: packing allows us to do optimize for locality. In particular, 
   I write out the pack file in "recency" order, ie the top-most objects 
   go first, and in particular, the "commit" objects go at the very top of 
   the file. Why? Because it means that the commit objects (which are 
   heavily used for the history generation by pretty much anything, since 
   "git-rev-list" will access them) are packed together, and in the right 
   order.
   Again, you can't do that if you do on-line updates as opposed to 
   offline packing.

So the usage pattern I envision is to pack stuff maybe once a month (depending on how much changes, of course), because then you really do get the best of both worlds: the simplicity of individual objects for recent work and the optimal packing and ordering that you can really work on for the longer range case. And your project never grows very big.

Btw, I'm not claiming that my current pack format is "optimal" of course. For example, while I write all objects in recency order, right now that means that if a recent object has been written as a delta that depends on an older one, I actually write the delta first (correct) but I won't write the older object until its recency ordering (wrong).

That kind of thing is trivial to fix (eventually), but it's an example of where ordering matters (ie if it's the other way around: the delta is the older object, it's probably better to leave it at the end of the file, since it's probably not going to be accessed much, making the effective packing at the head more efficicient). It's also an example of the kinds of things we can do exactly because we're doing the packing off-line.

			Linus
Previous: Christopher LiNext: Junio C Hamano
Message 11 of 38 in “CAREFUL! No more delta object support!”
  1. Linus TorvaldsJun 28, 2005
  2. Christopher LiJun 27, 2005
  3. Linus TorvaldsJun 28, 2005
  4. Junio C HamanoJun 28, 2005
  5. Christopher LiJun 28, 2005
  6. Petr BaudisJun 28, 2005
  7. Benjamin LaHaiseJun 28, 2005
  8. Petr BaudisJun 28, 2005
  9. Jan HarkesJun 28, 2005
  10. Christopher LiJun 28, 2005
  11. Linus TorvaldsJun 28, 2005
  12. Emit base objects of a delta chain when the delta is output.Junio C Hamano, Jun 29, 2005
  13. Junio C HamanoJun 28, 2005
  14. Skip writing out sha1 files for objects in packed git.Junio C Hamano, Jun 28, 2005
  15. Linus TorvaldsJun 28, 2005
  16. Junio C HamanoJun 28, 2005
  17. Linus TorvaldsJun 28, 2005
  18. Linus TorvaldsJun 28, 2005
  19. Junio C HamanoJun 28, 2005
  20. Adjust to git-init-db creating $GIT_OBJECT_DIRECTORY/packJunio C Hamano, Jun 28, 2005
  21. Linus TorvaldsJun 28, 2005
  22. Daniel BarkalowJun 28, 2005
  23. Linus TorvaldsJun 28, 2005
  24. Linus TorvaldsJun 28, 2005
  25. Daniel BarkalowJun 28, 2005
  26. Linus TorvaldsJun 28, 2005
  27. Linus TorvaldsJun 28, 2005
  28. Matthias UrlichsJun 28, 2005
  29. Matthias UrlichsJun 28, 2005
  30. Daniel BarkalowJun 28, 2005
  31. Linus TorvaldsJun 29, 2005
  32. Linus TorvaldsJun 29, 2005
  33. Daniel BarkalowJun 29, 2005
  34. Linus TorvaldsJun 29, 2005
  35. Daniel BarkalowJun 29, 2005
  36. Adjust fsck-cache to packed GIT and alternate object pool.Junio C Hamano, Jun 28, 2005
  37. Expose packed_git and alt_odb.Junio C Hamano, Jun 28, 2005
  38. 3/3 Update fsck-cache (take 2)Junio C Hamano, Jun 28, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.