git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git commit generation numbers

From
GSGeorge Spelvin <linux@horizon.com>
Date
Jul 20, 2011, 22:16 UTC
Message-ID
<20110720221632.14223.qmail@science.horizon.com>
In-Reply-To
<alpine.LFD.2.00.1107201538590.21187@xanadu.home>
Show 6 quoted lines
> The alternative of having to sometimes use the generation number, 
> sometimes use the possibly broken commit date, makes for much more 
> complicated code that has to be maintained forever.  Having a solution 
> that starts working only after a certain point in history doesn't look 
> eleguant to me at all.  It is not like having different pack formats 
> where back and forth conversions can be made for the _entire_ history.
It seemed like a pretty strong argument to me, too.
Show 6 quoted lines
> And if you don't care about graft/replace then the cached data is 
> immutable just like the in-commit version would, so there is no 
> consistency issues.  If you do care about graft/replace (or who knows 
> what other dag alteration scheme might be created in 5 years from now) 
> then a separate cache will be required _anyway_, regardless of any 
> in-commit gen number.

A possible workaround would be to keep track of the largest generation number skew introduced by any graft, and add that safety factor into the history-walking code, but that would be painful if you replace a single large commit with an equivalent long development history, such as adding a historical development tree behind a recently-cut-off one. or development history You can do a workaround at the expense of ine

> Neither did I think about the actual cache format (I don't think that
> adding it to the pack index is a good idea if grafts are to be honored)
> which certainly has bearing on that fundamental question too.

I was thinking of something very close to the V2 pack format. http://book.git-scm.com/7_the_packfile.html A magic number, a 256-entry fanout table, a sorted list of 20-byte hashes, followed by a matching list of 4-byte generation numbers.

Ending with a 20-byte hash of the replaces and grafts state that this cache is valid for, and a hash of the cache itself.

A bit of code factoring should make it easy to share much of the code.

It would certainly be possible to share the SHA1 table in an existing pack index and store the generation numbers of the base (no replacement) case, but you'd have to store null values for all the non-commit objects.

That takes 4 bytes per object, while a separate list of commits takes 24 bytes per commit. A separate list is better if commits are less than 1/6 of all objects.

Looking at git's own object database, we have:
 66125 blobs   (45.50%)
 49292 trees   (33.92%)
 29554 commits (20.33%)
   362 tags    ( 0.25%)
145333 total

So we're actually a bit over the 16.66% optimum. but it's not far enough to be a real efficiency problem.

Previous: Nicolas PitreNext: david@lang.hm
Message 10 of 35 in “Re: Git commit generation numbers”
  1. George SpelvinJul 17, 2011
  2. Long, MartinJul 17, 2011
  3. Linus TorvaldsJul 17, 2011
  4. George SpelvinJul 17, 2011
  5. Linus TorvaldsJul 17, 2011
  6. George SpelvinJul 18, 2011
  7. Anthony Van de GejuchteJul 18, 2011
  8. George SpelvinJul 18, 2011
  9. Nicolas PitreJul 20, 2011
  10. George SpelvinJul 20, 2011
  11. david@lang.hmJul 20, 2011
  12. Nicolas PitreJul 20, 2011
  13. Phil HordJul 21, 2011
  14. david@lang.hmJul 21, 2011
  15. Shawn PearceJul 21, 2011
  16. Phil HordJul 21, 2011
  17. david@lang.hmJul 21, 2011
  18. George SpelvinJul 21, 2011
  19. Jakub NarebskiJul 21, 2011
  20. George SpelvinJul 21, 2011
  21. Shawn PearceJul 21, 2011
  22. Jakub NarebskiJul 22, 2011
  23. Nicolas PitreJul 22, 2011
  24. david@lang.hmJul 22, 2011
  25. Jakub NarebskiJul 22, 2011
  26. Linus TorvaldsJul 22, 2011
  27. Jeff KingJul 22, 2011
  28. Felipe ContrerasJul 28, 2011
  29. Ramkumar RamachandraSep 6, 2011
  30. david@lang.hmJul 22, 2011
  31. Nicolas PitreJul 22, 2011
  32. david@lang.hmJul 22, 2011
  33. Phil HordJul 21, 2011
  34. Nicolas PitreJul 21, 2011
  35. Phil HordJul 21, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.