git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git commit generation numbers

From
GBGeert Bosch <bosch@adacore.com>
Date
Jul 15, 2011, 02:41 UTC
Message-ID
<186BDF84-7AE3-4E0F-8F6D-AA89A60C972C@adacore.com>
In-Reply-To
<CA+55aFyDzr+SfgSzWMr9pQuQUXTw9mcjZ-00NZof74PKZzbGPA@mail.gmail.com>
On Jul 14, 2011, at 21:19, Linus Torvalds wrote:
> But dammit, if you start using generation numbers, then they *are*
> required information. The fact that you then hide them in some
> unarchitected random file doesn't change anything! It just makes it
> ugly and random, for chrissake!

Generation numbers never will be required information, because we can always compute them. These numbers are really much more similar to other pack index information than anything else.

<aside> Sometimes I wish we'd have general "depth" information for each SHA1, which would be the maximum number of steps in the DAG to reach a leaf. This way, if we want to do something like "git log drivers/net/slip.c", we don't have to bother reading the majority of trees that have a depth less than two. The depth can also be used as a limiter for "contains" operations, where we want to see if commit X contains commit Y: depth (X) has to be at least depth (Y).

However, any such notion, wether generation or depth or whatever else we'll think of tomorrow, is something particular to a certain implementation of git. It does not add anything to the information we stored. </aside>

I don't think my commit should have a different SHA1 from yours, because your tree has a more generation numbers than mine.

The beauty and genius of GIT is that it just takes the minimum amount of data needed to uniquely identify the information to be stored, and stores that in a UNIQUE format. By allowing generation numbers to either be present or absent, that's all broken.

It's like computing the SHA1 of compressed data: it doesn't depend on the data we store, just about the particular representation we choose. Fortunately we have done away with the first mistake.

So, if you're going to add generation numbers, there has to be a flag day, after which generation numbers are required everywhere. Of course it would be possible to recognize "old style" commits and convert them on the fly, but that is true for pretty much any format change. However, adding redundant information seems like a poor excuse for having a flag day.

Storing generation data in pack indices on the other hand makes perfect sense: when we generate these indices, we do complete traversals and have all required information trivially at hand. We can never have that many loose objects, so lack of generation information there isn't a big deal. By storing generation information in the index, we can be sure it is consistent with the data contained in the pack, so there are no cache invalidation issues.

I know I must have missed some stupid and obvious reason why this is all wrong, I just don't quite see it yet.

  -Geert
Previous: Linus TorvaldsNext: Jeff King
Message 16 of 47 in “Git commit generation numbers”
  1. Linus TorvaldsJul 14, 2011
  2. Jeff KingJul 14, 2011
  3. Linus TorvaldsJul 14, 2011
  4. Linus TorvaldsJul 14, 2011
  5. Jeff KingJul 14, 2011
  6. Ted Ts'oJul 14, 2011
  7. Linus TorvaldsJul 14, 2011
  8. Jeff KingJul 14, 2011
  9. Ted Ts'oJul 14, 2011
  10. Jeff KingJul 14, 2011
  11. Linus TorvaldsJul 14, 2011
  12. Jeff KingJul 14, 2011
  13. Linus TorvaldsJul 14, 2011
  14. Jeff KingJul 14, 2011
  15. Linus TorvaldsJul 15, 2011
  16. Geert BoschJul 15, 2011
  17. Jeff KingJul 15, 2011
  18. Linus TorvaldsJul 15, 2011
  19. Shawn PearceJul 15, 2011
  20. Linus TorvaldsJul 15, 2011
  21. Ted Ts'oJul 15, 2011
  22. Linus TorvaldsJul 15, 2011
  23. Christian CouderJul 16, 2011
  24. Jeff KingJul 18, 2011
  25. Christian CouderJul 19, 2011
  26. Jeff KingJul 19, 2011
  27. Christian CouderJul 21, 2011
  28. Tony LuckJul 15, 2011
  29. Linus TorvaldsJul 15, 2011
  30. Jeff KingJul 15, 2011
  31. Jeff KingJul 15, 2011
  32. Linus TorvaldsJul 15, 2011
  33. Jeff KingJul 15, 2011
  34. Linus TorvaldsJul 15, 2011
  35. Linus TorvaldsJul 15, 2011
  36. Linus TorvaldsJul 15, 2011
  37. Jeff KingJul 16, 2011
  38. Jeff KingJul 16, 2011
  39. Jakub NarebskiJul 15, 2011
  40. Long, MartinJul 15, 2011
  41. Long, MartinJul 15, 2011
  42. Drew NorthupJul 15, 2011
  43. Linus TorvaldsJul 14, 2011
  44. Jakub NarebskiJul 14, 2011
  45. Junio C HamanoJul 14, 2011
  46. Jeff KingJul 14, 2011
  47. Junio C HamanoJul 14, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.