git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git commit generation numbers

From
Ramkumar Ramachandra <artagnon@gmail.com>
Date
Sep 6, 2011, 10:02 UTC
Message-ID
<CALkWK0kJ_0_MyUgQ+F+FgGun6vtk=VTh6Gsbb0u+EUrcLT5cGg@mail.gmail.com>
In-Reply-To
<CAMP44s2F429MG5DeRAULnSgNkCwrVGPfC2HeFw=iHXPXjkw0yA@mail.gmail.com>
Hi,

First, let me start out by saying that I'm a fairly new contributor to Git, and I'm far less experienced than the other people on this thread. I've read through all the discussions time and again, and thought about the problem for some time now - I can't say I understand it as fully as many of you do, but I think I may have a slightly different perspective to offer.

In what way is Git fundamentally different from Subversion? It's the simplicity of the data model. From the simplest building block, a key-value store, we have been able to compose and build things on top of it. The reason we built centralized version control systems earlier is because it was *easier* to address the composition problems. We dumped all related repository and problems into one central server. With so much information in one place, things are tightly coupled and problems are easier to solve. Still not convinced? What's the weakest component in Git today? Undoubtedly submodules. Ofcourse, a large part of the reason is that many people don't use submodules, and hence it doesn't improve -- but it's actually a circular problem. People don't use submodules, because it's so featureless and hard to develop. Why is it so hard? Back to the fundamental problem of composition from simple building blocks. In submodules, we have to take entire DAGs and build a composite DAG. The key pieces of information are deep inside Git's fundamnetals: Gitlinks. Other projects try like Gitslave try to attack the problem on a more superficial level, but they all hit a barrier when they discover that they can't compose big blocks of data: you need simple building blocks to compose.

It's the same story with C (and now, Haskell). Why does everyone like C so much? Because it only provides fundamental building blocks and gives people the freedom to compose the way they like. It doesn't provide big "template blocks" like Java, because they tend to be restrictive in the long run. Sure, Java is easier to start out with, but people soon realize that big blocks can't compose.

More than arguing about backward compatibility, and about how older versions of Git commits won't have generation numbers, I think this is what we should be focusing on. Sure, it'll additionally make sense to put in a cache to speed things up now, but we need to think about what Git will be 10~15 years from now. The fundamental pieces of information required for composition must be present in the fundamental building blocks.

The real question we should be asking is: "Should Git have had commit generation numbers in 2005?". If the answer is "yes", we should put them in now before it becomes even harder, bending over backwards for backward compatibility if necessary. Otherwise, we'll regret this decision 10~15 years later, when we're faced with deeper issues. If you want a concrete example, think about how you'd compose DAGs together (again, the submodules problem): where is the information required to prune each DAG and compose?

I wish I could write this in myself, but I'm afraid I don't have the engineering skill yet. I'll be happy to contribute whatever little I can, and participate in the review process.

Thanks.
-- Ram
Previous: Felipe ContrerasNext: david@lang.hm
Message 29 of 35 in “Re: Git commit generation numbers”
  1. George SpelvinJul 17, 2011
  2. Long, MartinJul 17, 2011
  3. Linus TorvaldsJul 17, 2011
  4. George SpelvinJul 17, 2011
  5. Linus TorvaldsJul 17, 2011
  6. George SpelvinJul 18, 2011
  7. Anthony Van de GejuchteJul 18, 2011
  8. George SpelvinJul 18, 2011
  9. Nicolas PitreJul 20, 2011
  10. George SpelvinJul 20, 2011
  11. david@lang.hmJul 20, 2011
  12. Nicolas PitreJul 20, 2011
  13. Phil HordJul 21, 2011
  14. david@lang.hmJul 21, 2011
  15. Shawn PearceJul 21, 2011
  16. Phil HordJul 21, 2011
  17. david@lang.hmJul 21, 2011
  18. George SpelvinJul 21, 2011
  19. Jakub NarebskiJul 21, 2011
  20. George SpelvinJul 21, 2011
  21. Shawn PearceJul 21, 2011
  22. Jakub NarebskiJul 22, 2011
  23. Nicolas PitreJul 22, 2011
  24. david@lang.hmJul 22, 2011
  25. Jakub NarebskiJul 22, 2011
  26. Linus TorvaldsJul 22, 2011
  27. Jeff KingJul 22, 2011
  28. Felipe ContrerasJul 28, 2011
  29. Ramkumar RamachandraSep 6, 2011
  30. david@lang.hmJul 22, 2011
  31. Nicolas PitreJul 22, 2011
  32. david@lang.hmJul 22, 2011
  33. Phil HordJul 21, 2011
  34. Nicolas PitreJul 21, 2011
  35. Phil HordJul 21, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.