{"thread":{"id":"27834","subject":"Re: Git commit generation numbers","startedAt":"2011-07-17T18:27:43Z","lastAt":"2011-09-06T10:02:03Z","messageCount":35,"participants":["George Spelvin","Long, Martin","Linus Torvalds","Anthony Van de Gejuchte","Nicolas Pitre","david@lang.hm","Phil Hord","Shawn Pearce","Jakub Narebski","Jeff King","Felipe Contreras","Ramkumar Ramachandra"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"171511","messageId":"20110717182743.14423.qmail@science.horizon.com","threadId":"27834","inReplyTo":null,"subject":"Re: Git commit generation numbers","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2011-07-17T18:27:43Z","receivedAt":"2011-07-17T18:27:43Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> The thing I hate about it is very fundamental: I think it's a hack around a basic git\n> design mistake. And it's a mistake we have known about for a long time.\n> \n> Now, I don't think it's a *fatal* mistake, but I do find it very broken to basically\n> say \"we made a mistake in the original commit design, and instead of fixing it we\n> create a separate workaround for it\".\n> \n> THAT I find distasteful. My reaction is that if we're going to add generation\n> numbers, then were should just do it the way we should have done them originally,\n> rather than as some separate hack.\n\nThere are a few design mistakes in git.  The way the object type\nand size are prefixed to the data for hasing purposes, which prevents\naligned fetching from memory-mapped data in the hashing code, isn't too\npretty either.\n\nBut git has generally preferred to avoid storing information that can\nbe recomputed.  File renames are the big example.  given this, why the\nheck store generation numbers?\n\nThey *can* be computed on demand, so arguably they *should*.  Cacheing is\nthen an optimization, just like packs, pack indexes, the hashed object\nstorage directories, and all that.\n\n\nI'm in the \"make it a cache\" camp, honestly.  \n\n\nFor example, here's a different possible generation number scheme.\nBy making the generation number a cache, it becomes a valid alternative\nto experiment with.\n\nSimply store a topologically sorted list of commits.  Each commit's\nposition can serve as a generation number, and is greater than the\npositions of all ancestors.  But by using the offset within the list,\nthe number is stored implicitly.\n\nGeneration numbers don't have to be consecutive as long as they're\ncorrectly ordered, so you could, e.g. choose to make them unique.\n\nI don't think this is actually worth it; I'm just using it as a\nnot-completely-insane example of a different design that nonetheless\nachieves the same goal.\n\nWhy freeze this in the object format?\n"},{"id":"171515","messageId":"CANfMb_8ctvPqAJ-mpAZKvcXLwV8QexFU-u48GC0CmNqxFeJ5YQ@mail.gmail.com","threadId":"27834","inReplyTo":"20110717182743.14423.qmail@science.horizon.com","subject":"Re: Git commit generation numbers","fromName":"Long, Martin","fromEmail":"martin@longhome.co.uk","sentAt":"2011-07-17T19:00:17Z","receivedAt":"2011-07-17T19:00:17Z","isPatch":false,"sender":{"key":"martin@longhome.co.uk","avatar":"https://gravatar.com/avatar/24daa6058f5db96c0590b65c055bc115a5f4a9ee9dcb67a8b9cce6c6469eac85?d=mp&s=160"},"body":"> Why freeze this in the object format?\n\nBecause if you put it in the object format, then it gets pushed and\npulled around, thereby putting generation numbers in every clone.\n\nI'm starting to think put them in the object store, for exactly that\nreason, and to start moving repositories in a direction where the look\nmore like they would if this had been done correctly from the start.\n\nThen, because some operations are still going to create a lot of\ntraversals, a cache is always an option to improve the performance in\nthat area.\n"},{"id":"171518","messageId":"CA+55aFwqFhzd_cmbFxkCyNXhF99igBqdr8p4J76hLz=m4=ZNWg@mail.gmail.com","threadId":"27834","inReplyTo":"20110717182743.14423.qmail@science.horizon.com","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-17T19:30:34Z","receivedAt":"2011-07-17T19:30:34Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Sun, Jul 17, 2011 at 11:27 AM, George Spelvin <linux@horizon.com> wrote:\n>\n> There are a few design mistakes in git.  The way the object type\n> and size are prefixed to the data for hasing purposes, which prevents\n> aligned fetching from memory-mapped data in the hashing code, isn't too\n> pretty either.\n\nWhy would you ever care? That makes no sense.\n\n> But git has generally preferred to avoid storing information that can\n> be recomputed.  File renames are the big example.  given this, why the\n> heck store generation numbers?\n\nGuys, please don't bring up file renames. I explained once already why\nbringing up file renames just makes you look like a f^&% moron.\n\nLet me explain one more time:\n\n - Storing file renames is STUPID. It's stupid for very fundamental\nreasons that have absolutely *NOTHING* to do with \"it can be computed\nlater\".\n\nIt's fundamentally stupid because it will FOREVER SCREW UP YOUR DATA,\nand because it will make merging an unmitigated disaster and make your\nrepository depend on how you *created* your data, rather than on what\nthe data is. It will totally break the situation of one person doing a\nrename, while another person does something else to the metadata (eg a\ncreate of the same filename).\n\nTrying to track file identities will leave to very fundamentally\nunsolvable issues like \"which file identity do we choose when two\ndifferent files get the same name\", or \"which file identity will we\nchoose when one file splits in two\".\n\nGit doesn't track renames, because unlike pretty much every other SCM\nout there, git really does have a good design, and because I damn well\nunderstood the real problems.\n\nSo bringing it up as an example of \"we don't store it because we can\ncompute it\" is really totally idiotic. It's a sign of not\nunderstanding the problems with renames. Stop doing it. That argument\nis totally irrelevant. Really.\n\nIt's like saying \"We shouldn't do generation numbers because fish\ndon't use bicycles\". The only thing that kind of argument does is to\nmake me convinced that you don't understand the problem enough to be\nworth even arguing with. It is not only a worthless argument, but it\nmakes your every other argument suspect.\n\nComprende? Stop it.\n\n> They *can* be computed on demand, so arguably they *should*.\n\nUmm, no.\n\nThat's actually a really bad argument.\n\nThere are valid things that we \"should\" do, but they have nothing to\ndo with \"if something can be done, it should be done\". That's just a\ncrazy argument.\n\nA thing we really *should* do is perform well. And be really reliable.\nAnd support a distributed workflow.\n\nThose are real arguments that aren't about \"just because it's there\".\n\nNow, some of those arguments can then be used to say \"don't bother\nstoring redundant data\". For example, redundant data takes disk space\nand network bandwidth, and if something can be recomputed cheaply (ie\nif it doesn't have a negative impact on performance), then redundant\ndata is just bad.\n\nAnd what appears like a much better argument (right now) is that some\ndata isn't needed AT ALL, because you can make do with other data\nentirely (ie dates).\n\nBut \"just because we could recompute it\" is a bad bad reason.\n\nThe thing is, the very basic design of git is all about *incomplete*\nDAG traversal. The DAG traversal part is pretty obvious and simple,\nbut the *partial* thing really is very very important. We absolutely\nneed it for reasonable scalability. We've spent a *lot* of time in git\ndevelopment on trying to perform really well by avoiding work. Not\njust in revision traversal, but in many other areas too (like making\ndiff and merge much faster by being able to handle whole identical\nrecursive subdirectories by just checking the SHA1, for example).\n\nThat's a *really* fundamental design issue in git. Performance was\nalways a primary goal. And by primary, I really mean primary. As in\n\"more important than just about anything else\".  There were other\nprimary goals, but really not very many.\n\nAnd there really aren't very good ways to limit DAG traversal.\nGeneration numbers are one of the very few fundamental ones. We hacked\naround it with dates, and it works pretty well in practice (well\nenough that I'm certainly ok with the hack), but it's definitely one\nof the areas where git simply does something \"wrong\". It's simply not\na entirely reliable algorithm, and that fact makes me a bit\nuncomfortable with it.\n\n(Now, in theory, a global *approximate* time is theoretically possible\nin a distributed environment, and as such it's arguable that \"global\ntime with a slop that is based on the speed of light and knowledge of\nlocation\" is at least theoretically sound. So the real problem with\ncommit dates is that people simply don't have good clocks. So it's a\npractical problem rather than a theoretical one, and it's a practical\nproblem that doesn't really cause enough problems in practice to not\nbe workable. But I'm making excuses for it, and I _know_ I'm making\nexcuses for it, so I'm not really happy about it)\n\nAnd it's just about the only area where I am aware of git doing\nsomething \"wrong\". Which is why I would like to have had generation\nnumbers even though the dates do work.\n\nAnyway, to get back to the actual issue of caching vs not caching: if\nyou think \"we could compute it dynamically\" means that we should, then\nwe damn well shouldn't cache it either - why cache it, when you could\njust compute it. And if it's worth it to waste resources on the cache\nin order to avoid performance issues, then it damn well would be ok to\nwaste (fewer) resources on just saving the generation number in the\nobject data base. And make that *fundamental* fix to a hack that git\nhas had since pretty much day one.\n\nAnd btw, git didn't have the date-based hack originally, because I\ndidn't think it would be problematic enough. I thought that we could\ndo universally efficient partial DAG traversal - not having to go all\nthe way to the root -  based purely on the DAG. The code in\n\"everybody_uninteresting()\" tries to be that \"limit DAG traversal by\nonly looking at the DAG itself\", and it works for many simple\nsituations. But it turns out that it does *not* work for many other\ncases.\n\nSo the generation number really is very very fundamnetal. It's\nabsolutely not some \"additional information that can be computed\",\nbecause the whole AND ONLY point of having the number is to not\ncompute it.\n\nWe are never interested in the generation number for its own sake. We\nare only interested in it in order to avoid having to look at the rest\nof the DAG.\n\nSo no, the number fundamentally isn't computable, because computing it\nobviates the need for it.\n\n                           Linus\n"},{"id":"171526","messageId":"20110717233959.3548.qmail@science.horizon.com","threadId":"27834","inReplyTo":"CA+55aFwqFhzd_cmbFxkCyNXhF99igBqdr8p4J76hLz=m4=ZNWg@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2011-07-17T23:39:59Z","receivedAt":"2011-07-17T23:39:59Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> So the generation number really is very very fundamnetal. It's\n> absolutely not some \"additional information that can be computed\",\n> because the whole AND ONLY point of having the number is to not\n> compute it.\n> \n> We are never interested in the generation number for its own sake. We\n> are only interested in it in order to avoid having to look at the rest\n> of the DAG.\n\nYou're making my point and somehow not seeing it.\n\nWhat you're describing here is the archetpical cache.\n\nThe only reason for having a memory cache is to avoid accessing memory!\nThe only reason for having a TLB is to avoid walking the page tables!\nThe only reason for having a page cache is to avoid hitting the disk!\nThe only reason for having a dcache is to avoid traversing the file\nsystem directories!\n\nAnd yes, the only reason for having a generation number cache is to avoid\ntraversing the DAG.  D'oh.  Do you think this is somehow news to anyone?\n\nThe fundamental nature of a cache is that it lets you look something up\nquickly that you could compute but don't want to.\n\nI'm slapping my forehead like Homer Simpson here.  The fact that computing\nthe generation number is expensive is why it's worth cacheing.  But the\nfact that it *can* be computed is a reason not to clutter the published\ncommit object format with it.\n\n\nThe generation number is NOT FUNDAMENTAL.  It contains no information\nthat's not already in the DAG.  The danger of putting it into a commit\nis that you'll do it wrong, and thereby screw everything up.\n\nIf we have broken code that generates a broken cache, we fix the code\nand the bugs magically go away.\n\nIf we have broken code that generates a broken commit object, we have\na huge problem.\n\nJust like we don't ship pack indexes around, but recompute them on arrival.\nThe index is essential for performance, but it's absolutely non-essential\nfor correctness.\n\n\nAs a general design principle, the exported data structures, like the\ncommits, should be as simple as possible.  Do not include extraneous\nor redundant data, because then you have to deal with the possibility\nof inconsistency.  This leads to bugs.  (Frequently buffer overflow bugs.)\n\nMaybe it would have been worth violating that principle during the initial\ngit design.  I still see a good argument for not doing that even if we\nhad a time machine.\n\nBut now that the commit format is established and widely used, the argument\nhas far more force.  Changing the commit format provides zero functionality\ngain, and the performance gain can be obtained a different way.\n\nMaybe a bit more code, but nothing extraordinary.\n\nTo me, the KISS principle says \"don't change the commit format!\"\n\nNow, you complain about code complexity.  But this is a read-only cache.\nThe generation number of a commit object never changes.  There's no update\noperation.  Like an I-cache, if there's ever any problem, throw it away.\n\nArguing that \"the patch to put it in the commit object is smaller\" is\nstupidly short-sighted.  Now every version of git from now until forever\nhas to support both kinds of commit objects.  (And browsing old git\ntrees will forever be slow.)\n\nYou only take on that sort of legacy support burden if you absolutely have to.\n\n> But \"just because we could recompute it\" is a bad bad reason.\n\nBull puckey.  You're ugly and stupid and WRONG.\n\nIt's an excellent reason.  I'm amazed that you're not seeing it.\nThe principle is \"don't include redundant data in a transport format.\"\nBecause it can be recomputed, it's redundant.  Therefore, it shouldn't\nbe included in the transport format.\n\nIt's exactly the same principle as \"don't store the indexes in the\ndatabase dump\" and \"don't store filename hashes in file system\narchives\".\n\nThis is a principle, not an iron-clad rule.  It can be violated for\ngood and sufficient reasons, notably performance.\n\nBut in this case, we can get the performance without it.  Without,\nin fact, changing the git transport format at all.\n\nAnd \"don't change a widely-used transport format\" is ANOTHER important\nprinciple.  Backward-compatible is much better than incompatible, but\nfar better to avoid changing it at all.\n\nBreaking two such principles without an absolutely iron-clad reason is\nugly and stupid and wrong.\n\n(As you well know, the more general principle is \"don't store redundant\ndata AT ALL unless you need to for performance\".  Redundant data is A\nBad Thing.  It can get out of sync.  But if you have to, a private cache\nis much better than a exchange format.)\n\n\nPut another way, it IS stupid, it IS expendable, and therefore it SHOULD go.\n"},{"id":"171529","messageId":"CA+55aFwt+RDRK_r=9CXbdzsLuGDqswvGTtJDKi9Q3DQwB_Ha5Q@mail.gmail.com","threadId":"27834","inReplyTo":"20110717233959.3548.qmail@science.horizon.com","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-17T23:58:59Z","receivedAt":"2011-07-17T23:58:59Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Sun, Jul 17, 2011 at 4:39 PM, George Spelvin <linux@horizon.com> wrote:\n>\n> I'm slapping my forehead like Homer Simpson here.  The fact that computing\n> the generation number is expensive is why it's worth cacheing.  But the\n> fact that it *can* be computed is a reason not to clutter the published\n> commit object format with it.\n\nAnd I'm slapping *my* forehead.\n\nNobody has *ever* given a reason why the cache would be better than\njust making it explicit.\n\nThat's my issue.\n\nWhy is that so hard for people to understand? The cache is just EXTRA WORK.\n\nTo take your TLB example: it's like having a TLB for a page table that\nwould be as easy to just create in a way that it's *faster* to look up\nin the actual data structure than it would be to look up in the cache.\n\nOr to take your disk cache example: wouldn't you say that a disk cache\nis a F&*&ING BAD IDEA if it is slower than the disk it caches?\n\nSeriously.\n\n                    Linus\n"},{"id":"171540","messageId":"20110718051347.28952.qmail@science.horizon.com","threadId":"27834","inReplyTo":"CA+55aFwt+RDRK_r=9CXbdzsLuGDqswvGTtJDKi9Q3DQwB_Ha5Q@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2011-07-18T05:13:47Z","receivedAt":"2011-07-18T05:13:47Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> Nobody has *ever* given a reason why the cache would be better than\n> just making it explicit.\n\nI thought I listed a few.  Let me be clearer.\n\n1) It involves changing the commit format.  Since the change is\n   backward-compatible, it's not too bad, but this is still fundamentally\n   A Bad Thing, to be avoided if possible.\n\n2) It can't be retrofitted to help historical browsing.\n\n3) You have to support commits without generation numbers forever.\n   This is a support burden.  If you can generate generation numbers for\n   an entire repository, including pre-existing commits, you can *throw\n   out* the commit date heuristic code entirely.\n\n4) It can't be made to work with grafts or replace objects.\n\n5) It includes information which is redundant, but hard to verify,\n   in git objects.  Leading to potentially bizarre and version-dependent\n   behaviour if it's wrong.  (Checking that the numbers are consistent\n   is the same work as regenerating a cache.)\n\n6) It makes git commits slightly larger.  (Okay, that's reaching.)\n\n> Why is that so hard for people to understand? The cache is just EXTRA WORK.\n\nThat's why it *might* have been a good idea to include the number in\nthe original design.  But now that the design is widely deployed, it's\nbetter to avoid changing the design if not necessary.\n\nWith a bit of extra work, it's not necessary.\n\n> To take your TLB example: it's like having a TLB for a page table that\n> would be as easy to just create in a way that it's *faster* to look up\n> in the actual data structure than it would be to look up in the cache.\n\nYou've subtly jumped points.  The original point was that it's worth\nprecomputing and storing the generation numbers.  I was trying to\nsay that this is fundamentally a caching operation.\n\nNow we're talking about *where* to store the cached generation numbers.\n\nYour point, which is a very valid one, is that they are to be stored\non disk, exactly one per commit, can be computed when the commit is\ngenerated, and are accessed at the same time as the commit, so it makes\nall kinds of sense to store them *with* the commits.  As part of them,\neven.\n\nThis has the huge benefit that it does away with the need for a *separate*\ndata structure.  (Kinda sorts like the way AMD stores instruction\nboundaries in the L1 I-cache, avoiding the need for a separate data\nstructure.)\n\nI'm arguing that, despite this annoying overhead, there are valid reasons\nto want to store it separately.  There are some practical ones, but the\nbasic one is an esthetic/maintainability judgement of \"less cruft in\nthe commit objects is worth more cruft in the code\".\n\nGit has done very well partly *because* of the minimality of its basic\npersistent object database format.  I think we should be very reluctant\nto add to that without a demonstrated need that *cannot* be met in\nanother way.\n\n\nIn this particular case, a TLB is not a transport format.  It's okay\nto add redundant cruft to make it faster, because it only lasts until\nthe next reboot.  (A more apropos, software-oriented analogy might be\n\"struct page\".)\n\nA git commit object *is* a transport format, one specifically designed\nfor transporting data a very long way forward in time, so it should be\ndesigned with considerable care, and cruft ruthlessly eradicated.\n\nWhatever you add to it has to be supported by every git implementation,\nforever.  As does every implementation bug ever produced.\n\nA cache, on the other hand, is purely a local implementation detail.\nIt can be changed between versions with much less effort.\n\nI agree it's more implementation work.  But the upside is a cleaner\nstruct commit.  Which is a very good thing.\n"},{"id":"171568","messageId":"A142F49B-FC91-410A-B3C8-15FC7E5C3C68@gmail.com","threadId":"27834","inReplyTo":"20110718051347.28952.qmail@science.horizon.com","subject":"Re: Git commit generation numbers","fromName":"Anthony Van de Gejuchte","fromEmail":"anthonyvdgent@gmail.com","sentAt":"2011-07-18T10:28:16Z","receivedAt":"2011-07-18T10:28:16Z","isPatch":false,"sender":{"key":"anthonyvdgent@gmail.com","avatar":null},"body":"On 18-jul-2011, at 07:13, George Spelvin wrote:\n\n>> Nobody has *ever* given a reason why the cache would be better than\n>> just making it explicit.\n> \n> I thought I listed a few.  Let me be clearer.\n> \n> 1) It involves changing the commit format.  Since the change is\n>   backward-compatible, it's not too bad, but this is still fundamentally\n>   A Bad Thing, to be avoided if possible.\n\nGit is designed to ignore data in this case afaik, so I do not see any\nreason why backwards-compatibility gets broken here.\n\n> \n> 2) It can't be retrofitted to help historical browsing.\n\nI like to see more (valid) arguments, as I do not see what you are\ntrying to explain.\n\n> \n> 3) You have to support commits without generation numbers forever.\n>   This is a support burden.  If you can generate generation numbers for\n>   an entire repository, including pre-existing commits, you can *throw\n>   out* the commit date heuristic code entirely.\n\nI'll give you a few months to rethink at this statement until this\nfeature does get used widely. I think there was never a moment where\nwe would ever think to rebuild older commits as this would break the\nhash of the commits where many people are potential looking for.\n\n> \n> 4) It can't be made to work with grafts or replace objects.\n> \n> 5) It includes information which is redundant, but hard to verify,\n>   in git objects.  Leading to potentially bizarre and version-dependent\n>   behaviour if it's wrong.  (Checking that the numbers are consistent\n>   is the same work as regenerating a cache.)\n\nThe data is *consistent* as long as the hash doesn't change, storing the\ndata in the commits *can* reduce resource and makes calculations cheaper.\nTherefore, I think there are enough reasons to add the generation number\nin the commit. Yes, many data can be calculated or can be an overhead,\nbut as Torvalds already said, it can be used as consistency check.\n\nIf the data does get wrong, then its probably caused by something stupid\nenough to break the rules. Yes, this is a problem but I think there are\nalready enough reasons given, look back to the archives of this topic.\n\nOk, there is one possible thing that *can* go wrong and that is when you\nare changing history with generation numbers with an older git client.\n\n(And thats a good reason to communicate with others as clear as possible\nabout this feature, but its still not version-dependent as it doesn't\nrequire a client to use it)\n\n> \n> 6) It makes git commits slightly larger.  (Okay, that's reaching.)\n> \n>> Why is that so hard for people to understand? The cache is just EXTRA WORK.\n> \n> That's why it *might* have been a good idea to include the number in\n> the original design.  But now that the design is widely deployed, it's\n> better to avoid changing the design if not necessary.\n> \n> With a bit of extra work, it's not necessary.\n> \n>> To take your TLB example: it's like having a TLB for a page table that\n>> would be as easy to just create in a way that it's *faster* to look up\n>> in the actual data structure than it would be to look up in the cache.\n> \n> You've subtly jumped points.  The original point was that it's worth\n> precomputing and storing the generation numbers.  I was trying to\n> say that this is fundamentally a caching operation.\n> \n> Now we're talking about *where* to store the cached generation numbers.\n> \n> Your point, which is a very valid one, is that they are to be stored\n> on disk, exactly one per commit, can be computed when the commit is\n> generated, and are accessed at the same time as the commit, so it makes\n> all kinds of sense to store them *with* the commits.  As part of them,\n> even.\n> \n> This has the huge benefit that it does away with the need for a *separate*\n> data structure.  (Kinda sorts like the way AMD stores instruction\n> boundaries in the L1 I-cache, avoiding the need for a separate data\n> structure.)\n> \n> I'm arguing that, despite this annoying overhead, there are valid reasons\n> to want to store it separately.  There are some practical ones, but the\n> basic one is an esthetic/maintainability judgement of \"less cruft in\n> the commit objects is worth more cruft in the code\".\n> \n> Git has done very well partly *because* of the minimality of its basic\n> persistent object database format.  I think we should be very reluctant\n> to add to that without a demonstrated need that *cannot* be met in\n> another way.\n> \n> \n> In this particular case, a TLB is not a transport format.  It's okay\n> to add redundant cruft to make it faster, because it only lasts until\n> the next reboot.  (A more apropos, software-oriented analogy might be\n> \"struct page\".)\n> \n> A git commit object *is* a transport format, one specifically designed\n> for transporting data a very long way forward in time, so it should be\n> designed with considerable care, and cruft ruthlessly eradicated.\n> \n> Whatever you add to it has to be supported by every git implementation,\n> forever.  As does every implementation bug ever produced.\n> \n> A cache, on the other hand, is purely a local implementation detail.\n> It can be changed between versions with much less effort.\n> \n> I agree it's more implementation work.  But the upside is a cleaner\n> struct commit.  Which is a very good thing.\n\nA cache would use more resources because they can become invalid at any\npoint and *should* be recalculated by every client. We are processing\ndata that *can* be reused by everybody with a git client which has this\nspecific feature, but does not break anything with an older client.\n\nSo please, calculate things only once as this may save a *lot* of time :-)\n\nI would see more advantage in a cache if the data could differs on\nevery client, but that still doesn't mean that you should use one.\n\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n\nMaybe I shouldn't even have responded to this as I tend not to agree with\nthe given opinions to use a cache, even when I think that Torvalds starts\nthrowing arguments as well for certain reasons, but thats probably my wrong\nthinking at it.\n"},{"id":"171573","messageId":"20110718114834.12406.qmail@science.horizon.com","threadId":"27834","inReplyTo":"A142F49B-FC91-410A-B3C8-15FC7E5C3C68@gmail.com","subject":"Re: Git commit generation numbers","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2011-07-18T11:48:34Z","receivedAt":"2011-07-18T11:48:34Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":">> 1) It involves changing the commit format.  Since the change is\n>>   backward-compatible, it's not too bad, but this is still fundamentally\n>>   A Bad Thing, to be avoided if possible.\n\n> Git is designed to ignore data in this case afaik, so I do not see any\n> reason why backwards-compatibility gets broken here.\n\nThat's what I just wrote.  \"The change is backward-compatible\"\nis a simpler and shorter way of writing \"it doesn't break\nbackwards-compatibility\" (to put the generation number in the commit\nobject).\n\nI just said that *any* change is still undesirable.\n\n>> 2) It can't be retrofitted to help historical browsing.\n\n> I like to see more (valid) arguments, as I do not see what you are\n> trying to explain.\n\nI apologize for being unclear.  I meant that if you store the generation\nin the commit, then you can't add generation numbers to an existing\nrepository (\"retrofit\") in order to speed up --contains and --topo-sort\noperations on pre-existing git repositories.\n\n(Without recomputing all the hashes and breaking the ability to merge\nwith people not using the feature.)\n\nAs Linus points out, this is not likely to be a major performance issue\nin practice, as operations like finding merge bases overwhelmingly\nuse recent objects (which will have generation numbers once the feature\ngoes in), but it is a measurable disadvantage.\n\n>> 3) You have to support commits without generation numbers forever.\n>>   This is a support burden.  If you can generate generation numbers for\n>>   an entire repository, including pre-existing commits, you can *throw\n>>   out* the commit date heuristic code entirely.\n\n> I'll give you a few months to rethink at this statement until this\n> feature does get used widely. I think there was never a moment where\n> we would ever think to rebuild older commits as this would break the\n> hash of the commits where many people are potential looking for.\n\nI'm afraid that your English grammar is sufficiently mangled here that\nI don't understand *your* point.  Which is a shame because it's\none of my more important points.\n\nStoring the generation number inside the commit means that a commit\nwith a generation number has a different hash than a commit without one.\nThis means that people won't want to break the hashes of existing commits\nby adding them.  In many cases, ever.\n\nWhich means that git will have to be able to work without the generation\nnumbers forever.\n\nIf the generation numbers are stored in a separate data structure that\ncan be added to an existing repository, then a new version of git can\ndo that when needed.  Which lets git depend on always having the the\ngeneration numbers to do all history walking and stop using commit date\nbased heuristics completely.\n\n>> 4) It can't be made to work with grafts or replace objects.\n>>\n>> 5) It includes information which is redundant, but hard to verify,\n>>   in git objects.  Leading to potentially bizarre and version-dependent\n>>   behaviour if it's wrong.  (Checking that the numbers are consistent\n>>   is the same work as regenerating a cache.)\n\n> The data is *consistent* as long as the hash doesn't change, storing the\n> data in the commits *can* reduce resource and makes calculations cheaper.\n\nYou're mixing up two issues.  Storing the generation number *anywhere*\ncan make calculations cheaper.  Storing them in the commit is indeed the\n*simplest* place, but the calculation cost point is equally true if the\nnumbers are stored somewhere else.\n\nAs for consistency...\n\nI'm defining \"consistent\" as consistency between the generation number\nand the parent pointers.  This is the property that the history-walking\noptimizations depend on.\n\nA commit's generation number is consistent if it is larger than the\ngeneration number of any of its parents.  (Optionally, you\nmay require that it be larger by exatly 1.)\n\nA generation number is *not* consistent if is less than or equal to the\ngeneration number of one of its parents.\n\nIf this happens, history walking code that uses the generation numbers\nwill not produce correct output.\n\nFurther, the nature of the incorrectness will depend on implementation\ndetails (\"potentially bizarre and version-dependent behaviour\") of the\nhistory-walking code.\n\nBy computing the generation numbers when needed, the entire \"what happens\nif someone makes a commit with an inconsistent generation number\"\nproblem goes away.  It goes from \"not likely to happen\" or \"somthing\nthat has to be checked for when receiving objects\" to \"can't happen\".\n\nThe computation to verify that an incoming commit's generation number\nis consistent is exactly the same computation needed to compute the\ngeneration number it should have: look up all parent commit generation\nnumbers and take the maximum.  The only question is whether we store\nthe result after computing it, or compare with the included generation\nnumber and possibly print an error message.\n\n\nFor example, suppose I generate a commit with a generation number of\nUINT_MAX.  Will this crash git?  That's a new error condition the code\nhas to worry about.  If I generate the generation number locally, I know\nthat can't happen in any repository that I can download in a reasonable\nperiod of time.\n\nIf we had generation numbers from day 1, we could just require that they\nalways be checked, and an inconsistent object could be always rejected.\n\nBut since old git versions ignore the generation number in commits, a\nbad generation number could spread a long way before someone notices it.\nIt becomes a visible problem.  Not a really big one (I'm pretty sure\nthat refusing to pull it introduces no security holes), but it's an\nerror condition that we have to actually think about.\n\n\n> A cache would use more resources because they can become invalid at any\n> point and *should* be recalculated by every client. We are processing\n> data that *can* be reused by everybody with a git client which has this\n> specific feature, but does not break anything with an older client.\n>\n> So please, calculate things only once as this may save a *lot* of time :-)\n\nThis is silly.  The cache can't become invalid except by disk corruption,\nwhich can corrupt numbers stored in the commit object just the same.\n(The corruption can be detected by git-fsck, but that's also true\nindependent of where the numbers are stored.)\n\nAnd the work to recalculate the numbers is far less than the work to\ngarbage collect, or repack, or generate the index of an incoming pack,\nor any of a dozen operations that are normally done by all clients.\n(Don't get me started on rename detection!)\n\nThis is a completely misplaced optimization.  Walking every commit in\nthe repository takes a few seconds and enough memory that we don't want\nto do it every \"git log\" operation, but it's barely perceptible compared\nto other repository maintenance operations.\n\nDo it once when you install a new git software version and then you can\nforget about it.\n\n> I would see more advantage in a cache if the data could differs on\n> every client, but that still doesn't mean that you should use one.\n\nIf you use grafts or replace objects, it can be.  That's my point 4)\nabove.  Supporting these makes maintaining a cache trickier, but it's\nsimply impossible to do with in-commit generation numbers.\n"},{"id":"171746","messageId":"alpine.LFD.2.00.1107201538590.21187@xanadu.home","threadId":"27834","inReplyTo":"20110718114834.12406.qmail@science.horizon.com","subject":"Re: Git commit generation numbers","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-07-20T20:51:40Z","receivedAt":"2011-07-20T20:51:40Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 18 Jul 2011, George Spelvin wrote:\n\n> Storing the generation number inside the commit means that a commit\n> with a generation number has a different hash than a commit without one.\n> This means that people won't want to break the hashes of existing commits\n> by adding them.  In many cases, ever.\n> \n> Which means that git will have to be able to work without the generation\n> numbers forever.\n\nI've been diverting myself from $day_job by reading through this thread.  \nStill, I couldn't make my mind between having the generation number \nstored in the commit object or in a separate cache by reading all the \narguments for each until now. Admittedly I'm not as involved in the \ndesign of Git as I once was, so my comments can be considered with the \nsame proportions.\n\nObviously, with a perfect design, we would have had gen numbers from the \nbeginning.  But we did mistakes, and now have to regret and live with \nthem (and yes I have my own share of responsibility for some of those \nregrets which are now embodied in the Git data format).\n\n> If the generation numbers are stored in a separate data structure that\n> can be added to an existing repository, then a new version of git can\n> do that when needed.  Which lets git depend on always having the the\n> generation numbers to do all history walking and stop using commit date\n> based heuristics completely.\n\nTo me this is the killer argument.  Being able to forget about the \nbroken date heuristics entirely and simplify the code is what makes the \nexternal cache so fundamentally better as it can be applied to any \nexisting repositories.  And it has no backward compatibility issues as \nold Git version won't work any worse if they can't make any usage of \nthat cache.\n\nThe alternative of having to sometimes use the generation number, \nsometimes use the possibly broken commit date, makes for much more \ncomplicated code that has to be maintained forever.  Having a solution \nthat starts working only after a certain point in history doesn't look \neleguant to me at all.  It is not like having different pack formats \nwhere back and forth conversions can be made for the _entire_ history.\n\nAnd if you don't care about graft/replace then the cached data is \nimmutable just like the in-commit version would, so there is no \nconsistency issues.  If you do care about graft/replace (or who knows \nwhat other dag alteration scheme might be created in 5 years from now) \nthen a separate cache will be required _anyway_, regardless of any \nin-commit gen number.\n\nSo to say that if a generation number is _really_ needed, then it should \ngo in a separate cache.  Saying that if we would have done it initially \nthen it would have been inside the commit object is not a good enough \njustification to do it today if it can't be applied to the whole of \nalready existing repositories and avoid special cases.\n\nI however have not formed any opinion on that fundamental question i.e. \nwhether or not gen numbers are worth it in today's conditions. Neither \ndid I think about the actual cache format (I don't think that adding it \nto the pack index is a good idea if grafts are to be honored) which \ncertainly has bearing on that fundamental question too.\n\nBut I don't see the point of starting to add them now to commit objects, \neven if we regret not doing it initially, simply because having them \nappear randomly based on the Git version/implementation being used is \nstill much uglier than some ad hoc cache or even not having them at all.\n\n\nNicolas\n"},{"id":"171751","messageId":"20110720221632.14223.qmail@science.horizon.com","threadId":"27834","inReplyTo":"alpine.LFD.2.00.1107201538590.21187@xanadu.home","subject":"Re: Git commit generation numbers","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2011-07-20T22:16:32Z","receivedAt":"2011-07-20T22:16:32Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> The alternative of having to sometimes use the generation number, \n> sometimes use the possibly broken commit date, makes for much more \n> complicated code that has to be maintained forever.  Having a solution \n> that starts working only after a certain point in history doesn't look \n> eleguant to me at all.  It is not like having different pack formats \n> where back and forth conversions can be made for the _entire_ history.\n\nIt seemed like a pretty strong argument to me, too.\n\n> And if you don't care about graft/replace then the cached data is \n> immutable just like the in-commit version would, so there is no \n> consistency issues.  If you do care about graft/replace (or who knows \n> what other dag alteration scheme might be created in 5 years from now) \n> then a separate cache will be required _anyway_, regardless of any \n> in-commit gen number.\n\nA possible workaround would be to keep track of the largest generation\nnumber skew introduced by any graft, and add that safety factor into\nthe history-walking code, but that would be painful if you replace a\nsingle large commit with an equivalent long development history, such\nas adding a historical development tree behind a recently-cut-off one.\nor development history You can do a workaround at the expense of ine\n\n> Neither did I think about the actual cache format (I don't think that\n> adding it to the pack index is a good idea if grafts are to be honored)\n> which certainly has bearing on that fundamental question too.\n\nI was thinking of something very close to the V2 pack format.\nhttp://book.git-scm.com/7_the_packfile.html\nA magic number, a 256-entry fanout table, a sorted list of 20-byte hashes,\nfollowed by a matching list of 4-byte generation numbers.\n\nEnding with a 20-byte hash of the replaces and grafts state that this\ncache is valid for, and a hash of the cache itself.\n\nA bit of code factoring should make it easy to share much of the code.\n\n\nIt would certainly be possible to share the SHA1 table in an existing\npack index and store the generation numbers of the base (no replacement)\ncase, but you'd have to store null values for all the non-commit objects.\n\nThat takes 4 bytes per object, while a separate list of commits\ntakes 24 bytes per commit.  A separate list is better if commits\nare less than 1/6 of all objects.\n\nLooking at git's own object database, we have:\n 66125 blobs   (45.50%)\n 49292 trees   (33.92%)\n 29554 commits (20.33%)\n   362 tags    ( 0.25%)\n145333 total\n\nSo we're actually a bit over the 16.66% optimum. but it's not far enough\nto be a real efficiency problem.\n"},{"id":"171761","messageId":"alpine.DEB.2.02.1107201624510.5222@asgard.lang.hm","threadId":"27834","inReplyTo":"20110720221632.14223.qmail@science.horizon.com","subject":"Re: Git commit generation numbers","fromName":"","fromEmail":"david@lang.hm","sentAt":"2011-07-20T23:26:38Z","receivedAt":"2011-07-20T23:26:38Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Wed, 20 Jul 2011, George Spelvin wrote:\n\n>> The alternative of having to sometimes use the generation number,\n>> sometimes use the possibly broken commit date, makes for much more\n>> complicated code that has to be maintained forever.  Having a solution\n>> that starts working only after a certain point in history doesn't look\n>> eleguant to me at all.  It is not like having different pack formats\n>> where back and forth conversions can be made for the _entire_ history.\n>\n> It seemed like a pretty strong argument to me, too.\n\nexcept that you then have different caches on different systems. If the \ngeneration number is part of the repository then it's going to be the same \nfor everyone.\n\nin either case, you still have the different heristics depending on what \nversion of git someone is running\n\nDavid Lang\n"},{"id":"171762","messageId":"alpine.LFD.2.00.1107201931510.21187@xanadu.home","threadId":"27834","inReplyTo":"alpine.DEB.2.02.1107201624510.5222@asgard.lang.hm","subject":"Re: Git commit generation numbers","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-07-20T23:36:55Z","receivedAt":"2011-07-20T23:36:55Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 20 Jul 2011, david@lang.hm wrote:\n\n> On Wed, 20 Jul 2011, George Spelvin wrote:\n> \n> > > The alternative of having to sometimes use the generation number,\n> > > sometimes use the possibly broken commit date, makes for much more\n> > > complicated code that has to be maintained forever.  Having a solution\n> > > that starts working only after a certain point in history doesn't look\n> > > eleguant to me at all.  It is not like having different pack formats\n> > > where back and forth conversions can be made for the _entire_ history.\n> > \n> > It seemed like a pretty strong argument to me, too.\n> \n> except that you then have different caches on different systems.\n\nSo what?\n\n> If the generation number is part of the repository then it's going to \n> be the same for everyone.\n\nThe actual generation number will be, and has to be, the same for \neveryone with the same repository content, regardless of the cache used.  \nIt is a well defined number with no room to interpretation.\n\n> in either case, you still have the different heristics depending on what\n> version of git someone is running\n\nIndeed.\n\n\nNicolas\n"},{"id":"171763","messageId":"4E276DF8.8030301@cisco.com","threadId":"27834","inReplyTo":"alpine.LFD.2.00.1107201931510.21187@xanadu.home","subject":"Re: Git commit generation numbers","fromName":"Phil Hord","fromEmail":"hordp@cisco.com","sentAt":"2011-07-21T00:08:24Z","receivedAt":"2011-07-21T00:08:24Z","isPatch":false,"sender":{"key":"phil.hord@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123908?v=4"},"body":"On 07/20/2011 07:36 PM, Nicolas Pitre wrote:\n> On Wed, 20 Jul 2011, david@lang.hm wrote:\n>\n>> If the generation number is part of the repository then it's going to\n>> be the same for everyone.\n> The actual generation number will be, and has to be, the same for\n> everyone with the same repository content, regardless of the cache used.\n> It is a well defined number with no room to interpretation.\n\nNonsense.\n\nEven if the generation number is well-defined and shared by all clients, \nthe only quasi-essential definition is \"for each A in ancestors_of(B), \ngen(A) < gen(B)\".\n\nIn practice, the actual generation number *will be the same* for \neveryone with the same repository content, unless and until someone \ndevelops a different calculation method.  But there is no reason to \nrequire that the number *has to be* the same for everyone unless you \nexpect (or require) everyone to share their gen-caches.\n\nSurely there will be a competent and efficient gen-cache API.  But most \ncode can just ask if B --contains A or even just use rev-list and \nbenefit from the increased speed of the answer.  Because most code \ndoesn't really care about the gen numbers themselves, but only the speed \nof determining ancestry.\n\nPhil\n"},{"id":"171764","messageId":"alpine.DEB.2.02.1107201714140.6412@asgard.lang.hm","threadId":"27834","inReplyTo":"4E276DF8.8030301@cisco.com","subject":"Re: Git commit generation numbers","fromName":"","fromEmail":"david@lang.hm","sentAt":"2011-07-21T00:18:28Z","receivedAt":"2011-07-21T00:18:28Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Wed, 20 Jul 2011, Phil Hord wrote:\n\n> On 07/20/2011 07:36 PM, Nicolas Pitre wrote:\n>> On Wed, 20 Jul 2011, david@lang.hm wrote:\n>> \n>>> If the generation number is part of the repository then it's going to\n>>> be the same for everyone.\n>> The actual generation number will be, and has to be, the same for\n>> everyone with the same repository content, regardless of the cache used.\n>> It is a well defined number with no room to interpretation.\n>\n> Nonsense.\n>\n> Even if the generation number is well-defined and shared by all clients, the \n> only quasi-essential definition is \"for each A in ancestors_of(B), gen(A) < \n> gen(B)\".\n>\n> In practice, the actual generation number *will be the same* for everyone \n> with the same repository content, unless and until someone develops a \n> different calculation method.  But there is no reason to require that the \n> number *has to be* the same for everyone unless you expect (or require) \n> everyone to share their gen-caches.\n\nand I think this is why Linus is not happy with a cache. He is seeing this \nas something that has significantly more value if it is going to be \nconsistant in a distributed manner than if it's just something calculated \nlocally that can be different from other systems.\n\nif it's just locally generated, then I could easily see generation numbers \nbeing different on different people's ssstems, dependin on the order that \nthey see commits (either locally generated or pulled from others)\n\nIf it's part of the commit, then as that commit gets propogated the \ngeneration number gets propogated as well, and every repository will agree \non what the generation number is for any commit that's shared.\n\nI agree that this consistancy guarantee seems to be valuable.\n\n> Surely there will be a competent and efficient gen-cache API.  But most code \n> can just ask if B --contains A or even just use rev-list and benefit from the \n> increased speed of the answer.  Because most code doesn't really care about \n> the gen numbers themselves, but only the speed of determining ancestry.\n\nin that case, why bother with generation numbers at all? the improved data \nbased heristic seems to solve that problem.\n\nDavid Lang\n"},{"id":"171765","messageId":"CAJo=hJuS_iYSS8iVWoJ1BiUANsGtYJoYm-WRa863isVNsq=5vw@mail.gmail.com","threadId":"27834","inReplyTo":"alpine.DEB.2.02.1107201714140.6412@asgard.lang.hm","subject":"Re: Git commit generation numbers","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-07-21T00:37:17Z","receivedAt":"2011-07-21T00:37:17Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Wed, Jul 20, 2011 at 17:18,  <david@lang.hm> wrote:\n>\n> if it's just locally generated, then I could easily see generation numbers\n> being different on different people's ssstems, dependin on the order that\n> they see commits (either locally generated or pulled from others)\n\nBut this should only happen if the user fudges with their Git sources\nand makes Git produce a different generation number.\n\nIf the algorithm is always \"gen(A) = max(gen(P) for each parent_of(A))\n+ 1\" then it doesn't matter who merged what commits, the same commit\nappears at the same part of the graph relative to all of its\nancestors, and therefore always has the same generation number. This\nis true whether or not the commit contains the generation number.\n\n> If it's part of the commit, then as that commit gets propogated the\n> generation number gets propogated as well, and every repository will agree\n> on what the generation number is for any commit that's shared.\n\nThis isn't really as beneficial as you are making it out to be. We\nalready can agree on what the generation number should be for any\ngiven commit, if you topo-sort the commit DAG, you get the same\nresult.\n\n> I agree that this consistancy guarantee seems to be valuable.\n\nIts valuable, but its consistent either with a cache, or not.\n\n-- \nShawn.\n"},{"id":"171766","messageId":"4E277540.7080408@cisco.com","threadId":"27834","inReplyTo":"alpine.DEB.2.02.1107201714140.6412@asgard.lang.hm","subject":"Re: Git commit generation numbers","fromName":"Phil Hord","fromEmail":"hordp@cisco.com","sentAt":"2011-07-21T00:39:28Z","receivedAt":"2011-07-21T00:39:28Z","isPatch":false,"sender":{"key":"phil.hord@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123908?v=4"},"body":"On 07/20/2011 08:18 PM, david@lang.hm wrote:\n> On Wed, 20 Jul 2011, Phil Hord wrote:\n>\n>> On 07/20/2011 07:36 PM, Nicolas Pitre wrote:\n>>> On Wed, 20 Jul 2011, david@lang.hm wrote:\n>>>\n>>>> If the generation number is part of the repository then it's going to\n>>>> be the same for everyone.\n>>> The actual generation number will be, and has to be, the same for\n>>> everyone with the same repository content, regardless of the cache \n>>> used.\n>>> It is a well defined number with no room to interpretation.\n>>\n>> Nonsense.\n>>\n>> Even if the generation number is well-defined and shared by all \n>> clients, the only quasi-essential definition is \"for each A in \n>> ancestors_of(B), gen(A) < gen(B)\".\n>>\n>> In practice, the actual generation number *will be the same* for \n>> everyone with the same repository content, unless and until someone \n>> develops a different calculation method.  But there is no reason to \n>> require that the number *has to be* the same for everyone unless you \n>> expect (or require) everyone to share their gen-caches.\n>\n> and I think this is why Linus is not happy with a cache. He is seeing \n> this as something that has significantly more value if it is going to \n> be consistant in a distributed manner than if it's just something \n> calculated locally that can be different from other systems.\n\nIt will only be used locally, so it needn't be consistent with anyone \nelse's.\n\n>\n> if it's just locally generated, then I could easily see generation \n> numbers being different on different people's ssstems, dependin on the \n> order that they see commits (either locally generated or pulled from \n> others)\n>\n> If it's part of the commit, then as that commit gets propogated the \n> generation number gets propogated as well, and every repository will \n> agree on what the generation number is for any commit that's shared.\n>\n> I agree that this consistancy guarantee seems to be valuable.\n\nI can't see why.\n\n>> Surely there will be a competent and efficient gen-cache API.  But \n>> most code can just ask if B --contains A or even just use rev-list \n>> and benefit from the increased speed of the answer.  Because most \n>> code doesn't really care about the gen numbers themselves, but only \n>> the speed of determining ancestry.\n>\n> in that case, why bother with generation numbers at all? the improved \n> data based heristic seems to solve that problem.\n\nDoes it?  Surely the ruckus would've died down in that case.  But I \nhaven't been reading pu.\n\nIt seems to me that the main drawback to a gen-cache is that it slows \ndown the first operation after even a local clone (with just hardlinks).\n\nOn the other hand, I see too many nails in the distributed-gen-numbers \ncoffin:  legacy commits can't catch up (and therefore suffer), and \nlegacy clients can trash or corrupt even \"new-style\" commits.\n\nPhil\n"},{"id":"171767","messageId":"4E27772B.60306@cisco.com","threadId":"27834","inReplyTo":"CAJo=hJuS_iYSS8iVWoJ1BiUANsGtYJoYm-WRa863isVNsq=5vw@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Phil Hord","fromEmail":"hordp@cisco.com","sentAt":"2011-07-21T00:47:39Z","receivedAt":"2011-07-21T00:47:39Z","isPatch":false,"sender":{"key":"phil.hord@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123908?v=4"},"body":"\nOn 07/20/2011 08:37 PM, Shawn Pearce wrote:\n> On Wed, Jul 20, 2011 at 17:18,<david@lang.hm>  wrote:\n>> if it's just locally generated, then I could easily see generation numbers\n>> being different on different people's ssstems, dependin on the order that\n>> they see commits (either locally generated or pulled from others)\n> But this should only happen if the user fudges with their Git sources\n> and makes Git produce a different generation number.\n>\n> If the algorithm is always \"gen(A) = max(gen(P) for each parent_of(A))\n> + 1\" then it doesn't matter who merged what commits, the same commit\n> appears at the same part of the graph relative to all of its\n> ancestors, and therefore always has the same generation number. This\n> is true whether or not the commit contains the generation number.\n\nInteresting.  I was going to disagree with the latter part of your \nstatement, but then I realized you're right.\n\nAnd that your algorithm allows duplicate generation numbers.\n\nAnd that there's nothing wrong with that.\n\nBecause it meets the one quasi-essential need, \"for each A in \nancestors_of(B), gen(A) < gen(B)\".\n\n>> If it's part of the commit, then as that commit gets propogated the\n>> generation number gets propogated as well, and every repository will agree\n>> on what the generation number is for any commit that's shared.\n> This isn't really as beneficial as you are making it out to be. We\n> already can agree on what the generation number should be for any\n> given commit, if you topo-sort the commit DAG, you get the same\n> result.\n>\n>> I agree that this consistancy guarantee seems to be valuable.\n> Its valuable, but its consistent either with a cache, or not.\n\nI still fail to see the value.\n\nPhil\n"},{"id":"171768","messageId":"alpine.LFD.2.00.1107202041480.21187@xanadu.home","threadId":"27834","inReplyTo":"4E276DF8.8030301@cisco.com","subject":"Re: Git commit generation numbers","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-07-21T00:58:38Z","receivedAt":"2011-07-21T00:58:38Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 20 Jul 2011, Phil Hord wrote:\n\n> On 07/20/2011 07:36 PM, Nicolas Pitre wrote:\n> > On Wed, 20 Jul 2011, david@lang.hm wrote:\n> > \n> > > If the generation number is part of the repository then it's going to\n> > > be the same for everyone.\n> > The actual generation number will be, and has to be, the same for\n> > everyone with the same repository content, regardless of the cache used.\n> > It is a well defined number with no room to interpretation.\n> \n> Nonsense.\n> \n> Even if the generation number is well-defined and shared by all clients, the\n> only quasi-essential definition is \"for each A in ancestors_of(B), gen(A) <\n> gen(B)\".\n\nSure.  But what do you gain by making holes in the sequence?\n\n> In practice, the actual generation number *will be the same* for everyone with\n> the same repository content, unless and until someone develops a different\n> calculation method.  But there is no reason to require that the number *has to\n> be* the same for everyone unless you expect (or require) everyone to share\n> their gen-caches.\n\nAnd with the above you clearly reinforced the argument _against_ storing \nthe generation number in the commit object.  If you can imagine a \ndifferent calculation method already, and if it is actually useful, then \nwho knows if something even better could be done eventually.\n\n\nNicolas\n"},{"id":"171769","messageId":"4E277C2E.2020402@cisco.com","threadId":"27834","inReplyTo":"alpine.LFD.2.00.1107202041480.21187@xanadu.home","subject":"Re: Git commit generation numbers","fromName":"Phil Hord","fromEmail":"hordp@cisco.com","sentAt":"2011-07-21T01:09:02Z","receivedAt":"2011-07-21T01:09:02Z","isPatch":false,"sender":{"key":"phil.hord@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123908?v=4"},"body":"On 07/20/2011 08:58 PM, Nicolas Pitre wrote:\n> On Wed, 20 Jul 2011, Phil Hord wrote:\n>\n>> On 07/20/2011 07:36 PM, Nicolas Pitre wrote:\n>>> On Wed, 20 Jul 2011, david@lang.hm wrote:\n>>>\n>>>> If the generation number is part of the repository then it's going to\n>>>> be the same for everyone.\n>>> The actual generation number will be, and has to be, the same for\n>>> everyone with the same repository content, regardless of the cache used.\n>>> It is a well defined number with no room to interpretation.\n>> Nonsense.\n>>\n>> Even if the generation number is well-defined and shared by all clients, the\n>> only quasi-essential definition is \"for each A in ancestors_of(B), gen(A)<\n>> gen(B)\".\n> Sure.  But what do you gain by making holes in the sequence?\n\nDepends on the algorithm.  Probably speed.  Possibly more efficient \nlimited-cache building (jit-style discovery in reverse, as-needed, for \nexample).\n\nWhat do you gain by enforcing contiguousness?  Why not require all gen \nnumbers to be even?  Or prime?  ;)\n\n>> In practice, the actual generation number *will be the same* for everyone with\n>> the same repository content, unless and until someone develops a different\n>> calculation method.  But there is no reason to require that the number *has to\n>> be* the same for everyone unless you expect (or require) everyone to share\n>> their gen-caches.\n> And with the above you clearly reinforced the argument _against_ storing\n> the generation number in the commit object.  If you can imagine a\n> different calculation method already, and if it is actually useful, then\n> who knows if something even better could be done eventually.\n\nGood.  Nice to see I'm being self-consistent, then.\n\nPhil\n"},{"id":"171784","messageId":"alpine.DEB.2.02.1107202119440.5355@asgard.lang.hm","threadId":"27834","inReplyTo":"CAJo=hJuS_iYSS8iVWoJ1BiUANsGtYJoYm-WRa863isVNsq=5vw@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"","fromEmail":"david@lang.hm","sentAt":"2011-07-21T04:26:43Z","receivedAt":"2011-07-21T04:26:43Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Wed, 20 Jul 2011, Shawn Pearce wrote:\n\n> On Wed, Jul 20, 2011 at 17:18,  <david@lang.hm> wrote:\n>>\n>> if it's just locally generated, then I could easily see generation numbers\n>> being different on different people's ssstems, dependin on the order that\n>> they see commits (either locally generated or pulled from others)\n>\n> But this should only happen if the user fudges with their Git sources\n> and makes Git produce a different generation number.\n>\n> If the algorithm is always \"gen(A) = max(gen(P) for each parent_of(A))\n> + 1\" then it doesn't matter who merged what commits, the same commit\n> appears at the same part of the graph relative to all of its\n> ancestors, and therefore always has the same generation number. This\n> is true whether or not the commit contains the generation number.\n\nI have to think about this more, but I'm wondering about cases where the \nsame result ia achieved via different methods, something along the lines \nof one person developing something with _many_ commits (creating a large \ngeneration number) that one person merges far sooner than another, causing \nthe commits that they do after the merge to have much larger generation \nnumbers than someone making the same changes, but doing the merge later\n\nsomething like\n\n   C9\n    \\\nC2 - C10 - C11 - C12\n\nvs\n                 C9\n                   \\\nC2 - C3 - C4 - C5 - C10\n\nwhere the C10-12 in the first set and C3-5 in the second set are \ncompletely unrelated to what's done in C9 and C12 in the first set and C10 \nin the sedond set are identical trees.\n\nnow I know that part of a commit is what it's parents are, so that is \ndifferent (and that may be enough to say that generations don't matter \nand this entire issue is moot), but I haven't thought about it long enough \nto convince myself what would (or should) happen in these cases.\n\nDavid Lang\n\n>> If it's part of the commit, then as that commit gets propogated the\n>> generation number gets propogated as well, and every repository will agree\n>> on what the generation number is for any commit that's shared.\n>\n> This isn't really as beneficial as you are making it out to be. We\n> already can agree on what the generation number should be for any\n> given commit, if you topo-sort the commit DAG, you get the same\n> result.\n>\n>> I agree that this consistancy guarantee seems to be valuable.\n>\n> Its valuable, but its consistent either with a cache, or not.\n>\n>\n"},{"id":"171803","messageId":"20110721124351.25143.qmail@science.horizon.com","threadId":"27834","inReplyTo":"alpine.DEB.2.02.1107202119440.5355@asgard.lang.hm","subject":"Re: Git commit generation numbers","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2011-07-21T12:43:51Z","receivedAt":"2011-07-21T12:43:51Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"On <david@lang.hm> wrote:\n> On Wed, 20 Jul 2011, Shawn Pearce wrote:\n>> If the algorithm is always \"gen(A) = max(gen(P) for each parent_of(A))\n>> + 1\" then it doesn't matter who merged what commits, the same commit\n>> appears at the same part of the graph relative to all of its\n>> ancestors, and therefore always has the same generation number. This\n>> is true whether or not the commit contains the generation number.\n\n> I have to think about this more, but I'm wondering about cases where the \n> same result ia achieved via different methods, something along the lines \n> of one person developing something with _many_ commits (creating a large \n> generation number) that one person merges far sooner than another, causing \n> the commits that they do after the merge to have much larger generation \n> numbers than someone making the same changes, but doing the merge later\n\nCan't happen.  Using the basic algorithm as Shawn described, the\ngeneration number is defined uniquely by the ancestor DAG.\n\nThe generation number is the length of the longest path to a\nroot (zero-ancestor) commit through the DAG.\n\nIf you look at past discussion, several people have thought it was\nokay to bake into the commit precsiely because it can be computed\nonce and will never change.\n\nHowever, git does have some ability to amend the history DAG after\nit's been written, using grafts and replace objects.  These can\nchange generation numbers, presisely because they change the DAG.\n\n> something like\n> \n>    C9\n>     \\\n> C2 - C10 - C11 - C12\n> \n> vs\n>                  C9\n>                    \\\n> C2 - C3 - C4 - C5 - C10\n>\n> where the C10-12 in the first set and C3-5 in the second set are\n> completely unrelated to what's done in C9 and C12 in the first set\n> and C10 in the second set are identical trees.\n\nThe generation numbers in the above are as follows:\nFirst example:\n\tC2 = C9 = 0\n\tC10 = 1 = max(C2, C9) + 1\n\tC11 = 2 = C10 + 1\n\tC12 = 3 = C11 + 1\n\nSecond example:\n\tC2 = C9 = 0\n\tC3 = 1 = C2 + 1\n\tC4 = 2 = C2 + 1\n\tC5 = 3 = C4 + 1\n\tC10 = 4 = max(C5, C9) + 1\n\nNow, the history pruning works fine if the \"+1\" is replaced my any other\nnon-zero increment, but it's not clear why you'd bother.\n"},{"id":"171812","messageId":"m3mxg7sasa.fsf@localhost.localdomain","threadId":"27834","inReplyTo":"20110721124351.25143.qmail@science.horizon.com","subject":"Re: Git commit generation numbers","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2011-07-21T19:19:52Z","receivedAt":"2011-07-21T19:19:52Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"George Spelvin, could you please try not mangle CC to include only\nemails, stripping names (e.g. \"spearce@spearce.org\" instead of\n\"Shawn Pearce <spearce@spearce.org>\")?\n\n\"George Spelvin\" <linux@horizon.com> writes:\n> On <david@lang.hm> wrote:\n>> On Wed, 20 Jul 2011, Shawn Pearce wrote:\n\n>>> If the algorithm is always \"gen(A) = max(gen(P) for each parent_of(A))\n>>> + 1\" then it doesn't matter who merged what commits, the same commit\n>>> appears at the same part of the graph relative to all of its\n>>> ancestors, and therefore always has the same generation number. This\n>>> is true whether or not the commit contains the generation number.\n> \n>> I have to think about this more, but I'm wondering about cases where the \n>> same result ia achieved via different methods, something along the lines \n>> of one person developing something with _many_ commits (creating a large \n>> generation number) that one person merges far sooner than another, causing \n>> the commits that they do after the merge to have much larger generation \n>> numbers than someone making the same changes, but doing the merge later\n> \n> Can't happen.  Using the basic algorithm as Shawn described, the\n> generation number is defined uniquely by the ancestor DAG.\n> \n> The generation number is the length of the longest path to a\n> root (zero-ancestor) commit through the DAG.\n> \n> If you look at past discussion, several people have thought it was\n> okay to bake into the commit precsiely because it can be computed\n> once and will never change.\n> \n> However, git does have some ability to amend the history DAG after\n> it's been written, using grafts and replace objects.  These can\n> change generation numbers, presisely because they change the DAG.\n\nThere is also another issue that I have mentioned, namely incomplete\nclones - which currently means shallow clone, without access to full\nhistory.\n\n\nNb. grafts are so horrible hack that I would be not against turning\noff generation numbers if they are used.\n\nIn the case of replace objects you need both non-replaced and replaced\nDAG generation numbers.\n\n-- \nJakub Narębski\n"},{"id":"171817","messageId":"20110721202722.3765.qmail@science.horizon.com","threadId":"27834","inReplyTo":"m3mxg7sasa.fsf@localhost.localdomain","subject":"Re: Git commit generation numbers","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2011-07-21T20:27:22Z","receivedAt":"2011-07-21T20:27:22Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> There is also another issue that I have mentioned, namely incomplete\n> clones - which currently means shallow clone, without access to full\n> history.\n\nAs far as history walking is concerned, you can just consider \"missing\nparent\" the same as \"no parent\" and start the generation numbers at 0.\nAs long as you recompute\n\n> Nb. grafts are so horrible hack that I would be not against turning\n> off generation numbers if they are used.\n\nYeah, but it's not too miserable to add support (the logic is very similar\nto replace objects), and then you would be able to have the history walking\ncode depend on the presence of generation numbers.  (The \"load the cache\"\nfunction would regenerate it if necessary.)\n\nOnly do this if you already have support for \"no generation numbers\" in\nthe history walking code for (say) loose objects.\n\n> In the case of replace objects you need both non-replaced and replaced\n> DAG generation numbers.\n\nYes, the cache validity/invalidation criteria are the tricky bit.\nHonestly, this is where the code gets ugly, not computing and storing\nthe generation numbers.\n\n\nOne thought on an expanded generation number cache:\n\nThere are many git operations that use ONLY the commit DAG, and do not\nactually use any information from the commits other than their hashes\nand parent pointers.  The ones that come to mind are rev-parse, rev-list,\ndescribe, name-rev, and merge-base.\n\nThese could be sped up if, instead of just generation numbers, we kept\na complete cached copy of the commit DAG, so the commit objects didn't\nhave to be uncompressed and parsed.\n\nThis could be provided by an extended form of generation number cache.\nIn addition to listing the generation number of each commit, it\nwould list all the ancestors (by file offset rather than hash, for\ncompactness).  Then simple commit walking could load this cache and\navoid unpacking commit objects from packs.\n\nA compact implementation would abuse the flexibility of generation numbers\nto make them serve double duty.  They would be used as offsets into a\ntable of parent pointers.  By keeping the table topologically sorted,\nthe offsets would satisfy the requirements for generation numbers, but\nwould be unique, and there would be additional gaps when a commit had\nmultiple parents.\n\nThe parent pointers would themselves be 31-bit offsets into the table of\nSHA-1 hashes, with the msbit meaning \"this commit has multiple parents,\nalso look at the following table entry\".  (If we use offset 0 to mean\n\"no parents\", it might be more convenient to have the offset point to\nthe *end* of the run of parents rather than the beginning, so \"following\"\nwould be earlier in the file, but that's an implementation detail.)\n\nI'm assuming that 2^31 commits having (in aggregate) 2^32 parents would\nbe enough for the time being.  As a local cache, it can be extended\nwith a software upgrade.  There's no need to ever have support for two\nformats in any given release; just notice that the cache format is wrong,\nblow it away, and regenerate it.\n"},{"id":"171819","messageId":"CAJo=hJsi-cVdNC+T4RiwXdsjypL3FPc8gPOL5qL9=tGj05xPGg@mail.gmail.com","threadId":"27834","inReplyTo":"20110721202722.3765.qmail@science.horizon.com","subject":"Re: Git commit generation numbers","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-07-21T20:33:21Z","receivedAt":"2011-07-21T20:33:21Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Thu, Jul 21, 2011 at 13:27, George Spelvin <linux@horizon.com> wrote:\n>\n> be enough for the time being.  As a local cache, it can be extended\n> with a software upgrade.  There's no need to ever have support for two\n> formats in any given release; just notice that the cache format is wrong,\n> blow it away, and regenerate it.\n\nDon't assume that. Consider a repository stored on NFS that is\nread-only to you. The NFS server has one version of Git installed, and\nis using cache format A. You have a newer version of Git installed on\nyour workstation, using cache format B. Now you cannot use this\nrepository as a local filesystem... its only available to you over the\nGit protocols. This breaks a number of people's environments.  :-)\n\nIts better if we can avoid having to change file formats very often,\neven if they are a local \"cache\".\n\n-- \nShawn.\n"},{"id":"171837","messageId":"201107221418.52414.jnareb@gmail.com","threadId":"27834","inReplyTo":"20110721202722.3765.qmail@science.horizon.com","subject":"Re: Git commit generation numbers","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2011-07-22T12:18:51Z","receivedAt":"2011-07-22T12:18:51Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Thu, 21 Jul 2011, George Spelvin wrote:\n\n> > There is also another issue that I have mentioned, namely incomplete\n> > clones - which currently means shallow clone, without access to full\n> > history.\n> \n> As far as history walking is concerned, you can just consider \"missing\n> parent\" the same as \"no parent\" and start the generation numbers at 0.\n> As long as you recompute.\n\nWell, shallow clone case can be considered both for putting 'true'\ngeneration numbers in commit header, and against it.\n\nFor, because with generation numbers in commits you can use true\ngeneration numbers.\n\nAgainst, because if there are commits without generation numbers in\nheader, you cannot assign true generation number, and you can only use\n\"shallow\" generation number, in generation numbers cache.\n\n> > Nb. grafts are so horrible hack that I would be not against turning\n> > off generation numbers if they are used.\n> \n> Yeah, but it's not too miserable to add support (the logic is very similar\n> to replace objects), and then you would be able to have the history walking\n> code depend on the presence of generation numbers.  (The \"load the cache\"\n> function would regenerate it if necessary.)\n> \n> Only do this if you already have support for \"no generation numbers\" in\n> the history walking code for (say) loose objects.\n\nGrafts are non-transferable, and if you use them to cull rather than add\nhistory they are unsafe against garbage collection... I think.\n \n> > In the case of replace objects you need both non-replaced and replaced\n> > DAG generation numbers.\n> \n> Yes, the cache validity/invalidation criteria are the tricky bit.\n> Honestly, this is where the code gets ugly, not computing and storing\n> the generation numbers.\n\nBTW. with storing generation number in commit header there is a problem\nwhat would old version of git, one which does not understand said header,\ndo during rebase.  Would it strip unknown headers, or would it copy\ngeneration number verbatim - which means that it can be incorrect?\n\nBTW2. code size comparing in-commit and external cache cases must take\ninto account yet to be written fsck for in-commit generation numbers.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"171838","messageId":"alpine.LFD.2.00.1107220907370.1762@xanadu.home","threadId":"27834","inReplyTo":"201107221418.52414.jnareb@gmail.com","subject":"Re: Git commit generation numbers","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-07-22T13:09:54Z","receivedAt":"2011-07-22T13:09:54Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 22 Jul 2011, Jakub Narebski wrote:\n\n> BTW. with storing generation number in commit header there is a problem\n> what would old version of git, one which does not understand said header,\n> do during rebase.  Would it strip unknown headers, or would it copy\n> generation number verbatim - which means that it can be incorrect?\n\nThey would indeed be copied verbatim and become incorrect.\n\n\nNicolas\n"},{"id":"171849","messageId":"alpine.DEB.2.02.1107221100560.6496@asgard.lang.hm","threadId":"27834","inReplyTo":"201107221418.52414.jnareb@gmail.com","subject":"Re: Git commit generation numbers","fromName":"","fromEmail":"david@lang.hm","sentAt":"2011-07-22T18:02:07Z","receivedAt":"2011-07-22T18:02:07Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 22 Jul 2011, Jakub Narebski wrote:\n\n>> Yes, the cache validity/invalidation criteria are the tricky bit.\n>> Honestly, this is where the code gets ugly, not computing and storing\n>> the generation numbers.\n>\n> BTW. with storing generation number in commit header there is a problem\n> what would old version of git, one which does not understand said header,\n> do during rebase.  Would it strip unknown headers, or would it copy\n> generation number verbatim - which means that it can be incorrect?\n\nLinus has already pointed out that this is safe.\n\nold versions won't create generation numbers, but they will ignore them if \nthey exist. Since commits are not modified after they are created, the old \nversions don't copy or modify them.\n\nDavid Lang\n"},{"id":"171850","messageId":"alpine.DEB.2.02.1107221102180.6496@asgard.lang.hm","threadId":"27834","inReplyTo":"alpine.LFD.2.00.1107220907370.1762@xanadu.home","subject":"Re: Git commit generation numbers","fromName":"","fromEmail":"david@lang.hm","sentAt":"2011-07-22T18:02:35Z","receivedAt":"2011-07-22T18:02:35Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 22 Jul 2011, Nicolas Pitre wrote:\n\n> On Fri, 22 Jul 2011, Jakub Narebski wrote:\n>\n>> BTW. with storing generation number in commit header there is a problem\n>> what would old version of git, one which does not understand said header,\n>> do during rebase.  Would it strip unknown headers, or would it copy\n>> generation number verbatim - which means that it can be incorrect?\n>\n> They would indeed be copied verbatim and become incorrect.\n\nhow would they become incorrect?\n\nDavid Lang\n"},{"id":"171851","messageId":"201107222034.20510.jnareb@gmail.com","threadId":"27834","inReplyTo":"alpine.DEB.2.02.1107221102180.6496@asgard.lang.hm","subject":"Re: Git commit generation numbers","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2011-07-22T18:34:19Z","receivedAt":"2011-07-22T18:34:19Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Fri, 22 Jul 2011, David Lang <david@lang.hm> wrote:\n> On Fri, 22 Jul 2011, Nicolas Pitre wrote:\n> > On Fri, 22 Jul 2011, Jakub Narebski wrote:\n> >\n> > > BTW. with storing generation number in commit header there is a problem\n> > > what would old version of git, one which does not understand said header,\n> > > do during rebase.  Would it strip unknown headers, or would it copy\n> > > generation number verbatim - which means that it can be incorrect?\n> >\n> > They would indeed be copied verbatim and become incorrect.\n> \n> how would they become incorrect?\n\nLet's assume that the following history was created with new git, one\nthat correcly adds generation number header to commits:\n\n\n  A(1)---B(2)---C(3)---D(4)---E(5)       <-- master\n          \\\n           \\----x(3)---y(4)---z(5)       <-- foo\n\nThe numbers are generation numbers in commit object.\n\nLet's assume that this repository is fetched into repository instance\nthat is managed by older git, one that doesn't understand generation\nheader.\n\nThen, if we do\n\n  [old]$ git rebase master foo\n\nand if old git _copies_ generation number header _verbatim_, we would\nget:\n\n  A(1)---B(2)---C(3)---D(4)---E(5)                         <-- master\n                               \\\n                                \\---x'(3)--y'(4)--z'(5)    <-- foo\n\nThose generation numbers are *incorrect*; they should be:\n\n  A(1)---B(2)---C(3)---D(4)---E(5)                         <-- master\n                               \\\n                                \\---x'(6)--y'(7)--z'(8)    <-- foo\n\n\nThat is IF unknown headers are copied verbatim during rebase.  For\n\"encoding\" header this is a good thing, for \"generation\" it isn't.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"171853","messageId":"CA+55aFzsZ6w_a_wPEuBjtDeSDYQviVfy9UmJMxPz4geu4CRthQ@mail.gmail.com","threadId":"27834","inReplyTo":"201107222034.20510.jnareb@gmail.com","subject":"Re: Git commit generation numbers","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2011-07-22T19:06:08Z","receivedAt":"2011-07-22T19:06:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Jul 22, 2011 at 11:34 AM, Jakub Narebski <jnareb@gmail.com> wrote:\n>\n> That is IF unknown headers are copied verbatim during rebase.  For\n> \"encoding\" header this is a good thing, for \"generation\" it isn't.\n\nAfaik, they aren't copied verbatim, and never have been. Afaik, the\nonly thing that has *ever* written commits is \"commit_tree()\"\n(originally \"main()\" in commit-tree.c). Why is this red herring even\nbeing discussed?\n\nOf course you can always generate bogus commits by writing them by\nhand. But that's irrelevant.\n\n                     Linus\n"},{"id":"171854","messageId":"alpine.DEB.2.02.1107221206370.11697@asgard.lang.hm","threadId":"27834","inReplyTo":"201107222034.20510.jnareb@gmail.com","subject":"Re: Git commit generation numbers","fromName":"","fromEmail":"david@lang.hm","sentAt":"2011-07-22T19:08:14Z","receivedAt":"2011-07-22T19:08:14Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 22 Jul 2011, Jakub Narebski wrote:\n\n> On Fri, 22 Jul 2011, David Lang <david@lang.hm> wrote:\n>> On Fri, 22 Jul 2011, Nicolas Pitre wrote:\n>>> On Fri, 22 Jul 2011, Jakub Narebski wrote:\n>>>\n>>>> BTW. with storing generation number in commit header there is a problem\n>>>> what would old version of git, one which does not understand said header,\n>>>> do during rebase.  Would it strip unknown headers, or would it copy\n>>>> generation number verbatim - which means that it can be incorrect?\n>>>\n>>> They would indeed be copied verbatim and become incorrect.\n>>\n>> how would they become incorrect?\n>\n> Let's assume that the following history was created with new git, one\n> that correcly adds generation number header to commits:\n>\n>\n>  A(1)---B(2)---C(3)---D(4)---E(5)       <-- master\n>          \\\n>           \\----x(3)---y(4)---z(5)       <-- foo\n>\n> The numbers are generation numbers in commit object.\n>\n> Let's assume that this repository is fetched into repository instance\n> that is managed by older git, one that doesn't understand generation\n> header.\n>\n> Then, if we do\n>\n>  [old]$ git rebase master foo\n>\n> and if old git _copies_ generation number header _verbatim_, we would\n> get:\n>\n>  A(1)---B(2)---C(3)---D(4)---E(5)                         <-- master\n>                               \\\n>                                \\---x'(3)--y'(4)--z'(5)    <-- foo\n>\n> Those generation numbers are *incorrect*; they should be:\n>\n>  A(1)---B(2)---C(3)---D(4)---E(5)                         <-- master\n>                               \\\n>                                \\---x'(6)--y'(7)--z'(8)    <-- foo\n>\n>\n> That is IF unknown headers are copied verbatim during rebase.  For\n> \"encoding\" header this is a good thing, for \"generation\" it isn't.\n\ncommit headers are _not_ copied during rebase\n\na rebase is not the exact same commit, it's a \"logically equivalent\" \ncommit.\n\nso when you do a rebase, you change the commit headers (you have to change \nthe parent headers in any case, and you would have to change the \ngeneration numbers as well)\n\nthis was discussed earlier in this thread.\n\nDavid Lang\n"},{"id":"171858","messageId":"alpine.LFD.2.00.1107221530270.1762@xanadu.home","threadId":"27834","inReplyTo":"alpine.DEB.2.02.1107221206370.11697@asgard.lang.hm","subject":"Re: Git commit generation numbers","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2011-07-22T19:40:14Z","receivedAt":"2011-07-22T19:40:14Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 22 Jul 2011, david@lang.hm wrote:\n\n> On Fri, 22 Jul 2011, Jakub Narebski wrote:\n> \n> > That is IF unknown headers are copied verbatim during rebase.  For\n> > \"encoding\" header this is a good thing, for \"generation\" it isn't.\n> \n> commit headers are _not_ copied during rebase\n\nYes, this turns out to be true as I forgot that rebase is constructed on \ntop of format-patch+am, and format-patch doesn't preserve the ancillary \nheaders such as the existing \"encoding\" header, or the hypothetical \n\"generation\" header.\n\n\nNicolas\n"},{"id":"171868","messageId":"20110722220216.GA14118@sigill.intra.peff.net","threadId":"27834","inReplyTo":"CA+55aFzsZ6w_a_wPEuBjtDeSDYQviVfy9UmJMxPz4geu4CRthQ@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-07-22T22:02:17Z","receivedAt":"2011-07-22T22:02:17Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jul 22, 2011 at 12:06:08PM -0700, Linus Torvalds wrote:\n\n> On Fri, Jul 22, 2011 at 11:34 AM, Jakub Narebski <jnareb@gmail.com> wrote:\n> >\n> > That is IF unknown headers are copied verbatim during rebase.  For\n> > \"encoding\" header this is a good thing, for \"generation\" it isn't.\n> \n> Afaik, they aren't copied verbatim, and never have been. Afaik, the\n> only thing that has *ever* written commits is \"commit_tree()\"\n> (originally \"main()\" in commit-tree.c). Why is this red herring even\n> being discussed?\n\nIn git.git, that is the case. There are other programs that may write\ngit commits, though. Try:\n\n  http://www.google.com/codesearch#search/&q=hash-object.*commit&type=cs\n\nMany uses seem OK (they are generating a commit from scratch). This one\nat least (the sixth result from the search above) would actually\ngenerate buggy generation headers (it modifies parents but passes other\nheaders through):\n\n  http://www.google.com/codesearch#XUVcT9DKB_U/replace&ct=rc&cd=7&q=hash-object.*commit\n\nIt may be worth saying that such code is stupid and ugly and wrong, or\nthat it is not deployed widely enough to care about.  But it's not\nentirely a red herring.\n\n-Peff\n"},{"id":"172262","messageId":"CAMP44s2F429MG5DeRAULnSgNkCwrVGPfC2HeFw=iHXPXjkw0yA@mail.gmail.com","threadId":"27834","inReplyTo":"CA+55aFzsZ6w_a_wPEuBjtDeSDYQviVfy9UmJMxPz4geu4CRthQ@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2011-07-28T15:00:48Z","receivedAt":"2011-07-28T15:00:48Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Fri, Jul 22, 2011 at 10:06 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n> On Fri, Jul 22, 2011 at 11:34 AM, Jakub Narebski <jnareb@gmail.com> wrote:\n>>\n>> That is IF unknown headers are copied verbatim during rebase.  For\n>> \"encoding\" header this is a good thing, for \"generation\" it isn't.\n>\n> Afaik, they aren't copied verbatim, and never have been. Afaik, the\n> only thing that has *ever* written commits is \"commit_tree()\"\n> (originally \"main()\" in commit-tree.c). Why is this red herring even\n> being discussed?\n>\n> Of course you can always generate bogus commits by writing them by\n> hand. But that's irrelevant.\n\nLet's suppose for a moment that the commits do have these wrong\ngeneration numbers, shouldn't a fetch on the newer client check these\nand show an error? But what if they are pushed to a central server\nthat has old version of git? It would be messy.\n\n-- \nFelipe Contreras\n"},{"id":"174908","messageId":"CALkWK0kJ_0_MyUgQ+F+FgGun6vtk=VTh6Gsbb0u+EUrcLT5cGg@mail.gmail.com","threadId":"27834","inReplyTo":"CAMP44s2F429MG5DeRAULnSgNkCwrVGPfC2HeFw=iHXPXjkw0yA@mail.gmail.com","subject":"Re: Git commit generation numbers","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2011-09-06T10:02:03Z","receivedAt":"2011-09-06T10:02:03Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Hi,\n\nFirst, let me start out by saying that I'm a fairly new contributor to\nGit, and I'm far less experienced than the other people on this\nthread.  I've read through all the discussions time and again, and\nthought about the problem for some time now - I can't say I understand\nit as fully as many of you do, but I think I may have a slightly\ndifferent perspective to offer.\n\nIn what way is Git fundamentally different from Subversion?  It's the\nsimplicity of the data model.  From the simplest building block, a\nkey-value store, we have been able to compose and build things on top\nof it.  The reason we built centralized version control systems\nearlier is because it was *easier* to address the composition\nproblems.  We dumped all related repository and problems into one\ncentral server.  With so much information in one place, things are\ntightly coupled and problems are easier to solve.  Still not\nconvinced?  What's the weakest component in Git today?  Undoubtedly\nsubmodules.  Ofcourse, a large part of the reason is that many people\ndon't use submodules, and hence it doesn't improve -- but it's\nactually a circular problem.  People don't use submodules, because\nit's so featureless and hard to develop.  Why is it so hard?  Back to\nthe fundamental problem of composition from simple building blocks.\nIn submodules, we have to take entire DAGs and build a composite DAG.\nThe key pieces of information are deep inside Git's fundamnetals:\nGitlinks.  Other projects try like Gitslave try to attack the problem\non a more superficial level, but they all hit a barrier when they\ndiscover that they can't compose big blocks of data: you need simple\nbuilding blocks to compose.\n\nIt's the same story with C (and now, Haskell).  Why does everyone like\nC so much?  Because it only provides fundamental building blocks and\ngives people the freedom to compose the way they like.  It doesn't\nprovide big \"template blocks\" like Java, because they tend to be\nrestrictive in the long run.  Sure, Java is easier to start out with,\nbut people soon realize that big blocks can't compose.\n\nMore than arguing about backward compatibility, and about how older\nversions of Git commits won't have generation numbers, I think this is\nwhat we should be focusing on.  Sure, it'll additionally make sense to\nput in a cache to speed things up now, but we need to think about what\nGit will be 10~15 years from now.  The fundamental pieces of\ninformation required for composition must be present in the\nfundamental building blocks.\n\nThe real question we should be asking is: \"Should Git have had commit\ngeneration numbers in 2005?\".  If the answer is \"yes\", we should put\nthem in now before it becomes even harder, bending over backwards for\nbackward compatibility if necessary.  Otherwise, we'll regret this\ndecision 10~15 years later, when we're faced with deeper issues.  If\nyou want a concrete example, think about how you'd compose DAGs\ntogether (again, the submodules problem): where is the information\nrequired to prune each DAG and compose?\n\nI wish I could write this in myself, but I'm afraid I don't have the\nengineering skill yet.  I'll be happy to contribute whatever little I\ncan, and participate in the review process.\n\nThanks.\n\n-- Ram\n"}]}