{"thread":{"id":"27875","subject":"Re: Git commit generation numbers","startedAt":"2011-07-21T12:03:09Z","lastAt":"2011-07-22T09:30:27Z","messageCount":7,"participants":["Drew Northup","George Spelvin","Phil Hord","Pēteris Kļaviņš","Christian Couder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"171799","messageId":"1311249789.9745.30.camel@drew-northup.unet.maine.edu","threadId":"27875","inReplyTo":"alpine.DEB.2.02.1107201624510.5222@asgard.lang.hm","subject":"Re: Git commit generation numbers","fromName":"Drew Northup","fromEmail":"drew.northup@maine.edu","sentAt":"2011-07-21T12:03:09Z","receivedAt":"2011-07-21T12:03:09Z","isPatch":false,"sender":{"key":"drew.northup@maine.edu","avatar":"https://avatars.githubusercontent.com/u/18331571?v=4"},"body":"\nOn Wed, 2011-07-20 at 16:26 -0700, david@lang.hm wrote:\n> On Wed, 20 Jul 2011, George Spelvin wrote:\n> \n> >> The alternative of having to sometimes use the generation number,\n> >> sometimes use the possibly broken commit date, makes for much more\n> >> complicated code that has to be maintained forever.  Having a solution\n> >> that starts working only after a certain point in history doesn't look\n> >> eleguant to me at all.  It is not like having different pack formats\n> >> where back and forth conversions can be made for the _entire_ history.\n> >\n> > It seemed like a pretty strong argument to me, too.\n> \n> except that you then have different caches on different systems. If the \n> generation number is part of the repository then it's going to be the same \n> for everyone.\n\nI keep hearing (reading) people stating this utterly unfounded argument.\nThe fact is that for any work not yet integrated back into a shared\nrepository it just isn't true--and even after upstream integration the\ntruth of such a statement may be limited.\n\nI have not read yet one discussion about how generation numbers [baked\ninto a commit] deal with rebasing, for instance. Do we assign one more\nthan the revision prior to the base of the rebase operation or do we\nstart with the revision one after the highest of those original commits\nincluded in the rebase? Depending on how that is done\n_drastically_different_ numbers can come out of different repository\ninstances for the same _final_ DAG. This is one major reason why, as I\nsee it, local storage is good for generation numbers and putting them in\nthe commit is bad. \n\nI have no problem with putting an _advisory_ \"revision number\" in the\ncommit. It would not be expected to have a proper \"1-to-1 and onto\"\nfunctional association with the _final_ DAG, but it could potentially\nget us some nice benefits. We would still need to answer questions like\nthe one I ask above, but it would hurt less to change if we need to.\n\nOne other sane option that was mentioned at least once in passing was to\nstore the generation number in some Git \"filesystem-level\" object. This\ncould then be reconciled with each \"git gc\" or \"git fsck\" operation if\nnot more often. This is less ad-hoc and messy than a separate cache,\nbecomes amenable to the standard tool-set, and always gets updated (no\ninvalid cache). If an _advisory_ revision number is available in commits\nthat are sent along those could conceivably be used to help build up the\nlocal git-fs generation numbers more quickly. (If a \"git pull\" is issued\nto our repo, or we push to another, we don't send the generation numbers\nlocally stored--we expect the git-fs machinery to regenerate those on\nthe fly.)\n\nI may not be one of the \"resident rocket scientists,\" but that's how I\nsee it.\n\n-- \n-Drew Northup\n________________________________________________\n\"As opposed to vegetable or mineral error?\"\n-John Pescatore, SANS NewsBites Vol. 12 Num. 59\n"},{"id":"171798","messageId":"20110721125544.26006.qmail@science.horizon.com","threadId":"27875","inReplyTo":"1311249789.9745.30.camel@drew-northup.unet.maine.edu","subject":"Re: Git commit generation numbers","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2011-07-21T12:55:44Z","receivedAt":"2011-07-21T12:55:44Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> I have not read yet one discussion about how generation numbers [baked\n> into a commit] deal with rebasing, for instance. Do we assign one more\n> than the revision prior to the base of the rebase operation or do we\n> start with the revision one after the highest of those original commits\n> included in the rebase? Depending on how that is done\n> _drastically_different_ numbers can come out of different repository\n> instances for the same _final_ DAG. This is one major reason why, as I\n> see it, local storage is good for generation numbers and putting them in\n> the commit is bad. \n\nEr, no.  Whenever a new commit object is generated (as the result\nof a rebase or not), its commit number is computed based on its\nparent commits.  It is NEVER copied.\n\nJust like the parent pointers themselves.  Remember, even though we talk\nabout \"the same commit\" after rebasing, it's really just an EQUIVALENT\ncommit according to some higher-level concept of similarity.  As far\nas the core git engine is concerned, it's always a DIFFERENT commit,\nwith different parent hashes and a different hash itself.\n\nThis point hasn't been mentioned explicltly precisely because it's\nso obvious; the history-walking code that the generation numbers are\nfor requires this property to function.\n"},{"id":"171781","messageId":"1311263869.9745.72.camel@drew-northup.unet.maine.edu","threadId":"27875","inReplyTo":"20110721125544.26006.qmail@science.horizon.com","subject":"Re: Git commit generation numbers","fromName":"Drew Northup","fromEmail":"drew.northup@maine.edu","sentAt":"2011-07-21T15:57:49Z","receivedAt":"2011-07-21T15:57:49Z","isPatch":false,"sender":{"key":"drew.northup@maine.edu","avatar":"https://avatars.githubusercontent.com/u/18331571?v=4"},"body":"\nOn Thu, 2011-07-21 at 08:55 -0400, George Spelvin wrote:\n> > I have not read yet one discussion about how generation numbers [baked\n> > into a commit] deal with rebasing, for instance. Do we assign one more\n> > than the revision prior to the base of the rebase operation or do we\n> > start with the revision one after the highest of those original commits\n> > included in the rebase? Depending on how that is done\n> > _drastically_different_ numbers can come out of different repository\n> > instances for the same _final_ DAG. This is one major reason why, as I\n> > see it, local storage is good for generation numbers and putting them in\n> > the commit is bad. \n> \n> Er, no.  Whenever a new commit object is generated (as the result\n> of a rebase or not), its commit number is computed based on its\n> parent commits.  It is NEVER copied.\n\nI don't see the word \"copy\" in my original. \n\nB-O1-O2-O3-O4-O5-O6\n \\\n  R1----R2-------R3\n\nWhat's the correct generation number for R3? I would say gen(B)+3. My\nreading of the posts made by some others was that they thought gen(O6)\nwas the correct answer. Still others seemed to indicate gen(O6)+1 was\nthe correct answer. I don't think everybody MEANT to be saying such\ndifferent things--that's just how they appeared on this end.\n\nNow, did you mean something different by \"commit number?\"\n\n-- \n-Drew Northup\n________________________________________________\n\"As opposed to vegetable or mineral error?\"\n-John Pescatore, SANS NewsBites Vol. 12 Num. 59\n"},{"id":"171791","messageId":"4E2852A1.30800@cisco.com","threadId":"27875","inReplyTo":"1311263869.9745.72.camel@drew-northup.unet.maine.edu","subject":"Re: Git commit generation numbers","fromName":"Phil Hord","fromEmail":"hordp@cisco.com","sentAt":"2011-07-21T16:24:01Z","receivedAt":"2011-07-21T16:24:01Z","isPatch":false,"sender":{"key":"phil.hord@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123908?v=4"},"body":"On 07/21/2011 11:57 AM, Drew Northup wrote:\n> On Thu, 2011-07-21 at 08:55 -0400, George Spelvin wrote:\n>>> I have not read yet one discussion about how generation numbers [baked\n>>> into a commit] deal with rebasing, for instance. Do we assign one more\n>>> than the revision prior to the base of the rebase operation or do we\n>>> start with the revision one after the highest of those original commits\n>>> included in the rebase? Depending on how that is done\n>>> _drastically_different_ numbers can come out of different repository\n>>> instances for the same _final_ DAG. This is one major reason why, as I\n>>> see it, local storage is good for generation numbers and putting them in\n>>> the commit is bad.\n>> Er, no.  Whenever a new commit object is generated (as the result\n>> of a rebase or not), its commit number is computed based on its\n>> parent commits.  It is NEVER copied.\n> I don't see the word \"copy\" in my original.\n>\n> B-O1-O2-O3-O4-O5-O6\n>   \\\n>    R1----R2-------R3\n>\n> What's the correct generation number for R3? I would say gen(B)+3.\nAnd you would be correct if you follow the SoP algorithm.\n\n> My\n> reading of the posts made by some others was that they thought gen(O6)\n> was the correct answer. Still others seemed to indicate gen(O6)+1 was\n> the correct answer.\nMaybe the confusion comes from the different storage mechanisms being \ndiscussed.  If the generation numbers are in a local cache and used by a \nsingle client, the determinism of the specific numbers doesn't much \nmatter.  If they are part of the commit, it still doesn't need to be \ncompletely deterministic. However, interoperability requires standards, \nand standards favor determinism, so dogmatic determinism may triumph in \nthat case.\n\n1. gen(06) might make sense if you mean to implement --date-order using \ngen-numbers, for example.  But I don't think it's practical in any case.\n\n2. gen(06)+1 might make sense if you mean to require that gen-numbers \nare unique per repo.  But this is both unsupportable and unnecessary, so \nit's a non-starter.\n\n3. gen(B)+1 is what you'd get from the the algorithm I saw proposed.\n\nAll three of these are provably correct by my definition of \"correct\": \n\"for each A in ancestors_of(B), gen(A) < gen(B)\".\n\nHowever, [1] and [2] have some extra features of dubious value.  Simpler \nis better for interoperability, so I like [3] for this purpose.\n\nEven [3] has an extra feature I think is unnecessary: determinism.  If \nthat \"requirement\" is dropped, I think all three of these algorithms are \n(functionally) roughly equivalent.\n\n> I don't think everybody MEANT to be saying such\n> different things--that's just how they appeared on this end.\n>\n> Now, did you mean something different by \"commit number?\"\n\nI remain unconvinced that there is value in gen-number distribution, so \nto my mind, the specific algorithm and whether or not it is \ndeterministic are unimportant.\n\nPhil ~ who wasn't really being asked, but felt like answering\n"},{"id":"171789","messageId":"20110721173633.21195.qmail@science.horizon.com","threadId":"27875","inReplyTo":"1311263869.9745.72.camel@drew-northup.unet.maine.edu","subject":"Re: Git commit generation numbers","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2011-07-21T17:36:33Z","receivedAt":"2011-07-21T17:36:33Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"Drew Northup wrote:\n> On Thu, 2011-07-21 at 08:55 -0400, George Spelvin wrote:\n>> I have not read yet one discussion about how generation numbers [baked\n>> into a commit] deal with rebasing, for instance. Do we assign one more\n>> than the revision prior to the base of the rebase operation or do we\n>> start with the revision one after the highest of those original commits\n>> included in the rebase? Depending on how that is done\n>> _drastically_different_ numbers can come out of different repository\n>> instances for the same _final_ DAG. This is one major reason why, as I\n>> see it, local storage is good for generation numbers and putting them in\n>> the commit is bad. \n> \n> Er, no.  Whenever a new commit object is generated (as the result\n> of a rebase or not), its commit number is computed based on its\n> parent commits.  It is NEVER copied.\n\n> I don't see the word \"copy\" in my original. \n\nIndeed, you didn't use it; it was my simplified mental model of your\nsuggestion that the rebased commits would have generation numbers that\nsomehow depended on the generation numbers before rebasing.\n\nAlthouugh you suggested something different, the mistake is the same:\nthe rebased commits' generation numbers have simply no relationship to\nthose of the original pre-rebase commits.  The generation numbers depend\nonly on the commits explicitly listed as parents in the commit objects.\n\nThat's why I went on to explain that the equivalence of the commits\nproduced by a rebase operation is a higher-level concept; the core git\nobject database just knows that they aren't identical, and therefore\nare different.\n\nThus, they would retain the same relative order as before the rebase\n(unless you permuted them with rebase -i), but start with the generation\nnumber of the rebase target.\n\n> B-O1-O2-O3-O4-O5-O6\n>  \\\n>   R1----R2-------R3\n\n> What's the correct generation number for R3? I would say gen(B)+3. My\n> reading of the posts made by some others was that they thought gen(O6)\n> was the correct answer. Still others seemed to indicate gen(O6)+1 was\n> the correct answer. I don't think everybody MEANT to be saying such\n> different things--that's just how they appeared on this end.\n\nAccording to the canonical algorithm, it's gen(B)+3 = gen(R2)+1.\n\nHowever, any non-decreasing series is equally permissible for\noptimizing history walking, so you could add jumps to (for example)\nmake the numbers unique if that simplified anything.\n\nI don't think it does simplify anything, so the issue hasn't been\ndiscussed much.\n\nFor the purpose of the optimization enabled by the generation\nnumbers, however, it doesn't actually matter.\n\nWhat matters is that if I am listing commits down multiple branches,\nonce I have walked back on each branch to commits of generation N or\nless, I know that I have found all possible descendants of all commits\nof generation N or more.\n\nThis lets me display the recent part of the commit DAG (back to generation\nN) without exploring the entire commit treem or worrying that I'll have to\n\"back up\" to insert a commit in its proper order.  Without precomputed\ngeneration numbers, the only way to be sure of this is to explore back\nto generation 0 (parentless commits) or to use date-based heuristics.\n\n> Now, did you mean something different by \"commit number?\"\n\nNo, just a bran fart I didn't catch before posting.\nI meant \"generation number\".\n"},{"id":"171825","messageId":"j0a9te$vcv$1@dough.gmane.org","threadId":"27875","inReplyTo":"4E2852A1.30800@cisco.com","subject":"Re: Git commit generation numbers","fromName":"Pēteris Kļaviņš","fromEmail":"klavins@netspace.net.au","sentAt":"2011-07-21T22:40:44Z","receivedAt":"2011-07-21T22:40:44Z","isPatch":false,"sender":{"key":"klavins@netspace.net.au","avatar":"https://gravatar.com/avatar/7bb2403e1c2330c5c199171858cb3b7c1e9780f0cf0dbf44f2468fe4a6a8b079?d=mp&s=160"},"body":"On 21/07/2011 5:24 PM, Phil Hord wrote:\n> Maybe the confusion comes from the different storage mechanisms being\n> discussed. If the generation numbers are in a local cache and used by a\n> single client, the determinism of the specific numbers doesn't much\n> matter. If they are part of the commit, it still doesn't need to be\n> completely deterministic. However, interoperability requires standards,\n> and standards favor determinism, so dogmatic determinism may triumph in\n> that case.\n>\n> 1. gen(06) might make sense if you mean to implement --date-order using\n> gen-numbers, for example. But I don't think it's practical in any case.\n>\n> 2. gen(06)+1 might make sense if you mean to require that gen-numbers\n> are unique per repo. But this is both unsupportable and unnecessary, so\n> it's a non-starter.\n>\n> 3. gen(B)+1 is what you'd get from the the algorithm I saw proposed.\n>\n> All three of these are provably correct by my definition of \"correct\":\n> \"for each A in ancestors_of(B), gen(A) < gen(B)\".\n>\n> However, [1] and [2] have some extra features of dubious value. Simpler\n> is better for interoperability, so I like [3] for this purpose.\n>\n> Even [3] has an extra feature I think is unnecessary: determinism. If\n> that \"requirement\" is dropped, I think all three of these algorithms are\n> (functionally) roughly equivalent.\n>\n>> I don't think everybody MEANT to be saying such\n>> different things--that's just how they appeared on this end.\n>>\n>> Now, did you mean something different by \"commit number?\"\n>\n> I remain unconvinced that there is value in gen-number distribution, so\n> to my mind, the specific algorithm and whether or not it is\n> deterministic are unimportant.\n>\n\nThe beauty of Git is that no two copies of a Git repository as a whole \nare the same:  some people make shallow copies;  others prune away all \nbranches except for the one they are interested in;  yet others graft \ntogether multiple original repositories.  The upshot is that two copies \nof the same repository may end up having different commits as their root \ncommits, and so the generation numbers computed for their repositories \nwould be different.  Indeed, the shallow repository copy could later be \nfilled out with additional underlying commits, and so on.\n\nGiven this context, I can't see the value in fixing generation numbers \nwithin commits.  In my mind generation numbers are extremely useful \ntransient helper objects in every Git repository but they have no \nmeaning outside that repository, sort of like GIT_WORK_TREE.\n\nPeter\n"},{"id":"171836","messageId":"CAP8UFD0FG47D0hV8ifj1p7bRfEXJ7Ez3mTw4Yv-KUWXPJnWLOg@mail.gmail.com","threadId":"27875","inReplyTo":"j0a9te$vcv$1@dough.gmane.org","subject":"Re: Git commit generation numbers","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2011-07-22T09:30:27Z","receivedAt":"2011-07-22T09:30:27Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Fri, Jul 22, 2011 at 12:40 AM, Pēteris Kļaviņš\n<klavins@netspace.net.au> wrote:\n>\n> The beauty of Git is that no two copies of a Git repository as a whole are\n> the same:  some people make shallow copies;  others prune away all branches\n> except for the one they are interested in;  yet others graft together\n> multiple original repositories.  The upshot is that two copies of the same\n> repository may end up having different commits as their root commits, and so\n> the generation numbers computed for their repositories would be different.\n>  Indeed, the shallow repository copy could later be filled out with\n> additional underlying commits, and so on.\n\nNot only people want different repos, but with their own repo they\nwant different \"views\" (or \"virtual graph\") of it.\n\n> Given this context, I can't see the value in fixing generation numbers\n> within commits.  In my mind generation numbers are extremely useful\n> transient helper objects in every Git repository but they have no meaning\n> outside that repository, sort of like GIT_WORK_TREE.\n\nIt's not even per repository that they have a meaning, it's per \"view\"\nof the commit graph.\n\nThanks,\nChristian.\n"}]}