# Re: Git commit generation numbers

7 messages from 2011-07-21 to 2011-07-22. Participants: Drew Northup, George Spelvin, Phil Hord, Pēteris Kļaviņš, Christian Couder.
Thread: https://gitlist.dev/t/27875

## Drew Northup, 2011-07-21 12:03

Subject: Re: Git commit generation numbers
Message-ID: <1311249789.9745.30.camel@drew-northup.unet.maine.edu>
URL: https://gitlist.dev/e/1311249789.9745.30.camel%40drew-northup.unet.maine.edu
In-Reply-To: <alpine.DEB.2.02.1107201624510.5222@asgard.lang.hm>

```

On Wed, 2011-07-20 at 16:26 -0700, david@lang.hm wrote:
> On Wed, 20 Jul 2011, George Spelvin wrote:
> 
> >> The alternative of having to sometimes use the generation number,
> >> sometimes use the possibly broken commit date, makes for much more
> >> complicated code that has to be maintained forever.  Having a solution
> >> that starts working only after a certain point in history doesn't look
> >> eleguant to me at all.  It is not like having different pack formats
> >> where back and forth conversions can be made for the _entire_ history.
> >
> > It seemed like a pretty strong argument to me, too.
> 
> except that you then have different caches on different systems. If the 
> generation number is part of the repository then it's going to be the same 
> for everyone.

I keep hearing (reading) people stating this utterly unfounded argument.
The fact is that for any work not yet integrated back into a shared
repository it just isn't true--and even after upstream integration the
truth of such a statement may be limited.

I have not read yet one discussion about how generation numbers [baked
into a commit] deal with rebasing, for instance. Do we assign one more
than the revision prior to the base of the rebase operation or do we
start with the revision one after the highest of those original commits
included in the rebase? Depending on how that is done
_drastically_different_ numbers can come out of different repository
instances for the same _final_ DAG. This is one major reason why, as I
see it, local storage is good for generation numbers and putting them in
the commit is bad. 

I have no problem with putting an _advisory_ "revision number" in the
commit. It would not be expected to have a proper "1-to-1 and onto"
functional association with the _final_ DAG, but it could potentially
get us some nice benefits. We would still need to answer questions like
the one I ask above, but it would hurt less to change if we need to.

One other sane option that was mentioned at least once in passing was to
store the generation number in some Git "filesystem-level" object. This
could then be reconciled with each "git gc" or "git fsck" operation if
not more often. This is less ad-hoc and messy than a separate cache,
becomes amenable to the standard tool-set, and always gets updated (no
invalid cache). If an _advisory_ revision number is available in commits
that are sent along those could conceivably be used to help build up the
local git-fs generation numbers more quickly. (If a "git pull" is issued
to our repo, or we push to another, we don't send the generation numbers
locally stored--we expect the git-fs machinery to regenerate those on
the fly.)

I may not be one of the "resident rocket scientists," but that's how I
see it.

-- 
-Drew Northup
________________________________________________
"As opposed to vegetable or mineral error?"
-John Pescatore, SANS NewsBites Vol. 12 Num. 59

```

## George Spelvin, 2011-07-21 12:55

Subject: Re: Git commit generation numbers
Message-ID: <20110721125544.26006.qmail@science.horizon.com>
URL: https://gitlist.dev/e/20110721125544.26006.qmail%40science.horizon.com
In-Reply-To: <1311249789.9745.30.camel@drew-northup.unet.maine.edu>

```
> I have not read yet one discussion about how generation numbers [baked
> into a commit] deal with rebasing, for instance. Do we assign one more
> than the revision prior to the base of the rebase operation or do we
> start with the revision one after the highest of those original commits
> included in the rebase? Depending on how that is done
> _drastically_different_ numbers can come out of different repository
> instances for the same _final_ DAG. This is one major reason why, as I
> see it, local storage is good for generation numbers and putting them in
> the commit is bad. 

Er, no.  Whenever a new commit object is generated (as the result
of a rebase or not), its commit number is computed based on its
parent commits.  It is NEVER copied.

Just like the parent pointers themselves.  Remember, even though we talk
about "the same commit" after rebasing, it's really just an EQUIVALENT
commit according to some higher-level concept of similarity.  As far
as the core git engine is concerned, it's always a DIFFERENT commit,
with different parent hashes and a different hash itself.

This point hasn't been mentioned explicltly precisely because it's
so obvious; the history-walking code that the generation numbers are
for requires this property to function.

```

## Drew Northup, 2011-07-21 15:57

Subject: Re: Git commit generation numbers
Message-ID: <1311263869.9745.72.camel@drew-northup.unet.maine.edu>
URL: https://gitlist.dev/e/1311263869.9745.72.camel%40drew-northup.unet.maine.edu
In-Reply-To: <20110721125544.26006.qmail@science.horizon.com>

```

On Thu, 2011-07-21 at 08:55 -0400, George Spelvin wrote:
> > I have not read yet one discussion about how generation numbers [baked
> > into a commit] deal with rebasing, for instance. Do we assign one more
> > than the revision prior to the base of the rebase operation or do we
> > start with the revision one after the highest of those original commits
> > included in the rebase? Depending on how that is done
> > _drastically_different_ numbers can come out of different repository
> > instances for the same _final_ DAG. This is one major reason why, as I
> > see it, local storage is good for generation numbers and putting them in
> > the commit is bad. 
> 
> Er, no.  Whenever a new commit object is generated (as the result
> of a rebase or not), its commit number is computed based on its
> parent commits.  It is NEVER copied.

I don't see the word "copy" in my original. 

B-O1-O2-O3-O4-O5-O6
 \
  R1----R2-------R3

What's the correct generation number for R3? I would say gen(B)+3. My
reading of the posts made by some others was that they thought gen(O6)
was the correct answer. Still others seemed to indicate gen(O6)+1 was
the correct answer. I don't think everybody MEANT to be saying such
different things--that's just how they appeared on this end.

Now, did you mean something different by "commit number?"

-- 
-Drew Northup
________________________________________________
"As opposed to vegetable or mineral error?"
-John Pescatore, SANS NewsBites Vol. 12 Num. 59

```

## Phil Hord, 2011-07-21 16:24

Subject: Re: Git commit generation numbers
Message-ID: <4E2852A1.30800@cisco.com>
URL: https://gitlist.dev/e/4E2852A1.30800%40cisco.com
In-Reply-To: <1311263869.9745.72.camel@drew-northup.unet.maine.edu>

```
On 07/21/2011 11:57 AM, Drew Northup wrote:
> On Thu, 2011-07-21 at 08:55 -0400, George Spelvin wrote:
>>> I have not read yet one discussion about how generation numbers [baked
>>> into a commit] deal with rebasing, for instance. Do we assign one more
>>> than the revision prior to the base of the rebase operation or do we
>>> start with the revision one after the highest of those original commits
>>> included in the rebase? Depending on how that is done
>>> _drastically_different_ numbers can come out of different repository
>>> instances for the same _final_ DAG. This is one major reason why, as I
>>> see it, local storage is good for generation numbers and putting them in
>>> the commit is bad.
>> Er, no.  Whenever a new commit object is generated (as the result
>> of a rebase or not), its commit number is computed based on its
>> parent commits.  It is NEVER copied.
> I don't see the word "copy" in my original.
>
> B-O1-O2-O3-O4-O5-O6
>   \
>    R1----R2-------R3
>
> What's the correct generation number for R3? I would say gen(B)+3.
And you would be correct if you follow the SoP algorithm.

> My
> reading of the posts made by some others was that they thought gen(O6)
> was the correct answer. Still others seemed to indicate gen(O6)+1 was
> the correct answer.
Maybe the confusion comes from the different storage mechanisms being 
discussed.  If the generation numbers are in a local cache and used by a 
single client, the determinism of the specific numbers doesn't much 
matter.  If they are part of the commit, it still doesn't need to be 
completely deterministic. However, interoperability requires standards, 
and standards favor determinism, so dogmatic determinism may triumph in 
that case.

1. gen(06) might make sense if you mean to implement --date-order using 
gen-numbers, for example.  But I don't think it's practical in any case.

2. gen(06)+1 might make sense if you mean to require that gen-numbers 
are unique per repo.  But this is both unsupportable and unnecessary, so 
it's a non-starter.

3. gen(B)+1 is what you'd get from the the algorithm I saw proposed.

All three of these are provably correct by my definition of "correct": 
"for each A in ancestors_of(B), gen(A) < gen(B)".

However, [1] and [2] have some extra features of dubious value.  Simpler 
is better for interoperability, so I like [3] for this purpose.

Even [3] has an extra feature I think is unnecessary: determinism.  If 
that "requirement" is dropped, I think all three of these algorithms are 
(functionally) roughly equivalent.

> I don't think everybody MEANT to be saying such
> different things--that's just how they appeared on this end.
>
> Now, did you mean something different by "commit number?"

I remain unconvinced that there is value in gen-number distribution, so 
to my mind, the specific algorithm and whether or not it is 
deterministic are unimportant.

Phil ~ who wasn't really being asked, but felt like answering

```

## George Spelvin, 2011-07-21 17:36

Subject: Re: Git commit generation numbers
Message-ID: <20110721173633.21195.qmail@science.horizon.com>
URL: https://gitlist.dev/e/20110721173633.21195.qmail%40science.horizon.com
In-Reply-To: <1311263869.9745.72.camel@drew-northup.unet.maine.edu>

```
Drew Northup wrote:
> On Thu, 2011-07-21 at 08:55 -0400, George Spelvin wrote:
>> I have not read yet one discussion about how generation numbers [baked
>> into a commit] deal with rebasing, for instance. Do we assign one more
>> than the revision prior to the base of the rebase operation or do we
>> start with the revision one after the highest of those original commits
>> included in the rebase? Depending on how that is done
>> _drastically_different_ numbers can come out of different repository
>> instances for the same _final_ DAG. This is one major reason why, as I
>> see it, local storage is good for generation numbers and putting them in
>> the commit is bad. 
> 
> Er, no.  Whenever a new commit object is generated (as the result
> of a rebase or not), its commit number is computed based on its
> parent commits.  It is NEVER copied.

> I don't see the word "copy" in my original. 

Indeed, you didn't use it; it was my simplified mental model of your
suggestion that the rebased commits would have generation numbers that
somehow depended on the generation numbers before rebasing.

Althouugh you suggested something different, the mistake is the same:
the rebased commits' generation numbers have simply no relationship to
those of the original pre-rebase commits.  The generation numbers depend
only on the commits explicitly listed as parents in the commit objects.

That's why I went on to explain that the equivalence of the commits
produced by a rebase operation is a higher-level concept; the core git
object database just knows that they aren't identical, and therefore
are different.

Thus, they would retain the same relative order as before the rebase
(unless you permuted them with rebase -i), but start with the generation
number of the rebase target.

> B-O1-O2-O3-O4-O5-O6
>  \
>   R1----R2-------R3

> What's the correct generation number for R3? I would say gen(B)+3. My
> reading of the posts made by some others was that they thought gen(O6)
> was the correct answer. Still others seemed to indicate gen(O6)+1 was
> the correct answer. I don't think everybody MEANT to be saying such
> different things--that's just how they appeared on this end.

According to the canonical algorithm, it's gen(B)+3 = gen(R2)+1.

However, any non-decreasing series is equally permissible for
optimizing history walking, so you could add jumps to (for example)
make the numbers unique if that simplified anything.

I don't think it does simplify anything, so the issue hasn't been
discussed much.

For the purpose of the optimization enabled by the generation
numbers, however, it doesn't actually matter.

What matters is that if I am listing commits down multiple branches,
once I have walked back on each branch to commits of generation N or
less, I know that I have found all possible descendants of all commits
of generation N or more.

This lets me display the recent part of the commit DAG (back to generation
N) without exploring the entire commit treem or worrying that I'll have to
"back up" to insert a commit in its proper order.  Without precomputed
generation numbers, the only way to be sure of this is to explore back
to generation 0 (parentless commits) or to use date-based heuristics.

> Now, did you mean something different by "commit number?"

No, just a bran fart I didn't catch before posting.
I meant "generation number".

```

## Pēteris Kļaviņš, 2011-07-21 22:40

Subject: Re: Git commit generation numbers
Message-ID: <j0a9te$vcv$1@dough.gmane.org>
URL: https://gitlist.dev/e/j0a9te%24vcv%241%40dough.gmane.org
In-Reply-To: <4E2852A1.30800@cisco.com>

```
On 21/07/2011 5:24 PM, Phil Hord wrote:
> Maybe the confusion comes from the different storage mechanisms being
> discussed. If the generation numbers are in a local cache and used by a
> single client, the determinism of the specific numbers doesn't much
> matter. If they are part of the commit, it still doesn't need to be
> completely deterministic. However, interoperability requires standards,
> and standards favor determinism, so dogmatic determinism may triumph in
> that case.
>
> 1. gen(06) might make sense if you mean to implement --date-order using
> gen-numbers, for example. But I don't think it's practical in any case.
>
> 2. gen(06)+1 might make sense if you mean to require that gen-numbers
> are unique per repo. But this is both unsupportable and unnecessary, so
> it's a non-starter.
>
> 3. gen(B)+1 is what you'd get from the the algorithm I saw proposed.
>
> All three of these are provably correct by my definition of "correct":
> "for each A in ancestors_of(B), gen(A) < gen(B)".
>
> However, [1] and [2] have some extra features of dubious value. Simpler
> is better for interoperability, so I like [3] for this purpose.
>
> Even [3] has an extra feature I think is unnecessary: determinism. If
> that "requirement" is dropped, I think all three of these algorithms are
> (functionally) roughly equivalent.
>
>> I don't think everybody MEANT to be saying such
>> different things--that's just how they appeared on this end.
>>
>> Now, did you mean something different by "commit number?"
>
> I remain unconvinced that there is value in gen-number distribution, so
> to my mind, the specific algorithm and whether or not it is
> deterministic are unimportant.
>

The beauty of Git is that no two copies of a Git repository as a whole 
are the same:  some people make shallow copies;  others prune away all 
branches except for the one they are interested in;  yet others graft 
together multiple original repositories.  The upshot is that two copies 
of the same repository may end up having different commits as their root 
commits, and so the generation numbers computed for their repositories 
would be different.  Indeed, the shallow repository copy could later be 
filled out with additional underlying commits, and so on.

Given this context, I can't see the value in fixing generation numbers 
within commits.  In my mind generation numbers are extremely useful 
transient helper objects in every Git repository but they have no 
meaning outside that repository, sort of like GIT_WORK_TREE.

Peter

```

## Christian Couder, 2011-07-22 09:30

Subject: Re: Git commit generation numbers
Message-ID: <CAP8UFD0FG47D0hV8ifj1p7bRfEXJ7Ez3mTw4Yv-KUWXPJnWLOg@mail.gmail.com>
URL: https://gitlist.dev/e/CAP8UFD0FG47D0hV8ifj1p7bRfEXJ7Ez3mTw4Yv-KUWXPJnWLOg%40mail.gmail.com
In-Reply-To: <j0a9te$vcv$1@dough.gmane.org>

```
On Fri, Jul 22, 2011 at 12:40 AM, Pēteris Kļaviņš
<klavins@netspace.net.au> wrote:
>
> The beauty of Git is that no two copies of a Git repository as a whole are
> the same:  some people make shallow copies;  others prune away all branches
> except for the one they are interested in;  yet others graft together
> multiple original repositories.  The upshot is that two copies of the same
> repository may end up having different commits as their root commits, and so
> the generation numbers computed for their repositories would be different.
>  Indeed, the shallow repository copy could later be filled out with
> additional underlying commits, and so on.

Not only people want different repos, but with their own repo they
want different "views" (or "virtual graph") of it.

> Given this context, I can't see the value in fixing generation numbers
> within commits.  In my mind generation numbers are extremely useful
> transient helper objects in every Git repository but they have no meaning
> outside that repository, sort of like GIT_WORK_TREE.

It's not even per repository that they have a meaning, it's per "view"
of the commit graph.

Thanks,
Christian.

```
