# RE: git pull on Linux/ACPI release tree

19 messages from 2006-01-09 to 2006-01-13. Participants: Luben Tuikov, Linus Torvalds, Martin Langhoff, Junio C Hamano, Kyle Moffett, Johannes Schindelin, Matthias Urlichs.
Thread: https://gitlist.dev/t/3023

## Linus Torvalds, 2006-01-09 16:57

Subject: RE: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.64.0601090850350.3169@g5.osdl.org>
URL: https://gitlist.dev/e/Pine.LNX.4.64.0601090850350.3169%40g5.osdl.org
In-Reply-To: <Pine.LNX.4.64.0601090835580.3169@g5.osdl.org>

```


On Mon, 9 Jan 2006, Linus Torvalds wrote:
>
> One thing we could do is to make it easier to apply a patch to a 
> _non_current_ branch.
>   [ ... ]
> Do you think that kind of workflow would be more palatable to you? It 
> shouldn't be /that/ hard to make git-apply branch-aware... (It was part of 
> my original plan, but it is more work than just using the working 
> directory, so I never finished the thought).

Btw, this is true in a bigger sense: the things "git" does have largely 
been driven by user needs. Initially mainly mine, but things like 
"git-rebase" were from people who wanted to work as "sub-maintainers" (eg 
Junio before he became the head honcho for git itself).

But if there are workflow problems, let's try to fix them. The "apply 
patches directly to another branch" suggestion may not be sane (maybe it's 
too confusing to apply a patch and not actually see it in the working 
tree), but workflow suggestions in general are appreciated.

We've made switching branches about as efficient as it can be (but if the 
differences are huge, the cost of re-writing the working directory is 
never going to be low). But switching branches has the "confusion factor" 
(ie you forget which branch you're on, and apply a patch to your working 
branch instead of your development branch), so maybe there are other ways 
of doing the same thing that might be sensible..

So send suggestions to the git lists. Maybe they're insane and can't be 
done, but while I designed git to work with _my_ case (ie mostly merging 
tons of different trees and then having occasional big batches of 
patches), it's certainly _supposed_ to support other maintainers too..

		Linus

```

## Luben Tuikov, 2006-01-09 22:51

Subject: RE: git pull on Linux/ACPI release tree
Message-ID: <20060109225143.60520.qmail@web31807.mail.mud.yahoo.com>
URL: https://gitlist.dev/e/20060109225143.60520.qmail%40web31807.mail.mud.yahoo.com
In-Reply-To: <Pine.LNX.4.64.0601090850350.3169@g5.osdl.org>

```
--- Linus Torvalds <torvalds@osdl.org> wrote:
> But if there are workflow problems, let's try to fix them. The "apply 
> patches directly to another branch" suggestion may not be sane (maybe it's 
> too confusing to apply a patch and not actually see it in the working 
> tree), but workflow suggestions in general are appreciated.

This is sensible, thank you.

A very general workflow I've seen people use is more/less as
I outlined in my previous email:

  tree A  (linus' or trunk)
     Project B  (Tree B)
        Project C  (Tree C, depending on stuff in Project B)

Now this could be how the "managers" see things, but development,
could've "cloned" from Tree B and Tree C further, as is often
customary to have a a) per user tree, or b) per bug tree.

So pull/merge/fetch/whatever follows Tree A->B->C.

It is sensible to have another tree say, called something
like "for_linus" or "upstream" or "product" which includes
what has accumulated in C from B and in B from A, (eq diff(C-A)).
I.e. a "push" tree.  So that I can tell you, "hey,
pull/fetch/merge/whatever the current verb en vogue is, from
here to get latest xyz".

What I also wanted to mention is that Tree B undeniably
depends on the _latest_ state of Tree A, since Project B
uses API/behaviour of the code in Tree A, so one cannot just
say they are independent.  Similarly for Tree C/Project C,
is dependent on B, and dependent on A.

Also sometimes a bugfix in C, prompts a bugfix in A,
so that the bugfix in A doesn't apply unless the bugfix in C.
(To get things more complicated.)

I think this is more/less the most easier to see, understand and
follow workflow approach, which is also the case for other SCMs.

What are the commands to follow to make everyone happy when
pulling from such a development process?

FWIW, "git diff A C | send to Linus" would get you the
"no merge messages/ancestors I want to see" idea, if I understand
this thread correctly.

> We've made switching branches about as efficient as it can be (but if the 
> differences are huge, the cost of re-writing the working directory is 
> never going to be low). But switching branches has the "confusion factor" 
> (ie you forget which branch you're on, and apply a patch to your working 
> branch instead of your development branch), so maybe there are other ways 
> of doing the same thing that might be sensible..

Yes.  Ever since I started used git, I never used branch
switching, but I do have git branches and I do use git branching.

I basically have a branch per directory, whereby the object db
is shared as is remotes/refs/etc, HEAD and index are not shared
of course.

This allows me to do a simple and fast "cd" to change/go to a
different branch, since they are in different directories.
So the time I wait to switch branches is the time the filesystem
takes to do a "cd".

This also allows me to build/test/patch/work on branches
simultaneously.

Thank you,
   Luben

```

## Linus Torvalds, 2006-01-09 23:07

Subject: RE: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.64.0601091502200.5588@g5.osdl.org>
URL: https://gitlist.dev/e/Pine.LNX.4.64.0601091502200.5588%40g5.osdl.org
In-Reply-To: <20060109225143.60520.qmail@web31807.mail.mud.yahoo.com>

```


On Mon, 9 Jan 2006, Luben Tuikov wrote:
> 
> Yes.  Ever since I started used git, I never used branch
> switching, but I do have git branches and I do use git branching.
> 
> I basically have a branch per directory, whereby the object db
> is shared as is remotes/refs/etc, HEAD and index are not shared
> of course.
> 
> This allows me to do a simple and fast "cd" to change/go to a
> different branch, since they are in different directories.
> So the time I wait to switch branches is the time the filesystem
> takes to do a "cd".
> 
> This also allows me to build/test/patch/work on branches
> simultaneously.

Yes. It has many advantages, and it's the approach I pushed pretty hard 
originally, but the "many branches in the same tree" approach seems to 
have become the more common one. Using many branches in the same tree is 
definitely the better approach for _distribution_, but that doesn't 
necessarily mean that it's the better one for development.

For example, you can have a git distribution tree with 20 different 
branches on kernel.org, but do development in 20 different trees with just 
one branch active - and when you do a "git push" to push out your branch 
in your development tree, it just updates that one branch on the 
distribution site.

So git certainly supports that kind of behaviour, but nobody I know 
actually does it that way (not even me, but since I tend to just merge 
other peoples code, I don't actually have multiple branches: I create 
temporary branches for one-off things, but don't maintain them that way).

			Linus

```

## Martin Langhoff, 2006-01-09 23:34

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <46a038f90601091534s7f4b36a5he05778f1ed82f34@mail.gmail.com>
URL: https://gitlist.dev/e/46a038f90601091534s7f4b36a5he05778f1ed82f34%40mail.gmail.com
In-Reply-To: <Pine.LNX.4.64.0601091502200.5588@g5.osdl.org>

```
On 1/10/06, Linus Torvalds <torvalds@osdl.org> wrote:
> Using many branches in the same tree is
> definitely the better approach for _distribution_, but that doesn't
> necessarily mean that it's the better one for development.
(...)
> So git certainly supports that kind of behaviour, but nobody I know
> actually does it that way

Hrm! We do. http://locke.catalyst.net.nz/gitweb?p=moodle.git;a=heads
shows a lot of heads that share 99% of the code. The repo is ~90MB --
and we check each head out with cogito, develop and push. It is a
shared team repo, using git+ssh and sticky gid and umask 002.

Works pretty well I have to add. The only odd thing is that the
fastest way to actually start working on a new branch is to ssh on to
the server and cp moodle.git/refs/heads/{foo,bar} and then cg-clone
that bar branch away. Perhaps I should code up an 'cg-branch-add
--in-server' patch.

regards,


martin

```

## Linus Torvalds, 2006-01-10 02:50

Subject: RE: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.64.0601091845160.5588@g5.osdl.org>
URL: https://gitlist.dev/e/Pine.LNX.4.64.0601091845160.5588%40g5.osdl.org
In-Reply-To: <20060109225143.60520.qmail@web31807.mail.mud.yahoo.com>

```


On Mon, 9 Jan 2006, Luben Tuikov wrote:
> 
> A very general workflow I've seen people use is more/less as
> I outlined in my previous email:
> 
>   tree A  (linus' or trunk)
>      Project B  (Tree B)
>         Project C  (Tree C, depending on stuff in Project B)
> 
> Now this could be how the "managers" see things, but development,
> could've "cloned" from Tree B and Tree C further, as is often
> customary to have a a) per user tree, or b) per bug tree.
> 
> So pull/merge/fetch/whatever follows Tree A->B->C.
> 
> It is sensible to have another tree say, called something
> like "for_linus" or "upstream" or "product" which includes
> what has accumulated in C from B and in B from A, (eq diff(C-A)).
> I.e. a "push" tree.  So that I can tell you, "hey,
> pull/fetch/merge/whatever the current verb en vogue is, from
> here to get latest xyz".
> 
> What I also wanted to mention is that Tree B undeniably
> depends on the _latest_ state of Tree A, since Project B
> uses API/behaviour of the code in Tree A, so one cannot just
> say they are independent.  Similarly for Tree C/Project C,
> is dependent on B, and dependent on A.

Note that in the case where the _latest_ state of the tre you are tracking 
really matters, then doing a "git pull" is absolutely and unquestionably 
the right thing to do. 

So if people thought that I don't want to have sub-maintainers pulling 
from my tree _at_all_, then that was a mis-communication. I don't in any 
way require a linear history, and criss-cross merges are supported 
perfectly well by git, and even encouraged in those situations.

After all, if tree B starts using features that are new to tree A, then 
the merge from A->B is required for functionality, and the synchronization 
is a fundamental part of the history of development. In that cases, the 
history complexity of the resulting tree is a result of real development 
complexity.

Now, obviously, for various reasons we want to avoid having those kinds of 
linkages as much as possible. We like to have develpment of different 
subsystems as independent as possible, not because it makes for a "more 
readable history", but because it makes it a lot easier to debug - if we 
have three independent features/development trees, they can be debugged 
independently too, while any linkages inevitably also mean that any bugs 
end up being interlinked..

		Linus

```

## Junio C Hamano, 2006-01-10 03:04

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <7v4q4cbx6l.fsf@assigned-by-dhcp.cox.net>
URL: https://gitlist.dev/e/7v4q4cbx6l.fsf%40assigned-by-dhcp.cox.net
In-Reply-To: <Pine.LNX.4.64.0601091845160.5588@g5.osdl.org>

```
Linus Torvalds <torvalds@osdl.org> writes:

> Now, obviously, for various reasons we want to avoid having those kinds of 
> linkages as much as possible. We like to have develpment of different 
> subsystems as independent as possible, not because it makes for a "more 
> readable history", but because it makes it a lot easier to debug - if we 
> have three independent features/development trees, they can be debugged 
> independently too, while any linkages inevitably also mean that any bugs 
> end up being interlinked..
>
> 		Linus

Yes.  If subproject B uses new features from A (either upstream
or sibling subproject), pulling A into B is inevitable.

On the other hand, if such merges becomes too frequent, it may
be a sign that A's feature set and interface is still changing
too rapidly for downstream use, but developers A and B are not
communicating well and B has not noticed that B might be better
off taking a break, addressing other non-overlapping areas while
giving a bit of time for A to settle things down.

An SCM is just _one_ of the ways for developers to communicate,
it will never be a replacement for developer communication.

```

## Kyle Moffett, 2006-01-10 06:33

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <99D82C29-4F19-4DD3-A961-698C3FC0631D@mac.com>
URL: https://gitlist.dev/e/99D82C29-4F19-4DD3-A961-698C3FC0631D%40mac.com
In-Reply-To: <Pine.LNX.4.64.0601091845160.5588-hNm40g4Ew95AfugRpC6u6w@public.gmane.org>

```
On Jan 09, 2006, at 21:50, Linus Torvalds wrote:
> if we  have three independent features/development trees, they can  
> be debugged independently too, while any linkages inevitably also  
> mean that any bugs end up being interlinked..

One example:

If I have ACPI, netdev, and swsusp trees change between an older  
version and a newer one, and my net driver starts breaking during  
suspend, I would be happiest debugging with the following set of  
patches/trees (Heavily simplified):

            ^
            |
           [5]
            |
          broken
         ^  ^   ^
       [2] [3]  [4]
       /    |     \
netdev3  acpi3   swsusp3
    ^       ^        ^
    |       |        |
netdev2  acpi2   swsusp2
    ^       ^        ^
    |       |        |
netdev1  acpi1   swsusp1
       ^    ^    ^
        \   |   /
         \  |  /
          \ | /
           \|/
            |
           [1]
            |
          works


If the old version [1] works and the new one [5] doesn't, then I can  
immediately test [2], [3], and [4].  If one of those doesn't work,  
I've identified the problematic patchset and cut the debugging by  
2/3.  If they all work, then we know precisely that it's the  
interactions between them, which also makes debugging a lot easier.

Cheers,
Kyle Moffett

--
There are two ways of constructing a software design. One way is to  
make it so simple that there are obviously no deficiencies. And the  
other way is to make it so complicated that there are no obvious  
deficiencies.  The first method is far more difficult.
   -- C.A.R. Hoare


-
To unsubscribe from this list: send the line "unsubscribe linux-acpi" in
the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

```

## Martin Langhoff, 2006-01-10 06:38

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <46a038f90601092238r3476556apf948bfe5247da484@mail.gmail.com>
URL: https://gitlist.dev/e/46a038f90601092238r3476556apf948bfe5247da484%40mail.gmail.com
In-Reply-To: <99D82C29-4F19-4DD3-A961-698C3FC0631D-ee4meeAH724@public.gmane.org>

```
On 1/10/06, Kyle Moffett <mrmacman_g4-ee4meeAH724@public.gmane.org> wrote:
> If they all work, then we know precisely that it's the
> interactions between them, which also makes debugging a lot easier.

The more complex your tree structure is, the more the interactions are
likely to be part of the problem. Is git-bisect not useful in this
scenario?

cheers,


martin
-
To unsubscribe from this list: send the line "unsubscribe linux-acpi" in
the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

```

## Kyle Moffett, 2006-01-10 18:05

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <252A408D-0B42-49F3-92BC-B80F94F19F40@mac.com>
URL: https://gitlist.dev/e/252A408D-0B42-49F3-92BC-B80F94F19F40%40mac.com
In-Reply-To: <46a038f90601092238r3476556apf948bfe5247da484@mail.gmail.com>

```
On Jan 10, 2006, at 01:38, Martin Langhoff wrote:
> On 1/10/06, Kyle Moffett <mrmacman_g4@mac.com> wrote:
>> If they all work, then we know precisely that it's the  
>> interactions between them, which also makes debugging a lot easier.
>
> The more complex your tree structure is, the more the interactions  
> are likely to be part of the problem. Is git-bisect not useful in  
> this scenario?

IIRC git-bisect just does an outright linearization of the whole tree  
anyways, which makes git-bisect work everywhere, even in the presence  
of difficult cross-merges.  On the other hand, if you are git- 
bisecting ACPI changes (perhaps due to some ACPI breakage), and ACPI  
has 10 pulls from mainline, you _also_ have to wade through the  
bisection of any other changes that occurred in mainline, even if  
they're totally irrelevant.  This is why it's useful to only pull  
mainline into your tree (EX: ACPI) when you functionally depend on  
changes there (as Linus so eloquently expounded upon).

Cheers,
Kyle Moffett

--
Q: Why do programmers confuse Halloween and Christmas?
A: Because OCT 31 == DEC 25.

```

## Linus Torvalds, 2006-01-10 18:27

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.64.0601101015260.4939@g5.osdl.org>
URL: https://gitlist.dev/e/Pine.LNX.4.64.0601101015260.4939%40g5.osdl.org
In-Reply-To: <252A408D-0B42-49F3-92BC-B80F94F19F40@mac.com>

```

On Tue, 10 Jan 2006, Kyle Moffett wrote:
>
> On Jan 10, 2006, at 01:38, Martin Langhoff wrote:
> > 
> > The more complex your tree structure is, the more the interactions are
> > likely to be part of the problem. Is git-bisect not useful in this scenario?
> 
> IIRC git-bisect just does an outright linearization of the whole tree anyways,
> which makes git-bisect work everywhere, even in the presence of difficult
> cross-merges.

It's not really a linearization - at no time does git-bisect _order_ the 
commits. After all, no linear order actually exists. 

Instead, it really cuts the tree up into successively smaller parts. 

Think of it as doing a binary search in a 2-dimensional surface - you 
can't linearize the plane, but you can decide to test first one half of 
the surface, and then depending on whether it was there, you can halve 
that surface etc.. 

> On the other hand, if you are git-bisecting ACPI changes
> (perhaps due to some ACPI breakage), and ACPI has 10 pulls from mainline, you
> _also_ have to wade through the bisection of any other changes that occurred
> in mainline, even if they're totally irrelevant.

Yes. Although if you _know_ that the problem happened in a specific file 
or specific subdirectory, you can actually tell "git bisect" to only 
bother with changes to that file/directory/set-of-directories to speed up 
the search.

IOW, if you absolutely know that it's ACPI-related, you can do something 
like

	git bisect start drivers/acpi arch/i386/kernel/acpi

to tell the bisect code that it should totally ignore anything that 
doesn't touch those two directories.

However, if it turns out that you were wrong (and the ACPI breakage was 
brought on by something that changed something else), "git bisect" will 
just get confused and report the wrong commit, so this is really something 
you should be careful with (and verify the end result by checking that 
undoing that _particular_ commit really fixes things).

And yes, "git bisect" _will_ work with bugs that depend on two branches of 
a merge: it will point to the merge commit itself as being the problem. 
Now, at that point you really are screwed, and you'll have to figure out 
why both branches work, but the combination of them do not.

Maybe it's as simple as just a merge done wrong (bad manual fixups), but 
maybe it's a perfectly executed merge that just happens to have one branch 
changing the assumptions that the other branch depended on.

Happily, that is not very common. I know people are using "git bisect", 
and I don't think anybody has ever reported it so far. It will happen 
eventually, but I'd actually expect it to be much more common that "git 
bisect" will hit other - worse - problems, like bugs that "come and go", 
and that a simple bisection simply cannot find because they aren't 
totally repeatable.

			Linus

```

## Johannes Schindelin, 2006-01-10 18:45

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.63.0601101938420.26999@wbgn013.biozentrum.uni-wuerzburg.de>
URL: https://gitlist.dev/e/Pine.LNX.4.63.0601101938420.26999%40wbgn013.biozentrum.uni-wuerzburg.de
In-Reply-To: <Pine.LNX.4.64.0601101015260.4939-hNm40g4Ew95AfugRpC6u6w@public.gmane.org>

```
Hi,

On Tue, 10 Jan 2006, Linus Torvalds wrote:

> 
> On Tue, 10 Jan 2006, Kyle Moffett wrote:
> >
> > On Jan 10, 2006, at 01:38, Martin Langhoff wrote:
> > > 
> > > The more complex your tree structure is, the more the interactions are
> > > likely to be part of the problem. Is git-bisect not useful in this scenario?
> > 
> > IIRC git-bisect just does an outright linearization of the whole tree anyways,
> > which makes git-bisect work everywhere, even in the presence of difficult
> > cross-merges.
> 
> It's not really a linearization - at no time does git-bisect _order_ the 
> commits. After all, no linear order actually exists. 
> 
> Instead, it really cuts the tree up into successively smaller parts. 
> 
> Think of it as doing a binary search in a 2-dimensional surface - you 
> can't linearize the plane, but you can decide to test first one half of 
> the surface, and then depending on whether it was there, you can halve 
> that surface etc.. 

How?

If you bisect, you test a commit. If the commit is bad, you assume *all* 
commits before that as bad. If it is good, you assume *all* commits after 
that as good.

Now, if you have a 2-dimensional surface, you don't have a *point*, but 
typically a *line* separating good from bad.

Further, the comparison with 2 dimensions is particularly bad. You 
*have* partially linear development lines, it got *nothing* to do with 
an area. The commits still make up a *list*, and it depends how you 
*order* that list for bisect. (And don't tell me they are not ordered: 
they are.)

If you order the commits by date, you don't get anything meaningful point 
before which it is bad, and after which it is good.

So, how is bisect supposed to work if you don't have one straight 
development line from bad to good?

Ciao,
Dscho


-
To unsubscribe from this list: send the line "unsubscribe linux-acpi" in
the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

```

## Linus Torvalds, 2006-01-10 19:01

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.64.0601101048440.4939@g5.osdl.org>
URL: https://gitlist.dev/e/Pine.LNX.4.64.0601101048440.4939%40g5.osdl.org
In-Reply-To: <Pine.LNX.4.63.0601101938420.26999@wbgn013.biozentrum.uni-wuerzburg.de>

```


On Tue, 10 Jan 2006, Johannes Schindelin wrote:
> > 
> > Think of it as doing a binary search in a 2-dimensional surface - you 
> > can't linearize the plane, but you can decide to test first one half of 
> > the surface, and then depending on whether it was there, you can halve 
> > that surface etc.. 
> 
> How?
> 
> If you bisect, you test a commit. If the commit is bad, you assume *all* 
> commits before that as bad. If it is good, you assume *all* commits after 
> that as good.

No, that's not how bisect works at all.

It's true that if a commit is bad, then all the commits _reachable_ from 
that commit are considered bad. 

And it's true that if a commit is good, then all commits that _reach_ that 
commit are considered good.

But that doesn't mean that there is an ordering. The commits that fall 
into the camp of being "neither good nor bad" are _not_ ordered. There are 
commits in there that are not directly reachable from the good commit.

> Now, if you have a 2-dimensional surface, you don't have a *point*, but 
> typically a *line* separating good from bad.

Exactly. 

And a git graph is not really a two-dimensional surface, but exactly was 
with a 2-dimensional surface, it is _not_ enough to have a *point* to 
separate the good from bad.

You need to have a _set of points_ to separate the good from the bad. You 
can think of it as a line that bisects the surface: if you were to print 
out the development graph, the set of points literally _do_ form a virtual 
line across the development surface.

(Actually, you can't in general print out the development graph on a 
2-dimensional paper without having development lines that cross each 
other, but you could actually do it in three dimensions, where the 
"boundary" between good and bad is actually a 2-dimensional surface in 
3-dimensional space).

But to describe the surface of "known good", you actually just need a list 
of known good commits, and the "commits reachable from those commits" 
_becomes_ the surface.

> Further, the comparison with 2 dimensions is particularly bad.

No it is not. It's a very good comparison.

In a linearized model (one-dimensional, fully ordered set), the only thing 
you need for bisection is two points: the beginning and the end.

In the git model, you need _many_ points to describe the area being 
bisected. Exactly the same way as if you were to bisect a 2-dimensional 
surface.

Now, the git history is _not_ really a two-dimensional surface, so it's 
just an analogy, not an exact identity. But from a visualization 
standpoint, it's a good way to think of each "git bisect" as adding a 
_line_ on the surface rather than a point on a linear line.

> So, how is bisect supposed to work if you don't have one straight 
> development line from bad to good?

Read the code.

I'm pretty proud of it. It's simple, and it's obvious once you think about 
it, but it is pretty novel as far as I know. BK certainly had nothing 
similar, not have I heard of anythign else that does it. Git _might_ be 
the first thing that has ever done it, although it's simple enough that I 
wouldn't be surprised if others have too.

			Linus

```

## Linus Torvalds, 2006-01-10 19:28

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.64.0601101111110.4939@g5.osdl.org>
URL: https://gitlist.dev/e/Pine.LNX.4.64.0601101111110.4939%40g5.osdl.org
In-Reply-To: <Pine.LNX.4.64.0601101048440.4939@g5.osdl.org>

```


On Tue, 10 Jan 2006, Linus Torvalds wrote:
> 
> Now, the git history is _not_ really a two-dimensional surface, so it's 
> just an analogy, not an exact identity. But from a visualization 
> standpoint, it's a good way to think of each "git bisect" as adding a 
> _line_ on the surface rather than a point on a linear line.

Actually, the way I think of it is akin to the "light cones" in physics. A 
point in space-time doesn't define a fully ordered "before and after": but 
it _does_ describe a "light cone" which tells you what is reachable from 
that point, and what that point reaches. Within those cones, that 
particular point ("commit") has a strict ordering.

And exactly as in physics, in git there's a lot of space that is _not_ 
ordered by that commit. And the way to bisect is basically to find the 
right points in "git space" to create the right "light cone" that you 
find the point where the git space that is reachable from that commit has 
the same volume as the git space that isn't reachable.

And maybe that makes more sense to you (if you're into physics), or maybe 
it makes less sense to you.

Now, since we always search the "git space" in the cone that is defined by 
"reachable from the bad commit, but not reachable from any good commit", 
the way we handle "bad" and "good" is actually not a mirror-image. If we 
fine a new _bad_ commit, we know that it was reachable from the old bad 
commit, and thus the old bad commit is now uninteresting: the new bad 
commit forms a "past light cone" that is a strict subset of the old one, 
so we can totally discard the old bad commit from any future 
consideration. It doesn't tell us anything new.

In contrast, if we find a new _good_ commit, the "past light cone" (aka 
"set of commits reachable from it") is -not- necessarily a proper superset 
of the previous set of good commits, so when we find a good commit, we 
still need to carry the _other_ good commits around, and the "known good" 
universe is the _union_ of all the "good commit past lightcones".

Then the "unknown space" is the set difference of the "past lightcone of 
the bad commit" and of this "union of past lightcones of good commits". 
It's the space that is reachable from the known-bad commit, but not 
reachable from any known-good commit.

So this means that when doing bisection, what we want to do is find the 
point in git space that has _new_ "reachability" within that unknown space 
that is as close to half that volume as space as possible. And that's 
exactly what "git-rev-list --bisect" calculates.

So every time, we try to either move the "known bad" light-cone down in 
time in the unknown space, _or_ we add a new "known good" light-cone. In 
either case, the "unknown git space" keeps shrinking by half each time.

("by half" is not exact, because git space is not only quanticized, it 
also has a rather strange "distance function". In other words, we're 
talking about a rather strange space. The good news is that the space is 
small enough that we can just enumerate every quantum and simply 
calculate the volume it defines in that space. IOW, we do a very 
brute-force thing, and it works fine).

			Linus

```

## Johannes Schindelin, 2006-01-10 19:38

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.63.0601102010100.27199@wbgn013.biozentrum.uni-wuerzburg.de>
URL: https://gitlist.dev/e/Pine.LNX.4.63.0601102010100.27199%40wbgn013.biozentrum.uni-wuerzburg.de
In-Reply-To: <Pine.LNX.4.64.0601101048440.4939@g5.osdl.org>

```
Hi,

[cut down the Cc: list, since this is getting special]

On Tue, 10 Jan 2006, Linus Torvalds wrote:

> > If you bisect, you test a commit. If the commit is bad, you assume *all* 
> > commits before that as bad. If it is good, you assume *all* commits after 
> > that as good.
> 
> No, that's not how bisect works at all.

Okay, so I got that wrong. But for a good reason: this is not the meaning 
of bisection in my lectures. Doesn't matter.

> It's true that if a commit is bad, then all the commits _reachable_ from 
> that commit are considered bad. 
> 
> And it's true that if a commit is good, then all commits that _reach_ that 
> commit are considered good.
> 
> But that doesn't mean that there is an ordering. The commits that fall 
> into the camp of being "neither good nor bad" are _not_ ordered. There are 
> commits in there that are not directly reachable from the good commit.

Those commits not reachable from the good commit are of no interest. Let's 
just ignore them.

> > Now, if you have a 2-dimensional surface, you don't have a *point*, but 
> > typically a *line* separating good from bad.
> 
> Exactly. 
> 
> And a git graph is not really a two-dimensional surface, but exactly was 
> with a 2-dimensional surface, it is _not_ enough to have a *point* to 
> separate the good from bad.
> 
> You need to have a _set of points_ to separate the good from the bad. You 
> can think of it as a line that bisects the surface: if you were to print 
> out the development graph, the set of points literally _do_ form a virtual 
> line across the development surface.

Okay, so there is a cut: Every directed path from good to bad has a single 
commit which is the first bad. Let's call the set of all such bad commits 
the cut set.

Is git-bisect capable of identifying all of the cut set, or just a single 
one?

> > Further, the comparison with 2 dimensions is particularly bad.
> 
> No it is not. It's a very good comparison.

>From your explanation I understand now why you like that comparison.

> > So, how is bisect supposed to work if you don't have one straight 
> > development line from bad to good?
> 
> Read the code.
> 
> I'm pretty proud of it.

I bet nobody can tell ;-)

Well, I read the code. And I answer my own question from 18--19 lines ago:

git-bisect is not capable of identifying the cut set, but pretends that 
there really is only one bad commit (see bisect_bad()).

That may be the best choice if all commits in the cut set except one are 
merges. (It is the best if the cut set contains only one element.)

But I see two problems with that:

- a problem can be introduced independently in two different branches, and
  occur in both of them before the merge (in which case bisect only 
  catches one of the commits), and

- AFAICT if the cut set is one merge and one regular commit, bisect could
  identify the merge by error.

Of course, all this makes only a difference if the bisect has to cross a 
merge.

BTW I think there is a thinko in git-rev-list.txt:

> Thus, if 'git-rev-list --bisect foo ^bar ^baz' outputs 'midpoint', the 
> output of 'git-rev-list foo ^midpoint' and 'git-rev-list midpoint ^bar 
             ^ this should be
	'git-rev-list foo ^midpoint ^bar ^baz'
> ^baz' would be of roughly the same length

Ciao,
Dscho

```

## Linus Torvalds, 2006-01-10 20:11

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.64.0601101151090.4939@g5.osdl.org>
URL: https://gitlist.dev/e/Pine.LNX.4.64.0601101151090.4939%40g5.osdl.org
In-Reply-To: <Pine.LNX.4.63.0601102010100.27199@wbgn013.biozentrum.uni-wuerzburg.de>

```


On Tue, 10 Jan 2006, Johannes Schindelin wrote:
> 
> Those commits not reachable from the good commit are of no interest. Let's 
> just ignore them.

Note that to avoid confusion, start talking about -multiple- good commits 
early.

So we have a list of "known good islands" in the git-space. And yes, we 
want to ignore anything that is reachable from them.

And here the magic part of"git-bisect.sh" is around line 133:

	... --not $(cd "$GIT_DIR" && ls refs/bisect/good-*) ...

It tells git-rev-parse to generate a list of commits that we're _not_ 
interested in, and that list will be one of the most critical parts of the 
stuff we give to "git-rev-list --bisect".

So that part of the script is literally the part that says "ignore all of 
git space that is reachable from the good commits", because we've listed 
all the good commits as refs named "refs/bisect/good-*".

> > You need to have a _set of points_ to separate the good from the bad. You 
> > can think of it as a line that bisects the surface: if you were to print 
> > out the development graph, the set of points literally _do_ form a virtual 
> > line across the development surface.
> 
> Okay, so there is a cut: Every directed path from good to bad has a single 
> commit which is the first bad. Let's call the set of all such bad commits 
> the cut set.

This set is uninteresting for two reasons:
 - it's hard to calculate
 - it's not the answer we want.

We want the _single_ commit that is the one that generates your "cut set".

Your "cut set" is really the "reachability border" from the single bad 
commit we're interested in to all the possible development lines.

In practice, the "cut set" is just the "bad commit" plus all the merges 
that merge that bad commit with somethign that wasn't reacable from it in 
the first place.

So the "cut set" isn't interesting.

> git-bisect is not capable of identifying the cut set, but pretends that 
> there really is only one bad commit (see bisect_bad()).

Not quite.

It could keep track of all bad commits (in fact, it does so in the log 
file), but the fact is, none but the lastest bad commit we have found 
matters.

By definition, "git bisect" is always going to test a commit that is 
reachable from the previously known bad commit. Agreed? Anything else 
would be insane - we know that we had a bad stat, and we're interested in 
finding out how _that_ bad state happened, so we're only ever interested 
in commits that are ancestors to that bad state.

So our search-space is _literally_ defined by two things:

 - the surface of "known good" commits (which defines the commits that 
   aren't interesting). 

   This is the "--not refs/bisect/good-*" part

 - the last "known bad" commit.

We'll always search the git commit space defined by these two knowns, 
agreed?

Now, realize that if we find a new bad commit, since that bad commit was 
by definition reachable from the _old_ bad commit (since we didn't even 
search outside its reachability), then equally by definition the 
reachability from that new bad commit is a strict superset of the 
reachability of the old bad commit.

So when we find a new bad commit, the old bad commit is no longer 
interesting.

So when you say "pretends that there really is only one bad commit", you 
didn't realize that it's not about "pretending". It's very fundamental: 
there is only ever _one_ bad commit that is interesting. It's the last one 
we found.

Even if we started out with two bad commits (ie some person reported two 
different versions as being bad), we're _still_ not interested in using 
them both. We should pick one of them, because the reachability area 
defined by two bad commits is always a superset of the reachability of 
either one.

So having multiple bad commits is _never_ interesting.

> But I see two problems with that:
> 
> - a problem can be introduced independently in two different branches, and
>   occur in both of them before the merge (in which case bisect only 
>   catches one of the commits), and

This is fine. Depending on whatever random factors, we'll test one of them 
first, and eventually find _one_ of the commits that fix it. If the exact 
same bug was introduced somewhere else, and merged, then undoing just the 
"one" bug will obviously undo the other one too.

If a _different_ bug was introduced (even if it had the same effects), 
yes, you now have two separate bugs. And bisecting two bugs is hard. You 
need to separate them out some way.

> - AFAICT if the cut set is one merge and one regular commit, bisect could
>   identify the merge by error.

It will never identify a commit without having done a full bisection, so 
if it ever had the choice of a "merge" and the "commit leading up to the 
merge", it will always have tried the "commit leading up to the merge", 
and decided that it was fundamentally more recent (had "smaller 
reachability") that the merge, and pinpoint it.

> BTW I think there is a thinko in git-rev-list.txt:
> 
> > Thus, if 'git-rev-list --bisect foo ^bar ^baz' outputs 'midpoint', the 
> > output of 'git-rev-list foo ^midpoint' and 'git-rev-list midpoint ^bar 
>              ^ this should be
> 	'git-rev-list foo ^midpoint ^bar ^baz'
> > ^baz' would be of roughly the same length

Yes.

			Linus

```

## Linus Torvalds, 2006-01-10 20:28

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.64.0601101221020.4939@g5.osdl.org>
URL: https://gitlist.dev/e/Pine.LNX.4.64.0601101221020.4939%40g5.osdl.org
In-Reply-To: <Pine.LNX.4.64.0601101151090.4939@g5.osdl.org>

```


On Tue, 10 Jan 2006, Linus Torvalds wrote:
> 
> If a _different_ bug was introduced (even if it had the same effects), 
> yes, you now have two separate bugs. And bisecting two bugs is hard. You 
> need to separate them out some way.

Side note: this is seldom a problem in practice. If it was effectively the 
same bug, just finding the one case that triggered it is sufficient: you 
then know what to look for, and if undoing that one commit isn't enough to 
fix it in the current tree (because the same bug existed in another form 
on another branch), you wouldn't actually start bisecting again. You'd 
start grepping the tree for other cases of that bug.

So the biggest advantage of "git bisect" is _not_ that you can just undo 
the buggy commit. In fact, usually you don't even want to undo it, because 
it probably had a raison-d'etre to begin with. The huge deal about "git 
bisect" is that it pinpoints what caused the bug, and then the fix is 
often something else.

Often it's a "Duh! I fixed one thing, but my fix didn't take Xyz into 
account, so it now broke for another reason" moment.

Most bugs are stupid, in other words.

The _real_ problem with git bisect is when you have a non-technical user 
(common) and there are silly bugs that you know of and already fixed that 
aren't really a problem, but that are show-stoppers for the user who isn't 
a kernel developer (or is, but doesn't know git). They're show-stoppers 
not because we care about them, but because they make the "purely 
mechanical" thing be one where you have to have some manual input.

Another problem (that I've not seen in practice yet, but that I bet _will_ 
be the worst issue) is non-reproducible bugs. They are the nastiest kind 
to debug in the first place, and sadly, "git bisect" simply doesn't help 
you with them. There, nothing but some luck and a lot of thinking and 
testing will help you.

			Linus

```

## Johannes Schindelin, 2006-01-10 20:47

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <Pine.LNX.4.63.0601102122001.30609@wbgn013.biozentrum.uni-wuerzburg.de>
URL: https://gitlist.dev/e/Pine.LNX.4.63.0601102122001.30609%40wbgn013.biozentrum.uni-wuerzburg.de
In-Reply-To: <Pine.LNX.4.64.0601101151090.4939@g5.osdl.org>

```
Hi,

On Tue, 10 Jan 2006, Linus Torvalds wrote:

> So having multiple bad commits is _never_ interesting.

Okay, I got it. A bug is supposed to be inherited by *all* its 
descendants. Good.

I have to keep in mind that a commit is not actually a patch set, but can 
be two or more (in case of a merge). So, a bug can be present in a 
development line for a long, long time, but be visible only after a merge. 
Since that commit can be compared to at least two trees, one of these 
diffs must show the bug.

Thanks,
Dscho

```

## Luben Tuikov, 2006-01-11 03:32

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <20060111033229.5590.qmail@web31809.mail.mud.yahoo.com>
URL: https://gitlist.dev/e/20060111033229.5590.qmail%40web31809.mail.mud.yahoo.com
In-Reply-To: <252A408D-0B42-49F3-92BC-B80F94F19F40-ee4meeAH724@public.gmane.org>

```
--- Kyle Moffett <mrmacman_g4-ee4meeAH724@public.gmane.org> wrote:
> they're totally irrelevant.  This is why it's useful to only pull  
> mainline into your tree (EX: ACPI) when you functionally depend on  
> changes there (as Linus so eloquently expounded upon).

Sometimes the dependency is _behavioural_.  For example certain
behaviour of other modules of the kernel changed and you want
to test that your module works ok with them under different
behaviour.  In which case you may or may not have to
change your code after the fact.

    Luben

-
To unsubscribe from this list: send the line "unsubscribe linux-acpi" in
the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

```

## Matthias Urlichs, 2006-01-13 23:35

Subject: Re: git pull on Linux/ACPI release tree
Message-ID: <pan.2006.01.13.23.34.58.269921@smurf.noris.de>
URL: https://gitlist.dev/e/pan.2006.01.13.23.34.58.269921%40smurf.noris.de
In-Reply-To: <Pine.LNX.4.64.0601101048440.4939@g5.osdl.org>

```
Hi, Linus Torvalds wrote:

> I'm pretty proud of it. It's simple, and it's obvious once you think about 
> it, but it is pretty novel as far as I know. BK certainly had nothing 
> similar, not have I heard of anythign else that does it.

Actually, I've written a hackish script that tries to do simple-minded
bisection (read: it searched for the 50% point on the shortest path
between any-of-GOOD and any-of-BAD, instead of considering the whole
graph) on BK trees. I haven't exactly published the thing anyplace though,
because, well, it was ugly. :-/

Besides, actually working with the current bisection point is no problem
at all for git. Doing the same thing in BK's world view is *painful*,
esp. given the size of the kernel tree.

-- 
Matthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de
Disclaimer: The quote was selected randomly. Really. | http://smurf.noris.de
 - -
Arthur felt at a bit of a loss. There was a whole Galaxy
of stuff out there for him, and he wondered if it was
churlish of him to complain to himself that it lacked just
two things: the world he was born on and the woman he loved.

```
