{"thread":{"id":"3023","subject":"RE: git pull on Linux/ACPI release tree","startedAt":"2006-01-09T16:57:03Z","lastAt":"2006-01-13T23:35:01Z","messageCount":19,"participants":["Luben Tuikov","Linus Torvalds","Martin Langhoff","Junio C Hamano","Kyle Moffett","Johannes Schindelin","Matthias Urlichs"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"15378","messageId":"Pine.LNX.4.64.0601090850350.3169@g5.osdl.org","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601090835580.3169@g5.osdl.org","subject":"RE: git pull on Linux/ACPI release tree","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-01-09T16:57:03Z","receivedAt":"2006-01-09T16:57:03Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 9 Jan 2006, Linus Torvalds wrote:\n>\n> One thing we could do is to make it easier to apply a patch to a \n> _non_current_ branch.\n>   [ ... ]\n> Do you think that kind of workflow would be more palatable to you? It \n> shouldn't be /that/ hard to make git-apply branch-aware... (It was part of \n> my original plan, but it is more work than just using the working \n> directory, so I never finished the thought).\n\nBtw, this is true in a bigger sense: the things \"git\" does have largely \nbeen driven by user needs. Initially mainly mine, but things like \n\"git-rebase\" were from people who wanted to work as \"sub-maintainers\" (eg \nJunio before he became the head honcho for git itself).\n\nBut if there are workflow problems, let's try to fix them. The \"apply \npatches directly to another branch\" suggestion may not be sane (maybe it's \ntoo confusing to apply a patch and not actually see it in the working \ntree), but workflow suggestions in general are appreciated.\n\nWe've made switching branches about as efficient as it can be (but if the \ndifferences are huge, the cost of re-writing the working directory is \nnever going to be low). But switching branches has the \"confusion factor\" \n(ie you forget which branch you're on, and apply a patch to your working \nbranch instead of your development branch), so maybe there are other ways \nof doing the same thing that might be sensible..\n\nSo send suggestions to the git lists. Maybe they're insane and can't be \ndone, but while I designed git to work with _my_ case (ie mostly merging \ntons of different trees and then having occasional big batches of \npatches), it's certainly _supposed_ to support other maintainers too..\n\n\t\tLinus\n"},{"id":"14368","messageId":"20060109225143.60520.qmail@web31807.mail.mud.yahoo.com","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601090850350.3169@g5.osdl.org","subject":"RE: git pull on Linux/ACPI release tree","fromName":"Luben Tuikov","fromEmail":"ltuikov@yahoo.com","sentAt":"2006-01-09T22:51:43Z","receivedAt":"2006-01-09T22:51:43Z","isPatch":false,"sender":{"key":"ltuikov@yahoo.com","avatar":null},"body":"--- Linus Torvalds <torvalds@osdl.org> wrote:\n> But if there are workflow problems, let's try to fix them. The \"apply \n> patches directly to another branch\" suggestion may not be sane (maybe it's \n> too confusing to apply a patch and not actually see it in the working \n> tree), but workflow suggestions in general are appreciated.\n\nThis is sensible, thank you.\n\nA very general workflow I've seen people use is more/less as\nI outlined in my previous email:\n\n  tree A  (linus' or trunk)\n     Project B  (Tree B)\n        Project C  (Tree C, depending on stuff in Project B)\n\nNow this could be how the \"managers\" see things, but development,\ncould've \"cloned\" from Tree B and Tree C further, as is often\ncustomary to have a a) per user tree, or b) per bug tree.\n\nSo pull/merge/fetch/whatever follows Tree A->B->C.\n\nIt is sensible to have another tree say, called something\nlike \"for_linus\" or \"upstream\" or \"product\" which includes\nwhat has accumulated in C from B and in B from A, (eq diff(C-A)).\nI.e. a \"push\" tree.  So that I can tell you, \"hey,\npull/fetch/merge/whatever the current verb en vogue is, from\nhere to get latest xyz\".\n\nWhat I also wanted to mention is that Tree B undeniably\ndepends on the _latest_ state of Tree A, since Project B\nuses API/behaviour of the code in Tree A, so one cannot just\nsay they are independent.  Similarly for Tree C/Project C,\nis dependent on B, and dependent on A.\n\nAlso sometimes a bugfix in C, prompts a bugfix in A,\nso that the bugfix in A doesn't apply unless the bugfix in C.\n(To get things more complicated.)\n\nI think this is more/less the most easier to see, understand and\nfollow workflow approach, which is also the case for other SCMs.\n\nWhat are the commands to follow to make everyone happy when\npulling from such a development process?\n\nFWIW, \"git diff A C | send to Linus\" would get you the\n\"no merge messages/ancestors I want to see\" idea, if I understand\nthis thread correctly.\n\n> We've made switching branches about as efficient as it can be (but if the \n> differences are huge, the cost of re-writing the working directory is \n> never going to be low). But switching branches has the \"confusion factor\" \n> (ie you forget which branch you're on, and apply a patch to your working \n> branch instead of your development branch), so maybe there are other ways \n> of doing the same thing that might be sensible..\n\nYes.  Ever since I started used git, I never used branch\nswitching, but I do have git branches and I do use git branching.\n\nI basically have a branch per directory, whereby the object db\nis shared as is remotes/refs/etc, HEAD and index are not shared\nof course.\n\nThis allows me to do a simple and fast \"cd\" to change/go to a\ndifferent branch, since they are in different directories.\nSo the time I wait to switch branches is the time the filesystem\ntakes to do a \"cd\".\n\nThis also allows me to build/test/patch/work on branches\nsimultaneously.\n\nThank you,\n   Luben\n"},{"id":"14371","messageId":"Pine.LNX.4.64.0601091502200.5588@g5.osdl.org","threadId":"3023","inReplyTo":"20060109225143.60520.qmail@web31807.mail.mud.yahoo.com","subject":"RE: git pull on Linux/ACPI release tree","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-01-09T23:07:45Z","receivedAt":"2006-01-09T23:07:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 9 Jan 2006, Luben Tuikov wrote:\n> \n> Yes.  Ever since I started used git, I never used branch\n> switching, but I do have git branches and I do use git branching.\n> \n> I basically have a branch per directory, whereby the object db\n> is shared as is remotes/refs/etc, HEAD and index are not shared\n> of course.\n> \n> This allows me to do a simple and fast \"cd\" to change/go to a\n> different branch, since they are in different directories.\n> So the time I wait to switch branches is the time the filesystem\n> takes to do a \"cd\".\n> \n> This also allows me to build/test/patch/work on branches\n> simultaneously.\n\nYes. It has many advantages, and it's the approach I pushed pretty hard \noriginally, but the \"many branches in the same tree\" approach seems to \nhave become the more common one. Using many branches in the same tree is \ndefinitely the better approach for _distribution_, but that doesn't \nnecessarily mean that it's the better one for development.\n\nFor example, you can have a git distribution tree with 20 different \nbranches on kernel.org, but do development in 20 different trees with just \none branch active - and when you do a \"git push\" to push out your branch \nin your development tree, it just updates that one branch on the \ndistribution site.\n\nSo git certainly supports that kind of behaviour, but nobody I know \nactually does it that way (not even me, but since I tend to just merge \nother peoples code, I don't actually have multiple branches: I create \ntemporary branches for one-off things, but don't maintain them that way).\n\n\t\t\tLinus\n"},{"id":"14374","messageId":"46a038f90601091534s7f4b36a5he05778f1ed82f34@mail.gmail.com","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601091502200.5588@g5.osdl.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-01-09T23:34:39Z","receivedAt":"2006-01-09T23:34:39Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 1/10/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> Using many branches in the same tree is\n> definitely the better approach for _distribution_, but that doesn't\n> necessarily mean that it's the better one for development.\n(...)\n> So git certainly supports that kind of behaviour, but nobody I know\n> actually does it that way\n\nHrm! We do. http://locke.catalyst.net.nz/gitweb?p=moodle.git;a=heads\nshows a lot of heads that share 99% of the code. The repo is ~90MB --\nand we check each head out with cogito, develop and push. It is a\nshared team repo, using git+ssh and sticky gid and umask 002.\n\nWorks pretty well I have to add. The only odd thing is that the\nfastest way to actually start working on a new branch is to ssh on to\nthe server and cp moodle.git/refs/heads/{foo,bar} and then cg-clone\nthat bar branch away. Perhaps I should code up an 'cg-branch-add\n--in-server' patch.\n\nregards,\n\n\nmartin\n"},{"id":"14384","messageId":"Pine.LNX.4.64.0601091845160.5588@g5.osdl.org","threadId":"3023","inReplyTo":"20060109225143.60520.qmail@web31807.mail.mud.yahoo.com","subject":"RE: git pull on Linux/ACPI release tree","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-01-10T02:50:14Z","receivedAt":"2006-01-10T02:50:14Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 9 Jan 2006, Luben Tuikov wrote:\n> \n> A very general workflow I've seen people use is more/less as\n> I outlined in my previous email:\n> \n>   tree A  (linus' or trunk)\n>      Project B  (Tree B)\n>         Project C  (Tree C, depending on stuff in Project B)\n> \n> Now this could be how the \"managers\" see things, but development,\n> could've \"cloned\" from Tree B and Tree C further, as is often\n> customary to have a a) per user tree, or b) per bug tree.\n> \n> So pull/merge/fetch/whatever follows Tree A->B->C.\n> \n> It is sensible to have another tree say, called something\n> like \"for_linus\" or \"upstream\" or \"product\" which includes\n> what has accumulated in C from B and in B from A, (eq diff(C-A)).\n> I.e. a \"push\" tree.  So that I can tell you, \"hey,\n> pull/fetch/merge/whatever the current verb en vogue is, from\n> here to get latest xyz\".\n> \n> What I also wanted to mention is that Tree B undeniably\n> depends on the _latest_ state of Tree A, since Project B\n> uses API/behaviour of the code in Tree A, so one cannot just\n> say they are independent.  Similarly for Tree C/Project C,\n> is dependent on B, and dependent on A.\n\nNote that in the case where the _latest_ state of the tre you are tracking \nreally matters, then doing a \"git pull\" is absolutely and unquestionably \nthe right thing to do. \n\nSo if people thought that I don't want to have sub-maintainers pulling \nfrom my tree _at_all_, then that was a mis-communication. I don't in any \nway require a linear history, and criss-cross merges are supported \nperfectly well by git, and even encouraged in those situations.\n\nAfter all, if tree B starts using features that are new to tree A, then \nthe merge from A->B is required for functionality, and the synchronization \nis a fundamental part of the history of development. In that cases, the \nhistory complexity of the resulting tree is a result of real development \ncomplexity.\n\nNow, obviously, for various reasons we want to avoid having those kinds of \nlinkages as much as possible. We like to have develpment of different \nsubsystems as independent as possible, not because it makes for a \"more \nreadable history\", but because it makes it a lot easier to debug - if we \nhave three independent features/development trees, they can be debugged \nindependently too, while any linkages inevitably also mean that any bugs \nend up being interlinked..\n\n\t\tLinus\n"},{"id":"14389","messageId":"7v4q4cbx6l.fsf@assigned-by-dhcp.cox.net","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601091845160.5588@g5.osdl.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-01-10T03:04:34Z","receivedAt":"2006-01-10T03:04:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Now, obviously, for various reasons we want to avoid having those kinds of \n> linkages as much as possible. We like to have develpment of different \n> subsystems as independent as possible, not because it makes for a \"more \n> readable history\", but because it makes it a lot easier to debug - if we \n> have three independent features/development trees, they can be debugged \n> independently too, while any linkages inevitably also mean that any bugs \n> end up being interlinked..\n>\n> \t\tLinus\n\nYes.  If subproject B uses new features from A (either upstream\nor sibling subproject), pulling A into B is inevitable.\n\nOn the other hand, if such merges becomes too frequent, it may\nbe a sign that A's feature set and interface is still changing\ntoo rapidly for downstream use, but developers A and B are not\ncommunicating well and B has not noticed that B might be better\noff taking a break, addressing other non-overlapping areas while\ngiving a bit of time for A to settle things down.\n\nAn SCM is just _one_ of the ways for developers to communicate,\nit will never be a replacement for developer communication.\n"},{"id":"14396","messageId":"99D82C29-4F19-4DD3-A961-698C3FC0631D@mac.com","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601091845160.5588-hNm40g4Ew95AfugRpC6u6w@public.gmane.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Kyle Moffett","fromEmail":"mrmacman_g4-ee4meeah724@public.gmane.org","sentAt":"2006-01-10T06:33:27Z","receivedAt":"2006-01-10T06:33:27Z","isPatch":false,"sender":{"key":"mrmacman_g4-ee4meeah724@public.gmane.org","avatar":null},"body":"On Jan 09, 2006, at 21:50, Linus Torvalds wrote:\n> if we  have three independent features/development trees, they can  \n> be debugged independently too, while any linkages inevitably also  \n> mean that any bugs end up being interlinked..\n\nOne example:\n\nIf I have ACPI, netdev, and swsusp trees change between an older  \nversion and a newer one, and my net driver starts breaking during  \nsuspend, I would be happiest debugging with the following set of  \npatches/trees (Heavily simplified):\n\n            ^\n            |\n           [5]\n            |\n          broken\n         ^  ^   ^\n       [2] [3]  [4]\n       /    |     \\\nnetdev3  acpi3   swsusp3\n    ^       ^        ^\n    |       |        |\nnetdev2  acpi2   swsusp2\n    ^       ^        ^\n    |       |        |\nnetdev1  acpi1   swsusp1\n       ^    ^    ^\n        \\   |   /\n         \\  |  /\n          \\ | /\n           \\|/\n            |\n           [1]\n            |\n          works\n\n\nIf the old version [1] works and the new one [5] doesn't, then I can  \nimmediately test [2], [3], and [4].  If one of those doesn't work,  \nI've identified the problematic patchset and cut the debugging by  \n2/3.  If they all work, then we know precisely that it's the  \ninteractions between them, which also makes debugging a lot easier.\n\nCheers,\nKyle Moffett\n\n--\nThere are two ways of constructing a software design. One way is to  \nmake it so simple that there are obviously no deficiencies. And the  \nother way is to make it so complicated that there are no obvious  \ndeficiencies.  The first method is far more difficult.\n   -- C.A.R. Hoare\n\n\n-\nTo unsubscribe from this list: send the line \"unsubscribe linux-acpi\" in\nthe body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org\nMore majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"14397","messageId":"46a038f90601092238r3476556apf948bfe5247da484@mail.gmail.com","threadId":"3023","inReplyTo":"99D82C29-4F19-4DD3-A961-698C3FC0631D-ee4meeAH724@public.gmane.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Martin Langhoff","fromEmail":"martin.langhoff-re5jqeeqqe8avxtiumwx3w@public.gmane.org","sentAt":"2006-01-10T06:38:07Z","receivedAt":"2006-01-10T06:38:07Z","isPatch":false,"sender":{"key":"martin.langhoff-re5jqeeqqe8avxtiumwx3w@public.gmane.org","avatar":null},"body":"On 1/10/06, Kyle Moffett <mrmacman_g4-ee4meeAH724@public.gmane.org> wrote:\n> If they all work, then we know precisely that it's the\n> interactions between them, which also makes debugging a lot easier.\n\nThe more complex your tree structure is, the more the interactions are\nlikely to be part of the problem. Is git-bisect not useful in this\nscenario?\n\ncheers,\n\n\nmartin\n-\nTo unsubscribe from this list: send the line \"unsubscribe linux-acpi\" in\nthe body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org\nMore majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"14418","messageId":"252A408D-0B42-49F3-92BC-B80F94F19F40@mac.com","threadId":"3023","inReplyTo":"46a038f90601092238r3476556apf948bfe5247da484@mail.gmail.com","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Kyle Moffett","fromEmail":"mrmacman_g4@mac.com","sentAt":"2006-01-10T18:05:56Z","receivedAt":"2006-01-10T18:05:56Z","isPatch":false,"sender":{"key":"mrmacman_g4@mac.com","avatar":null},"body":"On Jan 10, 2006, at 01:38, Martin Langhoff wrote:\n> On 1/10/06, Kyle Moffett <mrmacman_g4@mac.com> wrote:\n>> If they all work, then we know precisely that it's the  \n>> interactions between them, which also makes debugging a lot easier.\n>\n> The more complex your tree structure is, the more the interactions  \n> are likely to be part of the problem. Is git-bisect not useful in  \n> this scenario?\n\nIIRC git-bisect just does an outright linearization of the whole tree  \nanyways, which makes git-bisect work everywhere, even in the presence  \nof difficult cross-merges.  On the other hand, if you are git- \nbisecting ACPI changes (perhaps due to some ACPI breakage), and ACPI  \nhas 10 pulls from mainline, you _also_ have to wade through the  \nbisection of any other changes that occurred in mainline, even if  \nthey're totally irrelevant.  This is why it's useful to only pull  \nmainline into your tree (EX: ACPI) when you functionally depend on  \nchanges there (as Linus so eloquently expounded upon).\n\nCheers,\nKyle Moffett\n\n--\nQ: Why do programmers confuse Halloween and Christmas?\nA: Because OCT 31 == DEC 25.\n"},{"id":"14419","messageId":"Pine.LNX.4.64.0601101015260.4939@g5.osdl.org","threadId":"3023","inReplyTo":"252A408D-0B42-49F3-92BC-B80F94F19F40@mac.com","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-01-10T18:27:02Z","receivedAt":"2006-01-10T18:27:02Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nOn Tue, 10 Jan 2006, Kyle Moffett wrote:\n>\n> On Jan 10, 2006, at 01:38, Martin Langhoff wrote:\n> > \n> > The more complex your tree structure is, the more the interactions are\n> > likely to be part of the problem. Is git-bisect not useful in this scenario?\n> \n> IIRC git-bisect just does an outright linearization of the whole tree anyways,\n> which makes git-bisect work everywhere, even in the presence of difficult\n> cross-merges.\n\nIt's not really a linearization - at no time does git-bisect _order_ the \ncommits. After all, no linear order actually exists. \n\nInstead, it really cuts the tree up into successively smaller parts. \n\nThink of it as doing a binary search in a 2-dimensional surface - you \ncan't linearize the plane, but you can decide to test first one half of \nthe surface, and then depending on whether it was there, you can halve \nthat surface etc.. \n\n> On the other hand, if you are git-bisecting ACPI changes\n> (perhaps due to some ACPI breakage), and ACPI has 10 pulls from mainline, you\n> _also_ have to wade through the bisection of any other changes that occurred\n> in mainline, even if they're totally irrelevant.\n\nYes. Although if you _know_ that the problem happened in a specific file \nor specific subdirectory, you can actually tell \"git bisect\" to only \nbother with changes to that file/directory/set-of-directories to speed up \nthe search.\n\nIOW, if you absolutely know that it's ACPI-related, you can do something \nlike\n\n\tgit bisect start drivers/acpi arch/i386/kernel/acpi\n\nto tell the bisect code that it should totally ignore anything that \ndoesn't touch those two directories.\n\nHowever, if it turns out that you were wrong (and the ACPI breakage was \nbrought on by something that changed something else), \"git bisect\" will \njust get confused and report the wrong commit, so this is really something \nyou should be careful with (and verify the end result by checking that \nundoing that _particular_ commit really fixes things).\n\nAnd yes, \"git bisect\" _will_ work with bugs that depend on two branches of \na merge: it will point to the merge commit itself as being the problem. \nNow, at that point you really are screwed, and you'll have to figure out \nwhy both branches work, but the combination of them do not.\n\nMaybe it's as simple as just a merge done wrong (bad manual fixups), but \nmaybe it's a perfectly executed merge that just happens to have one branch \nchanging the assumptions that the other branch depended on.\n\nHappily, that is not very common. I know people are using \"git bisect\", \nand I don't think anybody has ever reported it so far. It will happen \neventually, but I'd actually expect it to be much more common that \"git \nbisect\" will hit other - worse - problems, like bugs that \"come and go\", \nand that a simple bisection simply cannot find because they aren't \ntotally repeatable.\n\n\t\t\tLinus\n"},{"id":"14420","messageId":"Pine.LNX.4.63.0601101938420.26999@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601101015260.4939-hNm40g4Ew95AfugRpC6u6w@public.gmane.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin-mmb7mzphnfy@public.gmane.org","sentAt":"2006-01-10T18:45:11Z","receivedAt":"2006-01-10T18:45:11Z","isPatch":false,"sender":{"key":"johannes.schindelin-mmb7mzphnfy@public.gmane.org","avatar":null},"body":"Hi,\n\nOn Tue, 10 Jan 2006, Linus Torvalds wrote:\n\n> \n> On Tue, 10 Jan 2006, Kyle Moffett wrote:\n> >\n> > On Jan 10, 2006, at 01:38, Martin Langhoff wrote:\n> > > \n> > > The more complex your tree structure is, the more the interactions are\n> > > likely to be part of the problem. Is git-bisect not useful in this scenario?\n> > \n> > IIRC git-bisect just does an outright linearization of the whole tree anyways,\n> > which makes git-bisect work everywhere, even in the presence of difficult\n> > cross-merges.\n> \n> It's not really a linearization - at no time does git-bisect _order_ the \n> commits. After all, no linear order actually exists. \n> \n> Instead, it really cuts the tree up into successively smaller parts. \n> \n> Think of it as doing a binary search in a 2-dimensional surface - you \n> can't linearize the plane, but you can decide to test first one half of \n> the surface, and then depending on whether it was there, you can halve \n> that surface etc.. \n\nHow?\n\nIf you bisect, you test a commit. If the commit is bad, you assume *all* \ncommits before that as bad. If it is good, you assume *all* commits after \nthat as good.\n\nNow, if you have a 2-dimensional surface, you don't have a *point*, but \ntypically a *line* separating good from bad.\n\nFurther, the comparison with 2 dimensions is particularly bad. You \n*have* partially linear development lines, it got *nothing* to do with \nan area. The commits still make up a *list*, and it depends how you \n*order* that list for bisect. (And don't tell me they are not ordered: \nthey are.)\n\nIf you order the commits by date, you don't get anything meaningful point \nbefore which it is bad, and after which it is good.\n\nSo, how is bisect supposed to work if you don't have one straight \ndevelopment line from bad to good?\n\nCiao,\nDscho\n\n\n-\nTo unsubscribe from this list: send the line \"unsubscribe linux-acpi\" in\nthe body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org\nMore majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"14422","messageId":"Pine.LNX.4.64.0601101048440.4939@g5.osdl.org","threadId":"3023","inReplyTo":"Pine.LNX.4.63.0601101938420.26999@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-01-10T19:01:18Z","receivedAt":"2006-01-10T19:01:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Jan 2006, Johannes Schindelin wrote:\n> > \n> > Think of it as doing a binary search in a 2-dimensional surface - you \n> > can't linearize the plane, but you can decide to test first one half of \n> > the surface, and then depending on whether it was there, you can halve \n> > that surface etc.. \n> \n> How?\n> \n> If you bisect, you test a commit. If the commit is bad, you assume *all* \n> commits before that as bad. If it is good, you assume *all* commits after \n> that as good.\n\nNo, that's not how bisect works at all.\n\nIt's true that if a commit is bad, then all the commits _reachable_ from \nthat commit are considered bad. \n\nAnd it's true that if a commit is good, then all commits that _reach_ that \ncommit are considered good.\n\nBut that doesn't mean that there is an ordering. The commits that fall \ninto the camp of being \"neither good nor bad\" are _not_ ordered. There are \ncommits in there that are not directly reachable from the good commit.\n\n> Now, if you have a 2-dimensional surface, you don't have a *point*, but \n> typically a *line* separating good from bad.\n\nExactly. \n\nAnd a git graph is not really a two-dimensional surface, but exactly was \nwith a 2-dimensional surface, it is _not_ enough to have a *point* to \nseparate the good from bad.\n\nYou need to have a _set of points_ to separate the good from the bad. You \ncan think of it as a line that bisects the surface: if you were to print \nout the development graph, the set of points literally _do_ form a virtual \nline across the development surface.\n\n(Actually, you can't in general print out the development graph on a \n2-dimensional paper without having development lines that cross each \nother, but you could actually do it in three dimensions, where the \n\"boundary\" between good and bad is actually a 2-dimensional surface in \n3-dimensional space).\n\nBut to describe the surface of \"known good\", you actually just need a list \nof known good commits, and the \"commits reachable from those commits\" \n_becomes_ the surface.\n\n> Further, the comparison with 2 dimensions is particularly bad.\n\nNo it is not. It's a very good comparison.\n\nIn a linearized model (one-dimensional, fully ordered set), the only thing \nyou need for bisection is two points: the beginning and the end.\n\nIn the git model, you need _many_ points to describe the area being \nbisected. Exactly the same way as if you were to bisect a 2-dimensional \nsurface.\n\nNow, the git history is _not_ really a two-dimensional surface, so it's \njust an analogy, not an exact identity. But from a visualization \nstandpoint, it's a good way to think of each \"git bisect\" as adding a \n_line_ on the surface rather than a point on a linear line.\n\n> So, how is bisect supposed to work if you don't have one straight \n> development line from bad to good?\n\nRead the code.\n\nI'm pretty proud of it. It's simple, and it's obvious once you think about \nit, but it is pretty novel as far as I know. BK certainly had nothing \nsimilar, not have I heard of anythign else that does it. Git _might_ be \nthe first thing that has ever done it, although it's simple enough that I \nwouldn't be surprised if others have too.\n\n\t\t\tLinus\n"},{"id":"14425","messageId":"Pine.LNX.4.64.0601101111110.4939@g5.osdl.org","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601101048440.4939@g5.osdl.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-01-10T19:28:58Z","receivedAt":"2006-01-10T19:28:58Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Jan 2006, Linus Torvalds wrote:\n> \n> Now, the git history is _not_ really a two-dimensional surface, so it's \n> just an analogy, not an exact identity. But from a visualization \n> standpoint, it's a good way to think of each \"git bisect\" as adding a \n> _line_ on the surface rather than a point on a linear line.\n\nActually, the way I think of it is akin to the \"light cones\" in physics. A \npoint in space-time doesn't define a fully ordered \"before and after\": but \nit _does_ describe a \"light cone\" which tells you what is reachable from \nthat point, and what that point reaches. Within those cones, that \nparticular point (\"commit\") has a strict ordering.\n\nAnd exactly as in physics, in git there's a lot of space that is _not_ \nordered by that commit. And the way to bisect is basically to find the \nright points in \"git space\" to create the right \"light cone\" that you \nfind the point where the git space that is reachable from that commit has \nthe same volume as the git space that isn't reachable.\n\nAnd maybe that makes more sense to you (if you're into physics), or maybe \nit makes less sense to you.\n\nNow, since we always search the \"git space\" in the cone that is defined by \n\"reachable from the bad commit, but not reachable from any good commit\", \nthe way we handle \"bad\" and \"good\" is actually not a mirror-image. If we \nfine a new _bad_ commit, we know that it was reachable from the old bad \ncommit, and thus the old bad commit is now uninteresting: the new bad \ncommit forms a \"past light cone\" that is a strict subset of the old one, \nso we can totally discard the old bad commit from any future \nconsideration. It doesn't tell us anything new.\n\nIn contrast, if we find a new _good_ commit, the \"past light cone\" (aka \n\"set of commits reachable from it\") is -not- necessarily a proper superset \nof the previous set of good commits, so when we find a good commit, we \nstill need to carry the _other_ good commits around, and the \"known good\" \nuniverse is the _union_ of all the \"good commit past lightcones\".\n\nThen the \"unknown space\" is the set difference of the \"past lightcone of \nthe bad commit\" and of this \"union of past lightcones of good commits\". \nIt's the space that is reachable from the known-bad commit, but not \nreachable from any known-good commit.\n\nSo this means that when doing bisection, what we want to do is find the \npoint in git space that has _new_ \"reachability\" within that unknown space \nthat is as close to half that volume as space as possible. And that's \nexactly what \"git-rev-list --bisect\" calculates.\n\nSo every time, we try to either move the \"known bad\" light-cone down in \ntime in the unknown space, _or_ we add a new \"known good\" light-cone. In \neither case, the \"unknown git space\" keeps shrinking by half each time.\n\n(\"by half\" is not exact, because git space is not only quanticized, it \nalso has a rather strange \"distance function\". In other words, we're \ntalking about a rather strange space. The good news is that the space is \nsmall enough that we can just enumerate every quantum and simply \ncalculate the volume it defines in that space. IOW, we do a very \nbrute-force thing, and it works fine).\n\n\t\t\tLinus\n"},{"id":"14426","messageId":"Pine.LNX.4.63.0601102010100.27199@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601101048440.4939@g5.osdl.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-01-10T19:38:40Z","receivedAt":"2006-01-10T19:38:40Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[cut down the Cc: list, since this is getting special]\n\nOn Tue, 10 Jan 2006, Linus Torvalds wrote:\n\n> > If you bisect, you test a commit. If the commit is bad, you assume *all* \n> > commits before that as bad. If it is good, you assume *all* commits after \n> > that as good.\n> \n> No, that's not how bisect works at all.\n\nOkay, so I got that wrong. But for a good reason: this is not the meaning \nof bisection in my lectures. Doesn't matter.\n\n> It's true that if a commit is bad, then all the commits _reachable_ from \n> that commit are considered bad. \n> \n> And it's true that if a commit is good, then all commits that _reach_ that \n> commit are considered good.\n> \n> But that doesn't mean that there is an ordering. The commits that fall \n> into the camp of being \"neither good nor bad\" are _not_ ordered. There are \n> commits in there that are not directly reachable from the good commit.\n\nThose commits not reachable from the good commit are of no interest. Let's \njust ignore them.\n\n> > Now, if you have a 2-dimensional surface, you don't have a *point*, but \n> > typically a *line* separating good from bad.\n> \n> Exactly. \n> \n> And a git graph is not really a two-dimensional surface, but exactly was \n> with a 2-dimensional surface, it is _not_ enough to have a *point* to \n> separate the good from bad.\n> \n> You need to have a _set of points_ to separate the good from the bad. You \n> can think of it as a line that bisects the surface: if you were to print \n> out the development graph, the set of points literally _do_ form a virtual \n> line across the development surface.\n\nOkay, so there is a cut: Every directed path from good to bad has a single \ncommit which is the first bad. Let's call the set of all such bad commits \nthe cut set.\n\nIs git-bisect capable of identifying all of the cut set, or just a single \none?\n\n> > Further, the comparison with 2 dimensions is particularly bad.\n> \n> No it is not. It's a very good comparison.\n\n>From your explanation I understand now why you like that comparison.\n\n> > So, how is bisect supposed to work if you don't have one straight \n> > development line from bad to good?\n> \n> Read the code.\n> \n> I'm pretty proud of it.\n\nI bet nobody can tell ;-)\n\nWell, I read the code. And I answer my own question from 18--19 lines ago:\n\ngit-bisect is not capable of identifying the cut set, but pretends that \nthere really is only one bad commit (see bisect_bad()).\n\nThat may be the best choice if all commits in the cut set except one are \nmerges. (It is the best if the cut set contains only one element.)\n\nBut I see two problems with that:\n\n- a problem can be introduced independently in two different branches, and\n  occur in both of them before the merge (in which case bisect only \n  catches one of the commits), and\n\n- AFAICT if the cut set is one merge and one regular commit, bisect could\n  identify the merge by error.\n\nOf course, all this makes only a difference if the bisect has to cross a \nmerge.\n\nBTW I think there is a thinko in git-rev-list.txt:\n\n> Thus, if 'git-rev-list --bisect foo ^bar ^baz' outputs 'midpoint', the \n> output of 'git-rev-list foo ^midpoint' and 'git-rev-list midpoint ^bar \n             ^ this should be\n\t'git-rev-list foo ^midpoint ^bar ^baz'\n> ^baz' would be of roughly the same length\n\nCiao,\nDscho\n"},{"id":"14431","messageId":"Pine.LNX.4.64.0601101151090.4939@g5.osdl.org","threadId":"3023","inReplyTo":"Pine.LNX.4.63.0601102010100.27199@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-01-10T20:11:02Z","receivedAt":"2006-01-10T20:11:02Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Jan 2006, Johannes Schindelin wrote:\n> \n> Those commits not reachable from the good commit are of no interest. Let's \n> just ignore them.\n\nNote that to avoid confusion, start talking about -multiple- good commits \nearly.\n\nSo we have a list of \"known good islands\" in the git-space. And yes, we \nwant to ignore anything that is reachable from them.\n\nAnd here the magic part of\"git-bisect.sh\" is around line 133:\n\n\t... --not $(cd \"$GIT_DIR\" && ls refs/bisect/good-*) ...\n\nIt tells git-rev-parse to generate a list of commits that we're _not_ \ninterested in, and that list will be one of the most critical parts of the \nstuff we give to \"git-rev-list --bisect\".\n\nSo that part of the script is literally the part that says \"ignore all of \ngit space that is reachable from the good commits\", because we've listed \nall the good commits as refs named \"refs/bisect/good-*\".\n\n> > You need to have a _set of points_ to separate the good from the bad. You \n> > can think of it as a line that bisects the surface: if you were to print \n> > out the development graph, the set of points literally _do_ form a virtual \n> > line across the development surface.\n> \n> Okay, so there is a cut: Every directed path from good to bad has a single \n> commit which is the first bad. Let's call the set of all such bad commits \n> the cut set.\n\nThis set is uninteresting for two reasons:\n - it's hard to calculate\n - it's not the answer we want.\n\nWe want the _single_ commit that is the one that generates your \"cut set\".\n\nYour \"cut set\" is really the \"reachability border\" from the single bad \ncommit we're interested in to all the possible development lines.\n\nIn practice, the \"cut set\" is just the \"bad commit\" plus all the merges \nthat merge that bad commit with somethign that wasn't reacable from it in \nthe first place.\n\nSo the \"cut set\" isn't interesting.\n\n> git-bisect is not capable of identifying the cut set, but pretends that \n> there really is only one bad commit (see bisect_bad()).\n\nNot quite.\n\nIt could keep track of all bad commits (in fact, it does so in the log \nfile), but the fact is, none but the lastest bad commit we have found \nmatters.\n\nBy definition, \"git bisect\" is always going to test a commit that is \nreachable from the previously known bad commit. Agreed? Anything else \nwould be insane - we know that we had a bad stat, and we're interested in \nfinding out how _that_ bad state happened, so we're only ever interested \nin commits that are ancestors to that bad state.\n\nSo our search-space is _literally_ defined by two things:\n\n - the surface of \"known good\" commits (which defines the commits that \n   aren't interesting). \n\n   This is the \"--not refs/bisect/good-*\" part\n\n - the last \"known bad\" commit.\n\nWe'll always search the git commit space defined by these two knowns, \nagreed?\n\nNow, realize that if we find a new bad commit, since that bad commit was \nby definition reachable from the _old_ bad commit (since we didn't even \nsearch outside its reachability), then equally by definition the \nreachability from that new bad commit is a strict superset of the \nreachability of the old bad commit.\n\nSo when we find a new bad commit, the old bad commit is no longer \ninteresting.\n\nSo when you say \"pretends that there really is only one bad commit\", you \ndidn't realize that it's not about \"pretending\". It's very fundamental: \nthere is only ever _one_ bad commit that is interesting. It's the last one \nwe found.\n\nEven if we started out with two bad commits (ie some person reported two \ndifferent versions as being bad), we're _still_ not interested in using \nthem both. We should pick one of them, because the reachability area \ndefined by two bad commits is always a superset of the reachability of \neither one.\n\nSo having multiple bad commits is _never_ interesting.\n\n> But I see two problems with that:\n> \n> - a problem can be introduced independently in two different branches, and\n>   occur in both of them before the merge (in which case bisect only \n>   catches one of the commits), and\n\nThis is fine. Depending on whatever random factors, we'll test one of them \nfirst, and eventually find _one_ of the commits that fix it. If the exact \nsame bug was introduced somewhere else, and merged, then undoing just the \n\"one\" bug will obviously undo the other one too.\n\nIf a _different_ bug was introduced (even if it had the same effects), \nyes, you now have two separate bugs. And bisecting two bugs is hard. You \nneed to separate them out some way.\n\n> - AFAICT if the cut set is one merge and one regular commit, bisect could\n>   identify the merge by error.\n\nIt will never identify a commit without having done a full bisection, so \nif it ever had the choice of a \"merge\" and the \"commit leading up to the \nmerge\", it will always have tried the \"commit leading up to the merge\", \nand decided that it was fundamentally more recent (had \"smaller \nreachability\") that the merge, and pinpoint it.\n\n> BTW I think there is a thinko in git-rev-list.txt:\n> \n> > Thus, if 'git-rev-list --bisect foo ^bar ^baz' outputs 'midpoint', the \n> > output of 'git-rev-list foo ^midpoint' and 'git-rev-list midpoint ^bar \n>              ^ this should be\n> \t'git-rev-list foo ^midpoint ^bar ^baz'\n> > ^baz' would be of roughly the same length\n\nYes.\n\n\t\t\tLinus\n"},{"id":"14435","messageId":"Pine.LNX.4.64.0601101221020.4939@g5.osdl.org","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601101151090.4939@g5.osdl.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-01-10T20:28:40Z","receivedAt":"2006-01-10T20:28:40Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 10 Jan 2006, Linus Torvalds wrote:\n> \n> If a _different_ bug was introduced (even if it had the same effects), \n> yes, you now have two separate bugs. And bisecting two bugs is hard. You \n> need to separate them out some way.\n\nSide note: this is seldom a problem in practice. If it was effectively the \nsame bug, just finding the one case that triggered it is sufficient: you \nthen know what to look for, and if undoing that one commit isn't enough to \nfix it in the current tree (because the same bug existed in another form \non another branch), you wouldn't actually start bisecting again. You'd \nstart grepping the tree for other cases of that bug.\n\nSo the biggest advantage of \"git bisect\" is _not_ that you can just undo \nthe buggy commit. In fact, usually you don't even want to undo it, because \nit probably had a raison-d'etre to begin with. The huge deal about \"git \nbisect\" is that it pinpoints what caused the bug, and then the fix is \noften something else.\n\nOften it's a \"Duh! I fixed one thing, but my fix didn't take Xyz into \naccount, so it now broke for another reason\" moment.\n\nMost bugs are stupid, in other words.\n\nThe _real_ problem with git bisect is when you have a non-technical user \n(common) and there are silly bugs that you know of and already fixed that \naren't really a problem, but that are show-stoppers for the user who isn't \na kernel developer (or is, but doesn't know git). They're show-stoppers \nnot because we care about them, but because they make the \"purely \nmechanical\" thing be one where you have to have some manual input.\n\nAnother problem (that I've not seen in practice yet, but that I bet _will_ \nbe the worst issue) is non-reproducible bugs. They are the nastiest kind \nto debug in the first place, and sadly, \"git bisect\" simply doesn't help \nyou with them. There, nothing but some luck and a lot of thinking and \ntesting will help you.\n\n\t\t\tLinus\n"},{"id":"14441","messageId":"Pine.LNX.4.63.0601102122001.30609@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601101151090.4939@g5.osdl.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-01-10T20:47:04Z","receivedAt":"2006-01-10T20:47:04Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 10 Jan 2006, Linus Torvalds wrote:\n\n> So having multiple bad commits is _never_ interesting.\n\nOkay, I got it. A bug is supposed to be inherited by *all* its \ndescendants. Good.\n\nI have to keep in mind that a commit is not actually a patch set, but can \nbe two or more (in case of a merge). So, a bug can be present in a \ndevelopment line for a long, long time, but be visible only after a merge. \nSince that commit can be compared to at least two trees, one of these \ndiffs must show the bug.\n\nThanks,\nDscho\n"},{"id":"14464","messageId":"20060111033229.5590.qmail@web31809.mail.mud.yahoo.com","threadId":"3023","inReplyTo":"252A408D-0B42-49F3-92BC-B80F94F19F40-ee4meeAH724@public.gmane.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Luben Tuikov","fromEmail":"ltuikov-/e1597as9lqavxtiumwx3w@public.gmane.org","sentAt":"2006-01-11T03:32:29Z","receivedAt":"2006-01-11T03:32:29Z","isPatch":false,"sender":{"key":"ltuikov-/e1597as9lqavxtiumwx3w@public.gmane.org","avatar":null},"body":"--- Kyle Moffett <mrmacman_g4-ee4meeAH724@public.gmane.org> wrote:\n> they're totally irrelevant.  This is why it's useful to only pull  \n> mainline into your tree (EX: ACPI) when you functionally depend on  \n> changes there (as Linus so eloquently expounded upon).\n\nSometimes the dependency is _behavioural_.  For example certain\nbehaviour of other modules of the kernel changed and you want\nto test that your module works ok with them under different\nbehaviour.  In which case you may or may not have to\nchange your code after the fact.\n\n    Luben\n\n-\nTo unsubscribe from this list: send the line \"unsubscribe linux-acpi\" in\nthe body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org\nMore majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"14642","messageId":"pan.2006.01.13.23.34.58.269921@smurf.noris.de","threadId":"3023","inReplyTo":"Pine.LNX.4.64.0601101048440.4939@g5.osdl.org","subject":"Re: git pull on Linux/ACPI release tree","fromName":"Matthias Urlichs","fromEmail":"smurf@smurf.noris.de","sentAt":"2006-01-13T23:35:01Z","receivedAt":"2006-01-13T23:35:01Z","isPatch":false,"sender":{"key":"matthias@urlichs.de","avatar":"https://gravatar.com/avatar/2708905af227313eba6f2b2ae0f7d0259b5ac5d71baef58fe5a13c699ce0bbf0?d=mp&s=160"},"body":"Hi, Linus Torvalds wrote:\n\n> I'm pretty proud of it. It's simple, and it's obvious once you think about \n> it, but it is pretty novel as far as I know. BK certainly had nothing \n> similar, not have I heard of anythign else that does it.\n\nActually, I've written a hackish script that tries to do simple-minded\nbisection (read: it searched for the 50% point on the shortest path\nbetween any-of-GOOD and any-of-BAD, instead of considering the whole\ngraph) on BK trees. I haven't exactly published the thing anyplace though,\nbecause, well, it was ugly. :-/\n\nBesides, actually working with the current bisection point is no problem\nat all for git. Doing the same thing in BK's world view is *painful*,\nesp. given the size of the kernel tree.\n\n-- \nMatthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de\nDisclaimer: The quote was selected randomly. Really. | http://smurf.noris.de\n - -\nArthur felt at a bit of a loss. There was a whole Galaxy\nof stuff out there for him, and he wondered if it was\nchurlish of him to complain to himself that it lacked just\ntwo things: the world he was born on and the woman he loved.\n"}]}