{"thread":{"id":"39","subject":"RE: Merge with git-pasky II.","startedAt":"2005-04-15T20:31:00Z","lastAt":"2005-04-16T00:32:13Z","messageCount":3,"participants":["Linus Torvalds","Barry Silverman"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"245","messageId":"Pine.LNX.4.58.0504151313450.7211@ppc970.osdl.org","threadId":"39","inReplyTo":"000d01c541ed$32241fd0$6400a8c0@gandalf","subject":"RE: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-15T20:31:00Z","receivedAt":"2005-04-15T20:31:00Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n[ I'm cc'ing the git list even though Barry's question wasn't cc'd. \n  Because I think his question is interesting and astute per se, even\n  if I disagree with the proposal ]\n\nOn Fri, 15 Apr 2005, Barry Silverman wrote:\n>\n> If git is totally project based, and each commit represents total state\n> of the project, then how important is the intermediate commit\n> information between two states. \n\nYou need it in order to do further merges.\n\n> IE, Area maintainer has A1->A2->A3->A4->A5 in a repository with 5\n> commits, and 5 comments. And I last synced with A1.\n> \n> A few days later I sync again. Couldn't I just pull the \"diff-tree A5\n> A1\" and then commit to my tree just the record A1->A5. Why does MY\n> repository need trees A2,A3,A4?\n\nBecause that second merge needs the first merge to work well. The first \nmerge might have had some small differences that ended up auto-merging (or \neven needing some manual help from you). The second time you sync, there \nmigth be some more automatic merging. And so on.\n\nIf you don't keep track of the incremental merges, you end up with one \nreally _difficult_ merge that may not be mergable at all. Not \nautomatically, and perhaps not even with help.\n\nSo in order to keep things mergable, you need to not diverge. And the \nintermediate merges are the \"anchor-points\" for the next merge, keeping \nthe divergences minimal. \n\nI'm personally convinced that one of the reasons CVS is a pain to merge is \nobviously that it doesn't do a good job of finding parents, but also \nexactly _because_ it makes merges so painful that people wait longer to do \nthem, so you never end up fixing the simple stuff. In contrast, if you\nhave all these small merges going on all the time, the hope is that \nthere's never any really painful nasty \"final merge\".\n\nSo you're right - the small merges do end up cluttering up the revision \nhistory. But it's a small price to pay if it means that you avoid having \nthe painful ones.\n\n> Isn't preserving the A1,A2,A3,A4,A5 a legacy of BK, which required all\n> the changesets to be loaded in order, and so is a completely \"file\"\n> driven concept? \n\nNope. In fact, to some degree git will need this even _more_, since the\ngit merger is likely to be _weaker_ than BK, and thus more easily\nconfused.\n\nI do believe that BK has these things for the same reason.\n\n\t\t\tLinus\n"},{"id":"254","messageId":"000701c5420e$e89177b0$6400a8c0@gandalf","threadId":"39","inReplyTo":"Pine.LNX.4.58.0504151313450.7211@ppc970.osdl.org","subject":"RE: Merge with git-pasky II.","fromName":"Barry Silverman","fromEmail":"barry@disus.com","sentAt":"2005-04-15T23:00:17Z","receivedAt":"2005-04-15T23:00:17Z","isPatch":false,"sender":{"key":"barry@disus.com","avatar":null},"body":"The issue I am trying to come to grips with in the current design, is\nthat the git repository of a number of interrelated projects will soon\nbecome the logical OR of all blobs, commits, and trees in ALL the\nprojects. \n\nThis will involve horrendous amounts of replication, as developers end\ninterchanging objects that originated from third parties who are not\nparty to the merge (and which happen to be in the repository because of\nprevious merge activity).\n\nIt would be really nice if the repositories for each project stay\ndistinct (and maybe even living on different servers). Merges should the\none of the few points of contact where the state be exchanged between\nthe repositories.\n\n>> If you don't keep track of the incremental merges, you end up with\none \n>> really _difficult_ merge that may not be mergable at all. Not \n>> automatically, and perhaps not even with help.\n\nNo argument from me.... In my example, I didn't intend A2,A3,A4 to be\nconsidered hash's of points where you did a merge - rather they are\nhashes of individual points where you did a commit BETWEEN the points\nwhere you merged. A1 and A5 are the \"merge points\"! \n\nIt makes perfect sense to define some subset of commits as \"merge-like\"\ncommits, and only have those copied over from one repository to the\nother. You could also only use merge-points in the common ancestor\ncalculation, and not worry about intermediate commits. \n\nOnly small changes to the existing logic are necessary to do a merge by\n\"distributing\" out the merge algorithm to each repository. This involves\nquerying each repository, and communicating the results, followed by\ncopying over only those blob objects necessary for the merge. \n\nAfter the merge, you would create a \"merge-point commit\" record that has\none of the parents pointing to a hash in the other repository!\n\nBut the BIG issue with this scheme, is that you will not be replicating\nover any of the intermediate commits, trees, or blobs (not really needed\nby the merge), but currently being traversed by various plumbing\ncomponents.\n\nHence my question....\n\n-----Original Message-----\nFrom: Linus Torvalds [mailto:torvalds@osdl.org] \nSent: Friday, April 15, 2005 4:31 PM\nTo: Barry Silverman\nCc: git@vger.kernel.org\nSubject: RE: Merge with git-pasky II.\n\n\n[ I'm cc'ing the git list even though Barry's question wasn't cc'd. \n  Because I think his question is interesting and astute per se, even\n  if I disagree with the proposal ]\n\nOn Fri, 15 Apr 2005, Barry Silverman wrote:\n>\n> If git is totally project based, and each commit represents total\nstate\n> of the project, then how important is the intermediate commit\n> information between two states. \n\nYou need it in order to do further merges.\n\n> IE, Area maintainer has A1->A2->A3->A4->A5 in a repository with 5\n> commits, and 5 comments. And I last synced with A1.\n> \n> A few days later I sync again. Couldn't I just pull the \"diff-tree A5\n> A1\" and then commit to my tree just the record A1->A5. Why does MY\n> repository need trees A2,A3,A4?\n\nBecause that second merge needs the first merge to work well. The first \nmerge might have had some small differences that ended up auto-merging\n(or \neven needing some manual help from you). The second time you sync, there\n\nmigth be some more automatic merging. And so on.\n\nIf you don't keep track of the incremental merges, you end up with one \nreally _difficult_ merge that may not be mergable at all. Not \nautomatically, and perhaps not even with help.\n\nSo in order to keep things mergable, you need to not diverge. And the \nintermediate merges are the \"anchor-points\" for the next merge, keeping \nthe divergences minimal. \n\nI'm personally convinced that one of the reasons CVS is a pain to merge\nis \nobviously that it doesn't do a good job of finding parents, but also \nexactly _because_ it makes merges so painful that people wait longer to\ndo \nthem, so you never end up fixing the simple stuff. In contrast, if you\nhave all these small merges going on all the time, the hope is that \nthere's never any really painful nasty \"final merge\".\n\nSo you're right - the small merges do end up cluttering up the revision \nhistory. But it's a small price to pay if it means that you avoid having\n\nthe painful ones.\n\n> Isn't preserving the A1,A2,A3,A4,A5 a legacy of BK, which required all\n> the changesets to be loaded in order, and so is a completely \"file\"\n> driven concept? \n\nNope. In fact, to some degree git will need this even _more_, since the\ngit merger is likely to be _weaker_ than BK, and thus more easily\nconfused.\n\nI do believe that BK has these things for the same reason.\n\n\t\t\tLinus\n\n\n"},{"id":"263","messageId":"Pine.LNX.4.58.0504151723530.7211@ppc970.osdl.org","threadId":"39","inReplyTo":"000701c5420e$e89177b0$6400a8c0@gandalf","subject":"RE: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T00:32:13Z","receivedAt":"2005-04-16T00:32:13Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, Barry Silverman wrote:\n>\n> The issue I am trying to come to grips with in the current design, is\n> that the git repository of a number of interrelated projects will soon\n> become the logical OR of all blobs, commits, and trees in ALL the\n> projects. \n\nNope. I'm actually against the notion of sharing object directories\nbetween projects. The git model _allows_ it, and I think it can be a valid \napproach, but it's absolutely not an approach that I would personally \nsuggest be used for the kernel.\n\nIn other words, the default git behaviour - and the one I personally am a \nproponent of - is to have one object directory per work tree, and really \nconsider all such trees independent. Then you can merge between trees if \nyou want to, and bring in objects that way, but normally you would _not_ \nhave tons of objects from other trees, and _especially_ not from other \nunrelated projects.\n\nThe reason git supports shared object archives is that (a) it falls out \ntrivially as part of the design, so not allowing it is silly and (b) it is \npart of a merge, where you _do_ want to get the objects of the trees you \nmerge, and in particular you need to generate a seperate tree that has all \nthose objects without having to copy them.\n\n(Before you do the merge, you need to bring the new objects into your \nrepository of course, but that I consider to be a separate issue, not \npart of the actual technical merge process).\n\nSo normally, you'd probably have a totally pruned tree, with only the \nobjects you need (and you might even consider the \"commit parent links\" \nless than necessary, especially if you're just a regular user and not a \ndeveloper who wants to merge).\n\nBut the ability to have extra objects is wonderful. It makes going\nbackwards in time basically free (while the equivalent \"bk undo\" in the BK\nworld is a very expensive operation), and it makes it easy to\nincrementally keep up-to-date with trees that you know you're _eventually_\ngoing to merge with. But it's not an excuse to put just any random crap in \nthat object directory..\n\n\t\tLinus\n"}]}