{"thread":{"id":"2386","subject":"Comments on recursive merge..","startedAt":"2005-11-07T16:48:06Z","lastAt":"2005-11-15T15:29:37Z","messageCount":58,"participants":["Linus Torvalds","Fredrik Kuivinen","Junio C Hamano","Johannes Schindelin","Petr Baudis","Ryan Anderson","Martin Langhoff","Catalin Marinas","Chuck Lever"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"11260","messageId":"Pine.LNX.4.64.0511070837530.3193@g5.osdl.org","threadId":"2386","inReplyTo":null,"subject":"Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-07T16:48:06Z","receivedAt":"2005-11-07T16:48:06Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nGuys,\n\n  I just hit my first real rename conflict, and very timidly tried the \n\"recursive\" strategy in the hopes that I wouldn't need to do things by \nhand.\n\nIt resolved things beautifully. Good job. \n\nMy only worry is that I don't read python, so I don't really know how it \ndoes what it does, which makes me nervous. Can somebody (Fredrik?) add \nsome documentation about the merge strategy and how it works.\n\nConsidering that the stupid resolve strategy really requires you to know \nhow git works when rename conflicts happen (things left in unmerged state \nare really quite hard to handle by hand unless you know exactly what \nyou're doing), I'd almost suggest making \"recursive\" the default. I'm a \nbit nervous about it, but knowing how it works would probably put most of \nthat to rest.\n\n\t\tLinus\n"},{"id":"11261","messageId":"Pine.LNX.4.64.0511070848440.3193@g5.osdl.org","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511070837530.3193@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-07T16:56:07Z","receivedAt":"2005-11-07T16:56:07Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 7 Nov 2005, Linus Torvalds wrote:\n> \n>   I just hit my first real rename conflict, and very timidly tried the \n> \"recursive\" strategy in the hopes that I wouldn't need to do things by \n> hand.\n> \n> It resolved things beautifully. Good job. \n\nBtw, one thing that it does is print out too much information.\n\nIn particular, I had renames on both sides of the merge (in case anybody \nwants to see which one I'm talking about: it's the current top-of-head \ncommit in the kernel archives: 333c47c847c90aaefde8b593054d9344106333b5).\n\nNow, renames that you've done yourself you really don't want to hear \nabout, at least if the other side didn't change anything in that file.\n\nRenames that the _other_ side has done (the one you're merging) you may or \nmay not want to know about, regardless of whether they happened to files \nthat are changed. But since \"git pull\" will do a \"git-apply --stat\" at the \nend and show the renames there, I'd argue that the merge strategy itself \nshould be quiet about any renames that are trivial.\n\nSo how about talking about renames only if you end up also doing a \nfile-level merge? As it is, doing the merge talked about renames that I \nhad merged earlier in my own branch, which is just confusing.\n\n\t\tLinus\n"},{"id":"11279","messageId":"20051107225807.GA10937@c165.ib.student.liu.se","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511070837530.3193@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Fredrik Kuivinen","fromEmail":"freku045@student.liu.se","sentAt":"2005-11-07T22:58:07Z","receivedAt":"2005-11-07T22:58:07Z","isPatch":false,"sender":{"key":"frekui@gmail.com","avatar":"https://avatars.githubusercontent.com/u/13770967?v=4"},"body":"On Mon, Nov 07, 2005 at 08:48:06AM -0800, Linus Torvalds wrote:\n> \n> Guys,\n> \n>   I just hit my first real rename conflict, and very timidly tried the \n> \"recursive\" strategy in the hopes that I wouldn't need to do things by \n> hand.\n> \n> It resolved things beautifully. Good job. \n\nI'm glad that it worked.\n\n> My only worry is that I don't read python, so I don't really know how it \n> does what it does, which makes me nervous. Can somebody (Fredrik?) add \n> some documentation about the merge strategy and how it works.\n\nI will write something up.\n\n> Considering that the stupid resolve strategy really requires you to know \n> how git works when rename conflicts happen (things left in unmerged state \n> are really quite hard to handle by hand unless you know exactly what \n> you're doing), I'd almost suggest making \"recursive\" the default. I'm a \n> bit nervous about it, but knowing how it works would probably put most of \n> that to rest.\n\nIt would be great if the recursive strategy could get some more\ntesting. I have tested it on a thousand commits or so in a few kernel\nrepositories and haven't found any bugs, but it could be due to errors\nin the test setup, testing the wrong repositories or just being lucky. Some\nreal-world testing would be great.\n\n- Fredrik\n"},{"id":"11281","messageId":"20051107231944.GA11327@c165.ib.student.liu.se","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511070848440.3193@g5.osdl.org","subject":"[PATCH] merge-recursive: Only print relevant rename messages","fromName":"Fredrik Kuivinen","fromEmail":"freku045@student.liu.se","sentAt":"2005-11-07T23:19:44Z","receivedAt":"2005-11-07T23:19:44Z","isPatch":true,"sender":{"key":"frekui@gmail.com","avatar":"https://avatars.githubusercontent.com/u/13770967?v=4"},"body":"On Mon, Nov 07, 2005 at 08:56:07AM -0800, Linus Torvalds wrote:\n> \n> Btw, one thing that it does is print out too much information.\n> \n> In particular, I had renames on both sides of the merge (in case anybody \n> wants to see which one I'm talking about: it's the current top-of-head \n> commit in the kernel archives: 333c47c847c90aaefde8b593054d9344106333b5).\n> \n> Now, renames that you've done yourself you really don't want to hear \n> about, at least if the other side didn't change anything in that file.\n> \n> Renames that the _other_ side has done (the one you're merging) you may or \n> may not want to know about, regardless of whether they happened to files \n> that are changed. But since \"git pull\" will do a \"git-apply --stat\" at the \n> end and show the renames there, I'd argue that the merge strategy itself \n> should be quiet about any renames that are trivial.\n> \n> So how about talking about renames only if you end up also doing a \n> file-level merge? As it is, doing the merge talked about renames that I \n> had merged earlier in my own branch, which is just confusing.\n> \n\nSounds like a good idea. How about something like the following?\n\n--\n\nIt isn't really interesting to know about the renames that have\nalready been committed to the branch you are working on. Furthermore,\nthe 'git-apply --stat' at the end of git-(merge|pull) will tell us\nabout any renames in the other branch.\n\nWith this commit only renames which require a file-level merge will\nbe printed.\n\nSigned-off-by: Fredrik Kuivinen <freku045@student.liu.se>\n\n\n---\n\n git-merge-recursive.py |   22 +++++++++++++++-------\n 1 files changed, 15 insertions(+), 7 deletions(-)\n\napplies-to: 5af1b5b93257ecfe993bb24975bf596faa342758\n89c029b439603630a53ee4e4d0cb7931111afd2a\ndiff --git a/git-merge-recursive.py b/git-merge-recursive.py\nindex 626d854..9983cd9 100755\n--- a/git-merge-recursive.py\n+++ b/git-merge-recursive.py\n@@ -162,10 +162,13 @@ def mergeTrees(head, merge, common, bran\n # Low level file merging, update and removal\n # ------------------------------------------\n \n+MERGE_NONE = 0\n+MERGE_TRIVIAL = 1\n+MERGE_3WAY = 2\n def mergeFile(oPath, oSha, oMode, aPath, aSha, aMode, bPath, bSha, bMode,\n               branch1Name, branch2Name):\n \n-    merge = False\n+    merge = MERGE_NONE\n     clean = True\n \n     if stat.S_IFMT(aMode) != stat.S_IFMT(bMode):\n@@ -178,7 +181,7 @@ def mergeFile(oPath, oSha, oMode, aPath,\n             sha = bSha\n     else:\n         if aSha != oSha and bSha != oSha:\n-            merge = True\n+            merge = MERGE_TRIVIAL\n \n         if aMode == oMode:\n             mode = bMode\n@@ -207,7 +210,8 @@ def mergeFile(oPath, oSha, oMode, aPath,\n             os.unlink(orig)\n             os.unlink(src1)\n             os.unlink(src2)\n-            \n+\n+            merge = MERGE_3WAY\n             clean = (code == 0)\n         else:\n             assert(stat.S_ISLNK(aMode) and stat.S_ISLNK(bMode))\n@@ -577,14 +581,16 @@ def processRenames(renamesA, renamesB, b\n                 updateFile(False, ren1.dstSha, ren1.dstMode, dstName1)\n                 updateFile(False, ren2.dstSha, ren2.dstMode, dstName2)\n             else:\n-                print 'Renaming', fmtRename(path, ren1.dstName)\n                 [resSha, resMode, clean, merge] = \\\n                          mergeFile(ren1.srcName, ren1.srcSha, ren1.srcMode,\n                                    ren1.dstName, ren1.dstSha, ren1.dstMode,\n                                    ren2.dstName, ren2.dstSha, ren2.dstMode,\n                                    branchName1, branchName2)\n \n-                if merge:\n+                if merge or not clean:\n+                    print 'Renaming', fmtRename(path, ren1.dstName)\n+\n+                if merge == MERGE_3WAY:\n                     print 'Auto-merging', ren1.dstName\n \n                 if not clean:\n@@ -653,14 +659,16 @@ def processRenames(renamesA, renamesB, b\n                 tryMerge = True\n \n             if tryMerge:\n-                print 'Renaming', fmtRename(ren1.srcName, ren1.dstName)\n                 [resSha, resMode, clean, merge] = \\\n                          mergeFile(ren1.srcName, ren1.srcSha, ren1.srcMode,\n                                    ren1.dstName, ren1.dstSha, ren1.dstMode,\n                                    ren1.srcName, srcShaOtherBranch, srcModeOtherBranch,\n                                    branchName1, branchName2)\n \n-                if merge:\n+                if merge or not clean:\n+                    print 'Renaming', fmtRename(ren1.srcName, ren1.dstName)\n+\n+                if merge == MERGE_3WAY:\n                     print 'Auto-merging', ren1.dstName\n \n                 if not clean:\n"},{"id":"11285","messageId":"7v64r4qai6.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"20051107231944.GA11327@c165.ib.student.liu.se","subject":"Re: [PATCH] merge-recursive: Only print relevant rename messages","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-07T23:54:57Z","receivedAt":"2005-11-07T23:54:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Fredrik Kuivinen <freku045@student.liu.se> writes:\n\n> @@ -178,7 +181,7 @@ def mergeFile(oPath, oSha, oMode, aPath,\n>              sha = bSha\n>      else:\n>          if aSha != oSha and bSha != oSha:\n> -            merge = True\n> +            merge = MERGE_TRIVIAL\n\nThe rest looks good to me, but are you sure about this part?  I\nhave a feeling that the above \"and\" should be \"or\", meaning, we\ncheck to see if there is _any_ change, and default to TRIVIAL,\nbut later we would find that we need a real merge and then\npromote it to MERGE_3WAY.\n"},{"id":"11286","messageId":"7vll00ov2l.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"20051107225807.GA10937@c165.ib.student.liu.se","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-08T00:13:38Z","receivedAt":"2005-11-08T00:13:38Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Fredrik Kuivinen <freku045@student.liu.se> writes:\n\n> On Mon, Nov 07, 2005 at 08:48:06AM -0800, Linus Torvalds wrote:\n>> \n>> Guys,\n>> \n>>   I just hit my first real rename conflict, and very timidly tried the \n>> \"recursive\" strategy in the hopes that I wouldn't need to do things by \n>> hand.\n>> \n>> It resolved things beautifully. Good job. \n>\n> I'm glad that it worked.\n\nThis is the first time I see you pleased by something in git\nthat was done without very close supervision from you.  All the\ncredits for this one goes to Fredrik, of course, but it is a\nsmall victory for me as the maintainer as well, and I am very\nhappy about it.\n\n>> ..., I'd almost suggest making \"recursive\" the default. I'm a\n>> bit nervous about it, but knowing how it works would probably\n>> put most of that to rest.\n\nAnother thing to consider is if it is fast enough for everyday\ntrivial merges.\n\nIn any case, I've been thinking about teaching git-merge to look\ninto .git/config to make it overridable which strategy to use by\ndefault.  This would eliminate the hardcoded 'resolve for\ntwo-head, octopus for more' rule from git-pull.  Then we could\nship git-merge with the default rule of 'recursive for two-head,\noctopus for more', and if it turns out to be premature, you can\nupdate your config file to use resolve for two-head case while\nwe sort things out.  That way, recursive would get wider test\ncoverage, and people who really need a working merge this minute\ncan choose to run resolve in emergency without specifying '-s\nresolve' on the command line of 'git pull' every time.\n"},{"id":"11287","messageId":"Pine.LNX.4.64.0511071629270.3247@g5.osdl.org","threadId":"2386","inReplyTo":"7vll00ov2l.fsf@assigned-by-dhcp.cox.net","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-08T00:33:56Z","receivedAt":"2005-11-08T00:33:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 7 Nov 2005, Junio C Hamano wrote:\n> \n> This is the first time I see you pleased by something in git\n> that was done without very close supervision from you.\n\nThat sounds like a backhanded way of saying that I'm micromanagering, \npicky and difficult to work with ;)\n\n> Another thing to consider is if it is fast enough for everyday\n> trivial merges.\n\nHmm. True. The _really_ trivial in-index case triggers for me pretty \noften, but I haven't done any statistics. It might be only 50% of the \ntime.\n\nIs the recursive thing noticeably slower for the \"easy\" cases (ie things \nthat the old regular resolve strategy does well)?\n\nIt's certainly an option to just do what I just did, namely use the \ndefault one until it breaks, and then just do \"git reset --hard\" and re-do \nthe pull with \"-s recursive\". A bit sad, and it would be good to have \ncoverage on the recursive strategy..\n\n\t\tLinus\n"},{"id":"11289","messageId":"7v4q6oosxp.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511071629270.3247@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-08T00:59:46Z","receivedAt":"2005-11-08T00:59:46Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> On Mon, 7 Nov 2005, Junio C Hamano wrote:\n>> \n>> This is the first time I see you pleased by something in git\n>> that was done without very close supervision from you.\n>\n> That sounds like a backhanded way of saying that I'm micromanagering, \n> picky and difficult to work with ;)\n\nSorry, that is not what I meant to say at all.\n\nYou used to micromanage, but it was _very_ good for git back\nthen.  I admit that only once I found you too picky and\ndifficult to work with while I was fixing a bad premature-free\nbug in the diffcore-rename code, but overall your attention to\ndetail well paid off.\n\nSince I inherited the project, we added quite a lot of stuff,\nbut I was still unsure if we are making good progress, or just\nstagnating with only small enhancements and obvious fixes.\n\nBut Fredrik merge turns out to be a spectacular success as you\nfound out, which is a triumph for Fredrik, but at the same time\nit means I was not doing too bad myself ;-).\n"},{"id":"11319","messageId":"Pine.LNX.4.63.0511081254520.2649@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511071629270.3247@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2005-11-08T11:58:50Z","receivedAt":"2005-11-08T11:58:50Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 7 Nov 2005, Linus Torvalds wrote:\n\n> Is the recursive thing noticeably slower for the \"easy\" cases (ie things \n> that the old regular resolve strategy does well)?\n\nIIRC recursive does nothing else than recursively merging the merge-bases \n(granted, in a clever way). So if there is only one merge-base, the only \nslow-down would be the startup of python (which is probably worth it, \nanyway).\n\n> It's certainly an option to just do what I just did, namely use the \n> default one until it breaks, and then just do \"git reset --hard\" and re-do \n> the pull with \"-s recursive\". A bit sad, and it would be good to have \n> coverage on the recursive strategy..\n\nWe already have a fallback list: after really-trivial, try automatic, ...,\ntry resolve. Why not just add recursive? So, if even resolve failed, just \ntry once more, with recursive.\n\nCiao,\nDscho\n"},{"id":"11325","messageId":"7vbr0vjek7.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511071629270.3247@g5.osdl.org","subject":"[RFC/PATCH] Make git-recursive the default strategy for git-pull.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-08T16:21:28Z","receivedAt":"2005-11-08T16:21:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This does two things:\n\n - It changes the hardcoded default merge strategy for two-head\n   git-pull from resolve to recursive.\n\n - .git/config file acquires two configuration items.\n   pull.twohead names the strategy for two-head case, and\n   pull.octopus names the strategy for octopus merge.\n\nIOW you are paranoid, you can have the following lines in your\n.git/config file and keep using git-merge-resolve when pulling\none remote:\n\n\t[pull]\n\t\ttwohead = resolve\n\nOTOH, you can say this:\n\n\t[pull]\n\t\ttwohead = resolve\n\t\ttwohead = recursive\n\nto try quicker resolve first, and when it fails, fall back to\nrecursive.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n\n---\n\n  Linus Torvalds <torvalds@osdl.org> writes:\n\n  > Hmm. True. The _really_ trivial in-index case triggers for me pretty \n  > often, but I haven't done any statistics. It might be only 50% of the \n  > time.\n  >...\n  > It's certainly an option to just do what I just did, namely use the \n  > default one until it breaks, and then just do \"git reset --hard\" and re-do \n  > the pull with \"-s recursive\". A bit sad, and it would be good to have \n  > coverage on the recursive strategy..\n\n  Hopefully something like this would make people aware of\n  recursive and give it a wider coverage and chance to mature.\n\n git-pull.sh |   16 ++++++++++++++--\n 1 files changed, 14 insertions(+), 2 deletions(-)\n\napplies-to: 75922cf23cc070e2d5220d961a8f645f1bc8bb60\n3acc20beaf0df9ce11a1b7aabf8c9dc7507a9b44\ndiff --git a/git-pull.sh b/git-pull.sh\nindex 2358af6..3b875ad 100755\n--- a/git-pull.sh\n+++ b/git-pull.sh\n@@ -79,10 +79,22 @@ case \"$merge_head\" in\n \texit 0\n \t;;\n ?*' '?*)\n-\tstrategy_default_args='-s octopus'\n+\tvar=`git-var -l | sed -ne 's/^pull\\.octopus=/-s /p'`\n+\tif test '' = \"$var\"\n+\tthen\n+\t\tstrategy_default_args='-s octopus'\n+\telse\n+\t\tstrategy_default_args=$var\n+\tfi\n \t;;\n *)\n-\tstrategy_default_args='-s resolve'\n+\tvar=`git-var -l | sed -ne 's/^pull\\.twohead=/-s /p'`\n+\tif test '' = \"$var\"\n+\tthen\n+\t\tstrategy_default_args='-s recursive'\n+\telse\n+\t\tstrategy_default_args=$var\n+\tfi\n \t;;\n esac\n \n---\n0.99.9.GIT\n"},{"id":"11346","messageId":"20051108210211.GA23265@c165.ib.student.liu.se","threadId":"2386","inReplyTo":"Pine.LNX.4.63.0511081254520.2649@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Comments on recursive merge..","fromName":"Fredrik Kuivinen","fromEmail":"freku045@student.liu.se","sentAt":"2005-11-08T21:02:11Z","receivedAt":"2005-11-08T21:02:11Z","isPatch":false,"sender":{"key":"frekui@gmail.com","avatar":"https://avatars.githubusercontent.com/u/13770967?v=4"},"body":"On Tue, Nov 08, 2005 at 12:58:50PM +0100, Johannes Schindelin wrote:\n> Hi,\n> \n> On Mon, 7 Nov 2005, Linus Torvalds wrote:\n> \n> > Is the recursive thing noticeably slower for the \"easy\" cases (ie things \n> > that the old regular resolve strategy does well)?\n> \n> IIRC recursive does nothing else than recursively merging the merge-bases \n> (granted, in a clever way). So if there is only one merge-base, the only \n> slow-down would be the startup of python (which is probably worth it, \n> anyway).\n> \n\nI haven't done any real measurements but my feeling is that the\nrecursive strategy is at least not very much slower than the resolve\nstrategy.\n\nIn the single-common-ancestor case I can think of the following things\nwhich may make a difference speed wise:\n\n* The recursive strategy is written in Python\n* The code for finding common ancestors is also written in Python and\n  is probably a bit slower than git-merge-base.\n* git-diff-tree -M --diff-filter=R <common ancestor> <branch> is\n  executed twice, once for each branch.\n\nOn the positive side the code which corresponds to git-merge-one-file\nin the git-resolve case is also written in python, we can therefore\navoid some forks and execs.\n\n> > It's certainly an option to just do what I just did, namely use the \n> > default one until it breaks, and then just do \"git reset --hard\" and re-do \n> > the pull with \"-s recursive\". A bit sad, and it would be good to have \n> > coverage on the recursive strategy..\n> \n> We already have a fallback list: after really-trivial, try automatic, ...,\n> try resolve. Why not just add recursive? So, if even resolve failed, just \n> try once more, with recursive.\n> \n\nI don't think this is a very good idea for two reasons. The first one\nis that there are some merge scenarios involving renames which should\nbe conflicts but are cleanly merged by git-resolve.\n\nThe second reason is that with the fall back list the recursive\nstrategy will only be used in the strange corner cases and will thus\nnot get nearly the same amount of testing it would get if it was the\nfirst choice (or directly after the really-trivial merge).\n\n- Fredrik\n"},{"id":"11350","messageId":"7vwtjierrx.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"20051108210211.GA23265@c165.ib.student.liu.se","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-08T21:47:14Z","receivedAt":"2005-11-08T21:47:14Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Fredrik Kuivinen <freku045@student.liu.se> writes:\n\n> The second reason is that with the fall back list the recursive\n> strategy will only be used in the strange corner cases and will thus\n> not get nearly the same amount of testing it would get if it was the\n> first choice (or directly after the really-trivial merge).\n\nThere are two reasons to avoid git-merge choose from more than\none strategy.\n\n1. The whole idea that git-merge implements \"goodness\" metric is\n   bogus.  It does not know what merge strategy is good and that\n   is the reason it punts and has the user choose his preferred\n   strategy.\n\n2. When it is going to loop over more than one strategy, it\n   stashes away the current working tree state, so that the\n   second and subsequent strategies can begin from a clean slate\n   (including local modifications since the current head).  If\n   we try only one, there is no such cost involved.\n\nI think the patch I sent out last night to change the recursive\nas the default strategy and make it overridable from the\nconfiguration mechanism would be a better way to give people\nmore exposure to the greatness of recursive while protecting\nthem from potential glitches if any.\n"},{"id":"11351","messageId":"Pine.LNX.4.64.0511081351020.3247@g5.osdl.org","threadId":"2386","inReplyTo":"20051108210211.GA23265@c165.ib.student.liu.se","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-08T21:52:06Z","receivedAt":"2005-11-08T21:52:06Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 8 Nov 2005, Fredrik Kuivinen wrote:\n>\n> * The code for finding common ancestors is also written in Python and\n>   is probably a bit slower than git-merge-base.\n\nBtw, what part of git-merge-bases is it that makes it not be practical?\n\n\t\t\tLinus\n"},{"id":"11354","messageId":"20051108223609.GA4805@c165.ib.student.liu.se","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511081351020.3247@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Fredrik Kuivinen","fromEmail":"freku045@student.liu.se","sentAt":"2005-11-08T22:36:09Z","receivedAt":"2005-11-08T22:36:09Z","isPatch":false,"sender":{"key":"frekui@gmail.com","avatar":"https://avatars.githubusercontent.com/u/13770967?v=4"},"body":"On Tue, Nov 08, 2005 at 01:52:06PM -0800, Linus Torvalds wrote:\n> \n> \n> On Tue, 8 Nov 2005, Fredrik Kuivinen wrote:\n> >\n> > * The code for finding common ancestors is also written in Python and\n> >   is probably a bit slower than git-merge-base.\n> \n> Btw, what part of git-merge-bases is it that makes it not be practical?\n> \n\nThe problem is in the multiple-common-ancestors case. If we have three\ncommon ancestors, A, B and C, we will start with merging A with B. The\nresult is a new 'virtual' commit object (not stored in the object\ndatabase), lets call it V. We are then going to merge V with C. To do\nthat we need to get the common ancestor(s) of V and C, and as V\ndoesn't exist in the database we can't use git-merge-base.\n\nI haven't given it a lot of thought though, it might be possible to\nuse git-merge-base in some way and get the same results as we get now.\n\nIt would certainly be possible to use git-merge-base in the first\niteration and use the python code only when we actually have any\n'virtual' commit objects.\n\n- Fredrik\n"},{"id":"11356","messageId":"Pine.LNX.4.63.0511090003280.27861@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"2386","inReplyTo":"20051108210211.GA23265@c165.ib.student.liu.se","subject":"Re: Comments on recursive merge..","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2005-11-08T23:04:26Z","receivedAt":"2005-11-08T23:04:26Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 8 Nov 2005, Fredrik Kuivinen wrote:\n\n> On Tue, Nov 08, 2005 at 12:58:50PM +0100, Johannes Schindelin wrote:\n> > \n> > We already have a fallback list: after really-trivial, try automatic, ...,\n> > try resolve. Why not just add recursive? So, if even resolve failed, just \n> > try once more, with recursive.\n> > \n> \n> I don't think this is a very good idea for two reasons. The first one\n> is that there are some merge scenarios involving renames which should\n> be conflicts but are cleanly merged by git-resolve.\n> \n> The second reason is that with the fall back list the recursive\n> strategy will only be used in the strange corner cases and will thus\n> not get nearly the same amount of testing it would get if it was the\n> first choice (or directly after the really-trivial merge).\n\nTwo very valid points. You convinced me.\n\nCiao,\nDscho\n"},{"id":"11357","messageId":"Pine.LNX.4.64.0511081450080.3247@g5.osdl.org","threadId":"2386","inReplyTo":"20051108223609.GA4805@c165.ib.student.liu.se","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-08T23:05:43Z","receivedAt":"2005-11-08T23:05:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 8 Nov 2005, Fredrik Kuivinen wrote:\n> \n> The problem is in the multiple-common-ancestors case. If we have three\n> common ancestors, A, B and C, we will start with merging A with B. The\n> result is a new 'virtual' commit object (not stored in the object\n> database), lets call it V. We are then going to merge V with C. To do\n> that we need to get the common ancestor(s) of V and C, and as V\n> doesn't exist in the database we can't use git-merge-base.\n\nHmm. That's really the same as the merge-base of \"C _and_ (A _or_ B)\", \nisn't it?\n\nSo we should be able to do that even without ever seeing the virtual \nmerge. In fact, it really is pretty trivial from a technical standpoint: \nthe \"A _or_ B\" part is really just inserting both A and B with the same \n\"flags\" value (see the \"merge_base()\" function in merge-base.c).\n\nSo in general, \"merge-base\" could trivially be extended to have any number \nof \"OR commits\" on either side, as long as there is just one \"and\".\n\nIt could also be extended to have multiple \"and\" cases, but that has \nactually already been done by \"git-show-branch\", I think. It's all the \nsame logic, except it uses more than just two bits.\n\nSo _technically_ it should be easy to do, it would just need some sane \ncommand line syntax to specify the grouping.\n\n> I haven't given it a lot of thought though, it might be possible to\n> use git-merge-base in some way and get the same results as we get now.\n> \n> It would certainly be possible to use git-merge-base in the first\n> iteration and use the python code only when we actually have any\n> 'virtual' commit objects.\n\nJust handling the first case specially would be sufficient for 99% of all \nuses.  And if the multi-parent cases are slightly slower, I don't think \nanybody cares.\n\nIn fact, we _always_ do the first git-merge-base in git-merge.sh anyway \n(actually, not git-merge-base, but \"git-show-branch --merge-base\"). So \nthat's really done already for the \"test if it's trivial\" case..\n\nJunio, that points out that \"git-merge-base\" is another program that could \njust be removed, since it's really supreceded by git-show-branch. Or did I \nmiss something?\n\n\n\t\t\tLinus\n"},{"id":"11358","messageId":"Pine.LNX.4.63.0511090017370.28256@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511081450080.3247@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2005-11-08T23:18:56Z","receivedAt":"2005-11-08T23:18:56Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 8 Nov 2005, Linus Torvalds wrote:\n\n> Junio, that points out that \"git-merge-base\" is another program that could \n> just be removed, since it's really supreceded by git-show-branch. Or did I \n> miss something?\n\nIIRC, git-show-branch has a limit on the number of refs it can take.\n\nCiao,\nDscho\n"},{"id":"11362","messageId":"Pine.LNX.4.64.0511081614140.3247@g5.osdl.org","threadId":"2386","inReplyTo":"Pine.LNX.4.63.0511090017370.28256@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-09T00:18:53Z","receivedAt":"2005-11-09T00:18:53Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 9 Nov 2005, Johannes Schindelin wrote:\n> \n> On Tue, 8 Nov 2005, Linus Torvalds wrote:\n> \n> > Junio, that points out that \"git-merge-base\" is another program that could \n> > just be removed, since it's really supreceded by git-show-branch. Or did I \n> > miss something?\n> \n> IIRC, git-show-branch has a limit on the number of refs it can take.\n\nWell, git-merge-base does too. git-merge-base only takes two refs ;)\n\nIn general, you need to keep track of one bit per ref, and since we have \na 32-bit \"flags\" word and need a couple of bits for other maintenance \ninfo, pretty much anything that figures out common heads will be limited \nsome way. \n\nThis is only a limit for the \"and\" logic - the \"or\" logic (if we implement \nit) will just share the same status bit for all the refs that are \"ored \ntogether\" and thus has no limits. \n\nOh, and the \"and\" logic can be extended by running the program multiple \ntimes, so it's not a \"hard\" limit, it's just an issue of convenience.\n\nThat said, anybody who ever does an octopus of more than just a few heads \ndeserves to be shot, so I don't think the limit should matter. The \nrecursive strategy should only add the \"or\" kind of refs, and it \nshouldn't be a problem (apart from just how to describe them).\n\n\t\tLinus\n"},{"id":"11364","messageId":"20051109003236.GA30496@pasky.or.cz","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511081450080.3247@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2005-11-09T00:32:36Z","receivedAt":"2005-11-09T00:32:36Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Wed, Nov 09, 2005 at 12:05:43AM CET, I got a letter\nwhere Linus Torvalds <torvalds@osdl.org> said that...\n> Junio, that points out that \"git-merge-base\" is another program that could \n> just be removed, since it's really supreceded by git-show-branch. Or did I \n> miss something?\n\nWow, I didn't know git-show-branch could do that (even though it's a bit\nunnatural to expect this from command named this way; then again,\nthere's git-rev-parse...).\n\nBTW, git-show-branch is also by orders of magnitude faster (not that\nthis would be any major timesaver). Median 0.006s vs. median 0.124s.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nVI has two modes: the one in which it beeps and the one in which\nit doesn't.\n"},{"id":"11365","messageId":"Pine.LNX.4.64.0511081646160.3247@g5.osdl.org","threadId":"2386","inReplyTo":"20051109003236.GA30496@pasky.or.cz","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-09T00:51:11Z","receivedAt":"2005-11-09T00:51:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 9 Nov 2005, Petr Baudis wrote:\n> \n> BTW, git-show-branch is also by orders of magnitude faster (not that\n> this would be any major timesaver). Median 0.006s vs. median 0.124s.\n\nOuch. That makes me suspicious. One reason git-merge-base is slow is \nbecause it's being pretty careful about some pathological examples of \ndates being just the wrong way around, and it might just be that the \nreason git-show-branch is faster is because it isn't doing that part \nright.\n\nSo yes, git-merge-base does extra work, but it does so because I think it \nneeds to.\n\nJunio? You even wrote the comment about the case in git-merge-base, I'm \nwondering whether it's a bug that we use the fast-and-cheap algorithm in \ngit-show-branch..\n\nOf course, arguably you can first try the fast-and-cheap thing, and if \nthat gives a merge parent that is acceptable, why not? So maybe it's the \nright thing for the \"let's see if this is trivial\" case, but I think it \nmight _think_ some cases are trivial that really shouldn't, because they \nactually have two merge parents.\n\n\t\tLinus\n"},{"id":"11366","messageId":"7vlkzyd4aq.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511081646160.3247@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-09T00:59:41Z","receivedAt":"2005-11-09T00:59:41Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Junio? You even wrote the comment about the case in git-merge-base, I'm \n> wondering whether it's a bug that we use the fast-and-cheap algorithm in \n> git-show-branch..\n\nI did show-branch soon after we worked on those pathlogical\nmerge-base fix, so I would be a bit surprised if I did it\nwithout using all the knowledge from that exercise, but I do not\nremember offhand.  The core logic should be simple\ngeneralization of two-head merge-base to N heads.\n"},{"id":"11370","messageId":"Pine.LNX.4.64.0511081716450.3247@g5.osdl.org","threadId":"2386","inReplyTo":"7vlkzyd4aq.fsf@assigned-by-dhcp.cox.net","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-09T01:22:15Z","receivedAt":"2005-11-09T01:22:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 8 Nov 2005, Junio C Hamano wrote:\n> \n> I did show-branch soon after we worked on those pathlogical\n> merge-base fix, so I would be a bit surprised if I did it\n> without using all the knowledge from that exercise, but I do not\n> remember offhand.  The core logic should be simple\n> generalization of two-head merge-base to N heads.\n\nHmm.\n\nLook at the \"join_revs()\" logic, and tell me I'm crazy.\n\nIt does:\n\n\tstruct commit *commit = pop_one_commit(list_p);\n\tint still_interesting = !!interesting(*list_p);\n\nin that order: it looks whether there are any interesting commits left \n_after_ it has popped the top-of-stack.\n\nWhich means that \"still_interesting\" can go down to zero if we just popped \nthe last interesting thing off the stack.\n\nWhich seems wrong, because the thing we just popped off the stack could \neasily itself be interesting (in fact, it should be so, 99% of the time), \nand can cause other interesting commits to be populated back onto the \nlist. So the \"still_interesting\" flag seems to be wrongly computed: the \nway it is computed now, it's meaningless.\n\nIn contrast, the \"merge_base()\" thing does\n\n\twhile (interesting(list)) {\n\t\t..\n\t}\n\nwhich means that we really will walk the list until there is nothing \ninteresting left. Which is admittedly expensive, but it was how we got rid \nof the pathological case.\n\nBut maybe I'm just missing something really subtle. Maybe git-show-branch \ndoes some really clever optimization that is valid.\n\n\t\tLinus\n"},{"id":"11374","messageId":"7v8xvyd2bh.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511081716450.3247@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-09T01:42:26Z","receivedAt":"2005-11-09T01:42:26Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> It does:\n>\n> \tstruct commit *commit = pop_one_commit(list_p);\n> \tint still_interesting = !!interesting(*list_p);\n>\n> in that order: it looks whether there are any interesting commits left \n> _after_ it has popped the top-of-stack.\n\nAhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhhh.  You are right.\n\nThe problem is most of the time hidden, because we usually do\none extra round (extra usually starts from 0 and we break out\nafter we say \"not interesting anymore\" and extra < 0).\n\nObviously, I was not thinking clearly.\n"},{"id":"11377","messageId":"7v7jbinyg9.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511081614140.3247@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-09T06:10:30Z","receivedAt":"2005-11-09T06:10:30Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> In general, you need to keep track of one bit per ref, and since we have \n> a 32-bit \"flags\" word and need a couple of bits for other maintenance \n> info, pretty much anything that figures out common heads will be limited \n> some way. \n>\n> This is only a limit for the \"and\" logic - the \"or\" logic (if we implement \n> it) will just share the same status bit for all the refs that are \"ored \n> together\" and thus has no limits. \n>\n> Oh, and the \"and\" logic can be extended by running the program multiple \n> times, so it's not a \"hard\" limit, it's just an issue of convenience.\n>\n> That said, anybody who ever does an octopus of more than just a few heads \n> deserves to be shot, so I don't think the limit should matter. The \n> recursive strategy should only add the \"or\" kind of refs, and it \n> shouldn't be a problem (apart from just how to describe them).\n\nCome to think of it, git-merge-octopus does AND.  If I am\nmerging topic branches 1, 2, 3,... N into my master, internally\nit does an equivalent of merging 1 into master, then 2 into the\nresult, then C into that result,..., and it uses merge-base of\nall the heads merged so far and the original master to pivot on.\n\nAnd I think this is *wrong*.  The merge base of each step when\nmerging head N does not have to be older than merge base of the\noriginal master and head N, but currently that is not what it\ndoes.  I should be ORing them ideally, but even if I do not, I\nshould be able to just use the merge base of head N and original\nmaster.\n"},{"id":"11389","messageId":"7v4q6mgm1l.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"7v8xvyd2bh.fsf@assigned-by-dhcp.cox.net","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-09T10:20:22Z","receivedAt":"2005-11-09T10:20:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> Linus Torvalds <torvalds@osdl.org> writes:\n>\n>> It does:\n>>\n>> \tstruct commit *commit = pop_one_commit(list_p);\n>> \tint still_interesting = !!interesting(*list_p);\n>>\n>> in that order: it looks whether there are any interesting commits left \n>> _after_ it has popped the top-of-stack.\n>\n> The problem is most of the time hidden,...\n\nAs you pointed out, still_interesting means \"after we are done\nwith this commit, do we still have something interesting to be\nprocessed?\", and the later \"extra < 0\" check compensates for\nthis.  After I pop the last interesting commit, I still look at\nits parents and push them back into the list.\n\nIt seems to be doing the right thing after all.  I hate to admit\nit, but I have been having hard time figuring out how this thing\nworks X-<.  In the meantime, I've checked commits from linux-2.6\nhistory that have more than one merge-base candidates.\n\"git-merge-base --all\" and \"git-show-branch --merge-base\" give\nthe same answer to all of them [*1*].\n\nI do not think \"git-show-branch --merge-base\" can be any more\nefficient than \"git-merge-base --all\".  It does _more_ things\n(probably unnecessary things as well).  Pasky's number could be\njust an artifact of hot/cold cache difference.\n\n[Footnote]\n\n*1* Here are the commits I used from linux-2.6 repository that\nhave more than one commits:\n\n    ba9b543d5bec0a7605952e2ba501fb8b0f3b6407\n    84ffa747520edd4556b136bdfc9df9eb1673ce12\n    da28c12089dfcfb8695b6b555cdb8e03dda2b690\n    3190186362466658f01b2e354e639378ce07e1a9\n    0c168775709faa74c1b87f1e61046e0c51ade7f3\n    0e396ee43e445cb7c215a98da4e76d0ce354d9d7\n    467ca22d3371f132ee225a5591a1ed0cd518cb3d\n"},{"id":"11383","messageId":"20051109103655.GB4960@c165.ib.student.liu.se","threadId":"2386","inReplyTo":"7v64r4qai6.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] merge-recursive: Only print relevant rename messages","fromName":"Fredrik Kuivinen","fromEmail":"freku045@student.liu.se","sentAt":"2005-11-09T10:36:55Z","receivedAt":"2005-11-09T10:36:55Z","isPatch":true,"sender":{"key":"frekui@gmail.com","avatar":"https://avatars.githubusercontent.com/u/13770967?v=4"},"body":"On Mon, Nov 07, 2005 at 03:54:57PM -0800, Junio C Hamano wrote:\n> Fredrik Kuivinen <freku045@student.liu.se> writes:\n> \n> > @@ -178,7 +181,7 @@ def mergeFile(oPath, oSha, oMode, aPath,\n> >              sha = bSha\n> >      else:\n> >          if aSha != oSha and bSha != oSha:\n> > -            merge = True\n> > +            merge = MERGE_TRIVIAL\n> \n> The rest looks good to me, but are you sure about this part?  I\n> have a feeling that the above \"and\" should be \"or\", meaning, we\n> check to see if there is _any_ change, and default to TRIVIAL,\n> but later we would find that we need a real merge and then\n> promote it to MERGE_3WAY.\n> \n\nYou are right. The code actually do the right thing, but it does it by\naccident. Please apply the following patch.\n\n---\n\nmerge-recursive: Fix limited output of rename messages\n\nThe previous code did the right thing, but it did it by accident.\n\nSigned-off-by: Fredrik Kuivinen <freku045@student.liu.se>\n\n\n---\n\n git-merge-recursive.py |   12 ++++--------\n 1 files changed, 4 insertions(+), 8 deletions(-)\n\napplies-to: bb7dd65e1d945edbe0137a761ebc388c7394067a\nf56613498cd7fb7013f532a04e63b580314ed957\ndiff --git a/git-merge-recursive.py b/git-merge-recursive.py\nindex 9983cd9..3657875 100755\n--- a/git-merge-recursive.py\n+++ b/git-merge-recursive.py\n@@ -162,13 +162,10 @@ def mergeTrees(head, merge, common, bran\n # Low level file merging, update and removal\n # ------------------------------------------\n \n-MERGE_NONE = 0\n-MERGE_TRIVIAL = 1\n-MERGE_3WAY = 2\n def mergeFile(oPath, oSha, oMode, aPath, aSha, aMode, bPath, bSha, bMode,\n               branch1Name, branch2Name):\n \n-    merge = MERGE_NONE\n+    merge = False\n     clean = True\n \n     if stat.S_IFMT(aMode) != stat.S_IFMT(bMode):\n@@ -181,7 +178,7 @@ def mergeFile(oPath, oSha, oMode, aPath,\n             sha = bSha\n     else:\n         if aSha != oSha and bSha != oSha:\n-            merge = MERGE_TRIVIAL\n+            merge = True\n \n         if aMode == oMode:\n             mode = bMode\n@@ -211,7 +208,6 @@ def mergeFile(oPath, oSha, oMode, aPath,\n             os.unlink(src1)\n             os.unlink(src2)\n \n-            merge = MERGE_3WAY\n             clean = (code == 0)\n         else:\n             assert(stat.S_ISLNK(aMode) and stat.S_ISLNK(bMode))\n@@ -590,7 +586,7 @@ def processRenames(renamesA, renamesB, b\n                 if merge or not clean:\n                     print 'Renaming', fmtRename(path, ren1.dstName)\n \n-                if merge == MERGE_3WAY:\n+                if merge:\n                     print 'Auto-merging', ren1.dstName\n \n                 if not clean:\n@@ -668,7 +664,7 @@ def processRenames(renamesA, renamesB, b\n                 if merge or not clean:\n                     print 'Renaming', fmtRename(ren1.srcName, ren1.dstName)\n \n-                if merge == MERGE_3WAY:\n+                if merge:\n                     print 'Auto-merging', ren1.dstName\n \n                 if not clean:\n---\n0.99.9.GIT\n"},{"id":"11392","messageId":"20051109145929.GE30496@pasky.or.cz","threadId":"2386","inReplyTo":"7v4q6mgm1l.fsf@assigned-by-dhcp.cox.net","subject":"Re: Comments on recursive merge..","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2005-11-09T14:59:29Z","receivedAt":"2005-11-09T14:59:29Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Wed, Nov 09, 2005 at 11:20:22AM CET, I got a letter\nwhere Junio C Hamano <junkio@cox.net> said that...\n> I do not think \"git-show-branch --merge-base\" can be any more\n> efficient than \"git-merge-base --all\".  It does _more_ things\n> (probably unnecessary things as well).  Pasky's number could be\n> just an artifact of hot/cold cache difference.\n\nCertainly not that. But I've fetched in the meantime and now show-branch\ntakes much longer - median 0.078s (git-merge-base's median still stays\naround 0.128s). So possibly git-show-branch did some smart optimization\nright away in the previous case. I can try to track down the particular\ncommits if there's any interest.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nVI has two modes: the one in which it beeps and the one in which\nit doesn't.\n"},{"id":"11393","messageId":"Pine.LNX.4.64.0511090800330.3247@g5.osdl.org","threadId":"2386","inReplyTo":"7v4q6mgm1l.fsf@assigned-by-dhcp.cox.net","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-09T16:30:49Z","receivedAt":"2005-11-09T16:30:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 9 Nov 2005, Junio C Hamano wrote:\n> \n> As you pointed out, still_interesting means \"after we are done\n> with this commit, do we still have something interesting to be\n> processed?\", and the later \"extra < 0\" check compensates for\n> this.  After I pop the last interesting commit, I still look at\n> its parents and push them back into the list.\n\nThat \"extra\" check only helps once. If we ever hit the \"extra--\", it's \ngone.\n\nIn other words, follow this:\n\n - we start out with \"extra = 0\" (default value)\n - we've got one \"interesting\" commit left, and we just popped it.\n - we now have \"still_interesting = 0\"\n - the commit has just one parent, and it's not something we've seen \n   before, so we add it to the seen list and decrement \"extra\", which is \n   now -1. We then insert it back to the list.\n - we go back up, pop the thing we just got, and now there are again no \n   interesting commits on the list any more, so \"still_interesting = 0\".\n - now \"extra\" is -1, and we break out of the loop without ever \n   percolating the flags of this commit to its parents.\n\nNo?\n\n> It seems to be doing the right thing after all.  I hate to admit it, but \n> I have been having hard time figuring out how this thing works X-<.  In \n> the meantime, I've checked commits from linux-2.6 history that have more \n> than one merge-base candidates.\n\nI'm not very impressed by \"it works for the seven cases I tried\".\n\nIt's entirely possible that there _is_ some reason it always works, but if \nso, I'd like to understand it. More likely, it works in _practice_ because \nthe only way to trigger anything else is likely such a perverse commit \nhistory that you'd never see it, but hey..\n\nAlso, I don't think this has necessarily anything to do with \"multiple \nmerge bases\". As far as I can tell, we can find a potential \"merge base\" \nthat starts the culling of uniniteresting things, but some other branch \n(that we haven't followed yet - perhaps the one we just broke out of \nearly) may end up causing an _earlier_ commit to turn out to also be a \nmerge-base, and the merge-base we found originally turns out to be a \nparent of the new one, and thus totally uninteresting.\n\nSee what I'm saying? Even with just _one_ well-defined merge base, we \nmight hit it.\n\nIt so happens that because we traverse the commit history in date order, \nwe almost never (but the keyword here is _almost_) hit the case where a \nchild of a commit ends up being parsed _after_ the commit that is its \nparent. That only happens when there are non-synchronized clocks etc, and \nthere are very few cases of that in the kernel tree.\n\nJust to see how rare that is, do this:\n\n\tgit-rev-list --pretty=raw HEAD |\n\t\tgrep '^committer' |\n\t\tcut -d'>' -f2 |\n\t\tcut -d' ' -f2 > date-list\n\nwhich basically generates the list of dates of commits in the kernel tree, \nsorted in the natural order that we always traverse the commits in.\n\nNow, do\n\n\tsort -nr date-list | diff -u date-list -\n\nto see how often the dates are off. I'm seeing only _three_ commits that \nhave time-warps (ie they were \"earlier\" than one of their parents). Out of \n13,000+.\n\nSo walking things in date order _almost_ always does the right thing just \nby mistake (well, it's not \"mistake\", of course. It's by design: it's the \nclosest we can get to a nice balanced walk. But the point is that it's \nstill just a heuristic, not something we can absolutely depend on).\n\nAnd THAT was the reason for the problem with the original git-merge-base \nalgorithm. Not multiple merge-bases (which was admittedly another \nproblem), but the fact that it didn't give the right merge-base at all due \nto time warps.\n\n(Again - it may be that there's something in show-branch that makes the \noptimization valid, but I just don't understand it).\n\n\t\t\tLinus\n"},{"id":"11401","messageId":"7virv1efzv.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511090800330.3247@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-09T20:13:56Z","receivedAt":"2005-11-09T20:13:56Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> That \"extra\" check only helps once. If we ever hit the \"extra--\", it's \n> gone.\n\nI think you are right here, but while digging into this I found\nan interesting case.\n\nThe current show-branch code does the same as merge-base in the\npathological example depicted in merge-base.c, but they seem to\ndo different things to this picture (commit grows from bottom to\ntop, time flows alphabetically; find base between G and H).\n\n                 H\n                / \\\n           G   A   \\\n           |\\ /     \\ \n           | B       \\\n           |  \\       \\\n            \\  C       F\n             \\  \\     / \n              \\  D   /   \n               \\ |  /\n                \\| /\n\t\t E\n\n\"git-merge-base --all\" says the merge bases are B and E, while\n\"show-branch --merge-base\" mentions only B.  In this case the\nlatter is probably the better answer.  Actually git-merge-base\nwithout --all only mentions E.  This is because we give up when\nwe find the list elements are all uninteresting.  And this is\nvery expensive to fix (I recall mentioning \"horizon effect\" last\ntime we worked on this --- around August 12th).\n\nG gets bit 1 and H gets bit 2.  Here is what happens in each\niteration:\n\n\tList\t\t\tA B C D E F G H\t\tResult\n\tG1 H2\t\t\t- - - - - - 1 2\n\tH2 E1 B1\t\t- 1 - - 1 - 1 2\n\tF2 E1 B1 A2\t\t2 1 - - 1 2 1 2\n\tE3 B1 A2\t\t2 1 - - 3 2 1 2\t\tE3\n\tB1 A2\t\t\t2 1 - - 3 2 1 2\t\tE3\n\tC1 A2\t\t\t2 1 1 - 3 2 1 2\t\tE3\n\tD1 A2\t\t\t2 1 1 1 3 2 1 2\t\tE3\n\tA2\t\t\t2 1 1 1 3 2 1 2\t\tE3\n\tB3\t\t\t2 3 1 1 3 2 1 2\t\tE3 B3\n\tC7\t\t\t2 3 7 1 3 2 1 2\t\tE3 B3\n\nWe popped B with flag 3, and started contaminating the well by\nreinjecting its parent C with flag 7.  That is all good, but\n\"while (interesting(list))\" check stops us from going further.\nIdeally the following two steps would have found out that E is\nalso uninteresting.\n\n\tD7\t\t\t2 3 7 7 3 2 1 2\n\tE7\t\t\t2 3 7 7 7 2 1 2\n\nBut that is expensive -- we would not know when to stop.\n\nA reproduction recipe is attached here, primarily so I do not\nhave to worry about losing it from /var/tmp/.\n\n-- >8 -- cut here -- >8 --\n#!/bin/sh\n\nrm -fr .git && git-init-db\nT=$(git-write-tree)\n\nM=1130000000\nZ=+0000\n\nexport GIT_COMMITTER_EMAIL=git@comm.iter.xz\nexport GIT_COMMITTER_NAME='C O Mmiter'\nexport GIT_AUTHOR_NAME='A U Thor'\nexport GIT_AUTHOR_EMAIL=git@au.thor.xz\n\ndoit() {\n\tOFFSET=$1; shift\n\tNAME=$1; shift\n\tPARENTS=\n\tfor P\n\tdo\n\t\tPARENTS=\"${PARENTS}-p $P \"\n\tdone\n\tGIT_COMMITTER_DATE=\"$(($M + $OFFSET)) $Z\"\n\tGIT_AUTHOR_DATE=$GIT_COMMITTER_DATE\n\texport GIT_COMMITTER_DATE GIT_AUTHOR_DATE\n\tcommit=$(echo $NAME | git-commit-tree $T $PARENTS)\n\techo $commit >.git/refs/tags/$NAME\n\techo $commit\n}\n\ncheckit() {\n    echo MB\n    git-merge-base --all \"$@\" | xargs git-name-rev\n    echo SB\n    git-show-branch --merge-base \"$@\" | xargs git-name-rev\n    git-show-branch --sha1-name --more=99 \"$@\"\n}\n\nE=$(doit 5 E)\nD=$(doit 4 D $E)\nF=$(doit 6 F $E)\nC=$(doit 3 C $D)\nB=$(doit 2 B $C)\nA=$(doit 1 A $B)\nG=$(doit 7 G $B $E)\nH=$(doit 8 H $A $F)\n\ncheckit $G $H\n\nexit\n"},{"id":"11409","messageId":"Pine.LNX.4.64.0511091348530.4627@g5.osdl.org","threadId":"2386","inReplyTo":"7virv1efzv.fsf@assigned-by-dhcp.cox.net","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-09T21:58:33Z","receivedAt":"2005-11-09T21:58:33Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 9 Nov 2005, Junio C Hamano wrote:\n> \n> The current show-branch code does the same as merge-base in the\n> pathological example depicted in merge-base.c, but they seem to\n> do different things to this picture (commit grows from bottom to\n> top, time flows alphabetically; find base between G and H).\n> \n>                  H\n>                 / \\\n>            G   A   \\\n>            |\\ /     \\ \n>            | B       \\\n>            |  \\       \\\n>             \\  C       F\n>              \\  \\     / \n>               \\  D   /   \n>                \\ |  /\n>                 \\| /\n> \t\t   E\n> \n> \"git-merge-base --all\" says the merge bases are B and E, while\n> \"show-branch --merge-base\" mentions only B.  In this case the\n> latter is probably the better answer.\n\nI don't agree.\n\nSure, B _may_ be the right answer for a particular merge strategt, but \nthere's no way of knowing. Maybe all the big changes came in through F, \nand H is the merge that sorted that out, and E actually ends up being the \nbetter base.\n\nSo I think from a correctness standpoint, the only thing that matters is \n\"git-merge-base --all\", and anything that doesn't know to return both E \nand B looks potentially buggy.\n\n> Actually git-merge-base without --all only mentions E.\n\nWell, we should really consider anything that doesn't take them all into \naccount to be a bug waiting to happen (or rather, a merge waiting for a \ndisaster), but E is the right one, since it's the more recent one).\n\nNow, this case obviously depends on history being almost maximally insane \n(ie pretty much _all_ the dates are wrong). So in practice we probably \ndon't care.\n\nSo maybe \"git-show-branch --merge-base\" ends up acceptable as a faster way \nto do the quick \"let's see if we can find _some_ merge-base to do the \nin-index merge with\", but personally I'd much rather always do a \n\"git-merge-base --all\", and only do the fast index merge if we only have \none potential parent.\n\nThat way there would never any question about what the \"quick merge\" does.\n\n\t\t\tLinus\n"},{"id":"11420","messageId":"7virv1a0ro.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511091348530.4627@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-09T22:56:27Z","receivedAt":"2005-11-09T22:56:27Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n>> \n>>                  H\n>>                 / \\\n>>            G   A   \\\n>>            |\\ /     \\ \n>>            | B       \\\n>>            |  \\       \\\n>>             \\  C       F\n>>              \\  \\     / \n>>               \\  D   /   \n>>                \\ |  /\n>>                 \\| /\n>> \t\t   E\n>> \n> So I think from a correctness standpoint, the only thing that matters is \n> \"git-merge-base --all\", and anything that doesn't know to return both E \n> and B looks potentially buggy.\n\nBut the point of well-poisoning you did in merge-base was to\ndetect that E is an ancestor of B and exclude it in the first\nplace.  If it matters what F does, it means checking ancestry\namong B C D E and declare that B is a better ancestor than C, D,\nE does not help or is sometimes harmful.  No question that B is\nalways superiour ancestor than C and D, but arguably the\npresence of F _might_ change situation for B vs E.\n\nI however do not see merge-base trying to take that into account\nand treat E differently from C and D in any way.  Only because F\nand E had newer timestamp than C and D, we ended up finding E\nfirst and did not poison E through B, and that's why you got\nboth B and E.  I think it was just an accident.  If F were older\nthan B, I suspect the result would have been very different.\n\n> Now, this case obviously depends on history being almost maximally insane \n> (ie pretty much _all_ the dates are wrong). So in practice we probably \n> don't care.\n\nI agree.  The above example was to answer my own question in\nthis message:\n\n\thttp://marc.theaimsgroup.com/?l=git&m=112382448222823\n\n> ... personally I'd much rather always do a \n> \"git-merge-base --all\", and only do the fast index merge if we only have \n> one potential parent.\n>\n> That way there would never any question about what the \"quick merge\" does.\n\nI agree we should try to stay away from \"heuristic\" and make\nthings safer, but after seeing the above, I'd need a bit more\ntime to convince myself that what 'git-merge-base --all' does is\n*the* safe approach.  Right now, it looks to me that both are\nheuristic that work most of the time (merge-base --all 99.99999%\nof the time, show-branch 99% of the time, or something like\nthat).\n"},{"id":"11431","messageId":"Pine.LNX.4.64.0511091518370.4627@g5.osdl.org","threadId":"2386","inReplyTo":"7virv1a0ro.fsf@assigned-by-dhcp.cox.net","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-09T23:34:49Z","receivedAt":"2005-11-09T23:34:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 9 Nov 2005, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n> >> \n> >>                  H\n> >>                 / \\\n> >>            G   A   \\\n> >>            |\\ /     \\ \n> >>            | B       \\\n> >>            |  \\       \\\n> >>             \\  C       F\n> >>              \\  \\     / \n> >>               \\  D   /   \n> >>                \\ |  /\n> >>                 \\| /\n> >> \t\t   E\n> >> \n> > So I think from a correctness standpoint, the only thing that matters is \n> > \"git-merge-base --all\", and anything that doesn't know to return both E \n> > and B looks potentially buggy.\n> \n> But the point of well-poisoning you did in merge-base was to\n> detect that E is an ancestor of B and exclude it in the first\n> place.\n\nAhh, you're right, and I'm wrong. That \"E\" is not a real merge-base, since \nthere _is_ a valid merge-base that is a direct descendant of it and thus \nobjectively better.\n\nAnd as to why git-merge-base returns E in the first place: it really \nshouldn't, but when it sees B, it can decide that C is uninteresting, and \nso there are no interesting commits left. So it never continues to walk D \nand thus never notices that D covers E and E is _also_ uninteresting.\n\nSo it thinks both E and B are interesting, and since E has a more recent \ndate, it will select that one when only showing one (and then show both \nwhen asked to).\n\n> I however do not see merge-base trying to take that into account\n> and treat E differently from C and D in any way.\n\nIt doesn't. git-merge-base simply walks the chain in as close to date \norder as it can, and when it decides that the rest of the chain is \nprovably uninteresting, it stops.\n\nWhich means that sometimes it can stop with too _many_ merge heads, just \nbecause it hasn't realized that they are reachable through a chain that is \notherwise provably uninteresting.\n\nThis is because we define \"uninteresting\" as meaning \"cannot reach any \nmore _new_ merge-heads\". Which is true. The fact that such a chain could \nreach some heads we found earlier and mark them as being pointless never \nenters the picture ;)\n\n> I agree we should try to stay away from \"heuristic\" and make\n> things safer, but after seeing the above, I'd need a bit more\n> time to convince myself that what 'git-merge-base --all' does is\n> *the* safe approach.\n\nWell, \"git-merge-base --all\" will be \"safer\" in the sense that it's \nguaranteed to give a superset of the merge-heads (which itself is \"safe\" \nin that it flags potentially interesting cases early).\n\nThen, the recursive merge strategy could notice (in fact, _will_ notice, \nif it tries to merge the merge-heads) that the merge of such a pair of \nmerge-heads is one of the heads itself (just a fast-forward), and thus the \nrecursive strategy should correctly have chosen \"B\" as the merge-head.\n\nSo yes, \"git-merge-base --all\" really is safe. Sometimes (under fairly odd \ncircumstances) a bit unnecessarily conservative, but always safe.\n\n> Right now, it looks to me that both are heuristic that work most of the \n> time (merge-base --all 99.99999% of the time, show-branch 99% of the \n> time, or something like that).\n\nThe thing is, I don't see what guarantees that the show-branch brhaviour \nis safe or conservative. It happened to pick B in this case, which was the \nright choice, but I don't see how that was anything but just luck and \nhappenstance.\n\nIOW, I can see that \"git-merge-base --all\" can return some unnecessary \nheads, but I can also argue for how that becomes safe and fixes itself. \nWith git-show-branch --merge-base, I don't know what that argument is, \nbecause I can't see how it _guarantees_ that it would always pick B over \nE.\n\n\t\t\tLinus\n"},{"id":"11543","messageId":"7vzmobuc00.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511091348530.4627@g5.osdl.org","subject":"merge-base: fully contaminate the well.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-11T02:58:07Z","receivedAt":"2005-11-11T02:58:07Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"The discussion on the list demonstrated a pathological case where\nan ancestor of a merge-base can be left interesting.  This commit\nintroduces a postprocessing phase to fix it.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n\n---\n\n  Linus Torvalds <torvalds@osdl.org> writes:\n\n  > On Wed, 9 Nov 2005, Junio C Hamano wrote:\n  >\n  >> But the point of well-poisoning you did in merge-base was to\n  >> detect that E is an ancestor of B and exclude it in the first\n  >> place.\n  >\n  > Ahh, you're right, and I'm wrong. That \"E\" is not a real merge-base, since \n  > there _is_ a valid merge-base that is a direct descendant of it and thus \n  > objectively better.\n  >\n  > Which means that sometimes it can stop with too _many_ merge heads, just \n  > because it hasn't realized that they are reachable through a chain that is \n  > otherwise provably uninteresting.\n\n  I am not particularly proud of this change, but here is an\n  attempt to fully contaminate the well without going all the\n  way down to root.  It adds a postprocessing phase which does\n  not parse any new commits.\n\n  > The thing is, I don't see what guarantees that the show-branch brhaviour \n  > is safe or conservative.\n\n  You are right about this; I have a separate patch to fix it.\n\n\n merge-base.c |   78 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++-\n 1 files changed, 77 insertions(+), 1 deletions(-)\n\napplies-to: 8eacf17303188e55375f76bea8051555ba1baf02\na2fd94f707128f3f390362725c8d8b0802940111\ndiff --git a/merge-base.c b/merge-base.c\nindex 286bf0e..43a6818 100644\n--- a/merge-base.c\n+++ b/merge-base.c\n@@ -80,6 +80,45 @@ static struct commit *interesting(struct\n  * Now, list does not have any interesting commit.  So we find the newest\n  * commit from the result list that is not marked uninteresting.  Which is\n  * commit B.\n+ *\n+ *\n+ * Another pathological example how this thing can fail to mark an ancestor\n+ * of a merge base as UNINTERESTING without the postprocessing phase.\n+ *\n+ *\t\t  2\n+ *\t\t  H\n+ *\t    1    / \\\n+ *\t    G   A   \\\n+ *\t    |\\ /     \\ \n+ *\t    | B       \\\n+ *\t    |  \\       \\\n+ *\t     \\  C       F\n+ *\t      \\  \\     / \n+ *\t       \\  D   /   \n+ *\t\t\\ |  /\n+ *\t\t \\| /\n+ *\t\t  E\n+ *\n+ *\t list\t\t\tA B C D E F G H\n+ *\t G1 H2\t\t\t- - - - - - 1 2\n+ *\t H2 E1 B1\t\t- 1 - - 1 - 1 2\n+ *\t F2 E1 B1 A2\t\t2 1 - - 1 2 1 2\n+ *\t E3 B1 A2\t\t2 1 - - 3 2 1 2\n+ *\t B1 A2\t\t\t2 1 - - 3 2 1 2\n+ *\t C1 A2\t\t\t2 1 1 - 3 2 1 2\n+ *\t D1 A2\t\t\t2 1 1 1 3 2 1 2\n+ *\t A2\t\t\t2 1 1 1 3 2 1 2\n+ *\t B3\t\t\t2 3 1 1 3 2 1 2\n+ *\t C7\t\t\t2 3 7 1 3 2 1 2\n+ *\n+ * At this point, unfortunately, everybody in the list is\n+ * uninteresting, so we fail to complete the following two\n+ * steps to fully marking uninteresting commits.\n+ *\n+ *\t D7\t\t\t2 3 7 7 3 2 1 2\n+ *\t E7\t\t\t2 3 7 7 7 2 1 2\n+ *\n+ * and we end up showing E as an interesting merge base.\n  */\n \n static int show_all = 0;\n@@ -88,6 +127,7 @@ static int merge_base(struct commit *rev\n {\n \tstruct commit_list *list = NULL;\n \tstruct commit_list *result = NULL;\n+\tstruct commit_list *tmp = NULL;\n \n \tif (rev1 == rev2) {\n \t\tprintf(\"%s\\n\", sha1_to_hex(rev1->object.sha1));\n@@ -104,9 +144,10 @@ static int merge_base(struct commit *rev\n \n \twhile (interesting(list)) {\n \t\tstruct commit *commit = list->item;\n-\t\tstruct commit_list *tmp = list, *parents;\n+\t\tstruct commit_list *parents;\n \t\tint flags = commit->object.flags & 7;\n \n+\t\ttmp = list;\n \t\tlist = list->next;\n \t\tfree(tmp);\n \t\tif (flags == 3) {\n@@ -130,6 +171,41 @@ static int merge_base(struct commit *rev\n \tif (!result)\n \t\treturn 1;\n \n+\t/*\n+\t * Postprocess to fully contaminate the well.\n+\t */\n+\tfor (tmp = result; tmp; tmp = tmp->next) {\n+\t\tstruct commit *c = tmp->item;\n+\t\t/* Reinject uninteresting ones to list,\n+\t\t * so we can scan their parents.\n+\t\t */\n+\t\tif (c->object.flags & UNINTERESTING)\n+\t\t\tcommit_list_insert(c, &list);\n+\t}\n+\twhile (list) {\n+\t\tstruct commit *c = list->item;\n+\t\tstruct commit_list *parents;\n+\n+\t\ttmp = list;\n+\t\tlist = list->next;\n+\t\tfree(tmp);\n+\n+\t\t/* Anything taken out of the list is uninteresting, so\n+\t\t * mark all its parents uninteresting.  We do not\n+\t\t * parse new ones (we already parsed all the relevant\n+\t\t * ones).\n+\t\t */\n+\t\tparents = c->parents;\n+\t\twhile (parents) {\n+\t\t\tstruct commit *p = parents->item;\n+\t\t\tparents = parents->next;\n+\t\t\tif (!(p->object.flags & UNINTERESTING)) {\n+\t\t\t\tp->object.flags |= UNINTERESTING;\n+\t\t\t\tcommit_list_insert(p, &list);\n+\t\t\t}\n+\t\t}\n+\t}\n+\n \twhile (result) {\n \t\tstruct commit *commit = result->item;\n \t\tresult = result->next;\n---\n0.99.9.GIT\n"},{"id":"11551","messageId":"Pine.LNX.4.64.0511102125510.4627@g5.osdl.org","threadId":"2386","inReplyTo":"7vzmobuc00.fsf@assigned-by-dhcp.cox.net","subject":"Re: merge-base: fully contaminate the well.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-11T05:36:03Z","receivedAt":"2005-11-11T05:36:03Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 10 Nov 2005, Junio C Hamano wrote:\n>\n> The discussion on the list demonstrated a pathological case where\n> an ancestor of a merge-base can be left interesting.  This commit\n> introduces a postprocessing phase to fix it.\n\nHmm. I'd suggest only doing this for the (relatively unlikely) case of \nthere being more than one merge-base..\n\nYou don't even have to be exact about it: just see if the result list is \nbigger than one. (ie \"result->next != NULL\"). Yeah, sometimes there are \nresult commits that get turned UNINTERESTING after they are added to the \nresult list, and your \"contaminate\" phase would be strictly not needed\nthen, but that's already a pretty unusual case.\n\nSo the cheap test is to just say\n\n\t/* Do we have multiple results? */\n\tif (result->next)\n\t\tcontaminate_well(result);\n\nno?\n\nBtw, I don't think your contamination logic is necessarily complete. We \nmay not even have parsed some of the commits that end up being on that \nstrange corner case. I think you catch the particular case you tried, but \nI think that in theory, with long chains of commits out of date order, you \ncould be in the situation of having determined that everything was \nuninteresting before you even parsed enough to see the chain from one \nmerge-base to another.\n\nIn fact, I think it would happen with your pathological example if it had \njust one more commit out-of-order in the E-D-C-B chain. But I didn't walk \nit through.\n\n\t\tLinus\n"},{"id":"11554","messageId":"7viruzu3du.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511102125510.4627@g5.osdl.org","subject":"Re: merge-base: fully contaminate the well.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-11T06:04:13Z","receivedAt":"2005-11-11T06:04:13Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> So the cheap test is to just say\n>\n> \t/* Do we have multiple results? */\n> \tif (result->next)\n> \t\tcontaminate_well(result);\n>\n> no?\n\nCorrect.  And this is only for really artificial corner case so\nwe should try to avoid for normal cases as much as possible,\ncheaply.\n\n> Btw, I don't think your contamination logic is necessarily complete. We \n> may not even have parsed some of the commits that end up being on that \n> strange corner case.\n\nI haven't tried walking any other test cases, but wouldn't that\nbe arguing that the our assumption that the current merge-base\nis at least complete if not optimum?\n"},{"id":"11565","messageId":"7v8xvvr3jr.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511102125510.4627@g5.osdl.org","subject":"Re: merge-base: fully contaminate the well.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-11T08:28:56Z","receivedAt":"2005-11-11T08:28:56Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Btw, I don't think your contamination logic is necessarily complete. We \n> may not even have parsed some of the commits that end up being on that \n> strange corner case....\n\nYou are right.  And the situation seems really bad.\n\nThe full-contaminator is not full at all, and fails miserably in\nnot so pathlogical case.  If we have something like this:\n\n\t1\t2\tList\t\tA B C D E F G\n\tF\tE\tF1 E2\t\t- - - - 2 1 -\n        |\\     /|\tG1 E2 D1 C1\t- - 1 1 2 1 1\n        \\ \\   / |\tE2 D1 C1\t- - 1 1 2 1 1\n        |\\  D  /|\tG3 D3 C3\t- - 3 3 2 1 3\n        | \\ | / |\tD3 C3\t\t- - 3 3 2 1 3\n        |   C   |\tC7\t\t- - 7 3 2 1 3\n        |   |   |\n        |   B   |\n        |   |   /\n         \\  A  /\n          \\ | /\n            G\n\nwe would end up finding D and G and stop there, without ever\nseeing A or B.  B _might_ be touched when we look at C at the\nlast round, but there is no way for us to find G is reachable\nfrom D (or C) without parsing more than what we parsed in the\nmain loop.\n\nThe worst part of this is that you can indefinitely extend C-B-A\nchain trivially, and all it takes is the one, initial commit G,\nthat has a screwed-up timestamp.  All the other commits in this\nexample are in the right time order.  Very sad.\n"},{"id":"11604","messageId":"Pine.LNX.4.64.0511110805350.4627@g5.osdl.org","threadId":"2386","inReplyTo":"7viruzu3du.fsf@assigned-by-dhcp.cox.net","subject":"Re: merge-base: fully contaminate the well.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-11T16:18:21Z","receivedAt":"2005-11-11T16:18:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 10 Nov 2005, Junio C Hamano wrote:\n> \n> > Btw, I don't think your contamination logic is necessarily complete. We \n> > may not even have parsed some of the commits that end up being on that \n> > strange corner case.\n> \n> I haven't tried walking any other test cases, but wouldn't that\n> be arguing that the our assumption that the current merge-base\n> is at least complete if not optimum?\n\nOh yes, it's always complete, even though it may not be optimal. And it's \ngoing to be optimal in all realistic cases.\n\nAnd even in the unrealistic cases if we end up returning a commit that is \nactually reachable by another one (through at least two levels of other \ncommits that were in the wrong date-order with the _other_ ways of \nreaching that commit), the recursive strategy of merging the merge-bases \nwill always end up boiling it doing to the optimal thing.\n\nSo I think this is absolutely 100% correct.\n\n\t\t\tLinus\n"},{"id":"11630","messageId":"7v4q6ilt3m.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511071629270.3247@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-11T22:25:49Z","receivedAt":"2005-11-11T22:25:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n>> Another thing to consider is if it is fast enough for everyday\n>> trivial merges.\n>\n> Hmm. True. The _really_ trivial in-index case triggers for me pretty \n> often, but I haven't done any statistics. It might be only 50% of the \n> time.\n\nJust for fun, I randomly picked two heads/master commits from\nlinux-2.6 repository (one was when I happened to have pulled the\nlast time, and the other was when I thought this might be an\ninteresting exercise and pulled again), and fed the commits\nbetween the two to a little script that looks at commits and\ntries to stat what they did (the script ignores renames so they\nappear as deletes and adds).\n\nHere is what the script spitted out:\n\n        Total commit objects: 3957\n        Trivial Merges: 72 (1.82%)\n        Merges: 225 (5.69%)\n        Number of paths touched by non-merge commits:\n                average 4.50, median 2, min 2, max 199\n        Number of merge parents:\n                average 2.00, median 2, min 2, max 2\n        Number of merge bases:\n                average 1.00, median 1, min 1, max 1\n        File level merges:\n                average 37.61, median 8, min 0, max 555\n        Number of changed paths from the first parent:\n                average 379.09, median 66, min 1, max 7553\n        File level 3-ways:\n                average 1.96, median 1, min 0, max 37\n        Paths deleted:\n                average 47.56, median 15, min 0, max 554\n\nThis counts what happened in individual devleoper's trees,\nsubsystem maintainer trees and your tree, not just what you saw\nyourself.\n\nSome observations.\n\n - Trivial Merges count is surprisingly high.  About 1/3 of\n   merges are pure in-index merges.\n\n - Most of the commits (developer commits, not merges) are\n   small and touches only a couple of paths.\n\n - Nobody does octopus ;-).\n\n - We did not have multi-base merge case during the period\n   looked at (but the sample count is very low).\n\n - merge-one-file was called for only a handful (median 8)\n   files, which is negligibly small compared to the total 17K\n   files in the kernel tree, and fairly small compared to the\n   number of changed paths from the first parent (meaning,\n   read-tree trivial collapsing helped majorly).  Among them,\n   the number of paths that needed real file-level 3-way merges\n   were even smaller (avg 1.96).\n\n   All three of these points together is a fine demonstration\n   that you designed git really right.\n\nThe samples were between these two commits:\n\ncommit 6693e74a16ef563960764bd963f1048392135c3c\nAuthor: Linus Torvalds <torvalds@g5.osdl.org>\nDate:   Tue Oct 25 20:40:09 2005 -0700\n\ncommit 388f7ef720a982f49925e7b4e96f216f208f8c03\nAuthor: Linus Torvalds <torvalds@g5.osdl.org>\nDate:   Fri Nov 11 09:26:39 2005 -0800\n"},{"id":"11633","messageId":"Pine.LNX.4.64.0511111437410.3228@g5.osdl.org","threadId":"2386","inReplyTo":"7v4q6ilt3m.fsf@assigned-by-dhcp.cox.net","subject":"Re: Comments on recursive merge..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-11T22:53:56Z","receivedAt":"2005-11-11T22:53:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 11 Nov 2005, Junio C Hamano wrote:\n> \n> Some observations.\n> \n>  - Trivial Merges count is surprisingly high.  About 1/3 of\n>    merges are pure in-index merges.\n\nI actually don't think that is surprisingly high, and would actually have \nexpected it to be closer to 50%.\n\nOn the other hand, the merges that end up being pure fast-forwards aren't \ncounted as merges at all (since they don't show up as commits), so maybe \nthat's what skews my preception of a big percentage of merges as being \nreally trivial.\n\n>  - Most of the commits (developer commits, not merges) are\n>    small and touches only a couple of paths.\n\nThis is something where I think the kernel is perhaps unusual, especially \nfor a big project. We really do encourage people to make lots of small and \nwell-defined changes, and the whole flow of development has been geared \ntowards it. \n\n>  - Nobody does octopus ;-).\n\nI do think octopus is really cool, and still think seeing that five-way \noctopus-merge in gitk in the git history was really really cool.\n\nIt doesn't look as good any more, btw: do \"gitk\" on the current git tree, \nand search for \"Octopus merge\", and you'll see some of the history lines \ncrossing each other. Paul?\n\nBut yeah, it's a pretty special thing. I think its coolness factor way \noutweighs its usefullness factor ;^p\n\n>  - We did not have multi-base merge case during the period\n>    looked at (but the sample count is very low).\n\nAgain, this is possibly because the kernel has already had a few years of \ndistributed SCM usage under its belt, and we've tried to not only merge \nin a timely manner, but also try to keep history reasonably clean and not \nhave a lot of cross-merging back and forth. That cuts down on multi-base \npossibilities.\n\n>  - merge-one-file was called for only a handful (median 8)\n>    files, which is negligibly small compared to the total 17K\n>    files in the kernel tree, and fairly small compared to the\n>    number of changed paths from the first parent (meaning,\n>    read-tree trivial collapsing helped majorly).  Among them,\n>    the number of paths that needed real file-level 3-way merges\n>    were even smaller (avg 1.96).\n\nI definitely think this is true for any big project.\n\nSmall projects will inevitably have changes that modify large portions of \nthe source base. But with small projects, it doesn't really matter _what_ \nyou do, you can do it fast.\n\nBig projects (at least the sane kind) will never have lots of changes that \nmodify a very big percentage of the source-tree. It's just too painful \n(and I'm not talking from a SCM angle, just from a developer angle).\n\n>    All three of these points together is a fine demonstration\n>    that you designed git really right.\n\nWell, it's self-re-inforcing. It was designed for the kernel usage \npatterns, so using the kernel to confirm that it's the \"right design\" is a \nbit self-serving. Sure, it's a good sign that my mental model of what the \nusage patters are does actually match reality, but at the same time it \nmight be more interesting to see if other projects that use git end up \nusing it the same way and/or have different statistics.\n\nI do expect that the size of the project will impact the statistics a lot. \n\n\t\tLinus\n"},{"id":"11642","messageId":"7v8xvuk876.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"Pine.LNX.4.64.0511111437410.3228@g5.osdl.org","subject":"Re: Comments on recursive merge..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-12T00:42:37Z","receivedAt":"2005-11-12T00:42:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> I actually don't think that is surprisingly high, and would actually have \n> expected it to be closer to 50%.\n>\n> On the other hand, the merges that end up being pure fast-forwards aren't \n> counted as merges at all (since they don't show up as commits), so maybe \n> that's what skews my preception of a big percentage of merges as being \n> really trivial.\n\nThat is what I meant by \"it does not show just what you saw\".\nIt shows the whole community experience by everybody who has\ncommit there.  Being _the_ integration point, I imagine that\nmost of the pure fast-forwards you saw were real merges for\nsomebody else, and that merge being fast/safe/convenient\ncounts, not for you but for your subsystem people.\n\n>>    All three of these points together is a fine demonstration\n>>    that you designed git really right.\n>\n> Well, it's self-re-inforcing. It was designed for the kernel\n> usage patterns, so using the kernel to confirm that it's the\n> \"right design\" is a bit self-serving.\n\nThat is true.  My point was that it's not like git was done\nright only for _you_, sacrificing subsystem people.  The sample\n4k commits show that the assumption of the usage pattern git is\noptimized for actually holds (commits being small, merges being\nmostly trivial) for kernel people other than you.\n\nIt is yet to be seen if the same assumption holds for other\nprojects.\n"},{"id":"11650","messageId":"43758D21.3060107@michonline.com","threadId":"2386","inReplyTo":"7v4q6ilt3m.fsf@assigned-by-dhcp.cox.net","subject":"Re: Comments on recursive merge..","fromName":"Ryan Anderson","fromEmail":"ryan@michonline.com","sentAt":"2005-11-12T06:35:13Z","receivedAt":"2005-11-12T06:35:13Z","isPatch":false,"sender":{"key":"ryan@michonline.com","avatar":null},"body":"Junio C Hamano wrote:\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n> \n>>>Another thing to consider is if it is fast enough for everyday\n>>>trivial merges.\n>>\n>>Hmm. True. The _really_ trivial in-index case triggers for me pretty \n>>often, but I haven't done any statistics. It might be only 50% of the \n>>time.\n> \n> \n> Just for fun, I randomly picked two heads/master commits from\n> linux-2.6 repository (one was when I happened to have pulled the\n> last time, and the other was when I thought this might be an\n> interesting exercise and pulled again), and fed the commits\n> between the two to a little script that looks at commits and\n> tries to stat what they did (the script ignores renames so they\n> appear as deletes and adds).\n\nMind sharing the script?\n\nIt'be nice to know if these stats are typical, or unusual when you get\nnumbers from a variety of other trees.\n\n\n"},{"id":"11652","messageId":"7v7jbeia3v.fsf_-_@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"43758D21.3060107@michonline.com","subject":"[PATCH] GIT commit statistics.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-12T07:44:20Z","receivedAt":"2005-11-12T07:44:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ryan Anderson <ryan@michonline.com> writes:\n\n> Junio C Hamano wrote:\n>\n>> Just for fun, I randomly picked two heads/master commits from\n>> linux-2.6 repository ... and fed the commits\n>> between the two to a little script that looks at commits and\n>> tries to stat what they did (the script ignores renames so they\n>> appear as deletes and adds).\n>\n> Mind sharing the script?\n>\n> It'be nice to know if these stats are typical, or unusual when you get\n> numbers from a variety of other trees.\n\nVery unpolished but here they are.\n\nI misread the trivial count in my original message.  Trivial and\nMerge are counted separately, so among 3957 commits, merges were\n297 (72 trivials and 225 others).\n\n-- >8 -- cut here -- >8 --\nSubject: [PATCH] GIT commit statistics\n\nA set of scripts that read the existing commit history, and\nshow various stats.\n\nSample usage:\n\n    # Arguments are given to git-rev-list; defaults to ORIG..\n    # if not given, to retrace what was just pulled.\n    $ ./contrib/jc-git-stat-1.sh v0.99.9g..maint |\n      ./contrib/jc-git-stat-1-log.perl\n    Total commit objects: 43\n    Trivial Merges: 1 (2.33%)\n    Merges: 1 (2.33%)\n    Number of paths touched by non-merge commits:\n\t    average 3.00, median 2, min 2, max 18\n    Number of merge parents:\n\t    average 2.50, median 3, min 2, max 3\n    Number of merge bases:\n\t    average 1.00, median 1, min 1, max 1\n    File level merges:\n\t    average 0.50, median 1, min 0, max 1\n    Number of changed paths from the first parent:\n\t    average 28.00, median 52, min 4, max 52\n    File level 3-ways:\n\t    average 1.00, median 1, min 1, max 1\n\n * \"Trivial Merges\" are the ones done by read-tree --trivial;\n * \"Merges\" are other merges;\n * \"File level merges\" are paths not collapsed by read-tree 3-way (i.e.\n   given to merge-one-file);\n * \"File level 3-ways\" are paths merge-one-file would have run 'merge';\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n\n---\n\n contrib/jc-git-stat-1-log.perl |   87 ++++++++++++++++++++++++++++++++++++++++\n contrib/jc-git-stat-1-mof.sh   |   59 +++++++++++++++++++++++++++\n contrib/jc-git-stat-1.sh       |   74 ++++++++++++++++++++++++++++++++++\n 3 files changed, 220 insertions(+), 0 deletions(-)\n create mode 100755 contrib/jc-git-stat-1-log.perl\n create mode 100755 contrib/jc-git-stat-1-mof.sh\n create mode 100755 contrib/jc-git-stat-1.sh\n\napplies-to: 9a0f0c748316751fbf593a21f2b16bcdd975095a\n2cb3da4b260ed82dc379a11d91f55fe774a2ea49\ndiff --git a/contrib/jc-git-stat-1-log.perl b/contrib/jc-git-stat-1-log.perl\nnew file mode 100755\nindex 0000000..b70af2b\n--- /dev/null\n+++ b/contrib/jc-git-stat-1-log.perl\n@@ -0,0 +1,87 @@\n+#!/usr/bin/perl\n+\n+my ($patches, $failures, $merges, $trivials) = (0, 0, 0, 0);\n+my (@patch_paths,\n+    @parent_counts,\n+    @base_counts,\n+    @merge_counts,\n+    @path_counts,\n+    @res_counts,\n+    @merge_m, \n+    @merge_a,\n+    @merge_d, \n+    @merge_c, \n+    @merge_u);\n+\n+sub avg_median {\n+    my ($ary) = shift;\n+    my ($msg) = shift;\n+    my @a = sort { $a <=> $b } @$ary;\n+    my $sum = 0;\n+    for (@a) { $sum += $_ }\n+    return unless (@a && $sum);\n+    my ($avg, $med) = ($sum/@a, $a[(@a/2)]);\n+    my ($min, $max) = ($a[0], $a[$#a]);\n+    printf \"%s:\\n\\taverage %.2f, median %d, min %d, max %d\\n\",\n+    \t$msg, $avg, $med, $min, $max;\n+}\n+\n+while (<>) {\n+    next unless (s/^([MCFT]) [0-9a-f]{40} //);\n+    chomp;\n+    my $type = $1;\n+    if ($type eq 'F') {\n+\t$failures++;\n+\tnext;\n+    }\n+    if ($type eq 'C') {\n+\t$patches++;\n+\tpush @patch_paths, $_;\n+\tnext;\n+    }\n+    if ($type eq 'M') {\n+\t$merges++;\n+    }\n+    elsif ($type eq 'T') {\n+\t$trivials++;\n+    }\n+    else {\n+\tdie \"?? $type\";\n+    }\n+    s/^(\\d+) (\\d+) (\\d+) (\\d+) (\\d+) *//;\n+    push @parent_counts, $1;\n+    push @base_counts, $2;\n+    push @merge_counts, $3;\n+    push @path_counts, $4;\n+    push @res_counts, $5;\n+    if ($type eq 'M') {\n+\t/M=(\\d+) A=(\\d+) D=(\\d+) C=(\\d+) U=(\\d+)/ or die;\n+\tpush @merge_m, $1;\n+\tpush @merge_a, $2;\n+\tpush @merge_d, $3;\n+\tpush @merge_c, $4;\n+\tpush @merge_u, $5;\n+    }\n+}\n+\n+my $total = ($failures+$patches+$merges+$trivials);\n+print \"Total commit objects: $total\\n\";\n+printf \"Trivial Merges: $trivials (%.2f%%)\\n\", ($trivials * 100.0/$total);\n+printf \"Merges: $merges (%.2f%%)\\n\", ($merges * 100.0/$total);\n+if ($failures) {\n+    print \"Failures: $failures\\n\";\n+}\n+\n+avg_median(\\@patch_paths, \"Number of paths touched by non-merge commits\");\n+avg_median(\\@parent_counts, \"Number of merge parents\");\n+avg_median(\\@base_counts, \"Number of merge bases\");\n+avg_median(\\@merge_counts, \"File level merges\");\n+avg_median(\\@path_counts, \"Number of changed paths from the first parent\");\n+#avg_median(\\@res_counts, \"\");\n+avg_median(\\@merge_m, \"File level 3-ways\");\n+avg_median(\\@merge_a, \"Paths added\");\n+avg_median(\\@merge_d, \"Paths deleted\");\n+avg_median(\\@merge_c, \"Paths identically added with wrong permission\");\n+avg_median(\\@merge_u, \"Paths added differently\");\n+\n+\ndiff --git a/contrib/jc-git-stat-1-mof.sh b/contrib/jc-git-stat-1-mof.sh\nnew file mode 100755\nindex 0000000..2be6d8b\n--- /dev/null\n+++ b/contrib/jc-git-stat-1-mof.sh\n@@ -0,0 +1,59 @@\n+#!/bin/sh\n+#\n+# Copyright (c) Linus Torvalds, 2005\n+# Copyright (c) Junio C Hamano, 2005\n+#\n+# This is modified from the git per-file merge script, called with\n+#\n+#   $1 - original file SHA1 (or empty)\n+#   $2 - file in branch1 SHA1 (or empty)\n+#   $3 - file in branch2 SHA1 (or empty)\n+#   $4 - pathname in repository\n+#   $5 - orignal file mode (or empty)\n+#   $6 - file in branch1 mode (or empty)\n+#   $7 - file in branch2 mode (or empty)\n+#\n+# Handle some trivial cases.. The _really_ trivial cases have\n+# been handled already by git-read-tree, but that one doesn't\n+# do any merges that might change the tree layout.\n+\n+case \"${1:-.}${2:-.}${3:-.}\" in\n+#\n+# Deleted in both or deleted in one and unchanged in the other\n+#\n+\"$1..\" | \"$1.$1\" | \"$1$1.\")\n+\techo D\n+\t;;\n+\n+#\n+# Added in one.\n+#\n+\".$2.\" | \"..$3\" )\n+\techo A\n+\t;;\n+\n+#\n+# Added in both (check for same permissions).\n+#\n+\".$3$2\")\n+\tif [ \"$6\" != \"$7\" ]; then\n+\t\techo C\n+\telse\n+\t\techo A\n+\tfi\n+\t;;\n+\n+#\n+# Modified in both, but differently.\n+#\n+\"$1$2$3\")\n+\techo M\n+\t;;\n+\n+\".$2$3\")\n+\techo U\n+\t;;\n+*)\n+\techo C\n+\t;;\n+esac\ndiff --git a/contrib/jc-git-stat-1.sh b/contrib/jc-git-stat-1.sh\nnew file mode 100755\nindex 0000000..b03c69d\n--- /dev/null\n+++ b/contrib/jc-git-stat-1.sh\n@@ -0,0 +1,74 @@\n+#!/bin/sh\n+\n+MOF=`dirname \"$0\"`/jc-git-stat-1-mof.sh\n+\n+GIT_INDEX_FILE=.tmp-index\n+export GIT_INDEX_FILE\n+LF='\n+'\n+\n+check_merge () {\n+\trm -f $GIT_INDEX_FILE\n+\tcommit=$1\n+\tshift\n+\n+\tcase \"$#\" in\n+\t2)\n+\t\tMB=$(git-merge-base --all \"$@\")\n+\t\t;;\n+\t*)\n+\t\tMB=$(git-show-branch --merge-base \"$@\")\n+\t\t;;\n+\tesac\n+\tbasecnt=$(echo $MB | wc -l)\n+\n+\tif git-read-tree --trivial -m $MB \"$@\" 2>/dev/null\n+\tthen\n+\t\ttype=T\n+\t\tpathcnt=$(git-diff-index --cached --name-status \"$1\" | wc -l)\n+\t\trescnt=$(git-diff-index --cached --name-status \"$commit\" | wc -l)\n+\t\techo \"T $commit $# $basecnt 0 $pathcnt $rescnt\"\n+\telif git-read-tree -m $MB \"$@\" 2>/dev/null\n+\tthen\n+\t        script='s/^ *\\([0-9]*\\) *\\([A-Z]\\)/\\2=\\1/'\n+\t\ttype=M\n+\t\tmergecnt=$(git-ls-files --unmerged | sort -k 4,4 -u | wc -l)\n+\t\tpathcnt=$(git-diff-index --cached --name-status \"$1\" | wc -l)\n+\n+\t\tC=0 A=0 M=0 U=0 D=0\n+\t\teval `git-merge-index -o \"$MOF\" -a |\n+\t\t\tsort |\n+\t\t\tuniq -c |\n+\t\t\tsed -e \"$script\"`\n+\t\trescnt=$(git-diff-index --cached --name-status \"$commit\" | wc -l)\n+\t\techo \"M $commit $# $basecnt $mergecnt $pathcnt $rescnt M=$M A=$A D=$D C=$C U=$U\"\n+\telse\n+\t\techo \"F $commit $# $basecnt\"\n+\tfi\n+}\n+\n+check_patch () {\n+\tpathcnt=$(git-diff-tree --name-status -r \"$1\" | wc -l)\n+\techo \"C $1 $pathcnt\"\n+}\n+\n+case \"$#\" in\n+0)\n+\tset ORIG_HEAD.. ;;\n+esac\n+\n+git-rev-list --parents \"$@\" |\n+while read commit parents\n+do\n+\tcase \"$parents\" in\n+\t?*' '?*)\n+\t\t# Merge\n+\t\tcheck_merge $commit $parents\n+\t\t;;\n+\t*)\n+\t\t# Change\n+\t\tcheck_patch $commit $parents\n+\t\t;;\n+\tesac\n+done\n+\n---\n0.99.9.GIT\n"},{"id":"11666","messageId":"46a038f90511120419v70166c60t93d58b7544e03e3b@mail.gmail.com","threadId":"2386","inReplyTo":"7v7jbeia3v.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2005-11-12T12:19:45Z","receivedAt":"2005-11-12T12:19:45Z","isPatch":true,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 11/12/05, Junio C Hamano <junkio@cox.net> wrote:\n> Ryan Anderson <ryan@michonline.com> writes:\n>\n> > Junio C Hamano wrote:\n> >\n> >> Just for fun, I randomly picked two heads/master commits from\n> >> linux-2.6 repository ... and fed the commits\n> >> between the two to a little script that looks at commits and\n> >> tries to stat what they did (the script ignores renames so they\n> >> appear as deletes and adds).\n\nRelated to this, I've been wondering whether it'd be possible to teach\ngit to rebase local patches, even if that means rewriting local\nhistory. When you are dealing with team shared repo, the sequences of\npull/push end up being quite messy, full of little meaningless merges.\nSimilarly, when dealing with an upstream, my tree gets slowly out of\nsync and slightly messy. Eventually I get a new checkout, and rebase\nany pending patches with git-format-patch and git-am.\n\nThe same process would be much easier if I could just cg-update from\nthe repo and get it to try and actually rebase my local commits --\nrewriting history as if I had committed them after the update. Of\ncourse, it'd be cheating... but we cheat all the time anyway, we only\nsweat harder at it ;-)\n\ncheers,\n\n\nmartin\n"},{"id":"11673","messageId":"20051112125331.GB30496@pasky.or.cz","threadId":"2386","inReplyTo":"46a038f90511120419v70166c60t93d58b7544e03e3b@mail.gmail.com","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2005-11-12T12:53:31Z","receivedAt":"2005-11-12T12:53:31Z","isPatch":true,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Nov 12, 2005 at 01:19:45PM CET, I got a letter\nwhere Martin Langhoff <martin.langhoff@gmail.com> said that...\n> The same process would be much easier if I could just cg-update from\n> the repo and get it to try and actually rebase my local commits --\n> rewriting history as if I had committed them after the update. Of\n> course, it'd be cheating... but we cheat all the time anyway, we only\n> sweat harder at it ;-)\n\nI'm a bit reluctant about this functionality available in cg-update, but\nthen people will start to want commit stack and stuff, while they should\nbe just already long using StGIT for tracking their patches.\n\nActually, I wanted to also implement e-mail functionality to cg-mkpatch,\nbut I'm not sure now - perhaps people wanting that should really just\nuse StGIT. Cogito or GIT core is not very suitable for keeping your\npatches against someone else's tree if he is not going to GIT-merge with\nyou, exactly because it's not really very convenient to update your\npatches.\n\nOn the same note, I would like StGIT to drop functionality not really\nbelonging to patch stack manager (stg add, stg rm, stg status, ...) so\nthat its commandset gets smaller and more focused - but before I would\nsuggest dropping stg status, cg-status must be able to do conflicts\ntracking, so I will dedicate another mail to this sometime in the\nfuture, with a more detailed proposal.\n\nSo, is there any reason why you want this in GIT/Cogito and don't want\nto use StGIT?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nVI has two modes: the one in which it beeps and the one in which\nit doesn't.\n"},{"id":"11689","messageId":"Pine.LNX.4.63.0511121947050.31652@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"2386","inReplyTo":"46a038f90511120419v70166c60t93d58b7544e03e3b@mail.gmail.com","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2005-11-12T19:04:24Z","receivedAt":"2005-11-12T19:04:24Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 13 Nov 2005, Martin Langhoff wrote:\n\n> [...] I've been wondering whether it'd be possible to teach git to \n> rebase local patches, even if that means rewriting local history. When \n> you are dealing with team shared repo, the sequences of pull/push end up \n> being quite messy, full of little meaningless merges. Similarly, when \n> dealing with an upstream, my tree gets slowly out of sync and slightly \n> messy. Eventually I get a new checkout, and rebase any pending patches \n> with git-format-patch and git-am.\n\nI thought about this problem for a while. You described a typical use case \n(in a small team with one central repo), where pushes are done only from a \ncleaned up branch, but the work branches stay.\n\nIf the tree corresponding to the central HEAD matches the tree \ncorresponding of your private merge branch at a given stage, a simple \ngraft should be your solution.\n\nExample: You pull and push from/to origin on a central server. Your \n(dirty) work branch is master. Now, at a given time, \norigin^{tree}==master~15^{tree}. You could then generate a graft which \ntells git that all parents of origin are also parents of master~15.\n\nIf at a given stage, origin^{tree}==master^{tree}, that graft would make \nyour next merge a fast forward, effectively cleaning up your branch \nwithout loosing your history.\n\nThe graft could be generated by a simple script.\n\nWould this help you?\n\nCiao,\nDscho\n"},{"id":"11723","messageId":"7vy83s95k0.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"46a038f90511120419v70166c60t93d58b7544e03e3b@mail.gmail.com","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-13T10:59:43Z","receivedAt":"2005-11-13T10:59:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com> writes:\n\n> Similarly, when dealing with an upstream, my tree gets slowly out of\n> sync and slightly messy. Eventually I get a new checkout, and rebase\n> any pending patches with git-format-patch and git-am.\n\nThe key is not to let your tree go \"slowly\" out of sync, from my\nexperience.  When Linus was the maintainer, I used to do the\nequivalent of the following all the time to keep up with his\ntree while keeping my history clean [*1*].\n\nFrequently [*2*], I tried to see if Linus made something new and\ninteresting.  My \"origin\" branch was always copy of Linus head.\n\n\t$ git fetch origin\n\nThe command would say \"Fast forward\".  So what did he do?\n\n\t$ git show-branch master origin\n        ! [origin] Separate LDFLAGS and CFLAGS.\n         * [master] Rename lost+found to lost-found.\n        --\n         + [master] Rename lost+found to lost-found.\n         + [master^] Fix compilation warnings in pack-redundant.c\n         + [master~2] Debian: build-depend on libexpat-dev.\n        +  [origin] Separate LDFLAGS and CFLAGS.\n        ++ [master~3] Split gitk into seperate RPM package\n\nAh, the commit master~3 was what he had the last time I pulled\nfrom him, and since then he made a commit while I did three.  I\ncould do \"git pull . origin\" at this point, but that would\nresult in a useless mini-merge.  My tree is not public so I can\nfreely rebase to clean things up.\n\n\t$ git rebase origin\n\t$ git show-branch\n        ! [origin] Separate LDFLAGS and CFLAGS.\n         * [master] Rename lost+found to lost-found.\n        --\n         + [master] Rename lost+found to lost-found.\n         + [master^] Fix compilation warnings in pack-redundant.c\n         + [master~2] Debian: build-depend on libexpat-dev.\n        ++ [origin] Separate LDFLAGS and CFLAGS.\n\nNow I am fast-forward, so I could ask him to pull from me [*3*].\n\nI think each of your developers can do the same, treating the\n\"project shared repository\" as \"Linus repository\" and pull that\ninto the \"origin\" branch, and when the \"master\" is ready, push\nit back into the shared repository (which is equivalent of Linus\npulling everything from me while doing nothing else in his\nrepository).\n\nFor a sizable change that deserves a topic branch with a long\nsequence of commits, rebasing is not always the optimum\nsolution; and you may want to keep the full merge history of\nsuch a branch pushed into the public repository as is.  But for\nsimpler cases that 'git rebase' can handle easily without\nconflicts, the above procedure would help you keeping the\nhistory of your shared repository less cluttered.\n\n[Footnotes]\n\n*1* Back then we did not have multi-head fetch, show-branch nor\nrebase, so I did these using a homebrew Porcelain.\n\n*2* Unlike CVS which always mucks with the working tree, 'git\nfetch' into a branch that is not current one is an operation and\ncan be done even when I am in the middle of a heavy hackery.\nBeing able to peek into what others are up even when your tree\nis in a messy state (the fetch is often followed by log and\ndiff) helps you to avoid doing duplicated work or going in a\nwrong direction, which was great.\n\n*3* Even back then almost all changes were fed via e-mail to\nthe maintainer.\n"},{"id":"11725","messageId":"20051113111130.GN30496@pasky.or.cz","threadId":"2386","inReplyTo":"46a038f90511120419v70166c60t93d58b7544e03e3b@mail.gmail.com","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2005-11-13T11:11:30Z","receivedAt":"2005-11-13T11:11:30Z","isPatch":true,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Nov 12, 2005 at 01:19:45PM CET, I got a letter\nwhere Martin Langhoff <martin.langhoff@gmail.com> said that...\n> Similarly, when dealing with an upstream, my tree gets slowly out of\n> sync and slightly messy.\n\nI've been replying only to this, but missed the team shared repo case,\nwhere StGIT obviously does not make much sense.\n\n> Related to this, I've been wondering whether it'd be possible to teach\n> git to rebase local patches, even if that means rewriting local\n> history. When you are dealing with team shared repo, the sequences of\n> pull/push end up being quite messy, full of little meaningless merges.\n\nWell, only pulls make up for merges. But what's wrong with that? This is\njust what you get when using distributed VCS, and what's the point in\nskewing the history to look different than how did it happen? I'd just\nget used to it. :-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nVI has two modes: the one in which it beeps and the one in which\nit doesn't.\n"},{"id":"11741","messageId":"46a038f90511131242p4692c74fn20c015998620b9f4@mail.gmail.com","threadId":"2386","inReplyTo":"7vy83s95k0.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2005-11-13T20:42:31Z","receivedAt":"2005-11-13T20:42:31Z","isPatch":true,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 11/13/05, Junio C Hamano <junkio@cox.net> wrote:\n> Ah, the commit master~3 was what he had the last time I pulled\n> from him, and since then he made a commit while I did three.  I\n> could do \"git pull . origin\" at this point, but that would\n> result in a useless mini-merge.  My tree is not public so I can\n> freely rebase to clean things up.\n>\n>         $ git rebase origin\n>         $ git show-branch\n\nWhat happens if there are conflicts during git-rebase? I'm thinking of\nadding an '-r' option to cg-update that will rebase instead of\nmerging, if the rebase is clean.\n\nIs there a cheap way to ask from a shell script whether the merge is\ntruly trivial? I thought git-diff-tree would help me here, but it\ndoesn't...\n\n\nmartin\n"},{"id":"11754","messageId":"7vlkzr6gzz.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"46a038f90511131242p4692c74fn20c015998620b9f4@mail.gmail.com","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-14T03:33:04Z","receivedAt":"2005-11-14T03:33:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com> writes:\n\n> On 11/13/05, Junio C Hamano <junkio@cox.net> wrote:\n>> ....  I\n>> could do \"git pull . origin\" at this point, but that would\n>> result in a useless mini-merge.  My tree is not public so I can\n>> freely rebase to clean things up.\n>>\n>>         $ git rebase origin\n>>         $ git show-branch\n>\n> What happens if there are conflicts during git-rebase?\n\nWell, obviously you could resolve them ;-).  But if you are\nrebasing just to reduce trivial mini-merges, it might make more\nsense to honestly record the merge if the rebase involves\nconflict resolution.  After all, the reason rebase got conflicts\nis because the development trail by somebody else that has been\nalready committed to the shared \"master\" branch overlapped what\nyou were doing in your \"master\" branch, isn't it?\n\nIn your message you indicated that you use \"format-patch\" piped\nto \"am\".  I think that is a better approach than \"rebase\" these\ndays; the conflict can be handled easier with that approach, and\nif you use \"--3way\" flag you do not even have to worry about\npatches in your branch that is already there in the shared\n\"master\" (your \"origin\") branch.\n\nSo instead of running \"git rebase origin\" at this point, I may\ndo something like this [*1*]:\n\n\t$ git-reset --hard origin\n        $ git-format-patch -k --stdout origin ORIG_HEAD | git am -3 -k\n\nThe first step rewinds my \"master\" (the original is stored in\nORIG_HEAD), and the second step extracts the commits that were\nin my master but not in origin in a patch form, an replay them\non top of the \"master\" (which was rewound to \"origin\").\n\n\"git-am\" would stop at the first unapplicable patch if there is\na conflict, leaving the conflicting patch in .dotest/patch.\nI have to fix it up before going further.  Here is how.\n\n1. \"git am\" 3-way fallback would have kicked in, because I have\n   all the blobs the patch is supposed to apply to, and my\n   working tree and index is in a state just like when I am\n   resolving a conflicting merge after a pull.  Clean up the\n   conflict in the working tree, build-test and all as usual.\n\n2. Run \"git diff HEAD >.dotest/patch\" to record what the patch\n   should have been if it were to apply cleanly on top of the\n   previous state.  If I did a noteworthy adjustment to the\n   patch, I might also edit .dotest/final-commit to update the\n   commit log message.\n\n3. Then reset the working tree and index before the failed\n   application of this patch with \"git reset --hard\".\n\nAfter that:\n\n\t$ git am -3\n\nwould let me restart from that commit that did not replay well.\n\n> Is there a cheap way to ask from a shell script whether the merge is\n> truly trivial? I thought git-diff-tree would help me here, but it\n> doesn't...\n\nThis was recently added by Linus to help git-merge do that:\n\n\tgit-read-tree --trivial -m -u $O $A $B\n\nThe command exits with a non-zero status, without touching index\nnor working tree, when the merge is not \"truly trivial\".\nOtherwise it does its thing -- the trivial in-index merge is\ndone, files in working tree updated and the only thing left for\nyou to do is to create a commit having parent $A and $B.\n\nWould that help?\n\n[Footnote]\n\n*1* This is what the \"make rebase restartable\" comment in TODO\nlist is about, and I wanted to rewrite \"rebase\" to do exactly\nthese two commands, but I got distracted ;-).\n"},{"id":"11755","messageId":"46a038f90511132001x6a9109fk17593b7ceaf3177e@mail.gmail.com","threadId":"2386","inReplyTo":"7vlkzr6gzz.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2005-11-14T04:01:43Z","receivedAt":"2005-11-14T04:01:43Z","isPatch":true,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 11/14/05, Junio C Hamano <junkio@cox.net> wrote:\n> In your message you indicated that you use \"format-patch\" piped\n> to \"am\".  I think that is a better approach than \"rebase\" these\n> days\n\nHmmm. But doesn't deal well with binary changes. We deal with a large\nset of projects, and while we don't manage that many binary files, it\nis just enough that I'll have to pass on only using format-patch.\n\n> This was recently added by Linus to help git-merge do that:\n>\n>         git-read-tree --trivial -m -u $O $A $B\n\nCool! Abusing that, perhaps I could teach git-rebase to take a\n'--trivial-only' flag, and then cg-update --rebase could try\ngit-rebase --trivial-only, and fall back on cg-merge...\n\ncheers,\n\n\nmartin\n"},{"id":"11759","messageId":"7vwtjb4vc4.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"46a038f90511132001x6a9109fk17593b7ceaf3177e@mail.gmail.com","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-14T06:06:19Z","receivedAt":"2005-11-14T06:06:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com> writes:\n\n> On 11/14/05, Junio C Hamano <junkio@cox.net> wrote:\n>> In your message you indicated that you use \"format-patch\" piped\n>> to \"am\".  I think that is a better approach than \"rebase\" these\n>> days\n>\n> Hmmm. But doesn't deal well with binary changes. We deal with a large\n> set of projects, and while we don't manage that many binary files, it\n> is just enough that I'll have to pass on only using format-patch.\n\nIt shouldn't be too tricky to enhance \"git am\" (git-apply called\nat around line 49 in it) to grok binary differences for this\npurpose, because you would have both pre- and post-image blob in\nyour object database, because the patch is being used only to\nreplay what you have in your reository, and it records their\nabbreviated SHA1 name.\n\nI've never felt need to \"merge\" the binary files myself and had\nnever got around doing this, but if you are interested, it would\ngo something like this:\n\n. Make \"index %.7s..%.7s\" that abbreviates pre- and post- image\n  blob SHA1s in diff.c configurable to spit out full 40 bytes.\n  Call that option --full-index-sha1.\n\n. The updated \"git rebase\" that uses the \"format-patch | am\" I\n  outlined would pass the --full-index-sha1 to format-patch\n  (which is pased onto underlying diff-tree -p).\n\n. In apply.c, check if all of the following holds:\n\n    * we have both the full 40-byte old_sha1_prefix[] and\n      new_sha1_prefix[]; and\n\n    * what the index records matches old_sha1_prefix[]; and\n\n    * the new blob is found in the object database;\n\n  and for such a path:\n\n    * change parse_chunk() not to barf even on a binary patch.\n\n    * change apply_data() to just declare the patch application\n      result is the new blob recorded in the patch.\n\nHmm.\n"},{"id":"11764","messageId":"46a038f90511140051o1fa5ef7cyb9dd723fb8161ef9@mail.gmail.com","threadId":"2386","inReplyTo":"7vwtjb4vc4.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2005-11-14T08:51:29Z","receivedAt":"2005-11-14T08:51:29Z","isPatch":true,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 11/14/05, Junio C Hamano <junkio@cox.net> wrote:\n> It shouldn't be too tricky to enhance \"git am\" (git-apply called\n> at around line 49 in it) to grok binary differences for this\n> purpose, because you would have both pre- and post-image blob in\n> your object database, because the patch is being used only to\n> replay what you have in your reository, and it records their\n> abbreviated SHA1 name.\n\nI'm curious. What would be the advantages of this over git-read-tree\n-m for use within a single repo? I keep thinking that I need an\nintra-repo way of doing it (arguably faster and more reliable),\ninstead of git-format-patch|git-am, which is less reliable and slower.\n\nOTOH,  if this is heading towards teaching git-am how to apply changes\nto binary files based on known SHA1s, this will give birth to a type\nof patch that applies only if you have the objects beforehand. Is that\nenough to get by? Perhaps we need a format to fully describe binary\nfiles?\n\n> I've never felt need to \"merge\" the binary files myself and had\n> never got around doing this, but if you are interested, it would\n> go something like this:\n\nThat's quite a bit of C hacking... I'm game for all the Perl and shell\nscripts in git, but I know better than start learning C in _this_\nproject with you, Linus and the whole list watching me make the fool\n;-)\n\nSo noone hacks this bit of C you are proposing I'll eventually get\nsomething done with cg-update and git-rebase.\n\nIn the meantime, the user experience of working with a small team and\na shared repo has improved significantly by switching from cg-update\nto cg-fetch && git-rebase origin. So much so that I suspect that it'd\nbe a big win for cg-update to default to rebase on merges from\n'origin'.\n\ncheers,\n\n\nmartin\n"},{"id":"11768","messageId":"20051114092554.GR30496@pasky.or.cz","threadId":"2386","inReplyTo":"46a038f90511140051o1fa5ef7cyb9dd723fb8161ef9@mail.gmail.com","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2005-11-14T09:25:54Z","receivedAt":"2005-11-14T09:25:54Z","isPatch":true,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Mon, Nov 14, 2005 at 09:51:29AM CET, I got a letter\nwhere Martin Langhoff <martin.langhoff@gmail.com> said that...\n> In the meantime, the user experience of working with a small team and\n> a shared repo has improved significantly by switching from cg-update\n> to cg-fetch && git-rebase origin. So much so that I suspect that it'd\n> be a big win for cg-update to default to rebase on merges from\n> 'origin'.\n\nI still don't understand what's the point. And it is confusing for\nthe user, since the history suddenly doesn't reflect how the GIT\ndevelopment happened (which is what the history is about), and it is\ndownright deadly as soon as more than one branch gets around, which\nmakes it even more confusing for the user (you can do cg-update, yeah\n- oh wait, unless you do X, then you need to remember to always do\ncg-update -R or whatever).\n\nAll the woes just to get rid of the merge commits. What's wrong on merge\ncommits? If they irritate you so much, cg-log -M. ;-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nVI has two modes: the one in which it beeps and the one in which\nit doesn't.\n"},{"id":"11769","messageId":"7vd5l3zij7.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"46a038f90511140051o1fa5ef7cyb9dd723fb8161ef9@mail.gmail.com","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-14T09:27:08Z","receivedAt":"2005-11-14T09:27:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com> writes:\n\n> I'm curious. What would be the advantages of this over git-read-tree\n> -m for use within a single repo?\n\nOn a large tree I had an impression that applying patch and then\nfalling back on 3-way is faster, but other than that, nothing,\nreally.  Just being able to use a single mechanism, which does\nnot buy us much.\n\n> OTOH,  if this is heading towards teaching git-am how to apply changes\n> to binary files based on known SHA1s, this will give birth to a type\n> of patch that applies only if you have the objects beforehand. Is that\n> enough to get by? Perhaps we need a format to fully describe binary\n> files?\n\nI'd rather not see our \"patch\" go in the direction of recording\nboth pre- and post-image of blob for binary files, which is what\nwe would end up doing if we really want to do binary flexibly.\n\nWell, that may be nice as an option, but not by default.\n\nAn option halfway in between would be to record the pre-image\nSHA1 and post- blob, perhaps compressed-uuencoded.  This would\nlimit us to the case that the recipient has not touched the\nbinary file and replacing it, but in practice that might be\nenough.\n\nI'm willing to do the C hackery myself if there is enough\ninterest in it.\n"},{"id":"11842","messageId":"46a038f90511141325p23531982kf6a230929bf6d202@mail.gmail.com","threadId":"2386","inReplyTo":"20051114092554.GR30496@pasky.or.cz","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2005-11-14T21:25:35Z","receivedAt":"2005-11-14T21:25:35Z","isPatch":true,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 11/14/05, Petr Baudis <pasky@suse.cz> wrote:\n> All the woes just to get rid of the merge commits. What's wrong on merge\n> commits? If they irritate you so much, cg-log -M. ;-)\n\nWell, if you have a team of 5 working closely, doing commit/update\nseveral times a day, soon the following things happen:\n\n - everyone has sightly different histories, even if they all have the same tree\n - in the repo, there are 'update' merges galore.\n - as soon as you _actually_ have branches, it's really hard to\ndistinguish branch merges from 'same branch update before commit'\nmerges.\n\nby replacing the daily/hourly intra-team, self-branch `cg-update` with\n`cg-fetch && git-rebase` our history makes much more sense _and_\nlooking at gitk I can see clearly again the interesting merges (ie:\nmerges with other branches).\n\nIn practice, it's exactly what happens with git's history -- the\nmerges are actually from rebased patches, that's why the history is\n\"readable\".\n\ndoes that make sense?\n\ncheers,\n\n\nmartin\n"},{"id":"11864","messageId":"7vy83qty2v.fsf@assigned-by-dhcp.cox.net","threadId":"2386","inReplyTo":"46a038f90511140051o1fa5ef7cyb9dd723fb8161ef9@mail.gmail.com","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-15T03:00:08Z","receivedAt":"2005-11-15T03:00:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com> writes:\n\n>> I've never felt need to \"merge\" the binary files myself and had\n>> never got around doing this, but if you are interested, it would\n>> go something like this:\n>\n> That's quite a bit of C hacking...\n\nI'll be pushing out something along the lines I said in the\nproposed updates branch tonight.  I've even run 1 (one) test\n;-).  The relevant commits are these:\n\n+  [pu~5^2] rebase: make it usable for binary files as well.\n+  [pu~5^2^] diff: --full-index\n+  [pu~5^2~2] apply: allow-binary-replacement.\n+  [pu~5^2~3] Rewrite rebase to use git-format-patch piped to git-am.\n\nNote that this does not make generated \"binary patch\" usable\nacross repositories yet.  Since our patches are reversible out\nof principle, making it usable across repositories involves\nrecording both the pre- and post-image of binary blobs in the\npatch output, which may be a useful option in some cases but not\nnecessary for the immediate application of intra-repo rebasing.\n"},{"id":"11883","messageId":"b0943d9e0511150204h25417993l@mail.gmail.com","threadId":"2386","inReplyTo":"20051112125331.GB30496@pasky.or.cz","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Catalin Marinas","fromEmail":"catalin.marinas@gmail.com","sentAt":"2005-11-15T10:04:32Z","receivedAt":"2005-11-15T10:04:32Z","isPatch":true,"sender":{"key":"catalin.marinas@gmail.com","avatar":null},"body":"On 12/11/05, Petr Baudis <pasky@suse.cz> wrote:\n> On the same note, I would like StGIT to drop functionality not really\n> belonging to patch stack manager (stg add, stg rm, stg status, ...) so\n> that its commandset gets smaller and more focused\n\nThis was the case with the first StGIT implementations but I slowly\nbegan to want to only use StGIT and not switch to something else for\ntrivial SCM operations. I eventually added 'stg commit' which stores\nthe patches permanently into the base of the stack to enable some kind\nof maintainer mode for StGIT. My main use for this was to import\npatches directly into the main branch and not keep a separate one and\npull between them.\n\n> - but before I would\n> suggest dropping stg status, cg-status must be able to do conflicts\n> tracking, so I will dedicate another mail to this sometime in the\n> future, with a more detailed proposal.\n\nThe gitmergeonefile.py script in StGIT adds every conflict to the\n.git/conflicts file which is read by 'stg status'. My goal is not to\nleave any unmerged entried in the index even if there are conflicts.\nMaybe this could be changed and .git/conflicts file avoided entirely.\n\nAnyway, while I'll try not to add more SCM functionality to StGIT, I\ndon't think I should remove the existing add/rm/status functionality.\nIt's just handy not to use a different command when you want a new\nfile added to a patch.\n\n--\nCatalin\n"},{"id":"11898","messageId":"4379FEE1.9080407@citi.umich.edu","threadId":"2386","inReplyTo":"b0943d9e0511150204h25417993l@mail.gmail.com","subject":"Re: [PATCH] GIT commit statistics.","fromName":"Chuck Lever","fromEmail":"cel@citi.umich.edu","sentAt":"2005-11-15T15:29:37Z","receivedAt":"2005-11-15T15:29:37Z","isPatch":true,"sender":{"key":"cel@citi.umich.edu","avatar":null},"body":"Catalin Marinas wrote:\n> On 12/11/05, Petr Baudis <pasky@suse.cz> wrote:\n> \n>>On the same note, I would like StGIT to drop functionality not really\n>>belonging to patch stack manager (stg add, stg rm, stg status, ...) so\n>>that its commandset gets smaller and more focused\n> \n> \n> This was the case with the first StGIT implementations but I slowly\n> began to want to only use StGIT and not switch to something else for\n> trivial SCM operations. I eventually added 'stg commit' which stores\n> the patches permanently into the base of the stack to enable some kind\n> of maintainer mode for StGIT. My main use for this was to import\n> patches directly into the main branch and not keep a separate one and\n> pull between them.\n\n  ...\n\n> Anyway, while I'll try not to add more SCM functionality to StGIT, I\n> don't think I should remove the existing add/rm/status functionality.\n> It's just handy not to use a different command when you want a new\n> file added to a patch.\n\npetr,\n\ncurrently it isn't recommended to use StGIT with other porcelains.\n\nso either: make it completely safe to use StGIT with other porcelains, \nor add a minimal amount of SCM-like functionality so users don't miss it \nand try to use other porcelains and trash their repositories.\n\noverall i agree with catalin-- it's just easier to use a single tool.\n\n\nbegin:vcard\nfn:Chuck Lever\nn:Lever;Charles\norg:Network Appliance, Incorporated;Linux NFS Client Development\nadr:535 West William Street, Suite 3100;;Center for Information Technology Integration;Ann Arbor;MI;48103-4943;USA\nemail;internet:cel@citi.umich.edu\ntitle:Member of Technical Staff\ntel;work:+1 734 763-4415\ntel;fax:+1 734 763 4434\ntel;home:+1 734 668-1089\nx-mozilla-html:FALSE\nurl:http://www.monkey.org/~cel/\nversion:2.1\nend:vcard\n\n"}]}