{"thread":{"id":"9","subject":"Merge with git-pasky II.","startedAt":"2005-04-13T21:25:46Z","lastAt":"2005-04-18T07:42:32Z","messageCount":130,"participants":["Petr Baudis","Christopher Li","Linus Torvalds","Paul Jackson","Junio C Hamano","Barry Silverman","Erik van Konijnenburg","David Woodhouse","Ingo Molnar","Johannes Schindelin","Theodore Ts'o","C. Scott Ananian","Daniel Barkalow","Simon Fowler","David Lang","Sanjoy Mahajan","Brad Roberts","Herbert Xu","Kenneth Johansson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"29","messageId":"20050413212546.GA17236@64m.dyndns.org","threadId":"9","inReplyTo":"20050414002902.GU25711@pasky.ji.cz","subject":"Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-13T21:25:46Z","receivedAt":"2005-04-13T21:25:46Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"While you are there, do you mind to move the shell script\nto a sub directory? Let's try how rename works.\n\nChris\n\nOn Thu, Apr 14, 2005 at 02:29:02AM +0200, Petr Baudis wrote:\n>   Hello Linus,\n> \n>   I think my tree should be ready for merging with you. It is the final\n> tree and I've already switched my main branch for it, so it's what\n> people doing git pull are getting for some time already.\n> \n>   Its main contents are all of my shell scripts. Apart of that, some\n> tiny fixes scattered all around can be found there, as well as some\n> patches which went through the mailing list. My last merge with you\n> concerned your commit 39021759c903a943a33a28cfbd5070d36d851581.\n> \n>   It's again\n> \n> \trsync://pasky.or.cz/git/\n> \n> this time my HEAD is fba83970090ef54c6eb86dcc2c2d5087af5ac637.\n> \n>   Note that my rsync tree still contains even my old branch; I thought\n> I'd leave it around in the public objects database for some time, shall\n> anyone want to have a look at the history of some of the scripts. But if\n> you want it gone, tell me and I will prune it (and perhaps offer it in\n> /git-old/ or whatever). I'm using the following:\n> \n> \tfsck-cache --unreachable $(commit-id) | grep unreachable \\\n> \t\t| cut -d ' ' -f 2 | sed 's/^\\(..\\)/.git\\/objects\\/\\1\\//' \\\n> \t\t| xargs rm\n> \n>   Thanks,\n> \n> -- \n> \t\t\t\tPetr \"Pasky\" Baudis\n> Stuff: http://pasky.or.cz/\n> C++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"31","messageId":"20050413220053.GB17236@64m.dyndns.org","threadId":"9","inReplyTo":"20050414004504.GW25711@pasky.ji.cz","subject":"Re: Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-13T22:00:53Z","receivedAt":"2005-04-13T22:00:53Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"\nOn Thu, Apr 14, 2005 at 02:45:04AM +0200, Petr Baudis wrote:\n> Dear diary, on Wed, Apr 13, 2005 at 11:25:46PM CEST, I got a letter\n> where Christopher Li <git@chrisli.org> told me that...\n> Well, unless Linus will want me otherwise, I'd like to postpone this\n> until I'm finally done with the damn merge - enough things already got\n> into my way today, so I would really like to focus on this tomorrow. So\n> I'll be probably merging only (or mostly) bugfixes until I have that\n> finished.\n\nSure, whenever you are ready.\n\n> P.S.: Just staring at\n> http://www.theregister.co.uk/2005/04/11/torvalds_attack/ ... I'm nothing\n> like a regular reader of (R), but I thought the guys have at least a bit\n> of sense. Duh. :/ Or is April 11 now yet another joke day after April 1?\n\nWhatever, is the news site. They never mention git though.\n\nChris\n\n"},{"id":"27","messageId":"20050414002902.GU25711@pasky.ji.cz","threadId":"9","inReplyTo":null,"subject":"Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-14T00:29:02Z","receivedAt":"2005-04-14T00:29:02Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"  Hello Linus,\n\n  I think my tree should be ready for merging with you. It is the final\ntree and I've already switched my main branch for it, so it's what\npeople doing git pull are getting for some time already.\n\n  Its main contents are all of my shell scripts. Apart of that, some\ntiny fixes scattered all around can be found there, as well as some\npatches which went through the mailing list. My last merge with you\nconcerned your commit 39021759c903a943a33a28cfbd5070d36d851581.\n\n  It's again\n\n\trsync://pasky.or.cz/git/\n\nthis time my HEAD is fba83970090ef54c6eb86dcc2c2d5087af5ac637.\n\n  Note that my rsync tree still contains even my old branch; I thought\nI'd leave it around in the public objects database for some time, shall\nanyone want to have a look at the history of some of the scripts. But if\nyou want it gone, tell me and I will prune it (and perhaps offer it in\n/git-old/ or whatever). I'm using the following:\n\n\tfsck-cache --unreachable $(commit-id) | grep unreachable \\\n\t\t| cut -d ' ' -f 2 | sed 's/^\\(..\\)/.git\\/objects\\/\\1\\//' \\\n\t\t| xargs rm\n\n  Thanks,\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"28","messageId":"20050414003045.GV25711@pasky.ji.cz","threadId":"9","inReplyTo":"20050414002902.GU25711@pasky.ji.cz","subject":"Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-14T00:30:45Z","receivedAt":"2005-04-14T00:30:45Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, Apr 14, 2005 at 02:29:02AM CEST, I got a letter\nwhere Petr Baudis <pasky@ucw.cz> told me that...\n>   Its main contents are all of my shell scripts. Apart of that, some\n> tiny fixes scattered all around can be found there, as well as some\n> patches which went through the mailing list. My last merge with you\n> concerned your commit 39021759c903a943a33a28cfbd5070d36d851581.\n> \n>   It's again\n> \n> \trsync://pasky.or.cz/git/\n> \n> this time my HEAD is fba83970090ef54c6eb86dcc2c2d5087af5ac637.\n\nI forgot to add that after merging, you will probably want to change the\nVERSION file (to contain whatever you want).\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"30","messageId":"20050414004504.GW25711@pasky.ji.cz","threadId":"9","inReplyTo":"20050413212546.GA17236@64m.dyndns.org","subject":"Re: Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-14T00:45:04Z","receivedAt":"2005-04-14T00:45:04Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Wed, Apr 13, 2005 at 11:25:46PM CEST, I got a letter\nwhere Christopher Li <git@chrisli.org> told me that...\n> While you are there, do you mind to move the shell script\n> to a sub directory? Let's try how rename works.\n\nWell, unless Linus will want me otherwise, I'd like to postpone this\nuntil I'm finally done with the damn merge - enough things already got\ninto my way today, so I would really like to focus on this tomorrow. So\nI'll be probably merging only (or mostly) bugfixes until I have that\nfinished.\n\nP.S.: Just staring at\nhttp://www.theregister.co.uk/2005/04/11/torvalds_attack/ ... I'm nothing\nlike a regular reader of (R), but I thought the guys have at least a bit\nof sense. Duh. :/ Or is April 11 now yet another joke day after April 1?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"36","messageId":"20050414012352.GA17700@64m.dyndns.org","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504132020550.7211@ppc970.osdl.org","subject":"Re: Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-14T01:23:52Z","receivedAt":"2005-04-14T01:23:52Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Wed, Apr 13, 2005 at 08:51:50PM -0700, Linus Torvalds wrote:\n> \n> \n> On Thu, 14 Apr 2005, Petr Baudis wrote:\n> \n> Thick skin is the name of the game. I'd not get any work done otherwise.\n> \n> On that note - I've been avoiding doing the merge-tree thing, in the hope \n> that somebody else does what I've described. I really do suck at scripting \n> things, yet this is clearly something where using C to do a lot of the \n> stuff is pointless.\n> \n> Almost all the parts do seem to be there, ie Daniel did the \"common \n> parent\" part, and the rest really does seem to be more about scripting \n> than writing more C plumbing stuff..\n\nDo you have preference about what language of script we used? I actually\nhesitated to introduce my Python script to git.\n\nI can build some script extension for git just like the one I did for\nsparse, is that some thing you want to see? \n\nChris\n\n"},{"id":"38","messageId":"20050414021602.GA18655@64m.dyndns.org","threadId":"9","inReplyTo":"20050413220341.13e5ce0f.pj@engr.sgi.com","subject":"Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-14T02:16:02Z","receivedAt":"2005-04-14T02:16:02Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Wed, Apr 13, 2005 at 10:03:41PM -0700, Paul Jackson wrote:\n> \n> If you have a thin skin or tend to annoy others with a bit too much\n> attitude or can't pass up a good language war (which is my failing, and\n> why I am responding to a discussion that I've not been involved in for\n> days) then the resulting flamage could be distracting.\n\nOh, my bad. I am not trying to start a language war here.\nThat is why I am hesitated about Python.\nJust try to find out the acceptability. No pushing.\n\nChris \n\n"},{"id":"33","messageId":"Pine.LNX.4.58.0504132020550.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"20050414004504.GW25711@pasky.ji.cz","subject":"Re: Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-14T03:51:50Z","receivedAt":"2005-04-14T03:51:50Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 14 Apr 2005, Petr Baudis wrote:\n>\n> http://www.theregister.co.uk/2005/04/11/torvalds_attack/ ... I'm nothing\n> like a regular reader of (R), but I thought the guys have at least a bit\n> of sense. Duh. :/ Or is April 11 now yet another joke day after April 1?\n\nI actually _am_ a fairly regular reader, and hey, being opinionated and a \nbit over the top is what makes the site worthwhile. It's obviously what \nmotivates the people. \n\nAnd then, occasionally, when they bite you, hey, that's the price of\nhaving a high profile. I worry more about sometimes not listening to\ncritics than I do about the critics themselves.\n\nThick skin is the name of the game. I'd not get any work done otherwise.\n\nOn that note - I've been avoiding doing the merge-tree thing, in the hope \nthat somebody else does what I've described. I really do suck at scripting \nthings, yet this is clearly something where using C to do a lot of the \nstuff is pointless.\n\nAlmost all the parts do seem to be there, ie Daniel did the \"common \nparent\" part, and the rest really does seem to be more about scripting \nthan writing more C plumbing stuff..\n\n\t\tLinus\n"},{"id":"37","messageId":"20050413220341.13e5ce0f.pj@engr.sgi.com","threadId":"9","inReplyTo":"20050414012352.GA17700@64m.dyndns.org","subject":"Re: Merge with git-pasky II.","fromName":"Paul Jackson","fromEmail":"pj@engr.sgi.com","sentAt":"2005-04-14T05:03:41Z","receivedAt":"2005-04-14T05:03:41Z","isPatch":false,"sender":{"key":"pj@engr.sgi.com","avatar":null},"body":"> Do you have preference about what language of script we used? \n\nDo you have a thick skin? <grin>\n\nCan you easily ignore language wars with an amused wave of the hand\nand a happy chuckle at the oh so predictable weaknesses of humans?\n\nThen I'd wager it will be fine.\n\nIf you have a thin skin or tend to annoy others with a bit too much\nattitude or can't pass up a good language war (which is my failing, and\nwhy I am responding to a discussion that I've not been involved in for\ndays) then the resulting flamage could be distracting.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"42","messageId":"20050413231652.62be96d8.pj@engr.sgi.com","threadId":"9","inReplyTo":"20050414021602.GA18655@64m.dyndns.org","subject":"Re: Merge with git-pasky II.","fromName":"Paul Jackson","fromEmail":"pj@engr.sgi.com","sentAt":"2005-04-14T06:16:52Z","receivedAt":"2005-04-14T06:16:52Z","isPatch":false,"sender":{"key":"pj@engr.sgi.com","avatar":null},"body":"> Oh, my bad. I am not trying to start a language war here.\n\nNeither am I - no problem what so ever. <chuckle ...>\n\nBesides, I think we'd be on the same side.\n\nMy point was only a gentle one -- as is often the case when dealing with\nthe strange species called human, whether or not you can get away with\nsomething is often a simple matter of ones attitude.\n\nThere is one Python script already in the kernel: scripts/show_delta.\nBut it's too small a sample to mean much.\n\nI think first thing is \"get it right.\"  Python is good for that\nin the hands of someone who enjoys coding in it.\n\nI wish you well.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"47","messageId":"7vfyxtsurd.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504132020550.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-14T07:05:42Z","receivedAt":"2005-04-14T07:05:42Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> On that note - I've been avoiding doing the merge-tree thing, in the hope \nLT> that somebody else does what I've described.\n\nI now have a Perl script that uses rev-tree, cat-file,\ndiff-tree, show-files (with one modification so that it can deal\nwith pathnames with embedded newlines), update-cache (with one\nmodification so that I can add an entry for a file that does not\nexist to the dircache) and merge (from RCS).  Quick and dirty.\n\nThe changes to show-files is to give it an optional '-z' flag,\nwhich chanegs record terminator to NUL character instead of LF.\n\nThe script git-merge.perl takes two head commits.  It basically\nfollows what you described as I remember ;-):\n\n 1. runs rev-tree with --edges to find the common anscestor.\n\n 2. creates a temporary directory \"./,,merge-temp\"; create a\n    symlink ./,,merge-temp/.git/objects that points at\n    .git/objects.\n\n 3. sets up dircache there, initially populated with this common\n    ancestor tree.  No files are checked out.  Just set up\n    .git/index and that's it.\n\n 4. runs diff-tree to find what has been changed in each head.\n\n 5. for each path involved:\n\n  5.0 if neither heads change it, leave it as is;\n  5.1 if only one head changes a path and the other does not, just\n      get the changed version;\n  5.2 if both heads change it, check all three out and run merge.\n\nIt does not currently commit.  You can go to ./,,merge-temp/ and\nsee show-diff to see the result of the merge.  Files added in\none head has already been run \"update-cache\" when the script\nends, but changed and merged files are not---dircache still has\nthe common ancestor view.  So show-diff you will be seeing may\nbe enormous and not very useful if two forks were done in the\ndistant past.  After reviewing the merge result, you can\nupdate-cache, write-tree and commit-tree as usual, but with one\ncaveat:  do not run \"show-files | xargs update-cache\" if you are\nrunning git-merge.perl without -f flag!\n\nBy default, git-merge.perl creates absolute minimum number of\nfiles in ./,,merge-temp---only the merged files are left there\nso that you can inspect them.  You will not see unmodified\nfiles nor files changed only by one side of the merge.\n\nIf you give '-o' (oneside checkout) flag to git-merge.perl, then\nthe files only one side of the merge changed are also checked\nout in ./,,merge-temp.  If you give '-f' (full checkout) flag to\ngit-merge.perl, then in addition to what '-o' checks out,\nunchanged files are checked out in ./,,merge-temp.  This default\nis geared towards a huge tree with small merges (favorite case\nof Linus, if I understand correctly).\n\nRunning 'show-diff' in such a sparsely populated merge result\ntree gives you huge results because recent show-diff shows diffs\nwith empty files.  I added a '-r' flag to show-diff, which\nsquelches diffs with empty files.\n\nAlso to implement 'changed only by one-side' without actually\nchecking the file out, I needed to add one option to\n'update-cache'.  --cacheinfo flag is used this way:\n\n    $ update-cache --cacheinfo mode sha1 path\n\nand adds the pathname with mode and sha1 to the .git/index\nwithout actually requiring you to have such a file there.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n\n---\n\n show-diff.c    |   11 ++-\n show-files.c   |   12 ++-\n update-cache.c |   25 +++++++\n git-merge.perl |  193 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 4 files changed, 234 insertions(+), 7 deletions(-)\n\n\nshow-diff.c:  a531ca4078525d1c8dcf84aae0bfa89fed6e5d96\n--- show-diff.c\n+++ show-diff.c\t2005-04-13 22:47:33.000000000 -0700\n@@ -58,15 +58,20 @@\n int main(int argc, char **argv)\n {\n \tint silent = 0;\n+\tint silent_on_nonexisting_files = 0;\n \tint entries = read_cache();\n \tint i;\n \n \twhile (argc-- > 1) {\n \t\tif (!strcmp(argv[1], \"-s\")) {\n-\t\t\tsilent = 1;\n+\t\t\tsilent_on_nonexisting_files = silent = 1;\n \t\t\tcontinue;\n \t\t}\n-\t\tusage(\"show-diff [-s]\");\n+\t\tif (!strcmp(argv[1], \"-r\")) {\n+\t\t\tsilent_on_nonexisting_files = 1;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tusage(\"show-diff [-s] [-r]\");\n \t}\n \n \tif (entries < 0) {\n@@ -83,7 +88,7 @@\n \n \t\tif (stat(ce->name, &st) < 0) {\n \t\t\tprintf(\"%s: %s\\n\", ce->name, strerror(errno));\n-\t\t\tif (errno == ENOENT && !silent)\n+\t\t\tif (errno == ENOENT && !silent_on_nonexisting_files)\n \t\t\t\tshow_diff_empty(ce);\n \t\t\tcontinue;\n \t\t}\nshow-files.c:  a9fa6767a418f870a34b39379f417bf37b17ee18\n--- show-files.c\n+++ show-files.c\t2005-04-13 21:18:40.000000000 -0700\n@@ -14,6 +14,7 @@\n static int show_cached = 0;\n static int show_others = 0;\n static int show_ignored = 0;\n+static int line_terminator = '\\n';\n \n static const char **dir;\n static int nr_dir;\n@@ -105,12 +106,12 @@\n \t}\n \tif (show_others) {\n \t\tfor (i = 0; i < nr_dir; i++)\n-\t\t\tprintf(\"%s\\n\", dir[i]);\n+\t\t\tprintf(\"%s%c\", dir[i], line_terminator);\n \t}\n \tif (show_cached) {\n \t\tfor (i = 0; i < active_nr; i++) {\n \t\t\tstruct cache_entry *ce = active_cache[i];\n-\t\t\tprintf(\"%s\\n\", ce->name);\n+\t\t\tprintf(\"%s%c\", ce->name, line_terminator);\n \t\t}\n \t}\n \tif (show_deleted) {\n@@ -119,7 +120,7 @@\n \t\t\tstruct stat st;\n \t\t\tif (!stat(ce->name, &st))\n \t\t\t\tcontinue;\n-\t\t\tprintf(\"%s\\n\", ce->name);\n+\t\t\tprintf(\"%s%c\", ce->name, line_terminator);\n \t\t}\n \t}\n \tif (show_ignored) {\n@@ -134,6 +135,11 @@\n \tfor (i = 1; i < argc; i++) {\n \t\tchar *arg = argv[i];\n \n+\t\tif (!strcmp(arg, \"-z\")) {\n+\t\t\tline_terminator = 0;\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tif (!strcmp(arg, \"--cached\")) {\n \t\t\tshow_cached = 1;\n \t\t\tcontinue;\nupdate-cache.c:  8f149d5a4ab60e030a0ab19fdb59b8ee2576ee71\n--- update-cache.c\n+++ update-cache.c\t2005-04-13 23:27:54.000000000 -0700\n@@ -203,6 +203,8 @@\n {\n \tint i, newfd, entries;\n \tint allow_options = 1;\n+\tconst char *sha1_force = NULL;\n+\tconst char *mode_force = NULL;\n \n \tnewfd = open(\".git/index.lock\", O_RDWR | O_CREAT | O_EXCL, 0600);\n \tif (newfd < 0)\n@@ -235,14 +237,35 @@\n \t\t\t\trefresh_cache();\n \t\t\t\tcontinue;\n \t\t\t}\n+\t\t\tif (!strcmp(path, \"--cacheinfo\")) {\n+\t\t\t\tmode_force = argv[++i];\n+\t\t\t\tsha1_force = argv[++i];\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tdie(\"unknown option %s\", path);\n \t\t}\n \t\tif (!verify_path(path)) {\n \t\t\tfprintf(stderr, \"Ignoring path %s\\n\", argv[i]);\n \t\t\tcontinue;\n \t\t}\n-\t\tif (add_file_to_cache(path))\n+\t\tif (sha1_force && mode_force) {\n+\t\t\tstruct cache_entry *ce;\n+\t\t\tint namelen = strlen(path);\n+\t\t\tint mode;\n+\t\t\tint size = cache_entry_size(namelen);\n+\t\t\tsscanf(mode_force, \"%o\", &mode);\n+\t\t\tce = malloc(size);\n+\t\t\tmemset(ce, 0, size);\n+\t\t\tmemcpy(ce->name, path, namelen);\n+\t\t\tce->namelen = namelen;\n+\t\t\tce->st_mode = mode;\n+\t\t\tget_sha1_hex(sha1_force, ce->sha1);\n+\n+\t\t\tadd_cache_entry(ce, 1);\n+\t\t}\n+\t\telse if (add_file_to_cache(path))\n \t\t\tdie(\"Unable to add %s to database\", path);\n+\t\tmode_force = sha1_force = NULL;\n \t}\n \tif (write_cache(newfd, active_cache, active_nr) ||\n \t    rename(\".git/index.lock\", \".git/index\"))\n\n--- /dev/null\t2005-03-19 15:28:25.000000000 -0800\n+++ git-merge.perl\t2005-04-13 23:45:23.000000000 -0700\n@@ -0,0 +1,193 @@\n+#!/usr/bin/perl -w\n+\n+use Getopt::Long;\n+\n+my $full_checkout = 0;\n+my $oneside_checkout = 0;\n+GetOptions(\"full\" => \\$full_checkout,\n+\t   \"oneside\" => \\$oneside_checkout)\n+    or die;\n+\n+if ($full_checkout) {\n+    $oneside_checkout = 1;\n+}\n+\n+sub read_rev_tree {\n+    my (@head) = @_;\n+    my ($fhi);\n+    open $fhi, '-|', 'rev-tree', '--edges', @head\n+\tor die \"$!: rev-tree --edges @head\";\n+    my $common;\n+    while (<$fhi>) {\n+\tchomp;\n+\t(undef, undef, $common) = split(/ /, $_);\n+\tif ($common =~ s/^([a-f0-f]{40}):\\d+$/$1/) {\n+\t    last;\n+\t}\n+    }\n+    close $fhi;\n+    return $common;\n+}\n+\n+sub read_commit_tree {\n+    my ($commit) = @_;\n+    my ($fhi);\n+    open $fhi, '-|', 'cat-file', 'commit', $commit\n+\tor die \"$!: cat-file commit $commit\";\n+    my $tree = <$fhi>;\n+    close $fhi;\n+    $tree =~ s/^tree //;\n+    return $tree;\n+}\n+\n+sub read_diff_tree {\n+    my (@tree) = @_;\n+    my ($fhi);\n+    local ($_, $/);\n+    $/ = \"\\0\"; \n+    my %path;\n+    open $fhi, '-|', 'diff-tree', '-r', @tree\n+\tor die \"$!: diff-tree -r @tree\";\n+    while (<$fhi>) {\n+\tchomp;\n+\tif (/^\\*[0-7]+->([0-7]+)\\tblob\\t[0-9a-f]+->([0-9a-f]{40})\\t(.*)$/s) {\n+\t    # mode newsha path\n+\t    $path{$3} = [$1, $2];\n+\t}\n+\telsif (/^\\+([0-7]+)\\tblob\\t([0-9a-f]{40})\\t(.*)$/s) {\n+\t    # mode newsha path\n+\t    $path{$3} = [$1, $2];\n+\t}\n+\telse {\n+\t    print STDERR \"$_??\";\n+\t}\n+    }\n+    close $fhi;\n+    return %path;\n+}\n+\n+sub read_show_files {\n+    my ($fhi);\n+    local ($_, $/);\n+    $/ = \"\\0\"; \n+    open $fhi, '-|', 'show-files', '-z'\n+\tor die \"$!: show-files -z\";\n+    my (@path) = map { chomp; $_ } <$fhi>;\n+    close $fhi;\n+    return @path;\n+}\n+\n+sub checkout_file {\n+    my ($path, $info) = @_;\n+    my (@elt) = split(/\\//, $path);\n+    my $j = '';\n+    my $tail = pop @elt;\n+    my ($fhi, $fho);\n+    for (@elt) {\n+\tmkdir \"$j$_\";\n+\t$j = \"$j$_/\";\n+    }\n+    open $fho, '>', \"$path\";\n+    open $fhi, '-|', 'cat-file', 'blob', $info->[1]\n+\tor die \"$!: cat-file blob $info->[1]\";\n+    while (<$fhi>) {\n+\tprint $fho $_;\n+    }\n+    close $fhi;\n+    close $fho;\n+    chmod oct(\"0$info->[0]\"), \"$path\";\n+}\n+\n+sub record_file {\n+    my ($path, $info) = @_;\n+    system 'update-cache', '--cacheinfo', @$info, $path;\n+}\n+\n+sub merge_tree {\n+    my ($path, $info0, $info1) = @_;\n+    print STDERR \"M - $path\\n\";\n+    checkout_file(',,merge-0', $info0);\n+    checkout_file(',,merge-1', $info1);\n+    system 'checkout-cache', $path;\n+    my ($fhi, $fho);\n+    open $fhi, '-|', 'merge', '-p', ',,merge-0', $path, ',,merge-1';\n+    open $fho, '>', \"$path+\";\n+    local ($/);\n+    while (<$fhi>) { print $fho $_; }\n+    close $fhi;\n+    close $fho;\n+    unlink ',,merge-0', ',,merge-1';\n+    rename \"$path+\", $path;\n+    # There is no reason to prefer info0 over info1 but\n+    # we need to pick one.\n+    chmod oct(\"0$info0->[0]\"), \"$path\";\n+}\n+\n+# Find common ancestor of two trees.\n+my $common = read_rev_tree(@ARGV);\n+print \"Common ancestor: $common\\n\";\n+\n+# Create a temporary directory and go there.\n+system 'rm', '-rf', ',,merge-temp';\n+for ((',,merge-temp', '.git')) { mkdir $_; chdir $_; }\n+symlink \"../../.git/objects\", \"objects\";\n+chdir '..';\n+\n+my $ancestor_tree = read_commit_tree($common);\n+system 'read-tree', $ancestor_tree;\n+\n+my %tree0 = read_diff_tree($ancestor_tree, read_commit_tree($ARGV[0]));\n+my %tree1 = read_diff_tree($ancestor_tree, read_commit_tree($ARGV[1]));\n+\n+my @ancestor_file = read_show_files();\n+my %ancestor_file = map { $_ => 1 } @ancestor_file;\n+\n+for (@ancestor_file) {\n+    if (! exists $tree0{$_} && ! exists $tree1{$_}) {\n+\tif ($full_checkout) {\n+\t    system 'checkout-cache', $_;\n+\t}\n+\tprint STDERR \"O - $_\\n\";\n+    }\n+}\n+\n+my %need_merge = ();\n+\n+for $path (keys %tree0) {\n+    if (! exists $tree1{$path}) {\n+\t# Only changed in tree 0 --- take his version\n+\tprint STDERR \"0 - $path\\n\";\n+\tif (! exists $ancestor_file{$path}) {\n+\t    checkout_file($path, $tree0{$path});\n+\t    system 'update-cache', '--add', \"$path\";\n+\t}\n+\telsif ($oneside_checkout) {\n+\t    checkout_file($path, $tree0{$path});\n+\t}\n+\telse {\n+\t    record_file($path, $tree0{$path});\n+\t}\n+    }\n+    else {\n+\tmerge_tree($path, $tree0{$path}, $tree1{$path});\n+    }\n+}\n+\n+for $path (keys %tree1) {\n+    if (! exists $tree0{$path}) {\n+\t# Only changed in tree 1 --- take his version\n+\tprint STDERR \"1 - $path\\n\";\n+\tif (! exists $ancestor_file{$path}) {\n+\t    checkout_file($path, $tree1{$path});\n+\t    system 'update-cache', '--add', \"$path\";\n+\t}\n+\telsif ($oneside_checkout) {\n+\t    checkout_file($path, $tree1{$path});\n+\t}\n+\telse {\n+\t    record_file($path, $tree1{$path});\n+\t}\n+    }\n+}\n+\n+# system 'show-diff';\n\n\n"},{"id":"50","messageId":"Pine.LNX.4.58.0504140051550.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"7vfyxtsurd.fsf@assigned-by-dhcp.cox.net","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-14T08:06:56Z","receivedAt":"2005-04-14T08:06:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 14 Apr 2005, Junio C Hamano wrote:\n> \n> I now have a Perl script that uses rev-tree, cat-file,\n> diff-tree, show-files (with one modification so that it can deal\n> with pathnames with embedded newlines), update-cache (with one\n> modification so that I can add an entry for a file that does not\n> exist to the dircache) and merge (from RCS).  Quick and dirty.\n\nThat's exactly what I wanted. Q'n'D is how the ball gets rolling.\n\nIn the meantime I wrote a very stupid \"merge-tree\" which does things\nslightly differently, but I really think your approach (aka my original\napproach) is actually a lot faster. I was just starting to worry that the \nball didn't start, so I wrote an even hackier one.\n\nMy really hacky one is called \"merge-tree\", and it really only merges one \ndirectory. For each entry in the directory it says either\n\n\tselect <mode> <sha1> path\n\nor\n\n\tmerge <mode>-><mode>,<mode> <sha1>-><sha1>,<sha1> path\n\ndepending on whether it could directly select the right object or not.\n\nIt's actually exactly the same algorithm as the first one, but I was \nafraid the first one would be so abstract that it (a) might not work and \n(b) wouldn't get people to work it out. This \"one directory at a time with \nvery explicit output\" thing is much more down-to-earth, but it's also\nlikely slower because it will need script help more often.\n\nThat said, I don't know. MOST of the time there will be just a single \n\"directory\" entry that needs merging, and then the script would just need \nto recurse into that directory with the new \"tree\" objects. So it might \nnot be too horrible.\n\nBut I'm really happy that you seem to have implemented my first \nsuggestion and I seem to have been wasting my time. \n\n>  5. for each path involved:\n> \n>   5.0 if neither heads change it, leave it as is;\n>   5.1 if only one head changes a path and the other does not, just\n>       get the changed version;\n>   5.2 if both heads change it, check all three out and run merge.\n\nYou missed one case: \n\n    5.0.1 if both heads change it to the same thing, take the new thing\n\nbut maybe you counted that as 5.0 (it _should_ fall out automatically from\nthe fact that \"diff-tree\" between the two destination trees shows no\ndifference for such a file).\n\nNow, arguably, your 5.2 will do things right, but the thing is, it's \nactually fairly _common_ that both heads have changed something to the \nsame thing. Namely if there was a previous merge that already handled that \ncase, but that previous merge may not be a proper parent of the new \ncommits.  So from a performance standpoint you really don't want to \nconsider that to be a merge - you just pick up the new contents directly.\n\nSee?\n\n(My stupid \"merge-tree\" should show the algorithm in painful obviousity. \nOf course, my stipid merge-tree may also be painfully buggy. You be the \njudge).\n\n> It does not currently commit.  You can go to ./,,merge-temp/ and\n> see show-diff to see the result of the merge.  Files added in\n> one head has already been run \"update-cache\" when the script\n> ends, but changed and merged files are not---dircache still has\n> the common ancestor view.\n\nThat sounds good.\n\n> Also to implement 'changed only by one-side' without actually\n> checking the file out, I needed to add one option to\n> 'update-cache'.  --cacheinfo flag is used this way:\n> \n>     $ update-cache --cacheinfo mode sha1 path\n\nYes. My \"merge-tree\" needs the exact same thing.\n\nLooks good from your explanation, but I'm too tired to look at the code. \nIt's 1AM, and the kids get up at 7.\n\nI'm not much of a hacker, I usually crash by 10PM these days ;^)\n\n\t\t\tLinus\n"},{"id":"55","messageId":"7v64ypsqev.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504140051550.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-14T08:39:36Z","receivedAt":"2005-04-14T08:39:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> But I'm really happy that you seem to have implemented my first \nLT> suggestion and I seem to have been wasting my time. \n\nThanks for the kind words.\n\n>> 5. for each path involved:\n>> \n>> 5.0 if neither heads change it, leave it as is;\n>> 5.1 if only one head changes a path and the other does not, just\n>> get the changed version;\n>> 5.2 if both heads change it, check all three out and run merge.\n\nLT> You missed one case: \n\nLT>     5.0.1 if both heads change it to the same thing, take the new thing\n\nLT> but maybe you counted that as 5.0 (it _should_ fall out automatically from\nLT> the fact that \"diff-tree\" between the two destination trees shows no\nLT> difference for such a file).\n\nActually I am not handling that.  It really is 5.1a---the exact\nsame code path as 5.1 can be used for this case, and as you\npoint out it is really a quite important optimization.\n\nI have to handle the following cases.  I think I currently do\nwrong things to them:\n\n  5.1a both head modify to the same thing.\n  5.1b one head removes, the other does not do anything.\n  5.1c both head remove.\n  5.3 one head removes, the other head modifies.\n\nHandling of 5.1a, 5.1b and 5.1c are obvious.\n\n  5.1a Update dircache to the same new thing.  Without -f or -o\n       flag do not touch ,,merge-temp/. directory; with -f or\n       -o, leave the new file in ,,merge-temp/.\n\n  5.1b Remove the path from dircache and do not have the file in\n       ,,merge-temp/. directory regardless of -f or -o flags.\n\n  5.1c Same as 5.1b\n\nI am not sure what to do with 5.3.  My knee-jerk reaction is to\nleave the modified result in ,,merge-temp/$path~ without\ntouching dircache.  If the merger wants to pick it up, he can\nrename $path~ to $path temporarily, run show-diff on it (I think\ngiving an option to show-diff to specify paths would be helpful\nfor this workflow), to decide if he wants to keep the file or\nnot.  Suggestions?\n\n"},{"id":"58","messageId":"Pine.LNX.4.58.0504140201130.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"7v64ypsqev.fsf@assigned-by-dhcp.cox.net","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-14T09:10:22Z","receivedAt":"2005-04-14T09:10:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 14 Apr 2005, Junio C Hamano wrote:\n> \n> I have to handle the following cases.  I think I currently do\n> wrong things to them:\n> \n>   5.1a both head modify to the same thing.\n>   5.1b one head removes, the other does not do anything.\n>   5.1c both head remove.\n>   5.3 one head removes, the other head modifies.\n\nThere's another interesting set of cases: one side creates a file, and the\nother one creates a directory. \n\n> I am not sure what to do with 5.3.\n\nMy very _strong_ preference is to just inform the user about a merge that\ncannot be performed, and not let it be automated. BIG warning, with some \nway for the user to specify the end result.\n\nThe thing is, these are pretty rare cases. But in order to make people \nfeel good about the _common_ case, it's important that they feel safe \nabout the rare one.\n\nPut another way: if git tells me when it can't do something (with some\nspecificity), I can then fix the situation up and try again. I might curse\na while, and maybe it ends up being so common that I might even automate\nit, but at least I'll be able to trust the end result.\n\nIn contrast, if git does something that _may_ be nonsensical, then I'll\nworry all the time, and not trust git. That's much worse than an \noccasional curse.\n\nSo the rule should be: only merge when it's \"obviously the right thing\".  \nIf it's not obvious, the merge should _not_ try to guess what the right\nthing is. It's much better to fail loudly.\n\n(That's especially true early on. There may be cases that end up being\nobvious after some usage. But I'd rather find them by having git be too\nstupid, than find out the hard way that git lost some data because it\nthought it was ok to remove a file that had been modified)\n\n\t\t\tLinus\n"},{"id":"66","messageId":"7vvf6pr4oq.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504140201130.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-14T11:14:13Z","receivedAt":"2005-04-14T11:14:13Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Here is a diff to update the git-merge.perl script I showed you\nearlier today ;-).  It contains the following updates against\nyour HEAD (bb95843a5a0f397270819462812735ee29796fb4).\n\n * git-merge.perl command we talked about on the git list.  I've\n   covered the changed-to-the-same case etc.  I still haven't done\n   anything about file-vs-directory case yet.\n\n   It does warn when it needed to run merge to automerge and let\n   merge give a warning message about conflicts if any.  In\n   modify/remove cases, modified in one but removed in the other\n   files are left in either $path~A~ or $path~B~ in the merge\n   temporary directory, and the script issues a warning at the\n   end.\n\n * show-files and ls-tree updates to add -z flag to NUL terminate records;\n   this is needed for git-merge.perl to work.\n\n * show-diff updates to add -r flag to squelch diffs for files not in\n   the working directory.  This is mainly useful when verifying the\n   result of an automated merge.\n\n * update-cache updates to add \"--cacheinfo mode sha1\" flag to register\n   a file that is not in the current working directory.  Needed for\n   minimum-checkout merging by git-merge.perl.\n\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n\n---\n\n git-merge.perl |  247 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n ls-tree.c      |    9 +-\n show-diff.c    |   11 +-\n show-files.c   |   12 ++\n update-cache.c |   25 +++++\n 5 files changed, 296 insertions(+), 8 deletions(-)\n\ndiff -x .git -Nru ,,1/git-merge.perl ,,2/git-merge.perl\n--- ,,1/git-merge.perl\t1969-12-31 16:00:00.000000000 -0800\n+++ ,,2/git-merge.perl\t2005-04-14 04:00:14.000000000 -0700\n@@ -0,0 +1,247 @@\n+#!/usr/bin/perl -w\n+\n+use Getopt::Long;\n+\n+my $full_checkout = 0;\n+my $oneside_checkout = 0;\n+GetOptions(\"full\" => \\$full_checkout,\n+\t   \"oneside\" => \\$oneside_checkout)\n+    or die;\n+\n+if ($full_checkout) {\n+    $oneside_checkout = 1;\n+}\n+\n+sub read_rev_tree {\n+    my (@head) = @_;\n+    my ($fhi);\n+    open $fhi, '-|', 'rev-tree', '--edges', @head\n+\tor die \"$!: rev-tree --edges @head\";\n+    my %common;\n+    while (<$fhi>) {\n+\tchomp;\n+\t(undef, undef, my @common) = split(/ /, $_);\n+\tfor (@common) {\n+\t    if (s/^([a-f0-f]{40}):3$/$1/) {\n+\t\t$common{$_}++;\n+\t    }\n+\t}\n+    }\n+    close $fhi;\n+\n+    my @common = (map { $_->[1] }\n+\t\t  sort { $b->[0] <=> $a->[0] }\n+\t\t  map { [ $common{$_} => $_ ] }\n+\t\t  keys %common);\n+\n+    return $common[0];\n+}\n+\n+sub read_commit_tree {\n+    my ($commit) = @_;\n+    my ($fhi);\n+    open $fhi, '-|', 'cat-file', 'commit', $commit\n+\tor die \"$!: cat-file commit $commit\";\n+    my $tree = <$fhi>;\n+    close $fhi;\n+    $tree =~ s/^tree //;\n+    return $tree;\n+}\n+\n+# Reads diff-tree -r output and gives a hash that maps a path\n+# to 3-tuple (old-mode new-mode new-sha).\n+# When creating, old-mode is undef.  When removing, new-* are undef.\n+sub read_diff_tree {\n+    my (@tree) = @_;\n+    my ($fhi);\n+    local ($_, $/);\n+    $/ = \"\\0\"; \n+    my %path;\n+    open $fhi, '-|', 'diff-tree', '-r', @tree\n+\tor die \"$!: diff-tree -r @tree\";\n+    while (<$fhi>) {\n+\tchomp;\n+\tif (/^\\*([0-7]+)->([0-7]+)\\tblob\\t[0-9a-f]+->([0-9a-f]{40})\\t(.*)$/s) {\n+\t    $path{$4} = [$1, $2, $3];\n+\t}\n+\telsif (/^\\+([0-7]+)\\tblob\\t([0-9a-f]{40})\\t(.*)$/s) {\n+\t    $path{$3} = [undef, $1, $2];\n+\t}\n+\telsif (/^\\-([0-7]+)\\tblob\\t[0-9a-f]{40}\\t(.*)$/s) {\n+\t    $path{$2} = [$1, undef, undef];\n+\t}\n+\telse {\n+\t    die \"cannot parse diff-tree output: $_\";\n+\t}\n+    }\n+    close $fhi;\n+    return %path;\n+}\n+\n+sub read_show_files {\n+    my ($fhi);\n+    local ($_, $/);\n+    $/ = \"\\0\"; \n+    open $fhi, '-|', 'show-files', '-z'\n+\tor die \"$!: show-files -z\";\n+    my (@path) = map { chomp; $_ } <$fhi>;\n+    close $fhi;\n+    return @path;\n+}\n+\n+sub checkout_file {\n+    my ($path, $info) = @_;\n+    my (@elt) = split(/\\//, $path);\n+    my $j = '';\n+    my $tail = pop @elt;\n+    my ($fhi, $fho);\n+    for (@elt) {\n+\tmkdir \"$j$_\";\n+\t$j = \"$j$_/\";\n+    }\n+    open $fho, '>', \"$path\";\n+    open $fhi, '-|', 'cat-file', 'blob', $info->[2]\n+\tor die \"$!: cat-file blob $info->[2]\";\n+    while (<$fhi>) {\n+\tprint $fho $_;\n+    }\n+    close $fhi;\n+    close $fho;\n+    chmod oct(\"0$info->[1]\"), \"$path\";\n+}\n+\n+sub record_file {\n+    my ($path, $info) = @_;\n+    system ('update-cache', '--add', '--cacheinfo',\n+\t    $info->[1], $info->[2], $path);\n+}\n+\n+sub merge_tree {\n+    my ($path, $info0, $info1) = @_;\n+    checkout_file(',,merge-0', $info0);\n+    checkout_file(',,merge-1', $info1);\n+    system 'checkout-cache', $path;\n+    my ($fhi, $fho);\n+    open $fhi, '-|', 'merge', '-p', ',,merge-0', $path, ',,merge-1';\n+    open $fho, '>', \"$path+\";\n+    local ($/);\n+    while (<$fhi>) { print $fho $_; }\n+    close $fhi;\n+    close $fho;\n+    unlink ',,merge-0', ',,merge-1';\n+    rename \"$path+\", $path;\n+    # There is no reason to prefer info0 over info1 but\n+    # we need to pick one.\n+    chmod oct(\"0$info0->[1]\"), \"$path\";\n+}\n+\n+# Find common ancestor of two trees.\n+my $common = read_rev_tree(@ARGV);\n+print \"Common ancestor: $common\\n\";\n+\n+# Create a temporary directory and go there.\n+system 'rm', '-rf', ',,merge-temp';\n+for ((',,merge-temp', '.git')) { mkdir $_; chdir $_; }\n+symlink \"../../.git/objects\", \"objects\";\n+chdir '..';\n+\n+my $ancestor_tree = read_commit_tree($common);\n+system 'read-tree', $ancestor_tree;\n+\n+my %tree0 = read_diff_tree($ancestor_tree, read_commit_tree($ARGV[0]));\n+my %tree1 = read_diff_tree($ancestor_tree, read_commit_tree($ARGV[1]));\n+\n+my @ancestor_file = read_show_files();\n+my %ancestor_file = map { $_ => 1 } @ancestor_file;\n+\n+for (@ancestor_file) {\n+    if (! exists $tree0{$_} && ! exists $tree1{$_}) {\n+\tif ($full_checkout) {\n+\t    system 'checkout-cache', $_;\n+\t}\n+\tprint STDERR \"O - $_\\n\";\n+    }\n+}\n+\n+for my $set ([\\%tree0, \\%tree1, 'A'], [\\%tree1, \\%tree0, 'B']) {\n+    my ($treeA, $treeB, $side) = @$set;\n+    while (my ($path, $info) = each %$treeA) {\n+\t# In this loop we do not deal with overlaps.\n+\tnext if (exists $treeB->{$path});\n+\n+\tif (! defined $info->[1]) {\n+\t    # deleted in this tree only.\n+\t    unlink $path;\n+\t    system 'update-cache', '--remove', $path;\n+\t    print STDERR \"$side D $path\\n\";\n+\t}\n+\telse {\n+\t    # modified or created in this tree only.\n+\t    print STDERR \"$side M $path\\n\";\n+\t    if ($oneside_checkout) {\n+\t\tcheckout_file($path, $info);\n+\t\tsystem 'update-cache', '--add', \"$path\";\n+\t    } else {\n+\t\trecord_file($path, $info);\n+\t    }\n+\t}\n+    }\n+}\n+\n+my @warning = ();\n+\n+while (my ($path, $info0) = each %tree0) {\n+    # We need to deal only with overlaps.\n+    next if (!exists $tree1{$path});\n+\n+    my $info1 = $tree1{$path};\n+    if (! defined $info0->[1]) {\n+\t# deleted in this tree.\n+\tif (! defined $info1->[1]) {\n+\t    # deleted in both trees.  Obvious.\n+\t    print STDERR \"*DD $path\\n\";\n+\t    unlink $path;\n+\t    system 'update-cache', '--remove', $path;\n+\t}\n+\telse {\n+\t    # oops.  tree0 wants to remove but tree1 wants to modify it.\n+\t    print STDERR \"*DM $path\\n\";\n+\t    checkout_file(\"$path~B~\", $info1);\n+\t    push @warning, $path;\n+\t}\n+    }\n+    else {\n+\t# modified or created in tree0\n+\tif (! defined $info1->[1]) {\n+\t    # oops.  tree0 wants to modify but tree1 wants to remove it.\n+\t    print STDERR \"*MD $path\\n\";\n+\t    checkout_file(\"$path~A~\", $info0);\n+\t    push @warning, $path;\n+\t}\n+\telse {\n+\t    # modified both in tree0 and tree1\n+\t    # are they modifying to the same contents?\n+\t    if ($info0->[2] eq $info1->[2]) {\n+\t\t# just mode changes (or no changes)\n+\t\t# we prefer tree0 over tree1 for no particular reason.\n+\t\tprint STDERR \"*MM $path\\n\";\n+\t\trecord_file($path, $info0);\n+\t    }\n+\t    else {\n+\t\t# modified in both.  Needs merge.\n+\t\tprint STDERR \"MRG $path\\n\";\n+\t\tmerge_tree($path, $info0, $info1);\n+\t    }\n+\t}\n+    }\n+}\n+\n+if (@warning) {\n+    print \"\\nThere are some files that were deleted in one branch and\\n\"\n+\t. \"modified in another.  Please examine them carefully:\\n\";\n+    for (@warning) {\n+\tprint \"$_\\n\";\n+    }\n+}\n+\n+# system 'show-diff';\n\n\ndiff -x .git -Nru ,,1/ls-tree.c ,,2/ls-tree.c\n--- ,,1/ls-tree.c\t2005-04-14 03:47:18.000000000 -0700\n+++ ,,2/ls-tree.c\t2005-04-14 04:00:14.000000000 -0700\n@@ -5,6 +5,8 @@\n  */\n #include \"cache.h\"\n \n+int line_termination = '\\n';\n+\n static int list(unsigned char *sha1)\n {\n \tvoid *buffer;\n@@ -31,7 +33,8 @@\n \t\t * It seems not worth it to read each file just to get this\n \t\t * and the file size. -- pasky@ucw.cz */\n \t\ttype = S_ISDIR(mode) ? \"tree\" : \"blob\";\n-\t\tprintf(\"%03o\\t%s\\t%s\\t%s\\n\", mode, type, sha1_to_hex(sha1), path);\n+\t\tprintf(\"%03o\\t%s\\t%s\\t%s%c\", mode, type, sha1_to_hex(sha1),\n+\t\t       path, line_termination);\n \t}\n \treturn 0;\n }\n@@ -40,6 +43,10 @@\n {\n \tunsigned char sha1[20];\n \n+\tif (argc == 3 && !strcmp(argv[1], \"-z\")) {\n+\t  line_termination = 0;\n+\t  argc--; argv++;\n+\t}\n \tif (argc != 2)\n \t\tusage(\"ls-tree <key>\");\n \tif (get_sha1_hex(argv[1], sha1) < 0)\n\n\ndiff -x .git -Nru ,,1/show-diff.c ,,2/show-diff.c\n--- ,,1/show-diff.c\t2005-04-14 03:47:18.000000000 -0700\n+++ ,,2/show-diff.c\t2005-04-14 04:00:14.000000000 -0700\n@@ -58,15 +58,20 @@\n int main(int argc, char **argv)\n {\n \tint silent = 0;\n+\tint silent_on_nonexisting_files = 0;\n \tint entries = read_cache();\n \tint i;\n \n \twhile (argc-- > 1) {\n \t\tif (!strcmp(argv[1], \"-s\")) {\n-\t\t\tsilent = 1;\n+\t\t\tsilent_on_nonexisting_files = silent = 1;\n \t\t\tcontinue;\n \t\t}\n-\t\tusage(\"show-diff [-s]\");\n+\t\tif (!strcmp(argv[1], \"-r\")) {\n+\t\t\tsilent_on_nonexisting_files = 1;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tusage(\"show-diff [-s] [-r]\");\n \t}\n \n \tif (entries < 0) {\n@@ -83,7 +88,7 @@\n \n \t\tif (stat(ce->name, &st) < 0) {\n \t\t\tprintf(\"%s: %s\\n\", ce->name, strerror(errno));\n-\t\t\tif (errno == ENOENT && !silent)\n+\t\t\tif (errno == ENOENT && !silent_on_nonexisting_files)\n \t\t\t\tshow_diff_empty(ce);\n \t\t\tcontinue;\n \t\t}\n\n\ndiff -x .git -Nru ,,1/show-files.c ,,2/show-files.c\n--- ,,1/show-files.c\t2005-04-14 03:47:18.000000000 -0700\n+++ ,,2/show-files.c\t2005-04-14 04:00:14.000000000 -0700\n@@ -14,6 +14,7 @@\n static int show_cached = 0;\n static int show_others = 0;\n static int show_ignored = 0;\n+static int line_terminator = '\\n';\n \n static const char **dir;\n static int nr_dir;\n@@ -105,12 +106,12 @@\n \t}\n \tif (show_others) {\n \t\tfor (i = 0; i < nr_dir; i++)\n-\t\t\tprintf(\"%s\\n\", dir[i]);\n+\t\t\tprintf(\"%s%c\", dir[i], line_terminator);\n \t}\n \tif (show_cached) {\n \t\tfor (i = 0; i < active_nr; i++) {\n \t\t\tstruct cache_entry *ce = active_cache[i];\n-\t\t\tprintf(\"%s\\n\", ce->name);\n+\t\t\tprintf(\"%s%c\", ce->name, line_terminator);\n \t\t}\n \t}\n \tif (show_deleted) {\n@@ -119,7 +120,7 @@\n \t\t\tstruct stat st;\n \t\t\tif (!stat(ce->name, &st))\n \t\t\t\tcontinue;\n-\t\t\tprintf(\"%s\\n\", ce->name);\n+\t\t\tprintf(\"%s%c\", ce->name, line_terminator);\n \t\t}\n \t}\n \tif (show_ignored) {\n@@ -134,6 +135,11 @@\n \tfor (i = 1; i < argc; i++) {\n \t\tchar *arg = argv[i];\n \n+\t\tif (!strcmp(arg, \"-z\")) {\n+\t\t\tline_terminator = 0;\n+\t\t\tcontinue;\n+\t\t}\n+\n \t\tif (!strcmp(arg, \"--cached\")) {\n \t\t\tshow_cached = 1;\n \t\t\tcontinue;\n\n\ndiff -x .git -Nru ,,1/update-cache.c ,,2/update-cache.c\n--- ,,1/update-cache.c\t2005-04-14 03:47:18.000000000 -0700\n+++ ,,2/update-cache.c\t2005-04-14 04:00:14.000000000 -0700\n@@ -250,6 +250,8 @@\n {\n \tint i, newfd, entries;\n \tint allow_options = 1;\n+\tconst char *sha1_force = NULL;\n+\tconst char *mode_force = NULL;\n \n \tnewfd = open(\".git/index.lock\", O_RDWR | O_CREAT | O_EXCL, 0600);\n \tif (newfd < 0)\n@@ -282,14 +284,35 @@\n \t\t\t\trefresh_cache();\n \t\t\t\tcontinue;\n \t\t\t}\n+\t\t\tif (!strcmp(path, \"--cacheinfo\")) {\n+\t\t\t\tmode_force = argv[++i];\n+\t\t\t\tsha1_force = argv[++i];\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tdie(\"unknown option %s\", path);\n \t\t}\n \t\tif (!verify_path(path)) {\n \t\t\tfprintf(stderr, \"Ignoring path %s\\n\", argv[i]);\n \t\t\tcontinue;\n \t\t}\n-\t\tif (add_file_to_cache(path))\n+\t\tif (sha1_force && mode_force) {\n+\t\t\tstruct cache_entry *ce;\n+\t\t\tint namelen = strlen(path);\n+\t\t\tint mode;\n+\t\t\tint size = cache_entry_size(namelen);\n+\t\t\tsscanf(mode_force, \"%o\", &mode);\n+\t\t\tce = malloc(size);\n+\t\t\tmemset(ce, 0, size);\n+\t\t\tmemcpy(ce->name, path, namelen);\n+\t\t\tce->namelen = namelen;\n+\t\t\tce->st_mode = mode;\n+\t\t\tget_sha1_hex(sha1_force, ce->sha1);\n+\n+\t\t\tadd_cache_entry(ce, 1);\n+\t\t}\n+\t\telse if (add_file_to_cache(path))\n \t\t\tdie(\"Unable to add %s to database\", path);\n+\t\tmode_force = sha1_force = NULL;\n \t}\n \tif (write_cache(newfd, active_cache, active_nr) ||\n \t    rename(\".git/index.lock\", \".git/index\"))\n\n"},{"id":"74","messageId":"20050414121624.GZ25711@pasky.ji.cz","threadId":"9","inReplyTo":"7vvf6pr4oq.fsf@assigned-by-dhcp.cox.net","subject":"Re: Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-14T12:16:24Z","receivedAt":"2005-04-14T12:16:24Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, Apr 14, 2005 at 01:14:13PM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> told me that...\n> Here is a diff to update the git-merge.perl script I showed you\n> earlier today ;-).  It contains the following updates against\n> your HEAD (bb95843a5a0f397270819462812735ee29796fb4).\n\nBah, you outran me. ;-)\n\n>  * git-merge.perl command we talked about on the git list.  I've\n>    covered the changed-to-the-same case etc.  I still haven't done\n>    anything about file-vs-directory case yet.\n> \n>    It does warn when it needed to run merge to automerge and let\n>    merge give a warning message about conflicts if any.  In\n>    modify/remove cases, modified in one but removed in the other\n>    files are left in either $path~A~ or $path~B~ in the merge\n>    temporary directory, and the script issues a warning at the\n>    end.\n\nI think I will take it rather my working git merge implementation - it's\ngetting insane in bash. ;-)\n\nI'll change it to use the cool git-pasky stuff (commit-id etc) and its\nstyle of committing - that is, it will merely record the update-caches\nto be done upon commit, and it will read-tree the branch we are merging\nto instead of the ancestor. (So that git diff gives useful output.)\n\n>  * show-files and ls-tree updates to add -z flag to NUL terminate records;\n>    this is needed for git-merge.perl to work.\n> \n>  * show-diff updates to add -r flag to squelch diffs for files not in\n>    the working directory.  This is mainly useful when verifying the\n>    result of an automated merge.\n\n-r traditionally means recursive - what's the reasoning behind the\nchoice of this letter?\n\n>  * update-cache updates to add \"--cacheinfo mode sha1\" flag to register\n>    a file that is not in the current working directory.  Needed for\n>    minimum-checkout merging by git-merge.perl.\n> \n> \n> Signed-off-by: Junio C Hamano <junkio@cox.net>\n> \n> ---\n> \n>  git-merge.perl |  247 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n>  ls-tree.c      |    9 +-\n>  show-diff.c    |   11 +-\n>  show-files.c   |   12 ++\n>  update-cache.c |   25 +++++\n>  5 files changed, 296 insertions(+), 8 deletions(-)\n> \n> diff -x .git -Nru ,,1/git-merge.perl ,,2/git-merge.perl\n> --- ,,1/git-merge.perl\t1969-12-31 16:00:00.000000000 -0800\n> +++ ,,2/git-merge.perl\t2005-04-14 04:00:14.000000000 -0700\n> @@ -0,0 +1,247 @@\n> +#!/usr/bin/perl -w\n> +\n> +use Getopt::Long;\n\nuse strict?\n\n> +\n> +my $full_checkout = 0;\n> +my $oneside_checkout = 0;\n> +GetOptions(\"full\" => \\$full_checkout,\n> +\t   \"oneside\" => \\$oneside_checkout)\n> +    or die;\n> +\n> +if ($full_checkout) {\n> +    $oneside_checkout = 1;\n> +}\n> +\n> +sub read_rev_tree {\n> +    my (@head) = @_;\n> +    my ($fhi);\n> +    open $fhi, '-|', 'rev-tree', '--edges', @head\n> +\tor die \"$!: rev-tree --edges @head\";\n> +    my %common;\n> +    while (<$fhi>) {\n> +\tchomp;\n> +\t(undef, undef, my @common) = split(/ /, $_);\n> +\tfor (@common) {\n> +\t    if (s/^([a-f0-f]{40}):3$/$1/) {\n> +\t\t$common{$_}++;\n> +\t    }\n> +\t}\n> +    }\n> +    close $fhi;\n> +\n> +    my @common = (map { $_->[1] }\n> +\t\t  sort { $b->[0] <=> $a->[0] }\n> +\t\t  map { [ $common{$_} => $_ ] }\n> +\t\t  keys %common);\n> +\n> +    return $common[0];\n> +}\n\nIt'd be simpler to do just\n\n\tmy @common = (map { $common{$_} }\n\t              sort { $b <=> $a }\n\t              keys %common)\n\nBut I really think this is a horrible heuristic. I believe you should\ntake the latest commit in the --edges output, and from that choose the\nbase whose rev-tree --edges the_base merged_branch has the least lines\non output. (That is, the path to it is shortest - ideally it's already\npart of the merged_branch.)\n\n> +\n> +sub read_commit_tree {\n> +    my ($commit) = @_;\n> +    my ($fhi);\n> +    open $fhi, '-|', 'cat-file', 'commit', $commit\n> +\tor die \"$!: cat-file commit $commit\";\n> +    my $tree = <$fhi>;\n> +    close $fhi;\n> +    $tree =~ s/^tree //;\n> +    return $tree;\n> +}\n> +\n> +# Reads diff-tree -r output and gives a hash that maps a path\n> +# to 3-tuple (old-mode new-mode new-sha).\n> +# When creating, old-mode is undef.  When removing, new-* are undef.\n\nWhat about\n\nsub OLDMODE { 0 }\nsub NEWMODE { 1 }\nsub NEWSHA { 2 }\n\nand then using that when accessing the tuple? Would make the code\nmuch more readable.\n\n> +sub read_diff_tree {\n> +    my (@tree) = @_;\n> +    my ($fhi);\n> +    local ($_, $/);\n> +    $/ = \"\\0\"; \n> +    my %path;\n> +    open $fhi, '-|', 'diff-tree', '-r', @tree\n> +\tor die \"$!: diff-tree -r @tree\";\n> +    while (<$fhi>) {\n> +\tchomp;\n> +\tif (/^\\*([0-7]+)->([0-7]+)\\tblob\\t[0-9a-f]+->([0-9a-f]{40})\\t(.*)$/s) {\n> +\t    $path{$4} = [$1, $2, $3];\n> +\t}\n> +\telsif (/^\\+([0-7]+)\\tblob\\t([0-9a-f]{40})\\t(.*)$/s) {\n> +\t    $path{$3} = [undef, $1, $2];\n> +\t}\n> +\telsif (/^\\-([0-7]+)\\tblob\\t[0-9a-f]{40}\\t(.*)$/s) {\n> +\t    $path{$2} = [$1, undef, undef];\n> +\t}\n> +\telse {\n> +\t    die \"cannot parse diff-tree output: $_\";\n> +\t}\n> +    }\n> +    close $fhi;\n> +    return %path;\n> +}\n> +\n> +sub read_show_files {\n> +    my ($fhi);\n> +    local ($_, $/);\n> +    $/ = \"\\0\"; \n> +    open $fhi, '-|', 'show-files', '-z'\n> +\tor die \"$!: show-files -z\";\n> +    my (@path) = map { chomp; $_ } <$fhi>;\n> +    close $fhi;\n> +    return @path;\n> +}\n> +\n> +sub checkout_file {\n> +    my ($path, $info) = @_;\n> +    my (@elt) = split(/\\//, $path);\n> +    my $j = '';\n> +    my $tail = pop @elt;\n> +    my ($fhi, $fho);\n> +    for (@elt) {\n> +\tmkdir \"$j$_\";\n> +\t$j = \"$j$_/\";\n> +    }\n> +    open $fho, '>', \"$path\";\n> +    open $fhi, '-|', 'cat-file', 'blob', $info->[2]\n> +\tor die \"$!: cat-file blob $info->[2]\";\n> +    while (<$fhi>) {\n> +\tprint $fho $_;\n> +    }\n> +    close $fhi;\n> +    close $fho;\n> +    chmod oct(\"0$info->[1]\"), \"$path\";\n> +}\n> +\n> +sub record_file {\n> +    my ($path, $info) = @_;\n> +    system ('update-cache', '--add', '--cacheinfo',\n> +\t    $info->[1], $info->[2], $path);\n> +}\n> +\n> +sub merge_tree {\n> +    my ($path, $info0, $info1) = @_;\n> +    checkout_file(',,merge-0', $info0);\n> +    checkout_file(',,merge-1', $info1);\n> +    system 'checkout-cache', $path;\n> +    my ($fhi, $fho);\n> +    open $fhi, '-|', 'merge', '-p', ',,merge-0', $path, ',,merge-1';\n> +    open $fho, '>', \"$path+\";\n> +    local ($/);\n> +    while (<$fhi>) { print $fho $_; }\n> +    close $fhi;\n> +    close $fho;\n> +    unlink ',,merge-0', ',,merge-1';\n> +    rename \"$path+\", $path;\n> +    # There is no reason to prefer info0 over info1 but\n> +    # we need to pick one.\n> +    chmod oct(\"0$info0->[1]\"), \"$path\";\n> +}\n\nIt is a good idea to check merge's exit code and give a notice at the\nend if there were any conflicts.\n\n> +\n> +# Find common ancestor of two trees.\n> +my $common = read_rev_tree(@ARGV);\n> +print \"Common ancestor: $common\\n\";\n> +\n> +# Create a temporary directory and go there.\n> +system 'rm', '-rf', ',,merge-temp';\n\nCan't we call it just ,,merge?\n\n> +for ((',,merge-temp', '.git')) { mkdir $_; chdir $_; }\n> +symlink \"../../.git/objects\", \"objects\";\n> +chdir '..';\n> +\n> +my $ancestor_tree = read_commit_tree($common);\n> +system 'read-tree', $ancestor_tree;\n> +\n> +my %tree0 = read_diff_tree($ancestor_tree, read_commit_tree($ARGV[0]));\n> +my %tree1 = read_diff_tree($ancestor_tree, read_commit_tree($ARGV[1]));\n> +\n> +my @ancestor_file = read_show_files();\n> +my %ancestor_file = map { $_ => 1 } @ancestor_file;\n> +\n> +for (@ancestor_file) {\n> +    if (! exists $tree0{$_} && ! exists $tree1{$_}) {\n> +\tif ($full_checkout) {\n> +\t    system 'checkout-cache', $_;\n> +\t}\n> +\tprint STDERR \"O - $_\\n\";\n\nHuh, what are you trying to do here? I think you should just record\nremove, no? (And I wouldn't do anything with my read-tree. ;-)\n\n> +    }\n> +}\n> +\n> +for my $set ([\\%tree0, \\%tree1, 'A'], [\\%tree1, \\%tree0, 'B']) {\n> +    my ($treeA, $treeB, $side) = @$set;\n> +    while (my ($path, $info) = each %$treeA) {\n> +\t# In this loop we do not deal with overlaps.\n> +\tnext if (exists $treeB->{$path});\n> +\n> +\tif (! defined $info->[1]) {\n> +\t    # deleted in this tree only.\n> +\t    unlink $path;\n> +\t    system 'update-cache', '--remove', $path;\n> +\t    print STDERR \"$side D $path\\n\";\n> +\t}\n> +\telse {\n> +\t    # modified or created in this tree only.\n> +\t    print STDERR \"$side M $path\\n\";\n> +\t    if ($oneside_checkout) {\n> +\t\tcheckout_file($path, $info);\n> +\t\tsystem 'update-cache', '--add', \"$path\";\n> +\t    } else {\n> +\t\trecord_file($path, $info);\n> +\t    }\n> +\t}\n> +    }\n> +}\n..snip..\n\nHmm, I think I will just need to play with the script a lot. ;-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"98","messageId":"7vll7lqlbg.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"20050414121624.GZ25711@pasky.ji.cz","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-14T18:12:35Z","receivedAt":"2005-04-14T18:12:35Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n\nPB> Bah, you outran me. ;-)\n\nJust being in a different timezone, I guess.\n\nPB> I'll change it to use the cool git-pasky stuff (commit-id etc) and its\nPB> style of committing - that is, it will merely record the update-caches\nPB> to be done upon commit, and it will read-tree the branch we are merging\nPB> to instead of the ancestor. (So that git diff gives useful output.)\n\nSorry, I have not seen what you have been doing since pasky 0.3,\nand I have not even started to understand the mental model of\nthe world your tool is building.  That said, my gut feeling is\nthat telling this script about git-pasky's world model might be\na mistake.  I'd rather see you consider the script as mere \"part\nof the plumbing\".  Maybe adding an extra parameter to the script\nto let the user explicitly specify the common ancestor to use\nwould be needed, but I would prefer git-pasky-merge to do its\nown magic (converting symbolic commit names into raw commit\nnames and such) before calling this low level script.\n\nThat way people like me who have not migrated to your framework\ncan still keep using it.  All the script currently needs is a\nbare git object database; i.e., nothing other than what is in\n.git/objects and a couple of commit record SHA1s as its\nparameters.  No .git/heads/, no .git/HEAD.local, no .git/tags,\nare involved for it to work, and I would prefer to keep things\nthat way if possible.\n\n>> * show-diff updates to add -r flag to squelch diffs for files not in\n>> the working directory.  This is mainly useful when verifying the\n>> result of an automated merge.\n\nPB> -r traditionally means recursive - what's the reasoning behind the\nPB> choice of this letter?\n\nWell, '-r' is not necessarily recursive. \"ls -r\" is reverse, \"sort\n-r\" is reverse.  \"less -r\" is raw.  \"cat -r\" is reversible.\n\"nethack -r\" is race ;-).  You are thinking as an SCM person so\nit may look that way.  \"diff -r\" is recursive.  \"darcs add -r\"\nis recursive.  But even in the SCM world, \"cvs add -r\" is not\n(it means read-only) neither \"co -r\" (explicit revision) ;-).\n\nI would rather pick '-q' if I were doing the patch today, but I\nwas too tired and did not think of a letter when I wrote it.  I\nguess '-r' stood for removed, but I agree it is a bad choice.\nAny objections to '-q'?\n\nPB> use strict?\n\nNot in this iteration but eventually yes.\n\nPB> It'd be simpler to do just\n\nPB> \tmy @common = (map { $common{$_} }\nPB> \t              sort { $b <=> $a }\nPB> \t              keys %common)\n\nWell, actually you spotted a bug between the implementation and\nwhat I wanted to do.  It should have been:\n\nmap { $_->[0] }\n    sort { $b->[1] <=> $a->[1] }\n        map { [ $common{$_} => $_ ] } keys %common\n\nThat is, sort [anscestor => number of times it appears] tuple by\nthe \"number of times it appears\" in decreasing order, and\nproject the resulting list to a list of ancestors.  It is trying\nto deal with the following pattern in rev-tree output:\n\nTIMESTAMP1 EDGE1:1 ANCESTOR1:3 ANCESTOR2:3\nTIMESTAMP2 EDGE2:2 ANCESTOR1:3\n\nand when the above happens I wanted to pick up ANCESTOR1, but\nthat was without no sound reason.\n\nPB> But I really think this is a horrible heuristic. I believe you should\nPB> take the latest commit in the --edges output, and from that choose the\nPB> base whose rev-tree --edges the_base merged_branch has the least lines\nPB> on output. (That is, the path to it is shortest - ideally it's already\nPB> part of the merged_branch.)\n\nI'll try something along that line.  Honestly the ancestor\nselection part was what I had most trouble with.  Thanks.\n\nPB> What about\n\nPB> sub OLDMODE { 0 }\nPB> sub NEWMODE { 1 }\nPB> sub NEWSHA { 2 }\n\nPB> and then using that when accessing the tuple? Would make the code\nPB> much more readable.\n\nTotally agreed; readability cleanup is needed, just as \"use\nstrict\" you mentioned, before it is ready for public\nconsumption.  Remember, however, the primary purpose of the\nmessage was to share it with Linus so that I can ask his opinion\nwhile the script was still slushy; the contents that array\ncontained was still changing then and was too early for symbolic\nconstants.  I'll do that in the next round.\n\nPB> It is a good idea to check merge's exit code and give a notice at the\nPB> end if there were any conflicts.\n\nIn principle yes, but I noticed that merge already gave me a\nnice warning message when it found conflicts, so there was no\nneed to do so myself in this case.  See sample output:\n\n    $ perl ./git-merge.perl \\\n        71796686221a0a56ccc25b02386ed8ea648da14d \\\n        bb95843a5a0f397270819462812735ee29796fb4 \n    Common ancestor: 9f02d4d233223462d3f6217b5837b786e6286ba4\n    O - COPYING\n    O - README\n    ...\n    O - write-tree.c\n    A M write-blob.c\n    A M show-diff.c\n    ...\n    A M update-cache.c\n    A M git-merge.perl\n    B M merge-tree.c\n    MRG Makefile\n    merge: warning: conflicts during merge\n    $ \n\n>> +# Create a temporary directory and go there.\n>> +system 'rm', '-rf', ',,merge-temp';\n\nPB> Can't we call it just ,,merge?\n\nI'd rather have a command line option '-o' (scrapping the\ncurrent '-o' and renaming it to something else; as you can see I\nam terrible at picking option names ;-)) to mean \"output to this\ndirectory\".  I am not really an Arch person so I do not\nparticulary care about /^,,/.  How about \"git~merge~$$\"?\n\n>> +for ((',,merge-temp', '.git')) { mkdir $_; chdir $_; }\n>> +symlink \"../../.git/objects\", \"objects\";\n>> +chdir '..';\n>> +\n>> +my $ancestor_tree = read_commit_tree($common);\n>> +system 'read-tree', $ancestor_tree;\n>> +\n>> +my %tree0 = read_diff_tree($ancestor_tree, read_commit_tree($ARGV[0]));\n>> +my %tree1 = read_diff_tree($ancestor_tree, read_commit_tree($ARGV[1]));\n>> +\n>> +my @ancestor_file = read_show_files();\n>> +my %ancestor_file = map { $_ => 1 } @ancestor_file;\n>> +\n>> +for (@ancestor_file) {\n>> +    if (! exists $tree0{$_} && ! exists $tree1{$_}) {\n>> +\tif ($full_checkout) {\n>> +\t    system 'checkout-cache', $_;\n>> +\t}\n>> +\tprint STDERR \"O - $_\\n\";\n\nPB> Huh, what are you trying to do here? I think you should just record\nPB> remove, no? (And I wouldn't do anything with my read-tree. ;-)\n\nAt this moment in the script, we have run \"read-tree\" the\nancestor so the dircache has the original.  %tree0 and %tree1\nboth did not touch the path ($_ here) so it is the same as\nancestor.  When '-f' is specified we are populating the output\nworking tree with the merge result so that is what that\n'checkout-cache' is about.  \"O - $path\" means \"we took the\noriginal\".\n\nThe idea is to populate the dircache of merge-temp with the\nmerge result and leave uncertain stuff as in the common ancestor\nstate, so that the user can fix them starting from there.\n\nMaybe it is a good time for me to summarize the output somewhere\nin a document.\n\n    O - $path\tTree-A and tree-B did not touch this; the result\n                is taken from the ancestor (O for original).\n\n    A D $path\tOnly tree-A (or tree-B) deleted this and the other\n    B D $path   branch did not touch this; the result is to delete.\n\n    A M $path\tOnly tree-A (or tree-B) modified this and the other\n    B M $path   branch did not touch this; the result is to use one\n                from tree-A (or tree-B).  This includes file\n                creation case.\n\n    *DD $path\tBoth tree-A and tree-B deleted this; the result\n                is to delete.\n\n    *DM $path   Tree-A deleted while tree-B modified this (or\n    *MD $path   vice versa), and manual conflict resolution is\n                needed; dircache is left as in the ancestor, and\n                the modified file is saved as $path~A~ in the\n                working directory.  The user can rename it to $path\n                and run show-diff to see what Tree-A wanted to do\n                and decide before running update-cache.\n\n    *MM $path   Tree-A and tree-B did the exact same\n                modification; the result is to use that.\n\n    MRG $path   Tree-A and tree-B have different modifications;\n                run \"merge\" and the merge result is left as\n                $path in the working directory.\n\nIn cases other than *DM, *MD, and MRG, the result is trivial and\nis recorded in the dircache.  Without '-o' (to be renamed ;-)\nnor '-f' there will not be a file checked out in the working\ndirectory for them.  The three merge cases need human attention.\nThe dircache is not touched in these cases and left as the\nancestor version, and the working directory gets some file as\ndescribed above.\n\nNOTE NOTE NOTE: I am not dealing with a case where both branches\ncreate the same file but with different contents.  In such a\ncase the current code falls into MRG path without having a\ncommon ancestor, which is nonsense---I can use /dev/null as the\ncommon ancestor, I guess.  Also NOTE NOTE NOTE I need to detect\nthe case where one branch creates a directory while the other\ncreates a file.  There is nothing an automated tool can do in\nthat case but it needs to be detected and be told the user\nloudly.\n\n"},{"id":"103","messageId":"Pine.LNX.4.58.0504141133260.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"7vll7lqlbg.fsf@assigned-by-dhcp.cox.net","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-14T18:36:52Z","receivedAt":"2005-04-14T18:36:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 14 Apr 2005, Junio C Hamano wrote:\n> \n> Sorry, I have not seen what you have been doing since pasky 0.3,\n> and I have not even started to understand the mental model of\n> the world your tool is building.  That said, my gut feeling is\n> that telling this script about git-pasky's world model might be\n> a mistake.  I'd rather see you consider the script as mere \"part\n> of the plumbing\". \n\nI agree. Having separate abstraction layers is good.  I'm actually very \nhappy with Pasky's cleaned-up-tree, exactly because unlike the first one, \nPasky did a great job of maintaining the abstraction between \"plumbing\" \nand user interfaces.\n\nThe plumbing should take user interface needs into account, but the more\nconceptually separate it is (\"does it makes sense on its own?\") the better\noff we'll be. And \"merge these two trees\" (which works on a _tree_ level)\nor \"find the common commit\" (which works on a _commit_ level) look like \nplumbing to me - the kind of things I should have written, if I weren't \nsuch a lazy slob.\n\n\t\tLinus\n"},{"id":"133","messageId":"20050414185122.GA25468@64m.dyndns.org","threadId":"9","inReplyTo":"7vll7lqlbg.fsf@assigned-by-dhcp.cox.net","subject":"Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-14T18:51:22Z","receivedAt":"2005-04-14T18:51:22Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Thu, Apr 14, 2005 at 11:12:35AM -0700, Junio C Hamano wrote:\n> >>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n> \n> At this moment in the script, we have run \"read-tree\" the\n> ancestor so the dircache has the original.  %tree0 and %tree1\n> both did not touch the path ($_ here) so it is the same as\n> ancestor.  When '-f' is specified we are populating the output\n> working tree with the merge result so that is what that\n> 'checkout-cache' is about.  \"O - $path\" means \"we took the\n> original\".\n> \n> The idea is to populate the dircache of merge-temp with the\n> merge result and leave uncertain stuff as in the common ancestor\n> state, so that the user can fix them starting from there.\n> \n> Maybe it is a good time for me to summarize the output somewhere\n> in a document.\n> \n>     O - $path\tTree-A and tree-B did not touch this; the result\n>                 is taken from the ancestor (O for original).\n> \n>     A D $path\tOnly tree-A (or tree-B) deleted this and the other\n>     B D $path   branch did not touch this; the result is to delete.\n> \n>     A M $path\tOnly tree-A (or tree-B) modified this and the other\n>     B M $path   branch did not touch this; the result is to use one\n>                 from tree-A (or tree-B).  This includes file\n>                 creation case.\n> \n>     *DD $path\tBoth tree-A and tree-B deleted this; the result\n>                 is to delete.\n> \n>     *DM $path   Tree-A deleted while tree-B modified this (or\n>     *MD $path   vice versa), and manual conflict resolution is\n>                 needed; dircache is left as in the ancestor, and\n>                 the modified file is saved as $path~A~ in the\n>                 working directory.  The user can rename it to $path\n>                 and run show-diff to see what Tree-A wanted to do\n>                 and decide before running update-cache.\n> \n>     *MM $path   Tree-A and tree-B did the exact same\n>                 modification; the result is to use that.\n> \n>     MRG $path   Tree-A and tree-B have different modifications;\n>                 run \"merge\" and the merge result is left as\n>                 $path in the working directory.\n> \n> In cases other than *DM, *MD, and MRG, the result is trivial and\n\nI believe there is simpler way to do it as in my demo python script.\nI start it easier but you bits me in time. It is a demo script, it\nonly print the action instead of actually going out to do it.\nchange that to corresponding os.system(\"\") call leaves to the reader.\n\nAgain, this is a demo how it can be done. Not python vs perl thing\nI did not chose perl only because I am not good at it.\n\n#!/usr/bin/env python\n\nimport re\nimport sys\nimport os\nfrom pprint import pprint\n\ndef get_tree(commit):\n    data = os.popen(\"cat-file commit %s\"%commit).read()\n    return re.findall(r\"(?m)^tree (\\w+)\", data)[0]\n\nPREFIX = 0\nPATH = -1\nSHA = -2\nORIGSHA = -3\n\ndef get_difftree(old, new):\n    lines = os.popen(\"diff-tree %s %s\"%(old, new)).read().split(\"\\x00\")\n    patterns = (r\"(\\*)(\\d+)->(\\d+)\\s(\\w+)\\s(\\w+)->(\\w+)\\s(.*)\",\n\t\tr\"([+-])(\\d+)\\s(\\w+)\\s(\\w+)\\s(.*)\")\n    res = {}\n    for l in lines:\n\tif not l: continue\n\tfor p in patterns:\n\t    m = re.findall(p, l)\n\t    if m:\n\t\tm = m[0]\n\t\tres[m[-1]] = m\n\t\tbreak\n\telse:\n\t    raise \"difftree: unknow line\", l\n    return res\n\ndef analyze(diff1, diff2):\n    diff1only = [ diff1[k] for k in diff1 if k not in diff2 ]\n    diff2only = [ diff2[k] for k in diff2 if k not in diff1 ]\n    both = [ (diff1[k],diff2[k]) for k in diff2 if k in diff1 ]\n\n    action(diff1only)\n    action(diff2only)\n    action_two(both)\n\ndef action(diffs):\n    for act in diffs:\n\tif act[PREFIX] == \"*\":\n\t    print \"modify\", act[PATH], act[SHA]\n\telif act[PREFIX] == '-':\n\t    print \"remove\", act[PATH], act[SHA]\n\telif act[PREFIX] == '+':\n\t    print \"add\", \"remove\", act[PATH], act[SHA]\n\telse:\n\t    raise \"unknow action\"\n\ndef action_two(diffs):\n    for act1, act2 in diffs:\n\tif len(act1) == len(act2):\t# same kind type\n\t    if act1[PREFIX] == act2[PREFIX]:\n\t\tif act1[SHA] == act2[SHA] or act1[PREFIX] == '-': \n\t\t    return action(act1)\n\t    \tif act1[PREFIX]=='*':\n\t\t    print \"3way-merge\", act1[PATH], act1[ORIGSHA], act1[SHA], act2[SHA]\n\t\t    return\n\tprint \"unable to handle\", act[PATH]\n\tprint \"one side wants\", act1[PREFIX]\n\tprint \"the other side wants\", act2[PREFIX]\n\t\n    \nargs = sys.argv[1:]\ntrees = map(get_tree, args)\nprint \"check out tree\", trees[0]\ndiff1 = get_difftree(trees[0], trees[1])\ndiff2 = get_difftree(trees[0], trees[2])\nanalyze(diff1, diff2)\n\n"},{"id":"116","messageId":"20050414193507.GA22699@pasky.ji.cz","threadId":"9","inReplyTo":"7vll7lqlbg.fsf@assigned-by-dhcp.cox.net","subject":"Re: Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-14T19:35:07Z","receivedAt":"2005-04-14T19:35:07Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, Apr 14, 2005 at 08:12:35PM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> told me that...\n> >>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n> \n> PB> Bah, you outran me. ;-)\n> \n> Just being in a different timezone, I guess.\n> \n> PB> I'll change it to use the cool git-pasky stuff (commit-id etc) and its\n> PB> style of committing - that is, it will merely record the update-caches\n> PB> to be done upon commit, and it will read-tree the branch we are merging\n> PB> to instead of the ancestor. (So that git diff gives useful output.)\n> \n> Sorry, I have not seen what you have been doing since pasky 0.3,\n> and I have not even started to understand the mental model of\n> the world your tool is building.  That said, my gut feeling is\n> that telling this script about git-pasky's world model might be\n> a mistake.  I'd rather see you consider the script as mere \"part\n> of the plumbing\".  Maybe adding an extra parameter to the script\n> to let the user explicitly specify the common ancestor to use\n> would be needed, but I would prefer git-pasky-merge to do its\n> own magic (converting symbolic commit names into raw commit\n> names and such) before calling this low level script.\n> \n> That way people like me who have not migrated to your framework\n> can still keep using it.  All the script currently needs is a\n> bare git object database; i.e., nothing other than what is in\n> .git/objects and a couple of commit record SHA1s as its\n> parameters.  No .git/heads/, no .git/HEAD.local, no .git/tags,\n> are involved for it to work, and I would prefer to keep things\n> that way if possible.\n\nI see, and I actually agree with it. However, I'll want merge-tree.pl to\ndo a little less than it does now for that, though. The mechanics in\n\"kernel\" is fine as long as I can control policy in my \"userspace\". ;-)\n\nBTW, the git* name sorta imply my toilet instead of the core plumbing, and\nit'd be more consistent with the current plumbnaming; and could we have\nit with the .pl extension, please? :-)\n\nWhat I would like your script to do is therefore just do the merge in a\ngiven already prepared (including built index) directory, with a passed\nbase. The base should be determined by a separate tool (I already saw\nsome patches); most future \"science\" will probably go to a clever\nselection of this base, anyway.\n\nThis will give the tool maximal flexibility. E.g., then someone who\nwants to can just merge with his working copy (if you don't give\ncheckout-cache -f - but why would you anyway), or do whatever other\ncleverness he wants.\n\n> >> * show-diff updates to add -r flag to squelch diffs for files not in\n> >> the working directory.  This is mainly useful when verifying the\n> >> result of an automated merge.\n..snip..\n> was too tired and did not think of a letter when I wrote it.  I\n> guess '-r' stood for removed, but I agree it is a bad choice.\n> Any objections to '-q'?\n\nNone here.\n\n> >> +# Create a temporary directory and go there.\n> >> +system 'rm', '-rf', ',,merge-temp';\n> \n> PB> Can't we call it just ,,merge?\n> \n> I'd rather have a command line option '-o' (scrapping the\n> current '-o' and renaming it to something else; as you can see I\n> am terrible at picking option names ;-)) to mean \"output to this\n> directory\".  I am not really an Arch person so I do not\n> particulary care about /^,,/.  How about \"git~merge~$$\"?\n\nI'm all for an -o, and I don't mind ,, - I just don't want it uselessly\nlong. I hope \"git~merge~$$\" was a joke... :-)\n\n> >> +for ((',,merge-temp', '.git')) { mkdir $_; chdir $_; }\n> >> +symlink \"../../.git/objects\", \"objects\";\n> >> +chdir '..';\n> >> +\n> >> +my $ancestor_tree = read_commit_tree($common);\n> >> +system 'read-tree', $ancestor_tree;\n> >> +\n> >> +my %tree0 = read_diff_tree($ancestor_tree, read_commit_tree($ARGV[0]));\n> >> +my %tree1 = read_diff_tree($ancestor_tree, read_commit_tree($ARGV[1]));\n> >> +\n> >> +my @ancestor_file = read_show_files();\n> >> +my %ancestor_file = map { $_ => 1 } @ancestor_file;\n> >> +\n> >> +for (@ancestor_file) {\n> >> +    if (! exists $tree0{$_} && ! exists $tree1{$_}) {\n\nBy the way, what about indentation with tabs? If you have a strong\nopinion about this, I don't insist - but if you really don't mind/care\neither way, it'd be great to use tabs as in the rest of the git code.\n\n> >> +\tif ($full_checkout) {\n> >> +\t    system 'checkout-cache', $_;\n> >> +\t}\n> >> +\tprint STDERR \"O - $_\\n\";\n> \n> PB> Huh, what are you trying to do here? I think you should just record\n> PB> remove, no? (And I wouldn't do anything with my read-tree. ;-)\n> \n> At this moment in the script, we have run \"read-tree\" the\n> ancestor so the dircache has the original.  %tree0 and %tree1\n> both did not touch the path ($_ here) so it is the same as\n> ancestor.  When '-f' is specified we are populating the output\n> working tree with the merge result so that is what that\n> 'checkout-cache' is about.  \"O - $path\" means \"we took the\n> original\".\n\nAha! Thanks.\n\nIs there a fundamental reason why the directory cache contains the\nancestor instead of the destination branch? It makes no sense to me and\nI think the script actually does not fundamentally depend on it. My main\nmotivation is that the user can then trivially see what is he actually\ngoing to commit to his destination branch, which would be bought for\nfree by that.\n\n> The idea is to populate the dircache of merge-temp with the\n> merge result and leave uncertain stuff as in the common ancestor\n> state, so that the user can fix them starting from there.\n\nAnd this is another thing I dislike a lot. I'd like merge-tree.pl to\nleave my directory cache alone, thank you very much. You know, I see\nwhat goes to the directory cache as actually part of the policy part.\n\nWhat you actually do is interfering with my different policy choice,\nwhich is to record stuff to index only at the time of commit (I've asked\nabout this and noone replied, so I assume it's an ok choice). show-diff\ndoes the right thing for me then, and I don't need to care about losing\n*any* information when replacing/rebuilding the index for any reason. I\nhave full control, and I like that. :-)\n\nI'd be happy with parsing merge-tree.pl output and doing the right thing\non my side. Of course I could then blast away the tediously modified\nindex with my one, but I didn't need to do any such hacking before and\nI'd prefer not to now either.\n\nActually, the only time I need to do explicit update-cache (with my\npolicy) when doing git merge is when deleting stuff or adding new stuff;\nboth of this is not so common as modifying, when I need not to do\nanything.\n\n> Maybe it is a good time for me to summarize the output somewhere\n> in a document.\n> \n>     O - $path\tTree-A and tree-B did not touch this; the result\n>                 is taken from the ancestor (O for original).\n> \n>     A D $path\tOnly tree-A (or tree-B) deleted this and the other\n>     B D $path   branch did not touch this; the result is to delete.\n> \n>     A M $path\tOnly tree-A (or tree-B) modified this and the other\n>     B M $path   branch did not touch this; the result is to use one\n>                 from tree-A (or tree-B).  This includes file\n>                 creation case.\n\nCould we please have the file creation case separately? Modification\nis much more common and creation has pretty different consequences\n(especially that it can't combine with anything else :-).\n\n>     *DD $path\tBoth tree-A and tree-B deleted this; the result\n>                 is to delete.\n> \n>     *DM $path   Tree-A deleted while tree-B modified this (or\n>     *MD $path   vice versa), and manual conflict resolution is\n>                 needed; dircache is left as in the ancestor, and\n>                 the modified file is saved as $path~A~ in the\n>                 working directory.  The user can rename it to $path\n>                 and run show-diff to see what Tree-A wanted to do\n>                 and decide before running update-cache.\n> \n>     *MM $path   Tree-A and tree-B did the exact same\n>                 modification; the result is to use that.\n> \n>     MRG $path   Tree-A and tree-B have different modifications;\n>                 run \"merge\" and the merge result is left as\n>                 $path in the working directory.\n\nHmm. I actually don't like this naming. I think it's not too consistent,\nis irregular, therefore parsing it would be ugly. What I propose:\n\n12c\\tname <- legend\n          <- original file\nD         <- tree #1 removed file\n D        <- tree #2 removed file\nDD        <- both trees removed file\nM         <- tree #1 modified file\n M\nDM*       <- conflict, tree #1 removed file, tree #2 modified file\nMD*\nMM        <- exact same modification\nMM*       <- different modifications, merging\n\nThis is generic, theoretically scales well even to more trees, is easy\nto parse trivially, still is human readable (actually the asterisk in\nthe 'conflict' column is there basically only for the humans), is\ncompletely regular and consistent.\n\nNow that we have the notion of tree A and tree B gone, I'd prefer to use\nnumbers instead of letters for the ~1~ and ~2~ suffixes. Not insisting,\nthough.\n\nWhat do you think?\n\n> In cases other than *DM, *MD, and MRG, the result is trivial and\n> is recorded in the dircache.  Without '-o' (to be renamed ;-)\n> nor '-f' there will not be a file checked out in the working\n> directory for them.  The three merge cases need human attention.\n> The dircache is not touched in these cases and left as the\n> ancestor version, and the working directory gets some file as\n> described above.\n> \n> NOTE NOTE NOTE: I am not dealing with a case where both branches\n> create the same file but with different contents.  In such a\n> case the current code falls into MRG path without having a\n> common ancestor, which is nonsense---I can use /dev/null as the\n> common ancestor, I guess.  Also NOTE NOTE NOTE I need to detect\n\nThat might be the best way at least for the start, although I suspect\nthat merge will fail horribly this way even in the case of slightest\ndifferences; still better than nothing.\n\n> the case where one branch creates a directory while the other\n> creates a file.  There is nothing an automated tool can do in\n> that case but it needs to be detected and be told the user\n> loudly.\n\nOr when both branches create directories... ;-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"121","messageId":"7v7jj5qgdz.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504141133260.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-14T19:59:04Z","receivedAt":"2005-04-14T19:59:04Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> On Thu, 14 Apr 2005, Junio C Hamano wrote:\n\n>> Sorry, I have not seen what you have been doing since pasky 0.3,\n>> and I have not even started to understand the mental model of\n>> the world your tool is building.  That said, my gut feeling is\n>> that telling this script about git-pasky's world model might be\n>> a mistake.  I'd rather see you consider the script as mere \"part\n>> of the plumbing\". \n\nLT> I agree. Having separate abstraction layers is good.  I'm actually very \nLT> happy with Pasky's cleaned-up-tree, exactly because unlike the first one, \nLT> Pasky did a great job of maintaining the abstraction between \"plumbing\" \nLT> and user interfaces.\n\nAgreed, not just with your agreeing with me, but with the\nstatement that Pasky did a good job (although I am ashamed to\nsay I have not caught up with the \"userland\" tools).\n\nLT> The plumbing should take user interface needs into account, but the more\nLT> conceptually separate it is (\"does it makes sense on its own?\") the better\nLT> off we'll be. And \"merge these two trees\" (which works on a _tree_ level)\nLT> or \"find the common commit\" (which works on a _commit_ level) look like \nLT> plumbing to me - the kind of things I should have written, if I weren't \nLT> such a lazy slob.\n\nI am planning drop the ancestor computation from the script, and\nmake it another command line parameter to the script.  Dan\nBarkalow's merge-base program should be used to compute it and\nhis result should drive the merge.  That sounds more UNIXy to\nme.  I even may want to make the script take three trees not\ncommits, since the merge script does not need commits (it only\nneeds trees).  As plumbing it would be cleaner interface to it\nto do so.  The wrapper SCM scripts can and should make sure it\nis fed trees when the user gives it commits (or symbolic\nrepresentation of it like .git/tags/blah, or `cat .git/HEAD`).\n\nBut one different thing to note here.\n\nYou say \"merge these two trees\" above (I take it that you mean\n\"merge these two trees, taking account of this tree as their\ncommon ancestor\", so actually you are dealing with three trees),\nand I am tending to agree with the notion of merging trees not\ncommits.  However you might get richer context and more sensible\nresulting merge if you say \"merge these two commits\".  Since\ncommit chaining is part of the fundamental git object model you\nmay as well use it.\n\nThis however opens up another set of can of worms---it would\ninvolve not just three trees but all the trees in the commit\nchain in between.  That's when you start wondering if it would\nbe better to add renames in the git object model, which is the\ntopic of another thread.  I have not formed an opinion on that\none myself yet.\n\n"},{"id":"120","messageId":"IGEMLBGAECDFPIKMIMLCCEELCHAA.barry@disus.com","threadId":"9","inReplyTo":"20050414193507.GA22699@pasky.ji.cz","subject":"Live Merging from remote repositories","fromName":"Barry Silverman","fromEmail":"barry@disus.com","sentAt":"2005-04-14T20:01:32Z","receivedAt":"2005-04-14T20:01:32Z","isPatch":false,"sender":{"key":"barry@disus.com","avatar":null},"body":"If you are merging from many distributed developers, than you would need to\nreplicate every one of their repositories into your own. Is this necessary?\n\nI have been looking at Junio's code for merging, and it looks like it would\nbe (relatively) easy change to make it run live across two remote\nrepositories - assuming \"future\" science to develop remote common ancestor\nlookup...\n\nIE, merge.pl $COMMON-BASE $LOCAL-CHANGESET remote::$REMOTE-CHANGESET\n\nTo make this work, only a couple of things need to happen:\n1) be able to remotely run \"remote::diff-tree $BASE $REMOTE-CHANGESET\", and\ncopy the results over the net to the place in the script where it is done\nlocally. This is not a LOT of data, and is bounded by the number of total\nnumber of blobs in the resulting tree.\n\n2) When a remote blob is required (for merging, or copying), then copy it\nfrom the remote .git/objects to the local one. You only copy the blobs that\nwill end up in the merged result (or be used for file merge).\n\nThe way Junio has done it, no intermediate trees or commits are used...\n\nYou don't copy remote's tree of commits between $BASE and $REMOTE-CHANGESET,\nor any of their associated trees and blobs (unless used to merge).\n\nIs this a bug or a feature?\n\n\nBarry Silverman\n\n"},{"id":"122","messageId":"20050414202016.GC22699@pasky.ji.cz","threadId":"9","inReplyTo":"7v7jj5qgdz.fsf@assigned-by-dhcp.cox.net","subject":"Re: Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-14T20:20:16Z","receivedAt":"2005-04-14T20:20:16Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, Apr 14, 2005 at 09:59:04PM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> told me that...\n> >>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n> \n> LT> On Thu, 14 Apr 2005, Junio C Hamano wrote:\n> \n> >> Sorry, I have not seen what you have been doing since pasky 0.3,\n> >> and I have not even started to understand the mental model of\n> >> the world your tool is building.  That said, my gut feeling is\n> >> that telling this script about git-pasky's world model might be\n> >> a mistake.  I'd rather see you consider the script as mere \"part\n> >> of the plumbing\". \n> \n> LT> I agree. Having separate abstraction layers is good.  I'm actually very \n> LT> happy with Pasky's cleaned-up-tree, exactly because unlike the first one, \n\n(Just a side-note - functionally and even organizationally, the cleaned\nup tree does not differ significantly from the original one.)\n\n> LT> Pasky did a great job of maintaining the abstraction between \"plumbing\" \n> LT> and user interfaces.\n> \n> Agreed, not just with your agreeing with me, but with the\n> statement that Pasky did a good job (although I am ashamed to\n> say I have not caught up with the \"userland\" tools).\n\nThanks. :-)\n\n> LT> The plumbing should take user interface needs into account, but the more\n> LT> conceptually separate it is (\"does it makes sense on its own?\") the better\n> LT> off we'll be. And \"merge these two trees\" (which works on a _tree_ level)\n> LT> or \"find the common commit\" (which works on a _commit_ level) look like \n> LT> plumbing to me - the kind of things I should have written, if I weren't \n> LT> such a lazy slob.\n> \n> I am planning drop the ancestor computation from the script, and\n> make it another command line parameter to the script.  Dan\n> Barkalow's merge-base program should be used to compute it and\n> his result should drive the merge.  That sounds more UNIXy to\n> me.\n\nGood move, I say!\n\n> I even may want to make the script take three trees not\n> commits, since the merge script does not need commits (it only\n> needs trees).  As plumbing it would be cleaner interface to it\n> to do so.  The wrapper SCM scripts can and should make sure it\n> is fed trees when the user gives it commits (or symbolic\n> representation of it like .git/tags/blah, or `cat .git/HEAD`).\n\nAgreed.\n\n> But one different thing to note here.\n> \n> You say \"merge these two trees\" above (I take it that you mean\n> \"merge these two trees, taking account of this tree as their\n> common ancestor\", so actually you are dealing with three trees),\n> and I am tending to agree with the notion of merging trees not\n> commits.  However you might get richer context and more sensible\n> resulting merge if you say \"merge these two commits\".  Since\n> commit chaining is part of the fundamental git object model you\n> may as well use it.\n\nCould you be more particular on the richer context etc?\n\nI think this script should stay strictly on the level of trees. When\nsomeone invents it, there could be a merge-commits script which does\nsomething very smart about two commits, traversing the graph between\nthem etc, and doing a set of merge-tree invocations, possibly preparing\nthe staging area for them etc.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"123","messageId":"20050414222326.E2442@banaan.localdomain","threadId":"9","inReplyTo":"20050414193507.GA22699@pasky.ji.cz","subject":"Re: Re: Merge with git-pasky II.","fromName":"Erik van Konijnenburg","fromEmail":"ekonijn@xs4all.nl","sentAt":"2005-04-14T20:23:26Z","receivedAt":"2005-04-14T20:23:26Z","isPatch":false,"sender":{"key":"ekonijn@xs4all.nl","avatar":null},"body":"On Thu, Apr 14, 2005 at 09:35:07PM +0200, Petr Baudis wrote:\n> Hmm. I actually don't like this naming. I think it's not too consistent,\n> is irregular, therefore parsing it would be ugly. What I propose:\n> \n> 12c\\tname <- legend\n>           <- original file\n> D         <- tree #1 removed file\n>  D        <- tree #2 removed file\n> DD        <- both trees removed file\n> M         <- tree #1 modified file\n>  M\n> DM*       <- conflict, tree #1 removed file, tree #2 modified file\n> MD*\n> MM        <- exact same modification\n> MM*       <- different modifications, merging\n> \n> This is generic, theoretically scales well even to more trees, is easy\n> to parse trivially, still is human readable (actually the asterisk in\n> the 'conflict' column is there basically only for the humans), is\n> completely regular and consistent.\n\nDetail: perhaps use underscore instead of space, to avoid space/tab typos\nthat are invisible on paper and user friendly mail clients?\n\nRegards,\nErik\n"},{"id":"164","messageId":"20050414202421.GC25468@64m.dyndns.org","threadId":"9","inReplyTo":"7vmzs1osv1.fsf@assigned-by-dhcp.cox.net","subject":"Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-14T20:24:21Z","receivedAt":"2005-04-14T20:24:21Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"Hi Junio,\n\nI think if the merge tree belong to plumbing, you can do\neven less in the merge.perl. You can just print out the\ninstruction for the upper level SCM what to to without\nactually doing it yourself.\n\nSo you don't have to do touch anything in the tree.\nThat is the way I use in my previous python script.\nYou just print out some easy to modify  \n\ne.g. in my python script it prints: (BTW, poor choice of print out name)\n\ncheck out tree 253290af8b9ebc8565dd8de4cda24d0432a92b57\nmodify pre-process.c 7684c115a87e41a9226ce79478101c746cf22c34\n3way-merge check.c dcb970cc1c5a83284dc5986abf07b6da76a8758c f77bfe119c19d928879091e0e3ee6debe3f1e1bf d315b43b025350d0107568a4d42cc2494d38621d\n\nYour merge tree can do the smae.\n\nThen the supper level SCM can easily follow instruction.\nSave your effort and make no assumption what SCM module is.\n\nChris\n\nOn Thu, Apr 14, 2005 at 04:12:34PM -0700, Junio C Hamano wrote:\n> >>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n> \n> I think you are contradicting yourself for saying the above\n> after agreeing with me that the script should just work on trees\n> not commits.  My understanding is that the tools is just to\n> merge two related trees relative to another ancestor tree,\n> nothing more.  Especially, it should not care what is in the\n> working directory---that is SCM person's business.\n> \n> I am just trying to follow my understanding of what Linus\n> wanted.  One of the guiding principle is to do as much things as\n> in dircache without ever checking things out or touching working\n> files unnecessarily.\n> \n> PB> This will give the tool maximal flexibility.\n> \n> I suspect it would force me to have a working directory\n> populated with files, just to do a merge.\n> \n> PB> I'm all for an -o, and I don't mind ,, - I just don't want it uselessly\n> PB> long. I hope \"git~merge~$$\" was a joke... :-)\n> \n> Which part do you object to?  PID part?  or tilde?  Would\n> git~merge do, perhaps?  It probably would not matter to you\n> because as an SCM you would always give an explicit --output\n> parameter to the script anyway.\n> \n> PB> By the way, what about indentation with tabs? If you have a\n> PB> strong opinion about this, I don't insist - but if you\n> PB> really don't mind/care either way, it'd be great to use tabs\n> PB> as in the rest of the git code.\n> \n> I do not have a strong opinion, but it is more trouble for me\n> only because I am lazy and am used to the indentation my Emacs\n> gives me.  I write code other than git, so changing Perl-mode\n> indentation setting globally for all .pl files is not an option\n> for me.  I'll see what I can do when I have time.\n> \n> PB> Is there a fundamental reason why the directory cache\n> PB> contains the ancestor instead of the destination branch?\n> \n> Because you are thinking as an SCM person where there are\n> distinction between tree-A and tree-B, two heads being merged.\n> There is no \"destination branch\" nor \"source branch\" in what I\n> am doing.  It is a merge of two equals derived from the same\n> ancestor.\n> \n> PB> I think the script actually does not fundamentally depend on it. My main\n> PB> motivation is that the user can then trivially see what is he actually\n> PB> going to commit to his destination branch, which would be bought for\n> PB> free by that.\n> \n> And again the user is *not* commiting to his \"destination\n> branch\".  At the level I am working at, the merge result should\n> be commited with two -p parameters to commit-tree --- tree-A and\n> tree-B, both being equal parents from the POV of git object\n> storage.\n> \n> PB> And this is another thing I dislike a lot. I'd like merge-tree.pl to\n> PB> leave my directory cache alone, thank you very much. You know, I see\n> PB> what goes to the directory cache as actually part of the policy part.\n> \n> Remember I am not touching *your* dircache.  It is a dircache in\n> the temporary merge area, specifically set up to help you review\n> the merge.  \n> \n> Can't the SCM driver do things along this line, perhaps?\n> \n>  - You have your working files and your dircache.  They may not\n>    match because you have uncommitted changes to your\n>    environment.  You want to merge with Linus head.  You know\n>    its SHA1 (call it COMMIT-Linus).  Your SCM knows which commit\n>    you started with (call it COMMIT-Current).\n> \n>  - First you merge the tree associated with COMMIT-Current.  Use\n>    it and COMMIT-Linus to find the common ancestor to use.\n> \n>  - Now use the tree SHA of COMMIT-Current, tree SHA1 of\n>    COMMIT-Linus, and tree SHA1 of the common ancestor commit to\n>    drive git-merge.perl (to be renamed ;-).  You will get a\n>    temporary directory.  Have your user examine what is in\n>    there, and fix the merge and have them tell you they are\n>    happy.\n> \n>  - You go to that temporary directory, do write-tree and\n>    commit-tree with -p parameter of COMMIT-Linus and\n>    COMMIT-Current.  This will result in a new commit.  Call that\n>    COMMIT-Merge.\n> \n>  - You, as an SCM, should know what your user have done in the\n>    working directory relative to COMMIT-Current.  Especially you\n>    should know the set of paths involved in that change.  Go in\n>    to the temporary area, checkout-cache those files if you have\n>    not done so.  Apply the changes you have there.  Optionally\n>    have the user examine the changes and have him confirm.  Lift\n>    those files into the user's working directory.\n> \n>  - Do your bookkeeping like \"echo COMMIT-Merge >.git/Head\", to\n>    make the user's working files based on COMMIT-Merge, and run\n>    read-tree using the COMMIT-Merge in the user's working\n>    directory.  At this point, show-diff output should show what\n>    the changes your user have had made if he had started working\n>    based on COMMIT-Merge instead of starting from\n>    COMMIT-Current.\n> \n> I think the above would result in what SCM person would call\n> \"merge upstream/sidestream changes into my working directory\".\n> \n"},{"id":"124","messageId":"20050414202451.GD22699@pasky.ji.cz","threadId":"9","inReplyTo":"20050414222326.E2442@banaan.localdomain","subject":"Re: Re: Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-14T20:24:51Z","receivedAt":"2005-04-14T20:24:51Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, Apr 14, 2005 at 10:23:26PM CEST, I got a letter\nwhere Erik van Konijnenburg <ekonijn@xs4all.nl> told me that...\n> On Thu, Apr 14, 2005 at 09:35:07PM +0200, Petr Baudis wrote:\n> > Hmm. I actually don't like this naming. I think it's not too consistent,\n> > is irregular, therefore parsing it would be ugly. What I propose:\n> > \n> > 12c\\tname <- legend\n> >           <- original file\n> > D         <- tree #1 removed file\n> >  D        <- tree #2 removed file\n> > DD        <- both trees removed file\n> > M         <- tree #1 modified file\n> >  M\n> > DM*       <- conflict, tree #1 removed file, tree #2 modified file\n> > MD*\n> > MM        <- exact same modification\n> > MM*       <- different modifications, merging\n> > \n> > This is generic, theoretically scales well even to more trees, is easy\n> > to parse trivially, still is human readable (actually the asterisk in\n> > the 'conflict' column is there basically only for the humans), is\n> > completely regular and consistent.\n> \n> Detail: perhaps use underscore instead of space, to avoid space/tab typos\n> that are invisible on paper and user friendly mail clients?\n\nI'd go for dots in that case. Looks less intrusive. :^)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"166","messageId":"20050414203013.GD25468@64m.dyndns.org","threadId":"9","inReplyTo":"20050414233159.GX22699@pasky.ji.cz","subject":"Re: Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-14T20:30:13Z","receivedAt":"2005-04-14T20:30:13Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Fri, Apr 15, 2005 at 01:31:59AM +0200, Petr Baudis wrote:\n> > I am just trying to follow my understanding of what Linus\n> > wanted.  One of the guiding principle is to do as much things as\n> > in dircache without ever checking things out or touching working\n> > files unnecessarily.\n> \n> I'm just arguing that instead of directly touching the directory cache,\n> you should just list what would you do there - and you already do this,\n\nThat is exactly what I suggest in the previous email. And my python script\ndoes exactly that ;-)\n\nChris\n\n"},{"id":"167","messageId":"20050414203717.GE25468@64m.dyndns.org","threadId":"9","inReplyTo":"20050414203013.GD25468@64m.dyndns.org","subject":"Re: Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-14T20:37:17Z","receivedAt":"2005-04-14T20:37:17Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"Is that some thing you want to see? Maybe clean up the error printing.\n\n\nChris\n\n--- /dev/null\t2003-01-30 05:24:37.000000000 -0500\n+++ merge.py\t2005-04-14 16:34:39.000000000 -0400\n@@ -0,0 +1,76 @@\n+#!/usr/bin/env python\n+\n+import re\n+import sys\n+import os\n+from pprint import pprint\n+\n+def get_tree(commit):\n+    data = os.popen(\"cat-file commit %s\"%commit).read()\n+    return re.findall(r\"(?m)^tree (\\w+)\", data)[0]\n+\n+PREFIX = 0\n+PATH = -1\n+SHA = -2\n+ORIGSHA = -3\n+\n+def get_difftree(old, new):\n+    lines = os.popen(\"diff-tree %s %s\"%(old, new)).read().split(\"\\x00\")\n+    patterns = (r\"(\\*)(\\d+)->(\\d+)\\s(\\w+)\\s(\\w+)->(\\w+)\\s(.*)\",\n+\t\tr\"([+-])(\\d+)\\s(\\w+)\\s(\\w+)\\s(.*)\")\n+    res = {}\n+    for l in lines:\n+\tif not l: continue\n+\tfor p in patterns:\n+\t    m = re.findall(p, l)\n+\t    if m:\n+\t\tm = m[0]\n+\t\tres[m[-1]] = m\n+\t\tbreak\n+\telse:\n+\t    raise \"difftree: unknow line\", l\n+    return res\n+\n+def analyze(diff1, diff2):\n+    diff1only = [ diff1[k] for k in diff1 if k not in diff2 ]\n+    diff2only = [ diff2[k] for k in diff2 if k not in diff1 ]\n+    both = [ (diff1[k],diff2[k]) for k in diff2 if k in diff1 ]\n+\n+    action(diff1only)\n+    action(diff2only)\n+    action_two(both)\n+\n+def action(diffs):\n+    for act in diffs:\n+\tif act[PREFIX] == \"*\":\n+\t    print \"modify\", act[PATH], act[SHA]\n+\telif act[PREFIX] == '-':\n+\t    print \"remove\", act[PATH], act[SHA]\n+\telif act[PREFIX] == '+':\n+\t    print \"add\", act[PATH], act[SHA]\n+\telse:\n+\t    raise \"unknow action\"\n+\n+def action_two(diffs):\n+    for act1, act2 in diffs:\n+\tif len(act1) == len(act2):\t# same kind type\n+\t    if act1[PREFIX] == act2[PREFIX]:\n+\t\tif act1[SHA] == act2[SHA] or act1[PREFIX] == '-': \n+\t\t    return action(act1)\n+\t    \tif act1[PREFIX]=='*':\n+\t\t    print \"do_merge\", act1[PATH], act1[ORIGSHA], act1[SHA], act2[SHA]\n+\t\t    return\n+\tprint \"unable to handle\", act[PATH]\n+\tprint \"one side wants\", act1[PREFIX]\n+\tprint \"the other side wants\", act2[PREFIX]\n+\t\n+\n+args = sys.argv[1:]\n+if len(args)!=3:\n+    print \"Usage merge.py <common> <rev1> <rev2>\"\n+trees = map(get_tree, args)\n+print \"checkout-tree\", trees[0]\n+diff1 = get_difftree(trees[0], trees[1])\n+diff2 = get_difftree(trees[0], trees[2])\n+analyze(diff1, diff2)\n+\n"},{"id":"170","messageId":"20050414205021.GA28082@64m.dyndns.org","threadId":"9","inReplyTo":"20050414203717.GE25468@64m.dyndns.org","subject":"Re: Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-14T20:50:21Z","receivedAt":"2005-04-14T20:50:21Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"BTW, I am not competing with Junio script. If that is the way\nwe all agree on. It is should be very easy for Junio to fix his\nperl script. right?\n\nChris\n\nOn Thu, Apr 14, 2005 at 04:37:17PM -0400, Christopher Li wrote:\n> Is that some thing you want to see? Maybe clean up the error printing.\n> \n> \n> Chris\n> \n> --- /dev/null\t2003-01-30 05:24:37.000000000 -0500\n> +++ merge.py\t2005-04-14 16:34:39.000000000 -0400\n> @@ -0,0 +1,76 @@\n> +#!/usr/bin/env python\n> +\n> +import re\n> +import sys\n> +import os\n> +from pprint import pprint\n> +\n> +def get_tree(commit):\n> +    data = os.popen(\"cat-file commit %s\"%commit).read()\n> +    return re.findall(r\"(?m)^tree (\\w+)\", data)[0]\n> +\n> +PREFIX = 0\n> +PATH = -1\n> +SHA = -2\n> +ORIGSHA = -3\n> +\n> +def get_difftree(old, new):\n> +    lines = os.popen(\"diff-tree %s %s\"%(old, new)).read().split(\"\\x00\")\n> +    patterns = (r\"(\\*)(\\d+)->(\\d+)\\s(\\w+)\\s(\\w+)->(\\w+)\\s(.*)\",\n> +\t\tr\"([+-])(\\d+)\\s(\\w+)\\s(\\w+)\\s(.*)\")\n> +    res = {}\n> +    for l in lines:\n> +\tif not l: continue\n> +\tfor p in patterns:\n> +\t    m = re.findall(p, l)\n> +\t    if m:\n> +\t\tm = m[0]\n> +\t\tres[m[-1]] = m\n> +\t\tbreak\n> +\telse:\n> +\t    raise \"difftree: unknow line\", l\n> +    return res\n> +\n> +def analyze(diff1, diff2):\n> +    diff1only = [ diff1[k] for k in diff1 if k not in diff2 ]\n> +    diff2only = [ diff2[k] for k in diff2 if k not in diff1 ]\n> +    both = [ (diff1[k],diff2[k]) for k in diff2 if k in diff1 ]\n> +\n> +    action(diff1only)\n> +    action(diff2only)\n> +    action_two(both)\n> +\n> +def action(diffs):\n> +    for act in diffs:\n> +\tif act[PREFIX] == \"*\":\n> +\t    print \"modify\", act[PATH], act[SHA]\n> +\telif act[PREFIX] == '-':\n> +\t    print \"remove\", act[PATH], act[SHA]\n> +\telif act[PREFIX] == '+':\n> +\t    print \"add\", act[PATH], act[SHA]\n> +\telse:\n> +\t    raise \"unknow action\"\n> +\n> +def action_two(diffs):\n> +    for act1, act2 in diffs:\n> +\tif len(act1) == len(act2):\t# same kind type\n> +\t    if act1[PREFIX] == act2[PREFIX]:\n> +\t\tif act1[SHA] == act2[SHA] or act1[PREFIX] == '-': \n> +\t\t    return action(act1)\n> +\t    \tif act1[PREFIX]=='*':\n> +\t\t    print \"do_merge\", act1[PATH], act1[ORIGSHA], act1[SHA], act2[SHA]\n> +\t\t    return\n> +\tprint \"unable to handle\", act[PATH]\n> +\tprint \"one side wants\", act1[PREFIX]\n> +\tprint \"the other side wants\", act2[PREFIX]\n> +\t\n> +\n> +args = sys.argv[1:]\n> +if len(args)!=3:\n> +    print \"Usage merge.py <common> <rev1> <rev2>\"\n> +trees = map(get_tree, args)\n> +print \"checkout-tree\", trees[0]\n> +diff1 = get_difftree(trees[0], trees[1])\n> +diff2 = get_difftree(trees[0], trees[2])\n> +analyze(diff1, diff2)\n> +\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"134","messageId":"20050414221127.GI22699@pasky.ji.cz","threadId":"9","inReplyTo":"20050414002902.GU25711@pasky.ji.cz","subject":"git merge","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-14T22:11:28Z","receivedAt":"2005-04-14T22:11:28Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"  Hi,\n\n  note that in my git tree there is a git merge implementation which\ndoes out-of-tree merges now. It is still very trivial, and basically\njust does something along the lines of (symbolically written)\n\n\tcheckout-cache $(diff-tree)\n\tgit diff $base $mergedbranch | git apply\n\t.. fix rejects etc ..\n\tgit commit\n\n  It seems to work, but it is only very lightly tested - it is likely\nthere are various tiny mistakes and typos in various unusual code paths\nand other weird corners of the scripts. Testing is encouraged, and\nespecially patches fixing bugs you come over.\n\n  It is designed in a way to make it possible to just replace the\ncheckout-cache and git diff | git apply steps with the merge-tree.pl\ntool when it is finished.\n\n  Thanks,\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"181","messageId":"20050414223039.GB28082@64m.dyndns.org","threadId":"9","inReplyTo":"7v7jj4q2j2.fsf@assigned-by-dhcp.cox.net","subject":"Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-14T22:30:39Z","receivedAt":"2005-04-14T22:30:39Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Thu, Apr 14, 2005 at 05:58:25PM -0700, Junio C Hamano wrote:\n> \n> I do like, however, the idea of separating the step of doing any\n> checkout/merge etc. and actually doing them.  So the command set\n> of parse-your-output needs to be defined.  Based on what I have\n> done so far, it would consist of the following:\n> \n>  - Result is this object $SHA1 with mode $mode at $path (takes\n>    one of the trees); you can do update-cache --cacheinfo (if\n>    you want to muck with dircache) or cat-file blob (if you want\n>    to get the file) or both.\n\nIs that SHA1 for tree or the file object? If it is tree it don't\nneed the $mode any more.  If it is file you might need to emit\nentry for it's parent directory, including the modes of directory.\n\n> \n>  - Result is to delete $path.\n> \n>  - Result is a merge between object $SHA1-1 and $SHA1-2 with\n>    mode $mode-1 or $mode-2 at $path.\n>\n> Would this be a good enough command set?\n\nAnd of course error/command for the files that unable to perform\nauto merge. including information of both revisions. That needs\nto be defined as well.\n\nChris\n\n"},{"id":"157","messageId":"7vmzs1osv1.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"20050414193507.GA22699@pasky.ji.cz","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-14T23:12:34Z","receivedAt":"2005-04-14T23:12:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n\nPB> What I would like your script to do is therefore just do the\nPB> merge in a given already prepared (including built index)\nPB> directory, with a passed base. The base should be determined\nPB> by a separate tool (I already saw some patches); most future\nPB> \"science\" will probably go to a clever selection of this\nPB> base, anyway.\n\nI think you are contradicting yourself for saying the above\nafter agreeing with me that the script should just work on trees\nnot commits.  My understanding is that the tools is just to\nmerge two related trees relative to another ancestor tree,\nnothing more.  Especially, it should not care what is in the\nworking directory---that is SCM person's business.\n\nI am just trying to follow my understanding of what Linus\nwanted.  One of the guiding principle is to do as much things as\nin dircache without ever checking things out or touching working\nfiles unnecessarily.\n\nPB> This will give the tool maximal flexibility.\n\nI suspect it would force me to have a working directory\npopulated with files, just to do a merge.\n\nPB> I'm all for an -o, and I don't mind ,, - I just don't want it uselessly\nPB> long. I hope \"git~merge~$$\" was a joke... :-)\n\nWhich part do you object to?  PID part?  or tilde?  Would\ngit~merge do, perhaps?  It probably would not matter to you\nbecause as an SCM you would always give an explicit --output\nparameter to the script anyway.\n\nPB> By the way, what about indentation with tabs? If you have a\nPB> strong opinion about this, I don't insist - but if you\nPB> really don't mind/care either way, it'd be great to use tabs\nPB> as in the rest of the git code.\n\nI do not have a strong opinion, but it is more trouble for me\nonly because I am lazy and am used to the indentation my Emacs\ngives me.  I write code other than git, so changing Perl-mode\nindentation setting globally for all .pl files is not an option\nfor me.  I'll see what I can do when I have time.\n\nPB> Is there a fundamental reason why the directory cache\nPB> contains the ancestor instead of the destination branch?\n\nBecause you are thinking as an SCM person where there are\ndistinction between tree-A and tree-B, two heads being merged.\nThere is no \"destination branch\" nor \"source branch\" in what I\nam doing.  It is a merge of two equals derived from the same\nancestor.\n\nPB> I think the script actually does not fundamentally depend on it. My main\nPB> motivation is that the user can then trivially see what is he actually\nPB> going to commit to his destination branch, which would be bought for\nPB> free by that.\n\nAnd again the user is *not* commiting to his \"destination\nbranch\".  At the level I am working at, the merge result should\nbe commited with two -p parameters to commit-tree --- tree-A and\ntree-B, both being equal parents from the POV of git object\nstorage.\n\nPB> And this is another thing I dislike a lot. I'd like merge-tree.pl to\nPB> leave my directory cache alone, thank you very much. You know, I see\nPB> what goes to the directory cache as actually part of the policy part.\n\nRemember I am not touching *your* dircache.  It is a dircache in\nthe temporary merge area, specifically set up to help you review\nthe merge.  \n\nCan't the SCM driver do things along this line, perhaps?\n\n - You have your working files and your dircache.  They may not\n   match because you have uncommitted changes to your\n   environment.  You want to merge with Linus head.  You know\n   its SHA1 (call it COMMIT-Linus).  Your SCM knows which commit\n   you started with (call it COMMIT-Current).\n\n - First you merge the tree associated with COMMIT-Current.  Use\n   it and COMMIT-Linus to find the common ancestor to use.\n\n - Now use the tree SHA of COMMIT-Current, tree SHA1 of\n   COMMIT-Linus, and tree SHA1 of the common ancestor commit to\n   drive git-merge.perl (to be renamed ;-).  You will get a\n   temporary directory.  Have your user examine what is in\n   there, and fix the merge and have them tell you they are\n   happy.\n\n - You go to that temporary directory, do write-tree and\n   commit-tree with -p parameter of COMMIT-Linus and\n   COMMIT-Current.  This will result in a new commit.  Call that\n   COMMIT-Merge.\n\n - You, as an SCM, should know what your user have done in the\n   working directory relative to COMMIT-Current.  Especially you\n   should know the set of paths involved in that change.  Go in\n   to the temporary area, checkout-cache those files if you have\n   not done so.  Apply the changes you have there.  Optionally\n   have the user examine the changes and have him confirm.  Lift\n   those files into the user's working directory.\n\n - Do your bookkeeping like \"echo COMMIT-Merge >.git/Head\", to\n   make the user's working files based on COMMIT-Merge, and run\n   read-tree using the COMMIT-Merge in the user's working\n   directory.  At this point, show-diff output should show what\n   the changes your user have had made if he had started working\n   based on COMMIT-Merge instead of starting from\n   COMMIT-Current.\n\nI think the above would result in what SCM person would call\n\"merge upstream/sidestream changes into my working directory\".\n\n"},{"id":"162","messageId":"7vfyxtose8.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"IGEMLBGAECDFPIKMIMLCCEELCHAA.barry@disus.com","subject":"Re: Live Merging from remote repositories","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-14T23:22:39Z","receivedAt":"2005-04-14T23:22:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"BS\" == Barry Silverman <barry@disus.com> writes:\n\nI have not thought about remote issues at all, other than the\ndistribution mechanism vaguely outlined in my previous mail (not\ncc'ed to git list but I would not mind if you reproduced it here\nif somebody asked), so I am not qualified to comment on that\npart of your message.\n\nBS> The way Junio has done it, no intermediate trees or commits\nBS> are used...\n\nBS> Is this a bug or a feature?\n\nI would call that a feature in that there is no need to look at\nintermediate state.  I also might call that a misfeature in that\nit may have resulted in a better merge if it looked at\nintermediate state.\n\nI just have this fuzzy feeling that, when doing this merge:\n\n                     A-1 --- A-2 --- A-3\n                    /                   \\ \n    Common Ancestor                      Merge Result\n                    \\                   /\n                     B-1 --- B-2 --- B-3\n\nlooking at diff(Common Ancestor, A-1), diff(Common Ancestor,\nB-1), diff(A-1, A-2), ... might give you richer context than\njust merging 3-way using Common Ancestor, A-3, and B-3 to derive\nthe Merge Result.  It might not.  I honestly do not know.\n\nBTW, Pasky, the above paragraph is my answer to your question in\nthe other message <20050414202016.GC22699@pasky.ji.cz>:\n\n> But one different thing to note here.\n> \n> You say \"merge these two trees\" above (I take it that you mean\n> \"merge these two trees, taking account of this tree as their\n> common ancestor\", so actually you are dealing with three trees),\n> and I am tending to agree with the notion of merging trees not\n> commits.  However you might get richer context and more sensible\n> resulting merge if you say \"merge these two commits\".  Since\n> commit chaining is part of the fundamental git object model you\n> may as well use it.\n\nPasky> Could you be more particular on the richer context etc?\n\n"},{"id":"163","messageId":"20050414233159.GX22699@pasky.ji.cz","threadId":"9","inReplyTo":"7vmzs1osv1.fsf@assigned-by-dhcp.cox.net","subject":"Re: Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-14T23:31:59Z","receivedAt":"2005-04-14T23:31:59Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Apr 15, 2005 at 01:12:34AM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> told me that...\n> >>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n> \n> PB> What I would like your script to do is therefore just do the\n> PB> merge in a given already prepared (including built index)\n> PB> directory, with a passed base. The base should be determined\n> PB> by a separate tool (I already saw some patches); most future\n> PB> \"science\" will probably go to a clever selection of this\n> PB> base, anyway.\n> \n> I think you are contradicting yourself for saying the above\n> after agreeing with me that the script should just work on trees\n> not commits.  My understanding is that the tools is just to\n> merge two related trees relative to another ancestor tree,\n> nothing more.  Especially, it should not care what is in the\n> working directory---that is SCM person's business.\n\nYes. Isn't this exactly what I'm saying?\n\nI'm arguing for doing less in my paragraph, you are arguing for doing\nless in your paragraph, and we even seem to agree on the direction in\nwhich we should do less.\n\n> I am just trying to follow my understanding of what Linus\n> wanted.  One of the guiding principle is to do as much things as\n> in dircache without ever checking things out or touching working\n> files unnecessarily.\n\nI'm just arguing that instead of directly touching the directory cache,\nyou should just list what would you do there - and you already do this,\nI think. So I'd be happy with a switch which would just do that and not\ntouch the directory cache. I'll parse your output and do the right thing\nfor me.\n\n> PB> This will give the tool maximal flexibility.\n> \n> I suspect it would force me to have a working directory\n> populated with files, just to do a merge.\n\nWhy would that be so?\n\n> PB> I'm all for an -o, and I don't mind ,, - I just don't want it uselessly\n> PB> long. I hope \"git~merge~$$\" was a joke... :-)\n> \n> Which part do you object to?  PID part?  or tilde?  Would\n> git~merge do, perhaps?  It probably would not matter to you\n> because as an SCM you would always give an explicit --output\n> parameter to the script anyway.\n\nYes. I'll just override it with ,,merge, I think. So, do whatever you\nwant. ;-))\n\n> PB> By the way, what about indentation with tabs? If you have a\n> PB> strong opinion about this, I don't insist - but if you\n> PB> really don't mind/care either way, it'd be great to use tabs\n> PB> as in the rest of the git code.\n> \n> I do not have a strong opinion, but it is more trouble for me\n> only because I am lazy and am used to the indentation my Emacs\n> gives me.  I write code other than git, so changing Perl-mode\n> indentation setting globally for all .pl files is not an option\n> for me.  I'll see what I can do when I have time.\n\nDoesn't Emacs have something equivalent to ./.vimrc? I've also seen\nthose funny -*- strings.\n\nWell, if it would mean a lot of trouble for you, just forget about it.\n\n> PB> Is there a fundamental reason why the directory cache\n> PB> contains the ancestor instead of the destination branch?\n> \n> Because you are thinking as an SCM person where there are\n> distinction between tree-A and tree-B, two heads being merged.\n> There is no \"destination branch\" nor \"source branch\" in what I\n> am doing.  It is a merge of two equals derived from the same\n> ancestor.\n\nThat's a valid point of view too.\n\nActually, when you would have a mode in which you would not write to the\ndirectory cache, do you need to read from it? You could do just direct\ncat-files like for the other trees, and it would be even faster. Then,\nyou could do without a directory cache altogether in this mode.\n\n> PB> And this is another thing I dislike a lot. I'd like merge-tree.pl to\n> PB> leave my directory cache alone, thank you very much. You know, I see\n> PB> what goes to the directory cache as actually part of the policy part.\n> \n> Remember I am not touching *your* dircache.  It is a dircache in\n> the temporary merge area, specifically set up to help you review\n> the merge.  \n\nYes, but I want to have a control over its dircache too. :-) That is\nbecause I want the user to be able to use the regular git commands like\n\"git diff\" there.\n\n> Can't the SCM driver do things along this line, perhaps?\n> \n>  - You have your working files and your dircache.  They may not\n>    match because you have uncommitted changes to your\n>    environment.  You want to merge with Linus head.  You know\n>    its SHA1 (call it COMMIT-Linus).  Your SCM knows which commit\n>    you started with (call it COMMIT-Current).\n> \n>  - First you merge the tree associated with COMMIT-Current.  Use\n>    it and COMMIT-Linus to find the common ancestor to use.\n> \n>  - Now use the tree SHA of COMMIT-Current, tree SHA1 of\n>    COMMIT-Linus, and tree SHA1 of the common ancestor commit to\n>    drive git-merge.perl (to be renamed ;-).  You will get a\n>    temporary directory.  Have your user examine what is in\n>    there, and fix the merge and have them tell you they are\n>    happy.\n> \n>  - You go to that temporary directory, do write-tree and\n>    commit-tree with -p parameter of COMMIT-Linus and\n>    COMMIT-Current.  This will result in a new commit.  Call that\n>    COMMIT-Merge.\n> \n>  - You, as an SCM, should know what your user have done in the\n>    working directory relative to COMMIT-Current.  Especially you\n>    should know the set of paths involved in that change.  Go in\n>    to the temporary area, checkout-cache those files if you have\n>    not done so.  Apply the changes you have there.  Optionally\n>    have the user examine the changes and have him confirm.  Lift\n>    those files into the user's working directory.\n> \n>  - Do your bookkeeping like \"echo COMMIT-Merge >.git/Head\", to\n>    make the user's working files based on COMMIT-Merge, and run\n>    read-tree using the COMMIT-Merge in the user's working\n>    directory.  At this point, show-diff output should show what\n>    the changes your user have had made if he had started working\n>    based on COMMIT-Merge instead of starting from\n>    COMMIT-Current.\n> \n> I think the above would result in what SCM person would call\n> \"merge upstream/sidestream changes into my working directory\".\n\nAnd that's exactly what I'm doing now with git merge. ;-) In fact,\nideally the whole change in my scripts when your script is finished\nwould be replacing\n\n\tcheckout-cache `diff-tree` # symbolic\n\tgit diff $base $merged | git apply\n\nwith\n\n\tmerge-tree.pl -b $base $(tree-id) $merged | parse-your-output\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"174","messageId":"Pine.LNX.4.58.0504141728590.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"7v7jj5qgdz.fsf@assigned-by-dhcp.cox.net","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-15T00:42:30Z","receivedAt":"2005-04-15T00:42:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 14 Apr 2005, Junio C Hamano wrote:\n>\n> You say \"merge these two trees\" above (I take it that you mean\n> \"merge these two trees, taking account of this tree as their\n> common ancestor\", so actually you are dealing with three trees),\n\nYes. We're definitely talking three trees.\n\n> and I am tending to agree with the notion of merging trees not\n> commits.  However you might get richer context and more sensible\n> resulting merge if you say \"merge these two commits\".  Since\n> commit chaining is part of the fundamental git object model you\n> may as well use it.\n\nYes and no. There are real advantages to using the commit state to just \nfigure out the trees, and then at least have the _option_ to do the merge \nat a pure tree object.\n\nIn particular, if you ever find yourself wanting to graft together two\ndifferent commit histories, that almost certainly is what you'd want to\ndo. Somebody might have arrived at the exact same tree some other way,\nstarting with a 2.6.12 tar.ball or something, and I think we should at\nleast support the notion of saying \"these two totally unrelated commits\nactually have the same base tree, so let's merge them in \"space\" (ie data)\neven if we can't really sanely join them in \"time\" (ie \"commits\").\n\nI dunno.\n\nAnd it's also a question of sanity. The fact is, we know how to make tree \nmerges unambiguous, by just totally ignoring the history between them. Ie \nwe know how to merge data. I am pretty damn sure that _nobody_ knows how \nto merge \"data over time\". Maybe BK does. I'm pretty sure it actually \ntakes the \"over time\" into account. But My goal is to get something that \nworks, and something that is reliable because it is simple and it has \nsimple rules.\n\nAs you say:\n\n> This however opens up another set of can of worms---it would\n> involve not just three trees but all the trees in the commit\n> chain in between.\n\nExactly.  I seriously believe that the model is _broken_, simply because \nit gets too complicated. At some point it boils down to \"keep it simple, \nstupid\".\n\n>  That's when you start wondering if it would\n> be better to add renames in the git object model, which is the\n> topic of another thread.  I have not formed an opinion on that\n> one myself yet.\n\nI've not even been convinved that renames are worth it. Nobody has really \ngiven a good reason why.\n\nThere are two reasons for renames I can think of:\n\n - space efficiency in delta-based trees. This is a total non-issue for \n   git, and trying to explicitly track renames is going to cause _more_\n   space to be wasted rather than less.\n\n - \"annotate\". Something git doesn't really handle anyway, and it has \n   little to do with renames. You can fake an annotate, but let's face it, \n   it's _always_ going to be depending on interpreting a diff. In fact, \n   that ends up how traditional SCM's do it too - they don't really \n   annotate lines, they just interpret the diff.\n\n   I think you might as well interpret the whole object thing. Git _does_ \n   tell you how the objects changed, and I actually believe that a diff \n   that works in between objects (ie can show \"these lines moved from this\n   file X to tjhat file Y\") is a _hell_ of a lot more powerful than\n   \"rename\"  is.\n\n   So I'd seriously suggest that instead of worryign about renames, people \n   think about global diffs that aren't per-file. Git is good at limiting \n   the changes to a set of objects, and it should be entirely possible to \n   think of diffs as ways of moving lines _between_ objects and not just\n   within objects. It's quite common to move a function from one file to \n   another - certainly more so than renaming the whole file.\n\n   In other words, I really believe renames are just a meaningless special \n   case of a much more interesting problem. Which is just one reason why \n   I'm not at all interested in bothering with them other than as a \"data \n   moved\" thing, which git already handles very well indeed.\n\nSo there,\n\n\t\tLinus\n"},{"id":"176","messageId":"7v7jj4q2j2.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"20050414233159.GX22699@pasky.ji.cz","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T00:58:25Z","receivedAt":"2005-04-15T00:58:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n>> I think the above would result in what SCM person would call\n>> \"merge upstream/sidestream changes into my working directory\".\n\nPB> And that's exactly what I'm doing now with git merge. ;-) In fact,\nPB> ideally the whole change in my scripts when your script is finished\nPB> would be replacing\n\nPB> \tcheckout-cache `diff-tree` # symbolic\nPB> \tgit diff $base $merged | git apply\n\nPB> with\n\nPB> \tmerge-tree.pl -b $base $(tree-id) $merged | parse-your-output\n\nIn the above I presume by $merged you mean the tree ID (or\ncommit ID) the user's working directory is based upon?  Well,\nmerge-trees (Linus has a single directory merge-tree already)\nlooks at tree IDs (or commit IDs); it would never involve\nworking files in random state that is not recorded as part of a\ntree (committed or not).  Given that constraints I am not sure\nhow well that would pan out.  I have to think about this a bit.\n\nI do like, however, the idea of separating the step of doing any\ncheckout/merge etc. and actually doing them.  So the command set\nof parse-your-output needs to be defined.  Based on what I have\ndone so far, it would consist of the following:\n\n - Result is this object $SHA1 with mode $mode at $path (takes\n   one of the trees); you can do update-cache --cacheinfo (if\n   you want to muck with dircache) or cat-file blob (if you want\n   to get the file) or both.\n\n - Result is to delete $path.\n\n - Result is a merge between object $SHA1-1 and $SHA1-2 with\n   mode $mode-1 or $mode-2 at $path.\n\nWould this be a good enough command set?\n\nPB> Doesn't Emacs have something equivalent to ./.vimrc? I've also seen\nPB> those funny -*- strings.\n\nThe former is global per user (that is me including other Perl\nfiles I work outside of git context), which is exactly what I\nsaid is unacceptable to me.  The latter is per file (applying to\neverybody else who touch the file), so if it is short and sweet\nI should use one.\n\n"},{"id":"178","messageId":"002201c54157$80d32a90$6400a8c0@gandalf","threadId":"9","inReplyTo":"7vfyxtose8.fsf@assigned-by-dhcp.cox.net","subject":"Question about git process model","fromName":"Barry Silverman","fromEmail":"barry@disus.com","sentAt":"2005-04-15T01:07:25Z","receivedAt":"2005-04-15T01:07:25Z","isPatch":false,"sender":{"key":"barry@disus.com","avatar":null},"body":"JH->Junio Hamano, LT->Linus Torvalds\n\nJH>>I just have this fuzzy feeling that, when doing this merge:\n\n                     A-1 --- A-2 --- A-3\n                    /                   \\ \n    Common Ancestor                      Merge Result\n                    \\                   /\n                     B-1 --- B-2 --- B-3\n\nJH>>looking at diff(Common Ancestor, A-1), diff(Common Ancestor,\nJH>>B-1), diff(A-1, A-2), ... might give you richer context than\nJH>>just merging 3-way using Common Ancestor, A-3, and B-3 to derive\nJH>>the Merge Result.  It might not.  I honestly do not know.\n\nIn the distributed git model, with lots of parallel development, and\nlots of merging - Is it important at the \"business process\" level that\nintermediate history (and content) be present for any forks off the\nmainline (in particular, forks maintained by someone else)?\n\nGit has the property that it is NOT delta based. In Junio's example\nabove, if branch A were your mainline, and branch B were imported\nchanges from elsewhere, Is it necessary to have B1, B2 available to you,\nwhen all that was required for you to merge successfully was B3?\n\nWhich leads me to....\n\nLT>>In particular, if you ever find yourself wanting to graft together\ntwo LT>>different commit histories, that almost certainly is what you'd\nwant to LT>>do. Somebody might have arrived at the exact same tree some\nother way, LT>>starting with a 2.6.12 tar.ball or something, and I think\nwe should at LT>>least support the notion of saying \"these two totally\nunrelated commits LT>>actually have the same base tree, so let's merge\nthem in \"space\" (ie LT>>data) even if we can't really sanely join them\nin \"time\" (ie \"commits\").\n\nSo in the case of only merging B3 - we would write into the commit\nrecord of\nMerge-Result, that it was parented by A1, and another new commit record\n-> B3-Prime.\n\nB3-Prime is space-wise identical to B3, but \"time-wise\" different.\n\nB3-Prime would be different than B3 because it would necessarily not\nhave the same SHA1 as B3. Why? B3 has B2 as a parent, B3-Prime has the\nCommon Ancestor as a parent - thus the Commit record is different, and\nso is the SHA1. \n\nWe would need a facility to recognize that B3 and B3-Prime were\nspace-wise the same, (and maybe have the SHA for B3 kept in somewhere in\nthe SCM portion of B3-Prime's commit record???)\n\nLinus, is that what you were saying?\n\n\n\n"},{"id":"182","messageId":"7vzmw0ok45.fsf_-_@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504141133260.7211@ppc970.osdl.org","subject":"[Patch] ls-tree enhancements","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T02:21:30Z","receivedAt":"2005-04-15T02:21:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This adds '-r' (recursive) option and '-z' (NUL terminated)\noption to ls-tree.  I need it so that the merge-trees (formerly\nknown as git-merge.perl) script does not need to create any\ntemporary dircache while merging.  It used to use show-files on\na temporary dircache to get the list of files in the ancestor\ntree, and also used the dircache to store the result of its\nautomerge.  I probably still need it for the latter reason, but\nwith this patch not for the former reason anymore.\n\nIt is relative to bb95843a5a0f397270819462812735ee29796fb4\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n\n---\n\n ls-tree.c |  108 +++++++++++++++++++++++++++++++++++++++++++++++++++-----------\n 1 files changed, 90 insertions(+), 18 deletions(-)\n\n\n--- ,,Linus/ls-tree.c\t2005-04-14 19:08:17.000000000 -0700\n+++ ,,Siam/ls-tree.c\t2005-04-14 19:11:23.000000000 -0700\n@@ -5,45 +5,117 @@\n  */\n #include \"cache.h\"\n \n-static int list(unsigned char *sha1)\n+int line_termination = '\\n';\n+int recursive = 0;\n+\n+struct path_prefix {\n+\tstruct path_prefix *prev;\n+\tconst char *name;\n+};\n+\n+static void print_path_prefix(struct path_prefix *prefix)\n {\n-\tvoid *buffer;\n-\tunsigned long size;\n-\tchar type[20];\n+\tif (prefix) {\n+\t\tif (prefix->prev)\n+\t\t\tprint_path_prefix(prefix->prev);\n+\t\tfputs(prefix->name, stdout);\n+\t\tputchar('/');\n+\t}\n+}\n+\n+static void list_recursive(void *buffer,\n+\t\t\t  unsigned char *type,\n+\t\t\t  unsigned long size,\n+\t\t\t  struct path_prefix *prefix)\n+{\n+\tstruct path_prefix this_prefix;\n+\tthis_prefix.prev = prefix;\n \n-\tbuffer = read_sha1_file(sha1, type, &size);\n-\tif (!buffer)\n-\t\tdie(\"unable to read sha1 file\");\n \tif (strcmp(type, \"tree\"))\n \t\tdie(\"expected a 'tree' node\");\n+\n \twhile (size) {\n-\t\tint len = strlen(buffer)+1;\n-\t\tunsigned char *sha1 = buffer + len;\n-\t\tchar *path = strchr(buffer, ' ')+1;\n+\t\tint namelen = strlen(buffer)+1;\n+\t\tvoid *eltbuf;\n+\t\tchar elttype[20];\n+\t\tunsigned long eltsize;\n+\t\tunsigned char *sha1 = buffer + namelen;\n+\t\tchar *path = strchr(buffer, ' ') + 1;\n \t\tunsigned int mode;\n-\t\tunsigned char *type;\n \n-\t\tif (size < len + 20 || sscanf(buffer, \"%o\", &mode) != 1)\n+\t\tif (size < namelen + 20 || sscanf(buffer, \"%o\", &mode) != 1)\n \t\t\tdie(\"corrupt 'tree' file\");\n \t\tbuffer = sha1 + 20;\n-\t\tsize -= len + 20;\n+\t\tsize -= namelen + 20;\n+\n \t\t/* XXX: We do some ugly mode heuristics here.\n \t\t * It seems not worth it to read each file just to get this\n-\t\t * and the file size. -- pasky@ucw.cz */\n-\t\ttype = S_ISDIR(mode) ? \"tree\" : \"blob\";\n-\t\tprintf(\"%03o\\t%s\\t%s\\t%s\\n\", mode, type, sha1_to_hex(sha1), path);\n+\t\t * and the file size. -- pasky@ucw.cz\n+\t\t * ... that is, when we are not recursive -- junkio@cox.net\n+\t\t */\n+\t\teltbuf = (recursive ? read_sha1_file(sha1, elttype, &eltsize) :\n+\t\t\t  NULL);\n+\t\tif (! eltbuf) {\n+\t\t\tif (recursive)\n+\t\t\t\terror(\"cannot read %s\", sha1_to_hex(sha1));\n+\t\t\ttype = S_ISDIR(mode) ? \"tree\" : \"blob\";\n+\t\t}\n+\t\telse\n+\t\t\ttype = elttype;\n+\n+\t\tprintf(\"%03o\\t%s\\t%s\\t\", mode, type, sha1_to_hex(sha1));\n+\t\tprint_path_prefix(prefix);\n+\t\tfputs(path, stdout);\n+\t\tputchar(line_termination);\n+\n+\t\tif (eltbuf && !strcmp(type, \"tree\")) {\n+\t\t\tthis_prefix.name = path;\n+\t\t\tlist_recursive(eltbuf, elttype, eltsize, &this_prefix);\n+\t\t}\n+\t\tfree(eltbuf);\n \t}\n+}\n+\n+static int list(unsigned char *sha1)\n+{\n+\tvoid *buffer;\n+\tunsigned long size;\n+\tchar type[20];\n+\n+\tbuffer = read_sha1_file(sha1, type, &size);\n+\tif (!buffer)\n+\t\tdie(\"unable to read sha1 file\");\n+\tlist_recursive(buffer, type, size, NULL);\n \treturn 0;\n }\n \n+static void _usage(void)\n+{\n+\tusage(\"ls-tree [-r] [-z] <key>\");\n+}\n+\n int main(int argc, char **argv)\n {\n \tunsigned char sha1[20];\n \n+\twhile (1 < argc && argv[1][0] == '-') {\n+\t\tswitch (argv[1][1]) {\n+\t\tcase 'z':\n+\t\t\tline_termination = 0;\n+\t\t\tbreak;\n+\t\tcase 'r':\n+\t\t\trecursive = 1;\n+\t\t\tbreak;\n+\t\tdefault:\n+\t\t\t_usage();\n+\t\t}\n+\t\targc--; argv++;\n+\t}\n+\n \tif (argc != 2)\n-\t\tusage(\"ls-tree <key>\");\n+\t\t_usage();\n \tif (get_sha1_hex(argv[1], sha1) < 0)\n-\t\tusage(\"ls-tree <key>\");\n+\t\t_usage();\n \tsha1_file_directory = getenv(DB_ENVIRONMENT);\n \tif (!sha1_file_directory)\n \t\tsha1_file_directory = DEFAULT_DB_ENVIRONMENT;\n\n"},{"id":"183","messageId":"000001c54163$85540150$6400a8c0@gandalf","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504141728590.7211@ppc970.osdl.org","subject":"RE: Merge with git-pasky II.","fromName":"Barry Silverman","fromEmail":"barry@disus.com","sentAt":"2005-04-15T02:33:26Z","receivedAt":"2005-04-15T02:33:26Z","isPatch":false,"sender":{"key":"barry@disus.com","avatar":null},"body":">>In particular, if you ever find yourself wanting to graft together two\n>>different commit histories, that almost certainly is what you'd want\nto >>do. Somebody might have arrived at the exact same tree some other\nway, >>starting with a 2.6.12 tar.ball or something, and I think we\nshould at >>least support the notion of saying \"these two totally\nunrelated commits >>actually have the same base tree, so let's merge\nthem in \"space\" (ie data) >>even if we can't really sanely join them in\n\"time\" (ie \"commits\").\n\nIf this is true - then the tree-id's of the two commits would be\nidentical, but the commit-id's wouldn't.\n\nDoes this imply that common ancestor lookup should work by comparing the\ntree-id's (space-wise the same) rather than the commit-ids (time-wise\nthe same)?\n\n-----Original Message-----\nFrom: git-owner@vger.kernel.org [mailto:git-owner@vger.kernel.org] On\nBehalf Of Linus Torvalds\nSent: Thursday, April 14, 2005 8:43 PM\nTo: Junio C Hamano\nCc: Petr Baudis; git@vger.kernel.org\nSubject: Re: Merge with git-pasky II.\n\n\n\nOn Thu, 14 Apr 2005, Junio C Hamano wrote:\n>\n> You say \"merge these two trees\" above (I take it that you mean\n> \"merge these two trees, taking account of this tree as their\n> common ancestor\", so actually you are dealing with three trees),\n\nYes. We're definitely talking three trees.\n\n> and I am tending to agree with the notion of merging trees not\n> commits.  However you might get richer context and more sensible\n> resulting merge if you say \"merge these two commits\".  Since\n> commit chaining is part of the fundamental git object model you\n> may as well use it.\n\nYes and no. There are real advantages to using the commit state to just \nfigure out the trees, and then at least have the _option_ to do the\nmerge \nat a pure tree object.\n\nIn particular, if you ever find yourself wanting to graft together two\ndifferent commit histories, that almost certainly is what you'd want to\ndo. Somebody might have arrived at the exact same tree some other way,\nstarting with a 2.6.12 tar.ball or something, and I think we should at\nleast support the notion of saying \"these two totally unrelated commits\nactually have the same base tree, so let's merge them in \"space\" (ie\ndata)\neven if we can't really sanely join them in \"time\" (ie \"commits\").\n\nI dunno.\n\nAnd it's also a question of sanity. The fact is, we know how to make\ntree \nmerges unambiguous, by just totally ignoring the history between them.\nIe \nwe know how to merge data. I am pretty damn sure that _nobody_ knows how\n\nto merge \"data over time\". Maybe BK does. I'm pretty sure it actually \ntakes the \"over time\" into account. But My goal is to get something that\n\nworks, and something that is reliable because it is simple and it has \nsimple rules.\n\nAs you say:\n\n> This however opens up another set of can of worms---it would\n> involve not just three trees but all the trees in the commit\n> chain in between.\n\nExactly.  I seriously believe that the model is _broken_, simply because\n\nit gets too complicated. At some point it boils down to \"keep it simple,\n\nstupid\".\n\n>  That's when you start wondering if it would\n> be better to add renames in the git object model, which is the\n> topic of another thread.  I have not formed an opinion on that\n> one myself yet.\n\nI've not even been convinved that renames are worth it. Nobody has\nreally \ngiven a good reason why.\n\nThere are two reasons for renames I can think of:\n\n - space efficiency in delta-based trees. This is a total non-issue for \n   git, and trying to explicitly track renames is going to cause _more_\n   space to be wasted rather than less.\n\n - \"annotate\". Something git doesn't really handle anyway, and it has \n   little to do with renames. You can fake an annotate, but let's face\nit, \n   it's _always_ going to be depending on interpreting a diff. In fact, \n   that ends up how traditional SCM's do it too - they don't really \n   annotate lines, they just interpret the diff.\n\n   I think you might as well interpret the whole object thing. Git\n_does_ \n   tell you how the objects changed, and I actually believe that a diff \n   that works in between objects (ie can show \"these lines moved from\nthis\n   file X to tjhat file Y\") is a _hell_ of a lot more powerful than\n   \"rename\"  is.\n\n   So I'd seriously suggest that instead of worryign about renames,\npeople \n   think about global diffs that aren't per-file. Git is good at\nlimiting \n   the changes to a set of objects, and it should be entirely possible\nto \n   think of diffs as ways of moving lines _between_ objects and not just\n   within objects. It's quite common to move a function from one file to\n\n   another - certainly more so than renaming the whole file.\n\n   In other words, I really believe renames are just a meaningless\nspecial \n   case of a much more interesting problem. Which is just one reason why\n\n   I'm not at all interested in bothering with them other than as a\n\"data \n   moved\" thing, which git already handles very well indeed.\n\nSo there,\n\n\t\tLinus\n-\nTo unsubscribe from this list: send the line \"unsubscribe git\" in\nthe body of a message to majordomo@vger.kernel.org\nMore majordomo info at  http://vger.kernel.org/majordomo-info.html\n\n\n"},{"id":"200","messageId":"20050415062807.GA29841@64m.dyndns.org","threadId":"9","inReplyTo":"7vfyxsmqmk.fsf@assigned-by-dhcp.cox.net","subject":"Re: Merge with git-pasky II.","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-15T06:28:07Z","receivedAt":"2005-04-15T06:28:07Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Fri, Apr 15, 2005 at 12:43:47AM -0700, Junio C Hamano wrote:\n> >>>>> \"CL\" == Christopher Li <git@chrisli.org> writes:\n> \n> CL> Is that SHA1 for tree or the file object?\n> \n> I am talking about a single file here.\n>\nThen do you emit the entry for it's parents directory?\n\ne.g. /foo/bar get created. foo doesn't exists. You have\nto create foo first. You don't have mode information for\nfoo yet. If it give the top level tree, the SCM can check it\nout by tree. hopefully have the mode on directory correctly.\nWell, if they care about those little details.\n\nChris\n \n"},{"id":"194","messageId":"7vfyxsmqmk.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"20050414223039.GB28082@64m.dyndns.org","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T07:43:47Z","receivedAt":"2005-04-15T07:43:47Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"CL\" == Christopher Li <git@chrisli.org> writes:\n\n>> - Result is this object $SHA1 with mode $mode at $path (takes\n>> one of the trees); you can do update-cache --cacheinfo (if\n>> you want to muck with dircache) or cat-file blob (if you want\n>> to get the file) or both.\n\nCL> Is that SHA1 for tree or the file object?\n\nI am talking about a single file here.\n\n\n"},{"id":"196","messageId":"1113556448.12012.269.camel@baythorne.infradead.org","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504141133260.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-15T09:14:08Z","receivedAt":"2005-04-15T09:14:08Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Thu, 2005-04-14 at 11:36 -0700, Linus Torvalds wrote:\n> And \"merge these two trees\" (which works on a _tree_ level)\n> or \"find the common commit\" (which works on a _commit_ level)\n\nI suspect that finding the common commit is actually a per-file thing;\nit's not just something you do for the _commit_ graph, then use for\nmerging each file in the two branches you're trying to merge.\n\nConsider a simple repository which contains two files A and B. We start\noff with the first version of each ('A1B1'), and the owner of each file\ntakes a branch and modifies their own file. There is cross-pulling\nbetween the two, and then each modifies the _other's_ file as well as\ntheir own...\n\n   (A1B2)--(A2B2)--(A2'B3)\n    /  \\   /            \\\n   /    \\ /              \\\n (A1B1)  X               (...)\n   \\    / \\              /\n    \\  /   \\            /\n   (A2B1)--(A2B2)--(A3B2')\n\nNow, we're trying to merge the two branches. It appears that the most\nuseful common ancestor to use for a three-way merge of file A is the\nversion from tree 'A2B1', while the most useful common ancestor for\nmerging file B is that in 'A1B2'.\n\n(I think it's a coincidence that in my example the useful files 'A2' and\n'B2' actually do end up in a single tree together at some point.)\n\n-- \ndwmw2\n\n\n"},{"id":"199","messageId":"20050415093649.GA28077@elte.hu","threadId":"9","inReplyTo":"1113556448.12012.269.camel@baythorne.infradead.org","subject":"Re: Merge with git-pasky II.","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-15T09:36:49Z","receivedAt":"2005-04-15T09:36:49Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* David Woodhouse <dwmw2@infradead.org> wrote:\n\n> Consider a simple repository which contains two files A and B. We \n> start off with the first version of each ('A1B1'), and the owner of \n> each file takes a branch and modifies their own file. There is \n> cross-pulling between the two, and then each modifies the _other's_ \n> file as well as their own...\n> \n>    (A1B2)--(A2B2)--(A2'B3)\n>     /  \\   /            \\\n>    /    \\ /              \\\n>  (A1B1)  X               (...)\n>    \\    / \\              /\n>     \\  /   \\            /\n>    (A2B1)--(A2B2)--(A3B2')\n> \n> Now, we're trying to merge the two branches. It appears that the most \n> useful common ancestor to use for a three-way merge of file A is the \n> version from tree 'A2B1', while the most useful common ancestor for \n> merging file B is that in 'A1B2'.\n\ndo such cases occur frequently? In the kernel at least it's not too \ntypical. Would it be a problem to go for the simple solution of using \n(A1B1) as the common ancestor (based on the tree graph), and then to do \na 3-way merge of all changes from that point on?\n\n\tIngo\n"},{"id":"201","messageId":"1113559330.12012.292.camel@baythorne.infradead.org","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504141728590.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-15T10:02:10Z","receivedAt":"2005-04-15T10:02:10Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Thu, 2005-04-14 at 17:42 -0700, Linus Torvalds wrote:\n> I've not even been convinved that renames are worth it. Nobody has\n> really given a good reason why.\n> \n> There are two reasons for renames I can think of:\n> \n>  - space efficiency in delta-based trees.\n>  - \"annotate\".\n\nNeither of those were my motivation for looking at renames. The reasons\nI wanted to track renames were:\n   - Per-file revision history which doesn't stop dead at a rename.\n   - Merging where files have been renamed in one branch and modified in\n     another. Which is basically a special case of the above; we need to\n     see the per-file revision history.\n\n>    So I'd seriously suggest that instead of worryign about renames, people \n>    think about global diffs that aren't per-file. Git is good at limiting \n>    the changes to a set of objects, and it should be entirely possible to \n>    think of diffs as ways of moving lines _between_ objects and not just\n>    within objects. It's quite common to move a function from one file to \n>    another - certainly more so than renaming the whole file.\n>\n>    In other words, I really believe renames are just a meaningless special \n>    case of a much more interesting problem. Which is just one reason why \n>    I'm not at all interested in bothering with them other than as a \"data \n>    moved\" thing, which git already handles very well indeed.\n\nGit doesn't handle 'data moved' except at a whole-tree level. For each\ncommit, it says \"these are the old trees; this is the new tree\".\n\nGit doesn't actually look hard into the contents of tree; certainly it\nhas no business looking at the contents of individual files; that is\nsomething that the SCM or possibly only the user should do. The storage\nof 'rename' information in the commit object is another kind of 'xattr'\nstorage which git would provides but not directly interpret.\n\nAnd you're right; it shouldn't have to be for renames only. There's no\nneed for us to limit it to one \"source\" and one \"destination\"; the SCM\ncan use it to track content as it sees fit.\n\nAs I said, the main aim of this is to track revision history of given\ncontent, for displaying to the user and for performing merges. So when a\nfile is split up, or a function is moved from it to another file, a\n'rename' xattr can be included to mark that files 'foo' and 'bar' in the\nnew tree are both associated with file 'wibble' in the parent.\n\nThat's as much as we need to provide for content tracking, and it _does_\nhandle the general case as well as we should be attempting to. We don't\nwant to get into dealing with file contents ourselves; we just want to\nstore the hint for the SCM or the user that \"your data went thataway\".\n\n-- \ndwmw2\n\n\n"},{"id":"202","messageId":"1113559533.12012.296.camel@baythorne.infradead.org","threadId":"9","inReplyTo":"20050415093649.GA28077@elte.hu","subject":"Re: Merge with git-pasky II.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-15T10:05:33Z","receivedAt":"2005-04-15T10:05:33Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Fri, 2005-04-15 at 11:36 +0200, Ingo Molnar wrote:\n> do such cases occur frequently? In the kernel at least it's not too \n> typical. \n\nIsn't it? I thought it was a fairly accurate representation of the\nprocess \"I make a whole bunch of changes to files I maintain, pulling\nfrom Linus while occasionally asking him to pull from my tree. Sometimes\nmy files are changed by someone else in Linus' tree, and sometimes I\nchange files that I don't actually own.\".\n\n-- \ndwmw2\n\n\n"},{"id":"211","messageId":"20050415102224.GA8924@thunk.org","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504151340221.27162@wgmdd8.biozentrum.uni-wuerzburg.de","subject":"Re: Merge with git-pasky II.","fromName":"Theodore Ts'o","fromEmail":"tytso@thunk.org","sentAt":"2005-04-15T10:22:24Z","receivedAt":"2005-04-15T10:22:24Z","isPatch":false,"sender":{"key":"tytso@thunk.org","avatar":"https://gravatar.com/avatar/bc16cd364de8c963cca27953f33db8cd94de51b67922a7697620d8512a05ab01?d=mp&s=160"},"body":"On Fri, Apr 15, 2005 at 02:03:08PM +0200, Johannes Schindelin wrote:\n> I disagree. In order to be trusted, this thing has to catch the following\n> scenario:\n> \n> Skywalker and Solo start from the same base. They commit quite a lot to\n> their trees. In between, Skywalker commits a tree, where the function\n> \"kazoom()\" has been added to the file \"deathstar.c\", but Solo also added\n> this function, but to the file \"moon.c\". A file-based merge would have no\n> problem merging each file, such that in the end, \"kazoom()\" is defined\n> twice.\n> \n> The same problems arise when one tries to merge line-wise, i.e. when for\n> each line a (possibly different) merge-parent is sought.\n\nBe careful.  There is a very big tradeoff between 100% perfections in\ncatching these sorts of errors, and usability.  There exists SCM's\nwhere you are not allowed to do commit such merges until you do a test\ncompile, or run a regression test suite (that being the only way to\ncatch these sorts of problems when we merge two branches like this).  \n\nBitKeeper never caught this sort of thing, and we trusted it.  In\npractice it was also rarely a problem.\n\nI'll also note that BitKeeper doesn't restrict you from doing a\ncommitting a changeset when you have modified files that have yet to\nbe checked in to the tree.  Same issue; you can accidentally check in\nchangesets that in trees that won't build, but if we added this kind\nof SCM-by-straightjacket philosophy it would decrease our productivity\nand people would simply not use such an SCM, thus negating its\neffectiveness.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"203","messageId":"7vwtr4ibkt.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"20050414233159.GX22699@pasky.ji.cz","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T10:22:26Z","receivedAt":"2005-04-15T10:22:26Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"After I re-read [*R1*], in which Linus talks about dircache,\nespecially this section:\n\n - The \"current directory cache\" describes some baseline. In particular,\n   note the \"some\" part. It's not tied to any special baseline, and you\n   can change your baseline any way you please.\n\n   So it does NOT have to track any particular state in either the object \n   database _or_ in your actual current working tree. In fact, all real \n   interactions with \"git\" are really about updating this staging area one \n   way or the other: you might check out the state from it into your \n   working area (partially or fully), you can push your working area into \n   the staging area (again, partially or fully).\n\n   And if you want to, you can write the thing that the staging area \n   represents as a \"tree\" into the object database, or you can merge a \n   tree from the object database into the staging area.\n\n   In other words: the staging area aka \"current directory cache\" is \n   really how all interaction takes place. The object database never \n   interacts directly with your working directory contents. ALL \n   interactions go through the current directory cache.\n\nI started to have more doubts on the approach of *not*\nperforming the merge in the dircache I set up specifically for\nmerging, which is the direction in which you are pushing if I\nunderstand you correctly.  Maybe I completely misunderstand what\nyou want.  This message is long but I need a clear understanding\nof what is expected to be useful to you, so please bear with me.\n\nPB> \tmerge-tree.pl -b $base $(tree-id) $merged | parse-your-output\n\nPlease help me understand this example you have given earlier.\nHere is my understanding of your assumption when the above\npipeline takes place.  Correct me if I am mistaken.\n\n * The user is in a working directory $W.  It is controlled by\n   git-tools and there are $W/.git/. directory and $W/.git/index\n   dircache.\n\n * The dircache $W/.git/index started its life as a read-tree\n   from some commit.  The git-tools is keeping track of which\n   commit it is somewhere, presumably in $W/.git/ directory.\n   Let's call it $C (commit).\n\n ? Question.  Is the $(tree-id) in your example the same as $C\n   above?\n\n * The user have run [*1*] (see Footnote below) checkout-cache\n   on $W/.git/index some time in the past and $W is full of\n   working files.  Some of them may or may not have modified.\n   There may be some additions or deletions.  So the contents of\n   the working directory may not match the tree associated with\n   $C.\n\n * The user may or may not have run [*1*] update-cache in $W.\n   The contents of the dircache $W/.git/index may not match the\n   tree associated with $C.\n\n ? Question.  Are you forbidding the user to run update-cache by\n   hand, and keeping track of the changes yourself, to be\n   applied all at once at \"git commit\" time, thereby\n   guaranteeing the $W/.git/index to match the tree associated\n   with $C all times?  From the description of The \"GIT toolkit\"\n   section in README, it is not clear to me which part of his\n   repository an end user is not supposed to muck with himself.\n\n * Now the user has some changes in his working directory and\n   notices upstream or a side branch has notable changes\n   desireble to be picked up.  So he runs some git-tools command\n   to cause the above quoted pipeline to run.\n\n ? Question.  Does $merged in your example mean such an upstream\n   or side branch?  Is $base in your example the common ancestor\n   between $C and $merged?\n\nAssuming that my above understanding of your model is correct,\nhere are my \"thinking aloud\".\n\n - \"merge-trees $base $C $merged\" looks only at the git object\n   database for those three trees named.  The data structure of\n   git object database is optimized to distinguish differences\n   in those recorded trees (and hence recorded blobs they point\n   at) without unpacking most of the files if the changes are\n   small, because all the blobs involved are already hashed.  It\n   is not very good at comparing things in git object store and\n   working files in random states, which would involve unpacking\n   blobs and comparing, so \"merge-trees\" does not bother.\n\n - What can come out from merge-trees is therefore one of the\n   following for each path from the union of paths contained in\n   $base, $C, and $merged:\n\n   (a) Neither $C nor $merged changed it --- merge result is what\n       is in $C.\n\n   (b) $C changed it but $merged did not --- merge result is what\n       is in $C.\n\n   (c) Both $C and $merged changed it in the same way --- merge\n       result is what is in $C.\n\n   (d) $C did not change it but $merged did --- merge result is\n       what is in $merged.\n\n   (e) Both $C and $merged changed it differently --- merge is\n       needed and automatically succeeds between $C and $merge.\n\n   (f) Both $C and $merged changed it differently --- merge is\n       needed but have conflicts.\n\n - Assuming we are dealing with the case where working files are\n   dirty and do not match what is in $C, among the above,\n   (a)-(c) can be ignored by SCM.  What the user has in his\n   working files is exactly what he would have got if he started\n   working from the merge result, although in reality the work\n   was started from $C.\n\n   Handling (d), (e) and (f) from SCM's point of view would be\n   the same.  They all involve 3-way merges between the file in\n   the working directory, and the file from $merged, pivoting on\n   the file from $base.  In order to help SCM, merge-trees\n   therefore should output SHA1 of blobs for such a file from\n   $base and $merged and expect SCM to run \"cat-file blob\" on\n   them and then merge or diff3.  Up to the point of giving\n   those two SHA1 out is the business of merge-trees and after\n   that it is up to SCM.\n\n   That would work.  So I should base the design of output from\n   merge-trees on the above analysis, which probably needs to be\n   extended to cover differences between creation, modification,\n   and deletion.\n\n - However, the above is quite different from the way Linus\n   envisioned initially, on which my current implementation is\n   based [*3*].\n\n   My current implementation is to record the merge outcome in\n   the temporary dircache $W/,,merge/.git/index for cases\n   (a)-(e).  The last case (f) is problematic and needs human\n   validation [*2*], so it is not recorded in that temporary\n   dircache, but the files to be merged are left in that\n   temporary directory and merge-trees stops there.  It is\n   expected that the end-user or SCM would merge the resulting\n   file and run update-cache to update $W/,,merge/.git/index.\n   After that happens, $W/,,merge/.git/index has the tree\n   representing the desired result of the merge.  It is expected\n   that the end-user or SCM would write-tree, commit-tree there\n   in the temporary directory, creating a new commit $C1.\n\n   Then, it is expected that the SCM would make a patch file\n   between $C and the user working directory, checks out $C1\n   (either in the user's working directory or another temporary\n   directory; at this point merge-trees does not care because it\n   has already done its job and exited), applies that patch to\n   bring the user edits over to $C1.  Then that directory would\n   contain the desired merge of user edits.\n\n   That is my understanding of how Linus originally wanted the\n   tool to do his kernel work with to work.  My hesitation to\n   suggestions from you to change it not to keep its own merge\n   dircache is coming from here.  Not doing what I am currently\n   doing to $W/,,merge/.git/index dircache would mean that SCM\n   would have to do more, not less, to arrive at $C1 (the result\n   of the clean $merge and $C merge pivoted at $base), where the\n   real SCM merge begins.\n\nAlthough I suspect I am misunderstanding what you want, your\nmessages so far suggest that what you want might be quite\ndifferent from what Linus wants.  Please do not misunderstand\nwhat I mean by saying this.  I am not saying that Linus is\nalways right [*4*] and therefore you are wrong for wanting\nsomething else.  It is just that, if what I started writing\nneeds to support both of those quite different needs, I need to\nknow what they are.  I think I understand what Linus wants well\nenough [*5*], but I am not certain about yours.\n\n\n[Footnotes]\n\n*1* By \"The user have run\" I mean either the user directly used\nthe low-level plumbing command himself, or used git-tools to\ncause such command to run.\n\n*2* Strictly speaking, case (e) needs human validation as\nwell, because successful textual merge does not guarantee\nsensible semantic merge.\n\n*3* See [*R2*] for descriptions on the way Linus wanted merge\nin git to happen.  Especially around \"5) At this point you need\nto MERGE\" onwards.  The current implementation handles (or\nattempts to handle) the `your working directory was fully\ncommitted' case described there.\n\n*4* According to Linus himself, he is always right ;-). [*R3*]\n\n*5* I consider [*R1*] and [*R2*] essential read for anybody\nwanting to understand merging operation in git object model (I\nam saying this for others; not for Pasky --- it would be like\npreaching to the choir ;-)).\n\n\n[References]\n\n*R1* <Pine.LNX.4.58.0504110928360.1267@ppc970.osdl.org>\nhttp://marc.theaimsgroup.com/?i=%3CPine.LNX.4.58.0504110928360.1267%20()%20ppc970%20!%20osdl%20!%20org%3E\n\n*R2* <Pine.LNX.4.58.0504121606580.4501@ppc970.osdl.org> \nhttp://marc.theaimsgroup.com/?i=%3CPine.LNX.4.58.0504121606580.4501%20()%20ppc970%20!%20osdl%20!%20org%3E\n\n*R3*\nhttp://www.uwsg.indiana.edu/hypermail/linux/kernel/0008.3/0555.html\n\n"},{"id":"204","messageId":"7vfyxsi9bq.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"20050415062807.GA29841@64m.dyndns.org","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T11:11:05Z","receivedAt":"2005-04-15T11:11:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"CL\" == Christopher Li <git@chrisli.org> writes:\n\nCL> Then do you emit the entry for it's parents directory?\n\nIn GIT object model, directory modes do not matter.  It is not\ndesigned to record directories, and running \"update-cache --add\nfoo\" when foo is a directory fails.\n\nThe data model of GIT is that it associates file datablob to a\nstring called \"pathname\" that happen to contain slashes in them.\nIt is kinda wierd.  When you externalize it with checkout-cache,\nthese slashes are mapped to hierarchical UNIX filesystem paths,\nrelative to whereever you happened to run checkout-cache.  The\nhierarchical \"tree\" representation in the GIT database was\nstarted as just a space optimization thing.\n\nCL> e.g. /foo/bar get created. foo doesn't exists. You have\nCL> to create foo first. You don't have mode information for\nCL> foo yet.\n\nAnd you will never have that information, since it is not\nrecorded anywhere.  If I say you should have foo/bar (by the\nway, no leading slashes are placed in the dircache either), and\nif it so happens that you do not have foo yet, you'd better\ncreate one without waiting to be told, because I will never tell\nyou to just create a directory.\n\nBy the way, Linus, while I was studying how the new hierarchical\ntrees are written out, I think I have found one small funny (I\nwould not call this a *bug*) there.  Here is an excerpt from\nwrite-tree (around ll. 56; I am basing on pasky-0.4 so your line\nnumbers may have some offsets):\n\n        sha1 = ce->sha1;\n        mode = ntohl(ce->st_mode);\n\n        /* Do we have _further_ subdirectories? */\n        filename = pathname + baselen;\n        dirname = strchr(filename, '/');\n        if (dirname) {\n                int subdir_written;\n\n                subdir_written = write_tree(cachep + nr, maxentries - nr, pathname, dirname-pathname+1, subdir_sha1);\n                nr += subdir_written;\n\n                /* Now we need to write out the directory entry into this tree.. */\n                mode |= S_IFDIR;\n                pathlen = dirname - pathname;\n\n                /* ..but the directory entry doesn't count towards the total count */\n                nr--;\n                sha1 = subdir_sha1;\n        }\n\nThis code is going through a flat list of cache entries sorted\nby pathnames.  The list is flat in the sense that the pathnames\nare like \"foo/bar\" i.e. with slashes inside.  The if() statement\nthere, upon seeing \"foo/bar\", slurps all the entries in foo/\nsubhierarchy and writes into a separate tree, recursively, to\n\"represent\" foo/.\n\nNotice what mode the \"tree\" object gets in this case?  File mode\nfor foo/bar (or whatever happens to be sorted the first among\nthe stuff in dircache from foo/ directory) ORed with S_IFDIR.  I\nthink this is nonsense, and we should just store constant\nS_IFDIR.\n\nAnother option, probably better from the SCM purist's POV, would\nbe to start recording directories in dircaches, so that people\ncan actually keep track of directory modes.  Does it matter? ---\nI would say not.  GIT does not have to be tar or cpio.  \n\n"},{"id":"208","messageId":"Pine.LNX.4.58.0504151340221.27162@wgmdd8.biozentrum.uni-wuerzburg.de","threadId":"9","inReplyTo":"1113556448.12012.269.camel@baythorne.infradead.org","subject":"Re: Merge with git-pasky II.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2005-04-15T12:03:08Z","receivedAt":"2005-04-15T12:03:08Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 15 Apr 2005, David Woodhouse wrote:\n\n> On Thu, 2005-04-14 at 11:36 -0700, Linus Torvalds wrote:\n> > And \"merge these two trees\" (which works on a _tree_ level)\n> > or \"find the common commit\" (which works on a _commit_ level)\n>\n> I suspect that finding the common commit is actually a per-file thing;\n> it's not just something you do for the _commit_ graph, then use for\n> merging each file in the two branches you're trying to merge.\n\nI disagree. In order to be trusted, this thing has to catch the following\nscenario:\n\nSkywalker and Solo start from the same base. They commit quite a lot to\ntheir trees. In between, Skywalker commits a tree, where the function\n\"kazoom()\" has been added to the file \"deathstar.c\", but Solo also added\nthis function, but to the file \"moon.c\". A file-based merge would have no\nproblem merging each file, such that in the end, \"kazoom()\" is defined\ntwice.\n\nThe same problems arise when one tries to merge line-wise, i.e. when for\neach line a (possibly different) merge-parent is sought.\n\nThe concept here is a *transaction*: when going from one tree to the next\ntree via a commit, a sort of integrity is maintained, which is breached\nwhen only looking at files and commits.\n\nCiao,\nDscho\n\n"},{"id":"213","messageId":"Pine.LNX.4.58.0504150740310.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"1113556448.12012.269.camel@baythorne.infradead.org","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-15T14:53:16Z","receivedAt":"2005-04-15T14:53:16Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, David Woodhouse wrote:\n> \n> I suspect that finding the common commit is actually a per-file thing;\n> it's not just something you do for the _commit_ graph, then use for\n> merging each file in the two branches you're trying to merge.\n\nI disagree.\n\nConceptually, you should never do _anything_ on a file level. Why? Because\nindividual files don't matter. You shouldn't merge two files cleanly just\nbecause they look fine - they _depend_ on the other files in the archive, \nand that's quite fundamentally why per-file tracking is really wrong from \na project standpoint.\n\nSo if you can't merge two files cleanly because the \"project\" history \nended up being further back than the \"file\" history, then that's a _good_ \nthing. You don't know what the hell happened to the other files that this \nfile depended on. Merging one file independently of the others is WRONG.\n\nAlso, I suspect that you'll find that if you do cross-merges, you'll \nbasically always end up in:\n\n> (I think it's a coincidence that in my example the useful files 'A2' and\n> 'B2' actually do end up in a single tree together at some point.)\n\nnope, I don't think that's coincidence. I think that's the normal case. \nYour file-based history is the one that can _incorrectly_ and \ncoincidentally happen to have a single file at some point, but since that \nfile doesn't stand alone, that's really not a fundamentally good reason to \nmerge it.\n\nReally, this \"individual files matter\" approach is a _disease_. They \ndon't. Individual files DO NOT EXIST. Files always exist as part of the \nproject, and the _only_ time you track a single file is when the project \nis a single file (and then that will be very very obvious in a git \narchive, thank you very much).\n\nSo the single-file mentality is a disease brought on by decades of _crap_. \nAnd by the fact that it ends up limiting the problem scope, so you can do \ncertain things easier.\n\nFor example, just doing intra-file diffs is a lot _easier_ and less \ntime-consuming than doing inter-file diffs. Bit it is _absolutely_ not \nbetter. In fact, it is clearly inferior to anybody who spends even five \nseconds thinking about it - yet we still do it, because of the historical \n(and INCORRECT) mindset that \"files matter\".\n\nFiles DO NOT matter. Never have. It's an implementation limitation to \nthink they do. You'll screw yourself up, and when somebody comes up with a \nhalf-way efficient way to generate inter-fiel diffs, your architecture is \ntotally and utterly unable to handle it.\n\nI don't care what you do at an SCM level, and if the crud you put on top\nof git wants to perpetuate mistakes of yesteryear, that's _your_ issue.  \nBut dammit, git is designed to do the right thing, and I will fight tooth\nand nail against anybody who thinks individual files matter.\n\n\t\tLinus\n"},{"id":"214","messageId":"20050415145324.GA4677@elte.hu","threadId":"9","inReplyTo":"1113559533.12012.296.camel@baythorne.infradead.org","subject":"Re: Merge with git-pasky II.","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-15T14:53:24Z","receivedAt":"2005-04-15T14:53:24Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* David Woodhouse <dwmw2@infradead.org> wrote:\n\n> On Fri, 2005-04-15 at 11:36 +0200, Ingo Molnar wrote:\n> > do such cases occur frequently? In the kernel at least it's not too \n> > typical. \n> \n> Isn't it? I thought it was a fairly accurate representation of the \n> process \"I make a whole bunch of changes to files I maintain, pulling \n> from Linus while occasionally asking him to pull from my tree. \n> Sometimes my files are changed by someone else in Linus' tree, and \n> sometimes I change files that I don't actually own.\".\n\nbut the specific scenario you described would require _Linus'_ tree to \nbe in limbo for a long time, and have uncommitted half-done edits. I.e.:\n\n   (A1B2)--(A2B2)--(A2'B3)\n    /  \\   /            \\\n   /    \\ /              \\\n (A1B1)  X               (...)\n   \\    / \\              /\n    \\  /   \\            /\n   (A2B1)--(A2B2)--(A3B2')\n\nin the above scenario Linus' tree needs to 'cross' with a maintainer's \ntree.  (maintainer's tree wont cross with another maintainer's tree, as \nmaintainer-to-maintainer merges rare.)\n\nbut for the scenario to occur, i think there needs to be a prolongued \n\"limbo\" period in Linus' tree for a 'crossing' to happen. But Linus' \nmerges are typically almost atomic: they are done then they are pushed \nout. It's definitely not in the 'days, sometimes weeks' timescale as \nmaintainer trees are.\n\nso for the scenario to occur, a maintainer, from whom Linus has just \npulled an update and Linus is merging the tree manually without \ncomitting, has to pull a file from the earlier Linus tree, and then \nLinus has to modify that same file again. This does not seem to be a \ncommon scenario.\n\nso i think to avoid the scenario, maintainers should not pull from each \nother - they should only pull/push to/from Linus' tree. Maybe this is an \nunacceptable limitation?\n\n\tIngo\n"},{"id":"215","messageId":"1113577744.27227.53.camel@hades.cambridge.redhat.com","threadId":"9","inReplyTo":"20050415145324.GA4677@elte.hu","subject":"Re: Merge with git-pasky II.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-15T15:09:03Z","receivedAt":"2005-04-15T15:09:03Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Fri, 2005-04-15 at 16:53 +0200, Ingo Molnar wrote:\n> but the specific scenario you described would require _Linus'_ tree to\n> be in limbo for a long time, and have uncommitted half-done edits.\n> I.e.:\n> \n>    (A1B2)--(A2B2)--(A2'B3)\n>     /  \\   /            \\\n>    /    \\ /              \\\n>  (A1B1)  X               (...)\n>    \\    / \\              /\n>     \\  /   \\            /\n>    (A2B1)--(A2B2)--(A3B2')\n> \n> in the above scenario Linus' tree needs to 'cross' with a maintainer's\n> tree.  (maintainer's tree wont cross with another maintainer's tree,\n> as maintainer-to-maintainer merges rare.)\n\nIs that true? Consider (A2B1) to be a bugfixes-only tree which I make\navailable for Linus to pull from. I keep doing more experimental stuff\nin my own private copy of the tree along the bottom branch, while Linus\n_eventually_ responds to my pull request and moves on, stopping only to\nadd a 'static' to one of my new functions. I move on too but don't pull\nfrom Linus again for a little while; the final merge happens when I _do_\npull again.\n\n-- \ndwmw2\n\n"},{"id":"216","messageId":"1113578964.27227.65.camel@hades.cambridge.redhat.com","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504150740310.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-15T15:29:24Z","receivedAt":"2005-04-15T15:29:24Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Fri, 2005-04-15 at 07:53 -0700, Linus Torvalds wrote:\n> Files DO NOT matter. Never have. It's an implementation limitation to \n> think they do. You'll screw yourself up, and when somebody comes up with a \n> half-way efficient way to generate inter-fiel diffs, your architecture is \n> totally and utterly unable to handle it.\n> \n> I don't care what you do at an SCM level, and if the crud you put on top\n> of git wants to perpetuate mistakes of yesteryear, that's _your_ issue.  \n> But dammit, git is designed to do the right thing, and I will fight tooth\n> and nail against anybody who thinks individual files matter.\n\nNo, really: individual files _DO_ matter. There's a reason we split\nstuff up into separate files, and if you look closely you'll find that\nwe don't just randomly put different functions into different files with\nneither rhyme nor reason -- there's a pattern to it; usually some kind\nof functional grouping.\n\nAnd when I'm looking for the change that broke something, I can almost\nalways tell which file it's in and go looking in _that_ file. It's a\n_whole_ lot easier to use the equivalent of 'bk revtool' than it is to\nsift through all the unrelated commits in the whole tree. If that's an\nimplementation limitation, then it's an implementation limitation in my\n_brain_ not just in my tools.\n\nOK, in fact it shouldn't be 'show me the history of this file'; it's\noften really 'show me the history of this function' which I want. But\nthat's fine. All I'm suggesting is that we should include the metadata\nwhich says \"content moved from file XXX to file YYY\" along with the\ncommit objects.\n\nI'm certainly not suggesting that we should implement jejb's idea of\nexplicit 'file revision history' objects -- the tree-based philosophy is\nperfectly sane and sufficient. But we do _also_ need a little\ninformation which allows us to track content as it moves around within\nthe tree, and the SCM has to have a sane way to filter out the noise\nwhen we're looking for what broke. Yes, that's part of the SCM\nfunctionality, and can live in an xattr-type field in the commit object\n-- but it does need to be stored, and in practice I suspect it _will_ be\nuseful for merging too.\n\nIt's not about ditching the per-tree tracking and doing per-file\ntracking instead. I agree that would be wrong. It's about storing enough\ninformation to track what happened to given content as it moved around\nwithin the tree.\n\n-- \ndwmw2\n\n"},{"id":"217","messageId":"Pine.LNX.4.58.0504150753440.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"1113559330.12012.292.camel@baythorne.infradead.org","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-15T15:32:46Z","receivedAt":"2005-04-15T15:32:46Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, David Woodhouse wrote:\n> \n> And you're right; it shouldn't have to be for renames only. There's no\n> need for us to limit it to one \"source\" and one \"destination\"; the SCM\n> can use it to track content as it sees fit.\n\nListen to yourself, and think about the problem for a second.\n\nFirst off, let's just posit that \"files\" do not matter. The only thing\nthat matters is how \"content\" moved in the tree. Ok? If I copy a function\nfrom one fiel to another, the perfect SCM will notice that, and show it as\na diff that removes it from one file and adds it to another, and is\n_still_ able to track authorship past the move. Agreed?\n\nNow, you basically propose to put that information in the \"commit\" log, \nand that's certainly valid. You can have the commit log say \"lines 50-89 \nin file kernel/sched.c moved to lines 100-139 in kernel/timer.c\", and then \nrenames fall out of that as one very small special case.\n\nYou can even say \"lines 50-89 in file kernel/sched.c copied to..\" and \nallow data to be tracked past not just movement, but also duplication.\n\nDo you agree that this is kind of what you'd want to aim for? That's a \nwinning SCM concept.\n\nHow do you think the SCM _gets_ at this information? In particular, how \nare you proposing that we determine this, especially since 90% of all \nstuff comes in as patches etc? \n\nYou propose that we spend time when generating the tree on doing so. I'm \ntelling you that that is wrong, for several reasons:\n\n - you're ignoring different paths for the same data. For example, you \n   will make it impossible to merge two trees that have done exactly the \n   same thing, except one did it as a patch (create/delete) and one did it \n   using some other heuristic.\n\n - you're doing the work at the wrong point. Doing it _well_ is quite \n   expensive. So if you do it at commit time, you cannot _afford_ to do it \n   well, and you'll always fall back to doing an ass-backwards job that \n   doesn't really get you to the good state, and only gets you to a \n   not-very-interesting easy 1% of the solution (ie full file renames).\n\n - you're doing the work at the wrong point for _another_ reason. You're \n   freezing your (crappy) algorithm at tree creation time, and basically \n   making it pointless to ever create something better later, because even \n   if hardware and software improves, you've codified that \"we have to\n   have crappy information\".\n\nNow, look at my proposal: \n\n - the actual information tracking tracks _nothing_ but information. You \n   have an SCM that tracks what changed at the only level that really \n   matters, namely the whole project. None of the information actually \n   makes any sense at all at a smaller granularity, since by definition, a\n   \"project\" depends on the other files, or it wouldn't be a project, it\n   would be _two_ projects or more.\n\n - When you're interested in the history of the information, you actually \n   track it, and you try to be _intelligent_ about it. You can actually do \n   a HELL of a lot better than whet you propose if you go the extra mile. \n   For example, let's say that you have a visualization tool that you can \n   use for finding out where a line of code came from. You start out at \n   some arbitrary point in the tree, and you drill down. That's how it \n   works, right?\n\n   So how do you drill down? You simply go backwards in history for that \n   project, tracking when that file+line changed (a \"file+line\" thing is \n   actually a \"sensible\" tracking unit at this point, because it makes\n   sense within the query you're doing - it's _not_ a sensible thing to\n   track at \"commit\" time, but when you ask yourself \"where did this line\n   come from\", that _question_ makes it sensible. Also note that \"where \n   did this _file_ come from is not a sensible question, since the file \n   may have been the combination (or split) of several files, so there is\n   no _answer_ to that question\"\n\n   So the question then becomes: \"how can you reasonably _efficiently_\n   find the history of one particular line\", and in fact it turns out that \n   by asking the question that way, it's pretty obvious: now that you\n   don't have to track the whole repository, you can always try to \n   minimize the thing you're looking for.\n\n   So what you do is walk back the history, and look at the tree objects \n   (both sides when you hit a merge), eand see if that file ever changes. \n   That's actually a very efficient operation in GIT - it matches\n   _exactly_ how git tracks things anyway. So it's not expensive at all.\n\n   When that file changes, you need to look if that _line_ changed (and \n   here is where it comes down to usability: from a practical standpoint\n   you probably don't care about a single line, you really _probably_ want\n   to see changes around it too). So you diff the old state and the new \n   state, and you see if you can still find where you were. If you still \n   can, and the line (and a few lines around it) is still the same, you \n   just continue to drill down. So that's not the interesting case.\n\n   So what happens when you found \"ok, that area changed\"? Your \n   visualization tool now shows it to the user, AND BECAUSE IT SEES THE \n   WHOLE TREE DIFF, it also shows where it probably came from. At _that_ \n   point, it is actually very trivial to use a modest amount of CPU time, \n   and look for probable sources within that diff. You can do it on modern \n   hardware in basically no time, so your visualization tool can actually \n   notice that\n \n\t\"oops, that line didn't even exist in the previous version, BUT I\n\t FOUND FIVE PLACES that matched almost perfectly in the same diff,\n\t and here they are\"\n\n   and voila, your tool now very efficiently showed the programmer that\n   the source of the line in question was actually that we had merged 5 \n   copies of the same code in different archtiectures into one common\n   helper function.\n\n   And if you didn't find some source that matched, or if the old file was\n   actually very similar around that line, and that line hadn't been\n   \"totally new\"? That's the easy case again - you show the programmer the\n   diff at that point in time, and you let him decide whether that diff \n   was what he was looking for, or whether he wants to continue to \"zoom\n   down\" into the history.\n\nThe above tool is (a) fairly easy to write for git (if you can do \nvisualization tools and (b) _exactly_ what I think most programmers \nactually want. Tell me I'm wrong. Honestly..\n\nAnd notice? My clearly _superior_ algorithm never needed any rename\ninformation at all. It would have been a total waste of time. It would\nalso have hidden the _real_ pattern, which was that a piece of code was\nmerged from several other matching pieces of code into one new helper\nfunction. But if it _had_ been a pure rename, my superior tool would have\ntrivially found that _too_. So rename infomation really really doesn't\nmatter.\n\nSo I'm claiming that any SCM that tries to track renames is fundamentally\nbroken unless it does so for internal reasons (ie to allow efficient\ndeltas), exactly because renames do not matter. They don't help you, and \nthey aren't what you were interested in _anyway_.\n\nWhat matters is finding \"where did this come from\", and the git\narchitecture does that very well indeed - much better than anything else\nout there. I outlined a simple algorithm that can be fairly trivially\ncoded up by somebody who really cares. Sure, pattern matching isn't\ntrivial, but you start out with just saying \"let's find that exact line,\nand two lines on each side\", and then you start improving on that.\n\nAnd that \"where did this come from\" decision should be done at _search_ \ntime, not commit time. Because at that time it's not only trivial to do, \nbut at that time you can _dynamically_ change your search criteria. For \nexample, you can make the \"match\" algorithm be dependent on what you are \nlooking at.\n\nIf it's C source code, it might want to ignore vairable names when it\nsearches for matching code. And if it's a OpenOffice document, you might\nhave some open-office-specific tools to do so. See? Also, the person doing \nthe searches can say whether he is interested in that particular line (or \neven that particial _identifier_ on a line), or whether he wants to see \nthe changes \"around\" that line.\n\nAll of which are very valid things to do, and all of which my world-view\nsupports very well indeed. And all of which your pitiful \"files matter\" \nworld-view totally doesn't get at all.\n\nIn other words, I'm right. I'm always right, but sometimes I'm more right \nthan other times. And dammit, when I say \"files don't matter\", I'm really \nreally Right(tm).\n\nPlease stop this \"track files\" crap. Git tracks _exactly_ what matters, \nnamely \"collections of files\". Nothing else is relevant, and even \n_thinking_ that it is relevant only limits your world-view. Notice how the \nnotion of CVS \"annotate\" always inevitably ends up limiting how people use \nit. I think it's a totally useless piece of crap, and I've described \nsomething that I think is a million times more useful, and it all fell out \n_exactly_ because I'm not limiting my thinking to the wrong model of the \nworld.\n\n\t\t\tLinus\n"},{"id":"218","messageId":"Pine.LNX.4.58.0504150836410.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"1113578964.27227.65.camel@hades.cambridge.redhat.com","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-15T15:51:54Z","receivedAt":"2005-04-15T15:51:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, David Woodhouse wrote:\n> \n> And when I'm looking for the change that broke something, I can almost\n> always tell which file it's in and go looking in _that_ file.\n\nRead my email about finding \"what changed\" that I sent out a minute ago.\n\nI claim that my algorithm for finding \"what changed\" handles your \"single\nfile\" case as a very small (and usually quite uninteresting) special case.\n\nI claim (and if you just look at my proposal I think you'll agree) that I\ncan track single functions, and do it efficiently. WITHOUT adding any\nmeta-data at all.\n\nThe thing is, if the question is \"I have this piece of code, and I want to \nsee what changed\", you fundamentally _can_ do that efficiently. That's \nreally what git was designed for. It's the whole _point_ of having history \nin the first place. If git didn't care, it wouldn't have a back-pointer to \nthe tree it came from, and we'd all be just merging pure trees.\n\nBut you mix that question up with \"how do I save that information in the \ncommit\", which is a totally unnecessary mix-up, and which makes things \nMUCH more complicated, for absolutely zero gain.\n\nIn fact, because you mixed up those two issues, the problem now became so\ncomplicated that you can no longer solve it, so you start doing hacks like\n\"the user has to tell us what he did\" (aka \"bk mv\" or \"svn rename\"), and \nyou start mentally to limit yourself to files, because you realize that \nyou _have_ to limit your intractable problem to make it at all solvable.\n\nAnd I'm telling you that your problem is STUPID. You made it stupid by \nthinking that every question about the source tree should be answered at \ncommit time. Which just clearly isn't true!\n\nIf you just drop the tying-together, and accept that \"what changed\" is a\nvalid question _regardless_ of trying to track it at commit time, now your\nwhole world opens up. Birds sing, the sun is shining on you, and beautiful\nscantily clad women (or men) dance around you. The world is suddenly a \ngood place, just _filled_ with possibilities.\n\nSuddenly you realize that if the question is just \"what changed in this\npiece of code\" (and let's face it, that _is_ the question), you can track \nit afterwards. Trying to tie in \"commit time\" into the question was what \nmade it hard. If you do _not_ due that (totally unnecessary) tie-in, the \nquestion suddenly becomes easy to answer, and several obvious and simple \nanswers spring to mind pretty immediately.\n\n> It's not about ditching the per-tree tracking and doing per-file\n> tracking instead. I agree that would be wrong. It's about storing enough\n> information to track what happened to given content as it moved around\n> within the tree.\n\nNo. Git absolutely does have everything you need already. You just aren't \nrealizing that it's already there - in the data - and that you can do much \nmore intelligent searches for changes if you accept that undeniable fact.\n\nThe fact that you can NOT do those searches at commit-time (which is a \nglobal op), and can only do them if you have a specific question in mind \n(\"what changed _here_\"), is the big issue.\n\nThe thing is, at commit-time you'd need to answer every possible question\n(\"what changed here, and here, and here, and in this function, and in this\nfile, and in this directory and why did this identifier get renamed and \nwhy is the sky blue\"). AND YOU FUNDAMENTALLY CANNOT DO THAT. It's \nimpossible.\n\nBut once you _know_ the question (which is the only time when the answer\nis actually relevant, so why care about if before that time?), you can\nfind out the answer by just automating the job of looking at the _data_. \nIt's easy. The question makes it obvious by its nature. The question is \nthe thing that gives you the specifics that makes the search possible in \nthe first place.\n\nAnd _this_ is why the data matters. Renames and file boundaries do not.\nAnd until you accept that, you just limit yourself.\n\n\t\tLinus\n"},{"id":"219","messageId":"20050415085424.5f61b80b.pj@engr.sgi.com","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504150740310.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Paul Jackson","fromEmail":"pj@engr.sgi.com","sentAt":"2005-04-15T15:54:24Z","receivedAt":"2005-04-15T15:54:24Z","isPatch":false,"sender":{"key":"pj@engr.sgi.com","avatar":null},"body":"Linus wrote:\n> For example, just doing intra-file diffs is a lot _easier_ and less \n> time-consuming than doing inter-file diffs. \n\nUm ah ... could you explain what you mean by inter and intra file diffs?\n\nGoogle found a three year old message by Andrew Morton, discussing\ninter and intra file fragmentation on ext2/ext3 file systems and the\nfind_group_dir() routine.  I don't think that's what you had in mind ;).\n\nWhen I run the 'diff' command, it usually between two files, not between\ntwo parts of a file.  So I'd have thought inter file diffs were easier.\n\nClearly, I don't git it.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"220","messageId":"1113580881.27227.73.camel@hades.cambridge.redhat.com","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504150753440.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-15T16:01:21Z","receivedAt":"2005-04-15T16:01:21Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Fri, 2005-04-15 at 08:32 -0700, Linus Torvalds wrote:\n>  - you're doing the work at the wrong point. Doing it _well_ is quite \n>    expensive. So if you do it at commit time, you cannot _afford_ to do it \n>    well, and you'll always fall back to doing an ass-backwards job that \n>    doesn't really get you to the good state, and only gets you to a \n>    not-very-interesting easy 1% of the solution (ie full file renames).\n> \n>  - you're doing the work at the wrong point for _another_ reason. You're \n>    freezing your (crappy) algorithm at tree creation time, and basically \n>    making it pointless to ever create something better later, because even \n>    if hardware and software improves, you've codified that \"we have to\n>    have crappy information\".\n\nOK, I'm inclined to agree. The only thing that prevents me from\ncapitulating entirely and resubscribing to the \"Torvalds is always\nright\" school is the concern that it _is_ expensive, and that's why I\noriginally wanted to do it at commit time because then it's a one-off\ncost rather than recurring every time we want to track the history of a\ngiven piece of content. Also because we actually have the developer's\nattention at commit time, and we can get _real_ answers from the user\nabout what she was doing, instead of having to guess.\n\nBut if it can be done cheaply enough at a later date even though we end\nup repeating ourselves, and if it can be done _well_ enough that we\nshouldn't have just asked the user in the first place, then yes, OK I\nagree.\n\n-- \ndwmw2\n\n"},{"id":"221","messageId":"20050415161330.GH19078@pasky.ji.cz","threadId":"9","inReplyTo":"7vzmw0ok45.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: ls-tree enhancements","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-15T16:13:30Z","receivedAt":"2005-04-15T16:13:30Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Apr 15, 2005 at 04:21:30AM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> told me that...\n> +static void _usage(void)\n> +{\n> +\tusage(\"ls-tree [-r] [-z] <key>\");\n> +}\n\n(namespace-nazi-hat\n This infriges the system namespaces. FWIW, I prefer to add the\nunderscore at the end of the identifier if wanting to do stuff like\nthis. Or just call it my_usage().\n)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"223","messageId":"Pine.LNX.4.61.0504151227590.27637@cag.csail.mit.edu","threadId":"9","inReplyTo":"20050415085424.5f61b80b.pj@engr.sgi.com","subject":"Re: Merge with git-pasky II.","fromName":"C. Scott Ananian","fromEmail":"cscott@cscott.net","sentAt":"2005-04-15T16:30:05Z","receivedAt":"2005-04-15T16:30:05Z","isPatch":false,"sender":{"key":"cscott@cscott.net","avatar":"https://gravatar.com/avatar/3551c2aefb299a0c45807f7677f5b26d8a5be4a4af359b4bf4fabbdd1f2b990e?d=mp&s=160"},"body":"On Fri, 15 Apr 2005, Paul Jackson wrote:\n\n> Um ah ... could you explain what you mean by inter and intra file diffs?\n\nintra file diffs: here are two versions of the same file.  what changed? \ninter file diffs: here is a new file, and here are *all the files in the \ncurrent committed version*.  Where did the contents of this new file come \nfrom?  (Note that the new file is often a slightly changed version of an \nexisting file in the current committed version.  But we don't assume that \nmust be true.)\n  --scott\n\nsupercomputer Pakistan WSHOOFS SECANT LCPANGS SDI assassination ZPSECANT \nSEQUIN AEBARMAN ESCOBILLA bomb mustard STANDEL ESGAIN Nazi FJDEFLECT\n                          ( http://cscott.net/ )\n"},{"id":"224","messageId":"Pine.LNX.4.61.0504151230180.27637@cag.csail.mit.edu","threadId":"9","inReplyTo":"1113580881.27227.73.camel@hades.cambridge.redhat.com","subject":"Re: Merge with git-pasky II.","fromName":"C. Scott Ananian","fromEmail":"cscott@cscott.net","sentAt":"2005-04-15T16:31:54Z","receivedAt":"2005-04-15T16:31:54Z","isPatch":false,"sender":{"key":"cscott@cscott.net","avatar":"https://gravatar.com/avatar/3551c2aefb299a0c45807f7677f5b26d8a5be4a4af359b4bf4fabbdd1f2b990e?d=mp&s=160"},"body":"On Fri, 15 Apr 2005, David Woodhouse wrote:\n\n> given piece of content. Also because we actually have the developer's\n> attention at commit time, and we can get _real_ answers from the user\n> about what she was doing, instead of having to guess.\n\nYes, but it's still hard to get *accurate* information.  And developers \ntend to use very short commit messages already...\n\n> But if it can be done cheaply enough at a later date even though we end\n> up repeating ourselves, and if it can be done _well_ enough that we\n> shouldn't have just asked the user in the first place, then yes, OK I\n> agree.\n\nI think examining the rsync algorithms should convince you that finding \ncommon chunks can be fairly efficient.  (See my next message for a more \nconcrete proposal.)\n  --scott\n\nRijndael AMLASH Moscow Ft. Bragg shotgun HTKEEPER SHERWOOD overthrow \nUzi anthrax Yeltsin Indonesia Suharto LITEMPO Dictionary Yakima KUBARK\n                          ( http://cscott.net/ )\n"},{"id":"225","messageId":"Pine.LNX.4.58.0504150950420.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"Pine.LNX.4.61.0504151230180.27637@cag.csail.mit.edu","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-15T17:11:32Z","receivedAt":"2005-04-15T17:11:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, C. Scott Ananian wrote:\n> \n> I think examining the rsync algorithms should convince you that finding \n> common chunks can be fairly efficient.\n\nNote that \"efficient\" really depends on how good a job you want to do, so \nyou can tune it to how much CPU you can afford to waste on the problem.\n\nFor example, my example had this thing where we merged five different\nfunctions into one function, and it is truly pretty efficient to find\nthings like that _IF_ we only look at the files that changed (since the\nset of files that change in any one particular commit tends to be small,\nrelative to the whole repository). There are many good algorithms for \nfinding \"common code\", and with modern hardware that is basically \ninstantaneous if you look at a few tens of files.\n\nFor example, people wrote efficient things to compare _millions_ of lines \nof code for the whole SCO saga - you can do quite well. Some googling \ncomes up with for example\n\n\thttp://minnie.tuhs.org/Programs\n\nand applying those to a smallish set of files is quite efficient.\n\nWhat is _not_ necessarily as easy is the situation where you notice that a \nnew set of lines appeared, but you don't see any place that matches that \nset of lines in the set of CHANGED files. That's actually quite common, ie \nlet's say that you have a new filesystem or a new driver, and almost \nalways it's based on a template or something, and you _would_ be able to \nsee where that template came from, except it's not in that _changed_ set.\n\nAnd that is still doable, but now you really have to compare against the\nwhole tree if you want to do it. Even _that_ is actually efficient if you\ncache the hashes - that's how the comparison tools compare two totally\nindependent trees against each other, and it makes it practically possible\nto do even that expensive O(n**2) operation in reasonable time. It's\ncertainly possible to do exactly the same thing for the \"new code got\nadded, does it bear any similarity to old code\" case.\n\nNote! This is a question that is relevant and actually is in the realm of\nthe \"possible to find the answer interactively\".  It may fairly expensive, \nbut the point is that this is the kind of relevant question that really \ndoes depend on the fundamental notion that \"data matters more than any \nlocal changes\". And when you think about the problem in that form, you \nfind these kinds of interesting questions that you _can_ answer.\n\nBecause the way git identifies data, the example \"is there any other\nrelevant code that may actually be similar to the newly added code\" is\nactually not that hard to do in git. Remember: the way to answer that\nquestion is to have a cache of hashes of the contents. Guess what git\n_is_? You can now index your line-based hashes of contents against the\n_object_ hashes that git keeps track of, and you suddenly have an\nefficient way to actually look up those hashes.\n\nNOTE! All of this is outside the scope of git itself. This is all\n\"visualization and comparison tools\" built up on top of git. And I'm not\nat all interested in writing those tools myself, and I'm absolutely not\nsigning up for that part. All I'm arguing for is that the git architecture\nis actually a very good architecture for doing these kinds of very very\ncool tools.\n\n\t\t\tLinus\n"},{"id":"228","messageId":"7v64ynj3rh.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"20050415161330.GH19078@pasky.ji.cz","subject":"Re: ls-tree enhancements","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T18:25:54Z","receivedAt":"2005-04-15T18:25:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n\n>> +static void _usage(void)\n\nPB>  This infriges the system namespaces. FWIW, I prefer to add the\nPB> underscore at the end of the identifier if wanting to do stuff like\nPB> this. Or just call it my_usage().\n\nThanks.  My bad.  Noted.\n\n\n"},{"id":"230","messageId":"20050415112927.72a7004f.pj@engr.sgi.com","threadId":"9","inReplyTo":"Pine.LNX.4.61.0504151227590.27637@cag.csail.mit.edu","subject":"Re: Merge with git-pasky II.","fromName":"Paul Jackson","fromEmail":"pj@engr.sgi.com","sentAt":"2005-04-15T18:29:27Z","receivedAt":"2005-04-15T18:29:27Z","isPatch":false,"sender":{"key":"pj@engr.sgi.com","avatar":null},"body":"> intra file diffs: here are two versions of the same file. \n\nAh so.  Linus faked me out.\n\nI was _sure_ that by \"file\" he meant \"file\" -- as in a bucket of bits\nwith a unique identifying <sha1>.\n\nIn that message, I guess by \"file\" he meant \"a version controlled\nfile, consisting of a series of content versions and meta-data\"\n\nThat's what I get for trusting Linus to always speak as a kernel\nhacker, not an SCM hacker.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"231","messageId":"Pine.LNX.4.58.0504151138490.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"7vaco0i3t9.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: write-tree is pasky-0.4","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-15T18:44:02Z","receivedAt":"2005-04-15T18:44:02Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, Junio C Hamano wrote:\n> \n> Linus, sorry for bothering you with a false alarm.  The problem\n> turns out to be introduced in pasky-0.4 and does not exist in\n> your HEAD.\n\nHey, all the code I write is always perfect, of course ;)\n\nThat said, I'm having some trouble merging with your perfect code,\nespecially since I decided that Russell's \"always big-endian\" thing was\ndefinitely the right way to go (but ended up doing it slightly\ndifferently).\n\nI did my own version of \"upcate-cache --cacheinfo\", although mine is a bit \nmore anal, and if you add a new filename it wants that \"--add\" flag in \nthere first (why? I really like to make sure that people who add or remove \nfiles from the cache say so explicitly, so that there are no surprises). \nOtherwise it should be compatible with yours.\n\nAnd I merged your \"Add -z option to show-files\", but you had based your \nother patches on Petr's tree which due to my other changes is not going to \nmerge totally cleanly with mine, so I'm wondering if you might want to try \nto re-merge your mergepoint stuff against my current tree? That way I can \ncontinue to maintain a set of \"core files\", and Pasky can maintain the \n\"usable interfaces\" part..\n\n\t\tLinus\n"},{"id":"234","messageId":"20050415185624.GB7417@pasky.ji.cz","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504151138490.7211@ppc970.osdl.org","subject":"Re: Re: write-tree is pasky-0.4","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-15T18:56:24Z","receivedAt":"2005-04-15T18:56:24Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Apr 15, 2005 at 08:44:02PM CEST, I got a letter\nwhere Linus Torvalds <torvalds@osdl.org> told me that...\n> And I merged your \"Add -z option to show-files\", but you had based your \n> other patches on Petr's tree which due to my other changes is not going to \n> merge totally cleanly with mine, so I'm wondering if you might want to try \n> to re-merge your mergepoint stuff against my current tree? That way I can \n> continue to maintain a set of \"core files\", and Pasky can maintain the \n> \"usable interfaces\" part..\n\nActually, I wanted to ask about this. :-)\n\nSo, I assume that you don't want to merge my \"SCM layer\" (which is\nperfectly fine by me). However, I also apply plenty of patches\nconcerning the \"core git\" - be it portability, leak fixes, argument\nparsing fixes and so on.\n\nWould it be of any benefit if I maintained two trees, one with just your\ncore git but what I merge (I think I'd call this branch git-pb), and one\nwith my git-pasky (to be renamed to Cogito) layer. I'd then put the\n\"core git\" changes to the git-pb branch and pull from it to the Cogito\nbranch regularily, but it should be safe for you to pull from it too.\n\nIn fact, in that case I might even end up entirely separating the Cogito\ntools from the core git and distributing them independently.\n\nBTW, just out of interest, are you personally planning to use Cogito for\nyour kernel and sparse (and possibly even git) work, or will you stay\nwith your lowlevel plumbing for that?\n\nThanks,\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"237","messageId":"20050415122051.751c2a53.pj@engr.sgi.com","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504150753440.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Paul Jackson","fromEmail":"pj@engr.sgi.com","sentAt":"2005-04-15T19:20:51Z","receivedAt":"2005-04-15T19:20:51Z","isPatch":false,"sender":{"key":"pj@engr.sgi.com","avatar":null},"body":"These notions that one can always best answer questions by looking at\nthe content, and that \"Individual files DO NOT EXIST\" seem over stated,\nto me.\n\nGranted, overstated for a good reason.  A couple sticks of dynamite are\nneeded to shake loose some old SCM thinking habits.\n\n===\n\nIngo has a point when he states:\n\n> i believe the fundamental thing to think about is not file or line or \n> namespace, but 'tracking developer intent'.\n\nHe too overstates - it's not _the_ (as in one and only) thing.\nBut it's useful.  Given the traditional terseness of many engineers,\nit's certainly not the _only_ thing.  The code speaks too.\n\n===\n\nThe above two are related in this way.  Traditional SCM uses per\nfile (versioned controlled file, as in s.* or *,v files) metadata\nto track 'developer intent'.\n\nI'm afraid we are at risk for confusing baby (developer intent)\nand bathwater (version controlled file structure of classic SCM's).\n\n===\n\nBut we already have a pretty damn good way of tracking developer\nintent that needs to fit naturally with whatever we build on top\nof git.\n\n     Mr. McGuire: I just want to say one word to you - just one word.\n     Ben: Yes sir.\n     Mr. McGuire: Are you listening?\n     Ben: Yes I am.\n     Mr. McGuire: 'Patches.'\n     # the original word was 'Plastics' - The Graduate (1967)\n\nAndrew and the other maintainers do a pretty good job of 'encouraging'\ndevelopers to provide useful statements of 'intent' in their patch\nheaders.\n\nThe patch series in something like *-mm, including per-patch\ncommentary, are a valuable part of this project.\n\n===\n\nI have not looked closely at what is being done here, on top of\ngit, for SCM like capabilities.  Hopefully the next two questions\nare not too stupid:\n\n 1) How do we track the patch header commentary?\n\n 2) Why can't we have a quilt like SCM, not bk/rcs/cvs/sccs/... like?\n\nFor (2), anyone publishing a Linux source would periodically announce an\n<sha1> value, attached to some name suitable for public consumption.\n\nFor example, sometime in the next month or so, Linus would announce\nthat the <sha1> of 2.6.12 is so-and-so.  That would identify the\nresult of applying a specific set of patches, resulting in a specific\nsource tree contents.  He would announce a few 2.6.12-rc* <sha1>'s\nbetween now and then.\n\nBetween now and then, Andrew would (if using these tools) have published\nseveral <sha1> values, one each for various 2.6.12-rc*-mm* versions.\n\nIf you explode such a <sha1> all out into a working directory, you get\nboth the source contents in the appropriately named files, and the\nquilt-style patches subdirectory, of the patch series that gets you\nhere, starting from some Time Zero for that series of published kernel\nversions.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"239","messageId":"20050415195451.GF7417@pasky.ji.cz","threadId":"9","inReplyTo":"7v7jj4q2j2.fsf@assigned-by-dhcp.cox.net","subject":"Re: Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-15T19:54:52Z","receivedAt":"2005-04-15T19:54:52Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Apr 15, 2005 at 02:58:25AM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> told me that...\n> >>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n> >> I think the above would result in what SCM person would call\n> >> \"merge upstream/sidestream changes into my working directory\".\n> \n> PB> And that's exactly what I'm doing now with git merge. ;-) In fact,\n> PB> ideally the whole change in my scripts when your script is finished\n> PB> would be replacing\n> \n> PB> \tcheckout-cache `diff-tree` # symbolic\n> PB> \tgit diff $base $merged | git apply\n> \n> PB> with\n> \n> PB> \tmerge-tree.pl -b $base $(tree-id) $merged | parse-your-output\n> \n> In the above I presume by $merged you mean the tree ID (or\n> commit ID) the user's working directory is based upon?  Well,\n> merge-trees (Linus has a single directory merge-tree already)\n> looks at tree IDs (or commit IDs); it would never involve\n> working files in random state that is not recorded as part of a\n> tree (committed or not).  Given that constraints I am not sure\n> how well that would pan out.  I have to think about this a bit.\n\nNo, $(tree-id) is the \"destination branhc\", what the user directory is\nbased upon; $merged is the branch you are merging now, relative to\n$base. When I throw away the useless \"-b\" argument, in practice it would\nlook like\n\n\tmerge-trees abcd 1234 5678\n\nfor doing\n\n      /------ 1234 -+-\nabcd <             /\n      \\------ 5678\n\n(not that the order of 1234 and 5678 would actually really matter)\n\nI fear I don't understand the rest of your paragraph. :-(\n\n> I do like, however, the idea of separating the step of doing any\n> checkout/merge etc. and actually doing them.  So the command set\n> of parse-your-output needs to be defined.  Based on what I have\n> done so far, it would consist of the following:\n> \n>  - Result is this object $SHA1 with mode $mode at $path (takes\n>    one of the trees); you can do update-cache --cacheinfo (if\n>    you want to muck with dircache) or cat-file blob (if you want\n>    to get the file) or both.\n> \n>  - Result is to delete $path.\n> \n>  - Result is a merge between object $SHA1-1 and $SHA1-2 with\n>    mode $mode-1 or $mode-2 at $path.\n> \n> Would this be a good enough command set?\n\nWhat about the conflicts? Like one tree deleting, other tree modifying?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"240","messageId":"7vr7hbhky9.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504140051550.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T19:57:34Z","receivedAt":"2005-04-15T19:57:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> In the meantime I wrote a very stupid \"merge-tree\" which\nLT> does things slightly differently, but I really think your\nLT> approach (aka my original approach) is actually a lot\nLT> faster. I was just starting to worry that the ball didn't\nLT> start, so I wrote an even hackier one.\n\nLT> ... This \"one directory at a time with very explicit output\"\nLT> thing is much more down-to-earth, but it's also likely\nLT> slower because it will need script help more often.\n\nI was looking at merge-tree.c last night to add recursive\nbehaviour (my favorite these days ;-) to it [*1*].\n\nBut then I started thinking.\n\nLT> ... For each entry in the directory it says either\nLT> \tselect <mode> <sha1> path\nLT> or\nLT> \tmerge <mode>-><mode>,<mode> <sha1>-><sha1>,<sha1> path\nLT> depending on whether it could directly select the right object or not.\n\nGiven that the case you are primarily interested in is the one\nthat affects only small parts of a huge tree (i.e. common kernel\nmerge pattern I understand from your previous messages), your\n\"hacky version\" [*2*], extended for recursive operation, would\nspit out 98% select and 2% merge, and probably the origin of\nthese selects are distributed across ancestor=90%, his=4%,\nmy=4%, or something similar.  Am I misestimating grossly?\n\nAssuming I am correct in the above, this would not scale for a\nhuge project.  We need to cut down the number of \"90% select\"\npart of the output to make it manageable.\n\nI am thinking about:\n\n - adding recursive behaviour (I am almost done with this);\n\n - adding another command line argument to merge-tree.c, to \n   tell \"do not output anything for the path if the resulting\n   merge is the same as what is in this tree\";\n\n - adding another output type, \"delete\" to make the output type\n   repertoire these three:\n\n    delete path\n    select <mode> <sha1> path\n    merge <mode>-><mode>,<mode> <sha1>-><sha1>,<sha1> path\n\nWhen the user of the output of \n\n  $ merge-tree <ancestor-sha1> <my-sha1> <his-sha1> <result-base-sha1>\n\nwant to get a dircache populated with the merged result, he can:\n\n  1. read-tree <result-base-sha1>\n  2. for each output:\n     a) \"delete\" -- delete path from dircache\n     b) \"select\" -- register mode-sha1 at path\n     c) \"merge\"  -- do the 3-way merge and register result at path\n\nDo you think this is sensible?\n\nThe reason I have the separate <result-base-sha1> instead of\nalways using <ancestor-sha1> is because the user may be thinking\nof patching an existing base which is different from \"my\" or\n\"his\" or \"ancestor\" and doing it in place.  That way, probably\nPasky's SCM can use it to patch the dircache it creates in its\nown ,,merge/ directory, which would most likely be initially\npopulated from the dircache in the user's working directory---\nwhich may or may not match \"my-sha1\" if the user has uncommitted\nupdate-cache there.\n\nPasky, do you think this is workable?  If so do you think this\nwould make your life easier?\n\n\n[Footnotes]\n\n*1* That's how I found the S_IFDIR problem (not in your tree but\nin the copy I had).\n\n*2* I did not find it quite \"hacky\".  It was a pleasant read.\nEspecially I liked \"smaller()\" part.\n\n\n"},{"id":"241","messageId":"7vmzrzhkd3.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504151138490.7211@ppc970.osdl.org","subject":"Re: write-tree is pasky-0.4","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T20:10:16Z","receivedAt":"2005-04-15T20:10:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> Hey, all the code I write is always perfect, of course ;)\n\nAnd you are always right ;-)  Liked that blast-from-the-past?\n\nLT> That said, I'm having some trouble merging with your perfect code,\nLT> especially since I decided that Russell's \"always big-endian\" thing was\nLT> definitely the right way to go (but ended up doing it slightly\nLT> differently).\n\nLT> I did my own version of \"upcate-cache --cacheinfo\", although\nLT> mine is a bit more anal, and if you add a new filename it\nLT> wants that \"--add\" flag in there first (why? I really like\nLT> to make sure that people who add or remove files from the\nLT> cache say so explicitly, so that there are no surprises).\nLT> Otherwise it should be compatible with yours.\n\nThanks.  Not honoring \"--add\" was an oversight on my part.\n\nLT> And I merged your \"Add -z option to show-files\", but you had\nLT> based your other patches on Petr's tree which due to my\nLT> other changes is not going to merge totally cleanly with\nLT> mine, so I'm wondering if you might want to try to re-merge\nLT> your mergepoint stuff against my current tree?\n\nMy pleasure.  I am currently not that interested in toilet part\nthan I am interested in plumbing part, so rebasing my tree back\nto yours is no problem for me.  Currently I see your HEAD is at\n461aef08823a18a6c69d472499ef5257f8c7f6c8, so I will generate a\nset of patches against it.\n\n"},{"id":"242","messageId":"Pine.LNX.4.58.0504151212160.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"20050415185624.GB7417@pasky.ji.cz","subject":"Re: Re: write-tree is pasky-0.4","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-15T20:13:21Z","receivedAt":"2005-04-15T20:13:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, Petr Baudis wrote:\n> \n> So, I assume that you don't want to merge my \"SCM layer\" (which is\n> perfectly fine by me). However, I also apply plenty of patches\n> concerning the \"core git\" - be it portability, leak fixes, argument\n> parsing fixes and so on.\n\nI'm actually perfectly happy to merge your SCM layer too eventually, but \nI'm nervous at this point.  Especially while people are discussing some \nSCM options that I'm personally very leery of, and think that may make \nsense for others, but that I personally distrust.\n\n> BTW, just out of interest, are you personally planning to use Cogito for\n> your kernel and sparse (and possibly even git) work, or will you stay\n> with your lowlevel plumbing for that?\n\nI'm really really hoping I'd use cogito, and that it ends up being just \none project. In particular, I'm hoping that in a few days, I'll have done \nenough plumbing that I don't even care any more, and then I'd not even \nmaintain a tree of my own. \n\nI'm really not that much of an SCM guy. I detest pretty much all SCM's out\nthere, and while it's been interesting to do 'git', I've done it because I\nwas forced to, and because I really wanted to put _my_ needs and opinions\nfirst in an SCM, and see how that works. That's why I've been so adamant\nabout having a \"philosophy\", because otherwise I'd probably just end up\nwith yet another SCM that I'd despise.\n\nSo for me, the \"optimal\" situation really ends up that you guys end up as\nthe maintainers. I don't even _want_ to maintain it, although I'd be more\nthan happy to be part of the engineering team. I just want to mark out the\ndirection well enough and get it to a point where I can _use_ it, that I\nfeel like I'm done.\n\nBut before I can do that, I need to feel like I can live with the end \nresult. The only missing part is merges, and I think you and Junio are \ngetting pretty close (with Daniel's parent finder, Junio's merger etc). \n\n\t\t\tLinus\n"},{"id":"246","messageId":"20050415204033.GG7417@pasky.ji.cz","threadId":"9","inReplyTo":"7vwtr4ibkt.fsf@assigned-by-dhcp.cox.net","subject":"Re: Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-15T20:40:33Z","receivedAt":"2005-04-15T20:40:33Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Apr 15, 2005 at 12:22:26PM CEST, I got a letter\nwhere Junio C Hamano <junkio@cox.net> told me that...\n> After I re-read [*R1*], in which Linus talks about dircache,\n> especially this section:\n> \n>  - The \"current directory cache\" describes some baseline. In particular,\n>    note the \"some\" part. It's not tied to any special baseline, and you\n>    can change your baseline any way you please.\n> \n>    So it does NOT have to track any particular state in either the object \n>    database _or_ in your actual current working tree. In fact, all real \n>    interactions with \"git\" are really about updating this staging area one \n>    way or the other: you might check out the state from it into your \n>    working area (partially or fully), you can push your working area into \n>    the staging area (again, partially or fully).\n> \n>    And if you want to, you can write the thing that the staging area \n>    represents as a \"tree\" into the object database, or you can merge a \n>    tree from the object database into the staging area.\n> \n>    In other words: the staging area aka \"current directory cache\" is \n>    really how all interaction takes place. The object database never \n>    interacts directly with your working directory contents. ALL \n>    interactions go through the current directory cache.\n> \n> I started to have more doubts on the approach of *not*\n> performing the merge in the dircache I set up specifically for\n> merging, which is the direction in which you are pushing if I\n> understand you correctly.  Maybe I completely misunderstand what\n> you want.  This message is long but I need a clear understanding\n> of what is expected to be useful to you, so please bear with me.\n\n\n\n> PB> \tmerge-tree.pl -b $base $(tree-id) $merged | parse-your-output\n> \n> Please help me understand this example you have given earlier.\n> Here is my understanding of your assumption when the above\n> pipeline takes place.  Correct me if I am mistaken.\n> \n>  * The user is in a working directory $W.  It is controlled by\n>    git-tools and there are $W/.git/. directory and $W/.git/index\n>    dircache.\n> \n>  * The dircache $W/.git/index started its life as a read-tree\n>    from some commit.  The git-tools is keeping track of which\n>    commit it is somewhere, presumably in $W/.git/ directory.\n>    Let's call it $C (commit).\n> \n>  ? Question.  Is the $(tree-id) in your example the same as $C\n>    above?\n\nYes. Actually $(tree-id) returns ID of the tree object, not the commit\nobject; but that doesn't matter here, probably - let's ignore that\ndistinction for simplicity.\n\n>  * The user have run [*1*] (see Footnote below) checkout-cache\n>    on $W/.git/index some time in the past and $W is full of\n>    working files.  Some of them may or may not have modified.\n>    There may be some additions or deletions.  So the contents of\n>    the working directory may not match the tree associated with\n>    $C.\n> \n>  * The user may or may not have run [*1*] update-cache in $W.\n>    The contents of the dircache $W/.git/index may not match the\n>    tree associated with $C.\n> \n>  ? Question.  Are you forbidding the user to run update-cache by\n>    hand, and keeping track of the changes yourself, to be\n>    applied all at once at \"git commit\" time, thereby\n>    guaranteeing the $W/.git/index to match the tree associated\n>    with $C all times?  From the description of The \"GIT toolkit\"\n>    section in README, it is not clear to me which part of his\n>    repository an end user is not supposed to muck with himself.\n\nIdeally, he shouldn't be using *any* of the low-level plumbing by now.\nThe only exception is update-cache --refresh, which he can do at will\n(I'm yet thinking what to do with it :-).\n\nThe git-tools always assume that index basically contains the state as\nof the last commit. (Actually the only time when this matters *now*\nmight be git diff - the user would get confused from the results.)\n\n>  * Now the user has some changes in his working directory and\n>    notices upstream or a side branch has notable changes\n>    desireble to be picked up.  So he runs some git-tools command\n>    to cause the above quoted pipeline to run.\n> \n>  ? Question.  Does $merged in your example mean such an upstream\n>    or side branch?  Is $base in your example the common ancestor\n>    between $C and $merged?\n\nCorrect.\n\n*HOWEVER* what is not correct is that git-tools would let you merge in\nyour working directory while you have local changes there.\n\nIn the past, the merge would happen in your working tree, but git-tools\nwouldn't let you go for it unless your working tree has no local\nchanges. It would complain loudly and refuse to, since it's *not* what\nyou want to do and it was most likely a mistake.\n\nCurrently, git merge just creates a ,,merge/ subdirectory sharing the\nobject database with your working tree, but with an independent checkout\nof it; it will do the merge there, and when you commit it there, it will\nupdate your working tree with the merged changes.\n\nI'm describing both behaviors since I might revert back to the first\none, based on what (if anything) will Linus reply to my mail about\nout-of-tree merges.\n\nBut either way, when a merge is about to happen upon us, the working\ntree is \"clean\".\n\n> Assuming that my above understanding of your model is correct,\n> here are my \"thinking aloud\".\n> \n>  - \"merge-trees $base $C $merged\" looks only at the git object\n>    database for those three trees named.  The data structure of\n>    git object database is optimized to distinguish differences\n>    in those recorded trees (and hence recorded blobs they point\n>    at) without unpacking most of the files if the changes are\n>    small, because all the blobs involved are already hashed.  It\n>    is not very good at comparing things in git object store and\n>    working files in random states, which would involve unpacking\n>    blobs and comparing, so \"merge-trees\" does not bother.\n> \n>  - What can come out from merge-trees is therefore one of the\n>    following for each path from the union of paths contained in\n>    $base, $C, and $merged:\n> \n>    (a) Neither $C nor $merged changed it --- merge result is what\n>        is in $C.\n\n(Or in $base, if you don't want to give $C \"unfair advantage\", since it\ndoes not matter. ;-)\n\n>    (b) $C changed it but $merged did not --- merge result is what\n>        is in $C.\n> \n>    (c) Both $C and $merged changed it in the same way --- merge\n>        result is what is in $C.\n> \n>    (d) $C did not change it but $merged did --- merge result is\n>        what is in $merged.\n> \n>    (e) Both $C and $merged changed it differently --- merge is\n>        needed and automatically succeeds between $C and $merge.\n> \n>    (f) Both $C and $merged changed it differently --- merge is\n>        needed but have conflicts.\n> \n>  - Assuming we are dealing with the case where working files are\n>    dirty and do not match what is in $C, among the above,\n>    (a)-(c) can be ignored by SCM.  What the user has in his\n>    working files is exactly what he would have got if he started\n>    working from the merge result, although in reality the work\n>    was started from $C.\n\nYes. Actually they can be ignored by git-tools in any case since what is\nin the directory cache is $C. So it never needs to do any special\naction.\n\n>    Handling (d), (e) and (f) from SCM's point of view would be\n>    the same.  They all involve 3-way merges between the file in\n>    the working directory, and the file from $merged, pivoting on\n>    the file from $base.  In order to help SCM, merge-trees\n>    therefore should output SHA1 of blobs for such a file from\n>    $base and $merged and expect SCM to run \"cat-file blob\" on\n>    them and then merge or diff3.  Up to the point of giving\n>    those two SHA1 out is the business of merge-trees and after\n>    that it is up to SCM.\n> \n>    That would work.  So I should base the design of output from\n>    merge-trees on the above analysis, which probably needs to be\n>    extended to cover differences between creation, modification,\n>    and deletion.\n\nYes, it sounds sensible.\n\nActually, you don't even need to make $C more special than $merged; I\ncan filter out only the $merged changes on the SCM level. I guess that\nwould add no complexity to your tool and make it usable even for more\nexotic kinds of merges (like the floating-in-the-void merge of two\n\"equally important\" trees).\n\n>  - However, the above is quite different from the way Linus\n>    envisioned initially, on which my current implementation is\n>    based [*3*].\n> \n>    My current implementation is to record the merge outcome in\n>    the temporary dircache $W/,,merge/.git/index for cases\n>    (a)-(e).  The last case (f) is problematic and needs human\n>    validation [*2*], so it is not recorded in that temporary\n>    dircache, but the files to be merged are left in that\n>    temporary directory and merge-trees stops there.  It is\n>    expected that the end-user or SCM would merge the resulting\n>    file and run update-cache to update $W/,,merge/.git/index.\n>    After that happens, $W/,,merge/.git/index has the tree\n>    representing the desired result of the merge.  It is expected\n>    that the end-user or SCM would write-tree, commit-tree there\n>    in the temporary directory, creating a new commit $C1.\n> \n>    Then, it is expected that the SCM would make a patch file\n>    between $C and the user working directory, checks out $C1\n>    (either in the user's working directory or another temporary\n>    directory; at this point merge-trees does not care because it\n>    has already done its job and exited), applies that patch to\n>    bring the user edits over to $C1.  Then that directory would\n>    contain the desired merge of user edits.\n> \n>    That is my understanding of how Linus originally wanted the\n>    tool to do his kernel work with to work.  My hesitation to\n>    suggestions from you to change it not to keep its own merge\n>    dircache is coming from here.  Not doing what I am currently\n>    doing to $W/,,merge/.git/index dircache would mean that SCM\n>    would have to do more, not less, to arrive at $C1 (the result\n>    of the clean $merge and $C merge pivoted at $base), where the\n>    real SCM merge begins.\n\nWell. Currently, apart from the directory cache part, I do it like you\ndescribe. I create a new directory, after commit I apply the diff back\nto the original tree etc.\n\nThe only problem is really the dircache, and that's because it would be\ndone totally differently than in the original tree, and I would\nunnecessarily have to introduce crowds of special cases to my tools in\norder for them to be usable in the merge tree (I call the \",,merge\"\ntemporary directory a \"merge tree\").\n\nAnd the user would still lose the capability of easily seeing the\nchanges being committed. I admit that I'm using this largely as an\nexcuse and there *could* be a tool made which would compare the given\ntree with the cache, but it would be clumsy to use, violate Linus' \"ALL\ninteractions go through the current directory cache\" paradigm (whew, the\nfirst time in my life I used this word), and we could do just fine with\nour current tools.\n\n> Although I suspect I am misunderstanding what you want, your\n> messages so far suggest that what you want might be quite\n> different from what Linus wants.  Please do not misunderstand\n> what I mean by saying this.  I am not saying that Linus is\n> always right [*4*] and therefore you are wrong for wanting\n> something else.  It is just that, if what I started writing\n> needs to support both of those quite different needs, I need to\n> know what they are.  I think I understand what Linus wants well\n> enough [*5*], but I am not certain about yours.\n\nI can't see the conflicts between what I want and what Linus wants.\nAfter all, Linus says that I can use the directory cache in any way I\nplease (well, the user can, but I'm speaking for him ;-). So I'm doing\nso, and with your tool I would get into problems, since it is suddenly\nimposing a policy on what should be in the index.\n\n..snip..\n> *2* Strictly speaking, case (e) needs human validation as\n> well, because successful textual merge does not guarantee\n> sensible semantic merge.\n..snip..\n\nActually, I think _all_ the caches should make the human validation\n_possible_ (by showing the diff of what would be merged), and it is\ntrivial to do so by having pristine index.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"247","messageId":"Pine.LNX.4.58.0504151334350.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"7vr7hbhky9.fsf@assigned-by-dhcp.cox.net","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-15T20:45:56Z","receivedAt":"2005-04-15T20:45:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, Junio C Hamano wrote:\n> \n> I was looking at merge-tree.c last night to add recursive\n> behaviour (my favorite these days ;-) to it [*1*].\n> \n> But then I started thinking.\n\nAlways good.\n\n> LT> ... For each entry in the directory it says either\n> LT> \tselect <mode> <sha1> path\n> LT> or\n> LT> \tmerge <mode>-><mode>,<mode> <sha1>-><sha1>,<sha1> path\n> LT> depending on whether it could directly select the right object or not.\n> \n> Given that the case you are primarily interested in is the one\n> that affects only small parts of a huge tree (i.e. common kernel\n> merge pattern I understand from your previous messages), your\n> \"hacky version\" [*2*], extended for recursive operation, would\n> spit out 98% select and 2% merge, and probably the origin of\n> these selects are distributed across ancestor=90%, his=4%,\n> my=4%, or something similar.  Am I misestimating grossly?\n\nNo. That's _exactly_ right. You do not want a recursive merge-tree. \n\nThe \"diff-tree\" thing is different, exactly because it prunes out all the \ndifferences early on.\n\n> I am thinking about:\n> \n>  - adding recursive behaviour (I am almost done with this);\n\nI think your suggestion sounds perfectly reasonable.\n\n\t\tLinus\n"},{"id":"248","messageId":"Pine.LNX.4.61.0504151617170.27637@cag.csail.mit.edu","threadId":"9","inReplyTo":"7vmzrzhkd3.fsf@assigned-by-dhcp.cox.net","subject":"Re: write-tree is pasky-0.4","fromName":"C. Scott Ananian","fromEmail":"cscott@cscott.net","sentAt":"2005-04-15T20:58:10Z","receivedAt":"2005-04-15T20:58:10Z","isPatch":false,"sender":{"key":"cscott@cscott.net","avatar":"https://gravatar.com/avatar/3551c2aefb299a0c45807f7677f5b26d8a5be4a4af359b4bf4fabbdd1f2b990e?d=mp&s=160"},"body":"On Fri, 15 Apr 2005, Junio C Hamano wrote:\n\n> to yours is no problem for me.  Currently I see your HEAD is at\n> 461aef08823a18a6c69d472499ef5257f8c7f6c8, so I will generate a\n> set of patches against it.\n\nHave you considered using an s/key-like system to make these hashes more \nhuman-readable?  Using the S/Key translation (11-bit chunks map to a 1-4 \nletter word), Linus' HEAD is at:\n   WOW-SCAN-NAVE-AUK-JILL-BASH-HI-LACE-LID-RIDE-RUSE-LINE-GLEE-WICK-A\n...which is a little longer, but speaking of branch \"wow-scan\" (which \ngives 22 bits of disambiguation) is probably less error-prone than \ndiscussing branch '461...' (only 12 bits).\n\nYou could supercharge this algorithm by using (say) \n/usr/dict/american-english-large (>2^17 words; 160 bits of hash = 10 \ndictionary words), or mixing upper and lower case (likely to reduce the 15 \nword s/key phrase to ~11 words) to give something like\n    RiDe-Rift-rIMe-rOSy-ScaR-sCat-ShiN-sIde-Sine-seeK-TIEd-TINT\nMy personal feeling is that case is likely to be dropped in casual \nconversation, so speaking of branch 'wow', 'wow-scan', or 'wow-scan-nave' \nis likely to be significantly more useful than trying to pronounce \nmixed-cased versions of these.\n\nThis is obviously a cogito issue, rather than a git-fs thing.\n  --scott\n\n[More info is in RFCs 2289 and 1760, although all I'm really using from \nthese is the word dictionary in the appendix.]\n     http://www.faqs.org/rfcs/rfc1760.html\n     http://www.faqs.org/rfcs/rfc2289.html\n\nSKIMMER MKOFTEN Ft. Bragg Sabana Seca ESMERALDITE NORAD HTAUTOMAT \nradar interception Pakistan BOND Kennedy postcard corporate globalization\n                          ( http://cscott.net/ )\n"},{"id":"249","messageId":"20050415212255.GJ7417@pasky.ji.cz","threadId":"9","inReplyTo":"Pine.LNX.4.61.0504151617170.27637@cag.csail.mit.edu","subject":"Re: Re: write-tree is pasky-0.4","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-15T21:22:55Z","receivedAt":"2005-04-15T21:22:55Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Apr 15, 2005 at 10:58:10PM CEST, I got a letter\nwhere \"C. Scott Ananian\" <cscott@cscott.net> told me that...\n> On Fri, 15 Apr 2005, Junio C Hamano wrote:\n> \n> >to yours is no problem for me.  Currently I see your HEAD is at\n> >461aef08823a18a6c69d472499ef5257f8c7f6c8, so I will generate a\n> >set of patches against it.\n> \n> Have you considered using an s/key-like system to make these hashes more \n> human-readable?  Using the S/Key translation (11-bit chunks map to a 1-4 \n> letter word), Linus' HEAD is at:\n>   WOW-SCAN-NAVE-AUK-JILL-BASH-HI-LACE-LID-RIDE-RUSE-LINE-GLEE-WICK-A\n> ...which is a little longer, but speaking of branch \"wow-scan\" (which \n> gives 22 bits of disambiguation) is probably less error-prone than \n> discussing branch '461...' (only 12 bits).\n> \n> You could supercharge this algorithm by using (say) \n> /usr/dict/american-english-large (>2^17 words; 160 bits of hash = 10 \n> dictionary words), or mixing upper and lower case (likely to reduce the 15 \n> word s/key phrase to ~11 words) to give something like\n>    RiDe-Rift-rIMe-rOSy-ScaR-sCat-ShiN-sIde-Sine-seeK-TIEd-TINT\n> My personal feeling is that case is likely to be dropped in casual \n> conversation, so speaking of branch 'wow', 'wow-scan', or 'wow-scan-nave' \n> is likely to be significantly more useful than trying to pronounce \n> mixed-cased versions of these.\n> \n> This is obviously a cogito issue, rather than a git-fs thing.\n\nI kind of like it, the only thing I fear is possible conflict with\nbranch names; it is not very likely though, I think. I believe (at\nleast) the first three words should be used if possible.\n\nI'm not sure in what cases do you think we should use those \"verbal\"\nnames, though. Of course we should accept them as IDs, but I don't think\nwe should ever show them automatically. Probably provide a trivial to\nuse tool to convert to them, and parameters for *-id tools to show them.\n\nI assume we would have a custom tool for the translation?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"250","messageId":"7vfyxrhfsw.fsf_-_@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"7vmzrzhkd3.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH 1/2] merge-trees script for Linus git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T21:48:47Z","receivedAt":"2005-04-15T21:48:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus,\n\n    what you have in 461aef08823a18a6c69d472499ef5257f8c7f6c8 is fine\nby me for the essential support for merge-trees (sorry for the\nconfusing name, but this is a stop-gap Q&D script until I do the real\nmerge-tree.c conversion).\n\nThis patch contains the merge-trees script itself and Makefile entry\nfor it.  I have some more fixes to merge-trees in the works but that\nwill follow later.\n\nI have an optional patch to add '-q' option to show-diff so that\ncomplaints for missing files can be squelched, which I will be sending\nyou in a separate message.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n Makefile    |    2 \n merge-trees |  302 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 303 insertions(+), 1 deletion(-)\n\nMakefile:  b39b4ea37586693dd707d1d0750a9b580350ec50\n--- Makefile\n+++ Makefile\t2005-04-15 13:32:06.000000000 -0700\n@@ -14,7 +14,7 @@\n \n PROG=   update-cache show-diff init-db write-tree read-tree commit-tree \\\n \tcat-file fsck-cache checkout-cache diff-tree rev-tree show-files \\\n-\tcheck-files ls-tree merge-tree\n+\tcheck-files ls-tree merge-tree merge-trees\n \n all: $(PROG)\n \n\n--- /dev/null\t2005-03-19 15:28:25.000000000 -0800\n+++ merge-trees\t2005-04-15 13:32:20.000000000 -0700\n@@ -0,0 +1,302 @@\n+#!/usr/bin/perl -w\n+\n+use strict;\n+use Cwd;\n+use Getopt::Long;\n+\n+my $full_checkout = 0;\n+my $partial_checkout = 0;\n+my $output_directory = ',,merge~tree';\n+\n+GetOptions(\"full-checkout\" => \\$full_checkout,\n+\t   \"partial-checkout\" => \\$partial_checkout,\n+\t   \"output-directory=s\" => \\$output_directory)\n+    or die;\n+\n+\n+if (@ARGV != 3) {\n+    die \"Usage: $0 -o [output-directory] [-f] [-p] ancestor A B\\n\";\n+}\n+\n+if ($full_checkout) {\n+    $partial_checkout = 1;\n+}\n+\n+################################################################\n+# UI helper -- although it is encouraged to give tree ID, \n+# it is OK to give commit ID.\n+sub possibly_commit_to_tree {\n+    my ($commit_or_tree_id) = @_;\n+    my $type = read_cat_file_t($commit_or_tree_id);\n+    if ($type eq 'tree') { return $commit_or_tree_id }\n+    if ($type ne 'commit') {\n+\tdie \"Tree ID (or commit ID) required, given $type.\";\n+    }\n+\n+    my ($fhi);\n+    open $fhi, '-|', 'cat-file', 'commit', $commit_or_tree_id\n+\tor die \"$!: cat-file commit $commit_or_tree_id\";\n+    my ($tree) = <$fhi>;\n+    close $fhi;\n+    ($tree =~ s/^tree (.*)$/$1/)\n+\tor die \"$tree: Linus says the first line is guaranteed to be tree.\";\n+    return $tree;\n+}\n+\n+sub read_cat_file_t {\n+    my ($id) = @_;\n+    my ($fhi);\n+    open $fhi, '-|', 'cat-file', '-t', $id\n+\tor die \"$!: cat-file -t $id\";\n+    my ($t) = <$fhi>;\n+    close $fhi;\n+    chomp($t);\n+    return $t;\n+}\n+\n+################################################################\n+# Reads diff-tree -r output and gives a hash that maps a path\n+# to 4-tuple (old-mode new-mode old-oid new-oid).\n+# When creating, old-* are undef.  When removing, new-* are undef.\n+\n+sub OLD_MODE () { 0 }\n+sub NEW_MODE () { 1 }\n+sub OLD_OID ()  { 2 }\n+sub NEW_OID ()  { 3 }\n+\n+sub read_diff_tree {\n+    my (@tree) = @_;\n+    my ($fhi);\n+\n+    # Regular expression piece for mode\n+    my $reM  = '[0-7]+';\n+\n+    # Regular expression piece for object ID.\n+    # There is a talk about base-64 so better make it easier to modify...\n+    my $reID = '[0-9a-f]{40}';\n+\n+    local ($_, $/);\n+    $/ = \"\\0\"; \n+    my %path;\n+    open $fhi, '-|', 'diff-tree', '-r', @tree\n+\tor die \"$!: diff-tree -r @tree\";\n+    while (<$fhi>) {\n+\tchomp;\n+\tif (/^\\*($reM)->($reM)\\tblob\\t($reID)->($reID)\\t(.*)$/so) {\n+\t    $path{$5} = [$1, $2, $3, $4]; # modified\n+\t}\n+\telsif (/^\\+($reM)\\tblob\\t($reID)\\t(.*)$/so) {\n+\t    $path{$3} = [undef, $1, undef, $2]; # added\n+\t}\n+\telsif (/^\\-($reM)\\tblob\\t($reID)\\t(.*)$/so) {\n+\t    $path{$3} = [$1, undef, $2, undef]; # deleted\n+\t}\n+\telse {\n+\t    die \"cannot parse diff-tree output: $_\";\n+\t}\n+    }\n+    close $fhi;\n+    return %path;\n+}\n+\n+################################################################\n+# Read show-files output to figure out the set of files contained\n+# in the tree.  This is used to figure out what ancestor had.\n+sub read_show_files {\n+    my ($fhi);\n+    local ($_, $/);\n+    $/ = \"\\0\"; \n+    open $fhi, '-|', 'show-files', '-z', '--cached'\n+\tor die \"$!: show-files -z --cached\";\n+    my (@path) = map { chomp; $_ } <$fhi>;\n+    close $fhi;\n+    return @path;\n+}\n+\n+################################################################\n+# Given path and info (typically returned from read_diff_tree),\n+# create the file in the working directory to match the NEW tree.\n+# This does not touch dircache.\n+sub checkout_file {\n+    my ($path, $info) = @_;\n+    my (@elt) = split(/\\//, $path);\n+    my $j = '';\n+    my $tail = pop @elt;\n+    my ($fhi, $fho);\n+    for (@elt) {\n+\tmkdir \"$j$_\";\n+\t$j = \"$j$_/\";\n+    }\n+    open $fho, '>', \"$path\";\n+    open $fhi, '-|', 'cat-file', 'blob', $info->[NEW_OID]\n+\tor die \"$!: cat-file blob $info->[NEW_OID]\";\n+    while (<$fhi>) {\n+\tprint $fho $_;\n+    }\n+    close $fhi;\n+    close $fho;\n+    chmod oct(\"0$info->[NEW_MODE]\"), \"$path\";\n+}\n+\n+################################################################\n+# Given path and info record the file in the dircache without\n+# affecting working directory.\n+sub record_file {\n+    my ($path, $info) = @_;\n+    system ('update-cache', '--add', '--cacheinfo',\n+\t    $info->[NEW_MODE], $info->[NEW_OID], $path);\n+}\n+\n+################################################################\n+# Merge info from two trees and leave it in path, without\n+# affecting dircache.\n+sub merge_tree {\n+    my ($path, $infoA, $infoB) = @_;\n+    checkout_file(\"$path~A~\", $infoA);\n+    checkout_file(\"$path~B~\", $infoB);\n+    system 'checkout-cache', $path;\n+    rename $path, \"$path~O~\";\n+    my ($fhi, $fho);\n+    open $fhi, '-|', 'merge', '-p', \"$path~A~\", \"$path~O~\", \"$path~B~\";\n+    open $fho, '>', $path;\n+    local ($/);\n+    while (<$fhi>) { print $fho $_; }\n+    close $fhi;\n+    close $fho;\n+    # There is no reason to prefer infoA over infoB but\n+    # we need to pick one.\n+    chmod oct(\"0$infoA->[NEW_MODE]\"), $path;\n+}\n+\n+################################################################\n+\n+# O stands for \"the original\".  A and B are being merged.\n+my ($treeO, $treeA, $treeB) = map { possibly_commit_to_tree $_ } @ARGV;\n+\n+# Create a temporary directory and go there.\n+system('rm', '-rf', $output_directory) == 0 &&\n+system('mkdir', '-p', \"$output_directory/.git\") == 0 &&\n+symlink(Cwd::getcwd . \"/.git/objects\", \"$output_directory/.git/objects\") &&\n+chdir $output_directory &&\n+system('read-tree', $treeO) == 0\n+    or die \"$!: Failed to set up merge working area $output_directory\";\n+\n+# Find out edits done in each branch.\n+my %treeA = read_diff_tree($treeO, $treeA);\n+my %treeB = read_diff_tree($treeO, $treeB);\n+\n+# The list of files that was in the ancestor.\n+my @ancestor_file = read_show_files();\n+my %ancestor_file = map { $_ => 1 } @ancestor_file;\n+\n+# Report output is formated as follows:\n+#\n+# The first letter shows the origin of the result.\n+#   O - original\n+#   A - treeA\n+#   B - treeB\n+#   M - both treeA and treeB\n+#   * - treeA and treeB conflicts; needs human action.\n+#\n+# The second and third letter shows what each tree did.\n+#   . - no change\n+#   A - created\n+#   M - modified\n+#   D - deleted\n+\n+for (@ancestor_file) {\n+    if (! exists $treeA{$_} && ! exists $treeB{$_}) {\n+\tif ($full_checkout) {\n+\t    system 'checkout-cache', $_;\n+\t}\n+\tprint STDERR \"O.. $_\\n\"; # keep original\n+    }\n+}\n+\n+for my $set ([\\%treeA, \\%treeB, 'A'], [\\%treeB, \\%treeA, 'B']) {\n+    my ($this, $other, $side) = @$set;\n+    my $delete_sign = ($side eq 'A') ? 'D.' : '.D';\n+    my $create_sign = ($side eq 'A') ? 'A.' : '.A';\n+    my $modify_sign = ($side eq 'A') ? 'M.' : '.M';\n+    while (my ($path, $info) = each %$this) {\n+\t# In this loop we do not deal with overlaps.\n+\tnext if (exists $other->{$path});\n+\n+\tif (! defined $info->[NEW_OID]) {\n+\t    # deleted in this tree only.\n+\t    unlink $path;\n+\t    system 'update-cache', '--remove', $path;\n+\t    print STDERR \"${side}${delete_sign} $path\\n\";\n+\t}\n+\telse {\n+\t    # modified or created in this tree only.\n+\t    my $create_or_modify =\n+\t\t(! defined $info->[OLD_OID]) ? $create_sign : $modify_sign;\n+\t    print STDERR \"${side}${create_or_modify} $path\\n\";\n+\t    if ($partial_checkout) {\n+\t\tcheckout_file($path, $info);\n+\t\tsystem 'update-cache', '--add', $path;\n+\t    } else {\n+\t\trecord_file($path, $info);\n+\t    }\n+\t}\n+    }\n+}\n+\n+my @warning = ();\n+\n+while (my ($path, $infoA) = each %treeA) {\n+    # We need to deal only with overlaps.\n+    next if (!exists $treeB{$path});\n+\n+    my $infoB = $treeB{$path};\n+    if (! defined $infoA->[NEW_OID]) {\n+\t# Deleted in tree A.\n+\tif (! defined $infoB->[NEW_OID]) {\n+\t    # Deleted in both trees (obvious).\n+\t    print STDERR \"MDD $path\\n\";\n+\t    unlink $path;\n+\t    system 'update-cache', '--remove', $path;\n+\t}\n+\telse {\n+\t    # TreeA wants to remove but TreeB wants to modify it.\n+\t    print STDERR \"*DM $path\\n\";\n+\t    checkout_file(\"$path~B~\", $infoB);\n+\t    push @warning, $path;\n+\t}\n+    }\n+    else {\n+\t# Modified or created in tree A\n+\tif (! defined $infoB->[NEW_OID]) {\n+\t    # TreeA wants to modify but treeB wants to remove it.\n+\t    print STDERR \"*MD $path\\n\";\n+\t    checkout_file(\"$path~A~\", $infoA);\n+\t    push @warning, $path;\n+\t}\n+\telse {\n+\t    # Modified both in treeA and treeB.\n+\t    # Are they modifying to the same contents?\n+\t    if ($infoA->[NEW_OID] eq $infoB->[NEW_OID]) {\n+\t\t# No changes or just the mode.\n+\t\t# we prefer TreeA over TreeB for no particular reason.\n+\t\tprint STDERR \"MMM $path\\n\";\n+\t\trecord_file($path, $infoA);\n+\t    }\n+\t    else {\n+\t\t# Modified in both.  Needs merge.\n+\t\tprint STDERR \"*MM $path\\n\";\n+\t\tmerge_tree($path, $infoA, $infoB);\n+\t    }\n+\t}\n+    }\n+}\n+\n+if (@warning) {\n+    print \"\\nThere are some files that were deleted in one branch and\\n\"\n+\t. \"modified in another.  Please examine them carefully:\\n\";\n+    for (@warning) {\n+\tprint \"$_\\n\";\n+    }\n+}\n+\n+# system 'show-diff', '-q';\n\n"},{"id":"251","messageId":"7vbr8fhfjv.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"7vfyxrhfsw.fsf_-_@assigned-by-dhcp.cox.net","subject":"[PATCH 2/2] merge-trees script for Linus git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T21:54:12Z","receivedAt":"2005-04-15T21:54:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus,\n\nThis is the '-q' option for show-diff.c to squelch complaints for\nmissing files.  It is handy if you want to run it in the merge\ntemporary directory after running merge-trees with its minimum\ncheckout mode, which is the default, because you would not find any\nfiles other than the ones that needs human validation after the merge\nthere.\n\nIt also fixes the argument parsing bug Paul Mackerras noticed in\n<16991.42305.118284.139777@cargo.ozlabs.ibm.com> but slightly\ndifferently.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n \n show-diff.c |   17 ++++++++++++-----\n 1 files changed, 12 insertions(+), 5 deletions(-)\n\nshow-diff.c:  3f7acd2a692a03026784a18f28521b9af322b71e\n--- show-diff.c\n+++ show-diff.c\t2005-04-15 14:14:53.000000000 -0700\n@@ -58,15 +58,20 @@\n int main(int argc, char **argv)\n {\n \tint silent = 0;\n+\tint silent_on_nonexisting_files = 0;\n \tint entries = read_cache();\n \tint i;\n \n-\twhile (argc-- > 1) {\n-\t\tif (!strcmp(argv[1], \"-s\")) {\n-\t\t\tsilent = 1;\n+\tfor (i = 1; i < argc; i++) {\n+\t\tif (!strcmp(argv[i], \"-s\")) {\n+\t\t\tsilent_on_nonexisting_files = silent = 1;\n \t\t\tcontinue;\n \t\t}\n-\t\tusage(\"show-diff [-s]\");\n+\t\tif (!strcmp(argv[i], \"-q\")) {\n+\t\t\tsilent_on_nonexisting_files = 1;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tusage(\"show-diff [-s] [-q]\");\n \t}\n \n \tif (entries < 0) {\n@@ -82,8 +87,10 @@\n \t\tvoid *new;\n \n \t\tif (stat(ce->name, &st) < 0) {\n+\t\t\tif (errno == ENOENT && silent_on_nonexisting_files)\n+\t\t\t\tcontinue;\n \t\t\tprintf(\"%s: %s\\n\", ce->name, strerror(errno));\n-\t\t\tif (errno == ENOENT && !silent)\n+\t\t\tif (errno == ENOENT)\n \t\t\t\tshow_diff_empty(ce);\n \t\t\tcontinue;\n \t\t}\n\n\n"},{"id":"252","messageId":"20050415223648.GP7417@pasky.ji.cz","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504151212160.7211@ppc970.osdl.org","subject":"Re: Re: Re: write-tree is pasky-0.4","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-15T22:36:48Z","receivedAt":"2005-04-15T22:36:48Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Fri, Apr 15, 2005 at 10:13:21PM CEST, I got a letter\nwhere Linus Torvalds <torvalds@osdl.org> told me that...\n> \n> \n> On Fri, 15 Apr 2005, Petr Baudis wrote:\n> > \n> > So, I assume that you don't want to merge my \"SCM layer\" (which is\n> > perfectly fine by me). However, I also apply plenty of patches\n> > concerning the \"core git\" - be it portability, leak fixes, argument\n> > parsing fixes and so on.\n> \n> I'm actually perfectly happy to merge your SCM layer too eventually, but \n> I'm nervous at this point.  Especially while people are discussing some \n> SCM options that I'm personally very leery of, and think that may make \n> sense for others, but that I personally distrust.\n\nYou mean the renames tracking and similar yet mostly theoretical\ndiscussions? Or do you dislike something already implemented? I'd be\nhappy to hear about it in that case. (To argue about it and likely get\npersuaded... ;-)\n\nBut otherwise it is great news to me. Actually, in that case, is it\nworth renaming it to Cogito and using cg to invoke it? Wouldn't be that\nactually more confusing after it gets merged? IOW, should I stick to\n\"git\" or feel free to rename it to \"cg\"?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"253","messageId":"7vvf6nfyt8.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"20050415204033.GG7417@pasky.ji.cz","subject":"Re: Merge with git-pasky II.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T22:41:07Z","receivedAt":"2005-04-15T22:41:07Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n\nPB> I can't see the conflicts between what I want and what Linus wants.\nPB> After all, Linus says that I can use the directory cache in any way I\nPB> please (well, the user can, but I'm speaking for him ;-). So I'm doing\nPB> so, and with your tool I would get into problems, since it is suddenly\nPB> imposing a policy on what should be in the index.\n\nI think our misunderstanding is coming from the use of the word\n\"merge tree\".  I think you have been assuming that I wanted you\nto run \"merge-trees -o ,,merge\" --- which would certainly cause\nme to muck with your dircache there.  I totally agree with you\nthat that is a *BAD* *THING*.  No question there.\n\nHowever, my assumption has been different.  I was assuming that\nyou would run \"merge-trees -o merge~tree\" (i.e. different from\nyour \"merge tree\"), so that you can get the merge results in a\nform parsable by you.  And then, using that information, you can\nmake your changes in ,,merge.  After you are done with that\ninformation, you can remove \"merge~trees\", of course.\n\nThe format I chose for the \"merge result in a form parsable by\nyou\" happens to be a dircache in \"merge~tree\", with minimum\nnumber of files checked out when merge cannot be automatically\ndone safely.  In the simplest case of not having any conflicting\nmerge between $C and $merged, Cogito can immediately run\nwrite-tree in \"merge~tree\" (not ,,merge) to obtain its tree-ID\n$T, so that it can feed it to diff-tree to compare it with\nwhatever tree state Cogito wants to apply the merges between $C\nand $merged to.\n\nI still do not understand what you do in ,,merge directory, but\nhere is one way you can update the user working directory\nin-place without having a ,,merge directory [*2*].  You can run\nyour \"git diff\" between $C and $T [*1*].  The result is the diff\nyou need to apply on top of your user's working files.  If the\nuser does not like the result of running that diff, it can\neasily be reversed.\n\nIf a manual merge were needed between $C and $merged, Cogito\ncould guide the user through that manual edit in \"merge~tree\",\nand run update-cache on those hand merged files in \"merge~tree\",\nbefore running write-tree in \"merge~tree\" to obtain $T; after\nthat, everything else is the same.\n\nYou make interesting points in other parts of your message I\nneed to regurgitate for a while, so I would not comment on them\nin this message.\n\n[Footnote]\n\n*1* I really like the convenience of being able to use tree-ID\nand commit-ID interchangeably there.  Thanks.\n\n*2* I understand that this would change the user's \"git-tools\"\nexperience a bit.  The user will not be told to \"go to ,,merge\nand commit there which will reflected back to your working tree\"\nanymore.  Instead the merge happens in-place.  Committing, not\ncommitting, or further hand-fixing the merge is up to the user.\nI suspect this change might even be for the better.\n\n"},{"id":"255","messageId":"7vr7hbfx66.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.61.0504151617170.27637@cag.csail.mit.edu","subject":"Re: write-tree is pasky-0.4","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T23:16:33Z","receivedAt":"2005-04-15T23:16:33Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"CSA\" == C Scott Ananian <cscott@cscott.net> writes:\n\nCSA> On Fri, 15 Apr 2005, Junio C Hamano wrote:\n>> to yours is no problem for me.  Currently I see your HEAD is at\n>> 461aef08823a18a6c69d472499ef5257f8c7f6c8, so I will generate a\n>> set of patches against it.\n\nCSA> Have you considered using an s/key-like system to make these hashes\nCSA> more human-readable?  Using the S/Key translation (11-bit chunks map\nCSA> to a 1-4\nCSA> letter word), Linus' HEAD is at:\nCSA>    WOW-SCAN-NAVE-AUK-JILL-BASH-HI-LACE-LID-RIDE-RUSE-LINE-GLEE-WICK-A\nCSA> ...which is a little longer, but speaking of branch \"wow-scan\" (which\nCSA> gives 22 bits of disambiguation) is probably less error-prone than\nCSA> discussing branch '461...' (only 12 bits).\n\nI understand monotone folks have the same issue and they let you\nuse unambiguous prefix string.  And why do you stop counting at\n\"461\" in your example?  To my eyes, \"461aef\" in this particular\nstring stands out and is easily typable, which gives me 24 bits\n;-).\n\nBut seriously I doubt the hex format is needed to be shown to\nhumans very often.  E-mail communications like this one being a\nvery special exception.  I do not expect for people to be\ntalking about \"Hey, Junio's patch against 461aef... from Linus\nis a total crap\" like that.\n\nThe only reason I mentioned his then-HEAD by hex is because I do\nnot have a public archive for him to pull from, and I wanted to\nmake it easy for him to do:\n\n $ export SHA1_FILE_DIRECTORY\n $ mkdir junk && cd junk && mkdir .git &&\n   read-tree `cat-file commit 461aef... | sed -e 's/^tree //;q'`\n $ patch < ../stupid-patch-from-junio-01\n $ show-diff\n\n(it might have been better if I used the tree ID for this purpose).\n\nFor Cogito users the hex format does not matter.  \"git pull\"\nwill get whatever HEAD recorded in the file on the sending end\nand the end user does not even have to know about it.\n\nCSA> This is obviously a cogito issue, rather than a git-fs thing.\n\nYes.\n\n"},{"id":"256","messageId":"7vmzrzfwe4.fsf_-_@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"7vfyxrhfsw.fsf_-_@assigned-by-dhcp.cox.net","subject":"[PATCH 3/2] merge-trees script for Linus git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-15T23:33:23Z","receivedAt":"2005-04-15T23:33:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus,\n\n    the merge-trees I sent you earlier was expecting the old\ndiff-tree behaviour, and I did not realize that I need an\nexplicit -z flag now.  Here is a fix.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n merge-trees |    4 ++--\n 1 files changed, 2 insertions(+), 2 deletions(-)\n\n--- merge-trees\t2005-04-15 13:21:35.000000000 -0700\n+++ merge-trees+\t2005-04-15 16:27:34.000000000 -0700\n@@ -78,8 +78,8 @@\n     local ($_, $/);\n     $/ = \"\\0\"; \n     my %path;\n-    open $fhi, '-|', 'diff-tree', '-r', @tree\n-\tor die \"$!: diff-tree -r @tree\";\n+    open $fhi, '-|', 'diff-tree', '-r', '-z', @tree\n+\tor die \"$!: diff-tree -r -z @tree\";\n     while (<$fhi>) {\n \tchomp;\n \tif (/^\\*($reM)->($reM)\\tblob\\t($reID)->($reID)\\t(.*)$/so) {\n\n"},{"id":"262","messageId":"Pine.LNX.4.58.0504151709180.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"20050415223648.GP7417@pasky.ji.cz","subject":"Re: Re: Re: write-tree is pasky-0.4","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T00:22:45Z","receivedAt":"2005-04-16T00:22:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Apr 2005, Petr Baudis wrote:\n> \n> But otherwise it is great news to me. Actually, in that case, is it\n> worth renaming it to Cogito and using cg to invoke it? Wouldn't be that\n> actually more confusing after it gets merged? IOW, should I stick to\n> \"git\" or feel free to rename it to \"cg\"?\n\nI'm perfectly happy for it to stay as \"git\", and in general I don't have\nany huge preferences either way. You guys can discuss names as much as you\nlike, it's the \"tracking renames\" and \"how to merge\" things that worry me.\n\nI think I've explained my name tracking worries.  When it comes to \"how to \nmerge\", there's three issues:\n\n - we do commonly have merge clashes where both trees have applied the \n   exact same patch. That should merge perfectly well using the 3-way\n   merge from a common parent that Junio has, but not your current \"bring\n   patches forward\" kind of strategy.\n - I _do_ actually sometimes merge with dirty state in my working \n   directory, which is why I want the merge to take place in a separate \n   (and temporary) directory, which allows for a failed merge without \n   having any major cleanup. If the merge fails, it's not a big deal, and \n   I can just blow the merge directory away without losing the work I had \n   in my \"real\" working directory.\n - reliability. I care much less for \"clever\" than I care for \"guaranteed \n   to never do the wrong thing\". If I have to fix up some stuff by hand, \n   I'll happily do so. But if I can't trust the merge and have to _check_ \n   things by hand afterwards, that will make me leery of the merges, and\n   _that_ is bad.\n\nThe third point is why I'm going to the ultra-conservative \"three-way \nmerge from the common parent\". It's not fancy, but it's something I feel \ncomfortable with as a merge strategy. For example, arch (and in particular \ndarcs) seems to want to try to be \"clever\" about the merges, and I'd \nalways live in fear. \n\nAnd, finally, there's obviously performance. I _think_ a normal merge with\nnary a conflict and just a few tens of files changed should be possible in\na second. I realize that sounds crazy to some people, but I think it's\nentirely doable. Half of that is writing the new tree out (that is a\nrelative costly op due to the compression). The other half is the \"work\".\n\n\t\tLinus\n"},{"id":"266","messageId":"Pine.LNX.4.58.0504151755590.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"7vmzrzfwe4.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 3/2] merge-trees script for Linus git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T01:02:07Z","receivedAt":"2005-04-16T01:02:07Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, Junio C Hamano wrote:\n> \n>     the merge-trees I sent you earlier was expecting the old\n> diff-tree behaviour, and I did not realize that I need an\n> explicit -z flag now.\n\nYou didn't need one - I just didn't want to merge your \"ls-tree\" change\nwithout making things be consistent. Once we started using the \"-z\" flag \nfor ls-tree, it just didn't make any sense not to do the same thing for \ndiff-tree.\n\nJust a heads-up - I'd really want to do the same thing to \"merge-tree.c\" \ntoo, but since you said that you were working on extending that to do \nrecursion etc, I decided to hold off. So if you're working on it, maybe \nyou can add the \"-z\" flag there too? \n\nI'm actually holding off merging the perl version exactly because you \nseemed to be working on the C version. I don't mind perl per se, but if \nthere's a real solution coming down the line..\n\n\t\tLinus\n"},{"id":"267","messageId":"Pine.LNX.4.21.0504152029410.30848-100000@iabervon.org","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504151709180.7211@ppc970.osdl.org","subject":"Re: Re: Re: write-tree is pasky-0.4","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-04-16T01:13:05Z","receivedAt":"2005-04-16T01:13:05Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 15 Apr 2005, Linus Torvalds wrote:\n\n> I think I've explained my name tracking worries.  When it comes to \"how to \n> merge\", there's three issues:\n> \n>  - we do commonly have merge clashes where both trees have applied the \n>    exact same patch. That should merge perfectly well using the 3-way\n>    merge from a common parent that Junio has, but not your current \"bring\n>    patches forward\" kind of strategy.\n\nI think 3-way merge is probably the best starting point, but I think that\nthere might be value in being able to identify the commits of each side\ninvolved in a conflict. I think this would help with cases where both\nsides pick up an identical patch, and then each side makes a further\nchange to a different part of the changed region (you find out that the\nother guy's change was supposed to follow the patch, and don't conflict\nwith it).\n\n>  - I _do_ actually sometimes merge with dirty state in my working \n>    directory, which is why I want the merge to take place in a separate \n>    (and temporary) directory, which allows for a failed merge without \n>    having any major cleanup. If the merge fails, it's not a big deal, and \n>    I can just blow the merge directory away without losing the work I had \n>    in my \"real\" working directory.\n\nIs there some reason you don't commit before merging? All of the current\nmerge theory seems to want to merge two commits, using the information git\nkeeps about them. It should be cheap to get a new clean working directory\nto merge in, too, particularly if we add a cache of hardlinkable expanded\nblobs.\n\n>  - reliability. I care much less for \"clever\" than I care for \"guaranteed \n>    to never do the wrong thing\". If I have to fix up some stuff by hand, \n>    I'll happily do so. But if I can't trust the merge and have to _check_ \n>    things by hand afterwards, that will make me leery of the merges, and\n>    _that_ is bad.\n> \n> The third point is why I'm going to the ultra-conservative \"three-way \n> merge from the common parent\". It's not fancy, but it's something I feel \n> comfortable with as a merge strategy. For example, arch (and in particular \n> darcs) seems to want to try to be \"clever\" about the merges, and I'd \n> always live in fear. \n\nHow much do you care about the situation where there is no best common\nancestor (which can happen if you're merging two main lines, each of which\nhas merged with both of a pair of minor trees)? I think that arch is even\nmore conservative, in that it doesn't look for a common ancestor, and\nreports conflicts whenever changes overlap at all. Of course, reliability\nby virtue of never working without help is not a big win over living in\nfear; you always have to check over it, not because you're afraid, but\nbecause it needs you to.\n\n> And, finally, there's obviously performance. I _think_ a normal merge with\n> nary a conflict and just a few tens of files changed should be possible in\n> a second. I realize that sounds crazy to some people, but I think it's\n> entirely doable. Half of that is writing the new tree out (that is a\n> relative costly op due to the compression). The other half is the \"work\".\n\nI think that the time spent on I/O will be overwhelmed by the time spent\nissuing the command at that rate. It might matter if you start getting\ninto merging lots of things at once, but that's more like a minute for a\nmerge group with 600 changes rather than a second per merge; we could\npotentially save a lot of time based of having a bunch of information left\nover from the previous merge when starting merge number 2. So 15 seconds\nplus half a second per merge might be better than a second per merge in\nthe case that matters.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"269","messageId":"20050416014442.GW4488@himi.org","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504150753440.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Simon Fowler","fromEmail":"simon@himi.org","sentAt":"2005-04-16T01:44:42Z","receivedAt":"2005-04-16T01:44:42Z","isPatch":false,"sender":{"key":"simon@himi.org","avatar":null},"body":"On Fri, Apr 15, 2005 at 08:32:46AM -0700, Linus Torvalds wrote:\n> In other words, I'm right. I'm always right, but sometimes I'm more right \n> than other times. And dammit, when I say \"files don't matter\", I'm really \n> really Right(tm).\n> \nYou're right, of course (All Hail Linus!), if you can make it work\nefficiently enough.\n\nJust to put something else on the table, here's how I'd go about\ntracking renames and the like, in another world where Linus /does/\nmake the odd mistake - it's basically a unique id for files in the\nrepository, added when the file is first recognised and updated when\nupdate-cache adds a new version to the cache. Renames copy the id\nacross to the new name, and add it into the cache.\n\nThis gives you an O(n) way to tell what file was what across\nrenames, and it might even be useful in Linus' world, or if someone\nwanted to build a traditional SCM on top of a git-a-like.\n\nAttached is a patch, and a rename-file.c to use it.\n\nSimon\n\n-- \nPGP public key Id 0x144A991C, or http://himi.org/stuff/himi.asc\n(crappy) Homepage: http://himi.org\ndoe #237 (see http://www.lemuria.org/DeCSS) \nMy DeCSS mirror: ftp://himi.org/pub/mirrors/css/ \n\n\nCOPYING:  fe2a4177a760fd110e78788734f167bd633be8de\nMakefile:  ca50293c4f211452d999b81f122e99babb9f2987\n--- Makefile\n+++ Makefile\t2005-04-15 22:17:49.000000000 +1000\n@@ -14,7 +14,7 @@\n \n PROG=   update-cache show-diff init-db write-tree read-tree commit-tree \\\n \tcat-file fsck-cache checkout-cache diff-tree rev-tree show-files \\\n-\tcheck-files ls-tree\n+\tcheck-files ls-tree rename-file\n \n SCRIPT=\tparent-id tree-id git gitXnormid.sh gitadd.sh gitaddremote.sh \\\n \tgitcommit.sh gitdiff-do gitdiff.sh gitlog.sh gitls.sh gitlsobj.sh \\\n@@ -73,6 +73,9 @@\n ls-tree: ls-tree.o read-cache.o\n \t$(CC) $(CFLAGS) -o ls-tree ls-tree.o read-cache.o $(LIBS)\n \n+rename-file: rename-file.o read-cache.o\n+\t$(CC) $(CFLAGS) -o rename-file rename-file.o read-cache.o $(LIBS)\n+\n read-cache.o: cache.h\n show-diff.o: cache.h\n \nREADME:  ded1a3b20e9bbe1f40e487ba5f9361719a1b6b85\nVERSION:  c27bd67cd632cc15dd520fbfbf807d482efa2dcf\ncache.h:  4d382549041d3281f8d44aa2e52f9f8ec47dd420\n--- cache.h\n+++ cache.h\t2005-04-14 22:35:59.000000000 +1000\n@@ -55,6 +55,7 @@\n \tunsigned int st_gid;\n \tunsigned int st_size;\n \tunsigned char sha1[20];\n+\tunsigned char guid[20];\n \tunsigned short namelen;\n \tchar name[0];\n };\ncat-file.c:  45be1badaa8517d4e3a69e0bf1cac2e90191e475\ncheck-files.c:  927b0b9aca742183fc8e7ccd73d73d8d5427e98f\ncheckout-cache.c:  f06871cdbc1b18ea93bdf4e17126aeb4cca1373e\ncommit-id:  65c81756c8f10d513d073ecbd741a3244663c4c9\ncommit-tree.c:  12196c79f31d004dff0df1f50dda67d8204f5568\ndiff-tree.c:  7dcc9eb7782fa176e27f1677b161ce78ac1d2070\n--- diff-tree.c\n+++ diff-tree.c\t2005-04-16 10:46:52.000000000 +1000\n@@ -1,33 +1,144 @@\n+#include <sys/param.h>\n #include \"cache.h\"\n \n-static int recursive = 0;\n+enum diff_type {\n+\tREMOVE,\n+\tADD,\n+\tRENAME,\n+\tMODIFY,\n+};\n+\n+struct guid_cache_entry {\n+\tenum diff_type diff;\n+\tunsigned char guid[20];\n+\tunsigned char sha1[20];\n+\tstruct guid_cache_entry *old;\n+\tunsigned int mode;\n+\tunsigned int pathlen;\n+\tunsigned char path[0];\n+};\t\n+\n+struct guid_cache {\n+\tunsigned int nr;\n+\tunsigned int alloc;\n+\tstruct guid_cache_entry **cache;\n+};\n \n-static int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base);\n+struct guid_cache guid_cache;\n+struct guid_cache *cache = &guid_cache;\n \n-static void update_tree_entry(void **bufp, unsigned long *sizep)\n+int guid_cache_pos(const char *guid)\n {\n-\tvoid *buf = *bufp;\n-\tunsigned long size = *sizep;\n-\tint len = strlen(buf) + 1 + 20;\n+\tint first, last;\n \n-\tif (size < len)\n-\t\tdie(\"corrupt tree file\");\n-\t*bufp = buf + len;\n-\t*sizep = size - len;\n+\tfirst = 0;\n+\tlast = cache->nr;\n+\twhile (last > first) {\n+\t\tint next = (last + first) >> 1;\n+\t\tstruct guid_cache_entry *gce = cache->cache[next];\n+\t\tint cmp = memcmp(guid, gce->guid, 20);\n+\t\tif (!cmp)\n+\t\t\treturn next;\n+\t\tif (cmp < 0) {\n+\t\t\tlast = next;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tfirst = next + 1;\n+\t}\n+\treturn - first-1;\n }\n \n-static const unsigned char *extract(void *tree, unsigned long size, const char **pathp, unsigned int *modep)\n+int add_guid_cache_entry(struct guid_cache_entry *gce)\n+{\n+\tint pos;\n+\t\n+\tpos = guid_cache_pos(gce->guid);\n+\t\n+\t/* if this is a rename or modify, the guid will show up a\n+\t * second time */\n+\tif (pos >= 0) {\n+\t\tstruct guid_cache_entry *old = cache->cache[pos];\n+\t\tint cmp = cache_name_compare(old->path, old->pathlen, gce->path, gce->pathlen);\n+\n+\t\tif (!cmp) {\n+\t\t\t/* pathname matches, so this must be a\n+\t\t\t * modify. */\n+\t\t\tgce->old = old;\n+\t\t\tgce->diff = MODIFY;\n+\t\t\tcache->cache[pos] = gce;\n+\t\t} else {\n+\t\t\t/* the pathnames are different, so the file\n+\t\t\t * must have been renamed somewhere along the\n+\t\t\t * line.  */\n+\t\t\tgce->old = old;\n+\t\t\tgce->diff = RENAME;\n+\t\t\tcache->cache[pos] = gce;\n+\t\t}\n+\t\treturn 0;\n+\t}\n+\tpos = -pos-1;\n+\n+\tif (cache->nr == cache->alloc) {\n+\t\tcache->alloc = alloc_nr(cache->alloc);\n+\t\tcache->cache = realloc(cache->cache, cache->alloc * sizeof(struct guid_cache_entry *));\n+\t}\n+\n+\tcache->nr++;\n+\tif (cache->nr > pos)\n+\t\tmemmove(cache->cache + pos + 1, cache->cache + pos, (cache->nr - pos - 1) * sizeof(struct guid_cache_entry *));\n+\tcache->cache[pos] = gce;\n+\treturn 0;\n+}\n+\n+static const unsigned char *extract(void *tree, unsigned long size, const char **pathp, unsigned int *modep, const unsigned char **guid)\n {\n \tint len = strlen(tree)+1;\n \tconst unsigned char *sha1 = tree + len;\n \tconst char *path = strchr(tree, ' ');\n \n-\tif (!path || size < len + 20 || sscanf(tree, \"%o\", modep) != 1)\n+\tif (!path || size < len + 40 || sscanf(tree, \"%o\", modep) != 1)\n \t\tdie(\"corrupt tree file\");\n \t*pathp = path+1;\n+\t*guid = tree + len + 20;\n \treturn sha1;\n }\n \n+static void guid_cache_tree_entry(void *buf, unsigned int len, const char *base, enum diff_type diff)\n+{\n+\tunsigned mode;\n+\tconst char *path;\n+\tconst unsigned char *guid;\n+\tconst unsigned char *sha1 = extract(buf, len, &path, &mode, &guid);\n+\tstruct guid_cache_entry *gce;\n+\tint baselen = strlen(base);\n+\t\n+\tgce = calloc(1, sizeof(struct guid_cache_entry) + baselen + strlen(path) + 1);\n+\tmemcpy(gce->guid, guid, 20);\n+\tmemcpy(gce->sha1, sha1, 20);\n+\tgce->diff = diff;\n+\tgce->mode = mode;\n+\tgce->pathlen = snprintf(gce->path, MAXPATHLEN, \"%s%s\", base, path);\n+\tgce->path[gce->pathlen + 1] = '\\0';\n+\n+\tadd_guid_cache_entry(gce);\n+}\n+\t\n+static int recursive = 0;\n+\n+static int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base);\n+\n+static void update_tree_entry(void **bufp, unsigned long *sizep)\n+{\n+\tvoid *buf = *bufp;\n+\tunsigned long size = *sizep;\n+\tint len = strlen(buf) + 1 + 40;\n+\n+\tif (size < len)\n+\t\tdie(\"corrupt tree file\");\n+\t*bufp = buf + len;\n+\t*sizep = size - len;\n+}\n+\n static char *malloc_base(const char *base, const char *path, int pathlen)\n {\n \tint baselen = strlen(base);\n@@ -38,23 +149,24 @@\n \treturn newbase;\n }\n \n-static void show_file(const char *prefix, void *tree, unsigned long size, const char *base);\n+static void changed_file(void *tree, unsigned long size, const char *base, enum diff_type diff);\n \n /* A whole sub-tree went away or appeared */\n-static void show_tree(const char *prefix, void *tree, unsigned long size, const char *base)\n+static void changed_tree(void *tree, unsigned long size, const char *base, enum diff_type diff)\n {\n \twhile (size) {\n-\t\tshow_file(prefix, tree, size, base);\n+\t\tchanged_file(tree, size, base, diff);\n \t\tupdate_tree_entry(&tree, &size);\n \t}\n }\n \n /* A file entry went away or appeared */\n-static void show_file(const char *prefix, void *tree, unsigned long size, const char *base)\n+static void changed_file(void *tree, unsigned long size, const char *base, enum diff_type diff)\n {\n \tunsigned mode;\n \tconst char *path;\n-\tconst unsigned char *sha1 = extract(tree, size, &path, &mode);\n+\tconst unsigned char *guid;\n+\tconst unsigned char *sha1 = extract(tree, size, &path, &mode, &guid);\n \n \tif (recursive && S_ISDIR(mode)) {\n \t\tchar type[20];\n@@ -66,38 +178,96 @@\n \t\tif (!tree || strcmp(type, \"tree\"))\n \t\t\tdie(\"corrupt tree sha %s\", sha1_to_hex(sha1));\n \n-\t\tshow_tree(prefix, tree, size, newbase);\n+\t\tchanged_tree(tree, size, newbase, diff);\n \t\t\n \t\tfree(tree);\n \t\tfree(newbase);\n \t\treturn;\n \t}\n \n-\tprintf(\"%s%o\\t%s\\t%s\\t%s%s%c\", prefix, mode,\n-\t       S_ISDIR(mode) ? \"tree\" : \"blob\",\n-\t       sha1_to_hex(sha1), base, path, 0);\n+\tguid_cache_tree_entry(tree, size, base, diff);\n }\n \n+static void show_one_file(struct guid_cache_entry *gce)\n+{\n+\tstruct guid_cache_entry *old;\n+\tchar old_sha1[50];\n+\tchar old_sha2[50];\n+\n+\tswitch(gce->diff) {\n+\tcase REMOVE:\n+\t\tsprintf(old_sha1, \"%s\", sha1_to_hex(gce->sha1));\n+\t\tprintf(\"-%o\\t%s\\t%s\\t%s\\t%s%c\", gce->mode,\n+\t\t       S_ISDIR(gce->mode) ? \"tree\" : \"blob\",\n+\t\t       old_sha1, sha1_to_hex(gce->guid), gce->path, 0);\n+\t\tbreak;\n+\tcase ADD:\n+\t\tsprintf(old_sha1, \"%s\", sha1_to_hex(gce->sha1));\n+\t\tprintf(\"+%o\\t%s\\t%s\\t%s\\t%s%c\", gce->mode,\n+\t\t       S_ISDIR(gce->mode) ? \"tree\" : \"blob\",\n+\t\t       old_sha1, sha1_to_hex(gce->guid), gce->path, 0);\n+\t\tbreak;\n+\tcase MODIFY:\n+\t\told = gce->old;\n+\t\tif (old) {\n+\t\t\tsprintf(old_sha1, \"%s\", sha1_to_hex(old->sha1));\n+\t\t\tsprintf(old_sha2, \"%s\", sha1_to_hex(gce->sha1));\n+\t\t\t\n+\t\t\tprintf(\"*%o->%o\\t%s\\t%s->%s\\t%s\\t%s%c\", old->mode, gce->mode,\n+\t\t\t       S_ISDIR(old->mode) ? \"tree\" : \"blob\",\n+\t\t\t       old_sha1, old_sha2, sha1_to_hex(gce->guid), gce->path, 0);\n+\t\t} else {\n+\t\t\tdie(\"diff-tree: internal error\");\n+\t\t}\n+\t\tbreak;\n+\tcase RENAME:\n+\t\told = gce->old;\n+\t\tif (old) {\n+\t\t\tsprintf(old_sha1, \"%s\", sha1_to_hex(gce->sha1));\n+\t\t\tsprintf(old_sha2, \"%s\", sha1_to_hex(old->sha1));\n+\t\t\t\n+\t\t\tprintf(\"r%o->%o\\t%s\\t%s->%s\\t%s\\t%s%c\", gce->mode, old->mode,\n+\t\t\t       S_ISDIR(old->mode) ? \"tree\" : \"blob\",\n+\t\t\t       old_sha1, old_sha2, sha1_to_hex(old->guid), old->path, 0);\n+\t\t} else {\n+\t\t\tdie(\"diff-tree: internal error\");\n+\t\t}\n+\t\tbreak;\n+\tdefault:\n+\t\tdie(\"diff-tree: internal error\");\n+\t}\n+}\n+\n+/* simply iterate over both caches looking for matching guids,\n+ * showing all files in both caches */\n+static void show_cache(void)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i < cache->nr; i++)\n+\t\tshow_one_file(cache->cache[i]);\n+}\n+\t\n static int compare_tree_entry(void *tree1, unsigned long size1, void *tree2, unsigned long size2, const char *base)\n {\n \tunsigned mode1, mode2;\n \tconst char *path1, *path2;\n \tconst unsigned char *sha1, *sha2;\n+\tconst unsigned char *guid1, *guid2;\n \tint cmp, pathlen1, pathlen2;\n-\tchar old_sha1_hex[50];\n \n-\tsha1 = extract(tree1, size1, &path1, &mode1);\n-\tsha2 = extract(tree2, size2, &path2, &mode2);\n+\tsha1 = extract(tree1, size1, &path1, &mode1, &guid1);\n+\tsha2 = extract(tree2, size2, &path2, &mode2, &guid2);\n \n \tpathlen1 = strlen(path1);\n \tpathlen2 = strlen(path2);\n \tcmp = cache_name_compare(path1, pathlen1, path2, pathlen2);\n \tif (cmp < 0) {\n-\t\tshow_file(\"-\", tree1, size1, base);\n+\t\tchanged_file(tree1, size1, base, REMOVE);\n \t\treturn -1;\n \t}\n \tif (cmp > 0) {\n-\t\tshow_file(\"+\", tree2, size2, base);\n+\t\tchanged_file(tree2, size2, base, ADD);\n \t\treturn 1;\n \t}\n \tif (!memcmp(sha1, sha2, 20) && mode1 == mode2)\n@@ -108,8 +278,8 @@\n \t * file, we need to consider it a remove and an add.\n \t */\n \tif (S_ISDIR(mode1) != S_ISDIR(mode2)) {\n-\t\tshow_file(\"-\", tree1, size1, base);\n-\t\tshow_file(\"+\", tree2, size2, base);\n+\t\tchanged_file(tree1, size1, base, REMOVE);\n+\t\tchanged_file(tree2, size2, base, ADD);\n \t\treturn 0;\n \t}\n \n@@ -121,10 +291,14 @@\n \t\treturn retval;\n \t}\n \n-\tstrcpy(old_sha1_hex, sha1_to_hex(sha1));\n-\tprintf(\"*%o->%o\\t%s\\t%s->%s\\t%s%s%c\", mode1, mode2,\n-\t       S_ISDIR(mode1) ? \"tree\" : \"blob\",\n-\t       old_sha1_hex, sha1_to_hex(sha2), base, path1, 0);\n+\tif (!memcmp(guid1, guid2, 20)) {\n+\t\tchanged_file(tree1, size1, base, MODIFY);\n+\t\tchanged_file(tree2, size2, base, MODIFY);\n+\t\treturn 0;\n+\t}\n+\t\n+\tchanged_file(tree1, size1, base, REMOVE);\n+\tchanged_file(tree2, size2, base, ADD);\n \treturn 0;\n }\n \n@@ -132,12 +306,12 @@\n {\n \twhile (size1 | size2) {\n \t\tif (!size1) {\n-\t\t\tshow_file(\"+\", tree2, size2, base);\n+\t\t\tchanged_file(tree2, size2, base, ADD);\n \t\t\tupdate_tree_entry(&tree2, &size2);\n \t\t\tcontinue;\n \t\t}\n \t\tif (!size2) {\n-\t\t\tshow_file(\"-\", tree1, size1, base);\n+\t\t\tchanged_file(tree1, size1, base, REMOVE);\n \t\t\tupdate_tree_entry(&tree1, &size1);\n \t\t\tcontinue;\n \t\t}\n@@ -179,6 +353,7 @@\n int main(int argc, char **argv)\n {\n \tunsigned char old[20], new[20];\n+\tint retval;\n \n \twhile (argc > 3) {\n \t\tchar *arg = argv[1];\n@@ -193,5 +368,7 @@\n \n \tif (argc != 3 || get_sha1_hex(argv[1], old) || get_sha1_hex(argv[2], new))\n \t\tusage(\"diff-tree <tree sha1> <tree sha1>\");\n-\treturn diff_tree_sha1(old, new, \"\");\n+\tretval =  diff_tree_sha1(old, new, \"\");\n+\tshow_cache();\n+\treturn retval;\n }\nfsck-cache.c:  9c900fe458cecd2bdb4c4571a584115b5cf24f22\n--- fsck-cache.c\n+++ fsck-cache.c\t2005-04-15 20:39:49.000000000 +1000\n@@ -165,9 +165,10 @@\n \twhile (size) {\n \t\tint len = 1+strlen(data);\n \t\tunsigned char *file_sha1 = data + len;\n+\t\tunsigned char *guid = file_sha1 + 20;\n \t\tchar *path = strchr(data, ' ');\n \t\tunsigned int mode;\n-\t\tif (size < len + 20 || !path || sscanf(data, \"%o\", &mode) != 1)\n+\t\tif (size < len + 40 || !path || sscanf(data, \"%o\", &mode) != 1)\n \t\t\treturn -1;\n \n \t\t/* Warn about trees that don't do the recursive thing.. */\n@@ -176,8 +177,8 @@\n \t\t\twarn_old_tree = 0;\n \t\t}\n \n-\t\tdata += len + 20;\n-\t\tsize -= len + 20;\n+\t\tdata += len + 40;\n+\t\tsize -= len + 40;\n \t\tmark_needs_sha1(sha1, S_ISDIR(mode) ? \"tree\" : \"blob\", file_sha1);\n \t}\n \treturn 0;\ngit:  2c557dcf2032325acc265b577ee104e605fdaede\ngitXnormid.sh:  a5d7a9f4a6e8d4860f35f69500965c2a493d80de\ngitadd.sh:  3ed93ea0fcb995673ba9ee1982e0e7abdbe35982\ngitaddremote.sh:  bf1f28823da5b5270aa8fa05b321faa514a57a11\ngitapply.sh:  d0e3c46e2ce1ee74e1a87ee6137955fa9b35c27b\ngitcancel.sh:  ec58f7444a42cd3cbaae919fc68c70a3866420c0\ngitcommit.sh:  3629f67bbd3f171d091552814908b67af7537f4d\ngitdiff-do:  d6174abceab34d22010c36a8453a6c3f3f184fe0\ngitdiff.sh:  5e47c4779d73c3f2f39f6be714c0145175933197\ngitexport.sh:  dad00bf251b38ce522c593ea9631f842d8ccc934\ngitlntree.sh:  17c4966ea64aeced96ae4f1b00f3775c1904b0f1\ngitlog.sh:  177c6d12dd9fa4b4920b08451ffe4badde544a39\ngitls.sh:  b6f15d82f16c1e9982c5031f3be22eb5430273af\ngitlsobj.sh:  128461d3de6a42cfaaa989fc6401bebdfa885b3f\ngitmerge.sh:  23e4a3ff342c6005928ceea598a2f52de6fb9817\ngitpull.sh:  0883898dda579e3fa44944b7b1d909257f6dc63e\ngitrm.sh:  5c18c38a890c9fd9ad2b866ee7b529539d2f3f8f\ngittag.sh:  c8cb31385d5a9622e95a4e0b2d6a4198038a659c\ngittrack.sh:  03d6db1fb3a70605ef249c632c04e542457f0808\ninit-db.c:  aa00fbb1b95624f6c30090a17354c9c08a6ac596\nls-tree.c:  3e2a6c7d183a42e41f1073dfec6794e8f8a5e75c\n--- ls-tree.c\n+++ ls-tree.c\t2005-04-15 15:55:40.000000000 +1000\n@@ -10,6 +10,7 @@\n \tvoid *buffer;\n \tunsigned long size;\n \tchar type[20];\n+\tchar old_sha1[50];\n \n \tbuffer = read_sha1_file(sha1, type, &size);\n \tif (!buffer)\n@@ -19,19 +20,21 @@\n \twhile (size) {\n \t\tint len = strlen(buffer)+1;\n \t\tunsigned char *sha1 = buffer + len;\n+\t\tunsigned char *guid = buffer + len + 20;\n \t\tchar *path = strchr(buffer, ' ')+1;\n \t\tunsigned int mode;\n \t\tunsigned char *type;\n \n-\t\tif (size < len + 20 || sscanf(buffer, \"%o\", &mode) != 1)\n+\t\tif (size < len + 40 || sscanf(buffer, \"%o\", &mode) != 1)\n \t\t\tdie(\"corrupt 'tree' file\");\n-\t\tbuffer = sha1 + 20;\n-\t\tsize -= len + 20;\n+\t\tbuffer = sha1 + 40;\n+\t\tsize -= len + 40;\n \t\t/* XXX: We do some ugly mode heuristics here.\n \t\t * It seems not worth it to read each file just to get this\n \t\t * and the file size. -- pasky@ucw.cz */\n \t\ttype = S_ISDIR(mode) ? \"tree\" : \"blob\";\n-\t\tprintf(\"%03o\\t%s\\t%s\\t%s\\n\", mode, type, sha1_to_hex(sha1), path);\n+\t\tsprintf(old_sha1, sha1_to_hex(guid));\n+\t\tprintf(\"%03o\\t%s\\t%s\\t%s\\t%s\\n\", mode, type, sha1_to_hex(sha1), old_sha1, path);\n \t}\n \treturn 0;\n }\nparent-id:  1801c6fe426592832e7250f8b760fb9d2e65220f\nread-cache.c:  7a6ae8b9b489f6b67c82e065dedd5716a6bfc0ef\n--- read-cache.c\n+++ read-cache.c\t2005-04-16 10:52:51.000000000 +1000\n@@ -4,6 +4,8 @@\n  * Copyright (C) Linus Torvalds, 2005\n  */\n #include <stdarg.h>\n+#include <time.h>\n+#include <sys/param.h>\n #include \"cache.h\"\n \n const char *sha1_file_directory = NULL;\n@@ -233,6 +235,22 @@\n \treturn 0;\n }\n \n+void new_guid(const char *filename, int namelen, unsigned char *returnguid)\n+{\n+\tsize_t size;\n+\ttime_t now = time(NULL);\n+\tchar buf[MAXPATHLEN + 20];\n+\tunsigned char guid[20];\n+\t\n+\tsize = snprintf(buf, MAXPATHLEN + 20, \"%ld%s\", now, filename) + 1;\n+\t\n+\tSHA1(buf, size, guid);\n+\n+\tif (returnguid)\n+\t\tmemcpy(returnguid, guid, 20);\n+\treturn;\n+}\n+\n static inline int collision_check(char *filename, void *buf, unsigned int size)\n {\n #ifdef COLLISION_CHECK\n@@ -363,11 +381,14 @@\n int add_cache_entry(struct cache_entry *ce, int ok_to_add)\n {\n \tint pos;\n+\tunsigned char guid[20];\n \n \tpos = cache_name_pos(ce->name, ce->namelen);\n \n \t/* existing match? Just replace it */\n \tif (pos >= 0) {\n+\t\tstruct cache_entry *old_ce = active_cache[pos];\n+\t\tmemcpy(ce->guid, old_ce->guid, 20);\n \t\tactive_cache[pos] = ce;\n \t\treturn 0;\n \t}\n@@ -376,6 +397,12 @@\n \tif (!ok_to_add)\n \t\treturn -1;\n \n+\tmemset(guid, 0, 20);\n+\tif (!memcmp(ce->guid, guid, 20)) {\n+\t\tnew_guid(ce->name, ce->namelen, guid);\n+\t\tmemcpy(ce->guid, guid, 20);\n+\t}\n+\n \t/* Make sure the array is big enough .. */\n \tif (active_nr == active_alloc) {\n \t\tactive_alloc = alloc_nr(active_alloc);\nread-tree.c:  eb548148aa6d212f05c2c622ffbe62a06cd072f9\n--- read-tree.c\n+++ read-tree.c\t2005-04-16 10:41:46.000000000 +1000\n@@ -5,7 +5,9 @@\n  */\n #include \"cache.h\"\n \n-static int read_one_entry(unsigned char *sha1, const char *base, int baselen, const char *pathname, unsigned mode)\n+static int read_one_entry(unsigned char *sha1, unsigned char *guid, \n+\t\t\t  const char *base, int baselen, \n+\t\t\t  const char *pathname, unsigned mode)\n {\n \tint len = strlen(pathname);\n \tunsigned int size = cache_entry_size(baselen + len);\n@@ -18,6 +20,7 @@\n \tmemcpy(ce->name, base, baselen);\n \tmemcpy(ce->name + baselen, pathname, len+1);\n \tmemcpy(ce->sha1, sha1, 20);\n+\tmemcpy(ce->guid, guid, 20);\n \treturn add_cache_entry(ce, 1);\n }\n \n@@ -35,14 +38,15 @@\n \twhile (size) {\n \t\tint len = strlen(buffer)+1;\n \t\tunsigned char *sha1 = buffer + len;\n+\t\tunsigned char *guid = buffer + len + 20;\n \t\tchar *path = strchr(buffer, ' ')+1;\n \t\tunsigned int mode;\n-\n-\t\tif (size < len + 20 || sscanf(buffer, \"%o\", &mode) != 1)\n+\t\t\n+\t\tif (size < len + 40 || sscanf(buffer, \"%o\", &mode) != 1)\n \t\t\treturn -1;\n \n-\t\tbuffer = sha1 + 20;\n-\t\tsize -= len + 20;\n+\t\tbuffer = sha1 + 40;\n+\t\tsize -= len + 40;\n \n \t\tif (S_ISDIR(mode)) {\n \t\t\tint retval;\n@@ -57,7 +61,7 @@\n \t\t\t\treturn -1;\n \t\t\tcontinue;\n \t\t}\n-\t\tif (read_one_entry(sha1, base, baselen, path, mode) < 0)\n+\t\tif (read_one_entry(sha1, guid, base, baselen, path, mode) < 0)\n \t\t\treturn -1;\n \t}\n \treturn 0;\nrev-tree.c:  395b0b3bfadb0537ae0c62744b25ead4b487f3f6\nshow-diff.c:  a531ca4078525d1c8dcf84aae0bfa89fed6e5d96\nshow-files.c:  a9fa6767a418f870a34b39379f417bf37b17ee18\ntree-id:  cb70e2c508a18107abe305633612ed702aa3ee4f\nupdate-cache.c:  62d0a6c41560d40863c44599355af10d9e089312\nwrite-tree.c:  1534477c91169ebddcf953e3f4d2872495477f6b\n--- write-tree.c\n+++ write-tree.c\t2005-04-15 13:46:05.000000000 +1000\n@@ -47,6 +47,7 @@\n \t\tconst char *pathname = ce->name, *filename, *dirname;\n \t\tint pathlen = ce->namelen, entrylen;\n \t\tunsigned char *sha1;\n+\t\tunsigned char *guid;\n \t\tunsigned int mode;\n \n \t\t/* Did we hit the end of the directory? Return how many we wrote */\n@@ -54,6 +55,7 @@\n \t\t\tbreak;\n \n \t\tsha1 = ce->sha1;\n+\t\tguid = ce->guid;\n \t\tmode = ce->st_mode;\n \n \t\t/* Do we have _further_ subdirectories? */\n@@ -86,6 +88,8 @@\n \t\tbuffer[offset++] = 0;\n \t\tmemcpy(buffer + offset, sha1, 20);\n \t\toffset += 20;\n+\t\tmemcpy(buffer + offset, guid, 20);\n+\t\toffset += 20;\n \t\tnr++;\n \t} while (nr < maxentries);\n \n\n\n/*\n * rename files in a git repository, keeping the guid.\n * \n * Copyright Simon Fowler <simon@dreamcraft.com.au>, 2005.\n */\n#include <unistd.h>\n#include <sys/stat.h>\n#include <errno.h>\n#include \"cache.h\"\n\nstatic int remove_lock = 0;\n\nstatic void remove_lock_file(void)\n{\n\tif (remove_lock)\n\t\tunlink(\".git/index.lock\");\n}\n\nint main(int argc, char *argv[])\n{\n\tstruct stat stats;\n\tstruct cache_entry *ce, *new;\n\tint newfd, entries, pos, pos2;\n\n\tif (argc != 3)\n\t\tusage(\"rename-file <old> <new>\");\n\tif (stat(argv[1], &stats)) {\n\t\tperror(\"rename-file: \");\n\t\texit(1);\n\t}\n\tif (!stat(argv[2], &stats))\n\t\tdie(\"rename-file: destination file already exists\");\n\t\n\tnewfd = open(\".git/index.lock\", O_RDWR | O_CREAT | O_EXCL, 0600);\n\tif (newfd < 0)\n\t\tdie(\"unable to create new cachefile\");\n\n\tatexit(remove_lock_file);\n\tremove_lock = 1;\n\n\tentries = read_cache();\n\tif (entries < 0)\n\t\tdie(\"cache corrupted\");\n\n\tpos = cache_name_pos(argv[1], strlen(argv[1]));\n\tpos2 = cache_name_pos(argv[2], strlen(argv[2]));\n\t\t\n\tif (pos < 0) \n\t\tdie(\"original file not in cache\");\n\tif (pos2 >= 0)\n\t\tdie(\"destination file already in cache\");\n\tce = active_cache[pos];\n\tnew = malloc(sizeof(struct cache_entry) + strlen(argv[2]) + 1);\n\tmemcpy(new, ce, sizeof(struct cache_entry));\n\tnew->namelen = strlen(argv[2]);\n\tmemcpy(new->name, argv[2], new->namelen);\n\t\n\tif (rename(argv[1], argv[2])) {\n\t\tperror(\"rename-file: \");\n\t\texit(1);\n\t}\n\n\tremove_file_from_cache(argv[1]);\n\tadd_cache_entry(new, 1);\n\n\tif (write_cache(newfd, active_cache, active_nr) ||\n\t    rename(\".git/index.lock\", \".git/index\"))\n\t\tdie(\"Unable to write new cachefile\");\n      \n\tremove_lock = 0;\n\treturn 0;\n}\n"},{"id":"271","messageId":"Pine.LNX.4.58.0504151913180.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"Pine.LNX.4.21.0504152029410.30848-100000@iabervon.org","subject":"Re: Re: Re: write-tree is pasky-0.4","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T02:18:51Z","receivedAt":"2005-04-16T02:18:51Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, Daniel Barkalow wrote:\n> \n> Is there some reason you don't commit before merging? All of the current\n> merge theory seems to want to merge two commits, using the information git\n> keeps about them.\n\nNote that the 3-way merge would _only_ merge the committed state. The \nthing is, 99% of all merges end up touching files that I never touch \nmyself (ie other architectures), so me being able to merge them even when \n_I_ am in the middle of something is a good thing.\n\nSo even when I have dirty state, the \"merge\" would only merge the clean\nstate. And then before the merge information is put back into my working\ndirectory, I'd do a \"check-files\" on the result, making sure that nothing\nthat got changed by the merge isn't up-to-date.\n\n> How much do you care about the situation where there is no best common\n> ancestor\n\nI care. Even if the best common parent is 3 months ago, I care. I'd much \nrather get a big explicit conflict than a \"clean merge\" that ends up being \ndebatable because people played games with per-file merging or something \nquestionable like that.\n\n> I think that the time spent on I/O will be overwhelmed by the time spent\n> issuing the command at that rate.\n\nThere is no time at all spent on IO.\n\nAll my email is local, and if this all ends up working out well, I can \ntrack the other peoples object trees in local subdirectories with some \ndaily rsyncs. And I have enough memory in my machines that there is \nbasically no disk IO - the only tree I normally touch is the kernel trees, \nthey all stay in cache.\n\n\t\tLinus\n"},{"id":"273","messageId":"Pine.LNX.4.21.0504152221070.30848-100000@iabervon.org","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504151913180.7211@ppc970.osdl.org","subject":"Re: Re: Re: write-tree is pasky-0.4","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-04-16T02:49:46Z","receivedAt":"2005-04-16T02:49:46Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 15 Apr 2005, Linus Torvalds wrote:\n\n> On Fri, 15 Apr 2005, Daniel Barkalow wrote:\n> > \n> > Is there some reason you don't commit before merging? All of the current\n> > merge theory seems to want to merge two commits, using the information git\n> > keeps about them.\n> \n> Note that the 3-way merge would _only_ merge the committed state. The \n> thing is, 99% of all merges end up touching files that I never touch \n> myself (ie other architectures), so me being able to merge them even when \n> _I_ am in the middle of something is a good thing.\n> \n> So even when I have dirty state, the \"merge\" would only merge the clean\n> state. And then before the merge information is put back into my working\n> directory, I'd do a \"check-files\" on the result, making sure that nothing\n> that got changed by the merge isn't up-to-date.\n\nSo you want to merge someone else's tree into your committed state, and\nthen merge the result with your working directory to get the working\ndirectory you continue with, provided that the second merge is trivial?\n\n> > How much do you care about the situation where there is no best common\n> > ancestor\n> \n> I care. Even if the best common parent is 3 months ago, I care. I'd much \n> rather get a big explicit conflict than a \"clean merge\" that ends up being \n> debatable because people played games with per-file merging or something \n> questionable like that.\n\nAre you thinking that the best common ancestor is the one that ties up\nabsolutely all of the chains of commits, or the closest one that the sides\nhave in common? I have the feeling that the former isn't going to be\nuseful, because there will be lines you're considering merging which go\nback to ancient kernels, where they keep merging in your changes, but\nthey still have a lineage back to 2.6.0 or something.\n\nFor the latter, there are sometimes multiple ancestors which fit this\ncriterion, and different ones of them are most helpful for different\nportions of the merge. I think this primarily happens when a branch you\nwant to merge has accepted multiple patches that you've also\naccepted (and the history identifies this fact); this may or may not be a\nsituation you want to allow on a regular basis.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"278","messageId":"Pine.LNX.4.58.0504152000570.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"Pine.LNX.4.21.0504152221070.30848-100000@iabervon.org","subject":"Re: Re: Re: write-tree is pasky-0.4","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T03:13:51Z","receivedAt":"2005-04-16T03:13:51Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, Daniel Barkalow wrote:\n> \n> So you want to merge someone else's tree into your committed state, and\n> then merge the result with your working directory to get the working\n> directory you continue with, provided that the second merge is trivial?\n\nNo, you don't even \"merge\" the working directory.\n\nThe low-level tools should entirely ignore the working directory. To a\nlow-level merge, the working directory doesn't even exist. It just gets\nthree commits (or trees) and merges two of them with the third as a\nparent, and does all of it in it's own temporary \"merge working\ndirectory\".\n\nSo on a technical level, the \"plumbing\" part really really doesn't care at \nall.\n\nHowever, from a _usability_ part, you expect after a merge that your \nworking directory has been updated to be the merged tree. And that's where \nthe \"if I have a working tree that is dirty, I want that part to fail\" \ncomes in. In other words, the final phase (after the \"tree-merge\" has \nactually successfully already finished) is to go back to the working \ndirectory, and check out the merged results.\n\nBut that checkout would be a variation on \"checkout-cache -a\" which first\nchecks that none of the files it is going to overwrite are dirty.\n\nDon't worry about this part. It's really totally separate from the true\nmerge itself. The \"real work\" has already been done by the time we notice\nthat \"oops, we can't actually show him the newly merged tree, because he\nhas got dirty data where we want to show it\".\n\n> > I care. Even if the best common parent is 3 months ago, I care. I'd much \n> > rather get a big explicit conflict than a \"clean merge\" that ends up being \n> > debatable because people played games with per-file merging or something \n> > questionable like that.\n> \n> Are you thinking that the best common ancestor is the one that ties up\n> absolutely all of the chains of commits, or the closest one that the sides\n> have in common?\n\nThe closest common one.\n\n> For the latter, there are sometimes multiple ancestors which fit this\n> criterion\n\nYes. Let's just pick one at random (or more likely, the latest one by \ndate - let's not actually be _random_ random) at first. \n\nThere are other heuristics we can try, ie if it turns out that it's common\nto have a couple of alternatives (but no more than some small number, say\nfive or so), we can literally just -try- to do a tree-only merge, and see\nhow many lines out common output you get from \"diff-tree\".\n\nBecause that \"how mnay files do we need to merge\" is the number you want\nto minimize, and doing a couple of extra \"diff-tree\" + \"join\"  operations\nshould be so fast that nobody will notice that we actually tried five\ndifferent merges to see which one looked the best.\n\nBut hey, especially if the merge fails with real clashes (ie there are\nchanges in common and running \"merge\" leaves conflicts), and there were\nother alternate parents to choose, there's nothing wrong with just\nprinting them out and saying \"you might try to specify one of these\nmanually\".\n\nI really don't think we should worry too much about this until we've \nactually used the system for a while and seen what it does. So just start \nwith \"nearest common parent with most recent date\". Which I think you \nalready implemented, no?\n\n\t\tLinus\n"},{"id":"281","messageId":"Pine.LNX.4.21.0504152318230.30848-100000@iabervon.org","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504152000570.7211@ppc970.osdl.org","subject":"Re: Re: Re: write-tree is pasky-0.4","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-04-16T03:56:42Z","receivedAt":"2005-04-16T03:56:42Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 15 Apr 2005, Linus Torvalds wrote:\n\n> On Fri, 15 Apr 2005, Daniel Barkalow wrote:\n> > \n> > So you want to merge someone else's tree into your committed state, and\n> > then merge the result with your working directory to get the working\n> > directory you continue with, provided that the second merge is trivial?\n> \n> No, you don't even \"merge\" the working directory.\n> \n> The low-level tools should entirely ignore the working directory. To a\n> low-level merge, the working directory doesn't even exist. It just gets\n> three commits (or trees) and merges two of them with the third as a\n> parent, and does all of it in it's own temporary \"merge working\n> directory\".\n\nIt seems like users won't expect there to be a new working directory for\nthe merge in which they are supposed to resolve te conflicts, but where\nthey don't see their uncommited changes. In any case, the low-level tools\nhave to care about *some* working directory, even if it isn't the parent\nof .git, and the parent of .git seems like where other similar things\nhappen. If we're being conservative about merging, we're likely to report\na lot of conflicts, at least until we work out better techniques than a\nsimple 3-way merge.\n\n> > For the latter, there are sometimes multiple ancestors which fit this\n> > criterion\n> \n> Yes. Let's just pick one at random (or more likely, the latest one by \n> date - let's not actually be _random_ random) at first. \n\nOkay; I've currently got the one where the number of generations it is\naway from the further head is the smallest, and of equal ones, an\narbitrary choice. If people are generally similar in the amount they\ndiverge before commiting, this should be the most similar ancestor.\n\n> There are other heuristics we can try, ie if it turns out that it's common\n> to have a couple of alternatives (but no more than some small number, say\n> five or so), we can literally just -try- to do a tree-only merge, and see\n> how many lines out common output you get from \"diff-tree\".\n> \n> Because that \"how mnay files do we need to merge\" is the number you want\n> to minimize, and doing a couple of extra \"diff-tree\" + \"join\"  operations\n> should be so fast that nobody will notice that we actually tried five\n> different merges to see which one looked the best.\n> \n> But hey, especially if the merge fails with real clashes (ie there are\n> changes in common and running \"merge\" leaves conflicts), and there were\n> other alternate parents to choose, there's nothing wrong with just\n> printing them out and saying \"you might try to specify one of these\n> manually\".\n\nI think we should be able to get good results out of doing the 5 merges\nand reporting a conflict only if there's a conflict in all of them; it\nshouldn't be possible for two to succeed but give different results (if it\ndid, clearly our current algorithm is unsafe, since it would give some\nundesired output if it happened to use the wrong ancestor).\n\nI'm thinking of not actually calling \"merge(1)\" for this at all; it just\ncalls diff3, and diff3 is only 1745 lines including option parsing. We can\nprobably arrange to look around for better ancestors in case of conflicts\nwe'd otherwise have to report, and get this all tidy and more efficient\nthan having diff3 re-read files. And if we only go to other ancestors in\ncase of conflicts, we're going to be a lot faster total than getting a\nreaction from the user, almost no matter what we do.\n\n> I really don't think we should worry too much about this until we've \n> actually used the system for a while and seen what it does. So just start \n> with \"nearest common parent with most recent date\". Which I think you \n> already implemented, no?\n\nI've got something like that (see above); did you want it in some form\nother than the patch I sent you?\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"282","messageId":"7v7jj3fjky.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504151755590.7211@ppc970.osdl.org","subject":"Re: [PATCH 3/2] merge-trees script for Linus git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T04:10:05Z","receivedAt":"2005-04-16T04:10:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> Just a heads-up - I'd really want to do the same thing to \"merge-tree.c\" \nLT> too, but since you said that you were working on extending that to do \nLT> recursion etc, I decided to hold off. So if you're working on it, maybe \nLT> you can add the \"-z\" flag there too? \n\nSent as a separate patch already.\n\nLT> I'm actually holding off merging the perl version exactly because you \nLT> seemed to be working on the C version. I don't mind perl per se, but if \nLT> there's a real solution coming down the line..\n\nI'd take the hint, but I would say the current Perl version\nwould be far more usable than the C version I would come up with\nby the end of this weekend because:\n\n - the Perl version creates a new temporary directory and leaves\n   a ready-to-use dircache there---the only thing needed from\n   that point for you is to fix it up any conflicts and do\n   update-cache on that dircache.  In that sense it is already\n   usable (Linus-usable, but probably not Pasky-usable due to\n   differences in phylosophy).\n\n - the enhancement I am planning on the C version does not do\n   the real work itself, as you have originally written (the\n   workings and the output from it are outlined in [*R1*]).\n   Somebody has to write the executor part that does read-tree\n   the base, update-cache --cacheinfo --add for the selects,\n   runs 3-way merge on conflicting files and runs update-cache\n   for the merges, update-cache --remove for the deletes, before\n   it matches the usability of the Perl version.  I do not\n   expect to have enough time this weekend to finish this.\n\nI know of one case in Perl version I need to see if it does the\nright thing but other than that it would be far better than the\nC version I'm toying with.\n\nJust to let you know, here is the plan I have for my part.\n\n 1. I am currently writing some test cases.  The plan is first\n    to make sure the Perl version works OK with the test cases\n    to flush initial problems out.\n\n 2. After that I'll see if a dumb but recursive C version I\n    already have spits out the right instructions.  This step is\n    to make sure that the test cases are sane, and by making\n    that sure, we will be able to say that we have something\n    usable in extremely short run (i.e. the Perl version) after\n    this step.\n\n 3. After that is done, I'll add the fourth argument to\n    merge-tree.c to specify the base so that it can cut down 90%\n    of trivial selects.  Only after this happens the executioner\n    script would be useful performance wise.\n\n[References]\n\n*R1* <7vr7hbhky9.fsf@assigned-by-dhcp.cox.net>\n\n"},{"id":"284","messageId":"Pine.LNX.4.58.0504152152580.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"7v7jj3fjky.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 3/2] merge-trees script for Linus git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T05:02:36Z","receivedAt":"2005-04-16T05:02:36Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, Junio C Hamano wrote:\n> \n> I'd take the hint, but I would say the current Perl version\n> would be far more usable than the C version I would come up with\n> by the end of this weekend because:\n\nActually, it turns out that I have a cunning plan.\n\nI'm full of cunning plans, in fact. It turns out that I can do merges even\nmore simply, if I just allow the notion of \"state\" into an index entry,\nand allow multiple index entries with the same name as long as they differ\nin \"state\".\n\nAnd that means that I can do all the merging in the regular index tree, \nusing very simple rules.\n\nLet's see how that works out. I'm writing the code now.\n\n\t\tLinus\n"},{"id":"286","messageId":"Pine.LNX.4.58.0504152256520.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504152152580.7211@ppc970.osdl.org","subject":"Re: [PATCH 3/2] merge-trees script for Linus git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T06:26:17Z","receivedAt":"2005-04-16T06:26:17Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 15 Apr 2005, Linus Torvalds wrote:\n> \n> Actually, it turns out that I have a cunning plan.\n\nDamn, my cunning plan is some good stuff. \n\nOr maybe it is _so_ cunning that I just confuse even myself. But it looks \nlike it is actually working, and that it allows pretty much instantaenous \nmerges.\n\nThe plan goes like this:\n\n - each \"index\" entry has two bits worth of \"stage\" state. stage 0 is the \n   normal one, and is the only one you'd see in any kind of normal use.\n\n - however, when you do \"read-tree\" with multiple trees, the \"stage\" \n   starts out at 0, but increments for each tree you read. And in \n   particular, the old \"-m\" flag (which used to be \"merge with old state\")  \n   has a new meaning: it now means \"start at stage 1\" instead.\n\n - this means that you can do\n\n\tread-tree -m <tree1> <tree2> <tree3>\n\n   and you will end up with an index with all of the <tree1> entries in \n   \"stage1\", all of the <tree2> entries in \"stage2\" and all of the <tree3>\n   entries in \"stage3\".\n\n - furthermore, \"read-tree\" has this special-case logic that says: if you \n   see a file that matches in all respects in all three states, it \n   \"collapses\" back to \"stage0\".\n\n - write-tree refuses to write a nonsensical tree, so write-tree will \n   complain about unmerged entries if it sees a single entry that is not\n   stage 0\".\n\nOk, this all sounds like a collection of totally nonsensical rules, but\nit's actually exactly what you want in order to do a fast merge. The \ndiffernt stages represent the \"result tree\" (stage 0, aka \"merged\"), the \noriginal tree (stage 1, aka \"orig\"), and the two trees you are trying to \nmerge (stage 2 and 3 respectively).\n\nIn fact, the way \"read-tree\" works, it's entirely agnostic about how you\nassign the stages, and you could really assign them any which way, and the\nabove is just a suggested way to do it (except since \"write-tree\" refuses\nto write anything but stage0 entries, it makes sense to always consider\nstage 0 to be the \"full merge\" state).\n\nSo what happens? Try it out. Select the original tree, and two trees to \nmerge, and look how it works:\n\n - if a file exists in identical format in all three trees, it will \n   automatically collapse to \"merged\" state by the new read-tree.\n\n - a file that has _any_ difference what-so-ever in the three trees will \n   stay as separate entries in the index. It's up to \"script policy\" to \n   determine how to remove the non-0 stages, and insert a merged version. \n   But since the index is always sorted, they're easy to find: they'll be\n   clustered together.\n\n - the index file saves and restores with all this information, so you can \n   merge things incrementally, but as long as it has entries in stages\n   1/2/3 (ie \"unmerged entries\") you can't write the result.\n\nSo now the merge algorithm ends up being really simple:\n\n - you walk the index in order, and ignore all entries of stage 0, since \n   they've already been done.\n - if you find a \"stage1\", but no matching \"stage2\" or \"stage3\", you know \n   it's been removed from both trees (it only existed in the original \n   tree), and you remove that entry.\n - if you find a matching \"stage2\" and \"stage3\" tree, you remove one of \n   them, and turn the other into a \"stage0\" entry. Remove any matching\n   \"stage1\" entry if it exists too.\n  .. all the normal trivial rules ..\n\nNOTE NOTE NOTE! I could make \"read-tree\" do some of these nontrivial \nmerges, but I ended up deciding that only the \"matches in all three \nstates\" thing collapses by default. Why? Because even though there are \nother trivial cases (\"matches in both merge trees but not in the original \none\"), those cases might actually be interesting for the merge logic to \nknow about, so I thought I'd leave all that information around. I expect \nit to be fairly rare anyway, so writing out a few extra index entries to \ndisk so that others can decide to annotate the merge a bit more sounded \nlike a fair deal.\n\nI should make \"ls-files\" have a \"-l\" format, which shows the index and the \nmode for each file too. Right now it's very hard to see what the contents \nof the index is. But all my tests seem to say that not only does this \nwork, it's pretty efficient too. And it's dead _simple_, thanks to having \nall the merge information in just one place, the same index we always use \nanyway.\n\nBtw, it also means that you don't even have to have a separate \nsubdirectory for this. All the information literally is in the index file, \nwhich is a temporary thing anyway. We don't need to worry about what is in \nthe working directory, since we'll never show it, and we'll never need to \nuse it.\n\nDamn, I'm good.\n\n(On the other hand, it is Friday evening at 11PM, and I'm sitting in front\nof the computer. I'm a sad case. I will now go take a beer, and relax. I\nthink this is another of my \"Really Good Ideas\" (tm), and is worth the\nbeer.  This \"feels\" right).\n\n\t\tLinus\n"},{"id":"287","messageId":"20050415235941.73f8a007.pj@engr.sgi.com","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504152000570.7211@ppc970.osdl.org","subject":"Re: write-tree is pasky-0.4","fromName":"Paul Jackson","fromEmail":"pj@engr.sgi.com","sentAt":"2005-04-16T06:59:41Z","receivedAt":"2005-04-16T06:59:41Z","isPatch":false,"sender":{"key":"pj@engr.sgi.com","avatar":null},"body":"One trick I've used to separate good automatic merges from ones that\nneed human interaction is to run both the 'patch' and 'merge' commands,\nwhich use different approaches to determining the result.\n\nIf they agree, take it.  To apply the changes between file1 and file2\nto filez:\n\n\tdiff -au file1 file2 | patch -f filez\n\tmerge -q filez file1 file2\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"288","messageId":"7vis2ncf8j.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504152152580.7211@ppc970.osdl.org","subject":"Re: [PATCH 3/2] merge-trees script for Linus git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T08:12:12Z","receivedAt":"2005-04-16T08:12:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> Damn, my cunning plan is some good stuff. \n\nI really like this a lot.  It is *so* *simple*, clear, flexible\nand an example of elegance.  This is one of the things I would\nhappily say \"Sheeeeeeeeeeeeeesh!  Why didn't *I* think of *THAT*\nfirst!!!\" to.\n\nLT> NOTE NOTE NOTE! I could make \"read-tree\" do some of these nontrivial \nLT> merges, but I ended up deciding that only the \"matches in all three \nLT> states\" thing collapses by default.\n\n * Understood and agreed.\n\nLT> Damn, I'm good.\n\n * Agreed ;-). Wholeheartedly.\n\nSo what's next?  Certainly I'd immediately drop (and I would\nimagine you would as well) both C or Perl version of\nmerge-tree(s).\n\nThe userland merge policies need ways to extract the stage\ninformation and manipulate them.  Am I correct to say that you\nmean by \"ls-files -l\" the extracting part?\n\nLT> I should make \"ls-files\" have a \"-l\" format, which shows the\nLT> index and the mode for each file too.\n\nYou probably meant \"ls-tree\".  You used the word \"mode\" but it\nalready shows the mode so I take it to mean \"stage\".  Perhaps\nsomething like this?\n\n$ ls-tree -l -r 49c200191ba2e3cd61978672a59c90e392f54b8b\n100644\tblob\tfe2a4177a760fd110e78788734f167bd633be8de\tCOPYING\n100644\tblob\tb39b4ea37586693dd707d1d0750a9b580350ec50:1\tman/frotz.6\n100644\tblob\tb39b4ea37586693dd707d1d0750a9b580350ec50:2\tman/frotz.6\n100664\tblob\teeed997e557fb079f38961354473113ca0d0b115:3\tman/frotz.6\n ...\n\nThe above example shows that COPYING has merged successfully,\nand O and A have the same contents and B has something different\nat man/frotz.6.\n\nAssuming that you would be working on that, I'd like to take the\ndircache manipulation part.  Let's think about the minimally\nnecessary set of operations:\n\n * The merge policy decides to take one of the existing stage.\n\n   In this case we need a way to register a known mode/sha1 at a\n   path.  We already have this as \"update-cache --cacheinfo\".\n   We just need to make sure that when \"update-cache\" puts\n   things at stage 0 it clears other stages as well.\n\n * The merge policy comes up with a desired blob somewhere on\n   the filesystem (perhaps by running an external merge\n   program).  It wants to register it as the result of the\n   merge.\n\n   We could do this today by first storing the \"desired blob\"\n   in a temporary file somewhere in the path the dircache\n   controls, \"update-cache --add\" the temporary file, ls-tree to\n   find its mode/sha1, \"update-cache --remove\" the temporary\n   file and finally \"update-cache --cacheinfo\" the mode/sha1.\n   This is workable but clumsy.  How about:\n\n   $ update-cache --graft [--add] desired-blob path\n\n   to say \"I want to register mode/sha1 from desired-blob, which\n   may not be of verify_path() satisfying name, at path in the\n   dircache\"?\n\n * The merge policy decides to delete the path.\n\n   We could do this today by first stashing away the file at the\n   path if it exists, \"update-cache --remove\" it, and restore\n   if necessary.  This is again workable but clumsy.  How about:\n\n   $ update-cache --force-remove path\n\n   to mean \"I want to remove the path from dircache even though\n   it may exist in my working tree\"?\n\nSo it all boils down to update-cache.  The new things to be\nintroduced are:\n\n * An explicit update-cache always removes stage 1/2/3 entries\n   associated with the named path.\n\n * update-cache --graft\n\n * update-cache --force-remove\n\nAm I on the right track?\n\nYou might want to go even lower level by letting them say\nsomething like:\n\n * update-cache --register-stage mode sha1 stage path\n\n   Registers the mode/sha1 at stage for path.  Does not look at\n   the working tree.  stage is [0-3]\n \n * update-cache --delete-stage stage-list path\n\n   Removes the entry at named stages for path.  Does not look at\n   the working tree.  stage-list is either [0-3](,[0-3])+ or\n   bitmask (i.e. (1 << stage-number) ORed together).  The former\n   would probably be easier to work with by scripts\n\n * write-blob path\n\n   Hashes and registers the file at path (regardless of what\n   verify_path() says) and writes the resulting blob's mode/sha1\n   to the standard output.\n\nIf you take this lower-level approach, an explicit update-cache\nwould not clear stage1/2/3.\n\nMy preference is the former, not so low-level, interface.\nGuidance?\n\n"},{"id":"289","messageId":"7vacnzcbrn.fsf_-_@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"7vis2ncf8j.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH] Byteorder fix for read-tree, new -m semantics version.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T09:27:08Z","receivedAt":"2005-04-16T09:27:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"The ce_namelen field has been renamed to ce_flags and split into\nthe top 2-bit unused, next 2-bit stage number and the lowest\n12-bit name-length, stored in the network byte order.  A new\nmacro create_ce_flags() is defined to synthesize this value from\nlength and stage, but it forgets to turn the value into the\nnetwork byte order.  Here is a fix.\n\nThe patch is against 9c03bd47892d11d0bb28c442184786db3c189978.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n cache.h |    2 +-\n 1 files changed, 1 insertion(+), 1 deletion(-)\n\n--- cache.h\n+++ cache.h\t2005-04-16 02:22:05.000000000 -0700\n@@ -66,7 +66,7 @@\n #define CE_NAMEMASK  (0x0fff)\n #define CE_STAGEMASK (0x3000)\n \n-#define create_ce_flags(len, stage) ((len) | ((stage) << 12))\n+#define create_ce_flags(len, stage) htons((len) | ((stage) << 12))\n \n const char *sha1_file_directory;\n struct cache_entry **active_cache;\n\n\n"},{"id":"291","messageId":"7vfyxrau23.fsf_-_@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"7vis2ncf8j.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH 1/2] Add --stage to show-files for new stage dircache.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T10:35:00Z","receivedAt":"2005-04-16T10:35:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"JNH\" == Junio C Hamano <junkio@cox.net> writes:\n>>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> I should make \"ls-files\" have a \"-l\" format, which shows the\nLT> index and the mode for each file too.\n\nJNH> You probably meant \"ls-tree\".  You used the word \"mode\" but it\nJNH> already shows the mode so I take it to mean \"stage\".\n\nI was *wrong*.  Of course you meant \"show-files\".\n\nInstead of sending you an apology, I am sending you the one I\nwrote myself.  Please find it in the next message ;-).\n\nHere is its sample output.  It shows file-mode, SHA1, stage and\npathname.  I am attaching this one because this is a\nverification that your read-tree -m passed the test.\n\n$ ../show-files --stage\n100664 578cc900ed980b72acfbdd1eea63e688a893c458 2 AA\n100664 f355077379fce072c210628691da232b59b6f25c 3 AA\n100664 d698ebc45d0edfe6e5b95aebb5983cb5c760960b 2 AN\n100664 0fa6a8e41814531679e1c76e968a9066fceb689d 1 DD\n100664 aff448a9467a4d83b164ef969cfe92ff18eb96be 1 DM\n100664 4bfe111723f11cb4a4deec7c837e12601030285f 3 DM\n100664 9b0f86e5cded99b9de3bd9d234747ec2d1a4cddd 1 DN\n100664 9b0f86e5cded99b9de3bd9d234747ec2d1a4cddd 3 DN\n100664 a6772f2a2c15bac796d8c7bb55885891956534cf 1 MD\n100664 dc2088ce13f659f2bd554b2c1b343f4966143b9b 2 MD\n100664 e4310204563a9059828644464779874c3a406fee 1 MM\n100664 fe5ddcd7618d26384cf98c6fcd15780c7125e6d6 2 MM\n100664 53a9d14868dbe346a9f0cf01fcda742545b55987 3 MM\n100664 f48f37ea0205a7e5591777b4d3ae0d153d3ef131 1 MN\n100664 d7600381b69b92f61bad50c5f8408e831b622ef0 2 MN\n100664 f48f37ea0205a7e5591777b4d3ae0d153d3ef131 3 MN\n100664 67fb1517ea8d59949a8e4f5f07f0422b212f64dc 3 NA\n100664 0e5842253af8881b2c9f579029d7b50a8e03d7f6 1 ND\n100664 0e5842253af8881b2c9f579029d7b50a8e03d7f6 2 ND\n100664 0d45c04c9d05fa9c21edf95fc2c1a43519a8c440 1 NM\n100664 0d45c04c9d05fa9c21edf95fc2c1a43519a8c440 2 NM\n100664 849bfa41d15951f5e97cb93e22cbcc2924ce4517 3 NM\n100664 83d94b8fd056921f22ad2ca0122dd7f64974be7c 0 NN\n\nThis is taken from the dircache after I ran\n\n    $ read-tree -m O A B\n\nusing the merge testcase I prepared earlier.  Very trivial,\nsingle ancestor O, with two branches A & B merge case.  This\ncovers all possible patterns, except file vs directory\nconflicts.  The filenames are all two letters, first letter\nbeing what the first branch does to that file while the second\none encodes what the second branch does to it.  The actions are:\n\n - A means \"Added in this branch --- did not exist in the ancestor.\"\n\n - N means \"No change in this branch.\"\n\n - D means \"Deleted in this branch.\"\n\n - M means \"Modified in this branch.\"\n\nSo, for example, the first branch modified file MN while the\nsecond one did not touch it.  Of course it existed in the\nancestor.  You can see that read-tree did the right thing\nbecause SHA1 for stage 1 and stage 3 match, and stage 2 is\ndifferent.\n\n    100664 f48f37ea0205a7e5591777b4d3ae0d153d3ef131 1 MN\n    100664 d7600381b69b92f61bad50c5f8408e831b622ef0 2 MN\n    100664 f48f37ea0205a7e5591777b4d3ae0d153d3ef131 3 MN\n\nI verified all of the above result and it shows your algorithm\nis doing exactly what is expected.\n\nFor those of you who are interested, this is the recipe to\nreproduce this merge testcase.  NOTE! NOTE! NOTE!  Do not run\nthis in your working tree, because it trashes .git in its\nworking directory.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n--- /dev/null\n+++ generate-merge-test.sh\n@@ -0,0 +1,163 @@\n+#!/bin/sh\n+\n+: Skip execution up to <<\\End_of_Commentary\n+\n+This directory is to hold a test case for merges.\n+\n+There is one ancestor (called O for Original) and two branches A\n+and B derived from it.  We want to do 3-way merge between A and\n+B, using O as the common ancestor.\n+\n+    merge A O B\n+    diff3 A O B\n+\n+Decisions are made by comparing contents of O, A and B pathname\n+by pathname.  The result is determined by the following guiding\n+principle:\n+\n+ - If only A does something to it and B does not touch it, take\n+   whatever A does.\n+\n+ - If only B does something to it and A does not touch it, take\n+   whatever B does.\n+\n+ - If both A and B does something but in the same way, take\n+   whatever they do.\n+\n+ - If A and B does something but different things, we need a\n+   3-way merge:\n+\n+   - We cannot do anything about the following cases:\n+\n+     * O does not have it.  A and B both must be adding to the\n+       same path independently.\n+\n+     * A deletes it.  B must be modifying.\n+\n+   - Otherwise, A and B are modifying.  Run 3-way merge.\n+\n+\n+First, the case matrix.\n+\n+ - Vertical axis is for A's actions.\n+ - Horizontal axis is for B's actions.\n+\n+.----------------------------------------------------------------.\n+| A        B | No Action  |   Delete   |   Modify   |    Add     |\n+|------------+------------+------------+------------+------------|\n+| No Action  |            |            |            |            |\n+|            | select O   | delete     | select B   | select B   |\n+|            |            |            |            |            |\n+|------------+------------+------------+------------+------------|\n+| Delete     |            |            | ********** |    can     |\n+|            | delete     | delete     | merge      |    not     |\n+|            |            |            |            |  happen    |\n+|------------+------------+------------+------------+------------|\n+| Modify     |            | ********** | ?????????? |    can     |\n+|            | select A   | merge      | select A=B |    not     |\n+|            |            |            | merge      |  happen    |\n+|------------+------------+------------+------------+------------|\n+| Add        |            |    can     |    can     | ?????????? |\n+|            | select A   |    not     |    not     | select A=B |\n+|            |            |  happen    |  happen    | merge      |\n+.----------------------------------------------------------------.\n+\n+End_of_Commentary\n+\n+rm -fr [NDMA][NDMA] S .git Trivial\n+init-db\n+\n+# Original tree.\n+mkdir S\n+for a in N D M\n+do\n+    for b in N D M\n+    do\n+        p=$a$b\n+\techo This is $p from the original tree. >$p\n+\techo This is S/$p from the original tree. >S/$p\n+\tupdate-cache --add $p || exit\n+\tupdate-cache --add S/$p || exit\n+    done\n+done\n+cat >Trivial <<\\EOF\n+This is a trivial merge sample text.\n+Branch A is expected to upcase this word.\n+There are some filler words to foil diff contexts here,\n+like this one,\n+and this one,\n+and this one is yet another one of them.\n+At the very end, here comes another line, that is\n+the word, expected to be upcased by Branch B.\n+This concludes the trivial merge sample file.\n+EOF\n+update-cache --add Trivial || exit\n+tree_O=$(write-tree)\n+commit_O=$(echo 'Original tree for the merge test.' | commit-tree $tree_O)\n+\n+# Branch A and B makes the changes according to the above matrix.\n+# Branch A\n+to_remove=$(echo D? S/D?)\n+rm -f $to_remove\n+update-cache --remove $to_remove || exit\n+\n+for p in M? S/M?\n+do\n+    echo This is modified $p in the branch A. >$p\n+    update-cache $p || exit\n+done\n+\n+for p in AN AA\n+do\n+    echo This is added $p in the branch A. >$p\n+    update-cache --add $p || exit\n+done\n+mv Trivial ,,Trivial\n+sed -e '/Branch A/s/word/WORD/g' <,,Trivial >Trivial\n+rm -f ,,Trivial\n+update-cache Trivial || exit\n+\n+tree_A=$(write-tree)\n+commit_A=$(echo 'Branch A for the merge test.' |\n+           commit-tree $tree_A -p $commit_O)\n+\t   \n+\n+# Branch B\n+# Start from O\n+rm -rf [NDMA][NDMA] S Trivial\n+mkdir S\n+../read-tree $tree_O\n+checkout-cache -a\n+\n+to_remove=$(echo ?D S/?D)\n+rm -f $to_remove\n+update-cache --remove $to_remove || exit\n+\n+for p in ?M S/?M\n+do\n+    echo This is modified $p in the branch B. >$p\n+    update-cache $p || exit\n+done\n+\n+for p in NA AA\n+do\n+    echo This is added $p in the branch B. >$p\n+    update-cache --add $p || exit\n+done\n+mv Trivial ,,Trivial\n+sed -e '/Branch B/s/word/WORD/g' <,,Trivial >Trivial\n+rm -f ,,Trivial\n+update-cache Trivial || exit\n+\n+tree_B=$(write-tree)\n+commit_B=$(echo 'Branch B for the merge test.' |\n+           commit-tree $tree_B -p $commit_O)\n+\n+for commit in $commit_O $commit_A $commit_B\n+do\n+    echo ================\n+    echo commit $commit\n+    cat-file commit $commit\n+done\n+echo ================\n+\n\n\n\n"},{"id":"292","messageId":"7vbr8fatq4.fsf_-_@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"7vfyxrau23.fsf_-_@assigned-by-dhcp.cox.net","subject":"[PATCH 2/2] Add --stage to show-files for new stage dircache.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T10:42:11Z","receivedAt":"2005-04-16T10:42:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"This adds --stage option to show-files command.  It shows\nfile-mode, SHA1, stage and pathname.  Record separator follows\nthe usual convention of -z option as before.\n\nThe patch is on top of the byte order fix for create_ce_flags in my\nprevious message.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n cache.h      |   12 +++++++-----\n show-files.c |   22 ++++++++++++++++++----\n 2 files changed, 25 insertions(+), 9 deletions(-)\n\n--- cache.h\t2005-04-16 03:02:36.000000000 -0700\n+++ cache.h=show-files-stage-flags\t2005-04-16 02:48:47.000000000 -0700\n@@ -65,8 +65,14 @@\n \n #define CE_NAMEMASK  (0x0fff)\n #define CE_STAGEMASK (0x3000)\n+#define CE_STAGESHIFT 12\n \n-#define create_ce_flags(len, stage) htons((len) | ((stage) << 12))\n+#define create_ce_flags(len, stage) htons((len) | ((stage) << CE_STAGESHIFT))\n+#define ce_namelen(ce) (CE_NAMEMASK & ntohs((ce)->ce_flags))\n+#define ce_size(ce) cache_entry_size(ce_namelen(ce))\n+#define ce_stage(ce) ((CE_STAGEMASK & ntohs((ce)->ce_flags)) >> CE_STAGESHIFT)\n+\n+#define cache_entry_size(len) ((offsetof(struct cache_entry,name) + (len) + 8) & ~7)\n \n const char *sha1_file_directory;\n struct cache_entry **active_cache;\n@@ -75,10 +81,6 @@\n #define DB_ENVIRONMENT \"SHA1_FILE_DIRECTORY\"\n #define DEFAULT_DB_ENVIRONMENT \".git/objects\"\n \n-#define cache_entry_size(len) ((offsetof(struct cache_entry,name) + (len) + 8) & ~7)\n-#define ce_namelen(ce) (CE_NAMEMASK & ntohs((ce)->ce_flags))\n-#define ce_size(ce) cache_entry_size(ce_namelen(ce))\n-\n #define alloc_nr(x) (((x)+16)*3/2)\n \n /* Initialize and use the cache information */\n\n\n\n--- show-files.c\n+++ show-files.c\t2005-04-16 02:58:32.000000000 -0700\n@@ -14,6 +14,7 @@\n static int show_cached = 0;\n static int show_others = 0;\n static int show_ignored = 0;\n+static int show_stage = 0;\n static int line_terminator = '\\n';\n \n static const char **dir;\n@@ -108,10 +109,19 @@\n \t\tfor (i = 0; i < nr_dir; i++)\n \t\t\tprintf(\"%s%c\", dir[i], line_terminator);\n \t}\n-\tif (show_cached) {\n+\tif (show_cached | show_stage) {\n \t\tfor (i = 0; i < active_nr; i++) {\n \t\t\tstruct cache_entry *ce = active_cache[i];\n-\t\t\tprintf(\"%s%c\", ce->name, line_terminator);\n+\t\t\tif (!show_stage)\n+\t\t\t\tprintf(\"%s%c\", ce->name, line_terminator);\n+\t\t\telse\n+\t\t\t\tprintf(/* \"%06o %s %d %10d %s%c\", */\n+\t\t\t\t       \"%06o %s %d %s%c\",\n+\t\t\t\t       ntohl(ce->ce_mode),\n+\t\t\t\t       sha1_to_hex(ce->sha1),\n+\t\t\t\t       ce_stage(ce),\n+\t\t\t\t       /* ntohl(ce->ce_size), */\n+\t\t\t\t       ce->name, line_terminator); \n \t\t}\n \t}\n \tif (show_deleted) {\n@@ -156,12 +166,16 @@\n \t\t\tshow_ignored = 1;\n \t\t\tcontinue;\n \t\t}\n+\t\tif (!strcmp(arg, \"--stage\")) {\n+\t\t\tshow_stage = 1;\n+\t\t\tcontinue;\n+\t\t}\n \n-\t\tusage(\"show-files (--[cached|deleted|others|ignored])*\");\n+\t\tusage(\"show-files [-z] (--[cached|deleted|others|ignored|stage])*\");\n \t}\n \n \t/* With no flags, we default to showing the cached files */\n-\tif (!(show_cached | show_deleted | show_others | show_ignored))\n+\tif (!(show_stage | show_deleted | show_others | show_ignored))\n \t\tshow_cached = 1;\n \n \tread_cache();\n\n"},{"id":"294","messageId":"Pine.LNX.4.62.0504160518310.21837@qynat.qvtvafvgr.pbz","threadId":"9","inReplyTo":"20050416014442.GW4488@himi.org","subject":"Re: Merge with git-pasky II.","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2005-04-16T12:19:24Z","receivedAt":"2005-04-16T12:19:24Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"\nOn Fri, Apr 15, 2005 at 08:32:46AM -0700, Linus Torvalds wrote:\n> In other words, I'm right. I'm always right, but sometimes I'm more \nright\n> than other times. And dammit, when I say \"files don't matter\", I'm \nreally\n> really Right(tm).\n>\nYou're right, of course (All Hail Linus!), if you can make it work\nefficiently enough.\n\nJust to put something else on the table, here's how I'd go about\ntracking renames and the like, in another world where Linus /does/\nmake the odd mistake - it's basically a unique id for files in the\nrepository, added when the file is first recognised and updated when\nupdate-cache adds a new version to the cache. Renames copy the id\nacross to the new name, and add it into the cache.\n\nThis gives you an O(n) way to tell what file was what across\nrenames, and it might even be useful in Linus' world, or if someone\nwanted to build a traditional SCM on top of a git-a-like.\n\nAttached is a patch, and a rename-file.c to use it.\n\nSimon\n\ngiven that you have multiple machines creating files, how do you deal with \nthe idea of the same 'unique id' being assigned to different files by \ndifferent machines?\n\nDavid Lang\n\n\n\n-- \nThere are two ways of constructing a software design. One way is to make it so simple that there are obviously no deficiencies. And the other way is to make it so complicated that there are no obvious deficiencies.\n  -- C.A.R. Hoare\n"},{"id":"305","messageId":"7vll7i95u1.fsf_-_@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"7vis2ncf8j.fsf@assigned-by-dhcp.cox.net","subject":"Issues with higher-order stages in dircache","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T14:03:34Z","receivedAt":"2005-04-16T14:03:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"JCH\" == Junio C Hamano <junkio@cox.net> writes:\n\nJCH> So what's next?\n\nHere is my current thinking on the impact your higher-order\nstage dircache entries would have to the rest of the system and\nhow to deal with them.\n\n * read-tree\n\n   - When merging two trees, i.e. \"read-tree -m A B\", shouldn't\n     we collapse identical stage-1/2 into stage-0?\n\n * update-cache\n\n   - An explicit \"update-cache [--add] [--remove] path\" should\n     be taken as a signal from the user (or Cogito) to tell the\n     dircache layer \"the merge is done and here is the result\".\n     So just delete higher-order stages for the path and record\n     the specified path at stage 0 (or remove it altogether).\n\n   - \"update-cache --refresh\" should just ignore a path that has\n     not been merged,  Maybe say \"needs merge\", just like \"needs\n     update\" [*1*].\n\n   - \"update-cache --cacheinfo\" should get an extra \"stage\"\n     argument.  Unmerged state is typically produced by running\n     \"read-tree -m\", but the user or Cogito can do it by hand\n     with this if he wanted to.\n\n   - I do not think we need a separate \"remove the entry for\n     this path at this stage\" thing.  That is only necessary if\n     the user or Cogito is doing things by hand (as opposed to\n     \"read-tree -m\"), which should be a very rare case.  He can\n     always do \"update-cache --remove\" followed by \"update-cache\n     --cacheinfo\" to obtain the desired result if he really\n     wanted to.  For that, \"update-cache --force-remove\" may\n     come in handy.\n\n * show-diff\n\n   - What should we do about unmerged paths?  Showing diffs\n     between the combinations (1->2), (1->3), and (2->3) that\n     exist may not be a bad idea.  It would not be confusing\n     because by definition dircache with higher-order stages is\n     a merge temporary directory and the user should not have a\n     working file there to begin with.\n\n     I think the current implementation does a very bad thing:\n     repeating the same diff as many times as it has\n     higher-order stages for the same path.\n\n * checkout-cache\n\n   - When checkout-cache is run with explicit paths that are\n     unmerged, what should we do?  What does that mean in the\n     first place?  One use scenario I can think of is that the\n     user or Cogito wants the contents at all three stages, in\n     order to run a merge tool on them.  From this point of\n     view, checking out all the available stages for the path\n     makes sense.\n\n     My \"cunning plan\" is to drop \".1-$file\", \".2-$file\", and\n     \".3-$file\" in the working directory.  How does that sound?\n\n   - When checkout-cache -a is run, presumably the user wants to\n     check out everything to verify (e.g. build-test) the\n     result.  In this case, we should skip unmerged paths, give\n     a warning, and check out only the merged ones.\n\n\n[Footnotes]\n\n*1* Unrelated note.  Who is the intended consumer of this \"needs\n    update\" message?  Should we make it machine readable with\n    '-z' flag as well?  Otherwise, shouldn't it go to stderr?\n    Currently it goes to stdout.\n\n"},{"id":"315","messageId":"Pine.LNX.4.58.0504160820320.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"7vis2ncf8j.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 3/2] merge-trees script for Linus git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T15:28:50Z","receivedAt":"2005-04-16T15:28:50Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Apr 2005, Junio C Hamano wrote:\n> \n> LT> NOTE NOTE NOTE! I could make \"read-tree\" do some of these nontrivial \n> LT> merges, but I ended up deciding that only the \"matches in all three \n> LT> states\" thing collapses by default.\n> \n>  * Understood and agreed.\n\nHaving slept on it, I think I'll merge all the trivial cases that don't \ninvolve a file going away or being added. Ie if the file is in all three \ntrees, but it's the same in two of them, we know what to do.\n\nThat way we'll leave thigns where the tree itself changed (files added or \nremoved at any point) and/or cases where you actually need a 3-way merge.\n\n> The userland merge policies need ways to extract the stage\n> information and manipulate them.  Am I correct to say that you\n> mean by \"ls-files -l\" the extracting part?\n\nNo, I meant \"show-files\", since we need to show the index, not a tree (no \nvalid tree can ever have the \"modes\" information, since (a) it doesn't \nhave the space for it anyway and (b) we refuse to write out a dirty index \nfile.\n\n\n\n> \n> LT> I should make \"ls-files\" have a \"-l\" format, which shows the\n> LT> index and the mode for each file too.\n> \n> You probably meant \"ls-tree\".  You used the word \"mode\" but it\n> already shows the mode so I take it to mean \"stage\".  Perhaps\n> something like this?\n> \n> $ ls-tree -l -r 49c200191ba2e3cd61978672a59c90e392f54b8b\n> 100644\tblob\tfe2a4177a760fd110e78788734f167bd633be8de\tCOPYING\n> 100644\tblob\tb39b4ea37586693dd707d1d0750a9b580350ec50:1\tman/frotz.6\n> 100644\tblob\tb39b4ea37586693dd707d1d0750a9b580350ec50:2\tman/frotz.6\n> 100664\tblob\teeed997e557fb079f38961354473113ca0d0b115:3\tman/frotz.6\n\nApart from the fact that it would be\n\n\tshow-files -l\n\nsince there are no tree objects that can have anything but fully merged\nstate, yes.\n\n> Assuming that you would be working on that, I'd like to take the\n> dircache manipulation part.  Let's think about the minimally\n> necessary set of operations:\n> \n>  * The merge policy decides to take one of the existing stage.\n> \n>    In this case we need a way to register a known mode/sha1 at a\n>    path.  We already have this as \"update-cache --cacheinfo\".\n>    We just need to make sure that when \"update-cache\" puts\n>    things at stage 0 it clears other stages as well.\n> \n>  * The merge policy comes up with a desired blob somewhere on\n>    the filesystem (perhaps by running an external merge\n>    program).  It wants to register it as the result of the\n>    merge.\n> \n>    We could do this today by first storing the \"desired blob\"\n>    in a temporary file somewhere in the path the dircache\n>    controls, \"update-cache --add\" the temporary file, ls-tree to\n>    find its mode/sha1, \"update-cache --remove\" the temporary\n>    file and finally \"update-cache --cacheinfo\" the mode/sha1.\n>    This is workable but clumsy.  How about:\n> \n>    $ update-cache --graft [--add] desired-blob path\n> \n>    to say \"I want to register mode/sha1 from desired-blob, which\n>    may not be of verify_path() satisfying name, at path in the\n>    dircache\"?\n> \n>  * The merge policy decides to delete the path.\n> \n>    We could do this today by first stashing away the file at the\n>    path if it exists, \"update-cache --remove\" it, and restore\n>    if necessary.  This is again workable but clumsy.  How about:\n> \n>    $ update-cache --force-remove path\n> \n>    to mean \"I want to remove the path from dircache even though\n>    it may exist in my working tree\"?\n\nYes.\n\n> Am I on the right track?\n\nExactly.\n\n> You might want to go even lower level by letting them say\n> something like:\n> \n>  * update-cache --register-stage mode sha1 stage path\n> \n>    Registers the mode/sha1 at stage for path.  Does not look at\n>    the working tree.  stage is [0-3]\n\nI'd prefer not. I'd avoid playing games with the stages at any other level\nthan the \"full tree\" level until we show a real need for it.\n\nLet's go with the known-needed minimal cases that are high-level enough to\nmake the scripting simple, and see if there is any reason to ever touch\nthe tree any other way.\n\n\t\tLinus\n"},{"id":"316","messageId":"Pine.LNX.4.58.0504161733110.31775@wgmdd8.biozentrum.uni-wuerzburg.de","threadId":"9","inReplyTo":"1113580881.27227.73.camel@hades.cambridge.redhat.com","subject":"Re: Merge with git-pasky II.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2005-04-16T15:33:52Z","receivedAt":"2005-04-16T15:33:52Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 15 Apr 2005, David Woodhouse wrote:\n\n> But if it can be done cheaply enough at a later date even though we end\n> up repeating ourselves, and if it can be done _well_ enough that we\n> shouldn't have just asked the user in the first place, then yes, OK I\n> agree.\n\nThe repetition could be helped by using a cache.\n\nCiao,\nDscho\n"},{"id":"317","messageId":"20050416153401.GD19099@pasky.ji.cz","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504151709180.7211@ppc970.osdl.org","subject":"Re: Re: Re: Re: write-tree is pasky-0.4","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-16T15:34:01Z","receivedAt":"2005-04-16T15:34:01Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Apr 16, 2005 at 02:22:45AM CEST, I got a letter\nwhere Linus Torvalds <torvalds@osdl.org> told me that...\n> \n> \n> On Sat, 16 Apr 2005, Petr Baudis wrote:\n> > \n> > But otherwise it is great news to me. Actually, in that case, is it\n> > worth renaming it to Cogito and using cg to invoke it? Wouldn't be that\n> > actually more confusing after it gets merged? IOW, should I stick to\n> > \"git\" or feel free to rename it to \"cg\"?\n> \n> I'm perfectly happy for it to stay as \"git\", and in general I don't have\n> any huge preferences either way. You guys can discuss names as much as you\n> like, it's the \"tracking renames\" and \"how to merge\" things that worry me.\n\n:-)\n\n> I think I've explained my name tracking worries.  When it comes to \"how to \n> merge\", there's three issues:\n> \n>  - we do commonly have merge clashes where both trees have applied the \n>    exact same patch. That should merge perfectly well using the 3-way\n>    merge from a common parent that Junio has, but not your current \"bring\n>    patches forward\" kind of strategy.\n\nMy current \"bring patches forward\" strategy is only very interim, to\nhave something working well enough for me to merge with you. I will\ngladly change it to use merge-tree*, when it is done. (Or read-tree -m -\nI will yet have to have a look, but it looks extremely promising.)\n\n>  - I _do_ actually sometimes merge with dirty state in my working \n>    directory, which is why I want the merge to take place in a separate \n>    (and temporary) directory, which allows for a failed merge without \n>    having any major cleanup. If the merge fails, it's not a big deal, and \n>    I can just blow the merge directory away without losing the work I had \n>    in my \"real\" working directory.\n\nOk. But still, especially when you do some nontrivial conflicts\nresolving, how do you check if it even compiles after the merge? Or do\nyou just commit it and possibly fix the compilation in another commit?\n\n>  - reliability. I care much less for \"clever\" than I care for \"guaranteed \n>    to never do the wrong thing\". If I have to fix up some stuff by hand, \n>    I'll happily do so. But if I can't trust the merge and have to _check_ \n>    things by hand afterwards, that will make me leery of the merges, and\n>    _that_ is bad.\n> \n> The third point is why I'm going to the ultra-conservative \"three-way \n> merge from the common parent\". It's not fancy, but it's something I feel \n> comfortable with as a merge strategy. For example, arch (and in particular \n> darcs) seems to want to try to be \"clever\" about the merges, and I'd \n> always live in fear. \n\nI agree and I would like to achieve the same. I too think the three-way\nmerge from the common parent is the best way to go for now.\n\n> And, finally, there's obviously performance. I _think_ a normal merge with\n> nary a conflict and just a few tens of files changed should be possible in\n> a second. I realize that sounds crazy to some people, but I think it's\n> entirely doable. Half of that is writing the new tree out (that is a\n> relative costly op due to the compression). The other half is the \"work\".\n\nBeing written in shell, there is plenty of space for optimization - from\nusing bash internals instead of textutils to rewriting parts of it in C.\nMy priority now is to get it right first, though. :-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"320","messageId":"20050416155536.GX4488@himi.org","threadId":"9","inReplyTo":"Pine.LNX.4.62.0504160518310.21837@qynat.qvtvafvgr.pbz","subject":"Re: Merge with git-pasky II.","fromName":"Simon Fowler","fromEmail":"simon@himi.org","sentAt":"2005-04-16T15:55:37Z","receivedAt":"2005-04-16T15:55:37Z","isPatch":false,"sender":{"key":"simon@himi.org","avatar":null},"body":"On Sat, Apr 16, 2005 at 05:19:24AM -0700, David Lang wrote:\n> Simon\n> \n> given that you have multiple machines creating files, how do you deal with \n> the idea of the same 'unique id' being assigned to different files by \n> different machines?\n> \nThe id is a sha1 hash of the current time and the full path of the\nfile being added - the chances of that being replicated without\nmalicious intent is extremely small. There are other things that\ncould be used, like the hostname, username of the person running the\nprogram, etc, but I don't really see them being necessary.\n\nSimon\n\n-- \nPGP public key Id 0x144A991C, or http://himi.org/stuff/himi.asc\n(crappy) Homepage: http://himi.org\ndoe #237 (see http://www.lemuria.org/DeCSS) \nMy DeCSS mirror: ftp://himi.org/pub/mirrors/css/ \n"},{"id":"321","messageId":"20050416160333.GF19099@pasky.ji.cz","threadId":"9","inReplyTo":"20050416155536.GX4488@himi.org","subject":"Re: Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-16T16:03:33Z","receivedAt":"2005-04-16T16:03:33Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Apr 16, 2005 at 05:55:37PM CEST, I got a letter\nwhere Simon Fowler <simon@himi.org> told me that...\n> On Sat, Apr 16, 2005 at 05:19:24AM -0700, David Lang wrote:\n> > Simon\n> > \n> > given that you have multiple machines creating files, how do you deal with \n> > the idea of the same 'unique id' being assigned to different files by \n> > different machines?\n> > \n> The id is a sha1 hash of the current time and the full path of the\n> file being added - the chances of that being replicated without\n> malicious intent is extremely small. There are other things that\n> could be used, like the hostname, username of the person running the\n> program, etc, but I don't really see them being necessary.\n\nWhy not just use UUID?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"323","messageId":"20050416162614.GY4488@himi.org","threadId":"9","inReplyTo":"20050416160333.GF19099@pasky.ji.cz","subject":"Re: Re: Merge with git-pasky II.","fromName":"Simon Fowler","fromEmail":"simon@himi.org","sentAt":"2005-04-16T16:26:14Z","receivedAt":"2005-04-16T16:26:14Z","isPatch":false,"sender":{"key":"simon@himi.org","avatar":null},"body":"On Sat, Apr 16, 2005 at 06:03:33PM +0200, Petr Baudis wrote:\n> Dear diary, on Sat, Apr 16, 2005 at 05:55:37PM CEST, I got a letter\n> where Simon Fowler <simon@himi.org> told me that...\n> > On Sat, Apr 16, 2005 at 05:19:24AM -0700, David Lang wrote:\n> > > Simon\n> > > \n> > > given that you have multiple machines creating files, how do you deal with \n> > > the idea of the same 'unique id' being assigned to different files by \n> > > different machines?\n> > > \n> > The id is a sha1 hash of the current time and the full path of the\n> > file being added - the chances of that being replicated without\n> > malicious intent is extremely small. There are other things that\n> > could be used, like the hostname, username of the person running the\n> > program, etc, but I don't really see them being necessary.\n> \n> Why not just use UUID?\n> \nHey, everything else in git seems to use sha1, so I just copied\nLinus' sha1 code ;-)\n\nAll I wanted was something that had a good chance of being unique\nacross any potential set of distributed repositories, to avoid the\nchance of accidental clashes. A sha1 hash of something that's not\nlikely to be replicated is a simple way to do that.\n\nSimon\n\n-- \nPGP public key Id 0x144A991C, or http://himi.org/stuff/himi.asc\n(crappy) Homepage: http://himi.org\ndoe #237 (see http://www.lemuria.org/DeCSS) \nMy DeCSS mirror: ftp://himi.org/pub/mirrors/css/ \n"},{"id":"322","messageId":"Pine.LNX.4.58.0504160913180.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"20050416160333.GF19099@pasky.ji.cz","subject":"Re: Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T16:26:57Z","receivedAt":"2005-04-16T16:26:57Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Apr 2005, Petr Baudis wrote:\n\n> Dear diary, on Sat, Apr 16, 2005 at 05:55:37PM CEST, I got a letter\n> where Simon Fowler <simon@himi.org> told me that...\n>\n> > The id is a sha1 hash of the current time and the full path of the\n> > file being added - the chances of that being replicated without\n> > malicious intent is extremely small. There are other things that\n> > could be used, like the hostname, username of the person running the\n> > program, etc, but I don't really see them being necessary.\n> \n> Why not just use UUID?\n\nNote that using anything that isn't data-related totally destroys the \nwhole point of the object database. Remember: any time we don't uniquely \ngenerate the same name for the same object, we'll waste disk-space.\n\nSo adding in user/machine/uuid's to the thing is always a mistake. The \nwhole thing depends on the hash being as close to 1:1 with the contents as \nhumanly possible. \n\nThere's also the issue of size. Yes, I could have chosen sha256 instead of\nsha1. But the keys would be almost twice as big, which in turn means that \nthe \"tree\" objects would be bigger, and that the \"index\" file would be \nbigger.\n\nIs that a huge problem? No. We can certainly move to it if sha1 ever shows\nitself to be weak. But I really think we are much better off just\nre-generating the whole tree and history at that point, rather than try to \npredict the future.\n\nThe fact is, with current knowledge, sha1 _is_ safe for what git uses it \nfor, for the forseeable future. And we have a migration strategy if I'm \nwrong. Don't worry about it.\n\nAlmost all attacks on sha1 will depend on _replacing_ a file with a bogus\nnew one. So guys, instead of using sha256 or going overboard, just make \nsure that when you synchronize, you NEVER import a file you already have.\n\nIt's really that simple. Add \"--ignore-existing\" to your rsync scripts,\nand you're pretty much done. That guarantees that a new evil blob by the\nnext mad scientist out to take over the world will never touch your\nrepository, and if we make this part of the _standard_ scripts, then\ndammit, security is in good _practices_ rather than just relying blindly\non the hash being secure.\n\nIn other words, I think we could have used md5's as the hash, if we just\nmake sure we have good practices. And it wouldn't have been \"insecure\".\n\nThe fact is, you don't merge with people you don't trust. If you don't\ntrust them, they have a much easier time corrupting your repository by\njust creating bugs in the code and checking that thing in. Who cares about\nhash collisions, when you can generate a kernel root vulnerability by just\nadding a single line of code and use the _correct_ hash for it.\n\nSo the sha1 hash does not replace _trust_. That comes from something else \naltogether.\n\n\t\t\tLinus\n"},{"id":"324","messageId":"Pine.LNX.4.58.0504160928250.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504160820320.7211@ppc970.osdl.org","subject":"Re: [PATCH 3/2] merge-trees script for Linus git","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T16:36:25Z","receivedAt":"2005-04-16T16:36:25Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Apr 2005, Linus Torvalds wrote:\n> \n> Having slept on it, I think I'll merge all the trivial cases that don't \n> involve a file going away or being added. Ie if the file is in all three \n> trees, but it's the same in two of them, we know what to do.\n\nJunio, I pushed this out, along with the two patches from you. It's still\nmore anal than my original \"tree-diff\" algorithm, in that it refuses to\ntouch anything where the name isn't the same in all three versions\n(original, new1 and new2), but now it does the \"if two of them match, just\nselect the result directly\" trivial merges.\n\nI really cannot see any sane case where user policy might dictate doing\nanything else, but if somebody can come up with an argument for a merge\nalgorithm that wouldn't do what that trivial merge does, we can make a\nflag for \"don't merge at all\".\n\nThe reason I do want to merge at all in \"read-tree\" is that I want to\navoid having to write out a huge index-file (it's 1.6MB on the kernel, so\nif you don't do _any_ trivial merges, it would be 4.8MB after reading\nthree trees) and then having people read it and parse it just to do stuff\nthat is obvious. Touching 5MB of data isn't cheap, even if you don't do a \nwhole lot to it.\n\nAnyway, with the modified read-tree, as far as I can tell it will now \nmerge all the cases where one side has done something to a file, and the \nother side has left it alone (or where both sides have done the exact same \nmodification). That should _really_ cut down the cases to just a few files \nfor most of the kernel merges I can think of. \n\nDoes it do the right thing for your tests?\n\n\t\tLinus\n"},{"id":"328","messageId":"7v64ym8wzu.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504160928250.7211@ppc970.osdl.org","subject":"Re: [PATCH 3/2] merge-trees script for Linus git","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T17:14:29Z","receivedAt":"2005-04-16T17:14:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> Anyway, with the modified read-tree, as far as I can tell it will now \nLT> merge all the cases where one side has done something to a file, and the \nLT> other side has left it alone (or where both sides have done the exact same \nLT> modification). That should _really_ cut down the cases to just a few files \nLT> for most of the kernel merges I can think of. \n\nLT> Does it do the right thing for your tests?\n\nYes.\n\n\n"},{"id":"354","messageId":"E1DMtuY-0002fL-3b@approximate.corpus.cam.ac.uk","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504150753440.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Sanjoy Mahajan","fromEmail":"sanjoy@mrao.cam.ac.uk","sentAt":"2005-04-16T20:29:02Z","receivedAt":"2005-04-16T20:29:02Z","isPatch":false,"sender":{"key":"sanjoy@mrao.cam.ac.uk","avatar":null},"body":"> And that \"where did this come from\" decision should be done at _search_\n> time, not commit time.\n\nI like this elegant approach, but clever pattern matching can help even\nat commit time.  Suppose hello.c is simply:\n\n  printf (\"Hello %d\\n\", year);\n\nAnd then developer A updates hello.c to:\n\n  printf (\"Hello %d\\n\", year);\n  printf (\"And   %d\\n\", year+1);\n\nMeanwhile developer B updates hello.c to:\n\n  printf (\"Hello %d\\n\", yyyy);\n\nHow to merge these two changes?  The psychic solution is\n\n  printf (\"Hello %d\\n\", yyyy);\n  printf (\"And   %d\\n\", yyyy+1);\n\nDarcs handles token renames specially, but it's not a general solution\nso let's leave it aside.  The example does not have enough information\nto make the psychic solution unique or reliable, but imagine that the\nexample were longer to solve that problem.  You'd want to describe the\ndelta A(hello.c) as\n   \n  1. duplicated message line\n  2. changed 2nd line a bit\n\nAnd B(hello.c) as\n\n  1. Changed year to yyyy\n\nIn that representation, merging the two deltas becomes\n\n  1. duplicated message line\n  2. changed 2nd line a bit\n  3. Changed year to yyyy in both lines\n\nOr, by commuting the merge operations and adjusting for their\nnon-commutativity (in terminology like darcs's -- I'm also a physicist):\n\n  1. Changed year to yyyy\n  2. duplicated message line\n  3. changed 2nd line a bit\n\nSo here some of the computation that Linus wants only at question time\n(e.g. 'how did that line get here??') is also useful at merge time.\nIt's difficult (expensive, unreliable) to describe deltas in the form\nabove or, worse, to merge two such descriptions, but I hope it\nillustrates the point.  And perhaps a robust and easier-to-compute\nchange-description language can be dreamt up, even if the general\nproblem of describing changes compactly is not computable -- it's almost\nthe same, or is the same, as finding the Kolmogorov complexity of a data\nset.\n\nOr have I missed a fundamental point?\n\n-Sanjoy\n\n`A society of sheep must in time beget a government of wolves.'\n   - Bertrand de Jouvenal\n"},{"id":"355","messageId":"Pine.LNX.4.58.0504161334540.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"E1DMtuY-0002fL-3b@approximate.corpus.cam.ac.uk","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T20:41:31Z","receivedAt":"2005-04-16T20:41:31Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Apr 2005, Sanjoy Mahajan wrote:\n> \n> I like this elegant approach, but clever pattern matching can help even\n> at commit time.  Suppose hello.c is simply:\n\nHere, what you're talking about is not \"commit\", but \"merge\".\n\nThe git model very much separates the two events. You first generate a \nmerged tree, an dyou commit that merge as a separate and largely totally \nindependent phase.\n\nAnd yes, I agree that with merging, you do end up potentially wanting to \ntry different things. When I've done my \"git\" merges, all I've really done \nis to make sure that the trivial parts basically merge in zero time, so \nthat you can afford to perhaps spend some effort on handling the _real_ \nmerge conflicts.\n\nMany systems seem to be designed around a \"clever merge\" algorithm, with \ndarcs perhaps being the most extreme example. The problem with that design \nis that 99.9% of all the work is not at all about being clever, and if you \ntry to base your design around the clever things, your performance will \ndefinitely suck.\n\nSo I think that with git, you can actually really try to be clever,\nbecause when you get a merge conflict, you're now only worrying about one\nfile out of 17,000, and then you can go wild on that one and try different\nmerge algorithms (token merge, character-merge, line-based merge, you name\nit).\n\nOf course, I might not actually personally want to depend on any clever \nmerges, but the git infrastructure really doesn't care. My plumbing \ndoesn't merge the conflicts that arise within one single object, or the \nfilename differences - you can do anything you want on that.\n\n\t\tLinus\n"},{"id":"376","messageId":"Pine.LNX.4.62.0504161559140.22652@qynat.qvtvafvgr.pbz","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504160913180.7211@ppc970.osdl.org","subject":"Re: Re: Merge with git-pasky II.","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2005-04-16T23:02:04Z","receivedAt":"2005-04-16T23:02:04Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Sat, 16 Apr 2005, Linus Torvalds wrote:\n\n> Almost all attacks on sha1 will depend on _replacing_ a file with a bogus\n> new one. So guys, instead of using sha256 or going overboard, just make\n> sure that when you synchronize, you NEVER import a file you already have.\n>\n> It's really that simple. Add \"--ignore-existing\" to your rsync scripts,\n> and you're pretty much done. That guarantees that a new evil blob by the\n> next mad scientist out to take over the world will never touch your\n> repository, and if we make this part of the _standard_ scripts, then\n> dammit, security is in good _practices_ rather than just relying blindly\n> on the hash being secure.\n>\n> In other words, I think we could have used md5's as the hash, if we just\n> make sure we have good practices. And it wouldn't have been \"insecure\".\n>\n> The fact is, you don't merge with people you don't trust. If you don't\n> trust them, they have a much easier time corrupting your repository by\n> just creating bugs in the code and checking that thing in. Who cares about\n> hash collisions, when you can generate a kernel root vulnerability by just\n> adding a single line of code and use the _correct_ hash for it.\n>\n> So the sha1 hash does not replace _trust_. That comes from something else\n> altogether.\n\nWhat I am bringing up is not intended to be a trust thing, but instead a \nsafety thing, accidents, not evil intent. makeing the rsync scripts \n--ignore-existing will avoid corrupting local data when pulling remotely, \nbut it won't solve the problem of running into a collision locally (and \nwon't do much to help you figure out what's wrong when you run into a \nremote collision)\n\nDavid Lang\n\n-- \nThere are two ways of constructing a software design. One way is to make it so simple that there are obviously no deficiencies. And the other way is to make it so complicated that there are no obvious deficiencies.\n  -- C.A.R. Hoare\n"},{"id":"439","messageId":"7v64ym2dju.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"7vll7i95u1.fsf_-_@assigned-by-dhcp.cox.net","subject":"Re: Issues with higher-order stages in dircache","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-17T05:11:01Z","receivedAt":"2005-04-17T05:11:01Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus,\n\n    earlier I wrote [*R1*]:\n\n   - An explicit \"update-cache [--add] [--remove] path\" should\n     be taken as a signal from the user (or Cogito) to tell the\n     dircache layer \"the merge is done and here is the result\".\n     So just delete higher-order stages for the path and record\n     the specified path at stage 0 (or remove it altogether).\n\nand I think this commit of yours implements the adding half.\n\n    commit be7b1f05cea8e5213ffef8f74ebdefed2aacb6fc:1\n    author Linus Torvalds <torvalds@ppc970.osdl.org> 1113678345 -0700\n    committer Linus Torvalds <torvalds@ppc970.osdl.org> 1113678345 -0700\n\n    When inserting a index entry of stage 0, remove all old unmerged entries.\n\nI am wondering if you have a particular reason not to do the\nsame for the removing half.  Without it, currently I do not see\na way for the user or Cogito to tell dircache layer that the\nmerge should result in removal.  That is, other than first\nadding a phony entry there (which brings the entry down to stage\n0) and then immediately doing a regular update-cache --remove.\nThat is two instead of one reading of 1.6MB index file for the\nkernel case.\n\nAlso do you have any comments on this one from the same message?\n\n * read-tree\n\n   - When merging two trees, i.e. \"read-tree -m A B\", shouldn't\n     we collapse identical stage-1/2 into stage-0?\n\n\n[References]\n\n*R1* http://marc.theaimsgroup.com/?l=git&m=111366023126466&w=2\n\n"},{"id":"440","messageId":"Pine.LNX.4.58.0504162228300.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"7v64ym2dju.fsf@assigned-by-dhcp.cox.net","subject":"Re: Issues with higher-order stages in dircache","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-17T05:31:26Z","receivedAt":"2005-04-17T05:31:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Apr 2005, Junio C Hamano wrote:\n> \n> I am wondering if you have a particular reason not to do the\n> same for the removing half.\n\nNo. Except for me being silly.\n\nPlease just make it so.\n\n> Also do you have any comments on this one from the same message?\n> \n>  * read-tree\n> \n>    - When merging two trees, i.e. \"read-tree -m A B\", shouldn't\n>      we collapse identical stage-1/2 into stage-0?\n\nHow do you actually intend to merge two trees? \n\nThat sounds like a total special case, and better done with \"diff-tree\".  \nBut regardless, since I assume the result is the later tree, why do a \n\"read-tree -m A B\", since what you really want is \"read-tree B\"?\n\nThe real merge always needs the base tree, and I'd hate to complicate the \nreal merge with some special-case that isn't relevant for that real case.\n\n\t\tLinus\n"},{"id":"448","messageId":"7vll7i0wmq.fsf@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504162228300.7211@ppc970.osdl.org","subject":"Re: Issues with higher-order stages in dircache","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-17T06:01:49Z","receivedAt":"2005-04-17T06:01:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\n>> - When merging two trees, i.e. \"read-tree -m A B\", shouldn't\n>> we collapse identical stage-1/2 into stage-0?\n\nLT> How do you actually intend to merge two trees? \n\nHow silly of me.  *BLUSH*\n\n"},{"id":"461","messageId":"7v4qe5yb7t.fsf_-_@assigned-by-dhcp.cox.net","threadId":"9","inReplyTo":"7vll7i95u1.fsf_-_@assigned-by-dhcp.cox.net","subject":"Summary of \"read-tree -m O A B\" mechanism","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-17T10:00:22Z","receivedAt":"2005-04-17T10:00:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Earlier I wrote down a list of issues your recent \"merge\nstage\" changes have introduced to the rest of the plumbing, with\na set of suggested adaptions.  I think all of them are cleared\nnow (you have a pile of patches from me in your mailbox).\n\nI do not know what percentage of people on this list are using\ngit without the Cogito part, but I suspect that the number might\nbe quite small.  I also suspect, from the description Petr gave\nus on how the merging in Cogito works, Cogito does not currently\nuse the \"read-tree -m O A B\" mechanism, and those majority who\ndo not deal with the low level tools themselves would not have\nto know about the merge issues yet.  But I think it is a good\ntime, now things have started to settle down, to summarize how\nvarious commands work when they see those \"funny\" dircache\nentries created after \"read-tree -m O A B\" has run.  Of course,\npeople working on Cogito needs to know them, once they decide to\nuse the \"reed-tree -m O A B\" mechanism.\n\n * read-tree -m O A B\n\n   - For description on how this works, the definitive reading\n     is [*R1*].  In short:\n\n     - unlike ordinary read-tree, \"-m\" form reads up to three\n       trees and creates paths that are \"unmerged\".  \n\n     - trivial merges are done by read-tree itself.  only\n       conflicting paths will be in unmerged state when\n       read-tree returns.\n\n * write-tree\n\n     - write-tree refuses to give you a tree until all the\n       unmerged paths are resolved.\n\n * show-files\n\n   - \"show-files --unmerged\" and \"show-files --stage\" can be\n     used to examine detailed information on unmerged paths.\n     For an unmerged path, instead of recording a single\n     mode/SHA1 pair, the dircache records up to three such\n     pairs; one from tree O in stage 1, A in stage 2, and B in\n     stage 3.  This information can be used by the user (or\n     Cogito) to see what should eventually be recorded at the\n     path.\n\n * update-cache\n\n   - An explicit \"update-cache [--add] path\" or \"update-cache\n     [--add] --cacheinfo mode SHA1 path\" tells the plumbing that\n     the user (or Cogito) wants to resolve it by storing\n     mode/SHA1 of the given working file or mode SHA1 specified\n     on the command line.  The path ceases to be in unmerged\n     state after this happens.\n\n     Similarly, \"update-cache --remove path\" resolves the\n     unmerged state and the merge result is not having anything\n     at that path.\n\n   - \"update-cache --refresh\", in addition to the \"needs update\"\n     message people are now familiar with, says \"needs merge\"\n     for unmerged paths.\n\n * show-diff\n\n   - show-diff on an unmerged path simply says \"unmerged\" (the\n     plumbing would not know what to diff with what among three\n     stages and the working file).  \n\n * checkout-cache\n\n   - \"checkout-cache -a\" warns about unmerged paths and checks\n     out only the merged paths.\n\n   - \"checkout-cache [-f] path\" on an unmerged path says\n     \"Unmerged\", just like the same command on non-existent path\n     says \"not in the cache\", and does not touch the working\n     file.\n \n\nI hope the descriptions in this summary is correct enough to be\nuseful to somebody.\n\n\n[Reference]\n\n*R1* http://marc.theaimsgroup.com/?l=git&m=111363270608902&w=2\n\n"},{"id":"471","messageId":"1113743652.3884.2.camel@localhost.localdomain","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504161733110.31775@wgmdd8.biozentrum.uni-wuerzburg.de","subject":"Re: Merge with git-pasky II.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-17T13:14:11Z","receivedAt":"2005-04-17T13:14:11Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Sat, 2005-04-16 at 17:33 +0200, Johannes Schindelin wrote:\n> > But if it can be done cheaply enough at a later date even though we end\n> > up repeating ourselves, and if it can be done _well_ enough that we\n> > shouldn't have just asked the user in the first place, then yes, OK I\n> > agree.\n> \n> The repetition could be helped by using a cache.\n\nPerhaps. Since neither such a cache nor even the commit comments are\nstrictly part of the git data, they probably shouldn't be included in\nthe sha1 hash of the commit object. However, I don't see a fundamental\nreason why we couldn't store them in the same file but omit them from\nthe hash calculations. That also allows us to retrospectively edit\ncommit comments without completely changing the entire subsequent\nhistory.\n\nOr is that a little too heretical a suggestion?\n\n-- \ndwmw2\n\n"},{"id":"475","messageId":"20050417145232.GA5289@elte.hu","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504160913180.7211@ppc970.osdl.org","subject":"Re: Re: Merge with git-pasky II.","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-17T14:52:32Z","receivedAt":"2005-04-17T14:52:32Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* Linus Torvalds <torvalds@osdl.org> wrote:\n\n> Almost all attacks on sha1 will depend on _replacing_ a file with a \n> bogus new one. So guys, instead of using sha256 or going overboard, \n> just make sure that when you synchronize, you NEVER import a file you \n> already have.\n\nhere is a bit complex, but still practical attack that doesnt rely on \nreplacement and which can only be detected if we check the sha1 \nuniqueness assumptions.\n\nIf you can generate a duplicate sha1 key for an arbitrary 'target' file, \nand Malice sends you a GIT-generated patch that introduces a new file \n(which doesnt exist in the current tree) which you review (in the email) \nand which looks safe to apply & harmless. Maybe the patch has a bit \nweird formatting and some weird comments (which in reality Malice used \nto generate the proper sha1 key) but otherwise the patch is for some \nseldom used arcane driver that no-one used for quite some time and \nno-one really cares about, so you are happy to apply the patch.\n\nThe compromise occurs when you apply the patch: the seemingly harmless \npatch has an sha1 key that Malice manufacured to match that of an \nalready existing, 'dangerous' object in your database.\n\nWith tens of thousands (or hundreds of thousands) of objects expected in \nthe repository sooner or later, there's quite a selection to pick from.  \nOnce you apply the patch, instead of the expected new file that you \nreviewed and found safe, the attacker has the other object included in \nthe official kernel.\n\nA dangerous object can be anything: e.g. a debugging hack that allows \narbitrary kernel-space writes. Or a known-insecure module (which since \nthen got fixed, but the buggy code still exists in the DB). The module \nis in a single file and is self-installing (e.g. it has __init code to \nregister itself as some driver.)\n\nMalice might even previously plant a dangerous object as some 'firmware \nmodule' in another arcane driver, which doesnt get compiled by default, \nbut still shows up in the DB. Or Malice might plant a dangerous object \nvia an innocent-looking documentation file.  (which contains some sample \ncode and is called sample.txt)\n\nthis type of 'false sharing attack' can only be prevented if an object \nis only 'shared' with another object if it has been memcmp-ed with the \nobject in the repository. I.e. if we trust the sharing decision! Once \nthe attack has occured it cannot be detected automatically: only people \nwill notice it. (why did that weird unrelated module show up in that old \ndriver?)\n\nThe compromise relies on you having reviewed something harmless, while \nin reality what happened within the DB was far less harmless. And the DB \nremains self-consistent: neither fsck, nor others importing your tree \nwill be able to detect the compromise. This attack can only be detected \nwhen you apply the patch, after that point all the information (except \nMalice's message in your inbox) is gone.\n\nso unless we actively check for collisions, once an sha1 key can be \ngenerated at will on near-arbitrary input, it's not a secure system \nanymore. We might be lucky and safe, but we wont be secure.\n\n\tIngo\n"},{"id":"476","messageId":"Pine.LNX.4.44.0504170804130.2625-100000@bellevue.puremagic.com","threadId":"9","inReplyTo":"20050417145232.GA5289@elte.hu","subject":"Re: Re: Merge with git-pasky II.","fromName":"Brad Roberts","fromEmail":"braddr@puremagic.com","sentAt":"2005-04-17T15:08:35Z","receivedAt":"2005-04-17T15:08:35Z","isPatch":false,"sender":{"key":"braddr@puremagic.com","avatar":null},"body":"On Sun, 17 Apr 2005, Ingo Molnar wrote:\n\n> Date: Sun, 17 Apr 2005 16:52:32 +0200\n> From: Ingo Molnar <mingo@elte.hu>\n> To: Linus Torvalds <torvalds@osdl.org>\n> Cc: Petr Baudis <pasky@ucw.cz>, Simon Fowler <simon@himi.org>,\n>      David Lang <david.lang@digitalinsight.com>, git@vger.kernel.org\n> Subject: Re: Re: Merge with git-pasky II.\n>\n>\n> * Linus Torvalds <torvalds@osdl.org> wrote:\n>\n> > Almost all attacks on sha1 will depend on _replacing_ a file with a\n> > bogus new one. So guys, instead of using sha256 or going overboard,\n> > just make sure that when you synchronize, you NEVER import a file you\n> > already have.\n>\n> With tens of thousands (or hundreds of thousands) of objects expected in\n> the repository sooner or later, there's quite a selection to pick from.\n> Once you apply the patch, instead of the expected new file that you\n> reviewed and found safe, the attacker has the other object included in\n> the official kernel.\n>\n> A dangerous object can be anything: e.g. a debugging hack that allows\n> arbitrary kernel-space writes. Or a known-insecure module (which since\n> then got fixed, but the buggy code still exists in the DB). The module\n> is in a single file and is self-installing (e.g. it has __init code to\n> register itself as some driver.)\n\nWhile I agree that a hash collision is bad and certainly worth preventing\nduring new object creation, for it to actually implant a trojan in a build\nsuccessfully it'd have to meet even more criteria than you've layed out.\nIt'd have to...\n\n  - be shadowing an object that's part of an active tree\n  - provide all the public symbols the shadowed object provided so that it\n    would still build and link successfully\n\nShadowing an object that's not part of the working tree means something on\nanother branch or obsoleted some time in the past is still db corruption,\nbut not nearly as big an issue from a trojan standpoint.\n\nLater,\nBrad\n\n\n"},{"id":"477","messageId":"20050417151858.GA5745@elte.hu","threadId":"9","inReplyTo":"Pine.LNX.4.44.0504170804130.2625-100000@bellevue.puremagic.com","subject":"Re: Re: Merge with git-pasky II.","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-17T15:18:58Z","receivedAt":"2005-04-17T15:18:58Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* Brad Roberts <braddr@puremagic.com> wrote:\n\n> While I agree that a hash collision is bad and certainly worth \n> preventing during new object creation, for it to actually implant a \n> trojan in a build successfully it'd have to meet even more criteria \n> than you've layed out. It'd have to...\n\n>   - provide all the public symbols the shadowed object provided so that it\n>     would still build and link successfully\n\nthat's not a problem. Most modules dont provide public symbols.  \nEspecially not drivers. Generally it's the modules that _dont_ have any \nglobal impact that get reviewed less stringently - an attacker would \nthus choose them for psychological reasons anyway.\n\n>   - be shadowing an object that's part of an active tree\n>\n> Shadowing an object that's not part of the working tree means \n> something on another branch or obsoleted some time in the past is \n> still db corruption, but not nearly as big an issue from a trojan \n> standpoint.\n\nit's not DB corruption, it's a feature of GIT: it's a content _cache_, \nnew and old alike. Nothing in GIT says that old objects in the \nrepository (which are still very much part of history) cannot be revived \nin newer trees. (in fact it regularly happens - e.g. if a fix is undone \nmanually.)\n\n\tIngo\n"},{"id":"482","messageId":"20050417152841.GA6157@elte.hu","threadId":"9","inReplyTo":"20050417145232.GA5289@elte.hu","subject":"Re: Re: Merge with git-pasky II.","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-17T15:28:41Z","receivedAt":"2005-04-17T15:28:41Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* Ingo Molnar <mingo@elte.hu> wrote:\n\n> The compromise relies on you having reviewed something harmless, while \n> in reality what happened within the DB was far less harmless. And the \n> DB remains self-consistent: neither fsck, nor others importing your \n> tree will be able to detect the compromise. This attack can only be \n> detected when you apply the patch, after that point all the \n> information (except Malice's message in your inbox) is gone.\n\nin fact, this attack cannot even be proven to be malicious, purely via \nthe email from Malice: it could be incredible bad luck that caused that \ngood-looking patch to be mistakenly matching a dangerous object.\n\nIn fact this could happen even today, _accidentally_. (but i'm willing \nto bet that hell will be freezing over first, and i'll have some really \ngood odds ;) There's probably a much higher likelyhood of Linus' tree \ngetting corrupted in some old fashioned way and introducing a security \nhole by accident)\n\n\tIngo\n"},{"id":"505","messageId":"Pine.LNX.4.58.0504171014430.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"20050417152841.GA6157@elte.hu","subject":"Re: Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-17T17:34:57Z","receivedAt":"2005-04-17T17:34:57Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 17 Apr 2005, Ingo Molnar wrote:\n> \n> in fact, this attack cannot even be proven to be malicious, purely via \n> the email from Malice: it could be incredible bad luck that caused that \n> good-looking patch to be mistakenly matching a dangerous object.\n\nI really hate theoretical discussions. \n\nThe fact is, a lot of _crap_ engineering gets done because of the question\n\"what if?\". It results in over-engineering, often to the point where the \nend result is quite a lot measurably worse than the sane results.\n\nYou are _literally_ arguing for the equivalent of \"what if a meteorite hit\nmy plane while it was in flight - maybe I should add three inches of\nhigh-tension armored steel around the plane, so that my passengers would\nbe protected\".\n\nThat's not engineering. That's five-year-olds discussing building their\nimaginary forts (\"I want gun-turrets and a mechanical horse one mile high,\nand my command center is 5 miles under-ground and totally encased in 5\nmeters of lead\").\n\nI absolutely _hate_ doing engineering on the principle of \"this might be\npossible in theory\", and I'm violently opposed to it. So far, I have not\nheard a single argument that I consider even _remotely_ likely.\n\nThe thing is, even if you can force a hash collission by sending somebody \na patch, it's really pretty much almost guaranteed that the patch is not \njust \"a few strange characters\", unless sha1 is really broken to the point \nwhere it's not cryptographically secure _at_all_.\n\nIn other words, unless somebody finds a way to make sha1 appear as nothing\nmore than a complicated set of parity bits, all brute-force \"get the same\nsha1\" is likely to be about generating a really strange blob based on the\nthing you want to replace - and by \"really strange\" I mean total binary\ncrap. And likely _much_ bigger too. And by \"much bigger\" I mean \"possibly\ngigabytes of data\".\n\nAnd the thing is, _if_ somebody finds a way to make sha1 act as just a\ncomplex parity bit, and comes up with generating a clashing object that\nactually makes sense, then going to sha256 is likely pointless too - I\nthink the algorithm is basically the same, just with more bits. If you've\nbroken sha1 to the point where it's _that_ breakable, then you've likely\nbroken sha256 too. Nobody has ever proven that you couldn't break sha256 \nwith some really clever algorithm...\n\nSo if you start playing \"what if?\" games, dammit, I can play mine.\n\nIf we want to have any kind of confidence that the hash is reall\nyunbreakable, we should make it not just longer than 160 bits, we should\nmake sure that it's two or more hashes, and that they are based on totally\ndifferent principles.\n\nAnd we should all digitally sign every single object too, and we should\nuse 4096-bit PGP keys and unguessable passphrases that are at least 20\nwords in length. And we should then build a bunker 5 miles underground,\nencased in lead, so that somebody cannot flip a few bits with a ray-gun, \nand make us believe that the sha1's match when they don't. Oh, and we need \nto all wear aluminum propeller beanies to make sure that they don't use \nthat ray-gun to make us do the modification _outselves_.\n\nAnd the thing is, that's just crazy talk. The difference between a crazy\nperson and an intelligent one is that the crazy one doesn't realize what\nmakes sense in the world. The goal of good engineering is not to ask \"what\nif?\", but to ask \"how do I make this work as well as possible\".\n\nSo please stop with the theoretical sha1 attacks. It is simply NOT TRUE\nthat you can generate an object that looks halfway sane and still gets you\nthe sha1 you want. Even the \"breakage\" doesn't actually do that.  And if\nit ever _does_ become true, it will quite possibly be thanks to some\ntechnology that breaks other hashes too.\n\nSo until proven otherwise, I worry about accidental hashes, and in 160\nbits of good hashing, that just isn't an issue either. Anybody who\ncompares a 128-bit md5-sum to a 160-bit sha1 doesn't understand the math.  \nIt didn't get \"slightly less likely\" to happen. It got so _unbelievably_\nless likely to happen that it's not even funny.\n\n\t\t\t\tLinus\n"},{"id":"551","messageId":"E1DNI0G-0000bo-00@gondolin.me.apana.org.au","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504171014430.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Herbert Xu","fromEmail":"herbert@gondor.apana.org.au","sentAt":"2005-04-17T22:12:32Z","receivedAt":"2005-04-17T22:12:32Z","isPatch":false,"sender":{"key":"herbert@gondor.apana.org.au","avatar":null},"body":"Linus Torvalds <torvalds@osdl.org> wrote:\n> \n> If we want to have any kind of confidence that the hash is reall\n> yunbreakable, we should make it not just longer than 160 bits, we should\n> make sure that it's two or more hashes, and that they are based on totally\n> different principles.\n\nSorry, it has already been shown that combining two difference hashes\ndoesn't necessarily provide the security that you would hope.\n\nI think what hasn't been discussed here is the cost of actually doing\nthe comparisons.  In other words, what is the minimum number of\ncomparisons we can get away and still deal with hash collisions\nsuccessfully?\n\nOnce we know what the cost is then we can decide whether it's worthwhile\nconsidering the odds involved.\n-- \nVisit Openswan at http://www.openswan.org/\nEmail: Herbert Xu ~{PmV>HI~} <herbert@gondor.apana.org.au>\nHome Page: http://gondor.apana.org.au/~herbert/\nPGP Key: http://gondor.apana.org.au/~herbert/pubkey.txt\n"},{"id":"558","messageId":"Pine.LNX.4.58.0504171530150.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"E1DNI0G-0000bo-00@gondolin.me.apana.org.au","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-17T22:35:17Z","receivedAt":"2005-04-17T22:35:17Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 18 Apr 2005, Herbert Xu wrote:\n> \n> Sorry, it has already been shown that combining two difference hashes\n> doesn't necessarily provide the security that you would hope.\n\nSorry, that's not true.\n\nQuite the reverse. Again, you bring up totally theoretical arguments. In \n_practice_ it has indeed been shown that using two hashes _does_ catch \nhash colissions.\n\nThe trivial example is using md5 sums with a length. The \"length\" is a \nrally bad \"hash\" of the file contents too. And the fact is, that simple \ncombination of hashes has proven to be more resistant to attack than the \nhash itself. It clearly _does_ make a difference in practice.\n\nSo _please_, can we drop the obviously bogus \"in theory\" arguments. They \ndo not matter. What matters is practice.\n\nAnd the fact is, in _theory_ we don't know if somebody may be trivially\nable to break any particular hash. But in practice we do know that it's\nless likely that you can break a combination of two totally unrelated\nhashes than you break one particular one.\n\nNOTE! I'm not actually arguing that we should do that. I'm actually\narguing totally the reverse: I'm arguing that there is a fine line between\nbeing \"very very careful\" and being \"crazy to the point of being\nincompetent\".\n\n\t\tLinus\n"},{"id":"569","messageId":"20050417232905.GA2721@gondor.apana.org.au","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504171530150.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Herbert Xu","fromEmail":"herbert@gondor.apana.org.au","sentAt":"2005-04-17T23:29:05Z","receivedAt":"2005-04-17T23:29:05Z","isPatch":false,"sender":{"key":"herbert@gondor.apana.org.au","avatar":null},"body":"On Sun, Apr 17, 2005 at 03:35:17PM -0700, Linus Torvalds wrote:\n> \n> Quite the reverse. Again, you bring up totally theoretical arguments. In \n> _practice_ it has indeed been shown that using two hashes _does_ catch \n> hash colissions.\n> \n> The trivial example is using md5 sums with a length. The \"length\" is a \n> rally bad \"hash\" of the file contents too. And the fact is, that simple \n> combination of hashes has proven to be more resistant to attack than the \n> hash itself. It clearly _does_ make a difference in practice.\n\nI wasn't disputing that of course.  However, the same effect can be\nachieved in using a single hash with a bigger length, e.g., sha256\nor sha512.\n\n> So _please_, can we drop the obviously bogus \"in theory\" arguments. They \n> do not matter. What matters is practice.\n\nI agree.  However, what is the actual cost in practice of detecting\ncollisions?\n\nI get the feeling that it isn't that bad.  For example, if we did it\nat the points where the blobs actually entered the tree, then the cost\nis always proportional to the change size (the number of new blobs).\n\nIs this really that bad considering that the average blob isn't very\nbig?\n\nCheers,\n-- \nVisit Openswan at http://www.openswan.org/\nEmail: Herbert Xu ~{PmV>HI~} <herbert@gondor.apana.org.au>\nHome Page: http://gondor.apana.org.au/~herbert/\nPGP Key: http://gondor.apana.org.au/~herbert/pubkey.txt\n"},{"id":"571","messageId":"20050417233441.GU1461@pasky.ji.cz","threadId":"9","inReplyTo":"20050417232905.GA2721@gondor.apana.org.au","subject":"Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-17T23:34:41Z","receivedAt":"2005-04-17T23:34:41Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Mon, Apr 18, 2005 at 01:29:05AM CEST, I got a letter\nwhere Herbert Xu <herbert@gondor.apana.org.au> told me that...\n> I get the feeling that it isn't that bad.  For example, if we did it\n> at the points where the blobs actually entered the tree, then the cost\n> is always proportional to the change size (the number of new blobs).\n\nNo. The collision check is done in the opposite cache - when you want to\nwrite a blob and there is already a file of the same hash in the tree.\nSo either the blob is already in the database, or you have a collision.\n\nTherefore, the cost is proportional to the size of what stays unchanged.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"576","messageId":"Pine.LNX.4.58.0504171644480.7211@ppc970.osdl.org","threadId":"9","inReplyTo":"20050417232905.GA2721@gondor.apana.org.au","subject":"Re: Merge with git-pasky II.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-17T23:50:46Z","receivedAt":"2005-04-17T23:50:46Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 18 Apr 2005, Herbert Xu wrote:\n> \n> I wasn't disputing that of course.  However, the same effect can be\n> achieved in using a single hash with a bigger length, e.g., sha256\n> or sha512.\n\nNo it cannot.\n\nIf somebody actually literally totally breaks that hash, length won't \nmatter. There are (bad) hashes where you can literally edit the content of \nthe file, and make sure that the end result has the same hash.\n\nIn that case, when the hash algorithm has actually been broken, the length \nof the hash ends up being not very relevant. \n\nFor example, you might \"hash\" your file by blocking it up in 16-byte\nblocks, and xoring all blocks together - the result is a 16-byte hash.  \nIt's a terrible hash, and obviously trivially breakable, and once broken\nit does _not_ help to make it use its 32-byte cousin. Not at all. You can \njust modify the breaking thing to equally cheaply make modifications to a \nfile and get the 32-byte hash \"right\" again.\n\nIs that kind of breakage likely for sha1? Hell no. Is it possible? In your \n\"in theory\" world where practice doesn't matter, yes.\n\n\t\tLinus\n"},{"id":"577","messageId":"4262F6EA.2010005@kenjo.org","threadId":"9","inReplyTo":"20050417233441.GU1461@pasky.ji.cz","subject":"Re: Merge with git-pasky II.","fromName":"Kenneth Johansson","fromEmail":"ken@kenjo.org","sentAt":"2005-04-17T23:53:14Z","receivedAt":"2005-04-17T23:53:14Z","isPatch":false,"sender":{"key":"ken@kenjo.org","avatar":null},"body":"Petr Baudis wrote:\n> Dear diary, on Mon, Apr 18, 2005 at 01:29:05AM CEST, I got a letter\n> where Herbert Xu <herbert@gondor.apana.org.au> told me that...\n> \n>>I get the feeling that it isn't that bad.  For example, if we did it\n>>at the points where the blobs actually entered the tree, then the cost\n>>is always proportional to the change size (the number of new blobs).\n> \n> \n> No. The collision check is done in the opposite cache - when you want to\n> write a blob and there is already a file of the same hash in the tree.\n> So either the blob is already in the database, or you have a collision.\n> \n> Therefore, the cost is proportional to the size of what stays unchanged.\n> \n\n?? now I'm confused. Surly the only cost involved is to never write over \na file that already exist in the cache and that is already done NOW as \nfar as I read the code. So there is NO extra cost in detecting an collision.\n\n\n\n\n\n"},{"id":"584","messageId":"20050418004906.GA3132@gondor.apana.org.au","threadId":"9","inReplyTo":"20050417233441.GU1461@pasky.ji.cz","subject":"Re: Merge with git-pasky II.","fromName":"Herbert Xu","fromEmail":"herbert@gondor.apana.org.au","sentAt":"2005-04-18T00:49:06Z","receivedAt":"2005-04-18T00:49:06Z","isPatch":false,"sender":{"key":"herbert@gondor.apana.org.au","avatar":null},"body":"On Mon, Apr 18, 2005 at 01:34:41AM +0200, Petr Baudis wrote:\n>\n> No. The collision check is done in the opposite cache - when you want to\n> write a blob and there is already a file of the same hash in the tree.\n> So either the blob is already in the database, or you have a collision.\n> Therefore, the cost is proportional to the size of what stays unchanged.\n\nThis is only true if we're calling update-cache on all unchanged files.\nIf that's what git is doing then we're in trouble anyway.\n\nRemember that prior to the collision check we've already spent the\neffort in\n\n1) Compressing the file.\n2) Computing a SHA1 hash on the result.\n\nThese two steps together (especially the first one) is much more\nexpensive than a file content comparison of the blob versus what's\nalready in the tree.\n\nSomehow I have a hard time seeing how this can be at all efficient if\nwe're compressing all checked out files including those which are\nunchanged.\n\nTherefore the only conclusion I can draw is that we're only calling\nupdate-cache on the set of changed files, or at most a small superset\nof them.  In that case, the cost of the collision check *is* proportional\nto the size of the change.\n\nCheers,\n-- \nVisit Openswan at http://www.openswan.org/\nEmail: Herbert Xu ~{PmV>HI~} <herbert@gondor.apana.org.au>\nHome Page: http://gondor.apana.org.au/~herbert/\nPGP Key: http://gondor.apana.org.au/~herbert/pubkey.txt\n"},{"id":"587","messageId":"20050418005529.GF1461@pasky.ji.cz","threadId":"9","inReplyTo":"20050418004906.GA3132@gondor.apana.org.au","subject":"Re: Merge with git-pasky II.","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-18T00:55:29Z","receivedAt":"2005-04-18T00:55:29Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Mon, Apr 18, 2005 at 02:49:06AM CEST, I got a letter\nwhere Herbert Xu <herbert@gondor.apana.org.au> told me that...\n> Therefore the only conclusion I can draw is that we're only calling\n> update-cache on the set of changed files, or at most a small superset\n> of them.  In that case, the cost of the collision check *is* proportional\n> to the size of the change.\n\nYes, of course, sorry for the confusion.  We only consider files you\neither specify manually or which have their stat metadata changed\nrelative to the directory cache. (That is from the git-pasky\nperspective; from the plumbing perspective, the user just does\nupdate-cache on whatever he picks.)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"622","messageId":"E1DNNgy-0003He-00@skye.ra.phy.cam.ac.uk","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504171014430.7211@ppc970.osdl.org","subject":"Re: Merge with git-pasky II.","fromName":"Sanjoy Mahajan","fromEmail":"sanjoy@mrao.cam.ac.uk","sentAt":"2005-04-18T04:16:59Z","receivedAt":"2005-04-18T04:16:59Z","isPatch":false,"sender":{"key":"sanjoy@mrao.cam.ac.uk","avatar":null},"body":"> So until proven otherwise, I worry about accidental hashes, and in\n> 160 bits of good hashing, that just isn't an issue either...[Going\n> from 128 bits to 160 bits made it] so _unbelievably_ less likely to\n> happen that it's not even funny.\n\nYou are right.  Here's how I learnt to stop worrying and love the 160\nbits.\n\nA 160-bit hash requires 2^80=10^24 files before the collision\nprobability is roughly 0.5 (actually 1-e^{-1/2}).  Now be very\nconservative: Instead of tolerating a 0.5 probability, worry about\neven a 10^-8 probability of a collision anywhere, anytime.\n\nThe magic number of files for that probability is 10^20 (roughly 10^40\npairs for 2^160=10^48 boxes).\n\nGiven 10 billion people using git, each producing 1 source file per\nsecond -- busy beavers all -- they would need 300 years to produce\n10^20 files.  And to reach the 10^-8 collision probability, all 10^20\nfiles must belong to the same project, and even OpenOffice will not be\nthat bloated.\n\n-Sanjoy\n"},{"id":"633","messageId":"20050418074232.GA20119@elte.hu","threadId":"9","inReplyTo":"Pine.LNX.4.58.0504171014430.7211@ppc970.osdl.org","subject":"Re: Re: Merge with git-pasky II.","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-18T07:42:32Z","receivedAt":"2005-04-18T07:42:32Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* Linus Torvalds <torvalds@osdl.org> wrote:\n\n> On Sun, 17 Apr 2005, Ingo Molnar wrote:\n> > \n> > in fact, this attack cannot even be proven to be malicious, purely via \n> > the email from Malice: it could be incredible bad luck that caused that \n> > good-looking patch to be mistakenly matching a dangerous object.\n> \n> I really hate theoretical discussions.\n\ni was only replying to your earlier point:\n\n> > > Almost all attacks on sha1 will depend on _replacing_ a file with \n> > > a bogus new one. So guys, instead of using sha256 or going \n> > > overboard, just make sure that when you synchronize, you NEVER \n> > > import a file you already have.\n\nwhich point i still believe is subtly wrong. You were suggesting to \nconcentrate on file replacement to counter most of the practical \nattacks, while i pointed out an attack _using the same basic mechanism \nthat your point above supposed_.\n\n[ if you can replace a file with a known hash, with a bogus new one, and \n  you still have enough control over the contents of your bogus new file \n  that it is 1) a valid file that builds 2) compromises the kernel, then \n  you likely have the same amount of control my 'theoretical' attack\n  requires. ]\n\n> And the thing is, _if_ somebody finds a way to make sha1 act as just a \n> complex parity bit, and comes up with generating a clashing object \n> that actually makes sense, then going to sha256 is likely pointless \n> too [...]\n\nyes, that's why i suggested to not actually trust the hash to be \ncryptographically secure, but to just assume it's a good generic hash we \ncan design a DB around, and to turn -DCOLLISION_CHECK on and enforce \nconsistency rules on boundaries.\n\n[ it's not bad to keep sha1 because even my suggested enhancement still\n  leaves 'content-less trust-pointers to untrusted content via email'\n  vectors open against attack (maintainer sends you an email that commit\n  X in Malice's repository Y is fine to pull, and you pull it blindly,\n  while the attacker has replaced his content with the compromised one\n  meanwhile), but it at least validates the bulk traffic that goes into\n  the DB: patches via emails and trusted repositories. ]\n\nso all i was suggesting was to extend your suggested 'overwrite \ncollision check' to a stricter 'content we throw away and use the sha1 \nshortcut for needs to be checked against the in-DB content as well'.\n\nin other words, your suggested 'rename check' is checking for 'positive \nduplicate content', while my addition would also check for 'negative \nduplicate content' as well.\n\nbut as usual, i could be wrong, so dont take this too serious :-)\n\n\tIngo\n"}]}