{"thread":{"id":"13345","subject":"detecting rename->commit->modify->commit","startedAt":"2008-05-01T14:10:24Z","lastAt":"2008-05-08T18:17:24Z","messageCount":49,"participants":["Ittay Dror","Jeff King","Avery Pennarun","David Tweed","Jakub Narebski","Sitaram Chamarty","Steven Grimm","Teemu Likonen","Junio C Hamano","Robin Rosenberg","Linus Torvalds","Shawn O. Pearce","Theodore Tso"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"75755","messageId":"4819CF50.2020509@tikalk.com","threadId":"13345","inReplyTo":null,"subject":"detecting rename->commit->modify->commit","fromName":"Ittay Dror","fromEmail":"ittayd@tikalk.com","sentAt":"2008-05-01T14:10:24Z","receivedAt":"2008-05-01T14:10:24Z","isPatch":false,"sender":{"key":"ittayd@tikalk.com","avatar":null},"body":"Hi,\n\nSay I have a file A, I rename to 'B', commit, then change file B and \ncommit. Does 'git diff -M HEAD^^..' detect that? From what I see now, it \nwill show 'B' as new (all of it with '+' prefix in the output). Am I right?\n\nThank you,\nIttay\n\n-- \nIttay Dror <ittayd@tikalk.com>\nTikal <http://www.tikalk.com>\nTikal Project <http://tikal.sourceforge.net>\n"},{"id":"75756","messageId":"20080501144524.GA10876@sigill.intra.peff.net","threadId":"13345","inReplyTo":"4819CF50.2020509@tikalk.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-01T14:45:24Z","receivedAt":"2008-05-01T14:45:24Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 01, 2008 at 05:10:24PM +0300, Ittay Dror wrote:\n\n> Say I have a file A, I rename to 'B', commit, then change file B and  \n> commit. Does 'git diff -M HEAD^^..' detect that? From what I see now, it  \n> will show 'B' as new (all of it with '+' prefix in the output). Am I \n> right?\n\nYes, it should find it, assuming the changes to B leave it recognizable.\nTry:\n\n  mkdir repo && cd repo && git init\n  cp /usr/share/dict/words A\n  git add . && git commit -m added\n  mv A B && git add B && git commit -a -m rename\n  echo change >>B && git commit -a -m change\n  git diff -M HEAD^^.. | head -n 7\n\nYou should see something like:\n\n  diff --git a/A b/B\n  similarity index 99%\n  rename from A\n  rename to B\n  index 8e50f11..6525618 100644\n  --- a/A\n  +++ b/B\n\nHowever, note the similarity index. If you change B so much that it\ndoesn't look close to the original A, then the rename is not detected\n(and intentionally so -- the argument is that it is no longer a rename\nin that context, but a rewritten file).\n\n-Peff\n"},{"id":"75757","messageId":"4819D98E.1040004@tikalk.com","threadId":"13345","inReplyTo":"4819CF50.2020509@tikalk.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Ittay Dror","fromEmail":"ittayd@tikalk.com","sentAt":"2008-05-01T14:54:06Z","receivedAt":"2008-05-01T14:54:06Z","isPatch":false,"sender":{"key":"ittayd@tikalk.com","avatar":null},"body":"Also, would anyone like to comment on: \nhttp://www.markshuttleworth.com/archives/123 (Renaming is the killer app \nof distributed version control \n<http://www.markshuttleworth.com/archives/123>)?\n\nThank you,\nIttay\n\nIttay Dror wrote:\n> Hi,\n>\n> Say I have a file A, I rename to 'B', commit, then change file B and \n> commit. Does 'git diff -M HEAD^^..' detect that? From what I see now, \n> it will show 'B' as new (all of it with '+' prefix in the output). Am \n> I right?\n>\n> Thank you,\n> Ittay\n>\n\n-- \nIttay Dror <ittayd@tikalk.com>\nTikal <http://www.tikalk.com>\nTikal Project <http://tikal.sourceforge.net>\n"},{"id":"75758","messageId":"4819DCF1.7090504@tikalk.com","threadId":"13345","inReplyTo":"20080501144524.GA10876@sigill.intra.peff.net","subject":"Re: detecting rename->commit->modify->commit","fromName":"Ittay Dror","fromEmail":"ittayd@tikalk.com","sentAt":"2008-05-01T15:08:33Z","receivedAt":"2008-05-01T15:08:33Z","isPatch":false,"sender":{"key":"ittayd@tikalk.com","avatar":null},"body":"But it doesn't work across directories :-(.\n\nTry:\n >mkdir foo\n >echo \"hello\" > foo/A\n >git add foo/A\n >git commit -m 'foo/A'\n >mkdir bar\n >git mv foo/A bar\n >git commit -m 'bar/A'\n >echo \"world\" >> bar/A\n >git add bar/A\n >git commit -m 'bar/A world'\n >git diff HEAD^^..HEAD^ | cat\ndiff --git a/foo/A b/bar/A\nsimilarity index 100%\nrename from foo/A\nrename to bar/A\n > git diff HEAD^^.. | cat\ndiff --git a/bar/A b/bar/A\nnew file mode 100644\nindex 0000000..94954ab\n--- /dev/null\n+++ b/bar/A\n@@ -0,0 +1,2 @@\n+hello\n+world\ndiff --git a/foo/A b/foo/A\ndeleted file mode 100644\nindex ce01362..0000000\n--- a/foo/A\n+++ /dev/null\n@@ -1 +0,0 @@\n-hello\n\n\n\n\n\nJeff King wrote:\n> On Thu, May 01, 2008 at 05:10:24PM +0300, Ittay Dror wrote:\n>\n>   \n>> Say I have a file A, I rename to 'B', commit, then change file B and  \n>> commit. Does 'git diff -M HEAD^^..' detect that? From what I see now, it  \n>> will show 'B' as new (all of it with '+' prefix in the output). Am I \n>> right?\n>>     \n>\n> Yes, it should find it, assuming the changes to B leave it recognizable.\n> Try:\n>\n>   mkdir repo && cd repo && git init\n>   cp /usr/share/dict/words A\n>   git add . && git commit -m added\n>   mv A B && git add B && git commit -a -m rename\n>   echo change >>B && git commit -a -m change\n>   git diff -M HEAD^^.. | head -n 7\n>\n> You should see something like:\n>\n>   diff --git a/A b/B\n>   similarity index 99%\n>   rename from A\n>   rename to B\n>   index 8e50f11..6525618 100644\n>   --- a/A\n>   +++ b/B\n>\n> However, note the similarity index. If you change B so much that it\n> doesn't look close to the original A, then the rename is not detected\n> (and intentionally so -- the argument is that it is no longer a rename\n> in that context, but a rewritten file).\n>\n> -Peff\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n>   \n\n-- \nIttay Dror <ittayd@tikalk.com>\nTikal <http://www.tikalk.com>\nTikal Project <http://tikal.sourceforge.net>\n"},{"id":"75759","messageId":"20080501150958.GA11145@sigill.intra.peff.net","threadId":"13345","inReplyTo":"4819D98E.1040004@tikalk.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-01T15:09:58Z","receivedAt":"2008-05-01T15:09:58Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 01, 2008 at 05:54:06PM +0300, Ittay Dror wrote:\n\n> Also, would anyone like to comment on:  \n> http://www.markshuttleworth.com/archives/123 (Renaming is the killer app  \n> of distributed version control  \n> <http://www.markshuttleworth.com/archives/123>)?\n\nMy two cents:\n\n1. I think he is overly obsessed with renaming. He seems concerned that\nsomebody will show up, make a big renaming patch, and then break your\nsystem. Guess what? They can also show up, make a big code change patch,\nand then break your system. In either case you have to review the\nchanges before accepting them, and it is up to the version control\nsystem to show you the changes in a way you can understand.\n\n2. I see the same old \"git developers decided renaming wasn't\nimportant\" argument. I think this is bogus. I think renaming _is_\nimportant, but I actually prefer git's approach of deducing renames,\nbecause it reflects a fundamental property of git: we track states, not\nchanges, and git doesn't care how you arrive at each state. So I am free\nto use a combination of git commands, editors, patch application tools,\nor anything else to get my tree to the right place.\n\n3. He doesn't like that git doesn't track _directory_ renames. This is\nnot a fundamental problem with git's approach (which could deduce\ndirectory renames after the fact), but rather comes from the fact that\ndirectory renames are controversial. That is, even if you know (through\ndeduction or because an explicit rename was recorded) that \"subdir1\"\nmoved to \"subdir2\", that doesn't necessarily mean that new files added\ninto \"subdir1\" should make that move, as well.\n\n-Peff\n"},{"id":"75760","messageId":"4819DFB0.5030401@tikalk.com","threadId":"13345","inReplyTo":"20080501150958.GA11145@sigill.intra.peff.net","subject":"Re: detecting rename->commit->modify->commit","fromName":"Ittay Dror","fromEmail":"ittayd@tikalk.com","sentAt":"2008-05-01T15:20:16Z","receivedAt":"2008-05-01T15:20:16Z","isPatch":false,"sender":{"key":"ittayd@tikalk.com","avatar":null},"body":"\nJeff King wrote:\n> My two cents:\n>\n> 1. I think he is overly obsessed with renaming. He seems concerned that\n> somebody will show up, make a big renaming patch, and then break your\n> system. Guess what? They can also show up, make a big code change patch,\n> and then break your system. In either case you have to review the\n> changes before accepting them, and it is up to the version control\n> system to show you the changes in a way you can understand\nI think he was more concerned that merges will break after such a change.\n\nIttay\n\n-- \nIttay Dror <ittayd@tikalk.com>\nTikal <http://www.tikalk.com>\nTikal Project <http://tikal.sourceforge.net>\n"},{"id":"75761","messageId":"20080501152035.GB11145@sigill.intra.peff.net","threadId":"13345","inReplyTo":"4819DCF1.7090504@tikalk.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-01T15:20:35Z","receivedAt":"2008-05-01T15:20:35Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 01, 2008 at 06:08:33PM +0300, Ittay Dror wrote:\n\n> But it doesn't work across directories :-(.\n\nYes, it does.\n\n> Try:\n> >mkdir foo\n> >echo \"hello\" > foo/A\n> >git add foo/A\n> >git commit -m 'foo/A'\n> >mkdir bar\n> >git mv foo/A bar\n> >git commit -m 'bar/A'\n> >echo \"world\" >> bar/A\n> >git add bar/A\n> >git commit -m 'bar/A world'\n> >git diff HEAD^^..HEAD^ | cat\n> diff --git a/foo/A b/bar/A\n> similarity index 100%\n> rename from foo/A\n> rename to bar/A\n\nSee, it just worked across directories.\n\n> > git diff HEAD^^.. | cat\n> diff --git a/bar/A b/bar/A\n> new file mode 100644\n> index 0000000..94954ab\n> --- /dev/null\n> +++ b/bar/A\n> @@ -0,0 +1,2 @@\n> +hello\n> +world\n> diff --git a/foo/A b/foo/A\n> deleted file mode 100644\n> index ce01362..0000000\n> --- a/foo/A\n> +++ /dev/null\n> @@ -1 +0,0 @@\n> -hello\n\nOf course it doesn't work here. You have two files, one containing\n\"hello\\n\" and one containing \"hello\\nworld\\n\". Their similarity is 50%,\nwhich is not enough to consider it a rename. And I would argue that's\nreasonable, since the files have only one line in common. The problem is\nthat you are using a toy example (which is why my example used\n/usr/share/dict/words, which has enough content to definitively call it\na rename).\n\n...\n\nHmm, looking at the code, though, 50% is supposed to be the default\nminimum. So there might actually be a bug.\n\n-Peff\n"},{"id":"75762","messageId":"4819E0AE.40602@tikalk.com","threadId":"13345","inReplyTo":"4819DCF1.7090504@tikalk.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Ittay Dror","fromEmail":"ittayd@tikalk.com","sentAt":"2008-05-01T15:24:30Z","receivedAt":"2008-05-01T15:24:30Z","isPatch":false,"sender":{"key":"ittayd@tikalk.com","avatar":null},"body":"Btw, this happened to me in a real use case. I wanted to restructure a \nsource tree. So I put it under git and started to happily move things \naround, always committing after a move. I thought that git will \ncorrectly identify these moves and show me the differences I made after \n(in a separate commit). But it doesn't, and now that I want to prepare a \nsummary of the changes I've made, I'm stuck with a huge diff that is \nhard to make sense of.\n\nIttay\n\nIttay Dror wrote:\n> But it doesn't work across directories :-(.\n>\n> Try:\n> >mkdir foo\n> >echo \"hello\" > foo/A\n> >git add foo/A\n> >git commit -m 'foo/A'\n> >mkdir bar\n> >git mv foo/A bar\n> >git commit -m 'bar/A'\n> >echo \"world\" >> bar/A\n> >git add bar/A\n> >git commit -m 'bar/A world'\n> >git diff HEAD^^..HEAD^ | cat\n> diff --git a/foo/A b/bar/A\n> similarity index 100%\n> rename from foo/A\n> rename to bar/A\n> > git diff HEAD^^.. | cat\n> diff --git a/bar/A b/bar/A\n> new file mode 100644\n> index 0000000..94954ab\n> --- /dev/null\n> +++ b/bar/A\n> @@ -0,0 +1,2 @@\n> +hello\n> +world\n> diff --git a/foo/A b/foo/A\n> deleted file mode 100644\n> index ce01362..0000000\n> --- a/foo/A\n> +++ /dev/null\n> @@ -1 +0,0 @@\n> -hello\n>\n>\n>\n>\n>\n> Jeff King wrote:\n>> On Thu, May 01, 2008 at 05:10:24PM +0300, Ittay Dror wrote:\n>>\n>>  \n>>> Say I have a file A, I rename to 'B', commit, then change file B \n>>> and  commit. Does 'git diff -M HEAD^^..' detect that? From what I \n>>> see now, it  will show 'B' as new (all of it with '+' prefix in the \n>>> output). Am I right?\n>>>     \n>>\n>> Yes, it should find it, assuming the changes to B leave it recognizable.\n>> Try:\n>>\n>>   mkdir repo && cd repo && git init\n>>   cp /usr/share/dict/words A\n>>   git add . && git commit -m added\n>>   mv A B && git add B && git commit -a -m rename\n>>   echo change >>B && git commit -a -m change\n>>   git diff -M HEAD^^.. | head -n 7\n>>\n>> You should see something like:\n>>\n>>   diff --git a/A b/B\n>>   similarity index 99%\n>>   rename from A\n>>   rename to B\n>>   index 8e50f11..6525618 100644\n>>   --- a/A\n>>   +++ b/B\n>>\n>> However, note the similarity index. If you change B so much that it\n>> doesn't look close to the original A, then the rename is not detected\n>> (and intentionally so -- the argument is that it is no longer a rename\n>> in that context, but a rewritten file).\n>>\n>> -Peff\n>> -- \n>> To unsubscribe from this list: send the line \"unsubscribe git\" in\n>> the body of a message to majordomo@vger.kernel.org\n>> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>>\n>>   \n>\n\n-- \nIttay Dror <ittayd@tikalk.com>\nTikal <http://www.tikalk.com>\nTikal Project <http://tikal.sourceforge.net>\n"},{"id":"75763","messageId":"32541b130805010827r22169651s37c707071f3448f2@mail.gmail.com","threadId":"13345","inReplyTo":"4819D98E.1040004@tikalk.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-05-01T15:27:34Z","receivedAt":"2008-05-01T15:27:34Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On 5/1/08, Ittay Dror <ittayd@tikalk.com> wrote:\n> Also, would anyone like to comment on:\n> http://www.markshuttleworth.com/archives/123 (Renaming is\n> the killer app of distributed version control\n> <http://www.markshuttleworth.com/archives/123>)?\n\nOne of the comments linked to this:\nhttp://automatthias.wordpress.com/2007/06/07/directory-renaming-in-scm/\n\nWhich points out that git doesn't really handle directory renames at\nall.  If someone creates file A/X then renames A to B, then merges\nwith someone who both added the file A/Y and modified A/X, git will\nproduce a tree containing (modified) B/Y and (new) A/Y.\n\nTechnically this is \"correct\" in that no data is lost and there are no\nconflicts, but it is obviously not what was \"intended\", which was that\nthe new file Y should have ended up in folder B.\n\nBefore you say this is not a realistic use case, I've personally had\nthis exact problem:\n\n- I had a project with all of my work in a folder \"src\"\n- I decided that the 'src' folder was redundant, so I moved it all to\nthe root folder\n- Someone else was working on an old maintenance branch which still had 'src'\n- When I merged from that person, some new files were created under\n'src', and of course didn't work.\n\nSince the maintenance branch was long-lived, this problem happened\nrepeatedly.  That said, it's also pretty easy to work around, so it's\nnot the end of the world.\n\nHave fun,\n\nAvery\n"},{"id":"75764","messageId":"20080501152859.GA11469@sigill.intra.peff.net","threadId":"13345","inReplyTo":"4819E0AE.40602@tikalk.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-01T15:28:59Z","receivedAt":"2008-05-01T15:28:59Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 01, 2008 at 06:24:30PM +0300, Ittay Dror wrote:\n\n> Btw, this happened to me in a real use case. I wanted to restructure a  \n> source tree. So I put it under git and started to happily move things  \n> around, always committing after a move. I thought that git will correctly \n> identify these moves and show me the differences I made after (in a \n> separate commit). But it doesn't, and now that I want to prepare a  \n> summary of the changes I've made, I'm stuck with a huge diff that is hard \n> to make sense of.\n\nIf you have a specific case where you think renames should have been\ndetected but they weren't, by all means, please share it. It's possible\nthat there is a bug in the rename detection, or that the limits are not\nset correctly, and we could improve it.\n\n-Peff\n"},{"id":"75765","messageId":"e1dab3980805010830l49e7d126pa4de831c174eda0c@mail.gmail.com","threadId":"13345","inReplyTo":"20080501150958.GA11145@sigill.intra.peff.net","subject":"Re: detecting rename->commit->modify->commit","fromName":"David Tweed","fromEmail":"david.tweed@gmail.com","sentAt":"2008-05-01T15:30:24Z","receivedAt":"2008-05-01T15:30:24Z","isPatch":false,"sender":{"key":"david.tweed@gmail.com","avatar":null},"body":"On Thu, May 1, 2008 at 4:09 PM, Jeff King <peff@peff.net> wrote:\n> On Thu, May 01, 2008 at 05:54:06PM +0300, Ittay Dror wrote:\n>\n>  > Also, would anyone like to comment on:\n>  > http://www.markshuttleworth.com/archives/123 (Renaming is the killer app\n>  > of distributed version control\n>  > <http://www.markshuttleworth.com/archives/123>)?\n\nI'll just make the obvious point that he's talking about a problem and\nan underlying cause:\n\nThe problem is not being able to successfully merge branches as time\ngoes by when one branch has had some renaming. He's decided the root\ncause is not have an explicit representation of renames which would\nenable the merges to succeed. So there are two questions:\n\n1. Does development often happen where files get renamed and then\nmodified significantly in a distributed fashion but it is still\nsensible to automatically merge the results?\n\n2. Do you need explicit rename tracking to do an automatic merge in those cases?\n\nI suspect that for 2 you don't in theory but considering all the\nnon-obvious possibilities would slow down the normal case of a\nstandard merge.\n\n-- \ncheers, dave tweed__________________________\ndavid.tweed@gmail.com\nRm 124, School of Systems Engineering, University of Reading.\n\"while having code so boring anyone can maintain it, use Python.\" --\nattempted insult seen on slashdot\n"},{"id":"75766","messageId":"4819E226.6000404@tikalk.com","threadId":"13345","inReplyTo":"20080501152035.GB11145@sigill.intra.peff.net","subject":"Re: detecting rename->commit->modify->commit","fromName":"Ittay Dror","fromEmail":"ittayd@tikalk.com","sentAt":"2008-05-01T15:30:46Z","receivedAt":"2008-05-01T15:30:46Z","isPatch":false,"sender":{"key":"ittayd@tikalk.com","avatar":null},"body":"Jeff King wrote:\n> Of course it doesn't work here. You have two files, one containing\n> \"hello\\n\" and one containing \"hello\\nworld\\n\". Their similarity is 50%,\n> which is not enough to consider it a rename. And I would argue that's\n> reasonable, since the files have only one line in common. The problem is\n> that you are using a toy example (which is why my example used\n> /usr/share/dict/words, which has enough content to definitively call it\n> a rename).\n>\n>   \nWell, I would have expected git to notice that the file was renamed in \none commit and keep tracking changes afterwards.\n\nAlso, as I wrote in another post, this happened to me with real files of \na real source tree, and with very small changes (and sometimes not at \nall) to these files.\n\nIttay\n\n-- \nIttay Dror <ittayd@tikalk.com>\nTikal <http://www.tikalk.com>\nTikal Project <http://tikal.sourceforge.net>\n"},{"id":"75767","messageId":"20080501153457.GB11469@sigill.intra.peff.net","threadId":"13345","inReplyTo":"32541b130805010827r22169651s37c707071f3448f2@mail.gmail.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-01T15:34:57Z","receivedAt":"2008-05-01T15:34:57Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 01, 2008 at 11:27:34AM -0400, Avery Pennarun wrote:\n\n> Before you say this is not a realistic use case, I've personally had\n> this exact problem:\n> \n> - I had a project with all of my work in a folder \"src\"\n> - I decided that the 'src' folder was redundant, so I moved it all to\n> the root folder\n> - Someone else was working on an old maintenance branch which still had 'src'\n> - When I merged from that person, some new files were created under\n> 'src', and of course didn't work.\n\nSure. But we've also had the exact case of:\n\n  - there are some files in subdir/, but that is not a good name, and\n    there is something else that you are going to add that would be\n    better named as subdir/.\n  - you rename subdir/ to bettername/\n  - you create subdir/newfile\n\nbut you _don't_ want newfile to go into bettername/. It's _replacing_\nwhat went into bettername/.\n\nSo I don't think you can always track the intent automatically.\n\nThough if you could specify the intent to the SCM, you could\ndifferentiate at the time of move between these two cases, and the merge\ncould do the right thing later. Or alternatively, you could specify at\ntime of merge which to do.  It's just that nobody has implemented it.\n\n-Peff\n"},{"id":"75768","messageId":"20080501153804.GC11469@sigill.intra.peff.net","threadId":"13345","inReplyTo":"4819E226.6000404@tikalk.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-01T15:38:05Z","receivedAt":"2008-05-01T15:38:05Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 01, 2008 at 06:30:46PM +0300, Ittay Dror wrote:\n\n> Well, I would have expected git to notice that the file was renamed in  \n> one commit and keep tracking changes afterwards.\n\nThat's not how git works, and that's not what you asked it to do. You\ngave it two states and asked it to diff between them. It never even\nlooked at the intermediate steps (and that's generally why git is so\nfast). If you want to follow the history and look at every commit, then\nthat is something that _can_ be done, and does get done with things like\n\"git log --follow\". But there is a diff mode currently implemented that\nwill crawl the history looking for interesting things.\n\n-Peff\n"},{"id":"75769","messageId":"m3hcdi6n7r.fsf@localhost.localdomain","threadId":"13345","inReplyTo":"4819E226.6000404@tikalk.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-05-01T15:47:33Z","receivedAt":"2008-05-01T15:47:33Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Ittay Dror <ittayd@tikalk.com> writes:\n\n> Jeff King wrote:\n> >\n> > Of course it doesn't work here. You have two files, one containing\n> > \"hello\\n\" and one containing \"hello\\nworld\\n\". Their similarity is 50%,\n> > which is not enough to consider it a rename. And I would argue that's\n> > reasonable, since the files have only one line in common. The problem is\n> > that you are using a toy example (which is why my example used\n> > /usr/share/dict/words, which has enough content to definitively call it\n> > a rename).\n> >\n> >\n> Well, I would have expected git to notice that the file was renamed in\n> one commit and keep tracking changes afterwards.\n> \n> Also, as I wrote in another post, this happened to me with real files\n> of a real source tree, and with very small changes (and sometimes not\n> at all) to these files.\n\nThe idea of rename detection is to help with merges.  If the files are\ndifferent enough that content based (similarity based) rename\ndetection doesn't detect rename, they are usually too different to\nmerge automatically anyway.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"75770","messageId":"32541b130805010850q165fe1d6me05e670ca93b0892@mail.gmail.com","threadId":"13345","inReplyTo":"20080501153457.GB11469@sigill.intra.peff.net","subject":"Re: detecting rename->commit->modify->commit","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-05-01T15:50:31Z","receivedAt":"2008-05-01T15:50:31Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On 5/1/08, Jeff King <peff@peff.net> wrote:\n> On Thu, May 01, 2008 at 11:27:34AM -0400, Avery Pennarun wrote:\n>\n>  > Before you say this is not a realistic use case, I've personally had\n>  > this exact problem:\n>  >\n>  > - I had a project with all of my work in a folder \"src\"\n>  > - I decided that the 'src' folder was redundant, so I moved it all to\n>  > the root folder\n>  > - Someone else was working on an old maintenance branch which still had 'src'\n>  > - When I merged from that person, some new files were created under\n>  > 'src', and of course didn't work.\n>\n>\n> Sure. But we've also had the exact case of:\n>\n>   - there are some files in subdir/ [1], but that is not a good name, and\n>     there is something else that you are going to add that would be\n>     better named as subdir/.\n>   - you rename subdir/ to bettername/ [2]\n>   - you create subdir/newfile [3]\n>\n>  but you _don't_ want newfile to go into bettername/. It's _replacing_\n>  what went into bettername/.\n\nI would argue that this is a sort of \"directory splitting\" operation.\nThat is, all anyone ever did was add some files to a subdir/ that\nalready existed [1], *or* move all the files from subdir/ to a\npreviously-empty bettername/ [2], *or* create a new subdir/ and add\nfiles to it [3]. In each case, no merge operation was necessary and it\nis completely obvious by comparing \"before and after\" trees which case\nit was.\n\nI guess my argument here is just that it should be *possible* to\ndeduce and implement both cases at merge time just fine using git's\nexisting storage model.  It just hasn't been implemented yet.  (And\nincidentally, I think that's totally awesome and I'd never want to go\nback to an explicit rename tracking model.)\n\nI should shut up now because the actual merge machinery scares me and\nI'm not willing to volunteer to write a patch for this one :)\n\nHave fun,\n\nAvery\n"},{"id":"75771","messageId":"2e24e5b90805010939g182de387i59722605ff93d72e@mail.gmail.com","threadId":"13345","inReplyTo":"4819D98E.1040004@tikalk.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2008-05-01T16:39:46Z","receivedAt":"2008-05-01T16:39:46Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"On Thu, May 1, 2008 at 8:24 PM, Ittay Dror <ittayd@tikalk.com> wrote:\n> Also, would anyone like to comment on:\n> http://www.markshuttleworth.com/archives/123 (Renaming is the killer app of\n> distributed version control <http://www.markshuttleworth.com/archives/123>)?\n\nsomeone already did, albeit in just discussion form rather than\nexamples, in a comment on that same page:\n\nhttp://www.markshuttleworth.com/archives/123#comment-118655\n"},{"id":"75772","messageId":"20080501164829.GA11636@sigill.intra.peff.net","threadId":"13345","inReplyTo":"32541b130805010850q165fe1d6me05e670ca93b0892@mail.gmail.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-01T16:48:29Z","receivedAt":"2008-05-01T16:48:29Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 01, 2008 at 11:50:31AM -0400, Avery Pennarun wrote:\n\n> I would argue that this is a sort of \"directory splitting\" operation.\n> That is, all anyone ever did was add some files to a subdir/ that\n> already existed [1], *or* move all the files from subdir/ to a\n> previously-empty bettername/ [2], *or* create a new subdir/ and add\n> files to it [3]. In each case, no merge operation was necessary and it\n> is completely obvious by comparing \"before and after\" trees which case\n> it was.\n\nI don't see it. I think the steps are exactly the same as in your\nexample. Consider:\n\n  1. You have some files in src/\n  2. All of the files from src/ get moved away\n  3. You merge in somebody else's work which adds a file in src/, but\n     their work is based on a commit which predates 2.\n\nThe question is: if they had seen 2., would they have put the file into\nsrc/, or into the new location? I think the answer depends on the\nsemantics of the file. If it is semantically an addition to the source\ncode that got moved, then yes. If it is a _replacement_ for the\nsource code that got moved, then no.\n\n> I guess my argument here is just that it should be *possible* to\n> deduce and implement both cases at merge time just fine using git's\n> existing storage model.  It just hasn't been implemented yet.  (And\n> incidentally, I think that's totally awesome and I'd never want to go\n> back to an explicit rename tracking model.)\n\nI think you lack information to decide automatically between the two\ncases listed above. But I think in most cases it would be sufficient for\nthe tool to say \"this directory seems to have moved, but this new file\nwas added in it\" and let the user decide which makes sense.\n\n> I should shut up now because the actual merge machinery scares me and\n> I'm not willing to volunteer to write a patch for this one :)\n\nIt would probably start not with merge machinery, but with diff\nmachinery to detect \"directory has moved\". But that is also scary. :)\n\nYou could also do this totally _outside_ of git, similar to\ngit-mergetool. Wait until you get a conflict, and then run a script\nwhich looks at the two endpoints and the merge base and says \"Oh, maybe\nthis is a good way of resolving.\"\n\n-Peff\n"},{"id":"75773","messageId":"481A12C6.6060900@tikalk.com","threadId":"13345","inReplyTo":"2e24e5b90805010939g182de387i59722605ff93d72e@mail.gmail.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Ittay Dror","fromEmail":"ittayd@tikalk.com","sentAt":"2008-05-01T18:58:14Z","receivedAt":"2008-05-01T18:58:14Z","isPatch":false,"sender":{"key":"ittayd@tikalk.com","avatar":null},"body":"Sitaram Chamarty wrote:\n> http://www.markshuttleworth.com/archives/123#comment-118655\n>\nHere is the comment from the thread, my comment on it is below:\n\n > This is a very strong point for renaming, but it is not necessarily \nan universal one.\n\n > Here is one example of the issue: one developer renaming a directory \nin his branch, and another adding a file to the original directory in \nhis branch. What happens at the merge ?\n > - Bazaar renames the directory and puts the new file in the _renamed_ \ndirectory.\n > - Git renames the directory with its files, but keeps the old \ndirectory too and adds the new file there.\n\n > Bazaar’s behavior certainly is better for C. However it is not \nuniversally better.\n\n > For example in Java you cannot rename a file without changing its \ncontents. So, moving a file to a directory different from where its \nauthor put it will almost certainly break the build.\n\n > The bottom line is, both behaviors can seem valid or broken, \ndepending on the case. Neither is perfect. At the very abstract level \nfile renames are _not_ a first-class operation. This is especially \napparent in a language like Java.\n\n > Content movement is the first class operation. Things like moving \nfunctions, etc. The question is how one can handle that and whether the \ncurrent strategy has a path for improvement. It could be > argued that \nonce you commit yourself to explicitly tracking file renames, you are \ngiving up a slew of opportunities for handling the more general cases.\n\n > One thing is for certain, a 100% ideal solution is impossible. It \nwould have to be aware of the target programming language _and_ the \nbuild environment.\n\nAnd my comment is that in this example, about Java, I think that \nmanually fixing the package name in the file (after noticing the build \nis broken) is easy. On the other hand, if the other developer changed \none of the renamed file, then manually merging the change in the file in \nthe old location to the file in the new location is not so easy: you \nfirst need to discover that this happened, then merge the two files (and \nyou still need to fix the package name).\n\nittay\n\n-- \nIttay Dror <ittayd@tikalk.com>\nTikal <http://www.tikalk.com>\nTikal Project <http://tikal.sourceforge.net>\n"},{"id":"75777","messageId":"D0968007-2A38-44DB-B26F-3D273F20D428@midwinter.com","threadId":"13345","inReplyTo":"20080501153457.GB11469@sigill.intra.peff.net","subject":"Re: detecting rename->commit->modify->commit","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2008-05-01T19:12:33Z","receivedAt":"2008-05-01T19:12:33Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"On May 1, 2008, at 8:34 AM, Jeff King wrote:\n> So I don't think you can always track the intent automatically.\n\nThat is absolutely true. You have to pick one case or the other as the  \ndefault unless there's some way to tell the system your intent either  \nat merge time or at move time.\n\nHowever, that leaves the question of which default will be wrong the  \nleast often.\n\nIn my personal experience, I think a directory rename has almost  \nalways meant that I would want new files to appear in the new  \ndirectory rather than to recreate the old directory. I can't think of  \na single time when I've wanted git's current behavior (though maybe  \nit's happened on occasion) but the current behavior has tripped me up  \nmore than once and forced me to do extra work shuffling things around  \nby hand post-merge. I acknowledge that there exist cases where the  \ncurrent behavior is correct -- but in my experience they're the  \nminority.\n\nOf course, the discussion is moot anyway until someone writes code to  \ndetect the situation; my impression is the current behavior is the way  \nit is simply because it's what naturally happens in the absence of  \nmerge-time detection of a directory getting renamed.\n\n-Steve\n"},{"id":"75780","messageId":"32541b130805011245j76421635me55947cf7869f31f@mail.gmail.com","threadId":"13345","inReplyTo":"20080501164829.GA11636@sigill.intra.peff.net","subject":"Re: detecting rename->commit->modify->commit","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-05-01T19:45:07Z","receivedAt":"2008-05-01T19:45:07Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Thu, May 1, 2008 at 12:48 PM, Jeff King <peff@peff.net> wrote:\n>  I don't see it. I think the steps are exactly the same as in your\n>  example. Consider:\n>\n>   1. You have some files in src/\n>   2. All of the files from src/ get moved away\n>   3. You merge in somebody else's work which adds a file in src/, but\n>      their work is based on a commit which predates 2.\n>\n>  The question is: if they had seen 2., would they have put the file into\n>  src/, or into the new location? I think the answer depends on the\n>  semantics of the file. If it is semantically an addition to the source\n>  code that got moved, then yes. If it is a _replacement_ for the\n>  source code that got moved, then no.\n\nI promised I would shut up, and I apparently didn't.  Sorry :)\n\nI think this case isn't so hard.  Basically, a merge involves three\ncommits; the merge-base, my branch, and your branch.\n\nIn your example above, we compare the merge-base to the new version;\nin that case, the new file is in an *existing* directory which\ndefinitely corresponds to src/ in #1, because the the new version has\nnever even heard about src/ being deleted.  Thus, the file must be\nintended to be part of the original src/, wherever it may now be.\n\nIn contrast, if the merge-base already had src/ being renamed, and\nsomeone put something into src/, we'd know that they're putting it\ninto a fundamentally different directory than the moved src/.\n\nExactly how you track the \"identity\" of a directory without breaking\nthings down by individual commit sounds a little complicated, but it\nfeels to me like it should be possible.\n\nI suspect this is a generalization of the earlier discussion (a few\nmonths ago) that I read in the archive about git's handling of empty\ndirectories.  Right now git does weird things with directory\ncreation/deletion because directories are not first-class citizens.\n\nAnyway, as with the empty directory stuff, if I occasionally have to\nmkdir/rmdir a couple things and rename a few files after doing a\nmerge, I'm not going to cry too much.  It sure beats explicitly\ntracking renames and then having an oops-I-forgot-to-explicitly-track\nrename throw a monkey wrench into my merges, which svn has saddled me\nwith lots of times.\n\nHave fun,\n\nAvery\n"},{"id":"75782","messageId":"20080501203940.GA3524@mithlond.arda.local","threadId":"13345","inReplyTo":"20080501152035.GB11145@sigill.intra.peff.net","subject":"Re: detecting rename->commit->modify->commit","fromName":"Teemu Likonen","fromEmail":"tlikonen@iki.fi","sentAt":"2008-05-01T20:39:40Z","receivedAt":"2008-05-01T20:39:40Z","isPatch":false,"sender":{"key":"tlikonen@iki.fi","avatar":null},"body":"Jeff King wrote (2008-05-01 11:20 -0400):\n\n> Hmm, looking at the code, though, 50% is supposed to be the default\n> minimum. So there might actually be a bug.\n\nI did some testing... A file, containing 10 lines (about 200 bytes),\nrenamed and then modified (similarity index being a bit over 50%). Git\ndetected the rename just fine with \"git diff -M\" over the rename and\nchange. When I edited the file even more (similarity only 40%) \"git diff\n-M\" didn't detect the rename but \"git diff -M4\" did. To me it looks like\nthis works nicely, better than I expected, actually.\n\nSmaller files than that do not seem to work with \"git diff -M\" over the\nrename and changes. They can be followed with \"git log --follow -p\"\nwhich works even with the two-line \"hello\\nworld\". And of course there\nis always\n\n  git diff commit1:path1/file1 commit2:path2/file2\n\nI'd conclude that for logs and diffs renames are detected very nicely\nand there's no problem at all to get wanted information from the repo.\nI wonder how this rename detection/tracking has become such a big thing,\na debate even. But maybe merges are different.\n"},{"id":"75787","messageId":"20080501224215.GB21731@sigill.intra.peff.net","threadId":"13345","inReplyTo":"32541b130805011245j76421635me55947cf7869f31f@mail.gmail.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-01T22:42:16Z","receivedAt":"2008-05-01T22:42:16Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 01, 2008 at 03:45:07PM -0400, Avery Pennarun wrote:\n\n> In your example above, we compare the merge-base to the new version;\n> in that case, the new file is in an *existing* directory which\n> definitely corresponds to src/ in #1, because the the new version has\n> never even heard about src/ being deleted.  Thus, the file must be\n> intended to be part of the original src/, wherever it may now be.\n\nI disagree with the final statement of the quoted paragraph above.\n\nJust because you didn't build on the commit that moved src/* doesn't\nmean the thing you put in src/ was intended to be moved along with src/.\nFor example:\n\n  - it might have been a new work unrelated to the existing work in src/\n    that got moved\n\n  - it might have been a replacement for the work in src/ that was\n    started before the movement. E.g., developer1 begins the replacement\n    work. developer2 moves the old work out of the way. When the\n    branches are merged, you don't want developer1's work moved.\n\nAnd yes, I think those are probably less common than \"it should be moved\nalong with src/*\". My point isn't that this isn't a valuable construct,\nbut that we should stop short of mind-reading, and focus on making it\n_easy_ to see what happened and to concisely specify the choice and\nproceed.\n\n-Peff\n"},{"id":"75788","messageId":"20080501230925.GC21731@sigill.intra.peff.net","threadId":"13345","inReplyTo":"20080501203940.GA3524@mithlond.arda.local","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-01T23:09:25Z","receivedAt":"2008-05-01T23:09:25Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"[cc'd Junio for comments on this rename optimization]\n\nOn Thu, May 01, 2008 at 11:39:40PM +0300, Teemu Likonen wrote:\n\n> > Hmm, looking at the code, though, 50% is supposed to be the default\n> > minimum. So there might actually be a bug.\n> \n> I did some testing... A file, containing 10 lines (about 200 bytes),\n> renamed and then modified (similarity index being a bit over 50%). Git\n\nAh, OK. The problem comes because the toy example is so tiny. It hits\nthis code chunk:\n\n  if (base_size * (MAX_SCORE-minimum_score) < delta_size * MAX_SCORE)\n          return 0;\n\nwhere base_size is the size of the smaller file in bytes, and delta_size\nis the difference between the size of the two files. This is an\noptimization so that we don't even have to look at the contents.\n\nBut it is basing the percentage off of the smaller file, so even though\nfile B (\"hello\\nworld\\n\") is 50% made up of file A (\"hello\\n\"), we\nactually end up saying \"there must be at least as much content added to\nmake B as there is in A already\". IOW, the \"percentage similarity\" is\nbased off of the smaller file for this optimization.\n\nObviously this is a toy case, but I wonder if there are other larger\ncases where you end up with a file which has substantial copied content,\nbut also _grows_ a lot (not just changes). For example, consider the\nfile:\n\n  1\n  2\n  3\n  4\n  5\n  6\n  7\n  8\n  9\n\nthat is, ten lines each with a number. Now rename it, and start adding\nmore numbers. We detect the addition of 10, 11, 12. But adding 13 means\nwe no longer match. So even with only 4 lines added, we fail to match.\n\nBut again, this is a bit of a toy case. It relies on the line length\nbeing a significant factor compared to number of lines.\n\n-Peff\n"},{"id":"75789","messageId":"20080501231427.GD21731@sigill.intra.peff.net","threadId":"13345","inReplyTo":"D0968007-2A38-44DB-B26F-3D273F20D428@midwinter.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-01T23:14:27Z","receivedAt":"2008-05-01T23:14:27Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 01, 2008 at 12:12:33PM -0700, Steven Grimm wrote:\n\n> However, that leaves the question of which default will be wrong the  \n> least often.\n>\n> In my personal experience, I think a directory rename has almost always \n> meant that I would want new files to appear in the new directory rather \n\nI do agree that the rename is probably more often desired.\n\n> Of course, the discussion is moot anyway until someone writes code to  \n> detect the situation; my impression is the current behavior is the way it \n> is simply because it's what naturally happens in the absence of  \n> merge-time detection of a directory getting renamed.\n\nYes, I think that is largely a correct impression (although I think\nLinus has spoken out against directory renaming in the past, so there is\nat least a little bit of conscious effort). I suspect the right sequence\nof steps to implement this would be:\n\n  1. write a proof-of-concept that shows directory renaming after the\n    fact (e.g., take a conflicted merge, scan the diff for directory\n    renames, and then fix up the files). That way it is available, but\n    doesn't impact git at all.\n\n  2. If people think it is useful, build it into the diff and merge\n     machinery so that it can happen automagically, but make it\n     optional. Thus git fully supports it, but the policy decision is\n     left up to the user.\n\n  3. Make it the default if it is the common choice.\n\nSo we just need somebody to volunteer to work on 1. ;)\n\n-Peff\n"},{"id":"75791","messageId":"2e24e5b90805011906g769723f0g3ffbbe6588cf23d0@mail.gmail.com","threadId":"13345","inReplyTo":"20080501203940.GA3524@mithlond.arda.local","subject":"Re: detecting rename->commit->modify->commit","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2008-05-02T02:06:18Z","receivedAt":"2008-05-02T02:06:18Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"On Fri, May 2, 2008 at 2:09 AM, Teemu Likonen <tlikonen@iki.fi> wrote:\n\n>  -M\" didn't detect the rename but \"git diff -M4\" did. To me it looks like\n>  this works nicely, better than I expected, actually.\n\nerr... I didn't realise -M had an option, and I just double checked\nthe man pages for diff, diff-files, diff-index, and diff-tree.  What\ndoes the 4 mean?\n\nSitaram\n"},{"id":"75793","messageId":"7vtzhhxwep.fsf@gitster.siamese.dyndns.org","threadId":"13345","inReplyTo":"2e24e5b90805011906g769723f0g3ffbbe6588cf23d0@mail.gmail.com","subject":"Re: detecting rename->commit->modify->commit","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-05-02T02:38:38Z","receivedAt":"2008-05-02T02:38:38Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Sitaram Chamarty\" <sitaramc@gmail.com> writes:\n\n> On Fri, May 2, 2008 at 2:09 AM, Teemu Likonen <tlikonen@iki.fi> wrote:\n>\n>>  -M\" didn't detect the rename but \"git diff -M4\" did. To me it looks like\n>>  this works nicely, better than I expected, actually.\n>\n> err... I didn't realise -M had an option, and I just double checked\n> the man pages for diff, diff-files, diff-index, and diff-tree.  What\n> does the 4 mean?\n\nThe option to -M<num>, -C<num>, -B<num>/<num> are \"raise or lower the\nsimilarity threshold to <num> / 10^N\" where N is the number of digits in\n<num>.  IOW, you will always be expressing number between 0 and 1.\n\nYou should also be able to say -M40% but that is an ancient part of the\ncode base so I might be misremembering things.\n"},{"id":"75855","messageId":"2e24e5b90805020959h42258110vfd6fb4957643e6fc@mail.gmail.com","threadId":"13345","inReplyTo":"7vtzhhxwep.fsf@gitster.siamese.dyndns.org","subject":"Re: detecting rename->commit->modify->commit","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2008-05-02T16:59:29Z","receivedAt":"2008-05-02T16:59:29Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"On Fri, May 2, 2008 at 8:08 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>  The option to -M<num>, -C<num>, -B<num>/<num> are \"raise or lower the\n>  similarity threshold to <num> / 10^N\" where N is the number of digits in\n>  <num>.  IOW, you will always be expressing number between 0 and 1.\n\nThanks.  The only mention of this I find (now) is in a file called\ndiffcore.txt, which appears to exist only in the HTML documentation,\nbut not in the \"man\" pages anywhere, as of 1.5.5.\n\n[ I pulled a few hairs out trying to find it in the man pages :-) ]\n\nI'd submit a patch, but a guy who takes the easy way out even to get\nthe documentation (essentially doing a checkout of the \"man\" branch)\nwould certainly not be able to test it :-(\n"},{"id":"75950","messageId":"481CA742.4080909@tikalk.com","threadId":"13345","inReplyTo":"20080501231427.GD21731@sigill.intra.peff.net","subject":"merge renamed files/directories? (was: Re: detecting rename->commit->modify->commit)","fromName":"Ittay Dror","fromEmail":"ittayd@tikalk.com","sentAt":"2008-05-03T17:56:18Z","receivedAt":"2008-05-03T17:56:18Z","isPatch":false,"sender":{"key":"ittayd@tikalk.com","avatar":null},"body":"Can someone comment whether supporting merges after renames will be on \nthe Git roadmap?\n\nAs a Java developer, I can say that refactoring of class names and \npackages happens quite often. Having to remember I've made this change \nthroughout the lifetime of a branch (or master, until pushed to a \ncentral repository), and needing to manually merge changes to files / \npackages (directories) I've refactored is something that I want my VCS \nto do.\n\nThank you,\nIttay\n\nJeff King wrote:\n> On Thu, May 01, 2008 at 12:12:33PM -0700, Steven Grimm wrote:\n>\n>   \n>> However, that leaves the question of which default will be wrong the  \n>> least often.\n>>\n>> In my personal experience, I think a directory rename has almost always \n>> meant that I would want new files to appear in the new directory rather \n>>     \n>\n> I do agree that the rename is probably more often desired.\n>\n>   \n>> Of course, the discussion is moot anyway until someone writes code to  \n>> detect the situation; my impression is the current behavior is the way it \n>> is simply because it's what naturally happens in the absence of  \n>> merge-time detection of a directory getting renamed.\n>>     \n>\n> Yes, I think that is largely a correct impression (although I think\n> Linus has spoken out against directory renaming in the past, so there is\n> at least a little bit of conscious effort). I suspect the right sequence\n> of steps to implement this would be:\n>\n>   1. write a proof-of-concept that shows directory renaming after the\n>     fact (e.g., take a conflicted merge, scan the diff for directory\n>     renames, and then fix up the files). That way it is available, but\n>     doesn't impact git at all.\n>\n>   2. If people think it is useful, build it into the diff and merge\n>      machinery so that it can happen automagically, but make it\n>      optional. Thus git fully supports it, but the policy decision is\n>      left up to the user.\n>\n>   3. Make it the default if it is the common choice.\n>\n> So we just need somebody to volunteer to work on 1. ;)\n>\n> -Peff\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n>   \n\n-- \nIttay Dror <ittayd@tikalk.com>\nTikal <http://www.tikalk.com>\nTikal Project <http://tikal.sourceforge.net>\n"},{"id":"75952","messageId":"32541b130805031111r4cbea8e1l19c34ac05016a89b@mail.gmail.com","threadId":"13345","inReplyTo":"481CA742.4080909@tikalk.com","subject":"Re: merge renamed files/directories? (was: Re: detecting rename->commit->modify->commit)","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-05-03T18:11:45Z","receivedAt":"2008-05-03T18:11:45Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On 5/3/08, Ittay Dror <ittayd@tikalk.com> wrote:\n> Can someone comment whether supporting merges after renames will be on the\n> Git roadmap?\n>\n>  As a Java developer, I can say that refactoring of class names and packages\n> happens quite often. Having to remember I've made this change throughout the\n> lifetime of a branch (or master, until pushed to a central repository), and\n> needing to manually merge changes to files / packages (directories) I've\n> refactored is something that I want my VCS to do.\n\nGit already works fine for renames.  The only situation where\nsomething funny happens is if you rename a whole directory and someone\nelse creates a file in the old directory.  (In that case, the new file\nends up in the old place instead of the new place.)  However, even in\nthat case, there is still no conflict and no manual merging necessary.\n\nIn fact, as someone else pointed out, renaming a java file requires\nyou to modify the file anyhow, so having git auto-move the file to\nanother directory *still* wouldn't make it work any better.\n\nHave fun,\n\nAvery\n"},{"id":"75990","messageId":"481D52CC.1030503@tikalk.com","threadId":"13345","inReplyTo":"32541b130805031111r4cbea8e1l19c34ac05016a89b@mail.gmail.com","subject":"Re: merge renamed files/directories?","fromName":"Ittay Dror","fromEmail":"ittayd@tikalk.com","sentAt":"2008-05-04T06:08:12Z","receivedAt":"2008-05-04T06:08:12Z","isPatch":false,"sender":{"key":"ittayd@tikalk.com","avatar":null},"body":"\n\nAvery Pennarun wrote:\n> Git already works fine for renames.  The only situation where\n> something funny happens is if you rename a whole directory and someone\n> else creates a file in the old directory.  (In that case, the new file\n> ends up in the old place instead of the new place.)  However, even in\n> that case, there is still no conflict and no manual merging necessary.\n>\n>   \nSorry, but this is not the situation as I have experienced it with a \nlocal repository I have. I renamed a directory (without changing any \nfiles in it). 'git diff <commit>^ <commit>' shows the rename fine, but \n'git log -p -M -C <initial commit>..' does not (that is, the history for \nfiles in that directory is shown from the rename commit only). Obviously \ngit-diff is not any better.\n> In fact, as someone else pointed out, renaming a java file requires\n> you to modify the file anyhow, so having git auto-move the file to\n> another directory *still* wouldn't make it work any better.\n>\n>   \nSure it will, because otherwise I need to move it and still need to fix \nit. And there are many other file formats and languages where such a \nmove will not require any change (I think it is funny that Java is a \njustification for not doing something for a tool primarily used by C \npeople). Also, what happens if I change the file in the new location and \nsomeone else changes it in the old location? Will I need to do a manual \nmerge?\n\n-- \nIttay Dror <ittayd@tikalk.com>\nTikal <http://www.tikalk.com>\nTikal Project <http://tikal.sourceforge.net>\n"},{"id":"75998","messageId":"m3zlr65s6j.fsf@localhost.localdomain","threadId":"13345","inReplyTo":"481D52CC.1030503@tikalk.com","subject":"Re: merge renamed files/directories?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-05-04T09:34:23Z","receivedAt":"2008-05-04T09:34:23Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Ittay Dror <ittayd@tikalk.com> writes:\n> Avery Pennarun wrote:\n\n> > Git already works fine for renames.  The only situation where\n> > something funny happens is if you rename a whole directory and someone\n> > else creates a file in the old directory.  (In that case, the new file\n> > ends up in the old place instead of the new place.)  However, even in\n> > that case, there is still no conflict and no manual merging necessary.\n>\n> Sorry, but this is not the situation as I have experienced it with a\n> local repository I have. I renamed a directory (without changing any\n> files in it). 'git diff <commit>^ <commit>' shows the rename fine, but\n> 'git log -p -M -C <initial commit>..' does not (that is, the history\n> for files in that directory is shown from the rename commit\n> only). Obviously git-diff is not any better.\n\nThis is one thing where git differs from other SCMs.  In \"git log --\n<path>\" (that is what I assume you have used) the <path> argument is\npath limiter.  It allows to specify more than one directory or a file.\n\nUnfortunately currently \"git log --follow=<file>\" works only for single\nfiles, and doesn't yet work for directories; which is caused, among\nother things, by the lack of directory rename detection in git.\n\n> [...] Also, what happens if I change the file in the new location\n> and someone else changes it in the old location? Will I need to do a\n> manual merge?\n\nNo, rename detection should make automatic merge possible.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"76140","messageId":"32541b130805050940x1297e907ofc67ee65494897eb@mail.gmail.com","threadId":"13345","inReplyTo":"481D52CC.1030503@tikalk.com","subject":"Re: merge renamed files/directories?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-05-05T16:40:24Z","receivedAt":"2008-05-05T16:40:24Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On 5/4/08, Ittay Dror <ittayd@tikalk.com> wrote:\n>  Avery Pennarun wrote:\n> > In fact, as someone else pointed out, renaming a java file requires\n> > you to modify the file anyhow, so having git auto-move the file to\n> > another directory *still* wouldn't make it work any better.\n>\n> Sure it will, because otherwise I need to move it and still need to fix it.\n> And there are many other file formats and languages where such a move will\n> not require any change (I think it is funny that Java is a justification for\n> not doing something for a tool primarily used by C people).\n\nI mentioned Java because you mentioned you were working in java.\n\nThe particular problem with Java doesn't happen to C people.  Imagine,\nfor example, that I add a new file, lib/foo.c, to lib/lib.a (thus they\nhave to modify lib/Makefile), while someone else renames \"lib\" to\n\"bettername\".\n\nWhen I merge, if git would create bettername/foo.c (it currently\nwon't) and properly automerge bettername/Makefile (it will), then the\nprogram would still compile correctly.  However this doesn't work in\nJava: lib/foo.java would include the word \"lib\" in its contents (in\nthe namespace declaration) and so there's no way automatic merging\nwould have resulted in a version that compiles correctly.\n\nSo what I said isn't to *justify* git's behaviour, merely to point out\nthat in java's case, there seems to be no way to get fully automatic\nmerging that would work.  In C, this case would have worked, if only\ngit supported directory renames.\n\nIn neither case is it very much work to fix by hand, though :)\n\nHave fun,\n\nAvery\n"},{"id":"76152","messageId":"200805052349.35867.robin.rosenberg.lists@dewire.com","threadId":"13345","inReplyTo":"32541b130805050940x1297e907ofc67ee65494897eb@mail.gmail.com","subject":"Re: merge renamed files/directories?","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2008-05-05T21:49:35Z","receivedAt":"2008-05-05T21:49:35Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"måndagen den 5 maj 2008 18.40.24 skrev Avery Pennarun:\n> On 5/4/08, Ittay Dror <ittayd@tikalk.com> wrote:\n> >  Avery Pennarun wrote:\n> > > In fact, as someone else pointed out, renaming a java file requires\n> > > you to modify the file anyhow, so having git auto-move the file to\n> > > another directory *still* wouldn't make it work any better.\n> >\n> > Sure it will, because otherwise I need to move it and still need to fix\n> > it. And there are many other file formats and languages where such a move\n> > will not require any change (I think it is funny that Java is a\n> > justification for not doing something for a tool primarily used by C\n> > people).\n>\n> I mentioned Java because you mentioned you were working in java.\n>\n> The particular problem with Java doesn't happen to C people.  Imagine,\n> for example, that I add a new file, lib/foo.c, to lib/lib.a (thus they\n> have to modify lib/Makefile), while someone else renames \"lib\" to\n> \"bettername\".\n>\n> When I merge, if git would create bettername/foo.c (it currently\n> won't) and properly automerge bettername/Makefile (it will), then the\n> program would still compile correctly.  However this doesn't work in\n> Java: lib/foo.java would include the word \"lib\" in its contents (in\n> the namespace declaration) and so there's no way automatic merging\n> would have resulted in a version that compiles correctly.\n\nYou will always find corner cases. Line-by line merge happens to\nwork, not because it is the theoretically correct way, but because we\nhave discovered that it nearly always works so our need for more\nspecialized merging is not huge. We have also adapted our development\npractices to the way line-by-line merging works, i.e. we avoid binary\nfiles and funny text file formats.\n\n> So what I said isn't to *justify* git's behaviour, merely to point out\n> that in java's case, there seems to be no way to get fully automatic\n> merging that would work.  In C, this case would have worked, if only\n> git supported directory renames.\n\nSure, a merge that understands this is java and does the correct thing. Evn\nyour case for C (with hypotetical directory rename detection) would fail if \nthe renamed directory was used in an #include-statement (like #include \n<lib/foo.h>) Say someone thinks xxdiff should move to lib/xxdiff, while \nsomeone else adds a new reference to <xxdiff/xxdiff.h>. To resolve all cases \nyou must have tools that understand what they are doing. Directyry rename\ndetection only solves a few cases, but it may be easy enough to implement to \nwarrant the effort to get the tick in the box.\n\n>\n> In neither case is it very much work to fix by hand, though :),\n\nI agree on that.\n\n-- robin\n"},{"id":"76154","messageId":"alpine.LFD.1.10.0805051512060.32269@woody.linux-foundation.org","threadId":"13345","inReplyTo":"200805052349.35867.robin.rosenberg.lists@dewire.com","subject":"Re: merge renamed files/directories?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-05-05T22:20:14Z","receivedAt":"2008-05-05T22:20:14Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 5 May 2008, Robin Rosenberg wrote:\n> \n> You will always find corner cases.\n\n.. and btw, this is why merging should always \n\n - be predictable (which implies \"simple\": overly clever merging, and \n   especially merging that takes complex history into account is *bad*, \n   because it's still going to do the wrong thing, but now it's going to \n   do so much less predictable)\n\n - be amenable to manual fixes even when it succeeds (ie even if an \n   automatic merge completes without errors, a subsequent build may find \n   problems, and a \"git commit --amend\" may well be the right thing to \n   do!)\n\n - aim for (preferrably easily-handled) conflicts when the unusual cases \n   happen.\n\n   Conflicts for *common* things are bad, because they just cause more \n   work, and people get too complacent about fixing them. But similarly, \n   thinking that the unusual cases should be handled automatically is also \n   wrong - because the unusual cases are likely the ones that need some \n   manual resolution anyway.\n\nGit will never do merges \"perfectly\", if only because it's fundamentally \nimpossible to do that. But one thing git *does* do is to make it pretty \ndamn easy to handle it.\n\nI really don't understand why people expect a directory rename to be \nhandled automatically, when it is (a) not that common and (b) not obvious \nwhat the solution is, but MOST OF ALL (c) so damn _easy_ to handle it \nmanually after-the-fact when you notice that something doesn't compile!\n\nReally. If you have a file that was created in the wrong subdirectory (and \nplease admit that this is not common - it requires not just a directory \nrename, but also a file create in another branch at the same time), what's \nso hard with just doing\n\n\tmake\n\t.. oh, oops, that was pretty obviousm, the expected source file \n\t   didn't exist ..\n\tgit mv olddir/file newdir/file\n\tgit commit --amend\n\nand \"Tadaa! All done\". Your merge that was *fundamentally impossible* to \ndo automatically, was trivially done manually, with no actual big \nhead-scratiching involved.\n\n\t\tLinus\n"},{"id":"76157","messageId":"ADDE27A8-6329-4C09-BC07-8EB023BA6D48@midwinter.com","threadId":"13345","inReplyTo":"alpine.LFD.1.10.0805051512060.32269@woody.linux-foundation.org","subject":"Re: merge renamed files/directories?","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2008-05-05T23:07:57Z","receivedAt":"2008-05-05T23:07:57Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"On May 5, 2008, at 3:20 PM, Linus Torvalds wrote:\n> I really don't understand why people expect a directory rename to be\n> handled automatically, when it is (a) not that common and (b) not  \n> obvious\n> what the solution is, but MOST OF ALL (c) so damn _easy_ to handle it\n> manually after-the-fact when you notice that something doesn't  \n> compile!\n\nAssuming all you track with git is source code that has dependencies  \nsuch that a compile command fails cleanly when things end up in the  \nwrong directory, sure.\n\nIf you're using git to, say, track a tree of documentation files or  \nimages that are referred to using relative URLs in HTML pages,  \ndetecting the breakage is less trivial unless you have a really solid  \nautomated QA process that can check for dangling references.\n\nAre directory renames as common as file renames? Certainly not. But  \nthey happen often enough that it's annoying to have to manually clean  \nup after them. Note that I did not say it is difficult or impossible  \nto manually clean up after them. I think the number of people who've  \nmentioned this on the list should stand as some kind of refutation of  \nthe idea that directory renames are so vanishingly rare as to not be  \nworth mentioning. I've run into the problem a few times myself.\n\n> and \"Tadaa! All done\". Your merge that was *fundamentally  \n> impossible* to\n> do automatically, was trivially done manually, with no actual big\n> head-scratiching involved.\n\n$ mkdir parent\n$ cd parent\n$ hg init\n$ mkdir subdir1\n$ echo \"I am the walrus\" > subdir1/file1\n$ hg add subdir1/file1\n$ hg commit -m 'initial commit'\n$ cd ..\n$ hg clone parent child\n$ cd child\n$ hg mv subdir1 subdir2\n$ hg commit -m 'rename subdir1 to subdir2'\n$ cd ../parent\n$ echo 'I love prunes' > subdir1/file2\n$ hg add subdir1/file2\n$ hg commit -m 'new file in subdir'\n$ cd ../child\n$ hg pull\n$ hg merge\n$ ls subdir2\nfile1   file2\n\nDoesn't seem *fundamentally* impossible to produce the results that  \nare most likely to be what people want. (Which doesn't equal  \n\"guaranteed to be 100% correct 100% of the time or your money back\" --  \nas you say, merging is an inexact science.)\n\n-Steve\n"},{"id":"76159","messageId":"alpine.LFD.1.10.0805051724510.32269@woody.linux-foundation.org","threadId":"13345","inReplyTo":"ADDE27A8-6329-4C09-BC07-8EB023BA6D48@midwinter.com","subject":"Re: merge renamed files/directories?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-05-06T00:29:12Z","receivedAt":"2008-05-06T00:29:12Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 5 May 2008, Steven Grimm wrote:\n>\n> Doesn't seem *fundamentally* impossible to produce the results that are most\n> likely to be what people want.\n\nYou didn't understand what was fundamentally impossible.\n\nAnd btw, this has nothing to do with directory renames either. There are \ntons of these kinds of merge issues that bad SCM developes have been \nmasturbating over for YEARS. There's a whole science of making idiotic new \nmerging models, one fancier than the other. The fact is, you cannot do a \nperfect job, the best thing you can do is pick a simple model, and try to \nmake it repeatable and easy to fix up.\n\nMaybe somebody bothers to implement some directory rename heuristic some \nday. Quite frankly, I personally cannot care less. It really is mental \nmasturbation, and has absolutely no relevance for any real-world problem.\n\n\t\tLinus\n"},{"id":"76161","messageId":"alpine.LFD.1.10.0805051737180.32269@woody.linux-foundation.org","threadId":"13345","inReplyTo":"alpine.LFD.1.10.0805051724510.32269@woody.linux-foundation.org","subject":"Re: merge renamed files/directories?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-05-06T00:40:25Z","receivedAt":"2008-05-06T00:40:25Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 5 May 2008, Linus Torvalds wrote:\n> \n> There are tons of these kinds of merge issues that bad SCM developes \n> have been masturbating over for YEARS.\n\n.. and if I sound rather less than enthused about these kinds of issues, \nit's because of having seen years and years of people talking about merge \nstrategies, and then at the same time using SVN which doesn't even record \nthe parenthood of the resulting merges, or thinking that code always \nmoves with whole files.\n\nIn other words, the details don't even matter. What matters is not being a \ntotal piece of sh*t in the big picture. \n\n\t\tLinus\n"},{"id":"76162","messageId":"32541b130805051838k367c44bau715774b46f7894cb@mail.gmail.com","threadId":"13345","inReplyTo":"alpine.LFD.1.10.0805051512060.32269@woody.linux-foundation.org","subject":"Re: merge renamed files/directories?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-05-06T01:38:26Z","receivedAt":"2008-05-06T01:38:26Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On 5/5/08, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>  I really don't understand why people expect a directory rename to be\n>  handled automatically, when it is (a) not that common and (b) not obvious\n>  what the solution is, but MOST OF ALL (c) so damn _easy_ to handle it\n>  manually after-the-fact when you notice that something doesn't compile!\n\nI general I agree with your point here, but I still find it surprising\nhow hard the directory-rename problem is made out to be.  As far as I\ncan see, the right implementation exactly parallels the single-file\nrename implementation.\n\nI think the same problem that prevents git from knowing the difference\nbetween empty and nonexistent directories (eg.\nhttp://kerneltrap.org/mailarchive/git/2007/7/18/251976) is the one\nthat prevents it from handling directory renames: git doesn't\nacknowledge that it's *already* treating directories as first-class\nobjects.\n\nWhat if you thought of a directory as simply a list of filenames?\n(This is more or less what unix does anyway.)  Then an *empty*\ndirectory is a tree of zero length; a nonexistent (or not tracked)\ndirectory is simply not listed in the parent; a directory with\nuntracked files is like a file with patches not yet added to the\nindex(*); and trying to merge a file into a nonexistent directory\n(when the original patch *didn't* create the directory fresh) would\ntrigger similar logic to the existing rename handling.  That is, put\nthe new file with the content that used to be next to it, by looking\nfor a tree with contents (names, not so much sha1's) similar to the\none it was expected to be in.\n\n> It really is mental\n> masturbation, and has absolutely no relevance for any real-world problem.\n\nI personally don't get very interested in non-real-world problems.\nHere's the actual case I tried to use a few months ago, but couldn't,\nbecause git doesn't track directory renames.  (Note that I was quite\nhappily able to do this in svn, as much as you can do anything happily\nin svn.)\n\nI have a branch called 'mylib' with my library project in its root\ndirectory.  What I wanted was to maintain my library in the 'mylib'\nbranch, then merge my library into the \"libs/mylib\" directory of my\napplication, which is in the 'myapp' branch.  (Of course, in real\nlife, there's more than one app using mylib in more than one\nrepository, and I'm actually doing 'git pull' of the mylib branch from\nelsewhere.)\n\nThis actually works like magic in git - except when you create a file\nin the 'mylib' branch, in which case it gets merged to the wrong path\nevery single time.  It seems to me like it should be very easy to put\nit in the right place instead, making one more interesting use case\npossible.\n\nI realize git-submodule is the way you're supposed to do something\nlike this, but git-submodule doesn't really do what I want (yet) for\nreasons discussed in other threads.\n\nHave fun,\n\nAvery\n\n(*) Applying the same metaphor in reverse, operations that are valid\non directories are also valid for file contents.  I can think of\nimmediate uses for a .gitignore-style list that talks about file\n*contents*.  Imagine if I could make a local patch to my Makefile,\nmark that one patch as \"ignored\", and never accidentally check it in.\n"},{"id":"76163","messageId":"20080506014636.GM29038@spearce.org","threadId":"13345","inReplyTo":"32541b130805051838k367c44bau715774b46f7894cb@mail.gmail.com","subject":"Re: merge renamed files/directories?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-05-06T01:46:37Z","receivedAt":"2008-05-06T01:46:37Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Avery Pennarun <apenwarr@gmail.com> wrote:\n> \n> I have a branch called 'mylib' with my library project in its root\n> directory.  What I wanted was to maintain my library in the 'mylib'\n> branch, then merge my library into the \"libs/mylib\" directory of my\n> application, which is in the 'myapp' branch. [...]\n> \n> This actually works like magic in git - except when you create a file\n> in the 'mylib' branch, in which case it gets merged to the wrong path\n> every single time.  It seems to me like it should be very easy to put\n> it in the right place instead, making one more interesting use case\n> possible.\n> \n> I realize git-submodule is the way you're supposed to do something\n> like this, but git-submodule doesn't really do what I want (yet) for\n> reasons discussed in other threads.\n\n`git pull -s subtree mylib` ?\n\nThis is how git-gui and gitk are merged into git.git, and it avoids\nthis case by looking for a subdirectory rename, more specifically\na rename of \"/\" to \"mylib/\".\n\nIt also can go the other way, that is rename \"mylib/\" to \"/\", but\nthis path is never used as far as I know as git-gui and gitk don't\never merge in the git.git history.\n\n-- \nShawn.\n"},{"id":"76164","messageId":"32541b130805051858u7b8f1cd7qd34fdf50c1f849d0@mail.gmail.com","threadId":"13345","inReplyTo":"20080506014636.GM29038@spearce.org","subject":"Re: merge renamed files/directories?","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2008-05-06T01:58:50Z","receivedAt":"2008-05-06T01:58:50Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On 5/5/08, Shawn O. Pearce <spearce@spearce.org> wrote:\n> Avery Pennarun <apenwarr@gmail.com> wrote:\n>  >\n>  > I have a branch called 'mylib' with my library project in its root\n>  > directory.  What I wanted was to maintain my library in the 'mylib'\n>  > branch, then merge my library into the \"libs/mylib\" directory of my\n>\n> > application, which is in the 'myapp' branch. [...]\n>\n> >\n>  > This actually works like magic in git - except when you create a file\n>  > in the 'mylib' branch, in which case it gets merged to the wrong path\n>  > every single time.  It seems to me like it should be very easy to put\n>  > it in the right place instead, making one more interesting use case\n>  > possible.\n>  >\n>  > I realize git-submodule is the way you're supposed to do something\n>  > like this, but git-submodule doesn't really do what I want (yet) for\n>  > reasons discussed in other threads.\n>\n> `git pull -s subtree mylib` ?\n\nFirst, I thought: wow!  How can that possibly work?  These guys are geniuses!\n\nThen I found out that git-merge-subtree is a git builtin, and git.c says this:\n\n  { \"merge-recursive\", cmd_merge_recursive, RUN_SETUP | NEED_WORK_TREE },\n  { \"merge-subtree\", cmd_merge_recursive, RUN_SETUP | NEED_WORK_TREE },\n\nAnd then my head exploded. :)\n\nStill scraping the pieces of my brain back off the floor... but does\nthis mean the subtree merge strategy would fail exactly like\nmerge-recursive when new files are created?\n\nHave fun,\n\nAvery\n"},{"id":"76165","messageId":"20080506021202.GN29038@spearce.org","threadId":"13345","inReplyTo":"32541b130805051858u7b8f1cd7qd34fdf50c1f849d0@mail.gmail.com","subject":"Re: merge renamed files/directories?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-05-06T02:12:02Z","receivedAt":"2008-05-06T02:12:02Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Avery Pennarun <apenwarr@gmail.com> wrote:\n> On 5/5/08, Shawn O. Pearce <spearce@spearce.org> wrote:\n> >\n> > `git pull -s subtree mylib` ?\n> \n> First, I thought: wow!  How can that possibly work?  These guys are geniuses!\n> \n> Then I found out that git-merge-subtree is a git builtin, and git.c says this:\n> \n>   { \"merge-recursive\", cmd_merge_recursive, RUN_SETUP | NEED_WORK_TREE },\n>   { \"merge-subtree\", cmd_merge_recursive, RUN_SETUP | NEED_WORK_TREE },\n> \n> And then my head exploded. :)\n> \n> Still scraping the pieces of my brain back off the floor... but does\n> this mean the subtree merge strategy would fail exactly like\n> merge-recursive when new files are created?\n\nNope.  If you go look at cmd_merge_recursive you will see it has\ndifferent behavior based upon the name it was invoked as, even\nthough it is the same C function and has the same implementation.\n\nIf it is started with the name \"merge-subtree\" it tries to find\na matching subtree prefix to insert in front of all names, or\nto remove from all names, such that a merge will correctly fully\ninclude a set of files in a subdirectory, or full pull out a set\nof files from a subdirectory.\n\nJunio is the genius that implemented this.  Works quite well for\nthis library->application merge case that I think you were trying\nto describe.\n\n-- \nShawn.\n"},{"id":"76166","messageId":"alpine.LFD.1.10.0805051918200.32269@woody.linux-foundation.org","threadId":"13345","inReplyTo":"32541b130805051838k367c44bau715774b46f7894cb@mail.gmail.com","subject":"Re: merge renamed files/directories?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-05-06T02:19:24Z","receivedAt":"2008-05-06T02:19:24Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 5 May 2008, Avery Pennarun wrote:\n>\n> I general I agree with your point here, but I still find it surprising\n> how hard the directory-rename problem is made out to be.\n\nI do agree that it's probably not that hard.\n\nBut I disagree with people who whine about pointless stuff, and don't send \npatches.\n\n\t\tLinus\n"},{"id":"76205","messageId":"20080506154709.GF6918@mit.edu","threadId":"13345","inReplyTo":"alpine.LFD.1.10.0805051724510.32269@woody.linux-foundation.org","subject":"Re: merge renamed files/directories?","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-05-06T15:47:09Z","receivedAt":"2008-05-06T15:47:09Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, May 05, 2008 at 05:29:12PM -0700, Linus Torvalds wrote:\n> \n> Maybe somebody bothers to implement some directory rename heuristic some \n> day. Quite frankly, I personally cannot care less. It really is mental \n> masturbation, and has absolutely no relevance for any real-world problem.\n> \n\nActually, the directory rename hueristic *does* have relevance in at\nleast some real-world cases.  For example, MySQL has plugin\ndirectories, and occasionally the plugins get renamed, for whatever\nreason.  If a plugin gets renamed, so does its directory, and if the\nrename operation happens in an experimental (or devel) branch, but\nthen for whatever reason, a new file is created in the devel (or\nmaint) branch, without the directory rename hueristic, when the\nchangeset is pulled into the experimental (or devel) branch, the file\nwill be created in the wrong directory.\n\nSo it may be rare, but this kind of thing does happen in the real\nworld.\n\n\t\t\t\t\t\t\t- Ted\n"},{"id":"76207","messageId":"alpine.LFD.1.10.0805060851470.32269@woody.linux-foundation.org","threadId":"13345","inReplyTo":"20080506154709.GF6918@mit.edu","subject":"Re: merge renamed files/directories?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-05-06T16:10:06Z","receivedAt":"2008-05-06T16:10:06Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 6 May 2008, Theodore Tso wrote:\n> \n> Actually, the directory rename hueristic *does* have relevance in at\n> least some real-world cases.  For example, MySQL has plugin\n> directories, and occasionally the plugins get renamed, for whatever\n> reason.\n\nI'm not saying that directory renames don't happen.\n\nI don't even say that merges across directory renames don't happen.\n\nI *am* saying that it's not a problem.\n\nIt's like data conflicts. Do they happen? Sure as hell. I can pretty much \nguarantee that any sane project will have more data conflicts than they \nwill have rename conflicts (whether single-file or directory), and it's \nnot only a problem, it's something that is absolutely *required* from a \nsource control management system!\n\nSo are data conflicts a problem?\n\nI claim that they aren't. They are a *positive* resource that you need to \nhandle. Some of the \"handling\" is obviously going to be to try to avoid \nthem, and if you get too much of them, the real \"problem\" is that you \nmerge too seldom, or more commonly that you have a piece of code that is \nsimply not done well enough, so many different people have to muck around \nin that area.\n\nBut fundamentally, you should always have data conflicts, and they aren't \na problem in themselves. They are a problem only\n\n - If they are hard to understand and see, and *unexpected*. The SCM\n   should explain what is going on, and explain why a conflict happens \n   (and that may perhaps mean after-the fact! I love \"gitk --merge\" \n   exactly because it tends to be very good at explaining what was going \n   on!).\n\n - If they are hard to fix.\n\n   For example, one of the main problems I had with BK merging was the \n   fact that while the megetool was wonderful, you effectively *had* to \n   merge using it, and you couldn't sanely do an \"incremental\" merge \n   where you first did a first merge job, then checked that it at \n   least compiles, then tested it, and finally looked at the diffs from \n   both parents and looked at whether those all made sense, and you could \n   \"refine\" or fix the merge along the different phases.\n\n   Of course, you hope that all merges are pretty obvious, and you can do \n   it right in one go, but no, they're not. They'll never be. They'll \n   never be fully automtic, but even when they aren't automatic, they'll \n   not even be trivially to do manually. But that's OK, as long as the \n   tool at least doesn't fight you, and lets you do whatever you want to \n   do a part of fixing things up.\n\nNow, take a look back at directory renames.\n\nDo they happen?\n\nYes.\n\nDo they potentially mis-merge?\n\nYes.\n\nBut are they common and/or hard to fix and handle?\n\nNo.\n\nAnd that's why I don't think people should call them \"problems\". The only \n_real_ issue here, I think, is that git just does things differently from \nother SCM's. Git does a _lot_ of things differently. You get used to it.\n\n\t\t\tLinus\n"},{"id":"76209","messageId":"alpine.LFD.1.10.0805060914190.32269@woody.linux-foundation.org","threadId":"13345","inReplyTo":"alpine.LFD.1.10.0805060851470.32269@woody.linux-foundation.org","subject":"Re: merge renamed files/directories?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-05-06T16:15:22Z","receivedAt":"2008-05-06T16:15:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 6 May 2008, Linus Torvalds wrote:\n> \n>\t\t\t\t\t\t\tI can pretty much \n> guarantee that any sane project will have more data conflicts than they \n> will have rename conflicts (whether single-file or directory), and it's \n> not only a problem, it's something that is absolutely *required* from a \n         ^^^-- not\n> source control management system!\n\nOops. That didn't read well.\n\n\t\t\tLinus\n"},{"id":"76210","messageId":"48208817.4060005@tikalk.com","threadId":"13345","inReplyTo":"alpine.LFD.1.10.0805060851470.32269@woody.linux-foundation.org","subject":"Re: merge renamed files/directories?","fromName":"Ittay Dror","fromEmail":"ittayd@tikalk.com","sentAt":"2008-05-06T16:32:23Z","receivedAt":"2008-05-06T16:32:23Z","isPatch":false,"sender":{"key":"ittayd@tikalk.com","avatar":null},"body":"\n\nLinus Torvalds wrote:\n>\n>  - If they are hard to understand and see, and *unexpected*. The SCM\n>    should explain what is going on, and explain why a conflict happens \n>    (and that may perhaps mean after-the fact! I love \"gitk --merge\" \n>    exactly because it tends to be very good at explaining what was going \n>    on!).\n>\n>   \nSo does git tell me what is going on with directory renames? Or should I \njust discover them when I try to compile (assuming that when the old \ndirectory name appears it will even get compiled, and that the file in \nit is something that gets compiled)\n\nAnd no, it's not a common problem, but I don't like the fact that a \nmerge conflict happens and the SCM doesn't tell me about it.\n\n-- \nIttay Dror <ittayd@tikalk.com>\nTikal <http://www.tikalk.com>\nTikal Project <http://tikal.sourceforge.net>\n"},{"id":"76211","messageId":"alpine.LFD.1.10.0805060936270.32269@woody.linux-foundation.org","threadId":"13345","inReplyTo":"48208817.4060005@tikalk.com","subject":"Re: merge renamed files/directories?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-05-06T16:39:01Z","receivedAt":"2008-05-06T16:39:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 6 May 2008, Ittay Dror wrote:\n> \n> And no, it's not a common problem, but I don't like the fact that a merge\n> conflict happens and the SCM doesn't tell me about it.\n\nI do agree that the most irritating feature of it is the silent clean \nmerge. When it's not obvious what the right thing to do is, generally a \nmerge strategy should try to warn, or even generate a conflict.\n\nThat said, anybody who thinks that \"merge was automatic and successful\" \nmeans that the mege was _correct_ is sadly mistaken. So you really \nshouldn't depend on it, and yeah, I strongly suggest building and testing \nafter a merge (and before you push the result out), so that you can fix \nany issues.\n\n\t\t\tLinus\n"},{"id":"76390","messageId":"20080508181723.GA30449@sigill.intra.peff.net","threadId":"13345","inReplyTo":"20080501231427.GD21731@sigill.intra.peff.net","subject":"Re: detecting rename->commit->modify->commit","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-05-08T18:17:24Z","receivedAt":"2008-05-08T18:17:24Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 01, 2008 at 07:14:27PM -0400, Jeff King wrote:\n\n>   1. write a proof-of-concept that shows directory renaming after the\n>     fact (e.g., take a conflicted merge, scan the diff for directory\n>     renames, and then fix up the files). That way it is available, but\n>     doesn't impact git at all.\n\nHere's a toy script that finds directory renames. I'm sure there are a\nton of corner cases it doesn't handle (like directory renames inside of\ndirectory renames). My test case was the very trivial:\n\n  mkdir repo && cd repo && git init\n\n  mkdir subdir\n  for i in 1 2 3; do\n    echo content $i >subdir/file$i\n  done\n  git add subdir\n  git commit -m initial\n\n  git mv subdir new\n  git commit -m move\n\n  git checkout -b other HEAD^\n  echo content 4 >subdir/file4\n  git add subdir\n  git commit -m new\n\n  git merge --no-commit master\n  perl ../find-dir-rename.pl\n  git commit\n\nAt which point you should see the merged commit with new/file4.\n\nScript is below.\n\n-- >8 --\n#!/usr/bin/perl\n#\n# Find renamed directories, and move any files in the \"old\"\n# directory into the \"new\".\n#\n# usage:\n#   git merge --no-commit <whatever>\n#   find-dir-rename\n#   git commit\n\nuse strict;\n\nforeach my $r (renamed_dirs()) {\n  move_dir_contents($r->{from}, $r->{to});\n}\nexit 0;\n\nsub renamed_dirs {\n  my $base = `git merge-base HEAD MERGE_HEAD`;\n  chomp $base;\n  return grep {\n    $_->{score} == 1\n  } (renamed_dirs_between($base, 'HEAD'),\n     renamed_dirs_between($base, 'MERGE_HEAD'));\n}\n\nsub renamed_dirs_between {\n  my ($base, $commit) = @_;\n\n  my %sources;\n  foreach my $pair (renamed_files($base, $commit)) {\n    my $d1 = dir_of($pair->[0]);\n    my $d2 = dir_of($pair->[1]);\n    next unless defined($d1) && defined($d2);\n\n    $sources{$d1}->{total}++;\n    $sources{$d1}->{dests}->{$d2}++;\n  }\n\n  return map {\n    my $from = $_;\n    map {\n      {\n        from => $from,\n        to => $_,\n        score => $sources{$from}->{dests}->{$_} / $sources{$from}->{total},\n      }\n    } keys(%{$sources{$from}->{dests}});\n  } removed_directories($base, $commit);\n}\n\nsub dir_of {\n  local $_ = shift;\n  s{/[^/]+$}{} or return undef;\n  return $_;\n}\n\nsub renamed_files {\n  my ($from, $to) = @_;\n  open(my $fh, '-|', qw(git diff-tree -r -M), $from, $to)\n    or die \"unable to open diff-tree: $!\";\n  return map {\n    chomp;\n    m/ R\\d+\\t([^\\t]+)\\t(.*)/ ? [$1 => $2] : ()\n  } <$fh>;\n}\n\nsub removed_directories {\n  my ($base, $commit) = @_;\n  my %new_dirs = map { $_ => 1 } directories($commit);\n  return grep { !exists $new_dirs{$_} } directories($base);\n}\n\nsub directories {\n  my $commit = shift;\n  return uniq(\n    map {\n      s{/[^/]+$}{} ? $_ : ()\n    } files($commit)\n  );\n}\n\nsub files {\n  my $commit = shift;\n  open(my $fh, '-|', qw(git ls-tree -r), $commit)\n    or die \"unable to open ls-tree: $!\";\n  return map {\n    chomp;\n    s/^[^\\t]*\\t//;\n    $_\n  } <$fh>;\n}\n\nsub uniq {\n  my %seen;\n  return grep { !$seen{$_}++ } @_;\n}\n\nsub move_dir_contents {\n  my ($from, $to) = @_;\n\n  my @files = glob(\"$from/*\");\n  return unless @files;\n\n  system(qw(git mv), @files, \"$to/\")\n    and die \"unable to move $from/* to $to\";\n  rmdir($from); # ignore error since there may be untracked files\n}\n"}]}