{"thread":{"id":"22508","subject":"git-mv redux: there must be something else going on","startedAt":"2010-02-03T18:25:54Z","lastAt":"2010-02-04T00:48:20Z","messageCount":20,"participants":["Ron Garret","Avery Pennarun","Nicolas Pitre","Pete Harlan","Thomas Rast","Junio C Hamano","Jay Soffian"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"133492","messageId":"ron1-32BD5F.10255403022010@news.gmane.org","threadId":"22508","inReplyTo":null,"subject":"git-mv redux: there must be something else going on","fromName":"Ron Garret","fromEmail":"ron1@flownet.com","sentAt":"2010-02-03T18:25:54Z","receivedAt":"2010-02-03T18:25:54Z","isPatch":false,"sender":{"key":"ron1@flownet.com","avatar":null},"body":"Based on my current understanding of git there should be no difference \nbetween a git mv and a git rm followed by a git add.  But empirically \nthere is a difference.  git log --M --follow is able to track files \nthrough git mvs even if their content changes completely.  Likewise, it \ndoes *not* track files through rm/add combinations even if the content \ndidn't change at all.  (See experiment transcript below.)\n\nSo something in my understanding of how git works must be wrong.  Git \nmust be keeping a separate record of file renames somewhere.  But where?\n\nJust for the record, I'm not complaining about this behavior.  In fact, \nwhat git does is exactly what I want.  I just want to understand how it \nworks.\n\nThanks,\nrg\n\n\n---\n\n[ron@mickey:~/devel/gittest]$ cat>file1\n1\n2\n3\n4\n5\n[ron@mickey:~/devel/gittest]$ cat>file2\na\nb\nc\nd\ne\n[ron@mickey:~/devel/gittest]$ git init\nInitialized empty Git repository in /Users/ron/devel/gittest/.git/\n[ron@mickey:~/devel/gittest]$ git add file1\n[ron@mickey:~/devel/gittest]$ git commit -m 'Add numbers'\n[master (root-commit) 54c2e4a] Add numbers\n 1 files changed, 5 insertions(+), 0 deletions(-)\n create mode 100644 file1\n[ron@mickey:~/devel/gittest]$ git rm file1\nrm 'file1'\n[ron@mickey:~/devel/gittest]$ git add file2\n[ron@mickey:~/devel/gittest]$ git commit -m 'numbers->letters'\n[master fe05d12] numbers->letters\n 2 files changed, 5 insertions(+), 5 deletions(-)\n delete mode 100644 file1\n create mode 100644 file2\n[ron@mickey:~/devel/gittest]$ git log --name-status -M --follow file2\ncommit fe05d1233be1bb11f4ed0e8496e4191795d515a0\nAuthor: rongarret <ron@mickey>\nDate:   Wed Feb 3 10:13:38 2010 -0800\n\n   numbers->letters\n\nA       file2\n[ron@mickey:~/devel/gittest]$ ls\nfile2 git/\n[ron@mickey:~/devel/gittest]$ cat>file2\n6\n7\n8\n9\n10\n[ron@mickey:~/devel/gittest]$ git mv file2 file3\n[ron@mickey:~/devel/gittest]$ git commit -m 'letters->numbers'\n[master ae3f6d4] letters->numbers\n 1 files changed, 0 insertions(+), 0 deletions(-)\n rename file2 => file3 (100%)\n[ron@mickey:~/devel/gittest]$ git log --name-status -M --follow file3\ncommit ae3f6d440483fa41cf08819237e87d567ac3a31d\nAuthor: rongarret <ron@mickey>\nDate:   Wed Feb 3 10:15:00 2010 -0800\n\n    letters->numbers\n\nR100    file2   file3\n\ncommit fe05d1233be1bb11f4ed0e8496e4191795d515a0\nAuthor: rongarret <ron@mickey>\nDate:   Wed Feb 3 10:13:38 2010 -0800\n\n   numbers->letters\n\nA       file2\n[ron@mickey:~/devel/gittest]$ ls\nfile3 git/\n[ron@mickey:~/devel/gittest]$ mv file3 file4\n[ron@mickey:~/devel/gittest]$ git rm file3\nrm 'file3'\n[ron@mickey:~/devel/gittest]$ git add file4\n[ron@mickey:~/devel/gittest]$ git commit -m 'rm/add identical content'\n[master a3d7227] rm/add identical content\n 2 files changed, 5 insertions(+), 5 deletions(-)\n delete mode 100644 file3\n create mode 100644 file4\n[ron@mickey:~/devel/gittest]$ git log --name-status -M --follow file4\ncommit a3d7227fc2edca75fff8894acd5b077d1788bb36\nAuthor: rongarret <ron@mickey>\nDate:   Wed Feb 3 10:17:23 2010 -0800\n\n    rm/add identical content\n\nA       file4\n[ron@mickey:~/devel/gittest]$\n"},{"id":"133496","messageId":"32541b131002031048i26d166d9w3567a60515235c34@mail.gmail.com","threadId":"22508","inReplyTo":"ron1-32BD5F.10255403022010@news.gmane.org","subject":"Re: git-mv redux: there must be something else going on","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-02-03T18:48:06Z","receivedAt":"2010-02-03T18:48:06Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Feb 3, 2010 at 1:25 PM, Ron Garret <ron1@flownet.com> wrote:\n> So something in my understanding of how git works must be wrong.  Git\n> must be keeping a separate record of file renames somewhere.  But where?\n\nIt doesn't.  Your experiment is wrong.\n\n> [ron@mickey:~/devel/gittest]$ cat>file2\n> 6\n> 7\n> 8\n> 9\n> 10\n> [ron@mickey:~/devel/gittest]$ git mv file2 file3\n> [ron@mickey:~/devel/gittest]$ git commit -m 'letters->numbers'\n> [master ae3f6d4] letters->numbers\n>  1 files changed, 0 insertions(+), 0 deletions(-)\n>  rename file2 => file3 (100%)\n\nWhoops.  You didn't 'git add file2' (before the mv) or 'git add file3'\n(after the mv), or use commit -a, so what you've committed is the\n*old* content of file2 under the name file3.  The *new* content of\nfile2 is still uncommitted in your work tree under the name file3.\nThis is why git can detect the move.  (The 100% is a good clue: it\nmeans the old and new files are 100% identical.)\n\nArtificial tests like this are useless anyway.  If you renamed file2\nto file3 *and* changed all the contents, did you *really* rename it?\nIf so, who cares?  What good does it do you to know this?  If someone\nelse tries to patch the old file2 and you merge it into a (totally\ndifferent) file3 vs a (now missing) file2, how is that any better?\n\nOn the other hand, if one guy moves file2 to file3 and changes a few\nlines, you want the other guy's patch to go into file3, whether the\nfirst guy used 'git mv' or add+rm or anything else.\n\nAs long as only a few lines changed, git does the right thing.  If\nmost/all of the lines have changed, then there is no right thing,\nbecause you'll get a nasty merge conflict either way.\n\nHave fun,\n\nAvery\n"},{"id":"133501","messageId":"ron1-5F71CB.11234903022010@news.gmane.org","threadId":"22508","inReplyTo":"32541b131002031048i26d166d9w3567a60515235c34@mail.gmail.com","subject":"Re: git-mv redux: there must be something else going on","fromName":"Ron Garret","fromEmail":"ron1@flownet.com","sentAt":"2010-02-03T19:23:49Z","receivedAt":"2010-02-03T19:23:49Z","isPatch":false,"sender":{"key":"ron1@flownet.com","avatar":null},"body":"In article \n<32541b131002031048i26d166d9w3567a60515235c34@mail.gmail.com>,\n Avery Pennarun <apenwarr@gmail.com> wrote:\n\n> On Wed, Feb 3, 2010 at 1:25 PM, Ron Garret <ron1@flownet.com> wrote:\n> > So something in my understanding of how git works must be wrong.  Git\n> > must be keeping a separate record of file renames somewhere.  But where?\n> \n> It doesn't.  Your experiment is wrong.\n> \n> > [ron@mickey:~/devel/gittest]$ cat>file2\n> > 6\n> > 7\n> > 8\n> > 9\n> > 10\n> > [ron@mickey:~/devel/gittest]$ git mv file2 file3\n> > [ron@mickey:~/devel/gittest]$ git commit -m 'letters->numbers'\n> > [master ae3f6d4] letters->numbers\n> >  1 files changed, 0 insertions(+), 0 deletions(-)\n> >  rename file2 => file3 (100%)\n> \n> Whoops.  You didn't 'git add file2' (before the mv) or 'git add file3'\n> (after the mv), or use commit -a, so what you've committed is the\n> *old* content of file2 under the name file3.  The *new* content of\n> file2 is still uncommitted in your work tree under the name file3.\n> This is why git can detect the move.  (The 100% is a good clue: it\n> means the old and new files are 100% identical.)\n\nAh.  That explains everything.  Thanks.  (I thought git mv was \nequivalent to git rm followed by git add.  But it's not.)\n\n> Artificial tests like this are useless anyway.\n\nYes, I know.  This was not intended to be a real-world example.  I was \njust trying to understand the heuristics that git uses to track filename \nchanges, and in particular, how much a file could change before git \ndecided it was a different file.  When I got to zero shared lines \nbetween old and new it was clear that I was missing something \nfundamental :-)\n\nSo... how *does* git decide when two blobs are different blobs and when \nthey are the same blob with mods?  I asked this question before and was \npointed to the diffcore docs, but that didn't really clear things up.  \nThat just describes all the different ways git can do diffs, not the \nactual heuristics that git uses to track content.\n\nrg\n"},{"id":"133509","messageId":"32541b131002031147r367ee08fxc64c4c54165953a3@mail.gmail.com","threadId":"22508","inReplyTo":"ron1-5F71CB.11234903022010@news.gmane.org","subject":"Re: git-mv redux: there must be something else going on","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-02-03T19:47:33Z","receivedAt":"2010-02-03T19:47:33Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Feb 3, 2010 at 2:23 PM, Ron Garret <ron1@flownet.com> wrote:\n> In article\n> Ah.  That explains everything.  Thanks.  (I thought git mv was\n> equivalent to git rm followed by git add.  But it's not.)\n\nI suppose in this case it's not.  The only difference is when your\nwork tree differs from your index, though, and it's to be expected\nthat 'git rm', in removing things from the index, would lose your\nability to track those differences.\n\n> So... how *does* git decide when two blobs are different blobs and when\n> they are the same blob with mods?  I asked this question before and was\n> pointed to the diffcore docs, but that didn't really clear things up.\n> That just describes all the different ways git can do diffs, not the\n> actual heuristics that git uses to track content.\n\nIf you really want to know the details, looking at the code really is\nprobably the best solution; it's not even that long.\n\nThe short version is that git chooses a set of candidate blobs, then\ndiffs them and figures out a percentage similarity between each pair.\n(A simple way to think of the similarity index is \"how long is the\ndiff compared to the file itself?\"  If the diff is of length zero, the\nsimilarity is 100%, and so on.) If the similarity is greater than a\ncertain threshold, then it's considered to be the same file.\n\nChoosing the set of candidates is actually the more interesting\nproblem, since detecting moves using the above algorithm is O(n^2)\nwith the number of candidates.  That's why 'git diff' and 'git log'\ndon't do it at all by default.\n\nIf you provide -M, the set of candidates is the set of files that were\nremoved/modified and the set of files that were added.  (Added files\nare compared against removed/modified files, iirc.)  Normally that's a\nvery short list.  With -C, you need to compare all\nadded/removed/modified files with all others, which is slightly more\nwork.  With --find-copies-harder, it becomes potentially a *lot* of\nwork.\n\nHave fun,\n\nAvery\n"},{"id":"133514","messageId":"alpine.LFD.2.00.1002031436490.1681@xanadu.home","threadId":"22508","inReplyTo":"ron1-5F71CB.11234903022010@news.gmane.org","subject":"Re: git-mv redux: there must be something else going on","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-02-03T19:53:23Z","receivedAt":"2010-02-03T19:53:23Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 3 Feb 2010, Ron Garret wrote:\n\n> So... how *does* git decide when two blobs are different blobs and when \n> they are the same blob with mods?  I asked this question before and was \n> pointed to the diffcore docs, but that didn't really clear things up.  \n> That just describes all the different ways git can do diffs, not the \n> actual heuristics that git uses to track content.\n\nYes, those same heuristics are used to make the decision.\n\n|The second transformation in the chain is diffcore-break, and is\n|controlled by the -B option to the 'git diff-{asterisk}' commands.  \n|This is used to detect a filepair that represents \"complete rewrite\" \n|and break such filepair into two filepairs that represent delete and\n|create.\n|[...]\n\n|This transformation is used to detect renames and copies, and is\n|controlled by the -M option (to detect renames) and the -C option\n|(to detect copies as well) to the 'git diff-{asterisk}' commands.  \n|[...]\n\nNote that you may use the -B, -C, -M and --find-copies-harder arguments \nwith log as well as diff commands even if there is no actual diff \noutput.  So the explanation is really in that document even if simple \nrename detection is concerned only by a fraction of what is said there.\n\nAnd Git can detect copied files too.\n\nThose semantics are not stored in the repository so they can be improved \nor even changed after the facts.\n\n\nNicolas\n"},{"id":"133518","messageId":"4B69D897.2060908@pcharlan.com","threadId":"22508","inReplyTo":"32541b131002031048i26d166d9w3567a60515235c34@mail.gmail.com","subject":"Re: git-mv redux: there must be something else going on","fromName":"Pete Harlan","fromEmail":"pgit@pcharlan.com","sentAt":"2010-02-03T20:12:07Z","receivedAt":"2010-02-03T20:12:07Z","isPatch":false,"sender":{"key":"pgit@pcharlan.com","avatar":null},"body":"On 02/03/2010 10:48 AM, Avery Pennarun wrote:\n>> [ron@mickey:~/devel/gittest]$ git mv file2 file3\n>> [ron@mickey:~/devel/gittest]$ git commit -m 'letters->numbers'\n>> [master ae3f6d4] letters->numbers\n>>  1 files changed, 0 insertions(+), 0 deletions(-)\n>>  rename file2 => file3 (100%)\n> \n> Whoops.  You didn't 'git add file2' (before the mv) or 'git add file3'\n> (after the mv), or use commit -a, so what you've committed is the\n> *old* content of file2 under the name file3.  The *new* content of\n> file2 is still uncommitted in your work tree under the name file3.\n\nIt may be reasonable for \"git mv foo bar\" to print a helpful message to\nthe user if foo has un-checked-in changes, similarly to what \"git rm\" does.\n\nUnlike \"git rm\", \"git mv\" could still perform the operation even without\n\"-f\", but the semantics of \"git mv\" differ enough from plain \"mv\" that a\nshort blurb from Git in that case might help.\n\n--Pete\n"},{"id":"133520","messageId":"ron1-34F9C6.12273203022010@news.gmane.org","threadId":"22508","inReplyTo":"alpine.LFD.2.00.1002031436490.1681@xanadu.home","subject":"Re: git-mv redux: there must be something else going on","fromName":"Ron Garret","fromEmail":"ron1@flownet.com","sentAt":"2010-02-03T20:27:32Z","receivedAt":"2010-02-03T20:27:32Z","isPatch":false,"sender":{"key":"ron1@flownet.com","avatar":null},"body":"In article <alpine.LFD.2.00.1002031436490.1681@xanadu.home>,\n Nicolas Pitre <nico@fluxnic.net> wrote:\n\n> On Wed, 3 Feb 2010, Ron Garret wrote:\n> \n> > So... how *does* git decide when two blobs are different blobs and when \n> > they are the same blob with mods?  I asked this question before and was \n> > pointed to the diffcore docs, but that didn't really clear things up.  \n> > That just describes all the different ways git can do diffs, not the \n> > actual heuristics that git uses to track content.\n> \n> Yes, those same heuristics are used to make the decision.\n> \n> |The second transformation in the chain is diffcore-break, and is\n> |controlled by the -B option to the 'git diff-{asterisk}' commands.  \n> |This is used to detect a filepair that represents \"complete rewrite\" \n> |and break such filepair into two filepairs that represent delete and\n> |create.\n> |[...]\n> \n> |This transformation is used to detect renames and copies, and is\n> |controlled by the -M option (to detect renames) and the -C option\n> |(to detect copies as well) to the 'git diff-{asterisk}' commands.  \n> |[...]\n> \n> Note that you may use the -B, -C, -M and --find-copies-harder arguments \n> with log as well as diff commands even if there is no actual diff \n> output.  So the explanation is really in that document even if simple \n> rename detection is concerned only by a fraction of what is said there.\n> \n> And Git can detect copied files too.\n> \n> Those semantics are not stored in the repository so they can be improved \n> or even changed after the facts.\n\nOK, on closer reading I see that the information is there, but it's well \nhidden :-)  (For example, the -M option takes an optional numerical \nargument so you can tweak how much similarity is needed to be considered \na move.  But the docs for git log don't mention this.  It's buried deep \nin the git diffcore docs.  But yes, it's there.)\n\nSo I think I'm beginning to understand how this works, but that leads me \nto another question: it seems to me that there are potential screw cases \nfor this purely content-based system of tracking files.  For example, \nsuppose I have a directory full of sample config files, all of which are \nsimilar to each other.  Will that cause diffcore to get confused?\n\nFeel free to treat that as a rhetorical question because obviously I can \n(and probably should) get the answer by trying it.\n\nThanks!\nrg\n"},{"id":"133523","messageId":"ron1-DFA9D6.12301403022010@news.gmane.org","threadId":"22508","inReplyTo":"32541b131002031147r367ee08fxc64c4c54165953a3@mail.gmail.com","subject":"Re: git-mv redux: there must be something else going on","fromName":"Ron Garret","fromEmail":"ron1@flownet.com","sentAt":"2010-02-03T20:30:14Z","receivedAt":"2010-02-03T20:30:14Z","isPatch":false,"sender":{"key":"ron1@flownet.com","avatar":null},"body":"In article \n<32541b131002031147r367ee08fxc64c4c54165953a3@mail.gmail.com>,\n Avery Pennarun <apenwarr@gmail.com> wrote:\n\n> On Wed, Feb 3, 2010 at 2:23 PM, Ron Garret <ron1@flownet.com> wrote:\n> > In article\n> > Ah.  That explains everything.  Thanks.  (I thought git mv was\n> > equivalent to git rm followed by git add.  But it's not.)\n> \n> I suppose in this case it's not.  The only difference is when your\n> work tree differs from your index, though, and it's to be expected\n> that 'git rm', in removing things from the index, would lose your\n> ability to track those differences.\n> \n> > So... how *does* git decide when two blobs are different blobs and when\n> > they are the same blob with mods?  I asked this question before and was\n> > pointed to the diffcore docs, but that didn't really clear things up.\n> > That just describes all the different ways git can do diffs, not the\n> > actual heuristics that git uses to track content.\n> \n> If you really want to know the details, looking at the code really is\n> probably the best solution; it's not even that long.\n> \n> The short version is that git chooses a set of candidate blobs, then\n> diffs them and figures out a percentage similarity between each pair.\n> (A simple way to think of the similarity index is \"how long is the\n> diff compared to the file itself?\"  If the diff is of length zero, the\n> similarity is 100%, and so on.) If the similarity is greater than a\n> certain threshold, then it's considered to be the same file.\n> \n> Choosing the set of candidates is actually the more interesting\n> problem, since detecting moves using the above algorithm is O(n^2)\n> with the number of candidates.  That's why 'git diff' and 'git log'\n> don't do it at all by default.\n> \n> If you provide -M, the set of candidates is the set of files that were\n> removed/modified and the set of files that were added.  (Added files\n> are compared against removed/modified files, iirc.)  Normally that's a\n> very short list.  With -C, you need to compare all\n> added/removed/modified files with all others, which is slightly more\n> work.  With --find-copies-harder, it becomes potentially a *lot* of\n> work.\n\nThanks!  That clarifies a lot.\n\nrg\n"},{"id":"133526","messageId":"ron1-176898.12310803022010@news.gmane.org","threadId":"22508","inReplyTo":"ron1-34F9C6.12273203022010@news.gmane.org","subject":"Re: git-mv redux: there must be something else going on","fromName":"Ron Garret","fromEmail":"ron1@flownet.com","sentAt":"2010-02-03T20:31:08Z","receivedAt":"2010-02-03T20:31:08Z","isPatch":false,"sender":{"key":"ron1@flownet.com","avatar":null},"body":"In article <ron1-34F9C6.12273203022010@news.gmane.org>,\n Ron Garret <ron1@flownet.com> wrote:\n\n> In article <alpine.LFD.2.00.1002031436490.1681@xanadu.home>,\n>  Nicolas Pitre <nico@fluxnic.net> wrote:\n> \n> > On Wed, 3 Feb 2010, Ron Garret wrote:\n> > \n> > > So... how *does* git decide when two blobs are different blobs and when \n> > > they are the same blob with mods?  I asked this question before and was \n> > > pointed to the diffcore docs, but that didn't really clear things up.  \n> > > That just describes all the different ways git can do diffs, not the \n> > > actual heuristics that git uses to track content.\n> > \n> > Yes, those same heuristics are used to make the decision.\n> > \n> > |The second transformation in the chain is diffcore-break, and is\n> > |controlled by the -B option to the 'git diff-{asterisk}' commands.  \n> > |This is used to detect a filepair that represents \"complete rewrite\" \n> > |and break such filepair into two filepairs that represent delete and\n> > |create.\n> > |[...]\n> > \n> > |This transformation is used to detect renames and copies, and is\n> > |controlled by the -M option (to detect renames) and the -C option\n> > |(to detect copies as well) to the 'git diff-{asterisk}' commands.  \n> > |[...]\n> > \n> > Note that you may use the -B, -C, -M and --find-copies-harder arguments \n> > with log as well as diff commands even if there is no actual diff \n> > output.  So the explanation is really in that document even if simple \n> > rename detection is concerned only by a fraction of what is said there.\n> > \n> > And Git can detect copied files too.\n> > \n> > Those semantics are not stored in the repository so they can be improved \n> > or even changed after the facts.\n> \n> OK, on closer reading I see that the information is there, but it's well \n> hidden :-)  (For example, the -M option takes an optional numerical \n> argument so you can tweak how much similarity is needed to be considered \n> a move.  But the docs for git log don't mention this.  It's buried deep \n> in the git diffcore docs.  But yes, it's there.)\n> \n> So I think I'm beginning to understand how this works, but that leads me \n> to another question: it seems to me that there are potential screw cases \n> for this purely content-based system of tracking files.  For example, \n> suppose I have a directory full of sample config files, all of which are \n> similar to each other.  Will that cause diffcore to get confused?\n> \n> Feel free to treat that as a rhetorical question because obviously I can \n> (and probably should) get the answer by trying it.\n\nActually, I think the answer is in Avery's post in another branch of \nthis thread.\n\nrg\n"},{"id":"133522","messageId":"ron1-A681F2.12340503022010@news.gmane.org","threadId":"22508","inReplyTo":"4B69D897.2060908@pcharlan.com","subject":"Re: git-mv redux: there must be something else going on","fromName":"Ron Garret","fromEmail":"ron1@flownet.com","sentAt":"2010-02-03T20:34:05Z","receivedAt":"2010-02-03T20:34:05Z","isPatch":false,"sender":{"key":"ron1@flownet.com","avatar":null},"body":"In article <4B69D897.2060908@pcharlan.com>,\n Pete Harlan <pgit@pcharlan.com> wrote:\n\n> On 02/03/2010 10:48 AM, Avery Pennarun wrote:\n> >> [ron@mickey:~/devel/gittest]$ git mv file2 file3\n> >> [ron@mickey:~/devel/gittest]$ git commit -m 'letters->numbers'\n> >> [master ae3f6d4] letters->numbers\n> >>  1 files changed, 0 insertions(+), 0 deletions(-)\n> >>  rename file2 => file3 (100%)\n> > \n> > Whoops.  You didn't 'git add file2' (before the mv) or 'git add file3'\n> > (after the mv), or use commit -a, so what you've committed is the\n> > *old* content of file2 under the name file3.  The *new* content of\n> > file2 is still uncommitted in your work tree under the name file3.\n> \n> It may be reasonable for \"git mv foo bar\" to print a helpful message to\n> the user if foo has un-checked-in changes, similarly to what \"git rm\" does.\n> \n> Unlike \"git rm\", \"git mv\" could still perform the operation even without\n> \"-f\", but the semantics of \"git mv\" differ enough from plain \"mv\" that a\n> short blurb from Git in that case might help.\n\nI think that a simple tweak to the docs would be enough.  Right now it \nsays:\n\n\"The index is updated after successful completion, but the change must \nstill be committed.\"\n\nI'm pretty sure I would have been less confused if it had said something \nlike:\n\n\"The index is updated to reflect the new name of the file, but NOT any \nnew content that file may contain.  Changed content must be added to the \nindex separately with git add, and all changes must still be commited.\"\n\nrg\n"},{"id":"133527","messageId":"32541b131002031240p6b67536ame6b69c6d662a7968@mail.gmail.com","threadId":"22508","inReplyTo":"ron1-34F9C6.12273203022010@news.gmane.org","subject":"Re: git-mv redux: there must be something else going on","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-02-03T20:40:02Z","receivedAt":"2010-02-03T20:40:02Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Feb 3, 2010 at 3:27 PM, Ron Garret <ron1@flownet.com> wrote:\n> So I think I'm beginning to understand how this works, but that leads me\n> to another question: it seems to me that there are potential screw cases\n> for this purely content-based system of tracking files.  For example,\n> suppose I have a directory full of sample config files, all of which are\n> similar to each other.  Will that cause diffcore to get confused?\n\nCases like that are always confusing, even to humans.  Person A\nrenames X to Y, but at the same time creates Z which is almost\nidentical.  Person B patches X, then merges in person A's changes.\n\nWhat do you expect to happen?  Should Y be changed, because that's the\nfile X was moved from?  Or should we change Z, because it's almost the\nsame content anyway?  Or maybe we should change both, since a change\nto the old X is probably intended to affect the copied *content* that\nended up in both Y and Z?\n\nSimply storing whether person A has renamed vs. copied vs. added a\nfile makes the answer to the \"what do you expect to happen\" question\nmore obvious, but fails to answer the \"what *should* happen\" question.\n Thus it's more of a distraction than a feature.  It took a while for\nme to accept this, but once I did, I realized that git's behaviour has\nstill never caused me a problem in real life, despite repeated file\nrenames and complicated merges.\n\nIn contrast, svn's explicit rename tracking has shot me in the foot\nnumerous times.  (svn remembers when I delete file X and then\nsubsequently re-add it with the same content.  So if I merge in\nsomeone's change to the *old* file X, it barfs because omg omg that's\na totally different file X and it can't possibly figure out what to\ndo.  Gee, thanks.  It's also hopelessly incompetent at handling\n\"renames\" in which a newbie developer didn't know to use svn mv, but\ninstead used svn rm, mv, and svn add.)\n\nHave fun,\n\nAvery\n"},{"id":"133529","messageId":"alpine.LFD.2.00.1002031533560.1681@xanadu.home","threadId":"22508","inReplyTo":"ron1-34F9C6.12273203022010@news.gmane.org","subject":"Re: git-mv redux: there must be something else going on","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2010-02-03T20:44:44Z","receivedAt":"2010-02-03T20:44:44Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 3 Feb 2010, Ron Garret wrote:\n\n> OK, on closer reading I see that the information is there, but it's well \n> hidden :-)  (For example, the -M option takes an optional numerical \n> argument so you can tweak how much similarity is needed to be considered \n> a move.  But the docs for git log don't mention this.  It's buried deep \n> in the git diffcore docs.  But yes, it's there.)\n\nThe doc is indeed not perfect.  Probably the -M option and friends could \nbe listed again in the git-log and git-diff pages with a more casual \nexplanation.\n\n> So I think I'm beginning to understand how this works, but that leads me \n> to another question: it seems to me that there are potential screw cases \n> for this purely content-based system of tracking files.  For example, \n> suppose I have a directory full of sample config files, all of which are \n> similar to each other.  Will that cause diffcore to get confused?\n\nThere are ways to fool the heuristics indeed.  But overall it is still \nmore reliable than manually having to record the rename into the tool \nsince humans are known for screwing these things up more often than \nmachines.  And again the heuristics can be modified after the fact if \nneeded, unlike the manually recorded false renames (or lack of rename \nrecord) which will remain wrong unless another manual correction is \napplied to the database.\n\n\nNicolas\n"},{"id":"133536","messageId":"c43166fa73391a40b43c27153ec142121fdb71d1.1265231310.git.trast@student.ethz.ch","threadId":"22508","inReplyTo":"ron1-A681F2.12340503022010@news.gmane.org","subject":"[PATCH] Documentation: clarify git-mv behaviour wrt dirty files","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2010-02-03T21:12:12Z","receivedAt":"2010-02-03T21:12:12Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Clearly point out that the rename happens separately for worktree and\nindex.  This confused users, as they are apparently told that git-mv\n== git-rm && mv && git-add, which it is not.\n\nWhile there, move the synposis to the synopsis section, which so far\nwas rather useless, and reword the first sentence to eliminate the\nmentions of 'script'.\n\nSigned-off-by: Thomas Rast <trast@student.ethz.ch>\n---\n\nRon, please don't drop the Cc lists, it's customary around here to Cc\neveryone involved so far.\n\nOn Wednesday 03 February 2010 21:34:05 you wrote:\n> In article <4B69D897.2060908@pcharlan.com>,\n>  Pete Harlan <pgit@pcharlan.com> wrote:\n> > Unlike \"git rm\", \"git mv\" could still perform the operation even without\n> > \"-f\", but the semantics of \"git mv\" differ enough from plain \"mv\" that a\n> > short blurb from Git in that case might help.\n> \n> I think that a simple tweak to the docs would be enough.  Right now it \n> says:\n> \n> \"The index is updated after successful completion, but the change must \n> still be committed.\"\n> \n> I'm pretty sure I would have been less confused if it had said something \n> like:\n> \n> \"The index is updated to reflect the new name of the file, but NOT any \n> new content that file may contain.  Changed content must be added to the \n> index separately with git add, and all changes must still be commited.\"\n\nHow about this change instead, which formulates it in terms of what\ndoes happen, instead of what does not.\n\nBTW, I'm wondering whether the \"move or rename\" distinction is really\nworth it.  Does the user care?  I always figured it was a technical\ndetail whether rename() works or you actually need to move anything.\n\n\n Documentation/git-mv.txt |   14 +++++++-------\n 1 files changed, 7 insertions(+), 7 deletions(-)\n\ndiff --git a/Documentation/git-mv.txt b/Documentation/git-mv.txt\nindex bdcb585..eff11b7 100644\n--- a/Documentation/git-mv.txt\n+++ b/Documentation/git-mv.txt\n@@ -8,22 +8,22 @@ git-mv - Move or rename a file, a directory, or a symlink\n \n SYNOPSIS\n --------\n-'git mv' <options>... <args>...\n+'git mv' [-f] [-n] <source> <destination>\n+'git mv' [-f] [-n] [-k] <source>... <destination directory>\n \n DESCRIPTION\n -----------\n-This script is used to move or rename a file, directory or symlink.\n-\n- git mv [-f] [-n] <source> <destination>\n- git mv [-f] [-n] [-k] <source> ... <destination directory>\n+'git-mv' renames files, directories, and symlinks in worktree and\n+index.\n \n In the first form, it renames <source>, which must exist and be either\n a file, symlink or directory, to <destination>.\n In the second form, the last argument has to be an existing\n directory; the given sources will be moved into this directory.\n \n-The index is updated after successful completion, but the change must still be\n-committed.\n+For every renamed file or symlink, the worktree and index contents are\n+renamed separately, preserving both staged and unstaged changes.  You\n+will still have to commit the rename.\n \n OPTIONS\n -------\n-- \n1.7.0.rc1.166.g7cae7\n"},{"id":"133541","messageId":"7v3a1idlvg.fsf@alter.siamese.dyndns.org","threadId":"22508","inReplyTo":"c43166fa73391a40b43c27153ec142121fdb71d1.1265231310.git.trast@student.ethz.ch","subject":"Re: [PATCH] Documentation: clarify git-mv behaviour wrt dirty files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2010-02-03T21:56:19Z","receivedAt":"2010-02-03T21:56:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Rast <trast@student.ethz.ch> writes:\n\n> Clearly point out that the rename happens separately for worktree and\n> index.  This confused users, as they are apparently told that git-mv\n> == git-rm && mv && git-add, which it is not.\n\nI may be confused too as I had to read these three lines three times and I\ndo not think these two sentences mesh well together.\n\nWhat happens with \"git mv A B\" is that it moves a work tree file A to B\nand moves the index entry for A to B, hence all of:\n\n (1) the fact that you do not have A anymore;\n\n (2) the fact that you now have B instead; and\n\n (3) the fact that your work tree file B (which used to be A) has changes\n     from its corresponding index entry\n\nare _consistently_ kept between the work tree and the index.\n\nI don't think \"happens separately for\" makes sense.  At best, it is an\nimplementation detail that doesn't help users understand what the command\ndoes and what it is used for better.\n\nOf course, it is different from\n\n    \"git rm -f --cached A && mv A B && git add B\"\n\nwhich would add changes that you were not prepared to add (i.e. you had\noutput from \"git diff A\" before you started).  I think that was a buggy\nway old scripted version of \"git mv\" used to work, by the way.\n\n> While there, move the synposis to the synopsis section, which so far\n> was rather useless, and reword the first sentence to eliminate the\n> mentions of 'script'.\n\nThat's a good change regardless.\n\n> +For every renamed file or symlink, the worktree and index contents are\n> +renamed separately, preserving both staged and unstaged changes....\n\nI'd just say:\n\n    While renaming paths, changes in the files in the work tree that you\n    have not added are preserved.\n\n> +....  You\n> +will still have to commit the rename.\n\nI don't understand why you want to say \"You will still have to commit the\nrename\" here.  It is like saying in \"git add\" manpage that \"You will still\nhave to commit the added contents\" because \"add\" only affects the index\nand does not make a commit.  Drop it.\n"},{"id":"133547","messageId":"ron1-9FA846.14332803022010@news.gmane.org","threadId":"22508","inReplyTo":"32541b131002031240p6b67536ame6b69c6d662a7968@mail.gmail.com","subject":"Re: git-mv redux: there must be something else going on","fromName":"Ron Garret","fromEmail":"ron1@flownet.com","sentAt":"2010-02-03T22:33:28Z","receivedAt":"2010-02-03T22:33:28Z","isPatch":false,"sender":{"key":"ron1@flownet.com","avatar":null},"body":"In article \n<32541b131002031240p6b67536ame6b69c6d662a7968@mail.gmail.com>,\n Avery Pennarun <apenwarr@gmail.com> wrote:\n\n> On Wed, Feb 3, 2010 at 3:27 PM, Ron Garret <ron1@flownet.com> wrote:\n> > So I think I'm beginning to understand how this works, but that leads me\n> > to another question: it seems to me that there are potential screw cases\n> > for this purely content-based system of tracking files.  For example,\n> > suppose I have a directory full of sample config files, all of which are\n> > similar to each other.  Will that cause diffcore to get confused?\n> \n> Cases like that are always confusing, even to humans.  Person A\n> renames X to Y, but at the same time creates Z which is almost\n> identical.  Person B patches X, then merges in person A's changes.\n> \n> What do you expect to happen?  Should Y be changed, because that's the\n> file X was moved from?  Or should we change Z, because it's almost the\n> same content anyway?  Or maybe we should change both, since a change\n> to the old X is probably intended to affect the copied *content* that\n> ended up in both Y and Z?\n> \n> Simply storing whether person A has renamed vs. copied vs. added a\n> file makes the answer to the \"what do you expect to happen\" question\n> more obvious, but fails to answer the \"what *should* happen\" question.\n>  Thus it's more of a distraction than a feature.  It took a while for\n> me to accept this, but once I did, I realized that git's behaviour has\n> still never caused me a problem in real life, despite repeated file\n> renames and complicated merges.\n> \n> In contrast, svn's explicit rename tracking has shot me in the foot\n> numerous times.  (svn remembers when I delete file X and then\n> subsequently re-add it with the same content.  So if I merge in\n> someone's change to the *old* file X, it barfs because omg omg that's\n> a totally different file X and it can't possibly figure out what to\n> do.  Gee, thanks.  It's also hopelessly incompetent at handling\n> \"renames\" in which a newbie developer didn't know to use svn mv, but\n> instead used svn rm, mv, and svn add.)\n\nHere's a realistic case where keeping explicit track of renames could be \nuseful.\n\nA and B start with a file named config.  A and B both make edits.  In \naddition, B renames config to be config1 and creates a new, very similar \nfile called config2.  B then merges from A with the expectation that B's \nedits to config would end up in config1 and not config2.  It seems to me \nthat without tracking renames, it would be luck of the draw which file \nthe patch got applied to.\n\nrg\n"},{"id":"133551","messageId":"32541b131002031518t1017d351xcf9071f0a937474e@mail.gmail.com","threadId":"22508","inReplyTo":"ron1-9FA846.14332803022010@news.gmane.org","subject":"Re: git-mv redux: there must be something else going on","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-02-03T23:18:59Z","receivedAt":"2010-02-03T23:18:59Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Feb 3, 2010 at 5:33 PM, Ron Garret <ron1@flownet.com> wrote:\n> Here's a realistic case where keeping explicit track of renames could be\n> useful.\n>\n> A and B start with a file named config.  A and B both make edits.  In\n> addition, B renames config to be config1 and creates a new, very similar\n> file called config2.  B then merges from A with the expectation that B's\n> edits to config would end up in config1 and not config2.  It seems to me\n> that without tracking renames, it would be luck of the draw which file\n> the patch got applied to.\n\nThe problem is that this single \"realistic case\" is not actually very\ncommon, and it's dwarfed by the other realistic cases: developer\nforgets to use 'git mv' to rename the file; developer accidentally\ndeletes a file, commits, and then readds it later; etc.\n\nHave I been bitten by exactly your example?  Yup.  But I've been\nbitten by lots of other related things too, and explicit rename\ntracking (at least in svn) has quite frequently made the problems\n*worse*.  In my personal experience, git screws up less often.  The\nfact that it's also elegant is a nice bonus too :)\n\nMore about this: http://marc.info/?l=git&m=114123702826251\n\nHave fun,\n\nAvery\n"},{"id":"133553","messageId":"76718491002031555i2c1558f9qe0c97d07ceb86bb6@mail.gmail.com","threadId":"22508","inReplyTo":"32541b131002031518t1017d351xcf9071f0a937474e@mail.gmail.com","subject":"Re: git-mv redux: there must be something else going on","fromName":"Jay Soffian","fromEmail":"jaysoffian@gmail.com","sentAt":"2010-02-03T23:55:32Z","receivedAt":"2010-02-03T23:55:32Z","isPatch":false,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"On Wed, Feb 3, 2010 at 6:18 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n> More about this: http://marc.info/?l=git&m=114123702826251\n\nI think the canonical email on the subject is this one:\n\nhttp://article.gmane.org/gmane.comp.version-control.git/217\n\n:-)\n\nj.\n"},{"id":"133554","messageId":"ron1-C4BB38.16101603022010@news.gmane.org","threadId":"22508","inReplyTo":"76718491002031555i2c1558f9qe0c97d07ceb86bb6@mail.gmail.com","subject":"Re: git-mv redux: there must be something else going on","fromName":"Ron Garret","fromEmail":"ron1@flownet.com","sentAt":"2010-02-04T00:10:16Z","receivedAt":"2010-02-04T00:10:16Z","isPatch":false,"sender":{"key":"ron1@flownet.com","avatar":null},"body":"In article \n<76718491002031555i2c1558f9qe0c97d07ceb86bb6@mail.gmail.com>,\n Jay Soffian <jaysoffian@gmail.com> wrote:\n\n> On Wed, Feb 3, 2010 at 6:18 PM, Avery Pennarun <apenwarr@gmail.com> wrote:\n> > More about this: http://marc.info/?l=git&m=114123702826251\n> \n> I think the canonical email on the subject is this one:\n> \n> http://article.gmane.org/gmane.comp.version-control.git/217\n> \n\nThe upshot seems to be this:\n\n> And that \"where did this come from\" decision should be done at _search_ \n> time, not commit time.\n\nAnd I'm mostly convinced, except for the one screw case that I outlined \nabove.  In that case the search-time result is ambiguous, and file \ntracking information could be used to resolve the ambiguity.  But it \ncertainly does seem like a rare enough situation that it's not worth \nworrying about.\n\nI think I'm starting to git it :-)\n\nrg\n"},{"id":"133555","messageId":"ron1-1F86D7.16103803022010@news.gmane.org","threadId":"22508","inReplyTo":"32541b131002031518t1017d351xcf9071f0a937474e@mail.gmail.com","subject":"Re: git-mv redux: there must be something else going on","fromName":"Ron Garret","fromEmail":"ron1@flownet.com","sentAt":"2010-02-04T00:10:38Z","receivedAt":"2010-02-04T00:10:38Z","isPatch":false,"sender":{"key":"ron1@flownet.com","avatar":null},"body":"In article \n<32541b131002031518t1017d351xcf9071f0a937474e@mail.gmail.com>,\n Avery Pennarun <apenwarr@gmail.com> wrote:\n\n> On Wed, Feb 3, 2010 at 5:33 PM, Ron Garret <ron1@flownet.com> wrote:\n> > Here's a realistic case where keeping explicit track of renames could be\n> > useful.\n> >\n> > A and B start with a file named config.  A and B both make edits.  In\n> > addition, B renames config to be config1 and creates a new, very similar\n> > file called config2.  B then merges from A with the expectation that B's\n> > edits to config would end up in config1 and not config2.  It seems to me\n> > that without tracking renames, it would be luck of the draw which file\n> > the patch got applied to.\n> \n> The problem is that this single \"realistic case\" is not actually very\n> common, and it's dwarfed by the other realistic cases: developer\n> forgets to use 'git mv' to rename the file; developer accidentally\n> deletes a file, commits, and then readds it later; etc.\n\nMakes sense.\n\n> Have I been bitten by exactly your example?  Yup.  But I've been\n> bitten by lots of other related things too, and explicit rename\n> tracking (at least in svn) has quite frequently made the problems\n> *worse*.  In my personal experience, git screws up less often.  The\n> fact that it's also elegant is a nice bonus too :)\n> \n> More about this: http://marc.info/?l=git&m=114123702826251\n\nThanks, that's a great read!\n\nrg\n"},{"id":"133558","messageId":"7vvded4yi3.fsf@alter.siamese.dyndns.org","threadId":"22508","inReplyTo":"ron1-9FA846.14332803022010@news.gmane.org","subject":"Re: git-mv redux: there must be something else going on","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2010-02-04T00:48:20Z","receivedAt":"2010-02-04T00:48:20Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ron Garret <ron1@flownet.com> writes:\n\n> A and B start with a file named config.  A and B both make edits.  In \n> addition, B renames config to be config1 and creates a new, very similar \n> file called config2.  B then merges from A with the expectation that B's \n> edits to config would end up in config1 and not config2.  It seems to me \n> that without tracking renames, it would be luck of the draw which file \n> the patch got applied to.\n\nI don't think the above is necessarily \"rename\" issue, but touches an\ninteresting point -- it is so \"interesting\" to the point that no sane SCM\nwould even consider that is a problem they need to solve.\n\nIf config1 and config2 are about two different ways to configure the\nsoftware (e.g. two different build for different customers), and change\nmade by A was to accomodate new configuration option made in the upstream,\nB might even want to have that addition reflected in _both_ of his\nconfiguration files, config1 and config2.\n\nEarlier in this message, I said that this is not an issue SCM should even\nbe solving, because a sane way to handle this would _not_ be to copy and\nedit config1/config2 and keep track of them in SCM; instead, saner people\nwould maintain a build procedure (e.g. Makefile target) to transform the\ntemplate \"config\" into necessary \"config1\" and \"config2\" customized\nvariants.\n"}]}