{"thread":{"id":"8668","subject":"Basename matching during rename/copy detection","startedAt":"2007-06-21T03:06:22Z","lastAt":"2007-06-24T22:23:23Z","messageCount":44,"participants":["Shawn O. Pearce","Junio C Hamano","Linus Torvalds","Andy Parkins","Johannes Schindelin","Matthieu Moy","Jeff King","Steven Grimm","Johannes Sixt","David Kastrup","Aidan Van Dyk","René Scharfe"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"45457","messageId":"20070621030622.GD8477@spearce.org","threadId":"8668","inReplyTo":null,"subject":"Basename matching during rename/copy detection","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-06-21T03:06:22Z","receivedAt":"2007-06-21T03:06:22Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"So Govind Salinas has found an interesting case in the rename\ndetection code:\n\n  $ git clone git://repo.or.cz/Widgit.git\n  $ git diff -M --raw -r 192e^ 192e | grep .resx\n  :100755 000000 4c8ab79... 0000000... D  Form1.resx\n  :100755 100755 9e70146... 9e70146... R100       CommitViewer.resx       UI/CommitViewer.resx\n  :100755 100755 90929fd... b40ff98... C091       RepoManager.resx        UI/Form1.resx\n  :100755 100755 90929fd... 90929fd... C100       PreferencesEditor.resx  UI/PreferencesEditor.resx\n  :100755 100755 90929fd... 90929fd... R100       PreferencesEditor.resx  UI/RepoManager.resx\n  :100755 100755 90929fd... 8535007... R097       RepoManager.resx        UI/RepoTreeView.resx\n\nIn this case several files had identical old images, and some\nkept that old image during the rename.  Unfortunately because of\nthe ordering of the files in the tree Git has decided to \"rename\"\nthe PreferencesEditor.resx file to UI/RepoManager.resx, rather than\nrenaming RepoManager.resx to UI/RepoManager.resx.  Go Git.\n\nI'm wondering if we shouldn't play the game of trying to match\ndelete/add pairs up by not only similarity, but also by path\nbasename.  In the case above its exactly what Govind thought should\nhappen; he moved the file from one directory to another, and didn't\neven change its content during the move.  But Git decided \"better\"\nto use a totally different file in the \"rename\".\n\n-- \nShawn.\n"},{"id":"45458","messageId":"7vsl8m3sph.fsf@assigned-by-dhcp.pobox.com","threadId":"8668","inReplyTo":"20070621030622.GD8477@spearce.org","subject":"Re: Basename matching during rename/copy detection","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-06-21T03:13:46Z","receivedAt":"2007-06-21T03:13:46Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> So Govind Salinas has found an interesting case in the rename\n> detection code:\n>\n>   $ git clone git://repo.or.cz/Widgit.git\n>   $ git diff -M --raw -r 192e^ 192e | grep .resx\n>   :100755 000000 4c8ab79... 0000000... D  Form1.resx\n>   :100755 100755 9e70146... 9e70146... R100       CommitViewer.resx       UI/CommitViewer.resx\n>   :100755 100755 90929fd... b40ff98... C091       RepoManager.resx        UI/Form1.resx\n>   :100755 100755 90929fd... 90929fd... C100       PreferencesEditor.resx  UI/PreferencesEditor.resx\n>   :100755 100755 90929fd... 90929fd... R100       PreferencesEditor.resx  UI/RepoManager.resx\n>   :100755 100755 90929fd... 8535007... R097       RepoManager.resx        UI/RepoTreeView.resx\n>\n> In this case several files had identical old images, and some\n> kept that old image during the rename.  Unfortunately because of\n> the ordering of the files in the tree Git has decided to \"rename\"\n> the PreferencesEditor.resx file to UI/RepoManager.resx, rather than\n> renaming RepoManager.resx to UI/RepoManager.resx.  Go Git.\n>\n> I'm wondering if we shouldn't play the game of trying to match\n> delete/add pairs up by not only similarity, but also by path\n> basename.  In the case above its exactly what Govind thought should\n> happen; he moved the file from one directory to another, and didn't\n> even change its content during the move.  But Git decided \"better\"\n> to use a totally different file in the \"rename\".\n\nActually, git did not decide anything, and certainly not better.\n\nHaving many \"identical files\" in the preimage is just stupid to\nbegin with (if you know they are identical, why are you storing\ncopies, instead of your build procedure to reuse the same file),\nso the algorithm did not bother finding a better match among\n\"equals\".\n\nI am not opposed to a patch that says \"Ok, these two preimages\nhave identical similarity score, *AND* indeed the preimages have\nthe same contents --- we tiebreak them with other heuristics to\nhelp stupid projects better\".  And I can see basename similarity\none of the useful heuristics you could use.\n"},{"id":"45461","messageId":"alpine.LFD.0.98.0706202031200.3593@woody.linux-foundation.org","threadId":"8668","inReplyTo":"20070621030622.GD8477@spearce.org","subject":"Re: Basename matching during rename/copy detection","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-06-21T03:42:38Z","receivedAt":"2007-06-21T03:42:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 20 Jun 2007, Shawn O. Pearce wrote:\n> \n> I'm wondering if we shouldn't play the game of trying to match\n> delete/add pairs up by not only similarity, but also by path\n> basename.\n\nI think we should just consider the basename as an \"added \nsimilarity bonus\".\n\nIOW, we currently sort purely by data similarity, but how about just \nadding a small increment for \"same base name\".\n\nWe could make it actually use the similarity of the filename itself as the \nbasis for the increment, which would be even better, but the trivial thing \nis to do something like\n\n\t--- a/diffcore-rename.c\n\t+++ b/diffcore-rename.c\n\t@@ -186,8 +186,11 @@ static int estimate_similarity(struct diff_filespec *src,\n\t \t */\n\t \tif (!dst->size)\n\t \t\tscore = 0; /* should not happen */\n\t-\telse\n\t+\telse {\n\t \t\tscore = (int)(src_copied * MAX_SCORE / max_size);\n\t+\t\tif (basename_same(src, dst))\n\t+\t\t\tscore++;\n\t+\t}\n\t \treturn score;\n\t }\n \nand just implement that \"basename_same()\" function.\n\nOr something.\n\nI do agree that the filename logically can and probably _should_ count \ntowards the \"similarity\". The filename _is_ part of the data in the global \nnotion of \"content\", after all. It's the \"index\" to the data.\n\n\t\tLinus\n"},{"id":"45478","messageId":"200706210900.49702.andyparkins@gmail.com","threadId":"8668","inReplyTo":"7vsl8m3sph.fsf@assigned-by-dhcp.pobox.com","subject":"Re: Basename matching during rename/copy detection","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-06-21T08:00:47Z","receivedAt":"2007-06-21T08:00:47Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Thursday 2007 June 21, Junio C Hamano wrote:\n\n> Having many \"identical files\" in the preimage is just stupid to\n> begin with (if you know they are identical, why are you storing\n> copies, instead of your build procedure to reuse the same file),\n> so the algorithm did not bother finding a better match among\n> \"equals\".\n\nThat's a really poor argument; it's not git's place to impose restrictions on \nwhat is stored in it.\n\nWhat if it's not a build environment at all but a home directory that's being \nstored - should no one be allowed to store copies of files because \nit's \"stupid\"?  What if it's a collection of images that all started out the \nsame, but have gradually had detail added (which is actually what I do in my \nGUI programs for toolbar images)?  What about files that are used as flags, \nand are all identically empty.\n\nNone of those seems like an abuse of a VCS to me.  In fact, I'd say it's one \nof git's strengths that a duplicate file in the working tree doesn't take up \nany extra space in the repository.\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"45479","messageId":"7vtzt1ybld.fsf@assigned-by-dhcp.pobox.com","threadId":"8668","inReplyTo":"200706210900.49702.andyparkins@gmail.com","subject":"Re: Basename matching during rename/copy detection","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-06-21T08:07:42Z","receivedAt":"2007-06-21T08:07:42Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Andy Parkins <andyparkins@gmail.com> writes:\n\n> On Thursday 2007 June 21, Junio C Hamano wrote:\n>\n>> Having many \"identical files\" in the preimage is just stupid to\n>> begin with (if you know they are identical, why are you storing\n>> copies, instead of your build procedure to reuse the same file),\n>> so the algorithm did not bother finding a better match among\n>> \"equals\".\n>\n> That's a really poor argument; it's not git's place to impose restrictions on \n> what is stored in it.\n\nIt's not even an argument, nor an attempt to justify it.  It was\njust an explanation of historical fact \"It did not bother\".\nPlease re-read the final part of the message, which you omitted\nfrom your quote.\n"},{"id":"45482","messageId":"200706211050.03519.andyparkins@gmail.com","threadId":"8668","inReplyTo":"7vtzt1ybld.fsf@assigned-by-dhcp.pobox.com","subject":"Re: Basename matching during rename/copy detection","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-06-21T09:50:01Z","receivedAt":"2007-06-21T09:50:01Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Thursday 2007 June 21, Junio C Hamano wrote:\n\n> It's not even an argument, nor an attempt to justify it.  It was\n> just an explanation of historical fact \"It did not bother\".\n> Please re-read the final part of the message, which you omitted\n> from your quote.\n\nI omitted it because I didn't object to that part :-)\n\nI appreciated (as always) your practicality in that what you proposed would \nlet people keep their copies.  What I was objecting to was the idea that any \nrepository with duplicate files was \"stupid\".\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"45486","messageId":"Pine.LNX.4.64.0706211248420.4059@racer.site","threadId":"8668","inReplyTo":"alpine.LFD.0.98.0706202031200.3593@woody.linux-foundation.org","subject":"[PATCH] diffcore-rename: favour identical basenames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-21T11:52:11Z","receivedAt":"2007-06-21T11:52:11Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nWhen there are several candidates for a rename source, and one of them\nhas an identical basename to the rename target, take that one.\n\nNoticed by Govind Salinas, posted by Shawn O. Pearce, partial patch\nby Linus Torvalds.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n\n\tOn Wed, 20 Jun 2007, Linus Torvalds wrote:\n\n\t> I think we should just consider the basename as an \"added \n\t> similarity  bonus\".\n\t> \n\t> IOW, we currently sort purely by data similarity, but how about \n\t> just adding a small increment for \"same base name\".\n\t> \n\t> [patch suggestion snipped, since it is identical what is below]\n\n\tHow 'bout this?\n\n diffcore-rename.c      |   33 ++++++++++++++++++++++++++++++++-\n t/t4001-diff-rename.sh |   13 +++++++++++++\n 2 files changed, 45 insertions(+), 1 deletions(-)\n\ndiff --git a/diffcore-rename.c b/diffcore-rename.c\nindex 93c40d9..79c984c 100644\n--- a/diffcore-rename.c\n+++ b/diffcore-rename.c\n@@ -119,6 +119,21 @@ static int is_exact_match(struct diff_filespec *src,\n \treturn 0;\n }\n \n+static int basename_same(struct diff_filespec *src, struct diff_filespec *dst)\n+{\n+\tint src_len = strlen(src->path), dst_len = strlen(dst->path);\n+\twhile (src_len && dst_len) {\n+\t\tchar c1 = src->path[--src_len];\n+\t\tchar c2 = dst->path[--dst_len];\n+\t\tif (c1 != c2)\n+\t\t\treturn 0;\n+\t\tif (c1 == '/')\n+\t\t\treturn 1;\n+\t}\n+\treturn (!src_len || src->path[src_len - 1] == '/') &&\n+\t\t(!dst_len || dst->path[dst_len - 1] == '/');\n+}\n+\n struct diff_score {\n \tint src; /* index in rename_src */\n \tint dst; /* index in rename_dst */\n@@ -186,8 +201,11 @@ static int estimate_similarity(struct diff_filespec *src,\n \t */\n \tif (!dst->size)\n \t\tscore = 0; /* should not happen */\n-\telse\n+\telse {\n \t\tscore = (int)(src_copied * MAX_SCORE / max_size);\n+\t\tif (basename_same(src, dst))\n+\t\t\tscore++;\n+\t}\n \treturn score;\n }\n \n@@ -295,9 +313,22 @@ void diffcore_rename(struct diff_options *options)\n \t\t\tif (rename_dst[i].pair)\n \t\t\t\tcontinue; /* dealt with an earlier round */\n \t\t\tfor (j = 0; j < rename_src_nr; j++) {\n+\t\t\t\tint k;\n \t\t\t\tstruct diff_filespec *one = rename_src[j].one;\n \t\t\t\tif (!is_exact_match(one, two, contents_too))\n \t\t\t\t\tcontinue;\n+\n+\t\t\t\t/* see if there is a basename match, too */\n+\t\t\t\tfor (k = j; k < rename_src_nr; k++) {\n+\t\t\t\t\tone = rename_src[k].one;\n+\t\t\t\t\tif (basename_same(one, two) &&\n+\t\t\t\t\t\tis_exact_match(one, two,\n+\t\t\t\t\t\t\tcontents_too)) {\n+\t\t\t\t\t\tj = k;\n+\t\t\t\t\t\tbreak;\n+\t\t\t\t\t}\n+\t\t\t\t}\n+\n \t\t\t\trecord_rename_pair(i, j, (int)MAX_SCORE);\n \t\t\t\trename_count++;\n \t\t\t\tbreak; /* we are done with this entry */\ndiff --git a/t/t4001-diff-rename.sh b/t/t4001-diff-rename.sh\nindex 2e3c20d..90c085f 100755\n--- a/t/t4001-diff-rename.sh\n+++ b/t/t4001-diff-rename.sh\n@@ -64,4 +64,17 @@ test_expect_success \\\n     'validate the output.' \\\n     'compare_diff_patch current expected'\n \n+test_expect_success 'favour same basenames over different ones' '\n+\tcp path1 another-path &&\n+\tgit add another-path &&\n+\tgit commit -m 1 &&\n+\tgit rm path1 &&\n+\tmkdir subdir &&\n+\tgit mv another-path subdir/path1 &&\n+\tgit runstatus | grep \"renamed: .*path1 -> subdir/path1\"'\n+\n+test_expect_success  'favour same basenames even with minor differences' '\n+\tgit show HEAD:path1 | sed \"s/15/16/\" > subdir/path1 &&\n+\tgit runstatus | grep \"renamed: .*path1 -> subdir/path1\"'\n+\n test_done\n-- \n1.5.2.2.2822.g027a6-dirty\n"},{"id":"45487","messageId":"Pine.LNX.4.64.0706211252190.4059@racer.site","threadId":"8668","inReplyTo":"200706211050.03519.andyparkins@gmail.com","subject":"Re: Basename matching during rename/copy detection","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-21T11:52:38Z","receivedAt":"2007-06-21T11:52:38Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 21 Jun 2007, Andy Parkins wrote:\n\n> On Thursday 2007 June 21, Junio C Hamano wrote:\n> \n> > It's not even an argument, nor an attempt to justify it.  It was just \n> > an explanation of historical fact \"It did not bother\". Please re-read \n> > the final part of the message, which you omitted from your quote.\n> \n> I appreciated (as always) your practicality in that what you proposed \n> would let people keep their copies.  What I was objecting to was the \n> idea that any repository with duplicate files was \"stupid\".\n\nFWIW I find it stupid, too.\n\nCiao,\nDscho\n"},{"id":"45490","messageId":"200706211344.47560.andyparkins@gmail.com","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706211252190.4059@racer.site","subject":"Re: Basename matching during rename/copy detection","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-06-21T12:44:46Z","receivedAt":"2007-06-21T12:44:46Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Thursday 2007 June 21, Johannes Schindelin wrote:\n\n> > would let people keep their copies.  What I was objecting to was the\n> > idea that any repository with duplicate files was \"stupid\".\n>\n> FWIW I find it stupid, too.\n\nThanks very much.  Okay, as I've been put in the position of defending this, \nlet me give you the use case that has cropped up for me to do this stupid \nthing.\n\nI've got a GUI program with a load of tool buttons.  Each of those buttons \nwill, in the final product, be unique images.  When I write the program, I \nwant to be able to refer to each of the button images in the correct place.  \ne.g.\n\n setImage( NewButton, \"path/to/new-button.png\" );\n setImage( OpenButton, \"path/to/open-button.png\" );\n setImage( SaveButton, \"path/to/save-button.png\" );\n\nUnfortunately, I can't draw.  So, I open up gimp, draw a big red X and save it \nas new-button.png.  Then I copy that file to open-button.png and \nsave-button.png, knowing that at some point in the future, someone will come \nand replace those red-X images with something appropriate.\n\nAll those images now go in the repository.  Symbolic links are not an option, \nas it's got to be checkable out on Windows.\n\nTell me what part of that is stupid?\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"45491","messageId":"vpqodj9zcxf.fsf@bauges.imag.fr","threadId":"8668","inReplyTo":"200706211344.47560.andyparkins@gmail.com","subject":"Re: Basename matching during rename/copy detection","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-06-21T12:53:32Z","receivedAt":"2007-06-21T12:53:32Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Andy Parkins <andyparkins@gmail.com> writes:\n\n> On Thursday 2007 June 21, Johannes Schindelin wrote:\n>\n>> > would let people keep their copies.  What I was objecting to was the\n>> > idea that any repository with duplicate files was \"stupid\".\n>>\n>> FWIW I find it stupid, too.\n>\n> Thanks very much.  Okay, as I've been put in the position of defending this, \n> let me give you the use case that has cropped up for me to do this stupid \n> thing.\n\nWell, why look so far to find an example of people having identical\nfiles in their tree?\n\n$ cd git\n$ git-ls-files -z | xargs -0 md5sum | cut -f 1 -d ' ' | wc -l              \n973\n$ git-ls-files -z | xargs -0 md5sum | cut -f 1 -d ' ' | sort | uniq | wc -l\n964\n$ \n\n-- \nMatthieu\n"},{"id":"45493","messageId":"20070621131047.GC4487@coredump.intra.peff.net","threadId":"8668","inReplyTo":"vpqodj9zcxf.fsf@bauges.imag.fr","subject":"Re: Basename matching during rename/copy detection","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-06-21T13:10:47Z","receivedAt":"2007-06-21T13:10:47Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jun 21, 2007 at 02:53:32PM +0200, Matthieu Moy wrote:\n\n> Well, why look so far to find an example of people having identical\n> files in their tree?\n> \n> $ cd git\n> $ git-ls-files -z | xargs -0 md5sum | cut -f 1 -d ' ' | wc -l              \n> 973\n> $ git-ls-files -z | xargs -0 md5sum | cut -f 1 -d ' ' | sort | uniq | wc -l\n> 964\n\nmd5? What is this, CVS? How about:\n\ngit-ls-files -s | cut -d' ' -f2 | sort | uniq -d | wc -l\n\nYour pipeline will also list files in the working directory, which can\ninflate the number of duplicates (note that git-foo.sh and git-foo will\nhave the same content).\n\n-Peff\n\nPS Please don't take this to mean I think duplicate files are stupid; I\nthink they can be quite useful. I just wanted to nitpick your shell\ncommand. :)\n"},{"id":"45494","messageId":"Pine.LNX.4.64.0706211417090.4059@racer.site","threadId":"8668","inReplyTo":"vpqodj9zcxf.fsf@bauges.imag.fr","subject":"Re: Basename matching during rename/copy detection","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-21T13:18:37Z","receivedAt":"2007-06-21T13:18:37Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 21 Jun 2007, Matthieu Moy wrote:\n\n> Well, why look so far to find an example of people having identical\n> files in their tree?\n> \n> $ cd git\n> $ git-ls-files -z | xargs -0 md5sum | cut -f 1 -d ' ' | wc -l              \n> 973\n> $ git-ls-files -z | xargs -0 md5sum | cut -f 1 -d ' ' | sort | uniq | wc -l\n> 964\n> $ \n\nHave you checked the files? They are all some blobs in the test scripts. \n\nCiao,\nDscho\n"},{"id":"45495","messageId":"20070621131915.GD4487@coredump.intra.peff.net","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706211248420.4059@racer.site","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-06-21T13:19:15Z","receivedAt":"2007-06-21T13:19:15Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jun 21, 2007 at 12:52:11PM +0100, Johannes Schindelin wrote:\n\n> When there are several candidates for a rename source, and one of them\n> has an identical basename to the rename target, take that one.\n\nThat's a reasonable heuristic, but it unfortunately won't match simple\nthings like:\n\n  i386_widget.c -> arch/i386/widget.c\n\nYou really don't care about \"is this a good match\" as much as providing\nan order to potential matches. I think something like a Levenshtein\ndistance between the full pathnames would give good results, and would\ncover almost every situation that the basename heuristic would (there\nare a few exceptions, like getting \"file.c\" from either \"file2.c\" or\n\"foo/file.c\", but that seems kind of pathological).\n\nSorry to post without a patch, but I don't have time right this second.\nI'll add it to the end of my (ever-growing) todo list if you think it's\na good idea and don't do it yourself. :)\n\n-Peff\n"},{"id":"45496","messageId":"Pine.LNX.4.64.0706211418430.4059@racer.site","threadId":"8668","inReplyTo":"200706211344.47560.andyparkins@gmail.com","subject":"Re: Basename matching during rename/copy detection","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-21T13:22:28Z","receivedAt":"2007-06-21T13:22:28Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 21 Jun 2007, Andy Parkins wrote:\n\n> I open up gimp, draw a big red X and save it as new-button.png.  Then I \n> copy that file to open-button.png and save-button.png, knowing that at \n> some point in the future, someone will come and replace those red-X \n> images with something appropriate.\n\nSo you have a couple of identical files in your repo, which are \nplaceholders. That is quite different from what I criticised of being \nstupid.\n\nWell, I would not have checked in the files in your place, but only one:\n\n\tdumb-red-X.png\n\nThen, my Makefile would have checked for the existence of, say, \nmy-wonderful-ok-button.png, and if it does not exist yet, copy \ndumb-red-X.png to it.\n\nNow, when somebody comes along, paining the prettiest ok button I ever \nsaw, I copy that over the copy of dumb-red-X.png, and check it in.\n\nIt has the further bonus that I know exactly which buttons I have to find \na suck^Wgifted artist for.\n\nCiao,\nDscho\n"},{"id":"45497","messageId":"vpqfy4lxwvl.fsf@bauges.imag.fr","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706211417090.4059@racer.site","subject":"Re: Basename matching during rename/copy detection","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-06-21T13:25:34Z","receivedAt":"2007-06-21T13:25:34Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Have you checked the files? They are all some blobs in the test scripts. \n\nYes, but how does it make any difference? You still want git to manage\nthem properly, don't you?\n\n-- \nMatthieu\n"},{"id":"45498","messageId":"Pine.LNX.4.64.0706211451480.4059@racer.site","threadId":"8668","inReplyTo":"vpqfy4lxwvl.fsf@bauges.imag.fr","subject":"Re: Basename matching during rename/copy detection","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-21T13:52:51Z","receivedAt":"2007-06-21T13:52:51Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 21 Jun 2007, Matthieu Moy wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > Have you checked the files? They are all some blobs in the test scripts. \n> \n> Yes, but how does it make any difference? You still want git to manage\n> them properly, don't you?\n\nYes. And Git explicitely allows what I call stupid. And yes, those \n_identical_ files in the test suit should probably all be folded into \nsingle files, and the places where they are used should reference _that_ \nsingle instance.\n\nCiao,\nDscho\n"},{"id":"45500","messageId":"Pine.LNX.4.64.0706211459140.4059@racer.site","threadId":"8668","inReplyTo":"20070621131915.GD4487@coredump.intra.peff.net","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-21T14:03:47Z","receivedAt":"2007-06-21T14:03:47Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 21 Jun 2007, Jeff King wrote:\n\n> On Thu, Jun 21, 2007 at 12:52:11PM +0100, Johannes Schindelin wrote:\n> \n> > When there are several candidates for a rename source, and one of them\n> > has an identical basename to the rename target, take that one.\n> \n> That's a reasonable heuristic, but it unfortunately won't match simple\n> things like:\n> \n>   i386_widget.c -> arch/i386/widget.c\n\nThat's right. But every heuristic falls down eventually. Personally, I \nthink basename_same() is good enough, even if the technical challenge to \nimplement a small enough Levenshtein, which still respects directory \nboundaries somehow (and not just throws them away).\n\nBesides, Levenshtein would introduce a ranking, not a boolean value like \nbasename_same(). And that complicates the code.\n\nAll in all, I'd say Levenshtein is not worth the _result_.\n\nCiao,\nDscho\n"},{"id":"45505","messageId":"467A9B2C.2060907@midwinter.com","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706211451480.4059@racer.site","subject":"Re: Basename matching during rename/copy detection","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-06-21T15:37:16Z","receivedAt":"2007-06-21T15:37:16Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Johannes Schindelin wrote:\n> Yes. And Git explicitely allows what I call stupid. And yes, those \n> _identical_ files in the test suit should probably all be folded into \n> single files, and the places where they are used should reference _that_ \n> single instance.\n>   \n\nTwo files that are identical in the current revision have not \nnecessarily been identical from the beginning. Doing what you suggest \nwill cause you to lose the history of all but one of those files.\n\nFiles can absolutely become identical in the real world. I know that for \na fact because it happened to me just this week (see my \"Directory \nrenames\" message from a few days ago.) Are you seriously suggesting that \nevery time I unpack an update from a third party, I should go through it \nand see if they have changed any files such that the contents now match \nanother file in my repository, and if so, I should remove all but one of \nthe copies from my repository and have a build system create it instead? \nThen undo that work when I unpack another update and the files are no \nlonger identical?\n\nWell, no, I know you're not suggesting that, but it's the logical \nconclusion of the \"it's stupid to ever have duplicate files\" philosophy. \nWhile that approach certainly makes life easier for the version control \nsystem, it doesn't exactly make life easier for the *developer*, which \nis kind of the whole point of why we're here.\n\n-Steve\n"},{"id":"45508","messageId":"Pine.LNX.4.64.0706211649520.4059@racer.site","threadId":"8668","inReplyTo":"467A9B2C.2060907@midwinter.com","subject":"Re: Basename matching during rename/copy detection","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-21T15:53:31Z","receivedAt":"2007-06-21T15:53:31Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 21 Jun 2007, Steven Grimm wrote:\n\n> Johannes Schindelin wrote:\n> > Yes. And Git explicitely allows what I call stupid. And yes, those\n> > _identical_ files in the test suit should probably all be folded into\n> > single files, and the places where they are used should reference _that_\n> > single instance.\n> >   \n> \n> Two files that are identical in the current revision have not necessarily\n> been identical from the beginning. Doing what you suggest will cause you to\n> lose the history of all but one of those files.\n> \n> Files can absolutely become identical in the real world. I know that for a\n> fact because it happened to me just this week (see my \"Directory renames\"\n> message from a few days ago.)\n\nNo, that message did not convince me. It was way too short on the side of \nfacts.\n\nAnd no, I do not think that two unrelated files can get exactly the same \ncontent.\n\nBe that as may, even _if_ there were such a case, I'd still try to reuse \nthe same file in the working directory. Just because Git can deal \nefficiently with millions of identical files does not mean that a working \ndirectory can, or worse, human developers.\n\nCiao,\nDscho\n"},{"id":"45512","messageId":"alpine.LFD.0.98.0706210910390.3593@woody.linux-foundation.org","threadId":"8668","inReplyTo":"20070621131915.GD4487@coredump.intra.peff.net","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-06-21T16:20:51Z","receivedAt":"2007-06-21T16:20:51Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 21 Jun 2007, Jeff King wrote:\n\n> On Thu, Jun 21, 2007 at 12:52:11PM +0100, Johannes Schindelin wrote:\n> \n> > When there are several candidates for a rename source, and one of them\n> > has an identical basename to the rename target, take that one.\n> \n> That's a reasonable heuristic, but it unfortunately won't match simple\n> things like:\n> \n>   i386_widget.c -> arch/i386/widget.c\n\nWe'e also had things like\n\n\tarch/i386/kernel/pci-pc.c -> arch/i386/kernel/pci/common.c\n\nso it's not always the ending of a file that is unchanged, but you still \noften have some \"similarity\" of the name (ie the \"pci\" substring is still \ncommon there).\n\nSo I agree that we can be even better about the heuristics. I don't know \nhow much it *matters* in practice.\n\nI do agree with the people who argue that you simply shouldn't depend on \nthese kinds of things, and if you have identical files, and move them \naround, you really are getting behaviour that doesn't matter.\n\nThe files are *identical* for christ sake! Following their history, it \ndoesn't matter *which* base you follow, since regardless, they've come to \nthe same point!\n\nSo in that sense, the current git behaviour is actually perfectly fine.\n\nAt the same time, I'll argue from a totally theoretical point that the \n\"filename\" is obviously part of the data in the tree, and as such, a \nsimilarity comparison that takes only the data into account is a bit \nlimited. So while I don't think a user should really care, I also think \nthat keeping the filename as part of the similarity analysis is actually \na perfectly logical and valid thing to do withing the git policy of \n\"content is king\".\n\nThe filename *is* part of the content, and it's doubly so when you think \nabout a rename or copy operation, where the whole point of the exercise is \nas much about the filename as about the data inside the file.\n\n\t\t\tLinus\n"},{"id":"45516","messageId":"467AADFA.9040804@midwinter.com","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706211649520.4059@racer.site","subject":"Re: Basename matching during rename/copy detection","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-06-21T16:57:30Z","receivedAt":"2007-06-21T16:57:30Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Johannes Schindelin wrote:\n> No, that message did not convince me. It was way too short on the side of \n> facts.\n>   \n\nShort of posting multiple historical versions of the third-party source \ncode in question, I'm not sure what I can do to convince you. And I'd \nrather not violate the license agreement on that code. I would have \nthought, though, that the fact that I supplied a detailed, reproducible \ntest case with obviously broken behavior would itself have been pretty \nconvincing.\n\nThe fact that not all projects contain any short files, or any files \nwhose contents have ever been identical, does not cause git's behavior \nin that test case to be correct. \"It's broken and unfixable\" is one \nthing; \"It's broken and we don't care\" is another; and \"It's broken and \nwe care but it's not at the top of anyone's priority list to fix\" is \nsomething else again. All of those are fine, but \"If it's broken, you \nare stupid\" and \"If it's broken, it's a sign your project isn't real\" \nare not.\n\nOr, to take another tack on this entirely, it is not the proper function \nof a version control system to dictate the contents of the projects \nunder its control. It should take whatever we humans throw at it and \nreproduce those contents faithfully with coherent, non-jumbled history. \nIt should do so even if what we're throwing at it is completely stupid.\n\nBy the way, I'll toss out one more example of legitimate duplicate \nfiles, though admittedly one where you might not care so much about \nhistory jumbling: if you have a project that makes use of two GPL \nlibraries or utilities whose source you want to keep locally, e.g. \nbecause you are making local modifications, you will have two copies of \nthe GNU \"COPYING\" file. Neither one produced by a build system (or at \nleast, not by *your* build system) and you are not permitted by the \nterms of the GPL to publish a copy of either piece of software without a \nverbatim copy of its license -- it says so right in section 1 of the GPL \n(the \"keep intact\" wording.) Removing one of those copies and expecting \na build system to reconstruct it after someone clones your repository \nwould arguably be a violation of the GPL.\n\n-Steve\n"},{"id":"45521","messageId":"7vfy4lw5yk.fsf@assigned-by-dhcp.pobox.com","threadId":"8668","inReplyTo":"alpine.LFD.0.98.0706210910390.3593@woody.linux-foundation.org","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-06-21T17:52:19Z","receivedAt":"2007-06-21T17:52:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> We'e also had things like\n>\n> \tarch/i386/kernel/pci-pc.c -> arch/i386/kernel/pci/common.c\n>\n> so it's not always the ending of a file that is unchanged, but you still \n> often have some \"similarity\" of the name (ie the \"pci\" substring is still \n> common there).\n\nThis is not an example to draw very useful conclusions, is it?\n\nThe heuristics to say '-pc => common' is a more likely rename\nthan '-obscure-arch => common' heavily depends on human\nintelligence in the context of a particular project, the kernel,\nwhere there are rules such as \"peripherals are tested most\nwidely on PC architectures, so assume that the vendors might\nhave tested their stuff only on PCs\".\n\nBut I do agree that not limiting to basename has values.\nTaking example from the \"I cannot draw so here is a red big X\",\nit is quite possible that two red big X's are replaced with\nproperly rendered icons, while their format modified, like so:\n\n    images/ok-button.gif => images/buttons/ok.png\n    images/cancel-button.gif => images/buttons/cancel.png\n\nThis suggests that we might be able to look around to see what\nother rename src/target candidate files there are, so that we\ncan figure out if there is a common pattern (i.e. in the above\nexample, \"patsubst images/%-button.gif,images/buttons/%.png\" is\nwhat is going on).  If we find such a pattern, we can base the\nassignment of \"basename similarity bonus\" on that pattern.\n"},{"id":"45522","messageId":"alpine.LFD.0.98.0706211118420.3593@woody.linux-foundation.org","threadId":"8668","inReplyTo":"7vfy4lw5yk.fsf@assigned-by-dhcp.pobox.com","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-06-21T18:24:50Z","receivedAt":"2007-06-21T18:24:50Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 21 Jun 2007, Junio C Hamano wrote:\n> \n> This is not an example to draw very useful conclusions, is it?\n> \n> The heuristics to say '-pc => common' is a more likely rename\n> than '-obscure-arch => common' heavily depends on human\n> intelligence in the context\n\nOh, absolutely.\n\nI'm just saying that *if* you see two equally weighed content moves, if \nyou then prefer the one that has more in common with the name, that's \nlikely the right choice. \n\nIn the actual example I gave, there was no ambiguity: the file contents \nwere very obvious. But let's sat that you happened to have an example of \ntwo files with 100% identical content that moved, and you had the files\n\n\t-arch/i386/kernel/pci-pc.c\n\t-arch/alpha/kernel/pci-pc.c\n\t+arch/i386/kernel/pci/common.c\n\t+arch/alpha/kernel/pci/common.c\n\nto match up, how would you do it? Again: they're all identical files: we \ncan obviously agree that two files got renamed, but what is the pairing.\n\nI'd suggest that if you do it by matching up the similarity of the \nfilenames (not necessarily \"exact same basename\"), you'd actually catch \nit. In this case, they all have \"pci\" in them, but the \"alpha\" similarity \nwould make you select the right one.\n\nSimilarly, in some other cases, the \"pci\" might be the thing they have in \ncommon, and might be the thing that decides that \"oh, those two filenames \nlook like they might be more of a better pair\".\n\nAnd yes, all of this would trigger only if the file data content match is \nnon-conclusive. The file data is *more* important, but that doesn't mean \nthat the file name similarity is *totally* unimportant either.\n\n\t\t\tLinus\n"},{"id":"45534","messageId":"Pine.LNX.4.64.0706220214250.4059@racer.site","threadId":"8668","inReplyTo":"20070621131915.GD4487@coredump.intra.peff.net","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-22T01:14:43Z","receivedAt":"2007-06-22T01:14:43Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 21 Jun 2007, Jeff King wrote:\n\n> I think something like a Levenshtein distance between the full pathnames \n> would give good results, and would cover almost every situation that the \n> basename heuristic would (there are a few exceptions, like getting \n> \"file.c\" from either \"file2.c\" or \"foo/file.c\", but that seems kind of \n> pathological).\n\nWell, now you only have to test if it makes sense:\n\n-- snipsnap --\n[PATCH] diffcore-rename: replace basename_same() heuristics by Levenshtein\n\nInstead of insisting on identical basenames, try the levenshtein\ndistance.\n\nBasically, if there are multiple rename source candidates, take the\none with the smallest Levenshtein distance.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n\n\tThe dangerous thing is that the score can get negative now.\n\n Makefile          |    4 ++--\n diffcore-rename.c |   42 +++++++++++++++---------------------------\n levenshtein.c     |   39 +++++++++++++++++++++++++++++++++++++++\n levenshtein.h     |    6 ++++++\n 4 files changed, 62 insertions(+), 29 deletions(-)\n create mode 100644 levenshtein.c\n create mode 100644 levenshtein.h\n\ndiff --git a/Makefile b/Makefile\nindex 74b69fb..e015833 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -303,12 +303,12 @@ LIB_H = \\\n \trun-command.h strbuf.h tag.h tree.h git-compat-util.h revision.h \\\n \ttree-walk.h log-tree.h dir.h path-list.h unpack-trees.h builtin.h \\\n \tutf8.h reflog-walk.h patch-ids.h attr.h decorate.h progress.h \\\n-\tmailmap.h remote.h\n+\tmailmap.h remote.h levenshtein.h\n \n DIFF_OBJS = \\\n \tdiff.o diff-lib.o diffcore-break.o diffcore-order.o \\\n \tdiffcore-pickaxe.o diffcore-rename.o tree-diff.o combine-diff.o \\\n-\tdiffcore-delta.o log-tree.o\n+\tdiffcore-delta.o log-tree.o levenshtein.o\n \n LIB_OBJS = \\\n \tblob.o commit.o connect.o csum-file.o cache-tree.o base85.o \\\ndiff --git a/diffcore-rename.c b/diffcore-rename.c\nindex 79c984c..41448c9 100644\n--- a/diffcore-rename.c\n+++ b/diffcore-rename.c\n@@ -4,6 +4,7 @@\n #include \"cache.h\"\n #include \"diff.h\"\n #include \"diffcore.h\"\n+#include \"levenshtein.h\"\n \n /* Table of rename/copy destinations */\n \n@@ -119,21 +120,6 @@ static int is_exact_match(struct diff_filespec *src,\n \treturn 0;\n }\n \n-static int basename_same(struct diff_filespec *src, struct diff_filespec *dst)\n-{\n-\tint src_len = strlen(src->path), dst_len = strlen(dst->path);\n-\twhile (src_len && dst_len) {\n-\t\tchar c1 = src->path[--src_len];\n-\t\tchar c2 = dst->path[--dst_len];\n-\t\tif (c1 != c2)\n-\t\t\treturn 0;\n-\t\tif (c1 == '/')\n-\t\t\treturn 1;\n-\t}\n-\treturn (!src_len || src->path[src_len - 1] == '/') &&\n-\t\t(!dst_len || dst->path[dst_len - 1] == '/');\n-}\n-\n struct diff_score {\n \tint src; /* index in rename_src */\n \tint dst; /* index in rename_dst */\n@@ -201,11 +187,9 @@ static int estimate_similarity(struct diff_filespec *src,\n \t */\n \tif (!dst->size)\n \t\tscore = 0; /* should not happen */\n-\telse {\n-\t\tscore = (int)(src_copied * MAX_SCORE / max_size);\n-\t\tif (basename_same(src, dst))\n-\t\t\tscore++;\n-\t}\n+\telse\n+\t\tscore = (int)(src_copied * MAX_SCORE / max_size)\n+\t\t\t- levenshtein(src->path, dst->path);\n \treturn score;\n }\n \n@@ -313,20 +297,24 @@ void diffcore_rename(struct diff_options *options)\n \t\t\tif (rename_dst[i].pair)\n \t\t\t\tcontinue; /* dealt with an earlier round */\n \t\t\tfor (j = 0; j < rename_src_nr; j++) {\n-\t\t\t\tint k;\n+\t\t\t\tint k, distance;\n \t\t\t\tstruct diff_filespec *one = rename_src[j].one;\n \t\t\t\tif (!is_exact_match(one, two, contents_too))\n \t\t\t\t\tcontinue;\n \n+\t\t\t\tdistance = levenshtein(one->path, two->path);\n \t\t\t\t/* see if there is a basename match, too */\n \t\t\t\tfor (k = j; k < rename_src_nr; k++) {\n+\t\t\t\t\tint d2;\n \t\t\t\t\tone = rename_src[k].one;\n-\t\t\t\t\tif (basename_same(one, two) &&\n-\t\t\t\t\t\tis_exact_match(one, two,\n-\t\t\t\t\t\t\tcontents_too)) {\n-\t\t\t\t\t\tj = k;\n-\t\t\t\t\t\tbreak;\n-\t\t\t\t\t}\n+\t\t\t\t\tif (!is_exact_match(one, two,\n+\t\t\t\t\t\t\t\tcontents_too))\n+\t\t\t\t\t\tcontinue;\n+\t\t\t\t\td2 = levenshtein(one->path, two->path);\n+\t\t\t\t\tif (d2 > distance)\n+\t\t\t\t\t\tcontinue;\n+\t\t\t\t\tdistance = d2;\n+\t\t\t\t\tj = k;\n \t\t\t\t}\n \n \t\t\t\trecord_rename_pair(i, j, (int)MAX_SCORE);\ndiff --git a/levenshtein.c b/levenshtein.c\nnew file mode 100644\nindex 0000000..80ef860\n--- /dev/null\n+++ b/levenshtein.c\n@@ -0,0 +1,39 @@\n+#include \"cache.h\"\n+#include \"levenshtein.h\"\n+\n+int levenshtein(const char *string1, const char *string2)\n+{\n+\tint len1 = strlen(string1), len2 = strlen(string2);\n+\tint *row1 = xmalloc(sizeof(int) * (len2 + 1));\n+\tint *row2 = xmalloc(sizeof(int) * (len2 + 1));\n+\tint i, j;\n+\n+\tfor (j = 1; j <= len2; j++)\n+\t\trow1[j] = j;\n+\tfor (i = 0; i < len1; i++) {\n+\t\tint *dummy;\n+\n+\t\trow2[0] = i + 1;\n+\t\tfor (j = 0; j < len2; j++) {\n+\t\t\t/* substitution */\n+\t\t\trow2[j + 1] = row1[j] + (string1[i] != string2[j]);\n+\t\t\t/* insertion */\n+\t\t\tif (row2[j + 1] > row1[j + 1] + 1)\n+\t\t\t\trow2[j + 1] = row1[j + 1] + 1;\n+\t\t\t/* deletion */\n+\t\t\tif (row2[j + 1] > row2[j] + 1)\n+\t\t\t\trow2[j + 1] = row2[j] + 1;\n+\t\t}\n+\n+\t\tdummy = row1;\n+\t\trow1 = row2;\n+\t\trow2 = dummy;\n+\t}\n+\n+\ti = row1[len2];\n+\tfree(row1);\n+\tfree(row2);\n+\n+\treturn i;\n+}\n+\ndiff --git a/levenshtein.h b/levenshtein.h\nnew file mode 100644\nindex 0000000..74a6626\n--- /dev/null\n+++ b/levenshtein.h\n@@ -0,0 +1,6 @@\n+#ifndef LEVENSHTEIN_H\n+#define LEVENSHTEIN_H\n+\n+int levenshtein(const char *string1, const char *string2);\n+\n+#endif\n-- \n1.5.2.2.2822.g027a6-dirty\n"},{"id":"45551","messageId":"20070622054142.GA7699@coredump.intra.peff.net","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706220214250.4059@racer.site","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-06-22T05:41:42Z","receivedAt":"2007-06-22T05:41:42Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jun 22, 2007 at 02:14:43AM +0100, Johannes Schindelin wrote:\n\n> @@ -313,20 +297,24 @@ void diffcore_rename(struct diff_options *options)\n>  \t\t\tif (rename_dst[i].pair)\n>  \t\t\t\tcontinue; /* dealt with an earlier round */\n>  \t\t\tfor (j = 0; j < rename_src_nr; j++) {\n> -\t\t\t\tint k;\n> +\t\t\t\tint k, distance;\n>  \t\t\t\tstruct diff_filespec *one = rename_src[j].one;\n>  \t\t\t\tif (!is_exact_match(one, two, contents_too))\n>  \t\t\t\t\tcontinue;\n>  \n> +\t\t\t\tdistance = levenshtein(one->path, two->path);\n>  \t\t\t\t/* see if there is a basename match, too */\n>  \t\t\t\tfor (k = j; k < rename_src_nr; k++) {\n\nThis loop can start at k = j+1, since otherwise we are just checking\nrename_src[j] against itself.\n\n> +int levenshtein(const char *string1, const char *string2)\n> +{\n> +\tint len1 = strlen(string1), len2 = strlen(string2);\n> +\tint *row1 = xmalloc(sizeof(int) * (len2 + 1));\n> +\tint *row2 = xmalloc(sizeof(int) * (len2 + 1));\n> +\tint i, j;\n> +\n> +\tfor (j = 1; j <= len2; j++)\n> +\t\trow1[j] = j;\n\nThis loop must start at j=0, not j=1; otherwise you have an undefined\nvalue in row1[0], which gets read when setting row2[1], and you get\na totally meaningless distance (I got -1209667248 on my test case!).\n\n-Peff\n"},{"id":"45552","messageId":"467B777D.C47BFE0E@eudaptics.com","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706220214250.4059@racer.site","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Johannes Sixt","fromEmail":"j.sixt@eudaptics.com","sentAt":"2007-06-22T07:17:17Z","receivedAt":"2007-06-22T07:17:17Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Johannes Schindelin wrote:\n>         The dangerous thing is that the score can get negative now.\n>  ...\n> +               score = (int)(src_copied * MAX_SCORE / max_size)\n> +                       - levenshtein(src->path, dst->path);\n\nDoes that also mean that you can't ever have a rename with a score of\n100%?\n\n(I haven't studied the algorithms and assume that levenshtein(a,b) == 0\nonly if a==b, and that without the -levenshtein(...) the score can grow\nto 100%.)\n\n-- Hannes\n"},{"id":"45558","messageId":"Pine.LNX.4.64.0706221042551.4059@racer.site","threadId":"8668","inReplyTo":"20070622054142.GA7699@coredump.intra.peff.net","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-22T10:22:09Z","receivedAt":"2007-06-22T10:22:09Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 22 Jun 2007, Jeff King wrote:\n\n> On Fri, Jun 22, 2007 at 02:14:43AM +0100, Johannes Schindelin wrote:\n> \n> > @@ -313,20 +297,24 @@ void diffcore_rename(struct diff_options *options)\n> >  \t\t\tif (rename_dst[i].pair)\n> >  \t\t\t\tcontinue; /* dealt with an earlier round */\n> >  \t\t\tfor (j = 0; j < rename_src_nr; j++) {\n> > -\t\t\t\tint k;\n> > +\t\t\t\tint k, distance;\n> >  \t\t\t\tstruct diff_filespec *one = rename_src[j].one;\n> >  \t\t\t\tif (!is_exact_match(one, two, contents_too))\n> >  \t\t\t\t\tcontinue;\n> >  \n> > +\t\t\t\tdistance = levenshtein(one->path, two->path);\n> >  \t\t\t\t/* see if there is a basename match, too */\n> >  \t\t\t\tfor (k = j; k < rename_src_nr; k++) {\n> \n> This loop can start at k = j+1, since otherwise we are just checking\n> rename_src[j] against itself.\n\nRight.\n\n> > +int levenshtein(const char *string1, const char *string2)\n> > +{\n> > +\tint len1 = strlen(string1), len2 = strlen(string2);\n> > +\tint *row1 = xmalloc(sizeof(int) * (len2 + 1));\n> > +\tint *row2 = xmalloc(sizeof(int) * (len2 + 1));\n> > +\tint i, j;\n> > +\n> > +\tfor (j = 1; j <= len2; j++)\n> > +\t\trow1[j] = j;\n> \n> This loop must start at j=0, not j=1; otherwise you have an undefined\n> value in row1[0], which gets read when setting row2[1], and you get\n> a totally meaningless distance (I got -1209667248 on my test case!).\n\nSorry for that. I originally had an xcalloc in there, and did not look at \nthat loop afterwards.\n\nAnd I completely forgot that on my laptop (on which I did this patch), I \nhad forgotten to add\n\n\tALL_CFLAGS += -DXMALLOC_POISON=1\n\nto config.mak.\n\nCiao,\nDscho\n"},{"id":"45559","messageId":"Pine.LNX.4.64.0706221122200.4059@racer.site","threadId":"8668","inReplyTo":"467B777D.C47BFE0E@eudaptics.com","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-22T10:39:46Z","receivedAt":"2007-06-22T10:39:46Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 22 Jun 2007, Johannes Sixt wrote:\n\n> Johannes Schindelin wrote:\n> >         The dangerous thing is that the score can get negative now.\n> >  ...\n> > +               score = (int)(src_copied * MAX_SCORE / max_size)\n> > +                       - levenshtein(src->path, dst->path);\n> \n> Does that also mean that you can't ever have a rename with a score of\n> 100%?\n> \n> (I haven't studied the algorithms and assume that levenshtein(a,b) == 0\n> only if a==b, and that without the -levenshtein(...) the score can grow\n> to 100%.)\n\nThere is a different code path for identical contents. So yes, you can \nstill hit 100%, but it is now much, much harder to hit a score close to \n100% [*1*].\n\nThe obviously correct way to do this is to have a subscore, and use it \n_strictly_ only when the score is identical.\n\nI see two ways to do this properly:\n\n- introduce a name_distance struct member, just below the score. This \n  means that estimate_similarity has to \"return\" two values instead of \n  one, and score_compare gets a bit more complex, too. Or\n\n- change the score to unsigned long, and shift the score to higher bits, \n  adding a constant minus the Levenshtein distance. It is safe to assume \n  that the filenames are shorter than 16384 bytes (PATH_MAX is actually \n  much smaller than that), and even if two filenames of that length are \n  completely different, the distance can not be larger than twice that \n  number, i.e. 16384 deletions + 16384 insertions. Therefore, you could \n  pick 32768 as that constant.\n\nHowever, I find both solutions ugly. Besides, I am not interested in the \nfeature myself, only the implementation of Levenshtein was interesting, \nand I thought I just post the code here. So I did only the minimal stuff \non top of the interesting one to make it sort of work.\n\nIf somebody wants to pick up the ball, be my guest, because I am out of \nthat game.\n\nCiao,\nDscho\n\nFootnote:\n\n*1* Actually, it is not _that_ bad. The score is not a value between 0 and \n    100, IOW it is _not_ what you see in the output of \"diff -M\". It is an \n    unsigned short between 0 and MAX_SCORE, which is defined in \n    diffcore.h as 60000.0.\n\n    The Levenshtein distance between two filenames cannot be larger than \n    the sum of their lengths, so it should be relatively safe. That is, if \n    you don't have such insanely long paths as e.g. egit. But even there, \n    the paths share most of their directories, and therefore the distances \n    should be much, much smaller in real life.\n"},{"id":"45563","messageId":"86ps3oi7ma.fsf_-_@lola.quinscape.zz","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706221122200.4059@racer.site","subject":"100% (was: [PATCH] diffcore-rename: favour identical basenames)","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-06-22T10:52:29Z","receivedAt":"2007-06-22T10:52:29Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\n> Footnote:\n>\n> *1* Actually, it is not _that_ bad. The score is not a value between 0 and \n>     100, IOW it is _not_ what you see in the output of \"diff -M\". It is an \n>     unsigned short between 0 and MAX_SCORE, which is defined in \n>     diffcore.h as 60000.0.\n>\n>     The Levenshtein distance between two filenames cannot be larger than \n>     the sum of their lengths, so it should be relatively safe. That is, if \n>     you don't have such insanely long paths as e.g. egit. But even there, \n>     the paths share most of their directories, and therefore the distances \n>     should be much, much smaller in real life.\n\nAs a note aside: would it be possible to always round downwards when\ncomputing similarities or converting between them?\n\nI very much would like to see the 100% figure reserved for identity.\nThis is particularly relevant when interpreting the output of git-diff\n--name-status with regard to R100, C100 and similar flags.\n\n-- \nDavid Kastrup\n"},{"id":"45566","messageId":"Pine.LNX.4.64.0706221347480.4059@racer.site","threadId":"8668","inReplyTo":"86ps3oi7ma.fsf_-_@lola.quinscape.zz","subject":"Re: 100% (was: [PATCH] diffcore-rename: favour identical basenames)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-22T12:49:02Z","receivedAt":"2007-06-22T12:49:02Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 22 Jun 2007, David Kastrup wrote:\n\n> As a note aside: would it be possible to always round downwards when \n> computing similarities or converting between them?\n\nI'd rather not. This would be counterintuitive. People expect rounded \nvalues.\n\n> I very much would like to see the 100% figure reserved for identity.\n> This is particularly relevant when interpreting the output of git-diff\n> --name-status with regard to R100, C100 and similar flags.\n\nYou should never depend on the output of --name-status if you're \ninterested in identifying identical files, but on the object names.\n\nCiao,\nDscho\n"},{"id":"45574","messageId":"200706221619.30521.andyparkins@gmail.com","threadId":"8668","inReplyTo":"alpine.LFD.0.98.0706210910390.3593@woody.linux-foundation.org","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-06-22T15:19:26Z","receivedAt":"2007-06-22T15:19:26Z","isPatch":true,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Thursday 2007 June 21, Linus Torvalds wrote:\n\n> The files are *identical* for christ sake! Following their history, it\n> doesn't matter *which* base you follow, since regardless, they've come to\n> the same point!\n>\n> So in that sense, the current git behaviour is actually perfectly fine.\n\nPerhaps not.  (Please don't read this as meaning I disagree with your \nfavour-the-identical-filename patch at all - in fact I think that would \naddress the case I give below).\n\nWhat if two files with different filenames and content converge at some point \nin history, then diverge again?  If git is tracking renames merely by content \nand picks the wrong one, then the history of fileA suddenly becomes the \nhistory of fileB.\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"45575","messageId":"Pine.LNX.4.64.0706221626450.4059@racer.site","threadId":"8668","inReplyTo":"200706221619.30521.andyparkins@gmail.com","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-22T15:28:47Z","receivedAt":"2007-06-22T15:28:47Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 22 Jun 2007, Andy Parkins wrote:\n\n> What if two files with different filenames and content converge at some \n> point in history, then diverge again?  If git is tracking renames merely \n> by content and picks the wrong one, then the history of fileA suddenly \n> becomes the history of fileB.\n\nThis is becoming highly ethereal. Like \"I could imagine that some day in \nfuture, some person could devise a device, that might allow you to do \nsomething that I can not explain, because I have not even thought of it\".\n\nIOW show me a reasonable example, and we'll talk business.\n\nCiao,\nDscho\n"},{"id":"45578","messageId":"20070622175117.GA16921@yugib.highrise.ca","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706221626450.4059@racer.site","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Aidan Van Dyk","fromEmail":"aidan@highrise.ca","sentAt":"2007-06-22T17:51:17Z","receivedAt":"2007-06-22T17:51:17Z","isPatch":true,"sender":{"key":"aidan@highrise.ca","avatar":"https://gravatar.com/avatar/853c50d90cce753dc1c390fdc6cbed558f5f969bd43fa4f5cb0118d8f71316f6?d=mp&s=160"},"body":"* Johannes Schindelin <Johannes.Schindelin@gmx.de> [070622 13:34]:\n> Hi,\n> \n> On Fri, 22 Jun 2007, Andy Parkins wrote:\n> \n> > What if two files with different filenames and content converge at some \n> > point in history, then diverge again?  If git is tracking renames merely \n> > by content and picks the wrong one, then the history of fileA suddenly \n> > becomes the history of fileB.\n> \n> This is becoming highly ethereal. Like \"I could imagine that some day in \n> future, some person could devise a device, that might allow you to do \n> something that I can not explain, because I have not even thought of it\".\n> \n> IOW show me a reasonable example, and we'll talk business.\n\nThe one time the \"content-only\" rename tracking bit me was the\nafter a merge, resulting in conflicts that were un-nessesary:\n\n-*-*-*-*-A-B-C-D\n\t  \\\n\t   *-E-*\n\nAt A, there were 2 files:\n\tdir1/foo\n\tdir2/foo\nThey were template files that happened to be the same in 2 themes.\n\nIn E, \"foo\" was renamed to \"foo-bar\" in all the template directories.\nGit detected this not as 2 renames, but as:\n\tdir1/foo-bar renamed from dir1/foo\n\tdir2/foo-bar copied from dir1/foo\n\tdir2/foo deleted\n\nMeanwhile, work was happening in B, C, and D, changing foo in both\ntemplates identically.\n\nWhen the branch with E was merged back into ABCD, there was a merge\nconflict with dir2/foo being deleted in one branch, and editit in the\nother.\n\nIn this case, the simple \"basename\" comparison wouldn't have even been\nenough.  \n\nBut the merge was easy enough (because no edits were made in the E\nbranch to those files, just the renames) that I could resolve it easily.\n\nI don't know if preventing this easy-to-fix merge conflict is worth the\nnecessary \"likeness of names\" necessary to avoid it...\n\na.\n\n-- \nAidan Van Dyk                                             Create like a god,\naidan@highrise.ca                                       command like a king,\nhttp://www.highrise.ca/                                   work like a slave.\n"},{"id":"45588","messageId":"Pine.LNX.4.64.0706230222330.4059@racer.site","threadId":"8668","inReplyTo":"86abusi1fw.fsf@lola.quinscape.zz","subject":"Re: 100%","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-23T01:31:28Z","receivedAt":"2007-06-23T01:31:28Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 22 Jun 2007, David Kastrup wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > On Fri, 22 Jun 2007, David Kastrup wrote:\n> >\n> >> As a note aside: would it be possible to always round downwards when \n> >> computing similarities or converting between them?\n> >\n> > I'd rather not. This would be counterintuitive. People expect rounded \n> > values.\n> \n> Which people?\n\nMe, for one. Thank you very much.\n\n> The people I know will expect \"100% identical\" or even \"100.0% \n> identical\" to mean identical, period.  They will be quite surprised to \n> hear that \"99.95%\" is supposed to be included.\n\nGranted, 100.0% means as close as you can get to \"completely\" with 4 \ndigits. But if you have an integer, you better use the complete range, \nrather than arbitrarily make one number more important than others.\n\nFor if you see an integer, you usually assume a rounded value. If you \ndon't, you're hopeless.\n\n> Also, for any kind of decision made upon percentages, it is much more \n> relevant to be able to draw a line at 50% rather than at 49.5%.\n\nI do see too many people in my day job who take the numbers they see for \nabsolute truths, so I cannot take that statement seriously, sorry.\n\n> Could you name a _single_ use case where rounding down could cause an \n> actual problem or even inconvenience for people?\n\nCould you name a _single_ use case where it does not?\n\nI mean, honestly, really. Really, really, really. A number is only a weak \n_indicator_, and an integer even more so, for what is _really_ going on.\n\n> >> I very much would like to see the 100% figure reserved for identity.  \n> >> This is particularly relevant when interpreting the output of \n> >> git-diff --name-status with regard to R100, C100 and similar flags.\n> >\n> > You should never depend on the output of --name-status if you're \n> > interested in identifying identical files, but on the object names.\n> \n> Which is rather inconvenient.\n\nFrankly, I am getting bored.\n\nThis argument crops up ever so often. \"If you did that, _I_ could be more \nlazy, and the _hell_ with other people who expect otherwise!\".\n\nNo, really.\n\n> I _know_ that one can't rely on the output of --name-status right now.\n\nAnd I _know_ that you can't rely on integer numbers. Or _any_ number which \nis not _completely_ precise.\n\nReally, I am getting bored with this discussion.\n\nCiao,\nDscho\n"},{"id":"45590","messageId":"7vd4zn9qdb.fsf@assigned-by-dhcp.pobox.com","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706211248420.4059@racer.site","subject":"Re: [PATCH] diffcore-rename: favour identical basenames","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-06-23T05:44:32Z","receivedAt":"2007-06-23T05:44:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> When there are several candidates for a rename source, and one of them\n> has an identical basename to the rename target, take that one.\n>\n> Noticed by Govind Salinas, posted by Shawn O. Pearce, partial patch\n> by Linus Torvalds.\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>\n> \tOn Wed, 20 Jun 2007, Linus Torvalds wrote:\n>\n> \t> I think we should just consider the basename as an \"added \n> \t> similarity  bonus\".\n\nThanks, I obviously agree with both of you.\n"},{"id":"45599","messageId":"467CF380.6060603@lsrfire.ath.cx","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706230222330.4059@racer.site","subject":"Re: 100%","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2007-06-23T10:18:40Z","receivedAt":"2007-06-23T10:18:40Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Johannes Schindelin schrieb:\n> On Fri, 22 Jun 2007, David Kastrup wrote:\n>> The people I know will expect \"100% identical\" or even \"100.0% \n>> identical\" to mean identical, period.  They will be quite surprised to \n>> hear that \"99.95%\" is supposed to be included.\n> \n> Granted, 100.0% means as close as you can get to \"completely\" with 4 \n> digits. But if you have an integer, you better use the complete range, \n> rather than arbitrarily make one number more important than others.\n> \n> For if you see an integer, you usually assume a rounded value. If you \n> don't, you're hopeless.\n\nWhy hopeless?  It's a useful convention to define \"100%\" as \"complete\n(not rounded)\".  See it this way: 50% of the time, a given percent value\nwill be shown as one point less than it's \"true\" value, but you gain the\nability to indicate full completeness.  And that's an interesting piece\nof information.  The price is small given that the needed accuracy is\nmore in the range of 10 percent points (I assume).\n\nIt's more a question of how to make sure everybody knows what the\nnumbers mean -- but that's why we have a directory named\n\"Documentation\". :-D  And even a person that hasn't read the docs is\nunlikely to really get harmed by inexact percentages, right?\n\nRené\n"},{"id":"45601","messageId":"Pine.LNX.4.64.0706231154300.4059@racer.site","threadId":"8668","inReplyTo":"467CF380.6060603@lsrfire.ath.cx","subject":"Re: 100%","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-23T10:56:19Z","receivedAt":"2007-06-23T10:56:19Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sat, 23 Jun 2007, René Scharfe wrote:\n\n> Johannes Schindelin schrieb:\n> > On Fri, 22 Jun 2007, David Kastrup wrote:\n> >> The people I know will expect \"100% identical\" or even \"100.0% \n> >> identical\" to mean identical, period.  They will be quite surprised to \n> >> hear that \"99.95%\" is supposed to be included.\n> > \n> > Granted, 100.0% means as close as you can get to \"completely\" with 4 \n> > digits. But if you have an integer, you better use the complete range, \n> > rather than arbitrarily make one number more important than others.\n> > \n> > For if you see an integer, you usually assume a rounded value. If you \n> > don't, you're hopeless.\n> \n> Why hopeless?  It's a useful convention to define \"100%\" as \"complete\n> (not rounded)\".\n\nBy the same reasoning, you could say \"never round down to 0%, because I \nwant to know when there is no similarity\".\n\nYou cannot be exact when you have to cut off fractions, so why try for \n_exactly_ one number?\n\nCiao,\nDscho\n"},{"id":"45604","messageId":"467D06D4.9050203@lsrfire.ath.cx","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706231154300.4059@racer.site","subject":"Re: 100%","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2007-06-23T11:41:08Z","receivedAt":"2007-06-23T11:41:08Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Johannes Schindelin schrieb:\n> Hi,\n> \n> On Sat, 23 Jun 2007, René Scharfe wrote:\n> \n>> Johannes Schindelin schrieb:\n>>> On Fri, 22 Jun 2007, David Kastrup wrote:\n>>>> The people I know will expect \"100% identical\" or even \"100.0% \n>>>> identical\" to mean identical, period.  They will be quite surprised to \n>>>> hear that \"99.95%\" is supposed to be included.\n>>> Granted, 100.0% means as close as you can get to \"completely\" with 4 \n>>> digits. But if you have an integer, you better use the complete range, \n>>> rather than arbitrarily make one number more important than others.\n>>>\n>>> For if you see an integer, you usually assume a rounded value. If you \n>>> don't, you're hopeless.\n>> Why hopeless?  It's a useful convention to define \"100%\" as \"complete\n>> (not rounded)\".\n> \n> By the same reasoning, you could say \"never round down to 0%, because I \n> want to know when there is no similarity\".\n> \n> You cannot be exact when you have to cut off fractions, so why try for \n> _exactly_ one number?\n\nBecause completeness is special.  If just one bit was available, I'd use\nit to indicate equality.  That's what the authors of cmp(1) did, too. :)\n\nAnd 0% is not special, at least not in a useful way that I can think of.\n  I.e. there is no practical difference between \"no two lines match\" and\n\"one percent of the lines match\".  If you're really interested in\nsimilarities with an index below 10% then you'd better work with\nabsolute numbers instead of rounded percentages.\n\nIf someone came around with an interest in those cases with exactly 0%\nsimilarity, then we might need to decide between rounding up or down.\nBut even in that hypothetical situation I think \"equality\" is still more\ninteresting a data point than \"really everything differs\".\n\nRené\n"},{"id":"45605","messageId":"Pine.LNX.4.64.0706231259021.4059@racer.site","threadId":"8668","inReplyTo":"467D06D4.9050203@lsrfire.ath.cx","subject":"Re: 100%","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-23T12:00:13Z","receivedAt":"2007-06-23T12:00:13Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sat, 23 Jun 2007, René Scharfe wrote:\n\n> Johannes Schindelin schrieb:\n>\n> > By the same reasoning, you could say \"never round down to 0%, because \n> > I want to know when there is no similarity\".\n> > \n> > You cannot be exact when you have to cut off fractions, so why try for \n> > _exactly_ one number?\n> \n> Because completeness is special.\n\nI am not convinced. My vote is still for the _common_ practice of just \nrounding. IOW keep it as is.\n\nCiao,\nDscho\n"},{"id":"45606","messageId":"467D0DE8.6030104@lsrfire.ath.cx","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706231259021.4059@racer.site","subject":"Re: 100%","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2007-06-23T12:11:20Z","receivedAt":"2007-06-23T12:11:20Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Johannes Schindelin schrieb:\n> Hi,\n> \n> On Sat, 23 Jun 2007, René Scharfe wrote:\n> \n>> Johannes Schindelin schrieb:\n>>\n>>> By the same reasoning, you could say \"never round down to 0%, because \n>>> I want to know when there is no similarity\".\n>>>\n>>> You cannot be exact when you have to cut off fractions, so why try for \n>>> _exactly_ one number?\n>> Because completeness is special.\n> \n> I am not convinced. My vote is still for the _common_ practice of just \n> rounding. IOW keep it as is.\n\nAs I already hinted at, the common result of comparing two files, as\ndone by e.g. cmp(1), is one bit that indicates equality.  This\ninformation is lost when using up/down rounding, but it is retained when\nrounding down.  It's _not_ common to be unable to determine equality\nfrom the result of a file compare.\n\nRené\n"},{"id":"45607","messageId":"Pine.LNX.4.64.0706231318180.4059@racer.site","threadId":"8668","inReplyTo":"467D0DE8.6030104@lsrfire.ath.cx","subject":"Re: 100%","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-23T12:21:00Z","receivedAt":"2007-06-23T12:21:00Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sat, 23 Jun 2007, René Scharfe wrote:\n\n> As I already hinted at, the common result of comparing two files, as \n> done by e.g. cmp(1), is one bit that indicates equality.  This \n> information is lost when using up/down rounding, but it is retained when \n> rounding down.  It's _not_ common to be unable to determine equality \n> from the result of a file compare.\n\nAnd as _I_ already hinted, this does not matter. The whole purpose to have \na number here instead of a bit is to have a larger range. In practice, I \nbet that the 100% are really uninteresting. At least here, they are.\n\nFor example, if you move a Java class from one package into another, you \nhave to change the package name in the file. Guess what, I am perfectly \nokay if the rename detector says \"100% similarity\" here. Because if it is \ncloser to 100% than to 99%, dammit, I want to see 100%, not 99%.\n\nNuff said about this subject.\n\nCiao,\nDscho\n"},{"id":"45619","messageId":"7vk5tu4gas.fsf@assigned-by-dhcp.pobox.com","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706231154300.4059@racer.site","subject":"Re: 100%","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-06-23T19:33:15Z","receivedAt":"2007-06-23T19:33:15Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> By the same reasoning, you could say \"never round down to 0%, because I \n> want to know when there is no similarity\".\n>\n> You cannot be exact when you have to cut off fractions, so why try for \n> _exactly_ one number?\n\nR0 or C0 would not happen in real life, so 0% is a moot issue.\n\nHowever, wasn't that you who did follow that \"certain numbers\nare special\" logic in diffstat?\n\nYou advocated \"diff --stat\" should draw at least one +/- for a\npatch that adds/removes lines.  And I (and others) agreed\nbecause zero is special in the context of that application.\n\nI think reserving R100 to mean \"identical byte sequences\" has\nvalue, when people look at --name-status output, in the context\nof \"similarity index\".\n"},{"id":"45623","messageId":"Pine.LNX.4.64.0706232130230.4059@racer.site","threadId":"8668","inReplyTo":"7vk5tu4gas.fsf@assigned-by-dhcp.pobox.com","subject":"Re: 100%","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-23T20:41:16Z","receivedAt":"2007-06-23T20:41:16Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sat, 23 Jun 2007, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > By the same reasoning, you could say \"never round down to 0%, because I \n> > want to know when there is no similarity\".\n> >\n> > You cannot be exact when you have to cut off fractions, so why try for \n> > _exactly_ one number?\n> \n> R0 or C0 would not happen in real life, so 0% is a moot issue.\n\nIt is, but not when you look at the formula.\n\n> However, wasn't that you who did follow that \"certain numbers\n> are special\" logic in diffstat?\n> \n> You advocated \"diff --stat\" should draw at least one +/- for a\n> patch that adds/removes lines.  And I (and others) agreed\n> because zero is special in the context of that application.\n\nActually, it was not me, but I implemented the version that we have now. \nI was reasonably scared that a non-linear diffstat would end up in git, \ntherefore I wrote a linear one.\n\nThe important thing to not here is that the diffstat as-is makes _better_ \nuse of the limited scale that is available.\n\nAnd as you pointed out, the low end of the scale is not really \ninteresting. The interesting parts are those around 100%. By rounding down \nyou make less use of the available scale.\n\n> I think reserving R100 to mean \"identical byte sequences\" has value, \n> when people look at --name-status output, in the context of \"similarity \n> index\".\n\nAh, whatever. You do what you want.\n\nYes, this interpretation has value. No, it is not the only one that has \nvalue. I am much more used to rounding, since at the end of the day it \nmakes better use of the scale, it is commonly used, and therefore _I_ \nexpect it (and no, I will not read the documentation when I expect to know \nwhat it means).\n\nBut hey, I don't care any more. AFAIAC you can change it from rounding to \nrounding down, and next year to rounding-up. We could even have an \nalgorithm which rounds down only in odd years, and I still would not care.\n\nCiao,\nDscho\n"},{"id":"45705","messageId":"467EEEDB.9030301@lsrfire.ath.cx","threadId":"8668","inReplyTo":"Pine.LNX.4.64.0706231318180.4059@racer.site","subject":"Re: 100%","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2007-06-24T22:23:23Z","receivedAt":"2007-06-24T22:23:23Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Johannes Schindelin schrieb:\n> Hi,\n> \n> On Sat, 23 Jun 2007, René Scharfe wrote:\n> \n>> As I already hinted at, the common result of comparing two files, as \n>> done by e.g. cmp(1), is one bit that indicates equality.  This \n>> information is lost when using up/down rounding, but it is retained when \n>> rounding down.  It's _not_ common to be unable to determine equality \n>> from the result of a file compare.\n> \n> And as _I_ already hinted, this does not matter. The whole purpose to have \n> a number here instead of a bit is to have a larger range. In practice, I \n> bet that the 100% are really uninteresting. At least here, they are.\n\nYou would lose your bet since both David and me expressed interest in\nthat pure 100% thing.\n\nRounding down instead of up/down doesn't affect the size of neither the\ninput nor the output range.  It affects the boundary of the input range,\n (-0.499 .. 100.499 versus 0.000 .. 100.999), but I can't find a problem\nwith that.\n\n> For example, if you move a Java class from one package into another, you \n> have to change the package name in the file. Guess what, I am perfectly \n> okay if the rename detector says \"100% similarity\" here. Because if it is \n> closer to 100% than to 99%, dammit, I want to see 100%, not 99%.\n\nThat uses a side effect of rounding and won't work for small files.  And\nof course (if the file is large enough) there could be other changes\n\"hidden\" in a similarity index value of 100% that was rounded up.\n\n> Nuff said about this subject.\n\nYes, let's advance this topic to the coding stage.\n\nRené\n"}]}