{"thread":{"id":"65571","subject":"git rename/moved status unreliable in ruby","startedAt":"2026-05-01T05:06:10Z","lastAt":"2026-05-05T00:46:24Z","messageCount":9,"participants":["sebastien.stettler","Phillip Wood","Johannes Sixt","Chris Torek","Junio C Hamano","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"542539","messageId":"OsOzcjEwvHCQSghLE8LD_wHb_jDlil9I88OUuhpiRONnVd1o9p3gStbK1mx4q7OwY3ePtbZO-BBgTNOCeJ2DMyvBsdlMhRmDrTP894KP5xo=@proton.me","threadId":"65571","inReplyTo":null,"subject":"git rename/moved status unreliable in ruby","fromName":"sebastien.stettler","fromEmail":"sebastien.stettler@proton.me","sentAt":"2026-05-01T05:05:59Z","receivedAt":"2026-05-01T05:06:10Z","isPatch":false,"body":"1. What did you do before the bug happened? (Steps to reproduce your issue)\n\nwhen moving ruby classes between namespaces they are marked as new files and the old\nones are marked as deleted\n\nif i only change the class name it will mark it as renamed\n\n2. What did you expect to happen? (Expected behavior)\n\nin the namespace state i would expected it to be marked as moved since\nnothing has fundementally changed\n\n3. What happened instead? (Actual behavior)\n\nthe file was marked as new file and the old file was marked as deleted\n\nWhat's different between what you expected and what actually happened\n\n4. Anything else you want to add:\n\nI have demonstrated the behavior here https://github.com/billybonks/git-rename\n\nMostly i would like to understand what is the expectation from gits point of view in these mutations.\nIf this is considered something that can be improved i am happy to build out more test cases, and help with implementation.\n\nif not, understanding the reasoning would be great\n\nThank you.\n\n\n\n[System Info]\ngit version:\ngit version 2.47.1\ncpu: arm64\nno commit associated with this build\nsizeof-long: 8\nsizeof-size_t: 8\nshell-path: /bin/sh\nfeature: fsmonitor--daemon\nlibcurl: 8.7.1\nzlib: 1.2.12\nuname: Darwin 25.3.0 Darwin Kernel Version 25.3.0: Wed Jan 28 20:51:28 PST 2026; root:xnu-12377.91.3~2/RELEASE_ARM64_T6041 arm64\ncompiler info: clang: 16.0.0 (clang-1600.0.26.4)\nlibc info: no libc information available\n$SHELL (typically, interactive shell): /bin/zsh\n\n\n[Enabled Hooks]\n\n\nSent with Proton Mail secure email.\n"},{"id":"542553","messageId":"026b84f4-7052-4d5b-a9ae-c2487569d1ee@gmail.com","threadId":"65571","inReplyTo":"OsOzcjEwvHCQSghLE8LD_wHb_jDlil9I88OUuhpiRONnVd1o9p3gStbK1mx4q7OwY3ePtbZO-BBgTNOCeJ2DMyvBsdlMhRmDrTP894KP5xo=@proton.me","subject":"Re: git rename/moved status unreliable in ruby","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-05-01T15:30:16Z","receivedAt":"2026-05-01T15:30:19Z","isPatch":false,"body":"Hi Sebastien\n\nOn 01/05/2026 06:05, sebastien.stettler wrote:\n> 1. What did you do before the bug happened? (Steps to reproduce your issue)\n> \n> when moving ruby classes between namespaces they are marked as new files and the old\n> ones are marked as deleted\n> \n> if i only change the class name it will mark it as renamed\n\nRename detection is based on how similar the two files are. Looking at \nthe example you linked to below you're changing a file that looks like\n\nmodule Math\n   class Calculator\n     def add(a, b)\n       a + b\n       ...\n     end\n   end\nend\n\nto\n\nmodule Math\n   module Calculators\n     class Calculator\n       def add(a, b)\n         a + b\n         ...\n       end\n     end\n   end\nend\n\nWhich means that git sees that every line has changed because the \nindentation has changed. If you want git to realize that the file has \nbeen renamed you could move it in one commit and then add modify it in \nthe next commit.\n\nThanks\n\nPhillip\n\n> 2. What did you expect to happen? (Expected behavior)\n> \n> in the namespace state i would expected it to be marked as moved since\n> nothing has fundementally changed\n> \n> 3. What happened instead? (Actual behavior)\n> \n> the file was marked as new file and the old file was marked as deleted\n> \n> What's different between what you expected and what actually happened\n> \n> 4. Anything else you want to add:\n> \n> I have demonstrated the behavior here https://github.com/billybonks/git-rename\n> \n> Mostly i would like to understand what is the expectation from gits point of view in these mutations.\n> If this is considered something that can be improved i am happy to build out more test cases, and help with implementation.\n> \n> if not, understanding the reasoning would be great\n> \n> Thank you.\n> \n> \n> \n> [System Info]\n> git version:\n> git version 2.47.1\n> cpu: arm64\n> no commit associated with this build\n> sizeof-long: 8\n> sizeof-size_t: 8\n> shell-path: /bin/sh\n> feature: fsmonitor--daemon\n> libcurl: 8.7.1\n> zlib: 1.2.12\n> uname: Darwin 25.3.0 Darwin Kernel Version 25.3.0: Wed Jan 28 20:51:28 PST 2026; root:xnu-12377.91.3~2/RELEASE_ARM64_T6041 arm64\n> compiler info: clang: 16.0.0 (clang-1600.0.26.4)\n> libc info: no libc information available\n> $SHELL (typically, interactive shell): /bin/zsh\n> \n> \n> [Enabled Hooks]\n> \n> \n> Sent with Proton Mail secure email.\n> \n\n"},{"id":"542579","messageId":"157dd2e5-ce64-4654-a66a-86a552c7c9f1@kdbg.org","threadId":"65571","inReplyTo":"026b84f4-7052-4d5b-a9ae-c2487569d1ee@gmail.com","subject":"Re: git rename/moved status unreliable in ruby","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2026-05-02T07:25:16Z","receivedAt":"2026-05-02T07:25:20Z","isPatch":false,"body":"Am 01.05.26 um 17:30 schrieb Phillip Wood:\n> Rename detection is based on how similar the two files are. Looking at\n> the example you linked to below you're changing a file that looks like\n> \n> ...\n> \n> Which means that git sees that every line has changed because the\n> indentation has changed.\n\nThat's correct, of course, but...\n\n> If you want git to realize that the file has\n> been renamed you could move it in one commit and then add modify it in\n> the next commit.\n\n... this is a fallacy. Splitting into two commits helps only certain\ncases, in particular, when the commit that moves the files is compared\nto an earlier commit, such as `git log` does. However, if a commit after\nthe change is compared to a commit before the move, the rename is still\nnot detected.\n\n-- Hannes\n\n"},{"id":"542580","messageId":"CAPx1Gvd_VEWHrBWtUjNeWZ+wfmsAOTamKmL6fhBSQi=MbmXRcw@mail.gmail.com","threadId":"65571","inReplyTo":"OsOzcjEwvHCQSghLE8LD_wHb_jDlil9I88OUuhpiRONnVd1o9p3gStbK1mx4q7OwY3ePtbZO-BBgTNOCeJ2DMyvBsdlMhRmDrTP894KP5xo=@proton.me","subject":"Re: git rename/moved status unreliable in ruby","fromName":"Chris Torek","fromEmail":"chris.torek@gmail.com","sentAt":"2026-05-02T08:06:58Z","receivedAt":"2026-05-02T08:07:12Z","isPatch":false,"body":"On Thu, Apr 30, 2026 at 10:06 PM sebastien.stettler\n<sebastien.stettler@proton.me> wrote:\n> ... understanding the reasoning would be great\n\nIn my opinion, the key to understanding here is this:\n\nGit Stores Snapshots.\n\nWhat this means is that every commit is a full snapshot of all of the\nfiles for that commit. There are no \"changes\" at all, there is only a\nfull snapshot, every time.\n\nNow, internally, the storage format is more complicated (and\ncompressive, ultimately using the concept of changes as well, though\nnot exactly the way one might expect). But from the \"what things look\nlike\" point of view, and how you should think about what Git sees,\neach commit is simply a full and complete snapshot of every file. So\nif you have one commit where `foo.rb` exists, and `bar.br` does not,\nthat snapshot has a `foo.rb` but no `bar.rb`. If you make a second\nsnapshot, in which `foo.rb` no longer exists but `bar.rb` does now,\nthat second snapshot, well, has those files.\n\nThe tricky part is that you normally ask Git to *compare* two\nsnapshots (at least for \"what changed\" purposes). When you do that,\nGit extracts both snapshots and, well, compares them. If `foo.rb` has\nbeen removed and `bar.rb` has been added, Git then goes on to compare\nthe *contents* of those two files.\n\nIf the contents match exactly, and you've asked Git to \"find renames\",\nGit will always say that the file that vanished from the first commit,\nonly to be created identically under a new name in the second, was\n\"renamed\", rather than the one file being deleted and the second\nadded.\n\nIf the contents match \"fuzzily\" (for some value and algorithm of\nfuzz-factor), Git may also say \"renamed\". You can control this with\n`--find-renames=<value>`. The key idea here is that Git is *finding*\nrenames: either exact-same-contents, or \"sufficiently similar\"\ncontents, based on remove-and-add pairs.\n\nSince Git only *stores* snapshots, you can get two different results\nfrom comparing the same two commits. All you have to do to get this is\nto adjust whether Git checks for renames at all, and if so, to what\nextent.\n\nThese rules apply to `git show`, `git diff`, `git merge`, and even the\ndiffstat that `git commit` optionally shows after a commit. For this\nreason, all the \"compare some commits\" commands -- including `git\nmerge` -- take this `--find-renames=<value>` option. Detection of\nrenames can be countermanded entirely with `--no-renames`.\n\nThis is why -- and when -- making two separate commits, one with\n\"exact same content for deleted-file-D vs added-file-A\", followed by\nlater changes to new file A, helps: if you compare the commit that has\nfile D to the middle commit, the two files match exactly, and any\nrename detection you have turned on finds that rename. If you then\ncompare the middle commit to the final commit, file A exists in both,\nso Git shows changes to file A. But as soon as you compare the\noriginal file-D-containing commit to the final file-A-updated commit,\nyou run into the original issue again: to detect this as a rename, you\nmay need to allow rather generous rename detection.\n\nIf, in the future, Git gets fancier rename detection, comparing the\noriginal commit directly against the final one could find the rename\nautomatically. So:\n\n> If this is considered something that can be improved ...\n\nIt *could* be improved. Doing so in a way that works for more than\njust some special cases -- e.g., in a way that works for ordinary\ntext, or graphical images, for instance, rather than just for Ruby\nsources (or just C sources, or C++, or Swift, or Python, or whatever)\n-- seems particularly tricky. Some degree of ignoring white-space\nchanges would probably help multiple cases, though.\n\nChris\n"},{"id":"542582","messageId":"IC7a4NnSKMdvXlVyaSDYEtU7iRlKdJGzCwrXNCFKrtFfnBJTMrwY522rHF8PfzYxFs43huo0KFGrqB6f4IQjmvYi2B8Ehh0cwfjHHOYW_RU=@proton.me","threadId":"65571","inReplyTo":"CAPx1Gvd_VEWHrBWtUjNeWZ+wfmsAOTamKmL6fhBSQi=MbmXRcw@mail.gmail.com","subject":"Re: git rename/moved status unreliable in ruby","fromName":"sebastien.stettler","fromEmail":"sebastien.stettler@proton.me","sentAt":"2026-05-02T09:34:18Z","receivedAt":"2026-05-02T09:34:30Z","isPatch":false,"body":"Thanks for all the responses and thoughts thus far.\n\n> Which means that git sees that every line has changed because the\n> indentation has changed. If you want git to realize that the file has\n> been renamed you could move it in one commit and then add modify it in\n> the next commit.\n\nThis is the minimal solution to the problem but as Johannes illustrated,\nthere cases where that \"move commit doesn't solve the problem\"\n\n> Splitting into two commits helps only certain\n> cases, in particular, when the commit that moves the files is compared\n> to an earlier commit, such as `git log` does. However, if a commit after\n> the change is compared to a commit before the move, the rename is still\n> not detected.\n\n\nFurther examples:\n\nThe move commit helps if i want to blame a file if i use the -W option, but if it doesnt\nhelp with git log since the changes are now considered modifications.\n\ngit blame -w ruby-example/lib/calculator.rb\n(shows previous files changes and not the indentation)\n```\n303f25f5 ruby-example/lib/calculator.rb (nothing    2026-05-02 16:40:51 +0800  1) module Lib\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800  2)   class Calculator\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800  3)     def add(a, b)\n3d22d303 ruby-example/calculator.rb     (billybonks 2026-05-02 16:34:14 +0800  4)       a + b\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800  5)     end\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800  6)\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800  7)     def subtract(a, b)\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800  8)       a - b\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800  9)     end\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 10)\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 11)     def multiply(a, b)\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 12)       a * b\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 13)     end\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 14)\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 15)     def divide(a, b)\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 16)       raise ZeroDivisionError, 'Cannot divide by zero' if b.zero?\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 17)\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 18)       a / b\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 19)     end\n^1cf274f ruby-example/calculator.rb     (billybonks 2026-05-01 12:36:47 +0800 20)   end\n303f25f5 ruby-example/lib/calculator.rb (nothing    2026-05-02 16:40:51 +0800 21) end\n```\n\n\ngit log --follow --diff-filter=ra -- ruby-example/lib/calculator.rb\n( it does not show the move commit )\n```\ncommit 303f25f50cde83c46f83bc3c337cd52b87b63d52 (HEAD -> example-2)\nAuthor: nothing <nothing@contributed.com>\nDate:   Sat May 2 16:40:51 2026 +0800\n\n    update lib and require\n\ncommit 3d22d30304d836c07b3274689e5e2c536b29e1bb (example-3)\nAuthor: billybonks <sebastienstettler@gmail.com>\nDate:   Sat May 2 16:34:14 2026 +0800\n\n    fix: sum was doing subtraction instead of addition\n```\nUsing this approach does incur a heavy burnder on the user since most tools that do renaming etc \nwill move and rename, and that cost does not give a complete solution as there are still many thigns\nthat don't work in the most ideal sense.\n\n\nAs mentioned by Chris\n\n> It *could* be improved. Doing so in a way that works for more than\n> just some special cases -- e.g., in a way that works for ordinary\n> text, or graphical images, for instance, rather than just for Ruby\n> sources (or just C sources, or C++, or Swift, or Python, or whatever)\n> -- seems particularly tricky. Some degree of ignoring white-space\n> changes would probably help multiple cases, though.\n\n\nHas there been explorations of ignoring white space for the similarity checker, i would \nassume that majority of white space movements across many languages would result in a \nsemantically similar document in most cases.\n \n- Sebastien\n\n\n\n\n\n\n\n"},{"id":"542623","messageId":"xmqqo6iwrw6k.fsf@gitster.g","threadId":"65571","inReplyTo":"026b84f4-7052-4d5b-a9ae-c2487569d1ee@gmail.com","subject":"Re: git rename/moved status unreliable in ruby","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-03T21:59:47Z","receivedAt":"2026-05-03T21:59:49Z","isPatch":false,"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n> Hi Sebastien\n>\n> On 01/05/2026 06:05, sebastien.stettler wrote:\n>> 1. What did you do before the bug happened? (Steps to reproduce your issue)\n>> \n>> when moving ruby classes between namespaces they are marked as new files and the old\n>> ones are marked as deleted\n>> \n>> if i only change the class name it will mark it as renamed\n>\n> Rename detection is based on how similar the two files are. Looking at \n> the example you linked to below you're changing a file that looks like\n> ...\n> Which means that git sees that every line has changed because the \n> indentation has changed. If you want git to realize that the file has \n> been renamed you could move it in one commit and then add modify it in \n> the next commit.\n\nI've seen this repeated many times, but it is misleading to give it\nwithout qualifying when that \"works\" and when it does not.  Such a\n\"stick to pure rename and make huge changes elsewhere\" strategy\nwould help your \"git log -M [--follow]\" and possibly \"git rebase\",\nbut it would not help all that much if you are doing \"git diff\" or\n\"git merge\".\n\n\n    \n"},{"id":"542655","messageId":"20260504100056.GB599780@coredump.intra.peff.net","threadId":"65571","inReplyTo":"IC7a4NnSKMdvXlVyaSDYEtU7iRlKdJGzCwrXNCFKrtFfnBJTMrwY522rHF8PfzYxFs43huo0KFGrqB6f4IQjmvYi2B8Ehh0cwfjHHOYW_RU=@proton.me","subject":"Re: git rename/moved status unreliable in ruby","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-05-04T10:00:56Z","receivedAt":"2026-05-04T10:00:58Z","isPatch":false,"body":"On Sat, May 02, 2026 at 09:34:18AM +0000, sebastien.stettler wrote:\n\n> Has there been explorations of ignoring white space for the similarity checker, i would \n> assume that majority of white space movements across many languages would result in a \n> semantically similar document in most cases.\n\nI don't think anybody has ever looked into it. We do have \"-w\" and\nfriends for diffs, and it makes sense that there might be some mode to\nsoften renames in the same way (especially if you are doing a \"-w\"\ndiff, or a merge that ignores whitespace).\n\nThe line you need to touch is probably this:\n\ndiff --git a/diffcore-delta.c b/diffcore-delta.c\nindex 2b7db39983..379f6010d3 100644\n--- a/diffcore-delta.c\n+++ b/diffcore-delta.c\n@@ -147,6 +147,8 @@ static struct spanhash_top *hash_chars(struct repository *r,\n \t\t/* Ignore CR in CRLF sequence if text */\n \t\tif (is_text && c == '\\r' && sz && *buf == '\\n')\n \t\t\tcontinue;\n+\t\tif (is_text && (c == ' ' || c == '\\t'))\n+\t\t\tcontinue;\n \n \t\taccum1 = (accum1 << 7) ^ (accum2 >> 25);\n \t\taccum2 = (accum2 << 7) ^ (old_1 >> 25);\n\nbut:\n\n  1. The option to ignore whitespace would need to be plumbed through\n     the rest of the diffcore code.\n\n  2. This concept probably throws off some other rename heuristics.\n     E.g., I think we do a rough check that the sizes of the objects are\n     not too far apart before even looking at the content. So you could\n     construct a pathological case where the line \"a\\n\" was changed to\n     have a million spaces, and the files would look like they couldn't\n     possibly be similar, even though they are identical when ignoring\n     whitespace. I think in practice you could just ignore this, as sane\n     cases would tend to have a reasonable ratio of content to\n     whitespace changes.\n\n-Peff\n"},{"id":"542734","messageId":"xmqqecjqpvhw.fsf@gitster.g","threadId":"65571","inReplyTo":"CAPx1Gvd_VEWHrBWtUjNeWZ+wfmsAOTamKmL6fhBSQi=MbmXRcw@mail.gmail.com","subject":"Re: git rename/moved status unreliable in ruby","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-05T00:09:47Z","receivedAt":"2026-05-05T00:09:49Z","isPatch":false,"body":"Chris Torek <chris.torek@gmail.com> writes:\n\n> This is why -- and when -- making two separate commits, one with\n> \"exact same content for deleted-file-D vs added-file-A\", followed by\n> later changes to new file A, helps: if you compare the commit that has\n\n\"helps\" -> \"somtimes helps\".  Only when comparison is done step-wise\n(e.g., \"git log -M/--follow\" and \"git rebase\"), it may help, but in\ngeneral, when comparison between only two endpoints matter (e.g.,\n\"git diff\" and \"git merge\"), such an artificial breaking of a\nlogically single change into two does not help.\n\n>> If this is considered something that can be improved ...\n>\n> It *could* be improved. Doing so in a way that works for more than\n> just some special cases -- e.g., in a way that works for ordinary\n> text, or graphical images, for instance, rather than just for Ruby\n> sources (or just C sources, or C++, or Swift, or Python, or whatever)\n> -- seems particularly tricky. Some degree of ignoring white-space\n> changes would probably help multiple cases, though.\n\nYou could tie it with the attributes system to allow logic\nspecialized for the nature of the contents.  The beauty of the\ndesign decision to store \"snapshots\" is that these heuristics can be\nimproved without having to change anything in the history that are\ncast in stone.\n"},{"id":"542737","messageId":"CAPx1GveSn30Ua6fD3ZhiRHiN+-DpcN=9FbUcY3GstiXz9UYZ_Q@mail.gmail.com","threadId":"65571","inReplyTo":"xmqqecjqpvhw.fsf@gitster.g","subject":"Re: git rename/moved status unreliable in ruby","fromName":"Chris Torek","fromEmail":"chris.torek@gmail.com","sentAt":"2026-05-05T00:46:10Z","receivedAt":"2026-05-05T00:46:24Z","isPatch":false,"body":"On Mon, May 4, 2026 at 5:09 PM Junio C Hamano <gitster@pobox.com> wrote:\n> Chris Torek <chris.torek@gmail.com> writes:\n>\n> > This is why -- and when -- making two separate commits ... helps\n>\n> \"helps\" -> \"somtimes helps\".  Only when comparison is done step-wise\n> (e.g., \"git log -M/--follow\" and \"git rebase\"), it may help,\n\nThat's why I said \"and when\". :-)\n\n > > ... Some degree of ignoring white-space\n> > changes would probably help multiple cases, though.\n>\n> You could tie it with the attributes system to allow logic\n> specialized for the nature of the contents.  The beauty of the\n> design decision to store \"snapshots\" is that these heuristics can be\n> improved without having to change anything in the history that are\n> cast in stone.\n\nIndeed. Something like Peff's suggestion might work, although I\nsee some danger in ignoring white space completely. It would\nprobably be better to compress \"all leading but non-empty white\nspace\" to either nothing or a single space, eliminate all trailing\nwhite space, and compress other white space to a single blank.\n\n(Though at the same time, when we're dealing with slugs extracted\nfrom very long single lines, this is probably wrong, so perhaps this\nshould only be done for \"intact single line\" slugs. Then again it\nmight not matter at this point.)\n\nDoing this on binary files and programs written in Whitespace[1]\nwould be wrong, of course. ;-)\n\nChris\n\n[1]: https://en.wikipedia.org/wiki/Whitespace_(programming_language)\n"}]}