From: Jeff King Date: Mon, 04 May 2026 10:00:56 GMT Subject: Re: git rename/moved status unreliable in ruby Message-ID: <20260504100056.GB599780@coredump.intra.peff.net> In-Reply-To: On Sat, May 02, 2026 at 09:34:18AM +0000, sebastien.stettler wrote: > Has there been explorations of ignoring white space for the similarity checker, i would > assume that majority of white space movements across many languages would result in a > semantically similar document in most cases. I don't think anybody has ever looked into it. We do have "-w" and friends for diffs, and it makes sense that there might be some mode to soften renames in the same way (especially if you are doing a "-w" diff, or a merge that ignores whitespace). The line you need to touch is probably this: diff --git a/diffcore-delta.c b/diffcore-delta.c index 2b7db39983..379f6010d3 100644 --- a/diffcore-delta.c +++ b/diffcore-delta.c @@ -147,6 +147,8 @@ static struct spanhash_top *hash_chars(struct repository *r, /* Ignore CR in CRLF sequence if text */ if (is_text && c == '\r' && sz && *buf == '\n') continue; + if (is_text && (c == ' ' || c == '\t')) + continue; accum1 = (accum1 << 7) ^ (accum2 >> 25); accum2 = (accum2 << 7) ^ (old_1 >> 25); but: 1. The option to ignore whitespace would need to be plumbed through the rest of the diffcore code. 2. This concept probably throws off some other rename heuristics. E.g., I think we do a rough check that the sizes of the objects are not too far apart before even looking at the content. So you could construct a pathological case where the line "a\n" was changed to have a million spaces, and the files would look like they couldn't possibly be similar, even though they are identical when ignoring whitespace. I think in practice you could just ignore this, as sane cases would tend to have a reasonable ratio of content to whitespace changes. -Peff