Re: [BUG] "git diff --word-diff" gives a diff while they are only space changes
- From
- Phillip Wood <phillip.wood123@gmail.com>
- Date
- May 15, 2026, 13:22 UTC
- Message-ID
- <d888553a-8a99-4c79-b720-a562ac9900e2@gmail.com>
- In-Reply-To
- <20260514095522.GA159111@qaa.vinc17.org>
On 14/05/2026 10:55, Vincent Lefevre wrote:
Show 36 quoted lines
> On 2026-05-14 16:37:39 +0900, Junio C Hamano wrote: >> Michael Montalbo <mmontalbo@gmail.com> writes: >> >>> @@ -457,6 +457,11 @@ endif::git-diff[] >>> + >>> Note that despite the name of the first mode, color is used to >>> highlight the changed parts in all modes if enabled. >>> ++ >>> +Word diff works by finding word-level changes within each hunk of >>> +the line-level diff. The line-level alignment determines which >>> +changed lines are compared to each other, which can affect the >>> +word-level output. >> >> The added text may not say anything wrong, but I am not sure how it >> helps the end user to know the way machinery works internally. > > Perhaps only the first sentence should be kept and that the following > should be added: "Because of that, using the --ignore-space-change > option is recommended." > > Note: Earlier in the discussion, Johannes Sixt suggested -w > (--ignore-all-space), but this is wrong, as > > git diff --word-diff -w <(printf foo) <(printf "f o o") > > gives no differences while one has 1 word "foo" vs 3 words "f o o". > > However, --ignore-space-change is actually not even sufficient > since > > git diff --ignore-space-change <(printf "foo bar") <(printf "foo\nbar") > > finds differences though there are only space changes (thus this > may affect hunks in case --word-diff would be used too). However, > I suppose that the cases where --word-diff --ignore-space-change > would not give a "real word diff" would be quite rare in practice.
I'm a bit wary of recommending -w unconditionally in case it gives unexpected results. I've not really found there to be a problem using --word-diff when reviewing code patches. In the examples you gave we'd ideally fix the problem by computing a single word-diff per hunk from the line based diff rather than splitting the hunk at each context line. I think we'd probably want to exclude the leading and trailing context to keep the hunk header accurate but we'd get better results by calculating the word diff of everything in the hunk between the first changed line and the last changed line.
Thanks
Phillip