{"thread":{"id":"58868","subject":"[bug] git diff --word-diff gives wrong result for utf-8 chinese","startedAt":"2022-11-29T03:46:59Z","lastAt":"2022-12-01T20:07:07Z","messageCount":12,"participants":["Ping Yin","Bagas Sanjaya","Ævar Arnfjörð Bjarmason","Junio C Hamano","Jeff King","Phillip Wood"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"468153","messageId":"CACSwcnQfTOYHxSJQqc+viiqkCqt=WZieuCw70PqOdvo88XdeOQ@mail.gmail.com","threadId":"58868","inReplyTo":null,"subject":"[bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Ping Yin","fromEmail":"pkufranky@gmail.com","sentAt":"2022-11-29T03:46:43Z","receivedAt":"2022-11-29T03:46:59Z","isPatch":false,"sender":{"key":"pkufranky@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5346?v=4"},"body":"Result of \"git diff\"\n\n-  为1\n+  为2\n\nor (if chinese can not be displayed correctly)\n\n-  <E4><B8><BA>1\n+  <E4><B8><BA>2\n\nActual result of \"git diff --color-words\"\n\n<E4><B8>[-<BA>1-]{+<BA>2+}\n\nExpected result of \"git diff --color-words\"\n\n为[-1-]{+2+}\n\nor (if chinese can not be displayed correctly)\n\n<E4><B8><BA>[-1-]{+2+}\n\n\nPing Yin\n"},{"id":"468154","messageId":"CACSwcnT9Pz3snq4Jp6K5qxHFiE_zo41bKVUjJ_LJ39WN7h=gbQ@mail.gmail.com","threadId":"58868","inReplyTo":"CACSwcnQfTOYHxSJQqc+viiqkCqt=WZieuCw70PqOdvo88XdeOQ@mail.gmail.com","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Ping Yin","fromEmail":"pkufranky@gmail.com","sentAt":"2022-11-29T03:49:09Z","receivedAt":"2022-11-29T03:49:25Z","isPatch":false,"sender":{"key":"pkufranky@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5346?v=4"},"body":"sorry, typo, s/--color-words/--word-diff/g\n\nPing Yin\n\nOn Tue, Nov 29, 2022 at 11:46 AM Ping Yin <pkufranky@gmail.com> wrote:\n>\n> Result of \"git diff\"\n>\n> -  为1\n> +  为2\n>\n> or (if chinese can not be displayed correctly)\n>\n> -  <E4><B8><BA>1\n> +  <E4><B8><BA>2\n>\n> Actual result of \"git diff --color-words\"\n>\n> <E4><B8>[-<BA>1-]{+<BA>2+}\n>\n> Expected result of \"git diff --color-words\"\n>\n> 为[-1-]{+2+}\n>\n> or (if chinese can not be displayed correctly)\n>\n> <E4><B8><BA>[-1-]{+2+}\n>\n>\n> Ping Yin\n"},{"id":"468160","messageId":"455e9d7e-7394-bad8-c8f2-3ddf3958f1a2@gmail.com","threadId":"58868","inReplyTo":"CACSwcnT9Pz3snq4Jp6K5qxHFiE_zo41bKVUjJ_LJ39WN7h=gbQ@mail.gmail.com","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Bagas Sanjaya","fromEmail":"bagasdotme@gmail.com","sentAt":"2022-11-29T08:18:37Z","receivedAt":"2022-11-29T08:18:49Z","isPatch":false,"sender":{"key":"bagasdotme@gmail.com","avatar":"https://avatars.githubusercontent.com/u/40219486?v=4"},"body":"On 11/29/22 10:49, Ping Yin wrote:\n> sorry, typo, s/--color-words/--word-diff/g\n> \n\nHi, welcome to Git mailing list!\n\nPlease remind yourself:\n\n  * Do not send HTML mails, send plain-text ones instead. Many mailing\n    lists (including vger.kernel.org that powers Git ML) reject HTML\n    emails for these are likely spam. Make sure your email isn't mangled\n    (tabs and spaces as-is, no line wrapping).\n\n  * Do not top-post, reply inline with appropriate context instead. I\n    have to cut the reply context as a result.\n\n  * When you submit a patch and people reply with their reviews, engage\n    with them (either sending revised patch addressing the reviews or\n    reply with justification). They will ignore you if you ignore them.\n\nThanks.\n\n-- \nAn old man doll... just what I always wanted! - Clara\n\n"},{"id":"468165","messageId":"221129.867czejabi.gmgdl@evledraar.gmail.com","threadId":"58868","inReplyTo":"CACSwcnQfTOYHxSJQqc+viiqkCqt=WZieuCw70PqOdvo88XdeOQ@mail.gmail.com","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-11-29T10:52:38Z","receivedAt":"2022-11-29T10:59:20Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Nov 29 2022, Ping Yin wrote:\n\n> Result of \"git diff\"\n>\n> -  为1\n> +  为2\n>\n> or (if chinese can not be displayed correctly)\n>\n> -  <E4><B8><BA>1\n> +  <E4><B8><BA>2\n>\n> Actual result of \"git diff --color-words\"\n>\n> <E4><B8>[-<BA>1-]{+<BA>2+}\n>\n> Expected result of \"git diff --color-words\"\n>\n> 为[-1-]{+2+}\n>\n> or (if chinese can not be displayed correctly)\n\nI think we could provide new ways to do per-language diffs, right now\nyou can use --word-diff-regex, but it would be handy to e.g. have a\nbuilt-in collection of those (or other non-regex boundary algorithms)\nfor Chinese etc.\n\nBut as for considering this a bug, or changing the existing behavior I\nthink we'd need to deal with:\n\n * We (approximately) split on space now, which is certainly\n   ASCII-biased, and outside of CJK fairly somewhat universal.\n\n * If we're going to split on \"real words\" in some cross-language aware\n   way, are we going to run into conflicts between what different\n   languages would consider sensible rules?\n\n * We probably don't want to make the \"diff\" dependent on the user's\n   locale, but e.g. saying \"I want a Chinese diff\" via a CLI option\n   would be OK.\n\n * Even for say Chinese, there's probably interesting edge cases when\n   it's combined with other languages or character sets (e.g. Chinese +\n   HTML).\n\n\n\n"},{"id":"468166","messageId":"xmqqlenu2dxx.fsf@gitster.g","threadId":"58868","inReplyTo":"221129.867czejabi.gmgdl@evledraar.gmail.com","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-11-29T11:32:58Z","receivedAt":"2022-11-29T11:33:30Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n>> or (if chinese can not be displayed correctly)\n>>\n>> -  <E4><B8><BA>1\n>> +  <E4><B8><BA>2\n>>\n>> Actual result of \"git diff --color-words\"\n>>\n>> <E4><B8>[-<BA>1-]{+<BA>2+}\n>> ...\n> I think we could provide new ways to do per-language diffs, right now\n> you can use --word-diff-regex, but it would be handy to e.g. have a\n> built-in collection of those (or other non-regex boundary algorithms)\n> for Chinese etc.\n\nI think you are thinking it with unnecessaarily complexity.  \n\nThe only thing that needs noticing in the above example, I think is,\nthat the three-byte sequence E4-B8-BA in the example is supposed to\nbe a single unicode character, and the actual result depicted can\nhappen only if we (incorrectly) chomp that single character in the\nmiddle.\n\nNo matter what language we are using, we shouldn't do that.\n\nI suspect that \"--word-diff\" internal is not even aware what a\ncharacter is, but if you assume UTF-8 (precomposed), then you should\nbe able to tell where the character boundary is by only looking at\nthe high-bit patterns to avoid producing such an output.\n"},{"id":"468199","messageId":"Y4ZOHwwgtztwhbhr@coredump.intra.peff.net","threadId":"58868","inReplyTo":"xmqqlenu2dxx.fsf@gitster.g","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2022-11-29T18:23:27Z","receivedAt":"2022-11-29T18:23:38Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Nov 29, 2022 at 08:32:58PM +0900, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n> \n> >> or (if chinese can not be displayed correctly)\n> >>\n> >> -  <E4><B8><BA>1\n> >> +  <E4><B8><BA>2\n> >>\n> >> Actual result of \"git diff --color-words\"\n> >>\n> >> <E4><B8>[-<BA>1-]{+<BA>2+}\n> >> ...\n> > I think we could provide new ways to do per-language diffs, right now\n> > you can use --word-diff-regex, but it would be handy to e.g. have a\n> > built-in collection of those (or other non-regex boundary algorithms)\n> > for Chinese etc.\n> \n> I think you are thinking it with unnecessaarily complexity.  \n> \n> The only thing that needs noticing in the above example, I think is,\n> that the three-byte sequence E4-B8-BA in the example is supposed to\n> be a single unicode character, and the actual result depicted can\n> happen only if we (incorrectly) chomp that single character in the\n> middle.\n> \n> No matter what language we are using, we shouldn't do that.\n> \n> I suspect that \"--word-diff\" internal is not even aware what a\n> character is, but if you assume UTF-8 (precomposed), then you should\n> be able to tell where the character boundary is by only looking at\n> the high-bit patterns to avoid producing such an output.\n\nAgreed that we should probably avoid breaking characters. But what\npuzzles me more is that we break it between B8 and BA, and not\nelsewhere. Why not between E4 and B8? Why not between BA and \"1\"?\n\nIf the rule is \"break on ascii whitespace\", then I'd have expected the\nwhole four-character sequence to be taken as a unit. In other words, it\ndoes should not have to care that a character is, as long as the bytes\nfor space characters cannot appear inside other characters (which is\ntrue of utf8).\n\n-Peff\n"},{"id":"468201","messageId":"Y4ZVXWNHO25IFYQL@coredump.intra.peff.net","threadId":"58868","inReplyTo":"Y4ZOHwwgtztwhbhr@coredump.intra.peff.net","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2022-11-29T18:54:21Z","receivedAt":"2022-11-29T18:54:27Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Nov 29, 2022 at 01:23:27PM -0500, Jeff King wrote:\n\n> > I suspect that \"--word-diff\" internal is not even aware what a\n> > character is, but if you assume UTF-8 (precomposed), then you should\n> > be able to tell where the character boundary is by only looking at\n> > the high-bit patterns to avoid producing such an output.\n> \n> Agreed that we should probably avoid breaking characters. But what\n> puzzles me more is that we break it between B8 and BA, and not\n> elsewhere. Why not between E4 and B8? Why not between BA and \"1\"?\n> \n> If the rule is \"break on ascii whitespace\", then I'd have expected the\n> whole four-character sequence to be taken as a unit. In other words, it\n> does should not have to care that a character is, as long as the bytes\n> for space characters cannot appear inside other characters (which is\n> true of utf8).\n\nEven more puzzling is that it produces the expected output for me:\n\n  [note that \\x is a bash-ism]\n  $ printf '\\xe4\\xb8\\xba1' >one\n  $ printf '\\xe4\\xb8\\xba2' >two\n  $ git diff --no-index --word-diff one two\n  diff --git a/one b/two\n  index 9ae469fc41..576e6e32d8 100644\n  --- a/one\n  +++ b/two\n  @@ -1 +1 @@\n  [-为1-]{+为2+}\n\nI wonder if OP has diff.wordRegex config (or attributes triggering a\ndiff.*.wordRegex) that is doing something else.\n\n-Peff\n"},{"id":"468304","messageId":"CACSwcnTj8kiM83+x3V69avohqW4ABOBVe4uG5n3giY-tQEP_Vg@mail.gmail.com","threadId":"58868","inReplyTo":"Y4ZVXWNHO25IFYQL@coredump.intra.peff.net","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Ping Yin","fromEmail":"pkufranky@gmail.com","sentAt":"2022-12-01T07:08:48Z","receivedAt":"2022-12-01T07:09:11Z","isPatch":false,"sender":{"key":"pkufranky@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5346?v=4"},"body":"> I wonder if OP has diff.wordRegex config (or attributes triggering a\n> diff.*.wordRegex) that is doing something else.\n\nWow, you are right, sorry for the noise.\n\n$ git config -l | grep word\ndiff.wordregex=[[:alnum:]_]+|[^[:space:]]\n"},{"id":"468305","messageId":"CACSwcnRDmiiJU8hL+ON6c+b4Q8UtLVbtku_rHSD+c+BwcNEX+Q@mail.gmail.com","threadId":"58868","inReplyTo":"Y4ZVXWNHO25IFYQL@coredump.intra.peff.net","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Ping Yin","fromEmail":"pkufranky@gmail.com","sentAt":"2022-12-01T07:33:06Z","receivedAt":"2022-12-01T07:33:22Z","isPatch":false,"sender":{"key":"pkufranky@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5346?v=4"},"body":"> > If the rule is \"break on ascii whitespace\",\n\nIs there a way to achieve this: break english by word, and break\nchinese by utf-8 character\n"},{"id":"468323","messageId":"4dac768f-1104-a565-db4c-8a1b7eb2870d@dunelm.org.uk","threadId":"58868","inReplyTo":"CACSwcnRDmiiJU8hL+ON6c+b4Q8UtLVbtku_rHSD+c+BwcNEX+Q@mail.gmail.com","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2022-12-01T14:51:29Z","receivedAt":"2022-12-01T14:51:37Z","isPatch":false,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Ping\n\nOn 01/12/2022 07:33, Ping Yin wrote:\n>>> If the rule is \"break on ascii whitespace\",\n> \n> Is there a way to achieve this: break english by word, and break\n> chinese by utf-8 character\n\nYou could extend your current regex so that it matches whole utf-8 \ncodepoints which is what git does for the builtin userdiff regexes. I've \nnot tested it but I think\n\ngit config --global diff.wordregex \"[[:alnum:]_]+|[^[:space:]]|$(printf \n'[\\xc0-\\xff][\\x80-\\xbf]+')\"\n\nshould work. The downside is that you end up with a .gitconfig that is \nnot valid utf-8. Perhaps someone else has a clever idea to get around that.\n\nBest Wishes\n\nPhillip\n"},{"id":"468326","messageId":"CACSwcnTBKHO249HMhps_M629hJiAftQr0BbS50DczMThX9DDHw@mail.gmail.com","threadId":"58868","inReplyTo":"4dac768f-1104-a565-db4c-8a1b7eb2870d@dunelm.org.uk","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Ping Yin","fromEmail":"pkufranky@gmail.com","sentAt":"2022-12-01T15:51:40Z","receivedAt":"2022-12-01T15:51:57Z","isPatch":false,"sender":{"key":"pkufranky@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5346?v=4"},"body":"Ping Yin\n\nOn Thu, Dec 1, 2022 at 10:51 PM Phillip Wood <phillip.wood123@gmail.com> wrote:\n>\n> Hi Ping\n>\n> On 01/12/2022 07:33, Ping Yin wrote:\n>\n> git config --global diff.wordregex \"[[:alnum:]_]+|[^[:space:]]|$(printf\n> '[\\xc0-\\xff][\\x80-\\xbf]+')\"\n>\n> should work. The downside is that you end up with a .gitconfig that is\n> not valid utf-8. Perhaps someone else has a clever idea to get around that.\n\nWow, it works. Thanks very much.\n"},{"id":"468350","messageId":"Y4kJWJObB5Er2CXZ@coredump.intra.peff.net","threadId":"58868","inReplyTo":"4dac768f-1104-a565-db4c-8a1b7eb2870d@dunelm.org.uk","subject":"Re: [bug] git diff --word-diff gives wrong result for utf-8 chinese","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2022-12-01T20:06:48Z","receivedAt":"2022-12-01T20:07:07Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 01, 2022 at 02:51:29PM +0000, Phillip Wood wrote:\n\n> On 01/12/2022 07:33, Ping Yin wrote:\n> > > > If the rule is \"break on ascii whitespace\",\n> > \n> > Is there a way to achieve this: break english by word, and break\n> > chinese by utf-8 character\n> \n> You could extend your current regex so that it matches whole utf-8\n> codepoints which is what git does for the builtin userdiff regexes. I've not\n> tested it but I think\n> \n> git config --global diff.wordregex \"[[:alnum:]_]+|[^[:space:]]|$(printf\n> '[\\xc0-\\xff][\\x80-\\xbf]+')\"\n> \n> should work. The downside is that you end up with a .gitconfig that is not\n> valid utf-8. Perhaps someone else has a clever idea to get around that.\n\nI think in more advanced regular expression engines you can do stuff\nlike matching \"[\\x{4e00}-\\x{9fcc}]\", or even \"\\p{Han}\". But I don't know\nthat the stock libc regex is capable of anything like this, even with\nEREs. That's the only option Git provides for matching word regexes, but\nin theory we could support libpcre. We already can optionally build\nagainst it; we would just need config/plumbing to get it into\ndiff.c:find_word_boundaries().\n\n-Peff\n"}]}