{"thread":{"id":"47642","subject":"UTF-8-safe way for char-level-diff","startedAt":"2018-01-19T14:13:30Z","lastAt":"2018-01-19T14:13:30Z","messageCount":1,"participants":["Danny Lin"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"336889","messageId":"CAMbsUu7P6RGpyDm+tKbBhgPyxoTYP3Zs5rjt302O8SH-H8ujeQ@mail.gmail.com","threadId":"47642","inReplyTo":null,"subject":"UTF-8-safe way for char-level-diff","fromName":"Danny Lin","fromEmail":"danny0838@gmail.com","sentAt":"2018-01-19T14:13:20Z","receivedAt":"2018-01-19T14:13:30Z","isPatch":false,"sender":{"key":"danny0838@gmail.com","avatar":"https://avatars.githubusercontent.com/u/531417?v=4"},"body":"Git has a diff.wordRegex config that allows the user to specify a\nregex that defines a word. Setting diff.wordRegex to \".\" works well\nfor a char-level diff for ASCII chars, but not for UTF-8 chars.\n\nFor example, if a file (encoded by UTF-8) with text \"一人\" is changed to\n\"丁人\", \"git diff --word-diff=color\" gets \"<E4><B8><80><81>人\" (where\n\"<80>\" is red and \"<81>\" is green) instead of desired \"一丁人\" (where \"一\"\nis red and \"丁\" is green). This could be very annoying when diff-ing\nfiles containing CJK chars.\n\nGit diff.wordRegex seems to implement a very basic regex that doesn't\nsupport matching char range by encoding such as \"\\x41\" for \"a\". Is\nthere a way to make the char-level diff work correctly? If not, maybe\nwe should implement a way to allow it.\n"}]}