{"thread":{"id":"64806","subject":"[GSoC PATCH 1/1] diff: improve scaling of filenames in diffstat to handle UTF-8 chars","startedAt":"2026-01-14T22:27:29Z","lastAt":"2026-01-17T17:52:38Z","messageCount":7,"participants":["LorenzoPegorari","Junio C Hamano","Lorenzo Pegorari"],"isPatch":true,"patchVersion":1,"patchTotal":1},"messages":[{"id":"533908","messageId":"aWgYRkv-YsuekdR_@lorenzo-VM","threadId":"64806","inReplyTo":null,"subject":"[GSoC PATCH 1/1] diff: improve scaling of filenames in diffstat to handle UTF-8 chars","fromName":"LorenzoPegorari","fromEmail":"lorenzo.pegorari2002@gmail.com","sentAt":"2026-01-14T22:27:18Z","receivedAt":"2026-01-14T22:27:29Z","isPatch":true,"sender":{"key":"lorenzo.pegorari2002@gmail.com","avatar":"https://avatars.githubusercontent.com/u/132087553?v=4"},"body":"The `show_stats()` function tries to scale the filenames in the diffstat to\nensure they don't exceed the given `name-width`. It does so by calculating\nthe \"display width\" of the characters to be dropped, but then advances the\nfilename pointer by that number of bytes.\n\nHowever, the \"display width\" of a character is not always equal to its byte\ncount. The result is that sometimes, when displaying UTF-8 characters,\nfilenames exceed the given `name-width`, and frequently the bytes of the\nUTF-8 characters are truncated.\n\nThe following is an example of the issue, where the 2 files are \"HelloHi\" and\n\"Hello你好\", and `name-width=6`:\n\n    ...oHi | 0\n    ...<BD><A0>好 | 0\n\nMake the filename pointer move by the actual number of bytes of the\ncharacters to drop from the filename, rather than their display width, using\nthe `utf8_width()` function.\n\nSigned-off-by: LorenzoPegorari <lorenzo.pegorari2002@gmail.com>\n---\n diff.c | 15 ++++-----------\n 1 file changed, 4 insertions(+), 11 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex a68ddd2168..271ace5728 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -2859,17 +2859,10 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)\n \t\t\tchar *slash;\n \t\t\tprefix = \"...\";\n \t\t\tlen -= 3;\n-\t\t\t/*\n-\t\t\t * NEEDSWORK: (name_len - len) counts the display\n-\t\t\t * width, which would be shorter than the byte\n-\t\t\t * length of the corresponding substring.\n-\t\t\t * Advancing \"name\" by that number of bytes does\n-\t\t\t * *NOT* skip over that many columns, so it is\n-\t\t\t * very likely that chomping the pathname at the\n-\t\t\t * slash we will find starting from \"name\" will\n-\t\t\t * leave the resulting string still too long.\n-\t\t\t */\n-\t\t\tname += name_len - len;\n+\n+\t\t\twhile (name_len > len)\n+\t\t\t\tname_len -= utf8_width((const char**)&name, NULL);\n+\n \t\t\tslash = strchr(name, '/');\n \t\t\tif (slash)\n \t\t\t\tname = slash;\n-- \n2.43.0\n\n"},{"id":"533911","messageId":"xmqqikd3ermt.fsf@gitster.g","threadId":"64806","inReplyTo":"aWgYRkv-YsuekdR_@lorenzo-VM","subject":"Re: [GSoC PATCH 1/1] diff: improve scaling of filenames in diffstat to handle UTF-8 chars","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-01-14T22:50:02Z","receivedAt":"2026-01-14T22:50:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"LorenzoPegorari <lorenzo.pegorari2002@gmail.com> writes:\n\n> The `show_stats()` function tries to scale the filenames in the diffstat to\n> ensure they don't exceed the given `name-width`. It does so by calculating\n> the \"display width\" of the characters to be dropped, but then advances the\n> filename pointer by that number of bytes.\n>\n> However, the \"display width\" of a character is not always equal to its byte\n> count. The result is that sometimes, when displaying UTF-8 characters,\n> filenames exceed the given `name-width`, and frequently the bytes of the\n> UTF-8 characters are truncated.\n>\n> The following is an example of the issue, where the 2 files are \"HelloHi\" and\n> \"Hello你好\", and `name-width=6`:\n>\n>     ...oHi | 0\n>     ...<BD><A0>好 | 0\n>\n> Make the filename pointer move by the actual number of bytes of the\n> characters to drop from the filename, rather than their display width, using\n> the `utf8_width()` function.\n>\n> Signed-off-by: LorenzoPegorari <lorenzo.pegorari2002@gmail.com>\n> ---\n>  diff.c | 15 ++++-----------\n>  1 file changed, 4 insertions(+), 11 deletions(-)\n\nTwo comments and a half.\n\n * The change needed for this is surprisingly simple.\n\n * You already know about samples that may exhibit the issue you are\n   addressing.  Can we add it as a test case somewhere in t/\n   directory?\n\n * The NEEDSWORK item addressed by this patch is one of the two\n   NEEDSWORK items added by ce8529b2 (diff: leave NEEDWORK notes in\n   show_stats() function, 2022-10-21).  Makes me wonder how involved\n   the changes would need to be to solve the other one?\n\nThanks.\n\n\n> diff --git a/diff.c b/diff.c\n> index a68ddd2168..271ace5728 100644\n> --- a/diff.c\n> +++ b/diff.c\n> @@ -2859,17 +2859,10 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)\n>  \t\t\tchar *slash;\n>  \t\t\tprefix = \"...\";\n>  \t\t\tlen -= 3;\n> -\t\t\t/*\n> -\t\t\t * NEEDSWORK: (name_len - len) counts the display\n> -\t\t\t * width, which would be shorter than the byte\n> -\t\t\t * length of the corresponding substring.\n> -\t\t\t * Advancing \"name\" by that number of bytes does\n> -\t\t\t * *NOT* skip over that many columns, so it is\n> -\t\t\t * very likely that chomping the pathname at the\n> -\t\t\t * slash we will find starting from \"name\" will\n> -\t\t\t * leave the resulting string still too long.\n> -\t\t\t */\n> -\t\t\tname += name_len - len;\n> +\n> +\t\t\twhile (name_len > len)\n> +\t\t\t\tname_len -= utf8_width((const char**)&name, NULL);\n> +\n>  \t\t\tslash = strchr(name, '/');\n>  \t\t\tif (slash)\n>  \t\t\t\tname = slash;\n"},{"id":"534004","messageId":"aWl_oeJJxtwsUyR3@lorenzo-VM","threadId":"64806","inReplyTo":"xmqqikd3ermt.fsf@gitster.g","subject":"Re: [GSoC PATCH 1/1] diff: improve scaling of filenames in diffstat to handle UTF-8 chars","fromName":"Lorenzo Pegorari","fromEmail":"lorenzo.pegorari2002@gmail.com","sentAt":"2026-01-16T00:00:33Z","receivedAt":"2026-01-16T00:00:37Z","isPatch":true,"sender":{"key":"lorenzo.pegorari2002@gmail.com","avatar":"https://avatars.githubusercontent.com/u/132087553?v=4"},"body":"On Wed, Jan 14, 2026 at 02:50:02PM -0800, Junio C Hamano wrote:\n> LorenzoPegorari <lorenzo.pegorari2002@gmail.com> writes:\n> \n> > The `show_stats()` function tries to scale the filenames in the diffstat to\n> > ensure they don't exceed the given `name-width`. It does so by calculating\n> > the \"display width\" of the characters to be dropped, but then advances the\n> > filename pointer by that number of bytes.\n> >\n> > However, the \"display width\" of a character is not always equal to its byte\n> > count. The result is that sometimes, when displaying UTF-8 characters,\n> > filenames exceed the given `name-width`, and frequently the bytes of the\n> > UTF-8 characters are truncated.\n> >\n> > The following is an example of the issue, where the 2 files are \"HelloHi\" and\n> > \"Hello你好\", and `name-width=6`:\n> >\n> >     ...oHi | 0\n> >     ...<BD><A0>好 | 0\n> >\n> > Make the filename pointer move by the actual number of bytes of the\n> > characters to drop from the filename, rather than their display width, using\n> > the `utf8_width()` function.\n> >\n> > Signed-off-by: LorenzoPegorari <lorenzo.pegorari2002@gmail.com>\n> > ---\n> >  diff.c | 15 ++++-----------\n> >  1 file changed, 4 insertions(+), 11 deletions(-)\n> \n> Two comments and a half.\n> \n>  * The change needed for this is surprisingly simple.\n\nIt is indeed surprisingly simple, I agree!\n\n>  * You already know about samples that may exhibit the issue you are\n>    addressing.  Can we add it as a test case somewhere in t/\n>    directory?\n\nYeah, we should add a test case. I will do it in the next reroll.\n\n>  * The NEEDSWORK item addressed by this patch is one of the two\n>    NEEDSWORK items added by ce8529b2 (diff: leave NEEDWORK notes in\n>    show_stats() function, 2022-10-21).  Makes me wonder how involved\n>    the changes would need to be to solve the other one?\n\nMmh, I see. I'll take a closer look, but at a first glance it doesn't\nseem too involved.\n\n>\n> Thanks.\n>\n\nThank you!\n\n> \n> > diff --git a/diff.c b/diff.c\n> > index a68ddd2168..271ace5728 100644\n> > --- a/diff.c\n> > +++ b/diff.c\n> > @@ -2859,17 +2859,10 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)\n> >  \t\t\tchar *slash;\n> >  \t\t\tprefix = \"...\";\n> >  \t\t\tlen -= 3;\n> > -\t\t\t/*\n> > -\t\t\t * NEEDSWORK: (name_len - len) counts the display\n> > -\t\t\t * width, which would be shorter than the byte\n> > -\t\t\t * length of the corresponding substring.\n> > -\t\t\t * Advancing \"name\" by that number of bytes does\n> > -\t\t\t * *NOT* skip over that many columns, so it is\n> > -\t\t\t * very likely that chomping the pathname at the\n> > -\t\t\t * slash we will find starting from \"name\" will\n> > -\t\t\t * leave the resulting string still too long.\n> > -\t\t\t */\n> > -\t\t\tname += name_len - len;\n> > +\n> > +\t\t\twhile (name_len > len)\n> > +\t\t\t\tname_len -= utf8_width((const char**)&name, NULL);\n> > +\n> >  \t\t\tslash = strchr(name, '/');\n> >  \t\t\tif (slash)\n> >  \t\t\t\tname = slash;\n"},{"id":"534005","messageId":"cover.1768520441.git.lorenzo.pegorari2002@gmail.com","threadId":"64806","inReplyTo":"aWgYRkv-YsuekdR_@lorenzo-VM","subject":"[GSoC PATCH v2 0/2] diff: improve scaling of filenames in diffstat to handle UTF-8 chars","fromName":"LorenzoPegorari","fromEmail":"lorenzo.pegorari2002@gmail.com","sentAt":"2026-01-16T00:04:22Z","receivedAt":"2026-01-16T00:04:26Z","isPatch":true,"sender":{"key":"lorenzo.pegorari2002@gmail.com","avatar":"https://avatars.githubusercontent.com/u/132087553?v=4"},"body":"Added a test (as Junio Hamano suggested) to check how the generated\ndiffstat handles UTF-8 characters when given various `name-width`s.\n\nThis allowed me to notice a bug where, if the given `name-width` was 2\nor less, the `len` variable would become negative, entering an infinite\nloop. So I also fixed this bug.\n\nLorenzoPegorari (2):\n  diff: improve scaling of filenames in diffstat to handle UTF-8 chars\n  t4073: add test for diffstat paths length when containing UTF-8 chars\n\n diff.c                          | 17 ++++-----\n t/meson.build                   |  1 +\n t/t4073-diff-stat-name-width.sh | 61 +++++++++++++++++++++++++++++++++\n 3 files changed, 68 insertions(+), 11 deletions(-)\n create mode 100755 t/t4073-diff-stat-name-width.sh\n\nRange-diff against v1:\n1:  63e73122d1 ! 1:  abeb8d3439 diff: improve scaling of filenames in diffstat to handle UTF-8 chars\n    @@ Commit message\n         characters to drop from the filename, rather than their display width, using\n         the `utf8_width()` function.\n     \n    +    Force `len` to not be less than 0 (this happens if the given `name-width` is\n    +    2 or less), otherwise an infinite loop is entered.\n    +\n         Signed-off-by: LorenzoPegorari <lorenzo.pegorari2002@gmail.com>\n     \n      ## diff.c ##\n    @@ diff.c: static void show_stats(struct diffstat_t *data, struct diff_options *opt\n     -\t\t\t * leave the resulting string still too long.\n     -\t\t\t */\n     -\t\t\tname += name_len - len;\n    ++\t\t\tif (len < 0)\n    ++\t\t\t\tlen = 0;\n     +\n     +\t\t\twhile (name_len > len)\n     +\t\t\t\tname_len -= utf8_width((const char**)&name, NULL);\n-:  ---------- > 2:  ee088ea6ef t4073: add test for diffstat paths length when containing UTF-8 chars\n-- \n2.43.0\n\n"},{"id":"534006","messageId":"abeb8d3439de6569fd73617de580fa510e19466b.1768520441.git.lorenzo.pegorari2002@gmail.com","threadId":"64806","inReplyTo":"cover.1768520441.git.lorenzo.pegorari2002@gmail.com","subject":"[GSoC PATCH v2 1/2] diff: improve scaling of filenames in diffstat to handle UTF-8 chars","fromName":"LorenzoPegorari","fromEmail":"lorenzo.pegorari2002@gmail.com","sentAt":"2026-01-16T00:05:03Z","receivedAt":"2026-01-16T00:05:07Z","isPatch":true,"sender":{"key":"lorenzo.pegorari2002@gmail.com","avatar":"https://avatars.githubusercontent.com/u/132087553?v=4"},"body":"The `show_stats()` function tries to scale the filenames in the diffstat to\nensure they don't exceed the given `name-width`. It does so by calculating\nthe \"display width\" of the characters to be dropped, but then advances the\nfilename pointer by that number of bytes.\n\nHowever, the \"display width\" of a character is not always equal to its byte\ncount. The result is that sometimes, when displaying UTF-8 characters,\nfilenames exceed the given `name-width`, and frequently the bytes of the\nUTF-8 characters are truncated.\n\nThe following is an example of the issue, where the 2 files are \"HelloHi\" and\n\"Hello你好\", and `name-width=6`:\n\n    ...oHi | 0\n    ...<BD><A0>好 | 0\n\nMake the filename pointer move by the actual number of bytes of the\ncharacters to drop from the filename, rather than their display width, using\nthe `utf8_width()` function.\n\nForce `len` to not be less than 0 (this happens if the given `name-width` is\n2 or less), otherwise an infinite loop is entered.\n\nSigned-off-by: LorenzoPegorari <lorenzo.pegorari2002@gmail.com>\n---\n diff.c | 17 ++++++-----------\n 1 file changed, 6 insertions(+), 11 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex a68ddd2168..452fc69775 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -2859,17 +2859,12 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)\n \t\t\tchar *slash;\n \t\t\tprefix = \"...\";\n \t\t\tlen -= 3;\n-\t\t\t/*\n-\t\t\t * NEEDSWORK: (name_len - len) counts the display\n-\t\t\t * width, which would be shorter than the byte\n-\t\t\t * length of the corresponding substring.\n-\t\t\t * Advancing \"name\" by that number of bytes does\n-\t\t\t * *NOT* skip over that many columns, so it is\n-\t\t\t * very likely that chomping the pathname at the\n-\t\t\t * slash we will find starting from \"name\" will\n-\t\t\t * leave the resulting string still too long.\n-\t\t\t */\n-\t\t\tname += name_len - len;\n+\t\t\tif (len < 0)\n+\t\t\t\tlen = 0;\n+\n+\t\t\twhile (name_len > len)\n+\t\t\t\tname_len -= utf8_width((const char**)&name, NULL);\n+\n \t\t\tslash = strchr(name, '/');\n \t\t\tif (slash)\n \t\t\t\tname = slash;\n-- \n2.43.0\n\n"},{"id":"534007","messageId":"ee088ea6ef91f0c349ed4940feab807d421dde66.1768520441.git.lorenzo.pegorari2002@gmail.com","threadId":"64806","inReplyTo":"cover.1768520441.git.lorenzo.pegorari2002@gmail.com","subject":"[GSoC PATCH v2 2/2] t4073: add test for diffstat paths length when containing UTF-8 chars","fromName":"LorenzoPegorari","fromEmail":"lorenzo.pegorari2002@gmail.com","sentAt":"2026-01-16T00:05:38Z","receivedAt":"2026-01-16T00:05:43Z","isPatch":true,"sender":{"key":"lorenzo.pegorari2002@gmail.com","avatar":"https://avatars.githubusercontent.com/u/132087553?v=4"},"body":"Add test checking the length of filepaths containing UTF-8 chars when\ngenerating a diffstat with various `name-width`s.\n\nSigned-off-by: LorenzoPegorari <lorenzo.pegorari2002@gmail.com>\n---\n t/meson.build                   |  1 +\n t/t4073-diff-stat-name-width.sh | 61 +++++++++++++++++++++++++++++++++\n 2 files changed, 62 insertions(+)\n create mode 100755 t/t4073-diff-stat-name-width.sh\n\ndiff --git a/t/meson.build b/t/meson.build\nindex 459c52a489..f2ad6d2f12 100644\n--- a/t/meson.build\n+++ b/t/meson.build\n@@ -498,6 +498,7 @@ integration_tests = [\n   't4070-diff-pairs.sh',\n   't4071-diff-minimal.sh',\n   't4072-diff-max-depth.sh',\n+  't4073-diff-stat.sh',\n   't4100-apply-stat.sh',\n   't4101-apply-nonl.sh',\n   't4102-apply-rename.sh',\ndiff --git a/t/t4073-diff-stat-name-width.sh b/t/t4073-diff-stat-name-width.sh\nnew file mode 100755\nindex 0000000000..ec5d3c3c1f\n--- /dev/null\n+++ b/t/t4073-diff-stat-name-width.sh\n@@ -0,0 +1,61 @@\n+#!/bin/sh\n+\n+test_description='git-diff check diffstat filepaths length when containing UTF-8 chars'\n+\n+. ./test-lib.sh\n+\n+\n+create_files () {\n+\tmkdir -p \"d你好\" &&\n+\ttouch \"d你好/f再见\"\n+}\n+\n+test_expect_success 'setup' '\n+\tgit init &&\n+\tgit config core.quotepath off &&\n+\tgit commit -m \"Initial commit\" --allow-empty &&\n+\tcreate_files &&\n+\tgit add . &&\n+\tgit commit -m \"Added files\"\n+'\n+\n+test_expect_success 'test name-width long enough for filepath' '\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=12 >out &&\n+\tgrep \"d你好/f再见 |\" out &&\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=11 >out &&\n+\tgrep \"d你好/f再见 |\" out\n+'\n+\n+test_expect_success 'test name-width not long enough for dir name' '\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=10 >out &&\n+\tgrep \".../f再见  |\" out &&\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=9 >out &&\n+\tgrep \".../f再见 |\" out\n+'\n+\n+test_expect_success 'test name-width not long enough for slash' '\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=8 >out &&\n+\tgrep \"...f再见 |\" out\n+'\n+\n+test_expect_success 'test name-width not long enough for file name' '\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=7 >out &&\n+\tgrep \"...再见 |\" out &&\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=6 >out &&\n+\tgrep \"...见  |\" out &&\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=5 >out &&\n+\tgrep \"...见 |\" out &&\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=4 >out &&\n+\tgrep \"...  |\" out\n+'\n+\n+test_expect_success 'test name-width minimum length' '\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=3 >out &&\n+\tgrep \"... |\" out &&\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=2 >out &&\n+\tgrep \"... |\" out &&\n+\tgit diff HEAD~1 HEAD --stat --stat-name-width=1 >out &&\n+\tgrep \"... |\" out\n+'\n+\n+test_done\n-- \n2.43.0\n\n"},{"id":"534121","messageId":"xmqqecno6s9q.fsf@gitster.g","threadId":"64806","inReplyTo":"ee088ea6ef91f0c349ed4940feab807d421dde66.1768520441.git.lorenzo.pegorari2002@gmail.com","subject":"Re: [GSoC PATCH v2 2/2] t4073: add test for diffstat paths length when containing UTF-8 chars","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-01-17T17:52:33Z","receivedAt":"2026-01-17T17:52:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"LorenzoPegorari <lorenzo.pegorari2002@gmail.com> writes:\n\n> diff --git a/t/meson.build b/t/meson.build\n> index 459c52a489..f2ad6d2f12 100644\n> --- a/t/meson.build\n> +++ b/t/meson.build\n> @@ -498,6 +498,7 @@ integration_tests = [\n>    't4070-diff-pairs.sh',\n>    't4071-diff-minimal.sh',\n>    't4072-diff-max-depth.sh',\n> +  't4073-diff-stat.sh',\n\nThis name ...\n\n>    't4100-apply-stat.sh',\n>    't4101-apply-nonl.sh',\n>    't4102-apply-rename.sh',\n> diff --git a/t/t4073-diff-stat-name-width.sh b/t/t4073-diff-stat-name-width.sh\n> new file mode 100755\n\n... must match this one.  I already locally updated the former to\nmatch, so no need to resend, but if you need to reroll the patch in\nthe future, please make sure to correct this part.  Thanks.\n"}]}