{"thread":{"id":"65507","subject":"[PATCH] diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation","startedAt":"2026-04-17T16:26:07Z","lastAt":"2026-04-20T16:42:00Z","messageCount":9,"participants":["Elijah Newren via GitGitGadget","Junio C Hamano","Elijah Newren","Lorenzo Pegorari"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"541827","messageId":"pull.2093.git.1776443163041.gitgitgadget@gmail.com","threadId":"65507","inReplyTo":null,"subject":"[PATCH] diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation","fromName":"Elijah Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-17T16:26:03Z","receivedAt":"2026-04-17T16:26:07Z","isPatch":true,"body":"From: Elijah Newren <newren@gmail.com>\n\nf85b49f3d4a (diff: improve scaling of filenames in diffstat to handle\nUTF-8 chars, 2024-10-27) introduced a loop in show_stats() that calls\nutf8_width() repeatedly to skip leading characters until the displayed\nwidth fits.  However, utf8_width() can return problematic values:\n\n  - For invalid UTF-8 sequences, pick_one_utf8_char() sets the name\n    pointer to NULL and utf8_width() returns 0.  Since name_len does\n    not change, the loop iterates once more and pick_one_utf8_char()\n    dereferences the NULL pointer, crashing.\n\n  - For control characters, utf8_width() returns -1, so name_len\n    grows when it is expected to shrink.  This can cause the loop to\n    consume more characters than the string contains, reading past\n    the trailing NUL.\n\nBy default, fill_print_name() will C-quotes filenames which escapes\ncontrol characters and invalid bytes to printable text.  That avoids\nthis bug from being triggered; however, with core.quotePath=false,\nraw bytes can reach this code.\n\nAdd tests exercising both failure modes with core.quotePath=false and\na narrow --stat-name-width to force truncation: one with a bare 0xC0\nbyte (invalid UTF-8 lead byte, triggers NULL deref) and one with a\n0x01 byte (control character, causes the loop to read past the end\nof the string).\n\nFix the bug by:\n  - Adding a *name check to terminate the loop at end-of-string\n  - Detecting the NULL pointer from invalid UTF-8 and falling back to\n    showing the full untruncated name\n  - Breaking on negative width (control characters)\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n    diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8\n    truncation\n    \n    Maintainer note: This is a new bug from the v2.54 cycle\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2093%2Fnewren%2Ffix%2Fdiffstat-utf8-loop-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2093/newren/fix/diffstat-utf8-loop-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/2093\n\n diff.c                 | 13 +++++++++++--\n t/t4052-stat-output.sh | 25 +++++++++++++++++++++++++\n 2 files changed, 36 insertions(+), 2 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex 397e38b41c..7b27241733 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -3093,8 +3093,17 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)\n \t\t\tif (len < 0)\n \t\t\t\tlen = 0;\n \n-\t\t\twhile (name_len > len)\n-\t\t\t\tname_len -= utf8_width((const char**)&name, NULL);\n+\t\t\twhile (name_len > len && *name) {\n+\t\t\t\tint w = utf8_width((const char **)&name, NULL);\n+\t\t\t\tif (!name) { /* Invalid UTF-8 */\n+\t\t\t\t\tname = file->print_name;\n+\t\t\t\t\tname_len = utf8_strwidth(name);\n+\t\t\t\t\tbreak;\n+\t\t\t\t}\n+\t\t\t\tif (w < 0)  /* control character */\n+\t\t\t\t\tbreak;\n+\t\t\t\tname_len -= w;\n+\t\t\t}\n \n \t\t\tslash = strchr(name, '/');\n \t\t\tif (slash)\ndiff --git a/t/t4052-stat-output.sh b/t/t4052-stat-output.sh\nindex 7c749062e2..84c53c1a51 100755\n--- a/t/t4052-stat-output.sh\n+++ b/t/t4052-stat-output.sh\n@@ -445,4 +445,29 @@ test_expect_success 'diffstat where line_prefix contains ANSI escape codes is co\n \ttest_grep \"<RED>|<RESET>  ${FILENAME_TRIMMED} | 0\" out\n '\n \n+test_expect_success 'diffstat truncation with invalid UTF-8 does not crash' '\n+\tempty_blob=$(git hash-object -w --stdin </dev/null) &&\n+\tprintf \"100644 blob $empty_blob\\taaa-\\300-aaa\\n\" |\n+\tgit mktree >tree_file &&\n+\ttree=$(cat tree_file) &&\n+\tempty_tree=$(git mktree </dev/null) &&\n+\tc1=$(git commit-tree -m before $empty_tree) &&\n+\tc2=$(git commit-tree -m after -p $c1 $tree) &&\n+\tgit -c core.quotepath=false diff --stat --stat-name-width=5 $c1..$c2 >output &&\n+\ttest_grep \"| 0\" output\n+'\n+\n+test_expect_success FUNNYNAMES 'diffstat truncation with control chars does not crash' '\n+\tFNAME=$(printf \"aaa-\\x01-aaa\") &&\n+\tgit commit --allow-empty -m setup &&\n+\t>$FNAME &&\n+\tgit add -- $FNAME &&\n+\tgit commit -m \"add file with control char name\" &&\n+\tgit -c core.quotepath=false diff --stat --stat-name-width=5 HEAD~1..HEAD >output &&\n+\ttest_grep \"| 0\" output &&\n+\trm -- $FNAME &&\n+\tgit rm -- $FNAME &&\n+\tgit commit -m \"remove test file\"\n+'\n+\n test_done\n\nbase-commit: 9f223ef1c026d91c7ac68cc0211bde255dda6199\n-- \ngitgitgadget\n"},{"id":"541836","messageId":"xmqqv7dpwfy5.fsf@gitster.g","threadId":"65507","inReplyTo":"pull.2093.git.1776443163041.gitgitgadget@gmail.com","subject":"Re: [PATCH] diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-17T19:21:38Z","receivedAt":"2026-04-17T19:21:41Z","isPatch":true,"body":"\"Elijah Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Elijah Newren <newren@gmail.com>\n>\n> f85b49f3d4a (diff: improve scaling of filenames in diffstat to handle\n> UTF-8 chars, 2024-10-27) introduced a loop in show_stats() that calls\n> utf8_width() repeatedly to skip leading characters until the displayed\n> width fits.\n\nA tangent, but I get a datestamp for the same f85b49f3 (diff:\nimprove scaling of filenames in diffstat to handle UTF-8 chars,\n2026-01-16) that is different from what you showed above.  Did you\nfind a bug in \"git show -s --pretty=reference\"?\n\n> diff --git a/diff.c b/diff.c\n> index 397e38b41c..7b27241733 100644\n> --- a/diff.c\n> +++ b/diff.c\n> @@ -3093,8 +3093,17 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)\n>  \t\t\tif (len < 0)\n>  \t\t\t\tlen = 0;\n>  \n> -\t\t\twhile (name_len > len)\n> -\t\t\t\tname_len -= utf8_width((const char**)&name, NULL);\n> +\t\t\twhile (name_len > len && *name) {\n\n\n\n> +\t\t\t\tint w = utf8_width((const char **)&name, NULL);\n> +\t\t\t\tif (!name) { /* Invalid UTF-8 */\n> +\t\t\t\t\tname = file->print_name;\n> +\t\t\t\t\tname_len = utf8_strwidth(name);\n> +\t\t\t\t\tbreak;\n> +\t\t\t\t}\n\nIOW, we punt on \"scaling\" and instead use the full string?  I was\nwondering if we can punt on only this segment by replacing this\nsegment with just \"...\" and resync at the next slash.\n\n> +\t\t\t\tif (w < 0)  /* control character */\n> +\t\t\t\t\tbreak;\n\nWhen we have a control characer, we instead chomp immediately before\nthat byte, which sounds good.  But then wouldn't the loop that found\nan Invalid UTF-8 sequence in the middle of a name want to do the\nsame, i.e., take the good bits found so far and chomp at the broken\nbyte?\n\n> +\t\t\t\tname_len -= w;\n> +\t\t\t}\n>  \n>  \t\t\tslash = strchr(name, '/');\n>  \t\t\tif (slash)\n\nThanks.\n"},{"id":"541841","messageId":"CABPp-BHt-O=CCnGHjoXBOHCe5CbD7beyrd_gX51g9Xg7cn_eFg@mail.gmail.com","threadId":"65507","inReplyTo":"xmqqv7dpwfy5.fsf@gitster.g","subject":"Re: [PATCH] diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2026-04-17T22:00:11Z","receivedAt":"2026-04-17T22:00:25Z","isPatch":true,"body":"On Fri, Apr 17, 2026 at 12:21 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Elijah Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > From: Elijah Newren <newren@gmail.com>\n> >\n> > f85b49f3d4a (diff: improve scaling of filenames in diffstat to handle\n> > UTF-8 chars, 2024-10-27) introduced a loop in show_stats() that calls\n> > utf8_width() repeatedly to skip leading characters until the displayed\n> > width fits.\n>\n> A tangent, but I get a datestamp for the same f85b49f3 (diff:\n> improve scaling of filenames in diffstat to handle UTF-8 chars,\n> 2026-01-16) that is different from what you showed above.  Did you\n> find a bug in \"git show -s --pretty=reference\"?\n\nHmm, indeed I get 2026-01-16 as well; I'm not sure what happened there.\n\n> > diff --git a/diff.c b/diff.c\n> > index 397e38b41c..7b27241733 100644\n> > --- a/diff.c\n> > +++ b/diff.c\n> > @@ -3093,8 +3093,17 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)\n> >                       if (len < 0)\n> >                               len = 0;\n> >\n> > -                     while (name_len > len)\n> > -                             name_len -= utf8_width((const char**)&name, NULL);\n> > +                     while (name_len > len && *name) {\n>\n>\n>\n> > +                             int w = utf8_width((const char **)&name, NULL);\n> > +                             if (!name) { /* Invalid UTF-8 */\n> > +                                     name = file->print_name;\n> > +                                     name_len = utf8_strwidth(name);\n> > +                                     break;\n> > +                             }\n>\n> IOW, we punt on \"scaling\" and instead use the full string?  I was\n> wondering if we can punt on only this segment by replacing this\n> segment with just \"...\" and resync at the next slash.\n\nGood point.  Alternatively, perhaps I could just add a wrapper around\nutf8_width() which never sets name to NULL and never returns a\nnegative value, and then use the original loop as-is other than\ncalling the new function?\n\n>\n> > +                             if (w < 0)  /* control character */\n> > +                                     break;\n>\n> When we have a control characer, we instead chomp immediately before\n> that byte, which sounds good.  But then wouldn't the loop that found\n> an Invalid UTF-8 sequence in the middle of a name want to do the\n> same, i.e., take the good bits found so far and chomp at the broken\n> byte?\n\nMakes sense, though I think my simpler alternative might be easier.\nI'll send in a re-roll.\n\n>\n> > +                             name_len -= w;\n> > +                     }\n> >\n> >                       slash = strchr(name, '/');\n> >                       if (slash)\n>\n> Thanks.\n"},{"id":"541843","messageId":"xmqq4il9w7ls.fsf@gitster.g","threadId":"65507","inReplyTo":"CABPp-BHt-O=CCnGHjoXBOHCe5CbD7beyrd_gX51g9Xg7cn_eFg@mail.gmail.com","subject":"Re: [PATCH] diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-17T22:21:51Z","receivedAt":"2026-04-17T22:21:53Z","isPatch":true,"body":"Elijah Newren <newren@gmail.com> writes:\n\n> Makes sense, though I think my simpler alternative might be easier.\n> I'll send in a re-roll.\n\nAs long as \"an invalid UTF-8\" and \"a control character\" behaves more\nor less the same (i.e., \"eek, we cannot measure the width of the\nUTF-8 character at this byte position, so let's do X as a fallback\",\nwhere X is the same regardless of the exact reason why we cannot\nmeasure the width), I'll be happy.  If we see a slash after the\nproblematic position, advancing to that slash might be the simplest,\nas that is in line with how the code works when there is no such\nproblem, but we also need to be prepared for a filename whose last\ncomponent is sufficiently long that we see no such slash after the\nproblematic byte.\n"},{"id":"541845","messageId":"pull.2093.v2.git.1776465910538.gitgitgadget@gmail.com","threadId":"65507","inReplyTo":"pull.2093.git.1776443163041.gitgitgadget@gmail.com","subject":"[PATCH v2] diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation","fromName":"Elijah Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-17T22:45:10Z","receivedAt":"2026-04-17T22:45:13Z","isPatch":true,"body":"From: Elijah Newren <newren@gmail.com>\n\nf85b49f3d4a (diff: improve scaling of filenames in diffstat to handle\nUTF-8 chars, 2026-01-16) introduced a loop in show_stats() that calls\nutf8_width() repeatedly to skip leading characters until the displayed\nwidth fits.  However, utf8_width() can return problematic values:\n\n  - For invalid UTF-8 sequences, pick_one_utf8_char() sets the name\n    pointer to NULL and utf8_width() returns 0.  Since name_len does\n    not change, the loop iterates once more and pick_one_utf8_char()\n    dereferences the NULL pointer, crashing.\n\n  - For control characters, utf8_width() returns -1, so name_len\n    grows when it is expected to shrink.  This can cause the loop to\n    consume more characters than the string contains, reading past\n    the trailing NUL.\n\nBy default, fill_print_name() will C-quotes filenames which escapes\ncontrol characters and invalid bytes to printable text.  That avoids\nthis bug from being triggered; however, with core.quotePath=false,\nraw bytes can reach this code.\n\nAdd tests exercising both failure modes with core.quotePath=false and\na narrow --stat-name-width to force truncation: one with a bare 0xC0\nbyte (invalid UTF-8 lead byte, triggers NULL deref) and one with a\n0x01 byte (control character, causes the loop to read past the end\nof the string).\n\nFix both issues by introducing utf8_ish_width(), a thin wrapper\naround utf8_width() that guarantees the pointer always advances and\nthe returned width is never negative:\n\n  - On invalid UTF-8 it restores the pointer, advances by one byte,\n    and returns width 1 (matching the strlen()-based fallback used\n    by utf8_strwidth()).\n  - On a control character it returns 0 (matching utf8_strnwidth()\n    which skips them).\n\nAlso add a \"&& *name\" guard to the while-loop condition so it\nterminates at end-of-string even when utf8_strwidth()'s strlen()\nfallback causes name_len to exceed the sum of per-character widths.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n    diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8\n    truncation\n    \n    Changes since v1:\n    \n     * Simplified the loop to almost what we had before via a wrapper\n       function that always succeeds in advancing the string and never\n       returns a negative width. (Which, as a consequence, treats invalid\n       UTF-8 and control characters the roughly the same, unlike v1.)\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2093%2Fnewren%2Ffix%2Fdiffstat-utf8-loop-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2093/newren/fix/diffstat-utf8-loop-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/2093\n\nRange-diff vs v1:\n\n 1:  fcd44d6cf8 ! 1:  4a72647ce2 diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation\n     @@ Commit message\n          diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation\n      \n          f85b49f3d4a (diff: improve scaling of filenames in diffstat to handle\n     -    UTF-8 chars, 2024-10-27) introduced a loop in show_stats() that calls\n     +    UTF-8 chars, 2026-01-16) introduced a loop in show_stats() that calls\n          utf8_width() repeatedly to skip leading characters until the displayed\n          width fits.  However, utf8_width() can return problematic values:\n      \n     @@ Commit message\n          0x01 byte (control character, causes the loop to read past the end\n          of the string).\n      \n     -    Fix the bug by:\n     -      - Adding a *name check to terminate the loop at end-of-string\n     -      - Detecting the NULL pointer from invalid UTF-8 and falling back to\n     -        showing the full untruncated name\n     -      - Breaking on negative width (control characters)\n     +    Fix both issues by introducing utf8_ish_width(), a thin wrapper\n     +    around utf8_width() that guarantees the pointer always advances and\n     +    the returned width is never negative:\n     +\n     +      - On invalid UTF-8 it restores the pointer, advances by one byte,\n     +        and returns width 1 (matching the strlen()-based fallback used\n     +        by utf8_strwidth()).\n     +      - On a control character it returns 0 (matching utf8_strnwidth()\n     +        which skips them).\n     +\n     +    Also add a \"&& *name\" guard to the while-loop condition so it\n     +    terminates at end-of-string even when utf8_strwidth()'s strlen()\n     +    fallback causes name_len to exceed the sum of per-character widths.\n      \n          Signed-off-by: Elijah Newren <newren@gmail.com>\n      \n       ## diff.c ##\n     +@@ diff.c: void print_stat_summary(FILE *fp, int files,\n     + \tprint_stat_summary_inserts_deletes(&o, files, insertions, deletions);\n     + }\n     + \n     ++/*\n     ++ * Like utf8_width(), but guaranteed safe for use in loops that subtract\n     ++ * per-character widths:\n     ++ *\n     ++ *   - utf8_width() sets *start to NULL on invalid UTF-8 and returns 0;\n     ++ *     we restore the pointer and advance by one byte, returning width 1\n     ++ *     (matching the strlen()-based fallback in utf8_strwidth()).\n     ++ *\n     ++ *   - utf8_width() returns -1 for control characters; we return 0\n     ++ *     (matching utf8_strnwidth() which skips them).\n     ++ */\n     ++static int utf8_ish_width(const char **start)\n     ++{\n     ++\tconst char *old = *start;\n     ++\tint w = utf8_width(start, NULL);\n     ++\tif (!*start) {\n     ++\t\t*start = old + 1;\n     ++\t\treturn 1;\n     ++\t}\n     ++\treturn (w < 0) ? 0 : w;\n     ++}\n     ++\n     + static void show_stats(struct diffstat_t *data, struct diff_options *options)\n     + {\n     + \tint i, len, add, del, adds = 0, dels = 0;\n      @@ diff.c: static void show_stats(struct diffstat_t *data, struct diff_options *options)\n       \t\t\tif (len < 0)\n       \t\t\t\tlen = 0;\n       \n      -\t\t\twhile (name_len > len)\n      -\t\t\t\tname_len -= utf8_width((const char**)&name, NULL);\n     -+\t\t\twhile (name_len > len && *name) {\n     -+\t\t\t\tint w = utf8_width((const char **)&name, NULL);\n     -+\t\t\t\tif (!name) { /* Invalid UTF-8 */\n     -+\t\t\t\t\tname = file->print_name;\n     -+\t\t\t\t\tname_len = utf8_strwidth(name);\n     -+\t\t\t\t\tbreak;\n     -+\t\t\t\t}\n     -+\t\t\t\tif (w < 0)  /* control character */\n     -+\t\t\t\t\tbreak;\n     -+\t\t\t\tname_len -= w;\n     -+\t\t\t}\n     ++\t\t\twhile (name_len > len && *name)\n     ++\t\t\t\tname_len -= utf8_ish_width((const char**)&name);\n       \n       \t\t\tslash = strchr(name, '/');\n       \t\t\tif (slash)\n\n\n diff.c                 | 26 ++++++++++++++++++++++++--\n t/t4052-stat-output.sh | 25 +++++++++++++++++++++++++\n 2 files changed, 49 insertions(+), 2 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex 397e38b41c..1a3b19f71f 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -2927,6 +2927,28 @@ void print_stat_summary(FILE *fp, int files,\n \tprint_stat_summary_inserts_deletes(&o, files, insertions, deletions);\n }\n \n+/*\n+ * Like utf8_width(), but guaranteed safe for use in loops that subtract\n+ * per-character widths:\n+ *\n+ *   - utf8_width() sets *start to NULL on invalid UTF-8 and returns 0;\n+ *     we restore the pointer and advance by one byte, returning width 1\n+ *     (matching the strlen()-based fallback in utf8_strwidth()).\n+ *\n+ *   - utf8_width() returns -1 for control characters; we return 0\n+ *     (matching utf8_strnwidth() which skips them).\n+ */\n+static int utf8_ish_width(const char **start)\n+{\n+\tconst char *old = *start;\n+\tint w = utf8_width(start, NULL);\n+\tif (!*start) {\n+\t\t*start = old + 1;\n+\t\treturn 1;\n+\t}\n+\treturn (w < 0) ? 0 : w;\n+}\n+\n static void show_stats(struct diffstat_t *data, struct diff_options *options)\n {\n \tint i, len, add, del, adds = 0, dels = 0;\n@@ -3093,8 +3115,8 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)\n \t\t\tif (len < 0)\n \t\t\t\tlen = 0;\n \n-\t\t\twhile (name_len > len)\n-\t\t\t\tname_len -= utf8_width((const char**)&name, NULL);\n+\t\t\twhile (name_len > len && *name)\n+\t\t\t\tname_len -= utf8_ish_width((const char**)&name);\n \n \t\t\tslash = strchr(name, '/');\n \t\t\tif (slash)\ndiff --git a/t/t4052-stat-output.sh b/t/t4052-stat-output.sh\nindex 7c749062e2..84c53c1a51 100755\n--- a/t/t4052-stat-output.sh\n+++ b/t/t4052-stat-output.sh\n@@ -445,4 +445,29 @@ test_expect_success 'diffstat where line_prefix contains ANSI escape codes is co\n \ttest_grep \"<RED>|<RESET>  ${FILENAME_TRIMMED} | 0\" out\n '\n \n+test_expect_success 'diffstat truncation with invalid UTF-8 does not crash' '\n+\tempty_blob=$(git hash-object -w --stdin </dev/null) &&\n+\tprintf \"100644 blob $empty_blob\\taaa-\\300-aaa\\n\" |\n+\tgit mktree >tree_file &&\n+\ttree=$(cat tree_file) &&\n+\tempty_tree=$(git mktree </dev/null) &&\n+\tc1=$(git commit-tree -m before $empty_tree) &&\n+\tc2=$(git commit-tree -m after -p $c1 $tree) &&\n+\tgit -c core.quotepath=false diff --stat --stat-name-width=5 $c1..$c2 >output &&\n+\ttest_grep \"| 0\" output\n+'\n+\n+test_expect_success FUNNYNAMES 'diffstat truncation with control chars does not crash' '\n+\tFNAME=$(printf \"aaa-\\x01-aaa\") &&\n+\tgit commit --allow-empty -m setup &&\n+\t>$FNAME &&\n+\tgit add -- $FNAME &&\n+\tgit commit -m \"add file with control char name\" &&\n+\tgit -c core.quotepath=false diff --stat --stat-name-width=5 HEAD~1..HEAD >output &&\n+\ttest_grep \"| 0\" output &&\n+\trm -- $FNAME &&\n+\tgit rm -- $FNAME &&\n+\tgit commit -m \"remove test file\"\n+'\n+\n test_done\n\nbase-commit: 9f223ef1c026d91c7ac68cc0211bde255dda6199\n-- \ngitgitgadget\n"},{"id":"541894","messageId":"aeVqqsdq9B7GE9gS@lorenzo-VM","threadId":"65507","inReplyTo":"pull.2093.v2.git.1776465910538.gitgitgadget@gmail.com","subject":"Re: [PATCH v2] diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation","fromName":"Lorenzo Pegorari","fromEmail":"lorenzo.pegorari2002@gmail.com","sentAt":"2026-04-19T23:52:10Z","receivedAt":"2026-04-19T23:52:14Z","isPatch":true,"body":"On Fri, Apr 17, 2026 at 10:45:10PM +0000, Elijah Newren via GitGitGadget wrote:\n> From: Elijah Newren <newren@gmail.com>\n> \n> f85b49f3d4a (diff: improve scaling of filenames in diffstat to handle\n> UTF-8 chars, 2026-01-16) introduced a loop in show_stats() that calls\n> utf8_width() repeatedly to skip leading characters until the displayed\n> width fits.  However, utf8_width() can return problematic values:\n> \n>   - For invalid UTF-8 sequences, pick_one_utf8_char() sets the name\n>     pointer to NULL and utf8_width() returns 0.  Since name_len does\n>     not change, the loop iterates once more and pick_one_utf8_char()\n>     dereferences the NULL pointer, crashing.\n> \n>   - For control characters, utf8_width() returns -1, so name_len\n>     grows when it is expected to shrink.  This can cause the loop to\n>     consume more characters than the string contains, reading past\n>     the trailing NUL.\n> \n> By default, fill_print_name() will C-quotes filenames which escapes\n> control characters and invalid bytes to printable text.  That avoids\n> this bug from being triggered; however, with core.quotePath=false,\n> raw bytes can reach this code.\n> \n> Add tests exercising both failure modes with core.quotePath=false and\n> a narrow --stat-name-width to force truncation: one with a bare 0xC0\n> byte (invalid UTF-8 lead byte, triggers NULL deref) and one with a\n> 0x01 byte (control character, causes the loop to read past the end\n> of the string).\n> \n> Fix both issues by introducing utf8_ish_width(), a thin wrapper\n> around utf8_width() that guarantees the pointer always advances and\n> the returned width is never negative:\n> \n>   - On invalid UTF-8 it restores the pointer, advances by one byte,\n>     and returns width 1 (matching the strlen()-based fallback used\n>     by utf8_strwidth()).\n>   - On a control character it returns 0 (matching utf8_strnwidth()\n>     which skips them).\n> \n> Also add a \"&& *name\" guard to the while-loop condition so it\n> terminates at end-of-string even when utf8_strwidth()'s strlen()\n> fallback causes name_len to exceed the sum of per-character widths.\ni> \n> Signed-off-by: Elijah Newren <newren@gmail.com>\n\nHi, thanks for CCing me and thanks for improving on my previous work.\n\nAll of these changes make a lot of sense, and indeed they fix issues\nthat I didn't consider in f85b49f3d4a (diff: improve scaling of\nfilenames in diffstat to handle UTF-8 chars, 2026-01-16).\n\n[...]\n\n> diff --git a/t/t4052-stat-output.sh b/t/t4052-stat-output.sh\n> index 7c749062e2..84c53c1a51 100755\n> --- a/t/t4052-stat-output.sh\n> +++ b/t/t4052-stat-output.sh\n> @@ -445,4 +445,29 @@ test_expect_success 'diffstat where line_prefix contains ANSI escape codes is co\n\n[...]\n\n>\n> +test_expect_success FUNNYNAMES 'diffstat truncation with control chars does not crash' '\n> +\tFNAME=$(printf \"aaa-\\x01-aaa\") &&\n> +\tgit commit --allow-empty -m setup &&\n> +\t>$FNAME &&\n> +\tgit add -- $FNAME &&\n> +\tgit commit -m \"add file with control char name\" &&\n> +\tgit -c core.quotepath=false diff --stat --stat-name-width=5 HEAD~1..HEAD >output &&\n> +\ttest_grep \"| 0\" output &&\n> +\trm -- $FNAME &&\n> +\tgit rm -- $FNAME &&\n> +\tgit commit -m \"remove test file\"\n> +'\n> +\n>  test_done\n\nThe only thing that I don't quite understand is this second test.\n\nFrom my tests, the previous code using:\n\n```\n[...]\nwhile (name_len > len)\n\tname_len -= utf8_width((const char**)&name, NULL);\n[...]\n```\n\npasses this second test just fine, while I believe it's supposed to\nfail.\n\nAm I missing something?\n\n\nThanks,\nLorenzo\n"},{"id":"541975","messageId":"CABPp-BHgnyS_SB6SX1dzAezfExomHts1t02+qr+duCPW6sk1nQ@mail.gmail.com","threadId":"65507","inReplyTo":"aeVqqsdq9B7GE9gS@lorenzo-VM","subject":"Re: [PATCH v2] diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2026-04-20T14:51:31Z","receivedAt":"2026-04-20T14:51:44Z","isPatch":true,"body":"On Sun, Apr 19, 2026 at 4:52 PM Lorenzo Pegorari\n<lorenzo.pegorari2002@gmail.com> wrote:\n>\n> > +test_expect_success FUNNYNAMES 'diffstat truncation with control chars does not crash' '\n> > +     FNAME=$(printf \"aaa-\\x01-aaa\") &&\n> > +     git commit --allow-empty -m setup &&\n> > +     >$FNAME &&\n> > +     git add -- $FNAME &&\n> > +     git commit -m \"add file with control char name\" &&\n> > +     git -c core.quotepath=false diff --stat --stat-name-width=5 HEAD~1..HEAD >output &&\n> > +     test_grep \"| 0\" output &&\n> > +     rm -- $FNAME &&\n> > +     git rm -- $FNAME &&\n> > +     git commit -m \"remove test file\"\n> > +'\n> > +\n> >  test_done\n>\n> The only thing that I don't quite understand is this second test.\n>\n> From my tests, the previous code using:\n>\n> ```\n> [...]\n> while (name_len > len)\n>         name_len -= utf8_width((const char**)&name, NULL);\n> [...]\n> ```\n>\n> passes this second test just fine, while I believe it's supposed to\n> fail.\n>\n> Am I missing something?\n\nSorry, I did two things wrong -- I forgot to specify that the second\ntest only fails under ASan, and I simplified the test too much such\nthat it doesn't fail under ASan without the fixes (and simplified in\nthree wrong ways: not enough control characters, wrong kind of control\ncharacter, attempting to use hex control code to printf instead of\noctal) and apparently forgot to re-check afterwards.  Using the\nfilename\n    FNAME=$(printf \"aaa-\\302\\237\\302\\237\\302\\237-aaa\") &&\nwill trigger the out-of-bounds read under ASan before the fixes;\nremoving the final \\302\\237 will make it pass with or without the code\nfixes.  I'll correct the patch and send in a new round.\n\nThanks for checking closely.\n"},{"id":"541979","messageId":"pull.2093.v3.git.1776699778177.gitgitgadget@gmail.com","threadId":"65507","inReplyTo":"pull.2093.v2.git.1776465910538.gitgitgadget@gmail.com","subject":"[PATCH v3] diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation","fromName":"Elijah Newren via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-20T15:42:58Z","receivedAt":"2026-04-20T15:43:01Z","isPatch":true,"body":"From: Elijah Newren <newren@gmail.com>\n\nf85b49f3d4a (diff: improve scaling of filenames in diffstat to handle\nUTF-8 chars, 2026-01-16) introduced a loop in show_stats() that calls\nutf8_width() repeatedly to skip leading characters until the displayed\nwidth fits.  However, utf8_width() can return problematic values:\n\n  - For invalid UTF-8 sequences, pick_one_utf8_char() sets the name\n    pointer to NULL and utf8_width() returns 0.  Since name_len does\n    not change, the loop iterates once more and pick_one_utf8_char()\n    dereferences the NULL pointer, crashing.\n\n  - For control characters, utf8_width() returns -1, so name_len\n    grows when it is expected to shrink.  This can cause the loop to\n    consume more characters than the string contains, reading past\n    the trailing NUL.\n\nBy default, fill_print_name() will C-quote filenames which escapes\ncontrol characters and invalid bytes to printable text.  That avoids\nthis bug from being triggered; however, with core.quotePath=false,\nmost characters are no longer escaped (though some control characters\nstill are) and raw bytes can reach this code.\n\nAdd tests exercising both failure modes with core.quotePath=false and\na narrow --stat-name-width to force truncation: one with a bare 0xC0\nbyte (invalid UTF-8 lead byte, triggers NULL deref) and one with\nseveral C1 control characters (repeats of 0xC2 0x9F, causing\nthe loop to read past the end of the string).  The second test\nreliably catches the out-of-bounds read when run under ASan, though\nit may pass silently without sanitizers.\n\nFix both issues by introducing utf8_ish_width(), a thin wrapper\naround utf8_width() that guarantees the pointer always advances and\nthe returned width is never negative:\n\n  - On invalid UTF-8 it restores the pointer, advances by one byte,\n    and returns width 1 (matching the strlen()-based fallback used\n    by utf8_strwidth()).\n  - On a control character it returns 0 (matching utf8_strnwidth()\n    which skips them).\n\nAlso add a \"&& *name\" guard to the while-loop condition so it\nterminates at end-of-string even when utf8_strwidth()'s strlen()\nfallback causes name_len to exceed the sum of per-character widths.\n\nSigned-off-by: Elijah Newren <newren@gmail.com>\n---\n    diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8\n    truncation\n    \n    Changes since v2:\n    \n     * Fixed the filename in the final test such that it will trigger the\n       out-of-bounds read under ASan, and updated the commit message to\n       point out that ASan is needed to notice the out-of-bounds read.\n    \n    Changes since v1:\n    \n     * Simplified the loop to almost what we had before via a wrapper\n       function that always succeeds in advancing the string and never\n       returns a negative width. (Which, as a consequence, treats invalid\n       UTF-8 and control characters the roughly the same, unlike v1.)\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2093%2Fnewren%2Ffix%2Fdiffstat-utf8-loop-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2093/newren/fix/diffstat-utf8-loop-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/2093\n\nRange-diff vs v2:\n\n 1:  4a72647ce2 ! 1:  4a3126720b diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation\n     @@ Commit message\n              consume more characters than the string contains, reading past\n              the trailing NUL.\n      \n     -    By default, fill_print_name() will C-quotes filenames which escapes\n     +    By default, fill_print_name() will C-quote filenames which escapes\n          control characters and invalid bytes to printable text.  That avoids\n          this bug from being triggered; however, with core.quotePath=false,\n     -    raw bytes can reach this code.\n     +    most characters are no longer escaped (though some control characters\n     +    still are) and raw bytes can reach this code.\n      \n          Add tests exercising both failure modes with core.quotePath=false and\n          a narrow --stat-name-width to force truncation: one with a bare 0xC0\n     -    byte (invalid UTF-8 lead byte, triggers NULL deref) and one with a\n     -    0x01 byte (control character, causes the loop to read past the end\n     -    of the string).\n     +    byte (invalid UTF-8 lead byte, triggers NULL deref) and one with\n     +    several C1 control characters (repeats of 0xC2 0x9F, causing\n     +    the loop to read past the end of the string).  The second test\n     +    reliably catches the out-of-bounds read when run under ASan, though\n     +    it may pass silently without sanitizers.\n      \n          Fix both issues by introducing utf8_ish_width(), a thin wrapper\n          around utf8_width() that guarantees the pointer always advances and\n     @@ t/t4052-stat-output.sh: test_expect_success 'diffstat where line_prefix contains\n      +\ttest_grep \"| 0\" output\n      +'\n      +\n     -+test_expect_success FUNNYNAMES 'diffstat truncation with control chars does not crash' '\n     -+\tFNAME=$(printf \"aaa-\\x01-aaa\") &&\n     ++test_expect_success FUNNYNAMES 'diffstat truncation with control chars does not read out of bounds' '\n     ++\tFNAME=$(printf \"aaa-\\302\\237\\302\\237\\302\\237-aaa\") &&\n      +\tgit commit --allow-empty -m setup &&\n      +\t>$FNAME &&\n      +\tgit add -- $FNAME &&\n\n\n diff.c                 | 26 ++++++++++++++++++++++++--\n t/t4052-stat-output.sh | 25 +++++++++++++++++++++++++\n 2 files changed, 49 insertions(+), 2 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex 397e38b41c..1a3b19f71f 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -2927,6 +2927,28 @@ void print_stat_summary(FILE *fp, int files,\n \tprint_stat_summary_inserts_deletes(&o, files, insertions, deletions);\n }\n \n+/*\n+ * Like utf8_width(), but guaranteed safe for use in loops that subtract\n+ * per-character widths:\n+ *\n+ *   - utf8_width() sets *start to NULL on invalid UTF-8 and returns 0;\n+ *     we restore the pointer and advance by one byte, returning width 1\n+ *     (matching the strlen()-based fallback in utf8_strwidth()).\n+ *\n+ *   - utf8_width() returns -1 for control characters; we return 0\n+ *     (matching utf8_strnwidth() which skips them).\n+ */\n+static int utf8_ish_width(const char **start)\n+{\n+\tconst char *old = *start;\n+\tint w = utf8_width(start, NULL);\n+\tif (!*start) {\n+\t\t*start = old + 1;\n+\t\treturn 1;\n+\t}\n+\treturn (w < 0) ? 0 : w;\n+}\n+\n static void show_stats(struct diffstat_t *data, struct diff_options *options)\n {\n \tint i, len, add, del, adds = 0, dels = 0;\n@@ -3093,8 +3115,8 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)\n \t\t\tif (len < 0)\n \t\t\t\tlen = 0;\n \n-\t\t\twhile (name_len > len)\n-\t\t\t\tname_len -= utf8_width((const char**)&name, NULL);\n+\t\t\twhile (name_len > len && *name)\n+\t\t\t\tname_len -= utf8_ish_width((const char**)&name);\n \n \t\t\tslash = strchr(name, '/');\n \t\t\tif (slash)\ndiff --git a/t/t4052-stat-output.sh b/t/t4052-stat-output.sh\nindex 7c749062e2..e009585925 100755\n--- a/t/t4052-stat-output.sh\n+++ b/t/t4052-stat-output.sh\n@@ -445,4 +445,29 @@ test_expect_success 'diffstat where line_prefix contains ANSI escape codes is co\n \ttest_grep \"<RED>|<RESET>  ${FILENAME_TRIMMED} | 0\" out\n '\n \n+test_expect_success 'diffstat truncation with invalid UTF-8 does not crash' '\n+\tempty_blob=$(git hash-object -w --stdin </dev/null) &&\n+\tprintf \"100644 blob $empty_blob\\taaa-\\300-aaa\\n\" |\n+\tgit mktree >tree_file &&\n+\ttree=$(cat tree_file) &&\n+\tempty_tree=$(git mktree </dev/null) &&\n+\tc1=$(git commit-tree -m before $empty_tree) &&\n+\tc2=$(git commit-tree -m after -p $c1 $tree) &&\n+\tgit -c core.quotepath=false diff --stat --stat-name-width=5 $c1..$c2 >output &&\n+\ttest_grep \"| 0\" output\n+'\n+\n+test_expect_success FUNNYNAMES 'diffstat truncation with control chars does not read out of bounds' '\n+\tFNAME=$(printf \"aaa-\\302\\237\\302\\237\\302\\237-aaa\") &&\n+\tgit commit --allow-empty -m setup &&\n+\t>$FNAME &&\n+\tgit add -- $FNAME &&\n+\tgit commit -m \"add file with control char name\" &&\n+\tgit -c core.quotepath=false diff --stat --stat-name-width=5 HEAD~1..HEAD >output &&\n+\ttest_grep \"| 0\" output &&\n+\trm -- $FNAME &&\n+\tgit rm -- $FNAME &&\n+\tgit commit -m \"remove test file\"\n+'\n+\n test_done\n\nbase-commit: 9f223ef1c026d91c7ac68cc0211bde255dda6199\n-- \ngitgitgadget\n"},{"id":"541986","messageId":"xmqq8qahr3ca.fsf@gitster.g","threadId":"65507","inReplyTo":"pull.2093.v3.git.1776699778177.gitgitgadget@gmail.com","subject":"Re: [PATCH v3] diff: fix out-of-bounds reads and NULL deref in diffstat UTF-8 truncation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-20T16:41:57Z","receivedAt":"2026-04-20T16:42:00Z","isPatch":true,"body":"\"Elijah Newren via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Elijah Newren <newren@gmail.com>\n>\n> f85b49f3d4a (diff: improve scaling of filenames in diffstat to handle\n> UTF-8 chars, 2026-01-16) introduced a loop in show_stats() that calls\n> utf8_width() repeatedly to skip leading characters until the displayed\n> width fits.  However, utf8_width() can return problematic values:\n>\n>   - For invalid UTF-8 sequences, pick_one_utf8_char() sets the name\n>     pointer to NULL and utf8_width() returns 0.  Since name_len does\n>     not change, the loop iterates once more and pick_one_utf8_char()\n>     dereferences the NULL pointer, crashing.\n>\n>   - For control characters, utf8_width() returns -1, so name_len\n>     grows when it is expected to shrink.  This can cause the loop to\n>     consume more characters than the string contains, reading past\n>     the trailing NUL.\n>\n> By default, fill_print_name() will C-quote filenames which escapes\n> control characters and invalid bytes to printable text.  That avoids\n> this bug from being triggered; however, with core.quotePath=false,\n> most characters are no longer escaped (though some control characters\n> still are) and raw bytes can reach this code.\n>\n> Add tests exercising both failure modes with core.quotePath=false and\n> a narrow --stat-name-width to force truncation: one with a bare 0xC0\n> byte (invalid UTF-8 lead byte, triggers NULL deref) and one with\n> several C1 control characters (repeats of 0xC2 0x9F, causing\n> the loop to read past the end of the string).  The second test\n> reliably catches the out-of-bounds read when run under ASan, though\n> it may pass silently without sanitizers.\n>\n> Fix both issues by introducing utf8_ish_width(), a thin wrapper\n> around utf8_width() that guarantees the pointer always advances and\n> the returned width is never negative:\n>\n>   - On invalid UTF-8 it restores the pointer, advances by one byte,\n>     and returns width 1 (matching the strlen()-based fallback used\n>     by utf8_strwidth()).\n>   - On a control character it returns 0 (matching utf8_strnwidth()\n>     which skips them).\n>\n> Also add a \"&& *name\" guard to the while-loop condition so it\n> terminates at end-of-string even when utf8_strwidth()'s strlen()\n> fallback causes name_len to exceed the sum of per-character widths.\n\nOK, that does sounds sensible.\n\nIf we start from a valid UTF-8 string, chomp a few bytes from the\ntail end of it, and feed it into this loop, the initial part of the\nlast character is fed to utf8_width(), which hopefully is already\nprepared to honor the NUL termination to avoid an OOB read while\nreturning an error.  And eventually we would see that NUL that\ntruncated the last UTF-8 multi-byte letter ourselves in the loop and\nthat is where this new loop terminating condition would help.\n\n"}]}