Re: [PATCH 1/2] t/unit-tests: add UTF-8 width tests for CJK chars
On Sat, Nov 15, 2025 at 4:17 AM Junio C Hamano <gitster@pobox.com> wrote:
Show 35 quoted lines
>
> Jiang Xin <worldhello.net@gmail.com> writes:
>
> [jc: the same question about the choice of Cc addresses applies]
>
> > This commit adds a new test suite (u-utf8-width.c) to test the UTF-8
> > width functions in Git, particularly focusing on multi-byte characters
> > from East Asian languages like Chinese, Japanese, and Korean that
> > typically require 2 display columns per character.
> >
> > The test suite includes:
> > - Tests for utf8_strnwidth with Chinese strings
> > - Tests for utf8_strwidth with Chinese strings
> > - Tests for Japanese and Korean characters
> > - Edge case tests with invalid UTF-8 sequences
> > - Proper test function naming following the Clar framework convention
> >
> > Also updated the build configuration in Makefile and meson.build to
> > include the new test suite in the build process.
>
> The usual way to compose a log message of this project is to
>
> - Give an observation on how the current system works in the
> present tense (so no need to say "Currently X is Y", or
> "Previously X was Y" to describe the state before your change;
> just "X is Y" is enough), and discuss what you perceive as a
> problem in it.
>
> - Propose a solution (optional---often, problem description
> trivially leads to an obvious solution in reader's minds).
>
> - Give commands to somebody editing the codebase to "make it so",
> instead of saying "This commit does X".
>
> in this order.
Will document the purpose in commit message of next reroll.
Show 8 quoted lines
> > +/*
> > + * Test edge cases with partial UTF-8 sequences
> > + */
>
> All tests before these make sense, but I am not sure if we want to
> hold utf8_strnwidth() to the requirement that it will tolerate "len"
> to end in the middle of a single character, as such a requirement by
> itself does not do application any good.
Will remove unnecessary test cases.