git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 1/2] t/unit-tests: add UTF-8 width tests for CJK chars

From
Jiang Xin <worldhello.net@gmail.com>
Date
Nov 15, 2025, 12:38 UTC
Message-ID
<CANYiYbEhGo2Z1Y+YaXfOs35+2nTOLYW4C=yXR65b4wGfe2nFgw@mail.gmail.com>
In-Reply-To
<xmqqzf8ogyhw.fsf@gitster.g>
On Sat, Nov 15, 2025 at 4:17 AM Junio C Hamano <gitster@pobox.com> wrote:
Show 35 quoted lines
>
> Jiang Xin <worldhello.net@gmail.com> writes:
>
> [jc: the same question about the choice of Cc addresses applies]
>
> > This commit adds a new test suite (u-utf8-width.c) to test the UTF-8
> > width functions in Git, particularly focusing on multi-byte characters
> > from East Asian languages like Chinese, Japanese, and Korean that
> > typically require 2 display columns per character.
> >
> > The test suite includes:
> > - Tests for utf8_strnwidth with Chinese strings
> > - Tests for utf8_strwidth with Chinese strings
> > - Tests for Japanese and Korean characters
> > - Edge case tests with invalid UTF-8 sequences
> > - Proper test function naming following the Clar framework convention
> >
> > Also updated the build configuration in Makefile and meson.build to
> > include the new test suite in the build process.
>
> The usual way to compose a log message of this project is to
>
>  - Give an observation on how the current system works in the
>    present tense (so no need to say "Currently X is Y", or
>    "Previously X was Y" to describe the state before your change;
>    just "X is Y" is enough), and discuss what you perceive as a
>    problem in it.
>
>  - Propose a solution (optional---often, problem description
>    trivially leads to an obvious solution in reader's minds).
>
>  - Give commands to somebody editing the codebase to "make it so",
>    instead of saying "This commit does X".
>
> in this order.
Will document the purpose in commit message of next reroll.
Show 8 quoted lines
> > +/*
> > + * Test edge cases with partial UTF-8 sequences
> > + */
>
> All tests before these make sense, but I am not sure if we want to
> hold utf8_strnwidth() to the requirement that it will tolerate "len"
> to end in the middle of a single character, as such a requirement by
> itself does not do application any good.
Will remove unnecessary test cases.
Previous: Junio C HamanoNext: Jiang Xin
Message 4 of 22 in “Fix misaligned output of git repo structure”
  1. 0/2 Fix misaligned output of git repo structureJiang Xin, Nov 14, 2025
  2. 1/2 t/unit-tests: add UTF-8 width tests for CJK charsJiang Xin, Nov 14, 2025
  3. Junio C HamanoNov 14, 2025
  4. Jiang XinNov 15, 2025
  5. 2/2 builtin/repo: fix table alignment for UTF-8 charactersJiang Xin, Nov 14, 2025
  6. Justin ToblerNov 14, 2025
  7. Jiang XinNov 15, 2025
  8. Junio C HamanoNov 14, 2025
  9. Jiang XinNov 15, 2025
  10. Junio C HamanoNov 15, 2025
  11. Jiang XinNov 16, 2025
  12. Junio C HamanoNov 16, 2025
  13. Kristoffer HaugsbakkNov 14, 2025
  14. Jiang XinNov 14, 2025
  15. Junio C HamanoNov 14, 2025
  16. Jiang XinNov 15, 2025
  17. Junio C HamanoNov 14, 2025
  18. 0/2 Fix misaligned output of git repo structureJiang Xin, Nov 15, 2025
  19. 1/2 t/unit-tests: add UTF-8 width tests for CJK charsJiang Xin, Nov 15, 2025
  20. 2/2 builtin/repo: fix table alignment for UTF-8 charactersJiang Xin, Nov 15, 2025
  21. Phillip WoodNov 15, 2025
  22. Junio C HamanoNov 15, 2025

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.