{"thread":{"id":"64478","subject":"[PATCH 0/2] Fix misaligned output of git repo structure","startedAt":"2025-11-14T05:52:54Z","lastAt":"2025-11-16T16:51:38Z","messageCount":22,"participants":["Jiang Xin","Kristoffer Haugsbakk","Junio C Hamano","Justin Tobler","Phillip Wood"],"isPatch":true,"patchVersion":1,"patchTotal":2},"messages":[{"id":"530678","messageId":"cover.1763098804.git.worldhello.net@gmail.com","threadId":"64478","inReplyTo":null,"subject":"[PATCH 0/2] Fix misaligned output of git repo structure","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-14T05:52:43Z","receivedAt":"2025-11-14T05:52:54Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"While localizing Git 2.52.0, I noticed that the output table from git\nrepo structure becomes misaligned when displaying UTF-8 characters. For\nexample:\n\n    | 仓库结构   | 值  |\n    | -------------- | ---- |\n    | * 引用       |      |\n    |   * 计数     |   67 |\n    |     * 分支   |    6 |\n    |     * 标签   |   30 |\n    |     * 远程   |   19 |\n    |     * 其它   |   12 |\n    |                |      |\n    | * 可达对象 |      |\n    |   * 计数     | 2217 |\n    |     * 提交   |  279 |\n    |     * 树      |  740 |\n    |     * 数据对象 | 1168 |\n    |     * 标签   |   30 |\n\nThe previous implementation used simple width formatting with printf()\nwhich didn't properly handle multi-byte UTF-8 characters, causing\nmisaligned table columns when displaying repository structure\ninformation.\n\nThis change modifies the stats_table_print_structure function to use\nstrbuf_utf8_align() instead of basic printf width specifiers. This\nensures proper column alignment regardless of the character encoding of\nthe content being displayed.\n\nBTW, I used two AI coding tools (Claude Code and Gemini-CLI) to generate\nthe commits, and added the \"Co-developed-by\" trailers in the commit\nmessages by using one of my opensource project:\n\n - https://github.com/ai-coding-workshop/commit-msg\n\n\n## Changes\n\nJiang Xin (2):\n  t/unit-tests: add UTF-8 width tests for CJK chars\n  builtin/repo: fix table alignment for UTF-8 characters\n\n Makefile                    |  1 +\n builtin/repo.c              | 22 ++++++++--\n t/meson.build               |  1 +\n t/unit-tests/u-utf8-width.c | 85 +++++++++++++++++++++++++++++++++++++\n 4 files changed, 105 insertions(+), 4 deletions(-)\n create mode 100644 t/unit-tests/u-utf8-width.c\n\n-- \n2.52.0.rc2.5.g4c20a63325.dirty\n\n"},{"id":"530679","messageId":"04ab347ff80e16d49524246a8923cc86cc7355be.1763098804.git.worldhello.net@gmail.com","threadId":"64478","inReplyTo":"cover.1763098804.git.worldhello.net@gmail.com","subject":"[PATCH 1/2] t/unit-tests: add UTF-8 width tests for CJK chars","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-14T05:52:44Z","receivedAt":"2025-11-14T05:52:55Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"This commit adds a new test suite (u-utf8-width.c) to test the UTF-8\nwidth functions in Git, particularly focusing on multi-byte characters\nfrom East Asian languages like Chinese, Japanese, and Korean that\ntypically require 2 display columns per character.\n\nThe test suite includes:\n- Tests for utf8_strnwidth with Chinese strings\n- Tests for utf8_strwidth with Chinese strings\n- Tests for Japanese and Korean characters\n- Edge case tests with invalid UTF-8 sequences\n- Proper test function naming following the Clar framework convention\n\nAlso updated the build configuration in Makefile and meson.build to\ninclude the new test suite in the build process.\n\nCo-developed-by: Claude <noreply@anthropic.com>\nSigned-off-by: Jiang Xin <worldhello.net@gmail.com>\n---\n Makefile                    |  1 +\n t/meson.build               |  1 +\n t/unit-tests/u-utf8-width.c | 85 +++++++++++++++++++++++++++++++++++++\n 3 files changed, 87 insertions(+)\n create mode 100644 t/unit-tests/u-utf8-width.c\n\ndiff --git a/Makefile b/Makefile\nindex 7e0f77e298..2a67546154 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1525,6 +1525,7 @@ CLAR_TEST_SUITES += u-string-list\n CLAR_TEST_SUITES += u-strvec\n CLAR_TEST_SUITES += u-trailer\n CLAR_TEST_SUITES += u-urlmatch-normalization\n+CLAR_TEST_SUITES += u-utf8-width\n CLAR_TEST_PROG = $(UNIT_TEST_BIN)/unit-tests$(X)\n CLAR_TEST_OBJS = $(patsubst %,$(UNIT_TEST_DIR)/%.o,$(CLAR_TEST_SUITES))\n CLAR_TEST_OBJS += $(UNIT_TEST_DIR)/clar/clar.o\ndiff --git a/t/meson.build b/t/meson.build\nindex a5531df415..dc43d69636 100644\n--- a/t/meson.build\n+++ b/t/meson.build\n@@ -24,6 +24,7 @@ clar_test_suites = [\n   'unit-tests/u-strvec.c',\n   'unit-tests/u-trailer.c',\n   'unit-tests/u-urlmatch-normalization.c',\n+  'unit-tests/u-utf8-width.c',\n ]\n \n clar_sources = [\ndiff --git a/t/unit-tests/u-utf8-width.c b/t/unit-tests/u-utf8-width.c\nnew file mode 100644\nindex 0000000000..455294ca90\n--- /dev/null\n+++ b/t/unit-tests/u-utf8-width.c\n@@ -0,0 +1,85 @@\n+#include \"unit-test.h\"\n+#include \"utf8.h\"\n+#include \"strbuf.h\"\n+\n+/*\n+ * Test utf8_strnwidth with various Chinese strings\n+ * Chinese characters typically have a width of 2 columns when displayed\n+ */\n+void test_utf8_width__strnwidth_chinese(void)\n+{\n+\tconst char *ansi_test;\n+\tconst char *str;\n+\n+\t/* Test basic ASCII - each character should have width 1 */\n+\tcl_assert_equal_i(5, utf8_strnwidth(\"hello\", 5, 0));\n+\tcl_assert_equal_i(5, utf8_strnwidth(\"hello\", 5, 1));  /* skip_ansi = 1 */\n+\n+\t/* Test simple Chinese characters - each should have width 2 */\n+\tcl_assert_equal_i(4, utf8_strnwidth(\"你好\", 6, 0));  /* \"你好\" is 6 bytes (3 bytes per char in UTF-8), 4 display columns */\n+\n+\t/* Test mixed ASCII and Chinese - ASCII = 1 column, Chinese = 2 columns */\n+\tcl_assert_equal_i(6, utf8_strnwidth(\"hi你好\", 8, 0));  /* \"h\"(1) + \"i\"(1) + \"你\"(2) + \"好\"(2) = 6 */\n+\n+\t/* Test longer Chinese string */\n+\tcl_assert_equal_i(10, utf8_strnwidth(\"你好世界！\", 15, 0));  /* 5 Chinese chars = 10 display columns */\n+\n+\t/* Test with skip_ansi = 1 to make sure it works with escape sequences */\n+\tansi_test = \"\\033[31m你好\\033[0m\";\n+\tcl_assert_equal_i(4, utf8_strnwidth(ansi_test, strlen(ansi_test), 1));  /* Skip escape sequences, just count \"你好\" which should be 4 columns */\n+\n+\t/* Test individual Chinese character width */\n+\tcl_assert_equal_i(2, utf8_strnwidth(\"中\", 3, 0));  /* Single Chinese char should be 2 columns */\n+\n+\t/* Test empty string */\n+\tcl_assert_equal_i(0, utf8_strnwidth(\"\", 0, 0));\n+\n+\t/* Test length limiting */\n+\tstr = \"你好世界\";\n+\tcl_assert_equal_i(2, utf8_strnwidth(str, 3, 0));  /* Only first char \"你\"(2 columns) within 3 bytes */\n+\tcl_assert_equal_i(4, utf8_strnwidth(str, 6, 0));  /* First two chars \"你好\"(4 columns) in 6 bytes */\n+}\n+\n+/*\n+ * Tests for utf8_strwidth (simpler version without length limit)\n+ */\n+void test_utf8_width__strwidth_chinese(void)\n+{\n+\t/* Test basic ASCII */\n+\tcl_assert_equal_i(5, utf8_strwidth(\"hello\"));\n+\n+\t/* Test Chinese characters */\n+\tcl_assert_equal_i(4, utf8_strwidth(\"你好\"));  /* 2 Chinese chars = 4 display columns */\n+\n+\t/* Test mixed ASCII and Chinese */\n+\tcl_assert_equal_i(9, utf8_strwidth(\"hello世界\"));  /* 5 ASCII (5 cols) + 2 Chinese (4 cols) = 9 */\n+\tcl_assert_equal_i(7, utf8_strwidth(\"hi世界!\"));   /* 2 ASCII (2 cols) + 2 Chinese (4 cols) + 1 ASCII (1 col) = 7 */\n+}\n+\n+/*\n+ * Additional tests with other East Asian characters\n+ */\n+void test_utf8_width__strnwidth_japanese_korean(void)\n+{\n+\t/* Japanese characters (should also be 2 columns each) */\n+\tcl_assert_equal_i(10, utf8_strnwidth(\"こんにちは\", 15, 0));  /* 5 Japanese chars @ 2 cols each = 10 display columns */\n+\n+\t/* Korean characters (should also be 2 columns each) */\n+\tcl_assert_equal_i(10, utf8_strnwidth(\"안녕하세요\", 15, 0));  /* 5 Korean chars @ 2 cols each = 10 display columns */\n+}\n+\n+/*\n+ * Test edge cases with partial UTF-8 sequences\n+ */\n+void test_utf8_width__strnwidth_edge_cases(void)\n+{\n+\tconst char *invalid;\n+\tunsigned char truncated_bytes[] = {0xe4, 0xbd, 0x00};  /* First 2 bytes of \"中\" + null */\n+\n+\t/* Test invalid UTF-8 - should fall back to byte count */\n+\tinvalid = \"\\xff\\xfe\";  /* Invalid UTF-8 sequence */\n+\tcl_assert_equal_i(2, utf8_strnwidth(invalid, 2, 0));  /* Should return length if invalid UTF-8 */\n+\n+\t/* Test partial UTF-8 character (truncated) */\n+\tcl_assert_equal_i(2, utf8_strnwidth((const char*)truncated_bytes, 2, 0));  /* Invalid UTF-8, returns byte count */\n+}\n-- \n2.52.0.rc2.5.g4c20a63325.dirty\n\n"},{"id":"530680","messageId":"a50bcde6446fbd87b4fb04b28c579a915457813a.1763098804.git.worldhello.net@gmail.com","threadId":"64478","inReplyTo":"cover.1763098804.git.worldhello.net@gmail.com","subject":"[PATCH 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-14T05:52:45Z","receivedAt":"2025-11-14T05:52:57Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"The output table from \"git repo structure\" is misaligned when displaying\nUTF-8 characters (e.g., non-ASCII glyphs). E.g.:\n\n    | 仓库结构   | 值  |\n    | -------------- | ---- |\n    | * 引用       |      |\n    |   * 计数     |   67 |\n    |     * 分支   |    6 |\n    |     * 标签   |   30 |\n    |     * 远程   |   19 |\n    |     * 其它   |   12 |\n    |                |      |\n    | * 可达对象 |      |\n    |   * 计数     | 2217 |\n    |     * 提交   |  279 |\n    |     * 树      |  740 |\n    |     * 数据对象 | 1168 |\n    |     * 标签   |   30 |\n\nThe previous implementation used simple width formatting with printf()\nwhich didn't properly handle multi-byte UTF-8 characters, causing\nmisaligned table columns when displaying repository structure\ninformation.\n\nThis change modifies the stats_table_print_structure function to use\nstrbuf_utf8_align() instead of basic printf width specifiers. This\nensures proper column alignment regardless of the character encoding of\nthe content being displayed.\n\nCo-developed-by: Gemini <noreply@developers.google.com>\nSigned-off-by: Jiang Xin <worldhello.net@gmail.com>\n---\n builtin/repo.c | 22 ++++++++++++++++++----\n 1 file changed, 18 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 9d4749f79b..d0b4a060b1 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -292,14 +292,21 @@ static void stats_table_print_structure(const struct stats_table *table)\n \tint name_col_width = utf8_strwidth(name_col_title);\n \tint value_col_width = utf8_strwidth(value_col_title);\n \tstruct string_list_item *item;\n+\tstruct strbuf buf = STRBUF_INIT;\n \n \tif (table->name_col_width > name_col_width)\n \t\tname_col_width = table->name_col_width;\n \tif (table->value_col_width > value_col_width)\n \t\tvalue_col_width = table->value_col_width;\n \n-\tprintf(\"| %-*s | %-*s |\\n\", name_col_width, name_col_title,\n-\t       value_col_width, value_col_title);\n+\tstrbuf_addstr(&buf, \"| \");\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n+\tstrbuf_addstr(&buf, \" | \");\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT, value_col_width, value_col_title);\n+\tstrbuf_addstr(&buf, \" |\");\n+\tprintf(\"%s\\n\", buf.buf);\n+\tstrbuf_reset(&buf);\n+\n \tprintf(\"| \");\n \tfor (int i = 0; i < name_col_width; i++)\n \t\tputchar('-');\n@@ -317,9 +324,16 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\t\tvalue = entry->value;\n \t\t}\n \n-\t\tprintf(\"| %-*s | %*s |\\n\", name_col_width, item->string,\n-\t\t       value_col_width, value);\n+\t\tstrbuf_reset(&buf);\n+\t\tstrbuf_addstr(&buf, \"| \");\n+\t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n+\t\tstrbuf_addstr(&buf, \" | \");\n+\t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n+\t\tstrbuf_addstr(&buf, \" |\");\n+\t\tprintf(\"%s\\n\", buf.buf);\n \t}\n+\n+\tstrbuf_release(&buf);\n }\n \n static void stats_table_clear(struct stats_table *table)\n-- \n2.52.0.rc2.5.g4c20a63325.dirty\n\n"},{"id":"530686","messageId":"8cb5d668-783f-4400-89b4-35054a6cbea0@app.fastmail.com","threadId":"64478","inReplyTo":"cover.1763098804.git.worldhello.net@gmail.com","subject":"Re: [PATCH 0/2] Fix misaligned output of git repo structure","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2025-11-14T07:41:13Z","receivedAt":"2025-11-14T07:41:35Z","isPatch":true,"sender":{"key":"kristofferhaugsbakk@fastmail.com","avatar":null},"body":"On Fri, Nov 14, 2025, at 06:52, Jiang Xin wrote:\n> While localizing Git 2.52.0, I noticed that the output table from git\n> repo structure becomes misaligned when displaying UTF-8 characters. For\n> example:\n>\n>[snip]\n>\n> BTW, I used two AI coding tools (Claude Code and Gemini-CLI) to generate\n> the commits, and added the \"Co-developed-by\" trailers in the commit\n> messages by using one of my opensource project:\n\nIs `Co-developed-by` supposed to have a different meaning than the more\ncommon `Co-authored-by`?\n\nhttps://lore.kernel.org/git/xmqq1pq7re7q.fsf@gitster.g/\n\n>\n>  - https://github.com/ai-coding-workshop/commit-msg\n>\n>\n> ## Changes\n>\n>[snip]\n"},{"id":"530687","messageId":"CANYiYbGyGKy=S6a3NJFyrv-bOZos+BXdR=nPXDT3W_dGxeiNPA@mail.gmail.com","threadId":"64478","inReplyTo":"8cb5d668-783f-4400-89b4-35054a6cbea0@app.fastmail.com","subject":"Re: [PATCH 0/2] Fix misaligned output of git repo structure","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-14T09:52:34Z","receivedAt":"2025-11-14T09:52:47Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Fri, Nov 14, 2025 at 3:41 PM Kristoffer Haugsbakk\n<kristofferhaugsbakk@fastmail.com> wrote:\n>\n> On Fri, Nov 14, 2025, at 06:52, Jiang Xin wrote:\n> > While localizing Git 2.52.0, I noticed that the output table from git\n> > repo structure becomes misaligned when displaying UTF-8 characters. For\n> > example:\n> >\n> >[snip]\n> >\n> > BTW, I used two AI coding tools (Claude Code and Gemini-CLI) to generate\n> > the commits, and added the \"Co-developed-by\" trailers in the commit\n> > messages by using one of my opensource project:\n>\n> Is `Co-developed-by` supposed to have a different meaning than the more\n> common `Co-authored-by`?\n\nThis is a very good question.\n\n**Background**\n\nAt Alibaba Cloud, our development team uses a variety of AI coding tools,\nincluding Cursor, Claude Code, Gemini-CLI, Lingma, and Qoder, etc. To\nmeasure adoption—specifically, how many developers are using AI coding\ntools and how much code is AI-generated—we needed a unified tracking\nmechanism compatible with all these tools. I chose to implement a git\ncommit-msg hook that automatically detects the AI coding tool responsible\nfor a commit based on environment variables at commit time.\n\n**Why choose the Co-developed-by trailer for AI developer?**\n\nGit repositories already use the Co-authored-by trailer to credit human\ncollaborators. Since any human developer, including co-authors, may use\nAI coding tools to assist their work, introducing a distinct trailer\nlike Co-developed-by allows us to clearly differentiate between human\ncontributors and the AI tools they used. For example, the following\ncommit trailers indicate two human engineers and the respective AI\ncoding tools they employed:\n\n    Co-developed-by: Cursor <noreply@cursor.com>\n    Co-authored-by: Real Person <real.person@example.com>\n    Co-developed-by: Gemini <noreply@developers.google.com>\n    Signed-off-by: Me <me@example.com>\n\nI noticed that Sasha Levin (NVIDIA) previously proposed adopting\nthe Co-developed-by trailer for the Linux kernel as well.\n\n - https://ostechnix.com/linux-kernel-ai-coding-assistants-rules-proposal/\n\n--\nJiang Xin\n"},{"id":"530691","messageId":"xmqqms4ok2xx.fsf@gitster.g","threadId":"64478","inReplyTo":"cover.1763098804.git.worldhello.net@gmail.com","subject":"Re: [PATCH 0/2] Fix misaligned output of git repo structure","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-14T16:13:14Z","receivedAt":"2025-11-14T16:13:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jiang Xin <worldhello.net@gmail.com> writes:\n\n> BTW, I used two AI coding tools (Claude Code and Gemini-CLI) to generate\n> the commits, and added the \"Co-developed-by\" trailers in the commit\n> messages by using one of my opensource project:\n\nWe had a mini-thread on this recently.\n\n  https://lore.kernel.org/git/xmqqo6p9zo8f.fsf@gitster.g/\n\n"},{"id":"530695","messageId":"wgxzx47nsro3h6ju3t2aatrygkr5g7i2dbl26fj53qh4f7jdxw@d233r7jflrke","threadId":"64478","inReplyTo":"a50bcde6446fbd87b4fb04b28c579a915457813a.1763098804.git.worldhello.net@gmail.com","subject":"Re: [PATCH 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2025-11-14T17:50:32Z","receivedAt":"2025-11-14T17:50:37Z","isPatch":true,"sender":{"key":"jltobler@gmail.com","avatar":"https://avatars.githubusercontent.com/u/53454972?v=4"},"body":"On 25/11/14 12:52AM, Jiang Xin wrote:\n> The output table from \"git repo structure\" is misaligned when displaying\n> UTF-8 characters (e.g., non-ASCII glyphs). E.g.:\n> \n>     | 仓库结构   | 值  |\n>     | -------------- | ---- |\n>     | * 引用       |      |\n>     |   * 计数     |   67 |\n>     |     * 分支   |    6 |\n>     |     * 标签   |   30 |\n>     |     * 远程   |   19 |\n>     |     * 其它   |   12 |\n>     |                |      |\n>     | * 可达对象 |      |\n>     |   * 计数     | 2217 |\n>     |     * 提交   |  279 |\n>     |     * 树      |  740 |\n>     |     * 数据对象 | 1168 |\n>     |     * 标签   |   30 |\n> \n> The previous implementation used simple width formatting with printf()\n> which didn't properly handle multi-byte UTF-8 characters, causing\n> misaligned table columns when displaying repository structure\n> information.\n\nThanks for finding this issue and submitting a fix! I failed to consider\nthe fact that the printf() format specifier width would be counting\nbytes. This causes the overall line width to fall short in some\nscenarios with multi-byte UTF-8 characters.\n\n> This change modifies the stats_table_print_structure function to use\n> strbuf_utf8_align() instead of basic printf width specifiers. This\n> ensures proper column alignment regardless of the character encoding of\n> the content being displayed.\n\nMakes sense.\n\n> Co-developed-by: Gemini <noreply@developers.google.com>\n> Signed-off-by: Jiang Xin <worldhello.net@gmail.com>\n> ---\n>  builtin/repo.c | 22 ++++++++++++++++++----\n>  1 file changed, 18 insertions(+), 4 deletions(-)\n> \n> diff --git a/builtin/repo.c b/builtin/repo.c\n> index 9d4749f79b..d0b4a060b1 100644\n> --- a/builtin/repo.c\n> +++ b/builtin/repo.c\n> @@ -292,14 +292,21 @@ static void stats_table_print_structure(const struct stats_table *table)\n>  \tint name_col_width = utf8_strwidth(name_col_title);\n>  \tint value_col_width = utf8_strwidth(value_col_title);\n>  \tstruct string_list_item *item;\n> +\tstruct strbuf buf = STRBUF_INIT;\n>  \n>  \tif (table->name_col_width > name_col_width)\n>  \t\tname_col_width = table->name_col_width;\n>  \tif (table->value_col_width > value_col_width)\n>  \t\tvalue_col_width = table->value_col_width;\n>  \n> -\tprintf(\"| %-*s | %-*s |\\n\", name_col_width, name_col_title,\n> -\t       value_col_width, value_col_title);\n> +\tstrbuf_addstr(&buf, \"| \");\n> +\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n> +\tstrbuf_addstr(&buf, \" | \");\n> +\tstrbuf_utf8_align(&buf, ALIGN_LEFT, value_col_width, value_col_title);\n> +\tstrbuf_addstr(&buf, \" |\");\n> +\tprintf(\"%s\\n\", buf.buf);\n\nOk, using strbuf_utf8_align() compensates the line width when using\nmulti-byte UTF-8 characters to ensure the correct length. Looks good.\n\n> +\tstrbuf_reset(&buf);\n\nDo we need to reset the buffer here? In the following loop we reset it\nat the start of each iteration.\n\n> +\n>  \tprintf(\"| \");\n>  \tfor (int i = 0; i < name_col_width; i++)\n>  \t\tputchar('-');\n> @@ -317,9 +324,16 @@ static void stats_table_print_structure(const struct stats_table *table)\n>  \t\t\tvalue = entry->value;\n>  \t\t}\n>  \n> -\t\tprintf(\"| %-*s | %*s |\\n\", name_col_width, item->string,\n> -\t\t       value_col_width, value);\n> +\t\tstrbuf_reset(&buf);\n> +\t\tstrbuf_addstr(&buf, \"| \");\n> +\t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n> +\t\tstrbuf_addstr(&buf, \" | \");\n> +\t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n> +\t\tstrbuf_addstr(&buf, \" |\");\n> +\t\tprintf(\"%s\\n\", buf.buf);\n\nHere we do the same thing for the values column. Looks good to me.\n\nThanks,\n-Justin\n"},{"id":"530696","messageId":"xmqqecq0ifld.fsf@gitster.g","threadId":"64478","inReplyTo":"CANYiYbGyGKy=S6a3NJFyrv-bOZos+BXdR=nPXDT3W_dGxeiNPA@mail.gmail.com","subject":"Re: [PATCH 0/2] Fix misaligned output of git repo structure","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-14T19:22:54Z","receivedAt":"2025-11-14T19:22:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jiang Xin <worldhello.net@gmail.com> writes:\n\n>> Is `Co-developed-by` supposed to have a different meaning than the more\n>> common `Co-authored-by`?\n>\n> This is a very good question.\n>\n> **Background**\n>\n> At Alibaba Cloud, our development team uses a variety of AI coding tools,\n> including Cursor, Claude Code, Gemini-CLI, Lingma, and Qoder, etc. To\n> measure adoption—specifically, how many developers are using AI coding\n> tools and how much code is AI-generated—we needed a unified tracking\n> mechanism compatible with all these tools. I chose to implement a git\n> commit-msg hook that automatically detects the AI coding tool responsible\n> for a commit based on environment variables at commit time.\n\nIn other words, addition of this is solely to help corporations like\nAlibaba to measure which AI tools are used (and what correlation\nthere are between success rate of the patches and the tools that\ngenerated them, etc..\n\nWhat is in it for us?  What benefit are we getting in exchange for\ntolerating these additional trailer lines in our log messages?\n\nA few random thoughts about generated contents:\n\n * Disclosing the tools that were used during the development of a\n   patch is a good practice in principle, but this is not limited to\n   use of AI tools.  We have fixes for issues found with existing\n   Coccinelle checks, sanitizers, static checkers, and it is the\n   usual practice for the patches that fix them to disclose how the\n   author discovered the issue.  When making mechanical replacement\n   changes en masse, it is the usual practice for the patches to\n   describe what scripts were used to make the changes in them.  But\n   we do not dedicate a trailer line for such a disclosure, and\n   there is no reason why AI tools has to be treated specially here.\n   Instead of \"Co-developed-by\" that only tells what tool was used,\n   why not disclose what prompts (again, somehow AI tools are\n   treated specially here, too---we call the input to these tools\n   \"scripts\" when the changes were made with sed or perl or\n   coccinelle) were used?\n\n * Whether some or all contents in a submitted patch were generated\n   by tools, it does not change the obligation of the person who\n   submits the patch.  They need to make sure that the changes are\n   reviewable, its goal and implementation are described in the\n   proposed log message appropriately, the updated code does what\n   the proposed log message claims to do.  They need to make sure\n   that they have the right to contribute the patch under DCO, and\n   sign off their patch accordingly.\n\n * What is made more difficult for a submitter with AI tools is that\n   it is often not obvious to the human developer how much of the\n   tools' generated output is parroting what the tools saw during\n   their training session, and what the licensing terms of these\n   training materials are.  Even if a hypothetical AI tool were\n   trained only with BSD licensed material, the output from such a\n   tool is likely to hold you under certain obligations like\n   including the original copyright notice, but without the tool\n   disclosing to you the human developer, you do not even know whose\n   copyright notice to include.\n\n * Worse yet, the above difficulty is only for the submitter of such\n   a patch, not the project that, trusting what the sign-off of the\n   submitter certifies, reviews and accepts such a patch.  It does\n   not make any difference if the original submitter copied and\n   pasted proprietary code of their employer in the patch, or\n   included code that AI tools \"borrowed\" from elsewhere without\n   following proper procedure to honor the licensing terms.  In\n   either case, the project may have accepted what was stolen\n   without knowing, and it is very likely that the submitter but not\n   the project is primarily held liable.  In a sense, the project\n   would be better off if the patch does not say it was generated\n   with AI tools---if the project does not know, it cannot possibly\n   held liable for it, even though the project will have to waste\n   engineering resources to rewrite or remove the remnant from such\n   a faulty contribution.\n"},{"id":"530697","messageId":"xmqqa50oiduy.fsf@gitster.g","threadId":"64478","inReplyTo":"a50bcde6446fbd87b4fb04b28c579a915457813a.1763098804.git.worldhello.net@gmail.com","subject":"Re: [PATCH 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-14T20:00:21Z","receivedAt":"2025-11-14T20:00:25Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jiang Xin <worldhello.net@gmail.com> writes:\n\nNot about the contents of the patch, but how was the list of\naddresses on CC produced?  Do they all have enough stakes in the\ncode being updated that they do not mind getting spammed like this?\n\nAlso, you had a non-address \"Gemini <noreply@developers.google.com>\",\nwhich forced me and anybody who will respond to the patch edit Cc\naddress list (or suffer bounces).  Please don't.\n\n> The output table from \"git repo structure\" is misaligned when displaying\n> UTF-8 characters (e.g., non-ASCII glyphs). E.g.:\n>\n>     | 仓库结构   | 值  |\n>     | -------------- | ---- |\n>     | * 引用       |      |\n>     |   * 计数     |   67 |\n>     |     * 分支   |    6 |\n>     |     * 标签   |   30 |\n>     |     * 远程   |   19 |\n>     |     * 其它   |   12 |\n>     |                |      |\n>     | * 可达对象 |      |\n>     |   * 计数     | 2217 |\n>     |     * 提交   |  279 |\n>     |     * 树      |  740 |\n>     |     * 数据对象 | 1168 |\n>     |     * 标签   |   30 |\n\nAs there is a concrete reproduction sample from a specific tool, ...\n\n>  builtin/repo.c | 22 ++++++++++++++++++----\n>  1 file changed, 18 insertions(+), 4 deletions(-)\n\n... it is a good idea to protect the change with a new test or two\nto make sure the expected alignment in the output.\n\nThanks.\n"},{"id":"530701","messageId":"xmqqzf8ogyhw.fsf@gitster.g","threadId":"64478","inReplyTo":"04ab347ff80e16d49524246a8923cc86cc7355be.1763098804.git.worldhello.net@gmail.com","subject":"Re: [PATCH 1/2] t/unit-tests: add UTF-8 width tests for CJK chars","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-14T20:17:31Z","receivedAt":"2025-11-14T20:17:34Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jiang Xin <worldhello.net@gmail.com> writes:\n\n[jc: the same question about the choice of Cc addresses applies]\n\n> This commit adds a new test suite (u-utf8-width.c) to test the UTF-8\n> width functions in Git, particularly focusing on multi-byte characters\n> from East Asian languages like Chinese, Japanese, and Korean that\n> typically require 2 display columns per character.\n>\n> The test suite includes:\n> - Tests for utf8_strnwidth with Chinese strings\n> - Tests for utf8_strwidth with Chinese strings\n> - Tests for Japanese and Korean characters\n> - Edge case tests with invalid UTF-8 sequences\n> - Proper test function naming following the Clar framework convention\n>\n> Also updated the build configuration in Makefile and meson.build to\n> include the new test suite in the build process.\n\nThe usual way to compose a log message of this project is to\n\n - Give an observation on how the current system works in the\n   present tense (so no need to say \"Currently X is Y\", or\n   \"Previously X was Y\" to describe the state before your change;\n   just \"X is Y\" is enough), and discuss what you perceive as a\n   problem in it.\n\n - Propose a solution (optional---often, problem description\n   trivially leads to an obvious solution in reader's minds).\n\n - Give commands to somebody editing the codebase to \"make it so\",\n   instead of saying \"This commit does X\".\n\nin this order.\n\n> +\t/* Test length limiting */\n> +\tstr = \"你好世界\";\n> +\tcl_assert_equal_i(2, utf8_strnwidth(str, 3, 0));  /* Only first char \"你\"(2 columns) within 3 bytes */\n> +\tcl_assert_equal_i(4, utf8_strnwidth(str, 6, 0));  /* First two chars \"你好\"(4 columns) in 6 bytes */\n\nWe also should test utf8_strwidth() on the same string here.\n\n> +/*\n> + * Test edge cases with partial UTF-8 sequences\n> + */\n\nAll tests before these make sense, but I am not sure if we want to\nhold utf8_strnwidth() to the requirement that it will tolerate \"len\"\nto end in the middle of a single character, as such a requirement by\nitself does not do application any good.\n\nA caller may have \"你好世界\" in str, learn that the first 4 bytes\nwould only need two display columns to show (i.e., 3-byte \"你\" plus\na single garbage byte, that would make UTF-8 encoded \"好\" if the\nremaining two bytes were included), and may want to learn how to\nshow only enough to fill the two display columns.  But there is not\nenough information given back by utf8_strnwidth() for such a caller\nto figure out that it needs to feed only the first three bytes (not\nfour) of str to printf() to do so.\n"},{"id":"530733","messageId":"CANYiYbGqLyZ9zvhR73z91yo2Yk-tT1LVcP--uEdaKL1hHJOEHA@mail.gmail.com","threadId":"64478","inReplyTo":"xmqqecq0ifld.fsf@gitster.g","subject":"Re: [PATCH 0/2] Fix misaligned output of git repo structure","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-15T12:25:51Z","receivedAt":"2025-11-15T12:26:03Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Sat, Nov 15, 2025 at 3:22 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Jiang Xin <worldhello.net@gmail.com> writes:\n>\n> >> Is `Co-developed-by` supposed to have a different meaning than the more\n> >> common `Co-authored-by`?\n> >\n> > This is a very good question.\n> >\n> > **Background**\n> >\n> > At Alibaba Cloud, our development team uses a variety of AI coding tools,\n> > including Cursor, Claude Code, Gemini-CLI, Lingma, and Qoder, etc. To\n> > measure adoption—specifically, how many developers are using AI coding\n> > tools and how much code is AI-generated—we needed a unified tracking\n> > mechanism compatible with all these tools. I chose to implement a git\n> > commit-msg hook that automatically detects the AI coding tool responsible\n> > for a commit based on environment variables at commit time.\n>\n> In other words, addition of this is solely to help corporations like\n> Alibaba to measure which AI tools are used (and what correlation\n> there are between success rate of the patches and the tools that\n> generated them, etc..\n>\n> What is in it for us?  What benefit are we getting in exchange for\n> tolerating these additional trailer lines in our log messages?\n>\n> A few random thoughts about generated contents:\n\nRegardless of whether the trailer used in commits to identify AI\ncoding tools was leaked intentionally or unintentionally, the\nfollowing insights are extremely valuable—thank you!\n\nI’ll add the appropriate configuration to disable the commit-msg\nhook’s automatic modification of commit messages for git.git\nrepositories, and I’ll review any AI-generated code more carefully,\nif present.\n\n--\nJiang Xin\n"},{"id":"530735","messageId":"CANYiYbEhGo2Z1Y+YaXfOs35+2nTOLYW4C=yXR65b4wGfe2nFgw@mail.gmail.com","threadId":"64478","inReplyTo":"xmqqzf8ogyhw.fsf@gitster.g","subject":"Re: [PATCH 1/2] t/unit-tests: add UTF-8 width tests for CJK chars","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-15T12:38:53Z","receivedAt":"2025-11-15T12:39:06Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Sat, Nov 15, 2025 at 4:17 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Jiang Xin <worldhello.net@gmail.com> writes:\n>\n> [jc: the same question about the choice of Cc addresses applies]\n>\n> > This commit adds a new test suite (u-utf8-width.c) to test the UTF-8\n> > width functions in Git, particularly focusing on multi-byte characters\n> > from East Asian languages like Chinese, Japanese, and Korean that\n> > typically require 2 display columns per character.\n> >\n> > The test suite includes:\n> > - Tests for utf8_strnwidth with Chinese strings\n> > - Tests for utf8_strwidth with Chinese strings\n> > - Tests for Japanese and Korean characters\n> > - Edge case tests with invalid UTF-8 sequences\n> > - Proper test function naming following the Clar framework convention\n> >\n> > Also updated the build configuration in Makefile and meson.build to\n> > include the new test suite in the build process.\n>\n> The usual way to compose a log message of this project is to\n>\n>  - Give an observation on how the current system works in the\n>    present tense (so no need to say \"Currently X is Y\", or\n>    \"Previously X was Y\" to describe the state before your change;\n>    just \"X is Y\" is enough), and discuss what you perceive as a\n>    problem in it.\n>\n>  - Propose a solution (optional---often, problem description\n>    trivially leads to an obvious solution in reader's minds).\n>\n>  - Give commands to somebody editing the codebase to \"make it so\",\n>    instead of saying \"This commit does X\".\n>\n> in this order.\n\nWill document the purpose in commit message of next reroll.\n\n> > +/*\n> > + * Test edge cases with partial UTF-8 sequences\n> > + */\n>\n> All tests before these make sense, but I am not sure if we want to\n> hold utf8_strnwidth() to the requirement that it will tolerate \"len\"\n> to end in the middle of a single character, as such a requirement by\n> itself does not do application any good.\n\nWill remove unnecessary test cases.\n"},{"id":"530736","messageId":"CANYiYbFjShuKULNeyQBKmAb07gwEt5PZp8S_62v9E=eVWGr9-w@mail.gmail.com","threadId":"64478","inReplyTo":"wgxzx47nsro3h6ju3t2aatrygkr5g7i2dbl26fj53qh4f7jdxw@d233r7jflrke","subject":"Re: [PATCH 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-15T12:41:49Z","receivedAt":"2025-11-15T12:42:01Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Sat, Nov 15, 2025 at 1:50 AM Justin Tobler <jltobler@gmail.com> wrote:\n> Ok, using strbuf_utf8_align() compensates the line width when using\n> multi-byte UTF-8 characters to ensure the correct length. Looks good.\n>\n> > +     strbuf_reset(&buf);\n>\n> Do we need to reset the buffer here? In the following loop we reset it\n> at the start of each iteration.\n\nWill remove this line in next reroll.\n\n--\nJiang Xin\n"},{"id":"530737","messageId":"CANYiYbEFN9BHtNh1PQ9C3gDJasq1PaKnkcH-Nq=FddUCAcMGqg@mail.gmail.com","threadId":"64478","inReplyTo":"xmqqa50oiduy.fsf@gitster.g","subject":"Re: [PATCH 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-15T12:54:16Z","receivedAt":"2025-11-15T12:54:28Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Sat, Nov 15, 2025 at 4:00 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Jiang Xin <worldhello.net@gmail.com> writes:\n>\n> Not about the contents of the patch, but how was the list of\n> addresses on CC produced?  Do they all have enough stakes in the\n> code being updated that they do not mind getting spammed like this?\n\nI’m cc’ing this patch series to all Git l10n team leads to inform them\nthat the issue has been identified and will be fixed.\n\n> Also, you had a non-address \"Gemini <noreply@developers.google.com>\",\n> which forced me and anybody who will respond to the patch edit Cc\n> address list (or suffer bounces).  Please don't.\n\nWill remove this trailer.\n\n> >     |     * 提交   |  279 |\n> >     |     * 树      |  740 |\n> >     |     * 数据对象 | 1168 |\n> >     |     * 标签   |   30 |\n>\n> As there is a concrete reproduction sample from a specific tool, ...\n\nThis output of the `git repo structure` command is based on the\nChinese translation for Git 2.52. The next reroll will retain only the\ntable header, which is sufficient to demonstrate the issue.\n\n>\n> >  builtin/repo.c | 22 ++++++++++++++++++----\n> >  1 file changed, 18 insertions(+), 4 deletions(-)\n>\n> ... it is a good idea to protect the change with a new test or two\n> to make sure the expected alignment in the output.\n\nWill add test cases for strbuf_utf8_align(), a function newly\nintroduced in builtin/repo.c.\n"},{"id":"530738","messageId":"cover.1763213290.git.worldhello.net@gmail.com","threadId":"64478","inReplyTo":"cover.1763098804.git.worldhello.net@gmail.com","subject":"[PATCH v2 0/2] Fix misaligned output of git repo structure","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-15T13:36:09Z","receivedAt":"2025-11-15T13:36:20Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"While localizing Git 2.52.0, I noticed that the output table from git\nrepo structure becomes misaligned when displaying UTF-8 characters. For\nexample:\n\n    | 仓库结构   | 值  |\n    | -------------- | ---- |\n\nThe previous implementation used simple width formatting with printf()\nwhich didn't properly handle multi-byte UTF-8 characters, causing\nmisaligned table columns when displaying repository structure\ninformation.\n\nThis change modifies the stats_table_print_structure function to use\nstrbuf_utf8_align() instead of basic printf width specifiers. This\nensures proper column alignment regardless of the character encoding of\nthe content being displayed.\n\nJiang Xin (2):\n  t/unit-tests: add UTF-8 width tests for CJK chars\n  builtin/repo: fix table alignment for UTF-8 characters\n\n Makefile                    |   1 +\n builtin/repo.c              |  21 ++++--\n t/meson.build               |   1 +\n t/unit-tests/u-utf8-width.c | 134 ++++++++++++++++++++++++++++++++++++\n 4 files changed, 153 insertions(+), 4 deletions(-)\n create mode 100644 t/unit-tests/u-utf8-width.c\n\n\n## Range-diff vs v1:\n\n1:  53c1e5219b ! 1:  72e73484d2 t/unit-tests: add UTF-8 width tests for CJK chars\n    @@ Metadata\n      ## Commit message ##\n         t/unit-tests: add UTF-8 width tests for CJK chars\n     \n    -    This commit adds a new test suite (u-utf8-width.c) to test the UTF-8\n    -    width functions in Git, particularly focusing on multi-byte characters\n    -    from East Asian languages like Chinese, Japanese, and Korean that\n    -    typically require 2 display columns per character.\n    -\n    -    The test suite includes:\n    -    - Tests for utf8_strnwidth with Chinese strings\n    -    - Tests for utf8_strwidth with Chinese strings\n    -    - Tests for Japanese and Korean characters\n    -    - Edge case tests with invalid UTF-8 sequences\n    -    - Proper test function naming following the Clar framework convention\n    +    The file \"builtin/repo.c\" uses utf8_strwidth() to calculate the display\n    +    width of UTF-8 characters in a table, but the resulting output is still\n    +    misaligned. Add test cases for both utf8_strwidth and utf8_strnwidth to\n    +    verify that they correctly compute the display width for UTF-8\n    +    characters.\n     \n         Also updated the build configuration in Makefile and meson.build to\n         include the new test suite in the build process.\n     \n    -    Co-developed-by: Claude <noreply@anthropic.com>\n         Signed-off-by: Jiang Xin <worldhello.net@gmail.com>\n     \n      ## Makefile ##\n    @@ t/unit-tests/u-utf8-width.c (new)\n     + */\n     +void test_utf8_width__strnwidth_chinese(void)\n     +{\n    -+\tconst char *ansi_test;\n     +\tconst char *str;\n     +\n     +\t/* Test basic ASCII - each character should have width 1 */\n    -+\tcl_assert_equal_i(5, utf8_strnwidth(\"hello\", 5, 0));\n    -+\tcl_assert_equal_i(5, utf8_strnwidth(\"hello\", 5, 1));  /* skip_ansi = 1 */\n    ++\tcl_assert_equal_i(5, utf8_strnwidth(\"Hello\", 5, 0));\n    ++\t/* skip_ansi = 1 */\n    ++\tcl_assert_equal_i(5, utf8_strnwidth(\"Hello\", 5, 1));\n     +\n     +\t/* Test simple Chinese characters - each should have width 2 */\n    -+\tcl_assert_equal_i(4, utf8_strnwidth(\"你好\", 6, 0));  /* \"你好\" is 6 bytes (3 bytes per char in UTF-8), 4 display columns */\n    ++\t/* \"你好\" is 6 bytes (3 bytes per char in UTF-8), 4 display columns */\n    ++\tcl_assert_equal_i(4, utf8_strnwidth(\"你好\", 6, 0));\n     +\n     +\t/* Test mixed ASCII and Chinese - ASCII = 1 column, Chinese = 2 columns */\n    -+\tcl_assert_equal_i(6, utf8_strnwidth(\"hi你好\", 8, 0));  /* \"h\"(1) + \"i\"(1) + \"你\"(2) + \"好\"(2) = 6 */\n    ++\t/* \"h\"(1) + \"i\"(1) + \"你\"(2) + \"好\"(2) = 6 */\n    ++\tcl_assert_equal_i(6, utf8_strnwidth(\"Hi你好\", 8, 0));\n     +\n     +\t/* Test longer Chinese string */\n    -+\tcl_assert_equal_i(10, utf8_strnwidth(\"你好世界！\", 15, 0));  /* 5 Chinese chars = 10 display columns */\n    -+\n    -+\t/* Test with skip_ansi = 1 to make sure it works with escape sequences */\n    -+\tansi_test = \"\\033[31m你好\\033[0m\";\n    -+\tcl_assert_equal_i(4, utf8_strnwidth(ansi_test, strlen(ansi_test), 1));  /* Skip escape sequences, just count \"你好\" which should be 4 columns */\n    ++\t/* 5 Chinese chars = 10 display columns */\n    ++\tcl_assert_equal_i(10, utf8_strnwidth(\"你好世界！\", 15, 0));\n     +\n     +\t/* Test individual Chinese character width */\n    -+\tcl_assert_equal_i(2, utf8_strnwidth(\"中\", 3, 0));  /* Single Chinese char should be 2 columns */\n    ++\tcl_assert_equal_i(2, utf8_strnwidth(\"中\", 3, 0));\n     +\n     +\t/* Test empty string */\n     +\tcl_assert_equal_i(0, utf8_strnwidth(\"\", 0, 0));\n     +\n     +\t/* Test length limiting */\n     +\tstr = \"你好世界\";\n    -+\tcl_assert_equal_i(2, utf8_strnwidth(str, 3, 0));  /* Only first char \"你\"(2 columns) within 3 bytes */\n    -+\tcl_assert_equal_i(4, utf8_strnwidth(str, 6, 0));  /* First two chars \"你好\"(4 columns) in 6 bytes */\n    ++\t/* Only first char \"你\"(2 columns) within 3 bytes */\n    ++\tcl_assert_equal_i(2, utf8_strnwidth(str, 3, 0));\n    ++\t/* First two chars \"你好\"(4 columns) in 6 bytes */\n    ++\tcl_assert_equal_i(4, utf8_strnwidth(str, 6, 0));\n     +}\n     +\n     +/*\n    @@ t/unit-tests/u-utf8-width.c (new)\n     +void test_utf8_width__strwidth_chinese(void)\n     +{\n     +\t/* Test basic ASCII */\n    -+\tcl_assert_equal_i(5, utf8_strwidth(\"hello\"));\n    ++\tcl_assert_equal_i(5, utf8_strwidth(\"Hello\"));\n     +\n     +\t/* Test Chinese characters */\n    -+\tcl_assert_equal_i(4, utf8_strwidth(\"你好\"));  /* 2 Chinese chars = 4 display columns */\n    ++\t/* 2 Chinese chars = 4 display columns */\n    ++\tcl_assert_equal_i(4, utf8_strwidth(\"你好\"));\n    ++\n    ++\t/* Test longer Chinese string */\n    ++\t/* 5 Chinese chars = 10 display columns */\n    ++\tcl_assert_equal_i(10, utf8_strwidth(\"你好世界！\"));\n     +\n     +\t/* Test mixed ASCII and Chinese */\n    -+\tcl_assert_equal_i(9, utf8_strwidth(\"hello世界\"));  /* 5 ASCII (5 cols) + 2 Chinese (4 cols) = 9 */\n    -+\tcl_assert_equal_i(7, utf8_strwidth(\"hi世界!\"));   /* 2 ASCII (2 cols) + 2 Chinese (4 cols) + 1 ASCII (1 col) = 7 */\n    ++\t/* 5 ASCII (5 cols) + 2 Chinese (4 cols) = 9 */\n    ++\tcl_assert_equal_i(9, utf8_strwidth(\"Hello世界\"));\n    ++\t/* 2 ASCII (2 cols) + 2 Chinese (4 cols) + 1 ASCII (1 col) = 7 */\n    ++\tcl_assert_equal_i(7, utf8_strwidth(\"Hi世界!\"));\n     +}\n     +\n     +/*\n    @@ t/unit-tests/u-utf8-width.c (new)\n     +void test_utf8_width__strnwidth_japanese_korean(void)\n     +{\n     +\t/* Japanese characters (should also be 2 columns each) */\n    -+\tcl_assert_equal_i(10, utf8_strnwidth(\"こんにちは\", 15, 0));  /* 5 Japanese chars @ 2 cols each = 10 display columns */\n    ++\t/* 5 Japanese chars x 2 cols each = 10 display columns */\n    ++\tcl_assert_equal_i(10, utf8_strnwidth(\"こんにちは\", 15, 0));\n     +\n     +\t/* Korean characters (should also be 2 columns each) */\n    -+\tcl_assert_equal_i(10, utf8_strnwidth(\"안녕하세요\", 15, 0));  /* 5 Korean chars @ 2 cols each = 10 display columns */\n    ++\t/* 5 Korean chars x 2 cols each = 10 display columns */\n    ++\tcl_assert_equal_i(10, utf8_strnwidth(\"안녕하세요\", 15, 0));\n     +}\n     +\n     +/*\n    -+ * Test edge cases with partial UTF-8 sequences\n    ++ * Test utf8_strnwidth with CJK strings and ANSI sequences\n     + */\n    -+void test_utf8_width__strnwidth_edge_cases(void)\n    ++void test_utf8_width__strnwidth_cjk_with_ansi(void)\n     +{\n    -+\tconst char *invalid;\n    -+\tunsigned char truncated_bytes[] = {0xe4, 0xbd, 0x00};  /* First 2 bytes of \"中\" + null */\n    -+\n    -+\t/* Test invalid UTF-8 - should fall back to byte count */\n    -+\tinvalid = \"\\xff\\xfe\";  /* Invalid UTF-8 sequence */\n    -+\tcl_assert_equal_i(2, utf8_strnwidth(invalid, 2, 0));  /* Should return length if invalid UTF-8 */\n    -+\n    -+\t/* Test partial UTF-8 character (truncated) */\n    -+\tcl_assert_equal_i(2, utf8_strnwidth((const char*)truncated_bytes, 2, 0));  /* Invalid UTF-8, returns byte count */\n    ++\t/* Test CJK with ANSI sequences */\n    ++\tconst char *ansi_test = \"\\033[1m你好\\033[0m\";\n    ++\tint width = utf8_strnwidth(ansi_test, strlen(ansi_test), 1);\n    ++\t/* Should skip ANSI sequences and count \"你好\" as 4 columns */\n    ++\tcl_assert_equal_i(4, width);\n    ++\n    ++\t/* Test mixed ASCII, CJK, and ANSI */\n    ++\tansi_test = \"Hello\\033[32m世界\\033[0m!\";\n    ++\twidth = utf8_strnwidth(ansi_test, strlen(ansi_test), 1);\n    ++\t/* \"Hello\"(5) + \"世界\"(4) + \"!\"(1) = 10 */\n    ++\tcl_assert_equal_i(10, width);\n     +}\n2:  65efad527f ! 2:  d0975427c9 builtin/repo: fix table alignment for UTF-8 characters\n    @@ Commit message\n             | -------------- | ---- |\n             | * 引用       |      |\n             |   * 计数     |   67 |\n    -        |     * 分支   |    6 |\n    -        |     * 标签   |   30 |\n    -        |     * 远程   |   19 |\n    -        |     * 其它   |   12 |\n    -        |                |      |\n    -        | * 可达对象 |      |\n    -        |   * 计数     | 2217 |\n    -        |     * 提交   |  279 |\n    -        |     * 树      |  740 |\n    -        |     * 数据对象 | 1168 |\n    -        |     * 标签   |   30 |\n     \n         The previous implementation used simple width formatting with printf()\n         which didn't properly handle multi-byte UTF-8 characters, causing\n    @@ Commit message\n         ensures proper column alignment regardless of the character encoding of\n         the content being displayed.\n     \n    -    Co-developed-by: Gemini <noreply@developers.google.com>\n    +    Also add test cases for strbuf_utf8_align(), a function newly introduced\n    +    in \"builtin/repo.c\".\n    +\n         Signed-off-by: Jiang Xin <worldhello.net@gmail.com>\n     \n      ## builtin/repo.c ##\n    @@ builtin/repo.c: static void stats_table_print_structure(const struct stats_table\n     +\tstrbuf_utf8_align(&buf, ALIGN_LEFT, value_col_width, value_col_title);\n     +\tstrbuf_addstr(&buf, \" |\");\n     +\tprintf(\"%s\\n\", buf.buf);\n    -+\tstrbuf_reset(&buf);\n     +\n      \tprintf(\"| \");\n      \tfor (int i = 0; i < name_col_width; i++)\n    @@ builtin/repo.c: static void stats_table_print_structure(const struct stats_table\n      }\n      \n      static void stats_table_clear(struct stats_table *table)\n    +\n    + ## t/unit-tests/u-utf8-width.c ##\n    +@@ t/unit-tests/u-utf8-width.c: void test_utf8_width__strnwidth_cjk_with_ansi(void)\n    + \t/* \"Hello\"(5) + \"世界\"(4) + \"!\"(1) = 10 */\n    + \tcl_assert_equal_i(10, width);\n    + }\n    ++\n    ++/*\n    ++ * Test the strbuf_utf8_align function with CJK characters\n    ++ */\n    ++void test_utf8_width__strbuf_utf8_align(void)\n    ++{\n    ++\tstruct strbuf buf = STRBUF_INIT;\n    ++\n    ++\t/* Test left alignment with CJK */\n    ++\tstrbuf_utf8_align(&buf, ALIGN_LEFT, 10, \"你好\");\n    ++\t/* Since \"你好\" is 4 display columns, we need 6 more spaces to reach 10 */\n    ++\tcl_assert_equal_s(\"你好      \", buf.buf);\n    ++\tstrbuf_reset(&buf);\n    ++\n    ++\t/* Test right alignment with CJK */\n    ++\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, 8, \"世界\");\n    ++\t/* \"世界\" is 4 display columns, so we need 4 leading spaces */\n    ++\tcl_assert_equal_s(\"    世界\", buf.buf);\n    ++\tstrbuf_reset(&buf);\n    ++\n    ++\t/* Test center alignment with CJK */\n    ++\tstrbuf_utf8_align(&buf, ALIGN_MIDDLE, 10, \"中\");\n    ++\t/* \"中\" is 2 display columns, so (10-2)/2 = 4 spaces on left, 4 on right */\n    ++\tcl_assert_equal_s(\"    中    \", buf.buf);\n    ++\tstrbuf_reset(&buf);\n    ++\n    ++\tstrbuf_utf8_align(&buf, ALIGN_MIDDLE, 5, \"中\");\n    ++\t/* \"中\" is 2 display columns, so (5-2)/2 = 1 spaces on left, 2 on right */\n    ++\tcl_assert_equal_s(\" 中  \", buf.buf);\n    ++\tstrbuf_reset(&buf);\n    ++\n    ++\t/* Test alignment that is smaller than string width */\n    ++\tstrbuf_utf8_align(&buf, ALIGN_LEFT, 2, \"你好\");\n    ++\t/* Since \"你好\" is 4 display columns, it should not be truncated */\n    ++\tcl_assert_equal_s(\"你好\", buf.buf);\n    ++\tstrbuf_release(&buf);\n    ++}\n\n--\nJiang Xin\n"},{"id":"530739","messageId":"72e73484d26442e71eadc992076a9a804acd5582.1763213290.git.worldhello.net@gmail.com","threadId":"64478","inReplyTo":"cover.1763213290.git.worldhello.net@gmail.com","subject":"[PATCH v2 1/2] t/unit-tests: add UTF-8 width tests for CJK chars","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-15T13:36:10Z","receivedAt":"2025-11-15T13:36:21Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"The file \"builtin/repo.c\" uses utf8_strwidth() to calculate the display\nwidth of UTF-8 characters in a table, but the resulting output is still\nmisaligned. Add test cases for both utf8_strwidth and utf8_strnwidth to\nverify that they correctly compute the display width for UTF-8\ncharacters.\n\nAlso updated the build configuration in Makefile and meson.build to\ninclude the new test suite in the build process.\n\nSigned-off-by: Jiang Xin <worldhello.net@gmail.com>\n---\n Makefile                    |  1 +\n t/meson.build               |  1 +\n t/unit-tests/u-utf8-width.c | 97 +++++++++++++++++++++++++++++++++++++\n 3 files changed, 99 insertions(+)\n create mode 100644 t/unit-tests/u-utf8-width.c\n\ndiff --git a/Makefile b/Makefile\nindex 7e0f77e298..2a67546154 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1525,6 +1525,7 @@ CLAR_TEST_SUITES += u-string-list\n CLAR_TEST_SUITES += u-strvec\n CLAR_TEST_SUITES += u-trailer\n CLAR_TEST_SUITES += u-urlmatch-normalization\n+CLAR_TEST_SUITES += u-utf8-width\n CLAR_TEST_PROG = $(UNIT_TEST_BIN)/unit-tests$(X)\n CLAR_TEST_OBJS = $(patsubst %,$(UNIT_TEST_DIR)/%.o,$(CLAR_TEST_SUITES))\n CLAR_TEST_OBJS += $(UNIT_TEST_DIR)/clar/clar.o\ndiff --git a/t/meson.build b/t/meson.build\nindex a5531df415..dc43d69636 100644\n--- a/t/meson.build\n+++ b/t/meson.build\n@@ -24,6 +24,7 @@ clar_test_suites = [\n   'unit-tests/u-strvec.c',\n   'unit-tests/u-trailer.c',\n   'unit-tests/u-urlmatch-normalization.c',\n+  'unit-tests/u-utf8-width.c',\n ]\n \n clar_sources = [\ndiff --git a/t/unit-tests/u-utf8-width.c b/t/unit-tests/u-utf8-width.c\nnew file mode 100644\nindex 0000000000..3766f19726\n--- /dev/null\n+++ b/t/unit-tests/u-utf8-width.c\n@@ -0,0 +1,97 @@\n+#include \"unit-test.h\"\n+#include \"utf8.h\"\n+#include \"strbuf.h\"\n+\n+/*\n+ * Test utf8_strnwidth with various Chinese strings\n+ * Chinese characters typically have a width of 2 columns when displayed\n+ */\n+void test_utf8_width__strnwidth_chinese(void)\n+{\n+\tconst char *str;\n+\n+\t/* Test basic ASCII - each character should have width 1 */\n+\tcl_assert_equal_i(5, utf8_strnwidth(\"Hello\", 5, 0));\n+\t/* skip_ansi = 1 */\n+\tcl_assert_equal_i(5, utf8_strnwidth(\"Hello\", 5, 1));\n+\n+\t/* Test simple Chinese characters - each should have width 2 */\n+\t/* \"你好\" is 6 bytes (3 bytes per char in UTF-8), 4 display columns */\n+\tcl_assert_equal_i(4, utf8_strnwidth(\"你好\", 6, 0));\n+\n+\t/* Test mixed ASCII and Chinese - ASCII = 1 column, Chinese = 2 columns */\n+\t/* \"h\"(1) + \"i\"(1) + \"你\"(2) + \"好\"(2) = 6 */\n+\tcl_assert_equal_i(6, utf8_strnwidth(\"Hi你好\", 8, 0));\n+\n+\t/* Test longer Chinese string */\n+\t/* 5 Chinese chars = 10 display columns */\n+\tcl_assert_equal_i(10, utf8_strnwidth(\"你好世界！\", 15, 0));\n+\n+\t/* Test individual Chinese character width */\n+\tcl_assert_equal_i(2, utf8_strnwidth(\"中\", 3, 0));\n+\n+\t/* Test empty string */\n+\tcl_assert_equal_i(0, utf8_strnwidth(\"\", 0, 0));\n+\n+\t/* Test length limiting */\n+\tstr = \"你好世界\";\n+\t/* Only first char \"你\"(2 columns) within 3 bytes */\n+\tcl_assert_equal_i(2, utf8_strnwidth(str, 3, 0));\n+\t/* First two chars \"你好\"(4 columns) in 6 bytes */\n+\tcl_assert_equal_i(4, utf8_strnwidth(str, 6, 0));\n+}\n+\n+/*\n+ * Tests for utf8_strwidth (simpler version without length limit)\n+ */\n+void test_utf8_width__strwidth_chinese(void)\n+{\n+\t/* Test basic ASCII */\n+\tcl_assert_equal_i(5, utf8_strwidth(\"Hello\"));\n+\n+\t/* Test Chinese characters */\n+\t/* 2 Chinese chars = 4 display columns */\n+\tcl_assert_equal_i(4, utf8_strwidth(\"你好\"));\n+\n+\t/* Test longer Chinese string */\n+\t/* 5 Chinese chars = 10 display columns */\n+\tcl_assert_equal_i(10, utf8_strwidth(\"你好世界！\"));\n+\n+\t/* Test mixed ASCII and Chinese */\n+\t/* 5 ASCII (5 cols) + 2 Chinese (4 cols) = 9 */\n+\tcl_assert_equal_i(9, utf8_strwidth(\"Hello世界\"));\n+\t/* 2 ASCII (2 cols) + 2 Chinese (4 cols) + 1 ASCII (1 col) = 7 */\n+\tcl_assert_equal_i(7, utf8_strwidth(\"Hi世界!\"));\n+}\n+\n+/*\n+ * Additional tests with other East Asian characters\n+ */\n+void test_utf8_width__strnwidth_japanese_korean(void)\n+{\n+\t/* Japanese characters (should also be 2 columns each) */\n+\t/* 5 Japanese chars x 2 cols each = 10 display columns */\n+\tcl_assert_equal_i(10, utf8_strnwidth(\"こんにちは\", 15, 0));\n+\n+\t/* Korean characters (should also be 2 columns each) */\n+\t/* 5 Korean chars x 2 cols each = 10 display columns */\n+\tcl_assert_equal_i(10, utf8_strnwidth(\"안녕하세요\", 15, 0));\n+}\n+\n+/*\n+ * Test utf8_strnwidth with CJK strings and ANSI sequences\n+ */\n+void test_utf8_width__strnwidth_cjk_with_ansi(void)\n+{\n+\t/* Test CJK with ANSI sequences */\n+\tconst char *ansi_test = \"\\033[1m你好\\033[0m\";\n+\tint width = utf8_strnwidth(ansi_test, strlen(ansi_test), 1);\n+\t/* Should skip ANSI sequences and count \"你好\" as 4 columns */\n+\tcl_assert_equal_i(4, width);\n+\n+\t/* Test mixed ASCII, CJK, and ANSI */\n+\tansi_test = \"Hello\\033[32m世界\\033[0m!\";\n+\twidth = utf8_strnwidth(ansi_test, strlen(ansi_test), 1);\n+\t/* \"Hello\"(5) + \"世界\"(4) + \"!\"(1) = 10 */\n+\tcl_assert_equal_i(10, width);\n+}\n-- \n2.52.0.rc2.5.g4c20a63325.dirty\n\n"},{"id":"530740","messageId":"d0975427c9002ed28e6bbf18403034709f286a2c.1763213290.git.worldhello.net@gmail.com","threadId":"64478","inReplyTo":"cover.1763213290.git.worldhello.net@gmail.com","subject":"[PATCH v2 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-15T13:36:11Z","receivedAt":"2025-11-15T13:36:22Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"The output table from \"git repo structure\" is misaligned when displaying\nUTF-8 characters (e.g., non-ASCII glyphs). E.g.:\n\n    | 仓库结构   | 值  |\n    | -------------- | ---- |\n    | * 引用       |      |\n    |   * 计数     |   67 |\n\nThe previous implementation used simple width formatting with printf()\nwhich didn't properly handle multi-byte UTF-8 characters, causing\nmisaligned table columns when displaying repository structure\ninformation.\n\nThis change modifies the stats_table_print_structure function to use\nstrbuf_utf8_align() instead of basic printf width specifiers. This\nensures proper column alignment regardless of the character encoding of\nthe content being displayed.\n\nAlso add test cases for strbuf_utf8_align(), a function newly introduced\nin \"builtin/repo.c\".\n\nSigned-off-by: Jiang Xin <worldhello.net@gmail.com>\n---\n builtin/repo.c              | 21 +++++++++++++++++----\n t/unit-tests/u-utf8-width.c | 37 +++++++++++++++++++++++++++++++++++++\n 2 files changed, 54 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 9d4749f79b..e3adb353a2 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -292,14 +292,20 @@ static void stats_table_print_structure(const struct stats_table *table)\n \tint name_col_width = utf8_strwidth(name_col_title);\n \tint value_col_width = utf8_strwidth(value_col_title);\n \tstruct string_list_item *item;\n+\tstruct strbuf buf = STRBUF_INIT;\n \n \tif (table->name_col_width > name_col_width)\n \t\tname_col_width = table->name_col_width;\n \tif (table->value_col_width > value_col_width)\n \t\tvalue_col_width = table->value_col_width;\n \n-\tprintf(\"| %-*s | %-*s |\\n\", name_col_width, name_col_title,\n-\t       value_col_width, value_col_title);\n+\tstrbuf_addstr(&buf, \"| \");\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, name_col_title);\n+\tstrbuf_addstr(&buf, \" | \");\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT, value_col_width, value_col_title);\n+\tstrbuf_addstr(&buf, \" |\");\n+\tprintf(\"%s\\n\", buf.buf);\n+\n \tprintf(\"| \");\n \tfor (int i = 0; i < name_col_width; i++)\n \t\tputchar('-');\n@@ -317,9 +323,16 @@ static void stats_table_print_structure(const struct stats_table *table)\n \t\t\tvalue = entry->value;\n \t\t}\n \n-\t\tprintf(\"| %-*s | %*s |\\n\", name_col_width, item->string,\n-\t\t       value_col_width, value);\n+\t\tstrbuf_reset(&buf);\n+\t\tstrbuf_addstr(&buf, \"| \");\n+\t\tstrbuf_utf8_align(&buf, ALIGN_LEFT, name_col_width, item->string);\n+\t\tstrbuf_addstr(&buf, \" | \");\n+\t\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, value_col_width, value);\n+\t\tstrbuf_addstr(&buf, \" |\");\n+\t\tprintf(\"%s\\n\", buf.buf);\n \t}\n+\n+\tstrbuf_release(&buf);\n }\n \n static void stats_table_clear(struct stats_table *table)\ndiff --git a/t/unit-tests/u-utf8-width.c b/t/unit-tests/u-utf8-width.c\nindex 3766f19726..86e09c3574 100644\n--- a/t/unit-tests/u-utf8-width.c\n+++ b/t/unit-tests/u-utf8-width.c\n@@ -95,3 +95,40 @@ void test_utf8_width__strnwidth_cjk_with_ansi(void)\n \t/* \"Hello\"(5) + \"世界\"(4) + \"!\"(1) = 10 */\n \tcl_assert_equal_i(10, width);\n }\n+\n+/*\n+ * Test the strbuf_utf8_align function with CJK characters\n+ */\n+void test_utf8_width__strbuf_utf8_align(void)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\n+\t/* Test left alignment with CJK */\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT, 10, \"你好\");\n+\t/* Since \"你好\" is 4 display columns, we need 6 more spaces to reach 10 */\n+\tcl_assert_equal_s(\"你好      \", buf.buf);\n+\tstrbuf_reset(&buf);\n+\n+\t/* Test right alignment with CJK */\n+\tstrbuf_utf8_align(&buf, ALIGN_RIGHT, 8, \"世界\");\n+\t/* \"世界\" is 4 display columns, so we need 4 leading spaces */\n+\tcl_assert_equal_s(\"    世界\", buf.buf);\n+\tstrbuf_reset(&buf);\n+\n+\t/* Test center alignment with CJK */\n+\tstrbuf_utf8_align(&buf, ALIGN_MIDDLE, 10, \"中\");\n+\t/* \"中\" is 2 display columns, so (10-2)/2 = 4 spaces on left, 4 on right */\n+\tcl_assert_equal_s(\"    中    \", buf.buf);\n+\tstrbuf_reset(&buf);\n+\n+\tstrbuf_utf8_align(&buf, ALIGN_MIDDLE, 5, \"中\");\n+\t/* \"中\" is 2 display columns, so (5-2)/2 = 1 spaces on left, 2 on right */\n+\tcl_assert_equal_s(\" 中  \", buf.buf);\n+\tstrbuf_reset(&buf);\n+\n+\t/* Test alignment that is smaller than string width */\n+\tstrbuf_utf8_align(&buf, ALIGN_LEFT, 2, \"你好\");\n+\t/* Since \"你好\" is 4 display columns, it should not be truncated */\n+\tcl_assert_equal_s(\"你好\", buf.buf);\n+\tstrbuf_release(&buf);\n+}\n-- \n2.52.0.rc2.5.g4c20a63325.dirty\n\n"},{"id":"530743","messageId":"0eee1597-3e83-4a47-90a5-60942da01673@gmail.com","threadId":"64478","inReplyTo":"d0975427c9002ed28e6bbf18403034709f286a2c.1763213290.git.worldhello.net@gmail.com","subject":"Re: [PATCH v2 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-11-15T15:04:11Z","receivedAt":"2025-11-15T15:04:18Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Jiang\n\nOn 15/11/2025 13:36, Jiang Xin wrote:\n> The output table from \"git repo structure\" is misaligned when displaying\n> UTF-8 characters (e.g., non-ASCII glyphs). E.g.:\n> \n>      | 仓库结构   | 值  |\n>      | -------------- | ---- |\n>      | * 引用       |      |\n>      |   * 计数     |   67 |\n> \n> The previous implementation used simple width formatting with printf()\n> which didn't properly handle multi-byte UTF-8 characters, causing\n> misaligned table columns when displaying repository structure\n> information.\n> \n> This change modifies the stats_table_print_structure function to use\n> strbuf_utf8_align() instead of basic printf width specifiers. This\n> ensures proper column alignment regardless of the character encoding of\n> the content being displayed.\n\nHow does it ensure proper column alignment for non-utf8 encodings? I\ndon't see how it is possible to calculate the display width without\nknowing the encoding.\n> Also add test cases for strbuf_utf8_align(), a function newly introduced\n> in \"builtin/repo.c\".\n\nNice.\n\nUsing strbuf_utf8_align ends up being quite verbose. An alternative\nwould be to keep using printf() but calculate the padding ourselves as\nshown below. Either way we end up calling utf8_strwidth() twice on the\nsame string which is a bit of a shame but probably doesn't matter too\nmuch in the grand scheme of things.\n\nThanks\n\nPhillip\n\n---- 8< ----\n\n  builtin/repo.c | 10 ++++++----\n  1 file changed, 6 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/repo.c b/builtin/repo.c\nindex 9d4749f79be..1b139b89672 100644\n--- a/builtin/repo.c\n+++ b/builtin/repo.c\n@@ -298,8 +298,9 @@ static void stats_table_print_structure(const struct stats_table *table)\n  \tif (table->value_col_width > value_col_width)\n  \t\tvalue_col_width = table->value_col_width;\n  \n-\tprintf(\"| %-*s | %-*s |\\n\", name_col_width, name_col_title,\n-\t       value_col_width, value_col_title);\n+\tprintf(\"| %s%*s | %s%*s |\\n\",\n+\t       name_col_title, name_col_width - utf8_strwidth(name_col_title), \"\",\n+\t       value_col_title, value_col_width - utf8_strwidth(value_col_title), \"\");\n  \tprintf(\"| \");\n  \tfor (int i = 0; i < name_col_width; i++)\n  \t\tputchar('-');\n@@ -317,8 +318,9 @@ static void stats_table_print_structure(const struct stats_table *table)\n  \t\t\tvalue = entry->value;\n  \t\t}\n  \n-\t\tprintf(\"| %-*s | %*s |\\n\", name_col_width, item->string,\n-\t\t       value_col_width, value);\n+\t\tprintf(\"| %s%*s | %*s%s |\\n\",\n+\t\titem->string, name_col_width - utf8_strwidth(item->string), \"\",\n+\t\tvalue_col_width - utf8_strwidth(value), \"\", value);\n  \t}\n  }\n  \n\n"},{"id":"530745","messageId":"xmqqtsyvfe2f.fsf@gitster.g","threadId":"64478","inReplyTo":"CANYiYbEFN9BHtNh1PQ9C3gDJasq1PaKnkcH-Nq=FddUCAcMGqg@mail.gmail.com","subject":"Re: [PATCH 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-15T16:36:24Z","receivedAt":"2025-11-15T16:36:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jiang Xin <worldhello.net@gmail.com> writes:\n\n>> >  builtin/repo.c | 22 ++++++++++++++++++----\n>> >  1 file changed, 18 insertions(+), 4 deletions(-)\n>>\n>> ... it is a good idea to protect the change with a new test or two\n>> to make sure the expected alignment in the output.\n>\n> Will add test cases for strbuf_utf8_align(), a function newly\n> introduced in builtin/repo.c.\n\nUnit tests are nice to make sure that building blocks like this\nhelper function works as expected.  To ensure that the application\nuses the building blocks correctly, you'd also need end-to-end test,\ngetting output out of the tool (\"repo struct\"?) and checking it.\n"},{"id":"530747","messageId":"xmqqjyzrfdgj.fsf@gitster.g","threadId":"64478","inReplyTo":"0eee1597-3e83-4a47-90a5-60942da01673@gmail.com","subject":"Re: [PATCH v2 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-15T16:49:32Z","receivedAt":"2025-11-15T16:49:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Phillip Wood <phillip.wood123@gmail.com> writes:\n\n> How does it ensure proper column alignment for non-utf8 encodings? I\n> don't see how it is possible to calculate the display width without\n> knowing the encoding.\n\nCorrect.  But for Git, pretty much the ship has sailed, I am afraid.\nAll tools that rely on utf8_strwidth() are \"broken\" in that way if\nyou feed latin-1 or ISO/IEC 2022, and that includes \"diff --stat\"\nwith pathnames in non-UTF8 encodings (I do not remember if we fully\nfixed the codepath for UTF-8---it used to be broken even for UTF-8).\n\n> Using strbuf_utf8_align ends up being quite verbose. An alternative\n> would be to keep using printf() but calculate the padding ourselves as\n> shown below.\n\nI think that has been the preferred way to do this, utf8_strwidth()\nto measure and decide how wide each column can be, then for each row,\nwe measure and make printf() fit, or truncate when the column we decide\nto allocate cannot accomodate the data on a particular row that is\noverly long.\n\nThanks.\n"},{"id":"530771","messageId":"CANYiYbFcap=c8xDy-=ZyaY3U4-jU9OEe18LPgTEAHi2wx2M0VQ@mail.gmail.com","threadId":"64478","inReplyTo":"xmqqtsyvfe2f.fsf@gitster.g","subject":"Re: [PATCH 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Jiang Xin","fromEmail":"worldhello.net@gmail.com","sentAt":"2025-11-16T13:32:52Z","receivedAt":"2025-11-16T13:33:04Z","isPatch":true,"sender":{"key":"worldhello.net@gmail.com","avatar":"https://avatars.githubusercontent.com/u/183860?v=4"},"body":"On Sun, Nov 16, 2025 at 12:36 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Jiang Xin <worldhello.net@gmail.com> writes:\n>\n> >> >  builtin/repo.c | 22 ++++++++++++++++++----\n> >> >  1 file changed, 18 insertions(+), 4 deletions(-)\n> >>\n> >> ... it is a good idea to protect the change with a new test or two\n> >> to make sure the expected alignment in the output.\n> >\n> > Will add test cases for strbuf_utf8_align(), a function newly\n> > introduced in builtin/repo.c.\n>\n> Unit tests are nice to make sure that building blocks like this\n> helper function works as expected.  To ensure that the application\n> uses the building blocks correctly, you'd also need end-to-end test,\n> getting output out of the tool (\"repo struct\"?) and checking it.\n\nt1901 already includes test cases to safeguard the output of the\n\"git repo structure\" command.  I could add a new test case to\nvalidate the output when localized in Chinese (as shown below),\nbut such a test would be inherently unstable, because it risks\nbreaking at the end of every release cycle whenever translations\nchange.\nTherefore, I feel it's better to fix the issue by using strbuf_utf8_align()\nand adding dedicated unit tests for it, rather than relying on\nfragile end-to-end localization tests.\n\n-------- 8< --------\n\ndiff --git a/t/t1901-repo-structure.sh b/t/t1901-repo-structure.sh\nindex 36a71a144e..fdab0a3d29 100755\n--- a/t/t1901-repo-structure.sh\n+++ b/t/t1901-repo-structure.sh\n@@ -34,6 +34,37 @@ test_expect_success 'empty repository' '\n        )\n '\n\n+test_expect_success 'output repo structure in non-ASCII glyphs' '\n+       test_when_finished \"rm -rf repo\" &&\n+       git init repo &&\n+       (\n+               cd repo &&\n+               cat >expect <<-\\EOF &&\n+               | 仓库结构       | 值 |\n+               | -------------- | -- |\n+               | * 引用         |    |\n+               |   * 计数       |  0 |\n+               |     * 分支     |  0 |\n+               |     * 标签     |  0 |\n+               |     * 远程     |  0 |\n+               |     * 其它     |  0 |\n+               |                |    |\n+               | * 可达的对象   |    |\n+               |   * 计数       |  0 |\n+               |     * 提交     |  0 |\n+               |     * 树       |  0 |\n+               |     * 数据对象 |  0 |\n+               |     * 标签     |  0 |\n+               EOF\n+\n+               env LC_ALL=zh_CN.utf-8 \\\n+                       git repo structure >out 2>err &&\n+\n+               test_cmp expect out &&\n+               test_line_count = 0 err\n+       )\n+'\n+\n test_expect_success 'repository with references and objects' '\n        test_when_finished \"rm -rf repo\" &&\n        git init repo &&\n"},{"id":"530773","messageId":"xmqqqztxdip4.fsf@gitster.g","threadId":"64478","inReplyTo":"CANYiYbFcap=c8xDy-=ZyaY3U4-jU9OEe18LPgTEAHi2wx2M0VQ@mail.gmail.com","subject":"Re: [PATCH 2/2] builtin/repo: fix table alignment for UTF-8 characters","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-16T16:51:35Z","receivedAt":"2025-11-16T16:51:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jiang Xin <worldhello.net@gmail.com> writes:\n\n> t1901 already includes test cases to safeguard the output of the\n> \"git repo structure\" command.  I could add a new test case to\n> validate the output when localized in Chinese (as shown below),\n> but such a test would be inherently unstable, because it risks\n> breaking at the end of every release cycle whenever translations\n> change.\n\nI haven't considered the i18n aspect.  We already compare program\noutput with expected output, so a change in a message has to be\nupdated together with the test that covers the code path, but po/\nupdates tend to come too late for test updates, so the problem is\nmuch more serious.\n\nOK.  Let's omit this feature from end-to-end testing at least for\nnow.  Thanks.\n"}]}