git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH v2 1/1] diff.c: When appropriate, use utf8_strwidth()

From
tboegi@web.de <tboegi@web.de>
Date
Aug 27, 2022, 08:50 UTC
Message-ID
<20220827085007.20030-1-tboegi@web.de>
In-Reply-To
<CA+VDVVVmi99i6ZY64tg8RkVXDc5gOzQP_SH12zhDKRkUnhWFgw@mail.gmail.com>
From: Torsten Bögershausen <tboegi@web.de>

When unicode filenames (encoded in UTF-8) are used, the visible width on the screen is not the same as strlen(filename).

For example, `git log --stat` may produce an output like this:
$ git log --stat
[snip the header]
 Arger.txt  | 1 +
 Ärger.txt | 1 +
 2 files changed, 2 insertions(+)
A side note: the original report was about cyrillic filenames.
After some investigations it turned out that
a) This is not a problem with "ambiguous characters" in unicode
b) The same problem exist for all unicode code points (so we
  can use Latin based Umlauts for demonstrations below)

The 'Ä' takes the same space on the screen as the 'A'. But needs one more byte in memory, so the the `git log --stat` output for "Arger.txt" (!) gets mis-aligned: The maximum length is derived from "Ärger.txt", 10 bytes in memory, 9 positions on the screen. That is why "Arger.txt" gets one extra ' ' for aligment, it needs 9 bytes in memory. If there was a file "Ö", it would be correctly aligned by chance, but "Öhö" would not.

The solution is of course, to use utf8_strwidth() instead of strlen() when dealing with the width on screen.

And then there is another problem: code like this strbuf_addf(&out, "%-*s", len, name);

(or using the underlying snprintf() function) does not align the buffer to a minimum of len measured in screen-width, but uses the memory count, if name is UTF-8 encoded.

We could be tempted to wish that snprintf() was UTF-8 aware. That doesn't seem to be the case anywhere (tested on Linux and Mac), probably snprintf() uses the "bytes in memory"/strlen() approach to be compatible with older versions and this will never change.

The choosen solution is to split code in diff.c like this
strbuf_addf(&out, "%-*s", len, name);
into something like this:
size_t num_padding_spaces = 0;
// [snip]
if (len > utf8_strwidth(name))
    num_padding_spaces = len - utf8_strwidth(name);
strbuf_addf(&out, "%s", name);
if (num_padding_spaces)
    strbuf_addchars(&out, ' ', num_padding_spaces);
Tests:
Two things need to be tested:
- The calculation of the maximum width
- The calculation of num_padding_spaces
The name "textfile" is changed into "textfilë", both have a width of 8.
If strlen() was used, to get the maximum width, the shorter "binfile" would
have been mis-aligned:
 binfile   |  [snip]
 textfilë | [snip]
If only "binfile" would be renamed into "binfilë":
 binfilë |  [snip]
 textfile | [snip]
In order to verify that the width is calculated correctly everywhere,
"binfile" is renamed into "binfïlë", giving 2 bytes more in strlen()
"textfile" is renamed into "textfilë", 1 byte more in strlen(),
and the updated t4012-diff-binary.sh checks the correct aligment:
 binfïlë  | [snip]
 textfilë | [snip]
Reported-by: Alexander Meshcheryakov <alexander.s.m@gmail.com>
Signed-off-by: Torsten Bögershausen <tboegi@web.de>
---
 diff.c                 | 37 +++++++++++++++++++++++--------------
 t/t4012-diff-binary.sh | 14 +++++++-------
 2 files changed, 30 insertions(+), 21 deletions(-)
diff --git a/diff.c b/diff.c
index 974626a621..cf38e1dc88 100644
--- a/diff.c
+++ b/diff.c
@@ -2591,7 +2591,7 @@ void print_stat_summary(FILE *fp, int files,
 static void show_stats(struct diffstat_t *data, struct diff_options *options)
 {
 	int i, len, add, del, adds = 0, dels = 0;
-	uintmax_t max_change = 0, max_len = 0;
+	uintmax_t max_change = 0, max_width = 0;
 	int total_files = data->nr, count;
 	int width, name_width, graph_width, number_width = 0, bin_width = 0;
 	const char *reset, *add_c, *del_c;
@@ -2620,9 +2620,9 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)
 			continue;
 		}
 		fill_print_name(file);
-		len = strlen(file->print_name);
-		if (max_len < len)
-			max_len = len;
+		len = utf8_strwidth(file->print_name);
+		if (max_width < len)
+			max_width = len;

 		if (file->is_unmerged) {
 			/* "Unmerged" is 8 characters */
@@ -2646,7 +2646,7 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)

 	/*
 	 * We have width = stat_width or term_columns() columns total.
-	 * We want a maximum of min(max_len, stat_name_width) for the name part.
+	 * We want a maximum of min(max_width, stat_name_width) for the name part.
 	 * We want a maximum of min(max_change, stat_graph_width) for the +- part.
 	 * We also need 1 for " " and 4 + decimal_width(max_change)
 	 * for " | NNNN " and one the empty column at the end, altogether
@@ -2701,8 +2701,8 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)
 		graph_width = options->stat_graph_width;

 	name_width = (options->stat_name_width > 0 &&
-		      options->stat_name_width < max_len) ?
-		options->stat_name_width : max_len;
+		      options->stat_name_width < max_width) ?
+		options->stat_name_width : max_width;

 	/*
 	 * Adjust adjustable widths not to exceed maximum width
@@ -2734,6 +2734,7 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)
 		char *name = file->print_name;
 		uintmax_t added = file->added;
 		uintmax_t deleted = file->deleted;
+		size_t num_padding_spaces = 0;
 		int name_len;

 		if (!file->is_interesting && (added + deleted == 0))
@@ -2743,7 +2744,7 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)
 		 * "scale" the filename
 		 */
 		len = name_width;
-		name_len = strlen(name);
+		name_len = utf8_strwidth(name);
 		if (name_width < name_len) {
 			char *slash;
 			prefix = "...";
@@ -2753,10 +2754,14 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)
 			if (slash)
 				name = slash;
 		}
+		if (len > utf8_strwidth(name))
+			num_padding_spaces = len - utf8_strwidth(name);

 		if (file->is_binary) {
-			strbuf_addf(&out, " %s%-*s |", prefix, len, name);
-			strbuf_addf(&out, " %*s", number_width, "Bin");
+			strbuf_addf(&out, " %s%s ", prefix,  name);
+			if (num_padding_spaces)
+				strbuf_addchars(&out, ' ', num_padding_spaces);
+			strbuf_addf(&out, "| %*s", number_width, "Bin");
 			if (!added && !deleted) {
 				strbuf_addch(&out, '\n');
 				emit_diff_symbol(options, DIFF_SYMBOL_STATS_LINE,
@@ -2776,8 +2781,10 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)
 			continue;
 		}
 		else if (file->is_unmerged) {
-			strbuf_addf(&out, " %s%-*s |", prefix, len, name);
-			strbuf_addstr(&out, " Unmerged\n");
+			strbuf_addf(&out, " %s%s ", prefix,  name);
+			if (num_padding_spaces)
+				strbuf_addchars(&out, ' ', num_padding_spaces);
+			strbuf_addstr(&out, "| Unmerged\n");
 			emit_diff_symbol(options, DIFF_SYMBOL_STATS_LINE,
 					 out.buf, out.len, 0);
 			strbuf_reset(&out);
@@ -2803,8 +2810,10 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)
 				add = total - del;
 			}
 		}
-		strbuf_addf(&out, " %s%-*s |", prefix, len, name);
-		strbuf_addf(&out, " %*"PRIuMAX"%s",
+		strbuf_addf(&out, " %s%s ", prefix,  name);
+		if (num_padding_spaces)
+			strbuf_addchars(&out, ' ', num_padding_spaces);
+		strbuf_addf(&out, "| %*"PRIuMAX"%s",
 			number_width, added + deleted,
 			added + deleted ? " " : "");
 		show_graph(&out, '+', add, add_c, reset);
diff --git a/t/t4012-diff-binary.sh b/t/t4012-diff-binary.sh
index c509143c81..2d49de01c8 100755
--- a/t/t4012-diff-binary.sh
+++ b/t/t4012-diff-binary.sh
@@ -113,20 +113,20 @@ test_expect_success 'diff --no-index with binary creation' '
 '

 cat >expect <<EOF
- binfile  |   Bin 0 -> 1026 bytes
- textfile | 10000 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
+ binfïlë  |   Bin 0 -> 1026 bytes
+ textfilë | 10000 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
 EOF

 test_expect_success 'diff --stat with binary files and big change count' '
-	printf "\01\00%1024d" 1 >binfile &&
-	git add binfile &&
+	printf "\01\00%1024d" 1 >binfïlë &&
+	git add binfïlë &&
 	i=0 &&
 	while test $i -lt 10000; do
 		echo $i &&
 		i=$(($i + 1)) || return 1
-	done >textfile &&
-	git add textfile &&
-	git diff --cached --stat binfile textfile >output &&
+	done >textfilë &&
+	git add textfilë &&
+	git -c core.quotepath=false diff --cached --stat binfïlë textfilë >output &&
 	grep " | " output >actual &&
 	test_cmp expect actual
 '
--
2.34.0
Previous: Junio C HamanoNext: Torsten Bögershausen
Message 16 of 42 in “[BUG] Unicode filenames handling in `git log --stat`”
  1. Alexander MeshcheryakovAug 9, 2022
  2. Calvin WanAug 9, 2022
  3. Alexander MeshcheryakovAug 9, 2022
  4. Calvin WanAug 9, 2022
  5. Junio C HamanoAug 10, 2022
  6. Torsten BögershausenAug 10, 2022
  7. Alexander MeshcheryakovAug 10, 2022
  8. Torsten BögershausenAug 10, 2022
  9. Torsten BögershausenAug 10, 2022
  10. Junio C HamanoAug 10, 2022
  11. Torsten BögershausenAug 10, 2022
  12. 1/1 diff.c: When appropriate, use utf8_strwidth()tboegi@web.de, Aug 14, 2022
  13. Junio C HamanoAug 14, 2022
  14. Torsten BögershausenAug 15, 2022
  15. Junio C HamanoAug 18, 2022
  16. 1/1 diff.c: When appropriate, use utf8_strwidth()tboegi@web.de, Aug 27, 2022
  17. Torsten BögershausenAug 27, 2022
  18. Eric SunshineAug 27, 2022
  19. Johannes SchindelinAug 29, 2022
  20. Torsten BögershausenAug 29, 2022
  21. Junio C HamanoAug 29, 2022
  22. Johannes SchindelinSep 2, 2022
  23. 2/2 diff.c: More changes and tests around utf8_strwidth()tboegi@web.de, Sep 2, 2022
  24. Johannes SchindelinSep 2, 2022
  25. 1/2 diff.c: When appropriate, use utf8_strwidth(), part1tboegi@web.de, Sep 2, 2022
  26. Johannes SchindelinSep 2, 2022
  27. 1/2 diff.c: When appropriate, use utf8_strwidth(), part1tboegi@web.de, Sep 3, 2022
  28. Junio C HamanoSep 5, 2022
  29. Torsten BögershausenSep 7, 2022
  30. Junio C HamanoSep 7, 2022
  31. 2/2 diff.c: More changes and tests around utf8_strwidth()tboegi@web.de, Sep 3, 2022
  32. Johannes SchindelinSep 5, 2022
  33. 1/1 diff.c: When appropriate, use utf8_strwidth()tboegi@web.de, Sep 14, 2022
  34. Junio C HamanoSep 14, 2022
  35. Torsten BögershausenSep 26, 2022
  36. Junio C HamanoOct 10, 2022
  37. Torsten BögershausenOct 20, 2022
  38. Junio C HamanoOct 20, 2022
  39. Torsten BögershausenOct 21, 2022
  40. Junio C HamanoOct 21, 2022
  41. Torsten BögershausenOct 23, 2022
  42. Junio C HamanoSep 15, 2022

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.