git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v2] utf8: use size_t for string width methods and callee sites.

From
Pablo Sabater <pabloosabaterr@gmail.com>
Date
Jul 27, 2026, 01:06 UTC
Message-ID
<DK8Y8F4650AW.1XN921ROZW70F@gmail.com>
In-Reply-To
<20260726195718.1914131-1-hardikxk@gmail.com>

[+cc Junio, who reviewed this while I was writing mine, to keep him in the thread]

Hi!

Note that you have sent v2 in-reply-to my review from last version, not to the v1.

This Patch does not compile with DEVELOPER=1:
  $ make DEVELOPER=1
  builtin/blame.c:681:20: error: comparison of integers of different signs: 'int' and 'size_t'
        (aka 'unsigned long') [-Werror,-Wsign-compare]
    681 |                 if (longest_file < num)
        |                     ~~~~~~~~~~~~ ^ ~~~
  builtin/blame.c:691:23: error: comparison of integers of different signs: 'int' and 'size_t'
        (aka 'unsigned long') [-Werror,-Wsign-compare]
    691 |                         if (longest_author < num)
        |                             ~~~~~~~~~~~~~~ ^ ~~~
  builtin/blame.c:696:25: error: comparison of integers of different signs: 'int' and 'size_t'
        (aka 'unsigned long') [-Werror,-Wsign-compare]
    696 |                 if (longest_src_lines < num)
        |                     ~~~~~~~~~~~~~~~~~ ^ ~~~
  builtin/blame.c:699:25: error: comparison of integers of different signs: 'int' and 'size_t'
        (aka 'unsigned long') [-Werror,-Wsign-compare]
    699 |                 if (longest_dst_lines < num)
        |                     ~~~~~~~~~~~~~~~~~ ^ ~~~

Also note that many files have DISABLE_SIGN_COMPARE_WARNINGS which hides from us this errors.

That's why we have to be extra careful when changing signatures of functions accross the codebase.

The title is being too explicit, when a function signature changes, it is expected that their callers will change, no need to add it to the title. What about:

  utf8: make utf8_strwidth() and utf8_strnwidth() return size_t
does it work?
On Sun Jul 26, 2026 at 9:57 PM CEST, Hardik Kumar wrote:
Show 5 quoted lines
> utf8_strwidth() and utf8_strnwidth() return int, even though the
> return value is always non-negative:
>
> - utf8_strnwidth() accumulates the width into a size_t and otherwise
>   returns its size_t len parameter,
nit: change the comma for a dot.
Show 6 quoted lines
> - utf8_strwidth() just forwards its result.
>
> Change their signatures to return size_t instead.
>
> Update the types of the variables the said method is used to avoid
> potential UB caused by implicit conversion from size_t to int.
This is not correct and it reads a bit off, what about:
  Update the types of the variables where these functions are used, to
  avoid the implicit conversion from size_t to int.

This is not correct because the implicit conversion from size_t to int is not undefined behavior.

>
> The returned values from `utf8_strwidth()` are casted to int at places
> where it was falling tests or required other changes.

nit: s/casted/cast/ nit: s/falling/failing/

Show 31 quoted lines
>
> Signed-off-by: Hardik Kumar <hardikxk@gmail.com>
> ---
> Changes in v2:
> - reworked types for utf8_strwidth and its sites of usage.
> - removed redundant parens around `string`.
> - updated commit message for better explaining the patch.
>
>  builtin/blame.c  |  4 ++--
>  builtin/branch.c |  2 +-
>  builtin/repo.c   | 10 +++++-----
>  column.c         |  2 +-
>  diff.c           |  7 ++++---
>  gettext.c        |  2 +-
>  gettext.h        |  2 +-
>  pretty.c         |  5 +++--
>  utf8.c           | 13 ++++---------
>  utf8.h           |  4 ++--
>  wt-status.c      |  8 ++++----
>  11 files changed, 28 insertions(+), 31 deletions(-)
>
> diff --git a/builtin/blame.c b/builtin/blame.c
> index 48d5251..2d24b63 100644
> --- a/builtin/blame.c
> +++ b/builtin/blame.c
> @@ -564,7 +564,7 @@ static void emit_other(struct blame_scoreboard *sb, struct blame_entry *ent,
>  					name = ci.author_mail.buf;
>  				else
>  					name = ci.author.buf;
> -				pad = longest_author - utf8_strwidth(name);
> +				pad = longest_author - cast_size_t_to_int(utf8_strwidth(name));
This one is fine.
Show 9 quoted lines
>  				printf(" (%s%*s %10s",
>  				       name, pad, "",
>  				       format_time(ci.author_time,
> @@ -668,7 +668,7 @@ static void find_alignment(struct blame_scoreboard *sb, int *option)
>
>  	for (e = sb->ent; e; e = e->next) {
>  		struct blame_origin *suspect = e->suspect;
> -		int num;
> +		size_t num;
Looking at how num is used, it is reused for multiple things:
- strlen()
- utf8_strwidth()
- line-number sums
The longest_* variables we compare num against are still int.
Can we split num into different variables?
Show 13 quoted lines
>  		size_t marks_count = count_marks(e, *option);
>
>  		if (max_marks_count < marks_count)
> diff --git a/builtin/branch.c b/builtin/branch.c
> index dede60d..514ba64 100644
> --- a/builtin/branch.c
> +++ b/builtin/branch.c
> @@ -354,7 +354,7 @@ static int calc_maxwidth(struct ref_array *refs, int remote_bonus)
>  	for (i = 0; i < refs->nr; i++) {
>  		struct ref_array_item *it = refs->items[i];
>  		const char *desc = it->refname;
> -		int w;
> +		size_t w;
w receives utf8_strwidth() but later we have:
  if (w > max)

This is now size_t > int. Here I would keep w int and cast.

Show 13 quoted lines
>
>  		skip_prefix(it->refname, "refs/heads/", &desc);
>  		skip_prefix(it->refname, "refs/remotes/", &desc);
> diff --git a/builtin/repo.c b/builtin/repo.c
> index 84e012f..47b9191 100644
> --- a/builtin/repo.c
> +++ b/builtin/repo.c
> @@ -367,7 +367,7 @@ static void stats_table_vaddf(struct stats_table *table,
>  	struct strbuf buf = STRBUF_INIT;
>  	struct string_list_item *item;
>  	char *formatted_name;
> -	int name_width;
> +	size_t name_width;
Same as above:
  if (name_width > table->name_col_width)
I think that these three fields can be promoted safely
  struct stats_table {
	  [snip]
	  int name_col_width;
	  int value_col_width;
	  int unit_col_width;
  };

but check every use of them afterwards for code that still expects an int.

Show 9 quoted lines
>
>  	strbuf_vaddf(&buf, format, ap);
>  	formatted_name = strbuf_detach(&buf, NULL);
> @@ -387,12 +387,12 @@ static void stats_table_vaddf(struct stats_table *table,
>  		string_list_append_nodup(&table->annotations, strbuf_detach(&buf, NULL));
>  	}
>  	if (entry->value) {
> -		int value_width = utf8_strwidth(entry->value);
> +		size_t value_width = utf8_strwidth(entry->value);

I feel this one is partially my fault, I wrote these as example output of the grep I sent last reroll. But they still need to be checked:

>  		if (value_width > table->value_col_width)
We are comparing size_t > int.
>  			table->value_col_width = value_width;
We are narrowing size_t to int.
Show 15 quoted lines
>  	}
>  	if (entry->unit) {
> -		int unit_width = utf8_strwidth(entry->unit);
> +		size_t unit_width = utf8_strwidth(entry->unit);
>  		if (unit_width > table->unit_col_width)
>  			table->unit_col_width = unit_width;
>  	}
> @@ -582,8 +582,8 @@ static void stats_table_print_structure(const struct stats_table *table)
>  {
>  	const char *name_col_title = _("Repository structure");
>  	const char *value_col_title = _("Value");
> -	int title_name_width = utf8_strwidth(name_col_title);
> -	int title_value_width = utf8_strwidth(value_col_title);
> +	size_t title_name_width = utf8_strwidth(name_col_title);
> +	size_t title_value_width = utf8_strwidth(value_col_title);
Same problem, these are compared against int *_col_width locals,
and:
  value_col_width = title_value_width - unit_col_width

below the context now mixes size_t and int. Promoting the struct fields as suggested above fixes all of this at once.

Show 16 quoted lines
>  	int name_col_width = table->name_col_width;
>  	int value_col_width = table->value_col_width;
>  	int unit_col_width = table->unit_col_width;
> diff --git a/column.c b/column.c
> index 93fae31..6b7f921 100644
> --- a/column.c
> +++ b/column.c
> @@ -24,7 +24,7 @@ struct column_data {
>  };
>
>  /* return length of 's' in letters, ANSI escapes stripped */
> -static int item_length(const char *s)
> +static size_t item_length(const char *s)
>  {
>  	return utf8_strnwidth(s, strlen(s), 1);
>  }

item_length() has only one caller, which stores the result into an int *, so the value gets narrowed right back to int and this change buys nothing. Keep returning int and cast inside.

Show 11 quoted lines
> diff --git a/diff.c b/diff.c
> index 589c196..4887958 100644
> --- a/diff.c
> +++ b/diff.c
> @@ -2952,7 +2952,8 @@ static int utf8_ish_width(const char **start)
>
>  static void show_stats(struct diffstat_t *data, struct diff_options *options)
>  {
> -	int i, len, add, del, adds = 0, dels = 0;
> +	int i, add, del, adds = 0, dels = 0;
> +	size_t len;
This will have problems below:
	[snip]
	len = name_width;
	name_len = utf8_strwidth(name);
	if (name_width < name_len) {
		prefix = "...";
		len -= 3;
		if (len < 0)
			len = 0;
		while (name_len > len && *name)
	[snip]

if (len < 0) will always be false for a size_t, becoming dead code. Also name_len > len is int > size_t.

Let's step back a bit. name_len receives utf8_strwidth(), let's make it size_t too, and compare against len so the types stay consistent. The dead code can become:

	len = len > 3 ? len - 3 : 0;
We would also have to change the check:
	if (name_width < name_len)
to:
	if (len < name_len)
Show 9 quoted lines
>  	uintmax_t max_change = 0, max_len = 0;
>  	int total_files = data->nr, count;
>  	int width, name_width, graph_width, number_width = 0, bin_width = 0;
> @@ -3037,7 +3038,7 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)
>  	 * making the line longer than the maximum width.
>  	 */
>  	if (options->stat_width == -1)
> -		width = term_columns() - utf8_strnwidth(line_prefix, strlen(line_prefix), 1);
> +		width = term_columns() - cast_size_t_to_int(utf8_strnwidth(line_prefix, strlen(line_prefix), 1));
This one is fine.
Show 11 quoted lines
>  	else
>  		width = options->stat_width ? options->stat_width : 80;
>  	number_width = decimal_width(max_change) > number_width ?
> @@ -3123,7 +3124,7 @@ static void show_stats(struct diffstat_t *data, struct diff_options *options)
>  			if (slash)
>  				name = slash;
>  		}
> -		padding = len - utf8_strwidth(name);
> +		padding = len - cast_size_t_to_int(utf8_strwidth(name));
>  		if (padding < 0)
>  			padding = 0;

The cast doesn't work here because len is also size_t. We could do this to be sure that there will be no problems:

	size_t name_disp = utf8_strwidth(name);
	if (name_disp > len)
		padding = 0;
	else
		padding = cast_size_t_to_int(len - name_disp);
Show 24 quoted lines
>
> diff --git a/gettext.c b/gettext.c
> index 8d08a61..4d5d05e 100644
> --- a/gettext.c
> +++ b/gettext.c
> @@ -129,7 +129,7 @@ void git_setup_gettext(void)
>  }
>
>  /* return the number of columns of string 's' in current locale */
> -int gettext_width(const char *s)
> +size_t gettext_width(const char *s)
>  {
>  	static int is_utf8 = -1;
>  	if (is_utf8 == -1)
> diff --git a/gettext.h b/gettext.h
> index 484cafa..f161a21 100644
> --- a/gettext.h
> +++ b/gettext.h
> @@ -31,7 +31,7 @@
>  #ifndef NO_GETTEXT
>  extern int git_gettext_enabled;
>  void git_setup_gettext(void);
> -int gettext_width(const char *s);
> +size_t gettext_width(const char *s);

Careful, this is inside an #ifndef, if we change the signature here, the other branch must follow.

Show 22 quoted lines
>  #else
>  #define git_gettext_enabled (0)
>  static inline void git_setup_gettext(void)
> diff --git a/pretty.c b/pretty.c
> index d8a9f37..f7d392d 100644
> --- a/pretty.c
> +++ b/pretty.c
> @@ -1805,11 +1805,12 @@ static size_t format_and_pad_commit(struct strbuf *sb, /* in UTF-8 */
>  {
>  	struct strbuf local_sb = STRBUF_INIT;
>  	size_t total_consumed = 0;
> -	int len, padding = c->padding;
> +	int padding = c->padding;
> +	size_t len;
>
>  	if (padding < 0) {
>  		const char *start = strrchr(sb->buf, '\n');
> -		int occupied;
> +		size_t occupied;
>  		if (!start)
>  			start = sb->buf;
>  		occupied = utf8_strnwidth(start, strlen(start), 1);

padding is signed on purpose, so we can't have len as size_t. Later we have len > padding. Keep len and occupied as int, casting the utf8_strnwidth() results instead.

Show 25 quoted lines
> diff --git a/utf8.c b/utf8.c
> index 96460cc..cefaefe 100644
> --- a/utf8.c
> +++ b/utf8.c
> @@ -208,7 +208,7 @@ int utf8_width(const char **start, size_t *remainder_p)
>   * string, assuming that the string is utf8.  Returns strlen() instead
>   * if the string does not look like a valid utf8 string.
>   */
> -int utf8_strnwidth(const char *string, size_t len, int skip_ansi)
> +size_t utf8_strnwidth(const char *string, size_t len, int skip_ansi)
>  {
>  	const char *orig = string;
>  	size_t width = 0;
> @@ -225,15 +225,10 @@ int utf8_strnwidth(const char *string, size_t len, int skip_ansi)
>  		if (glyph_width > 0)
>  			width += glyph_width;
>  	}
> -
> -	/*
> -	 * TODO: fix the interface of this function and `utf8_strwidth()` to
> -	 * return `size_t` instead of `int`.
> -	 */
> -	return cast_size_t_to_int(string ? width : len);
> +	return string ? width : len;
>  }
This is good.
Show 13 quoted lines
>
> -int utf8_strwidth(const char *string)
> +size_t utf8_strwidth(const char *string)
>  {
>  	return utf8_strnwidth(string, strlen(string), 0);
>  }
> @@ -821,7 +816,7 @@ void strbuf_utf8_align(struct strbuf *buf, align_type position, unsigned int wid
>  		       const char *s)
>  {
>  	size_t slen = strlen(s);
> -	int display_len = utf8_strnwidth(s, slen, 0);
> +	size_t display_len = utf8_strnwidth(s, slen, 0);
>  	int utf8_compensation = slen - display_len;
This is fine.
Show 33 quoted lines
>
>  	if (display_len >= width) {
> diff --git a/utf8.h b/utf8.h
> index cf8ecb0..531e968 100644
> --- a/utf8.h
> +++ b/utf8.h
> @@ -7,8 +7,8 @@ typedef unsigned int ucs_char_t;  /* assuming 32bit int */
>
>  size_t display_mode_esc_sequence_len(const char *s);
>  int utf8_width(const char **start, size_t *remainder_p);
> -int utf8_strnwidth(const char *string, size_t len, int skip_ansi);
> -int utf8_strwidth(const char *string);
> +size_t utf8_strnwidth(const char *string, size_t len, int skip_ansi);
> +size_t utf8_strwidth(const char *string);
>  int is_utf8(const char *text);
>  int is_encoding_utf8(const char *name);
>  int same_encoding(const char *, const char *);
> diff --git a/wt-status.c b/wt-status.c
> index 58461e0..0e1e32d 100644
> --- a/wt-status.c
> +++ b/wt-status.c
> @@ -331,9 +331,9 @@ static int maxwidth(const char *(*label)(int), int minval, int maxval)
>
>  	for (i = minval; i <= maxval; i++) {
>  		const char *s = label(i);
> -		int len = s ? utf8_strwidth(s) : 0;
> +		size_t len = s ? utf8_strwidth(s) : 0;
>  		if (len > result)
> -			result = len;
> +			result = cast_size_t_to_int(len);
>  	}
>  	return result;
>  }

This function returns non-negative, I would say that the TODO applies to it as well because it returns 0 or utf8_strwidth(). Cleaning it to return size_t would help in some steps below:

Show 16 quoted lines
> @@ -360,7 +360,7 @@ static void wt_longstatus_print_unmerged_data(struct wt_status *s,
>  	status_printf(s, color(WT_STATUS_HEADER, s), "\t");
>
>  	how = wt_status_unmerged_status_string(d->stagemask);
> -	len = label_width - utf8_strwidth(how);
> +	len = label_width - cast_size_t_to_int(utf8_strwidth(how));
>  	status_printf_more(s, c, "%s%.*s%s\n", how, len, padding, one);
>  	strbuf_release(&onebuf);
>  }
> @@ -429,7 +429,7 @@ static void wt_longstatus_print_change_data(struct wt_status *s,
>  	what = wt_status_diff_status_string(status);
>  	if (!what)
>  		BUG("unhandled diff status %c", status);
> -	len = label_width - utf8_strwidth(what);
> +	len = label_width - cast_size_t_to_int(utf8_strwidth(what));
>  	assert(len >= 0);

We can see here that len has to be non-negative. label_width comes from maxwidth(). With the cleanup I suggested, maxwidth() returns size_t and so does label_width.

We can have the same safety with:
	size_t what_width = utf8_strwidth(what);
	if (what_width > label_width)
		BUG("label wider than column");
	len = cast_size_t_to_int(label_width - what_width);

You need the cast at the end anyway because len is used as the precision in "%.*s".

But I think this way leaves the code cleaner.
>  	if (one_name != two_name)
>  		status_printf_more(s, c, "%s%.*s%s -> %s",

Please know that the hunks of code I suggest might not be the direct solution, so don't take them directly. I do this because I don't want to write whole functions and I think that if I just write what's relevant it will be easier to understand.

Also, part of your work as author is to verify what you add and be able to defend it, which means taking what others say (including this review) with a grain of salt, everyone can make mistakes. You can run the tests on your own to check before sending the next version.

Nice work, Pablo

Previous: Hardik KumarNext: Hardik Kumar
Message 10 of 23 in “change utf8_strwidth() return type to size_t”
  1. change utf8_strwidth() return type to size_tHardik Kumar, Jul 26, 2026
  2. René ScharfeJul 26, 2026
  3. Hardik KumarJul 26, 2026
  4. Pablo SabaterJul 26, 2026
  5. Hardik KumarJul 26, 2026
  6. utf8: use size_t for string width methods and callee sites.Hardik Kumar, Jul 26, 2026
  7. Junio C HamanoJul 27, 2026
  8. Junio C HamanoJul 27, 2026
  9. Hardik KumarJul 27, 2026
  10. Pablo SabaterJul 27, 2026
  11. Hardik KumarJul 27, 2026
  12. utf8: make utf8_strwidth() and utf8_strnwidth() return size_tHardik Kumar, Jul 27, 2026
  13. Hardik KumarJul 27, 2026
  14. Phillip WoodJul 27, 2026
  15. Junio C HamanoJul 27, 2026
  16. Hardik KumarJul 27, 2026
  17. Junio C HamanoJul 27, 2026
  18. Hardik KumarJul 27, 2026
  19. utf8: replace utf8_strwidth todo with descriptive commentHardik Kumar, Jul 27, 2026
  20. Phillip WoodJul 28, 2026
  21. Hardik KumarJul 28, 2026
  22. Junio C HamanoJul 28, 2026
  23. utf8: replace utf8_strwidth todo with descriptive commentHardik Kumar, Jul 28, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.