git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v1] alias: support UTF-8 characters via subsection syntax

From
Jeff King <peff@peff.net>
Date
Feb 10, 2026, 07:44 UTC
Message-ID
<20260210074441.GE1756549@coredump.intra.peff.net>
In-Reply-To
<20260209220115.461109-1-jonatan@jontes.page>
On Mon, Feb 09, 2026 at 11:01:15PM +0100, Jonatan Holmgren wrote:
Show 13 quoted lines
> Git aliases are currently restricted to ASCII characters due to
> config key syntax limitations. This prevents non-English speakers
> from creating aliases in their native languages.
> 
> Add support for UTF-8 alias names using config subsections:
> 
>     [alias "förgrena"]
>         command = branch
> 
> The subsection name is matched verbatim (case-sensitive), while the
> existing flat syntax (alias.name) remains case-insensitive for
> backward compatibility. This approach uses existing config
> infrastructure and avoids complex Unicode normalization.

Thanks, this mostly looks quite good. I have a few comments below that range from a few small bugs to project-specific gotchas to style/readability suggestions.

> diff --git a/Documentation/RelNotes/2.54.0.adoc b/Documentation/RelNotes/2.54.0.adoc
> [...]

Usually patches do not fill out their own release notes, as it would create many conflicts when merging topics in arbitrary order. Instead, the maintainer generally writes up a blurb that goes in the merge commits and is eventually placed in the release notes by script. Suggesting a blurb does save work for the maintainer, but it should go in the cover letter (so for a single patch, after the "---" lines).

There's a bit on this in Documentation/SubmittingPatches, under the section "the-topic-summary".

> diff --git a/Documentation/config/alias.adoc b/Documentation/config/alias.adoc
> index 80ce17d2de..feba1e2022 100644
> --- a/Documentation/config/alias.adoc
> +++ b/Documentation/config/alias.adoca

Yay, I was happy to see nice comprehensive documentation here. And you do not seem to have gotten caught by any of the (many) formatting gotchas.

> +After defining these, you can run `git co`, `git hämta`, `git gömma`,
> +and `git "my alias with whitespace"` (note the quotes for aliases with spaces).

The use of literal backticks here makes sense, and if you haven't disabled it with a build-time knob, they should be shown in bold. Curiously in my build of the docs this doesn't quite work, and I get "git h" in bold, but "ämta" unbolded. Weird, and presumably a bug in asciidoc (the generated docbook xml shows the same thing, as does html). I get the same with asciidoctor, too. I'm not sure if it is worth working around it (or even how we would do so).

Show 32 quoted lines
> diff --git a/alias.c b/alias.c
> index 1a1a141a0a..f0c5f12fdd 100644
> --- a/alias.c
> +++ b/alias.c
> @@ -17,19 +17,45 @@ static int config_alias_cb(const char *key, const char *value,
>  			   const struct config_context *ctx UNUSED, void *d)
>  {
>  	struct config_alias_data *data = d;
> -	const char *p;
> +	const char *cmd, *subkey;
> +	size_t cmd_len;
> +	int is_subsection;
>  
> -	if (!skip_prefix(key, "alias.", &p))
> +	/* Use parse_config_key() to handle both 2-level and 3-level keys */
> +	if (parse_config_key(key, "alias", &cmd, &cmd_len, &subkey) < 0)
>  		return 0;
>  
> +	/*
> +	 * Support two syntaxes:
> +	 * 1. alias.name = value (simple, 2-level key)
> +	 * 2. [alias "name"] command = value (new, 3-level key)
> +	 */
> +	if (cmd) {
> +		if (strcmp(subkey, "command"))
> +			return 0;
> +		is_subsection = 1;
> +	} else {
> +		cmd = subkey;
> +		cmd_len = strlen(cmd);
> +		is_subsection = 0;
> +	}

OK, so versus what I showed earlier, we are keeping an is_subsection flag in order to differentiate the cases. And then we use that below...

Show 10 quoted lines
>  	if (data->alias) {
> -		if (!strcasecmp(p, data->alias)) {
> +		int match;
> +		if (is_subsection) {
> +			match = (strlen(data->alias) == cmd_len &&
> +				 !strncmp(data->alias, cmd, cmd_len));
> +		} else {
> +			match = (strlen(data->alias) == cmd_len &&
> +				 !strncasecmp(data->alias, cmd, cmd_len));
> +		}
...to decide whether to be case-insensitive or not. Makes sense.

I was going to suggest some simplifications here: if is_subsection is false, then we know that cmd is a NUL-terminated string, and we could just use strcasecmp() as before. And in fact, with that conditional we do not need the "cmd" aliasing at all. Which also means we don't need a separate flag to distinguish the cases.

You could just write:
  if (cmd) {
	match = (strlen(data->alias) == cmd_len &&
	         !strncmp(data->alias, cmd, cmd_len));
  } else {
	match = !strcasecmp(data->alias, subkey);
  }
But I guess we do use the redirection in the listing path later anyway:
>  	} else if (data->list) {
> -		string_list_append(data->list, p);
> +		string_list_append_nodup(data->list, xmemdupz(cmd, cmd_len));
>  	}

So maybe it is better as-is. I dunno. I have a feeling this would all need even more refactoring even if we ever added any more alias.foo.* config entries. So we can probably just leave it until then for more rearranging.

One thing I did find a little funny is the use of the word "subkey" for one variable. Usually we just call this the "key", but that name is annoyingly already taken by the function parameter (well before your patch). Usually we call that "var" instead. I don't know if it's worth changing or not. It's probably only confusing to me because I've stared at so many other Git config callback functions.

> diff --git a/t/t0014-alias-utf8.sh b/t/t0014-alias-utf8.sh

Is there any reason to add a new script here, and not just put these in the existing t0014? If you were bailing from the script when the prereq was not met, we'd need a separate script. But since you are using per-test prereqs, they can just exist alongside the other alias tests.

And also, you can't have two scripts with the same number. ;) Running "make test" should complain about that (I'll assume you ran your tests, but manually as ./t0014).

Show 5 quoted lines
> +# Skip if filesystem/locale doesn't support UTF-8
> +test_lazy_prereq UTF8_LOCALE '
> +	test_have_prereq !MINGW &&
> +	test_set_prereq UTF8_LOCALE
> +'

This is a slight mis-use of set_prereq. You should just need to return success from the lazy-prereq function, and it will be set automatically. So more like:

  test_lazy_prereq UTF8_LOCALE '
	test_have_prereq !MINGW
  '

At which point I wonder if there is much value in having UTF8_LOCALE at all, and not just putting "!MINGW" in each of the tests.

Beyond that, my larger question is: what are we asking of the platform and why does MINGW not work? We do not care about actual utf8 filesystem support, since these are just aliases. As long as the config file can store them and we can pass them on the command-line, that should be enough.

Show 5 quoted lines
> +test_expect_success UTF8_LOCALE 'setup UTF-8 aliases' '
> +	git config alias."förgrena".command branch &&
> +	git config alias."分支".command "branch --list" &&
> +	git config alias."test name".command status
> +'

Usually we'd just define these in the tests that use them, which keeps each individual test a bit more self-contained. So more like:

  test_expect_success 'UTF-8 alias with Swedish characters' '
	test_config alias."förgrena".command branch &&
	git förgrena >output &&
	...
  '

But we use them in the listing test, too, so they are each used twice. I dunno. It might make sense for the listing test to just define its own expected set.

> +test_expect_success UTF8_LOCALE 'UTF-8 alias with Swedish characters' '
> +	git förgrena >output &&
> +	test_grep -E "^(\* )?(main|master)" output
> +'

So I see we're aliasing "branch" here, which gives us this complicated regex for checking the output. Would something like "!echo ran my alias" make the test easier to understand?

> +test_expect_success 'list UTF-8 aliases' '
> +	git config --get-regexp "^alias\\..*\\.command" >output &&
> +	test_line_count -ge 3 output
> +'

This one needs the UTF8_LOCALE prereq, too (since we'd only have set up those aliases if we had it).

Is doing --get-regexp here all that interesting? It's not testing your new code at all, and would have passed already. I expected us to look at the output of "git help -a" to see that the entries are listed.

And trying that with your patch:
  ./git -c alias.föo.command=branch help -a
shows some possible issues:
  - we show it as "föo.command"
  - the alignment with other aliases is not quite right. Probably it is
    using a byte-count instead of utf8_strwidth()

I was puzzled that we would not parse it as föo correctly, since your patch modified the listing code. It looks like help.c has its own separate parser for this. Yuck. :( I think this also means that "git help föo" will fail.

So probably we want a preparatory patch to teach help.c to rely on list_aliases() and alias_lookup() from alias.[ch].

Show 5 quoted lines
> +test_expect_success 'flat syntax still works' '
> +	git config alias.testlegacy status &&
> +	git testlegacy >output &&
> +	test_grep "On branch" output
> +'
This probably is covered elsewhere already, but OK. :)
Show 5 quoted lines
> +test_expect_success 'new subsection syntax works' '
> +	git config alias.testnew.command status &&
> +	git testnew >output &&
> +	test_grep "On branch" output
> +'
Nice basic test without utf8. Good.
Show 5 quoted lines
> +test_expect_success 'subsection syntax only accepts command key' '
> +	git config alias.invalid.notcommand "value" &&
> +	test_must_fail git invalid 2>error &&
> +	test_grep -i "not a git command" error
> +'
Good thinking.
Show 14 quoted lines
> +test_expect_success 'simple syntax is case-insensitive' '
> +	git config alias.LegacyCase status &&
> +	git legacycase >output 2>&1 &&
> +	test_grep "On branch" output
> +'
> +
> +test_expect_success 'subsection syntax is case-sensitive' '
> +	test_commit case-test &&
> +	git config alias.SubCase.command "log --oneline" &&
> +	git config alias.subcase.command status &&
> +	git SubCase >upper.out 2>&1 &&
> +	git subcase >lower.out 2>&1 &&
> +	! test_cmp upper.out lower.out
> +'
And these are both very nice to see.
-Peff
Previous: Jonatan HolmgrenNext: Torsten Bögershausen
Message 15 of 88 in “[RFC] Support UTF-8 characters in Git alias names”
  1. Jonatan HolmgrenFeb 8, 2026
  2. D. Ben KnobleFeb 8, 2026
  3. brian m. carlsonFeb 8, 2026
  4. Junio C HamanoFeb 9, 2026
  5. Jonatan HolmgrenFeb 9, 2026
  6. Junio C HamanoFeb 9, 2026
  7. brian m. carlsonFeb 9, 2026
  8. Junio C HamanoFeb 9, 2026
  9. Ben KnobleFeb 10, 2026
  10. Junio C HamanoFeb 10, 2026
  11. Jeff KingFeb 10, 2026
  12. Jeff KingFeb 9, 2026
  13. Theodore TsoFeb 9, 2026
  14. alias: support UTF-8 characters via subsection syntaxJonatan Holmgren, Feb 9, 2026
  15. Jeff KingFeb 10, 2026
  16. Torsten BögershausenFeb 10, 2026
  17. Junio C HamanoFeb 10, 2026
  18. 0/2 support UTF-8 in alias namesJonatan Holmgren, Feb 10, 2026
  19. 1/2 help: use list_aliases() for alias listing and lookupJonatan Holmgren, Feb 10, 2026
  20. Junio C HamanoFeb 10, 2026
  21. 2/2 alias: support non-alphanumeric names via subsection syntaxJonatan Holmgren, Feb 10, 2026
  22. Junio C HamanoFeb 10, 2026
  23. Jonatan HolmgrenFeb 10, 2026
  24. Kristoffer HaugsbakkFeb 23, 2026
  25. Kristoffer HaugsbakkFeb 23, 2026
  26. Junio C HamanoFeb 23, 2026
  27. Kristoffer HaugsbakkFeb 23, 2026
  28. Patrick SteinhardtFeb 24, 2026
  29. 0/3 support UTF-8 in alias namesJonatan Holmgren, Feb 10, 2026
  30. 1/3 help: use list_aliases() for alias listingJonatan Holmgren, Feb 10, 2026
  31. Junio C HamanoFeb 10, 2026
  32. 2/3 alias: prepare for subsection aliasesJonatan Holmgren, Feb 10, 2026
  33. 3/3 alias: support non-alphanumeric names via subsection syntaxJonatan Holmgren, Feb 10, 2026
  34. 0/3 support UTF-8 in alias namesJonatan Holmgren, Feb 11, 2026
  35. 2/3 alias: prepare for subsection aliasesJonatan Holmgren, Feb 11, 2026
  36. Junio C HamanoFeb 11, 2026
  37. 1/3 help: use list_aliases() for alias listingJonatan Holmgren, Feb 11, 2026
  38. Junio C HamanoFeb 11, 2026
  39. 3/3 alias: support non-alphanumeric names via subsection syntaxJonatan Holmgren, Feb 11, 2026
  40. Junio C HamanoFeb 11, 2026
  41. Richard KerryFeb 12, 2026
  42. Jonatan HolmgrenFeb 12, 2026
  43. Jonatan HolmgrenFeb 12, 2026
  44. Torsten BögershausenFeb 12, 2026
  45. Jonatan HolmgrenFeb 12, 2026
  46. 0/4 support uTF-8 in alias namesJonatan Holmgren, Feb 16, 2026
  47. 4/4 completion: fix zsh alias listing for subsection aliasesJonatan Holmgren, Feb 16, 2026
  48. D. Ben KnobleFeb 16, 2026
  49. Junio C HamanoFeb 17, 2026
  50. 2/4 alias: prepare for subsection aliasesJonatan Holmgren, Feb 16, 2026
  51. 1/4 help: use list_aliases() for alias listingJonatan Holmgren, Feb 16, 2026
  52. 3/4 alias: support non-alphanumeric names via subsection syntaxJonatan Holmgren, Feb 16, 2026
  53. 0/4 support UTF-8 in alias namesJonatan Holmgren, Feb 18, 2026
  54. 2/4 alias: prepare for subsection aliasesJonatan Holmgren, Feb 18, 2026
  55. Kristoffer HaugsbakkFeb 18, 2026
  56. 1/4 help: use list_aliases() for alias listingJonatan Holmgren, Feb 18, 2026
  57. 4/4 completion: fix zsh alias listing for subsection aliasesJonatan Holmgren, Feb 18, 2026
  58. 3/4 alias: support non-alphanumeric names via subsection syntaxJonatan Holmgren, Feb 18, 2026
  59. 0/4 support UTF-8 in alias namesJonatan Holmgren, Feb 18, 2026
  60. 1/4 help: use list_aliases() for alias listingJonatan Holmgren, Feb 18, 2026
  61. Jacob KellerFeb 24, 2026
  62. Junio C HamanoFeb 24, 2026
  63. Junio C HamanoFeb 25, 2026
  64. Jacob KellerFeb 26, 2026
  65. Jacob KellerFeb 24, 2026
  66. 2/4 alias: prepare for subsection aliasesJonatan Holmgren, Feb 18, 2026
  67. 3/4 alias: support non-alphanumeric names via subsection syntaxJonatan Holmgren, Feb 18, 2026
  68. Kristoffer HaugsbakkFeb 24, 2026
  69. Jonatan HolmgrenFeb 24, 2026
  70. Kristoffer HaugsbakkFeb 24, 2026
  71. 4/4 completion: fix zsh alias listing for subsection aliasesJonatan Holmgren, Feb 18, 2026
  72. Junio C HamanoFeb 19, 2026
  73. Jonatan HolmgrenFeb 19, 2026
  74. 0/2 Fix small issues in alias subsection handlingJonatan Holmgren, Feb 24, 2026
  75. 1/2 doc: fix list continuation in alias subsection exampleJonatan Holmgren, Feb 24, 2026
  76. Junio C HamanoFeb 24, 2026
  77. Kristoffer HaugsbakkFeb 24, 2026
  78. Junio C HamanoFeb 24, 2026
  79. 2/2 alias: treat empty subsection [alias ""] as plain [alias]Jonatan Holmgren, Feb 24, 2026
  80. Junio C HamanoFeb 26, 2026
  81. 0/3 Fix small issues in alias subsection handlingJonatan Holmgren, Feb 26, 2026
  82. 2/3 alias: treat empty subsection [alias ""] as plain [alias]Jonatan Holmgren, Feb 26, 2026
  83. 1/3 doc: fix list continuation in alias subsection exampleJonatan Holmgren, Feb 26, 2026
  84. Kristoffer HaugsbakkMar 3, 2026
  85. Jonatan HolmgrenMar 3, 2026
  86. 3/3 git, help: fix memory leaks in alias listingJonatan Holmgren, Feb 26, 2026
  87. Junio C HamanoFeb 26, 2026
  88. doc: fix list continuation in alias.adocJonatan Holmgren, Mar 3, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.