git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v2 1/1] Support working-tree-encoding "UTF-16LE-BOM"

From
Junio C Hamano <gitster@pobox.com>
Date
Jan 22, 2019, 20:13 UTC
Message-ID
<xmqqk1iwfo1v.fsf@gitster-ct.c.googlers.com>
In-Reply-To
<20190120164327.3234-1-tboegi@web.de>
tboegi@web.de writes:
> The unicode standard itself defines 3 possible ways how to encode UTF-16.
> a) UTF-16, without BOM, big endian:
> b) UTF-16, with BOM, little endian:
> c) UTF-16, with BOM, big endian:
Is it OK to interpret "possible" as "allowed" above?
Show 5 quoted lines
> iconv (and libiconv) can generate UTF-16, UTF-16LE or UTF-16BE:
>
> d) UTF-16
> $ printf 'git' | iconv -f UTF-8 -t UTF-16 | od -c
> 0000000  376 377  \0   g  \0   i  \0   t
So among three, encoder can only do "big endian with BOM" (c).

Lack of (a) "big endian without BOM" in the encoder is not a problem in practice, as you can ask UTF-16BE to produce the stream, tell the decoder that you have UTF-16 and the lack of the BOM would make the decoder take it as (a).

But lack of (b) "little endian with BOM" is a problem.

So the proposal is to invent UTF-16-[BL]E-BOM that prepends BOM in front of UTF-16-[BL]E output to allow those who want (b).

Which makes sense, I guess. I do find it a bit ugly in the sense that it is something iconv should learn to do, as the issue is shared with all applications that want to use libiconv and convert into UTF-16.

Do you add UTF-16-BE-BOM for consistency? It would be identical to telling iconv to encode to UTF-16, if I understood your problem description correctly.

Show 18 quoted lines
> diff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt
> index b8392fc330..4a88ab8be7 100644
> --- a/Documentation/gitattributes.txt
> +++ b/Documentation/gitattributes.txt
> @@ -343,13 +343,13 @@ automatic line ending conversion based on your platform.
>  ------------------------
>
>  Use the following attributes if your '*.ps1' files are UTF-16 little
> -endian encoded without BOM and you want Git to use Windows line endings
> +endian encoded with BOM and you want Git to use Windows line endings
>  in the working directory. Please note, it is highly recommended to
>  explicitly define the line endings with `eol` if the `working-tree-encoding`
>  attribute is used to avoid ambiguity.
>
>  ------------------------
> -*.ps1		text working-tree-encoding=UTF-16LE eol=CRLF
> +*.ps1		text working-tree-encoding=UTF-16LE-BOM eol=CRLF
>  ------------------------

This change is robbing from those who do want a file without BOM to give to those who do want a file with BOM. Are the latter class of people the majority of the intended readers (read: Windows folks)?

I wonder if the following, instead of the above hunk, would work better:
 endian encoded without BOM and you want Git to use Windows line endings
-in the working directory. Please note, it is highly recommended to
+in the working directory (use `UTF-16-LE-BOM` instead of `UTF-16LE` if
+you want UTF-16 little endian with BOM).
+Please note, it is highly recommended to
 explicitly define the line endings with `eol` if the `working-tree-encoding`
Show 26 quoted lines
> @@ -540,10 +546,30 @@ char *reencode_string_len(const char *in, size_t insz,
>  {
>  	iconv_t conv;
>  	char *out;
> +	const char *bom_str = NULL;
> +	size_t bom_len = 0;
>
>  	if (!in_encoding)
>  		return NULL;
>
> +	/* UTF-16LE-BOM is the same as UTF-16 for reading */
> +	if (same_utf_encoding("UTF-16LE-BOM", in_encoding))
> +		in_encoding = "UTF-16";
> +
> +	/*
> +	 * For writing, UTF-16 iconv typically creates "UTF-16BE-BOM"
> +	 * Some users under Windows want the little endian version
> +	 */
> +	if (same_utf_encoding("UTF-16LE-BOM", out_encoding)) {
> +		bom_str = utf16_le_bom;
> +		bom_len = sizeof(utf16_le_bom);
> +		out_encoding = "UTF-16LE";
> +	} else if (same_utf_encoding("UTF-16BE-BOM", out_encoding)) {
> +		bom_str = utf16_be_bom;
> +		bom_len = sizeof(utf16_be_bom);
> +		out_encoding = "UTF-16BE";

OK, you do allow BE-BOM and the code does not rely on the fact that iconv happens to produce it with "UTF-16", because the library is free to switch between the three possible output (a)-(c) and we do not want to get affected by such a switch. Makes sense.

Previous: tboegi@web.deNext: tboegi@web.de
Message 18 of 25 in “git-rebase is ignoring working-tree-encoding”
  1. Adrián Gimeno BalaguerNov 2, 2018
  2. brian m. carlsonNov 4, 2018
  3. Adrián Gimeno BalaguerNov 4, 2018
  4. brian m. carlsonNov 4, 2018
  5. Torsten BögershausenNov 4, 2018
  6. Adrián Gimeno BalaguerNov 5, 2018
  7. Torsten BögershausenNov 5, 2018
  8. Torsten BögershausenNov 6, 2018
  9. Adrián Gimeno BalaguerNov 7, 2018
  10. Torsten BögershausenNov 8, 2018
  11. Alexandre GrigorievDec 26, 2018
  12. brian m. carlsonDec 26, 2018
  13. Alexandre GrigorievDec 27, 2018
  14. Torsten BögershausenDec 27, 2018
  15. Alexandre GrigorievDec 23, 2018
  16. 1/1 Support working-tree-encoding "UTF-16LE-BOM"tboegi@web.de, Dec 29, 2018
  17. 1/1 Support working-tree-encoding "UTF-16LE-BOM"tboegi@web.de, Jan 20, 2019
  18. Junio C HamanoJan 22, 2019
  19. 1/1 Support working-tree-encoding "UTF-16LE-BOM"tboegi@web.de, Jan 30, 2019
  20. Jason PyeronJan 30, 2019
  21. Torsten BögershausenJan 30, 2019
  22. 1/1 gitattributes.txt: fix typotboegi@web.de, Mar 6, 2019
  23. Junio C HamanoMar 7, 2019
  24. Adrián Gimeno BalaguerDec 29, 2018
  25. Philip OakleyDec 29, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.