git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH 0/2] Improve documentation on UTF-16

From
brian m. carlson <sandals@crustytoothpaste.net>
Date
Dec 27, 2018, 02:17 UTC
Message-ID
<20181227021734.528629-1-sandals@crustytoothpaste.net>

We've recently fielded several reports from unhappy Windows users about our handling of UTF-16, UTF-16LE, and UTF-16BE, none of which seem to be suitable for certain Windows programs.

In an effort to communicate the reasons for our behavior more effectively, explain in the documentation that the UTF-16 variant that people have been asking for hasn't been standardized, and therefore hasn't been implemented in iconv(3). Mention what each of the variants do, so that people can make a decision which one meets their needs the best.

In addition, add a comment in the code about why we must, for correctness reasons, reject a UTF-16LE or UTF-16BE sequence that begins with U+FEFF, namely that such a codepoint semantically represents a ZWNBSP, not a BOM, but that that codepoint at the beginning of a UTF-8 sequence (as encoded in the object store) would be misinterpreted as a BOM instead.

This comment is in the code because I think it needs to be somewhere, but I'm not sure the documentation is the right place for it. If desired, I can add it to the documentation, although I feel the lurid details are not interesting to most users. If the wording is confusing, I'm very open to hearing suggestions for how to improve it.

I don't use Windows, so I don't know what MSVCRT does. If it requires a BOM but doesn't accept big-endian encoding, then perhaps we should report that as a bug to Microsoft so it can be fixed in a future version. That would probably make a lot more programs work right out of the box and dramatically improve the user experience.

As a note, I'm currently on vacation through the 2nd, so my responses may be slightly delayed.

brian m. carlson (2):
  Documentation: document UTF-16-related behavior
  utf8: add comment explaining why BOMs are rejected
 Documentation/gitattributes.txt | 5 +++++
 utf8.c                          | 7 +++++++
 2 files changed, 12 insertions(+)
Next: brian m. carlson
Message 1 of 12 in “Improve documentation on UTF-16”
  1. 0/2 Improve documentation on UTF-16brian m. carlson, Dec 27, 2018
  2. 1/2 Documentation: document UTF-16-related behaviorbrian m. carlson, Dec 27, 2018
  3. 2/2 utf8: add comment explaining why BOMs are rejectedbrian m. carlson, Dec 27, 2018
  4. Johannes SixtDec 27, 2018
  5. brian m. carlsonDec 27, 2018
  6. Johannes SixtDec 27, 2018
  7. brian m. carlsonDec 27, 2018
  8. Johannes SixtDec 28, 2018
  9. Philip OakleyDec 28, 2018
  10. Ævar Arnfjörð BjarmasonDec 28, 2018
  11. Philip OakleyDec 28, 2018
  12. brian m. carlsonDec 29, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.