git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] utf8: handle systems that don't write BOM for UTF-16

From
Eric Sunshine <sunshine@sunshineco.com>
Date
Feb 10, 2019, 01:45 UTC
Message-ID
<CAPig+cRyzZMOM19ztgR_wqvk68P_1eNNVBBj5pbY=MhQm08WAw@mail.gmail.com>
In-Reply-To
<20190209200802.277139-1-sandals@crustytoothpaste.net>

On Sat, Feb 9, 2019 at 3:08 PM brian m. carlson <sandals@crustytoothpaste.net> wrote:

Show 7 quoted lines
> [...]
> Add a Makefile and #define knob, ICONV_NEEDS_BOM, that can be set if the
> iconv implementation has this behavior. When set, Git will write a BOM
> manually for UTF-16 and UTF-32 and then force the data to be written in
> UTF-16BE or UTF-32BE. We choose big-endian behavior here because the
> tests use the raw "UTF-16" encoding, which will be big-endian when the
> implementation requires this knob to be set.

The name ICONV_NEEDS_BOM makes it sound as if we must feed a BOM _into_ 'iconv', which is quite confusing since the actual intention is that 'iconv' doesn't emit a BOM and we need to make up for the deficiency. Using a name such as ICONV_OMITS_BOM or ICONV_NEGLECTS_BOM makes it somewhat clearer that there is some deficiency with which we need to deal.

Show 6 quoted lines
> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>
> ---
> diff --git a/Makefile b/Makefile
> @@ -259,6 +259,9 @@ all::
> +# Define ICONV_NEEDS_BOM if your iconv implementation does not write a
> +# byte-order mark (BOM) when writing UTF-16 or UTF-32.

Not a big deal, but I wonder if it would be helpful to tack on "..., in which case it outputs big-endian unconditionally." or something.

Show 15 quoted lines
> diff --git a/t/t0028-working-tree-encoding.sh b/t/t0028-working-tree-encoding.sh
> @@ -6,6 +6,25 @@ test_description='working-tree-encoding conversion via gitattributes'
> +test_lazy_prereq NO_UTF16_BOM '
> +       test $(printf abc | iconv -f UTF-8 -t UTF-16 | wc -c) = 6
> +'
> +
> +test_lazy_prereq NO_UTF32_BOM '
> +       test $(printf abc | iconv -f UTF-8 -t UTF-32 | wc -c) = 12
> +'
> +
> +write_utf16 () {
> +       test_have_prereq NO_UTF16_BOM && printf '\xfe\xff'
> +       iconv -f UTF-8 -t UTF-16
> +
> +}
Stray blank line before the closing brace.
Show 5 quoted lines
> +
> +write_utf32 () {
> +       test_have_prereq NO_UTF32_BOM && printf '\x00\x00\xfe\xff'
> +       iconv -f UTF-8 -t UTF-32
> +}

It's probably doesn't matter much with these two tiny functions, but I was wondering if it would make sense to maintain the &&-chain, perhaps like this:

    if test test_have_prereq NO_UTF32_BOM
    then
        printf '\x00\x00\xfe\xff'
    fi &&
    iconv -f UTF-8 -t UTF-32
Previous: brian m. carlsonNext: brian m. carlson
Message 14 of 30 in “t0028-working-tree-encoding.sh failing on musl based systems (Alpine Linux)”
  1. Kevin DaudtFeb 7, 2019
  2. brian m. carlsonFeb 8, 2019
  3. Rich FelkerFeb 8, 2019
  4. brian m. carlsonFeb 8, 2019
  5. Kevin DaudtFeb 8, 2019
  6. brian m. carlsonFeb 8, 2019
  7. Junio C HamanoFeb 8, 2019
  8. Kevin DaudtFeb 8, 2019
  9. brian m. carlsonFeb 8, 2019
  10. Junio C HamanoFeb 8, 2019
  11. brian m. carlsonFeb 9, 2019
  12. Kevin DaudtFeb 9, 2019
  13. utf8: handle systems that don't write BOM for UTF-16brian m. carlson, Feb 9, 2019
  14. Eric SunshineFeb 10, 2019
  15. brian m. carlsonFeb 10, 2019
  16. Torsten BögershausenFeb 10, 2019
  17. brian m. carlsonFeb 10, 2019
  18. Junio C HamanoFeb 11, 2019
  19. utf8: handle systems that don't write BOM for UTF-16brian m. carlson, Feb 11, 2019
  20. Eric SunshineFeb 11, 2019
  21. brian m. carlsonFeb 11, 2019
  22. utf8: handle systems that don't write BOM for UTF-16brian m. carlson, Feb 11, 2019
  23. Kevin DaudtFeb 11, 2019
  24. brian m. carlsonFeb 11, 2019
  25. Junio C HamanoFeb 12, 2019
  26. brian m. carlsonFeb 12, 2019
  27. Junio C HamanoFeb 12, 2019
  28. utf8: handle systems that don't write BOM for UTF-16brian m. carlson, Feb 12, 2019
  29. Rich FelkerFeb 8, 2019
  30. Torsten BögershausenFeb 9, 2019

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.