{"thread":{"id":"50109","subject":"[PATCH 0/2] Improve documentation on UTF-16","startedAt":"2018-12-27T02:17:48Z","lastAt":"2018-12-29T23:18:49Z","messageCount":12,"participants":["brian m. carlson","Johannes Sixt","Ævar Arnfjörð Bjarmason","Philip Oakley"],"isPatch":true,"patchVersion":1,"patchTotal":2},"messages":[{"id":"365826","messageId":"20181227021734.528629-1-sandals@crustytoothpaste.net","threadId":"50109","inReplyTo":null,"subject":"[PATCH 0/2] Improve documentation on UTF-16","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-12-27T02:17:32Z","receivedAt":"2018-12-27T02:17:48Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"We've recently fielded several reports from unhappy Windows users about\nour handling of UTF-16, UTF-16LE, and UTF-16BE, none of which seem to be\nsuitable for certain Windows programs.\n\nIn an effort to communicate the reasons for our behavior more\neffectively, explain in the documentation that the UTF-16 variant that\npeople have been asking for hasn't been standardized, and therefore\nhasn't been implemented in iconv(3). Mention what each of the variants\ndo, so that people can make a decision which one meets their needs the\nbest.\n\nIn addition, add a comment in the code about why we must, for\ncorrectness reasons, reject a UTF-16LE or UTF-16BE sequence that begins\nwith U+FEFF, namely that such a codepoint semantically represents a\nZWNBSP, not a BOM, but that that codepoint at the beginning of a UTF-8\nsequence (as encoded in the object store) would be misinterpreted as a\nBOM instead.\n\nThis comment is in the code because I think it needs to be somewhere,\nbut I'm not sure the documentation is the right place for it. If\ndesired, I can add it to the documentation, although I feel the lurid\ndetails are not interesting to most users. If the wording is confusing,\nI'm very open to hearing suggestions for how to improve it.\n\nI don't use Windows, so I don't know what MSVCRT does. If it requires a\nBOM but doesn't accept big-endian encoding, then perhaps we should\nreport that as a bug to Microsoft so it can be fixed in a future\nversion. That would probably make a lot more programs work right out of\nthe box and dramatically improve the user experience.\n\nAs a note, I'm currently on vacation through the 2nd, so my responses\nmay be slightly delayed.\n\nbrian m. carlson (2):\n  Documentation: document UTF-16-related behavior\n  utf8: add comment explaining why BOMs are rejected\n\n Documentation/gitattributes.txt | 5 +++++\n utf8.c                          | 7 +++++++\n 2 files changed, 12 insertions(+)\n\n"},{"id":"365827","messageId":"20181227021734.528629-2-sandals@crustytoothpaste.net","threadId":"50109","inReplyTo":"20181227021734.528629-1-sandals@crustytoothpaste.net","subject":"[PATCH 1/2] Documentation: document UTF-16-related behavior","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-12-27T02:17:33Z","receivedAt":"2018-12-27T02:17:49Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"There are a number of broken Windows programs which want to process\nfiles in a UTF-16 variant that is always little endian and always\ncontains a BOM. Git cannot produce or accept such an encoding for the\nworking-tree-encoding because no such encoding has been defined with\nIANA or implemented in iconv(3).\n\nDocument this behavior since it is a frequent source of confusion for\nusers. Additionally, document that specifying \"UTF-16\" may produce bytes\nof either endianness, but will be sure to provide a BOM to distinguish.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n Documentation/gitattributes.txt | 5 +++++\n 1 file changed, 5 insertions(+)\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex b8392fc330..2b2c93afd1 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -330,6 +330,11 @@ That operation will fail and cause an error.\n - Reencoding content requires resources that might slow down certain\n   Git operations (e.g 'git checkout' or 'git add').\n \n+- It is not possible to specify a variant of UTF-16 with a BOM and a\n+  specified endianness, because no such variants have been standardized.\n+  Using \"UTF-16\" will produce a BOM with an unspecified endianness, and\n+  using \"UTF-16LE\" or \"UTF-16BE\" will prohibit a BOM from being used.\n+\n Use the `working-tree-encoding` attribute only if you cannot store a file\n in UTF-8 encoding and if you want Git to be able to process the content\n as text.\n"},{"id":"365828","messageId":"20181227021734.528629-3-sandals@crustytoothpaste.net","threadId":"50109","inReplyTo":"20181227021734.528629-1-sandals@crustytoothpaste.net","subject":"[PATCH 2/2] utf8: add comment explaining why BOMs are rejected","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-12-27T02:17:34Z","receivedAt":"2018-12-27T02:17:52Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"A source of confusion for many Git users is why UTF-16LE and UTF-16BE do\nnot allow a BOM, instead treating it as a ZWNBSP, according to the\nUnicode FAQ[0]. Explain in a comment why we cannot allow that to occur\ndue to our use of UTF-8 internally.\n\n[0] https://unicode.org/faq/utf_bom.html#bom9\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n utf8.c | 7 +++++++\n 1 file changed, 7 insertions(+)\n\ndiff --git a/utf8.c b/utf8.c\nindex eb78587504..22af2c485a 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -571,6 +571,13 @@ static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n \n+/*\n+ * We check here for a forbidden BOM. When using UTF-16BE or UTF-16LE, a BOM is\n+ * not allowed by RFC 2781, and any U+FEFF would be treated as a ZWNBSP, not a\n+ * BOM. However, because we encode into UTF-8 internally, we cannot allow that\n+ * character to occur as a ZWNBSP, since when encoded into UTF-8 it would be\n+ * interpreted as a BOM.\n+ */\n int has_prohibited_utf_bom(const char *enc, const char *data, size_t len)\n {\n \treturn (\n"},{"id":"365834","messageId":"93f0a854-9b8d-500c-b015-59c50ecdb0f3@kdbg.org","threadId":"50109","inReplyTo":"20181227021734.528629-1-sandals@crustytoothpaste.net","subject":"Re: [PATCH 0/2] Improve documentation on UTF-16","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2018-12-27T10:06:17Z","receivedAt":"2018-12-27T10:06:22Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 27.12.18 um 03:17 schrieb brian m. carlson:\n> We've recently fielded several reports from unhappy Windows users about\n> our handling of UTF-16, UTF-16LE, and UTF-16BE, none of which seem to be\n> suitable for certain Windows programs.\n> \n> In an effort to communicate the reasons for our behavior more\n> effectively, explain in the documentation that the UTF-16 variant that\n> people have been asking for hasn't been standardized, and therefore\n> hasn't been implemented in iconv(3). Mention what each of the variants\n> do, so that people can make a decision which one meets their needs the\n> best.\n> \n> In addition, add a comment in the code about why we must, for\n> correctness reasons, reject a UTF-16LE or UTF-16BE sequence that begins\n> with U+FEFF, namely that such a codepoint semantically represents a\n> ZWNBSP, not a BOM, but that that codepoint at the beginning of a UTF-8\n> sequence (as encoded in the object store) would be misinterpreted as a\n> BOM instead.\n> \n> This comment is in the code because I think it needs to be somewhere,\n> but I'm not sure the documentation is the right place for it. If\n> desired, I can add it to the documentation, although I feel the lurid\n> details are not interesting to most users. If the wording is confusing,\n> I'm very open to hearing suggestions for how to improve it.\n> \n> I don't use Windows, so I don't know what MSVCRT does. If it requires a\n> BOM but doesn't accept big-endian encoding, then perhaps we should\n> report that as a bug to Microsoft so it can be fixed in a future\n> version. That would probably make a lot more programs work right out of\n> the box and dramatically improve the user experience.\n\nIt worries me that theoretical correctness is regarded higher than \nexisting practice. I do not care a lot what some RFC tells what programs \nshould do if the majority of the software does something different and \nthat behavior has been proven useful in practice.\n\nMy understanding is that there is no such thing as a \"byte order \nmarker\". It just so happens that when the first character in some UTF-16 \ntext file begins with a ZWNBSP, then it is possible to derive the \nendianness of the file automatically. Other then that, that very first \ncode point U+FEFF *is part of the data* and must not be removed when the \ndata is reencoded. If Git does something different, it is bogus, IMO.\n\n-- Hannes\n"},{"id":"365868","messageId":"20181227164353.GC423984@genre.crustytoothpaste.net","threadId":"50109","inReplyTo":"93f0a854-9b8d-500c-b015-59c50ecdb0f3@kdbg.org","subject":"Re: [PATCH 0/2] Improve documentation on UTF-16","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-12-27T16:43:53Z","receivedAt":"2018-12-27T16:44:05Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Thu, Dec 27, 2018 at 11:06:17AM +0100, Johannes Sixt wrote:\n> It worries me that theoretical correctness is regarded higher than existing\n> practice. I do not care a lot what some RFC tells what programs should do if\n> the majority of the software does something different and that behavior has\n> been proven useful in practice.\n\nThe majority of OSes produce the behavior I document here, and they are\nthe majority of systems on the Internet. Windows is the outlier here,\nalthough a significant one. It is a common user of UTF-16 and its\nvariants, but so are Java and JavaScript, and they're present on a lot\nof devices. Swallowing the U+FEFF would break compatibility with those\nsystems.\n\nThe issue that Windows users are seeing is that libiconv always produces\nbig-endian data for UTF-16, and they always want little-endian. glibc\nproduces native-endian data, which is what Windows users want. Git for\nWindows could patch libiconv to do that (and that is the simple,\nfive-minute solution to this problem), but we'd still want to warn\npeople that they're relying on unspecified behavior, hence this series.\n\nI would even be willing to patch Git for Windows's libiconv if somebody\ncould point me to the repo (although I obviously cannot test it, not\nbeing a Windows user). I feel strongly, though, that fixing this is\noutside of the scope of Git proper, and it's not a thing we should be\nhandling here.\n\n> My understanding is that there is no such thing as a \"byte order marker\". It\n> just so happens that when the first character in some UTF-16 text file\n> begins with a ZWNBSP, then it is possible to derive the endianness of the\n> file automatically. Other then that, that very first code point U+FEFF *is\n> part of the data* and must not be removed when the data is reencoded. If Git\n> does something different, it is bogus, IMO.\n\nYou've got part of this. For UTF-16LE and UTF-16BE, a U+FEFF is part of\nthe text, as would a second one be if we had two at the beginning of a\nUTF-16 or UTF-8 sequence. If someone produces UTF-16LE and places a\nU+FEFF at the beginning of it, when we encode to UTF-8, we emit only one\nU+FEFF, which has the wrong semantics.\n\nTo be correct here and accept a U+FEFF, we'd need to check for a U+FEFF\nat the beginning of a UTF-16LE or UTF-16BE sequence and ensure we encode\nan extra U+FEFF at the beginning of the UTF-8 data (one for BOM and one\nfor the text) and then strip it off when we decode. That's kind of ugly,\nand since iconv doesn't do that itself, we'd have to.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"365879","messageId":"435b6870-379c-7183-da99-35aec5cf1137@kdbg.org","threadId":"50109","inReplyTo":"20181227164353.GC423984@genre.crustytoothpaste.net","subject":"Re: [PATCH 0/2] Improve documentation on UTF-16","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2018-12-27T19:55:27Z","receivedAt":"2018-12-27T19:55:32Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 27.12.18 um 17:43 schrieb brian m. carlson:\n> On Thu, Dec 27, 2018 at 11:06:17AM +0100, Johannes Sixt wrote:\n>> It worries me that theoretical correctness is regarded higher than existing\n>> practice. I do not care a lot what some RFC tells what programs should do if\n>> the majority of the software does something different and that behavior has\n>> been proven useful in practice.\n> \n> The majority of OSes produce the behavior I document here, and they are\n> the majority of systems on the Internet. Windows is the outlier here,\n> although a significant one. It is a common user of UTF-16 and its\n> variants, but so are Java and JavaScript, and they're present on a lot\n> of devices. Swallowing the U+FEFF would break compatibility with those\n> systems.\n> \n> The issue that Windows users are seeing is that libiconv always produces\n> big-endian data for UTF-16, and they always want little-endian. glibc\n> produces native-endian data, which is what Windows users want. Git for\n> Windows could patch libiconv to do that (and that is the simple,\n> five-minute solution to this problem), but we'd still want to warn\n> people that they're relying on unspecified behavior, hence this series.\n> \n> I would even be willing to patch Git for Windows's libiconv if somebody\n> could point me to the repo (although I obviously cannot test it, not\n> being a Windows user). I feel strongly, though, that fixing this is\n> outside of the scope of Git proper, and it's not a thing we should be\n> handling here.\n\nPlease appologize that I leave the majority of what you said uncommented \nas I am not deep in the matter and don't have a firm understanding of \nall the issues. I'll just trust what you said is sound.\n\nJust one thing: Please do the count by *users* (or existing files or \nnumber of charactes exchanged or something similar); do not just count \nOSs; I mean, Windows is *not* the outlier if it handles 90% of the \nUTF-16 data in the world. (I'm just making up numbers here, but I think \nyou get the point.)\n\n>> My understanding is that there is no such thing as a \"byte order marker\". It\n>> just so happens that when the first character in some UTF-16 text file\n>> begins with a ZWNBSP, then it is possible to derive the endianness of the\n>> file automatically. Other then that, that very first code point U+FEFF *is\n>> part of the data* and must not be removed when the data is reencoded. If Git\n>> does something different, it is bogus, IMO.\n> \n> You've got part of this. For UTF-16LE and UTF-16BE, a U+FEFF is part of\n> the text, as would a second one be if we had two at the beginning of a\n> UTF-16 or UTF-8 sequence. If someone produces UTF-16LE and places a\n> U+FEFF at the beginning of it, when we encode to UTF-8, we emit only one\n> U+FEFF, which has the wrong semantics.\n> \n> To be correct here and accept a U+FEFF, we'd need to check for a U+FEFF\n> at the beginning of a UTF-16LE or UTF-16BE sequence and ensure we encode\n> an extra U+FEFF at the beginning of the UTF-8 data (one for BOM and one\n> for the text) and then strip it off when we decode. That's kind of ugly,\n> and since iconv doesn't do that itself, we'd have to.\n\nBut why do you add another U+FEFF on the way to UTF-8? There is one in \nthe incoming UTF-16 data, and only *that* one must be converted. If \nthere is no U+FEFF in the UTF-16 data, the should not be one in UTF-8, \neither. Puzzled...\n\n-- Hannes\n"},{"id":"365892","messageId":"20181227234535.GD423984@genre.crustytoothpaste.net","threadId":"50109","inReplyTo":"435b6870-379c-7183-da99-35aec5cf1137@kdbg.org","subject":"Re: [PATCH 0/2] Improve documentation on UTF-16","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-12-27T23:45:35Z","receivedAt":"2018-12-27T23:45:45Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Thu, Dec 27, 2018 at 08:55:27PM +0100, Johannes Sixt wrote:\n> Am 27.12.18 um 17:43 schrieb brian m. carlson:\n> > You've got part of this. For UTF-16LE and UTF-16BE, a U+FEFF is part of\n> > the text, as would a second one be if we had two at the beginning of a\n> > UTF-16 or UTF-8 sequence. If someone produces UTF-16LE and places a\n> > U+FEFF at the beginning of it, when we encode to UTF-8, we emit only one\n> > U+FEFF, which has the wrong semantics.\n> > \n> > To be correct here and accept a U+FEFF, we'd need to check for a U+FEFF\n> > at the beginning of a UTF-16LE or UTF-16BE sequence and ensure we encode\n> > an extra U+FEFF at the beginning of the UTF-8 data (one for BOM and one\n> > for the text) and then strip it off when we decode. That's kind of ugly,\n> > and since iconv doesn't do that itself, we'd have to.\n> \n> But why do you add another U+FEFF on the way to UTF-8? There is one in the\n> incoming UTF-16 data, and only *that* one must be converted. If there is no\n> U+FEFF in the UTF-16 data, the should not be one in UTF-8, either.\n> Puzzled...\n\nSo for UTF-16, there must be a BOM. For UTF-16LE and UTF-16BE, there\nmust not be a BOM. So if we do this:\n\n  $ printf '\\xfe\\xff\\x00\\x0a' | iconv -f UTF-16BE -t UTF-16 | xxd -g1\n  00000000: ff fe ff fe 0a 00                                ......\n\nThat U+FEFF we have in the input is part of the text as a ZWNBSP; it is\nnot a BOM. We end up with two U+FEFF values. The first is the BOM that's\nrequired as part of UTF-16. The second is semantically part of the text\nand has the semantics of a zero-width non-breaking space.\n\nIn UTF-8, if the sequence starts with U+FEFF, it has the semantics of a\nBOM just like in UTF-16 (except that it's optional): it's not part of\nthe text, and should be stripped off. So when we receive a UTF-16LE or\nUTF-16BE sequence and it contains a U+FEFF (which is part of the text),\nwe need to insert a BOM in front of the sequence that's part of the text\nto keep the semantics.\n\nEssentially, we have this situation:\n\nText (in memory):  U+FEFF U+000A\nSemantics of text: ZWNBSP NL\nUTF-16BE:          FE FF  00 0A\nSemantics:         ZWNBSP NL\nUTF-16:            FE FF FE FF  00 0A\nSemantics:         BOM   ZWNBSP NL\nUTF-8:             EF BB BF EF BB BF 0A\nSemantics:         BOM      ZWNBSP   NL\n\nIf you don't have a U+FEFF, then things can be simpler:\n\nText (in memory):  U+0041 U+0042 U+0043\nSemantics of text: A      B      C\nUTF-16BE:          00 41 00 42 00 43\nSemantics:         A     B     C\nUTF-16:            FE FF 00 41 00 42 00 43\nSemantics:         BOM   A     B     C\nUTF-8:             41 42 43\nSemantics:         A  B  C\nUTF-8 (optional):  EF BB BF 41 42 43\nSemantics:         BOM      A  B  C\n\n(I have picked big-endian UTF-16 here, but little-endian is fine, too;\nthis is just easier for me to type.)\n\nThis is all a huge edge case involving correctly serializing code\npoints. By rejecting U=FEFF in UTF-16BE and UTF-16LE, we don't have to\ndeal with any of it.\n\nAs mentioned, I think patching Git for Windows's iconv is the smallest,\nmost achievable solution to this, because it means we don't have to\nhandle any of this edge case ourselves. Windows and WSL users can both\nwrite \"UTF-16\" and get a BOM and little-endian behavior, while we can\ndelegate all the rest of the encoding stuff to libiconv.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"365899","messageId":"87lg4adoo5.fsf@evledraar.gmail.com","threadId":"50109","inReplyTo":"20181227021734.528629-1-sandals@crustytoothpaste.net","subject":"Re: [PATCH 0/2] Improve documentation on UTF-16","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-12-28T08:46:18Z","receivedAt":"2018-12-28T08:46:25Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Dec 27 2018, brian m. carlson wrote:\n\n> We've recently fielded several reports from unhappy Windows users about\n> our handling of UTF-16, UTF-16LE, and UTF-16BE, none of which seem to be\n> suitable for certain Windows programs.\n\nJust for context, is \"we\" here $DAYJOB or a reference to some previous\nML thread(s) on this list, or something else?\n"},{"id":"365900","messageId":"34d4f912-2ec3-9dd1-f5fb-aad6a26e1464@kdbg.org","threadId":"50109","inReplyTo":"20181227234535.GD423984@genre.crustytoothpaste.net","subject":"Re: [PATCH 0/2] Improve documentation on UTF-16","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2018-12-28T08:59:05Z","receivedAt":"2018-12-28T08:59:09Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 28.12.18 um 00:45 schrieb brian m. carlson:\n> On Thu, Dec 27, 2018 at 08:55:27PM +0100, Johannes Sixt wrote:\n>> But why do you add another U+FEFF on the way to UTF-8? There is one in the\n>> incoming UTF-16 data, and only *that* one must be converted. If there is no\n>> U+FEFF in the UTF-16 data, the should not be one in UTF-8, either.\n>> Puzzled...\n> \n> So for UTF-16, there must be a BOM. For UTF-16LE and UTF-16BE, there\n> must not be a BOM. So if we do this:\n> \n>    $ printf '\\xfe\\xff\\x00\\x0a' | iconv -f UTF-16BE -t UTF-16 | xxd -g1\n>    00000000: ff fe ff fe 0a 00                                ......\n\nWhat sort of braindamage is this? Fix iconv.\n\nBut as I said, I'm not an expert. I just vented my worries that \nwidespread existing practice would be ignored under the excuse \"you are \nthe outlier\".\n\n-- Hannes\n"},{"id":"365930","messageId":"5c2ba2da-6bfc-bbbb-30e8-4eb86dd7136c@talktalk.net","threadId":"50109","inReplyTo":"87lg4adoo5.fsf@evledraar.gmail.com","subject":"Re: [PATCH 0/2] Improve documentation on UTF-16","fromName":"Philip Oakley","fromEmail":"philipoakley@talktalk.net","sentAt":"2018-12-28T20:35:03Z","receivedAt":"2018-12-28T20:37:34Z","isPatch":true,"sender":{"key":"philipoakley@talktalk.net","avatar":null},"body":"On 28/12/2018 08:46, Ævar Arnfjörð Bjarmason wrote:\n> On Thu, Dec 27 2018, brian m. carlson wrote:\n>\n>> We've recently fielded several reports from unhappy Windows users about\n>> our handling of UTF-16, UTF-16LE, and UTF-16BE, none of which seem to be\n>> suitable for certain Windows programs.\n> Just for context, is \"we\" here $DAYJOB or a reference to some previous\n> ML thread(s) on this list, or something else?\n\n\nI think \nhttps://public-inbox.org/git/CADN+U_PUfnYWb-wW6drRANv-ZaYBEk3gWHc7oJtxohA5Vc3NEg@mail.gmail.com/ \nwas the most recent on the Git list.\n\n-- \n\nPhilip\n\n"},{"id":"365932","messageId":"feeb0176-5bd5-229d-0ebb-d10120748aca@talktalk.net","threadId":"50109","inReplyTo":"34d4f912-2ec3-9dd1-f5fb-aad6a26e1464@kdbg.org","subject":"Re: [PATCH 0/2] Improve documentation on UTF-16","fromName":"Philip Oakley","fromEmail":"philipoakley@talktalk.net","sentAt":"2018-12-28T20:31:56Z","receivedAt":"2018-12-28T20:40:05Z","isPatch":true,"sender":{"key":"philipoakley@talktalk.net","avatar":null},"body":"On 28/12/2018 08:59, Johannes Sixt wrote:\n> Am 28.12.18 um 00:45 schrieb brian m. carlson:\n>> On Thu, Dec 27, 2018 at 08:55:27PM +0100, Johannes Sixt wrote:\n>>> But why do you add another U+FEFF on the way to UTF-8? There is one \n>>> in the\n>>> incoming UTF-16 data, and only *that* one must be converted. If \n>>> there is no\n>>> U+FEFF in the UTF-16 data, the should not be one in UTF-8, either.\n>>> Puzzled...\n>>\n>> So for UTF-16, there must be a BOM. For UTF-16LE and UTF-16BE, there\n>> must not be a BOM. So if we do this:\n>>\n>>    $ printf '\\xfe\\xff\\x00\\x0a' | iconv -f UTF-16BE -t UTF-16 | xxd -g1\n>>    00000000: ff fe ff fe 0a 00 ......\n>\n> What sort of braindamage is this? Fix iconv.\n>\n> But as I said, I'm not an expert. I just vented my worries that \n> widespread existing practice would be ignored under the excuse \"you \n> are the outlier\".\n>\n> -- Hannes\n\nFor ref, I dug out a Microsoft document [1] on its view of BOMs which \ncan be compared to the ref [0] Brian gave\n\n[1] \nhttps://docs.microsoft.com/en-us/windows/desktop/intl/using-byte-order-marks\n\n[0] https://unicode.org/faq/utf_bom.html#bom9\n\nMaybe the documentation patch ([PATCH 1/2] Documentation: document \nUTF-16-related behavior) should include the line \", because we encode \ninto UTF-8 internally,\", and a link to ref [0], and maybe [1]\n\n\nWhether the various Windows programs actually follow the Microsoft \nconvention is another matter altogether .\n\n-- \n\nPhilip\n\n\n"},{"id":"365981","messageId":"20181229231739.GE423984@genre.crustytoothpaste.net","threadId":"50109","inReplyTo":"87lg4adoo5.fsf@evledraar.gmail.com","subject":"Re: [PATCH 0/2] Improve documentation on UTF-16","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-12-29T23:17:39Z","receivedAt":"2018-12-29T23:18:49Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Fri, Dec 28, 2018 at 09:46:18AM +0100, Ævar Arnfjörð Bjarmason wrote:\n> \n> On Thu, Dec 27 2018, brian m. carlson wrote:\n> \n> > We've recently fielded several reports from unhappy Windows users about\n> > our handling of UTF-16, UTF-16LE, and UTF-16BE, none of which seem to be\n> > suitable for certain Windows programs.\n> \n> Just for context, is \"we\" here $DAYJOB or a reference to some previous\n> ML thread(s) on this list, or something else?\n\n\"We\" in this case is the Git list. I think the list has seen at least\nthree threads in recent months.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"}]}