{"thread":{"id":"49734","subject":"git-rebase is ignoring working-tree-encoding","startedAt":"2018-11-02T02:30:33Z","lastAt":"2019-03-07T00:24:37Z","messageCount":25,"participants":["Adrián Gimeno Balaguer","brian m. carlson","Torsten Bögershausen","Alexandre Grigoriev","tboegi@web.de","Philip Oakley","Junio C Hamano","Jason Pyeron"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"362228","messageId":"CADN+U_PUfnYWb-wW6drRANv-ZaYBEk3gWHc7oJtxohA5Vc3NEg@mail.gmail.com","threadId":"49734","inReplyTo":null,"subject":"git-rebase is ignoring working-tree-encoding","fromName":"Adrián Gimeno Balaguer","fromEmail":"adrigibal@gmail.com","sentAt":"2018-11-02T02:30:17Z","receivedAt":"2018-11-02T02:30:33Z","isPatch":false,"sender":{"key":"adrigibal@gmail.com","avatar":null},"body":"I’m attempting to perform fixups via git-rebase of UTF-16 LE files\n(the project I’m working on requires that exact encoding on certain\nfiles). When the rebase is complete, Git changes that file’s encoding\nto UTF-16 BE. I have been using the newer working-tree-encoding\nattribute in .gitattributes. I’m using Git for Windows.\n\n$ git version\ngit version 2.19.1.windows.1\n\nHere is a sample UTF-16 LE file (with BOM and LF endings) with\nfollowing atributes in .gitattributes:\n\ntest.txt eol=lf -text working-tree-encoding=UTF-16\n\nI put eol=lf and -text to tell Git to not change the encoding of the\nfile on checkout, but that doesn’t even help. Asides, the newer\nworking-tree-encoding allows me to view human-readable diffs of that\nfile (in GitHub Desktop and Git Bash). Now, note that doing for\nexample consecutive commits to the same file does not affect the\nUTF-16 LE encoding. And before I discovered this attribute, the whole\nthing was even worse when squash/fixup rebasing, as Git would modify\nthe file with Chinese characters (when manually setting it as text via\n.gitattributes).\n\nSo, again the problem with the exposed .gitattributes line is that\nafter fixup rebasing, UTF-16 LE files encoding change to UTF-16 BE.\n\nFor long, I have been working with the involved UTF-16 LE files set as\nbinary via .gitattributes (e.g. “test.txt binary”), so that Git would\nnot modify the file encoding, but this doesn’t allow me to view the\ndiffs upon changes in GitHub Desktop, which I want (and neither via\ngit diff).\n"},{"id":"362400","messageId":"20181104154744.GI731755@genre.crustytoothpaste.net","threadId":"49734","inReplyTo":"CADN+U_PUfnYWb-wW6drRANv-ZaYBEk3gWHc7oJtxohA5Vc3NEg@mail.gmail.com","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-11-04T15:47:44Z","receivedAt":"2018-11-04T15:47:52Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Fri, Nov 02, 2018 at 03:30:17AM +0100, Adrián Gimeno Balaguer wrote:\n> I’m attempting to perform fixups via git-rebase of UTF-16 LE files\n> (the project I’m working on requires that exact encoding on certain\n> files). When the rebase is complete, Git changes that file’s encoding\n> to UTF-16 BE. I have been using the newer working-tree-encoding\n> attribute in .gitattributes. I’m using Git for Windows.\n> \n> $ git version\n> git version 2.19.1.windows.1\n> \n> Here is a sample UTF-16 LE file (with BOM and LF endings) with\n> following atributes in .gitattributes:\n> \n> test.txt eol=lf -text working-tree-encoding=UTF-16\n\nDo things work for you if you write this as \"UTF-16LE\"?  When you use\nworking-tree-encoding, the file is stored internally as UTF-8, but it's\nserialized to the specified encoding when written out.\n\nAsking for \"UTF-16\" is ambiguous: there are two endiannesses, and so as\nlong as you get a BOM in the output, either one is an acceptable option.\nWhich one you get is dependent on what the underlying code thinks is the\ndefault, and traditionally for Unix systems and Unix tools that's been\nbig-endian.  If you want a particular endianness, you should specify it.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"362402","messageId":"CADN+U_Nw5wCyK1SPRgsxzFbJ-KKnOV2Ub8YA3_a80SZYwKC5FQ@mail.gmail.com","threadId":"49734","inReplyTo":"20181104154744.GI731755@genre.crustytoothpaste.net","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"Adrián Gimeno Balaguer","fromEmail":"adrigibal@gmail.com","sentAt":"2018-11-04T16:37:09Z","receivedAt":"2018-11-04T16:37:24Z","isPatch":false,"sender":{"key":"adrigibal@gmail.com","avatar":null},"body":"El dom., 4 nov. 2018 a las 16:48, brian m. carlson\n(<sandals@crustytoothpaste.net>) escribió:\n> Do things work for you if you write this as \"UTF-16LE\"?  When you use\n> working-tree-encoding, the file is stored internally as UTF-8, but it's\n> serialized to the specified encoding when written out.\n\nWhen I use UTF-16LE or UTF-16BE, then I can't commit or view diffs of\nspecified files, as Git prohibites BOM existance in these cases,\nshowing an error when attempting to commit. But BOM must also exist\nfor the project. I even experimented for fixing this issue within\nGit's source. It turns out that Git is following an Unicode rule that\nsays that BOM is not permitted when declaring exact UTF-16BE/UTF-16LE\nMIME (and UTF-32 variants) encoding types:\n\nhttps://github.com/git/git/blob/master/utf8.h#L87\n\n> Asking for \"UTF-16\" is ambiguous: there are two endiannesses, and so as\n> long as you get a BOM in the output, either one is an acceptable option.\n> Which one you get is dependent on what the underlying code thinks is the\n> default, and traditionally for Unix systems and Unix tools that's been\n> big-endian.  If you want a particular endianness, you should specify it.\n\nI wrote a \"counterpart\" easy fix which instead only prohibites BOM for\nthe opposite endianness (for example if\nworking-tree-encoding=UTF-16LE, then finding an UTF-16BE BOM in the\nfile would cause Git to signal the error right before committing,\ndiffing, etc.). That way user would be encouraged to modify the file's\nencoding to match the one specified in working-tree-encoding before\nallowing these actions, therefore preventing Git from encoding to the\nwrong endianness after file is written out. With few repository tests,\nthis new behaviour worked as expected. But then I realized this\nsolution would perhaps be unacceptable for Git's source code as it\nwould violate that Unicode standard. Anyways, here is a PR in my Git\nfork with the changes I did, for reference:\n\nhttps://github.com/AdRiAnIlloO/git/pull/1\n\nAh this point, the solution I came with recently for my project was\nwriting some code in Shell to fix the endianness of the re-encoded\nfiles to UTF-16BE after the Git's write out process (or a \"working\ntree refresh\" in my own words), within the same script that I use to\npack assets including the localization files.\n\n> brian m. carlson: Houston, Texas, US\n> OpenPGP: https://keybase.io/bk2204\n\n\n\n-- \nAdrián\n"},{"id":"362404","messageId":"20181104170729.GA21372@tor.lan","threadId":"49734","inReplyTo":"CADN+U_PUfnYWb-wW6drRANv-ZaYBEk3gWHc7oJtxohA5Vc3NEg@mail.gmail.com","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2018-11-04T17:07:29Z","receivedAt":"2018-11-04T17:07:33Z","isPatch":false,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On Fri, Nov 02, 2018 at 03:30:17AM +0100, Adrián Gimeno Balaguer wrote:\n> I’m attempting to perform fixups via git-rebase of UTF-16 LE files\n> (the project I’m working on requires that exact encoding on certain\n> files). When the rebase is complete, Git changes that file’s encoding\n> to UTF-16 BE. I have been using the newer working-tree-encoding\n> attribute in .gitattributes. I’m using Git for Windows.\n> \n> $ git version\n> git version 2.19.1.windows.1\n> \n> Here is a sample UTF-16 LE file (with BOM and LF endings) with\n> following atributes in .gitattributes:\n> \n> test.txt eol=lf -text working-tree-encoding=UTF-16\n> \n> I put eol=lf and -text to tell Git to not change the encoding of the\n> file on checkout, but that doesn’t even help. Asides, the newer\n> working-tree-encoding allows me to view human-readable diffs of that\n> file (in GitHub Desktop and Git Bash). Now, note that doing for\n> example consecutive commits to the same file does not affect the\n> UTF-16 LE encoding. And before I discovered this attribute, the whole\n> thing was even worse when squash/fixup rebasing, as Git would modify\n> the file with Chinese characters (when manually setting it as text via\n> .gitattributes).\n> \n> So, again the problem with the exposed .gitattributes line is that\n> after fixup rebasing, UTF-16 LE files encoding change to UTF-16 BE.\n> \n> For long, I have been working with the involved UTF-16 LE files set as\n> binary via .gitattributes (e.g. “test.txt binary”), so that Git would\n> not modify the file encoding, but this doesn’t allow me to view the\n> diffs upon changes in GitHub Desktop, which I want (and neither via\n> git diff).\n\nThanks for the report.\nI have tried to follow the problem from your verbal descriptions\n(and the PR) but I need to admit that I don't fully understand the\nproblem (yet).\n\nCould you try to create some instructions how to reproduce it?\nA numer of shell istructions would be great,\nin best case some kind of \"test case\", like the tests in\nthe t/ directory in Git.\n\nIt would be nice to be able to re-produce it.\nAnd if there is a bug, to get it fixed.\n"},{"id":"362410","messageId":"20181104183813.GJ731755@genre.crustytoothpaste.net","threadId":"49734","inReplyTo":"CADN+U_Nw5wCyK1SPRgsxzFbJ-KKnOV2Ub8YA3_a80SZYwKC5FQ@mail.gmail.com","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-11-04T18:38:13Z","receivedAt":"2018-11-04T18:38:26Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sun, Nov 04, 2018 at 05:37:09PM +0100, Adrián Gimeno Balaguer wrote:\n> I wrote a \"counterpart\" easy fix which instead only prohibites BOM for\n> the opposite endianness (for example if\n> working-tree-encoding=UTF-16LE, then finding an UTF-16BE BOM in the\n> file would cause Git to signal the error right before committing,\n> diffing, etc.). That way user would be encouraged to modify the file's\n> encoding to match the one specified in working-tree-encoding before\n> allowing these actions, therefore preventing Git from encoding to the\n> wrong endianness after file is written out. With few repository tests,\n> this new behaviour worked as expected. But then I realized this\n> solution would perhaps be unacceptable for Git's source code as it\n> would violate that Unicode standard. Anyways, here is a PR in my Git\n> fork with the changes I did, for reference:\n\nI actually think such a solution (although I haven't looked at your\npatch) would be fine, and I would encourage you to send it to the list.\nIt's my understanding that many people on Windows want to write things\nin UTF-16 encoding but only little-endian with a BOM.  Allowing them to\nwrite that, even if Git won't be able to guarantee producing that, would\nbe fine, as long as the data is what we expect.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"362452","messageId":"CADN+U_MgrGHLQ5QNa-HgzxLN4zJLJPu4PaT2MTRoc18=gET+5Q@mail.gmail.com","threadId":"49734","inReplyTo":"20181104170729.GA21372@tor.lan","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"Adrián Gimeno Balaguer","fromEmail":"adrigibal@gmail.com","sentAt":"2018-11-05T04:24:39Z","receivedAt":"2018-11-05T04:24:54Z","isPatch":false,"sender":{"key":"adrigibal@gmail.com","avatar":null},"body":"El dom., 4 nov. 2018 a las 18:07, Torsten Bögershausen\n(<tboegi@web.de>) escribió:\n>\n> Thanks for the report.\n> I have tried to follow the problem from your verbal descriptions\n> (and the PR) but I need to admit that I don't fully understand the\n> problem (yet).\n\nI have created a PR in the Git's repository. You can read an updated\ndescription there:\n\nhttps://github.com/git/git/pull/550\n\n> Could you try to create some instructions how to reproduce it?\n> A numer of shell instructions would be great,\n> in best case some kind of \"test case\", like the tests in\n> the t/ directory in Git.\n>\n> It would be nice to be able to re-produce it.\n> And if there is a bug, to get it fixed.\n\nThis is covered in the mentioned PR above. Thanks for feedback.\n\n-- \nAdrián\n"},{"id":"362496","messageId":"20181105181014.GA30777@tor.lan","threadId":"49734","inReplyTo":"CADN+U_MgrGHLQ5QNa-HgzxLN4zJLJPu4PaT2MTRoc18=gET+5Q@mail.gmail.com","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2018-11-05T18:10:14Z","receivedAt":"2018-11-05T18:10:19Z","isPatch":false,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On Mon, Nov 05, 2018 at 05:24:39AM +0100, Adrián Gimeno Balaguer wrote:\n\n[]\n\n> https://github.com/git/git/pull/550\n \n[]\n \n> This is covered in the mentioned PR above. Thanks for feedback.\n\nThanks for the code,\nI will have a look (the next days)\n\n> \n> -- \n> Adrián\n"},{"id":"362608","messageId":"20181106201618.GA30158@tor.lan","threadId":"49734","inReplyTo":"20181105181014.GA30777@tor.lan","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2018-11-06T20:16:18Z","receivedAt":"2018-11-06T20:16:22Z","isPatch":false,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On Mon, Nov 05, 2018 at 07:10:14PM +0100, Torsten Bögershausen wrote:\n> On Mon, Nov 05, 2018 at 05:24:39AM +0100, Adrián Gimeno Balaguer wrote:\n> \n> []\n> \n> > https://github.com/git/git/pull/550\n>  \n> []\n>  \n> > This is covered in the mentioned PR above. Thanks for feedback.\n> \n> Thanks for the code,\n> I will have a look (the next days)\n> \n> > \n> > -- \n> > Adrián\n\nHej Adrián,\n\nI still didn't manage to fully understand your problem.\nI tried to convert your test into my understanding,\nIt can be fetched here (or copied from this message, see below)\n\nhttps://github.com/tboegi/git/tree/tb.181106_UTF16LE_commit\n\nThe commit of an empty file seems to work for me, in the initial\nreport a \"rebase\" was mentioned, which is not in the TC ?\n\nIs the following what you intended to test ?\n\n#!/bin/sh\ntest_description='UTF-16 LE/BE file encoding using working-tree-encoding'\n\n\n. ./test-lib.sh\n\n# We specify the UTF-16LE BOM manually, to not depend on programs such as iconv.\nutf16leBOM=$(printf '\\377\\376')\n\ntest_expect_success 'Stage empty UTF-16LE file as binary' '\n\t>empty_0.txt &&\n\techo \"empty_0.txt binary\" >>.gitattributes &&\n\tgit add empty_0.txt\n'\n\n\ntest_expect_success 'Stage empty file with enc=UTF.16BL' '\n\t>utf16le_0.txt &&\n\techo \"utf16le_0.txt text working-tree-encoding=UTF-16BE\" >>.gitattributes &&\n\tgit add utf16le_0.txt\n'\n\n\ntest_expect_success 'Create and stage UTF-16LE file with only BOM' '\n\tprintf \"$utf16leBOM\" >utf16le_1.txt &&\n\techo \"utf16le_1.txt text working-tree-encoding=UTF-16\" >>.gitattributes &&\n\tgit add utf16le_1.txt\n'\n\ntest_expect_success 'Dont stage UTF-16LE file with only BOM with enc=UTF.16BE' '\n\tprintf \"$utf16leBOM\" >utf16le_2.txt &&\n\techo \"utf16le_2.txt text working-tree-encoding=UTF-16BE\" >>.gitattributes &&\n\ttest_must_fail git add utf16le_2.txt\n'\n\ntest_expect_success 'commit all files' '\n\ttest_tick &&\n\tgit commit -m \"Commit all 3 files\"\n'\n\ntest_expect_success 'All commited files have the same sha' '\n\tgit ls-files -s --eol >tmp1 &&\n\tsed -e \"s!\ti/none.*!!\" <tmp1 | uniq -u >actual &&\n\t>expect &&\n\ttest_cmp expect actual\n'\n\ntest_done\n"},{"id":"362630","messageId":"CADN+U_N345aMaiN4CT-_qsecw2gv=8-r+Hqq+CNz-xOx2KGYzg@mail.gmail.com","threadId":"49734","inReplyTo":"20181106201618.GA30158@tor.lan","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"Adrián Gimeno Balaguer","fromEmail":"adrigibal@gmail.com","sentAt":"2018-11-07T04:38:18Z","receivedAt":"2018-11-07T04:38:32Z","isPatch":false,"sender":{"key":"adrigibal@gmail.com","avatar":null},"body":"Hello Torsten,\n\nThanks for answering.\n\nAnswering to your question, I removed the comments with \"rebase\" since\nmy reported encoding issue happens on more simpler operations\n(described in the PR), and the problem is not directly related to\nrebasing, so I considered it better in order to avoid unrelated\nconfusions.\n\nLet's get back to the problem. Each system has a default endianness.\nAlso, in .gitattributes's working-tree-encoding, Git behaves\ndifferently depending on the attribute's value and the contents of the\nreferenced entry file. When I put the value \"UTF-16\", then the file\nmust have a BOM, or Git complains. Otherwise, if I put the value\n\"UTF-16BE\" or \"UTF-16LE\", then Git prohibites operations if file has a\nBOM for that main encoding (UTF-16 here), which can be relate to any\nendianness.\n\nMy very initial goal was, given a UTF-16LE file, to be able to view\nhuman-readable diffs whenever I make a change on it (and yes, it must\nbe Little Endian). Plus, this file had a BOM. Now, what are the\noptions with Git currently (consider only working-tree-encoding)? If I\nput working-tree-encoding=UTF-16, then I could view readable diffs and\ncommit the file, but here is the main problem: Git looses information\nabout what initial endianness the file had, therefore, after\nstaging/committing it re-encodes the file from UTF-8 (as stored\ninternally) to UTF-16 and the default system endianness. In my case it\ndid to Big Endian, thus affecting the project's requirement. That is\nwhy I ended up writing a fixup script to change the encoding back to\nUTF-16LE.\n\nOn the other hand, once I set working-tree-encoding=UTF-16LE, then Git\nprohibited me from committing the file and even viewing human-readable\ndiffs (the output simply tells it's a binary file). In this sense, the\ninternal location of these  errors is within the function of utf8.c I\nmade changes to in the PR. I hope I was clearer!\n\nFinally, Git behaviour around this is based on Unicode standards,\nwhich is why I acknowledged that my changes violated them after\nrefering to a link which is present in the ut8.h file.\nEl mar., 6 nov. 2018 a las 21:16, Torsten Bögershausen\n(<tboegi@web.de>) escribió:\n>\n> On Mon, Nov 05, 2018 at 07:10:14PM +0100, Torsten Bögershausen wrote:\n> > On Mon, Nov 05, 2018 at 05:24:39AM +0100, Adrián Gimeno Balaguer wrote:\n> >\n> > []\n> >\n> > > https://github.com/git/git/pull/550\n> >\n> > []\n> >\n> > > This is covered in the mentioned PR above. Thanks for feedback.\n> >\n> > Thanks for the code,\n> > I will have a look (the next days)\n> >\n> > >\n> > > --\n> > > Adrián\n>\n> Hej Adrián,\n>\n> I still didn't manage to fully understand your problem.\n> I tried to convert your test into my understanding,\n> It can be fetched here (or copied from this message, see below)\n>\n> https://github.com/tboegi/git/tree/tb.181106_UTF16LE_commit\n>\n> The commit of an empty file seems to work for me, in the initial\n> report a \"rebase\" was mentioned, which is not in the TC ?\n>\n> Is the following what you intended to test ?\n>\n> #!/bin/sh\n> test_description='UTF-16 LE/BE file encoding using working-tree-encoding'\n>\n>\n> . ./test-lib.sh\n>\n> # We specify the UTF-16LE BOM manually, to not depend on programs such as iconv.\n> utf16leBOM=$(printf '\\377\\376')\n>\n> test_expect_success 'Stage empty UTF-16LE file as binary' '\n>         >empty_0.txt &&\n>         echo \"empty_0.txt binary\" >>.gitattributes &&\n>         git add empty_0.txt\n> '\n>\n>\n> test_expect_success 'Stage empty file with enc=UTF.16BL' '\n>         >utf16le_0.txt &&\n>         echo \"utf16le_0.txt text working-tree-encoding=UTF-16BE\" >>.gitattributes &&\n>         git add utf16le_0.txt\n> '\n>\n>\n> test_expect_success 'Create and stage UTF-16LE file with only BOM' '\n>         printf \"$utf16leBOM\" >utf16le_1.txt &&\n>         echo \"utf16le_1.txt text working-tree-encoding=UTF-16\" >>.gitattributes &&\n>         git add utf16le_1.txt\n> '\n>\n> test_expect_success 'Dont stage UTF-16LE file with only BOM with enc=UTF.16BE' '\n>         printf \"$utf16leBOM\" >utf16le_2.txt &&\n>         echo \"utf16le_2.txt text working-tree-encoding=UTF-16BE\" >>.gitattributes &&\n>         test_must_fail git add utf16le_2.txt\n> '\n>\n> test_expect_success 'commit all files' '\n>         test_tick &&\n>         git commit -m \"Commit all 3 files\"\n> '\n>\n> test_expect_success 'All commited files have the same sha' '\n>         git ls-files -s --eol >tmp1 &&\n>         sed -e \"s!      i/none.*!!\" <tmp1 | uniq -u >actual &&\n>         >expect &&\n>         test_cmp expect actual\n> '\n>\n> test_done\n\n\n\n-- \nAdrián\n"},{"id":"362742","messageId":"20181108170230.GA6652@tor.lan","threadId":"49734","inReplyTo":"CADN+U_N345aMaiN4CT-_qsecw2gv=8-r+Hqq+CNz-xOx2KGYzg@mail.gmail.com","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2018-11-08T17:02:30Z","receivedAt":"2018-11-08T17:07:38Z","isPatch":false,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On Wed, Nov 07, 2018 at 05:38:18AM +0100, Adrián Gimeno Balaguer wrote:\n> Hello Torsten,\n> \n> Thanks for answering.\n> \n> Answering to your question, I removed the comments with \"rebase\" since\n> my reported encoding issue happens on more simpler operations\n> (described in the PR), and the problem is not directly related to\n> rebasing, so I considered it better in order to avoid unrelated\n> confusions.\n> \n> Let's get back to the problem. Each system has a default endianness.\n> Also, in .gitattributes's working-tree-encoding, Git behaves\n> differently depending on the attribute's value and the contents of the\n> referenced entry file. When I put the value \"UTF-16\", then the file\n> must have a BOM, or Git complains. Otherwise, if I put the value\n> \"UTF-16BE\" or \"UTF-16LE\", then Git prohibites operations if file has a\n> BOM for that main encoding (UTF-16 here), which can be relate to any\n> endianness.\n> \n> My very initial goal was, given a UTF-16LE file, to be able to view\n> human-readable diffs whenever I make a change on it (and yes, it must\n> be Little Endian). Plus, this file had a BOM. Now, what are the\n> options with Git currently (consider only working-tree-encoding)? If I\n> put working-tree-encoding=UTF-16, then I could view readable diffs and\n> commit the file, but here is the main problem: Git looses information\n> about what initial endianness the file had, therefore, after\n> staging/committing it re-encodes the file from UTF-8 (as stored\n> internally) to UTF-16 and the default system endianness. In my case it\n> did to Big Endian, thus affecting the project's requirement. That is\n> why I ended up writing a fixup script to change the encoding back to\n> UTF-16LE.\n\nOK, I think I understand your problem now.\nThe file format which you ask for could be named \"UTF-16-BOM-LE\",\nbut that does not exist in reality.\nIf you use UTF-16, then there must be a BOM, and if there is a BOM,\nthen a Unicode-aware application -should- be able to handle it.\n\nWhy does your project require such a format ?\n\n> \n> On the other hand, once I set working-tree-encoding=UTF-16LE, then Git\n> prohibited me from committing the file and even viewing human-readable\n> diffs (the output simply tells it's a binary file). In this sense, the\n> internal location of these  errors is within the function of utf8.c I\n> made changes to in the PR. I hope I was clearer!\n> \n> Finally, Git behaviour around this is based on Unicode standards,\n> which is why I acknowledged that my changes violated them after\n> refering to a link which is present in the ut8.h file.\n\n[]\n"},{"id":"365760","messageId":"009e01d49ace$4876c230$d9644690$@gmail.com","threadId":"49734","inReplyTo":"20181104170729.GA21372@tor.lan","subject":"RE: git-rebase is ignoring working-tree-encoding","fromName":"Alexandre Grigoriev","fromEmail":"alegrigoriev@gmail.com","sentAt":"2018-12-23T14:46:25Z","receivedAt":"2018-12-23T14:46:25Z","isPatch":false,"sender":{"key":"alegrigoriev@gmail.com","avatar":null},"body":"\n\n>On Fri, Nov 02, 2018 at 03:30:17AM +0100, Adrián Gimeno Balaguer wrote:\n>> I’m attempting to perform fixups via git-rebase of UTF-16 LE files\n>> (the project I’m working on requires that exact encoding on certain\n>> files). When the rebase is complete, Git changes that file’s encoding\n>> to UTF-16 BE. I have been using the newer working-tree-encoding\n>> attribute in .gitattributes. I’m using Git for Windows.\n>> \n>> $ git version\n>> git version 2.19.1.windows.1\n>> \n\n>Thanks for the report.\n>I have tried to follow the problem from your verbal descriptions\n>(and the PR) but I need to admit that I don't fully understand the\n>problem (yet).\n>Could you try to create some instructions how to reproduce it?\n>A numer of shell istructions would be great,\n>in best case some kind of \"test case\", like the tests in\n>the t/ directory in Git.\n>It would be nice to be able to re-produce it.\n>And if there is a bug, to get it fixed.\n\nThis is not as much Git issue (and not rebase issue at all), as libiconv issue.\n\nIconv program exhibits the same behavior. If you ask it to convert to UTF-16,\nIt will produce UTF-16BE with BOM.\n\nThat said, it appears that Centos (tested on 7.4) devs have seen the wrong in it and patched libiconv to produce UTF-16LE with BOM.\nGit on Centos does check out files as UTF-16LE, and iconv produces these files, as well.\nJust need to find out what patch they applied to libiconv.\n\n\n\n"},{"id":"365806","messageId":"002201d49cb5$cc554160$64ffc420$@gmail.com","threadId":"49734","inReplyTo":"20181108170230.GA6652@tor.lan","subject":"RE: git-rebase is ignoring working-tree-encoding","fromName":"Alexandre Grigoriev","fromEmail":"alegrigoriev@gmail.com","sentAt":"2018-12-26T00:56:11Z","receivedAt":"2018-12-26T00:56:13Z","isPatch":false,"sender":{"key":"alegrigoriev@gmail.com","avatar":null},"body":"> -----Original Message-----\n> From: git-owner@vger.kernel.org [mailto:git-owner@vger.kernel.org] On\n> Behalf Of Torsten Bogershausen\n> Sent: Thursday, November 8, 2018 9:03 AM\n> To: Adrián Gimeno Balaguer\n> Cc: git@vger.kernel.org\n> Subject: Re: git-rebase is ignoring working-tree-encoding\n> \n> On Wed, Nov 07, 2018 at 05:38:18AM +0100, Adrián Gimeno Balaguer wrote:\n> > Hello Torsten,\n> >\n> > Thanks for answering.\n> >\n> > Answering to your question, I removed the comments with \"rebase\" since\n> > my reported encoding issue happens on more simpler operations\n> > (described in the PR), and the problem is not directly related to\n> > rebasing, so I considered it better in order to avoid unrelated\n> > confusions.\n> >\n\n> OK, I think I understand your problem now.\n> The file format which you ask for could be named \"UTF-16-BOM-LE\",\n> but that does not exist in reality.\n> If you use UTF-16, then there must be a BOM, and if there is a BOM,\n> then a Unicode-aware application -should- be able to handle it.\n> \n> Why does your project require such a format ?\n> \n\nMany tools in Windows still do not understand UTF-8, although it's getting\nbetter. I think Windows is about the only OS where tools still require\nUTF-16 for full internationalization.\nMany tools written in C use MSVC RTL, where fopen(), unfortunately, doesn't\nunderstand UTF-16BE (though such a rudimentary program as Notepad does).\n\nFor this reason, it's very reasonable to ask that the programming tools\nproduce UTF-16 files with particular endianness, natural for the platform\nthey're running on.\n\nThe iconv programmers' boneheaded decision to always produce UTF-16BE with\nBOM for UTF-16 output doesn't make sense.\nAgain, git and iconv/libiconv in Centos on x86 do the right thing and\nproduce UTF-16LE with BOM in this case.\n\nAlso, iconv/libiconv should not be rejecting files with BOM for input\nencoding UTF-16BE or UTF-16LE.\nThe BOM is not some magic tag. It's just a zero-width space, with unique\nproperty that its 8 and 16 bit encoding variants can be recognized one from\nanother. It can appear anywhere in a file.\nIf it's a first character in the file, then the file encoding can be\nreliably detected. But it's just a character, and iconv should be accepting\nsuch files as valid.\n\n"},{"id":"365814","messageId":"20181226192525.GB423984@genre.crustytoothpaste.net","threadId":"49734","inReplyTo":"002201d49cb5$cc554160$64ffc420$@gmail.com","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-12-26T19:25:25Z","receivedAt":"2018-12-26T19:26:16Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Tue, Dec 25, 2018 at 04:56:11PM -0800, Alexandre Grigoriev wrote:\n> Many tools in Windows still do not understand UTF-8, although it's getting\n> better. I think Windows is about the only OS where tools still require\n> UTF-16 for full internationalization.\n> Many tools written in C use MSVC RTL, where fopen(), unfortunately, doesn't\n> understand UTF-16BE (though such a rudimentary program as Notepad does).\n> \n> For this reason, it's very reasonable to ask that the programming tools\n> produce UTF-16 files with particular endianness, natural for the platform\n> they're running on.\n> \n> The iconv programmers' boneheaded decision to always produce UTF-16BE with\n> BOM for UTF-16 output doesn't make sense.\n> Again, git and iconv/libiconv in Centos on x86 do the right thing and\n> produce UTF-16LE with BOM in this case.\n\nA program which claims to support \"UTF-16\" must support both\nendiannesses, according to RFC 2781. A program writing UTF-16-LE must\nnot write a BOM at the beginning. I realize this is inconvenient, but\nthe bad behavior of some Windows programs doesn't mean that Git should\nignore interoperability with non-Windows systems using UTF-16 correctly\nin favor of Windows.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"365830","messageId":"005601d49d8f$45c109b0$d1431d10$@gmail.com","threadId":"49734","inReplyTo":"20181226192525.GB423984@genre.crustytoothpaste.net","subject":"RE: git-rebase is ignoring working-tree-encoding","fromName":"Alexandre Grigoriev","fromEmail":"alegrigoriev@gmail.com","sentAt":"2018-12-27T02:52:56Z","receivedAt":"2018-12-27T02:52:59Z","isPatch":false,"sender":{"key":"alegrigoriev@gmail.com","avatar":null},"body":"\n> -----Original Message-----\n> From: brian m. carlson [mailto:sandals@crustytoothpaste.net]\n> Sent: Wednesday, December 26, 2018 11:25 AM\n> To: Alexandre Grigoriev\n> Cc: 'Torsten Bögershausen'; 'Adrián Gimeno Balaguer'; git@vger.kernel.org\n> Subject: Re: git-rebase is ignoring working-tree-encoding\n> \n> On Tue, Dec 25, 2018 at 04:56:11PM -0800, Alexandre Grigoriev wrote:\n> > Many tools in Windows still do not understand UTF-8, although it's\n> > getting better. I think Windows is about the only OS where tools still\n> > require\n> > UTF-16 for full internationalization.\n> > Many tools written in C use MSVC RTL, where fopen(), unfortunately,\n> > doesn't understand UTF-16BE (though such a rudimentary program as\n> Notepad does).\n> >\n> > For this reason, it's very reasonable to ask that the programming\n> > tools produce UTF-16 files with particular endianness, natural for the\n> > platform they're running on.\n> >\n> > The iconv programmers' boneheaded decision to always produce UTF-16BE\n> > with BOM for UTF-16 output doesn't make sense.\n> > Again, git and iconv/libiconv in Centos on x86 do the right thing and\n> > produce UTF-16LE with BOM in this case.\n> \n> A program which claims to support \"UTF-16\" must support both\n> endiannesses, according to RFC 2781. A program writing UTF-16-LE must not\n> write a BOM at the beginning. I realize this is inconvenient, but the bad\n> behavior of some Windows programs doesn't mean that Git should ignore\n> interoperability with non-Windows systems using UTF-16 correctly in favor of\n> Windows.\n\nOK, we have a choice either:\na) to live in that corner of the real world where you have to use available tools, some of which have historical reasons\nto only support UTF-16LE with BOM, because nobody ever throws a different flavor of UTF-16 at them;\nOr b) to live in an ivory tower where you don't really need to use UTF-16 LE or BE or any other flavor,\nbecause everything is just UTF-8, and tell all those other people using that lame OS to shut up and wait until their tools start to support\nthe formats you don't really have to care about;\n\n> behavior of some Windows programs doesn't mean that Git should ignore\n> interoperability with non-Windows systems using UTF-16 correctly in favor of\n> Windows.\n\nYes, Git (actually libiconv) should not ignore interoperability.\nThis means it should check out files on a *Windows* system in a format which *Windows* tools\ncan understand.\nAnd, by the way, Centos (or RedHat?) developers understood that.\nThere, on an x86 installation, when you ask for UTF-16, it produces UTF-16LE with BOM.\nJust as every user there would want.\n\n\n"},{"id":"365840","messageId":"20181227144525.GA2467@tor.lan","threadId":"49734","inReplyTo":"005601d49d8f$45c109b0$d1431d10$@gmail.com","subject":"Re: git-rebase is ignoring working-tree-encoding","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2018-12-27T14:45:25Z","receivedAt":"2018-12-27T14:45:36Z","isPatch":false,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On Wed, Dec 26, 2018 at 06:52:56PM -0800, Alexandre Grigoriev wrote:\n> \n> > -----Original Message-----\n> > From: brian m. carlson [mailto:sandals@crustytoothpaste.net]\n> > Sent: Wednesday, December 26, 2018 11:25 AM\n> > To: Alexandre Grigoriev\n> > Cc: 'Torsten Bögershausen'; 'Adrián Gimeno Balaguer'; git@vger.kernel.org\n> > Subject: Re: git-rebase is ignoring working-tree-encoding\n> > \n> > On Tue, Dec 25, 2018 at 04:56:11PM -0800, Alexandre Grigoriev wrote:\n> > > Many tools in Windows still do not understand UTF-8, although it's\n> > > getting better. I think Windows is about the only OS where tools still\n> > > require\n> > > UTF-16 for full internationalization.\n> > > Many tools written in C use MSVC RTL, where fopen(), unfortunately,\n> > > doesn't understand UTF-16BE (though such a rudimentary program as\n> > Notepad does).\n> > >\n> > > For this reason, it's very reasonable to ask that the programming\n> > > tools produce UTF-16 files with particular endianness, natural for the\n> > > platform they're running on.\n> > >\n> > > The iconv programmers' boneheaded decision to always produce UTF-16BE\n> > > with BOM for UTF-16 output doesn't make sense.\n> > > Again, git and iconv/libiconv in Centos on x86 do the right thing and\n> > > produce UTF-16LE with BOM in this case.\n> > \n> > A program which claims to support \"UTF-16\" must support both\n> > endiannesses, according to RFC 2781. A program writing UTF-16-LE must not\n> > write a BOM at the beginning. I realize this is inconvenient, but the bad\n> > behavior of some Windows programs doesn't mean that Git should ignore\n> > interoperability with non-Windows systems using UTF-16 correctly in favor of\n> > Windows.\n> \n> OK, we have a choice either:\n> a) to live in that corner of the real world where you have to use available tools, some of which have historical reasons\n> to only support UTF-16LE with BOM, because nobody ever throws a different flavor of UTF-16 at them;\n> Or b) to live in an ivory tower where you don't really need to use UTF-16 LE or BE or any other flavor,\n> because everything is just UTF-8, and tell all those other people using that lame OS to shut up and wait until their tools start to support\n> the formats you don't really have to care about;\n> \n> > behavior of some Windows programs doesn't mean that Git should ignore\n> > interoperability with non-Windows systems using UTF-16 correctly in favor of\n> > Windows.\n> \n> Yes, Git (actually libiconv) should not ignore interoperability.\n> This means it should check out files on a *Windows* system in a format which *Windows* tools\n> can understand.\n> And, by the way, Centos (or RedHat?) developers understood that.\n> There, on an x86 installation, when you ask for UTF-16, it produces UTF-16LE with BOM.\n> Just as every user there would want.\n> \n> \n\nSorry if I feel confused here - does the problem still exist ?\nIf yes, does the following patch help ?\n\n\n\ndiff --git a/utf8.c b/utf8.c\nindex eb78587504..2facef84d4 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -9,6 +9,23 @@ struct interval {\n \tucs_char_t last;\n };\n \n+static int has_bom_prefix(const char *data, size_t len,\n+\t\t\t  const char *bom, size_t bom_len)\n+{\n+\treturn data && bom && (len >= bom_len) && !memcmp(data, bom, bom_len);\n+}\n+\n+static const char utf16_be_bom[] = {'\\xFE', '\\xFF'};\n+static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n+static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n+static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n+\n+static inline uint16_t default_swab16(uint16_t val)\n+{\n+\treturn (((val & 0xff00) >>  8) |\n+\t\t((val & 0x00ff) <<  8));\n+}\n+\n size_t display_mode_esc_sequence_len(const char *s)\n {\n \tconst char *p = s;\n@@ -556,21 +573,19 @@ char *reencode_string_len(const char *in, size_t insz,\n \n \tout = reencode_string_iconv(in, insz, conv, outsz);\n \ticonv_close(conv);\n+\tif (has_bom_prefix(out, *outsz, utf16_be_bom, sizeof(utf16_be_bom))) {\n+\t\t/* UTF-16 should be little endian under Git */\n+\t\tsize_t    num_points = *outsz / sizeof(uint16_t);\n+\t\tuint16_t *point = (uint16_t*) out;\n+\t\twhile (num_points--) {\n+\t\t\t*point = default_swab16(*point);\n+\t\t\tpoint++;\n+\t\t}\n+\t}\n \treturn out;\n }\n #endif\n \n-static int has_bom_prefix(const char *data, size_t len,\n-\t\t\t  const char *bom, size_t bom_len)\n-{\n-\treturn data && bom && (len >= bom_len) && !memcmp(data, bom, bom_len);\n-}\n-\n-static const char utf16_be_bom[] = {'\\xFE', '\\xFF'};\n-static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n-static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n-static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n-\n int has_prohibited_utf_bom(const char *enc, const char *data, size_t len)\n {\n \treturn (\n\n\n"},{"id":"365946","messageId":"20181229110924.26598-1-tboegi@web.de","threadId":"49734","inReplyTo":"CADN+U_PUfnYWb-wW6drRANv-ZaYBEk3gWHc7oJtxohA5Vc3NEg@mail.gmail.com","subject":"[PATCH/RFC v1 1/1] Support working-tree-encoding \"UTF-16LE-BOM\"","fromName":"","fromEmail":"tboegi@web.de","sentAt":"2018-12-29T11:09:24Z","receivedAt":"2018-12-29T11:10:07Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"From: Torsten Bögershausen <tboegi@web.de>\n\nUsers who want UTF-16 files in the working tree set the .gitattributes\nlike this:\ntest.txt working-tree-encoding=UTF-16\n\nAfter a checkout, the resulting file has a BOM and is encoded in \"UTF-16\".\nThe unicode standard allows both little- and big-endianess (LE/BE) for\nthose files, the BOM will tell which one is used inside the file.\niconv seems to prefer the BE version.\nNot all users under Windows are happy with this when tools are not fully\nunicode aware and don't digest the BE version at all.\n\nToday there is no name for \"UTF-16 with BOM, little endian please\".\nIntroduce \"UTF-16LE-BOM\".\n\nRported-by: Adrián Gimeno Balaguer <adrigibal@gmail.com>\nSigned-off-by: Torsten Bögershausen <tboegi@web.de>\n---\n\nThis feels like an RFC at the moment - please comment.\nUsing UTF-16 in the way \"UTF-16LE-BOM\" is used in this patch\ncould be an alternative - simply produce UTF-16 in LE version\nunder Git - this could make people using Git happy as well.\n\nDocumentation/gitattributes.txt  |  4 +--\n compat/precompose_utf8.c         |  2 +-\n t/t0028-working-tree-encoding.sh | 12 ++++++++-\n utf8.c                           | 42 ++++++++++++++++++++++++--------\n utf8.h                           |  2 +-\n 5 files changed, 47 insertions(+), 15 deletions(-)\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex b8392fc330..4a88ab8be7 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -343,13 +343,13 @@ automatic line ending conversion based on your platform.\n ------------------------\n \n Use the following attributes if your '*.ps1' files are UTF-16 little\n-endian encoded without BOM and you want Git to use Windows line endings\n+endian encoded with BOM and you want Git to use Windows line endings\n in the working directory. Please note, it is highly recommended to\n explicitly define the line endings with `eol` if the `working-tree-encoding`\n attribute is used to avoid ambiguity.\n \n ------------------------\n-*.ps1\t\ttext working-tree-encoding=UTF-16LE eol=CRLF\n+*.ps1\t\ttext working-tree-encoding=UTF-16LE-BOM eol=CRLF\n ------------------------\n \n You can get a list of all available encodings on your platform with the\ndiff --git a/compat/precompose_utf8.c b/compat/precompose_utf8.c\nindex de61c15d34..136250fbf6 100644\n--- a/compat/precompose_utf8.c\n+++ b/compat/precompose_utf8.c\n@@ -79,7 +79,7 @@ void precompose_argv(int argc, const char **argv)\n \t\tsize_t namelen;\n \t\toldarg = argv[i];\n \t\tif (has_non_ascii(oldarg, (size_t)-1, &namelen)) {\n-\t\t\tnewarg = reencode_string_iconv(oldarg, namelen, ic_precompose, NULL);\n+\t\t\tnewarg = reencode_string_iconv(oldarg, namelen, ic_precompose, 0, NULL);\n \t\t\tif (newarg)\n \t\t\t\targv[i] = newarg;\n \t\t}\ndiff --git a/t/t0028-working-tree-encoding.sh b/t/t0028-working-tree-encoding.sh\nindex 7e87b5a200..e58ecbfc44 100755\n--- a/t/t0028-working-tree-encoding.sh\n+++ b/t/t0028-working-tree-encoding.sh\n@@ -11,9 +11,12 @@ test_expect_success 'setup test files' '\n \n \ttext=\"hallo there!\\ncan you read me?\" &&\n \techo \"*.utf16 text working-tree-encoding=utf-16\" >.gitattributes &&\n+\techo \"*.utf16lebom text working-tree-encoding=UTF-16LE-BOM\" >>.gitattributes &&\n \tprintf \"$text\" >test.utf8.raw &&\n \tprintf \"$text\" | iconv -f UTF-8 -t UTF-16 >test.utf16.raw &&\n \tprintf \"$text\" | iconv -f UTF-8 -t UTF-32 >test.utf32.raw &&\n+\tprintf \"\\377\\376\"                         >test.utf16lebom.raw &&\n+\tprintf \"$text\" | iconv -f UTF-8 -t UTF-32LE >>test.utf16lebom.raw &&\n \n \t# Line ending tests\n \tprintf \"one\\ntwo\\nthree\\n\" >lf.utf8.raw &&\n@@ -32,7 +35,8 @@ test_expect_success 'setup test files' '\n \t# Add only UTF-16 file, we will add the UTF-32 file later\n \tcp test.utf16.raw test.utf16 &&\n \tcp test.utf32.raw test.utf32 &&\n-\tgit add .gitattributes test.utf16 &&\n+\tcp test.utf16lebom.raw test.utf16lebom &&\n+\tgit add .gitattributes test.utf16 test.utf16lebom &&\n \tgit commit -m initial\n '\n \n@@ -51,6 +55,12 @@ test_expect_success 're-encode to UTF-16 on checkout' '\n \ttest_cmp_bin test.utf16.raw test.utf16\n '\n \n+test_expect_success 're-encode to UTF-16-LE-BOM on checkout' '\n+\trm test.utf16lebom &&\n+\tgit checkout test.utf16lebom &&\n+\ttest_cmp_bin test.utf16lebom.raw test.utf16lebom\n+'\n+\n test_expect_success 'check $GIT_DIR/info/attributes support' '\n \ttest_when_finished \"rm -f test.utf32.git\" &&\n \ttest_when_finished \"git reset --hard HEAD\" &&\ndiff --git a/utf8.c b/utf8.c\nindex eb78587504..83824dc2f4 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -4,6 +4,11 @@\n \n /* This code is originally from http://www.cl.cam.ac.uk/~mgk25/ucs/ */\n \n+static const char utf16_be_bom[] = {'\\xFE', '\\xFF'};\n+static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n+static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n+static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n+\n struct interval {\n \tucs_char_t first;\n \tucs_char_t last;\n@@ -470,16 +475,17 @@ int utf8_fprintf(FILE *stream, const char *format, ...)\n #else\n \ttypedef char * iconv_ibp;\n #endif\n-char *reencode_string_iconv(const char *in, size_t insz, iconv_t conv, size_t *outsz_p)\n+char *reencode_string_iconv(const char *in, size_t insz, iconv_t conv,\n+\t\t\t    size_t bom_len, size_t *outsz_p)\n {\n \tsize_t outsz, outalloc;\n \tchar *out, *outpos;\n \ticonv_ibp cp;\n \n \toutsz = insz;\n-\toutalloc = st_add(outsz, 1); /* for terminating NUL */\n+\toutalloc = st_add(outsz, 1 + bom_len); /* for terminating NUL */\n \tout = xmalloc(outalloc);\n-\toutpos = out;\n+\toutpos = out + bom_len;\n \tcp = (iconv_ibp)in;\n \n \twhile (1) {\n@@ -540,10 +546,30 @@ char *reencode_string_len(const char *in, size_t insz,\n {\n \ticonv_t conv;\n \tchar *out;\n+\tconst char *bom_str = NULL;\n+\tsize_t bom_len = 0;\n \n \tif (!in_encoding)\n \t\treturn NULL;\n \n+\t/* UTF-16LE-BOM is the same as UTF-16 for reading */\n+\tif (same_utf_encoding(\"UTF-16LE-BOM\", in_encoding))\n+\t\tin_encoding = \"UTF-16\";\n+\n+\t/*\n+\t * For writing, UTF-16 iconv typically creates \"UTF-16BE-BOM\"\n+\t * Some users under Windows want the little endian version\n+\t */\n+\tif (same_utf_encoding(\"UTF-16LE-BOM\", out_encoding)) {\n+\t\tbom_str = utf16_le_bom;\n+\t\tbom_len = sizeof(utf16_le_bom);\n+\t\tout_encoding = \"UTF-16LE\";\n+\t} else if (same_utf_encoding(\"UTF-16BE-BOM\", out_encoding)) {\n+\t\tbom_str = utf16_be_bom;\n+\t\tbom_len = sizeof(utf16_be_bom);\n+\t\tout_encoding = \"UTF-16BE\";\n+\t}\n+\n \tconv = iconv_open(out_encoding, in_encoding);\n \tif (conv == (iconv_t) -1) {\n \t\tin_encoding = fallback_encoding(in_encoding);\n@@ -553,9 +579,10 @@ char *reencode_string_len(const char *in, size_t insz,\n \t\tif (conv == (iconv_t) -1)\n \t\t\treturn NULL;\n \t}\n-\n-\tout = reencode_string_iconv(in, insz, conv, outsz);\n+\tout = reencode_string_iconv(in, insz, conv, bom_len, outsz);\n \ticonv_close(conv);\n+\tif (out && bom_str && bom_len)\n+\t\tmemcpy(out, bom_str, bom_len);\n \treturn out;\n }\n #endif\n@@ -566,11 +593,6 @@ static int has_bom_prefix(const char *data, size_t len,\n \treturn data && bom && (len >= bom_len) && !memcmp(data, bom, bom_len);\n }\n \n-static const char utf16_be_bom[] = {'\\xFE', '\\xFF'};\n-static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n-static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n-static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n-\n int has_prohibited_utf_bom(const char *enc, const char *data, size_t len)\n {\n \treturn (\ndiff --git a/utf8.h b/utf8.h\nindex edea55e093..84efbfcb1f 100644\n--- a/utf8.h\n+++ b/utf8.h\n@@ -27,7 +27,7 @@ void strbuf_utf8_replace(struct strbuf *sb, int pos, int width,\n \n #ifndef NO_ICONV\n char *reencode_string_iconv(const char *in, size_t insz,\n-\t\t\t    iconv_t conv, size_t *outsz);\n+\t\t\t    iconv_t conv, size_t bom_len, size_t *outsz);\n char *reencode_string_len(const char *in, size_t insz,\n \t\t\t  const char *out_encoding,\n \t\t\t  const char *in_encoding,\n-- \n2.20.1.2.gb21ebb671b\n\n"},{"id":"365949","messageId":"CADN+U_Mo4Ui-rmZe1+xoHOMA4koXGNpJ5XEGYoYZfYPGqP9VPQ@mail.gmail.com","threadId":"49734","inReplyTo":"CADN+U_OccLuLN7_0rjikDgLT+Zvt8hka-=xsnVVLJORjYzP78Q@mail.gmail.com","subject":"Re: [PATCH/RFC v1 1/1] Support working-tree-encoding \"UTF-16LE-BOM\"","fromName":"Adrián Gimeno Balaguer","fromEmail":"adrigibal@gmail.com","sentAt":"2018-12-29T15:48:25Z","receivedAt":"2018-12-29T15:48:39Z","isPatch":true,"sender":{"key":"adrigibal@gmail.com","avatar":null},"body":"Hello again.\n\nI appreciate the grown interest in this issue.\n\nTorsten, may I know what is the benefit on your code? My PR solved it\nby only tweaking the utf8.c's function 'has_prohibited_utf_bom', which\nis likely the shortest way:\n\nhttps://github.com/git/git/pull/550/files\n\nIn order to make sure everything is clear, here is a case list of\ncurrent Git behaviour and new one after my PR, regarding this issue.\n\nCurrent behaviour:\n\n- Placing 'test.txt working-tree-encoding=UTF-16' for a new test.txt\nfile with either UTF-16 BE or LE BOM, and comitting everything -> The\nfile gets re-encoded from UTF-8 (as stored internally), to UTF-16 and\nthe default system/libiconv endianness -> Problem (as long as user\nrequired the opposite endianness for any reason on his project). As a\nnote, user can see however human-readable diffs on that file.\n\n- Placing  'test.txt working-tree-encoding=UTF-16LE' or 'test.txt\nworking-tree-encoding=UTF-16BE' for a new test.txt file with either\nUTF-16 BE or LE BOM, and comitting everything: we assume user is doing\nthis because he requires that exact endianness, thus he writes it in\norder to attempt preserving it -> Git prohibites commiting it, also no\nhuman-readable diff is shown in the diff viewer/tool being used, but\nfile is simply shown as binary.\n\nNew behaviour:\n\n-  Just got too lazy to repeat it all over, read my PR description:\nhttps://github.com/git/git/pull/550\n\n- Git translations may need to be tweaked to in order to be consistent\nwith new behaviour.\n\nThanks for your attention.\n"},{"id":"365967","messageId":"27bb049b-265f-7fec-06ef-5e8e29c2d7f7@talktalk.net","threadId":"49734","inReplyTo":"CADN+U_Mo4Ui-rmZe1+xoHOMA4koXGNpJ5XEGYoYZfYPGqP9VPQ@mail.gmail.com","subject":"Re: [PATCH/RFC v1 1/1] Support working-tree-encoding \"UTF-16LE-BOM\"","fromName":"Philip Oakley","fromEmail":"philipoakley@talktalk.net","sentAt":"2018-12-29T17:54:49Z","receivedAt":"2018-12-29T17:54:53Z","isPatch":true,"sender":{"key":"philipoakley@talktalk.net","avatar":null},"body":"(adding Brian as cc who was in the original thread)\n\nOn 29/12/2018 15:48, Adrián Gimeno Balaguer wrote:\n> Hello again.\n> \n> I appreciate the grown interest in this issue.\n> \n> Torsten, may I know what is the benefit on your code? My PR solved it\n> by only tweaking the utf8.c's function 'has_prohibited_utf_bom', which\n> is likely the shortest way:\n> \n> https://github.com/git/git/pull/550/files\n\nMy main complaint with the PR would be the lack of documentation updates.\n\nAs the discussion has highlighted, whatever our solution, we will need \nto tell the users in plain and simple terms which parts of which \nstandards are being used, and why we need to be somehow 'different'.\n\nThat is because a revision control system must be able to recover the \noriginal, for use in the original software tool, not just interpret it \nis some alternate form. The standards generally abdicate responsibility \nfor the last step ;-)\n\nI did not fully understand the conversion process you proposed, as I \nassumed(?) that on receipt of the source file, the Git conversion to \nutf-8 would convert the 16-bit BOM to the three byte utf-8 BOM byte \nsequence `EF BB BF` which has lost any knowledge of the original BE/LE \ncoding.\n\nOr, are we saying that the the 16-bit BOM is being interpreted as, a) \nthe BE/LE indicator and b) a genuine \"ZERO WIDTH NON-BREAKING SPACE\" \nwhich is stored as the two byte utf-8 character code, again loosing \n(once stored as a blob object) the BE/LE indication.\n\nOr, we see the BOM, note the endianness and then loose the BOM character \nwhen converting to utf-8. My ignorance of this step is starting to show. \nRegular users are probably even more confused, hence my hope for some \ndocumentation.\n\nGiven the above confusions, and many more when exploring the internet, \nthe provision of a new, extra, clear, name for the encoding, as \nsuggested by Torsten does offer an advantage in that it explicitly \n(rather than implicitly) makes plain what we are trying to do, without \nsqueezing it in 'under the radar'.\n\nThat said, assuming an appropriate internal utf-8 Git coding that does \nremember the BE/LE state [if so how?] then the PR is a neat trick.\n\nTorsten's patch also suffers from the lack of user facing documentation.\n\n> \n> In order to make sure everything is clear, here is a case list of\n> current Git behaviour and new one after my PR, regarding this issue.\n> \n> Current behaviour:\n> \n> - Placing 'test.txt working-tree-encoding=UTF-16' for a new test.txt\n> file with either UTF-16 BE or LE BOM, and comitting everything -> The\n> file gets re-encoded from UTF-8 (as stored internally), to UTF-16 and\n> the default system/libiconv endianness -> Problem (as long as user\n> required the opposite endianness for any reason on his project). As a\n> note, user can see however human-readable diffs on that file.\n> \n> - Placing  'test.txt working-tree-encoding=UTF-16LE' or 'test.txt\n> working-tree-encoding=UTF-16BE' for a new test.txt file with either\n> UTF-16 BE or LE BOM, and comitting everything: we assume user is doing\n> this because he requires that exact endianness, thus he writes it in\n> order to attempt preserving it -> Git prohibites commiting it, also no\n> human-readable diff is shown in the diff viewer/tool being used, but\n> file is simply shown as binary.\n> \n> New behaviour:\n> \n> -  Just got too lazy to repeat it all over, read my PR description:\n> https://github.com/git/git/pull/550\n\n\"In this PR: Git only prohibites the opposite BOM than the one in \nworking-tree-encoding (e.g. if declared LE, then it denies BE BOM \npresence within the associated file, of the declared UTF-16/UTF-32). \nThis way the user can now make Git operations which were previously \nimpossible, with the only requisite being to match the endianness of \nworking-tree-encoding attribute with the associated file/s.\"\n\n> \n> - Git translations may need to be tweaked to in order to be consistent\n> with new behaviour.\n> \n> Thanks for your attention.\n> \n-- \nPhilip\n"},{"id":"367200","messageId":"20190120164327.3234-1-tboegi@web.de","threadId":"49734","inReplyTo":"CADN+U_PUfnYWb-wW6drRANv-ZaYBEk3gWHc7oJtxohA5Vc3NEg@mail.gmail.com","subject":"[PATCH v2 1/1] Support working-tree-encoding \"UTF-16LE-BOM\"","fromName":"","fromEmail":"tboegi@web.de","sentAt":"2019-01-20T16:43:27Z","receivedAt":"2019-01-20T16:45:04Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"From: Torsten Bögershausen <tboegi@web.de>\n\nUsers who want UTF-16 files in the working tree set the .gitattributes\nlike this:\ntest.txt working-tree-encoding=UTF-16\n\nThe unicode standard itself defines 3 possible ways how to encode UTF-16.\nThe following 3 versions convert all back to 'g' 'i' 't' in UTF-8:\n\na) UTF-16, without BOM, big endian:\n$ printf \"\\000g\\000i\\000t\" | iconv -f UTF-16 -t UTF-8 | od -c\n0000000    g   i   t\n\nb) UTF-16, with BOM, little endian:\n$ printf \"\\377\\376g\\000i\\000t\\000\" | iconv -f UTF-16 -t UTF-8 | od -c\n0000000    g   i   t\n\nc) UTF-16, with BOM, big endian:\n$ printf \"\\376\\377\\000g\\000i\\000t\" | iconv -f UTF-16 -t UTF-8 | od -c\n0000000    g   i   t\n\nGit uses libiconv to convert from UTF-8 in the index into ITF-16 in the\nworking tree.\nAfter a checkout, the resulting file has a BOM and is encoded in \"UTF-16\",\nin the version (c) above.\nThis is what iconv generates, more details follow below.\n\niconv (and libiconv) can generate UTF-16, UTF-16LE or UTF-16BE:\n\nd) UTF-16\n$ printf 'git' | iconv -f UTF-8 -t UTF-16 | od -c\n0000000  376 377  \\0   g  \\0   i  \\0   t\n\ne) UTF-16LE\n$ printf 'git' | iconv -f UTF-8 -t UTF-16LE | od -c\n0000000    g  \\0   i  \\0   t  \\0\n\nf)  UTF-16BE\n$ printf 'git' | iconv -f UTF-8 -t UTF-16BE | od -c\n0000000   \\0   g  \\0   i  \\0   t\n\nThere is no way to generate version (b) from above in a Git working tree,\nbut that is what some applications need.\n(All fully unicode aware applications should be able to read all 3 variants,\nbut in practise we are not there yet).\n\nWhen producing UTF-16 as an output, iconv generates the big endian version\nwith a BOM. (big endian is probably chosen for historical reasons).\n\niconv can produce UTF-16 files with little endianess by using \"UTF-16LE\"\nas encoding, and that file does not have a BOM.\n\nNot all users (especially under Windows) are happy with this.\nSome tools are not fully unicode aware and can only handle version (b).\n\nToday there is no way to produce version (b) with iconv (or libiconv).\nLooking into the history of iconv, it seems as if version (c) will\nbe used in all future iconv versions (for compatibility reasons).\n\nSolve this dilemma and introduce a Git-specific \"UTF-16LE-BOM\".\nlibiconv can not handle the encoding, so Git pick it up, handles the BOM\nand uses libiconv to convert the rest of the stream.\n\nRported-by: Adrián Gimeno Balaguer <adrigibal@gmail.com>\nSigned-off-by: Torsten Bögershausen <tboegi@web.de>\n---\n\nI still think it makes sense to support  UTF-16, little endian and\nwith BOM in Git.\nThis V2 should make more clear, what standards we follow, and why\nthe naming scheme of Unicode does not cover all use cases in real world.\n\n Documentation/gitattributes.txt  |  4 +--\n compat/precompose_utf8.c         |  2 +-\n t/t0028-working-tree-encoding.sh | 12 ++++++++-\n utf8.c                           | 42 ++++++++++++++++++++++++--------\n utf8.h                           |  2 +-\n 5 files changed, 47 insertions(+), 15 deletions(-)\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex b8392fc330..4a88ab8be7 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -343,13 +343,13 @@ automatic line ending conversion based on your platform.\n ------------------------\n\n Use the following attributes if your '*.ps1' files are UTF-16 little\n-endian encoded without BOM and you want Git to use Windows line endings\n+endian encoded with BOM and you want Git to use Windows line endings\n in the working directory. Please note, it is highly recommended to\n explicitly define the line endings with `eol` if the `working-tree-encoding`\n attribute is used to avoid ambiguity.\n\n ------------------------\n-*.ps1\t\ttext working-tree-encoding=UTF-16LE eol=CRLF\n+*.ps1\t\ttext working-tree-encoding=UTF-16LE-BOM eol=CRLF\n ------------------------\n\n You can get a list of all available encodings on your platform with the\ndiff --git a/compat/precompose_utf8.c b/compat/precompose_utf8.c\nindex de61c15d34..136250fbf6 100644\n--- a/compat/precompose_utf8.c\n+++ b/compat/precompose_utf8.c\n@@ -79,7 +79,7 @@ void precompose_argv(int argc, const char **argv)\n \t\tsize_t namelen;\n \t\toldarg = argv[i];\n \t\tif (has_non_ascii(oldarg, (size_t)-1, &namelen)) {\n-\t\t\tnewarg = reencode_string_iconv(oldarg, namelen, ic_precompose, NULL);\n+\t\t\tnewarg = reencode_string_iconv(oldarg, namelen, ic_precompose, 0, NULL);\n \t\t\tif (newarg)\n \t\t\t\targv[i] = newarg;\n \t\t}\ndiff --git a/t/t0028-working-tree-encoding.sh b/t/t0028-working-tree-encoding.sh\nindex 7e87b5a200..e58ecbfc44 100755\n--- a/t/t0028-working-tree-encoding.sh\n+++ b/t/t0028-working-tree-encoding.sh\n@@ -11,9 +11,12 @@ test_expect_success 'setup test files' '\n\n \ttext=\"hallo there!\\ncan you read me?\" &&\n \techo \"*.utf16 text working-tree-encoding=utf-16\" >.gitattributes &&\n+\techo \"*.utf16lebom text working-tree-encoding=UTF-16LE-BOM\" >>.gitattributes &&\n \tprintf \"$text\" >test.utf8.raw &&\n \tprintf \"$text\" | iconv -f UTF-8 -t UTF-16 >test.utf16.raw &&\n \tprintf \"$text\" | iconv -f UTF-8 -t UTF-32 >test.utf32.raw &&\n+\tprintf \"\\377\\376\"                         >test.utf16lebom.raw &&\n+\tprintf \"$text\" | iconv -f UTF-8 -t UTF-32LE >>test.utf16lebom.raw &&\n\n \t# Line ending tests\n \tprintf \"one\\ntwo\\nthree\\n\" >lf.utf8.raw &&\n@@ -32,7 +35,8 @@ test_expect_success 'setup test files' '\n \t# Add only UTF-16 file, we will add the UTF-32 file later\n \tcp test.utf16.raw test.utf16 &&\n \tcp test.utf32.raw test.utf32 &&\n-\tgit add .gitattributes test.utf16 &&\n+\tcp test.utf16lebom.raw test.utf16lebom &&\n+\tgit add .gitattributes test.utf16 test.utf16lebom &&\n \tgit commit -m initial\n '\n\n@@ -51,6 +55,12 @@ test_expect_success 're-encode to UTF-16 on checkout' '\n \ttest_cmp_bin test.utf16.raw test.utf16\n '\n\n+test_expect_success 're-encode to UTF-16-LE-BOM on checkout' '\n+\trm test.utf16lebom &&\n+\tgit checkout test.utf16lebom &&\n+\ttest_cmp_bin test.utf16lebom.raw test.utf16lebom\n+'\n+\n test_expect_success 'check $GIT_DIR/info/attributes support' '\n \ttest_when_finished \"rm -f test.utf32.git\" &&\n \ttest_when_finished \"git reset --hard HEAD\" &&\ndiff --git a/utf8.c b/utf8.c\nindex eb78587504..83824dc2f4 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -4,6 +4,11 @@\n\n /* This code is originally from http://www.cl.cam.ac.uk/~mgk25/ucs/ */\n\n+static const char utf16_be_bom[] = {'\\xFE', '\\xFF'};\n+static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n+static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n+static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n+\n struct interval {\n \tucs_char_t first;\n \tucs_char_t last;\n@@ -470,16 +475,17 @@ int utf8_fprintf(FILE *stream, const char *format, ...)\n #else\n \ttypedef char * iconv_ibp;\n #endif\n-char *reencode_string_iconv(const char *in, size_t insz, iconv_t conv, size_t *outsz_p)\n+char *reencode_string_iconv(const char *in, size_t insz, iconv_t conv,\n+\t\t\t    size_t bom_len, size_t *outsz_p)\n {\n \tsize_t outsz, outalloc;\n \tchar *out, *outpos;\n \ticonv_ibp cp;\n\n \toutsz = insz;\n-\toutalloc = st_add(outsz, 1); /* for terminating NUL */\n+\toutalloc = st_add(outsz, 1 + bom_len); /* for terminating NUL */\n \tout = xmalloc(outalloc);\n-\toutpos = out;\n+\toutpos = out + bom_len;\n \tcp = (iconv_ibp)in;\n\n \twhile (1) {\n@@ -540,10 +546,30 @@ char *reencode_string_len(const char *in, size_t insz,\n {\n \ticonv_t conv;\n \tchar *out;\n+\tconst char *bom_str = NULL;\n+\tsize_t bom_len = 0;\n\n \tif (!in_encoding)\n \t\treturn NULL;\n\n+\t/* UTF-16LE-BOM is the same as UTF-16 for reading */\n+\tif (same_utf_encoding(\"UTF-16LE-BOM\", in_encoding))\n+\t\tin_encoding = \"UTF-16\";\n+\n+\t/*\n+\t * For writing, UTF-16 iconv typically creates \"UTF-16BE-BOM\"\n+\t * Some users under Windows want the little endian version\n+\t */\n+\tif (same_utf_encoding(\"UTF-16LE-BOM\", out_encoding)) {\n+\t\tbom_str = utf16_le_bom;\n+\t\tbom_len = sizeof(utf16_le_bom);\n+\t\tout_encoding = \"UTF-16LE\";\n+\t} else if (same_utf_encoding(\"UTF-16BE-BOM\", out_encoding)) {\n+\t\tbom_str = utf16_be_bom;\n+\t\tbom_len = sizeof(utf16_be_bom);\n+\t\tout_encoding = \"UTF-16BE\";\n+\t}\n+\n \tconv = iconv_open(out_encoding, in_encoding);\n \tif (conv == (iconv_t) -1) {\n \t\tin_encoding = fallback_encoding(in_encoding);\n@@ -553,9 +579,10 @@ char *reencode_string_len(const char *in, size_t insz,\n \t\tif (conv == (iconv_t) -1)\n \t\t\treturn NULL;\n \t}\n-\n-\tout = reencode_string_iconv(in, insz, conv, outsz);\n+\tout = reencode_string_iconv(in, insz, conv, bom_len, outsz);\n \ticonv_close(conv);\n+\tif (out && bom_str && bom_len)\n+\t\tmemcpy(out, bom_str, bom_len);\n \treturn out;\n }\n #endif\n@@ -566,11 +593,6 @@ static int has_bom_prefix(const char *data, size_t len,\n \treturn data && bom && (len >= bom_len) && !memcmp(data, bom, bom_len);\n }\n\n-static const char utf16_be_bom[] = {'\\xFE', '\\xFF'};\n-static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n-static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n-static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n-\n int has_prohibited_utf_bom(const char *enc, const char *data, size_t len)\n {\n \treturn (\ndiff --git a/utf8.h b/utf8.h\nindex edea55e093..84efbfcb1f 100644\n--- a/utf8.h\n+++ b/utf8.h\n@@ -27,7 +27,7 @@ void strbuf_utf8_replace(struct strbuf *sb, int pos, int width,\n\n #ifndef NO_ICONV\n char *reencode_string_iconv(const char *in, size_t insz,\n-\t\t\t    iconv_t conv, size_t *outsz);\n+\t\t\t    iconv_t conv, size_t bom_len, size_t *outsz);\n char *reencode_string_len(const char *in, size_t insz,\n \t\t\t  const char *out_encoding,\n \t\t\t  const char *in_encoding,\n--\n2.20.1.2.gb21ebb671\n\n"},{"id":"367349","messageId":"xmqqk1iwfo1v.fsf@gitster-ct.c.googlers.com","threadId":"49734","inReplyTo":"20190120164327.3234-1-tboegi@web.de","subject":"Re: [PATCH v2 1/1] Support working-tree-encoding \"UTF-16LE-BOM\"","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-01-22T20:13:48Z","receivedAt":"2019-01-22T20:13:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"tboegi@web.de writes:\n\n> The unicode standard itself defines 3 possible ways how to encode UTF-16.\n> a) UTF-16, without BOM, big endian:\n> b) UTF-16, with BOM, little endian:\n> c) UTF-16, with BOM, big endian:\n\nIs it OK to interpret \"possible\" as \"allowed\" above?\n\n> iconv (and libiconv) can generate UTF-16, UTF-16LE or UTF-16BE:\n>\n> d) UTF-16\n> $ printf 'git' | iconv -f UTF-8 -t UTF-16 | od -c\n> 0000000  376 377  \\0   g  \\0   i  \\0   t\n\nSo among three, encoder can only do \"big endian with BOM\" (c).\n\nLack of (a) \"big endian without BOM\" in the encoder is not a problem\nin practice, as you can ask UTF-16BE to produce the stream, tell the\ndecoder that you have UTF-16 and the lack of the BOM would make the\ndecoder take it as (a).\n\nBut lack of (b) \"little endian with BOM\" is a problem.\n\nSo the proposal is to invent UTF-16-[BL]E-BOM that prepends BOM in\nfront of UTF-16-[BL]E output to allow those who want (b).\n\nWhich makes sense, I guess.  I do find it a bit ugly in the sense\nthat it is something iconv should learn to do, as the issue is\nshared with all applications that want to use libiconv and convert\ninto UTF-16.\n\nDo you add UTF-16-BE-BOM for consistency?  It would be identical to\ntelling iconv to encode to UTF-16, if I understood your problem\ndescription correctly.\n\n> diff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\n> index b8392fc330..4a88ab8be7 100644\n> --- a/Documentation/gitattributes.txt\n> +++ b/Documentation/gitattributes.txt\n> @@ -343,13 +343,13 @@ automatic line ending conversion based on your platform.\n>  ------------------------\n>\n>  Use the following attributes if your '*.ps1' files are UTF-16 little\n> -endian encoded without BOM and you want Git to use Windows line endings\n> +endian encoded with BOM and you want Git to use Windows line endings\n>  in the working directory. Please note, it is highly recommended to\n>  explicitly define the line endings with `eol` if the `working-tree-encoding`\n>  attribute is used to avoid ambiguity.\n>\n>  ------------------------\n> -*.ps1\t\ttext working-tree-encoding=UTF-16LE eol=CRLF\n> +*.ps1\t\ttext working-tree-encoding=UTF-16LE-BOM eol=CRLF\n>  ------------------------\n\nThis change is robbing from those who do want a file without BOM to\ngive to those who do want a file with BOM.  Are the latter class of\npeople the majority of the intended readers (read: Windows folks)?\n\nI wonder if the following, instead of the above hunk, would work better:\n\n endian encoded without BOM and you want Git to use Windows line endings\n-in the working directory. Please note, it is highly recommended to\n+in the working directory (use `UTF-16-LE-BOM` instead of `UTF-16LE` if\n+you want UTF-16 little endian with BOM).\n+Please note, it is highly recommended to\n explicitly define the line endings with `eol` if the `working-tree-encoding`\n\n> @@ -540,10 +546,30 @@ char *reencode_string_len(const char *in, size_t insz,\n>  {\n>  \ticonv_t conv;\n>  \tchar *out;\n> +\tconst char *bom_str = NULL;\n> +\tsize_t bom_len = 0;\n>\n>  \tif (!in_encoding)\n>  \t\treturn NULL;\n>\n> +\t/* UTF-16LE-BOM is the same as UTF-16 for reading */\n> +\tif (same_utf_encoding(\"UTF-16LE-BOM\", in_encoding))\n> +\t\tin_encoding = \"UTF-16\";\n> +\n> +\t/*\n> +\t * For writing, UTF-16 iconv typically creates \"UTF-16BE-BOM\"\n> +\t * Some users under Windows want the little endian version\n> +\t */\n> +\tif (same_utf_encoding(\"UTF-16LE-BOM\", out_encoding)) {\n> +\t\tbom_str = utf16_le_bom;\n> +\t\tbom_len = sizeof(utf16_le_bom);\n> +\t\tout_encoding = \"UTF-16LE\";\n> +\t} else if (same_utf_encoding(\"UTF-16BE-BOM\", out_encoding)) {\n> +\t\tbom_str = utf16_be_bom;\n> +\t\tbom_len = sizeof(utf16_be_bom);\n> +\t\tout_encoding = \"UTF-16BE\";\n\nOK, you do allow BE-BOM and the code does not rely on the fact that\niconv happens to produce it with \"UTF-16\", because the library is\nfree to switch between the three possible output (a)-(c) and we do\nnot want to get affected by such a switch.  Makes sense.\n\n"},{"id":"368137","messageId":"20190130150152.23040-1-tboegi@web.de","threadId":"49734","inReplyTo":"CADN+U_PUfnYWb-wW6drRANv-ZaYBEk3gWHc7oJtxohA5Vc3NEg@mail.gmail.com","subject":"[PATCH v3 1/1] Support working-tree-encoding \"UTF-16LE-BOM\"","fromName":"","fromEmail":"tboegi@web.de","sentAt":"2019-01-30T15:01:52Z","receivedAt":"2019-01-30T15:01:59Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"From: Torsten Bögershausen <tboegi@web.de>\n\nUsers who want UTF-16 files in the working tree set the .gitattributes\nlike this:\ntest.txt working-tree-encoding=UTF-16\n\nThe unicode standard itself defines 3 allowed ways how to encode UTF-16.\nThe following 3 versions convert all back to 'g' 'i' 't' in UTF-8:\n\na) UTF-16, without BOM, big endian:\n$ printf \"\\000g\\000i\\000t\" | iconv -f UTF-16 -t UTF-8 | od -c\n0000000    g   i   t\n\nb) UTF-16, with BOM, little endian:\n$ printf \"\\377\\376g\\000i\\000t\\000\" | iconv -f UTF-16 -t UTF-8 | od -c\n0000000    g   i   t\n\nc) UTF-16, with BOM, big endian:\n$ printf \"\\376\\377\\000g\\000i\\000t\" | iconv -f UTF-16 -t UTF-8 | od -c\n0000000    g   i   t\n\nGit uses libiconv to convert from UTF-8 in the index into ITF-16 in the\nworking tree.\nAfter a checkout, the resulting file has a BOM and is encoded in \"UTF-16\",\nin the version (c) above.\nThis is what iconv generates, more details follow below.\n\niconv (and libiconv) can generate UTF-16, UTF-16LE or UTF-16BE:\n\nd) UTF-16\n$ printf 'git' | iconv -f UTF-8 -t UTF-16 | od -c\n0000000  376 377  \\0   g  \\0   i  \\0   t\n\ne) UTF-16LE\n$ printf 'git' | iconv -f UTF-8 -t UTF-16LE | od -c\n0000000    g  \\0   i  \\0   t  \\0\n\nf)  UTF-16BE\n$ printf 'git' | iconv -f UTF-8 -t UTF-16BE | od -c\n0000000   \\0   g  \\0   i  \\0   t\n\nThere is no way to generate version (b) from above in a Git working tree,\nbut that is what some applications need.\n(All fully unicode aware applications should be able to read all 3 variants,\nbut in practise we are not there yet).\n\nWhen producing UTF-16 as an output, iconv generates the big endian version\nwith a BOM. (big endian is probably chosen for historical reasons).\n\niconv can produce UTF-16 files with little endianess by using \"UTF-16LE\"\nas encoding, and that file does not have a BOM.\n\nNot all users (especially under Windows) are happy with this.\nSome tools are not fully unicode aware and can only handle version (b).\n\nToday there is no way to produce version (b) with iconv (or libiconv).\nLooking into the history of iconv, it seems as if version (c) will\nbe used in all future iconv versions (for compatibility reasons).\n\nSolve this dilemma and introduce a Git-specific \"UTF-16LE-BOM\".\nlibiconv can not handle the encoding, so Git pick it up, handles the BOM\nand uses libiconv to convert the rest of the stream.\n(UTF-16BE-BOM is added for consistency)\n\nRported-by: Adrián Gimeno Balaguer <adrigibal@gmail.com>\nSigned-off-by: Torsten Bögershausen <tboegi@web.de>\n---\n\nChanges since v2:\n  Update the commit message (s/possible/allowed/)\n  Update the documentation, as suggested by Junio:\n  ...wonder if the following,\n     instead of the above hunk, would work better..\n  Yes, it does.\n\nDocumentation/gitattributes.txt  |  4 ++-\n compat/precompose_utf8.c         |  2 +-\n t/t0028-working-tree-encoding.sh | 12 ++++++++-\n utf8.c                           | 42 ++++++++++++++++++++++++--------\n utf8.h                           |  2 +-\n 5 files changed, 48 insertions(+), 14 deletions(-)\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex b8392fc330..a2310fb920 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -344,7 +344,9 @@ automatic line ending conversion based on your platform.\n\n Use the following attributes if your '*.ps1' files are UTF-16 little\n endian encoded without BOM and you want Git to use Windows line endings\n-in the working directory. Please note, it is highly recommended to\n+in the working directory (use `UTF-16-LE-BOM` instead of `UTF-16LE` if\n+you want UTF-16 little endian with BOM).\n+Please note, it is highly recommended to\n explicitly define the line endings with `eol` if the `working-tree-encoding`\n attribute is used to avoid ambiguity.\n\ndiff --git a/compat/precompose_utf8.c b/compat/precompose_utf8.c\nindex de61c15d34..136250fbf6 100644\n--- a/compat/precompose_utf8.c\n+++ b/compat/precompose_utf8.c\n@@ -79,7 +79,7 @@ void precompose_argv(int argc, const char **argv)\n \t\tsize_t namelen;\n \t\toldarg = argv[i];\n \t\tif (has_non_ascii(oldarg, (size_t)-1, &namelen)) {\n-\t\t\tnewarg = reencode_string_iconv(oldarg, namelen, ic_precompose, NULL);\n+\t\t\tnewarg = reencode_string_iconv(oldarg, namelen, ic_precompose, 0, NULL);\n \t\t\tif (newarg)\n \t\t\t\targv[i] = newarg;\n \t\t}\ndiff --git a/t/t0028-working-tree-encoding.sh b/t/t0028-working-tree-encoding.sh\nindex 7e87b5a200..e58ecbfc44 100755\n--- a/t/t0028-working-tree-encoding.sh\n+++ b/t/t0028-working-tree-encoding.sh\n@@ -11,9 +11,12 @@ test_expect_success 'setup test files' '\n\n \ttext=\"hallo there!\\ncan you read me?\" &&\n \techo \"*.utf16 text working-tree-encoding=utf-16\" >.gitattributes &&\n+\techo \"*.utf16lebom text working-tree-encoding=UTF-16LE-BOM\" >>.gitattributes &&\n \tprintf \"$text\" >test.utf8.raw &&\n \tprintf \"$text\" | iconv -f UTF-8 -t UTF-16 >test.utf16.raw &&\n \tprintf \"$text\" | iconv -f UTF-8 -t UTF-32 >test.utf32.raw &&\n+\tprintf \"\\377\\376\"                         >test.utf16lebom.raw &&\n+\tprintf \"$text\" | iconv -f UTF-8 -t UTF-32LE >>test.utf16lebom.raw &&\n\n \t# Line ending tests\n \tprintf \"one\\ntwo\\nthree\\n\" >lf.utf8.raw &&\n@@ -32,7 +35,8 @@ test_expect_success 'setup test files' '\n \t# Add only UTF-16 file, we will add the UTF-32 file later\n \tcp test.utf16.raw test.utf16 &&\n \tcp test.utf32.raw test.utf32 &&\n-\tgit add .gitattributes test.utf16 &&\n+\tcp test.utf16lebom.raw test.utf16lebom &&\n+\tgit add .gitattributes test.utf16 test.utf16lebom &&\n \tgit commit -m initial\n '\n\n@@ -51,6 +55,12 @@ test_expect_success 're-encode to UTF-16 on checkout' '\n \ttest_cmp_bin test.utf16.raw test.utf16\n '\n\n+test_expect_success 're-encode to UTF-16-LE-BOM on checkout' '\n+\trm test.utf16lebom &&\n+\tgit checkout test.utf16lebom &&\n+\ttest_cmp_bin test.utf16lebom.raw test.utf16lebom\n+'\n+\n test_expect_success 'check $GIT_DIR/info/attributes support' '\n \ttest_when_finished \"rm -f test.utf32.git\" &&\n \ttest_when_finished \"git reset --hard HEAD\" &&\ndiff --git a/utf8.c b/utf8.c\nindex eb78587504..83824dc2f4 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -4,6 +4,11 @@\n\n /* This code is originally from http://www.cl.cam.ac.uk/~mgk25/ucs/ */\n\n+static const char utf16_be_bom[] = {'\\xFE', '\\xFF'};\n+static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n+static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n+static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n+\n struct interval {\n \tucs_char_t first;\n \tucs_char_t last;\n@@ -470,16 +475,17 @@ int utf8_fprintf(FILE *stream, const char *format, ...)\n #else\n \ttypedef char * iconv_ibp;\n #endif\n-char *reencode_string_iconv(const char *in, size_t insz, iconv_t conv, size_t *outsz_p)\n+char *reencode_string_iconv(const char *in, size_t insz, iconv_t conv,\n+\t\t\t    size_t bom_len, size_t *outsz_p)\n {\n \tsize_t outsz, outalloc;\n \tchar *out, *outpos;\n \ticonv_ibp cp;\n\n \toutsz = insz;\n-\toutalloc = st_add(outsz, 1); /* for terminating NUL */\n+\toutalloc = st_add(outsz, 1 + bom_len); /* for terminating NUL */\n \tout = xmalloc(outalloc);\n-\toutpos = out;\n+\toutpos = out + bom_len;\n \tcp = (iconv_ibp)in;\n\n \twhile (1) {\n@@ -540,10 +546,30 @@ char *reencode_string_len(const char *in, size_t insz,\n {\n \ticonv_t conv;\n \tchar *out;\n+\tconst char *bom_str = NULL;\n+\tsize_t bom_len = 0;\n\n \tif (!in_encoding)\n \t\treturn NULL;\n\n+\t/* UTF-16LE-BOM is the same as UTF-16 for reading */\n+\tif (same_utf_encoding(\"UTF-16LE-BOM\", in_encoding))\n+\t\tin_encoding = \"UTF-16\";\n+\n+\t/*\n+\t * For writing, UTF-16 iconv typically creates \"UTF-16BE-BOM\"\n+\t * Some users under Windows want the little endian version\n+\t */\n+\tif (same_utf_encoding(\"UTF-16LE-BOM\", out_encoding)) {\n+\t\tbom_str = utf16_le_bom;\n+\t\tbom_len = sizeof(utf16_le_bom);\n+\t\tout_encoding = \"UTF-16LE\";\n+\t} else if (same_utf_encoding(\"UTF-16BE-BOM\", out_encoding)) {\n+\t\tbom_str = utf16_be_bom;\n+\t\tbom_len = sizeof(utf16_be_bom);\n+\t\tout_encoding = \"UTF-16BE\";\n+\t}\n+\n \tconv = iconv_open(out_encoding, in_encoding);\n \tif (conv == (iconv_t) -1) {\n \t\tin_encoding = fallback_encoding(in_encoding);\n@@ -553,9 +579,10 @@ char *reencode_string_len(const char *in, size_t insz,\n \t\tif (conv == (iconv_t) -1)\n \t\t\treturn NULL;\n \t}\n-\n-\tout = reencode_string_iconv(in, insz, conv, outsz);\n+\tout = reencode_string_iconv(in, insz, conv, bom_len, outsz);\n \ticonv_close(conv);\n+\tif (out && bom_str && bom_len)\n+\t\tmemcpy(out, bom_str, bom_len);\n \treturn out;\n }\n #endif\n@@ -566,11 +593,6 @@ static int has_bom_prefix(const char *data, size_t len,\n \treturn data && bom && (len >= bom_len) && !memcmp(data, bom, bom_len);\n }\n\n-static const char utf16_be_bom[] = {'\\xFE', '\\xFF'};\n-static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n-static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n-static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n-\n int has_prohibited_utf_bom(const char *enc, const char *data, size_t len)\n {\n \treturn (\ndiff --git a/utf8.h b/utf8.h\nindex edea55e093..84efbfcb1f 100644\n--- a/utf8.h\n+++ b/utf8.h\n@@ -27,7 +27,7 @@ void strbuf_utf8_replace(struct strbuf *sb, int pos, int width,\n\n #ifndef NO_ICONV\n char *reencode_string_iconv(const char *in, size_t insz,\n-\t\t\t    iconv_t conv, size_t *outsz);\n+\t\t\t    iconv_t conv, size_t bom_len, size_t *outsz);\n char *reencode_string_len(const char *in, size_t insz,\n \t\t\t  const char *out_encoding,\n \t\t\t  const char *in_encoding,\n--\n2.20.1.2.gb21ebb671\n\n"},{"id":"368141","messageId":"000901d4b8af$edaccf20$c9066d60$@pdinc.us","threadId":"49734","inReplyTo":"20190130150152.23040-1-tboegi@web.de","subject":"RE: [PATCH v3 1/1] Support working-tree-encoding \"UTF-16LE-BOM\"","fromName":"Jason Pyeron","fromEmail":"jpyeron@pdinc.us","sentAt":"2019-01-30T15:24:44Z","receivedAt":"2019-01-30T15:41:43Z","isPatch":true,"sender":{"key":"jpyeron@pdinc.us","avatar":"https://gravatar.com/avatar/c2e53452caa53d940768a1ffc9cf76196d851b9b534b7a39cd39852a70a0508f?d=mp&s=160"},"body":"> -----Original Message-----\n> From: git-owner@vger.kernel.org <git-owner@vger.kernel.org> On Behalf Of\n> tboegi@web.de\n> Sent: Wednesday, January 30, 2019 10:02 AM\n> To: git@vger.kernel.org; adrigibal@gmail.com\n> Cc: Torsten Bögershausen <tboegi@web.de>\n> Subject: [PATCH v3 1/1] Support working-tree-encoding \"UTF-16LE-BOM\"\n> \n> From: Torsten Bögershausen <tboegi@web.de>\n> \n> Users who want UTF-16 files in the working tree set the .gitattributes\n> like this:\n> test.txt working-tree-encoding=UTF-16\n> \n> The unicode standard itself defines 3 allowed ways how to encode UTF-16.\n> The following 3 versions convert all back to 'g' 'i' 't' in UTF-8:\n> \n> a) UTF-16, without BOM, big endian:\n> $ printf \"\\000g\\000i\\000t\" | iconv -f UTF-16 -t UTF-8 | od -c\n> 0000000    g   i   t\n> \n> b) UTF-16, with BOM, little endian:\n> $ printf \"\\377\\376g\\000i\\000t\\000\" | iconv -f UTF-16 -t UTF-8 | od -c\n> 0000000    g   i   t\n> \n> c) UTF-16, with BOM, big endian:\n> $ printf \"\\376\\377\\000g\\000i\\000t\" | iconv -f UTF-16 -t UTF-8 | od -c\n> 0000000    g   i   t\n> \n> Git uses libiconv to convert from UTF-8 in the index into ITF-16 in the\n> working tree.\n> After a checkout, the resulting file has a BOM and is encoded in \"UTF-16\",\n> in the version (c) above.\n> This is what iconv generates, more details follow below.\n> \n> iconv (and libiconv) can generate UTF-16, UTF-16LE or UTF-16BE:\n> \n> d) UTF-16\n> $ printf 'git' | iconv -f UTF-8 -t UTF-16 | od -c\n> 0000000  376 377  \\0   g  \\0   i  \\0   t\n> \n> e) UTF-16LE\n> $ printf 'git' | iconv -f UTF-8 -t UTF-16LE | od -c\n> 0000000    g  \\0   i  \\0   t  \\0\n> \n> f)  UTF-16BE\n> $ printf 'git' | iconv -f UTF-8 -t UTF-16BE | od -c\n> 0000000   \\0   g  \\0   i  \\0   t\n> \n> There is no way to generate version (b) from above in a Git working tree,\n> but that is what some applications need.\n> (All fully unicode aware applications should be able to read all 3\n> variants,\n> but in practise we are not there yet).\n> \n> When producing UTF-16 as an output, iconv generates the big endian version\n> with a BOM. (big endian is probably chosen for historical reasons).\n> \n> iconv can produce UTF-16 files with little endianess by using \"UTF-16LE\"\n> as encoding, and that file does not have a BOM.\n> \n> Not all users (especially under Windows) are happy with this.\n> Some tools are not fully unicode aware and can only handle version (b).\n> \n> Today there is no way to produce version (b) with iconv (or libiconv).\n> Looking into the history of iconv, it seems as if version (c) will\n> be used in all future iconv versions (for compatibility reasons).\n\n\nReading the RFC 2781 section 3.3:\n \n   Text in the \"UTF-16BE\" charset MUST be serialized with the octets\n   which make up a single 16-bit UTF-16 value in big-endian order.\n   Systems labelling UTF-16BE text MUST NOT prepend a BOM to the text.\n\n   Text in the \"UTF-16LE\" charset MUST be serialized with the octets\n   which make up a single 16-bit UTF-16 value in little-endian order.\n   Systems labelling UTF-16LE text MUST NOT prepend a BOM to the text.\n\nI opened a bug with libiconv... https://savannah.gnu.org/bugs/index.php?55609\n\n> \n> Solve this dilemma and introduce a Git-specific \"UTF-16LE-BOM\".\n> libiconv can not handle the encoding, so Git pick it up, handles the BOM\n> and uses libiconv to convert the rest of the stream.\n> (UTF-16BE-BOM is added for consistency)\n> \n> Rported-by: Adrián Gimeno Balaguer <adrigibal@gmail.com>\n> Signed-off-by: Torsten Bögershausen <tboegi@web.de>\n> ---\n> \n> Changes since v2:\n>   Update the commit message (s/possible/allowed/)\n>   Update the documentation, as suggested by Junio:\n>   ...wonder if the following,\n>      instead of the above hunk, would work better..\n>   Yes, it does.\n> \n> Documentation/gitattributes.txt  |  4 ++-\n>  compat/precompose_utf8.c         |  2 +-\n>  t/t0028-working-tree-encoding.sh | 12 ++++++++-\n>  utf8.c                           | 42 ++++++++++++++++++++++++--------\n>  utf8.h                           |  2 +-\n>  5 files changed, 48 insertions(+), 14 deletions(-)\n> \n> diff --git a/Documentation/gitattributes.txt\n> b/Documentation/gitattributes.txt\n> index b8392fc330..a2310fb920 100644\n> --- a/Documentation/gitattributes.txt\n> +++ b/Documentation/gitattributes.txt\n> @@ -344,7 +344,9 @@ automatic line ending conversion based on your\n> platform.\n> \n>  Use the following attributes if your '*.ps1' files are UTF-16 little\n>  endian encoded without BOM and you want Git to use Windows line endings\n> -in the working directory. Please note, it is highly recommended to\n> +in the working directory (use `UTF-16-LE-BOM` instead of `UTF-16LE` if\n> +you want UTF-16 little endian with BOM).\n> +Please note, it is highly recommended to\n>  explicitly define the line endings with `eol` if the `working-tree-\n> encoding`\n>  attribute is used to avoid ambiguity.\n> \n> diff --git a/compat/precompose_utf8.c b/compat/precompose_utf8.c\n> index de61c15d34..136250fbf6 100644\n> --- a/compat/precompose_utf8.c\n> +++ b/compat/precompose_utf8.c\n> @@ -79,7 +79,7 @@ void precompose_argv(int argc, const char **argv)\n>  \t\tsize_t namelen;\n>  \t\toldarg = argv[i];\n>  \t\tif (has_non_ascii(oldarg, (size_t)-1, &namelen)) {\n> -\t\t\tnewarg = reencode_string_iconv(oldarg, namelen,\n> ic_precompose, NULL);\n> +\t\t\tnewarg = reencode_string_iconv(oldarg, namelen,\n> ic_precompose, 0, NULL);\n>  \t\t\tif (newarg)\n>  \t\t\t\targv[i] = newarg;\n>  \t\t}\n> diff --git a/t/t0028-working-tree-encoding.sh b/t/t0028-working-tree-\n> encoding.sh\n> index 7e87b5a200..e58ecbfc44 100755\n> --- a/t/t0028-working-tree-encoding.sh\n> +++ b/t/t0028-working-tree-encoding.sh\n> @@ -11,9 +11,12 @@ test_expect_success 'setup test files' '\n> \n>  \ttext=\"hallo there!\\ncan you read me?\" &&\n>  \techo \"*.utf16 text working-tree-encoding=utf-16\" >.gitattributes &&\n> +\techo \"*.utf16lebom text working-tree-encoding=UTF-16LE-BOM\"\n> >>.gitattributes &&\n>  \tprintf \"$text\" >test.utf8.raw &&\n>  \tprintf \"$text\" | iconv -f UTF-8 -t UTF-16 >test.utf16.raw &&\n>  \tprintf \"$text\" | iconv -f UTF-8 -t UTF-32 >test.utf32.raw &&\n> +\tprintf \"\\377\\376\"                         >test.utf16lebom.raw &&\n> +\tprintf \"$text\" | iconv -f UTF-8 -t UTF-32LE >>test.utf16lebom.raw &&\n> \n>  \t# Line ending tests\n>  \tprintf \"one\\ntwo\\nthree\\n\" >lf.utf8.raw &&\n> @@ -32,7 +35,8 @@ test_expect_success 'setup test files' '\n>  \t# Add only UTF-16 file, we will add the UTF-32 file later\n>  \tcp test.utf16.raw test.utf16 &&\n>  \tcp test.utf32.raw test.utf32 &&\n> -\tgit add .gitattributes test.utf16 &&\n> +\tcp test.utf16lebom.raw test.utf16lebom &&\n> +\tgit add .gitattributes test.utf16 test.utf16lebom &&\n>  \tgit commit -m initial\n>  '\n> \n> @@ -51,6 +55,12 @@ test_expect_success 're-encode to UTF-16 on checkout' '\n>  \ttest_cmp_bin test.utf16.raw test.utf16\n>  '\n> \n> +test_expect_success 're-encode to UTF-16-LE-BOM on checkout' '\n> +\trm test.utf16lebom &&\n> +\tgit checkout test.utf16lebom &&\n> +\ttest_cmp_bin test.utf16lebom.raw test.utf16lebom\n> +'\n> +\n>  test_expect_success 'check $GIT_DIR/info/attributes support' '\n>  \ttest_when_finished \"rm -f test.utf32.git\" &&\n>  \ttest_when_finished \"git reset --hard HEAD\" &&\n> diff --git a/utf8.c b/utf8.c\n> index eb78587504..83824dc2f4 100644\n> --- a/utf8.c\n> +++ b/utf8.c\n> @@ -4,6 +4,11 @@\n> \n>  /* This code is originally from http://www.cl.cam.ac.uk/~mgk25/ucs/ */\n> \n> +static const char utf16_be_bom[] = {'\\xFE', '\\xFF'};\n> +static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n> +static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n> +static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n> +\n>  struct interval {\n>  \tucs_char_t first;\n>  \tucs_char_t last;\n> @@ -470,16 +475,17 @@ int utf8_fprintf(FILE *stream, const char *format,\n> ...)\n>  #else\n>  \ttypedef char * iconv_ibp;\n>  #endif\n> -char *reencode_string_iconv(const char *in, size_t insz, iconv_t conv,\n> size_t *outsz_p)\n> +char *reencode_string_iconv(const char *in, size_t insz, iconv_t conv,\n> +\t\t\t    size_t bom_len, size_t *outsz_p)\n>  {\n>  \tsize_t outsz, outalloc;\n>  \tchar *out, *outpos;\n>  \ticonv_ibp cp;\n> \n>  \toutsz = insz;\n> -\toutalloc = st_add(outsz, 1); /* for terminating NUL */\n> +\toutalloc = st_add(outsz, 1 + bom_len); /* for terminating NUL */\n>  \tout = xmalloc(outalloc);\n> -\toutpos = out;\n> +\toutpos = out + bom_len;\n>  \tcp = (iconv_ibp)in;\n> \n>  \twhile (1) {\n> @@ -540,10 +546,30 @@ char *reencode_string_len(const char *in, size_t\n> insz,\n>  {\n>  \ticonv_t conv;\n>  \tchar *out;\n> +\tconst char *bom_str = NULL;\n> +\tsize_t bom_len = 0;\n> \n>  \tif (!in_encoding)\n>  \t\treturn NULL;\n> \n> +\t/* UTF-16LE-BOM is the same as UTF-16 for reading */\n> +\tif (same_utf_encoding(\"UTF-16LE-BOM\", in_encoding))\n> +\t\tin_encoding = \"UTF-16\";\n> +\n> +\t/*\n> +\t * For writing, UTF-16 iconv typically creates \"UTF-16BE-BOM\"\n> +\t * Some users under Windows want the little endian version\n> +\t */\n> +\tif (same_utf_encoding(\"UTF-16LE-BOM\", out_encoding)) {\n> +\t\tbom_str = utf16_le_bom;\n> +\t\tbom_len = sizeof(utf16_le_bom);\n> +\t\tout_encoding = \"UTF-16LE\";\n> +\t} else if (same_utf_encoding(\"UTF-16BE-BOM\", out_encoding)) {\n> +\t\tbom_str = utf16_be_bom;\n> +\t\tbom_len = sizeof(utf16_be_bom);\n> +\t\tout_encoding = \"UTF-16BE\";\n> +\t}\n> +\n>  \tconv = iconv_open(out_encoding, in_encoding);\n>  \tif (conv == (iconv_t) -1) {\n>  \t\tin_encoding = fallback_encoding(in_encoding);\n> @@ -553,9 +579,10 @@ char *reencode_string_len(const char *in, size_t\n> insz,\n>  \t\tif (conv == (iconv_t) -1)\n>  \t\t\treturn NULL;\n>  \t}\n> -\n> -\tout = reencode_string_iconv(in, insz, conv, outsz);\n> +\tout = reencode_string_iconv(in, insz, conv, bom_len, outsz);\n>  \ticonv_close(conv);\n> +\tif (out && bom_str && bom_len)\n> +\t\tmemcpy(out, bom_str, bom_len);\n>  \treturn out;\n>  }\n>  #endif\n> @@ -566,11 +593,6 @@ static int has_bom_prefix(const char *data, size_t\n> len,\n>  \treturn data && bom && (len >= bom_len) && !memcmp(data, bom,\n> bom_len);\n>  }\n> \n> -static const char utf16_be_bom[] = {'\\xFE', '\\xFF'};\n> -static const char utf16_le_bom[] = {'\\xFF', '\\xFE'};\n> -static const char utf32_be_bom[] = {'\\0', '\\0', '\\xFE', '\\xFF'};\n> -static const char utf32_le_bom[] = {'\\xFF', '\\xFE', '\\0', '\\0'};\n> -\n>  int has_prohibited_utf_bom(const char *enc, const char *data, size_t len)\n>  {\n>  \treturn (\n> diff --git a/utf8.h b/utf8.h\n> index edea55e093..84efbfcb1f 100644\n> --- a/utf8.h\n> +++ b/utf8.h\n> @@ -27,7 +27,7 @@ void strbuf_utf8_replace(struct strbuf *sb, int pos, int\n> width,\n> \n>  #ifndef NO_ICONV\n>  char *reencode_string_iconv(const char *in, size_t insz,\n> -\t\t\t    iconv_t conv, size_t *outsz);\n> +\t\t\t    iconv_t conv, size_t bom_len, size_t *outsz);\n>  char *reencode_string_len(const char *in, size_t insz,\n>  \t\t\t  const char *out_encoding,\n>  \t\t\t  const char *in_encoding,\n> --\n> 2.20.1.2.gb21ebb671\n> \n\n\n"},{"id":"368145","messageId":"20190130174932.lhs3npztu5tusy3e@tb-raspi4","threadId":"49734","inReplyTo":"000901d4b8af$edaccf20$c9066d60$@pdinc.us","subject":"Re: [PATCH v3 1/1] Support working-tree-encoding \"UTF-16LE-BOM\"","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2019-01-30T17:49:32Z","receivedAt":"2019-01-30T17:49:39Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On Wed, Jan 30, 2019 at 10:24:44AM -0500, Jason Pyeron wrote:\n> > -----Original Message-----\n> > From: git-owner@vger.kernel.org <git-owner@vger.kernel.org> On Behalf Of\n> > tboegi@web.de\n> > Sent: Wednesday, January 30, 2019 10:02 AM\n> > To: git@vger.kernel.org; adrigibal@gmail.com\n> > Cc: Torsten Bögershausen <tboegi@web.de>\n> > Subject: [PATCH v3 1/1] Support working-tree-encoding \"UTF-16LE-BOM\"\n> >\n> > From: Torsten Bögershausen <tboegi@web.de>\n> >\n> > Users who want UTF-16 files in the working tree set the .gitattributes\n> > like this:\n> > test.txt working-tree-encoding=UTF-16\n> >\n> > The unicode standard itself defines 3 allowed ways how to encode UTF-16.\n> > The following 3 versions convert all back to 'g' 'i' 't' in UTF-8:\n> >\n> > a) UTF-16, without BOM, big endian:\n> > $ printf \"\\000g\\000i\\000t\" | iconv -f UTF-16 -t UTF-8 | od -c\n> > 0000000    g   i   t\n> >\n> > b) UTF-16, with BOM, little endian:\n> > $ printf \"\\377\\376g\\000i\\000t\\000\" | iconv -f UTF-16 -t UTF-8 | od -c\n> > 0000000    g   i   t\n> >\n> > c) UTF-16, with BOM, big endian:\n> > $ printf \"\\376\\377\\000g\\000i\\000t\" | iconv -f UTF-16 -t UTF-8 | od -c\n> > 0000000    g   i   t\n> >\n> > Git uses libiconv to convert from UTF-8 in the index into ITF-16 in the\n> > working tree.\n> > After a checkout, the resulting file has a BOM and is encoded in \"UTF-16\",\n> > in the version (c) above.\n> > This is what iconv generates, more details follow below.\n> >\n> > iconv (and libiconv) can generate UTF-16, UTF-16LE or UTF-16BE:\n> >\n> > d) UTF-16\n> > $ printf 'git' | iconv -f UTF-8 -t UTF-16 | od -c\n> > 0000000  376 377  \\0   g  \\0   i  \\0   t\n> >\n> > e) UTF-16LE\n> > $ printf 'git' | iconv -f UTF-8 -t UTF-16LE | od -c\n> > 0000000    g  \\0   i  \\0   t  \\0\n> >\n> > f)  UTF-16BE\n> > $ printf 'git' | iconv -f UTF-8 -t UTF-16BE | od -c\n> > 0000000   \\0   g  \\0   i  \\0   t\n> >\n> > There is no way to generate version (b) from above in a Git working tree,\n> > but that is what some applications need.\n> > (All fully unicode aware applications should be able to read all 3\n> > variants,\n> > but in practise we are not there yet).\n> >\n> > When producing UTF-16 as an output, iconv generates the big endian version\n> > with a BOM. (big endian is probably chosen for historical reasons).\n> >\n> > iconv can produce UTF-16 files with little endianess by using \"UTF-16LE\"\n> > as encoding, and that file does not have a BOM.\n> >\n> > Not all users (especially under Windows) are happy with this.\n> > Some tools are not fully unicode aware and can only handle version (b).\n> >\n> > Today there is no way to produce version (b) with iconv (or libiconv).\n> > Looking into the history of iconv, it seems as if version (c) will\n> > be used in all future iconv versions (for compatibility reasons).\n>\n>\n> Reading the RFC 2781 section 3.3:\n>\n>    Text in the \"UTF-16BE\" charset MUST be serialized with the octets\n>    which make up a single 16-bit UTF-16 value in big-endian order.\n>    Systems labelling UTF-16BE text MUST NOT prepend a BOM to the text.\n>\n>    Text in the \"UTF-16LE\" charset MUST be serialized with the octets\n>    which make up a single 16-bit UTF-16 value in little-endian order.\n>    Systems labelling UTF-16LE text MUST NOT prepend a BOM to the text.\n>\n> I opened a bug with libiconv... https://savannah.gnu.org/bugs/index.php?55609\n>\n\nUTF-16 may be a), b) or c) from above.\nEvery unicode compliant system should be able to read all 3 of them.\n\nWhen writing, the system/application/converter is free to choose one of those.\nProbably out of historical reason, big endian is preferred (in iconv),\nand to be helpful to systems/applications a BOM is written in the beginning.\nThis is according to the RFC, why do you think that this is a bug ?\n\n\n\n"},{"id":"370784","messageId":"20190306052310.31546-1-tboegi@web.de","threadId":"49734","inReplyTo":"CADN+U_PUfnYWb-wW6drRANv-ZaYBEk3gWHc7oJtxohA5Vc3NEg@mail.gmail.com","subject":"[PATCH v1 1/1] gitattributes.txt: fix typo","fromName":"","fromEmail":"tboegi@web.de","sentAt":"2019-03-06T05:23:10Z","receivedAt":"2019-03-06T05:23:17Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"From: Yash Bhatambare <ybhatambare@gmail.com>\n\n`UTF-16-LE-BOM` to `UTF-16LE-BOM`.\n\nthis closes https://github.com/git-for-windows/git/issues/2095\n\nSigned-off-by: Yash Bhatambare <ybhatambare@gmail.com>\nSigned-off-by: Torsten Bögershausen <tboegi@web.de>\n---\n\nThis patch already made it into Git for Windows,\nso I send it upstream \"as is\".\n\nDocumentation/gitattributes.txt | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 9b41f81c06..bdd11a2ddd 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -346,7 +346,7 @@ automatic line ending conversion based on your platform.\n\n Use the following attributes if your '*.ps1' files are UTF-16 little\n endian encoded without BOM and you want Git to use Windows line endings\n-in the working directory (use `UTF-16-LE-BOM` instead of `UTF-16LE` if\n+in the working directory (use `UTF-16LE-BOM` instead of `UTF-16LE` if\n you want UTF-16 little endian with BOM).\n Please note, it is highly recommended to\n explicitly define the line endings with `eol` if the `working-tree-encoding`\n--\n2.19.1.593.gc670b1f876\n\n"},{"id":"370850","messageId":"xmqqva0vo7jz.fsf@gitster-ct.c.googlers.com","threadId":"49734","inReplyTo":"20190306052310.31546-1-tboegi@web.de","subject":"Re: [PATCH v1 1/1] gitattributes.txt: fix typo","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-03-07T00:24:32Z","receivedAt":"2019-03-07T00:24:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"tboegi@web.de writes:\n\n>  Use the following attributes if your '*.ps1' files are UTF-16 little\n>  endian encoded without BOM and you want Git to use Windows line endings\n> -in the working directory (use `UTF-16-LE-BOM` instead of `UTF-16LE` if\n> +in the working directory (use `UTF-16LE-BOM` instead of `UTF-16LE` if\n\nThanks for your attention to detail ;-)\n"}]}