{"thread":{"id":"58232","subject":"Feature request: better error messages when UTF-8 bites","startedAt":"2022-07-27T20:21:51Z","lastAt":"2022-07-28T18:01:47Z","messageCount":4,"participants":["CH","Johannes Sixt","Thomas Guyot","Torsten Bögershausen"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"460037","messageId":"f5a49da29fd0e5577083f1006d394158@ch.pkts.ca","threadId":"58232","inReplyTo":null,"subject":"Feature request: better error messages when UTF-8 bites","fromName":"CH","fromEmail":"ch-and-git.vger.kernel.org@ch.pkts.ca","sentAt":"2022-07-27T20:21:43Z","receivedAt":"2022-07-27T20:21:51Z","isPatch":false,"sender":{"key":"ch-and-git.vger.kernel.org@ch.pkts.ca","avatar":null},"body":"Hi;\n\nJust found an annoyance in `git log` (and likely elsewhere) that may \nwarrant a change:\n\nSomehow when copying and pasting a commit from a website to the command \nline, a UTF-8 Byte Order Mark (BOM) \n[https://en.wikipedia.org/wiki/Byte_order_mark] was appended to one of \nthe commit ids.  BOMs are invisible, as are many other UTF-8 code \npoints.  The upshot was that Git didn't like it, and complained \nbitterly:\n\n> $ strace -etrace=execve -s 200 git diff \n> 038179704f0066aa815d5429221cf381ff4ef289  \n> 47346a462d8ba40b9a8b073e351c362522c46aa6\n> \n> execve(\"/usr/bin/git\", [\"git\", \"diff\", \n> \"038179704f0066aa815d5429221cf381ff4ef289\\357\\273\\277\", \n> \"47346a462d8ba40b9a8b073e351c362522c46aa6\"], 0x7fffec3c4bb0 /* 80 vars \n> */) = 0\n> \n> fatal: ambiguous argument '038179704f0066aa815d5429221cf381ff4ef289': \n> unknown revision or path not in the working tree.\n> Use '--' to separate paths from revisions, like this:\n> 'git <command> [<revision>...] -- [<file>...]'\n> +++ exited with 128 +++\n\nFeature request:\n================\n\nWhen printing the \"fatal: ambiguous argument '......': ....\", perhaps \nescape (url or otherwise) the ambiguous argument when printing it in the \nerror message, or maybe add a sentence about non-ASCII characters being \nfound.\n\nThis is sort of a difficult corner-case, in that it is perfectly legal \nto have UTF-8 characters in a branch or tag name (see \ngit-check-ref-format for the allowed characters), so someone could \nindeed create a branch named \n\"038179704f0066aa815d5429221cf381ff4ef289\\357\\273\\277\" if they were a \ntortured soul bent on overthrowing polite society.  Rejecting input \nbecause it has bytes with values above \\177 is therefore not a solution.\n\nSimilarly, scanning the input for invisible UTF-8 characters (or even \ninvalid UTF-8 sequences) is leaning too far the other way: git should \nnot be validating character encodings.  It should stay encoding-neutral, \nas the alternative leads to madness, driving developers into becoming \ntortured souls bent on rigidly enforcing polite society.  We have enough \nof those already.\n\nIt's unclear as to whether violent overthrow or rigid enforcement is the \nlesser of two evils, but let's not perform the experiment to find out.  \n:-)\n\nCheers!\n\n-- \nCH (ch-and-git.vger.kernel.org@ch.pkts.ca)\n"},{"id":"460065","messageId":"4b09bf98-dae2-491e-9858-801a9bcdd2fa@kdbg.org","threadId":"58232","inReplyTo":"f5a49da29fd0e5577083f1006d394158@ch.pkts.ca","subject":"Re: Feature request: better error messages when UTF-8 bites","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2022-07-28T05:42:28Z","receivedAt":"2022-07-28T05:42:40Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 27.07.22 um 22:21 schrieb CH:\n> Somehow when copying and pasting a commit from a website to the command\n> line, a UTF-8 Byte Order Mark (BOM)\n> [https://en.wikipedia.org/wiki/Byte_order_mark] was appended to one of\n> the commit ids.  BOMs are invisible, as are many other UTF-8 code\n> points.  The upshot was that Git didn't like it, and complained bitterly:\n> \n>> $ strace -etrace=execve -s 200 git diff\n>> 038179704f0066aa815d5429221cf381ff4ef289 \n>> 47346a462d8ba40b9a8b073e351c362522c46aa6\n>>\n>> execve(\"/usr/bin/git\", [\"git\", \"diff\",\n>> \"038179704f0066aa815d5429221cf381ff4ef289\\357\\273\\277\",\n>> \"47346a462d8ba40b9a8b073e351c362522c46aa6\"], 0x7fffec3c4bb0 /* 80 vars\n>> */) = 0\n>>\n>> fatal: ambiguous argument '038179704f0066aa815d5429221cf381ff4ef289':\n>> unknown revision or path not in the working tree.\n>> Use '--' to separate paths from revisions, like this:\n>> 'git <command> [<revision>...] -- [<file>...]'\n>> +++ exited with 128 +++\n> \n> Feature request:\n> ================\n> \n> When printing the \"fatal: ambiguous argument '......': ....\", perhaps\n> escape (url or otherwise) the ambiguous argument when printing it in the\n> error message, or maybe add a sentence about non-ASCII characters being\n> found.\n\nThat's not going to fly, IMHO, because when I type\n\n   git diff todo/René\n\nI would not want to see\n\nfatal: ambiguous argument 'todo/Ren\\303\\251': unknown ...\n\nI'm convinced that there are thousands of users who use non-ASCII branch\nand file names that they also frequently mis-type. They'd all be greeted\nwith unintelligible nerdy gibberish.\n\nI may be able to change my mind if ambiguous input (in the sense of \"is\nnot what it seems to be\") leads to a security hazard that is unique to Git.\n\n-- Hannes\n"},{"id":"460067","messageId":"1e454493-ca57-ca2e-7d82-7333a769817e@gmail.com","threadId":"58232","inReplyTo":"4b09bf98-dae2-491e-9858-801a9bcdd2fa@kdbg.org","subject":"Re: Feature request: better error messages when UTF-8 bites","fromName":"Thomas Guyot","fromEmail":"tguyot@gmail.com","sentAt":"2022-07-28T09:40:42Z","receivedAt":"2022-07-28T09:40:47Z","isPatch":false,"sender":{"key":"tguyot@gmail.com","avatar":"https://avatars.githubusercontent.com/u/403890?v=4"},"body":"On 2022-07-28 01:42, Johannes Sixt wrote:\n> Am 27.07.22 um 22:21 schrieb CH:\n>> Somehow when copying and pasting a commit from a website to the command\n>> line, a UTF-8 Byte Order Mark (BOM)\n>> [https://en.wikipedia.org/wiki/Byte_order_mark] was appended to one of\n>> the commit ids.  BOMs are invisible, as are many other UTF-8 code\n>> points.  The upshot was that Git didn't like it, and complained bitterly:\n>>\n>>> $ strace -etrace=execve -s 200 git diff\n>>> 038179704f0066aa815d5429221cf381ff4ef289\n>>> 47346a462d8ba40b9a8b073e351c362522c46aa6\n>>>\n>>> execve(\"/usr/bin/git\", [\"git\", \"diff\",\n>>> \"038179704f0066aa815d5429221cf381ff4ef289\\357\\273\\277\",\n>>> \"47346a462d8ba40b9a8b073e351c362522c46aa6\"], 0x7fffec3c4bb0 /* 80 vars\n>>> */) = 0\n>>>\n>>> fatal: ambiguous argument '038179704f0066aa815d5429221cf381ff4ef289':\n>>> unknown revision or path not in the working tree.\n>>> Use '--' to separate paths from revisions, like this:\n>>> 'git <command> [<revision>...] -- [<file>...]'\n>>> +++ exited with 128 +++\n>> Feature request:\n>> ================\n>>\n>> When printing the \"fatal: ambiguous argument '......': ....\", perhaps\n>> escape (url or otherwise) the ambiguous argument when printing it in the\n>> error message, or maybe add a sentence about non-ASCII characters being\n>> found.\n> That's not going to fly, IMHO, because when I type\n>\n>     git diff todo/René\n>\n> I would not want to see\n>\n> fatal: ambiguous argument 'todo/Ren\\303\\251': unknown ...\n\nThis is actually already MUCH better that the OP's example. In his \nexample he has a string that looks like a 40-char hash, and git \ncomplains without showing any of the Unicode gibberish attached to that \nsha1. It would be better if at least it printed something in ASCII with \nescaped bytes in the error message.\n\nMoreover this isn't even close to the issue above - he's talking about a \nno-op, non-printing Unicode marker that crept in. While I do think it \nshouldn't be an issue, it shouldn't even have been passed to git. IMHO \nit should have been stripped by the browser itself on copy, or by the \nterminal on paste... FWIW I'm using rxvt-unicode, and copying this from \nthe terminal doesn't copy the marker but pasting the marker copied from \nChrome is passed on to bash and git.\n\nNB: I also though what if the shell handled it, but that isn't even \nreally a character so not technically suitable for $IFS, and even if we \nconsidered that option it wouldn't really play well with POSIX's \ndefinition of $IFS - how to tell for example between a single Unicode \ncodepoint and a list of binary characters? There is just no definition \nof wide chars for $IFS, not in POSIX nor in recent versions of Bash AFAIK.\n\n\nTL;DR; the issue is IMHO on the browser side, which shouldn't include \nthe marker in the copied text, or maybe on the terminal, BUT when passed \non to git it should at least print the escaped Unicode chars in the \nerror, otherwise it's just too confusing for the user.\n\n\nBTW you actually raise another issue - I do think for file paths git \ncould either recompose (NFC) or decompose (NFD) the strings on storage \nand comparison (which should probably be an option... the current \ndefault for 2.30.2 is to treat them and print them as binary (escaped on \nprint). Consider the following when using core.quotePath=false:\n\n$ touch \"nfc_$(printf '\\xf4')\"\n$ touch \"nfd_$(printf '\\x6f\\xcc\\x82')\"\n$ git add nf[cd]*\n$ git status\nOn branch test\nChanges to be committed:\n   (use \"git restore --staged <file>...\" to unstage)\n     new file:   nfc_ô\n     new file:   nfd_ô\n\nI'm not sure how the Unicode will be translated here, it might depend on \nthe mail client if they even's get sent as-is, but both shows the exact \nsame file name, one in NFD and one in NFC format.\n\nBoth are canonically equivalent and reversible. It appears MacOS already \ndecompose (NFD?) filenames by default and git provides an option to \nrecompose the characters (core.precomposeUnicode) which, according to \nthe manual, is not even usable on Linux...\n\nMore on Unicode normalization: https://unicode.org/reports/tr15/\n\n--\nThomas\n"},{"id":"460139","messageId":"20220728180128.dlhhu7wlbubvnyph@tb-raspi4","threadId":"58232","inReplyTo":"1e454493-ca57-ca2e-7d82-7333a769817e@gmail.com","subject":"Re: Feature request: better error messages when UTF-8 bites","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2022-07-28T18:01:28Z","receivedAt":"2022-07-28T18:01:47Z","isPatch":false,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"[]\n\n>\n> BTW you actually raise another issue - I do think for file paths git could\n> either recompose (NFC) or decompose (NFD) the strings on storage and\n> comparison (which should probably be an option... the current default for\n> 2.30.2 is to treat them and print them as binary (escaped on print).\n> Consider the following when using core.quotePath=false:\n>\n> $ touch \"nfc_$(printf '\\xf4')\"\n\nThis is not valid unicode, isn't it ?\nProbably we need to use the octal version, since not all\nprintf() implementation support hex values starting with 0x,\nbut all of them support octal:\n\nauml=$(printf '\\303\\244')\naumlcdiar=$(printf '\\141\\314\\210')\n\n\n> $ touch \"nfd_$(printf '\\x6f\\xcc\\x82')\"\n> $ git add nf[cd]*\n> $ git status\n> On branch test\n> Changes to be committed:\n>   (use \"git restore --staged <file>...\" to unstage)\n>     new file:   nfc_ô\n>     new file:   nfd_ô\n\n>\n> I'm not sure how the Unicode will be translated here, it might depend on the\n> mail client if they even's get sent as-is, but both shows the exact same\n> file name, one in NFD and one in NFC format.\n\nTranslated by \"whom\" ?\nMost programs do no translate anything here.\n>\n> Both are canonically equivalent and reversible. It appears MacOS already\n> decompose (NFD?) filenames by default and git provides an option to\n> recompose the characters (core.precomposeUnicode) which, according to the\n> manual, is not even usable on Linux...\n\nYes. Technically you can have both under Linux, at least unless you\nare running ZFS, which may be created unicode-aware (or not, that is the default).\n\nBut why do you want to have 2 files on disk with different normalizations ?\n\n\n>\n> More on Unicode normalization: https://unicode.org/reports/tr15/\n>\n> --\n> Thomas\n"}]}