{"thread":{"id":"8588","subject":"Stupid quoting...","startedAt":"2007-06-13T11:30:49Z","lastAt":"2007-06-24T20:25:42Z","messageCount":35,"participants":["David Kastrup","Alex Riesen","Johannes Schindelin","Steven Grimm","Junio C Hamano","Jakub Narebski","Jeff King","Olivier Galibert","Jan Hudec","Robin Rosenberg"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"44920","messageId":"86ir9sw0pi.fsf@lola.quinscape.zz","threadId":"8588","inReplyTo":null,"subject":"Stupid quoting...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-06-13T11:30:49Z","receivedAt":"2007-06-13T11:30:49Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\nHi,\n\nwhat is the point in quoting file names and their characters in\ngit-diff's output?  And what is the recommended way of undoing the\ndamage?\n\nI have something like\n\ngit-diff -M -C --name-status -r master^ master | {\n    while read -r flag name\n    do\n\tcase \"$name\" in *\\\\[0-3][0-7][0-7]*)\n\t\tname=$(echo -e $(echo \"$name\"|sed 's/\\\\\\([0-3][0-7][0-7]\\)/\\\\0\\1/g;s/\\\\\\([^0]\\)/\\\\\\\\\\1/g'))\n\tesac\n        [...]\n\nin order to get through the worst with utf-8 file names, and it is a\ncomplete nuisance (double quotemarks are treated later).\n\nIs there any utility or pipe or invocation that can take a sequence of\nfilenames as printed by git and turn them back into what they actually\nwere in the first place?\n\n-- \nDavid Kastrup\n"},{"id":"44922","messageId":"81b0412b0706130506m8b35f9eje32e2c4e669b348d@mail.gmail.com","threadId":"8588","inReplyTo":"86ir9sw0pi.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-06-13T12:06:48Z","receivedAt":"2007-06-13T12:06:48Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 6/13/07, David Kastrup <dak@gnu.org> wrote:\n>\n> what is the point in quoting file names and their characters in\n> git-diff's output?  And what is the recommended way of undoing the\n> damage?\n>\n\nJust use \"-z\". Everything will be unquoted and separated by \\0\n"},{"id":"44925","messageId":"Pine.LNX.4.64.0706131316390.4059@racer.site","threadId":"8588","inReplyTo":"86ir9sw0pi.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-13T12:21:12Z","receivedAt":"2007-06-13T12:21:12Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 13 Jun 2007, David Kastrup wrote:\n\n> what is the point in quoting file names and their characters in\n> git-diff's output?  And what is the recommended way of undoing the\n> damage?\n\nThe recommended way is not using spaces to begin with. I mean, does \n\"David\" contain spaces? People seem not to see the problem, and fail to \nblame Microsoft for all the damage they have done, introducing that \nstupid, stupid concept of filenames containing spaces, and _enforcing_ it.\n\n> I have something like\n> \n> git-diff -M -C --name-status -r master^ master | {\n>     while read -r flag name\n>     do\n> \tcase \"$name\" in *\\\\[0-3][0-7][0-7]*)\n> \t\tname=$(echo -e $(echo \"$name\"|sed 's/\\\\\\([0-3][0-7][0-7]\\)/\\\\0\\1/g;s/\\\\\\([^0]\\)/\\\\\\\\\\1/g'))\n> \tesac\n>         [...]\n> \n> in order to get through the worst with utf-8 file names, and it is a\n> complete nuisance (double quotemarks are treated later).\n\nPlease understand that the quotes are not there for you, but for \nprocessing by other programs.\n\nHowever, I _suspect_ that you want to do something like\n\n\tname=\"$(echo $name)\"\n\nbecause \"echo\" is exactly one of the programs this quoting was invented \nfor.\n\nCiao,\nDscho\n"},{"id":"45022","messageId":"Pine.LNX.4.64.0706140145450.4059@racer.site","threadId":"8588","inReplyTo":"86ejkgvxmb.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-14T00:51:06Z","receivedAt":"2007-06-14T00:51:06Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[somehow I got the impression your mail did not make it to the list]\n\nOn Wed, 13 Jun 2007, David Kastrup wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > On Wed, 13 Jun 2007, David Kastrup wrote:\n> >\n> >> what is the point in quoting file names and their characters in\n> >> git-diff's output?  And what is the recommended way of undoing the\n> >> damage?\n> >\n> > The recommended way is not using spaces to begin with.\n> \n> Who is talking about spaces?\n\nThat is the common reason for quoting. I mean, really, how many files do \nyou have which contain newlines or backslashes or tabs? Huh?\n\n> > I mean, does \"David\" contain spaces?\n> \n> \"Günter\" contains non-ASCII characters.\n\nAnd \"Guenther\" (sorry, have problems with my mailer, so I simulate it in \nplain ASCII\" does not need quotes, _even_ if containing non-ASCII \ncharacters.\n\nSo what exactly was your point again?\n\n> > People seem not to see the problem, and fail to blame Microsoft for \n> > all the damage they have done, introducing that stupid, stupid concept \n> > of filenames containing spaces, and _enforcing_ it.\n> \n> The concept of UNIX file names is _any_ byte sequence not containing \"/\" \n> or an ASCII NUL.  Microsoft actually prohibits quite a few more \n> characters.  Filenames with spaces first came into serious use under \n> MacOS, the first graphical user interface where no shell and \n> metacharacters interfered with the choice of file names.\n> \n> Blaming Microsoft here is completely ridiculous.\n\nIt is completely unridiculous. Before Microsoft -- in its infinite wisdom \n-- decided to create folders like \"Program Files\", and \"Documents and \nSettings\", and made it the _default_ (of all things) to save its \nridiculous Word documents as \"New Document\", _nobody_ on this planet even \n_thought_ about including stupid whitespace in a filename.\n\nYou can tell that this is true by looking at now-ancient Unix scripts.\n\n> >> I have something like\n> >> \n> >> git-diff -M -C --name-status -r master^ master | {\n> >>     while read -r flag name\n> >>     do\n> >> \tcase \"$name\" in *\\\\[0-3][0-7][0-7]*)\n> >> \t\tname=$(echo -e $(echo \"$name\"|sed 's/\\\\\\([0-3][0-7][0-7]\\)/\\\\0\\1/g;s/\\\\\\([^0]\\)/\\\\\\\\\\1/g'))\n> >> \tesac\n> >>         [...]\n> >> \n> >> in order to get through the worst with utf-8 file names, and it is a\n> >> complete nuisance (double quotemarks are treated later).\n> >\n> > Please understand that the quotes are not there for you, but for \n> > processing by other programs.\n> >\n> > However, I _suspect_ that you want to do something like\n> >\n> > \tname=\"$(echo $name)\"\n> >\n> > because \"echo\" is exactly one of the programs this quoting was invented \n> > for.\n> \n> Only that it does not work with echo.  echo requires \\0NNN for octal\n> escapes, not \\NNN, and then only when \"echo -e\" is used.\n\nUm. How does that apply here? Git only does quoting so that programs like \necho get it right, when passed the name? No funny \\0NNN or \\NNN or \nwhatever?\n\n> You are really haphazard in distributing your blame.\n> \n> Can you actually name a program that would work with the default\n> output of git here?\n\necho.\n\nCiao,\nDscho\n"},{"id":"45023","messageId":"4670948B.7070407@midwinter.com","threadId":"8588","inReplyTo":"Pine.LNX.4.64.0706131316390.4059@racer.site","subject":"Re: Stupid quoting...","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-06-14T01:06:19Z","receivedAt":"2007-06-14T01:06:19Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Johannes Schindelin wrote:\n> The recommended way is not using spaces to begin with. I mean, does \n> \"David\" contain spaces? People seem not to see the problem, and fail to \n> blame Microsoft for all the damage they have done, introducing that \n> stupid, stupid concept of filenames containing spaces, and _enforcing_ it.\n>   \n\nTo be fair, Microsoft did not invent the concept of filenames with \nspaces. Even UNIX has, I believe, always allowed them, though you risked \nrunning into buggy scripts misbehaving if you used them. And really, \nfilenames with spaces are only a nuisance in a text-based command line \nscripting environment, and only then because someone early on decided to \nuse space rather than some other metacharacter as the only available \ndelimiter between command arguments in scripts.\n\nFor regular users who aren't writing shell scripts, long-form filenames \nmean a document's name is the same as its title, which is hugely helpful \nfrom a user interface point of view: \"February Marketing Budget\" is much \nmore self-documenting as a filename than \"febmkbgt\" or whatever they \nwould have had to choose in the old days (and remember, UNIX had a \n14-character filename size limit in the early days; long filenames \nweren't introduced until BSD.)\n\nI view filenames as primarily for human consumption, so they should act \nthe way humans expect names to act. Computers can just use inode numbers \nor file IDs to refer to files (like, say, SHA1 hashes) -- the pretty \nnames are all for the benefit of us meat-brained entities.\n\nThen again, that means I also think case sensitivity in filenames was a \nbad design choice. To use your \"people's names\" analogy, if you ask any \nrandom person whether \"Billy-bob Thornton\" and \"Billy-Bob Thornton\" are \nthe same name, they'll almost always say yes; most people don't consider \ncapital letters and lower-case letters to be different letters, just \ndifferent forms of the same letters. And I know how popular *that* \nopinion is around here...\n\n-Steve\n"},{"id":"45025","messageId":"Pine.LNX.4.64.0706140211070.4059@racer.site","threadId":"8588","inReplyTo":"4670948B.7070407@midwinter.com","subject":"Re: Stupid quoting...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-14T01:12:47Z","receivedAt":"2007-06-14T01:12:47Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 13 Jun 2007, Steven Grimm wrote:\n\n> Johannes Schindelin wrote:\n> > The recommended way is not using spaces to begin with. I mean, does \"David\"\n> > contain spaces? People seem not to see the problem, and fail to blame\n> > Microsoft for all the damage they have done, introducing that stupid,\n> > stupid concept of filenames containing spaces, and _enforcing_ it.\n> >   \n> \n> To be fair, Microsoft did not invent the concept of filenames with spaces.\n\nI didn't say that, did I?\n\nThey _forced_ the use onto the world. That's what I was complaining about.\n\n> Even UNIX has, I believe, always allowed them, though you risked running \n> into buggy scripts misbehaving if you used them. And really, filenames \n> with spaces are only a nuisance in a text-based command line scripting \n> environment, and only then because someone early on decided to use space \n> rather than some other metacharacter as the only available delimiter \n> between command arguments in scripts.\n\nOkay, Steven Grimm. How do you think _I_ can tell that Steven is your name \nfrom looking at your _full_ name \"Steven Grimm\"? Huh?\n\nExactly. I split at the space.\n\nCiao,\nDscho\n"},{"id":"45026","messageId":"467097B6.3030604@midwinter.com","threadId":"8588","inReplyTo":"Pine.LNX.4.64.0706140211070.4059@racer.site","subject":"Re: Stupid quoting...","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-06-14T01:19:50Z","receivedAt":"2007-06-14T01:19:50Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"Johannes Schindelin wrote:\n> Okay, Steven Grimm. How do you think _I_ can tell that Steven is your name \n> from looking at your _full_ name \"Steven Grimm\"? Huh?\n>\n> Exactly. I split at the space.\n>\n>   \n\nAt the risk of drawing the conversation way off topic: What's Mary Ann \nSummers' first name? (Hint: It's not \"Mary.\")\n\n-Steve\n"},{"id":"45027","messageId":"Pine.LNX.4.64.0706140233350.4059@racer.site","threadId":"8588","inReplyTo":"467097B6.3030604@midwinter.com","subject":"Re: Stupid quoting...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-14T01:34:55Z","receivedAt":"2007-06-14T01:34:55Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 13 Jun 2007, Steven Grimm wrote:\n\n> Johannes Schindelin wrote:\n> > Okay, Steven Grimm. How do you think _I_ can tell that Steven is your name\n> > from looking at your _full_ name \"Steven Grimm\"? Huh?\n> > \n> > Exactly. I split at the space.\n> \n> At the risk of drawing the conversation way off topic: What's Mary Ann\n> Summers' first name? (Hint: It's not \"Mary.\")\n\nThe first first name _is_ Mary. Maybe it is not the name you shout when \ncalling her.\n\nBut that is irrelevant. The names of Mary Ann Summers are separated by \nspaces. Period.\n\nCiao,\nDscho\n"},{"id":"45041","messageId":"86wsy76p4v.fsf@lola.quinscape.zz","threadId":"8588","inReplyTo":"Pine.LNX.4.64.0706140145450.4059@racer.site","subject":"Re: Stupid quoting...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-06-14T06:12:16Z","receivedAt":"2007-06-14T06:12:16Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Hi,\n>\n> [somehow I got the impression your mail did not make it to the list]\n>\n> On Wed, 13 Jun 2007, David Kastrup wrote:\n>\n>> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n>> \n>> > On Wed, 13 Jun 2007, David Kastrup wrote:\n>> >\n>> >> what is the point in quoting file names and their characters in\n>> >> git-diff's output?  And what is the recommended way of undoing the\n>> >> damage?\n>> >\n>> > The recommended way is not using spaces to begin with.\n>> \n>> Who is talking about spaces?\n>\n> That is the common reason for quoting. I mean, really, how many files do \n> you have which contain newlines or backslashes or tabs? Huh?\n\nI am talking about non-ASCII characters.\n\n>\n>> > I mean, does \"David\" contain spaces?\n>> \n>> \"Günter\" contains non-ASCII characters.\n>\n> And \"Guenther\" (sorry, have problems with my mailer, so I simulate\n> it in plain ASCII\" does not need quotes, _even_ if containing\n> non-ASCII characters.\n>\n> So what exactly was your point again?\n\nYou _are_ aware that git writes out \\303\\274 (8 characters: 2\nbackslashes and 6 digits) instead of ü in a file name?  And I am\ntalking about a pure utf-8 locale.\n\nLANG=en_US.UTF-8\nLC_CTYPE=\"en_US.UTF-8\"\nLC_NUMERIC=\"en_US.UTF-8\"\nLC_TIME=\"en_US.UTF-8\"\nLC_COLLATE=\"en_US.UTF-8\"\nLC_MONETARY=\"en_US.UTF-8\"\nLC_MESSAGES=\"en_US.UTF-8\"\nLC_PAPER=\"en_US.UTF-8\"\nLC_NAME=\"en_US.UTF-8\"\nLC_ADDRESS=\"en_US.UTF-8\"\nLC_TELEPHONE=\"en_US.UTF-8\"\nLC_MEASUREMENT=\"en_US.UTF-8\"\nLC_IDENTIFICATION=\"en_US.UTF-8\"\nLC_ALL=\n\nMy point was that these octal escape sequences are utterly pointless.\n\n>> > People seem not to see the problem, and fail to blame Microsoft for \n>> > all the damage they have done, introducing that stupid, stupid concept \n>> > of filenames containing spaces, and _enforcing_ it.\n>> \n>> The concept of UNIX file names is _any_ byte sequence not\n>> containing \"/\" or an ASCII NUL.  Microsoft actually prohibits quite\n>> a few more characters.  Filenames with spaces first came into\n>> serious use under MacOS, the first graphical user interface where\n>> no shell and metacharacters interfered with the choice of file\n>> names.\n>> \n>> Blaming Microsoft here is completely ridiculous.\n>\n> It is completely unridiculous. Before Microsoft -- in its infinite\n> wisdom -- decided to create folders like \"Program Files\", and\n> \"Documents and Settings\", and made it the _default_ (of all things)\n> to save its ridiculous Word documents as \"New Document\", _nobody_ on\n> this planet even _thought_ about including stupid whitespace in a\n> filename.\n>\n> You can tell that this is true by looking at now-ancient Unix\n> scripts.\n\nYou are making a spectacle of yourself.  Do you even read what you are\nreplying to?  When spaces became commonplace in _MacOS_, _MacOS_ was\nby no means Unix-based.  Microsoft only followed the trend (with a\ndelay of several years, by the way) when imitating the MacOS GUI.\n\n>> >> I have something like\n>> >> \n>> >> git-diff -M -C --name-status -r master^ master | {\n>> >>     while read -r flag name\n>> >>     do\n>> >> \tcase \"$name\" in *\\\\[0-3][0-7][0-7]*)\n>> >> \t\tname=$(echo -e $(echo \"$name\"|sed 's/\\\\\\([0-3][0-7][0-7]\\)/\\\\0\\1/g;s/\\\\\\([^0]\\)/\\\\\\\\\\1/g'))\n>> >> \tesac\n>> >>         [...]\n>> >> \n>> >> in order to get through the worst with utf-8 file names, and it is a\n>> >> complete nuisance (double quotemarks are treated later).\n>> >\n>> > Please understand that the quotes are not there for you, but for \n>> > processing by other programs.\n>> >\n>> > However, I _suspect_ that you want to do something like\n>> >\n>> > \tname=\"$(echo $name)\"\n>> >\n>> > because \"echo\" is exactly one of the programs this quoting was invented \n>> > for.\n>> \n>> Only that it does not work with echo.  echo requires \\0NNN for\n>> octal escapes, not \\NNN, and then only when \"echo -e\" is used.\n>\n> Um. How does that apply here? Git only does quoting so that programs\n> like echo get it right, when passed the name? No funny \\0NNN or \\NNN\n> or whatever?\n\ngit puts out funny \\NNN quotes.  That's what I am complaining about.\n\n>> You are really haphazard in distributing your blame.\n>> \n>> Can you actually name a program that would work with the default\n>> output of git here?\n>\n> echo.\n\nIt doesn't, since it does not interpret the \\NNN escape sequences that\ngit chooses to output.\n\n-- \nDavid Kastrup\n"},{"id":"45044","messageId":"81b0412b0706140006v601b345re7dc0e58488cf61e@mail.gmail.com","threadId":"8588","inReplyTo":"86wsy76p4v.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-06-14T07:06:30Z","receivedAt":"2007-06-14T07:06:30Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 6/14/07, David Kastrup <dak@gnu.org> wrote:\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> >> Can you actually name a program that would work with the default\n> >> output of git here?\n> >\n> > echo.\n>\n> It doesn't, since it does not interpret the \\NNN escape sequences that\n> git chooses to output.\n\nHave you tried that -z switch yet?\n"},{"id":"45052","messageId":"7vlkemapk8.fsf@assigned-by-dhcp.pobox.com","threadId":"8588","inReplyTo":"Pine.LNX.4.64.0706131316390.4059@racer.site","subject":"Re: Stupid quoting...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-06-14T08:49:27Z","receivedAt":"2007-06-14T08:49:27Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> On Wed, 13 Jun 2007, David Kastrup wrote:\n>\n>> what is the point in quoting file names and their characters in\n>> git-diff's output?  And what is the recommended way of undoing the\n>> damage?\n>\n> The recommended way is not using spaces to begin with. I mean, does \n> \"David\" contain spaces? People seem not to see the problem, and fail to \n> blame Microsoft for all the damage they have done, introducing that \n> stupid, stupid concept of filenames containing spaces, and _enforcing_ it.\n\nWhy are you talking about spaces ;-)?\n\nThere are a few things to note, but the first thing is that mere\nspaces do not trigger quoting.  A tab (HT) does, so do non ASCII\ncharacters.  The second thing is that we do this quoting for\nvarious good reasons, and it is not likely to change.\n\nAs Alex mentions, the most safe way for programs to read is to\nread from the -z format.  However, even if you are capable to do\nso, it may be inconvenient in some languages (mainstream\nlanguages like C and Perl are not among them).  Not quoting SP\nis a conscious decision, as SP in filenames are rather common,\nmore common than non ASCII and much more common than HT.\n\nThe \"raw\" formats \"ls-files -s\", \"ls-tree\" and \"diff --raw\"\nproduce are designed to put names at the end, and typically\ndelimited with a HT, so that \"lazy\" scripts can use cut (whose\ndefault delimiter is a HT) to pick out pieces from its output.\nAnd plumbing tools reading from the standard input (most\nnotably, \"update-index --stdin\") know how to unquote them.  In\npractice, not many people use non ASCII in pathnames and expect\nthem work sanely for everybody, so loosely written scripts, as\nlong as they cut at HT to pick out the pathname part, \"mostly\"\nwork (I think traditional core git scripts are safe, I suspect\nsome contributed ones shipped with git core may not be, Cogito\nused to be very unsafe but it was audited and became much safer\nbefore it got discontinued).\n\nThe pathname quoting rules in textual output was chosen\nprimarily to make diff output safer, as one of the most\nimportant workflow git supports is e-mailable patches.\n\nGNU patch treats HT on \"+++ name\"/\"--- name\" lines as the end of\nname (and after HT comes timestamp), but the timestamp part is\ntreated as optional, which introduces ambiguities and confusion.\nThe issue was discussed some time ago (check the list archive\nfor discussion among I, Linus and Paul Eggert -- the GNU diff\nand patch maintainer) and the quoting rules we use now is\nconsistent with what the diff and patch plan to use.  The update\non the GNU side may have already happened, it may not have.\n\nWhen a patch appears in an e-mail, you would need to be aware\nthat not everybody has the luxury of living in UTF-8 only world.\nYour commit message and cover letter may be in one encoding, the\npathnames that appear in diff headers may be in your filesystem\nencoding, and the patch text that appear as the diff payload may\nbe in another document specific encoding.  All three could be\ndifferent (worse, a patch that touch more than one file can\ncarry different encodings in the payload part), and mixing\ncharacter set in a single piece of e-mail confuses people's MUA\nand tends to mangle messages.  Quoting non ASCII characters in\npathnames, even they are perfectly valid and ordinary UTF-8\nstrings, is to eliminate one element in the above three as a\npossible source of worries.\n"},{"id":"45053","messageId":"81b0412b0706140151o4e0d3a2fkfde1cd2726f9b14e@mail.gmail.com","threadId":"8588","inReplyTo":"86hcpb6lr6.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-06-14T08:51:57Z","receivedAt":"2007-06-14T08:51:57Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 6/14/07, David Kastrup <dak@gnu.org> wrote:\n> \"Alex Riesen\" <raa.lkml@gmail.com> writes:\n>\n> > On 6/14/07, David Kastrup <dak@gnu.org> wrote:\n> >> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> >> >> Can you actually name a program that would work with the default\n> >> >> output of git here?\n> >> >\n> >> > echo.\n> >>\n> >> It doesn't, since it does not interpret the \\NNN escape sequences that\n> >> git chooses to output.\n> >\n> > Have you tried that -z switch yet?\n>\n> What has that to do with \"the default output of git\"?\n>\n> Yes, in my application I _will_ be using -z\n\nJust checking.\n\n>  (in connection with the rather hackish read -d '' name\n> command from bash which is not really documented) but that does not\n> change the fact that the default output is broken.  There is no reason\n> whatsoever to use octal quotes for non-ASCII characters.  Neither\n> programs nor humans are better off by that, and none of the derision\n> bestowed upon me changes that.\n\nWell, fix that. How do you think _should_ it be?\n\nIt's just up until now you are only complaining.\nNo _useful_ idea came from you.\n"},{"id":"45179","messageId":"f51irh$shq$1@sea.gmane.org","threadId":"8588","inReplyTo":"86ir9sw0pi.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-06-16T21:03:05Z","receivedAt":"2007-06-16T21:03:05Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"David Kastrup wrote:\n\n> what is the point in quoting file names and their characters in\n> git-diff's output? \n\n7-bit email.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"45259","messageId":"86ir9l1ylc.fsf@lola.quinscape.zz","threadId":"8588","inReplyTo":"f51irh$shq$1@sea.gmane.org","subject":"Re: Stupid quoting...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-06-18T08:00:31Z","receivedAt":"2007-06-18T08:00:31Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> David Kastrup wrote:\n>\n>> what is the point in quoting file names and their characters in\n>> git-diff's output? \n>\n> 7-bit email.\n\nI think it can be reasonably safely assumed that people using 8-bit\ncharacters in file names will not refrain from using them in the files\nthemselves: file names usually are chosen descriptive of the contents,\nand so rarely are in a different language.  So I don't see what\nquoting such characters in file names is supposed to buy with regard\nto diff output in 7-bit email.\n\n-- \nDavid Kastrup\n"},{"id":"45273","messageId":"20070618161933.GC4662@sigill.intra.peff.net","threadId":"8588","inReplyTo":"86ir9l1ylc.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-06-18T16:19:33Z","receivedAt":"2007-06-18T16:19:33Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 18, 2007 at 10:00:31AM +0200, David Kastrup wrote:\n\n> > 7-bit email.\n> \n> I think it can be reasonably safely assumed that people using 8-bit\n> characters in file names will not refrain from using them in the files\n\nNot to mention the commit messages.\n\nBut more importantly, diffs aren't necessarily going through mail. When\nI run 'git-show', this isn't useful to me:\n\ndiff --git \"a/ni\\303\\261o\" \"b/ni\\303\\261o\"\n\nI can only imagine how git-show might look to somebody using all-utf8\nfilenames (such as Japanese).\n\n-Peff\n"},{"id":"45304","messageId":"Pine.LNX.4.64.0706190156110.4059@racer.site","threadId":"8588","inReplyTo":"86ir9l1ylc.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-19T01:00:44Z","receivedAt":"2007-06-19T01:00:44Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 18 Jun 2007, David Kastrup wrote:\n\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n> > David Kastrup wrote:\n> >\n> >> what is the point in quoting file names and their characters in \n> >> git-diff's output?\n> >\n> > 7-bit email.\n> \n> I think it can be reasonably safely assumed that people using 8-bit\n> characters in file names will not refrain from using them in the files\n> themselves: [...]\n\nHowever, please realise that chances are very good that none of these \n8-bit unclean things show in the diff.\n\nBesides, the proper fix would probably involve making none-8-bit-clean \ndiffs binary diffs (for FORMAT_EMAIL only, of course).\n\n> So I don't see what quoting such characters in file names is supposed to \n> buy with regard to diff output in 7-bit email.\n\nBut isn't that obvious? Even if the diffs are not 7-bit clean, which I \nconsider as an error, quoting the file names is already half what is \nrequired.\n\nDon't just throw away backwards compatibility, only because it does not \nfit your wishes.\n\nCiao,\nDscho\n"},{"id":"45327","messageId":"86sl8owfqj.fsf@lola.quinscape.zz","threadId":"8588","inReplyTo":"Pine.LNX.4.64.0706190156110.4059@racer.site","subject":"Re: Stupid quoting...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-06-19T07:44:20Z","receivedAt":"2007-06-19T07:44:20Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> On Mon, 18 Jun 2007, David Kastrup wrote:\n>\n>> Jakub Narebski <jnareb@gmail.com> writes:\n>> \n>> > David Kastrup wrote:\n>> >\n>> >> what is the point in quoting file names and their characters in \n>> >> git-diff's output?\n>> >\n>> > 7-bit email.\n>> \n>> I think it can be reasonably safely assumed that people using 8-bit\n>> characters in file names will not refrain from using them in the files\n>> themselves: [...]\n>\n> However, please realise that chances are very good that none of these \n> 8-bit unclean things show in the diff.\n\nPuh-leaze.  So you prefer a behavior which makes it harder to notice\nproblems, on the chance that it may sometimes work by accident?\n\nIf you want to process diffs, you need an 8-bit clean (and\nspace-preserving) channel, period.  This is the task of mail\nencapsulation, not of the diff utility.\n\n> Besides, the proper fix would probably involve making\n> none-8-bit-clean diffs binary diffs (for FORMAT_EMAIL only, of\n> course).\n\nThis is so utterly absurd for people working on non-English documents\nthat I get the expression you are pulling people's legs considering\nyour Email address.\n\n>> So I don't see what quoting such characters in file names is\n>> supposed to buy with regard to diff output in 7-bit email.\n>\n> But isn't that obvious? Even if the diffs are not 7-bit clean, which\n> I consider as an error, quoting the file names is already half what\n> is required.\n\nWhat is required is a reliable mail channel, and there are a lot of\ntools for that, from uuencode to various MIME standards and\nencapsulation methods.  The right tool for the right job.  Everything\nelse is a mistake because it makes life harder for everyone, not just\nthose using mail, for no good purpose.\n\n> Don't just throw away backwards compatibility, only because it does\n> not fit your wishes.\n\nThere is no backwards compatibility involved here _at_ _all_.  No\ncurrent tool can process the quoted mess, not even humans (random\noctal escape sequences are not more readable than characters, or we\nnever would have progressed beyond ASCII).\n\nSo you are not talking about backward compatibility, but rather\ngratuitous forward _incompatibility_, and nobody is better off by the\nlatter.  There is no point in making life harder for people using\nnon-ASCII characters when there is absolutely no benefit whatsoever\ninvolved for those restricting themselves to ASCII characters.\n\n-- \nDavid Kastrup\n"},{"id":"45336","messageId":"Pine.LNX.4.64.0706191048570.4059@racer.site","threadId":"8588","inReplyTo":"86sl8owfqj.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-19T09:50:31Z","receivedAt":"2007-06-19T09:50:31Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 19 Jun 2007, David Kastrup wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > Don't just throw away backwards compatibility, only because it does \n> > not fit your wishes.\n> \n> There is no backwards compatibility involved here _at_ _all_.\n\nI was not talking about Git here. The specification for SMTP is not going \nto change just because you want it. There are still mail servers out there \nwhich speak 7-bit, and the standard requires you to cope with them.\n\nCiao,\nDscho\n"},{"id":"45379","messageId":"20070619205313.GA90303@dspnet.fr.eu.org","threadId":"8588","inReplyTo":"Pine.LNX.4.64.0706191048570.4059@racer.site","subject":"Re: Stupid quoting...","fromName":"Olivier Galibert","fromEmail":"galibert@pobox.com","sentAt":"2007-06-19T20:53:14Z","receivedAt":"2007-06-19T20:53:14Z","isPatch":false,"sender":{"key":"galibert@pobox.com","avatar":null},"body":"On Tue, Jun 19, 2007 at 10:50:31AM +0100, Johannes Schindelin wrote:\n> Hi,\n> \n> On Tue, 19 Jun 2007, David Kastrup wrote:\n> \n> > Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> > \n> > > Don't just throw away backwards compatibility, only because it does \n> > > not fit your wishes.\n> > \n> > There is no backwards compatibility involved here _at_ _all_.\n> \n> I was not talking about Git here. The specification for SMTP is not going \n> to change just because you want it. There are still mail servers out there \n> which speak 7-bit, and the standard requires you to cope with them.\n\nThere are standards to send 8-bit into 7-bit for email, and \\xxx is in\nnone of them.  And 8-to-7 encoding for email is not git's job in any\ncase unless git speaks SMTP directly.  8-to-7 is the mail client\nresponsability.\n\n  OG.\n"},{"id":"45393","messageId":"Pine.LNX.4.64.0706200307070.4059@racer.site","threadId":"8588","inReplyTo":"86645kutow.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-06-20T02:19:54Z","receivedAt":"2007-06-20T02:19:54Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[sorry for responding so late, your mail got stuck in the GWB-like spam \nfilter.]\n\nOn Tue, 19 Jun 2007, David Kastrup wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > Hi,\n> >\n> > On Tue, 19 Jun 2007, David Kastrup wrote:\n> >\n> >> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> >> \n> >> > Don't just throw away backwards compatibility, only because it does \n> >> > not fit your wishes.\n> >> \n> >> There is no backwards compatibility involved here _at_ _all_.\n> >\n> > I was not talking about Git here. The specification for SMTP is not\n> > going to change just because you want it. There are still mail\n> > servers out there which speak 7-bit, and the standard requires you\n> > to cope with them.\n> \n> Is there a reason you elide all the relevant material before replying?\n> I repeat: this is the task of MIME, uuencode or a number of other\n> mechanisms.\n\nThe problem there, of course, is that you still might want to reply to the \npatch, even if the name was chosen as non-ASCII (which is a sin, if you \nbelieve in UNIX).\n\nUsually, comments are not done on the filenames, so they can be as escaped \nas they want in an email, as long as the commenter still recognizes their \nnames.\n\n> git is not a mail transport system, and there are far too many other \n> problems in unarmored mail (like spaces, wrapping and other stuff) that \n> it would make any sense to mangle diffs and other material in a manner \n> that makes it quite unprocessable for _both_ human readers as well as \n> scripts intended to process them.\n\nThere you have a point. If the name is non-ASCII, it uses a specific \nencoding. if the human reader has a different encoding set in her display, \nis it any better to display garbled characters (possibly leaving the \nconsole in a corrupted state), or to display escaped characters?\n\nAnd scripts have been known to get encodings all wrong, so I think the \nescaping is the best way out, absent a perfect knowledge of what encoding \nthe file name was meant for.\n\n> Anyway, it has become quite clear from this exchange that you have \n> already made the decision not to be convinced by me and will not be \n> deterred from that, even though the problem is not the one you initially \n> tried deriding me for (spaces in filenames).\n\nI am sorry. No, really, I am sorry that you received it as derision. By \nall means, it was _not_ meant as that. The problem was on my side, not \nyours: I simply did not get that you were talking about non-ASCII \ncharacters, even if you were talking about them.\n\n> Hopefully some developer with less of an attitude towards non-ASCII \n> usage will find himself able to follow the arguments with some more \n> objectivity.\n> \n> I don't see our discourse leading anywhere: the points have been made.\n\nI would really, really, really like to see a solution. Alas, I cannot \nthink of one, other than _forcing_ the developers to use ASCII-only \nfilenames.\n\nNote that there is no convention yet in Git to state which encoding your \nfilenames are supposed to use. And in fact, we already had a fine example \nin git.git why this is particularly difficult. MacOSX is too clever to be \ntrue, in that it gladly takes filenames in one encoding, but reads those \nfilenames out in _another_ encoding. Thus, a \"git add <filename>\" can well \nend up in git-status saying that a file was deleted, and another file \n(actually the same, but in a different encoding) is untracked.\n\nAgain, I would be _so_ glad if you solved the problem, now that I actually \nunderstand it.\n\nCiao,\nDscho\n"},{"id":"45408","messageId":"7vd4zrw3k4.fsf@assigned-by-dhcp.pobox.com","threadId":"8588","inReplyTo":"Pine.LNX.4.64.0706200307070.4059@racer.site","subject":"Re: Stupid quoting...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-06-20T06:19:39Z","receivedAt":"2007-06-20T06:19:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> I don't see our discourse leading anywhere: the points have been made.\n>\n> I would really, really, really like to see a solution. Alas, I cannot \n> think of one, other than _forcing_ the developers to use ASCII-only \n> filenames.\n>\n> Note that there is no convention yet in Git to state which encoding your \n> filenames are supposed to use. And in fact, we already had a fine example \n> in git.git why this is particularly difficult. MacOSX is too clever to be \n> true, in that it gladly takes filenames in one encoding, but reads those \n> filenames out in _another_ encoding. Thus, a \"git add <filename>\" can well \n> end up in git-status saying that a file was deleted, and another file \n> (actually the same, but in a different encoding) is untracked.\n\nBy the way, the pathname quoting done by \"diff\" does not even\nattempt to tackle that.  I already explained why in the thread\nso I would not repeat myself.\n\nHaving said that, the absolute minimum that needs to be quoted\nare double-quote (because it is used by quoting as agreed with\nGNU diff/patch maintainer), backslash (used to introduce C-like\nquoting), newline and horizontal tab (makes \"patch\" confused, as\nit would make it ambiguous where the pathname ends), so I am not\nopposed to a patch that introduces a new mode, probably on by\ndefault _unless_ we are generating --format=email, that does not\nquote high byte values.  That would solve \"My UTF-8 filenames\nare unreadable on my terminal\" problem.\n"},{"id":"45418","messageId":"86y7ifoykt.fsf@lola.quinscape.zz","threadId":"8588","inReplyTo":"7vd4zrw3k4.fsf@assigned-by-dhcp.pobox.com","subject":"Re: Stupid quoting...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-06-20T07:49:06Z","receivedAt":"2007-06-20T07:49:06Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n>\n>>> I don't see our discourse leading anywhere: the points have been made.\n>>\n>> I would really, really, really like to see a solution. Alas, I\n>> cannot think of one, other than _forcing_ the developers to use\n>> ASCII-only filenames.\n\nAnd ASCII-only files.  Just eradicate that dreaded Bit 7 from the world.\n\n>> Note that there is no convention yet in Git to state which encoding\n>> your filenames are supposed to use. And in fact, we already had a\n>> fine example in git.git why this is particularly difficult. MacOSX\n>> is too clever to be true, in that it gladly takes filenames in one\n>> encoding, but reads those filenames out in _another_\n>> encoding. Thus, a \"git add <filename>\" can well end up in\n>> git-status saying that a file was deleted, and another file\n>> (actually the same, but in a different encoding) is untracked.\n>\n> Having said that, the absolute minimum that needs to be quoted are\n> double-quote (because it is used by quoting as agreed with GNU\n> diff/patch maintainer), backslash (used to introduce C-like\n> quoting),\n> newline and horizontal tab (makes \"patch\" confused, as it would make\n> it ambiguous where the pathname ends), so I am not opposed to a\n> patch that introduces a new mode, probably on by default _unless_ we\n> are generating --format=email, that does not quote high byte values.\n\nI think it would be ok to quote non-graphic characters with octal\nescape sequences.  On ASCII-based systems, those are the characters\n0x00 to 0x1f.  They don't have a visual representation of their own,\nanyway.  _IF_ they appear in filenames, it is certainly a case\ninvolved with excessive cleverness and/or garbage.  I'd leave the rest\nalone.\n\n> That would solve \"My UTF-8 filenames are unreadable on my terminal\"\n> problem.\n\nBut there is no point if the most primitive of mail readers does a\nbetter job than listing the directory will.\n\n7-Bit terminals are the wrong thing to use for manipulating\n8-bit-encoded files, period.  And the escape sequences for 8-bit\nterminals are quite certain to start with characters in the 0x00 to\n0x1f range.\n\n-- \nDavid Kastrup\n"},{"id":"45420","messageId":"f5ap5r$sj7$2@sea.gmane.org","threadId":"8588","inReplyTo":"86y7ifoykt.fsf@lola.quinscape.zz","subject":"Re: Stupid quoting...","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-06-20T08:40:26Z","receivedAt":"2007-06-20T08:40:26Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"David Kastrup wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n>> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n>>\n>>>> I don't see our discourse leading anywhere: the points have been made.\n>>>\n>>> I would really, really, really like to see a solution. Alas, I\n>>> cannot think of one, other than _forcing_ the developers to use\n>>> ASCII-only filenames.\n>>> Note that there is no convention yet in Git to state which encoding\n>>> your filenames are supposed to use. And in fact, we already had a\n>>> fine example in git.git why this is particularly difficult. MacOSX\n>>> is too clever to be true, in that it gladly takes filenames in one\n>>> encoding, but reads those filenames out in _another_\n>>> encoding. Thus, a \"git add <filename>\" can well end up in\n>>> git-status saying that a file was deleted, and another file\n>>> (actually the same, but in a different encoding) is untracked.\n>>\n>> Having said that, the absolute minimum that needs to be quoted are\n>> double-quote (because it is used by quoting as agreed with GNU\n>> diff/patch maintainer), backslash (used to introduce C-like\n>> quoting),\n>> newline and horizontal tab (makes \"patch\" confused, as it would make\n>> it ambiguous where the pathname ends), so I am not opposed to a\n>> patch that introduces a new mode, probably on by default _unless_ we\n>> are generating --format=email, that does not quote high byte values.\n> \n> I think it would be ok to quote non-graphic characters with octal\n> escape sequences.  On ASCII-based systems, those are the characters\n> 0x00 to 0x1f.  They don't have a visual representation of their own,\n> anyway.  _IF_ they appear in filenames, it is certainly a case\n> involved with excessive cleverness and/or garbage.  I'd leave the rest\n> alone.\n\nBy the way, ls(1) has its --quoting-style=WORD option, why shouldn't\ngit-diff and friends (including git-format-patch) have the same? And we\ncould change the default later on...\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"45423","messageId":"86tzt3ovbo.fsf@lola.quinscape.zz","threadId":"8588","inReplyTo":"f5ap5r$sj7$2@sea.gmane.org","subject":"Re: Stupid quoting...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-06-20T08:59:23Z","receivedAt":"2007-06-20T08:59:23Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> By the way, ls(1) has its --quoting-style=WORD option, why shouldn't\n> git-diff and friends (including git-format-patch) have the same? And\n> we could change the default later on...\n\nBecause interpreting a diff means interpreting both file names as well\nas contents.  It does not make much sense to use different forms of\nescaping (\\01a and similar) here, though in the diff command line,\nsome additional quoting might be called for.\n\nIt is also worth noting that bash's echo -e can interpret octal\nescapes only when they start with \\0, and the quoted 3-character forms\nof 0x00-0x1f incidentally do start in this manner.  There is still\npotential for misinterpretation if an escaped character is immediately\nfollowed by a digit.  Since octal ASCII digits are in the range 060 to\n067, one can get around this problem by continuing to escape\ncharacters until one hits a non-octal-digit.  So there is at least a\nreasonable builtin way for bash scripts to translate the three-digit\noctal escapes for 0x00 to 0x1f uniquely into the proper corresponding\nstrings.\n\nWith regard to escaping: unless used unarmored in Email (a bad idea)\nor on a terminal, it might be easiest (for post-processors) to\ncompletely refrain from escaping (in effect ignoring the\nnon-printability of characters) and just apply a minimal level of\nquoting on the file names.\n\n-- \nDavid Kastrup\n"},{"id":"45642","messageId":"20070624065008.GA6979@efreet.light.src","threadId":"8588","inReplyTo":"7vd4zrw3k4.fsf@assigned-by-dhcp.pobox.com","subject":"Re: Stupid quoting...","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2007-06-24T06:50:08Z","receivedAt":"2007-06-24T06:50:08Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Tue, Jun 19, 2007 at 23:19:39 -0700, Junio C Hamano wrote:\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> >> I don't see our discourse leading anywhere: the points have been made.\n> >\n> > I would really, really, really like to see a solution. Alas, I cannot \n> > think of one, other than _forcing_ the developers to use ASCII-only \n> > filenames.\n> >\n> > Note that there is no convention yet in Git to state which encoding your \n> > filenames are supposed to use. And in fact, we already had a fine example \n> > in git.git why this is particularly difficult. MacOSX is too clever to be \n> > true, in that it gladly takes filenames in one encoding, but reads those \n> > filenames out in _another_ encoding. Thus, a \"git add <filename>\" can well \n> > end up in git-status saying that a file was deleted, and another file \n> > (actually the same, but in a different encoding) is untracked.\n\nI saw bazaar folks discussing this MacOSX issue. Basically in MacOSX\nfilenames are *unicode* strings (just as they are in Windows, btw). Unicode,\nfor compatibility reasons allows expressing many characters in multiple forms\n-- composed and decomposed. For example 'á' can be expressed as '\\u00e1'\n('\\xc3\\xa1' in utf-8) or as 'a\\u0301' ('a\\xcc\\x81' in utf-8).\n\nMaxOSX opts to, in accord with unicode standard, treat such representations\nas equal and it does so by normalizing all filenames to one form. I don't\nknow whether it uses compatibility normalization and I believe it uses the\ndecomposed form (which makes the issue immediately obvious, because most\nprograms work in composed form).\n\n> By the way, the pathname quoting done by \"diff\" does not even\n> attempt to tackle that.  I already explained why in the thread\n> so I would not repeat myself.\n> \n> Having said that, the absolute minimum that needs to be quoted\n> are double-quote (because it is used by quoting as agreed with\n> GNU diff/patch maintainer), backslash (used to introduce C-like\n> quoting), newline and horizontal tab (makes \"patch\" confused, as\n> it would make it ambiguous where the pathname ends), so I am not\n> opposed to a patch that introduces a new mode, probably on by\n> default _unless_ we are generating --format=email, that does not\n> quote high byte values.  That would solve \"My UTF-8 filenames\n> are unreadable on my terminal\" problem.\n\nIMHO it should be the default even for email format. Most projects that use\nnon-ascii filenames probably have all members using same locale. And for\nsuch group, it will just work. Also usually the file names, content and\ncommit messages will usually be in the same (though project-specific)\nencoding, so if charset in content-type is set to that, people with different\nlocale able to represent the same characters will still see the names\ncorrectly. For other people, the MUA will probably print some escape anyway\n(it will not screw up the terminal -- it usually knows what it can safely\npass to it).\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"45664","messageId":"200706241314.46238.robin.rosenberg.lists@dewire.com","threadId":"8588","inReplyTo":"20070624065008.GA6979@efreet.light.src","subject":"Re: Stupid quoting...","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-06-24T11:14:45Z","receivedAt":"2007-06-24T11:14:45Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"söndag 24 juni 2007 skrev Jan Hudec:\n> IMHO it should be the default even for email format. Most projects that use\n> non-ascii filenames probably have all members using same locale. And for\n> such group, it will just work. Also usually the file names, content and\n> commit messages will usually be in the same (though project-specific)\n> encoding, so if charset in content-type is set to that, people with \ndifferent\n> locale able to represent the same characters will still see the names\n> correctly. For other people, the MUA will probably print some escape anyway\n> (it will not screw up the terminal -- it usually knows what it can safely\n> pass to it).\n\nI can't talk about \"most\" here, only local conditions, i.e. northern Europe \nwhere both the legacy ISO encodings are very common with a steady increase in \nUTF-8 usage, in the Linux community. People using OSS in windows almost \nexclusively get the windows-1252 (for most practical purposes the same as \nISO-8859-1).\n\nEven a *very* small set of random people you will wind up with people having \ndifferent locales.\n\n-- robin\n"},{"id":"45666","messageId":"7vzm2ptw04.fsf@assigned-by-dhcp.cox.net","threadId":"8588","inReplyTo":"200706241314.46238.robin.rosenberg.lists@dewire.com","subject":"Re: Stupid quoting...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-06-24T11:47:07Z","receivedAt":"2007-06-24T11:47:07Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Robin Rosenberg <robin.rosenberg.lists@dewire.com> writes:\n\n> I can't talk about \"most\" here, only local conditions, i.e. northern Europe \n> where both the legacy ISO encodings are very common with a steady increase in \n> UTF-8 usage, in the Linux community. People using OSS in windows almost \n> exclusively get the windows-1252 (for most practical purposes the same as \n> ISO-8859-1).\n>\n> Even a *very* small set of random people you will wind up with people having \n> different locales.\n\nMore problematic is the case where pathnames and contents are in\ndifferent encodings, even for the same language.\n\nFor example, my mbox files that store messages I receive from\npeople in Japan have contents in ISO-2022 as that is the\nlongstanding standard encoding used for e-mail over there, but\nthe pathname encoding used by the system I have that mbox file\non is EUC-JP.\n\nIf I were to create a patch between two versions of such a file,\nthe diff header would show the pathname encoded in one, and the\nchanged contents would ben shown in another.  As long as you\ntreat \"git diff\" output as binary blob, that would work just\nfine, but when you have to transmit such a diff in e-mail as an\nin-line patch, you would have troubles.\n"},{"id":"45667","messageId":"85myypef7p.fsf@lola.goethe.zz","threadId":"8588","inReplyTo":"7vzm2ptw04.fsf@assigned-by-dhcp.cox.net","subject":"Re: Stupid quoting...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-06-24T11:58:50Z","receivedAt":"2007-06-24T11:58:50Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> If I were to create a patch between two versions of such a file, the\n> diff header would show the pathname encoded in one, and the changed\n> contents would ben shown in another.  As long as you treat \"git\n> diff\" output as binary blob, that would work just fine, but when you\n> have to transmit such a diff in e-mail as an in-line patch, you\n> would have troubles.\n\nASCII-armoring of what amounts to binary files is the task of the mail\nsoftware.  Also working with encodings.  Escaping characters in the\ndiff headers but not in the file contents is not going to achieve\nanything useful, anyway.\n\nWith the proper mailing software, you can get your diff across the\nline in a manner where the other side can make use of it.  This is not\nthe case for unarmored mail with ^ escapes in them, since the\nreceiving side can't distinguish them from \"real\" ^ characters.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"45668","messageId":"7vsl8htuin.fsf@assigned-by-dhcp.cox.net","threadId":"8588","inReplyTo":"85myypef7p.fsf@lola.goethe.zz","subject":"Re: Stupid quoting...","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-06-24T12:19:12Z","receivedAt":"2007-06-24T12:19:12Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"David Kastrup <dak@gnu.org> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> If I were to create a patch between two versions of such a file, the\n>> diff header would show the pathname encoded in one, and the changed\n>> contents would ben shown in another.  As long as you treat \"git\n>> diff\" output as binary blob, that would work just fine, but when you\n>> have to transmit such a diff in e-mail as an in-line patch, you\n>> would have troubles.\n>\n> ASCII-armoring of what amounts to binary files is the task of the mail\n> software.  Also working with encodings.  Escaping characters in the\n> diff headers but not in the file contents is not going to achieve\n> anything useful, anyway.\n\nYou misunderstood me.  The issue is not about transmitting\nwithout corruption.  Armoring would make it impossible to\nCOMMENTING on the patch INLINE.\n\nAnd that is where the pathname quoting git diff does originally\ncomes from.\n"},{"id":"45669","messageId":"20070624124125.GA18803@coredump.intra.peff.net","threadId":"8588","inReplyTo":"7vsl8htuin.fsf@assigned-by-dhcp.cox.net","subject":"Re: Stupid quoting...","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-06-24T12:41:25Z","receivedAt":"2007-06-24T12:41:25Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Jun 24, 2007 at 05:19:12AM -0700, Junio C Hamano wrote:\n\n> > ASCII-armoring of what amounts to binary files is the task of the mail\n> > software.  Also working with encodings.  Escaping characters in the\n> > diff headers but not in the file contents is not going to achieve\n> > anything useful, anyway.\n> \n> You misunderstood me.  The issue is not about transmitting\n> without corruption.  Armoring would make it impossible to\n> COMMENTING on the patch INLINE.\n> \n> And that is where the pathname quoting git diff does originally\n> comes from.\n\nThen how about quoted-printable?\n\nThe point is that you're _already_ screwed by the fact that there can be\nup to three different encodings in a patch (commit message, pathnames,\nand file contents) but we only know one of them (the commit message).\nWith the other two, trying to convert encodings is pointless, since we\ndon't know the starting point. So we can either output them as-is as\nbinary, or use some sort of quoting mechanism.\n\nThe quoting that happens now is:\n  - sometimes unnecessary, and hurts people who are _not_ sending the\n    diff through the mail\n  - not recognized by any widely-used un-quoter. I can't comment on your\n    diff very well if it changes the file \"\\a/f\\303\\263\\303\\266\", and\n    there's no viewer that will let me read that in a sane way.\n    I think David's point is that by doing the quoting at the MIME\n    level (using 8bit, or 7bit with QP), the recipient's MUA can at\n    least show the binary characters.  Sure, that will totally break if\n    you are using a bad mismatch of encodings, but there's nothing we\n    can do to fix that, not knowing what the encodings are. At least it\n    _will_ work in the case that your encodings are the same.\n\nThe only argument I see _for_ the current quoting is for parsing by\nnon-mail programs (like patch or git-apply); in that case, it would seem\nonly necessary only to quote tab, newline, backslash, and double quote.\nBut at least those retain their human-readability.\n\n-Peff\n"},{"id":"45673","messageId":"20070624162559.GC6979@efreet.light.src","threadId":"8588","inReplyTo":"200706241314.46238.robin.rosenberg.lists@dewire.com","subject":"Re: Stupid quoting...","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2007-06-24T16:25:59Z","receivedAt":"2007-06-24T16:25:59Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Sun, Jun 24, 2007 at 13:14:45 +0200, Robin Rosenberg wrote:\n> söndag 24 juni 2007 skrev Jan Hudec:\n> > IMHO it should be the default even for email format. Most projects that use\n> > non-ascii filenames probably have all members using same locale. And for\n> > such group, it will just work. Also usually the file names, content and\n> > commit messages will usually be in the same (though project-specific)\n> > encoding, so if charset in content-type is set to that, people with \n> different\n> > locale able to represent the same characters will still see the names\n> > correctly. For other people, the MUA will probably print some escape anyway\n> > (it will not screw up the terminal -- it usually knows what it can safely\n> > pass to it).\n> \n> I can't talk about \"most\" here, only local conditions, i.e. northern Europe \n> where both the legacy ISO encodings are very common with a steady increase in \n> UTF-8 usage, in the Linux community. People using OSS in windows almost \n> exclusively get the windows-1252 (for most practical purposes the same as \n> ISO-8859-1).\n> \n> Even a *very* small set of random people you will wind up with people having \n> different locales.\n\nA small set of *random* people will likely have different locales. But\na project that would use non-ascii filenames would probably use some\nparticular language and thus be run by people that all speak that language --\nwhich means they are not random at all and probably will use the same locale.\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"45686","messageId":"200706242139.44708.robin.rosenberg.lists@dewire.com","threadId":"8588","inReplyTo":"20070624162559.GC6979@efreet.light.src","subject":"Re: Stupid quoting...","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-06-24T19:39:44Z","receivedAt":"2007-06-24T19:39:44Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"söndag 24 juni 2007 skrev Jan Hudec:\n> On Sun, Jun 24, 2007 at 13:14:45 +0200, Robin Rosenberg wrote:\n> > I can't talk about \"most\" here, only local conditions, i.e. northern \nEurope \n> > where both the legacy ISO encodings are very common with a steady increase \nin \n> > UTF-8 usage, in the Linux community. People using OSS in windows almost \n> > exclusively get the windows-1252 (for most practical purposes the same as \n> > ISO-8859-1).\n> > \n> > Even a *very* small set of random people you will wind up with people \nhaving \n> > different locales.\n> \n> A small set of *random* people will likely have different locales. But\n> a project that would use non-ascii filenames would probably use some\n> particular language and thus be run by people that all speak that \nlanguage --\n> which means they are not random at all and probably will use the same \nlocale.\n\nI was still in referernce to those \"local conditions\" at that point. It was \nnot meant as a universal statement. Substitutute that for \"A small bunch of \nswedish speaking people from Stockholm\".\n\n-- robin\n"},{"id":"45688","messageId":"85ejk1cexn.fsf@lola.goethe.zz","threadId":"8588","inReplyTo":"200706242139.44708.robin.rosenberg.lists@dewire.com","subject":"Re: Stupid quoting...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-06-24T19:47:48Z","receivedAt":"2007-06-24T19:47:48Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Robin Rosenberg <robin.rosenberg.lists@dewire.com> writes:\n\n> I was still in referernce to those \"local conditions\" at that\n> point. It was not meant as a universal statement. Substitutute that\n> for \"A small bunch of swedish speaking people from Stockholm\".\n\nA wrong encoding is a wrong encoding.  Escaping the characters in\ntransition will not magically make the encodings adapt.  Escaping\ncharacters buys us exactly zilch _unless_ the _channel_ is not 8-bit\nclean.  In which case we should use a normal\nmail-armoring/attachment/inline data wrapper.\n\nIn fact, when using editors with some heuristics regarding character\nsets (like Emacs), leaving 8-bit characters intact gives the editor a\nchance to guess the correct character set even if it is not the\ndefault on the receiving end.\n\nEscaping the characters, in contrast, just hides 8-bit usage away in\ntransition.  An escaped character in the wrong encoding will get\nreconstituted into the wrong encoding.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"45689","messageId":"200706242217.48284.robin.rosenberg.lists@dewire.com","threadId":"8588","inReplyTo":"85ejk1cexn.fsf@lola.goethe.zz","subject":"Re: Stupid quoting...","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-06-24T20:17:47Z","receivedAt":"2007-06-24T20:17:47Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"söndag 24 juni 2007 skrev David Kastrup:\n> Robin Rosenberg <robin.rosenberg.lists@dewire.com> writes:\n> \n> > I was still in referernce to those \"local conditions\" at that\n> > point. It was not meant as a universal statement. Substitutute that\n> > for \"A small bunch of swedish speaking people from Stockholm\".\n> \n> A wrong encoding is a wrong encoding.  Escaping the characters in\n> transition will not magically make the encodings adapt.  Escaping\n> characters buys us exactly zilch _unless_ the _channel_ is not 8-bit\n> clean.  In which case we should use a normal\n> mail-armoring/attachment/inline data wrapper.\n> \n> In fact, when using editors with some heuristics regarding character\n> sets (like Emacs), leaving 8-bit characters intact gives the editor a\n> chance to guess the correct character set even if it is not the\n> default on the receiving end.\n> \n> Escaping the characters, in contrast, just hides 8-bit usage away in\n> transition.  An escaped character in the wrong encoding will get\n> reconstituted into the wrong encoding.\n> \n\nPlease don't quote me when the content is not in reference to me or what\nI've written.\n\n-- robin\n"},{"id":"45690","messageId":"854pkxcd6h.fsf@lola.goethe.zz","threadId":"8588","inReplyTo":"200706242217.48284.robin.rosenberg.lists@dewire.com","subject":"Re: Stupid quoting...","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-06-24T20:25:42Z","receivedAt":"2007-06-24T20:25:42Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Robin Rosenberg <robin.rosenberg.lists@dewire.com> writes:\n\n> söndag 24 juni 2007 skrev David Kastrup:\n>> Robin Rosenberg <robin.rosenberg.lists@dewire.com> writes:\n>> \n>> > I was still in referernce to those \"local conditions\" at that\n>> > point. It was not meant as a universal statement. Substitutute that\n>> > for \"A small bunch of swedish speaking people from Stockholm\".\n>> \n>> A wrong encoding is a wrong encoding.  Escaping the characters in\n>> transition will not magically make the encodings adapt.\n>\n> Please don't quote me when the content is not in reference to me or what\n> I've written.\n\nNo idea how that has happened: I actually intended to refer to\nsomething different.  Sorry.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"}]}