{"thread":{"id":"2042","subject":"[RFC] embedded TAB and LF in pathnames","startedAt":"2005-10-07T19:35:19Z","lastAt":"2005-10-14T17:18:15Z","messageCount":33,"participants":["Junio C Hamano","Alex Riesen","Robert Fitzsimons","Paul Eggert","Linus Torvalds","Daniel Barkalow","H. Peter Anvin","Johannes Schindelin"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"9806","messageId":"7vu0ftyvbc.fsf@assigned-by-dhcp.cox.net","threadId":"2042","inReplyTo":null,"subject":"[RFC] embedded TAB and LF in pathnames","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-07T19:35:19Z","receivedAt":"2005-10-07T19:35:19Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"While I was reviewing git-status fix by Kai Ruemmler, it struck\nme that our barebone Porcelain-ish layer got a bit sloppier over\ntime.  The core layer does not care about any metacharacters in\nthe pathname, and it has provisions, primarily in the form of\n'-z' flag, for carefully written Porcelain layers to handle\npathnames with embedded metacharacters correctly.\n\nOne exception, however, is the interaction between the git-diff\nfamily output and git-apply.  We needed to be compatible with\nother people's diff, which meant that we should not have to\nworry too much about pathnames with embedded TABs and LFs\nbecause GNU diff would not produce usable diff for such things\nanyway.  But 'git-diff --names' barfing if a pathname contained\nthese characters when run without '-z' flag was too much.  This\nstill breaks 'git-status'.\n\nSo I am considering the following changes:\n\n - 'raw' output format without '-z', upon finding a TAB or LF,\n   would not die, but just issue a warning.  However, the paths\n   are \"munged\" in a way described later.\n\n - '--name-only' and '--name-status' format issue the same\n   warning when finding these characters and run without '-z'.\n   And the paths are \"munged\" as well.\n\n - 'patch' output format also issues a warning.  The paths are\n   \"munged\" but in a slightly different manner from the above.\n\n - 'git-apply' is taught about the path munging in the diff\n   input for git diffs (i.e. 'diff --git') and do sensible\n   things.\n\nOne possible way for path munging goes like this.  We could take\nadvantage of the fact that we do not ever output '//' ourselves,\nand '//' never appears in valid diffs by other people's tools,\nunless done deliberately by hand (\"diff -u a//foo. b//foo.c\"\nfrom the command line).  So we could use '//' as if it is a\nbackslash.  Examples.\n\n  \"foo/bar.c\" --> \"foo/bar.c\"\t(no funny letters - as before)\n  \"foo\\nbar\"  --> \"foo//0Abar\" (double slash followed by 2 hex)\n  \"foo\\tbar\"  --> \"foo//09bar\" (double slash followed by 2 hex)\n\nSo a diff output to rename \"foo/bar.c\" to \"foo\\nbar.c\" would\nbecome:\n\n  diff --git a/foo/bar.c b/foo//0Abar.c\n  similarity index 100%\n  rename from foo\n  rename to foo//0Abar.c\n\nThe byte-values subject to this munging is LF for patch output\n(because git-apply seems to grok TABs in pathnames just fine),\nand TAB and LF for 'raw', '--name-only', '--name-status' without\n'-z'.\n\nI have not made up my mind on the exact choice of the quoting\nconvention.  We could say '///' instead of '//', for example, or\neven '//{LF}//' instead of '//0A' proposed above.  One thing I\nam trying to avoid is \"foo\\nbar\", which I suspect would be\nunfriendly to the Cygwin folks.\n"},{"id":"9820","messageId":"20051007232909.GB8893@steel.home","threadId":"2042","inReplyTo":"7vu0ftyvbc.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] embedded TAB and LF in pathnames","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2005-10-07T23:29:09Z","receivedAt":"2005-10-07T23:29:09Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"Junio C Hamano, Fri, Oct 07, 2005 21:35:19 +0200:\n> I have not made up my mind on the exact choice of the quoting\n> convention.  We could say '///' instead of '//', for example, or\n> even '//{LF}//' instead of '//0A' proposed above.  One thing I\n> am trying to avoid is \"foo\\nbar\", which I suspect would be\n> unfriendly to the Cygwin folks.\n\nBeing unhappy one of them, I think I'd better manage (even if by\npostprocessing the output).\n\nPlease, don't make the common case ugly just because of that platform\n(insanely broken anyway).\n"},{"id":"9822","messageId":"7vpsqgyjrj.fsf@assigned-by-dhcp.cox.net","threadId":"2042","inReplyTo":"20051007232909.GB8893@steel.home","subject":"Re: [RFC] embedded TAB and LF in pathnames","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-07T23:44:48Z","receivedAt":"2005-10-07T23:44:48Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Alex Riesen <raa.lkml@gmail.com> writes:\n\n> Junio C Hamano, Fri, Oct 07, 2005 21:35:19 +0200:\n>> I have not made up my mind on the exact choice of the quoting\n>> convention.  We could say '///' instead of '//', for example, or\n>> even '//{LF}//' instead of '//0A' proposed above.  One thing I\n>> am trying to avoid is \"foo\\nbar\", which I suspect would be\n>> unfriendly to the Cygwin folks.\n>\n> Being unhappy one of them, I think I'd better manage (even if by\n> postprocessing the output).\n>\n> Please, don't make the common case ugly just because of that platform\n> (insanely broken anyway).\n\nYou really have to realize that having LF and TAB in filenames\nare *NOT* the common case, no matter which platform you are\ntalking about.\n"},{"id":"9828","messageId":"20051008064555.GA3831@steel.home","threadId":"2042","inReplyTo":"7vpsqgyjrj.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] embedded TAB and LF in pathnames","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2005-10-08T06:45:55Z","receivedAt":"2005-10-08T06:45:55Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"Junio C Hamano, Sat, Oct 08, 2005 01:44:48 +0200:\n> > Junio C Hamano, Fri, Oct 07, 2005 21:35:19 +0200:\n> >> I have not made up my mind on the exact choice of the quoting\n> >> convention.  We could say '///' instead of '//', for example, or\n> >> even '//{LF}//' instead of '//0A' proposed above.  One thing I\n> >> am trying to avoid is \"foo\\nbar\", which I suspect would be\n> >> unfriendly to the Cygwin folks.\n> >\n> > Being unhappy one of them, I think I'd better manage (even if by\n> > postprocessing the output).\n> >\n> > Please, don't make the common case ugly just because of that platform\n> > (insanely broken anyway).\n> \n> You really have to realize that having LF and TAB in filenames\n> are *NOT* the common case, no matter which platform you are\n> talking about.\n> \n\nYes, but \"//\" in a path is quite common. Even \"///\" is not uncommon.\n\nHow about copy ls' approach were possible?\n\n   -b, --escape, --quoting-style=escape\n          Quote  nongraphic  characters in file names using alphabetic and\n          octal backslash sequences like those used in C. This  option  is\n          the  same as -Q except that filenames are not surrounded by dou-\n          ble-quotes.\n"},{"id":"9829","messageId":"7vachks7aq.fsf@assigned-by-dhcp.cox.net","threadId":"2042","inReplyTo":"20051008064555.GA3831@steel.home","subject":"Re: [RFC] embedded TAB and LF in pathnames","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-08T09:10:37Z","receivedAt":"2005-10-08T09:10:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Alex Riesen <raa.lkml@gmail.com> writes:\n\n          Quote  nongraphic  characters in file names using alphabetic and\n          octal backslash sequences like those used in C. This  option  is\n          the  same as -Q except that filenames are not surrounded by dou-\n          ble-quotes.\n\nIf you have a file whose name is 'foo' + LF + 'bar', and if you\nuse backslash convention, your diff would start like this:\n\n    diff --git a/foo\\nbar b/foo\\nbar\n    @@ 1,2 3,4 @@\n     context\n    -deleted\n    ...\n\nwhich looks quite natural.\n\nI would, however, prefer this kind of funny pathnames to *stand*\n*out* more than usual, to make it really obvious that there is\nsomething really funky going on.  In that sense, the above is a\nbit too innocuous-looking to my taste.\n\nBut this \"embedded LF and TAB\" is a corner case.  I would not be\nusing such paths that would trigger the quoting myself anyway,\nand I do not particularly care as long as the tools do the right\nthing -- any quoting rule would do, as long as the generating\nside (git-diff) is consistent with accepting side (git-apply),\nand as long as there is no new ambiguity introduced.\n\nThe backslash proposal is introducing a small ambiguity.  You\ncannot tell if the file had an embedded LF between 'foo' and\n'bar' (and generated with your git-diff) or had an embedded\nbackslash between 'foo' and 'nbar' (and generated with existing\ngit-diff).  Since we never had a version of git-diff that\noutputs double-slashes '//' in paths, there is no ambiguity if\nwe use it as a quoting mechanism.\n\nJust as a concrete demonstration, here is how the git-status\noutput and git-diff output would look like for a file 'pqr' in a\ndirectory whose name is 'def' + LF + 'ghi' that uses the version\nof git-diff from the proposed updates branch:\n\n        # Changed but not updated:\n        #   (use git-update-index to mark for commit)\n        #\n        #\tmodified: def//{LF}//ghi/pqr\n\n        diff --git a/def//{LF}//ghi/pqr b/def//{LF}//ghi/pqr\n        index 9ee055c..47dbc3f 100644\n        --- a/def//{LF}//ghi/pqr\n        +++ b/def//{LF}//ghi/pqr\n        @@ -1 +1,2 @@\n         Fri Oct  7 23:19:04 PDT 2005\n        +foo\n\nI am not married to this quoting syntax -- I think it *is* ugly,\nbut as I said before, I'd prefer to have something ugly here.\n\nI would easily be persuaded otherwise, though.  A working patch\nwould probably be the most effective way of persuasion, but a\nmock output without the code to produce and/or parse it would\nalso be fine as a starting point for discussion.\n"},{"id":"9830","messageId":"20051008133032.GA32079@localhost","threadId":"2042","inReplyTo":"7vachks7aq.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Robert Fitzsimons","fromEmail":"robfitz@273k.net","sentAt":"2005-10-08T13:30:32Z","receivedAt":"2005-10-08T13:30:32Z","isPatch":true,"sender":{"key":"robfitz@273k.net","avatar":null},"body":"Instead of using //{LF}// and //{TAG}// to quote embedded tab and\nlinefeed characters in pathnames use URI quoting.\n\n'\\t' becomes %09\n'\\n' becomes %10\n'%' becomes %25\n\nSigned-off-by: Robert Fitzsimons <robfitz@273k.net>\n\n---\n\n> I am not married to this quoting syntax -- I think it *is* ugly,\n> but as I said before, I'd prefer to have something ugly here.\n> \n> I would easily be persuaded otherwise, though.  A working patch\n> would probably be the most effective way of persuasion, but a\n> mock output without the code to produce and/or parse it would\n> also be fine as a starting point for discussion.\n\nUsing URI encoding might be an option it's not a ugly and more peopel\nshould under stand what it means.  Heres a posible patch against pu.\n\nRobert\n\n\n apply.c       |   19 ++++++++++++-------\n diff.c        |   26 +++++++++++++++++---------\n git-status.sh |   10 ++++++----\n 3 files changed, 35 insertions(+), 20 deletions(-)\n\napplies-to: a9332b0c2bd80a182f946d22d4ec7511c32c55f4\n8029a957cab1a912562696fdce8beea5fc2c11c4\ndiff --git a/apply.c b/apply.c\n--- a/apply.c\n+++ b/apply.c\n@@ -75,21 +75,26 @@ static char *unmunge_name(char *name)\n \n \tif (!name)\n \t\treturn name;\n-\tcp = strstr(name, \"//\");\n+\tcp = strstr(name, \"%\");\n \tif (!cp)\n \t\treturn name;\n \tret_name = strdup(name);\n \tfor (cp = dp = ret_name; (ch = *cp); cp++) {\n-\t\tif (ch == '/' && cp[1] == '/' && cp[2] == '{') {\n-\t\t\t/* //{TAB}// or //{LF}// */\n-\t\t\tif (!strncmp(cp + 3, \"TAB}//\", 6)) {\n+\t\tif (ch == '%') {\n+\t\t\t/* %09 or %10 or %25 */\n+\t\t\tif (!strncmp(cp + 1, \"09\", 2)) {\n \t\t\t\t*dp++ = '\\t';\n-\t\t\t\tcp += 8;\n+\t\t\t\tcp += 2;\n \t\t\t\tcontinue;\n \t\t\t}\n-\t\t\telse if (!strncmp(cp + 3, \"LF}//\", 5)) {\n+\t\t\telse if (!strncmp(cp + 1, \"10\", 2)) {\n \t\t\t\t*dp++ = '\\n';\n-\t\t\t\tcp += 7;\n+\t\t\t\tcp += 2;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\telse if (!strncmp(cp + 1, \"25\", 2)) {\n+\t\t\t\t*dp++ = '%';\n+\t\t\t\tcp += 2;\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\terror(\"malformed munged name '%s' (looking at %s)\",\ndiff --git a/diff.c b/diff.c\n--- a/diff.c\n+++ b/diff.c\n@@ -13,7 +13,7 @@ static const char *path_munge(const char\n {\n \tconst char *cp;\n \tchar *retpath, *dp;\n-\tint ch, munge_inter_name = 0, munge_line_term = 0;\n+\tint ch, munge_inter_name = 0, munge_line_term = 0, munge_quote = 0;\n \n \tif (!path)\n \t\treturn path;\n@@ -23,23 +23,31 @@ static const char *path_munge(const char\n \t\t\tmunge_inter_name++;\n \t\tif (line_term && ch == '\\n')\n \t\t\tmunge_line_term++;\n+\t\tif (ch == '%')\n+\t\t\tmunge_quote++;\n \t}\n-\tif (!(munge_inter_name + munge_line_term))\n+\tif (!(munge_inter_name + munge_line_term + munge_quote))\n \t\treturn path;\n \n-\t/* need //{TAB}// and //{LF}// */\n+\t/* need %09 and %10 and %25 */\n \tretpath = xmalloc(cp - path +\n-\t\t\t  munge_inter_name * 8 +\n-\t\t\t  munge_line_term * 7 + 1);\n+\t\t\t  munge_inter_name * 3 +\n+\t\t\t  munge_line_term * 3 +\n+\t\t\t  munge_quote * 3 + 1);\n \tfor (cp = path, dp = retpath; (ch = *cp); cp++, dp++) {\n \t\tif (inter_name && ch == '\\t') {\n-\t\t\tmemcpy(dp, \"//{TAB}//\", 9);\n-\t\t\tdp += 8;\n+\t\t\tmemcpy(dp, \"%09\", 3);\n+\t\t\tdp += 2;\n \t\t\tcontinue;\n \t\t}\n \t\tif (line_term && ch == '\\n') {\n-\t\t\tmemcpy(dp, \"//{LF}//\", 8);\n-\t\t\tdp += 7;\n+\t\t\tmemcpy(dp, \"%10\", 3);\n+\t\t\tdp += 2;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (ch == '%') {\n+\t\t\tmemcpy(dp, \"%25\", 3);\n+\t\t\tdp += 2;\n \t\t\tcontinue;\n \t\t}\n \t\t*dp = ch;\ndiff --git a/git-status.sh b/git-status.sh\n--- a/git-status.sh\n+++ b/git-status.sh\n@@ -54,8 +54,9 @@ else\n \tperl -e '$/ = \"\\0\";\n \t\twhile (<>) {\n \t\t\tchomp;\n-\t\t\ts|\\t|//{TAB}//|g;\n-\t\t\ts|\\n|//{LF}//|g;\n+\t\t\ts|%([^021][^059])|%25\\1|g;\n+\t\t\ts|\\t|%09|g;\n+\t\t\ts|\\n|%10|g;\n \t\t\ts/ /\\\\ /g;\n \t\t\ts/^/A /;\n \t\t\tprint \"$_\\n\";\n@@ -84,8 +85,9 @@ perl -e '$/ = \"\\0\";\n \tmy $shown = 0;\n \twhile (<>) {\n \t\tchomp;\n-\t\ts|\\t|//{TAB}//|g;\n-\t\ts|\\n|//{LF}//|g;\n+\t\ts|%([^01][^09])|%25\\1|g;\n+\t\ts|\\t|%09|g;\n+\t\ts|\\n|%10|g;\n \t\ts/^/#\t/;\n \t\tif (!$shown) {\n \t\t\tprint \"#\\n# Ignored files:\\n\";\n---\n0.99.8.GIT\n"},{"id":"9840","messageId":"7v64s7svya.fsf@assigned-by-dhcp.cox.net","threadId":"2042","inReplyTo":"20051008133032.GA32079@localhost","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-08T18:30:21Z","receivedAt":"2005-10-08T18:30:21Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Robert Fitzsimons <robfitz@273k.net> writes:\n\n> Instead of using //{LF}// and //{TAG}// to quote embedded tab and\n> linefeed characters in pathnames use URI quoting.\n>\n> '\\t' becomes %09\n> '\\n' becomes %10\n> '%' becomes %25\n>\n> Signed-off-by: Robert Fitzsimons <robfitz@273k.net>\n\nThis would break existing setup where people *has* per-cent\nletter in their pathname -- which I think is worse than the\nbackslash proposal.\n"},{"id":"9846","messageId":"7vu0frpxs1.fsf@assigned-by-dhcp.cox.net","threadId":"2042","inReplyTo":"7v64s7svya.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-08T20:19:10Z","receivedAt":"2005-10-08T20:19:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> Robert Fitzsimons <robfitz@273k.net> writes:\n>\n>> '\\t' becomes %09\n>> '\\n' becomes %10\n>> '%' becomes %25\n>>\n>> Signed-off-by: Robert Fitzsimons <robfitz@273k.net>\n>\n> This would break existing setup where people *has* per-cent\n> letter in their pathname -- which I think is worse than the\n> backslash proposal.\n\nHaving said that, I think something along the lines of backslash\nor URI encoding is the cleanest way to go in the long run, with\none condition: diffs generated with git-diff should be\napplicable with 'GNU patch', especially if there is no funnies\nlike renames and the recipient does not mind losing mode\ninformation.\n\nAlthough 'GNU patch' has --quoting-style flag, it seems to be\nused only on its output side (i.e. reporting which file it is\npatching, etc.).  If we can sell changes to teach the filename\nencoding convention to its util.c::fetchname() upstream, we\ncould tell people that 'diff --git' can be applied with newer\n'GNU patch' when the patch is about a file whose name contains\n'%' character (which is not that unusual, compared to TAB and\nLF).  While we are selling those changes to 'GNU patch', we\nmight be even be able to sell the other extended 'diff --git'\nmetainformation support.\n\nThe same filename quoting rules change should probably be sold\nto 'GNU diff' as well, so that plain diff can natively quote\nfunny characters in its output without forcing us to fake it\nby using the -L flag.\n\nIf all of the above is what we aim for, I would say that is a\ngood direction to go in the longer term.  The double-slash hack\nwas just to avoid all these hassles of having to muck with other\npeople's tools.\n"},{"id":"9856","messageId":"7vk6gnezuu.fsf@assigned-by-dhcp.cox.net","threadId":"2042","inReplyTo":"20051008133032.GA32079@localhost","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-09T10:42:01Z","receivedAt":"2005-10-09T10:42:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Robert Fitzsimons <robfitz@273k.net> writes:\n\n> Instead of using //{LF}// and //{TAG}// to quote embedded tab and\n> linefeed characters in pathnames use URI quoting.\n\nI changed my mind, although I still do not particularly like the\ninnocuous looking C-style or URI-style quoting, simply because I\nfeel these funny characters should stand out loudly, but at the\nsame time I realize that is just a matter of personal taste.\nAlso the list does not seem to mind losing an extra unusual\ncharacter (either backslash or per-cent) for quoting too much.\n\nAfter futzing with this a bit more, I decided that C-style\nquoting is the cleanest way in the longer run.  I'll be talking\nwith the current maintainer of GNU patch about making it\nunderstand C-style quoting in its input, when the program\noperates under --quoting-style=c flag (or maybe some other\nflag).\n\nIn addition to git-diff output, git-ls-files output is also\nquoted for TAB and LF (and backslash) when not using '-z' in the\nversion I have in the \"pu\" branch.  I haven't converted\ngit-ls-tree yet, but that should also be done before these\nchanges can graduate out of \"pu\" branch.\n\nExisting Porcelains or people's scripts should not break (any\nmore than they currently are broken ;-).  Any self-respecting\nPorcelain should be either parsing '-z' output (in which case\nthere is no change), or parsing non '-z' output after declaring\nthat it does not support filenames with embedded TAB or LF (in\nwhich case there is no new breakage, except that they have one\nmore character that their users cannot have in the filename --\nbackslash).\n\nHere is an example output from my random repository, that has\nfiles with TAB, LF and backslash in their names (Note that the\nfile \"pc\" + one backslash + \"h.c\" is shown with two backslashes).\n\n\t: siamese; git status\n        # Updated but not checked in:\n        #   (will commit)\n        #\n        #\tnew file: ab\\n\\tc/mno\n        #\tmodified: abc/mno\n        #\trenamed: def\\nghi/pqr -> dee/pqr\n        #\tnew file: dee/www\n        #\tmodified: j  k l\n        #\n        #\n        # Changed but not updated:\n        #   (use git-update-index to mark for commit)\n        #\n        #\tdeleted:  abc/mno\n        #\n        #\n        # Ignored files:\n        #   (use \"git add\" to add to commit)\n        #\n        #\tdiff-sample\n        #\tpc\\\\h.c\n        #\tpch.c.orig\n        #\tquote\\targ.c\n        #\tquotearg.c.orig\n\n\t: siamese; git diff HEAD\n        diff --git a/abc/mno b/ab\\n\\tc/mno\n        similarity index 72%\n        rename from abc/mno\n        rename to ab\\n\\tc/mno\n        index 0ac2a8c..3deac99 100644\n        --- a/abc/mno\n        +++ b/ab\\n\\tc/mno\n        @@ -1 +1,3 @@\n         Fri Oct  7 23:18:45 PDT 2005\n        +foo\n        +foo\n        ...\n\n        : siamese; git diff HEAD | git apply --index-info\n        100644 0ac2a8c8cad088c3e843689dbd833aeabf6b1870\tabc/mno\n        100644 9ee055c103e84ffdd9ec15457481c92699d12fc8\tdef\\nghi/pqr\n\t...\n\nAnyway, I'll keep this in the \"pu\" branch a bit longer to let\nthe discussion simmer.\n"},{"id":"9949","messageId":"87mzlgh8xa.fsf@penguin.cs.ucla.edu","threadId":"2042","inReplyTo":"7vu0frpxs1.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2005-10-11T06:20:01Z","receivedAt":"2005-10-11T06:20:01Z","isPatch":true,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> Although 'GNU patch' has --quoting-style flag, it seems to be\n> used only on its output side\n\nYes, that's right.\n\nThe convention I had been thinking of adding is to have GNU diff\nuse shell-quoting style, e.g.,\n\n'three\no'\\''clock'\n\nto represent a file name with a newline and an apostrophe in it.\nThis sort of file name can be cut and pasted into the shell.\nThe quoting could be used with any file name containing a\ntroublesome character.\n\nPerhaps another quoting style would be better.\n\nAn issue I hadn't really had time to think about is the character\nencoding of file names.  E.g., suppose one file system uses UTF-8\nencoding for Japanese file names, and another file system uses EUC-JP.\nI suppose it would be nice to handle this problem too.  Perhaps GNU\n'diff' could standardize on using UTF-8 in its file names, regardless\nof what the underlying file system uses.  Another option is to pass\nthe bytes of the file name through, no matter what.  This might\nrequire a runtime flag to diff, or to patch, or both.\n"},{"id":"9952","messageId":"7vy850v6zu.fsf@assigned-by-dhcp.cox.net","threadId":"2042","inReplyTo":"87mzlgh8xa.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-11T07:37:57Z","receivedAt":"2005-10-11T07:37:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Paul Eggert <eggert@CS.UCLA.EDU> writes:\n\n> The convention I had been thinking of adding is to have GNU diff\n> use shell-quoting style, e.g.,\n>\n> 'three\n> o'\\''clock'\n>\n> to represent a file name with a newline and an apostrophe in it.\n> This sort of file name can be cut and pasted into the shell.\n> The quoting could be used with any file name containing a\n> troublesome character.\n>\n> Perhaps another quoting style would be better.\n\nA patch header (both \"diff --git\" line and ---/+++ lines) I've\nbeen considering, and have in the proposed updates branch, looks\nsomething like this:\n\n    diff --git a/def\\nghi/pqr b/dee/pqr\n    similarity index 72%\n    rename from def\\nghi/pqr\n    rename to dee/pqr\n    index 9ee055c..243fbbc 100644\n    --- a/def\\nghi/pqr\n    +++ b/dee/pqr\n    @@ -1 +1,3 @@\n     Fri Oct  7 23:19:04 PDT 2005\n    +foo\n    +foo\n\nIf we can keep things on one line, that would help parsing the\nstuff very simple, but more importantly, it is easier to see\nwhat's happening.  The pattern is the same whether you have\nfunny pathnames or not, and that helps the human consumer.\n\nAdjusting the \"git diff\" output to the style the GNU diff with\nyour shell quoting style would produce something like this:\n\n    diff --git 'a/def\n    ghi/pqr' b/dee/pqr\n    similarity index 72%\n    rename from 'def\n    ghi/pqr'\n    rename to dee/pqr\n    index 9ee055c..243fbbc 100644\n    --- 'a/def\n    ghi/pqr'\n    +++ b/dee/pqr\n    @@ -1 +1,3 @@\n     Fri Oct  7 23:19:04 PDT 2005\n    +foo\n    +foo\n\nWhich, while it is possible to make tools parse them, is very\ndistracting for humans to read and review.  Yes, LF is quoted,\nbut it still breaks the line, disrupting the pattern we are used\nto see.  If you are talking about a funny file, whose name is\n\"a\\ndiff --git a/b/c\", your diff would look like this:\n\n    diff --git 'a/\n    diff --git a/b/c' 'b/\n    diff --git a/b/c'\n    index 9ee055c..243fbbc 100644\n    --- 'a/\n    diff --git a/b/c'\n    +++ 'b/\n    diff --git a/b/c'\n    @@ -1 +1,3 @@\n     Fri Oct  7 23:19:04 PDT 2005\n    +foo\n    +foo\n\nWe are used to tell the \"less\" command to do \"/^diff --git .*\"\nwhile reviewing patches.  The shell quoting, while I admit I\nlearned its beauty from you, is a disaster for human consumption.\n\nFor diff output quoting purposes, LF is the only thing that\nmatters, as you mentioned in another message to me.  Our parsing\nside (\"GNU patch\" counterpart) checks two pathnames on \"diff\n--git\" line and makes sure what follows a/ and b/ are consistent\n(that is, they should be identical, or each are the same as\n\"rename from\" and \"rename to\"), so there is no ambiguity.  But\nagain for human consumption purposes, we cannot easily tell SP\nand TAB apart by just reading, and a TAB is so unusual character\nto have in pathname (as opposed to SP which is not that\nuncommon), we may be better off making them visible.\n\nQuoting TAB incidentally has an added benefit, which you as GNU\ndiff/patch person would probably not care too much about.  Our\nother tools sometimes need to show two paths in one record, and\nTAB is used as the field separator between two paths (LF is the\nrecord separator).  The tools do have '-z' mode to let us use\nanything but NUL in the pathname, and carefully written scripts\ntend to run them with '-z' flag and use Perl or Python to parse\npaths out, but it would be nicer if we did not always have to.\n\nFor example, the 'git commit' command prepares the log editor\nwith the status information about changes being committed, and\nneeds to mention paths.  This is purely for human consumption,\nand showing something like:\n\n\t# Type commit message to this file.  Lines that start\n        # with '#' are ignored.\n        #\n        # Updated but not checked in:\n        #   (will commit)\n        #\n        #\tnew file: ab\\n\\tc/mno\n        #\tmodified: abc/mno\n        #\trenamed: def\\nghi/pqr -> dee/pqr\n        ...\n\nis perfectly readable for human users, and can be done without\nrunning the tool in '-z' mode, if the tool output is quoted with\n'\\n' and '\\t' convention -- the parsing and formatting side can\njust split the field with TAB and show them, without worrying\nabout an embedded LF making the rest of the pathname spilling\nover to the next line.  And once we start teaching the user we\nrepresent funny characters in their paths this way, it becomes\nnicer to be consistent in the diff output as well.\n"},{"id":"9964","messageId":"Pine.LNX.4.64.0510110802470.14597@g5.osdl.org","threadId":"2042","inReplyTo":"87mzlgh8xa.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-11T15:17:25Z","receivedAt":"2005-10-11T15:17:25Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 10 Oct 2005, Paul Eggert wrote:\n> \n> An issue I hadn't really had time to think about is the character\n> encoding of file names.\n\nPlease don't. Use filenames as if they are just binary blobs of data, \nthat's the only thing that has a high chance of success. Yes, it too can \nbreak in the presense of something _else_ doing character translation \nand/or people moving a patch from one encoding to another , buthat's \njust true of anything.\n\nEventually everybody will hopefully use UTF-8, and nothing else really \nmatters, but the thing is, if you see filenames as just blobs of data, \nthat works with UTF-8 too, so it's not \"wrong\" even in the long run. And \nuntil everybody has one single encoding, you simply won't be able to tell, \nand the likelihood that you'd screw up is pretty high.\n\nThe happy part of the \"binary blob\" approach is that users _understand_ \nit. People who actively use different encoding formats are (painfully) \naware of conversions, and they may curse you for not doing the random \nencoding format of the day, but they will be able to handle it.\n\nIn contrast, if you start doing conversions, I guarantee you that people \nwill _not_ be able to handle it when you do something strange - you've \nchanged the data.\n\nPersonally, I'd like the normal C quoting the best. Leave space as-is, and \nquote TAB/NL as \\t and \\n respectively. It's pretty universally understood \nin programming circles even outside of C, and it's not like a very \nuncommon patch format like that really needs to be well-understood outside \nof those circles.\n\nIt also has a very obvious and ASCII-safe format for other characters (ie \njust the normal octal escapes: \\377 etc..\n\nThat said, I personally don't think it's necessarily even worth it. If \nsomebody wants to use names with tabs and newlines, is he really going to \nwork with diffs? Or is it just a driver error?\n\n\t\t\tLinus\n"},{"id":"9970","messageId":"87ek6s0w34.fsf@penguin.cs.ucla.edu","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510110802470.14597@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2005-10-11T18:03:59Z","receivedAt":"2005-10-11T18:03:59Z","isPatch":true,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Personally, I'd like the normal C quoting the best.\n\nThat would be fine with me too.  How about if we use the equivalent of\n--quoting-style=\"c\" for file names that contain funny bytes, and no\nquoting for other file names?  So, for example, something like this:\n\n    diff --git \"space tab\\tnewline\\nquote\\\"backslash\\\\\" b/dee/pqr\n    similarity index 72%\n    rename from \"space tab\\tnewline\\nquote\\\"backslash\\\\\"\n    rename to dee/pqr\n    index 9ee055c..243fbbc 100644\n    --- \"space tab\\tnewline\\nquote\\\"backslash\\\\\"\n    +++ b/dee/pqr\n    @@ -1 +1,3 @@\n     Fri Oct  7 23:19:04 PDT 2005\n    +foo\n    +foo\n\nThe surrounding double-quotes are an extra indication to the human\nreader that there is something weird about the quoted file name.\n\n> Use filenames as if they are just binary blobs of data, \n> that's the only thing that has a high chance of success.\n\nThanks for thinking those things through.  I agree mostly, but there's\nstill a technical problem, in that we have to decide what a \"funny\nbyte\" is if we are using C-style quoting.  For example, the simplest\napproach is to say a byte is funny if it is space, backslash, quote,\nan ASCII control character, or is non-ASCII.  But this will cause\nperfectly-reasonable UTF-8 file names to be presented in git format\nusing unreadable strings like \"a\\293\\203\\257b\" or whatever.\n\nPerhaps it would be better to say that a byte is \"funny\" if it is\nspace, backslash, quote, an ASCII control character, or a byte that is\nnot part of a valid UTF-8 encoding.  This will let UTF-8 file names\nthrough unscathed, while still warning the reader when funny business\nis going on.  File names with other encodings (e.g., Shift-JIS) will\ncontain lots of backslashes, but that's OK: we don't mind making\nnonstandard encodings hard-to-read, so long as we preserve the bytes\ncorrectly.\n\nWe could implement in other GNU applications by having a new quoting\nstyle that supports this quoting behavior.  I can arrange for that.\n\n\n> If somebody wants to use names with tabs and newlines, is he really\n> going to work with diffs? Or is it just a driver error?\n\nThe current-supported scheme with 'diff' and 'patch' should work for\neverything but newlines.  I like the idea of getting it to work even\nwith newlines, and I am willing to sacrifice old patches with file\nnames starting with '\"' (extremely rare, if any) to get newlines to\nwork.  Among other things I worry about people submitting\npurposely-malformed patches in non-git environments.\n"},{"id":"9971","messageId":"Pine.LNX.4.64.0510111121030.14597@g5.osdl.org","threadId":"2042","inReplyTo":"87ek6s0w34.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-11T18:37:46Z","receivedAt":"2005-10-11T18:37:46Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Oct 2005, Paul Eggert wrote:\n>\n> For example, the simplest approach is to say a byte is funny if it is \n> space, backslash, quote, an ASCII control character, or is non-ASCII.  \n> But this will cause perfectly-reasonable UTF-8 file names to be \n> presented in git format using unreadable strings like \"a\\293\\203\\257b\" \n> or whatever.\n\nI think the simplest question to ask is \"what are we protecting against?\"\n\nThere's only two characters that are _really_ special diff itself: \\n and \n\\t. The former is obvious, the latter just because the regular gnu diff \nformat puts a tab between the name and the date (and if you _knew_ the \ndate was always there you could just work backwards, but since not all \ndiffs even put a date, \\t ends up being special in practice).\n\nSo what else would you want to protect against? I hope not 8-bit \ncleanness: if some stupid protocol still isn't 8-bit clean, it should be \nfixed.\n\nAnd \\0 is already impossible, at least on sane systems.\n\nSo arguably you don't need to quote anything else than \\n and \\t (and that \nobviously means you have to quote \\ itself). That means that any filename \nalways shows \"sanely\" in its own byte locale, and everything is readable, \nregardless of whether it's UTF-8 or just plain byte-encoded Latin1, or \nanything else.\n\nSo I don't think you should quote invalid UTF-8: it's invalid UTF-8 \nwhether Ã­tis quoted or not.\n\n\t\tLinus\n\nPS. There _is_ something you may want to quote, namely the standard CSI \nterminal escapes. Not because they wouldn't pass through, but because some \npeople might just \"cat\" a patch. This is debatable. Now, they are in all \nin the range 0x00-0x1f and 0x80-0x9f, and since UTF-8 encoding is supposed \nto happen before it (but you don't know how many get that right), if you \nwant to quote those characters, you need to do so _both_ for the \"raw\" \nformat and for the UTF-8 format.\n\nNow, the UTF-8 format for that high range is actually the same character, \nexcept preceded by a 0xc2 (I think), so the simplest thing is to do \nquoting _purely_ on a byte-stream level (ignore any UTF-8 stuff), and \nscrew the fact that you end up with a non-UTF-8 sequence (character 0x0080 \nis UTF-8 sequence 0xC2 0x80, and would be quoted as 0xC2 + \"\\200\", which \nis no longer valid in UTF-8).\n\nIt gets quite nasty. For any UTF-8 quoting scheme you come up with, I'll \npoint out something that it does wrong or looks horrible for a Latin1 \nfilename ;)"},{"id":"9973","messageId":"87slv7zvqj.fsf@penguin.cs.ucla.edu","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510111121030.14597@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2005-10-11T19:42:12Z","receivedAt":"2005-10-11T19:42:12Z","isPatch":true,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> the simplest question to ask is \"what are we protecting against?\"\n\nI'd like to protect against:\n\n  1.  File names that cannot be handled correctly with the current\n      formats.  Newline is the obvious problem here, along with\n      (arguably) tab and space.\n\n  2.  Common transliterations of patches.  Many programs (and mailers,\n      alas) expand tabs to spaces, append CR to lines, prepend spaces\n      to lines, break lines at spaces, etc.  'patch' already deals\n      with this to some extent, but it'd be nice if the format\n      resisted these transliterations better.\n\n  3.  Humans misreading patches.  The patch format is intended to be\n      human-readable, after all.\n\n  4.  Reencoded patches.  Programs like Emacs can and will convert\n      patches from UTF-8 to EUC-JP, for example.\n\nYou convinced me that (4) is not worth the hassle, but I'd still like\nto address (1)-(3) when it's easy.\n\n> invalid UTF-8 [is] invalid UTF-8\n\nYes, but (2) and (3) can lose information about invalid UTF-8 if we\ndon't suitably protect the encoding errors.  I daresay that many\nmailers will mishandle invalid UTF-8, for example.\n\n> There _is_ something you may want to quote, namely the standard CSI\n> terminal escapes.\n\nIf I understand you aright, we could do that by modifying my previous\nproposal to escape all bytes in the UTF-8 representation of a control\ncharacter.  In Unicode, the characters 0080 through 009F are control\ncharacters, so that should suffice to quote the terminal escapes you\nmentioned.  (Perhaps we should also escape unassigned Unicode\ncharacters too, on the theory that they might become control\ncharacters in the future.)\n\n> For any UTF-8 quoting scheme you come up with, I'll point out\n> something that it does wrong or looks horrible for a Latin1 filename\n> ;)\n\nYes, quite true.  But we don't have to come up with something that's\nperfect in all cases, just something that's good enough to handle\ncases that we expect will be common in practice, in a world where\nUTF-8 is the preferred encoding for non-ASCII characters.\n"},{"id":"9979","messageId":"Pine.LNX.4.64.0510111346220.14597@g5.osdl.org","threadId":"2042","inReplyTo":"87slv7zvqj.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-11T20:56:12Z","receivedAt":"2005-10-11T20:56:12Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Oct 2005, Paul Eggert wrote:\n> \n> Yes, quite true.  But we don't have to come up with something that's\n> perfect in all cases, just something that's good enough to handle\n> cases that we expect will be common in practice, in a world where\n> UTF-8 is the preferred encoding for non-ASCII characters.\n\nThe thing is, I can almost guarantee you that any quoting in the high \ncharacters is going to be _worse_ than no quoting at all.\n\nExactly because quoting as UTF-8 is the wrong thing when it isn't actually \nUTF-8, and quoting as non-UTF-8 is the wrong thing when it _is_.\n\nNot quoting at all, on the other hand, is unambigious. If you have a \nmailer that corrupts your text stream (which-ever type it is), then it's \nclearly the mailers problem. The _mailer_ at least has a chance in hell to \nknow what character set it is getting mailed as.\n\nThe other alternative is to quote _everything_ non-ASCII. That's \ndefinitely reliable, but it's also unquestionably ugly as hell, especially \nin the long run.\n\nYes, there are some complex quoting approaches you can do, which quote \nthings \"correctly\" (ie at a byte stream level) _and_ keep it valid UTF-8 \nat the same time.\n\nFor example, you can read it as a UTF-8 stream, but then quote things at a \nbyte level (ie if you quote one \"character\", you quote _all_ bytes in that \ncharacter). And you quote if:\n\n - the UTF-8 _character_ is in the 0x80-0x9f control range\n - any _raw_byte_ is in the 0x80-0x9f range (it might not be UTF-8)\n - any _raw_byte_ is 0xfe-0xff (illegal UTF-8 character)\n - misformed UTF-8 (non-shortest sequence, or just generally invalid \n   sequences with missing or wrong high bits)\n\nbut quite frankly, that's a pretty painful thing to write. The upside is \nthat it's easy to decode: you can _unquote_ it just as a byte stream.\n\n\t\t\tLinus\n"},{"id":"10003","messageId":"877jcjmdmq.fsf@penguin.cs.ucla.edu","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510111346220.14597@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2005-10-12T06:51:41Z","receivedAt":"2005-10-12T06:51:41Z","isPatch":true,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> you can read it as a UTF-8 stream, but then quote things at a byte\n> level (ie if you quote one \"character\", you quote _all_ bytes in\n> that character).\n\nYes, that's what I had in mind.\n\n> And you quote if:\n>\n>  - the UTF-8 _character_ is in the 0x80-0x9f control range\n\nYes.  Or more generally, if it's any UTF-8 control character.\n\n>  - any _raw_byte_ is in the 0x80-0x9f range (it might not be UTF-8)\n\nWhy quote the raw bytes?  Is this for terminal escapes on older xterm\n(or xterm-like) implementations that don't understand UTF-8?  If so,\nI'm not sure I'd bother, as it would introduce a lot of annoying\nquoting with perfectly reasonable UTF-8, and (if we assume the world\nis moving to UTF-8) it addresses a problem that is going away.\n\n>  - any _raw_byte_ is 0xfe-0xff (illegal UTF-8 character)\n>  - misformed UTF-8 (non-shortest sequence, or just generally invalid \n>    sequences with missing or wrong high bits)\n\nYes, that makes sense.\n\n> quite frankly, that's a pretty painful thing to write.\n\nIt's not trivially short, yes.  But it shouldn't be that hard.\n\nAlso, I guess we don't have to write it, at least not at first.  As\nlong as we specify something like the C quoted-string format mentioned\nearlier, we can encode into that format using a naive algorithm (e.g.,\nquote any non-ASCII byte or ASCII control character), and beautify the\nencoding method later.\n\n> The upside is that it's easy to decode: you can _unquote_ it just as\n> a byte stream.\n\nYes, that's the idea.\n\nAlso, the interchange format is the most important thing.  We have to\ndecode anything that is in the format, and we must encode into the\nformat.  Encoding prettily is nice, but not necessary.\n"},{"id":"10025","messageId":"Pine.LNX.4.64.0510120749230.14597@g5.osdl.org","threadId":"2042","inReplyTo":"877jcjmdmq.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-12T14:59:55Z","receivedAt":"2005-10-12T14:59:55Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Oct 2005, Paul Eggert wrote:\n> \n> >  - any _raw_byte_ is in the 0x80-0x9f range (it might not be UTF-8)\n> \n> Why quote the raw bytes?  Is this for terminal escapes on older xterm\n> (or xterm-like) implementations that don't understand UTF-8?\n\nIt's not about \"understanding\" UTF-8.\n\nEven a perfectly modern xterm may simply not be in UTF-8 mode: if it \nwasn't in an UTF-8 locale, then it won't do UTF-8 decoding.\n\n> If so, I'm not sure I'd bother, as it would introduce a lot of annoying\n> quoting with perfectly reasonable UTF-8, and (if we assume the world\n> is moving to UTF-8) it addresses a problem that is going away.\n\nUTF-8 is only _now_ getting really widespread, and I think it's because \nRedHat bit the bullet and made UTF-8 the default locale a few years ago.\n\nThese things take _decades_.\n\nI don't know if you realize it, but it's only within the last couple of \nyears that the old 7-bit \"finnish ASCII\" went away. Finnish and Swedish \nhave three extra characters: åäö (latin1) and Ã¥Ã¤Ã¶ (utf-8). But only\nwithin the last few years has the really _old_ ASCII representation really \ngone away so much that I don't see it at all (the characters '{' '}' and \n'|' were taken over, so that if you had a Finnish ASCII font, programming \nin C was really funky - but it was common enough that I could do it \nwithout thinking much about it ;)\n\nSo lots of people still use the byte-wide encodings. Whether really old \nASCII only or some special locale-dependent one (of which latin1 and the \n\"win-latin1\" thing are obviously the most common by far). And in that \nlocale, it's not the UTF-8 control characters that matter, it's the _byte_ \ncontrol characters that do.\n\nSo if you want to support any other locale than UTF-8, you need to escape \nthem. Assuming you want to escape control characters at all, of course (I \nstill think it's perfectly fine to just let the raw mess through and \ndepend on escaping at higher levels)\n\n\t\t\tLinus"},{"id":"10036","messageId":"Pine.LNX.4.63.0510121452030.23242@iabervon.org","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510120749230.14597@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-10-12T19:07:49Z","receivedAt":"2005-10-12T19:07:49Z","isPatch":true,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Wed, 12 Oct 2005, Linus Torvalds wrote:\n\n> So if you want to support any other locale than UTF-8, you need to escape \n> them. Assuming you want to escape control characters at all, of course (I \n> still think it's perfectly fine to just let the raw mess through and \n> depend on escaping at higher levels)\n\nI think it's actually sufficient to escape 0x00-0x1f and 0x7f; those \nranges are both easy and, as far as I can tell, include all of the control \ncharacters that do annoying things. I think escape, backspace, delete, and \nbell are the only ones we'd rather the terminal not get; beyond that, \npatches with screwy filenames look screwy, but don't screw up anything \noutside of the filename.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"10037","messageId":"Pine.LNX.4.64.0510121220230.15297@g5.osdl.org","threadId":"2042","inReplyTo":"Pine.LNX.4.63.0510121452030.23242@iabervon.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-12T19:52:41Z","receivedAt":"2005-10-12T19:52:41Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 12 Oct 2005, Daniel Barkalow wrote:\n> \n> I think it's actually sufficient to escape 0x00-0x1f and 0x7f; those \n> ranges are both easy\n\nThey are indeed easy.\n\n>\t\t and, as far as I can tell, include all of the control \n> characters that do annoying things.\n\nNope. The traditional vt100 escape sequence is \"ESC\" followed by a \ncharacter to indicate the type of sequence (the most common one is '['). \nThat's all 7-bit and fine.\n\nHOWEVER, they made the 8-bit extension be such that any of these vt100 \nbegin sequences where the second character is in the appropriate range can \nbe instead shortened by one character, by instead using a single 8-bit \ncharacter of \"0x80+(char-0x40)\". Ie the traditional \"ESC + '['\" (\\x1b\\x5b) \ncan also be written as a single '\\x9b' character, aka CSI.\n\nIn other words, 0x80-0x9f are _all_ just vt100 shorthand for ESC+'@' \nthrough ESC+'_'.\n\n(I guess it's not strictly \"vt100\" any more - it's the extended vt220 \nformat).\n\n> I think escape, backspace, delete, and \n> bell are the only ones we'd rather the terminal not get; beyond that, \n> patches with screwy filenames look screwy, but don't screw up anything \n> outside of the filename.\n\nTry this on a (non-UTF-8) xterm:\n\n\techo -en '\\x9b5B---\\x9b1A---\\x9b4A\\r'\n\nand it should do:\n - move cursor 5 lines down\n - print \"---\"\n - move cursor 1 line up\n - print \"---\"\n - move cursor 4 lines up\n - return carriage to beginning.\n\nIn other words, your screen should end up looking something like this:\n\n\t[torvalds@g5 ~]$ echo -en '\\x9b5B---\\x9b1A---\\x9b4A\\r'\n\t[torvalds@g5 ~]$\n\t\n\t\n\t\n\t   ---\n\t---\n\nwhere that \"staircase\" of two \"---\" things was done with cursor movements.\n\nAnd that's a _benign_ sequence. You can do all kinds of funky stuff that \nreally screws up the user experience. Including have the thing echo keys \nto you that you didn't type:\n\n\techo -en '\\x9b5n'\n\nor lock the keyboard (I don't think any of the terminal emulators \nimplement the latter, or some of the other stranger sequences - things to \ndo double-wide characters etc).\n\n\t\t\tLinus\n\nPS. You can do all the same in UTF-8 one, but then you'll have to add a \n\\xc2 before the \\x9b:\n\n\techo -en '\\xc2\\x9b5B---\\xc2\\x9b1A---\\xc2\\x9b4A\\r'\n\netc..\n"},{"id":"10038","messageId":"434D702F.4060409@zytor.com","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510121220230.15297@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-10-12T20:21:03Z","receivedAt":"2005-10-12T20:21:03Z","isPatch":true,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> Nope. The traditional vt100 escape sequence is \"ESC\" followed by a \n> character to indicate the type of sequence (the most common one is '['). \n> That's all 7-bit and fine.\n> \n> HOWEVER, they made the 8-bit extension be such that any of these vt100 \n> begin sequences where the second character is in the appropriate range can \n> be instead shortened by one character, by instead using a single 8-bit \n> character of \"0x80+(char-0x40)\". Ie the traditional \"ESC + '['\" (\\x1b\\x5b) \n> can also be written as a single '\\x9b' character, aka CSI.\n> \n> In other words, 0x80-0x9f are _all_ just vt100 shorthand for ESC+'@' \n> through ESC+'_'.\n> \n> (I guess it's not strictly \"vt100\" any more - it's the extended vt220 \n> format).\n> \n\nActually, it's even trickier than that.\n\nCSI is character 0x1b of control code set C1; there are two \"windows\" \nfor control codes -- CL (0x00-0x1f) and CR (0x80-0x9f).  Normally CL is \nmapped to C0 and CR is mapped to CL, but ESC will temporarily map C1 \ninto CL.\n\nVT1xx didn't support this since they didn't support 8-bit anything.\n\nAnyway, a *lot* of character sets -- not just UTF-8 -- use the CR range \nof bytes for printables.\n\n\t-hpa\n"},{"id":"10039","messageId":"7virw2e9ef.fsf@assigned-by-dhcp.cox.net","threadId":"2042","inReplyTo":"87vf02qy79.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-12T21:02:32Z","receivedAt":"2005-10-12T21:02:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Paul Eggert <eggert@CS.UCLA.EDU> writes:\n\n> Linus Torvalds <torvalds@osdl.org> writes:\n>\n>> I don't know if you realize it, but it's only within the last couple of \n>> years that the old 7-bit \"finnish ASCII\" went away.\n>\n> Aach!  Those Finns!  Always on the trailing edge of technology!\n\nNah, Japanese are much worse.  We are so used to see Yen signs\nat the end of multi-line CPP macro definitions (backslashes are\ntaken over by it) and I do not foresee it going away anytime\nsoon.  I think windows people believe Yen signs are path\ncomponent separators ;-).\n"},{"id":"10040","messageId":"Pine.LNX.4.64.0510121355280.15297@g5.osdl.org","threadId":"2042","inReplyTo":"87vf02qy79.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-12T21:05:25Z","receivedAt":"2005-10-12T21:05:25Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 12 Oct 2005, Paul Eggert wrote:\n> \n> Your email message suggests that we need to be cautious here.\n> That message contained UTF-8 text but its header said \"Content-Type:\n> TEXT/PLAIN; charset=ISO-8859-1\".\n\nWell, my email message was wrong and evil, because it _mixed_ two \ndifferent encodings in the same text. No sane client could have shown them \nboth at the same time - but especially with a stupid client, you could \nhave changed your terminal to show either one or the other by switching \nfrom utf-8 to latin1 encoding and doing a refresh.\n\nIn other words, my email really was a nasty case of not one or the other, \nbut both.\n\nNow, I believe patches can actually be that way - it's not at all \nimpossible to have a diff where the _filename_ is utf-8, but the content \nof the patch itself is some byte-encoding like latin1. Or the other way \naround.\n\n> If we're still having problems like this in 2005 then I guess we need\n> to deal with them.  This suggests we should be escaping every\n> non-ASCII byte, at least for patches designed to be emailed robustly.\n\nI find that email is very robust - it's basically 8-bit clean. No \ncharacter encoding, no crap. Just a byte stream. It really _is_ the most \nreliable format.\n\nNow, a lot of email clients are really weak in _showing_ it, and as \nmentioned, the email that mixed both is fundamentally not something you \nreally even _can_ show sanely. But who cares? What matters is not what it \nlooks like, but what it _saves_ as. If you save the email message, it \nshould come out as the same reliable 8-bit byte stream, or your client is \nactively corrupting messages rather than just showing them.\n\nThis is really what my argument boils down to: character set encoding \nshould _not_ EVER affect the _transfer_ of the data. It doesn't matter if \nsomething is latin1 or utf-8, the only thing that matters is the byte \nsequence. Only when you _display_ it should you try to figure out what the \nbyte sequence possibly means.\n\nSo I repeat: \n - escape as little as possible\n - make the _viewer_ decide how to view it.\n\nYes, if people use \"cat\" to view patches, it can be dangerous. But that's \n_their_ problem.\n\n\t\tLinus\n"},{"id":"10041","messageId":"434D7B7E.6060403@zytor.com","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510121355280.15297@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-10-12T21:09:18Z","receivedAt":"2005-10-12T21:09:18Z","isPatch":true,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> Now, I believe patches can actually be that way - it's not at all \n> impossible to have a diff where the _filename_ is utf-8, but the content \n> of the patch itself is some byte-encoding like latin1. Or the other way \n> around.\n> \n\nOr both.  Trivial example: a patch to change names in comments from ISO \n8859-1 to UTF-8.\n\n\t-hpa\n"},{"id":"10042","messageId":"Pine.LNX.4.63.0510122314450.9599@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510121355280.15297@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2005-10-12T21:15:13Z","receivedAt":"2005-10-12T21:15:13Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 12 Oct 2005, Linus Torvalds wrote:\n\n> Yes, if people use \"cat\" to view patches, it can be dangerous. But that's \n> _their_ problem.\n\nNo, that is the cat's problem. Sorry, couldn't resist.\n\nCiao,\nDscho\n"},{"id":"10043","messageId":"Pine.LNX.4.64.0510121411550.15297@g5.osdl.org","threadId":"2042","inReplyTo":"87vf02qy79.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-12T21:24:31Z","receivedAt":"2005-10-12T21:24:31Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 12 Oct 2005, Paul Eggert wrote:\n> \n> Worse, when I used Emacs to copy your text into another file -- the\n> sort of thing that is likely to be done with an emailed patch -- the\n> file contained the UTF-8 encoding of the gibberish, rather than the\n> original bytes of your message.\n\nBtw, this is an example of where locale-based character translations just \nfundamentally suck.\n\ncut-and-paste quote naturally tries to translate between the source \nand destination locales, but it fundamentally cannot work. The only thing \nthat ever works is bit-for-bit copying.\n\nAny program that tries to do locale conversion is always going to be a bug \nwaiting to happen.\n\nIf GNU emacs does locale translations rather than just do a binary \ntransfer of the data, then that's a sign that GNU emavs is being really \nstupid. If the data was UTF-8 to begin with, then a binary copy is also \ngoing to be UTF-8. And if it wasn't UTF-8, then a binary copy is the only \nthing that is sensible.\n\nAnd this is the thing that makes UTF-8 so wonderful: exactly the fact that \nit makes bit-for-bit copying an acceptable policy again, and locales \nbecome a non-issue. In a truly UTF-8 world, you should _never_ convert \nanything at all (and that includes mis-formed UTF-8).\n\nAny non-binary file saving or transfer approach where characters have \n\"meaning\" is always mistake. It's why DOS/Windows \"binary\" vs \"text\" files \nwas wrong. It's why font-encoding locales are wrong (Mixed text with two \ntypes? Yet another metadata quoting scheme? No thank you! It's also why \nUCS-16 and UCS-32 were total disasters: they had \"context\" in their \nencoding).\n\nSay \"yes\" to binary transfer. Because text transfers are broken.\n\n\t\t\tLinus\n"},{"id":"10044","messageId":"7vwtkictdn.fsf@assigned-by-dhcp.cox.net","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510121355280.15297@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-12T21:33:56Z","receivedAt":"2005-10-12T21:33:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> This is really what my argument boils down to: character set encoding \n> should _not_ EVER affect the _transfer_ of the data. It doesn't matter if \n> something is latin1 or utf-8, the only thing that matters is the byte \n> sequence. Only when you _display_ it should you try to figure out what the \n> byte sequence possibly means.\n>\n> So I repeat: \n>  - escape as little as possible\n>  - make the _viewer_ decide how to view it.\n\nI think the same argument can be made about patch application,\nalthough strictly speaking it is not \"viewing\".  Let the patch\nprogram decide (or the user to tell her decision to the patch\nprogram) what the unescaped byte sequence in the patch that\nrepresents the path being affected is encoded in, and do\nsomething sensible while taking into account that the pathname\nencoding on the working tree may be different from what is\nrecorded in the patch.\n\nFor example, one of my partitions is ntfs mounted with\nnls=euc-jp, and I expect the tool to help me apply patches to a\nJapanese-named file when the patch is from a system with UTF-8\nencoded filenames.\n"},{"id":"10086","messageId":"87irw1q7eu.fsf@penguin.cs.ucla.edu","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510121411550.15297@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2005-10-14T00:16:57Z","receivedAt":"2005-10-14T00:16:57Z","isPatch":true,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> So I repeat: \n>  - escape as little as possible\n>  - make the _viewer_ decide how to view it.\n\nUnder my most recent proposal, the only bytes one must escape are \",\n\\, and LF.  Doesn't that satisfy these two main criteria?\n\n\n> If GNU emacs does locale translations rather than just do a binary\n> transfer of the data, then that's a sign that GNU emacs is being\n> really stupid.\n\nPerhaps so, but it has a lot of company.  I have even worse problems\nwith Mozilla Thunderbird.  And as we observed, Pine also has problems\nsending properly-formatted email containing arbitrary binary data.\n\nI suspect the vast majority of email clients will screw up in\nrelatively common cases involving unusual characters in file names.\nUsing attachments avoids many of the problems, but lots of patches are\nemailed inline and I'd rather not force people to use attachments to\nsend diffs.\n\n\n> I find that email is very robust - it's basically 8-bit clean. No \n> character encoding, no crap. Just a byte stream. It really _is_ the most \n> reliable format.\n\nHmm.  To test that theory, I just now sent plain-text email to myself,\ncontaining a carriage-return (CR) byte in the middle of a line.\n\nThe CR byte was transliterated into a LF.  Ooops.\n\nThis was the very first (and only) test I tried, which isn't a good\nsign for reliability.  If you're curious, I tracked the problem down\nto Exim, a popular mail transfer agent that is running on my personal\nDebian GNU/Linux (stable) box.  As to why Exim munges email, please see\n<http://www.exim.org/exim-html-4.40/doc/html/spec_44.html#SECT44.1>.\n(And I didn't know about the Exim glitch before trying my test.\nI'm normally a Sendmail man myself.)\n\nMore generally, I suspect inline patches with weird bytes will suffer\ngreatly from encoding and recoding by mail agents.\n\n\n> What matters is not what it looks like, but what it _saves_ as. If\n> you save the email message, it should come out as the same reliable\n> 8-bit byte stream\n\nUnfortunately this isn't true for Emacs, and I suspect other mailers\nwill have similar problems.  For example, with Emacs I can easily save\neither the exact byte-for-byte message body that my mail transfer\nagent gave me; or I can have Emacs decode the message into its\nconstituent characters, reencode the result as UTF-8, and put that\ninto a file.  In neither case, though, am I saving the original byte\nstream that you presented to your mail user agent.  Even if I save the\nbyte-for-byte message body, it is often in quoted-printable format so\nI'll have to decode strings like \"=EF\" to recover the original bytes.\nThis is doable, yes, but it's inconvenient in practice, at least with\nthe mail user agents I'm familiar with.  And even if I do it, I don't\nnecessarily have the same byte stream you gave your mail user agent; I\nmerely have the byte stream that your MUA gave to your MTA, and these\nmay not be the same thing (they certainly aren't always the same thing\nwith Emacs).\n\n\nThe simplest fix for git may be to say \"Don't use inline patches; use\nattachments if you must email anything with strange characters in it.\"\nThat's fine.  But I prefer a format that also allows GNU diff, if it\nchooses, to generate output that resists common inline-email botches.\n"},{"id":"10087","messageId":"87ek6ork3y.fsf@penguin.cs.ucla.edu","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510121355280.15297@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2005-10-14T00:57:21Z","receivedAt":"2005-10-14T00:57:21Z","isPatch":true,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> I find that email is very robust - it's basically 8-bit clean. No \n> character encoding, no crap. Just a byte stream. It really _is_ the most \n> reliable format.\n\nI found another amusing bit of info that tends to undercut this claim.\n\nThis discussion thread is archived at\n<http://marc.theaimsgroup.com/?t=112877773400002&r=1&w=2&n=22>.\nBut there's an item missing from the archive: my message with\nMessage-ID <87vf02qy79.fsf@penguin.cs.ucla.edu>.  This is the message\nwith the joke \"Aach!  Those Finns!  Always on the trailing edge of\ntechnology!\".\n\nAll my other messages are achived.  What was special about this\none?  Surely there's not a joke filter at theaimsgroup.com!\n\nI nosed around through the archive and here's my guess as to what\nhappened.  My message's email header contained this:\n\n   Content-Type: text/plain; charset=utf-8\n   Content-Transfer-Encoding: quoted-printable\n\nand my guess is that the web archiver can't handle that format.\n\nThis is just a guess.  I can't confirm it because (among other things)\nthe web archiver won't give me all the bytes of the messages that it\narchives.  Even its \"Download message RAW\" doesn't do that: it omits\nthe header.  But I have a strong suspicion.  Let's put it this way: I\nthink mine was the only message in the thread that said\n\"charset=utf-8\".\n\nIf my guess is right, the archiver dropped my email on the floor\nsimply because it contained UTF-8.  This is not a good sign for\nputting UTF-8 into email, or for relying on email to transmit byte\nstreams.\n"},{"id":"10089","messageId":"Pine.LNX.4.64.0510132203220.23590@g5.osdl.org","threadId":"2042","inReplyTo":"87irw1q7eu.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-14T05:20:53Z","receivedAt":"2005-10-14T05:20:53Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 13 Oct 2005, Paul Eggert wrote:\n> \n> Perhaps so, but it has a lot of company.  I have even worse problems\n> with Mozilla Thunderbird.  And as we observed, Pine also has problems\n> sending properly-formatted email containing arbitrary binary data.\n\nNo, pine does it right. Exactly because it sends _arbitraty_ binary data.\n\nThe fact that I turned the terminal into utf-8 mode in order to generate \nthe bytes (that end up being a garbage string in latin1) is not pine's \nfault. \n\nThe point being that because the transport was 8-bit clean, I could do \nthat. I could mix a latin1-encoding with a UTF-8 encoding, and the other \nside could see the mixed setting. Now, the other side had no way of \nknowing that I mixed things (unless it was a smart human and could read \nand understand what I wrote), so any email client would have trouble \nshowing it.\n\nBut it got _transferred_ right, and you could have saved the email, and \nturned the terminal into latin1 or utf-8 mode, and done a \"cat\" both ways, \nand you'd have seen both versions.\n\n> I suspect the vast majority of email clients will screw up in\n> relatively common cases involving unusual characters in file names.\n\nNot if they just save it.\n\nOh, sure, they can't _display_ it, since they don't know what it is, but \nwhen they save it, they'd _better_ save it bit-for-bit.\n\nWhich is the right thing to do. Then you apply it with \"patch\", and you \nget the right answer.\n\n> Using attachments avoids many of the problems, but lots of patches are\n> emailed inline and I'd rather not force people to use attachments to\n> send diffs.\n\ninline or attachment should not matter to any sane email client. If it \ndoes, then the email client isn't sane.\n\nThe point is, when you save it, it _has_ to be saved bit-for-bit. \n\nThe only difference between a binary attachment and a text thing is that \nan email client will _try_ to show the text thing to you as text. It has \nno other meaning.\n\nAnd trying is better than not trying. Attachments are _inferior_ to inline \nfor that reason.\n\n> > I find that email is very robust - it's basically 8-bit clean. No \n> > character encoding, no crap. Just a byte stream. It really _is_ the most \n> > reliable format.\n> \n> Hmm.  To test that theory, I just now sent plain-text email to myself,\n> containing a carriage-return (CR) byte in the middle of a line.\n> \n> The CR byte was transliterated into a LF.  Ooops.\n\nI'm not surprised, since CR/LF is special for a lot of (sad) reasons. Oh, \nwell.\n\nI agree that it makes sense to escape \\r, and obviously you _have_ to \nescape \\n. In general, escaping pretty much everything in the 0-31 range \nis likely the right approach, since those are never printable anyway.\n\nThat, btw, is probably true of the patch contents too, not just the \nfilename. The exception being \\t (and in patch contents, \\n is obviously \npart of the stream).\n\n> More generally, I suspect inline patches with weird bytes will suffer\n> greatly from encoding and recoding by mail agents.\n\nI've had pretty good luck. We do have 8-bit stuff occasionally, but it \nalmost always makes it through. \n\nSpaces and tabs are much worse (yes, they're more common too). That's \nclearly just crap mailers.\n\n> Unfortunately this isn't true for Emacs, and I suspect other mailers\n> will have similar problems.  For example, with Emacs I can easily save\n> either the exact byte-for-byte message body that my mail transfer\n> agent gave me; or I can have Emacs decode the message into its\n> constituent characters, reencode the result as UTF-8, and put that\n> into a file.\n\nWell, as long as there's a choice.\n\n> In neither case, though, am I saving the original byte\n> stream that you presented to your mail user agent.  Even if I save the\n> byte-for-byte message body, it is often in quoted-printable format so\n> I'll have to decode strings like \"=EF\" to recover the original bytes.\n\nYou have a broken mail client. Now, I'm not a big fan of QP (I think it \nwas making a stupid excuse for bad transport), but QP is a _mail_ level \nquoting protocol, and the same way a MUA uses QP to encode, the MUA should \nhave de-coded the QP. It shouldn't leave it to somebody else.\n\nI think GNU emacs is a horrible mistake (\"do everything - badly\"), but you \nmay be able to fix it by letting your mail transport agent do the un-QP \nfor you. A lot of them do, which makes it easier to then use weak MUA's.\n\nAnyway, it sounds like GNU emacs made the wrong choices (hey, I'm not \nsurprised). It should have decoded QP, not the character set. There are \nlots of tools that do charset conversions, that's not very email-specific.\n\n\t\t\tLinus\n"},{"id":"10090","messageId":"Pine.LNX.4.64.0510132229490.23590@g5.osdl.org","threadId":"2042","inReplyTo":"87ek6ork3y.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-14T05:43:19Z","receivedAt":"2005-10-14T05:43:19Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 13 Oct 2005, Paul Eggert wrote:\n\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n> > I find that email is very robust - it's basically 8-bit clean. No \n> > character encoding, no crap. Just a byte stream. It really _is_ the most \n> > reliable format.\n> \n> I found another amusing bit of info that tends to undercut this claim.\n\nNo, I think you found that email as a _transfer_ is mostly 8-bit clean \n(finally! Oh - has qmail gotten fixed?).\n\nBut the end-points aren't. They do strange things with encodings, \nsometimes. They see an encoding they don't know what to do with, and they \njust freak out.\n\n\t\tLinus\n"},{"id":"10095","messageId":"7vll0wvb2a.fsf@assigned-by-dhcp.cox.net","threadId":"2042","inReplyTo":"87vf02qy79.fsf@penguin.cs.ucla.edu","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-14T06:59:09Z","receivedAt":"2005-10-14T06:59:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Paul Eggert <eggert@CS.UCLA.EDU> writes:\n\n> Here is the proposed format.  Each file name is a string of bytes, in\n> one of the following two formats:\n>\n> A.  A nonempty sequence of ASCII graphic characters (i.e., bytes in\n>     the range '!' == '\\041' through '~' == '\\177').  The first byte\n>     cannot be '!' == '\\041' or '\"' == '\\042'.  Leading '\"' is used for\n>     (B) below, and leading '!' is reserved for future extensions.\n>\n> B.  A nonempty C-language character string literal, with the following\n>     restrictions and modifications:\n>\n>     B1.  No multibyte character processing is done.  Members of the\n>          string literal are treated as bytes, not characters.  Null\n>          bytes are not allowed, and '\"' == '\\042', '\\\\' == '\\134' and\n>          '\\n' == '\\012' are allowed only if properly escaped as shown\n>          below; but all other bytes are allowed.\n>\n>     B2.  No trigraph processing is done (e.g., ??/ stands for three\n>          bytes, not one).\n>\n>     B3.  No line-splicing is done (i.e., backslash-newline is not allowed).\n>\n>     B4.  Only the following escape sequences are allowed.\n>\n>            \\\" \\\\ \\a \\b \\f \\n \\r \\t \\v\n>            \\XYZ  (where X, Y, and Z are octal digits, X <= 3, and\n>                   at least one of the digits is nonzero)\n\nJust to let you know, I am slowly converting apply.c to accept\nthis format, and also diff.c to produce this.  I did not\npersonally like the missing double quotes around what I did\nanyway, although it was easier to code.\n"},{"id":"10107","messageId":"434FE857.4040201@zytor.com","threadId":"2042","inReplyTo":"Pine.LNX.4.64.0510132203220.23590@g5.osdl.org","subject":"Re: [PATCH] Try URI quoting for embedded TAB and LF in pathnames","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-10-14T17:18:15Z","receivedAt":"2005-10-14T17:18:15Z","isPatch":true,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> No, pine does it right. Exactly because it sends _arbitraty_ binary data.\n> \n> The fact that I turned the terminal into utf-8 mode in order to generate \n> the bytes (that end up being a garbage string in latin1) is not pine's \n> fault. \n> \n\nI would think a full-screen editor would need to know about multibyte \nencodings.\n\n\t-hpa\n"}]}