{"thread":{"id":"4123","subject":"git-mailinfo '-u' argument should be default.","startedAt":"2006-05-12T16:46:02Z","lastAt":"2007-01-10T03:24:53Z","messageCount":8,"participants":["David Woodhouse","Johannes Schindelin","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"19865","messageId":"1147452362.2794.452.camel@pmac.infradead.org","threadId":"4123","inReplyTo":null,"subject":"git-mailinfo '-u' argument should be default.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2006-05-12T16:46:02Z","receivedAt":"2006-05-12T16:46:02Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"If you apply a patch with 'git-am', it takes the raw content of the\nFrom: header, in whatever character set that happens to be, and puts\nthat content, untagged, into the commit object. \n\nThat is almost _never_ the right thing to do, surely? Raw data in\nuntagged character sets is marginally better than line noise.\n\nWe should default to the '-u' behaviour, which converts to the character\nset which the git repo is stored in. It's not just for converting to\nUTF-8, although it looks like it was in the past. Now, however, it\nconverts to the character set defined in the configuration. So even if\nit's a Luddite who for some reason is sticking with an obsolete\ncharacter set, it gets that right.\n\nThe only option which makes sense _other_ than that would be to just\nstick the RFC2047-encoded original From: header into the commit, surely?\n\n-- \ndwmw2\n"},{"id":"31241","messageId":"1168351405.14763.347.camel@shinybook.infradead.org","threadId":"4123","inReplyTo":"1147452362.2794.452.camel@pmac.infradead.org","subject":"[PATCH] Re: git-mailinfo '-u' argument should be default.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2007-01-09T14:03:25Z","receivedAt":"2007-01-09T14:03:25Z","isPatch":true,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Fri, 2006-05-12 at 17:46 +0100, David Woodhouse wrote:\n> If you apply a patch with 'git-am', it takes the raw content of the\n> From: header, in whatever character set that happens to be, and puts\n> that content, untagged, into the commit object. \n> \n> That is almost _never_ the right thing to do, surely? Raw data in\n> untagged character sets is marginally better than line noise.\n> \n> We should default to the '-u' behaviour, which converts to the character\n> set which the git repo is stored in. It's not just for converting to\n> UTF-8, although it looks like it was in the past. Now, however, it\n> converts to the character set defined in the configuration. So even if\n> it's a Luddite who for some reason is sticking with an obsolete\n> character set, it gets that right.\n\nThis patch:\n 1. Fixes the default not to throw away the MIME information.\n 2. Adds a '-n' option with the old behaviour, although I can't\n    actually imagine why someone might find that desirable.\n 3. Aborts if the conversion fails, allowing the user to fix it \n    rather than silently corrupting the input. There's always the\n    new '-n' option if the user _really_ wants it corrupted. :)\n\ndiff --git a/Documentation/git-mailinfo.txt b/Documentation/git-mailinfo.txt\nindex ea0a065..97b6c5b 100644\n--- a/Documentation/git-mailinfo.txt\n+++ b/Documentation/git-mailinfo.txt\n@@ -8,7 +8,7 @@ git-mailinfo - Extracts patch from a single e-mail message\n \n SYNOPSIS\n --------\n-'git-mailinfo' [-k] [-u | --encoding=<encoding>] <msg> <patch>\n+'git-mailinfo' [-k] [-n | --encoding=<encoding>] <msg> <patch>\n \n \n DESCRIPTION\n@@ -34,19 +34,20 @@ OPTIONS\n \n -u::\n \tBy default, the commit log message, author name and\n-\tauthor email are taken from the e-mail without any\n-\tcharset conversion, after minimally decoding MIME\n-\ttransfer encoding.  This flag causes the resulting\n-\tcommit to be encoded in the encoding specified by\n-\ti18n.commitencoding configuration (defaults to utf-8) by\n-\ttransliterating them. \n+\tauthor email are transliterated if necessary in order \n+\tto store them in the correct encoding as specified by\n+\tthe i18n.commitencoding configuration option (which\n+\tdefaults to UTF-8). This flag disables such character\n+\tset conversion; performing basic MIME transfer decoding\n+\tbut then intentionally discarding the character set\n+\tinformation and using the raw byte sequences.\n \tNote that the patch is always used as is without charset\n-\tconversion, even with this flag.\n+\tconversion, even without this flag.\n \n --encoding=<encoding>::\n-\tSimilar to -u but if the local convention is different\n-\tfrom what is specified by i18n.commitencoding, this flag\n-\tcan be used to override it.\n+\tIf the i18n.commitencoding configuration option is set\n+\tset incorrectly (or unset when the UTF-8 default is not\n+\tdesired), this flag can be used to override it.\n \n <msg>::\n \tThe commit log message extracted from e-mail, usually\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex a67f3eb..02bd183 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -518,8 +518,7 @@ static void convert_to_utf8(char *line, char *charset)\n \tif (!out) {\n \t\tfprintf(stderr, \"cannot convert from %s to %s\\n\",\n \t\t\tinput_charset, metainfo_charset);\n-\t\t*charset = 0;\n-\t\treturn;\n+\t\texit(1);\n \t}\n \tstrcpy(line, out);\n \tfree(out);\n@@ -793,7 +792,7 @@ int mailinfo(FILE *in, FILE *out, int ks, const char *encoding,\n }\n \n static const char mailinfo_usage[] =\n-\t\"git-mailinfo [-k] [-u | --encoding=<encoding>] msg patch <mail >info\";\n+\t\"git-mailinfo [-k] [-n | --encoding=<encoding>] msg patch <mail >info\";\n \n int cmd_mailinfo(int argc, const char **argv, const char *prefix)\n {\n@@ -802,12 +801,16 @@ int cmd_mailinfo(int argc, const char **argv, const char *prefix)\n \t */\n \tgit_config(git_default_config);\n \n+\tmetainfo_charset = (git_commit_encoding\n+\t\t\t    ? git_commit_encoding : \"utf-8\");\n+\n \twhile (1 < argc && argv[1][0] == '-') {\n \t\tif (!strcmp(argv[1], \"-k\"))\n \t\t\tkeep_subject = 1;\n-\t\telse if (!strcmp(argv[1], \"-u\"))\n-\t\t\tmetainfo_charset = (git_commit_encoding\n-\t\t\t\t\t    ? git_commit_encoding : \"utf-8\");\n+\t\telse if (!strcmp(argv[1], \"-u\")) \n+\t\t\t; /* this is now default behaviour */\n+\t\telse if (!strcmp(argv[1], \"-n\")) \n+\t\t\tmetainfo_charset = NULL;\n \t\telse if (!strncmp(argv[1], \"--encoding=\", 11))\n \t\t\tmetainfo_charset = argv[1] + 11;\n \t\telse\n\n-- \ndwmw2\n"},{"id":"31243","messageId":"Pine.LNX.4.63.0701091505500.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"4123","inReplyTo":"1168351405.14763.347.camel@shinybook.infradead.org","subject":"Re: [PATCH] Re: git-mailinfo '-u' argument should be default.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-01-09T14:08:07Z","receivedAt":"2007-01-09T14:08:07Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 9 Jan 2007, David Woodhouse wrote:\n\nin the Documentation,\n\n>  -u::\n\nneeds to be replaced by \"-n::\", and\n\n> +\tIf the i18n.commitencoding configuration option is set\n> +\tset incorrectly (or unset when the UTF-8 default is not\n\nneeds only one \"set\", not two.\n\nCiao,\nDscho\n"},{"id":"31244","messageId":"1168352489.14763.352.camel@shinybook.infradead.org","threadId":"4123","inReplyTo":"Pine.LNX.4.63.0701091505500.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: [PATCH] Re: git-mailinfo '-u' argument should be default.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2007-01-09T14:21:28Z","receivedAt":"2007-01-09T14:21:28Z","isPatch":true,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Tue, 2007-01-09 at 15:08 +0100, Johannes Schindelin wrote:\n> Hi,\n> \n> On Tue, 9 Jan 2007, David Woodhouse wrote:\n> \n> in the Documentation,\n> \n> >  -u::\n> \n> needs to be replaced by \"-n::\", and\n> \n> > +\tIf the i18n.commitencoding configuration option is set\n> > +\tset incorrectly (or unset when the UTF-8 default is not\n> \n> needs only one \"set\", not two.\n\nOops; thanks.\n\ndiff --git a/Documentation/git-mailinfo.txt b/Documentation/git-mailinfo.txt\nindex ea0a065..e267053 100644\n--- a/Documentation/git-mailinfo.txt\n+++ b/Documentation/git-mailinfo.txt\n@@ -8,7 +8,7 @@ git-mailinfo - Extracts patch from a single e-mail message\n \n SYNOPSIS\n --------\n-'git-mailinfo' [-k] [-u | --encoding=<encoding>] <msg> <patch>\n+'git-mailinfo' [-k] [-n | --encoding=<encoding>] <msg> <patch>\n \n \n DESCRIPTION\n@@ -32,21 +32,22 @@ OPTIONS\n \tmunging, and is most useful when used to read back 'git\n \tformat-patch --mbox' output.\n \n--u::\n+-n::\n \tBy default, the commit log message, author name and\n-\tauthor email are taken from the e-mail without any\n-\tcharset conversion, after minimally decoding MIME\n-\ttransfer encoding.  This flag causes the resulting\n-\tcommit to be encoded in the encoding specified by\n-\ti18n.commitencoding configuration (defaults to utf-8) by\n-\ttransliterating them. \n+\tauthor email are transliterated if necessary in order \n+\tto store them in the correct encoding as specified by\n+\tthe i18n.commitencoding configuration option (which\n+\tdefaults to UTF-8). This flag disables such character\n+\tset conversion; performing basic MIME transfer decoding\n+\tbut then intentionally discarding the character set\n+\tinformation and using the raw byte sequences.\n \tNote that the patch is always used as is without charset\n-\tconversion, even with this flag.\n+\tconversion, even without this flag.\n \n --encoding=<encoding>::\n-\tSimilar to -u but if the local convention is different\n-\tfrom what is specified by i18n.commitencoding, this flag\n-\tcan be used to override it.\n+\tIf the i18n.commitencoding configuration option is set\n+\tincorrectly (or unset when the UTF-8 default is not\n+\tdesired), this flag can be used to override it.\n \n <msg>::\n \tThe commit log message extracted from e-mail, usually\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex a67f3eb..02bd183 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -518,8 +518,7 @@ static void convert_to_utf8(char *line, char *charset)\n \tif (!out) {\n \t\tfprintf(stderr, \"cannot convert from %s to %s\\n\",\n \t\t\tinput_charset, metainfo_charset);\n-\t\t*charset = 0;\n-\t\treturn;\n+\t\texit(1);\n \t}\n \tstrcpy(line, out);\n \tfree(out);\n@@ -793,7 +792,7 @@ int mailinfo(FILE *in, FILE *out, int ks, const char *encoding,\n }\n \n static const char mailinfo_usage[] =\n-\t\"git-mailinfo [-k] [-u | --encoding=<encoding>] msg patch <mail >info\";\n+\t\"git-mailinfo [-k] [-n | --encoding=<encoding>] msg patch <mail >info\";\n \n int cmd_mailinfo(int argc, const char **argv, const char *prefix)\n {\n@@ -802,12 +801,16 @@ int cmd_mailinfo(int argc, const char **argv, const char *prefix)\n \t */\n \tgit_config(git_default_config);\n \n+\tmetainfo_charset = (git_commit_encoding\n+\t\t\t    ? git_commit_encoding : \"utf-8\");\n+\n \twhile (1 < argc && argv[1][0] == '-') {\n \t\tif (!strcmp(argv[1], \"-k\"))\n \t\t\tkeep_subject = 1;\n-\t\telse if (!strcmp(argv[1], \"-u\"))\n-\t\t\tmetainfo_charset = (git_commit_encoding\n-\t\t\t\t\t    ? git_commit_encoding : \"utf-8\");\n+\t\telse if (!strcmp(argv[1], \"-u\")) \n+\t\t\t; /* this is now default behaviour */\n+\t\telse if (!strcmp(argv[1], \"-n\")) \n+\t\t\tmetainfo_charset = NULL;\n \t\telse if (!strncmp(argv[1], \"--encoding=\", 11))\n \t\t\tmetainfo_charset = argv[1] + 11;\n \t\telse\n\n-- \ndwmw2\n"},{"id":"31266","messageId":"7vzm8skphz.fsf@assigned-by-dhcp.cox.net","threadId":"4123","inReplyTo":"1168351405.14763.347.camel@shinybook.infradead.org","subject":"Re: [PATCH] Re: git-mailinfo '-u' argument should be default.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-09T18:46:00Z","receivedAt":"2007-01-09T18:46:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"David Woodhouse <dwmw2@infradead.org> writes:\n\n> On Fri, 2006-05-12 at 17:46 +0100, David Woodhouse wrote:\n>>  ...\n> This patch:\n>  1. Fixes the default not to throw away the MIME information.\n>  2. Adds a '-n' option with the old behaviour, although I can't\n>     actually imagine why someone might find that desirable.\n>  3. Aborts if the conversion fails, allowing the user to fix it \n>     rather than silently corrupting the input. There's always the\n>     new '-n' option if the user _really_ wants it corrupted. :)\n\nThanks.\n\nDocumentation/SubmittingPatches, and Sign-off?\n\nI do not think you would want to make '-n' in the third point\nsound so negative and make people on projects that chose to use\nlegacy encoding for whatever reasons feel _dirty_.  If the\nnatural language in project's log is limited and a legacy\nencoding is sufficient, and if all the participants agree on a\nlegacy encoding to use because tools other than git they need to\nuse are more convenient with the legacy encoding rather than\nUTF-8, there is no need to give a lecture to them saying they\nshould switch to UTF-8 and/or what they have been doing is\nsub-par -- it isn't.\n\nIf the command allows straight-through (and I think it should)\nbut now defaults to UTF-8, you also need to update the existing\nPorcelain-level tools (i.e. callers of mailinfo) so that they\npass -n when the end-user says \"I want straight-through\"; asking\nfor UTF-8 can be done by either not passing anything or\nexplicitly passing -u, but the point is that the callers need to\nbe changed anyway.  And at that point, the default of mailinfo\ndoes not matter that much -- although it is good for consistency\nto make it also default to UTF-8.\n\nI've updated git-am yesterday to default to --utf8 (which is\noverridable with --no-utf8) but did not touch mailinfo during\nthat process; it further needs to be told about the -n option.\nI haven't touched git-applymbox yet but it should also be taught\nabout the new default and the override.\n"},{"id":"31313","messageId":"1168386544.14763.407.camel@shinybook.infradead.org","threadId":"4123","inReplyTo":"7vzm8skphz.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Re: git-mailinfo '-u' argument should be default.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2007-01-09T23:49:04Z","receivedAt":"2007-01-09T23:49:04Z","isPatch":true,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Tue, 2007-01-09 at 10:46 -0800, Junio C Hamano wrote:\n> I do not think you would want to make '-n' in the third point\n> sound so negative \n\nNo, I really _do_ want that.\n\n> and make people on projects that chose to use\n> legacy encoding for whatever reasons feel _dirty_. \n\n... but not that, because it wasn't aimed at them.\n\n>  If the natural language in project's log is limited and a legacy\n> encoding is sufficient, and if all the participants agree on a\n> legacy encoding to use...\n(...for git's own purely internal storage format).\n\nThat's not the use case for the -n option. Their case is what the\ni18n.commitencoding configuration option exists for.\n\nAlthough having said that, I don't actually know _why_ we let them\noverride the default, since it's _internal_ to git. As long as git\nitself is correctly doing the conversion on the way in and out, there's\nno reason for them to care whether we use UTF-8, UCS-4, EBCDIC or some\nother arbitrary encoding (as long as our encoding can represent anything\nthey choose to throw at us).\n\n>  because tools other than git they need to\n> use are more convenient with the legacy encoding rather than\n> UTF-8,\n\nThat makes about as much sense to me as letting them configure git to\nstore objects uncompressed \"because tools other than git are more\nconvenient without compression\". \n\nIf our choice of _internal_ storage affects their other tools, then\neither they're doing something very strange like poking at git objects\ndirectly, or there's a bug in the git tools.\n\n>  there is no need to give a lecture to them saying they\n> should switch to UTF-8 and/or what they have been doing is\n> sub-par -- it isn't. \n\nIf people, for whatever reason, want git to use a given legacy character\nset for its storage format, they just have to set i18n.commitencoding.\nThose people aren't being lecture. (Although perhaps they _should_ be;\neither they're poking at things which shouldn't concern them, or they\nshould be _reporting_ bugs instead of just working round them.)\n\nThe only people who would want the -n option would be those who _want_\nto intentionally throw away the character set encoding, and have one\ncommit¹ in EBCDIC, a second in UTF-8 and a third in BIG5 with no way of\ntelling which is which; each of them _labelled_ with the default\nencoding for the repository, which is probably UTF-8.\n\n-- \ndwmw2\n\n¹ Actually it's worse than that -- with RFC2047 you can have multiple \nencodings within the same _line_ of text. Evolution at least will do that;\nit uses ISO8859-1 for any character it can, and falls back to UTF-8 for\nother characters. Even within the same header. Importing with '-n' would\njust throw away the charset information and use the raw bytes. Even just \nimporting the RFC2047-encoded text as-is would be better than that.\n"},{"id":"31332","messageId":"7vejq3hf8m.fsf@assigned-by-dhcp.cox.net","threadId":"4123","inReplyTo":"1168386544.14763.407.camel@shinybook.infradead.org","subject":"Re: [PATCH] Re: git-mailinfo '-u' argument should be default.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-10T00:55:53Z","receivedAt":"2007-01-10T00:55:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"David Woodhouse <dwmw2@infradead.org> writes:\n\n> On Tue, 2007-01-09 at 10:46 -0800, Junio C Hamano wrote:\n>> I do not think you would want to make '-n' in the third point\n>> sound so negative \n>\n> No, I really _do_ want that.\n>\n>> and make people on projects that chose to use\n>> legacy encoding for whatever reasons feel _dirty_. \n>\n> ... but not that, because it wasn't aimed at them.\n>\n>>  If the natural language in project's log is limited and a legacy\n>> encoding is sufficient, and if all the participants agree on a\n>> legacy encoding to use...\n> (...for git's own purely internal storage format).\n>\n> That's not the use case for the -n option. Their case is what the\n> i18n.commitencoding configuration option exists for.\n\nI think you missed a subtle difference.\n\nFor a project that is latin-1 only (or ISO-2022 only, for that\nmatter -- 'only' is the real keyword here), users did not have\nto do anything, and happily kept using git, and that includes\nthat they did not have to set i18n.commitencoding to anything.\n\nDefaulting to -u now means disrupt their established workflows\nare disrupted by people coming from UTF-8 only world.  They now\nare forced to set i18n.commitencoding to latin-1 and/or use -n.\n\nWhen we inconvenience others by making changes to make our own\nlife easier, it is not a good idea to insult them at the same\ntime; rather, we should be asking forgiveness from them.\n"},{"id":"31357","messageId":"1168399493.14763.462.camel@shinybook.infradead.org","threadId":"4123","inReplyTo":"7vejq3hf8m.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Re: git-mailinfo '-u' argument should be default.","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2007-01-10T03:24:53Z","receivedAt":"2007-01-10T03:24:53Z","isPatch":true,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Tue, 2007-01-09 at 16:55 -0800, Junio C Hamano wrote:\n> > That's not the use case for the -n option. Their case is what the\n> > i18n.commitencoding configuration option exists for.\n> \n> I think you missed a subtle difference.\n> \n> For a project that is latin-1 only (or ISO-2022 only, for that\n> matter -- 'only' is the real keyword here), users did not have\n> to do anything, and happily kept using git, and that includes\n> that they did not have to set i18n.commitencoding to anything.\n\nWell, apart from the fact that gitweb and git-format-patch would be\nmislabelling their output, for example.\n\n> Defaulting to -u now means disrupt their established workflows\n> are disrupted by people coming from UTF-8 only world.  They now\n> are forced to set i18n.commitencoding to latin-1 and/or use -n.\n>\n> When we inconvenience others by making changes to make our own\n> life easier, it is not a good idea to insult them at the same\n> time; rather, we should be asking forgiveness from them.\n\nThis is true. Sometimes when we fix bugs, we find that people were\nrelying on the old behaviour and are inconvenienced by the fix. That's\nunfortunate, and we should make sure there's a simple way for those\npeople to adapt -- which in this case is setting 'i18n.commitencoding'.\n\nI'm not entirely sure where the 'insulting' comes in though -- I think\nyou're referring to a comment which was \n a) Only in my cover message rather than in the documentation, and\n b) Not even relevant to these people -- these people should be \n    setting i18n.commitencoding to match their status quo, and my\n    disparaging comment was about the '-n' option which should be\n    avoided.\n\n-- \ndwmw2\n"}]}