{"thread":{"id":"4161","subject":"[PATCH] CMIT_FMT_EMAIL: Q-encode Subject: and display-name part of From: fields.","startedAt":"2006-05-16T10:18:24Z","lastAt":"2006-05-16T10:49:50Z","messageCount":3,"participants":["Junio C Hamano","Jakub Narebski","Rocco Rutte"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"20056","messageId":"7vmzdi9ssv.fsf@assigned-by-dhcp.cox.net","threadId":"4161","inReplyTo":null,"subject":"[PATCH] CMIT_FMT_EMAIL: Q-encode Subject: and display-name part of From: fields.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-16T10:18:24Z","receivedAt":"2006-05-16T10:18:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"By convention, the commit message and the author/committer names\nin the commit objects are UTF-8 encoded.  When formatting for\ne-mails, Q-encode them according to RFC 2047.\n\nWhile we are at it, generate the content-type and\ncontent-transfer-encoding headers as well.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n\n---\n\n With this patch, the output formatted with\n\n\tgit show --pretty=email --patch-with-stat 9d7f73d4\n\n would start like this:\n\n   From 9d7f73d43fa49d0d2f5a8cfcce9d659e8ad2d265  Thu Apr 7 15:13:13 2005\n   From: =?utf-8?q?Lukas_Sandstr=C3=B6m?= <lukass@etek.chalmers.se>\n   Date: Sat, 25 Feb 2006 12:20:13 +0100\n   Subject: [PATCH] git-fetch: print the new and old ref when fast-forwarding\n   Content-Type: text/plain; charset=UTF-8\n   Content-Transfer-Encoding: 8bit\n\n This is marked RFC because I am not convinced if this kind of\n header formatting should be done by format-patch; we might be\n better off leaving the proper massaging to whatever downstream\n program that reads its output (e.g. send-email or imap-send).\n We produce the mbox format (and that is a requirement -- its\n output should be consumable by git-am), so the downstream needs\n to strip off the initial UNIX-From line at least anyway.\n\n Thoughts?\n\n If we decide to do the header formatting here, there are two\n further enhancements that need to be done:\n\n (1) The charset must be configurable for projects that use\n     encoding different from UTF-8, perhaps with the .git/config\n     [i18n] commitEncoding.  It is only a convention, not a hard\n     rule, to use UTF-8 for the metainformation.\n\n (2) Some projects, notably Wine, seem to prefer patches to be\n     sent as attachments, and we have support for that in the\n     script version of format-patch.  We would want to have the\n     same here.  This needs to be an option; define a new\n     format, CMIT_FMT_MIME, and invoke it with --pretty=mime.\n\n     Ideally we would want to say, in the body part header for\n     the attachment, that the type of the payload is a raw 8bit\n     text/patch without any specific charset (if the upstream\n     project has a UTF-8 encoded file, you should not send in a\n     patch in iso-8859-1 and expect somebody to automagically\n     transcode your patch -- the patch is applied as is and MTA\n     should not molest it).\n\n The RFC2047 q-encoding code definitely needs to be audited by\n an RFC lawyer.  I used to be one myself but I lost my edge and\n patience these days.\n\ndiff --git a/commit.c b/commit.c\nindex 93b3903..dee5756 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -413,6 +413,46 @@ static int get_one_line(const char *msg,\n \treturn ret;\n }\n \n+static int is_rfc2047_special(char ch)\n+{\n+\treturn ((ch & 0x80) || (ch == '=') || (ch == '?') || (ch == '_'));\n+}\n+\n+static int add_rfc2047(char *buf, const char *line, int len)\n+{\n+\tchar *bp = buf;\n+\tint i, needquote;\n+\tstatic const char q_utf8[] = \"=?utf-8?q?\";\n+\n+\tfor (i = needquote = 0; !needquote && i < len; i++) {\n+\t\tunsigned ch = line[i];\n+\t\tif (ch & 0x80)\n+\t\t\tneedquote++;\n+\t\tif ((i + 1 < len) &&\n+\t\t    (ch == '=' && line[i+1] == '?'))\n+\t\t\tneedquote++;\n+\t}\n+\tif (!needquote)\n+\t\treturn sprintf(buf, \"%.*s\", len, line);\n+\n+\tmemcpy(bp, q_utf8, sizeof(q_utf8)-1);\n+\tbp += sizeof(q_utf8)-1;\n+\tfor (i = 0; i < len; i++) {\n+\t\tunsigned ch = line[i];\n+\t\tif (is_rfc2047_special(ch)) {\n+\t\t\tsprintf(bp, \"=%02X\", ch);\n+\t\t\tbp += 3;\n+\t\t}\n+\t\telse if (ch == ' ')\n+\t\t\t*bp++ = '_';\n+\t\telse\n+\t\t\t*bp++ = ch;\n+\t}\n+\tmemcpy(bp, \"?=\", 2);\n+\tbp += 2;\n+\treturn bp - buf;\n+}\n+\n static int add_user_info(const char *what, enum cmit_fmt fmt, char *buf, const char *line)\n {\n \tchar *date;\n@@ -431,12 +471,26 @@ static int add_user_info(const char *wha\n \ttz = strtol(date, NULL, 10);\n \n \tif (fmt == CMIT_FMT_EMAIL) {\n-\t\twhat = \"From\";\n+\t\tchar *name_tail = strchr(line, '<');\n+\t\tint display_name_length;\n+\t\tif (!name_tail)\n+\t\t\treturn 0;\n+\t\twhile (line < name_tail && isspace(name_tail[-1]))\n+\t\t\tname_tail--;\n+\t\tdisplay_name_length = name_tail - line;\n \t\tfiller = \"\";\n+\t\tstrcpy(buf, \"From: \");\n+\t\tret = strlen(buf);\n+\t\tret += add_rfc2047(buf + ret, line, display_name_length);\n+\t\tmemcpy(buf + ret, name_tail, namelen - display_name_length);\n+\t\tret += namelen - display_name_length;\n+\t\tbuf[ret++] = '\\n';\n+\t}\n+\telse {\n+\t\tret = sprintf(buf, \"%s: %.*s%.*s\\n\", what,\n+\t\t\t      (fmt == CMIT_FMT_FULLER) ? 4 : 0,\n+\t\t\t      filler, namelen, line);\n \t}\n-\tret = sprintf(buf, \"%s: %.*s%.*s\\n\", what,\n-\t\t      (fmt == CMIT_FMT_FULLER) ? 4 : 0,\n-\t\t      filler, namelen, line);\n \tswitch (fmt) {\n \tcase CMIT_FMT_MEDIUM:\n \t\tret += sprintf(buf + ret, \"Date:   %s\\n\", show_date(time, tz));\n@@ -575,14 +629,24 @@ unsigned long pretty_print_commit(enum c\n \t\t\tint slen = strlen(subject);\n \t\t\tmemcpy(buf + offset, subject, slen);\n \t\t\toffset += slen;\n+\t\t\toffset += add_rfc2047(buf + offset, line, linelen);\n+\t\t}\n+\t\telse {\n+\t\t\tmemset(buf + offset, ' ', indent);\n+\t\t\tmemcpy(buf + offset + indent, line, linelen);\n+\t\t\toffset += linelen + indent;\n \t\t}\n-\t\tmemset(buf + offset, ' ', indent);\n-\t\tmemcpy(buf + offset + indent, line, linelen);\n-\t\toffset += linelen + indent;\n \t\tbuf[offset++] = '\\n';\n \t\tif (fmt == CMIT_FMT_ONELINE)\n \t\t\tbreak;\n-\t\tsubject = NULL;\n+\t\tif (subject) {\n+\t\t\tstatic const char header[] =\n+\t\t\t\t\"Content-Type: text/plain; charset=UTF-8\\n\"\n+\t\t\t\t\"Content-Transfer-Encoding: 8bit\\n\";\n+\t\t\tmemcpy(buf + offset, header, sizeof(header)-1);\n+\t\t\toffset += sizeof(header)-1;\n+\t\t\tsubject = NULL;\n+\t\t}\n \t}\n \twhile (offset && isspace(buf[offset-1]))\n \t\toffset--;\n"},{"id":"20058","messageId":"e4ca2p$ud5$1@sea.gmane.org","threadId":"4161","inReplyTo":"7vmzdi9ssv.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] CMIT_FMT_EMAIL: Q-encode Subject: and display-name part of From: fields.","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-16T10:38:37Z","receivedAt":"2006-05-16T10:38:37Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> By convention, the commit message and the author/committer names\n> in the commit objects are UTF-8 encoded.  When formatting for\n> e-mails, Q-encode them according to RFC 2047.\n> \n> While we are at it, generate the content-type and\n> content-transfer-encoding headers as well.\n> \n> Signed-off-by: Junio C Hamano <junkio@cox.net>\n> \n> ---\n> \n>  With this patch, the output formatted with\n> \n> git show --pretty=email --patch-with-stat 9d7f73d4\n> \n>  would start like this:\n> \n>    From 9d7f73d43fa49d0d2f5a8cfcce9d659e8ad2d265  Thu Apr 7 15:13:13 2005\n>    From: =?utf-8?q?Lukas_Sandstr=C3=B6m?= <lukass@etek.chalmers.se>\n>    Date: Sat, 25 Feb 2006 12:20:13 +0100\n>    Subject: [PATCH] git-fetch: print the new and old ref when fast-forwarding \n>    Content-Type: text/plain; charset=UTF-8 \n>    Content-Transfer-Encoding: 8bit\n\nI guess that we also need\n     \n     MIME-Version: 1.0\n\n(from what I remember of troubles with Eoutlook Express not sending all \nthe required headers, and tin not working properly).\n\nIf I remember correctly encoding headers using quoted-printable is needed\nonly because headers are before charset is set. IIRC there was proposal\nto use UTF-8 for headers regardless of the charset used for body of message.\n\nP.S. Should we set User-Agent header as well?\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"20059","messageId":"20060516104949.GA9641@klaus.daprodeges.fqdn.th-h.de","threadId":"4161","inReplyTo":"7vmzdi9ssv.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] CMIT_FMT_EMAIL: Q-encode Subject: and display-name part of From: fields.","fromName":"Rocco Rutte","fromEmail":"pdmef@gmx.net","sentAt":"2006-05-16T10:49:50Z","receivedAt":"2006-05-16T10:49:50Z","isPatch":true,"sender":{"key":"pdmef@gmx.net","avatar":null},"body":"Hi,\n\n* Junio C Hamano [06-05-16 03:18:24 -0700] wrote:\n\n[...]\n\n> Thoughts?\n\n> If we decide to do the header formatting here, there are two\n> further enhancements that need to be done:\n\n> (1) The charset must be configurable for projects that use\n>     encoding different from UTF-8, perhaps with the .git/config\n>     [i18n] commitEncoding.  It is only a convention, not a hard\n>     rule, to use UTF-8 for the metainformation.\n\nTo write an encoder really fully conforming to RfC2047 is a mess. Not so\nmuch because the algorithms are difficult but because there're many\nthings to take care of if you want to do it right.\n\nFor example, encoded words are required to be at most something below 80\ncharacters long. For names this maybe is not an issue, but for subjects.\nI didn't really check whether your patch produces only the minimum\nencoding (i.e. only those words that need it and not just all words with\n'_' or '=20' in between them) but if not, 80 isn't that much after all.\nAnd you may need to think about header folding (and unfolding for\nreading it back in).\n\nAlso, supporting any character set (via iconv()) blows up the\nimplementation. There're character sets for which other RfCs define the\nencoding method so only using quoted-printable is not fully correct in\nall possible cases.\n\nAnd, with the first point, several character sets really can become a\nmess as you need to produce several encoded words because the input\nwould exceed RfC limits otherwise. Because for multi-byte character sets\nyou musn't break within a multi-byte character sequence but only at\ntheir boundaries. So you need a generic way to detect the byte-size of\nsuch a character in any supported character set.\n\nWith just the UTF-8 encoding all of this is pretty simple though.\n\nI would rather try to find a way to implement this in a scripting\nlanguage that already has standard modules for this or makes it easy to\nwrite one. In C this gets quite lengthy...\n\n   bye, Rocco\n-- \n:wq!\n"}]}