{"thread":{"id":"31858","subject":"[PATCH v2 0/7] Cure some format-patch wrapping and encoding issues","startedAt":"2012-10-18T14:43:27Z","lastAt":"2012-10-18T14:43:34Z","messageCount":8,"participants":["Jan H. Schönherr"],"isPatch":true,"patchVersion":2,"patchTotal":7},"messages":[{"id":"201514","messageId":"1350571414-17907-1-git-send-email-schnhrr@cs.tu-berlin.de","threadId":"31858","inReplyTo":null,"subject":"[PATCH v2 0/7] Cure some format-patch wrapping and encoding issues","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-10-18T14:43:27Z","receivedAt":"2012-10-18T14:43:27Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"Hi all.\n\n[This is the second version of this series. If you still remember\nthe first version, you might want to jump directly to the summary\nof changes below.]\n\nThe main point of this series is to teach git to encode my name\ncorrectly, see patches 5+6, so that the decoded version is actually\nmy name, so that send-email does not insist on adding a wrong\nsuperfluous From: line to the mail body.\n\nThe other patches more mostly by-products that fix other issues\nI came across.\n\nPatch 1 fixes an old off-by-one error, so that wrapped text may\nnow use all available columns.\n\nPatches 2 and 3 make the wrapping of header lines more correct,\ni. e., neither too early nor too late.\n\nPatch 4 does some refactoring, which is too unrelated to be included\nin one of the later patches.\n\nPatch 5 improves RFC 2047 encoding; patch 6 removes an old non-RFC\nconform workaround.\n\nPatch 7 is more an RFC, which seems to be a good idea from my point\nof view. Indeed, I thought the current implementation is erroneous,\nuntil Junio C Hamano pointed out, that this might be desired behavior.\nThus, make up your mind about this one.\n\n\nThe series is currently based on the maint branch, but it applies\nto master as well. It does also apply to next, but then my\nimplementation of isprint() has to be dropped from patch 5.\n\n\nChanges in v2:\n- patch 1 is new and is a result of the v1 discussion\n- patch 5+6 split the old patch 4 into two patches\n- use of constants for maximum line lengths\n- even better adherence to RFC 2047 than v1\n- updated commit messages/comments\n\n\nRegards\nJan\n\nJan H. Schönherr (7):\n  utf8: fix off-by-one wrapping of text\n  format-patch: do not wrap non-rfc2047 headers too early\n  format-patch: do not wrap rfc2047 encoded headers too late\n  format-patch: introduce helper function last_line_length()\n  format-patch: make rfc2047 encoding more strict\n  format-patch: fix rfc2047 address encoding with respect to rfc822\n    specials\n  format-patch tests: check quoting/encoding in To: and Cc: headers\n\n git-compat-util.h       |   2 +\n pretty.c                | 149 +++++++++++++++++++++++--------\n t/t4014-format-patch.sh | 231 ++++++++++++++++++++++++++++++------------------\n t/t4202-log.sh          |   4 +-\n utf8.c                  |   2 +-\n 5 Dateien geändert, 262 Zeilen hinzugefügt(+), 126 Zeilen entfernt(-)\n\n-- \n1.7.12\n"},{"id":"201515","messageId":"1350571414-17907-2-git-send-email-schnhrr@cs.tu-berlin.de","threadId":"31858","inReplyTo":"1350571414-17907-1-git-send-email-schnhrr@cs.tu-berlin.de","subject":"[PATCH v2 1/7] utf8: fix off-by-one wrapping of text","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-10-18T14:43:28Z","receivedAt":"2012-10-18T14:43:28Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"From: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n\nThe wrapping logic in strbuf_add_wrapped_text() does currently not allow\nlines that entirely fill the allowed width, instead it wraps the line one\ncharacter too early.\n\nFor example, the text \"This is the sixth commit.\" formatted via\n\"%w(11,1,2)\" (wrap at 11 characters, 1 char indent of first line, 2 char\nindent of following lines) results in four lines: \" This is\", \"  the\",\n\"  sixth\", \"  commit.\" This is wrong, because \"  the sixth\" is exactly\n11 characters long, and thus allowed.\n\nFix this by allowing the (width+1) character of a line to be a valid\nwrapping point if it is a whitespace character.\n\nSigned-off-by: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n---\nv2: new patch, result of v1 discussion\n---\n t/t4202-log.sh | 4 ++--\n utf8.c         | 2 +-\n 2 Dateien geändert, 3 Zeilen hinzugefügt(+), 3 Zeilen entfernt(-)\n\ndiff --git a/t/t4202-log.sh b/t/t4202-log.sh\nindex b3ac6be..584e3d8 100755\n--- a/t/t4202-log.sh\n+++ b/t/t4202-log.sh\n@@ -72,9 +72,9 @@ cat > expect << EOF\n   commit.\n EOF\n \n-test_expect_success 'format %w(12,1,2)' '\n+test_expect_success 'format %w(11,1,2)' '\n \n-\tgit log -2 --format=\"%w(12,1,2)This is the %s commit.\" > actual &&\n+\tgit log -2 --format=\"%w(11,1,2)This is the %s commit.\" > actual &&\n \ttest_cmp expect actual\n '\n \ndiff --git a/utf8.c b/utf8.c\nindex a544f15..28791a7 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -353,7 +353,7 @@ retry:\n \n \t\tc = *text;\n \t\tif (!c || isspace(c)) {\n-\t\t\tif (w < width || !space) {\n+\t\t\tif (w <= width || !space) {\n \t\t\t\tconst char *start = bol;\n \t\t\t\tif (!c && text == start)\n \t\t\t\t\treturn w;\n-- \n1.7.12\n"},{"id":"201516","messageId":"1350571414-17907-3-git-send-email-schnhrr@cs.tu-berlin.de","threadId":"31858","inReplyTo":"1350571414-17907-1-git-send-email-schnhrr@cs.tu-berlin.de","subject":"[PATCH v2 2/7] format-patch: do not wrap non-rfc2047 headers too early","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-10-18T14:43:29Z","receivedAt":"2012-10-18T14:43:29Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"From: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n\nDo not wrap the second and later lines of non-rfc2047-encoded headers\nsubstantially before the 78 character limit.\n\nInstead of passing the remaining length of the first line as wrapping\nwidth, use the correct maximum length and tell strbuf_add_wrapped_bytes()\nhow many characters of the first line are already used.\n\nSigned-off-by: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n---\nv2:\n- removed off-by-one correction now handled by first patch\n- commit message clarifications\n---\n pretty.c                |  2 +-\n t/t4014-format-patch.sh | 60 ++++++++++++++++++++++++++++---------------------\n 2 Dateien geändert, 35 Zeilen hinzugefügt(+), 27 Zeilen entfernt(-)\n\ndiff --git a/pretty.c b/pretty.c\nindex 8b1ea9f..71e4024 100644\n--- a/pretty.c\n+++ b/pretty.c\n@@ -286,7 +286,7 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \t\tif ((i + 1 < len) && (ch == '=' && line[i+1] == '?'))\n \t\t\tgoto needquote;\n \t}\n-\tstrbuf_add_wrapped_bytes(sb, line, len, 0, 1, max_length - line_len);\n+\tstrbuf_add_wrapped_bytes(sb, line, len, -line_len, 1, max_length);\n \treturn;\n \n needquote:\ndiff --git a/t/t4014-format-patch.sh b/t/t4014-format-patch.sh\nindex 959aa26..d66e358 100755\n--- a/t/t4014-format-patch.sh\n+++ b/t/t4014-format-patch.sh\n@@ -752,16 +752,14 @@ M64=$M8$M8$M8$M8$M8$M8$M8$M8\n M512=$M64$M64$M64$M64$M64$M64$M64$M64\n cat >expect <<'EOF'\n Subject: [PATCH] foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo\n- bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar\n- foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo\n- bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar\n- foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo\n- bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar\n- foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo\n- bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar\n- foo bar foo bar foo bar foo bar\n+ bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar\n+ foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo\n+ bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar\n+ foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo\n+ bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar\n+ foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar foo bar\n EOF\n-test_expect_success 'format-patch wraps extremely long headers (ascii)' '\n+test_expect_success 'format-patch wraps extremely long subject (ascii)' '\n \techo content >>file &&\n \tgit add file &&\n \tgit commit -m \"$M512\" &&\n@@ -807,28 +805,12 @@ test_expect_success 'format-patch wraps extremely long headers (rfc2047)' '\n \ttest_cmp expect subject\n '\n \n-M8=\"foo_bar_\"\n-M64=$M8$M8$M8$M8$M8$M8$M8$M8\n-cat >expect <<EOF\n-From: $M64\n- <foobar@foo.bar>\n-EOF\n-test_expect_success 'format-patch wraps non-quotable headers' '\n-\trm -rf patches/ &&\n-\techo content >>file &&\n-\tgit add file &&\n-\tgit commit -mfoo --author \"$M64 <foobar@foo.bar>\" &&\n-\tgit format-patch --stdout -1 >patch &&\n-\tsed -n \"/^From: /p; /^ /p; /^$/q\" <patch >from &&\n-\ttest_cmp expect from\n-'\n-\n check_author() {\n \techo content >>file &&\n \tgit add file &&\n \tGIT_AUTHOR_NAME=$1 git commit -m author-check &&\n \tgit format-patch --stdout -1 >patch &&\n-\tgrep ^From: patch >actual &&\n+\tsed -n \"/^From: /p; /^ /p; /^$/q\" <patch >actual &&\n \ttest_cmp expect actual\n }\n \n@@ -853,6 +835,32 @@ test_expect_success 'rfc2047-encoded headers also double-quote 822 specials' '\n \tcheck_author \"Föo B. Bar\"\n '\n \n+cat >expect <<EOF\n+From: foo_bar_foo_bar_foo_bar_foo_bar_foo_bar_foo_bar_foo_bar_foo_bar_\n+ <author@example.com>\n+EOF\n+test_expect_success 'format-patch wraps moderately long from-header (ascii)' '\n+\tcheck_author \"foo_bar_foo_bar_foo_bar_foo_bar_foo_bar_foo_bar_foo_bar_foo_bar_\"\n+'\n+\n+cat >expect <<'EOF'\n+From: Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar\n+ Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo\n+ Bar Foo Bar Foo Bar Foo Bar <author@example.com>\n+EOF\n+test_expect_success 'format-patch wraps extremely long from-header (ascii)' '\n+\tcheck_author \"Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar\"\n+'\n+\n+cat >expect <<'EOF'\n+From: \"Foo.Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar\n+ Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo\n+ Bar Foo Bar Foo Bar Foo Bar\" <author@example.com>\n+EOF\n+test_expect_success 'format-patch wraps extremely long from-header (rfc822)' '\n+\tcheck_author \"Foo.Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar\"\n+'\n+\n cat >expect <<'EOF'\n Subject: header with . in it\n EOF\n-- \n1.7.12\n"},{"id":"201517","messageId":"1350571414-17907-4-git-send-email-schnhrr@cs.tu-berlin.de","threadId":"31858","inReplyTo":"1350571414-17907-1-git-send-email-schnhrr@cs.tu-berlin.de","subject":"[PATCH v2 3/7] format-patch: do not wrap rfc2047 encoded headers too late","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-10-18T14:43:30Z","receivedAt":"2012-10-18T14:43:30Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"From: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n\nEncoded characters add more than one character at once to an encoded\nheader. Include all characters that are about to be added in the length\ncalculation for wrapping.\n\nAdditionally, RFC 2047 imposes a maximum line length of 76 characters\nif that line contains an rfc2047 encoded word.\n\nSigned-off-by: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n---\nv2:\n- use constants for both, the 76 and 78 char limit\n- rephrase comment\n---\n pretty.c                | 26 +++++++++++++---------\n t/t4014-format-patch.sh | 58 +++++++++++++++++++++++++++++--------------------\n 2 Dateien geändert, 51 Zeilen hinzugefügt(+), 33 Zeilen entfernt(-)\n\ndiff --git a/pretty.c b/pretty.c\nindex 71e4024..da75879 100644\n--- a/pretty.c\n+++ b/pretty.c\n@@ -263,6 +263,9 @@ static void add_rfc822_quoted(struct strbuf *out, const char *s, int len)\n \n static int is_rfc2047_special(char ch)\n {\n+\tif (ch == ' ' || ch == '\\n')\n+\t\treturn 1;\n+\n \treturn (non_ascii(ch) || (ch == '=') || (ch == '?') || (ch == '_'));\n }\n \n@@ -270,6 +273,7 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \t\t       const char *encoding)\n {\n \tstatic const int max_length = 78; /* per rfc2822 */\n+\tstatic const int max_encoded_length = 76; /* per rfc2047 */\n \tint i;\n \tint line_len;\n \n@@ -295,23 +299,25 @@ needquote:\n \tline_len += strlen(encoding) + 5; /* 5 for =??q? */\n \tfor (i = 0; i < len; i++) {\n \t\tunsigned ch = line[i] & 0xFF;\n+\t\tint is_special = is_rfc2047_special(ch);\n+\n+\t\t/*\n+\t\t * According to RFC 2047, we could encode the special character\n+\t\t * ' ' (space) with '_' (underscore) for readability. But many\n+\t\t * programs do not understand this and just leave the\n+\t\t * underscore in place. Thus, we do nothing special here, which\n+\t\t * causes ' ' to be encoded as '=20', avoiding this problem.\n+\t\t */\n \n-\t\tif (line_len >= max_length - 2) {\n+\t\tif (line_len + 2 + (is_special ? 3 : 1) > max_encoded_length) {\n \t\t\tstrbuf_addf(sb, \"?=\\n =?%s?q?\", encoding);\n \t\t\tline_len = strlen(encoding) + 5 + 1; /* =??q? plus SP */\n \t\t}\n \n-\t\t/*\n-\t\t * We encode ' ' using '=20' even though rfc2047\n-\t\t * allows using '_' for readability.  Unfortunately,\n-\t\t * many programs do not understand this and just\n-\t\t * leave the underscore in place.\n-\t\t */\n-\t\tif (is_rfc2047_special(ch) || ch == ' ' || ch == '\\n') {\n+\t\tif (is_special) {\n \t\t\tstrbuf_addf(sb, \"=%02X\", ch);\n \t\t\tline_len += 3;\n-\t\t}\n-\t\telse {\n+\t\t} else {\n \t\t\tstrbuf_addch(sb, ch);\n \t\t\tline_len++;\n \t\t}\ndiff --git a/t/t4014-format-patch.sh b/t/t4014-format-patch.sh\nindex d66e358..1d5636d 100755\n--- a/t/t4014-format-patch.sh\n+++ b/t/t4014-format-patch.sh\n@@ -772,30 +772,31 @@ M8=\"föö bar \"\n M64=$M8$M8$M8$M8$M8$M8$M8$M8\n M512=$M64$M64$M64$M64$M64$M64$M64$M64\n cat >expect <<'EOF'\n-Subject: [PATCH] =?UTF-8?q?f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+Subject: [PATCH] =?UTF-8?q?f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n+ =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n+ =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n+ =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n+ =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n+ =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n+ =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n EOF\n-test_expect_success 'format-patch wraps extremely long headers (rfc2047)' '\n+test_expect_success 'format-patch wraps extremely long subject (rfc2047)' '\n \trm -rf patches/ &&\n \techo content >>file &&\n \tgit add file &&\n@@ -862,6 +863,17 @@ test_expect_success 'format-patch wraps extremely long from-header (rfc822)' '\n '\n \n cat >expect <<'EOF'\n+From: =?UTF-8?q?Fo=C3=B6=20Bar=20Foo=20Bar=20Foo=20Bar=20Foo=20Bar=20Foo?=\n+ =?UTF-8?q?=20Bar=20Foo=20Bar=20Foo=20Bar=20Foo=20Bar=20Foo=20Bar=20Foo=20?=\n+ =?UTF-8?q?Bar=20Foo=20Bar=20Foo=20Bar=20Foo=20Bar=20Foo=20Bar=20Foo=20Bar?=\n+ =?UTF-8?q?=20Foo=20Bar=20Foo=20Bar=20Foo=20Bar=20Foo=20Bar=20Foo=20Bar=20?=\n+ =?UTF-8?q?Foo=20Bar=20Foo=20Bar?= <author@example.com>\n+EOF\n+test_expect_success 'format-patch wraps extremely long from-header (rfc2047)' '\n+\tcheck_author \"Foö Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar Foo Bar\"\n+'\n+\n+cat >expect <<'EOF'\n Subject: header with . in it\n EOF\n test_expect_success 'subject lines do not have 822 atom-quoting' '\n-- \n1.7.12\n"},{"id":"201518","messageId":"1350571414-17907-5-git-send-email-schnhrr@cs.tu-berlin.de","threadId":"31858","inReplyTo":"1350571414-17907-1-git-send-email-schnhrr@cs.tu-berlin.de","subject":"[PATCH v2 4/7] format-patch: introduce helper function last_line_length()","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-10-18T14:43:31Z","receivedAt":"2012-10-18T14:43:31Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"From: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n\nCurrently, an open-coded loop to calculate the length of the last\nline of a string buffer is used in multiple places.\n\nMove that code into a function of its own.\n\nSigned-off-by: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n---\n pretty.c | 25 +++++++++++++------------\n 1 Datei geändert, 13 Zeilen hinzugefügt(+), 12 Zeilen entfernt(-)\n\ndiff --git a/pretty.c b/pretty.c\nindex da75879..482402d 100644\n--- a/pretty.c\n+++ b/pretty.c\n@@ -240,6 +240,17 @@ static int has_rfc822_specials(const char *s, int len)\n \treturn 0;\n }\n \n+static int last_line_length(struct strbuf *sb)\n+{\n+\tint i;\n+\n+\t/* How many bytes are already used on the last line? */\n+\tfor (i = sb->len - 1; i >= 0; i--)\n+\t\tif (sb->buf[i] == '\\n')\n+\t\t\tbreak;\n+\treturn sb->len - (i + 1);\n+}\n+\n static void add_rfc822_quoted(struct strbuf *out, const char *s, int len)\n {\n \tint i;\n@@ -275,13 +286,7 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \tstatic const int max_length = 78; /* per rfc2822 */\n \tstatic const int max_encoded_length = 76; /* per rfc2047 */\n \tint i;\n-\tint line_len;\n-\n-\t/* How many bytes are already used on the current line? */\n-\tfor (i = sb->len - 1; i >= 0; i--)\n-\t\tif (sb->buf[i] == '\\n')\n-\t\t\tbreak;\n-\tline_len = sb->len - (i+1);\n+\tint line_len = last_line_length(sb);\n \n \tfor (i = 0; i < len; i++) {\n \t\tint ch = line[i];\n@@ -346,7 +351,6 @@ void pp_user_info(const struct pretty_print_context *pp,\n \tif (pp->fmt == CMIT_FMT_EMAIL) {\n \t\tchar *name_tail = strchr(line, '<');\n \t\tint display_name_length;\n-\t\tint final_line;\n \t\tif (!name_tail)\n \t\t\treturn;\n \t\twhile (line < name_tail && isspace(name_tail[-1]))\n@@ -361,10 +365,7 @@ void pp_user_info(const struct pretty_print_context *pp,\n \t\t\tadd_rfc2047(sb, quoted.buf, quoted.len, encoding);\n \t\t\tstrbuf_release(&quoted);\n \t\t}\n-\t\tfor (final_line = 0; final_line < sb->len; final_line++)\n-\t\t\tif (sb->buf[sb->len - final_line - 1] == '\\n')\n-\t\t\t\tbreak;\n-\t\tif (namelen - display_name_length + final_line > 78) {\n+\t\tif (namelen - display_name_length + last_line_length(sb) > 78) {\n \t\t\tstrbuf_addch(sb, '\\n');\n \t\t\tif (!isspace(name_tail[0]))\n \t\t\t\tstrbuf_addch(sb, ' ');\n-- \n1.7.12\n"},{"id":"201519","messageId":"1350571414-17907-6-git-send-email-schnhrr@cs.tu-berlin.de","threadId":"31858","inReplyTo":"1350571414-17907-1-git-send-email-schnhrr@cs.tu-berlin.de","subject":"[PATCH v2 5/7] format-patch: make rfc2047 encoding more strict","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-10-18T14:43:32Z","receivedAt":"2012-10-18T14:43:32Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"From: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n\nRFC 2047 requires more characters to be encoded than it is currently done.\nEspecially, RFC 2047 distinguishes between allowed remaining characters\nin encoded words in addresses (From, To, etc.) and other headers, such\nas Subject.\n\nMake add_rfc2047() and is_rfc2047_special() location dependent and include\nall non-allowed characters to hopefully be RFC 2047 conform.\n\nThis especially fixes a problem, where RFC 822 specials (e. g. \".\") were\nleft unencoded in addresses, which was solved with a non-standard-conform\nworkaround in the past (which is going to be removed in a follow-up patch).\n\nSigned-off-by: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n---\nv2:\n- part of restructured patch 4 of v1\n- disallow even more characters in is_rfc2047_special()\n\nThe implementation of isprint() should later probably be substituted by\nthe one from Nguyen:\nhttp://article.gmane.org/gmane.comp.version-control.git/207666\n---\n git-compat-util.h       |  2 ++\n pretty.c                | 67 +++++++++++++++++++++++++++++++++++++++++++------\n t/t4014-format-patch.sh | 15 ++++++++---\n 3 Dateien geändert, 72 Zeilen hinzugefügt(+), 12 Zeilen entfernt(-)\n\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 000042d..d4ea446 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -475,6 +475,7 @@ extern const char tolower_trans_tbl[256];\n #undef isdigit\n #undef isalpha\n #undef isalnum\n+#undef isprint\n #undef islower\n #undef isupper\n #undef tolower\n@@ -492,6 +493,7 @@ extern unsigned char sane_ctype[256];\n #define isdigit(x) sane_istest(x,GIT_DIGIT)\n #define isalpha(x) sane_istest(x,GIT_ALPHA)\n #define isalnum(x) sane_istest(x,GIT_ALPHA | GIT_DIGIT)\n+#define isprint(x) ((x) >= 0x20 && (x) <= 0x7e)\n #define islower(x) sane_iscase(x, 1)\n #define isupper(x) sane_iscase(x, 0)\n #define is_glob_special(x) sane_istest(x,GIT_GLOB_SPECIAL)\ndiff --git a/pretty.c b/pretty.c\nindex 482402d..613e4ea 100644\n--- a/pretty.c\n+++ b/pretty.c\n@@ -272,16 +272,65 @@ static void add_rfc822_quoted(struct strbuf *out, const char *s, int len)\n \tstrbuf_addch(out, '\"');\n }\n \n-static int is_rfc2047_special(char ch)\n+enum rfc2047_type {\n+\tRFC2047_SUBJECT,\n+\tRFC2047_ADDRESS,\n+};\n+\n+static int is_rfc2047_special(char ch, enum rfc2047_type type)\n {\n-\tif (ch == ' ' || ch == '\\n')\n+\t/*\n+\t * rfc2047, section 4.2:\n+\t *\n+\t *    8-bit values which correspond to printable ASCII characters other\n+\t *    than \"=\", \"?\", and \"_\" (underscore), MAY be represented as those\n+\t *    characters.  (But see section 5 for restrictions.)  In\n+\t *    particular, SPACE and TAB MUST NOT be represented as themselves\n+\t *    within encoded words.\n+\t */\n+\n+\t/*\n+\t * rule out non-ASCII characters and non-printable characters (the\n+\t * non-ASCII check should be redundant as isprint() is not localized\n+\t * and only knows about ASCII, but be defensive about that)\n+\t */\n+\tif (non_ascii(ch) || !isprint(ch))\n+\t\treturn 1;\n+\n+\t/*\n+\t * rule out special printable characters (' ' should be the only\n+\t * whitespace character considered printable, but be defensive and use\n+\t * isspace())\n+\t */\n+\tif (isspace(ch) || ch == '=' || ch == '?' || ch == '_')\n \t\treturn 1;\n \n-\treturn (non_ascii(ch) || (ch == '=') || (ch == '?') || (ch == '_'));\n+\t/*\n+\t * rfc2047, section 5.3:\n+\t *\n+\t *    As a replacement for a 'word' entity within a 'phrase', for example,\n+\t *    one that precedes an address in a From, To, or Cc header.  The ABNF\n+\t *    definition for 'phrase' from RFC 822 thus becomes:\n+\t *\n+\t *    phrase = 1*( encoded-word / word )\n+\t *\n+\t *    In this case the set of characters that may be used in a \"Q\"-encoded\n+\t *    'encoded-word' is restricted to: <upper and lower case ASCII\n+\t *    letters, decimal digits, \"!\", \"*\", \"+\", \"-\", \"/\", \"=\", and \"_\"\n+\t *    (underscore, ASCII 95.)>.  An 'encoded-word' that appears within a\n+\t *    'phrase' MUST be separated from any adjacent 'word', 'text' or\n+\t *    'special' by 'linear-white-space'.\n+\t */\n+\n+\tif (type != RFC2047_ADDRESS)\n+\t\treturn 0;\n+\n+\t/* '=' and '_' are special cases and have been checked above */\n+\treturn !(isalnum(ch) || ch == '!' || ch == '*' || ch == '+' || ch == '-' || ch == '/');\n }\n \n static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n-\t\t       const char *encoding)\n+\t\t       const char *encoding, enum rfc2047_type type)\n {\n \tstatic const int max_length = 78; /* per rfc2822 */\n \tstatic const int max_encoded_length = 76; /* per rfc2047 */\n@@ -304,7 +353,7 @@ needquote:\n \tline_len += strlen(encoding) + 5; /* 5 for =??q? */\n \tfor (i = 0; i < len; i++) {\n \t\tunsigned ch = line[i] & 0xFF;\n-\t\tint is_special = is_rfc2047_special(ch);\n+\t\tint is_special = is_rfc2047_special(ch, type);\n \n \t\t/*\n \t\t * According to RFC 2047, we could encode the special character\n@@ -358,11 +407,13 @@ void pp_user_info(const struct pretty_print_context *pp,\n \t\tdisplay_name_length = name_tail - line;\n \t\tstrbuf_addstr(sb, \"From: \");\n \t\tif (!has_rfc822_specials(line, display_name_length)) {\n-\t\t\tadd_rfc2047(sb, line, display_name_length, encoding);\n+\t\t\tadd_rfc2047(sb, line, display_name_length,\n+\t\t\t\t\t\tencoding, RFC2047_ADDRESS);\n \t\t} else {\n \t\t\tstruct strbuf quoted = STRBUF_INIT;\n \t\t\tadd_rfc822_quoted(&quoted, line, display_name_length);\n-\t\t\tadd_rfc2047(sb, quoted.buf, quoted.len, encoding);\n+\t\t\tadd_rfc2047(sb, quoted.buf, quoted.len,\n+\t\t\t\t\t\tencoding, RFC2047_ADDRESS);\n \t\t\tstrbuf_release(&quoted);\n \t\t}\n \t\tif (namelen - display_name_length + last_line_length(sb) > 78) {\n@@ -1294,7 +1345,7 @@ void pp_title_line(const struct pretty_print_context *pp,\n \tstrbuf_grow(sb, title.len + 1024);\n \tif (pp->subject) {\n \t\tstrbuf_addstr(sb, pp->subject);\n-\t\tadd_rfc2047(sb, title.buf, title.len, encoding);\n+\t\tadd_rfc2047(sb, title.buf, title.len, encoding, RFC2047_SUBJECT);\n \t} else {\n \t\tstrbuf_addbuf(sb, &title);\n \t}\ndiff --git a/t/t4014-format-patch.sh b/t/t4014-format-patch.sh\nindex 1d5636d..727d606 100755\n--- a/t/t4014-format-patch.sh\n+++ b/t/t4014-format-patch.sh\n@@ -818,21 +818,28 @@ check_author() {\n cat >expect <<'EOF'\n From: \"Foo B. Bar\" <author@example.com>\n EOF\n-test_expect_success 'format-patch quotes dot in headers' '\n+test_expect_success 'format-patch quotes dot in from-headers' '\n \tcheck_author \"Foo B. Bar\"\n '\n \n cat >expect <<'EOF'\n From: \"Foo \\\"The Baz\\\" Bar\" <author@example.com>\n EOF\n-test_expect_success 'format-patch quotes double-quote in headers' '\n+test_expect_success 'format-patch quotes double-quote in from-headers' '\n \tcheck_author \"Foo \\\"The Baz\\\" Bar\"\n '\n \n cat >expect <<'EOF'\n-From: =?UTF-8?q?\"F=C3=B6o=20B.=20Bar\"?= <author@example.com>\n+From: =?UTF-8?q?F=C3=B6o=20Bar?= <author@example.com>\n EOF\n-test_expect_success 'rfc2047-encoded headers also double-quote 822 specials' '\n+test_expect_success 'format-patch uses rfc2047-encoded from-headers when necessary' '\n+\tcheck_author \"Föo Bar\"\n+'\n+\n+cat >expect <<'EOF'\n+From: =?UTF-8?q?F=C3=B6o=20B=2E=20Bar?= <author@example.com>\n+EOF\n+test_expect_failure 'rfc2047-encoded from-headers leave no rfc822 specials' '\n \tcheck_author \"Föo B. Bar\"\n '\n \n-- \n1.7.12\n"},{"id":"201520","messageId":"1350571414-17907-7-git-send-email-schnhrr@cs.tu-berlin.de","threadId":"31858","inReplyTo":"1350571414-17907-1-git-send-email-schnhrr@cs.tu-berlin.de","subject":"[PATCH v2 6/7] format-patch: fix rfc2047 address encoding with respect to rfc822 specials","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-10-18T14:43:33Z","receivedAt":"2012-10-18T14:43:33Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"From: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n\nAccording to RFC 2047 and RFC 822, rfc2047 encoded words and and rfc822\nquoted strings do not mix. Since add_rfc2047() no longer leaves RFC 822\nspecials behind, the quoting is also no longer necessary to create a\nstandard-conform mail.\n\nRemove the quoting, when RFC 2047 encoding takes place. This actually\nrequires to refactor add_rfc2047() a bit, so that the different cases\ncan be distinguished.\n\nWith this patch, my own name gets correctly decoded as Jan H. Schönherr\n(without quotes) and not as \"Jan H. Schönherr\" (with quotes).\n\nSigned-off-by: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n---\nv2:\n- part of restructured patch 4 of v1\n- use constants for both, the 76 and 78 char limit\n- select correct maximum length for possible final folding\n- removed off-by-one correction now handled by first patch\n---\n pretty.c                | 49 ++++++++++++++++++++++++++++++++-----------------\n t/t4014-format-patch.sh |  2 +-\n 2 Dateien geändert, 33 Zeilen hinzugefügt(+), 18 Zeilen entfernt(-)\n\ndiff --git a/pretty.c b/pretty.c\nindex 613e4ea..413e758 100644\n--- a/pretty.c\n+++ b/pretty.c\n@@ -231,7 +231,7 @@ static int is_rfc822_special(char ch)\n \t}\n }\n \n-static int has_rfc822_specials(const char *s, int len)\n+static int needs_rfc822_quoting(const char *s, int len)\n {\n \tint i;\n \tfor (i = 0; i < len; i++)\n@@ -329,25 +329,29 @@ static int is_rfc2047_special(char ch, enum rfc2047_type type)\n \treturn !(isalnum(ch) || ch == '!' || ch == '*' || ch == '+' || ch == '-' || ch == '/');\n }\n \n-static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n-\t\t       const char *encoding, enum rfc2047_type type)\n+static int needs_rfc2047_encoding(const char *line, int len,\n+\t\t\t\t  enum rfc2047_type type)\n {\n-\tstatic const int max_length = 78; /* per rfc2822 */\n-\tstatic const int max_encoded_length = 76; /* per rfc2047 */\n \tint i;\n-\tint line_len = last_line_length(sb);\n \n \tfor (i = 0; i < len; i++) {\n \t\tint ch = line[i];\n \t\tif (non_ascii(ch) || ch == '\\n')\n-\t\t\tgoto needquote;\n+\t\t\treturn 1;\n \t\tif ((i + 1 < len) && (ch == '=' && line[i+1] == '?'))\n-\t\t\tgoto needquote;\n+\t\t\treturn 1;\n \t}\n-\tstrbuf_add_wrapped_bytes(sb, line, len, -line_len, 1, max_length);\n-\treturn;\n \n-needquote:\n+\treturn 0;\n+}\n+\n+static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n+\t\t       const char *encoding, enum rfc2047_type type)\n+{\n+\tstatic const int max_encoded_length = 76; /* per rfc2047 */\n+\tint i;\n+\tint line_len = last_line_length(sb);\n+\n \tstrbuf_grow(sb, len * 3 + strlen(encoding) + 100);\n \tstrbuf_addf(sb, \"=?%s?q?\", encoding);\n \tline_len += strlen(encoding) + 5; /* 5 for =??q? */\n@@ -383,6 +387,7 @@ void pp_user_info(const struct pretty_print_context *pp,\n \t\t  const char *what, struct strbuf *sb,\n \t\t  const char *line, const char *encoding)\n {\n+\tint max_length = 78; /* per rfc2822 */\n \tchar *date;\n \tint namelen;\n \tunsigned long time;\n@@ -406,17 +411,21 @@ void pp_user_info(const struct pretty_print_context *pp,\n \t\t\tname_tail--;\n \t\tdisplay_name_length = name_tail - line;\n \t\tstrbuf_addstr(sb, \"From: \");\n-\t\tif (!has_rfc822_specials(line, display_name_length)) {\n+\t\tif (needs_rfc2047_encoding(line, display_name_length, RFC2047_ADDRESS)) {\n \t\t\tadd_rfc2047(sb, line, display_name_length,\n \t\t\t\t\t\tencoding, RFC2047_ADDRESS);\n-\t\t} else {\n+\t\t\tmax_length = 76; /* per rfc2047 */\n+\t\t} else if (needs_rfc822_quoting(line, display_name_length)) {\n \t\t\tstruct strbuf quoted = STRBUF_INIT;\n \t\t\tadd_rfc822_quoted(&quoted, line, display_name_length);\n-\t\t\tadd_rfc2047(sb, quoted.buf, quoted.len,\n-\t\t\t\t\t\tencoding, RFC2047_ADDRESS);\n+\t\t\tstrbuf_add_wrapped_bytes(sb, quoted.buf, quoted.len,\n+\t\t\t\t\t\t\t-6, 1, max_length);\n \t\t\tstrbuf_release(&quoted);\n+\t\t} else {\n+\t\t\tstrbuf_add_wrapped_bytes(sb, line, display_name_length,\n+\t\t\t\t\t\t\t-6, 1, max_length);\n \t\t}\n-\t\tif (namelen - display_name_length + last_line_length(sb) > 78) {\n+\t\tif (namelen - display_name_length + last_line_length(sb) > max_length) {\n \t\t\tstrbuf_addch(sb, '\\n');\n \t\t\tif (!isspace(name_tail[0]))\n \t\t\t\tstrbuf_addch(sb, ' ');\n@@ -1336,6 +1345,7 @@ void pp_title_line(const struct pretty_print_context *pp,\n \t\t   const char *encoding,\n \t\t   int need_8bit_cte)\n {\n+\tstatic const int max_length = 78; /* per rfc2047 */\n \tstruct strbuf title;\n \n \tstrbuf_init(&title, 80);\n@@ -1345,7 +1355,12 @@ void pp_title_line(const struct pretty_print_context *pp,\n \tstrbuf_grow(sb, title.len + 1024);\n \tif (pp->subject) {\n \t\tstrbuf_addstr(sb, pp->subject);\n-\t\tadd_rfc2047(sb, title.buf, title.len, encoding, RFC2047_SUBJECT);\n+\t\tif (needs_rfc2047_encoding(title.buf, title.len, RFC2047_SUBJECT))\n+\t\t\tadd_rfc2047(sb, title.buf, title.len,\n+\t\t\t\t\t\tencoding, RFC2047_SUBJECT);\n+\t\telse\n+\t\t\tstrbuf_add_wrapped_bytes(sb, title.buf, title.len,\n+\t\t\t\t\t -last_line_length(sb), 1, max_length);\n \t} else {\n \t\tstrbuf_addbuf(sb, &title);\n \t}\ndiff --git a/t/t4014-format-patch.sh b/t/t4014-format-patch.sh\nindex 727d606..e024eb8 100755\n--- a/t/t4014-format-patch.sh\n+++ b/t/t4014-format-patch.sh\n@@ -839,7 +839,7 @@ test_expect_success 'format-patch uses rfc2047-encoded from-headers when necessa\n cat >expect <<'EOF'\n From: =?UTF-8?q?F=C3=B6o=20B=2E=20Bar?= <author@example.com>\n EOF\n-test_expect_failure 'rfc2047-encoded from-headers leave no rfc822 specials' '\n+test_expect_success 'rfc2047-encoded from-headers leave no rfc822 specials' '\n \tcheck_author \"Föo B. Bar\"\n '\n \n-- \n1.7.12\n"},{"id":"201521","messageId":"1350571414-17907-8-git-send-email-schnhrr@cs.tu-berlin.de","threadId":"31858","inReplyTo":"1350571414-17907-1-git-send-email-schnhrr@cs.tu-berlin.de","subject":"[PATCH v2 7/7] format-patch tests: check quoting/encoding in To: and Cc: headers","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-10-18T14:43:34Z","receivedAt":"2012-10-18T14:43:34Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"From: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n\ngit-format-patch does currently not parse user supplied extra header\nvalues (e. g., --cc, --add-header) and just replays them. That forces\nusers to add them RFC 2822/2047 conform in encoded form, e. g.\n\n--cc '=?UTF-8?q?Jan=20H=2E=20Sch=C3=B6nherr?= <...>'\n\nwhich is inconvenient. We would want to update git-format-patch to\naccept human-readable input\n\n--cc 'Jan H. Schönherr <...>'\n\nand handle the encoding, wrapping and quoting internally in the future,\nsimilar to what is already done in git-send-email. The necessary code\nshould mostly exist in the code paths that handle the From: and Subject:\nheaders.\n\nWhether we want to do this only for the git-format-patch options\n--to and --cc (and the corresponding config options) or also for\nuser supplied headers via --add-header, is open for discussion.\n\nFor now, add test_expect_failure tests for To: and Cc: headers as a\nreminder and fix tests that would otherwise fail should this get\nimplemented.\n\nSigned-off-by: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\n---\nThis patch is RFC material. There are a few reasons, why this\nis a good idea and also a few, why it is bad:\n\nPro:\n- current git-format-patch behavior differs from git-send-email\n- we should be able to use the address format that git uses\n  elsewhere (e. g., author and committer info)\n- necessary code mostly exists\n\nCon:\n- changes current behavior\n- make code more complex\n\n(Feel free to add more.)\n\nThe first drawback can be mitigated by checking whether the\ninput is already properly encoded, so that we do not accidentally\ndouble-encode things. git-send-email does that, but that's written\nin Perl, so we would need even more code.\n\nFor now, this is only about _addresses_ supplied to git-format-patch,\nnot _headers_. We could also validate/encode/wrap user supplied headers.\nRFC 2822/2047 is specific enough to allow that. But there is no point\nthinking about that without the intention of encoding addresses.\n\nv2:\n- updated commit message as suggested by Junio C Hamano\n---\n t/t4014-format-patch.sh | 98 +++++++++++++++++++++++++++++++++----------------\n 1 Datei geändert, 66 Zeilen hinzugefügt(+), 32 Zeilen entfernt(-)\n\ndiff --git a/t/t4014-format-patch.sh b/t/t4014-format-patch.sh\nindex e024eb8..ad9f69e 100755\n--- a/t/t4014-format-patch.sh\n+++ b/t/t4014-format-patch.sh\n@@ -110,73 +110,107 @@ test_expect_success 'replay did not screw up the log message' '\n \n test_expect_success 'extra headers' '\n \n-\tgit config format.headers \"To: R. E. Cipient <rcipient@example.com>\n+\tgit config format.headers \"To: R E Cipient <rcipient@example.com>\n \" &&\n-\tgit config --add format.headers \"Cc: S. E. Cipient <scipient@example.com>\n+\tgit config --add format.headers \"Cc: S E Cipient <scipient@example.com>\n \" &&\n \tgit format-patch --stdout master..side > patch2 &&\n \tsed -e \"/^\\$/q\" patch2 > hdrs2 &&\n-\tgrep \"^To: R. E. Cipient <rcipient@example.com>\\$\" hdrs2 &&\n-\tgrep \"^Cc: S. E. Cipient <scipient@example.com>\\$\" hdrs2\n+\tgrep \"^To: R E Cipient <rcipient@example.com>\\$\" hdrs2 &&\n+\tgrep \"^Cc: S E Cipient <scipient@example.com>\\$\" hdrs2\n \n '\n \n test_expect_success 'extra headers without newlines' '\n \n-\tgit config --replace-all format.headers \"To: R. E. Cipient <rcipient@example.com>\" &&\n-\tgit config --add format.headers \"Cc: S. E. Cipient <scipient@example.com>\" &&\n+\tgit config --replace-all format.headers \"To: R E Cipient <rcipient@example.com>\" &&\n+\tgit config --add format.headers \"Cc: S E Cipient <scipient@example.com>\" &&\n \tgit format-patch --stdout master..side >patch3 &&\n \tsed -e \"/^\\$/q\" patch3 > hdrs3 &&\n-\tgrep \"^To: R. E. Cipient <rcipient@example.com>\\$\" hdrs3 &&\n-\tgrep \"^Cc: S. E. Cipient <scipient@example.com>\\$\" hdrs3\n+\tgrep \"^To: R E Cipient <rcipient@example.com>\\$\" hdrs3 &&\n+\tgrep \"^Cc: S E Cipient <scipient@example.com>\\$\" hdrs3\n \n '\n \n test_expect_success 'extra headers with multiple To:s' '\n \n-\tgit config --replace-all format.headers \"To: R. E. Cipient <rcipient@example.com>\" &&\n-\tgit config --add format.headers \"To: S. E. Cipient <scipient@example.com>\" &&\n+\tgit config --replace-all format.headers \"To: R E Cipient <rcipient@example.com>\" &&\n+\tgit config --add format.headers \"To: S E Cipient <scipient@example.com>\" &&\n \tgit format-patch --stdout master..side > patch4 &&\n \tsed -e \"/^\\$/q\" patch4 > hdrs4 &&\n-\tgrep \"^To: R. E. Cipient <rcipient@example.com>,\\$\" hdrs4 &&\n-\tgrep \"^ *S. E. Cipient <scipient@example.com>\\$\" hdrs4\n+\tgrep \"^To: R E Cipient <rcipient@example.com>,\\$\" hdrs4 &&\n+\tgrep \"^ *S E Cipient <scipient@example.com>\\$\" hdrs4\n '\n \n-test_expect_success 'additional command line cc' '\n+test_expect_success 'additional command line cc (ascii)' '\n \n-\tgit config --replace-all format.headers \"Cc: R. E. Cipient <rcipient@example.com>\" &&\n+\tgit config --replace-all format.headers \"Cc: R E Cipient <rcipient@example.com>\" &&\n+\tgit format-patch --cc=\"S E Cipient <scipient@example.com>\" --stdout master..side | sed -e \"/^\\$/q\" >patch5 &&\n+\tgrep \"^Cc: R E Cipient <rcipient@example.com>,\\$\" patch5 &&\n+\tgrep \"^ *S E Cipient <scipient@example.com>\\$\" patch5\n+'\n+\n+test_expect_failure 'additional command line cc (rfc822)' '\n+\n+\tgit config --replace-all format.headers \"Cc: R E Cipient <rcipient@example.com>\" &&\n \tgit format-patch --cc=\"S. E. Cipient <scipient@example.com>\" --stdout master..side | sed -e \"/^\\$/q\" >patch5 &&\n-\tgrep \"^Cc: R. E. Cipient <rcipient@example.com>,\\$\" patch5 &&\n-\tgrep \"^ *S. E. Cipient <scipient@example.com>\\$\" patch5\n+\tgrep \"^Cc: R E Cipient <rcipient@example.com>,\\$\" patch5 &&\n+\tgrep \"^ *\"S. E. Cipient\" <scipient@example.com>\\$\" patch5\n '\n \n test_expect_success 'command line headers' '\n \n \tgit config --unset-all format.headers &&\n-\tgit format-patch --add-header=\"Cc: R. E. Cipient <rcipient@example.com>\" --stdout master..side | sed -e \"/^\\$/q\" >patch6 &&\n-\tgrep \"^Cc: R. E. Cipient <rcipient@example.com>\\$\" patch6\n+\tgit format-patch --add-header=\"Cc: R E Cipient <rcipient@example.com>\" --stdout master..side | sed -e \"/^\\$/q\" >patch6 &&\n+\tgrep \"^Cc: R E Cipient <rcipient@example.com>\\$\" patch6\n '\n \n test_expect_success 'configuration headers and command line headers' '\n \n-\tgit config --replace-all format.headers \"Cc: R. E. Cipient <rcipient@example.com>\" &&\n-\tgit format-patch --add-header=\"Cc: S. E. Cipient <scipient@example.com>\" --stdout master..side | sed -e \"/^\\$/q\" >patch7 &&\n-\tgrep \"^Cc: R. E. Cipient <rcipient@example.com>,\\$\" patch7 &&\n-\tgrep \"^ *S. E. Cipient <scipient@example.com>\\$\" patch7\n+\tgit config --replace-all format.headers \"Cc: R E Cipient <rcipient@example.com>\" &&\n+\tgit format-patch --add-header=\"Cc: S E Cipient <scipient@example.com>\" --stdout master..side | sed -e \"/^\\$/q\" >patch7 &&\n+\tgrep \"^Cc: R E Cipient <rcipient@example.com>,\\$\" patch7 &&\n+\tgrep \"^ *S E Cipient <scipient@example.com>\\$\" patch7\n '\n \n-test_expect_success 'command line To: header' '\n+test_expect_success 'command line To: header (ascii)' '\n \n \tgit config --unset-all format.headers &&\n+\tgit format-patch --to=\"R E Cipient <rcipient@example.com>\" --stdout master..side | sed -e \"/^\\$/q\" >patch8 &&\n+\tgrep \"^To: R E Cipient <rcipient@example.com>\\$\" patch8\n+'\n+\n+test_expect_failure 'command line To: header (rfc822)' '\n+\n \tgit format-patch --to=\"R. E. Cipient <rcipient@example.com>\" --stdout master..side | sed -e \"/^\\$/q\" >patch8 &&\n-\tgrep \"^To: R. E. Cipient <rcipient@example.com>\\$\" patch8\n+\tgrep \"^To: \"R. E. Cipient\" <rcipient@example.com>\\$\" patch8\n+'\n+\n+test_expect_failure 'command line To: header (rfc2047)' '\n+\n+\tgit format-patch --to=\"R Ä Cipient <rcipient@example.com>\" --stdout master..side | sed -e \"/^\\$/q\" >patch8 &&\n+\tgrep \"^To: =?UTF-8?q?R=20=C3=84=20Cipient?= <rcipient@example.com>\\$\" patch8\n '\n \n-test_expect_success 'configuration To: header' '\n+test_expect_success 'configuration To: header (ascii)' '\n+\n+\tgit config format.to \"R E Cipient <rcipient@example.com>\" &&\n+\tgit format-patch --stdout master..side | sed -e \"/^\\$/q\" >patch9 &&\n+\tgrep \"^To: R E Cipient <rcipient@example.com>\\$\" patch9\n+'\n+\n+test_expect_failure 'configuration To: header (rfc822)' '\n \n \tgit config format.to \"R. E. Cipient <rcipient@example.com>\" &&\n \tgit format-patch --stdout master..side | sed -e \"/^\\$/q\" >patch9 &&\n-\tgrep \"^To: R. E. Cipient <rcipient@example.com>\\$\" patch9\n+\tgrep \"^To: \"R. E. Cipient\" <rcipient@example.com>\\$\" patch9\n+'\n+\n+test_expect_failure 'configuration To: header (rfc2047)' '\n+\n+\tgit config format.to \"R Ä Cipient <rcipient@example.com>\" &&\n+\tgit format-patch --stdout master..side | sed -e \"/^\\$/q\" >patch9 &&\n+\tgrep \"^To: =?UTF-8?q?R=20=C3=84=20Cipient?= <rcipient@example.com>\\$\" patch9\n '\n \n # check_patch <patch>: Verify that <patch> looks like a half-sane\n@@ -190,11 +224,11 @@ check_patch () {\n test_expect_success '--no-to overrides config.to' '\n \n \tgit config --replace-all format.to \\\n-\t\t\"R. E. Cipient <rcipient@example.com>\" &&\n+\t\t\"R E Cipient <rcipient@example.com>\" &&\n \tgit format-patch --no-to --stdout master..side |\n \tsed -e \"/^\\$/q\" >patch10 &&\n \tcheck_patch patch10 &&\n-\t! grep \"^To: R. E. Cipient <rcipient@example.com>\\$\" patch10\n+\t! grep \"^To: R E Cipient <rcipient@example.com>\\$\" patch10\n '\n \n test_expect_success '--no-to and --to replaces config.to' '\n@@ -212,21 +246,21 @@ test_expect_success '--no-to and --to replaces config.to' '\n test_expect_success '--no-cc overrides config.cc' '\n \n \tgit config --replace-all format.cc \\\n-\t\t\"C. E. Cipient <rcipient@example.com>\" &&\n+\t\t\"C E Cipient <rcipient@example.com>\" &&\n \tgit format-patch --no-cc --stdout master..side |\n \tsed -e \"/^\\$/q\" >patch12 &&\n \tcheck_patch patch12 &&\n-\t! grep \"^Cc: C. E. Cipient <rcipient@example.com>\\$\" patch12\n+\t! grep \"^Cc: C E Cipient <rcipient@example.com>\\$\" patch12\n '\n \n test_expect_success '--no-add-header overrides config.headers' '\n \n \tgit config --replace-all format.headers \\\n-\t\t\"Header1: B. E. Cipient <rcipient@example.com>\" &&\n+\t\t\"Header1: B E Cipient <rcipient@example.com>\" &&\n \tgit format-patch --no-add-header --stdout master..side |\n \tsed -e \"/^\\$/q\" >patch13 &&\n \tcheck_patch patch13 &&\n-\t! grep \"^Header1: B. E. Cipient <rcipient@example.com>\\$\" patch13\n+\t! grep \"^Header1: B E Cipient <rcipient@example.com>\\$\" patch13\n '\n \n test_expect_success 'multiple files' '\n-- \n1.7.12\n"}]}