{"thread":{"id":"17040","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","startedAt":"2008-12-26T18:38:41Z","lastAt":"2009-01-14T08:19:42Z","messageCount":12,"participants":["Kirill Smelkov","Junio C Hamano","Alexander Potashev"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"299215","messageId":"1230316721-14339-1-git-send-email-kirr@mns.spb.ru","threadId":"17040","inReplyTo":null,"subject":"[PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2008-12-26T18:38:41Z","receivedAt":"2008-12-26T18:38:41Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"When native language (RU) is in use, subject header usually contains several\nparts, e.g.\n\nSubject: [Navy-patches] [PATCH]\n\t=?utf-8?b?0JjQt9C80LXQvdGR0L0g0YHQv9C40YHQvtC6INC/0LA=?=\n\t=?utf-8?b?0LrQtdGC0L7QsiDQvdC10L7QsdGF0L7QtNC40LzRi9GFINC00LvRjyA=?=\n\t=?utf-8?b?0YHQsdC+0YDQutC4?=\n\nThis exposes several bugs in builtin-mailinfo.c that I try to fix:\n\n\n1. decode_b_segment: do not append explicit NUL -- explicit NUL was preventing\n   correct header construction on parts concatenation via strbuf_addbuf in\n   decode_header_bq. Fixes:\n\n-Subject: Изменён список пакетов необходимых для сборки\n+Subject: Изменён список па\n\n\nThen\n\n2. (hackish) do not emit '\\n' after processing of every header segment. It\n   seems we should emit previous part as-is only if it does not end with\n   '=?='. Fixes:\n\n-Subject: Изменён список пакетов необходимых для сборки\n+Subject: Изменён список па кетов необходимых для сборки\n\n\nSorry for low-quality patch and description. I did what I could and don't have\nenergy and time dig more into MIME.\n\nPlease help.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n\n---\n builtin-mailinfo.c  |   18 ++++++++++++++++-\n t/t5100-mailinfo.sh |    2 +-\n t/t5100/info0012    |    5 ++++\n t/t5100/msg0012     |    7 ++++++\n t/t5100/patch0012   |   30 +++++++++++++++++++++++++++++\n t/t5100/sample.mbox |   52 +++++++++++++++++++++++++++++++++++++++++++++++++++\n 6 files changed, 112 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex e890f7a..d138bc3 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -436,6 +436,14 @@ static struct strbuf *decode_b_segment(const struct strbuf *b_seg)\n \t\t\t * for now we just trust the data.\n \t\t\t */\n \t\t\tc = 0;\n+\n+\t\t\t/* XXX: the following is needed not to output NUL in\n+\t\t\t * the resulting string\n+\t\t\t *\n+\t\t\t * This seems to be ok, but I'm not 100% sure -- that's\n+\t\t\t * why this is an RFC.\n+\t\t\t */\n+\t\t\tcontinue;\n \t\t}\n \t\telse\n \t\t\tcontinue; /* garbage */\n@@ -513,7 +521,15 @@ static int decode_header_bq(struct strbuf *it)\n \t\tstrbuf_reset(&piecebuf);\n \t\trfc2047 = 1;\n \n-\t\tif (in != ep) {\n+\t\t/* XXX: the follwoing is needed not to output '\\n' on every\n+\t\t * multi-line segment in Subject.\n+\t\t *\n+\t\t * I suspect this is not 100% correct, but I'm not a MIME guy\n+\t\t * -- that's why this is an RFC.\n+\t\t */\n+\n+\t\t/* if in does not end with '=?=', we emit it as is */\n+\t\tif (in <= (ep-2) && !(ep[-1]=='\\n' && ep[-2]=='=')) {\n \t\t\tstrbuf_add(&outbuf, in, ep - in);\n \t\t\tin = ep;\n \t\t}\ndiff --git a/t/t5100-mailinfo.sh b/t/t5100-mailinfo.sh\nindex fe14589..6825f99 100755\n--- a/t/t5100-mailinfo.sh\n+++ b/t/t5100-mailinfo.sh\n@@ -11,7 +11,7 @@ test_expect_success 'split sample box' \\\n \t'git mailsplit -o. \"$TEST_DIRECTORY\"/t5100/sample.mbox >last &&\n \tlast=`cat last` &&\n \techo total is $last &&\n-\ttest `cat last` = 11'\n+\ttest `cat last` = 12'\n \n for mail in `echo 00*`\n do\ndiff --git a/t/t5100/info0012 b/t/t5100/info0012\nnew file mode 100644\nindex 0000000..ac1216f\n--- /dev/null\n+++ b/t/t5100/info0012\n@@ -0,0 +1,5 @@\n+Author: Dmitriy Blinov\n+Email: bda@mnsspb.ru\n+Subject: Изменён список пакетов необходимых для сборки\n+Date: Wed, 12 Nov 2008 17:54:41 +0300\n+\ndiff --git a/t/t5100/msg0012 b/t/t5100/msg0012\nnew file mode 100644\nindex 0000000..1dc2bf7\n--- /dev/null\n+++ b/t/t5100/msg0012\n@@ -0,0 +1,7 @@\n+textlive-* исправлены на texlive-*\n+docutils заменён на python-docutils\n+\n+Действительно, оказалось, что rest2web вытягивает за собой\n+python-docutils. В то время как сам rest2web не нужен.\n+\n+Signed-off-by: Dmitriy Blinov <bda@mnsspb.ru>\ndiff --git a/t/t5100/patch0012 b/t/t5100/patch0012\nnew file mode 100644\nindex 0000000..36a0b68\n--- /dev/null\n+++ b/t/t5100/patch0012\n@@ -0,0 +1,30 @@\n+---\n+ howto/build_navy.txt |    6 +++---\n+ 1 files changed, 3 insertions(+), 3 deletions(-)\n+\n+diff --git a/howto/build_navy.txt b/howto/build_navy.txt\n+index 3fd3afb..0ee807e 100644\n+--- a/howto/build_navy.txt\n++++ b/howto/build_navy.txt\n+@@ -119,8 +119,8 @@\n+    - libxv-dev\n+    - libusplash-dev\n+    - latex-make\n+-   - textlive-lang-cyrillic\n+-   - textlive-latex-extra\n++   - texlive-lang-cyrillic\n++   - texlive-latex-extra\n+    - dia\n+    - python-pyrex\n+    - libtool\n+@@ -128,7 +128,7 @@\n+    - sox\n+    - cython\n+    - imagemagick\n+-   - docutils\n++   - python-docutils\n+ \n+ #. на машине dinar: добавить свой открытый ssh-ключ в authorized_keys2 пользователя ddev\n+ #. на своей машине: отредактировать /etc/sudoers (команда ``visudo``) примерно следующим образом::\n+-- \n+1.5.6.5\ndiff --git a/t/t5100/sample.mbox b/t/t5100/sample.mbox\nindex 4bf7947..94da4da 100644\n--- a/t/t5100/sample.mbox\n+++ b/t/t5100/sample.mbox\n@@ -501,3 +501,55 @@ index 3e5fe51..aabfe5c 100644\n \n --=-=-=--\n \n+From bda@mnsspb.ru Wed Nov 12 17:54:41 2008\n+From: Dmitriy Blinov <bda@mnsspb.ru>\n+To: navy-patches@dinar.mns.mnsspb.ru\n+Date: Wed, 12 Nov 2008 17:54:41 +0300\n+Message-Id: <1226501681-24923-1-git-send-email-bda@mnsspb.ru>\n+X-Mailer: git-send-email 1.5.6.5\n+MIME-Version: 1.0\n+Content-Type: text/plain;\n+  charset=utf-8\n+Content-Transfer-Encoding: 8bit\n+Subject: [Navy-patches] [PATCH]\n+\t=?utf-8?b?0JjQt9C80LXQvdGR0L0g0YHQv9C40YHQvtC6INC/0LA=?=\n+\t=?utf-8?b?0LrQtdGC0L7QsiDQvdC10L7QsdGF0L7QtNC40LzRi9GFINC00LvRjyA=?=\n+\t=?utf-8?b?0YHQsdC+0YDQutC4?=\n+\n+textlive-* исправлены на texlive-*\n+docutils заменён на python-docutils\n+\n+Действительно, оказалось, что rest2web вытягивает за собой\n+python-docutils. В то время как сам rest2web не нужен.\n+\n+Signed-off-by: Dmitriy Blinov <bda@mnsspb.ru>\n+---\n+ howto/build_navy.txt |    6 +++---\n+ 1 files changed, 3 insertions(+), 3 deletions(-)\n+\n+diff --git a/howto/build_navy.txt b/howto/build_navy.txt\n+index 3fd3afb..0ee807e 100644\n+--- a/howto/build_navy.txt\n++++ b/howto/build_navy.txt\n+@@ -119,8 +119,8 @@\n+    - libxv-dev\n+    - libusplash-dev\n+    - latex-make\n+-   - textlive-lang-cyrillic\n+-   - textlive-latex-extra\n++   - texlive-lang-cyrillic\n++   - texlive-latex-extra\n+    - dia\n+    - python-pyrex\n+    - libtool\n+@@ -128,7 +128,7 @@\n+    - sox\n+    - cython\n+    - imagemagick\n+-   - docutils\n++   - python-docutils\n+ \n+ #. на машине dinar: добавить свой открытый ssh-ключ в authorized_keys2 пользователя ddev\n+ #. на своей машине: отредактировать /etc/sudoers (команда ``visudo``) примерно следующим образом::\n+-- \n+1.5.6.5\n-- \ntg: (2292ebd..) t/mailinfo-multiline-subject (depends on: tmp)\n"},{"id":"99624","messageId":"20090107224342.GB4946@roro3","threadId":"17040","inReplyTo":"1230316721-14339-1-git-send-email-kirr@mns.spb.ru","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Kirill Smelkov","fromEmail":"kirr@landau.phys.spbu.ru","sentAt":"2009-01-07T22:43:42Z","receivedAt":"2009-01-07T22:43:42Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Fri, Dec 26, 2008 at 09:38:41PM +0300, Kirill Smelkov wrote:\n> When native language (RU) is in use, subject header usually contains several\n> parts, e.g.\n> \n> Subject: [Navy-patches] [PATCH]\n> \t=?utf-8?b?0JjQt9C80LXQvdGR0L0g0YHQv9C40YHQvtC6INC/0LA=?=\n> \t=?utf-8?b?0LrQtdGC0L7QsiDQvdC10L7QsdGF0L7QtNC40LzRi9GFINC00LvRjyA=?=\n> \t=?utf-8?b?0YHQsdC+0YDQutC4?=\n\nWhich btw should be extracted by git-mailinfo to:\n\n    'Subject: Изменён список пакетов необходимых для сборки'\n\n> This exposes several bugs in builtin-mailinfo.c that I try to fix:\n> \n> \n> 1. decode_b_segment: do not append explicit NUL -- explicit NUL was preventing\n>    correct header construction on parts concatenation via strbuf_addbuf in\n>    decode_header_bq. Fixes:\n> \n> -Subject: Изменён список пакетов необходимых для сборки\n> +Subject: Изменён список па\n> \n> \n> Then\n> \n> 2. (hackish) do not emit '\\n' after processing of every header segment. It\n>    seems we should emit previous part as-is only if it does not end with\n>    '=?='. Fixes:\n> \n> -Subject: Изменён список пакетов необходимых для сборки\n> +Subject: Изменён список па кетов необходимых для сборки\n> \n> \n> Sorry for low-quality patch and description. I did what I could and don't have\n> energy and time dig more into MIME.\n> \n> Please help.\n> \n> Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n> \n> ---\n>  builtin-mailinfo.c  |   18 ++++++++++++++++-\n>  t/t5100-mailinfo.sh |    2 +-\n>  t/t5100/info0012    |    5 ++++\n>  t/t5100/msg0012     |    7 ++++++\n>  t/t5100/patch0012   |   30 +++++++++++++++++++++++++++++\n>  t/t5100/sample.mbox |   52 +++++++++++++++++++++++++++++++++++++++++++++++++++\n>  6 files changed, 112 insertions(+), 2 deletions(-)\n\nJunio, All,\n\nWhat about this patch?\n\nIt at least exposes bug in git-mailinfo wrt handling of multiline\nsubjects, and in very details documents it and adds a test for it.\n\n\nYes, my fixes are of 'low quality', but may I try to attract git\ncommunity attention one more time?\n\n\nThanks beforehand,\nKirill\n\n\nP.S. original post with patch:\n\nhttp://marc.info/?l=git&m=123031899307286&w=2\n"},{"id":"99680","messageId":"7vy6xm5i6h.fsf@gitster.siamese.dyndns.org","threadId":"17040","inReplyTo":"20090107224342.GB4946@roro3","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-08T08:13:42Z","receivedAt":"2009-01-08T08:13:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@landau.phys.spbu.ru> writes:\n\n> On Fri, Dec 26, 2008 at 09:38:41PM +0300, Kirill Smelkov wrote:\n>> When native language (RU) is in use, subject header usually contains several\n>> parts, e.g.\n> ...\n> Junio, All,\n>\n> What about this patch?\n\nWhat's most interesting is that I do not recall seeing this patch before.\nNeither gmane (which is my back-up interface to the mailing list) nor my\nmailbox seems to have a copy, and from the look of quoted parts (namely,\nsome Russian strings in the message), it is not implausible that my spam\nfilter (either on my receiving end or at the ISP) may have eaten it.\n\n> It at least exposes bug in git-mailinfo wrt handling of multiline\n> subjects, and in very details documents it and adds a test for it.\n>\n> ..., but may I try to attract git\n> community attention one more time?\n\nIt is very appreciated.\n\n> P.S. original post with patch:\n>\n> http://marc.info/?l=git&m=123031899307286&w=2\n\nI have not had chance to look at your patch at marc yet, but from the look\nof your problem description, I presume you could trigger this with any\nutf-8 b-encoded loooooong subject line?\n"},{"id":"99681","messageId":"7vy6xm42l3.fsf@gitster.siamese.dyndns.org","threadId":"17040","inReplyTo":"7vy6xm5i6h.fsf@gitster.siamese.dyndns.org","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-08T08:35:52Z","receivedAt":"2009-01-08T08:35:52Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Kirill Smelkov <kirr@landau.phys.spbu.ru> writes:\n> ...\n>> http://marc.info/?l=git&m=123031899307286&w=2\n>\n> I have not had chance to look at your patch at marc yet, but from the look\n> of your problem description, I presume you could trigger this with any\n> utf-8 b-encoded loooooong subject line?\n\nOk, I took a look at it after downloading from the marc archive.\n\n> diff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\n> index e890f7a..d138bc3 100644\n> --- a/builtin-mailinfo.c\n> +++ b/builtin-mailinfo.c\n> @@ -436,6 +436,14 @@ static struct strbuf *decode_b_segment(const struct strbuf *b_seg)\n>  \t\t\t * for now we just trust the data.\n>  \t\t\t */\n>  \t\t\tc = 0;\n> +\n> +\t\t\t/* XXX: the following is needed not to output NUL in\n> +\t\t\t * the resulting string\n> +\t\t\t *\n> +\t\t\t * This seems to be ok, but I'm not 100% sure -- that's\n> +\t\t\t * why this is an RFC.\n> +\t\t\t */\n> +\t\t\tcontinue;\n>  \t\t}\n>  \t\telse\n>  \t\t\tcontinue; /* garbage */\n\nB encoding (RFC 2045) encodes an octet stream into a sequence of groups of\n4 letters from 64-char alphabet, each of which encodes 6-bit, plus zero or\nmore padding char '=' to make the result multiple of 4.\n\n * If the length of the payload is a multiple of 3 octets, there is no\n   special handling.  Padding char '=' is not produced;\n\n * If it is a multiple of 3 octets plus one, the remaining one octet is\n   encoded with two letters, and two more padding char '=' is added;\n\n * If it is a multiple of 3 octets plus two, the remaining two octets are\n   encoded with three letters, and one padding char '=' is added.\n\nHence, a \"correct\" implementation should decode the input as if '=' were\nthe same as 'A' (which encodes 6 bits of 0) til the end, making sure that\nthe padding char '=' appears only at the end of the input, that no char\noutside the Base64 encoding alphabet appears in the input, and that the\nlength of the entire encoded string is multiple of 4.  Finally it would\ndiscard either one or two octets (depending on the number of padding chars\nit saw) from the end of the output.\n\nOur decode_b_segment() however emits each octet as it completes, without\nwaiting for the 24-bit group that contains it to complete.  When decoding\na correctly encoded input, by the time we see a padding '=', all the real\npayload octets are complete and we would not have any real information\nstill kept in the variable \"acc\" (accumulator), so ignoring '=' (you do\nnot even need to assign c = 0) like your patch did would work just fine.\nAn alternative would be to count the number of padding at the end and drop\nthe NULs from the output as necessary after the loop but that does not add\nany value to the current code.\n\nIdeally we should validate the encoded string a bit more carefully (see\nthe \"correct\" implementation about), and warn if a malformed input is\nfound (but probably not reject outright).  But as a low-impact fix for the\nmaintenance branches, I think your fix is very good.\n\n\tSide note: I suspect that the existing code was Ok before strbuf\n\tconversion as we assumed NUL terminated output buffer.\n\n> @@ -513,7 +521,15 @@ static int decode_header_bq(struct strbuf *it)\n>  \t\tstrbuf_reset(&piecebuf);\n>  \t\trfc2047 = 1;\n>  \n> -\t\tif (in != ep) {\n> +\t\t/* XXX: the follwoing is needed not to output '\\n' on every\n> +\t\t * multi-line segment in Subject.\n> +\t\t *\n> +\t\t * I suspect this is not 100% correct, but I'm not a MIME guy\n> +\t\t * -- that's why this is an RFC.\n> +\t\t */\n> +\n> +\t\t/* if in does not end with '=?=', we emit it as is */\n> +\t\tif (in <= (ep-2) && !(ep[-1]=='\\n' && ep[-2]=='=')) {\n>  \t\t\tstrbuf_add(&outbuf, in, ep - in);\n>  \t\t\tin = ep;\n> \n>  \t\t}\n\nI am not a MIME guy either (and mailinfo has a big comment that says we do\nnot really do MIME --- we just pretend to do), but let me give it a try.\n\nRFC2046 specifies that an encoded-word (\"=?charset?encoding?...?=\") may\nnot be more than 75 characters long, and multiple encoded-words, separated\nby CRLF SPACE can be used to encode more text if needed.\n\nIt further specifies that an encoded-word can appear next to ordinary text\nor another encoded-word but it must be separated by linear white space,\nand says that such linear white space is to be ignored when displaying.\n\nWhich means that we should be eating the CRLF SPACE we see if we have seen\nan encoded-word immediately before and we are about to process another\nencoded-word.\n\nBased on the above discussion, here is what I came up with.  It passes\nyour test, but I ran out of energy to try breaking it seriously in any\nother way than just running the existing test suite.  \n\nWe might want to steal some test cases from the \"8. Examples\" section of\nRFC2047 and add them to t5100.\n\nThanks.\n\n builtin-mailinfo.c |   27 +++++++++++++++++++--------\n 1 files changed, 19 insertions(+), 8 deletions(-)\n\ndiff --git c/builtin-mailinfo.c w/builtin-mailinfo.c\nindex e890f7a..fcb32c9 100644\n--- c/builtin-mailinfo.c\n+++ w/builtin-mailinfo.c\n@@ -430,13 +430,6 @@ static struct strbuf *decode_b_segment(const struct strbuf *b_seg)\n \t\t\tc -= 'a' - 26;\n \t\telse if ('0' <= c && c <= '9')\n \t\t\tc -= '0' - 52;\n-\t\telse if (c == '=') {\n-\t\t\t/* padding is almost like (c == 0), except we do\n-\t\t\t * not output NUL resulting only from it;\n-\t\t\t * for now we just trust the data.\n-\t\t\t */\n-\t\t\tc = 0;\n-\t\t}\n \t\telse\n \t\t\tcontinue; /* garbage */\n \t\tswitch (pos++) {\n@@ -514,7 +507,25 @@ static int decode_header_bq(struct strbuf *it)\n \t\trfc2047 = 1;\n \n \t\tif (in != ep) {\n-\t\t\tstrbuf_add(&outbuf, in, ep - in);\n+\t\t\t/*\n+\t\t\t * We are about to process an encoded-word\n+\t\t\t * that begins at ep, but there is something\n+\t\t\t * before the encoded word.\n+\t\t\t */\n+\t\t\tchar *scan;\n+\t\t\tfor (scan = in; scan < ep; scan++)\n+\t\t\t\tif (!isspace(*scan))\n+\t\t\t\t\tbreak;\n+\n+\t\t\tif (scan != ep || in == it->buf) {\n+\t\t\t\t/*\n+\t\t\t\t * We should not lose that \"something\",\n+\t\t\t\t * unless we have just processed an\n+\t\t\t\t * encoded-word, and there is only LWS\n+\t\t\t\t * before the one we are about to process.\n+\t\t\t\t */\n+\t\t\t\tstrbuf_add(&outbuf, in, ep - in);\n+\t\t\t}\n \t\t\tin = ep;\n \t\t}\n \t\t/* E.g.\n"},{"id":"99693","messageId":"20090108100813.GA15640@myhost","threadId":"17040","inReplyTo":"1230316721-14339-1-git-send-email-kirr@mns.spb.ru","subject":"Re: [PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Alexander Potashev","fromEmail":"aspotashev@gmail.com","sentAt":"2009-01-08T10:08:13Z","receivedAt":"2009-01-08T10:08:13Z","isPatch":true,"sender":{"key":"aspotashev@gmail.com","avatar":null},"body":"On 21:38 Fri 26 Dec     , Kirill Smelkov wrote:\n> When native language (RU) is in use, subject header usually contains several\n> parts, e.g.\n> \n> Subject: [Navy-patches] [PATCH]\n> \t=?utf-8?b?0JjQt9C80LXQvdGR0L0g0YHQv9C40YHQvtC6INC/0LA=?=\n> \t=?utf-8?b?0LrQtdGC0L7QsiDQvdC10L7QsdGF0L7QtNC40LzRi9GFINC00LvRjyA=?=\n> \t=?utf-8?b?0YHQsdC+0YDQutC4?=\n> \n\n>  t/t5100/info0012    |    5 ++++\n>  t/t5100/msg0012     |    7 ++++++\n>  t/t5100/patch0012   |   30 +++++++++++++++++++++++++++++\n>  t/t5100/sample.mbox |   52 +++++++++++++++++++++++++++++++++++++++++++++++++++\n>  6 files changed, 112 insertions(+), 2 deletions(-)\n\nThe testcases are too long, a minimal mbox with encoded \"Subject:\" would\nbe enough to test the mailinfo parser, it's all the you need to test\nhere.\n"},{"id":"99744","messageId":"20090108231135.GB4185@roro3","threadId":"17040","inReplyTo":"20090108100813.GA15640@myhost","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Kirill Smelkov","fromEmail":"kirr@landau.phys.spbu.ru","sentAt":"2009-01-08T23:11:35Z","receivedAt":"2009-01-08T23:11:35Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Thu, Jan 08, 2009 at 12:13:42AM -0800, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@landau.phys.spbu.ru> writes:\n> \n> > On Fri, Dec 26, 2008 at 09:38:41PM +0300, Kirill Smelkov wrote:\n> >> When native language (RU) is in use, subject header usually contains several\n> >> parts, e.g.\n> > ...\n> > Junio, All,\n> >\n> > What about this patch?\n> \n> What's most interesting is that I do not recall seeing this patch before.\n> Neither gmane (which is my back-up interface to the mailing list) nor my\n> mailbox seems to have a copy, and from the look of quoted parts (namely,\n> some Russian strings in the message), it is not implausible that my spam\n> filter (either on my receiving end or at the ISP) may have eaten it.\n> \n> > It at least exposes bug in git-mailinfo wrt handling of multiline\n> > subjects, and in very details documents it and adds a test for it.\n> >\n> > ..., but may I try to attract git\n> > community attention one more time?\n> \n> It is very appreciated.\n\nThanks!\n\n\nOn Thu, Jan 08, 2009 at 12:35:52AM -0800, Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n> > Kirill Smelkov <kirr@landau.phys.spbu.ru> writes:\n> > ...\n> >> http://marc.info/?l=git&m=123031899307286&w=2\n> >\n> > I have not had chance to look at your patch at marc yet, but from the look\n> > of your problem description, I presume you could trigger this with any\n> > utf-8 b-encoded loooooong subject line?\n> \n> Ok, I took a look at it after downloading from the marc archive.\n> \n> > diff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\n> > index e890f7a..d138bc3 100644\n> > --- a/builtin-mailinfo.c\n> > +++ b/builtin-mailinfo.c\n> > @@ -436,6 +436,14 @@ static struct strbuf *decode_b_segment(const struct strbuf *b_seg)\n> >  \t\t\t * for now we just trust the data.\n> >  \t\t\t */\n> >  \t\t\tc = 0;\n> > +\n> > +\t\t\t/* XXX: the following is needed not to output NUL in\n> > +\t\t\t * the resulting string\n> > +\t\t\t *\n> > +\t\t\t * This seems to be ok, but I'm not 100% sure -- that's\n> > +\t\t\t * why this is an RFC.\n> > +\t\t\t */\n> > +\t\t\tcontinue;\n> >  \t\t}\n> >  \t\telse\n> >  \t\t\tcontinue; /* garbage */\n> \n> B encoding (RFC 2045) encodes an octet stream into a sequence of groups of\n> 4 letters from 64-char alphabet, each of which encodes 6-bit, plus zero or\n> more padding char '=' to make the result multiple of 4.\n> \n>  * If the length of the payload is a multiple of 3 octets, there is no\n>    special handling.  Padding char '=' is not produced;\n> \n>  * If it is a multiple of 3 octets plus one, the remaining one octet is\n>    encoded with two letters, and two more padding char '=' is added;\n> \n>  * If it is a multiple of 3 octets plus two, the remaining two octets are\n>    encoded with three letters, and one padding char '=' is added.\n> \n> Hence, a \"correct\" implementation should decode the input as if '=' were\n> the same as 'A' (which encodes 6 bits of 0) til the end, making sure that\n> the padding char '=' appears only at the end of the input, that no char\n> outside the Base64 encoding alphabet appears in the input, and that the\n> length of the entire encoded string is multiple of 4.  Finally it would\n> discard either one or two octets (depending on the number of padding chars\n> it saw) from the end of the output.\n> \n> Our decode_b_segment() however emits each octet as it completes, without\n> waiting for the 24-bit group that contains it to complete.  When decoding\n> a correctly encoded input, by the time we see a padding '=', all the real\n> payload octets are complete and we would not have any real information\n> still kept in the variable \"acc\" (accumulator), so ignoring '=' (you do\n> not even need to assign c = 0) like your patch did would work just fine.\n> An alternative would be to count the number of padding at the end and drop\n> the NULs from the output as necessary after the loop but that does not add\n> any value to the current code.\n> \n> Ideally we should validate the encoded string a bit more carefully (see\n> the \"correct\" implementation about), and warn if a malformed input is\n> found (but probably not reject outright).  But as a low-impact fix for the\n> maintenance branches, I think your fix is very good.\n> \n> \tSide note: I suspect that the existing code was Ok before strbuf\n> \tconversion as we assumed NUL terminated output buffer.\n\nJunio, thanks for the explanation.\n\nI've updated the patch and included your analysis into description.\n\n> > @@ -513,7 +521,15 @@ static int decode_header_bq(struct strbuf *it)\n> >  \t\tstrbuf_reset(&piecebuf);\n> >  \t\trfc2047 = 1;\n> >  \n> > -\t\tif (in != ep) {\n> > +\t\t/* XXX: the follwoing is needed not to output '\\n' on every\n> > +\t\t * multi-line segment in Subject.\n> > +\t\t *\n> > +\t\t * I suspect this is not 100% correct, but I'm not a MIME guy\n> > +\t\t * -- that's why this is an RFC.\n> > +\t\t */\n> > +\n> > +\t\t/* if in does not end with '=?=', we emit it as is */\n> > +\t\tif (in <= (ep-2) && !(ep[-1]=='\\n' && ep[-2]=='=')) {\n> >  \t\t\tstrbuf_add(&outbuf, in, ep - in);\n> >  \t\t\tin = ep;\n> > \n> >  \t\t}\n> \n> I am not a MIME guy either (and mailinfo has a big comment that says we do\n> not really do MIME --- we just pretend to do), but let me give it a try.\n> \n> RFC2046 specifies that an encoded-word (\"=?charset?encoding?...?=\") may\n> not be more than 75 characters long, and multiple encoded-words, separated\n> by CRLF SPACE can be used to encode more text if needed.\n> \n> It further specifies that an encoded-word can appear next to ordinary text\n> or another encoded-word but it must be separated by linear white space,\n> and says that such linear white space is to be ignored when displaying.\n> \n> Which means that we should be eating the CRLF SPACE we see if we have seen\n> an encoded-word immediately before and we are about to process another\n> encoded-word.\n> \n> Based on the above discussion, here is what I came up with.  It passes\n> your test, but I ran out of energy to try breaking it seriously in any\n> other way than just running the existing test suite.  \n\nThanks again very much!\n\nI was once maintaining software, and I think I understand what you mean\nby saying 'ran out of energy', so I'll try to do my best to help improve\nthis patch and to get it merged.\n\n> We might want to steal some test cases from the \"8. Examples\" section of\n> RFC2047 and add them to t5100.\n\nGood idea. I took all the examples and incorporated them into our\ntestsuite.\n\n> \n> Thanks.\n> \n>  builtin-mailinfo.c |   27 +++++++++++++++++++--------\n>  1 files changed, 19 insertions(+), 8 deletions(-)\n> \n> diff --git c/builtin-mailinfo.c w/builtin-mailinfo.c\n> index e890f7a..fcb32c9 100644\n> --- c/builtin-mailinfo.c\n> +++ w/builtin-mailinfo.c\n> @@ -430,13 +430,6 @@ static struct strbuf *decode_b_segment(const struct strbuf *b_seg)\n>  \t\t\tc -= 'a' - 26;\n>  \t\telse if ('0' <= c && c <= '9')\n>  \t\t\tc -= '0' - 52;\n> -\t\telse if (c == '=') {\n> -\t\t\t/* padding is almost like (c == 0), except we do\n> -\t\t\t * not output NUL resulting only from it;\n> -\t\t\t * for now we just trust the data.\n> -\t\t\t */\n> -\t\t\tc = 0;\n> -\t\t}\n>  \t\telse\n>  \t\t\tcontinue; /* garbage */\n>  \t\tswitch (pos++) {\n> @@ -514,7 +507,25 @@ static int decode_header_bq(struct strbuf *it)\n>  \t\trfc2047 = 1;\n>  \n>  \t\tif (in != ep) {\n> -\t\t\tstrbuf_add(&outbuf, in, ep - in);\n> +\t\t\t/*\n> +\t\t\t * We are about to process an encoded-word\n> +\t\t\t * that begins at ep, but there is something\n> +\t\t\t * before the encoded word.\n> +\t\t\t */\n> +\t\t\tchar *scan;\n> +\t\t\tfor (scan = in; scan < ep; scan++)\n> +\t\t\t\tif (!isspace(*scan))\n> +\t\t\t\t\tbreak;\n> +\n> +\t\t\tif (scan != ep || in == it->buf) {\n> +\t\t\t\t/*\n> +\t\t\t\t * We should not lose that \"something\",\n> +\t\t\t\t * unless we have just processed an\n> +\t\t\t\t * encoded-word, and there is only LWS\n> +\t\t\t\t * before the one we are about to process.\n> +\t\t\t\t */\n> +\t\t\t\tstrbuf_add(&outbuf, in, ep - in);\n> +\t\t\t}\n>  \t\t\tin = ep;\n>  \t\t}\n>  \t\t/* E.g.\n\nBased on the above description the code looks good now. I've\nincorporated it into the patch and added tests from RFC2047 (see patch\nbelow).\n\nOn Thu, Jan 08, 2009 at 01:08:13PM +0300, Alexander Potashev wrote:\n> On 21:38 Fri 26 Dec     , Kirill Smelkov wrote:\n> > When native language (RU) is in use, subject header usually contains several\n> > parts, e.g.\n> > \n> > Subject: [Navy-patches] [PATCH]\n> > \t=?utf-8?b?0JjQt9C80LXQvdGR0L0g0YHQv9C40YHQvtC6INC/0LA=?=\n> > \t=?utf-8?b?0LrQtdGC0L7QsiDQvdC10L7QsdGF0L7QtNC40LzRi9GFINC00LvRjyA=?=\n> > \t=?utf-8?b?0YHQsdC+0YDQutC4?=\n> > \n> \n> >  t/t5100/info0012    |    5 ++++\n> >  t/t5100/msg0012     |    7 ++++++\n> >  t/t5100/patch0012   |   30 +++++++++++++++++++++++++++++\n> >  t/t5100/sample.mbox |   52 +++++++++++++++++++++++++++++++++++++++++++++++++++\n> >  6 files changed, 112 insertions(+), 2 deletions(-)\n> \n> The testcases are too long, a minimal mbox with encoded \"Subject:\" would\n> be enough to test the mailinfo parser, it's all the you need to test\n> here.\n\nThanks Alexander for pointing this out.\n\nI've based my testcase on already-in-there tests, which e.g. for\nt/t5100/{info,msg,patch}00{04,05,09,10,11} are of approximately the same\nsize and are based on real mails.\n\nIs this ok?\n\n\nAs to new RFC2047-examples based tests, I've tried to keep them to the\nbare minimum.\n\n\nChanges since v1:\n\n o incorporated Junio's description and code about padding\n o incorporated Junio's description and code about LWS between encoded\n   words\n o incorporated tests from RFC2047 examples  (one testresult is unclear\n   -- see patch description)\n\n\nFrom: Kirill Smelkov <kirr@landau.phys.spbu.ru>\nSubject: mailinfo: correctly handle multiline 'Subject:' header\n\nWhen native language (RU) is in use, subject header usually contains several\nparts, e.g.\n\nSubject: [Navy-patches] [PATCH]\n\t=?utf-8?b?0JjQt9C80LXQvdGR0L0g0YHQv9C40YHQvtC6INC/0LA=?=\n\t=?utf-8?b?0LrQtdGC0L7QsiDQvdC10L7QsdGF0L7QtNC40LzRi9GFINC00LvRjyA=?=\n\t=?utf-8?b?0YHQsdC+0YDQutC4?=\n\n( which btw should be extracted by git-mailinfo to:\n\n    'Subject: Изменён список пакетов необходимых для сборки' )\n\nThis exposes several bugs in builtin-mailinfo.c which we try to fix:\n\n1. decode_b_segment: do not append explicit NUL -- explicit NUL was preventing\n   correct header construction on parts concatenation via strbuf_addbuf in\n   decode_header_bq. Fixes:\n\n-Subject: Изменён список пакетов необходимых для сборки\n+Subject: Изменён список па\n\nJunio:\n\n> B encoding (RFC 2045) encodes an octet stream into a sequence of groups of\n> 4 letters from 64-char alphabet, each of which encodes 6-bit, plus zero or\n> more padding char '=' to make the result multiple of 4.\n>\n>  * If the length of the payload is a multiple of 3 octets, there is no\n>    special handling.  Padding char '=' is not produced;\n>\n>  * If it is a multiple of 3 octets plus one, the remaining one octet is\n>    encoded with two letters, and two more padding char '=' is added;\n>\n>  * If it is a multiple of 3 octets plus two, the remaining two octets are\n>    encoded with three letters, and one padding char '=' is added.\n>\n> Hence, a \"correct\" implementation should decode the input as if '=' were\n> the same as 'A' (which encodes 6 bits of 0) til the end, making sure that\n> the padding char '=' appears only at the end of the input, that no char\n> outside the Base64 encoding alphabet appears in the input, and that the\n> length of the entire encoded string is multiple of 4.  Finally it would\n> discard either one or two octets (depending on the number of padding chars\n> it saw) from the end of the output.\n>\n> Our decode_b_segment() however emits each octet as it completes, without\n> waiting for the 24-bit group that contains it to complete.  When decoding\n> a correctly encoded input, by the time we see a padding '=', all the real\n> payload octets are complete and we would not have any real information\n> still kept in the variable \"acc\" (accumulator), so ignoring '=' (you do\n> not even need to assign c = 0) like your patch did would work just fine.\n> An alternative would be to count the number of padding at the end and drop\n> the NULs from the output as necessary after the loop but that does not add\n> any value to the current code.\n>\n> Ideally we should validate the encoded string a bit more carefully (see\n> the \"correct\" implementation about), and warn if a malformed input is\n> found (but probably not reject outright).  But as a low-impact fix for the\n> maintenance branches, I think your fix is very good.\n>\n> \tSide note: I suspect that the existing code was Ok before strbuf\n> \tconversion as we assumed NUL terminated output buffer.\n\n\nThen\n\n2. whitespaces between encoded words should be removed\n\n-Subject: Изменён список пакетов необходимых для сборки\n+Subject: Изменён список па кетов необходимых для сборки\n\nJunio:\n\n> I am not a MIME guy either (and mailinfo has a big comment that says we do\n> not really do MIME --- we just pretend to do), but let me give it a try.\n>\n> RFC2046 specifies that an encoded-word (\"=?charset?encoding?...?=\") may\n> not be more than 75 characters long, and multiple encoded-words, separated\n> by CRLF SPACE can be used to encode more text if needed.\n>\n> It further specifies that an encoded-word can appear next to ordinary text\n> or another encoded-word but it must be separated by linear white space,\n> and says that such linear white space is to be ignored when displaying.\n>\n> Which means that we should be eating the CRLF SPACE we see if we have seen\n> an encoded-word immediately before and we are about to process another\n> encoded-word.\n\nAlso as suggested by Junio, in order to try to catch other MIME problems test\ncases from the \"8. Examples\" section of RFC2047 are added to t5100 testsuite as\nwell.\n\n    [but I'm not sure whether testresult with Nathaniel Borenstein\n     (םולש ןב ילטפנ) is correct -- see rfc2047-info-0004]\n\nBig-thanks-to: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Kirill Smelkov <kirr@landau.phys.spbu.ru>\n\n---\n builtin-mailinfo.c           |   27 +++++++++++++++------\n t/t5100-mailinfo.sh          |   24 ++++++++++++++++++-\n t/t5100/info0012             |    5 ++++\n t/t5100/msg0012              |    7 +++++\n t/t5100/patch0012            |   30 ++++++++++++++++++++++++\n t/t5100/rfc2047-info-0001    |    4 +++\n t/t5100/rfc2047-info-0002    |    4 +++\n t/t5100/rfc2047-info-0003    |    4 +++\n t/t5100/rfc2047-info-0004    |    5 ++++\n t/t5100/rfc2047-info-0005    |    2 +\n t/t5100/rfc2047-info-0006    |    2 +\n t/t5100/rfc2047-info-0007    |    2 +\n t/t5100/rfc2047-info-0008    |    2 +\n t/t5100/rfc2047-info-0009    |    2 +\n t/t5100/rfc2047-info-0010    |    2 +\n t/t5100/rfc2047-info-0011    |    2 +\n t/t5100/rfc2047-samples.mbox |   48 ++++++++++++++++++++++++++++++++++++++\n t/t5100/sample.mbox          |   52 ++++++++++++++++++++++++++++++++++++++++++\n 18 files changed, 215 insertions(+), 9 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex f7c8c08..77a7121 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -430,13 +430,6 @@ static struct strbuf *decode_b_segment(const struct strbuf *b_seg)\n \t\t\tc -= 'a' - 26;\n \t\telse if ('0' <= c && c <= '9')\n \t\t\tc -= '0' - 52;\n-\t\telse if (c == '=') {\n-\t\t\t/* padding is almost like (c == 0), except we do\n-\t\t\t * not output NUL resulting only from it;\n-\t\t\t * for now we just trust the data.\n-\t\t\t */\n-\t\t\tc = 0;\n-\t\t}\n \t\telse\n \t\t\tcontinue; /* garbage */\n \t\tswitch (pos++) {\n@@ -514,7 +507,25 @@ static int decode_header_bq(struct strbuf *it)\n \t\trfc2047 = 1;\n \n \t\tif (in != ep) {\n-\t\t\tstrbuf_add(&outbuf, in, ep - in);\n+\t\t\t/*\n+\t\t\t * We are about to process an encoded-word\n+\t\t\t * that begins at ep, but there is something\n+\t\t\t * before the encoded word.\n+\t\t\t */\n+\t\t\tchar *scan;\n+\t\t\tfor (scan = in; scan < ep; scan++)\n+\t\t\t\tif (!isspace(*scan))\n+\t\t\t\t\tbreak;\n+\n+\t\t\tif (scan != ep || in == it->buf) {\n+\t\t\t\t/*\n+\t\t\t\t * We should not lose that \"something\",\n+\t\t\t\t * unless we have just processed an\n+\t\t\t\t * encoded-word, and there is only LWS\n+\t\t\t\t * before the one we are about to process.\n+\t\t\t\t */\n+\t\t\t\tstrbuf_add(&outbuf, in, ep - in);\n+\t\t\t}\n \t\t\tin = ep;\n \t\t}\n \t\t/* E.g.\ndiff --git a/t/t5100-mailinfo.sh b/t/t5100-mailinfo.sh\nindex fe14589..625c204 100755\n--- a/t/t5100-mailinfo.sh\n+++ b/t/t5100-mailinfo.sh\n@@ -11,7 +11,7 @@ test_expect_success 'split sample box' \\\n \t'git mailsplit -o. \"$TEST_DIRECTORY\"/t5100/sample.mbox >last &&\n \tlast=`cat last` &&\n \techo total is $last &&\n-\ttest `cat last` = 11'\n+\ttest `cat last` = 12'\n \n for mail in `echo 00*`\n do\n@@ -26,6 +26,28 @@ do\n \t'\n done\n \n+\n+test_expect_success 'split box with rfc2047 samples' \\\n+\t'mkdir rfc2047 &&\n+\tgit mailsplit -orfc2047 \"$TEST_DIRECTORY\"/t5100/rfc2047-samples.mbox \\\n+\t  >rfc2047/last &&\n+\tlast=`cat rfc2047/last` &&\n+\techo total is $last &&\n+\ttest `cat rfc2047/last` = 11'\n+\n+for mail in `echo rfc2047/00*`\n+do\n+\ttest_expect_success \"mailinfo $mail\" '\n+\t\tgit mailinfo -u $mail-msg $mail-patch <$mail >$mail-info &&\n+\t\techo msg &&\n+\t\ttest_cmp \"$TEST_DIRECTORY\"/t5100/empty $mail-msg &&\n+\t\techo patch &&\n+\t\ttest_cmp \"$TEST_DIRECTORY\"/t5100/empty $mail-patch &&\n+\t\techo info &&\n+\t\ttest_cmp \"$TEST_DIRECTORY\"/t5100/rfc2047-info-$(basename $mail) $mail-info\n+\t'\n+done\n+\n test_expect_success 'respect NULs' '\n \n \tgit mailsplit -d3 -o. \"$TEST_DIRECTORY\"/t5100/nul-plain &&\ndiff --git a/t/t5100/empty b/t/t5100/empty\nnew file mode 100644\nindex 0000000..e69de29\ndiff --git a/t/t5100/info0012 b/t/t5100/info0012\nnew file mode 100644\nindex 0000000..ac1216f\n--- /dev/null\n+++ b/t/t5100/info0012\n@@ -0,0 +1,5 @@\n+Author: Dmitriy Blinov\n+Email: bda@mnsspb.ru\n+Subject: Изменён список пакетов необходимых для сборки\n+Date: Wed, 12 Nov 2008 17:54:41 +0300\n+\ndiff --git a/t/t5100/msg0012 b/t/t5100/msg0012\nnew file mode 100644\nindex 0000000..1dc2bf7\n--- /dev/null\n+++ b/t/t5100/msg0012\n@@ -0,0 +1,7 @@\n+textlive-* исправлены на texlive-*\n+docutils заменён на python-docutils\n+\n+Действительно, оказалось, что rest2web вытягивает за собой\n+python-docutils. В то время как сам rest2web не нужен.\n+\n+Signed-off-by: Dmitriy Blinov <bda@mnsspb.ru>\ndiff --git a/t/t5100/patch0012 b/t/t5100/patch0012\nnew file mode 100644\nindex 0000000..36a0b68\n--- /dev/null\n+++ b/t/t5100/patch0012\n@@ -0,0 +1,30 @@\n+---\n+ howto/build_navy.txt |    6 +++---\n+ 1 files changed, 3 insertions(+), 3 deletions(-)\n+\n+diff --git a/howto/build_navy.txt b/howto/build_navy.txt\n+index 3fd3afb..0ee807e 100644\n+--- a/howto/build_navy.txt\n++++ b/howto/build_navy.txt\n+@@ -119,8 +119,8 @@\n+    - libxv-dev\n+    - libusplash-dev\n+    - latex-make\n+-   - textlive-lang-cyrillic\n+-   - textlive-latex-extra\n++   - texlive-lang-cyrillic\n++   - texlive-latex-extra\n+    - dia\n+    - python-pyrex\n+    - libtool\n+@@ -128,7 +128,7 @@\n+    - sox\n+    - cython\n+    - imagemagick\n+-   - docutils\n++   - python-docutils\n+ \n+ #. на машине dinar: добавить свой открытый ssh-ключ в authorized_keys2 пользователя ddev\n+ #. на своей машине: отредактировать /etc/sudoers (команда ``visudo``) примерно следующим образом::\n+-- \n+1.5.6.5\ndiff --git a/t/t5100/rfc2047-info-0001 b/t/t5100/rfc2047-info-0001\nnew file mode 100644\nindex 0000000..0a383b0\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0001\n@@ -0,0 +1,4 @@\n+Author: Keith Moore\n+Email: moore@cs.utk.edu\n+Subject: If you can read this you understand the example.\n+\ndiff --git a/t/t5100/rfc2047-info-0002 b/t/t5100/rfc2047-info-0002\nnew file mode 100644\nindex 0000000..881be75\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0002\n@@ -0,0 +1,4 @@\n+Author: Olle Järnefors\n+Email: ojarnef@admin.kth.se\n+Subject: Time for ISO 10646?\n+\ndiff --git a/t/t5100/rfc2047-info-0003 b/t/t5100/rfc2047-info-0003\nnew file mode 100644\nindex 0000000..d0f7891\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0003\n@@ -0,0 +1,4 @@\n+Author: Patrik Fältström\n+Email: paf@nada.kth.se\n+Subject: RFC-HDR care and feeding\n+\ndiff --git a/t/t5100/rfc2047-info-0004 b/t/t5100/rfc2047-info-0004\nnew file mode 100644\nindex 0000000..850f831\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0004\n@@ -0,0 +1,5 @@\n+Author: Nathaniel Borenstein  \n+     (םולש ןב ילטפנ)\n+Email: nsb@thumper.bellcore.com\n+Subject: Test of new header generator\n+\ndiff --git a/t/t5100/rfc2047-info-0005 b/t/t5100/rfc2047-info-0005\nnew file mode 100644\nindex 0000000..c27be3b\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0005\n@@ -0,0 +1,2 @@\n+Subject: (a)\n+\ndiff --git a/t/t5100/rfc2047-info-0006 b/t/t5100/rfc2047-info-0006\nnew file mode 100644\nindex 0000000..9dad474\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0006\n@@ -0,0 +1,2 @@\n+Subject: (a b)\n+\ndiff --git a/t/t5100/rfc2047-info-0007 b/t/t5100/rfc2047-info-0007\nnew file mode 100644\nindex 0000000..294f195\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0007\n@@ -0,0 +1,2 @@\n+Subject: (ab)\n+\ndiff --git a/t/t5100/rfc2047-info-0008 b/t/t5100/rfc2047-info-0008\nnew file mode 100644\nindex 0000000..294f195\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0008\n@@ -0,0 +1,2 @@\n+Subject: (ab)\n+\ndiff --git a/t/t5100/rfc2047-info-0009 b/t/t5100/rfc2047-info-0009\nnew file mode 100644\nindex 0000000..294f195\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0009\n@@ -0,0 +1,2 @@\n+Subject: (ab)\n+\ndiff --git a/t/t5100/rfc2047-info-0010 b/t/t5100/rfc2047-info-0010\nnew file mode 100644\nindex 0000000..9dad474\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0010\n@@ -0,0 +1,2 @@\n+Subject: (a b)\n+\ndiff --git a/t/t5100/rfc2047-info-0011 b/t/t5100/rfc2047-info-0011\nnew file mode 100644\nindex 0000000..9dad474\n--- /dev/null\n+++ b/t/t5100/rfc2047-info-0011\n@@ -0,0 +1,2 @@\n+Subject: (a b)\n+\ndiff --git a/t/t5100/rfc2047-samples.mbox b/t/t5100/rfc2047-samples.mbox\nnew file mode 100644\nindex 0000000..3ca2470\n--- /dev/null\n+++ b/t/t5100/rfc2047-samples.mbox\n@@ -0,0 +1,48 @@\n+From nobody Mon Sep 17 00:00:00 2001\n+From: =?US-ASCII?Q?Keith_Moore?= <moore@cs.utk.edu>\n+To: =?ISO-8859-1?Q?Keld_J=F8rn_Simonsen?= <keld@dkuug.dk>\n+CC: =?ISO-8859-1?Q?Andr=E9?= Pirard <PIRARD@vm1.ulg.ac.be>\n+Subject: =?ISO-8859-1?B?SWYgeW91IGNhbiByZWFkIHRoaXMgeW8=?=\n+ =?ISO-8859-2?B?dSB1bmRlcnN0YW5kIHRoZSBleGFtcGxlLg==?=\n+\n+From nobody Mon Sep 17 00:00:00 2001\n+From: =?ISO-8859-1?Q?Olle_J=E4rnefors?= <ojarnef@admin.kth.se>\n+To: ietf-822@dimacs.rutgers.edu, ojarnef@admin.kth.se\n+Subject: Time for ISO 10646?\n+\n+From nobody Mon Sep 17 00:00:00 2001\n+To: Dave Crocker <dcrocker@mordor.stanford.edu>\n+Cc: ietf-822@dimacs.rutgers.edu, paf@comsol.se\n+From: =?ISO-8859-1?Q?Patrik_F=E4ltstr=F6m?= <paf@nada.kth.se>\n+Subject: Re: RFC-HDR care and feeding\n+\n+From nobody Mon Sep 17 00:00:00 2001\n+From: Nathaniel Borenstein <nsb@thumper.bellcore.com>\n+      (=?iso-8859-8?b?7eXs+SDv4SDp7Oj08A==?=)\n+To: Greg Vaudreuil <gvaudre@NRI.Reston.VA.US>, Ned Freed\n+   <ned@innosoft.com>, Keith Moore <moore@cs.utk.edu>\n+Subject: Test of new header generator\n+MIME-Version: 1.0\n+Content-type: text/plain; charset=ISO-8859-1\n+\n+From nobody Mon Sep 17 00:00:00 2001\n+Subject: (=?ISO-8859-1?Q?a?=)\n+\n+From nobody Mon Sep 17 00:00:00 2001\n+Subject: (=?ISO-8859-1?Q?a?= b)\n+\n+From nobody Mon Sep 17 00:00:00 2001\n+Subject: (=?ISO-8859-1?Q?a?= =?ISO-8859-1?Q?b?=)\n+\n+From nobody Mon Sep 17 00:00:00 2001\n+Subject: (=?ISO-8859-1?Q?a?=  =?ISO-8859-1?Q?b?=)\n+\n+From nobody Mon Sep 17 00:00:00 2001\n+Subject: (=?ISO-8859-1?Q?a?=\n+    =?ISO-8859-1?Q?b?=)\n+\n+From nobody Mon Sep 17 00:00:00 2001\n+Subject: (=?ISO-8859-1?Q?a_b?=)\n+\n+From nobody Mon Sep 17 00:00:00 2001\n+Subject: (=?ISO-8859-1?Q?a?= =?ISO-8859-2?Q?_b?=)\ndiff --git a/t/t5100/sample.mbox b/t/t5100/sample.mbox\nindex 4bf7947..94da4da 100644\n--- a/t/t5100/sample.mbox\n+++ b/t/t5100/sample.mbox\n@@ -501,3 +501,55 @@ index 3e5fe51..aabfe5c 100644\n \n --=-=-=--\n \n+From bda@mnsspb.ru Wed Nov 12 17:54:41 2008\n+From: Dmitriy Blinov <bda@mnsspb.ru>\n+To: navy-patches@dinar.mns.mnsspb.ru\n+Date: Wed, 12 Nov 2008 17:54:41 +0300\n+Message-Id: <1226501681-24923-1-git-send-email-bda@mnsspb.ru>\n+X-Mailer: git-send-email 1.5.6.5\n+MIME-Version: 1.0\n+Content-Type: text/plain;\n+  charset=utf-8\n+Content-Transfer-Encoding: 8bit\n+Subject: [Navy-patches] [PATCH]\n+\t=?utf-8?b?0JjQt9C80LXQvdGR0L0g0YHQv9C40YHQvtC6INC/0LA=?=\n+\t=?utf-8?b?0LrQtdGC0L7QsiDQvdC10L7QsdGF0L7QtNC40LzRi9GFINC00LvRjyA=?=\n+\t=?utf-8?b?0YHQsdC+0YDQutC4?=\n+\n+textlive-* исправлены на texlive-*\n+docutils заменён на python-docutils\n+\n+Действительно, оказалось, что rest2web вытягивает за собой\n+python-docutils. В то время как сам rest2web не нужен.\n+\n+Signed-off-by: Dmitriy Blinov <bda@mnsspb.ru>\n+---\n+ howto/build_navy.txt |    6 +++---\n+ 1 files changed, 3 insertions(+), 3 deletions(-)\n+\n+diff --git a/howto/build_navy.txt b/howto/build_navy.txt\n+index 3fd3afb..0ee807e 100644\n+--- a/howto/build_navy.txt\n++++ b/howto/build_navy.txt\n+@@ -119,8 +119,8 @@\n+    - libxv-dev\n+    - libusplash-dev\n+    - latex-make\n+-   - textlive-lang-cyrillic\n+-   - textlive-latex-extra\n++   - texlive-lang-cyrillic\n++   - texlive-latex-extra\n+    - dia\n+    - python-pyrex\n+    - libtool\n+@@ -128,7 +128,7 @@\n+    - sox\n+    - cython\n+    - imagemagick\n+-   - docutils\n++   - python-docutils\n+ \n+ #. на машине dinar: добавить свой открытый ssh-ключ в authorized_keys2 пользователя ddev\n+ #. на своей машине: отредактировать /etc/sudoers (команда ``visudo``) примерно следующим образом::\n+-- \n+1.5.6.5\n-- \ntg: (c123b7c..) t/mailinfo-multiline-subject (depends on: master)\n\nThanks,\nKirill\n"},{"id":"99861","messageId":"20090110101240.GA9048@roro3.zxlink","threadId":"17040","inReplyTo":"20090108231135.GB4185@roro3","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Kirill Smelkov","fromEmail":"kirr@landau.phys.spbu.ru","sentAt":"2009-01-10T10:12:40Z","receivedAt":"2009-01-10T10:12:40Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Fri, Jan 09, 2009 at 02:11:35AM +0300, Kirill Smelkov wrote:\n> Changes since v1:\n> \n>  o incorporated Junio's description and code about padding\n>  o incorporated Junio's description and code about LWS between encoded\n>    words\n>  o incorporated tests from RFC2047 examples  (one testresult is unclear\n>    -- see patch description)\n> \n> \n> From: Kirill Smelkov <kirr@landau.phys.spbu.ru>\n> Subject: mailinfo: correctly handle multiline 'Subject:' header\n\n[...]\n\nJunio, All, just in case this again got spam-detected:\n\nhttp://marc.info/?l=git&m=123145624611936&w=2\n\n\nThanks,\nKirill\n"},{"id":"99931","messageId":"7veizatxo9.fsf@gitster.siamese.dyndns.org","threadId":"17040","inReplyTo":"20090108231135.GB4185@roro3","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-11T01:54:14Z","receivedAt":"2009-01-11T01:54:14Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@landau.phys.spbu.ru> writes:\n\n>     [but I'm not sure whether testresult with Nathaniel Borenstein\n>      (םולש ןב ילטפנ) is correct -- see rfc2047-info-0004]\n> ...\n> diff --git a/t/t5100/rfc2047-info-0004 b/t/t5100/rfc2047-info-0004\n> new file mode 100644\n> index 0000000..850f831\n> --- /dev/null\n> +++ b/t/t5100/rfc2047-info-0004\n> @@ -0,0 +1,5 @@\n> +Author: Nathaniel Borenstein  \n> +     (םולש ןב ילטפנ)\n> +Email: nsb@thumper.bellcore.com\n> +Subject: Test of new header generator\n> +\n\nThat does look wrong.  If you can fix this, please do so; otherwise please\nmark the test that deals with this entry with test_expect_failure, until\nsomebody else does.\n"},{"id":"100184","messageId":"20090112223447.GA5948@roro3.zxlink","threadId":"17040","inReplyTo":"7veizatxo9.fsf@gitster.siamese.dyndns.org","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Kirill Smelkov","fromEmail":"kirr@landau.phys.spbu.ru","sentAt":"2009-01-12T22:34:47Z","receivedAt":"2009-01-12T22:34:47Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Sat, Jan 10, 2009 at 05:54:14PM -0800, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@landau.phys.spbu.ru> writes:\n> \n> >     [but I'm not sure whether testresult with Nathaniel Borenstein\n> >      (םולש ןב ילטפנ) is correct -- see rfc2047-info-0004]\n> > ...\n> > diff --git a/t/t5100/rfc2047-info-0004 b/t/t5100/rfc2047-info-0004\n> > new file mode 100644\n> > index 0000000..850f831\n> > --- /dev/null\n> > +++ b/t/t5100/rfc2047-info-0004\n> > @@ -0,0 +1,5 @@\n> > +Author: Nathaniel Borenstein  \n> > +     ([somethig that could be detected as spam])\n> > +Email: nsb@thumper.bellcore.com\n> > +Subject: Test of new header generator\n> > +\n> \n> That does look wrong.  If you can fix this, please do so; otherwise please\n> mark the test that deals with this entry with test_expect_failure, until\n> somebody else does.\n\nYes, I think I've dealt with it -- we weren't unfolding 'From' header,\nand we were not skipping comments in rfc822 headers, so:\n\nFrom: Kirill Smelkov <kirr@landau.phys.spbu.ru>\nSubject: [PATCH] mailinfo: 'From:' header should be unfold as well\n\nAt present we do headers unfolding (see RFC822 3.1.1. LONG HEADER FIELDS) for\nall fields except 'From' (always) and 'Subject' (when keep_subject is set)\n\nNot unfolding 'From' is a bug -- see above-mentioned RFC link.\n\nSigned-off-by: Kirill Smelkov <kirr@landau.phys.spbu.ru>\n\n---\n builtin-mailinfo.c  |    1 +\n t/t5100/sample.mbox |    5 ++++-\n 2 files changed, 5 insertions(+), 1 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex f7c8c08..6d72c1b 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -860,6 +860,7 @@ static void handle_info(void)\n \t\t\t}\n \t\t\toutput_header_lines(fout, \"Subject\", hdr);\n \t\t} else if (!memcmp(header[i], \"From\", 4)) {\n+\t\t\tcleanup_space(hdr);\n \t\t\thandle_from(hdr);\n \t\t\tfprintf(fout, \"Author: %s\\n\", name.buf);\n \t\t\tfprintf(fout, \"Email: %s\\n\", email.buf);\ndiff --git a/t/t5100/sample.mbox b/t/t5100/sample.mbox\nindex 4bf7947..d465685 100644\n--- a/t/t5100/sample.mbox\n+++ b/t/t5100/sample.mbox\n@@ -2,7 +2,10 @@\n \t\n     \n From nobody Mon Sep 17 00:00:00 2001\n-From: A U Thor <a.u.thor@example.com>\n+From: A\n+      U\n+      Thor\n+      <a.u.thor@example.com>\n Date: Fri, 9 Jun 2006 00:44:16 -0700\n Subject: [PATCH] a commit.\n \n-- \ntg: (1562445..) t/mail-from-unfold (depends on: master)\n\n\n\n\nFrom: Kirill Smelkov <kirr@landau.phys.spbu.ru>\nSubject: [PATCH] mailinfo: more smarter removal of rfc822 comments from 'From'\n\nAs described in RFC822 (3.4.3 COMMENTS, and  A.1.4.), comments, as e.g.\n\n    John (zzz) Doe <john.doe@xz> (Comment)\n\nshould \"NOT [be] included in the destination mailbox\"\n\nWe need this functionality to pass all RFC2047 based tests in the next commit.\n\nSigned-off-by: Kirill Smelkov <kirr@landau.phys.spbu.ru>\n\n---\n builtin-mailinfo.c  |   30 ++++++++++++++++++++++++++++++\n t/t5100/sample.mbox |    4 ++--\n 2 files changed, 32 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 6d72c1b..c0b1ab4 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -29,6 +29,9 @@ static struct strbuf **p_hdr_data, **s_hdr_data;\n #define MAX_HDR_PARSED 10\n #define MAX_BOUNDARIES 5\n \n+static void cleanup_space(struct strbuf *sb);\n+\n+\n static void get_sane_name(struct strbuf *out, struct strbuf *name, struct strbuf *email)\n {\n \tstruct strbuf *src = name;\n@@ -120,6 +123,33 @@ static void handle_from(const struct strbuf *from)\n \t\tstrbuf_setlen(&f, f.len - 1);\n \t}\n \n+\t/* This still could not be finished for emails like\n+\t *\n+\t *\t\"John (zzz) Doe <john.doe@xz> (Comment)\"\n+\t *\n+\t * The email part had already been removed, so let's kill comments as\n+\t * well -- RFC822 says comments should not be present in destination\n+\t * mailbox (3.4.3. Comments  and  A.1.4.)\n+\t */\n+\twhile (1) {\n+\t\tchar *ta;\n+\n+\t\tat = strchr(f.buf, '(');\n+\t\tif (!at)\n+\t\t\tbreak;\n+\t\tta = strchr(at, ')');\n+\t\tif (!ta)\n+\t\t\tbreak;\n+\n+\t\tstrbuf_remove(&f, at - f.buf, ta-at + (*ta ? 1 : 0));\n+\t}\n+\n+\t/* and let's finally cleanup spaces that were around (possibly\n+\t * internal) comments\n+\t */\n+\tcleanup_space(&f);\n+\tstrbuf_trim(&f);\n+\n \tget_sane_name(&name, &f, &email);\n \tstrbuf_release(&f);\n }\ndiff --git a/t/t5100/sample.mbox b/t/t5100/sample.mbox\nindex d465685..42e02f3 100644\n--- a/t/t5100/sample.mbox\n+++ b/t/t5100/sample.mbox\n@@ -2,10 +2,10 @@\n \t\n     \n From nobody Mon Sep 17 00:00:00 2001\n-From: A\n+From: A (zzz)\n       U\n       Thor\n-      <a.u.thor@example.com>\n+      <a.u.thor@example.com> (Comment)\n Date: Fri, 9 Jun 2006 00:44:16 -0700\n Subject: [PATCH] a commit.\n \n-- \ntg: (b798ad9..) t/mail-from-comments (depends on: t/mail-from-unfold)\n\n\n\nAll these patches + original one (trivially adapted) could be pulled from\n\n    git://repo.or.cz/git/kirr.git  for-junio\n\n\n\nKirill Smelkov (3):\n      mailinfo: 'From:' header should be unfold as well\n      mailinfo: more smarter removal of rfc822 comments from 'From'\n      mailinfo: correctly handle multiline 'Subject:' header\n\n\nbuiltin-mailinfo.c           |   58 ++++++++++++++++++++++++++++++++++++------\nt/t5100-mailinfo.sh          |   24 ++++++++++++++++-\nt/t5100/info0012             |    5 +++\nt/t5100/msg0012              |    7 +++++\nt/t5100/patch0012            |   30 +++++++++++++++++++++\nt/t5100/rfc2047-info-0001    |    4 +++\nt/t5100/rfc2047-info-0002    |    4 +++\nt/t5100/rfc2047-info-0003    |    4 +++\nt/t5100/rfc2047-info-0004    |    4 +++\nt/t5100/rfc2047-info-0005    |    2 +\nt/t5100/rfc2047-info-0006    |    2 +\nt/t5100/rfc2047-info-0007    |    2 +\nt/t5100/rfc2047-info-0008    |    2 +\nt/t5100/rfc2047-info-0009    |    2 +\nt/t5100/rfc2047-info-0010    |    2 +\nt/t5100/rfc2047-info-0011    |    2 +\nt/t5100/rfc2047-samples.mbox |   48 ++++++++++++++++++++++++++++++++++\nt/t5100/sample.mbox          |   57 ++++++++++++++++++++++++++++++++++++++++-\n18 files changed, 249 insertions(+), 10 deletions(-)\n\n\nThanks,\nKirill\n"},{"id":"100193","messageId":"7v63kkgl5b.fsf@gitster.siamese.dyndns.org","threadId":"17040","inReplyTo":"20090112223447.GA5948@roro3.zxlink","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-12T23:27:44Z","receivedAt":"2009-01-12T23:27:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@landau.phys.spbu.ru> writes:\n\n> diff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\n> index f7c8c08..6d72c1b 100644\n> --- a/builtin-mailinfo.c\n> +++ b/builtin-mailinfo.c\n> @@ -860,6 +860,7 @@ static void handle_info(void)\n>  \t\t\t}\n>  \t\t\toutput_header_lines(fout, \"Subject\", hdr);\n>  \t\t} else if (!memcmp(header[i], \"From\", 4)) {\n> +\t\t\tcleanup_space(hdr);\n>  \t\t\thandle_from(hdr);\n>  \t\t\tfprintf(fout, \"Author: %s\\n\", name.buf);\n>  \t\t\tfprintf(fout, \"Email: %s\\n\", email.buf);\n> diff --git a/t/t5100/sample.mbox b/t/t5100/sample.mbox\n> index 4bf7947..d465685 100644\n> --- a/t/t5100/sample.mbox\n> +++ b/t/t5100/sample.mbox\n> @@ -2,7 +2,10 @@\n>  \t\n>      \n>  From nobody Mon Sep 17 00:00:00 2001\n> -From: A U Thor <a.u.thor@example.com>\n> +From: A\n> +      U\n> +      Thor\n> +      <a.u.thor@example.com>\n>  Date: Fri, 9 Jun 2006 00:44:16 -0700\n>  Subject: [PATCH] a commit.\n\nI think this is a reasonable change.\n\nBut doesn't this\n\n>  From nobody Mon Sep 17 00:00:00 2001\n> -From: A\n> +From: A (zzz)\n>        U\n>        Thor\n> -      <a.u.thor@example.com>\n> +      <a.u.thor@example.com> (Comment)\n\nregress for people who spell their names like this?\n\n\tFrom: john.doe@email.xz (John Doe)\n"},{"id":"100259","messageId":"20090113093916.GA25471@landau.phys.spbu.ru","threadId":"17040","inReplyTo":"7v63kkgl5b.fsf@gitster.siamese.dyndns.org","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Kirill Smelkov","fromEmail":"kirr@landau.phys.spbu.ru","sentAt":"2009-01-13T09:39:16Z","receivedAt":"2009-01-13T09:39:16Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Mon, Jan 12, 2009 at 03:27:44PM -0800, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@landau.phys.spbu.ru> writes:\n> \n> > diff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\n> > index f7c8c08..6d72c1b 100644\n> > --- a/builtin-mailinfo.c\n> > +++ b/builtin-mailinfo.c\n> > @@ -860,6 +860,7 @@ static void handle_info(void)\n> >  \t\t\t}\n> >  \t\t\toutput_header_lines(fout, \"Subject\", hdr);\n> >  \t\t} else if (!memcmp(header[i], \"From\", 4)) {\n> > +\t\t\tcleanup_space(hdr);\n> >  \t\t\thandle_from(hdr);\n> >  \t\t\tfprintf(fout, \"Author: %s\\n\", name.buf);\n> >  \t\t\tfprintf(fout, \"Email: %s\\n\", email.buf);\n> > diff --git a/t/t5100/sample.mbox b/t/t5100/sample.mbox\n> > index 4bf7947..d465685 100644\n> > --- a/t/t5100/sample.mbox\n> > +++ b/t/t5100/sample.mbox\n> > @@ -2,7 +2,10 @@\n> >  \t\n> >      \n> >  From nobody Mon Sep 17 00:00:00 2001\n> > -From: A U Thor <a.u.thor@example.com>\n> > +From: A\n> > +      U\n> > +      Thor\n> > +      <a.u.thor@example.com>\n> >  Date: Fri, 9 Jun 2006 00:44:16 -0700\n> >  Subject: [PATCH] a commit.\n> \n> I think this is a reasonable change.\n\nThanks.\n\n\n> But doesn't this\n> \n> >  From nobody Mon Sep 17 00:00:00 2001\n> > -From: A\n> > +From: A (zzz)\n> >        U\n> >        Thor\n> > -      <a.u.thor@example.com>\n> > +      <a.u.thor@example.com> (Comment)\n> \n> regress for people who spell their names like this?\n> \n> \tFrom: john.doe@email.xz (John Doe)\n\nI think everything is ok:\n\nThere is an explicit handler for such emails before my comments removal\nin builtin-mailinfo.c:\n\n        /* The remainder is name.  It could be \"John Doe <john.doe@xz>\"\n         * or \"john.doe@xz (John Doe)\", but we have removed the\n         * email part, so trim from both ends, possibly removing\n         * the () pair at the end.\n         */\n        strbuf_trim(&f);\n        if (f.buf[0] == '(' && f.len && f.buf[f.len - 1] == ')') {\n                strbuf_remove(&f, 0, 1);\n                strbuf_setlen(&f, f.len - 1);\n        }\n\n\nhttp://repo.or.cz/w/git.git?a=blob;f=builtin-mailinfo.c;h=f7c8c08b320c99d8bf96443ae57aa33c1de7e8c0;hb=HEAD#l112\n\n\nAnd only a test for this is missing\n\n\n\nFrom 77316ad6db2c3b0f4be238c4ba855b2f785b50d6 Mon Sep 17 00:00:00 2001\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\nDate: Tue, 13 Jan 2009 12:33:48 +0300\nSubject: [PATCH] mailinfo: add explicit test for mails like '<a.u.thor@example.com> (A U Thor)'\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n t/t5100-mailinfo.sh |    2 +-\n t/t5100/info0013    |    5 +++++\n t/t5100/sample.mbox |    5 +++++\n 3 files changed, 11 insertions(+), 1 deletions(-)\n create mode 100644 t/t5100/info0013\n create mode 100644 t/t5100/msg0013\n create mode 100644 t/t5100/patch0013\n\ndiff --git a/t/t5100-mailinfo.sh b/t/t5100-mailinfo.sh\nindex 625c204..e70ea94 100755\n--- a/t/t5100-mailinfo.sh\n+++ b/t/t5100-mailinfo.sh\n@@ -11,7 +11,7 @@ test_expect_success 'split sample box' \\\n \t'git mailsplit -o. \"$TEST_DIRECTORY\"/t5100/sample.mbox >last &&\n \tlast=`cat last` &&\n \techo total is $last &&\n-\ttest `cat last` = 12'\n+\ttest `cat last` = 13'\n \n for mail in `echo 00*`\n do\ndiff --git a/t/t5100/info0013 b/t/t5100/info0013\nnew file mode 100644\nindex 0000000..bbe049e\n--- /dev/null\n+++ b/t/t5100/info0013\n@@ -0,0 +1,5 @@\n+Author: A U Thor\n+Email: a.u.thor@example.com\n+Subject: a patch\n+Date: Fri, 9 Jun 2006 00:44:16 -0700\n+\ndiff --git a/t/t5100/msg0013 b/t/t5100/msg0013\nnew file mode 100644\nindex 0000000..e69de29\ndiff --git a/t/t5100/patch0013 b/t/t5100/patch0013\nnew file mode 100644\nindex 0000000..e69de29\ndiff --git a/t/t5100/sample.mbox b/t/t5100/sample.mbox\nindex 4f80b82..c5ad206 100644\n--- a/t/t5100/sample.mbox\n+++ b/t/t5100/sample.mbox\n@@ -556,3 +556,8 @@ index 3fd3afb..0ee807e 100644\n  #. п╫п╟ я│п╡п╬п╣п╧ п╪п╟я┬п╦п╫п╣: п╬я┌я─п╣п╢п╟п╨я┌п╦я─п╬п╡п╟я┌я▄ /etc/sudoers (п╨п╬п╪п╟п╫п╢п╟ ``visudo``) п©я─п╦п╪п╣я─п╫п╬ я│п╩п╣п╢я┐я▌я┴п╦п╪ п╬п╠я─п╟п╥п╬п╪::\n -- \n 1.5.6.5\n+From nobody Mon Sep 17 00:00:00 2001\n+From: <a.u.thor@example.com> (A U Thor)\n+Date: Fri, 9 Jun 2006 00:44:16 -0700\n+Subject: [PATCH] a patch\n+\n-- \n1.6.1.101.g0335\n"},{"id":"100379","messageId":"20090114081942.GA6399@landau.phys.spbu.ru","threadId":"17040","inReplyTo":"20090113093916.GA25471@landau.phys.spbu.ru","subject":"Re: [BUG PATCH RFC] mailinfo: correctly handle multiline 'Subject:' header","fromName":"Kirill Smelkov","fromEmail":"kirr@landau.phys.spbu.ru","sentAt":"2009-01-14T08:19:42Z","receivedAt":"2009-01-14T08:19:42Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Tue, Jan 13, 2009 at 12:39:16PM +0300, Kirill Smelkov wrote:\n> On Mon, Jan 12, 2009 at 03:27:44PM -0800, Junio C Hamano wrote:\n> > Kirill Smelkov <kirr@landau.phys.spbu.ru> writes:\n\n[...]\n\n> > But doesn't this\n> > \n> > >  From nobody Mon Sep 17 00:00:00 2001\n> > > -From: A\n> > > +From: A (zzz)\n> > >        U\n> > >        Thor\n> > > -      <a.u.thor@example.com>\n> > > +      <a.u.thor@example.com> (Comment)\n> > \n> > regress for people who spell their names like this?\n> > \n> > \tFrom: john.doe@email.xz (John Doe)\n> \n> I think everything is ok:\n[...]\n\nJust in case it got spam-detected again:\n\nhttp://marc.info/?l=git&m=123183962105146&w=2\n\n\nThanks,\nKirill\n"}]}