{"thread":{"id":"26480","subject":"a bug about format-patch of multibyte characters comment","startedAt":"2011-02-12T10:13:15Z","lastAt":"2011-02-13T10:50:19Z","messageCount":13,"participants":["xiaozhu","Martin Krüger","Jeff King","Johannes Sixt","xzer"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"160943","messageId":"4D565D3B.7060808@gmail.com","threadId":"26480","inReplyTo":null,"subject":"a bug about format-patch of multibyte characters comment","fromName":"xiaozhu","fromEmail":"xiaozhu@gmail.com","sentAt":"2011-02-12T10:13:15Z","receivedAt":"2011-02-12T10:13:15Z","isPatch":false,"sender":{"key":"xiaozhu@gmail.com","avatar":null},"body":"Hi,\n\nI found a bug when I use format-patch to export a patch which contains comment with\nsome multibyte characters. I also found the relation source, but I can't understand\nthe source clearly, so I think I need a help to know how can I fix it.\n\nAt first, the symptom.\n\nI commit a fix to my repository with comment like following:\n-----------------------------------------------------\nXXXXXXXXXXXX\nYYYYYY\n-----------------------------------------------------\n\ntwo lines of multibyte language comment.\n\nthen I use format-patch to export this fix, I get a patch file like following:\n\n------------------------------------------------------------------------------\n From d3532c3263a02a2367a3aa5c9cc3f0bd738b79b1 Mon Sep 17 00:00:00 2001\nFrom: xz <xz>\nDate: Fri, 11 Feb 2011 21:30:35 +0900\nSubject: [PATCH] =?UTF-8?q?=E6=97=A5=E6=9C=AC=E8=AA=9E=E3=81=8C=E5=A4=A7=E4=B8=88=E5=A4=AB\n=20=E6=94=B9=E8=A1=8C=E3=81=99=E3=82=8B?=\nMIME-Version: 1.0\nContent-Type: text/plain; charset=UTF-8\nContent-Transfer-Encoding: 8bit\n\n---\n  testfile.txt |    4 +++-\n  1 files changed, 3 insertions(+), 1 deletions(-)\n\ndiff --git a/testfile.txt b/testfile.txt\nindex 1e5d832..da982fd 100644\n--- a/testfile.txt\n+++ b/testfile.txt\n@@ -1 +1,3 @@\n-sadfasdf\n\\ No newline at end of file\n+sadfasdf\n\n..........\n\n-------------------------------------------------------------------------------\n\nIf I use am to apply this patch, am can't analyze the comment correctly, then the\ncommitted comment will become\n\"=?UTF-8?q?=E6=97=A5=E6=9C=AC=E8=AA=9E=E3=81=8C=E5=A4=A7=E4=B8=88=E5=A4=AB\".\n\nAbove is the symptom.\n\nThen I did some try, I modify the comment to 3 lines:\n-----------------------------------------------------\nXXXXXXXXXXXX\n\nYYYYYY\n-----------------------------------------------------\n\nadd a empty line, then I get a patch like following:\n------------------------------------------------------------------------------\n From d3532c3263a02a2367a3aa5c9cc3f0bd738b79b1 Mon Sep 17 00:00:00 2001\nFrom: xz <xz>\nDate: Fri, 11 Feb 2011 21:30:35 +0900\nSubject: [PATCH] =?UTF-8?q?=E6=97=A5=E6=9C=AC=E8=AA=9E=E3=81=8C=E5=A4=A7=E4=B8=88=E5=A4=AB?=\nMIME-Version: 1.0\nContent-Type: text/plain; charset=UTF-8\nContent-Transfer-Encoding: 8bit\n\nYYYYYY\n---\n  testfile.txt |    4 +++-\n  1 files changed, 3 insertions(+), 1 deletions(-)\n\ndiff --git a/testfile.txt b/testfile.txt\nindex 1e5d832..da982fd 100644\n--- a/testfile.txt\n+++ b/testfile.txt\n@@ -1 +1,3 @@\n-sadfasdf\n\\ No newline at end of file\n+sadfasdf\n\n..........\n\n-------------------------------------------------------------------------------\n\nthis patch will be applied successfully. So I know the problem is about the subject creating.\nI search the source, then I found the following function at \"pretty.c:655\":\n\nconst char *format_subject(struct strbuf *sb, const char *msg,\n\t\t\t   const char *line_separator)\n{\n\tint first = 1;\n\n\tfor (;;) {\n\t\tconst char *line = msg;\n\t\tint linelen = get_one_line(line);\n\n\t\tmsg += linelen;\n\t\t\n\t\tif (!linelen || is_empty_line(line, &linelen))\n\t\t\tbreak;\n\n\t\tif (!sb)cat\n\t\t\tcontinue;\n\t\tstrbuf_grow(sb, linelen + 2);\n\t\tif (!first)\n\t\t\tstrbuf_addstr(sb, line_separator);\n\t\tstrbuf_add(sb, line, linelen);\n\t\tfirst = 0;\n\t}\n\treturn msg;\n}\n\nAt first I want to know: Does this function means that always add the first line\nof comment to the argument sb, then return the rest? Is there any other thing that I\ndidn't considered?\n\nI found 4 place where to call this function, I think there is no problem about 3\nof them, but I don't know is there any other problem to the rest one which is\nat \"pretty.c:931\".\n\nAt last, if what I think is correct, I plan to fix it as following:\n\nconst char *format_subject(struct strbuf *sb, const char *msg,\n\t\t\t   const char *line_separator)\n{\n\tint first = 1;\n\n\t//for (;;) {\n\t\tconst char *line = msg;\n\t\tint linelen = get_one_line(line);\n\n\t\tmsg += linelen;\n\t\t\n\t\tif (!linelen || is_empty_line(line, &linelen)) return msg;\n\t\t\t//break;\n\n\t\tif (!sb) return msg;\n\t\t\t//continue;\n\t\tstrbuf_grow(sb, linelen + 2);\n\t\tif (!first)\n\t\t\tstrbuf_addstr(sb, line_separator);\n\t\tstrbuf_add(sb, line, linelen);\n\t\tfirst = 0;\n\t//}\n\treturn msg;\n}\n\nI dont't think it is necessary to have a loop here, so I want to remove\nthe loop. Is there anybody can confirm my fix for me?\n"},{"id":"160947","messageId":"20110212123026.238490@gmx.net","threadId":"26480","inReplyTo":"4D565D3B.7060808@gmail.com","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"Martin Krüger","fromEmail":"martin.krueger@gmx.com","sentAt":"2011-02-12T12:30:26Z","receivedAt":"2011-02-12T12:30:26Z","isPatch":false,"sender":{"key":"martin.krueger@gmx.com","avatar":null},"body":"Hi\n\nI recently stumbled over the same problem:\n\nhttp://thread.gmane.org/gmane.comp.version-control.git/162792/focus=162797\n\n\nbest regards.\n\nmartin\n"},{"id":"160978","messageId":"20110213075337.GA12112@sigill.intra.peff.net","threadId":"26480","inReplyTo":"4D565D3B.7060808@gmail.com","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-02-13T07:53:37Z","receivedAt":"2011-02-13T07:53:37Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Feb 12, 2011 at 07:13:15PM +0900, xiaozhu wrote:\n\n> I commit a fix to my repository with comment like following:\n> -----------------------------------------------------\n> XXXXXXXXXXXX\n> YYYYYY\n> -----------------------------------------------------\n> \n> two lines of multibyte language comment.\n\nOK.\n\n> then I use format-patch to export this fix, I get a patch file like following:\n> \n> ------------------------------------------------------------------------------\n> From d3532c3263a02a2367a3aa5c9cc3f0bd738b79b1 Mon Sep 17 00:00:00 2001\n> From: xz <xz>\n> Date: Fri, 11 Feb 2011 21:30:35 +0900\n> Subject: [PATCH] =?UTF-8?q?=E6=97=A5=E6=9C=AC=E8=AA=9E=E3=81=8C=E5=A4=A7=E4=B8=88=E5=A4=AB\n> =20=E6=94=B9=E8=A1=8C=E3=81=99=E3=82=8B?=\n\nYeah, this is wrong. There should be a whitespace indentation in a\nmulti-line header, or the whole thing should be on one line. The newline\nin your commit subject is apparently leaking through, and it should be\nqp-encoded.\n\n> const char *format_subject(struct strbuf *sb, const char *msg,\n> \t\t\t   const char *line_separator)\n> {\n> \tint first = 1;\n> \n> \tfor (;;) {\n> \t\tconst char *line = msg;\n> \t\tint linelen = get_one_line(line);\n> \n> \t\tmsg += linelen;\n> \t\t\n> \t\tif (!linelen || is_empty_line(line, &linelen))\n> \t\t\tbreak;\n> \n> \t\tif (!sb)cat\n> \t\t\tcontinue;\n> \t\tstrbuf_grow(sb, linelen + 2);\n> \t\tif (!first)\n> \t\t\tstrbuf_addstr(sb, line_separator);\n> \t\tstrbuf_add(sb, line, linelen);\n> \t\tfirst = 0;\n> \t}\n> \treturn msg;\n> }\n> \n> At first I want to know: Does this function means that always add the first line\n> of comment to the argument sb, then return the rest? Is there any other thing that I\n> didn't considered?\n\nNo. It adds all lines in the first paragraph into sb. Which, if you\nfollow the usual git convention, is a single line. But we also see\ncommits imported from other version control systems where that\nconvention is not as prevalent. In practice, using the whole first\nparagraph seems to make the most sense as a subject.\n\n> At last, if what I think is correct, I plan to fix it as following:\n> \n> const char *format_subject(struct strbuf *sb, const char *msg,\n> \t\t\t   const char *line_separator)\n> {\n> \tint first = 1;\n> \n> \t//for (;;) {\n> \t\tconst char *line = msg;\n> \t\tint linelen = get_one_line(line);\n> \n> \t\tmsg += linelen;\n> \t\t\n> \t\tif (!linelen || is_empty_line(line, &linelen)) return msg;\n> \t\t\t//break;\n> \n> \t\tif (!sb) return msg;\n> \t\t\t//continue;\n> \t\tstrbuf_grow(sb, linelen + 2);\n> \t\tif (!first)\n> \t\t\tstrbuf_addstr(sb, line_separator);\n> \t\tstrbuf_add(sb, line, linelen);\n> \t\tfirst = 0;\n> \t//}\n> \treturn msg;\n> }\n\nNo, that's not right. It breaks the use-whole-paragraph feature. The\nright fix is to encode the embedded newline in the subject properly.\n\n-Peff\n"},{"id":"160979","messageId":"20110213083137.GB12112@sigill.intra.peff.net","threadId":"26480","inReplyTo":"20110213075337.GA12112@sigill.intra.peff.net","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-02-13T08:31:37Z","receivedAt":"2011-02-13T08:31:37Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 13, 2011 at 02:53:37AM -0500, Jeff King wrote:\n\n> No, that's not right. It breaks the use-whole-paragraph feature. The\n> right fix is to encode the embedded newline in the subject properly.\n\nHrm. It is actually a little more complex. Right now for a subject like:\n\n  one\n  two\n  three\n\nwe will convert that to the subject \"one two three\" for \"git log\n--oneline\" output. But for email messages, we actually return \"one\\n\ntwo\\n three\", which is conflating header folding with the actual\nconstruction of the title.\n\nShouldn't we still be generating \"one two three\", encoding it via\nrfc2047 if necessary, and _then_ deciding if folding is required? Yes,\nindividual lines in a multi-line subject are good candidates for\nfolding, but don't we need to be checking for and folding long lines\nanyway?\n\n-Peff\n"},{"id":"160981","messageId":"4D579A35.1000007@gmail.com","threadId":"26480","inReplyTo":"20110213083137.GB12112@sigill.intra.peff.net","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"xiaozhu","fromEmail":"xiaozhu@gmail.com","sentAt":"2011-02-13T08:45:41Z","receivedAt":"2011-02-13T08:45:41Z","isPatch":false,"sender":{"key":"xiaozhu@gmail.com","avatar":null},"body":"> Shouldn't we still be generating \"one two three\", encoding it via\n> rfc2047 if necessary, and _then_ deciding if folding is required? Yes,\n> individual lines in a multi-line subject are good candidates for\n> folding, but don't we need to be checking for and folding long lines\n> anyway?\n\nIt seems that by rfc2047 there is no multi-line subject spec. A subject\nwith multi-line will be always conflated to one single line. And also\nthat if we just generate the subject within multi-line just like the\ncurrent implemention, yes, we can modify the git-am to decode it correctly,\nbut most of the mail client will can not show it correctly.\n\nSo it seems that there is only one way that combining the whole first\nparagraph to a single line? But it will be a nightmare for some long comment.\n"},{"id":"160982","messageId":"20110213085236.GA2251@sigill.intra.peff.net","threadId":"26480","inReplyTo":"4D579A35.1000007@gmail.com","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-02-13T08:52:36Z","receivedAt":"2011-02-13T08:52:36Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 13, 2011 at 05:45:41PM +0900, xiaozhu wrote:\n\n> >Shouldn't we still be generating \"one two three\", encoding it via\n> >rfc2047 if necessary, and _then_ deciding if folding is required? Yes,\n> >individual lines in a multi-line subject are good candidates for\n> >folding, but don't we need to be checking for and folding long lines\n> >anyway?\n> \n> It seems that by rfc2047 there is no multi-line subject spec. A subject\n> with multi-line will be always conflated to one single line.\n\nSorry, I don't quite parse what you're saying. If the header takes up\nmultiple lines, then yes, that gets decoded as a single line by rfc822\nheader folding. I would then expect that result to be rfc2047-decoded if\nnecessary, and in theory it could contain encoded newlines.\n\n> And also that if we just generate the subject within multi-line just\n> like the current implemention, yes, we can modify the git-am to decode\n> it correctly, but most of the mail client will can not show it\n> correctly.\n\nAgain, I don't quite understand what you're saying. The output generated\nby format-patch now is _not_ valid according to rfc2822. Changing git-am\nto parse its bogus output won't help that.\n\n> So it seems that there is only one way that combining the whole first\n> paragraph to a single line? But it will be a nightmare for some long comment.\n\nIt's not the only way, but it is how we treat multi-line subjects in all\nother parts of git, so it is at least consistent (and that behavior was\nagreed upon after seeing what is worse: truncating to a single line, or\nmerging lines).\n\n-Peff\n"},{"id":"160987","messageId":"201102131048.58111.j6t@kdbg.org","threadId":"26480","inReplyTo":"20110213075337.GA12112@sigill.intra.peff.net","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2011-02-13T09:48:58Z","receivedAt":"2011-02-13T09:48:58Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"On Sonntag, 13. Februar 2011, Jeff King wrote:\n> On Sat, Feb 12, 2011 at 07:13:15PM +0900, xiaozhu wrote:\n> > Subject: [PATCH]\n> > =?UTF-8?q?=E6=97=A5=E6=9C=AC=E8=AA=9E=E3=81=8C=E5=A4=A7=E4=B8=88=E5=A4=AB\n> > =20=E6=94=B9=E8=A1=8C=E3=81=99=E3=82=8B?=\n>\n> Yeah, this is wrong. There should be a whitespace indentation in a\n> multi-line header, or the whole thing should be on one line. The newline\n> in your commit subject is apparently leaking through, and it should be\n> qp-encoded.\n\nIsn't it wrong that format-patch (and --pretty=email) does any quoting in the \nfirst place? Isn't it the task of the MUA (git-send-email) to do the quoting?\n\nFor example, when I import a format-patch generated patch into a mail message, \nI don't want to see:\n\n From: =?UTF-8?q?Joh=C3=A4nnes=20S=C3=BCxt?= <me@localhost>\n\nbut rather:\n\n From: Johännes Süxt <me@localhost>\n\nI know this is a bit late, but this really comes as a surprise. (I've never \nhad to pay attention to this behavior in the past...)\n\n-- Hannes\n"},{"id":"160988","messageId":"20110213100349.GA5483@sigill.intra.peff.net","threadId":"26480","inReplyTo":"201102131048.58111.j6t@kdbg.org","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-02-13T10:03:50Z","receivedAt":"2011-02-13T10:03:50Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 13, 2011 at 10:48:58AM +0100, Johannes Sixt wrote:\n\n> On Sonntag, 13. Februar 2011, Jeff King wrote:\n> > On Sat, Feb 12, 2011 at 07:13:15PM +0900, xiaozhu wrote:\n> > > Subject: [PATCH]\n> > > =?UTF-8?q?=E6=97=A5=E6=9C=AC=E8=AA=9E=E3=81=8C=E5=A4=A7=E4=B8=88=E5=A4=AB\n> > > =20=E6=94=B9=E8=A1=8C=E3=81=99=E3=82=8B?=\n> >\n> > Yeah, this is wrong. There should be a whitespace indentation in a\n> > multi-line header, or the whole thing should be on one line. The newline\n> > in your commit subject is apparently leaking through, and it should be\n> > qp-encoded.\n> \n> Isn't it wrong that format-patch (and --pretty=email) does any quoting in the \n> first place? Isn't it the task of the MUA (git-send-email) to do the quoting?\n\nI dunno. I guess it depends on how you view the output of format-patch.\nIs it an mbox containing rfc2822-valid messages? Then it ought to be\nquoted.\n\nCertainly when opening the output in an editor, it is prettier to have\nit unquoted. But what do MUAs expect when opening such an mbox? mutt\nseems to handle unencoded utf8 just fine, but I don't know about other\nMUAs.\n\n> For example, when I import a format-patch generated patch into a mail message, \n> I don't want to see:\n> \n>  From: =?UTF-8?q?Joh=C3=A4nnes=20S=C3=BCxt?= <me@localhost>\n> \n> but rather:\n> \n>  From: Johännes Süxt <me@localhost>\n> \n> I know this is a bit late, but this really comes as a surprise. (I've never \n> had to pay attention to this behavior in the past...)\n\nI can see how that would be annoying if your workflow is to paste\nformat-patch output into an email. But it has been that way literally\nfor years (I just tried v1.5.0, and format-patch produces\nrfc2047-encoded headers). So yes, you are a bit late. :)\n\n-Peff\n"},{"id":"160992","messageId":"4D57AEFC.10608@gmail.com","threadId":"26480","inReplyTo":"20110213085236.GA2251@sigill.intra.peff.net","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"xiaozhu","fromEmail":"xiaozhu@gmail.com","sentAt":"2011-02-13T10:14:20Z","receivedAt":"2011-02-13T10:14:20Z","isPatch":false,"sender":{"key":"xiaozhu@gmail.com","avatar":null},"body":"\n\nOn 2011/02/13 17:52, Jeff King wrote:\n> On Sun, Feb 13, 2011 at 05:45:41PM +0900, xiaozhu wrote:\n>\n>>> Shouldn't we still be generating \"one two three\", encoding it via\n>>> rfc2047 if necessary, and _then_ deciding if folding is required? Yes,\n>>> individual lines in a multi-line subject are good candidates for\n>>> folding, but don't we need to be checking for and folding long lines\n>>> anyway?\n>>\n>> It seems that by rfc2047 there is no multi-line subject spec. A subject\n>> with multi-line will be always conflated to one single line.\n>\n> Sorry, I don't quite parse what you're saying. If the header takes up\n> multiple lines, then yes, that gets decoded as a single line by rfc822\n> header folding. I would then expect that result to be rfc2047-decoded if\n> necessary, and in theory it could contain encoded newlines.\n\nI am not similar with mail format. I read the rfc2047 again, but I didn't\nsee any description about line separator encoding. Perhaps a base64 encoded-word\nwill contain the line separator involuntarily? I also found a sample in\nrfc2047 it show us a line broken subject mail, but it didn't say any thing\nabout line separator encoding.\n\n>> And also that if we just generate the subject within multi-line just\n>> like the current implemention, yes, we can modify the git-am to decode\n>> it correctly, but most of the mail client will can not show it\n>> correctly.\n>\n> Again, I don't quite understand what you're saying. The output generated\n> by format-patch now is _not_ valid according to rfc2822. Changing git-am\n> to parse its bogus output won't help that.\n>\n>> So it seems that there is only one way that combining the whole first\n>> paragraph to a single line? But it will be a nightmare for some long comment.\n>\n> It's not the only way, but it is how we treat multi-line subjects in all\n> other parts of git, so it is at least consistent (and that behavior was\n> agreed upon after seeing what is worse: truncating to a single line, or\n> merging lines).\n>\n> -Peff\n\nA sample of rfc2047 show us a legal line broken subject mail, like following:\n------------------------------------------------------------------\n  Subject: =?ISO-8859-1?B?SWYgeW91IGNhbiByZWFkIHRoaXMgeW8=?=\n     =?ISO-8859-2?B?dSB1bmRlcnN0YW5kIHRoZSBleGFtcGxlLg==?=\n------------------------------------------------------------------\n\nI understand that the current format-patch is not not valid to rfc2822/rfc2047,\nbut even a valid one just like above, most of the mail client will can not show it\ncorrectly, they show the first line only, I think that's a problem of user\nfriendliness.\n\n-xzer\n"},{"id":"160993","messageId":"AANLkTi=Ty22nzd6ja=XmMzMu+YzDKDSBMCOGRfKenhR4@mail.gmail.com","threadId":"26480","inReplyTo":"4D57AEFC.10608@gmail.com","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"xzer","fromEmail":"xiaozhu@gmail.com","sentAt":"2011-02-13T10:22:14Z","receivedAt":"2011-02-13T10:22:14Z","isPatch":false,"sender":{"key":"xiaozhu@gmail.com","avatar":null},"body":"> A sample of rfc2047 show us a legal line broken subject mail, like\n> following:\n> ------------------------------------------------------------------\n>  Subject: =?ISO-8859-1?B?SWYgeW91IGNhbiByZWFkIHRoaXMgeW8=?=\n>    =?ISO-8859-2?B?dSB1bmRlcnN0YW5kIHRoZSBleGFtcGxlLg==?=\n> ------------------------------------------------------------------\n>\n> I understand that the current format-patch is not not valid to\n> rfc2822/rfc2047,\n> but even a valid one just like above, most of the mail client will can not\n> show it\n> correctly, they show the first line only, I think that's a problem of user\n> friendliness.\n\nI am sorry I made a mistake, there is no problem of mail client, I just\ncreate a wrong format to test. So now I think if we can generate a valid\nrfc2047 patch file, and then make the am also analyze the patch file\ncorrectly, there is no problem. Isn't it?\n\n-xzer\n"},{"id":"160994","messageId":"20110213102307.GA7735@sigill.intra.peff.net","threadId":"26480","inReplyTo":"4D57AEFC.10608@gmail.com","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-02-13T10:23:08Z","receivedAt":"2011-02-13T10:23:08Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 13, 2011 at 07:14:20PM +0900, xiaozhu wrote:\n\n> I am not similar with mail format. I read the rfc2047 again, but I didn't\n> see any description about line separator encoding. Perhaps a base64 encoded-word\n> will contain the line separator involuntarily? I also found a sample in\n> rfc2047 it show us a line broken subject mail, but it didn't say any thing\n> about line separator encoding.\n\nIf the encoded-word contains the line separator, wouldn't it be a\nliteral character in the header value then? I.e., using rfc2047 you can\nembed a literal newline into your subject (or possibly, a 0x0a byte may\nbe part of a multi-byte character; that can't happen in utf8, but I\nbelieve it can in utf16).\n\n> A sample of rfc2047 show us a legal line broken subject mail, like following:\n> ------------------------------------------------------------------\n>  Subject: =?ISO-8859-1?B?SWYgeW91IGNhbiByZWFkIHRoaXMgeW8=?=\n>     =?ISO-8859-2?B?dSB1bmRlcnN0YW5kIHRoZSBleGFtcGxlLg==?=\n> ------------------------------------------------------------------\n> \n> I understand that the current format-patch is not not valid to rfc2822/rfc2047,\n> but even a valid one just like above, most of the mail client will can not show it\n> correctly, they show the first line only, I think that's a problem of user\n> friendliness.\n\nThen those mail clients are broken. That should first be unfolded to put\nboth encoded-words on the same line (separated by whitespace, I think,\nthough it doesn't matter for encoded words), and then each encoded word\nshould be decoded (the resulting text is \"If you can read this you\nunderstand the example.\").\n\nMutt and other MUAs do this just fine. If you have a MUA that doesn't,\ncomplain to the author.\n\n-Peff\n"},{"id":"160995","messageId":"20110213102627.GB7735@sigill.intra.peff.net","threadId":"26480","inReplyTo":"AANLkTi=Ty22nzd6ja=XmMzMu+YzDKDSBMCOGRfKenhR4@mail.gmail.com","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-02-13T10:26:27Z","receivedAt":"2011-02-13T10:26:27Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 13, 2011 at 07:22:14PM +0900, xzer wrote:\n\n> > A sample of rfc2047 show us a legal line broken subject mail, like\n> > following:\n> > ------------------------------------------------------------------\n> >  Subject: =?ISO-8859-1?B?SWYgeW91IGNhbiByZWFkIHRoaXMgeW8=?=\n> >    =?ISO-8859-2?B?dSB1bmRlcnN0YW5kIHRoZSBleGFtcGxlLg==?=\n> > ------------------------------------------------------------------\n> >\n> > I understand that the current format-patch is not not valid to\n> > rfc2822/rfc2047,\n> > but even a valid one just like above, most of the mail client will can not\n> > show it\n> > correctly, they show the first line only, I think that's a problem of user\n> > friendliness.\n> \n> I am sorry I made a mistake, there is no problem of mail client, I just\n> create a wrong format to test. So now I think if we can generate a valid\n> rfc2047 patch file, and then make the am also analyze the patch file\n> correctly, there is no problem. Isn't it?\n\nAh, OK, our mails just crossed paths. git-am already parses this\ncorrectly (actually, it is git-mailinfo that parses on behalf of\ngit-am). We just need format-patch to generate it (and it should also\nprobably be folding long lines in general).\n\n-Peff\n"},{"id":"160997","messageId":"4D57B76B.6000600@gmail.com","threadId":"26480","inReplyTo":"20110213102627.GB7735@sigill.intra.peff.net","subject":"Re: a bug about format-patch of multibyte characters comment","fromName":"xiaozhu","fromEmail":"xiaozhu@gmail.com","sentAt":"2011-02-13T10:50:19Z","receivedAt":"2011-02-13T10:50:19Z","isPatch":false,"sender":{"key":"xiaozhu@gmail.com","avatar":null},"body":"> Ah, OK, our mails just crossed paths. git-am already parses this\n> correctly (actually, it is git-mailinfo that parses on behalf of\n> git-am). We just need format-patch to generate it (and it should also\n> probably be folding long lines in general).\n\nI tested it, yes, the git-am parses it \"correctly\". There is only\none problem that we lost the line break after import a patch. Is\nthere a way that we can retain the line break after import the\npatch?\n\n-xzer\n"}]}