{"thread":{"id":"33085","subject":"[PATCH] format-patch: RFC 2047 says multi-octet character may not be split","startedAt":"2013-03-06T11:08:26Z","lastAt":"2013-03-10T07:05:26Z","messageCount":5,"participants":["Kirill Smelkov"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"210687","messageId":"1362568106-30741-1-git-send-email-kirr@mns.spb.ru","threadId":"33085","inReplyTo":null,"subject":"[PATCH] format-patch: RFC 2047 says multi-octet character may not be split","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2013-03-06T11:08:26Z","receivedAt":"2013-03-06T11:08:26Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Intro\n-----\n\nIn 'Subject:' characters are encoded in Q encoding, as per RFC 2047, e.g.\n\n    föö\n\nbecomes\n\n    =?UTF-8?q?f=C3=B6=C3=B6?=\n\n. Long encoded lines must be wrapped to be no longer than 76 bytes.\n\nAlso RFC 2047, section 5 (3) says:\n\n    Each 'encoded-word' MUST represent an integral number of\n    characters.  A multi-octet character may not be split across\n    adjacent 'encoded- word's.\n\nthat means that for\n\n    Subject: .... föö bar\n\nencoding\n\n    Subject: =?UTF-8?q?....=20f=C3=B6=C3=B6?=\n     =?UTF-8?q?=20bar?=\n\nis correct, and\n\n    Subject: =?UTF-8?q?....=20f=C3=B6=C3?=      <-- NOTE ö is broken here\n     =?UTF-8?q?=B6=20bar?=\n\nis not, because \"ö\" character UTF-8 encoding C3 B6 is split here across\nadjacent encoded words.\n\n~~~~\n\nAs it is now, format-patch does not respect \"multi-octet charactes may\nnot be split\" rule, and so sending patches with non-english subject has\nissues:\n\n    The problematic case shows in mail readers as \".... fö?? bar\".\n\nSolution\n--------\n\nI introduce mbs_chrlen() function to compute character length in bytes\nfor multi-byte text according to encoding, and use it appropriately in\nadd_rfc2047() in pretty.\n\nSo far it works correctly only for UTF-8 encoding, because we have\ninfrastructure for it in place already, but other encoding could be\nsupported too in the future with the help of iconv. For now they all, except\nUTF-8, are treated as being one-byte encodings, which was format-patch\ncurrent behaviour, with appropriate TODO put in mbs_chrlen().\n\nNot sure whether mbs_chrlen() is a good name, but otherwise my\nunderstanding is that the patch is ok to go in.\n\nThanks.\n\nCc: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n pretty.c                | 27 +++++++++++++++++++--------\n t/t4014-format-patch.sh | 27 ++++++++++++++-------------\n utf8.c                  | 39 +++++++++++++++++++++++++++++++++++++++\n utf8.h                  |  2 ++\n 4 files changed, 74 insertions(+), 21 deletions(-)\n\ndiff --git a/pretty.c b/pretty.c\nindex b57adef..c9c7ff5 100644\n--- a/pretty.c\n+++ b/pretty.c\n@@ -345,7 +345,7 @@ static int needs_rfc2047_encoding(const char *line, int len,\n \treturn 0;\n }\n \n-static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n+static void add_rfc2047(struct strbuf *sb, const char *line, size_t len,\n \t\t       const char *encoding, enum rfc2047_type type)\n {\n \tstatic const int max_encoded_length = 76; /* per rfc2047 */\n@@ -355,9 +355,18 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \tstrbuf_grow(sb, len * 3 + strlen(encoding) + 100);\n \tstrbuf_addf(sb, \"=?%s?q?\", encoding);\n \tline_len += strlen(encoding) + 5; /* 5 for =??q? */\n-\tfor (i = 0; i < len; i++) {\n-\t\tunsigned ch = line[i] & 0xFF;\n-\t\tint is_special = is_rfc2047_special(ch, type);\n+\n+\twhile (len) {\n+\t\t/*\n+\t\t * RFC 2047, section 5 (3):\n+\t\t *\n+\t\t * Each 'encoded-word' MUST represent an integral number of\n+\t\t * characters.  A multi-octet character may not be split across\n+\t\t * adjacent 'encoded- word's.\n+\t\t */\n+\t\tconst unsigned char *p = (const unsigned char *)line;\n+\t\tint chrlen = mbs_chrlen(&line, &len, encoding);\n+\t\tint is_special = (chrlen > 1) || is_rfc2047_special(*p, type);\n \n \t\t/*\n \t\t * According to RFC 2047, we could encode the special character\n@@ -367,16 +376,18 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \t\t * causes ' ' to be encoded as '=20', avoiding this problem.\n \t\t */\n \n-\t\tif (line_len + 2 + (is_special ? 3 : 1) > max_encoded_length) {\n+\t\tif (line_len + 2 + (is_special ? 3*chrlen : 1) > max_encoded_length) {\n \t\t\tstrbuf_addf(sb, \"?=\\n =?%s?q?\", encoding);\n \t\t\tline_len = strlen(encoding) + 5 + 1; /* =??q? plus SP */\n \t\t}\n \n \t\tif (is_special) {\n-\t\t\tstrbuf_addf(sb, \"=%02X\", ch);\n-\t\t\tline_len += 3;\n+\t\t\tfor (i = 0; i < chrlen; i++) {\n+\t\t\t\tstrbuf_addf(sb, \"=%02X\", p[i]);\n+\t\t\t\tline_len += 3;\n+\t\t\t}\n \t\t} else {\n-\t\t\tstrbuf_addch(sb, ch);\n+\t\t\tstrbuf_addch(sb, *p);\n \t\t\tline_len++;\n \t\t}\n \t}\ndiff --git a/t/t4014-format-patch.sh b/t/t4014-format-patch.sh\nindex 78633cb..b993dae 100755\n--- a/t/t4014-format-patch.sh\n+++ b/t/t4014-format-patch.sh\n@@ -837,25 +837,26 @@ Subject: [PATCH] =?UTF-8?q?f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n- =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar?=\n EOF\n test_expect_success 'format-patch wraps extremely long subject (rfc2047)' '\n \trm -rf patches/ &&\ndiff --git a/utf8.c b/utf8.c\nindex 8f6e84b..7911b58 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -531,3 +531,42 @@ char *reencode_string(const char *in, const char *out_encoding, const char *in_e\n \treturn out;\n }\n #endif\n+\n+/*\n+ * Returns first character length in bytes for multi-byte `text` according to\n+ * `encoding`.\n+ *\n+ * - The `text` pointer is updated to point at the next character.\n+ * - When `remainder_p` is not NULL, on entry `*remainder_p` is how much bytes\n+ *   we can consume from text, and on exit `*remainder_p` is reduced by returned\n+ *   character length. Otherwise `text` is treated as limited by NUL.\n+ */\n+int mbs_chrlen(const char **text, size_t *remainder_p, const char *encoding)\n+{\n+\tint chrlen;\n+\tconst char *p = *text;\n+\tsize_t r = (remainder_p ? *remainder_p : INT_MAX);\n+\n+\tif (r < 1)\n+\t\treturn 0;\n+\n+\tif (is_encoding_utf8(encoding)) {\n+\t\tpick_one_utf8_char(&p, &r);\n+\n+\t\tchrlen = p ? (p - *text)\n+\t\t\t   : 1 /* not valid UTF-8 -> raw byte sequence */;\n+\t}\n+\telse {\n+\t\t/* TODO use iconv to decode one char and obtain its chrlen\n+\t\t *\n+\t\t * for now, let's treat encodings != UTF-8 as one-byte\n+\t\t */\n+\t\tchrlen = 1;\n+\t}\n+\n+\t*text += chrlen;\n+\tif (remainder_p)\n+\t\t*remainder_p -= chrlen;\n+\n+\treturn chrlen;\n+}\ndiff --git a/utf8.h b/utf8.h\nindex 501b2bd..1f8ecad 100644\n--- a/utf8.h\n+++ b/utf8.h\n@@ -22,4 +22,6 @@ char *reencode_string(const char *in, const char *out_encoding, const char *in_e\n #define reencode_string(a,b,c) NULL\n #endif\n \n+int mbs_chrlen(const char **text, size_t *remainder_p, const char *encoding);\n+\n #endif\n-- \n1.8.2.rc2.353.gd2380b4\n"},{"id":"210761","messageId":"20130307105430.GA3049@tugrik.mns.mnsspb.ru","threadId":"33085","inReplyTo":"7vd2vcqv1y.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] format-patch: RFC 2047 says multi-octet character may not be split","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2013-03-07T10:55:07Z","receivedAt":"2013-03-07T10:55:07Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Junio,\n\nOn Wed, Mar 06, 2013 at 09:47:53AM -0800, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@mns.spb.ru> writes:\n> \n> > Intro\n> > -----\n> \n> Drop this.  We know the beginning part is \"intro\" already ;-)\n\n:)\n\n\n> >     Subject: .... föö bar\n> >\n> > encoding\n> >\n> >     Subject: =?UTF-8?q?....=20f=C3=B6=C3=B6?=\n> >      =?UTF-8?q?=20bar?=\n> >\n> > is correct, and\n> >\n> >     Subject: =?UTF-8?q?....=20f=C3=B6=C3?=      <-- NOTE ö is broken here\n> >      =?UTF-8?q?=B6=20bar?=\n> >\n> > is not, because \"ö\" character UTF-8 encoding C3 B6 is split here across\n> > adjacent encoded words.\n> \n> The above is an important part to keep in the log message.\n> Everything above that I snipped can be left out for brevity.\n> \n> > As it is now, format-patch does not respect \"multi-octet charactes may\n> > not be split\" rule, and so sending patches with non-english subject has\n> > issues:\n> >\n> >     The problematic case shows in mail readers as \".... fö?? bar\".\n> \n> But the log message lacks crucial bits of information before you\n> start talking about your solution.  Where does it go wrong?  What\n> did the earlier attempt bafc478..41dd00bad miss?  This can be fixed\n> trivially by replacing the above (and the \"solution\" section),\n> perhaps like this:\n> \n>     Even though an earlier attempt (bafc478..41dd00bad) cleaned\n>     up RFC 2047 encoding, pretty.c::add_rfc2047() still decides\n>     where to split the output line by going through the input\n>     one byte at a time, and potentially splits a character in\n>     the middle.  A subject line may end up showing like this:\n> \n>          The problematic case shows in mail readers as \".... fö?? bar\".\n> \n>     Instead, make the loop grab one _character_ at a time and\n>     determine its output length to see where to break the output\n>     line.  Note that this version only knows about UTF-8, but the\n>     logic to grab one character is abstracted out in mbs_chrlen()\n>     function to make it possible to extend it to other encodings.\n\nI agree my description was messy and thanks for reworking and clarifying\nit - your version is much better.\n\nI'll use its slight variation for the updated patch.\n\n\n> > +\twhile (len) {\n> > +\t\t/*\n> > +\t\t * RFC 2047, section 5 (3):\n> > +\t\t *\n> > +\t\t * Each 'encoded-word' MUST represent an integral number of\n> > +\t\t * characters.  A multi-octet character may not be split across\n> > +\t\t * adjacent 'encoded- word's.\n> > +\t\t */\n> > +\t\tconst unsigned char *p = (const unsigned char *)line;\n> > +\t\tint chrlen = mbs_chrlen(&line, &len, encoding);\n> > +\t\tint is_special = (chrlen > 1) || is_rfc2047_special(*p, type);\n> >  \n> >  \t\t/*\n> >  \t\t * According to RFC 2047, we could encode the special character\n> > @@ -367,16 +376,18 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n> >  \t\t * causes ' ' to be encoded as '=20', avoiding this problem.\n> >  \t\t */\n> >  \n> > +\t\tif (line_len + 2 + (is_special ? 3*chrlen : 1) > max_encoded_length) {\n> \n> Always have SP around binary operators such as '*' (multiplication).\n\nok, but note that's just a matter of style, and if one is used to code\nformulas, _not_ having SP is more convenient sometimes.\n\n\n> I would actually suggest adding an extra variable \"encoded_len\" and\n> do something like this:\n> \n> \t/* \"=%02X\" times num_char, or the byte itself */\n> \tencoded_len = is_special ? 3 * num_char : 1;\n>         if (max_encoded_length < line_len + 2 + encoded_len) {\n>         \t/* It will not fit---break the line */\n> \t\t...\n\nRight. Actually if we add encoded_len, adding encoded_fmt is tempting\n\n    const char *encoded_fmt = is_special ? \"=%02X\"    : \"%c\";\n\nand then encoding part simplifies to just unconditional\n\n    for (i = 0; i < chrlen; i++)\n            strbuf_addf(sb, encoded_fmt, p[i]);\n    line_len += encoded_len;\n\n\n> You may also want to say what the hardcoded \"2\" is about in the\n> comment there.\n\nok.\n\n\n> > diff --git a/utf8.c b/utf8.c\n> > index 8f6e84b..7911b58 100644\n> > --- a/utf8.c\n> > +++ b/utf8.c\n> > @@ -531,3 +531,42 @@ char *reencode_string(const char *in, const char *out_encoding, const char *in_e\n> >  \treturn out;\n> >  }\n> >  #endif\n> > +\n> > +/*\n> > + * Returns first character length in bytes for multi-byte `text` according to\n> > + * `encoding`.\n> > + *\n> > + * - The `text` pointer is updated to point at the next character.\n> > + * - When `remainder_p` is not NULL, on entry `*remainder_p` is how much bytes\n> > + *   we can consume from text, and on exit `*remainder_p` is reduced by returned\n> > + *   character length. Otherwise `text` is treated as limited by NUL.\n> > + */\n> > +int mbs_chrlen(const char **text, size_t *remainder_p, const char *encoding)\n> > +{\n> > +\tint chrlen;\n> > +\tconst char *p = *text;\n> > +\tsize_t r = (remainder_p ? *remainder_p : INT_MAX);\n> \n> Ugly, and more importantly I suspect this is wrong because size_t is\n> not signed and INT_MAX is.\n\nWhy is it ugly? There is similiar snippet in pick_one_utf8_char():\n\n        /*\n         * A caller that assumes NUL terminated text can choose\n         * not to bother with the remainder length.  We will\n         * stop at the first NUL.\n         */ \n        remainder = (remainder_p ? *remainder_p : 999);\n\nonly ad-hoc 999 is used there.\n\nI agree about INT_MAX being signed - my mistake - better change it to\nSIZE_MAX or ((size_t)-1) for portability, but otherwise the construct is\nimho ok. I'll change to SIZE_MAX since it is alredy used in Git.\n\nComputing r in the beginning simplifies following code.\n\n> > +\tif (r < 1)\n> > +\t\treturn 0;\n> > +\n> > +\tif (is_encoding_utf8(encoding)) {\n> > +\t\tpick_one_utf8_char(&p, &r);\n> > +\n> > +\t\tchrlen = p ? (p - *text)\n> > +\t\t\t   : 1 /* not valid UTF-8 -> raw byte sequence */;\n> > +\t}\n> > +\telse {\n> > +\t\t/* TODO use iconv to decode one char and obtain its chrlen\n> > +\t\t *\n> > +\t\t * for now, let's treat encodings != UTF-8 as one-byte\n> > +\t\t */\n> > +\t\tchrlen = 1;\n> \n> \t/*\n>          * We format our multi-line\n>          * comments like this\n>          */\n\nok, I agree.\n\n\n> Thanks.\n\nThanks too,\nKirill\n\n\nInterdiff and updated patch follows:\n\ndiff -u b/pretty.c b/pretty.c\n--- b/pretty.c\n+++ b/pretty.c\n@@ -369,4 +369,8 @@\n \t\tint is_special = (chrlen > 1) || is_rfc2047_special(*p, type);\n \n+\t\t/* \"=%02X\" * chrlen, or the byte itself */\n+\t\tconst char *encoded_fmt = is_special ? \"=%02X\"    : \"%c\";\n+\t\tint\t    encoded_len = is_special ? 3 * chrlen : 1;\n+\n \t\t/*\n \t\t * According to RFC 2047, we could encode the special character\n@@ -376,20 +380,15 @@\n \t\t * causes ' ' to be encoded as '=20', avoiding this problem.\n \t\t */\n \n-\t\tif (line_len + 2 + (is_special ? 3*chrlen : 1) > max_encoded_length) {\n+\t\tif (line_len + encoded_len + /* ?= */2 > max_encoded_length) {\n+\t\t\t/* It will not fit---break the line */\n \t\t\tstrbuf_addf(sb, \"?=\\n =?%s?q?\", encoding);\n \t\t\tline_len = strlen(encoding) + 5 + 1; /* =??q? plus SP */\n \t\t}\n \n-\t\tif (is_special) {\n-\t\t\tfor (i = 0; i < chrlen; i++) {\n-\t\t\t\tstrbuf_addf(sb, \"=%02X\", p[i]);\n-\t\t\t\tline_len += 3;\n-\t\t\t}\n-\t\t} else {\n-\t\t\tstrbuf_addch(sb, *p);\n-\t\t\tline_len++;\n-\t\t}\n+\t\tfor (i = 0; i < chrlen; i++)\n+\t\t\tstrbuf_addf(sb, encoded_fmt, p[i]);\n+\t\tline_len += encoded_len;\n \t}\n \tstrbuf_addstr(sb, \"?=\");\n }\ndiff -u b/utf8.c b/utf8.c\n--- b/utf8.c\n+++ b/utf8.c\n@@ -545,7 +545,7 @@\n {\n \tint chrlen;\n \tconst char *p = *text;\n-\tsize_t r = (remainder_p ? *remainder_p : INT_MAX);\n+\tsize_t r = (remainder_p ? *remainder_p : SIZE_MAX);\n \n \tif (r < 1)\n \t\treturn 0;\n@@ -557,8 +557,8 @@\n \t\t\t   : 1 /* not valid UTF-8 -> raw byte sequence */;\n \t}\n \telse {\n-\t\t/* TODO use iconv to decode one char and obtain its chrlen\n-\t\t *\n+\t\t/*\n+\t\t * TODO use iconv to decode one char and obtain its chrlen\n \t\t * for now, let's treat encodings != UTF-8 as one-byte\n \t\t */\n \t\tchrlen = 1;\n\n---- 8< ----\nFrom 46b9cddc63c07cb5513cfbf6d20aaaa98c66bcdf Mon Sep 17 00:00:00 2001\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\nDate: Wed, 6 Mar 2013 14:28:46 +0400\nSubject: [PATCH v2] format-patch: RFC 2047 says multi-octet character may not be split\n\nEven though an earlier attempt (bafc478..41dd00bad) cleaned\nup RFC 2047 encoding, pretty.c::add_rfc2047() still decides\nwhere to split the output line by going through the input\none byte at a time, and potentially splits a character in\nthe middle.  A subject line may end up showing like this:\n\n     \".... fö?? bar\".   (instead of  \".... föö bar\".)\n\nif split incorrectly.\n\nRFC 2047, section 5 (3) explicitly forbids such beaviour\n\n    Each 'encoded-word' MUST represent an integral number of\n    characters.  A multi-octet character may not be split across\n    adjacent 'encoded- word's.\n\nthat means that e.g. for\n\n    Subject: .... föö bar\n\nencoding\n\n    Subject: =?UTF-8?q?....=20f=C3=B6=C3=B6?=\n     =?UTF-8?q?=20bar?=\n\nis correct, and\n\n    Subject: =?UTF-8?q?....=20f=C3=B6=C3?=      <-- NOTE ö is broken here\n     =?UTF-8?q?=B6=20bar?=\n\nis not, because \"ö\" character UTF-8 encoding C3 B6 is split here across\nadjacent encoded words.\n\nTo fix the problem, make the loop grab one _character_ at a time and\ndetermine its output length to see where to break the output line.  Note\nthat this version only knows about UTF-8, but the logic to grab one\ncharacter is abstracted out in mbs_chrlen() function to make it possible\nto extend it to other encodings with the help of iconv in the future.\n\n(With help from Junio C Hamano <gitster@pobox.com>)\nCc: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n pretty.c                | 34 ++++++++++++++++++++++------------\n t/t4014-format-patch.sh | 27 ++++++++++++++-------------\n utf8.c                  | 39 +++++++++++++++++++++++++++++++++++++++\n utf8.h                  |  2 ++\n 4 files changed, 77 insertions(+), 25 deletions(-)\n\ndiff --git a/pretty.c b/pretty.c\nindex b57adef..c5fae69 100644\n--- a/pretty.c\n+++ b/pretty.c\n@@ -345,7 +345,7 @@ static int needs_rfc2047_encoding(const char *line, int len,\n \treturn 0;\n }\n \n-static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n+static void add_rfc2047(struct strbuf *sb, const char *line, size_t len,\n \t\t       const char *encoding, enum rfc2047_type type)\n {\n \tstatic const int max_encoded_length = 76; /* per rfc2047 */\n@@ -355,9 +355,22 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \tstrbuf_grow(sb, len * 3 + strlen(encoding) + 100);\n \tstrbuf_addf(sb, \"=?%s?q?\", encoding);\n \tline_len += strlen(encoding) + 5; /* 5 for =??q? */\n-\tfor (i = 0; i < len; i++) {\n-\t\tunsigned ch = line[i] & 0xFF;\n-\t\tint is_special = is_rfc2047_special(ch, type);\n+\n+\twhile (len) {\n+\t\t/*\n+\t\t * RFC 2047, section 5 (3):\n+\t\t *\n+\t\t * Each 'encoded-word' MUST represent an integral number of\n+\t\t * characters.  A multi-octet character may not be split across\n+\t\t * adjacent 'encoded- word's.\n+\t\t */\n+\t\tconst unsigned char *p = (const unsigned char *)line;\n+\t\tint chrlen = mbs_chrlen(&line, &len, encoding);\n+\t\tint is_special = (chrlen > 1) || is_rfc2047_special(*p, type);\n+\n+\t\t/* \"=%02X\" * chrlen, or the byte itself */\n+\t\tconst char *encoded_fmt = is_special ? \"=%02X\"    : \"%c\";\n+\t\tint\t    encoded_len = is_special ? 3 * chrlen : 1;\n \n \t\t/*\n \t\t * According to RFC 2047, we could encode the special character\n@@ -367,18 +380,15 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \t\t * causes ' ' to be encoded as '=20', avoiding this problem.\n \t\t */\n \n-\t\tif (line_len + 2 + (is_special ? 3 : 1) > max_encoded_length) {\n+\t\tif (line_len + encoded_len + /* ?= */2 > max_encoded_length) {\n+\t\t\t/* It will not fit---break the line */\n \t\t\tstrbuf_addf(sb, \"?=\\n =?%s?q?\", encoding);\n \t\t\tline_len = strlen(encoding) + 5 + 1; /* =??q? plus SP */\n \t\t}\n \n-\t\tif (is_special) {\n-\t\t\tstrbuf_addf(sb, \"=%02X\", ch);\n-\t\t\tline_len += 3;\n-\t\t} else {\n-\t\t\tstrbuf_addch(sb, ch);\n-\t\t\tline_len++;\n-\t\t}\n+\t\tfor (i = 0; i < chrlen; i++)\n+\t\t\tstrbuf_addf(sb, encoded_fmt, p[i]);\n+\t\tline_len += encoded_len;\n \t}\n \tstrbuf_addstr(sb, \"?=\");\n }\ndiff --git a/t/t4014-format-patch.sh b/t/t4014-format-patch.sh\nindex 78633cb..b993dae 100755\n--- a/t/t4014-format-patch.sh\n+++ b/t/t4014-format-patch.sh\n@@ -837,25 +837,26 @@ Subject: [PATCH] =?UTF-8?q?f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n- =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar?=\n EOF\n test_expect_success 'format-patch wraps extremely long subject (rfc2047)' '\n \trm -rf patches/ &&\ndiff --git a/utf8.c b/utf8.c\nindex 8f6e84b..7f64857 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -531,3 +531,42 @@ char *reencode_string(const char *in, const char *out_encoding, const char *in_e\n \treturn out;\n }\n #endif\n+\n+/*\n+ * Returns first character length in bytes for multi-byte `text` according to\n+ * `encoding`.\n+ *\n+ * - The `text` pointer is updated to point at the next character.\n+ * - When `remainder_p` is not NULL, on entry `*remainder_p` is how much bytes\n+ *   we can consume from text, and on exit `*remainder_p` is reduced by returned\n+ *   character length. Otherwise `text` is treated as limited by NUL.\n+ */\n+int mbs_chrlen(const char **text, size_t *remainder_p, const char *encoding)\n+{\n+\tint chrlen;\n+\tconst char *p = *text;\n+\tsize_t r = (remainder_p ? *remainder_p : SIZE_MAX);\n+\n+\tif (r < 1)\n+\t\treturn 0;\n+\n+\tif (is_encoding_utf8(encoding)) {\n+\t\tpick_one_utf8_char(&p, &r);\n+\n+\t\tchrlen = p ? (p - *text)\n+\t\t\t   : 1 /* not valid UTF-8 -> raw byte sequence */;\n+\t}\n+\telse {\n+\t\t/*\n+\t\t * TODO use iconv to decode one char and obtain its chrlen\n+\t\t * for now, let's treat encodings != UTF-8 as one-byte\n+\t\t */\n+\t\tchrlen = 1;\n+\t}\n+\n+\t*text += chrlen;\n+\tif (remainder_p)\n+\t\t*remainder_p -= chrlen;\n+\n+\treturn chrlen;\n+}\ndiff --git a/utf8.h b/utf8.h\nindex 501b2bd..1f8ecad 100644\n--- a/utf8.h\n+++ b/utf8.h\n@@ -22,4 +22,6 @@ char *reencode_string(const char *in, const char *out_encoding, const char *in_e\n #define reencode_string(a,b,c) NULL\n #endif\n \n+int mbs_chrlen(const char **text, size_t *remainder_p, const char *encoding);\n+\n #endif\n-- \n1.8.2.rc2.353.gd2380b4\n"},{"id":"210899","messageId":"20130309152722.GA32248@mini.zxlink","threadId":"33085","inReplyTo":"7vobevm6fp.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] format-patch: RFC 2047 says multi-octet character may not be split","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2013-03-09T15:27:23Z","receivedAt":"2013-03-09T15:27:23Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Thu, Mar 07, 2013 at 10:05:30AM -0800, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@mns.spb.ru> writes:\n> \n> >> > @@ -367,16 +376,18 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n> >> >  \t\t * causes ' ' to be encoded as '=20', avoiding this problem.\n> >> >  \t\t */\n> >> >  \n> >> > +\t\tif (line_len + 2 + (is_special ? 3*chrlen : 1) > max_encoded_length) {\n> >> \n> >> Always have SP around binary operators such as '*' (multiplication).\n> >\n> > ok, but note that's just a matter of style, and if one is used to code\n> > formulas,...\n> \n> Well, when working on a project with others, what _you_ are used to\n> does not matter.  Also please never call coding style \"just a matter\n> of\".  Keeping things consistent with the style around the area is a\n> prerequisite.\n> \n>    When you have time:\n>    Cf. https://www.youtube.com/watch?feature=player_embedded&v=fMeH7wqOwXA\n\nJunio, what Greg says here is all known and good and respected. I agree\ncoding style is not a \"just a matter of\" and is important to follow for\nproject to stay consisting. My note here was just a sentiment about\nspaces around operators, which I didn't know was in the coding style\nbecause it is not in Documentation/CodingGuidelines, and especially if\nthe project sometimes uses my style\n\n    *offset = 60*off;                                       date.c      5e2a78a4\n    diff += 7*n;                                            date.c      6b7b0427\n    sub_size < 2*window && i+1 < delta_search_threads       pack_objects.c  bf874896\n\n\nBut anyway, I'm ok with any style the project chooses - it's not so important\nfor me to insist here, so let it be \"3 * chrlen\" and lets forget about it.\n\n\n> > Actually if we add encoded_len, adding encoded_fmt is tempting\n> >\n> >     const char *encoded_fmt = is_special ? \"=%02X\"    : \"%c\";\n> >\n> > and then encoding part simplifies to just unconditional\n> >\n> >     for (i = 0; i < chrlen; i++)\n> >             strbuf_addf(sb, encoded_fmt, p[i]);\n> >     line_len += encoded_len;\n> \n> Sounds very sensible ;-)\n\nThanks.\n\n> >  \t\t * for now, let's treat encodings != UTF-8 as one-byte\n> >  \t\t */\n> >  \t\tchrlen = 1;\n> >\n> > ---- 8< ----\n> > From 46b9cddc63c07cb5513cfbf6d20aaaa98c66bcdf Mon Sep 17 00:00:00 2001\n> > From: Kirill Smelkov <kirr@mns.spb.ru>\n> > Date: Wed, 6 Mar 2013 14:28:46 +0400\n> > Subject: [PATCH v2] format-patch: RFC 2047 says multi-octet character may not be split\n> \n> Good use of scissors line; but please drop these four lines after\n> it.  The first is unwanted and the rest are redundant.\n> \n> > Even though an earlier attempt (bafc478..41dd00bad) cleaned\n> > up RFC 2047 encoding, pretty.c::add_rfc2047() still decides\n> > where to split the output line by going through the input\n> > ...\n> > @@ -367,18 +380,15 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n> >  \t\t * causes ' ' to be encoded as '=20', avoiding this problem.\n> >  \t\t */\n> >  \n> > -\t\tif (line_len + 2 + (is_special ? 3 : 1) > max_encoded_length) {\n> > +\t\tif (line_len + encoded_len + /* ?= */2 > max_encoded_length) {\n> > +\t\t\t/* It will not fit---break the line */\n> \n> It doesn't look much clearer with /* ?= */ unless we say something\n> that contains the word \"close\", e.g. \"?= to close the encoded part\".\n> Maybe it is just me.\n\nHow about\n\n        if (line_len + encoded_len + 2 > max_encoded_length) {\n                /* It won't fit with trailing \"?=\" --- break the line */\n\n?\n\n> \n> > -\t\tif (is_special) {\n> > -\t\t\tstrbuf_addf(sb, \"=%02X\", ch);\n> > -\t\t\tline_len += 3;\n> > -\t\t} else {\n> > -\t\t\tstrbuf_addch(sb, ch);\n> > -\t\t\tline_len++;\n> > -\t\t}\n> > +\t\tfor (i = 0; i < chrlen; i++)\n> > +\t\t\tstrbuf_addf(sb, encoded_fmt, p[i]);\n> > +\t\tline_len += encoded_len;\n> \n> Nice code reduction.\n\nThanks.\n\nInterdiff and updated patch follow. Note I'm sending this from home, so\n'From:' line after scissors is kept as necessary.\n\nKirill\n\nP.S. sorry for the delay - I harmed my arm yesterday.\n\n\ndiff --git a/pretty.c b/pretty.c\nindex 8fce619..41f04e6 100644\n--- a/pretty.c\n+++ b/pretty.c\n@@ -380,8 +380,8 @@ static void add_rfc2047(struct strbuf *sb, const char *line, size_t len,\n \t\t * causes ' ' to be encoded as '=20', avoiding this problem.\n \t\t */\n \n-\t\tif (line_len + encoded_len + /* ?= */2 > max_encoded_length) {\n-\t\t\t/* It will not fit---break the line */\n+\t\tif (line_len + encoded_len + 2 > max_encoded_length) {\n+\t\t\t/* It won't fit with trailing \"?=\" --- break the line */\n \t\t\tstrbuf_addf(sb, \"?=\\n =?%s?q?\", encoding);\n \t\t\tline_len = strlen(encoding) + 5 + 1; /* =??q? plus SP */\n \t\t}\n\n---- 8< ----\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\n split\n\nEven though an earlier attempt (bafc478..41dd00bad) cleaned\nup RFC 2047 encoding, pretty.c::add_rfc2047() still decides\nwhere to split the output line by going through the input\none byte at a time, and potentially splits a character in\nthe middle.  A subject line may end up showing like this:\n\n     \".... fö?? bar\".   (instead of  \".... föö bar\".)\n\nif split incorrectly.\n\nRFC 2047, section 5 (3) explicitly forbids such behaviour\n\n    Each 'encoded-word' MUST represent an integral number of\n    characters.  A multi-octet character may not be split across\n    adjacent 'encoded- word's.\n\nthat means that e.g. for\n\n    Subject: .... föö bar\n\nencoding\n\n    Subject: =?UTF-8?q?....=20f=C3=B6=C3=B6?=\n     =?UTF-8?q?=20bar?=\n\nis correct, and\n\n    Subject: =?UTF-8?q?....=20f=C3=B6=C3?=      <-- NOTE ö is broken here\n     =?UTF-8?q?=B6=20bar?=\n\nis not, because \"ö\" character UTF-8 encoding C3 B6 is split here across\nadjacent encoded words.\n\nTo fix the problem, make the loop grab one _character_ at a time and\ndetermine its output length to see where to break the output line.  Note\nthat this version only knows about UTF-8, but the logic to grab one\ncharacter is abstracted out in mbs_chrlen() function to make it possible\nto extend it to other encodings with the help of iconv in the future.\n\n(With help from Junio C Hamano <gitster@pobox.com>)\nCc: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n pretty.c                | 34 ++++++++++++++++++++++------------\n t/t4014-format-patch.sh | 27 ++++++++++++++-------------\n utf8.c                  | 39 +++++++++++++++++++++++++++++++++++++++\n utf8.h                  |  2 ++\n 4 files changed, 77 insertions(+), 25 deletions(-)\n\ndiff --git a/pretty.c b/pretty.c\nindex b57adef..41f04e6 100644\n--- a/pretty.c\n+++ b/pretty.c\n@@ -345,7 +345,7 @@ static int needs_rfc2047_encoding(const char *line, int len,\n \treturn 0;\n }\n \n-static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n+static void add_rfc2047(struct strbuf *sb, const char *line, size_t len,\n \t\t       const char *encoding, enum rfc2047_type type)\n {\n \tstatic const int max_encoded_length = 76; /* per rfc2047 */\n@@ -355,9 +355,22 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \tstrbuf_grow(sb, len * 3 + strlen(encoding) + 100);\n \tstrbuf_addf(sb, \"=?%s?q?\", encoding);\n \tline_len += strlen(encoding) + 5; /* 5 for =??q? */\n-\tfor (i = 0; i < len; i++) {\n-\t\tunsigned ch = line[i] & 0xFF;\n-\t\tint is_special = is_rfc2047_special(ch, type);\n+\n+\twhile (len) {\n+\t\t/*\n+\t\t * RFC 2047, section 5 (3):\n+\t\t *\n+\t\t * Each 'encoded-word' MUST represent an integral number of\n+\t\t * characters.  A multi-octet character may not be split across\n+\t\t * adjacent 'encoded- word's.\n+\t\t */\n+\t\tconst unsigned char *p = (const unsigned char *)line;\n+\t\tint chrlen = mbs_chrlen(&line, &len, encoding);\n+\t\tint is_special = (chrlen > 1) || is_rfc2047_special(*p, type);\n+\n+\t\t/* \"=%02X\" * chrlen, or the byte itself */\n+\t\tconst char *encoded_fmt = is_special ? \"=%02X\"    : \"%c\";\n+\t\tint\t    encoded_len = is_special ? 3 * chrlen : 1;\n \n \t\t/*\n \t\t * According to RFC 2047, we could encode the special character\n@@ -367,18 +380,15 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \t\t * causes ' ' to be encoded as '=20', avoiding this problem.\n \t\t */\n \n-\t\tif (line_len + 2 + (is_special ? 3 : 1) > max_encoded_length) {\n+\t\tif (line_len + encoded_len + 2 > max_encoded_length) {\n+\t\t\t/* It won't fit with trailing \"?=\" --- break the line */\n \t\t\tstrbuf_addf(sb, \"?=\\n =?%s?q?\", encoding);\n \t\t\tline_len = strlen(encoding) + 5 + 1; /* =??q? plus SP */\n \t\t}\n \n-\t\tif (is_special) {\n-\t\t\tstrbuf_addf(sb, \"=%02X\", ch);\n-\t\t\tline_len += 3;\n-\t\t} else {\n-\t\t\tstrbuf_addch(sb, ch);\n-\t\t\tline_len++;\n-\t\t}\n+\t\tfor (i = 0; i < chrlen; i++)\n+\t\t\tstrbuf_addf(sb, encoded_fmt, p[i]);\n+\t\tline_len += encoded_len;\n \t}\n \tstrbuf_addstr(sb, \"?=\");\n }\ndiff --git a/t/t4014-format-patch.sh b/t/t4014-format-patch.sh\nindex 78633cb..b993dae 100755\n--- a/t/t4014-format-patch.sh\n+++ b/t/t4014-format-patch.sh\n@@ -837,25 +837,26 @@ Subject: [PATCH] =?UTF-8?q?f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n- =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar?=\n EOF\n test_expect_success 'format-patch wraps extremely long subject (rfc2047)' '\n \trm -rf patches/ &&\ndiff --git a/utf8.c b/utf8.c\nindex 8f6e84b..7f64857 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -531,3 +531,42 @@ char *reencode_string(const char *in, const char *out_encoding, const char *in_e\n \treturn out;\n }\n #endif\n+\n+/*\n+ * Returns first character length in bytes for multi-byte `text` according to\n+ * `encoding`.\n+ *\n+ * - The `text` pointer is updated to point at the next character.\n+ * - When `remainder_p` is not NULL, on entry `*remainder_p` is how much bytes\n+ *   we can consume from text, and on exit `*remainder_p` is reduced by returned\n+ *   character length. Otherwise `text` is treated as limited by NUL.\n+ */\n+int mbs_chrlen(const char **text, size_t *remainder_p, const char *encoding)\n+{\n+\tint chrlen;\n+\tconst char *p = *text;\n+\tsize_t r = (remainder_p ? *remainder_p : SIZE_MAX);\n+\n+\tif (r < 1)\n+\t\treturn 0;\n+\n+\tif (is_encoding_utf8(encoding)) {\n+\t\tpick_one_utf8_char(&p, &r);\n+\n+\t\tchrlen = p ? (p - *text)\n+\t\t\t   : 1 /* not valid UTF-8 -> raw byte sequence */;\n+\t}\n+\telse {\n+\t\t/*\n+\t\t * TODO use iconv to decode one char and obtain its chrlen\n+\t\t * for now, let's treat encodings != UTF-8 as one-byte\n+\t\t */\n+\t\tchrlen = 1;\n+\t}\n+\n+\t*text += chrlen;\n+\tif (remainder_p)\n+\t\t*remainder_p -= chrlen;\n+\n+\treturn chrlen;\n+}\ndiff --git a/utf8.h b/utf8.h\nindex 501b2bd..1f8ecad 100644\n--- a/utf8.h\n+++ b/utf8.h\n@@ -22,4 +22,6 @@ char *reencode_string(const char *in, const char *out_encoding, const char *in_e\n #define reencode_string(a,b,c) NULL\n #endif\n \n+int mbs_chrlen(const char **text, size_t *remainder_p, const char *encoding);\n+\n #endif\n-- \n1.8.2.rc2.366.g3bc8dda\n"},{"id":"210901","messageId":"20130309153443.GB32248@mini.zxlink","threadId":"33085","inReplyTo":"20130309152722.GA32248@mini.zxlink","subject":"Re: [PATCH] format-patch: RFC 2047 says multi-octet character may not be split","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2013-03-09T15:34:43Z","receivedAt":"2013-03-09T15:34:43Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Sat, Mar 09, 2013 at 07:27:23PM +0400, Kirill Smelkov wrote:\n> ---- 8< ----\n> From: Kirill Smelkov <kirr@mns.spb.ru>\n>  split\n\nSorry for the confusion...\n\n---- 8< ----\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\n\nEven though an earlier attempt (bafc478..41dd00bad) cleaned\nup RFC 2047 encoding, pretty.c::add_rfc2047() still decides\nwhere to split the output line by going through the input\none byte at a time, and potentially splits a character in\nthe middle.  A subject line may end up showing like this:\n\n     \".... fö?? bar\".   (instead of  \".... föö bar\".)\n\nif split incorrectly.\n\nRFC 2047, section 5 (3) explicitly forbids such beaviour\n\n    Each 'encoded-word' MUST represent an integral number of\n    characters.  A multi-octet character may not be split across\n    adjacent 'encoded- word's.\n\nthat means that e.g. for\n\n    Subject: .... föö bar\n\nencoding\n\n    Subject: =?UTF-8?q?....=20f=C3=B6=C3=B6?=\n     =?UTF-8?q?=20bar?=\n\nis correct, and\n\n    Subject: =?UTF-8?q?....=20f=C3=B6=C3?=      <-- NOTE ö is broken here\n     =?UTF-8?q?=B6=20bar?=\n\nis not, because \"ö\" character UTF-8 encoding C3 B6 is split here across\nadjacent encoded words.\n\nTo fix the problem, make the loop grab one _character_ at a time and\ndetermine its output length to see where to break the output line.  Note\nthat this version only knows about UTF-8, but the logic to grab one\ncharacter is abstracted out in mbs_chrlen() function to make it possible\nto extend it to other encodings with the help of iconv in the future.\n\n(With help from Junio C Hamano <gitster@pobox.com>)\nCc: Jan H. Schönherr <schnhrr@cs.tu-berlin.de>\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n pretty.c                | 34 ++++++++++++++++++++++------------\n t/t4014-format-patch.sh | 27 ++++++++++++++-------------\n utf8.c                  | 39 +++++++++++++++++++++++++++++++++++++++\n utf8.h                  |  2 ++\n 4 files changed, 77 insertions(+), 25 deletions(-)\n\ndiff --git a/pretty.c b/pretty.c\nindex b57adef..41f04e6 100644\n--- a/pretty.c\n+++ b/pretty.c\n@@ -345,7 +345,7 @@ static int needs_rfc2047_encoding(const char *line, int len,\n \treturn 0;\n }\n \n-static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n+static void add_rfc2047(struct strbuf *sb, const char *line, size_t len,\n \t\t       const char *encoding, enum rfc2047_type type)\n {\n \tstatic const int max_encoded_length = 76; /* per rfc2047 */\n@@ -355,9 +355,22 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \tstrbuf_grow(sb, len * 3 + strlen(encoding) + 100);\n \tstrbuf_addf(sb, \"=?%s?q?\", encoding);\n \tline_len += strlen(encoding) + 5; /* 5 for =??q? */\n-\tfor (i = 0; i < len; i++) {\n-\t\tunsigned ch = line[i] & 0xFF;\n-\t\tint is_special = is_rfc2047_special(ch, type);\n+\n+\twhile (len) {\n+\t\t/*\n+\t\t * RFC 2047, section 5 (3):\n+\t\t *\n+\t\t * Each 'encoded-word' MUST represent an integral number of\n+\t\t * characters.  A multi-octet character may not be split across\n+\t\t * adjacent 'encoded- word's.\n+\t\t */\n+\t\tconst unsigned char *p = (const unsigned char *)line;\n+\t\tint chrlen = mbs_chrlen(&line, &len, encoding);\n+\t\tint is_special = (chrlen > 1) || is_rfc2047_special(*p, type);\n+\n+\t\t/* \"=%02X\" * chrlen, or the byte itself */\n+\t\tconst char *encoded_fmt = is_special ? \"=%02X\"    : \"%c\";\n+\t\tint\t    encoded_len = is_special ? 3 * chrlen : 1;\n \n \t\t/*\n \t\t * According to RFC 2047, we could encode the special character\n@@ -367,18 +380,15 @@ static void add_rfc2047(struct strbuf *sb, const char *line, int len,\n \t\t * causes ' ' to be encoded as '=20', avoiding this problem.\n \t\t */\n \n-\t\tif (line_len + 2 + (is_special ? 3 : 1) > max_encoded_length) {\n+\t\tif (line_len + encoded_len + 2 > max_encoded_length) {\n+\t\t\t/* It won't fit with trailing \"?=\" --- break the line */\n \t\t\tstrbuf_addf(sb, \"?=\\n =?%s?q?\", encoding);\n \t\t\tline_len = strlen(encoding) + 5 + 1; /* =??q? plus SP */\n \t\t}\n \n-\t\tif (is_special) {\n-\t\t\tstrbuf_addf(sb, \"=%02X\", ch);\n-\t\t\tline_len += 3;\n-\t\t} else {\n-\t\t\tstrbuf_addch(sb, ch);\n-\t\t\tline_len++;\n-\t\t}\n+\t\tfor (i = 0; i < chrlen; i++)\n+\t\t\tstrbuf_addf(sb, encoded_fmt, p[i]);\n+\t\tline_len += encoded_len;\n \t}\n \tstrbuf_addstr(sb, \"?=\");\n }\ndiff --git a/t/t4014-format-patch.sh b/t/t4014-format-patch.sh\nindex 78633cb..b993dae 100755\n--- a/t/t4014-format-patch.sh\n+++ b/t/t4014-format-patch.sh\n@@ -837,25 +837,26 @@ Subject: [PATCH] =?UTF-8?q?f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n  =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n  =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n  =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n- =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3?=\n- =?UTF-8?q?=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n- =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3?=\n- =?UTF-8?q?=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n- =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6?=\n+ =?UTF-8?q?=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6?=\n+ =?UTF-8?q?=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f?=\n+ =?UTF-8?q?=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar?=\n+ =?UTF-8?q?=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20bar=20f=C3=B6=C3=B6=20?=\n+ =?UTF-8?q?bar?=\n EOF\n test_expect_success 'format-patch wraps extremely long subject (rfc2047)' '\n \trm -rf patches/ &&\ndiff --git a/utf8.c b/utf8.c\nindex 8f6e84b..7f64857 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -531,3 +531,42 @@ char *reencode_string(const char *in, const char *out_encoding, const char *in_e\n \treturn out;\n }\n #endif\n+\n+/*\n+ * Returns first character length in bytes for multi-byte `text` according to\n+ * `encoding`.\n+ *\n+ * - The `text` pointer is updated to point at the next character.\n+ * - When `remainder_p` is not NULL, on entry `*remainder_p` is how much bytes\n+ *   we can consume from text, and on exit `*remainder_p` is reduced by returned\n+ *   character length. Otherwise `text` is treated as limited by NUL.\n+ */\n+int mbs_chrlen(const char **text, size_t *remainder_p, const char *encoding)\n+{\n+\tint chrlen;\n+\tconst char *p = *text;\n+\tsize_t r = (remainder_p ? *remainder_p : SIZE_MAX);\n+\n+\tif (r < 1)\n+\t\treturn 0;\n+\n+\tif (is_encoding_utf8(encoding)) {\n+\t\tpick_one_utf8_char(&p, &r);\n+\n+\t\tchrlen = p ? (p - *text)\n+\t\t\t   : 1 /* not valid UTF-8 -> raw byte sequence */;\n+\t}\n+\telse {\n+\t\t/*\n+\t\t * TODO use iconv to decode one char and obtain its chrlen\n+\t\t * for now, let's treat encodings != UTF-8 as one-byte\n+\t\t */\n+\t\tchrlen = 1;\n+\t}\n+\n+\t*text += chrlen;\n+\tif (remainder_p)\n+\t\t*remainder_p -= chrlen;\n+\n+\treturn chrlen;\n+}\ndiff --git a/utf8.h b/utf8.h\nindex 501b2bd..1f8ecad 100644\n--- a/utf8.h\n+++ b/utf8.h\n@@ -22,4 +22,6 @@ char *reencode_string(const char *in, const char *out_encoding, const char *in_e\n #define reencode_string(a,b,c) NULL\n #endif\n \n+int mbs_chrlen(const char **text, size_t *remainder_p, const char *encoding);\n+\n #endif\n-- \n1.8.2.rc2.366.g3bc8dda\n"},{"id":"210957","messageId":"20130310070525.GA4020@vasy.zxlink","threadId":"33085","inReplyTo":"7vzjyce6jc.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] format-patch: RFC 2047 says multi-octet character may not be split","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2013-03-10T07:05:26Z","receivedAt":"2013-03-10T07:05:26Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Sat, Mar 09, 2013 at 11:07:19AM -0800, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n> > P.S. sorry for the delay - I harmed my arm yesterday.\n> \n> Ouch. Take care and be well soon.\n\nThanks, and thanks fr accepting the patch.\n"}]}