{"thread":{"id":"36942","subject":"[PATCH v4] http: fix charset detection of extract_content_type()","startedAt":"2014-06-17T22:11:53Z","lastAt":"2014-06-18T07:40:26Z","messageCount":2,"participants":["Yi EungJun","Jeff King"],"isPatch":true,"patchVersion":4,"patchTotal":null},"messages":[{"id":"244493","messageId":"1403043113-12579-1-git-send-email-eungjun.yi@navercorp.com","threadId":"36942","inReplyTo":null,"subject":"[PATCH v4] http: fix charset detection of extract_content_type()","fromName":"Yi EungJun","fromEmail":"semtlenori@gmail.com","sentAt":"2014-06-17T22:11:53Z","receivedAt":"2014-06-17T22:11:53Z","isPatch":true,"sender":{"key":"semtlenori@gmail.com","avatar":"https://gravatar.com/avatar/8363435d2badb3450df0dd7c4ec2113f6e1d62c44dd3cf7dfe83c5a389b8a9bf?d=mp&s=160"},"body":"From: Yi EungJun <eungjun.yi@navercorp.com>\n\nextract_content_type() could not extract a charset parameter if the\nparameter is not the first one and there is a whitespace and a following\nsemicolon just before the parameter. For example:\n\n    text/plain; format=fixed ;charset=utf-8\n\nAnd it also could not handle correctly some other cases, such as:\n\n    text/plain; charset=utf-8; format=fixed\n    text/plain; some-param=\"a long value with ;semicolons;\"; charset=utf-8\n\nThanks-to: Jeff King <peff@peff.net>\nSigned-off-by: Yi EungJun <eungjun.yi@navercorp.com>\n---\n http.c                     | 4 ++--\n t/lib-httpd/error.sh       | 4 ++++\n t/t5550-http-fetch-dumb.sh | 5 +++++\n 3 files changed, 11 insertions(+), 2 deletions(-)\n\ndiff --git a/http.c b/http.c\nindex 2b4f6a3..3a28b21 100644\n--- a/http.c\n+++ b/http.c\n@@ -927,7 +927,7 @@ static int extract_param(const char *raw, const char *name,\n \t\treturn -1;\n \traw++;\n \n-\twhile (*raw && !isspace(*raw))\n+\twhile (*raw && !isspace(*raw) && *raw != ';')\n \t\tstrbuf_addch(out, *raw++);\n \treturn 0;\n }\n@@ -971,7 +971,7 @@ static void extract_content_type(struct strbuf *raw, struct strbuf *type,\n \n \tstrbuf_reset(charset);\n \twhile (*p) {\n-\t\twhile (isspace(*p))\n+\t\twhile (isspace(*p) || *p == ';')\n \t\t\tp++;\n \t\tif (!extract_param(p, \"charset\", charset))\n \t\t\treturn;\ndiff --git a/t/lib-httpd/error.sh b/t/lib-httpd/error.sh\nindex eafc9d2..a77b8e5 100755\n--- a/t/lib-httpd/error.sh\n+++ b/t/lib-httpd/error.sh\n@@ -19,6 +19,10 @@ case \"$PATH_INFO\" in\n \tprintf \"text/plain; charset=utf-16\"\n \tcharset=utf-16\n \t;;\n+*odd-spacing*)\n+\tprintf \"text/plain; foo=bar ;charset=utf-16; other=nonsense\"\n+\tcharset=utf-16\n+\t;;\n esac\n printf \"\\n\"\n \ndiff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\nindex 01b8aae..ac71418 100755\n--- a/t/t5550-http-fetch-dumb.sh\n+++ b/t/t5550-http-fetch-dumb.sh\n@@ -191,5 +191,10 @@ test_expect_success 'http error messages are reencoded' '\n \tgrep \"this is the error message\" stderr\n '\n \n+test_expect_success 'reencoding is robust to whitespace oddities' '\n+\ttest_must_fail git clone \"$HTTPD_URL/error/odd-spacing\" 2>stderr &&\n+\tgrep \"this is the error message\" stderr\n+'\n+\n stop_httpd\n test_done\n-- \n2.0.0.422.gb6302de\n\nOops, I fixed the whitespace error.\n"},{"id":"244504","messageId":"20140618074026.GB24769@sigill.intra.peff.net","threadId":"36942","inReplyTo":"1403043113-12579-1-git-send-email-eungjun.yi@navercorp.com","subject":"Re: [PATCH v4] http: fix charset detection of extract_content_type()","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2014-06-18T07:40:26Z","receivedAt":"2014-06-18T07:40:26Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jun 18, 2014 at 07:11:53AM +0900, Yi EungJun wrote:\n\n> From: Yi EungJun <eungjun.yi@navercorp.com>\n> \n> extract_content_type() could not extract a charset parameter if the\n> parameter is not the first one and there is a whitespace and a following\n> semicolon just before the parameter. For example:\n> \n>     text/plain; format=fixed ;charset=utf-8\n> \n> And it also could not handle correctly some other cases, such as:\n> \n>     text/plain; charset=utf-8; format=fixed\n>     text/plain; some-param=\"a long value with ;semicolons;\"; charset=utf-8\n> \n> Thanks-to: Jeff King <peff@peff.net>\n> Signed-off-by: Yi EungJun <eungjun.yi@navercorp.com>\n\nThanks, this version looks good to me.\n\n-Peff\n"}]}