{"thread":{"id":"14398","subject":"[PATCH] git-mailinfo: Fix getting the subject from the body","startedAt":"2008-07-10T21:41:33Z","lastAt":"2008-07-15T03:13:56Z","messageCount":14,"participants":["Lukas Sandström","Junio C Hamano","Don Zickus"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"82852","messageId":"4876820D.4070806@etek.chalmers.se","threadId":"14398","inReplyTo":null,"subject":"[PATCH] git-mailinfo: Fix getting the subject from the body","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2008-07-10T21:41:33Z","receivedAt":"2008-07-10T21:41:33Z","isPatch":true,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"\"Subject: \" isn't in the static array \"header\", and thus\nmemcmp(\"Subject: \", header[i], 7) will never match.\n\nSigned-off-by: Lukas Sandström <lukass@etek.chalmers.se>\n---\n\nThis has been broken since 2007-03-12, with commit\n87ab799234639c26ea10de74782fa511cb3ca606\nso it might not be very important.\n\n builtin-mailinfo.c |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 962aa34..2d1520f 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -334,7 +334,7 @@ static int check_header(char *line, unsigned linesize, char **hdr_data, int over\n \t\treturn 1;\n \tif (!memcmp(\"[PATCH]\", line, 7) && isspace(line[7])) {\n \t\tfor (i = 0; header[i]; i++) {\n-\t\t\tif (!memcmp(\"Subject: \", header[i], 9)) {\n+\t\t\tif (!memcmp(\"Subject\", header[i], 7)) {\n \t\t\t\tif (! handle_header(line, hdr_data[i], 0)) {\n \t\t\t\t\treturn 1;\n \t\t\t\t}\n-- \n1.5.4.5\n"},{"id":"82858","messageId":"48768F30.8070409@etek.chalmers.se","threadId":"14398","inReplyTo":"7vod55o0tx.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] git-mailinfo: Fix getting the subject from the body","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2008-07-10T22:37:36Z","receivedAt":"2008-07-10T22:37:36Z","isPatch":true,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Junio C Hamano wrote:\n> Lukas Sandström <lukass@etek.chalmers.se> writes:\n> \n>> \"Subject: \" isn't in the static array \"header\", and thus\n>> memcmp(\"Subject: \", header[i], 7) will never match.\n>>\n>> Signed-off-by: Lukas Sandström <lukass@etek.chalmers.se>\n>> ---\n>>\n>> This has been broken since 2007-03-12, with commit\n>> 87ab799234639c26ea10de74782fa511cb3ca606\n>> so it might not be very important.\n> \n> Wow, thanks.  Perhaps we would want some additional test scripts?\n\nYes, apparently.\n\nAfter looking at this part some more, I see that there is no guarantee\nthat hdr_data[i] != NULL in this codepath, and then we won't use the\nsubject anyway.\n\nI'm currently working on rewriting git-mailinfo to use strbuf's insted\nof the preallocated buffers currently used. Do you want me to send a\npatch allocating hdr_data[i], or can you wait for my strbuf-conversion\npatch? It'll propably be ready for review tonight.\n\n/Lukas\n"},{"id":"82893","messageId":"7v3amhnwy9.fsf@gitster.siamese.dyndns.org","threadId":"14398","inReplyTo":"48768F30.8070409@etek.chalmers.se","subject":"Re: [PATCH] git-mailinfo: Fix getting the subject from the body","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-07-10T23:25:34Z","receivedAt":"2008-07-10T23:25:34Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lukas Sandström <lukass@etek.chalmers.se> writes:\n\n> I'm currently working on rewriting git-mailinfo to use strbuf's insted\n> of the preallocated buffers currently used. Do you want me to send a\n> patch allocating hdr_data[i], or can you wait for my strbuf-conversion\n> patch? It'll propably be ready for review tonight.\n\nHeh, after getting burned by that NUL thingy, I was waiting for somebody\nto step up.  Thanks.\n"},{"id":"82896","messageId":"48769E40.8030303@etek.chalmers.se","threadId":"14398","inReplyTo":"7v3amhnwy9.fsf@gitster.siamese.dyndns.org","subject":"[PATCH] Add some useful functions for strbuf manipulation.","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2008-07-10T23:41:52Z","receivedAt":"2008-07-10T23:41:52Z","isPatch":true,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Signed-off-by: Lukas Sandström <lukass@etek.chalmers.se>\n---\n\nJunio C Hamano wrote:\n> Heh, after getting burned by that NUL thingy, I was waiting for somebody\n> to step up.  Thanks.\n\nHere we go then. Two freshly baked patches.\n\nNote that this is a pretty straight buffers -> strbuf's conversion,\nbut I think it will be a good start for further hardening of mailinfo.\n\n\n strbuf.c |   70 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n strbuf.h |    6 +++++\n 2 files changed, 76 insertions(+), 0 deletions(-)\n\ndiff --git a/strbuf.c b/strbuf.c\nindex 4aed752..28d6776 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -60,6 +60,18 @@ void strbuf_grow(struct strbuf *sb, size_t extra)\n \tALLOC_GROW(sb->buf, sb->len + extra + 1, sb->alloc);\n }\n \n+void strbuf_trim(struct strbuf *sb)\n+{\n+\tchar *b = sb->buf;\n+\twhile (sb->len > 0 && isspace((unsigned char)sb->buf[sb->len - 1]))\n+\t\tsb->len--;\n+\twhile(sb->len > 0 && isspace(*b)) {\n+\t\tb++;\n+\t\tsb->len--;\n+\t}\n+\tmemmove(sb->buf, b, sb->len);\n+\tsb->buf[sb->len] = '\\0';\n+}\n void strbuf_rtrim(struct strbuf *sb)\n {\n \twhile (sb->len > 0 && isspace((unsigned char)sb->buf[sb->len - 1]))\n@@ -67,6 +79,64 @@ void strbuf_rtrim(struct strbuf *sb)\n \tsb->buf[sb->len] = '\\0';\n }\n \n+void strbuf_ltrim(struct strbuf *sb)\n+{\n+\tchar *b = sb->buf;\n+\twhile(sb->len > 0 && isspace(*b)) {\n+\t\tb++;\n+\t\tsb->len--;\n+\t}\n+\tmemmove(sb->buf, b, sb->len);\n+\tsb->buf[sb->len] = '\\0';\n+}\n+\n+void strbuf_tolower(struct strbuf *sb)\n+{\n+\tint i;\n+\tfor (i = 0; i < sb->len; i++)\n+\t\tsb->buf[i] = tolower(sb->buf[i]);\n+}\n+\n+struct strbuf ** strbuf_split(struct strbuf *sb, int delim)\n+{\n+\tint alloc = 2, pos = 0;\n+\tchar *n, *p;\n+\tstruct strbuf **ret;\n+\tstruct strbuf *t;\n+\n+\tret = xcalloc(alloc, sizeof(struct strbuf *));\n+\tp = n = sb->buf;\n+\twhile (n < sb->buf + sb->len) {\n+\t\tint len;\n+\t\tn = memchr(n, delim, sb->len - (n - sb->buf));\n+\t\tif (pos + 1 >= alloc) {\n+\t\t\talloc = alloc * 2;\n+\t\t\tret = xrealloc(ret, sizeof(struct strbuf *) * alloc);\n+\t\t}\n+\t\tif (!n)\n+\t\t\tn = sb->buf + sb->len - 1;\n+\t\tlen = n - p + 1;\n+\t\tt = xmalloc(sizeof(struct strbuf));\n+\t\tstrbuf_init(t, len);\n+\t\tstrbuf_add(t, p, len);\n+\t\tret[pos] = t;\n+\t\tret[++pos] = NULL;\n+\t\tp = ++n;\n+\t}\n+\treturn ret;\n+}\n+\n+void strbuf_list_free(struct strbuf ** sbs)\n+{\n+\tstruct strbuf **s = sbs;\n+\n+\twhile(*s) {\n+\t\tstrbuf_release(*s);\n+\t\tfree(*s++);\n+\t}\n+\tfree(sbs);\n+}\n+\n int strbuf_cmp(struct strbuf *a, struct strbuf *b)\n {\n \tint cmp;\ndiff --git a/strbuf.h b/strbuf.h\nindex faec229..577d14e 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -77,8 +77,14 @@ static inline void strbuf_setlen(struct strbuf *sb, size_t len) {\n #define strbuf_reset(sb)  strbuf_setlen(sb, 0)\n \n /*----- content related -----*/\n+extern void strbuf_trim(struct strbuf *);\n extern void strbuf_rtrim(struct strbuf *);\n+extern void strbuf_ltrim(struct strbuf *);\n extern int strbuf_cmp(struct strbuf *, struct strbuf *);\n+extern void strbuf_tolower(struct strbuf *);\n+\n+extern struct strbuf ** strbuf_split(struct strbuf*, int delim);\n+extern void strbuf_list_free(struct strbuf **);\n \n /*----- add data in your buffer -----*/\n static inline void strbuf_addch(struct strbuf *sb, int c) {\n-- \n1.5.4.5\n"},{"id":"82897","messageId":"48769E91.60205@etek.chalmers.se","threadId":"14398","inReplyTo":"48769E40.8030303@etek.chalmers.se","subject":"[PATCH/RFC] git-mailinfo: use strbuf's instead of fixed buffers","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2008-07-10T23:43:13Z","receivedAt":"2008-07-10T23:43:13Z","isPatch":true,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Signed-off-by: Lukas Sandström <lukass@etek.chalmers.se>\n---\n builtin-mailinfo.c |  705 +++++++++++++++++++++++++---------------------------\n 1 files changed, 333 insertions(+), 372 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 2d1520f..254a97c 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -5,14 +5,15 @@\n #include \"cache.h\"\n #include \"builtin.h\"\n #include \"utf8.h\"\n+#include \"strbuf.h\"\n \n static FILE *cmitmsg, *patchfile, *fin, *fout;\n \n static int keep_subject;\n static const char *metainfo_charset;\n-static char line[1000];\n-static char name[1000];\n-static char email[1000];\n+static struct strbuf line = STRBUF_INIT;\n+static struct strbuf name = STRBUF_INIT;\n+static struct strbuf email = STRBUF_INIT;\n \n static enum  {\n \tTE_DONTCARE, TE_QP, TE_BASE64,\n@@ -21,74 +22,74 @@ static enum  {\n \tTYPE_TEXT, TYPE_OTHER,\n } message_type;\n \n-static char charset[256];\n+static struct strbuf charset = STRBUF_INIT;\n static int patch_lines;\n-static char **p_hdr_data, **s_hdr_data;\n+static struct strbuf **p_hdr_data, **s_hdr_data;\n \n #define MAX_HDR_PARSED 10\n #define MAX_BOUNDARIES 5\n \n-static char *sanity_check(char *name, char *email)\n+static void sanity_check(struct strbuf *out, struct strbuf *name, struct strbuf *email)\n {\n-\tint len = strlen(name);\n-\tif (len < 3 || len > 60)\n-\t\treturn email;\n-\tif (strchr(name, '@') || strchr(name, '<') || strchr(name, '>'))\n-\t\treturn email;\n-\treturn name;\n+\tstruct strbuf o = STRBUF_INIT;\n+\tif (name->len < 3 || name->len > 60)\n+\t\tstrbuf_addbuf(&o, email);\n+\tif (strchr(name->buf, '@') || strchr(name->buf, '<') ||\n+\t\tstrchr(name->buf, '>'))\n+\t\tstrbuf_addbuf(&o, email);\n+\tstrbuf_addbuf(&o, name);\n+\tstrbuf_reset(out);\n+\tstrbuf_addbuf(out, &o);\n+\tstrbuf_release(&o);\n }\n \n-static int bogus_from(char *line)\n+static int bogus_from(const struct strbuf *line)\n {\n \t/* John Doe <johndoe> */\n-\tchar *bra, *ket, *dst, *cp;\n \n+\tchar *bra, *ket;\n \t/* This is fallback, so do not bother if we already have an\n \t * e-mail address.\n \t */\n-\tif (*email)\n+\tif (email.len)\n \t\treturn 0;\n \n-\tbra = strchr(line, '<');\n+\tbra = strchr(line->buf, '<');\n \tif (!bra)\n \t\treturn 0;\n \tket = strchr(bra, '>');\n \tif (!ket)\n \t\treturn 0;\n \n-\tfor (dst = email, cp = bra+1; cp < ket; )\n-\t\t*dst++ = *cp++;\n-\t*dst = 0;\n-\tfor (cp = line; isspace(*cp); cp++)\n-\t\t;\n-\tfor (bra--; isspace(*bra); bra--)\n-\t\t*bra = 0;\n-\tcp = sanity_check(cp, email);\n-\tstrcpy(name, cp);\n+\tstrbuf_reset(&email);\n+\tstrbuf_add(&email, bra + 1, ket - bra - 1);\n+\n+\tstrbuf_reset(&name);\n+\tstrbuf_add(&name, line->buf, bra - line->buf);\n+\tstrbuf_trim(&name);\n+\tsanity_check(&name, &name, &email);\n \treturn 1;\n }\n \n-static int handle_from(char *in_line)\n+static int handle_from(struct strbuf *from)\n {\n-\tchar line[1000];\n \tchar *at;\n-\tchar *dst;\n+\tsize_t el;\n \n-\tstrcpy(line, in_line);\n-\tat = strchr(line, '@');\n+\tat = strchr(from->buf, '@');\n \tif (!at)\n-\t\treturn bogus_from(line);\n+\t\treturn bogus_from(from);\n \n \t/*\n \t * If we already have one email, don't take any confusing lines\n \t */\n-\tif (*email && strchr(at+1, '@'))\n+\tif (email.len && strchr(at + 1, '@'))\n \t\treturn 0;\n \n \t/* Pick up the string around '@', possibly delimited with <>\n-\t * pair; that is the email part.  White them out while copying.\n+\t * pair; that is the email part.\n \t */\n-\twhile (at > line) {\n+\twhile (at > from->buf) {\n \t\tchar c = at[-1];\n \t\tif (isspace(c))\n \t\t\tbreak;\n@@ -98,56 +99,35 @@ static int handle_from(char *in_line)\n \t\t}\n \t\tat--;\n \t}\n-\tdst = email;\n-\tfor (;;) {\n-\t\tunsigned char c = *at;\n-\t\tif (!c || c == '>' || isspace(c)) {\n-\t\t\tif (c == '>')\n-\t\t\t\t*at = ' ';\n-\t\t\tbreak;\n-\t\t}\n-\t\t*at++ = ' ';\n-\t\t*dst++ = c;\n-\t}\n-\t*dst++ = 0;\n-\n+\tel = strcspn(at, \" \\n\\t\\r\\v\\f>\");\n+\tstrbuf_reset(&email);\n+\tstrbuf_add(&email, at, el);\n+\tstrbuf_remove(from, at - from->buf, el + 1);\n \t/* The remainder is name.  It could be \"John Doe <john.doe@xz>\"\n \t * or \"john.doe@xz (John Doe)\", but we have whited out the\n \t * email part, so trim from both ends, possibly removing\n \t * the () pair at the end.\n \t */\n-\tat = line + strlen(line);\n-\twhile (at > line) {\n-\t\tunsigned char c = *--at;\n-\t\tif (!isspace(c)) {\n-\t\t\tat[(c == ')') ? 0 : 1] = 0;\n-\t\t\tbreak;\n-\t\t}\n-\t}\n \n-\tat = line;\n-\tfor (;;) {\n-\t\tunsigned char c = *at;\n-\t\tif (!c || !isspace(c)) {\n-\t\t\tif (c == '(')\n-\t\t\t\tat++;\n-\t\t\tbreak;\n-\t\t}\n-\t\tat++;\n-\t}\n-\tat = sanity_check(at, email);\n-\tstrcpy(name, at);\n+\tstrbuf_trim(from);\n+\tif (*from->buf == '(')\n+\t\tstrbuf_remove(&name, 0, 1);\n+\tif (*(from->buf + from->len - 1) == ')')\n+\t\tstrbuf_setlen(from, from->len - 1);\n+\n+\tsanity_check(&name, from, &email);\n \treturn 1;\n }\n \n-static int handle_header(char *line, char *data, int ofs)\n+static void handle_header(struct strbuf **out, const struct strbuf *line)\n {\n-\tif (!line || !data)\n-\t\treturn 1;\n-\n-\tstrcpy(data, line+ofs);\n+\tif (!*out) {\n+\t\t*out = xmalloc(sizeof(struct strbuf));\n+\t\tstrbuf_init(*out, line->len);\n+\t} else\n+\t\tstrbuf_reset(*out);\n \n-\treturn 0;\n+\tstrbuf_addbuf(*out, (struct strbuf *)line); /* const warning */\n }\n \n /* NOTE NOTE NOTE.  We do not claim we do full MIME.  We just attempt\n@@ -156,13 +136,13 @@ static int handle_header(char *line, char *data, int ofs)\n  * case insensitively.\n  */\n \n-static int slurp_attr(const char *line, const char *name, char *attr)\n+static int slurp_attr(const char *line, const char *name, struct strbuf *attr)\n {\n \tconst char *ends, *ap = strcasestr(line, name);\n \tsize_t sz;\n \n \tif (!ap) {\n-\t\t*attr = 0;\n+\t\tstrbuf_setlen(attr, 0);\n \t\treturn 0;\n \t}\n \tap += strlen(name);\n@@ -173,180 +153,176 @@ static int slurp_attr(const char *line, const char *name, char *attr)\n \telse\n \t\tends = \"; \\t\";\n \tsz = strcspn(ap, ends);\n-\tmemcpy(attr, ap, sz);\n-\tattr[sz] = 0;\n+\tstrbuf_add(attr, ap, sz);\n \treturn 1;\n }\n \n struct content_type {\n-\tchar *boundary;\n-\tint boundary_len;\n+\tstruct strbuf *boundary;\n };\n \n static struct content_type content[MAX_BOUNDARIES];\n \n static struct content_type *content_top = content;\n \n-static int handle_content_type(char *line)\n+static int handle_content_type(struct strbuf *line)\n {\n-\tchar boundary[256];\n+\tstruct strbuf *boundary = xmalloc(sizeof(struct strbuf));\n+\tstrbuf_init(boundary, line->len);\n \n-\tif (strcasestr(line, \"text/\") == NULL)\n+\tif (!strcasestr(line->buf, \"text/\"))\n \t\t message_type = TYPE_OTHER;\n-\tif (slurp_attr(line, \"boundary=\", boundary + 2)) {\n-\t\tmemcpy(boundary, \"--\", 2);\n+\tif (slurp_attr(line->buf, \"boundary=\", boundary)) {\n+\t\tstrbuf_insert(boundary, 0, \"--\", 2);\n \t\tif (content_top++ >= &content[MAX_BOUNDARIES]) {\n \t\t\tfprintf(stderr, \"Too many boundaries to handle\\n\");\n \t\t\texit(1);\n \t\t}\n-\t\tcontent_top->boundary_len = strlen(boundary);\n-\t\tcontent_top->boundary = xmalloc(content_top->boundary_len+1);\n-\t\tstrcpy(content_top->boundary, boundary);\n-\t}\n-\tif (slurp_attr(line, \"charset=\", charset)) {\n-\t\tint i, c;\n-\t\tfor (i = 0; (c = charset[i]) != 0; i++)\n-\t\t\tcharset[i] = tolower(c);\n+\t\tcontent_top->boundary = boundary;\n+\t} else {\n+\t\tstrbuf_release(boundary);\n+\t\tfree(boundary);\n \t}\n+\tif (slurp_attr(line->buf, \"charset=\", &charset))\n+\t\tstrbuf_tolower(&charset);\n \treturn 0;\n }\n \n-static int handle_content_transfer_encoding(char *line)\n+static int handle_content_transfer_encoding(struct strbuf *line)\n {\n-\tif (strcasestr(line, \"base64\"))\n+\tif (strcasestr(line->buf, \"base64\"))\n \t\ttransfer_encoding = TE_BASE64;\n-\telse if (strcasestr(line, \"quoted-printable\"))\n+\telse if (strcasestr(line->buf, \"quoted-printable\"))\n \t\ttransfer_encoding = TE_QP;\n \telse\n \t\ttransfer_encoding = TE_DONTCARE;\n \treturn 0;\n }\n \n-static int is_multipart_boundary(const char *line)\n-{\n-\treturn (!memcmp(line, content_top->boundary, content_top->boundary_len));\n-}\n-\n-static int eatspace(char *line)\n+static int is_multipart_boundary(struct strbuf *line)\n {\n-\tint len = strlen(line);\n-\twhile (len > 0 && isspace(line[len-1]))\n-\t\tline[--len] = 0;\n-\treturn len;\n+\treturn !strbuf_cmp(line, content_top->boundary);\n }\n \n-static char *cleanup_subject(char *subject)\n+static void cleanup_subject(struct strbuf *subject)\n {\n-\tfor (;;) {\n-\t\tchar *p;\n-\t\tint len, remove;\n-\t\tswitch (*subject) {\n+\tchar *pos;\n+\tsize_t remove;\n+\twhile (subject->len) {\n+\t\tswitch (*subject->buf) {\n \t\tcase 'r': case 'R':\n-\t\t\tif (!memcmp(\"e:\", subject+1, 2)) {\n-\t\t\t\tsubject += 3;\n+\t\t\tif (subject->len <= 3)\n+\t\t\t\tbreak;\n+\t\t\tif (!memcmp(subject->buf + 1, \"e:\", 2)) {\n+\t\t\t\tstrbuf_remove(subject, 0, 3);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tbreak;\n \t\tcase ' ': case '\\t': case ':':\n-\t\t\tsubject++;\n+\t\t\tstrbuf_remove(subject, 0, 1);\n \t\t\tcontinue;\n-\n+\t\t\tbreak;\n \t\tcase '[':\n-\t\t\tp = strchr(subject, ']');\n-\t\t\tif (!p) {\n-\t\t\t\tsubject++;\n-\t\t\t\tcontinue;\n-\t\t\t}\n-\t\t\tlen = strlen(p);\n-\t\t\tremove = p - subject;\n-\t\t\tif (remove <= len *2) {\n-\t\t\t\tsubject = p+1;\n-\t\t\t\tcontinue;\n-\t\t\t}\n+\t\t\tif ((pos = strchr(subject->buf, ']'))) {\n+\t\t\t\tremove = pos - subject->buf + 1;\n+\t\t\t\t/* Don't remove too much. */\n+\t\t\t\tif (remove <= (subject->len - remove + 1) * 2) {\n+\t\t\t\t\tstrbuf_remove(subject, 0, remove);\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n+\t\t\t} else\n+\t\t\t\tstrbuf_remove(subject, 0, 1);\n \t\t\tbreak;\n \t\t}\n-\t\teatspace(subject);\n-\t\treturn subject;\n+\t\tstrbuf_trim(subject);\n+\t\treturn;\n \t}\n }\n \n-static void cleanup_space(char *buf)\n+static void cleanup_space(struct strbuf *sb)\n {\n-\tunsigned char c;\n-\twhile ((c = *buf) != 0) {\n-\t\tbuf++;\n-\t\tif (isspace(c)) {\n-\t\t\tbuf[-1] = ' ';\n-\t\t\tc = *buf;\n-\t\t\twhile (isspace(c)) {\n-\t\t\t\tint len = strlen(buf);\n-\t\t\t\tmemmove(buf, buf+1, len);\n-\t\t\t\tc = *buf;\n-\t\t\t}\n+\tsize_t pos, cnt;\n+\tfor (pos = 0; pos < sb->len; pos++) {\n+\t\tif (isspace(sb->buf[pos])) {\n+\t\t\tsb->buf[pos] = ' ';\n+\t\t\tfor (cnt = 0; isspace(sb->buf[pos + cnt + 1]); cnt++);\n+\t\t\tstrbuf_remove(sb, pos + 1, cnt);\n \t\t}\n \t}\n }\n \n-static void decode_header(char *it, unsigned itsize);\n+static void decode_header(struct strbuf *line);\n static const char *header[MAX_HDR_PARSED] = {\n \t\"From\",\"Subject\",\"Date\",\n };\n \n-static int check_header(char *line, unsigned linesize, char **hdr_data, int overwrite)\n+static int cmp_header(const struct strbuf *line, const char *hdr)\n {\n-\tint i;\n+\tint len = strlen(hdr);\n+\treturn !strncasecmp(line->buf, hdr, len) && line->len > len &&\n+\t\t\tline->buf[len] == ':' && isspace(line->buf[len + 1]);\n+}\n \n+static int check_header(const struct strbuf *line, struct strbuf *hdr_data[], int overwrite)\n+{\n+\tint i, ret = 0, len;\n+\tstruct strbuf sb = STRBUF_INIT;\n \t/* search for the interesting parts */\n \tfor (i = 0; header[i]; i++) {\n \t\tint len = strlen(header[i]);\n-\t\tif ((!hdr_data[i] || overwrite) &&\n-\t\t    !strncasecmp(line, header[i], len) &&\n-\t\t    line[len] == ':' && isspace(line[len + 1])) {\n+\t\tif ((!hdr_data[i] || overwrite) && cmp_header(line, header[i])) {\n \t\t\t/* Unwrap inline B and Q encoding, and optionally\n \t\t\t * normalize the meta information to utf8.\n \t\t\t */\n-\t\t\tdecode_header(line + len + 2, linesize - len - 2);\n-\t\t\thdr_data[i] = xmalloc(1000 * sizeof(char));\n-\t\t\tif (! handle_header(line, hdr_data[i], len + 2)) {\n-\t\t\t\treturn 1;\n-\t\t\t}\n+\t\t\tstrbuf_add(&sb, line->buf + len + 2, line->len - len -2);\n+\t\t\tdecode_header(&sb);\n+\t\t\thandle_header(&hdr_data[i], &sb);\n+\t\t\tret = 1;\n+\t\t\tgoto check_header_out;\n \t\t}\n \t}\n \n \t/* Content stuff */\n-\tif (!strncasecmp(line, \"Content-Type\", 12) &&\n-\t\tline[12] == ':' && isspace(line[12 + 1])) {\n-\t\tdecode_header(line + 12 + 2, linesize - 12 - 2);\n-\t\tif (! handle_content_type(line)) {\n-\t\t\treturn 1;\n-\t\t}\n-\t}\n-\tif (!strncasecmp(line, \"Content-Transfer-Encoding\", 25) &&\n-\t\tline[25] == ':' && isspace(line[25 + 1])) {\n-\t\tdecode_header(line + 25 + 2, linesize - 25 - 2);\n-\t\tif (! handle_content_transfer_encoding(line)) {\n-\t\t\treturn 1;\n-\t\t}\n+\tif (cmp_header(line, \"Content-Type\")) {\n+\t\tlen = strlen(\"Content-Type: \");\n+\t\tstrbuf_reset(&sb);\n+\t\tstrbuf_add(&sb, line->buf + len, line->len - len);\n+\t\tdecode_header(&sb);\n+\t\tstrbuf_insert(&sb, 0, \"Content-Type: \", len);\n+\t\tif (!handle_content_type(&sb))\n+\t\t\tret = 1;\n+\t\t\tgoto check_header_out;\n+\t}\n+\tif (cmp_header(line, \"Content-Transfer-Encoding\")) {\n+\t\tlen = strlen(\"Content-Transfer-Encoding: \");\n+\t\tstrbuf_reset(&sb);\n+\t\tstrbuf_add(&sb, line->buf + len, line->len - len);\n+\t\tdecode_header(&sb);\n+\t\tif (!handle_content_transfer_encoding(&sb))\n+\t\t\tret = 1;\n+\t\t\tgoto check_header_out;\n \t}\n \n \t/* for inbody stuff */\n-\tif (!memcmp(\">From\", line, 5) && isspace(line[5]))\n-\t\treturn 1;\n-\tif (!memcmp(\"[PATCH]\", line, 7) && isspace(line[7])) {\n+\tif (!prefixcmp(line->buf, \">From\") && isspace(line->buf[5]))\n+\t\tret = 1;\n+\t\tgoto check_header_out;\n+\tif (!prefixcmp(line->buf, \"[PATCH]\") && isspace(line->buf[7])) {\n \t\tfor (i = 0; header[i]; i++) {\n \t\t\tif (!memcmp(\"Subject\", header[i], 7)) {\n-\t\t\t\tif (! handle_header(line, hdr_data[i], 0)) {\n-\t\t\t\t\treturn 1;\n-\t\t\t\t}\n+\t\t\t\thandle_header(&hdr_data[i], line);\n+\t\t\t\tret = 1;\n+\t\t\t\tgoto check_header_out;\n \t\t\t}\n \t\t}\n \t}\n \n-\t/* no match */\n-\treturn 0;\n+check_header_out:\n+\tstrbuf_release(&sb);\n+\treturn ret;\n }\n \n-static int is_rfc2822_header(char *line)\n+static int is_rfc2822_header(const struct strbuf *line)\n {\n \t/*\n \t * The section that defines the loosest possible\n@@ -357,15 +333,15 @@ static int is_rfc2822_header(char *line)\n \t * ftext = %d33-57 / %59-126\n \t */\n \tint ch;\n-\tchar *cp = line;\n+\tchar *cp = line->buf;\n \n \t/* Count mbox From headers as headers */\n-\tif (!memcmp(line, \"From \", 5) || !memcmp(line, \">From \", 6))\n+\tif (line->len >= 6 && (!memcmp(cp, \"From \", 5) || !memcmp(cp, \">From \", 6)))\n \t\treturn 1;\n \n \twhile ((ch = *cp++)) {\n \t\tif (ch == ':')\n-\t\t\treturn cp != line;\n+\t\t\treturn 1;\n \t\tif ((33 <= ch && ch <= 57) ||\n \t\t    (59 <= ch && ch <= 126))\n \t\t\tcontinue;\n@@ -374,34 +350,20 @@ static int is_rfc2822_header(char *line)\n \treturn 0;\n }\n \n-/*\n- * sz is size of 'line' buffer in bytes.  Must be reasonably\n- * long enough to hold one physical real-world e-mail line.\n- */\n-static int read_one_header_line(char *line, int sz, FILE *in)\n+static int read_one_header_line(struct strbuf *line, FILE *in)\n {\n-\tint len;\n-\n-\t/*\n-\t * We will read at most (sz-1) bytes and then potentially\n-\t * re-add NUL after it.  Accessing line[sz] after this is safe\n-\t * and we can allow len to grow up to and including sz.\n-\t */\n-\tsz--;\n-\n \t/* Get the first part of the line. */\n-\tif (!fgets(line, sz, in))\n+\tif (strbuf_getline(line, in, '\\n'))\n \t\treturn 0;\n \n \t/*\n \t * Is it an empty line or not a valid rfc2822 header?\n \t * If so, stop here, and return false (\"not a header\")\n \t */\n-\tlen = eatspace(line);\n-\tif (!len || !is_rfc2822_header(line)) {\n+\tstrbuf_rtrim(line);\n+\tif (!line->len || !is_rfc2822_header(line)) {\n \t\t/* Re-add the newline */\n-\t\tline[len] = '\\n';\n-\t\tline[len + 1] = '\\0';\n+\t\tstrbuf_addch(line, '\\n');\n \t\treturn 0;\n \t}\n \n@@ -410,65 +372,53 @@ static int read_one_header_line(char *line, int sz, FILE *in)\n \t * Yuck, 2822 header \"folding\"\n \t */\n \tfor (;;) {\n-\t\tint peek, addlen;\n-\t\tstatic char continuation[1000];\n+\t\tint peek;\n+\t\tstruct strbuf continuation = STRBUF_INIT;\n \n \t\tpeek = fgetc(in); ungetc(peek, in);\n \t\tif (peek != ' ' && peek != '\\t')\n \t\t\tbreak;\n-\t\tif (!fgets(continuation, sizeof(continuation), in))\n+\t\tif (strbuf_getline(&continuation, in, '\\n'))\n \t\t\tbreak;\n-\t\taddlen = eatspace(continuation);\n-\t\tif (len < sz - 1) {\n-\t\t\tif (addlen >= sz - len)\n-\t\t\t\taddlen = sz - len - 1;\n-\t\t\tmemcpy(line + len, continuation, addlen);\n-\t\t\tline[len] = '\\n';\n-\t\t\tlen += addlen;\n-\t\t}\n+\t\tcontinuation.buf[0] = '\\n';\n+\t\tstrbuf_rtrim(&continuation);\n+\t\tstrbuf_addbuf(line, &continuation);\n \t}\n-\tline[len] = 0;\n \n \treturn 1;\n }\n \n-static int decode_q_segment(char *in, char *ot, unsigned otsize, char *ep, int rfc2047)\n+static struct strbuf *decode_q_segment(const struct strbuf *q_seg, int rfc2047)\n {\n-\tchar *otbegin = ot;\n-\tchar *otend = ot + otsize;\n+\tconst char *in = q_seg->buf;\n \tint c;\n-\twhile ((c = *in++) != 0 && (in <= ep)) {\n-\t\tif (ot == otend) {\n-\t\t\t*--ot = '\\0';\n-\t\t\treturn -1;\n-\t\t}\n+\tstruct strbuf *out = xmalloc(sizeof(struct strbuf));\n+\tstrbuf_init(out, q_seg->len);\n+\n+\twhile ((c = *in++) != 0) {\n \t\tif (c == '=') {\n \t\t\tint d = *in++;\n \t\t\tif (d == '\\n' || !d)\n \t\t\t\tbreak; /* drop trailing newline */\n-\t\t\t*ot++ = ((hexval(d) << 4) | hexval(*in++));\n+\t\t\tstrbuf_addch(out, (hexval(d) << 4) | hexval(*in++));\n \t\t\tcontinue;\n \t\t}\n \t\tif (rfc2047 && c == '_') /* rfc2047 4.2 (2) */\n \t\t\tc = 0x20;\n-\t\t*ot++ = c;\n+\t\tstrbuf_addch(out, c);\n \t}\n-\t*ot = 0;\n-\treturn (ot - otbegin);\n+\treturn out;\n }\n \n-static int decode_b_segment(char *in, char *ot, unsigned otsize, char *ep)\n+static struct strbuf *decode_b_segment(const struct strbuf *b_seg)\n {\n \t/* Decode in..ep, possibly in-place to ot */\n \tint c, pos = 0, acc = 0;\n-\tchar *otbegin = ot;\n-\tchar *otend = ot + otsize;\n+\tconst char *in = b_seg->buf;\n+\tstruct strbuf *out = xmalloc(sizeof(struct strbuf));\n+\tstrbuf_init(out, b_seg->len);\n \n-\twhile ((c = *in++) != 0 && (in <= ep)) {\n-\t\tif (ot == otend) {\n-\t\t\t*--ot = '\\0';\n-\t\t\treturn -1;\n-\t\t}\n+\twhile ((c = *in++) != 0) {\n \t\tif (c == '+')\n \t\t\tc = 62;\n \t\telse if (c == '/')\n@@ -493,21 +443,20 @@ static int decode_b_segment(char *in, char *ot, unsigned otsize, char *ep)\n \t\t\tacc = (c << 2);\n \t\t\tbreak;\n \t\tcase 1:\n-\t\t\t*ot++ = (acc | (c >> 4));\n+\t\t\tstrbuf_addch(out, (acc | (c >> 4)));\n \t\t\tacc = (c & 15) << 4;\n \t\t\tbreak;\n \t\tcase 2:\n-\t\t\t*ot++ = (acc | (c >> 2));\n+\t\t\tstrbuf_addch(out, (acc | (c >> 2)));\n \t\t\tacc = (c & 3) << 6;\n \t\t\tbreak;\n \t\tcase 3:\n-\t\t\t*ot++ = (acc | c);\n+\t\t\tstrbuf_addch(out, (acc | c));\n \t\t\tacc = pos = 0;\n \t\t\tbreak;\n \t\t}\n \t}\n-\t*ot = 0;\n-\treturn (ot - otbegin);\n+\treturn out;\n }\n \n /*\n@@ -521,16 +470,16 @@ static int decode_b_segment(char *in, char *ot, unsigned otsize, char *ep)\n  * Otherwise, we default to assuming it is Latin1 for historical\n  * reasons.\n  */\n-static const char *guess_charset(const char *line, const char *target_charset)\n+static const char *guess_charset(const struct strbuf *line, const char *target_charset)\n {\n \tif (is_encoding_utf8(target_charset)) {\n-\t\tif (is_utf8(line))\n+\t\tif (is_utf8(line->buf))\n \t\t\treturn NULL;\n \t}\n \treturn \"latin1\";\n }\n \n-static void convert_to_utf8(char *line, unsigned linesize, const char *charset)\n+static void convert_to_utf8(struct strbuf *line, const char *charset)\n {\n \tchar *out;\n \n@@ -542,112 +491,119 @@ static void convert_to_utf8(char *line, unsigned linesize, const char *charset)\n \n \tif (!strcmp(metainfo_charset, charset))\n \t\treturn;\n-\tout = reencode_string(line, metainfo_charset, charset);\n+\tout = reencode_string(line->buf, metainfo_charset, charset);\n \tif (!out)\n \t\tdie(\"cannot convert from %s to %s\\n\",\n \t\t    charset, metainfo_charset);\n-\tstrlcpy(line, out, linesize);\n-\tfree(out);\n+\tstrbuf_attach(line, out, strlen(out), strlen(out));\n }\n \n-static int decode_header_bq(char *it, unsigned itsize)\n+static int decode_header_bq(struct strbuf *it)\n {\n \tchar *in, *out, *ep, *cp, *sp;\n-\tchar outbuf[1000];\n+\tstruct strbuf outbuf = STRBUF_INIT, *dec;\n+\tstruct strbuf charset_q = STRBUF_INIT, piecebuf = STRBUF_INIT;\n \tint rfc2047 = 0;\n \n-\tin = it;\n-\tout = outbuf;\n-\twhile ((ep = strstr(in, \"=?\")) != NULL) {\n-\t\tint sz, encoding;\n-\t\tchar charset_q[256], piecebuf[256];\n+\tin = it->buf;\n+\twhile (in - it->buf <= it->len && (ep = strstr(in, \"=?\")) != NULL) {\n+\t\tint encoding;\n+\t\tstrbuf_reset(&charset_q);\n+\t\tstrbuf_reset(&piecebuf);\n \t\trfc2047 = 1;\n \n \t\tif (in != ep) {\n-\t\t\tsz = ep - in;\n-\t\t\tmemcpy(out, in, sz);\n-\t\t\tout += sz;\n-\t\t\tin += sz;\n+\t\t\tstrbuf_add(&outbuf, in, ep - in);\n+\t\t\tin = ep;\n \t\t}\n \t\t/* E.g.\n \t\t * ep : \"=?iso-2022-jp?B?GyR...?= foo\"\n \t\t * ep : \"=?ISO-8859-1?Q?Foo=FCbar?= baz\"\n \t\t */\n \t\tep += 2;\n-\t\tcp = strchr(ep, '?');\n-\t\tif (!cp)\n-\t\t\treturn rfc2047; /* no munging */\n-\t\tfor (sp = ep; sp < cp; sp++)\n-\t\t\tcharset_q[sp - ep] = tolower(*sp);\n-\t\tcharset_q[cp - ep] = 0;\n+\n+\t\tif (ep - it->buf >= it->len || !(cp = strchr(ep, '?')))\n+\t\t\tgoto decode_header_bq_out;\n+\n+\t\tif (cp + 3 - it->buf > it->len)\n+\t\t\tgoto decode_header_bq_out;\n+\t\tstrbuf_add(&charset_q, ep, cp - ep);\n+\t\tstrbuf_tolower(&charset_q);\n+\n \t\tencoding = cp[1];\n \t\tif (!encoding || cp[2] != '?')\n-\t\t\treturn rfc2047; /* no munging */\n+\t\t\tgoto decode_header_bq_out;\n \t\tep = strstr(cp + 3, \"?=\");\n \t\tif (!ep)\n-\t\t\treturn rfc2047; /* no munging */\n+\t\t\tgoto decode_header_bq_out;\n+\t\tstrbuf_add(&piecebuf, cp + 3, ep - cp - 3);\n \t\tswitch (tolower(encoding)) {\n \t\tdefault:\n-\t\t\treturn rfc2047; /* no munging */\n+\t\t\tgoto decode_header_bq_out;\n \t\tcase 'b':\n-\t\t\tsz = decode_b_segment(cp + 3, piecebuf, sizeof(piecebuf), ep);\n+\t\t\tdec = decode_b_segment(&piecebuf);\n \t\t\tbreak;\n \t\tcase 'q':\n-\t\t\tsz = decode_q_segment(cp + 3, piecebuf, sizeof(piecebuf), ep, 1);\n+\t\t\tdec = decode_q_segment(&piecebuf, 1);\n \t\t\tbreak;\n \t\t}\n-\t\tif (sz < 0)\n-\t\t\treturn rfc2047;\n \t\tif (metainfo_charset)\n-\t\t\tconvert_to_utf8(piecebuf, sizeof(piecebuf), charset_q);\n+\t\t\tconvert_to_utf8(dec, charset_q.buf);\n \n-\t\tsz = strlen(piecebuf);\n-\t\tif (outbuf + sizeof(outbuf) <= out + sz)\n-\t\t\treturn rfc2047; /* no munging */\n-\t\tstrcpy(out, piecebuf);\n-\t\tout += sz;\n+\t\tstrbuf_addbuf(&outbuf, dec);\n+\t\tstrbuf_release(dec);\n+\t\tfree(dec);\n \t\tin = ep + 2;\n \t}\n-\tstrcpy(out, in);\n-\tstrlcpy(it, outbuf, itsize);\n+\tstrbuf_addstr(&outbuf, in);\n+\tstrbuf_reset(it);\n+\tstrbuf_addbuf(it, &outbuf);\n+decode_header_bq_out:\n+\tstrbuf_release(&outbuf);\n+\tstrbuf_release(&charset_q);\n+\tstrbuf_release(&piecebuf);\n \treturn rfc2047;\n }\n \n-static void decode_header(char *it, unsigned itsize)\n+static void decode_header(struct strbuf *it)\n {\n-\n-\tif (decode_header_bq(it, itsize))\n+\tif (decode_header_bq(it))\n \t\treturn;\n \t/* otherwise \"it\" is a straight copy of the input.\n \t * This can be binary guck but there is no charset specified.\n \t */\n \tif (metainfo_charset)\n-\t\tconvert_to_utf8(it, itsize, \"\");\n+\t\tconvert_to_utf8(it, \"\");\n }\n \n-static int decode_transfer_encoding(char *line, unsigned linesize, int inputlen)\n+static void decode_transfer_encoding(struct strbuf *line)\n {\n-\tchar *ep;\n+\tstruct strbuf *ret;\n+\tint len;\n \n \tswitch (transfer_encoding) {\n \tcase TE_QP:\n-\t\tep = line + inputlen;\n-\t\treturn decode_q_segment(line, line, linesize, ep, 0);\n+\t\tret = decode_q_segment(line, 0);\n+\t\tbreak;\n \tcase TE_BASE64:\n-\t\tep = line + inputlen;\n-\t\treturn decode_b_segment(line, line, linesize, ep);\n+\t\tret = decode_b_segment(line);\n+\t\tbreak;\n \tcase TE_DONTCARE:\n \tdefault:\n-\t\treturn inputlen;\n+\t\treturn;\n \t}\n+\tstrbuf_reset(line);\n+\tstrbuf_addbuf(line, ret);\n+\tstrbuf_release(ret);\n+\tfree(ret);\n }\n \n-static int handle_filter(char *line, unsigned linesize, int linelen);\n+static int handle_filter(struct strbuf *line);\n \n static int find_boundary(void)\n {\n-\twhile(fgets(line, sizeof(line), fin) != NULL) {\n-\t\tif (is_multipart_boundary(line))\n+\twhile(!strbuf_getline(&line, fin, '\\n')) {\n+\t\tif (is_multipart_boundary(&line))\n \t\t\treturn 1;\n \t}\n \treturn 0;\n@@ -655,11 +611,15 @@ static int find_boundary(void)\n \n static int handle_boundary(void)\n {\n-\tchar newline[]=\"\\n\";\n+\tstruct strbuf newline = STRBUF_INIT;\n+\n+\tstrbuf_addch(&newline, '\\n');\n again:\n-\tif (!memcmp(line+content_top->boundary_len, \"--\", 2)) {\n+\tif (line.len >= content_top->boundary->len + 2 &&\n+\t    !memcmp(line.buf + content_top->boundary->len, \"--\", 2)) {\n \t\t/* we hit an end boundary */\n \t\t/* pop the current boundary off the stack */\n+\t\tstrbuf_release(content_top->boundary);\n \t\tfree(content_top->boundary);\n \n \t\t/* technically won't happen as is_multipart_boundary()\n@@ -670,7 +630,8 @@ again:\n \t\t\t\t\t\"can't recover\\n\");\n \t\t\texit(1);\n \t\t}\n-\t\thandle_filter(newline, sizeof(newline), strlen(newline));\n+\t\thandle_filter(&newline);\n+\t\tstrbuf_release(&newline);\n \n \t\t/* skip to the next boundary */\n \t\tif (!find_boundary())\n@@ -680,39 +641,44 @@ again:\n \n \t/* set some defaults */\n \ttransfer_encoding = TE_DONTCARE;\n-\tcharset[0] = 0;\n+\tstrbuf_reset(&charset);\n \tmessage_type = TYPE_TEXT;\n \n \t/* slurp in this section's info */\n-\twhile (read_one_header_line(line, sizeof(line), fin))\n-\t\tcheck_header(line, sizeof(line), p_hdr_data, 0);\n+\twhile (read_one_header_line(&line, fin))\n+\t\tcheck_header(&line, p_hdr_data, 0);\n \n+\tstrbuf_release(&newline);\n \t/* eat the blank line after section info */\n-\treturn (fgets(line, sizeof(line), fin) != NULL);\n+\treturn (strbuf_getline(&line, fin, '\\n') == 0);\n }\n \n-static inline int patchbreak(const char *line)\n+static inline int patchbreak(const struct strbuf *line)\n {\n+\tsize_t i;\n+\n \t/* Beginning of a \"diff -\" header? */\n-\tif (!memcmp(\"diff -\", line, 6))\n+\tif (!prefixcmp(line->buf, \"diff -\"))\n \t\treturn 1;\n \n \t/* CVS \"Index: \" line? */\n-\tif (!memcmp(\"Index: \", line, 7))\n+\tif (!prefixcmp(line->buf, \"Index: \"))\n \t\treturn 1;\n \n \t/*\n \t * \"--- <filename>\" starts patches without headers\n \t * \"---<sp>*\" is a manual separator\n \t */\n-\tif (!memcmp(\"---\", line, 3)) {\n-\t\tline += 3;\n+\tif (line->len < 4)\n+\t\treturn 0;\n+\n+\tif (!prefixcmp(line->buf, \"---\")) {\n \t\t/* space followed by a filename? */\n-\t\tif (line[0] == ' ' && !isspace(line[1]))\n+\t\tif (line->buf[3] == ' ' && !isspace(line->buf[4]))\n \t\t\treturn 1;\n \t\t/* Just whitespace? */\n-\t\tfor (;;) {\n-\t\t\tunsigned char c = *line++;\n+\t\tfor (i = 3; i < line->len; i++) {\n+\t\t\tunsigned char c = line->buf[i];\n \t\t\tif (c == '\\n')\n \t\t\t\treturn 1;\n \t\t\tif (!isspace(c))\n@@ -723,32 +689,25 @@ static inline int patchbreak(const char *line)\n \treturn 0;\n }\n \n-\n-static int handle_commit_msg(char *line, unsigned linesize)\n+static int handle_commit_msg(struct strbuf *line)\n {\n \tstatic int still_looking = 1;\n-\tchar *endline = line + linesize;\n+\tchar *c;\n \n \tif (!cmitmsg)\n \t\treturn 0;\n \n \tif (still_looking) {\n-\t\tchar *cp = line;\n-\t\tif (isspace(*line)) {\n-\t\t\tfor (cp = line + 1; *cp; cp++) {\n-\t\t\t\tif (!isspace(*cp))\n-\t\t\t\t\tbreak;\n-\t\t\t}\n-\t\t\tif (!*cp)\n-\t\t\t\treturn 0;\n-\t\t}\n-\t\tif ((still_looking = check_header(cp, endline - cp, s_hdr_data, 0)) != 0)\n+\t\tstrbuf_ltrim(line);\n+\t\tif (!line->len)\n+\t\t\treturn 0;\n+\t\tif ((still_looking = check_header(line, s_hdr_data, 0)) != 0)\n \t\t\treturn 0;\n \t}\n \n \t/* normalize the log message to UTF-8. */\n \tif (metainfo_charset)\n-\t\tconvert_to_utf8(line, endline - line, charset);\n+\t\tconvert_to_utf8(line, charset.buf);\n \n \tif (patchbreak(line)) {\n \t\tfclose(cmitmsg);\n@@ -756,18 +715,18 @@ static int handle_commit_msg(char *line, unsigned linesize)\n \t\treturn 1;\n \t}\n \n-\tfputs(line, cmitmsg);\n+\tfputs(line->buf, cmitmsg);\n \treturn 0;\n }\n \n-static int handle_patch(char *line, int len)\n+static int handle_patch(const struct strbuf *line)\n {\n-\tfwrite(line, 1, len, patchfile);\n+\tfwrite(line->buf, 1, line->len, patchfile);\n \tpatch_lines++;\n \treturn 0;\n }\n \n-static int handle_filter(char *line, unsigned linesize, int linelen)\n+static int handle_filter(struct strbuf *line)\n {\n \tstatic int filter = 0;\n \n@@ -776,11 +735,11 @@ static int handle_filter(char *line, unsigned linesize, int linelen)\n \t */\n \tswitch (filter) {\n \tcase 0:\n-\t\tif (!handle_commit_msg(line, linesize))\n+\t\tif (!handle_commit_msg(line))\n \t\t\tbreak;\n \t\tfilter++;\n \tcase 1:\n-\t\tif (!handle_patch(line, linelen))\n+\t\tif (!handle_patch(line))\n \t\t\tbreak;\n \t\tfilter++;\n \tdefault:\n@@ -793,101 +752,105 @@ static int handle_filter(char *line, unsigned linesize, int linelen)\n static void handle_body(void)\n {\n \tint rc = 0;\n-\tstatic char newline[2000];\n-\tstatic char *np = newline;\n-\tint len = strlen(line);\n+\tint len = 0;\n+\tstruct strbuf prev = STRBUF_INIT;\n \n \t/* Skip up to the first boundary */\n \tif (content_top->boundary) {\n \t\tif (!find_boundary())\n-\t\t\treturn;\n+\t\t\tgoto handle_body_out;\n \t}\n \n \tdo {\n+\t\tstrbuf_setlen(&line, line.len + len);\n+\n \t\t/* process any boundary lines */\n-\t\tif (content_top->boundary && is_multipart_boundary(line)) {\n+\t\tif (content_top->boundary && is_multipart_boundary(&line)) {\n \t\t\t/* flush any leftover */\n-\t\t\tif (np != newline)\n-\t\t\t\thandle_filter(newline, sizeof(newline),\n-\t\t\t\t\t      np - newline);\n+\t\t\tif (line.len)\n+\t\t\t\thandle_filter(&line);\n+\n \t\t\tif (!handle_boundary())\n-\t\t\t\treturn;\n-\t\t\tlen = strlen(line);\n+\t\t\t\tgoto handle_body_out;\n \t\t}\n \n \t\t/* Unwrap transfer encoding */\n-\t\tlen = decode_transfer_encoding(line, sizeof(line), len);\n-\t\tif (len < 0) {\n-\t\t\terror(\"Malformed input line\");\n-\t\t\treturn;\n-\t\t}\n+\t\tdecode_transfer_encoding(&line);\n \n \t\tswitch (transfer_encoding) {\n \t\tcase TE_BASE64:\n \t\tcase TE_QP:\n \t\t{\n-\t\t\tchar *op = line;\n+\t\t\tstruct strbuf **lines, **it, *sb;\n+\n+\t\t\t/* Prepend any previous partial lines */\n+\t\t\tstrbuf_insert(&line, 0, prev.buf, prev.len);\n+\t\t\tstrbuf_reset(&prev);\n \n \t\t\t/* binary data most likely doesn't have newlines */\n \t\t\tif (message_type != TYPE_TEXT) {\n-\t\t\t\trc = handle_filter(line, sizeof(line), len);\n+\t\t\t\trc = handle_filter(&line);\n \t\t\t\tbreak;\n \t\t\t}\n-\n \t\t\t/*\n \t\t\t * This is a decoded line that may contain\n \t\t\t * multiple new lines.  Pass only one chunk\n \t\t\t * at a time to handle_filter()\n \t\t\t */\n-\t\t\tdo {\n-\t\t\t\twhile (op < line + len && *op != '\\n')\n-\t\t\t\t\t*np++ = *op++;\n-\t\t\t\t*np = *op;\n-\t\t\t\tif (*np != 0) {\n-\t\t\t\t\t/* should be sitting on a new line */\n-\t\t\t\t\t*(++np) = 0;\n-\t\t\t\t\top++;\n-\t\t\t\t\trc = handle_filter(newline, sizeof(newline), np - newline);\n-\t\t\t\t\tnp = newline;\n-\t\t\t\t}\n-\t\t\t} while (op < line + len);\n+\t\t\tlines = strbuf_split(&line, '\\n');\n+\t\t\tstrbuf_reset(&line);\n+\t\t\tfor (it = lines; (sb = *it); it++) {\n+\t\t\t\tif (*(it + 1) == NULL) /* The last token */\n+\t\t\t\t\tif (sb->buf[sb->len - 1] != '\\n') {\n+\t\t\t\t\t\t/* Partial line, save it for later. */\n+\t\t\t\t\t\tstrbuf_addbuf(&prev, sb);\n+\t\t\t\t\t\tbreak;\n+\t\t\t\t\t}\n+\t\t\t\trc = handle_filter(sb);\n+\t\t\t}\n \t\t\t/*\n-\t\t\t * The partial chunk is saved in newline and will be\n+\t\t\t * The partial chunk is saved in \"prev\" and will be\n \t\t\t * appended by the next iteration of read_line_with_nul().\n \t\t\t */\n+\t\t\tstrbuf_list_free(lines);\n \t\t\tbreak;\n \t\t}\n \t\tdefault:\n-\t\t\trc = handle_filter(line, sizeof(line), len);\n+\t\t\trc = handle_filter(&line);\n+\t\t\tstrbuf_reset(&line);\n \t\t}\n \t\tif (rc)\n \t\t\t/* nothing left to filter */\n \t\t\tbreak;\n-\t} while ((len = read_line_with_nul(line, sizeof(line), fin)));\n+\t\tif (strbuf_avail(&line) < 100)\n+\t\t\tstrbuf_grow(&line, 100);\n+\t} while ((len = read_line_with_nul(line.buf, strbuf_avail(&line), fin)));\n \n+handle_body_out:\n+\tstrbuf_release(&prev);\n \treturn;\n }\n \n-static void output_header_lines(FILE *fout, const char *hdr, char *data)\n+static void output_header_lines(FILE *fout, const char *hdr, const struct strbuf *data)\n {\n+\tchar *sp = data->buf;\n \twhile (1) {\n-\t\tchar *ep = strchr(data, '\\n');\n+\t\tchar *ep = strchr(sp, '\\n');\n \t\tint len;\n \t\tif (!ep)\n-\t\t\tlen = strlen(data);\n+\t\t\tlen = strlen(sp);\n \t\telse\n-\t\t\tlen = ep - data;\n-\t\tfprintf(fout, \"%s: %.*s\\n\", hdr, len, data);\n+\t\t\tlen = ep - sp;\n+\t\tfprintf(fout, \"%s: %.*s\\n\", hdr, len, sp);\n \t\tif (!ep)\n \t\t\tbreak;\n-\t\tdata = ep + 1;\n+\t\tsp = ep + 1;\n \t}\n }\n \n static void handle_info(void)\n {\n-\tchar *sub;\n-\tchar *hdr;\n+\tstruct strbuf *hdr;\n \tint i;\n \n \tfor (i = 0; header[i]; i++) {\n@@ -901,20 +864,18 @@ static void handle_info(void)\n \t\t\tcontinue;\n \n \t\tif (!memcmp(header[i], \"Subject\", 7)) {\n-\t\t\tif (keep_subject)\n-\t\t\t\tsub = hdr;\n-\t\t\telse {\n-\t\t\t\tsub = cleanup_subject(hdr);\n-\t\t\t\tcleanup_space(sub);\n+\t\t\tif (!keep_subject) {\n+\t\t\t\tcleanup_subject(hdr);\n+\t\t\t\tcleanup_space(hdr);\n \t\t\t}\n-\t\t\toutput_header_lines(fout, \"Subject\", sub);\n+\t\t\toutput_header_lines(fout, \"Subject\", hdr);\n \t\t} else if (!memcmp(header[i], \"From\", 4)) {\n \t\t\thandle_from(hdr);\n-\t\t\tfprintf(fout, \"Author: %s\\n\", name);\n-\t\t\tfprintf(fout, \"Email: %s\\n\", email);\n+\t\t\tfprintf(fout, \"Author: %s\\n\", name.buf);\n+\t\t\tfprintf(fout, \"Email: %s\\n\", email.buf);\n \t\t} else {\n \t\t\tcleanup_space(hdr);\n-\t\t\tfprintf(fout, \"%s: %s\\n\", header[i], hdr);\n+\t\t\tfprintf(fout, \"%s: %s\\n\", header[i], hdr->buf);\n \t\t}\n \t}\n \tfprintf(fout, \"\\n\");\n@@ -941,8 +902,8 @@ static int mailinfo(FILE *in, FILE *out, int ks, const char *encoding,\n \t\treturn -1;\n \t}\n \n-\tp_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(char *));\n-\ts_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(char *));\n+\tp_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(*p_hdr_data));\n+\ts_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(*s_hdr_data));\n \n \tdo {\n \t\tpeek = fgetc(in);\n@@ -950,8 +911,8 @@ static int mailinfo(FILE *in, FILE *out, int ks, const char *encoding,\n \tungetc(peek, in);\n \n \t/* process the email header */\n-\twhile (read_one_header_line(line, sizeof(line), fin))\n-\t\tcheck_header(line, sizeof(line), p_hdr_data, 1);\n+\twhile (read_one_header_line(&line, fin))\n+\t\tcheck_header(&line, p_hdr_data, 1);\n \n \thandle_body();\n \thandle_info();\n-- \n1.5.4.5\n"},{"id":"83059","messageId":"7vy747fx9x.fsf_-_@gitster.siamese.dyndns.org","threadId":"14398","inReplyTo":"48769E91.60205@etek.chalmers.se","subject":"Re:! [PATCH/RFC] git-mailinfo: use strbuf's instead of fixed buffers","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-07-12T06:10:18Z","receivedAt":"2008-07-12T06:10:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lukas Sandström <lukass@etek.chalmers.se> writes:\n\n> -static char *sanity_check(char *name, char *email)\n> +static void sanity_check(struct strbuf *out, struct strbuf *name, struct strbuf *email)\n>  {\n> -\tint len = strlen(name);\n> -\tif (len < 3 || len > 60)\n> -\t\treturn email;\n> -\tif (strchr(name, '@') || strchr(name, '<') || strchr(name, '>'))\n> -\t\treturn email;\n> -\treturn name;\n> +\tstruct strbuf o = STRBUF_INIT;\n> +\tif (name->len < 3 || name->len > 60)\n> +\t\tstrbuf_addbuf(&o, email);\n> +\tif (strchr(name->buf, '@') || strchr(name->buf, '<') ||\n> +\t\tstrchr(name->buf, '>'))\n> +\t\tstrbuf_addbuf(&o, email);\n> +\tstrbuf_addbuf(&o, name);\n> +\tstrbuf_reset(out);\n> +\tstrbuf_addbuf(out, &o);\n> +\tstrbuf_release(&o);\n\nThis does not look like a correct conversion.  When name is too short or\ntoo long, we do not even look at name and return email straight.  Perhaps\nthis would be more faithful conversion:\n\n\tstruct strbuf *src = name;\n\tif (name->len < 3 ||\n            60 < name->len ||\n\t    strchr(name->buf, '@') ||\n\t    strchr(name->buf, '<') ||\n\t    strchr(name->buf, '>'))\n\t\tsrc = email;\n\telse if (name == out)\n        \treturn;\n\tstrbuf_reset(out);\n\tstrbuf_addbuf(out, src);\n\nIt is not your fault, but sanity_check() is a grave misnomer for this\nfunction.  This does \"get_sane_name\" (i.e. we have name and email but if\nname does not look right, use email instead).\n\n> -static int bogus_from(char *line)\n> +static int bogus_from(const struct strbuf *line)\n>  {\n>  \t/* John Doe <johndoe> */\n> -\tchar *bra, *ket, *dst, *cp;\n>  \n> +\tchar *bra, *ket;\n>  \t/* This is fallback, so do not bother if we already have an\n>  \t * e-mail address.\n>  \t */\n> -\tif (*email)\n> +\tif (email.len)\n>  \t\treturn 0;\n>  \n> -\tbra = strchr(line, '<');\n> +\tbra = strchr(line->buf, '<');\n>  \tif (!bra)\n>  \t\treturn 0;\n>  \tket = strchr(bra, '>');\n>  \tif (!ket)\n>  \t\treturn 0;\n>  \n> -\tfor (dst = email, cp = bra+1; cp < ket; )\n> -\t\t*dst++ = *cp++;\n> -\t*dst = 0;\n> -\tfor (cp = line; isspace(*cp); cp++)\n> -\t\t;\n> -\tfor (bra--; isspace(*bra); bra--)\n> -\t\t*bra = 0;\n> -\tcp = sanity_check(cp, email);\n> -\tstrcpy(name, cp);\n> +\tstrbuf_reset(&email);\n> +\tstrbuf_add(&email, bra + 1, ket - bra - 1);\n> +\n> +\tstrbuf_reset(&name);\n> +\tstrbuf_add(&name, line->buf, bra - line->buf);\n> +\tstrbuf_trim(&name);\n> +\tsanity_check(&name, &name, &email);\n>  \treturn 1;\n>  }\n\nConversion looks correct but its return value does not make much sense\n(again, not your fault).  bogus_from() is given a bogus looking from line\n(it is not about checking if it is bogus), and returns 0 if we already\nhave e-mail address, if the from line does not have bra-ket for grabbing\ne-mail address for, but returns 1 if we managed to get name and email\npairs.  The inconsistency does not matter only because its sole caller\nhandle_from() returns its return value, and its caller discards it.  We\nmay be better off declaring this function and handle_from() as void.\n\n> -static int handle_from(char *in_line)\n> +static int handle_from(struct strbuf *from)\n> ...\n> +\tel = strcspn(at, \" \\n\\t\\r\\v\\f>\");\n> +\tstrbuf_reset(&email);\n> +\tstrbuf_add(&email, at, el);\n> +\tstrbuf_remove(from, at - from->buf, el + 1);\n>  \t/* The remainder is name.  It could be \"John Doe <john.doe@xz>\"\n>  \t * or \"john.doe@xz (John Doe)\", but we have whited out the\n>  \t * email part, so trim from both ends, possibly removing\n>  \t * the () pair at the end.\n>  \t */\n\nNow, it should read \"but we have removed the email part\", I think.\n\n> +\tstrbuf_trim(from);\n> +\tif (*from->buf == '(')\n> +\t\tstrbuf_remove(&name, 0, 1);\n> +\tif (*(from->buf + from->len - 1) == ')')\n\nCan from be empty at this point before this check?\n\n> +\t\tstrbuf_setlen(from, from->len - 1);\n> +\n> +\tsanity_check(&name, from, &email);\n>  \treturn 1;\n>  }\n\nWe used to copy the data from the argument (in_line) before munging it in\nthis function, but now we are modifying it in place (from).  Does this\nupset our caller, or the original code was just doing an extra unnecessary\ncopy?\n\n> -static int handle_header(char *line, char *data, int ofs)\n> +static void handle_header(struct strbuf **out, const struct strbuf *line)\n>  {\n> -\tif (!line || !data)\n> -\t\treturn 1;\n> -\n> -\tstrcpy(data, line+ofs);\n> +\tif (!*out) {\n> +\t\t*out = xmalloc(sizeof(struct strbuf));\n> +\t\tstrbuf_init(*out, line->len);\n> +\t} else\n> +\t\tstrbuf_reset(*out);\n>  \n> -\treturn 0;\n> +\tstrbuf_addbuf(*out, (struct strbuf *)line); /* const warning */\n>  }\n\nI think its second parameter can safely become \"const struct strbuf *\";\nperhaps we should fix the definition of strbuf_addbuf() in your first\npatch?\n\n> @@ -173,180 +153,176 @@ static int slurp_attr(const char *line, const char *name, char *attr)\n>  \telse\n>  \t\tends = \"; \\t\";\n>  \tsz = strcspn(ap, ends);\n> -\tmemcpy(attr, ap, sz);\n> -\tattr[sz] = 0;\n> +\tstrbuf_add(attr, ap, sz);\n>  \treturn 1;\n>  }\n>  \n>  struct content_type {\n> -\tchar *boundary;\n> -\tint boundary_len;\n> +\tstruct strbuf *boundary;\n>  };\n>  \n>  static struct content_type content[MAX_BOUNDARIES];\n\nWouldn't it make more sense to get rid of \"struct content_type\" altogether\nand use \"struct strbuf *content[MAX_BOUNDARIES]\" directly?\n\nI'll review from handle_content_type() til the rest of the file\nseparately, as my concentration is wearing out..\n"},{"id":"83066","messageId":"7v3amfxx3a.fsf@gitster.siamese.dyndns.org","threadId":"14398","inReplyTo":"4876820D.4070806@etek.chalmers.se","subject":"Re: [PATCH] git-mailinfo: Fix getting the subject from the body","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-07-12T09:36:57Z","receivedAt":"2008-07-12T09:36:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lukas Sandström <lukass@etek.chalmers.se> writes:\n\n> \"Subject: \" isn't in the static array \"header\", and thus\n> memcmp(\"Subject: \", header[i], 7) will never match.\n>\n> Signed-off-by: Lukas Sandström <lukass@etek.chalmers.se>\n> ---\n>\n> This has been broken since 2007-03-12, with commit\n> 87ab799234639c26ea10de74782fa511cb3ca606\n> so it might not be very important.\n>\n>  builtin-mailinfo.c |    2 +-\n>  1 files changed, 1 insertions(+), 1 deletions(-)\n>\n> diff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\n> index 962aa34..2d1520f 100644\n> --- a/builtin-mailinfo.c\n> +++ b/builtin-mailinfo.c\n> @@ -334,7 +334,7 @@ static int check_header(char *line, unsigned linesize, char **hdr_data, int over\n>  \t\treturn 1;\n>  \tif (!memcmp(\"[PATCH]\", line, 7) && isspace(line[7])) {\n>  \t\tfor (i = 0; header[i]; i++) {\n> -\t\t\tif (!memcmp(\"Subject: \", header[i], 9)) {\n> +\t\t\tif (!memcmp(\"Subject\", header[i], 7)) {\n>  \t\t\t\tif (! handle_header(line, hdr_data[i], 0)) {\n>  \t\t\t\t\treturn 1;\n>  \t\t\t\t}\n\nActually, I do not think your patch alone makes any difference, and the\noriginal code looks somewhat bogus.  If there is no \"Subject: \" in the\nsame section of the message (either in e-mail header in which case\nhdr_data == p_hdr_data[], or in the message body part in which case\nhdr_data == s_hdr_data[]), hdr_data[1] will be NULL, because the only\nplace that allocates the storage for the data is the first loop of this\nfunction that deals with real-RFC2822-header-looking lines.\n\nYou'd probably need something like this on top of your patch to actually\nactivate the code.\n\nAnother thing I noticed and found puzzling is the handling of \">From \"\nline that is shown in the context below.  check_header() is supposed to\nreturn true when it handled header (i.e. not part of the commit message)\nand return false when line is not part of the header.  As \">From \" is part\nof the commit log message, shouldn't it return zero?\n\nDon, this part was what you introduced.  Has this codepath ever been\nexercised in the real life?\n\n builtin-mailinfo.c  |    2 ++\n t/t5100-mailinfo.sh |    2 +-\n t/t5100/sample.mbox |   35 +++++++++++++++++++++++++++++++++++\n 3 files changed, 38 insertions(+), 1 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 2d1520f..13f0502 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -332,12 +332,14 @@ static int check_header(char *line, unsigned linesize, char **hdr_data, int over\n \t/* for inbody stuff */\n \tif (!memcmp(\">From\", line, 5) && isspace(line[5]))\n \t\treturn 1;\n \tif (!memcmp(\"[PATCH]\", line, 7) && isspace(line[7])) {\n \t\tfor (i = 0; header[i]; i++) {\n \t\t\tif (!memcmp(\"Subject\", header[i], 7)) {\n+\t\t\t\tif (!hdr_data[i])\n+\t\t\t\t\thdr_data[i] = xmalloc(linesize + 20);\n \t\t\t\tif (! handle_header(line, hdr_data[i], 0)) {\n \t\t\t\t\treturn 1;\n \t\t\t\t}\n \t\t\t}\n \t\t}\n \t}\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 2d1520f..13f0502 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -335,6 +335,8 @@ static int check_header(char *line, unsigned linesize, char **hdr_data, int over\n \tif (!memcmp(\"[PATCH]\", line, 7) && isspace(line[7])) {\n \t\tfor (i = 0; header[i]; i++) {\n \t\t\tif (!memcmp(\"Subject\", header[i], 7)) {\n+\t\t\t\tif (!hdr_data[i])\n+\t\t\t\t\thdr_data[i] = xmalloc(linesize + 20);\n \t\t\t\tif (! handle_header(line, hdr_data[i], 0)) {\n \t\t\t\t\treturn 1;\n \t\t\t\t}\ndiff --git a/t/t5100-mailinfo.sh b/t/t5100-mailinfo.sh\nindex 577ecc2..e9f3e72 100755\n--- a/t/t5100-mailinfo.sh\n+++ b/t/t5100-mailinfo.sh\n@@ -11,7 +11,7 @@ test_expect_success 'split sample box' \\\n \t'git mailsplit -o. ../t5100/sample.mbox >last &&\n \tlast=`cat last` &&\n \techo total is $last &&\n-\ttest `cat last` = 9'\n+\ttest `cat last` = 10'\n \n for mail in `echo 00*`\n do\ndiff --git a/t/t5100/sample.mbox b/t/t5100/sample.mbox\nindex 0476b96..aba57f9 100644\n--- a/t/t5100/sample.mbox\n+++ b/t/t5100/sample.mbox\n@@ -430,3 +430,38 @@ index b426a14..97756ec 100644\n =20\n =20\n  2. When the environment variable 'GIT_EXTERNAL_DIFF' is set, the\n+From b9704a518e21158433baa2cc2d591fea687967f6 Mon Sep 17 00:00:00 2001\n+From: =?UTF-8?q?Lukas=20Sandstr=C3=B6m?= <lukass@etek.chalmers.se>\n+Date: Thu, 10 Jul 2008 23:41:33 +0200\n+Subject: Re: discussion that lead to this patch\n+MIME-Version: 1.0\n+Content-Type: text/plain; charset=UTF-8\n+Content-Transfer-Encoding: 8bit\n+\n+[PATCH] git-mailinfo: Fix getting the subject from the body\n+\n+\"Subject: \" isn't in the static array \"header\", and thus\n+memcmp(\"Subject: \", header[i], 7) will never match.\n+\n+Signed-off-by: Lukas Sandström <lukass@etek.chalmers.se>\n+Signed-off-by: Junio C Hamano <gitster@pobox.com>\n+---\n+ builtin-mailinfo.c |    2 +-\n+ 1 files changed, 1 insertions(+), 1 deletions(-)\n+\n+diff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\n+index 962aa34..2d1520f 100644\n+--- a/builtin-mailinfo.c\n++++ b/builtin-mailinfo.c\n+@@ -334,7 +334,7 @@ static int check_header(char *line, unsigned linesize, char **hdr_data, int over\n+ \t\treturn 1;\n+ \tif (!memcmp(\"[PATCH]\", line, 7) && isspace(line[7])) {\n+ \t\tfor (i = 0; header[i]; i++) {\n+-\t\t\tif (!memcmp(\"Subject: \", header[i], 9)) {\n++\t\t\tif (!memcmp(\"Subject\", header[i], 7)) {\n+ \t\t\t\tif (! handle_header(line, hdr_data[i], 0)) {\n+ \t\t\t\t\treturn 1;\n+ \t\t\t\t}\n+-- \n+1.5.6.2.455.g1efb2\n+\n"},{"id":"83106","messageId":"487925FA.5020001@etek.chalmers.se","threadId":"14398","inReplyTo":"7v3amfxx3a.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] git-mailinfo: Fix getting the subject from the body","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2008-07-12T21:45:30Z","receivedAt":"2008-07-12T21:45:30Z","isPatch":true,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Junio C Hamano wrote:\n> Lukas Sandström <lukass@etek.chalmers.se> writes:\n> \n>> \"Subject: \" isn't in the static array \"header\", and thus\n>> memcmp(\"Subject: \", header[i], 7) will never match.\n>>\n>> Signed-off-by: Lukas Sandström <lukass@etek.chalmers.se>\n>> ---\n>>\n>> This has been broken since 2007-03-12, with commit\n>> 87ab799234639c26ea10de74782fa511cb3ca606\n>> so it might not be very important.\n>>\n>>  builtin-mailinfo.c |    2 +-\n>>  1 files changed, 1 insertions(+), 1 deletions(-)\n>>\n>> diff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\n>> index 962aa34..2d1520f 100644\n>> --- a/builtin-mailinfo.c\n>> +++ b/builtin-mailinfo.c\n>> @@ -334,7 +334,7 @@ static int check_header(char *line, unsigned linesize, char **hdr_data, int over\n>>  \t\treturn 1;\n>>  \tif (!memcmp(\"[PATCH]\", line, 7) && isspace(line[7])) {\n>>  \t\tfor (i = 0; header[i]; i++) {\n>> -\t\t\tif (!memcmp(\"Subject: \", header[i], 9)) {\n>> +\t\t\tif (!memcmp(\"Subject\", header[i], 7)) {\n>>  \t\t\t\tif (! handle_header(line, hdr_data[i], 0)) {\n>>  \t\t\t\t\treturn 1;\n>>  \t\t\t\t}\n> \n> Actually, I do not think your patch alone makes any difference, and the\n> original code looks somewhat bogus.  If there is no \"Subject: \" in the\n> same section of the message (either in e-mail header in which case\n> hdr_data == p_hdr_data[], or in the message body part in which case\n> hdr_data == s_hdr_data[]), hdr_data[1] will be NULL, because the only\n> place that allocates the storage for the data is the first loop of this\n> function that deals with real-RFC2822-header-looking lines.\n> \n> You'd probably need something like this on top of your patch to actually\n> activate the code.\n\nRight, I noticed that too. It's fixed in the strbuf conversion, I think.\n\nLukas Sandström <lukass@etek.chalmers.se> wrote:\n > After looking at this part some more, I see that there is no guarantee\n > that hdr_data[i] != NULL in this codepath, and then we won't use the\n > subject anyway.\n\nI'll be hiking the next week, in case you wonder why I'm not responding.\n\nI'll try to get another version of the patches out before I leave.\n\n/Lukas\n"},{"id":"83155","messageId":"487A46C5.6000503@etek.chalmers.se","threadId":"14398","inReplyTo":"7vy747fx9x.fsf_-_@gitster.siamese.dyndns.org","subject":"Re: ! [PATCH/RFC] git-mailinfo: use strbuf's instead of fixed buffers","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2008-07-13T18:17:41Z","receivedAt":"2008-07-13T18:17:41Z","isPatch":true,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Junio C Hamano wrote:\n> Lukas Sandström <lukass@etek.chalmers.se> writes:\n> \n>> -static char *sanity_check(char *name, char *email)\n>> +static void sanity_check(struct strbuf *out, struct strbuf *name, struct strbuf *email)\n>>  {\n>> -\tint len = strlen(name);\n>> -\tif (len < 3 || len > 60)\n>> -\t\treturn email;\n>> -\tif (strchr(name, '@') || strchr(name, '<') || strchr(name, '>'))\n>> -\t\treturn email;\n>> -\treturn name;\n>> +\tstruct strbuf o = STRBUF_INIT;\n>> +\tif (name->len < 3 || name->len > 60)\n>> +\t\tstrbuf_addbuf(&o, email);\n>> +\tif (strchr(name->buf, '@') || strchr(name->buf, '<') ||\n>> +\t\tstrchr(name->buf, '>'))\n>> +\t\tstrbuf_addbuf(&o, email);\n>> +\tstrbuf_addbuf(&o, name);\n>> +\tstrbuf_reset(out);\n>> +\tstrbuf_addbuf(out, &o);\n>> +\tstrbuf_release(&o);\n> \n> This does not look like a correct conversion.  When name is too short or\n> too long, we do not even look at name and return email straight.  Perhaps\n> this would be more faithful conversion:\n> \n> \tstruct strbuf *src = name;\n> \tif (name->len < 3 ||\n>             60 < name->len ||\n> \t    strchr(name->buf, '@') ||\n> \t    strchr(name->buf, '<') ||\n> \t    strchr(name->buf, '>'))\n> \t\tsrc = email;\n> \telse if (name == out)\n>         \treturn;\n> \tstrbuf_reset(out);\n> \tstrbuf_addbuf(out, src);\n> \n> It is not your fault, but sanity_check() is a grave misnomer for this\n> function.  This does \"get_sane_name\" (i.e. we have name and email but if\n> name does not look right, use email instead).\n\nRight. I changed the name and used your implementation.\n\n>> -static int bogus_from(char *line)\n>> +static int bogus_from(const struct strbuf *line)\n>>  {\n\n[ ... ]\n\n>>  \treturn 1;\n>>  }\n> \n> Conversion looks correct but its return value does not make much sense\n> (again, not your fault).  bogus_from() is given a bogus looking from line\n> (it is not about checking if it is bogus), and returns 0 if we already\n> have e-mail address, if the from line does not have bra-ket for grabbing\n> e-mail address for, but returns 1 if we managed to get name and email\n> pairs.  The inconsistency does not matter only because its sole caller\n> handle_from() returns its return value, and its caller discards it.  We\n> may be better off declaring this function and handle_from() as void.\n> \n\nDone.\n\n>>  \t/* The remainder is name.  It could be \"John Doe <john.doe@xz>\"\n>>  \t * or \"john.doe@xz (John Doe)\", but we have whited out the\n>>  \t * email part, so trim from both ends, possibly removing\n>>  \t * the () pair at the end.\n>>  \t */\n> \n> Now, it should read \"but we have removed the email part\", I think.\n\nDone.\n>> +\tstrbuf_trim(from);\n>> +\tif (*from->buf == '(')\n>> +\t\tstrbuf_remove(&name, 0, 1);\n>> +\tif (*(from->buf + from->len - 1) == ')')\n> \n> Can from be empty at this point before this check?\n> \n\nYes. Added a test for that.\n\n>> +\t\tstrbuf_setlen(from, from->len - 1);\n>> +\n>> +\tsanity_check(&name, from, &email);\n>>  \treturn 1;\n>>  }\n> \n> We used to copy the data from the argument (in_line) before munging it in\n> this function, but now we are modifying it in place (from).  Does this\n> upset our caller, or the original code was just doing an extra unnecessary\n> copy?\n\nNo, it's ok with the current code, since when handle_from() is called, it's\nthe last time we look at the header[i] field. I changed it to copy its argument\nanyway, for future-proofing.\n\n> \n>> -static int handle_header(char *line, char *data, int ofs)\n>> +static void handle_header(struct strbuf **out, const struct strbuf *line)\n>>  {\n>> -\tif (!line || !data)\n>> -\t\treturn 1;\n>> -\n>> -\tstrcpy(data, line+ofs);\n>> +\tif (!*out) {\n>> +\t\t*out = xmalloc(sizeof(struct strbuf));\n>> +\t\tstrbuf_init(*out, line->len);\n>> +\t} else\n>> +\t\tstrbuf_reset(*out);\n>>  \n>> -\treturn 0;\n>> +\tstrbuf_addbuf(*out, (struct strbuf *)line); /* const warning */\n>>  }\n> \n> I think its second parameter can safely become \"const struct strbuf *\";\n> perhaps we should fix the definition of strbuf_addbuf() in your first\n> patch?\n> \nDone.\n\n>> @@ -173,180 +153,176 @@ static int slurp_attr(const char *line, const char *name, char *attr)\n>>  \telse\n>>  \t\tends = \"; \\t\";\n>>  \tsz = strcspn(ap, ends);\n>> -\tmemcpy(attr, ap, sz);\n>> -\tattr[sz] = 0;\n>> +\tstrbuf_add(attr, ap, sz);\n>>  \treturn 1;\n>>  }\n>>  \n>>  struct content_type {\n>> -\tchar *boundary;\n>> -\tint boundary_len;\n>> +\tstruct strbuf *boundary;\n>>  };\n>>  \n>>  static struct content_type content[MAX_BOUNDARIES];\n> \n> Wouldn't it make more sense to get rid of \"struct content_type\" altogether\n> and use \"struct strbuf *content[MAX_BOUNDARIES]\" directly?\n\nSure.\n\nI'll send a patch updated with your comments shortly.\n\n/Lukas\n"},{"id":"83156","messageId":"487A4948.8080003@etek.chalmers.se","threadId":"14398","inReplyTo":"487A46C5.6000503@etek.chalmers.se","subject":"[PATCH] Make some strbuf_*() struct strbuf arguments const.","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2008-07-13T18:28:24Z","receivedAt":"2008-07-13T18:28:24Z","isPatch":true,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Signed-off-by: Lukas Sandström <lukass@etek.chalmers.se>\n---\n\n strbuf.c |    2 +-\n strbuf.h |    6 +++---\n 2 files changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git a/strbuf.c b/strbuf.c\nindex 4aed752..7767170 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -67,7 +67,7 @@ void strbuf_rtrim(struct strbuf *sb)\n \tsb->buf[sb->len] = '\\0';\n }\n \n-int strbuf_cmp(struct strbuf *a, struct strbuf *b)\n+int strbuf_cmp(const struct strbuf *a, const struct strbuf *b)\n {\n \tint cmp;\n \tif (a->len < b->len) {\ndiff --git a/strbuf.h b/strbuf.h\nindex faec229..a1b0143 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -61,7 +61,7 @@ static inline void strbuf_swap(struct strbuf *a, struct strbuf *b) {\n }\n \n /*----- strbuf size related -----*/\n-static inline size_t strbuf_avail(struct strbuf *sb) {\n+static inline size_t strbuf_avail(const struct strbuf *sb) {\n \treturn sb->alloc ? sb->alloc - sb->len - 1 : 0;\n }\n \n@@ -78,7 +78,7 @@ static inline void strbuf_setlen(struct strbuf *sb, size_t len) {\n \n /*----- content related -----*/\n extern void strbuf_rtrim(struct strbuf *);\n-extern int strbuf_cmp(struct strbuf *, struct strbuf *);\n+extern int strbuf_cmp(const struct strbuf *, const struct strbuf *);\n \n /*----- add data in your buffer -----*/\n static inline void strbuf_addch(struct strbuf *sb, int c) {\n@@ -98,7 +98,7 @@ extern void strbuf_add(struct strbuf *, const void *, size_t);\n static inline void strbuf_addstr(struct strbuf *sb, const char *s) {\n \tstrbuf_add(sb, s, strlen(s));\n }\n-static inline void strbuf_addbuf(struct strbuf *sb, struct strbuf *sb2) {\n+static inline void strbuf_addbuf(struct strbuf *sb, const struct strbuf *sb2) {\n \tstrbuf_add(sb, sb2->buf, sb2->len);\n }\n extern void strbuf_adddup(struct strbuf *sb, size_t pos, size_t len);\n-- \n1.5.4.5\n"},{"id":"83157","messageId":"487A497E.6030308@etek.chalmers.se","threadId":"14398","inReplyTo":"487A4948.8080003@etek.chalmers.se","subject":"[PATCH] Add some useful functions for strbuf manipulation.","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2008-07-13T18:29:18Z","receivedAt":"2008-07-13T18:29:18Z","isPatch":true,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Signed-off-by: Lukas Sandström <lukass@etek.chalmers.se>\n---\n strbuf.c |   70 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n strbuf.h |    6 +++++\n 2 files changed, 76 insertions(+), 0 deletions(-)\n\ndiff --git a/strbuf.c b/strbuf.c\nindex 7767170..6294940 100644\n--- a/strbuf.c\n+++ b/strbuf.c\n@@ -60,6 +60,18 @@ void strbuf_grow(struct strbuf *sb, size_t extra)\n \tALLOC_GROW(sb->buf, sb->len + extra + 1, sb->alloc);\n }\n \n+void strbuf_trim(struct strbuf *sb)\n+{\n+\tchar *b = sb->buf;\n+\twhile (sb->len > 0 && isspace((unsigned char)sb->buf[sb->len - 1]))\n+\t\tsb->len--;\n+\twhile(sb->len > 0 && isspace(*b)) {\n+\t\tb++;\n+\t\tsb->len--;\n+\t}\n+\tmemmove(sb->buf, b, sb->len);\n+\tsb->buf[sb->len] = '\\0';\n+}\n void strbuf_rtrim(struct strbuf *sb)\n {\n \twhile (sb->len > 0 && isspace((unsigned char)sb->buf[sb->len - 1]))\n@@ -67,6 +79,64 @@ void strbuf_rtrim(struct strbuf *sb)\n \tsb->buf[sb->len] = '\\0';\n }\n \n+void strbuf_ltrim(struct strbuf *sb)\n+{\n+\tchar *b = sb->buf;\n+\twhile(sb->len > 0 && isspace(*b)) {\n+\t\tb++;\n+\t\tsb->len--;\n+\t}\n+\tmemmove(sb->buf, b, sb->len);\n+\tsb->buf[sb->len] = '\\0';\n+}\n+\n+void strbuf_tolower(struct strbuf *sb)\n+{\n+\tint i;\n+\tfor (i = 0; i < sb->len; i++)\n+\t\tsb->buf[i] = tolower(sb->buf[i]);\n+}\n+\n+struct strbuf ** strbuf_split(const struct strbuf *sb, int delim)\n+{\n+\tint alloc = 2, pos = 0;\n+\tchar *n, *p;\n+\tstruct strbuf **ret;\n+\tstruct strbuf *t;\n+\n+\tret = xcalloc(alloc, sizeof(struct strbuf *));\n+\tp = n = sb->buf;\n+\twhile (n < sb->buf + sb->len) {\n+\t\tint len;\n+\t\tn = memchr(n, delim, sb->len - (n - sb->buf));\n+\t\tif (pos + 1 >= alloc) {\n+\t\t\talloc = alloc * 2;\n+\t\t\tret = xrealloc(ret, sizeof(struct strbuf *) * alloc);\n+\t\t}\n+\t\tif (!n)\n+\t\t\tn = sb->buf + sb->len - 1;\n+\t\tlen = n - p + 1;\n+\t\tt = xmalloc(sizeof(struct strbuf));\n+\t\tstrbuf_init(t, len);\n+\t\tstrbuf_add(t, p, len);\n+\t\tret[pos] = t;\n+\t\tret[++pos] = NULL;\n+\t\tp = ++n;\n+\t}\n+\treturn ret;\n+}\n+\n+void strbuf_list_free(struct strbuf ** sbs)\n+{\n+\tstruct strbuf **s = sbs;\n+\n+\twhile(*s) {\n+\t\tstrbuf_release(*s);\n+\t\tfree(*s++);\n+\t}\n+\tfree(sbs);\n+}\n+\n int strbuf_cmp(const struct strbuf *a, const struct strbuf *b)\n {\n \tint cmp;\ndiff --git a/strbuf.h b/strbuf.h\nindex a1b0143..5e65a3f 100644\n--- a/strbuf.h\n+++ b/strbuf.h\n@@ -77,8 +77,14 @@ static inline void strbuf_setlen(struct strbuf *sb, size_t len) {\n #define strbuf_reset(sb)  strbuf_setlen(sb, 0)\n \n /*----- content related -----*/\n+extern void strbuf_trim(struct strbuf *);\n extern void strbuf_rtrim(struct strbuf *);\n+extern void strbuf_ltrim(struct strbuf *);\n extern int strbuf_cmp(const struct strbuf *, const struct strbuf *);\n+extern void strbuf_tolower(struct strbuf *);\n+\n+extern struct strbuf ** strbuf_split(const struct strbuf*, int delim);\n+extern void strbuf_list_free(struct strbuf **);\n \n /*----- add data in your buffer -----*/\n static inline void strbuf_addch(struct strbuf *sb, int c) {\n-- \n1.5.4.5\n"},{"id":"83158","messageId":"487A49B4.6000407@etek.chalmers.se","threadId":"14398","inReplyTo":"487A497E.6030308@etek.chalmers.se","subject":"[PATCH] git-mailinfo: use strbuf's instead of fixed buffers","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2008-07-13T18:30:12Z","receivedAt":"2008-07-13T18:30:12Z","isPatch":true,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Signed-off-by: Lukas Sandström <lukass@etek.chalmers.se>\n---\n builtin-mailinfo.c |  750 ++++++++++++++++++++++++----------------------------\n 1 files changed, 348 insertions(+), 402 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 2d1520f..329ac43 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -5,14 +5,15 @@\n #include \"cache.h\"\n #include \"builtin.h\"\n #include \"utf8.h\"\n+#include \"strbuf.h\"\n \n static FILE *cmitmsg, *patchfile, *fin, *fout;\n \n static int keep_subject;\n static const char *metainfo_charset;\n-static char line[1000];\n-static char name[1000];\n-static char email[1000];\n+static struct strbuf line = STRBUF_INIT;\n+static struct strbuf name = STRBUF_INIT;\n+static struct strbuf email = STRBUF_INIT;\n \n static enum  {\n \tTE_DONTCARE, TE_QP, TE_BASE64,\n@@ -21,74 +22,77 @@ static enum  {\n \tTYPE_TEXT, TYPE_OTHER,\n } message_type;\n \n-static char charset[256];\n+static struct strbuf charset = STRBUF_INIT;\n static int patch_lines;\n-static char **p_hdr_data, **s_hdr_data;\n+static struct strbuf **p_hdr_data, **s_hdr_data;\n \n #define MAX_HDR_PARSED 10\n #define MAX_BOUNDARIES 5\n \n-static char *sanity_check(char *name, char *email)\n+static void get_sane_name(struct strbuf *out, struct strbuf *name, struct strbuf *email)\n {\n-\tint len = strlen(name);\n-\tif (len < 3 || len > 60)\n-\t\treturn email;\n-\tif (strchr(name, '@') || strchr(name, '<') || strchr(name, '>'))\n-\t\treturn email;\n-\treturn name;\n+\tstruct strbuf *src = name;\n+\tif (name->len < 3 || 60 < name->len || strchr(name->buf, '@') ||\n+\t\tstrchr(name->buf, '<') || strchr(name->buf, '>'))\n+\t\tsrc = email;\n+\telse if (name == out)\n+\t\treturn;\n+\tstrbuf_reset(out);\n+\tstrbuf_addbuf(out, src);\n }\n \n-static int bogus_from(char *line)\n+static void parse_bogus_from(const struct strbuf *line)\n {\n \t/* John Doe <johndoe> */\n-\tchar *bra, *ket, *dst, *cp;\n \n+\tchar *bra, *ket;\n \t/* This is fallback, so do not bother if we already have an\n \t * e-mail address.\n \t */\n-\tif (*email)\n-\t\treturn 0;\n+\tif (email.len)\n+\t\treturn;\n \n-\tbra = strchr(line, '<');\n+\tbra = strchr(line->buf, '<');\n \tif (!bra)\n-\t\treturn 0;\n+\t\treturn;\n \tket = strchr(bra, '>');\n \tif (!ket)\n-\t\treturn 0;\n+\t\treturn;\n \n-\tfor (dst = email, cp = bra+1; cp < ket; )\n-\t\t*dst++ = *cp++;\n-\t*dst = 0;\n-\tfor (cp = line; isspace(*cp); cp++)\n-\t\t;\n-\tfor (bra--; isspace(*bra); bra--)\n-\t\t*bra = 0;\n-\tcp = sanity_check(cp, email);\n-\tstrcpy(name, cp);\n-\treturn 1;\n+\tstrbuf_reset(&email);\n+\tstrbuf_add(&email, bra + 1, ket - bra - 1);\n+\n+\tstrbuf_reset(&name);\n+\tstrbuf_add(&name, line->buf, bra - line->buf);\n+\tstrbuf_trim(&name);\n+\tget_sane_name(&name, &name, &email);\n }\n \n-static int handle_from(char *in_line)\n+static void handle_from(const struct strbuf *from)\n {\n-\tchar line[1000];\n \tchar *at;\n-\tchar *dst;\n+\tsize_t el;\n+\tstruct strbuf f;\n \n-\tstrcpy(line, in_line);\n-\tat = strchr(line, '@');\n+\tstrbuf_init(&f, from->len);\n+\tstrbuf_addbuf(&f, from);\n+\n+\tat = strchr(f.buf, '@');\n \tif (!at)\n-\t\treturn bogus_from(line);\n+\t\treturn parse_bogus_from(from);\n \n \t/*\n \t * If we already have one email, don't take any confusing lines\n \t */\n-\tif (*email && strchr(at+1, '@'))\n-\t\treturn 0;\n+\tif (email.len && strchr(at + 1, '@')) {\n+\t\tstrbuf_release(&f);\n+\t\treturn;\n+\t}\n \n \t/* Pick up the string around '@', possibly delimited with <>\n-\t * pair; that is the email part.  White them out while copying.\n+\t * pair; that is the email part.\n \t */\n-\twhile (at > line) {\n+\twhile (at > f.buf) {\n \t\tchar c = at[-1];\n \t\tif (isspace(c))\n \t\t\tbreak;\n@@ -98,56 +102,35 @@ static int handle_from(char *in_line)\n \t\t}\n \t\tat--;\n \t}\n-\tdst = email;\n-\tfor (;;) {\n-\t\tunsigned char c = *at;\n-\t\tif (!c || c == '>' || isspace(c)) {\n-\t\t\tif (c == '>')\n-\t\t\t\t*at = ' ';\n-\t\t\tbreak;\n-\t\t}\n-\t\t*at++ = ' ';\n-\t\t*dst++ = c;\n-\t}\n-\t*dst++ = 0;\n+\tel = strcspn(at, \" \\n\\t\\r\\v\\f>\");\n+\tstrbuf_reset(&email);\n+\tstrbuf_add(&email, at, el);\n+\tstrbuf_remove(&f, at - f.buf, el + 1);\n \n \t/* The remainder is name.  It could be \"John Doe <john.doe@xz>\"\n-\t * or \"john.doe@xz (John Doe)\", but we have whited out the\n+\t * or \"john.doe@xz (John Doe)\", but we have removed the\n \t * email part, so trim from both ends, possibly removing\n \t * the () pair at the end.\n \t */\n-\tat = line + strlen(line);\n-\twhile (at > line) {\n-\t\tunsigned char c = *--at;\n-\t\tif (!isspace(c)) {\n-\t\t\tat[(c == ')') ? 0 : 1] = 0;\n-\t\t\tbreak;\n-\t\t}\n-\t}\n+\tstrbuf_trim(&f);\n+\tif (f.buf[0] == '(')\n+\t\tstrbuf_remove(&name, 0, 1);\n+\tif (f.len && f.buf[f.len - 1] == ')')\n+\t\tstrbuf_setlen(&f, f.len - 1);\n \n-\tat = line;\n-\tfor (;;) {\n-\t\tunsigned char c = *at;\n-\t\tif (!c || !isspace(c)) {\n-\t\t\tif (c == '(')\n-\t\t\t\tat++;\n-\t\t\tbreak;\n-\t\t}\n-\t\tat++;\n-\t}\n-\tat = sanity_check(at, email);\n-\tstrcpy(name, at);\n-\treturn 1;\n+\tget_sane_name(&name, &f, &email);\n+\tstrbuf_release(&f);\n }\n \n-static int handle_header(char *line, char *data, int ofs)\n+static void handle_header(struct strbuf **out, const struct strbuf *line)\n {\n-\tif (!line || !data)\n-\t\treturn 1;\n-\n-\tstrcpy(data, line+ofs);\n+\tif (!*out) {\n+\t\t*out = xmalloc(sizeof(struct strbuf));\n+\t\tstrbuf_init(*out, line->len);\n+\t} else\n+\t\tstrbuf_reset(*out);\n \n-\treturn 0;\n+\tstrbuf_addbuf(*out, line);\n }\n \n /* NOTE NOTE NOTE.  We do not claim we do full MIME.  We just attempt\n@@ -156,13 +139,13 @@ static int handle_header(char *line, char *data, int ofs)\n  * case insensitively.\n  */\n \n-static int slurp_attr(const char *line, const char *name, char *attr)\n+static int slurp_attr(const char *line, const char *name, struct strbuf *attr)\n {\n \tconst char *ends, *ap = strcasestr(line, name);\n \tsize_t sz;\n \n \tif (!ap) {\n-\t\t*attr = 0;\n+\t\tstrbuf_setlen(attr, 0);\n \t\treturn 0;\n \t}\n \tap += strlen(name);\n@@ -173,180 +156,171 @@ static int slurp_attr(const char *line, const char *name, char *attr)\n \telse\n \t\tends = \"; \\t\";\n \tsz = strcspn(ap, ends);\n-\tmemcpy(attr, ap, sz);\n-\tattr[sz] = 0;\n+\tstrbuf_add(attr, ap, sz);\n \treturn 1;\n }\n \n-struct content_type {\n-\tchar *boundary;\n-\tint boundary_len;\n-};\n-\n-static struct content_type content[MAX_BOUNDARIES];\n+static struct strbuf *content[MAX_BOUNDARIES];\n \n-static struct content_type *content_top = content;\n+static struct strbuf **content_top = content;\n \n-static int handle_content_type(char *line)\n+static void handle_content_type(struct strbuf *line)\n {\n-\tchar boundary[256];\n+\tstruct strbuf *boundary = xmalloc(sizeof(struct strbuf));\n+\tstrbuf_init(boundary, line->len);\n \n-\tif (strcasestr(line, \"text/\") == NULL)\n+\tif (!strcasestr(line->buf, \"text/\"))\n \t\t message_type = TYPE_OTHER;\n-\tif (slurp_attr(line, \"boundary=\", boundary + 2)) {\n-\t\tmemcpy(boundary, \"--\", 2);\n+\tif (slurp_attr(line->buf, \"boundary=\", boundary)) {\n+\t\tstrbuf_insert(boundary, 0, \"--\", 2);\n \t\tif (content_top++ >= &content[MAX_BOUNDARIES]) {\n \t\t\tfprintf(stderr, \"Too many boundaries to handle\\n\");\n \t\t\texit(1);\n \t\t}\n-\t\tcontent_top->boundary_len = strlen(boundary);\n-\t\tcontent_top->boundary = xmalloc(content_top->boundary_len+1);\n-\t\tstrcpy(content_top->boundary, boundary);\n+\t\t*content_top = boundary;\n+\t\tboundary = NULL;\n \t}\n-\tif (slurp_attr(line, \"charset=\", charset)) {\n-\t\tint i, c;\n-\t\tfor (i = 0; (c = charset[i]) != 0; i++)\n-\t\t\tcharset[i] = tolower(c);\n+\tif (slurp_attr(line->buf, \"charset=\", &charset))\n+\t\tstrbuf_tolower(&charset);\n+\n+\tif (boundary) {\n+\t\tstrbuf_release(boundary);\n+\t\tfree(boundary);\n \t}\n-\treturn 0;\n }\n \n-static int handle_content_transfer_encoding(char *line)\n+static void handle_content_transfer_encoding(const struct strbuf *line)\n {\n-\tif (strcasestr(line, \"base64\"))\n+\tif (strcasestr(line->buf, \"base64\"))\n \t\ttransfer_encoding = TE_BASE64;\n-\telse if (strcasestr(line, \"quoted-printable\"))\n+\telse if (strcasestr(line->buf, \"quoted-printable\"))\n \t\ttransfer_encoding = TE_QP;\n \telse\n \t\ttransfer_encoding = TE_DONTCARE;\n-\treturn 0;\n-}\n-\n-static int is_multipart_boundary(const char *line)\n-{\n-\treturn (!memcmp(line, content_top->boundary, content_top->boundary_len));\n }\n \n-static int eatspace(char *line)\n+static int is_multipart_boundary(const struct strbuf *line)\n {\n-\tint len = strlen(line);\n-\twhile (len > 0 && isspace(line[len-1]))\n-\t\tline[--len] = 0;\n-\treturn len;\n+\treturn !strbuf_cmp(line, *content_top);\n }\n \n-static char *cleanup_subject(char *subject)\n+static void cleanup_subject(struct strbuf *subject)\n {\n-\tfor (;;) {\n-\t\tchar *p;\n-\t\tint len, remove;\n-\t\tswitch (*subject) {\n+\tchar *pos;\n+\tsize_t remove;\n+\twhile (subject->len) {\n+\t\tswitch (*subject->buf) {\n \t\tcase 'r': case 'R':\n-\t\t\tif (!memcmp(\"e:\", subject+1, 2)) {\n-\t\t\t\tsubject += 3;\n+\t\t\tif (subject->len <= 3)\n+\t\t\t\tbreak;\n+\t\t\tif (!memcmp(subject->buf + 1, \"e:\", 2)) {\n+\t\t\t\tstrbuf_remove(subject, 0, 3);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tbreak;\n \t\tcase ' ': case '\\t': case ':':\n-\t\t\tsubject++;\n+\t\t\tstrbuf_remove(subject, 0, 1);\n \t\t\tcontinue;\n-\n \t\tcase '[':\n-\t\t\tp = strchr(subject, ']');\n-\t\t\tif (!p) {\n-\t\t\t\tsubject++;\n-\t\t\t\tcontinue;\n-\t\t\t}\n-\t\t\tlen = strlen(p);\n-\t\t\tremove = p - subject;\n-\t\t\tif (remove <= len *2) {\n-\t\t\t\tsubject = p+1;\n-\t\t\t\tcontinue;\n-\t\t\t}\n+\t\t\tif ((pos = strchr(subject->buf, ']'))) {\n+\t\t\t\tremove = pos - subject->buf + 1;\n+\t\t\t\t/* Don't remove too much. */\n+\t\t\t\tif (remove <= (subject->len - remove + 1) * 2) {\n+\t\t\t\t\tstrbuf_remove(subject, 0, remove);\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n+\t\t\t} else\n+\t\t\t\tstrbuf_remove(subject, 0, 1);\n \t\t\tbreak;\n \t\t}\n-\t\teatspace(subject);\n-\t\treturn subject;\n+\t\tstrbuf_trim(subject);\n+\t\treturn;\n \t}\n }\n \n-static void cleanup_space(char *buf)\n+static void cleanup_space(struct strbuf *sb)\n {\n-\tunsigned char c;\n-\twhile ((c = *buf) != 0) {\n-\t\tbuf++;\n-\t\tif (isspace(c)) {\n-\t\t\tbuf[-1] = ' ';\n-\t\t\tc = *buf;\n-\t\t\twhile (isspace(c)) {\n-\t\t\t\tint len = strlen(buf);\n-\t\t\t\tmemmove(buf, buf+1, len);\n-\t\t\t\tc = *buf;\n-\t\t\t}\n+\tsize_t pos, cnt;\n+\tfor (pos = 0; pos < sb->len; pos++) {\n+\t\tif (isspace(sb->buf[pos])) {\n+\t\t\tsb->buf[pos] = ' ';\n+\t\t\tfor (cnt = 0; isspace(sb->buf[pos + cnt + 1]); cnt++);\n+\t\t\tstrbuf_remove(sb, pos + 1, cnt);\n \t\t}\n \t}\n }\n \n-static void decode_header(char *it, unsigned itsize);\n+static void decode_header(struct strbuf *line);\n static const char *header[MAX_HDR_PARSED] = {\n \t\"From\",\"Subject\",\"Date\",\n };\n \n-static int check_header(char *line, unsigned linesize, char **hdr_data, int overwrite)\n+static inline int cmp_header(const struct strbuf *line, const char *hdr)\n {\n-\tint i;\n+\tint len = strlen(hdr);\n+\treturn !strncasecmp(line->buf, hdr, len) && line->len > len &&\n+\t\t\tline->buf[len] == ':' && isspace(line->buf[len + 1]);\n+}\n \n+static int check_header(const struct strbuf *line,\n+\t\t\t\tstruct strbuf *hdr_data[], int overwrite)\n+{\n+\tint i, ret = 0, len;\n+\tstruct strbuf sb = STRBUF_INIT;\n \t/* search for the interesting parts */\n \tfor (i = 0; header[i]; i++) {\n \t\tint len = strlen(header[i]);\n-\t\tif ((!hdr_data[i] || overwrite) &&\n-\t\t    !strncasecmp(line, header[i], len) &&\n-\t\t    line[len] == ':' && isspace(line[len + 1])) {\n+\t\tif ((!hdr_data[i] || overwrite) && cmp_header(line, header[i])) {\n \t\t\t/* Unwrap inline B and Q encoding, and optionally\n \t\t\t * normalize the meta information to utf8.\n \t\t\t */\n-\t\t\tdecode_header(line + len + 2, linesize - len - 2);\n-\t\t\thdr_data[i] = xmalloc(1000 * sizeof(char));\n-\t\t\tif (! handle_header(line, hdr_data[i], len + 2)) {\n-\t\t\t\treturn 1;\n-\t\t\t}\n+\t\t\tstrbuf_add(&sb, line->buf + len + 2, line->len - len -2);\n+\t\t\tdecode_header(&sb);\n+\t\t\thandle_header(&hdr_data[i], &sb);\n+\t\t\tret = 1;\n+\t\t\tgoto check_header_out;\n \t\t}\n \t}\n \n \t/* Content stuff */\n-\tif (!strncasecmp(line, \"Content-Type\", 12) &&\n-\t\tline[12] == ':' && isspace(line[12 + 1])) {\n-\t\tdecode_header(line + 12 + 2, linesize - 12 - 2);\n-\t\tif (! handle_content_type(line)) {\n-\t\t\treturn 1;\n-\t\t}\n-\t}\n-\tif (!strncasecmp(line, \"Content-Transfer-Encoding\", 25) &&\n-\t\tline[25] == ':' && isspace(line[25 + 1])) {\n-\t\tdecode_header(line + 25 + 2, linesize - 25 - 2);\n-\t\tif (! handle_content_transfer_encoding(line)) {\n-\t\t\treturn 1;\n-\t\t}\n+\tif (cmp_header(line, \"Content-Type\")) {\n+\t\tlen = strlen(\"Content-Type: \");\n+\t\tstrbuf_add(&sb, line->buf + len, line->len - len);\n+\t\tdecode_header(&sb);\n+\t\tstrbuf_insert(&sb, 0, \"Content-Type: \", len);\n+\t\thandle_content_type(&sb);\n+\t\tret = 1;\n+\t\tgoto check_header_out;\n+\t}\n+\tif (cmp_header(line, \"Content-Transfer-Encoding\")) {\n+\t\tlen = strlen(\"Content-Transfer-Encoding: \");\n+\t\tstrbuf_add(&sb, line->buf + len, line->len - len);\n+\t\tdecode_header(&sb);\n+\t\thandle_content_transfer_encoding(&sb);\n+\t\tret = 1;\n+\t\tgoto check_header_out;\n \t}\n \n \t/* for inbody stuff */\n-\tif (!memcmp(\">From\", line, 5) && isspace(line[5]))\n-\t\treturn 1;\n-\tif (!memcmp(\"[PATCH]\", line, 7) && isspace(line[7])) {\n+\tif (!prefixcmp(line->buf, \">From\") && isspace(line->buf[5]))\n+\t\tret = 1; /* Should this return 0? */\n+\t\tgoto check_header_out;\n+\tif (!prefixcmp(line->buf, \"[PATCH]\") && isspace(line->buf[7])) {\n \t\tfor (i = 0; header[i]; i++) {\n \t\t\tif (!memcmp(\"Subject\", header[i], 7)) {\n-\t\t\t\tif (! handle_header(line, hdr_data[i], 0)) {\n-\t\t\t\t\treturn 1;\n-\t\t\t\t}\n+\t\t\t\thandle_header(&hdr_data[i], line);\n+\t\t\t\tret = 1;\n+\t\t\t\tgoto check_header_out;\n \t\t\t}\n \t\t}\n \t}\n \n-\t/* no match */\n-\treturn 0;\n+check_header_out:\n+\tstrbuf_release(&sb);\n+\treturn ret;\n }\n \n-static int is_rfc2822_header(char *line)\n+static int is_rfc2822_header(const struct strbuf *line)\n {\n \t/*\n \t * The section that defines the loosest possible\n@@ -357,15 +331,15 @@ static int is_rfc2822_header(char *line)\n \t * ftext = %d33-57 / %59-126\n \t */\n \tint ch;\n-\tchar *cp = line;\n+\tchar *cp = line->buf;\n \n \t/* Count mbox From headers as headers */\n-\tif (!memcmp(line, \"From \", 5) || !memcmp(line, \">From \", 6))\n+\tif (!prefixcmp(cp, \"From \") || !prefixcmp(cp, \">From \"))\n \t\treturn 1;\n \n \twhile ((ch = *cp++)) {\n \t\tif (ch == ':')\n-\t\t\treturn cp != line;\n+\t\t\treturn 1;\n \t\tif ((33 <= ch && ch <= 57) ||\n \t\t    (59 <= ch && ch <= 126))\n \t\t\tcontinue;\n@@ -374,34 +348,20 @@ static int is_rfc2822_header(char *line)\n \treturn 0;\n }\n \n-/*\n- * sz is size of 'line' buffer in bytes.  Must be reasonably\n- * long enough to hold one physical real-world e-mail line.\n- */\n-static int read_one_header_line(char *line, int sz, FILE *in)\n+static int read_one_header_line(struct strbuf *line, FILE *in)\n {\n-\tint len;\n-\n-\t/*\n-\t * We will read at most (sz-1) bytes and then potentially\n-\t * re-add NUL after it.  Accessing line[sz] after this is safe\n-\t * and we can allow len to grow up to and including sz.\n-\t */\n-\tsz--;\n-\n \t/* Get the first part of the line. */\n-\tif (!fgets(line, sz, in))\n+\tif (strbuf_getline(line, in, '\\n'))\n \t\treturn 0;\n \n \t/*\n \t * Is it an empty line or not a valid rfc2822 header?\n \t * If so, stop here, and return false (\"not a header\")\n \t */\n-\tlen = eatspace(line);\n-\tif (!len || !is_rfc2822_header(line)) {\n+\tstrbuf_rtrim(line);\n+\tif (!line->len || !is_rfc2822_header(line)) {\n \t\t/* Re-add the newline */\n-\t\tline[len] = '\\n';\n-\t\tline[len + 1] = '\\0';\n+\t\tstrbuf_addch(line, '\\n');\n \t\treturn 0;\n \t}\n \n@@ -410,65 +370,53 @@ static int read_one_header_line(char *line, int sz, FILE *in)\n \t * Yuck, 2822 header \"folding\"\n \t */\n \tfor (;;) {\n-\t\tint peek, addlen;\n-\t\tstatic char continuation[1000];\n+\t\tint peek;\n+\t\tstruct strbuf continuation = STRBUF_INIT;\n \n \t\tpeek = fgetc(in); ungetc(peek, in);\n \t\tif (peek != ' ' && peek != '\\t')\n \t\t\tbreak;\n-\t\tif (!fgets(continuation, sizeof(continuation), in))\n+\t\tif (strbuf_getline(&continuation, in, '\\n'))\n \t\t\tbreak;\n-\t\taddlen = eatspace(continuation);\n-\t\tif (len < sz - 1) {\n-\t\t\tif (addlen >= sz - len)\n-\t\t\t\taddlen = sz - len - 1;\n-\t\t\tmemcpy(line + len, continuation, addlen);\n-\t\t\tline[len] = '\\n';\n-\t\t\tlen += addlen;\n-\t\t}\n+\t\tcontinuation.buf[0] = '\\n';\n+\t\tstrbuf_rtrim(&continuation);\n+\t\tstrbuf_addbuf(line, &continuation);\n \t}\n-\tline[len] = 0;\n \n \treturn 1;\n }\n \n-static int decode_q_segment(char *in, char *ot, unsigned otsize, char *ep, int rfc2047)\n+static struct strbuf *decode_q_segment(const struct strbuf *q_seg, int rfc2047)\n {\n-\tchar *otbegin = ot;\n-\tchar *otend = ot + otsize;\n+\tconst char *in = q_seg->buf;\n \tint c;\n-\twhile ((c = *in++) != 0 && (in <= ep)) {\n-\t\tif (ot == otend) {\n-\t\t\t*--ot = '\\0';\n-\t\t\treturn -1;\n-\t\t}\n+\tstruct strbuf *out = xmalloc(sizeof(struct strbuf));\n+\tstrbuf_init(out, q_seg->len);\n+\n+\twhile ((c = *in++) != 0) {\n \t\tif (c == '=') {\n \t\t\tint d = *in++;\n \t\t\tif (d == '\\n' || !d)\n \t\t\t\tbreak; /* drop trailing newline */\n-\t\t\t*ot++ = ((hexval(d) << 4) | hexval(*in++));\n+\t\t\tstrbuf_addch(out, (hexval(d) << 4) | hexval(*in++));\n \t\t\tcontinue;\n \t\t}\n \t\tif (rfc2047 && c == '_') /* rfc2047 4.2 (2) */\n \t\t\tc = 0x20;\n-\t\t*ot++ = c;\n+\t\tstrbuf_addch(out, c);\n \t}\n-\t*ot = 0;\n-\treturn (ot - otbegin);\n+\treturn out;\n }\n \n-static int decode_b_segment(char *in, char *ot, unsigned otsize, char *ep)\n+static struct strbuf *decode_b_segment(const struct strbuf *b_seg)\n {\n \t/* Decode in..ep, possibly in-place to ot */\n \tint c, pos = 0, acc = 0;\n-\tchar *otbegin = ot;\n-\tchar *otend = ot + otsize;\n+\tconst char *in = b_seg->buf;\n+\tstruct strbuf *out = xmalloc(sizeof(struct strbuf));\n+\tstrbuf_init(out, b_seg->len);\n \n-\twhile ((c = *in++) != 0 && (in <= ep)) {\n-\t\tif (ot == otend) {\n-\t\t\t*--ot = '\\0';\n-\t\t\treturn -1;\n-\t\t}\n+\twhile ((c = *in++) != 0) {\n \t\tif (c == '+')\n \t\t\tc = 62;\n \t\telse if (c == '/')\n@@ -493,21 +441,20 @@ static int decode_b_segment(char *in, char *ot, unsigned otsize, char *ep)\n \t\t\tacc = (c << 2);\n \t\t\tbreak;\n \t\tcase 1:\n-\t\t\t*ot++ = (acc | (c >> 4));\n+\t\t\tstrbuf_addch(out, (acc | (c >> 4)));\n \t\t\tacc = (c & 15) << 4;\n \t\t\tbreak;\n \t\tcase 2:\n-\t\t\t*ot++ = (acc | (c >> 2));\n+\t\t\tstrbuf_addch(out, (acc | (c >> 2)));\n \t\t\tacc = (c & 3) << 6;\n \t\t\tbreak;\n \t\tcase 3:\n-\t\t\t*ot++ = (acc | c);\n+\t\t\tstrbuf_addch(out, (acc | c));\n \t\t\tacc = pos = 0;\n \t\t\tbreak;\n \t\t}\n \t}\n-\t*ot = 0;\n-\treturn (ot - otbegin);\n+\treturn out;\n }\n \n /*\n@@ -521,16 +468,16 @@ static int decode_b_segment(char *in, char *ot, unsigned otsize, char *ep)\n  * Otherwise, we default to assuming it is Latin1 for historical\n  * reasons.\n  */\n-static const char *guess_charset(const char *line, const char *target_charset)\n+static const char *guess_charset(const struct strbuf *line, const char *target_charset)\n {\n \tif (is_encoding_utf8(target_charset)) {\n-\t\tif (is_utf8(line))\n+\t\tif (is_utf8(line->buf))\n \t\t\treturn NULL;\n \t}\n \treturn \"latin1\";\n }\n \n-static void convert_to_utf8(char *line, unsigned linesize, const char *charset)\n+static void convert_to_utf8(struct strbuf *line, const char *charset)\n {\n \tchar *out;\n \n@@ -542,112 +489,119 @@ static void convert_to_utf8(char *line, unsigned linesize, const char *charset)\n \n \tif (!strcmp(metainfo_charset, charset))\n \t\treturn;\n-\tout = reencode_string(line, metainfo_charset, charset);\n+\tout = reencode_string(line->buf, metainfo_charset, charset);\n \tif (!out)\n \t\tdie(\"cannot convert from %s to %s\\n\",\n \t\t    charset, metainfo_charset);\n-\tstrlcpy(line, out, linesize);\n-\tfree(out);\n+\tstrbuf_attach(line, out, strlen(out), strlen(out));\n }\n \n-static int decode_header_bq(char *it, unsigned itsize)\n+static int decode_header_bq(struct strbuf *it)\n {\n \tchar *in, *out, *ep, *cp, *sp;\n-\tchar outbuf[1000];\n+\tstruct strbuf outbuf = STRBUF_INIT, *dec;\n+\tstruct strbuf charset_q = STRBUF_INIT, piecebuf = STRBUF_INIT;\n \tint rfc2047 = 0;\n \n-\tin = it;\n-\tout = outbuf;\n-\twhile ((ep = strstr(in, \"=?\")) != NULL) {\n-\t\tint sz, encoding;\n-\t\tchar charset_q[256], piecebuf[256];\n+\tin = it->buf;\n+\twhile (in - it->buf <= it->len && (ep = strstr(in, \"=?\")) != NULL) {\n+\t\tint encoding;\n+\t\tstrbuf_reset(&charset_q);\n+\t\tstrbuf_reset(&piecebuf);\n \t\trfc2047 = 1;\n \n \t\tif (in != ep) {\n-\t\t\tsz = ep - in;\n-\t\t\tmemcpy(out, in, sz);\n-\t\t\tout += sz;\n-\t\t\tin += sz;\n+\t\t\tstrbuf_add(&outbuf, in, ep - in);\n+\t\t\tin = ep;\n \t\t}\n \t\t/* E.g.\n \t\t * ep : \"=?iso-2022-jp?B?GyR...?= foo\"\n \t\t * ep : \"=?ISO-8859-1?Q?Foo=FCbar?= baz\"\n \t\t */\n \t\tep += 2;\n-\t\tcp = strchr(ep, '?');\n-\t\tif (!cp)\n-\t\t\treturn rfc2047; /* no munging */\n-\t\tfor (sp = ep; sp < cp; sp++)\n-\t\t\tcharset_q[sp - ep] = tolower(*sp);\n-\t\tcharset_q[cp - ep] = 0;\n+\n+\t\tif (ep - it->buf >= it->len || !(cp = strchr(ep, '?')))\n+\t\t\tgoto decode_header_bq_out;\n+\n+\t\tif (cp + 3 - it->buf > it->len)\n+\t\t\tgoto decode_header_bq_out;\n+\t\tstrbuf_add(&charset_q, ep, cp - ep);\n+\t\tstrbuf_tolower(&charset_q);\n+\n \t\tencoding = cp[1];\n \t\tif (!encoding || cp[2] != '?')\n-\t\t\treturn rfc2047; /* no munging */\n+\t\t\tgoto decode_header_bq_out;\n \t\tep = strstr(cp + 3, \"?=\");\n \t\tif (!ep)\n-\t\t\treturn rfc2047; /* no munging */\n+\t\t\tgoto decode_header_bq_out;\n+\t\tstrbuf_add(&piecebuf, cp + 3, ep - cp - 3);\n \t\tswitch (tolower(encoding)) {\n \t\tdefault:\n-\t\t\treturn rfc2047; /* no munging */\n+\t\t\tgoto decode_header_bq_out;\n \t\tcase 'b':\n-\t\t\tsz = decode_b_segment(cp + 3, piecebuf, sizeof(piecebuf), ep);\n+\t\t\tdec = decode_b_segment(&piecebuf);\n \t\t\tbreak;\n \t\tcase 'q':\n-\t\t\tsz = decode_q_segment(cp + 3, piecebuf, sizeof(piecebuf), ep, 1);\n+\t\t\tdec = decode_q_segment(&piecebuf, 1);\n \t\t\tbreak;\n \t\t}\n-\t\tif (sz < 0)\n-\t\t\treturn rfc2047;\n \t\tif (metainfo_charset)\n-\t\t\tconvert_to_utf8(piecebuf, sizeof(piecebuf), charset_q);\n+\t\t\tconvert_to_utf8(dec, charset_q.buf);\n \n-\t\tsz = strlen(piecebuf);\n-\t\tif (outbuf + sizeof(outbuf) <= out + sz)\n-\t\t\treturn rfc2047; /* no munging */\n-\t\tstrcpy(out, piecebuf);\n-\t\tout += sz;\n+\t\tstrbuf_addbuf(&outbuf, dec);\n+\t\tstrbuf_release(dec);\n+\t\tfree(dec);\n \t\tin = ep + 2;\n \t}\n-\tstrcpy(out, in);\n-\tstrlcpy(it, outbuf, itsize);\n+\tstrbuf_addstr(&outbuf, in);\n+\tstrbuf_reset(it);\n+\tstrbuf_addbuf(it, &outbuf);\n+decode_header_bq_out:\n+\tstrbuf_release(&outbuf);\n+\tstrbuf_release(&charset_q);\n+\tstrbuf_release(&piecebuf);\n \treturn rfc2047;\n }\n \n-static void decode_header(char *it, unsigned itsize)\n+static void decode_header(struct strbuf *it)\n {\n-\n-\tif (decode_header_bq(it, itsize))\n+\tif (decode_header_bq(it))\n \t\treturn;\n \t/* otherwise \"it\" is a straight copy of the input.\n \t * This can be binary guck but there is no charset specified.\n \t */\n \tif (metainfo_charset)\n-\t\tconvert_to_utf8(it, itsize, \"\");\n+\t\tconvert_to_utf8(it, \"\");\n }\n \n-static int decode_transfer_encoding(char *line, unsigned linesize, int inputlen)\n+static void decode_transfer_encoding(struct strbuf *line)\n {\n-\tchar *ep;\n+\tstruct strbuf *ret;\n+\tint len;\n \n \tswitch (transfer_encoding) {\n \tcase TE_QP:\n-\t\tep = line + inputlen;\n-\t\treturn decode_q_segment(line, line, linesize, ep, 0);\n+\t\tret = decode_q_segment(line, 0);\n+\t\tbreak;\n \tcase TE_BASE64:\n-\t\tep = line + inputlen;\n-\t\treturn decode_b_segment(line, line, linesize, ep);\n+\t\tret = decode_b_segment(line);\n+\t\tbreak;\n \tcase TE_DONTCARE:\n \tdefault:\n-\t\treturn inputlen;\n+\t\treturn;\n \t}\n+\tstrbuf_reset(line);\n+\tstrbuf_addbuf(line, ret);\n+\tstrbuf_release(ret);\n+\tfree(ret);\n }\n \n-static int handle_filter(char *line, unsigned linesize, int linelen);\n+static void handle_filter(struct strbuf *line);\n \n static int find_boundary(void)\n {\n-\twhile(fgets(line, sizeof(line), fin) != NULL) {\n-\t\tif (is_multipart_boundary(line))\n+\twhile(!strbuf_getline(&line, fin, '\\n')) {\n+\t\tif (is_multipart_boundary(&line))\n \t\t\treturn 1;\n \t}\n \treturn 0;\n@@ -655,12 +609,17 @@ static int find_boundary(void)\n \n static int handle_boundary(void)\n {\n-\tchar newline[]=\"\\n\";\n+\tstruct strbuf newline = STRBUF_INIT;\n+\n+\tstrbuf_addch(&newline, '\\n');\n again:\n-\tif (!memcmp(line+content_top->boundary_len, \"--\", 2)) {\n+\tif (line.len >= (*content_top)->len + 2 &&\n+\t    !memcmp(line.buf + (*content_top)->len, \"--\", 2)) {\n \t\t/* we hit an end boundary */\n \t\t/* pop the current boundary off the stack */\n-\t\tfree(content_top->boundary);\n+\t\tstrbuf_release(*content_top);\n+\t\tfree(*content_top);\n+\t\t*content_top = NULL;\n \n \t\t/* technically won't happen as is_multipart_boundary()\n \t\t   will fail first.  But just in case..\n@@ -670,7 +629,8 @@ again:\n \t\t\t\t\t\"can't recover\\n\");\n \t\t\texit(1);\n \t\t}\n-\t\thandle_filter(newline, sizeof(newline), strlen(newline));\n+\t\thandle_filter(&newline);\n+\t\tstrbuf_release(&newline);\n \n \t\t/* skip to the next boundary */\n \t\tif (!find_boundary())\n@@ -680,39 +640,44 @@ again:\n \n \t/* set some defaults */\n \ttransfer_encoding = TE_DONTCARE;\n-\tcharset[0] = 0;\n+\tstrbuf_reset(&charset);\n \tmessage_type = TYPE_TEXT;\n \n \t/* slurp in this section's info */\n-\twhile (read_one_header_line(line, sizeof(line), fin))\n-\t\tcheck_header(line, sizeof(line), p_hdr_data, 0);\n+\twhile (read_one_header_line(&line, fin))\n+\t\tcheck_header(&line, p_hdr_data, 0);\n \n+\tstrbuf_release(&newline);\n \t/* eat the blank line after section info */\n-\treturn (fgets(line, sizeof(line), fin) != NULL);\n+\treturn (strbuf_getline(&line, fin, '\\n') == 0);\n }\n \n-static inline int patchbreak(const char *line)\n+static inline int patchbreak(const struct strbuf *line)\n {\n+\tsize_t i;\n+\n \t/* Beginning of a \"diff -\" header? */\n-\tif (!memcmp(\"diff -\", line, 6))\n+\tif (!prefixcmp(line->buf, \"diff -\"))\n \t\treturn 1;\n \n \t/* CVS \"Index: \" line? */\n-\tif (!memcmp(\"Index: \", line, 7))\n+\tif (!prefixcmp(line->buf, \"Index: \"))\n \t\treturn 1;\n \n \t/*\n \t * \"--- <filename>\" starts patches without headers\n \t * \"---<sp>*\" is a manual separator\n \t */\n-\tif (!memcmp(\"---\", line, 3)) {\n-\t\tline += 3;\n+\tif (line->len < 4)\n+\t\treturn 0;\n+\n+\tif (!prefixcmp(line->buf, \"---\")) {\n \t\t/* space followed by a filename? */\n-\t\tif (line[0] == ' ' && !isspace(line[1]))\n+\t\tif (line->buf[3] == ' ' && !isspace(line->buf[4]))\n \t\t\treturn 1;\n \t\t/* Just whitespace? */\n-\t\tfor (;;) {\n-\t\t\tunsigned char c = *line++;\n+\t\tfor (i = 3; i < line->len; i++) {\n+\t\t\tunsigned char c = line->buf[i];\n \t\t\tif (c == '\\n')\n \t\t\t\treturn 1;\n \t\t\tif (!isspace(c))\n@@ -723,32 +688,25 @@ static inline int patchbreak(const char *line)\n \treturn 0;\n }\n \n-\n-static int handle_commit_msg(char *line, unsigned linesize)\n+static int handle_commit_msg(struct strbuf *line)\n {\n \tstatic int still_looking = 1;\n-\tchar *endline = line + linesize;\n+\tchar *c;\n \n \tif (!cmitmsg)\n \t\treturn 0;\n \n \tif (still_looking) {\n-\t\tchar *cp = line;\n-\t\tif (isspace(*line)) {\n-\t\t\tfor (cp = line + 1; *cp; cp++) {\n-\t\t\t\tif (!isspace(*cp))\n-\t\t\t\t\tbreak;\n-\t\t\t}\n-\t\t\tif (!*cp)\n-\t\t\t\treturn 0;\n-\t\t}\n-\t\tif ((still_looking = check_header(cp, endline - cp, s_hdr_data, 0)) != 0)\n+\t\tstrbuf_ltrim(line);\n+\t\tif (!line->len)\n+\t\t\treturn 0;\n+\t\tif ((still_looking = check_header(line, s_hdr_data, 0)) != 0)\n \t\t\treturn 0;\n \t}\n \n \t/* normalize the log message to UTF-8. */\n \tif (metainfo_charset)\n-\t\tconvert_to_utf8(line, endline - line, charset);\n+\t\tconvert_to_utf8(line, charset.buf);\n \n \tif (patchbreak(line)) {\n \t\tfclose(cmitmsg);\n@@ -756,142 +714,132 @@ static int handle_commit_msg(char *line, unsigned linesize)\n \t\treturn 1;\n \t}\n \n-\tfputs(line, cmitmsg);\n+\tfputs(line->buf, cmitmsg);\n \treturn 0;\n }\n \n-static int handle_patch(char *line, int len)\n+static void handle_patch(const struct strbuf *line)\n {\n-\tfwrite(line, 1, len, patchfile);\n+\tfwrite(line->buf, 1, line->len, patchfile);\n \tpatch_lines++;\n-\treturn 0;\n }\n \n-static int handle_filter(char *line, unsigned linesize, int linelen)\n+static void handle_filter(struct strbuf *line)\n {\n \tstatic int filter = 0;\n \n-\t/* filter tells us which part we left off on\n-\t * a non-zero return indicates we hit a filter point\n-\t */\n+\t/* filter tells us which part we left off on */\n \tswitch (filter) {\n \tcase 0:\n-\t\tif (!handle_commit_msg(line, linesize))\n+\t\tif (!handle_commit_msg(line))\n \t\t\tbreak;\n \t\tfilter++;\n \tcase 1:\n-\t\tif (!handle_patch(line, linelen))\n-\t\t\tbreak;\n-\t\tfilter++;\n-\tdefault:\n-\t\treturn 1;\n+\t\thandle_patch(line);\n+\t\tbreak;\n \t}\n-\n-\treturn 0;\n }\n \n static void handle_body(void)\n {\n-\tint rc = 0;\n-\tstatic char newline[2000];\n-\tstatic char *np = newline;\n-\tint len = strlen(line);\n+\tint len = 0;\n+\tstruct strbuf prev = STRBUF_INIT;\n \n \t/* Skip up to the first boundary */\n-\tif (content_top->boundary) {\n+\tif (*content_top) {\n \t\tif (!find_boundary())\n-\t\t\treturn;\n+\t\t\tgoto handle_body_out;\n \t}\n \n \tdo {\n+\t\tstrbuf_setlen(&line, line.len + len);\n+\n \t\t/* process any boundary lines */\n-\t\tif (content_top->boundary && is_multipart_boundary(line)) {\n+\t\tif (*content_top && is_multipart_boundary(&line)) {\n \t\t\t/* flush any leftover */\n-\t\t\tif (np != newline)\n-\t\t\t\thandle_filter(newline, sizeof(newline),\n-\t\t\t\t\t      np - newline);\n+\t\t\tif (line.len)\n+\t\t\t\thandle_filter(&line);\n+\n \t\t\tif (!handle_boundary())\n-\t\t\t\treturn;\n-\t\t\tlen = strlen(line);\n+\t\t\t\tgoto handle_body_out;\n \t\t}\n \n \t\t/* Unwrap transfer encoding */\n-\t\tlen = decode_transfer_encoding(line, sizeof(line), len);\n-\t\tif (len < 0) {\n-\t\t\terror(\"Malformed input line\");\n-\t\t\treturn;\n-\t\t}\n+\t\tdecode_transfer_encoding(&line);\n \n \t\tswitch (transfer_encoding) {\n \t\tcase TE_BASE64:\n \t\tcase TE_QP:\n \t\t{\n-\t\t\tchar *op = line;\n+\t\t\tstruct strbuf **lines, **it, *sb;\n+\n+\t\t\t/* Prepend any previous partial lines */\n+\t\t\tstrbuf_insert(&line, 0, prev.buf, prev.len);\n+\t\t\tstrbuf_reset(&prev);\n \n \t\t\t/* binary data most likely doesn't have newlines */\n \t\t\tif (message_type != TYPE_TEXT) {\n-\t\t\t\trc = handle_filter(line, sizeof(line), len);\n+\t\t\t\thandle_filter(&line);\n \t\t\t\tbreak;\n \t\t\t}\n-\n \t\t\t/*\n \t\t\t * This is a decoded line that may contain\n \t\t\t * multiple new lines.  Pass only one chunk\n \t\t\t * at a time to handle_filter()\n \t\t\t */\n-\t\t\tdo {\n-\t\t\t\twhile (op < line + len && *op != '\\n')\n-\t\t\t\t\t*np++ = *op++;\n-\t\t\t\t*np = *op;\n-\t\t\t\tif (*np != 0) {\n-\t\t\t\t\t/* should be sitting on a new line */\n-\t\t\t\t\t*(++np) = 0;\n-\t\t\t\t\top++;\n-\t\t\t\t\trc = handle_filter(newline, sizeof(newline), np - newline);\n-\t\t\t\t\tnp = newline;\n-\t\t\t\t}\n-\t\t\t} while (op < line + len);\n+\t\t\tlines = strbuf_split(&line, '\\n');\n+\t\t\tfor (it = lines; (sb = *it); it++) {\n+\t\t\t\tif (*(it + 1) == NULL) /* The last line */\n+\t\t\t\t\tif (sb->buf[sb->len - 1] != '\\n') {\n+\t\t\t\t\t\t/* Partial line, save it for later. */\n+\t\t\t\t\t\tstrbuf_addbuf(&prev, sb);\n+\t\t\t\t\t\tbreak;\n+\t\t\t\t\t}\n+\t\t\t\thandle_filter(sb);\n+\t\t\t}\n \t\t\t/*\n-\t\t\t * The partial chunk is saved in newline and will be\n+\t\t\t * The partial chunk is saved in \"prev\" and will be\n \t\t\t * appended by the next iteration of read_line_with_nul().\n \t\t\t */\n+\t\t\tstrbuf_list_free(lines);\n \t\t\tbreak;\n \t\t}\n \t\tdefault:\n-\t\t\trc = handle_filter(line, sizeof(line), len);\n+\t\t\thandle_filter(&line);\n \t\t}\n-\t\tif (rc)\n-\t\t\t/* nothing left to filter */\n-\t\t\tbreak;\n-\t} while ((len = read_line_with_nul(line, sizeof(line), fin)));\n \n-\treturn;\n+\t\tstrbuf_reset(&line);\n+\t\tif (strbuf_avail(&line) < 100)\n+\t\t\tstrbuf_grow(&line, 100);\n+\t} while ((len = read_line_with_nul(line.buf, strbuf_avail(&line), fin)));\n+\n+handle_body_out:\n+\tstrbuf_release(&prev);\n }\n \n-static void output_header_lines(FILE *fout, const char *hdr, char *data)\n+static void output_header_lines(FILE *fout, const char *hdr, const struct strbuf *data)\n {\n+\tconst char *sp = data->buf;\n \twhile (1) {\n-\t\tchar *ep = strchr(data, '\\n');\n+\t\tchar *ep = strchr(sp, '\\n');\n \t\tint len;\n \t\tif (!ep)\n-\t\t\tlen = strlen(data);\n+\t\t\tlen = strlen(sp);\n \t\telse\n-\t\t\tlen = ep - data;\n-\t\tfprintf(fout, \"%s: %.*s\\n\", hdr, len, data);\n+\t\t\tlen = ep - sp;\n+\t\tfprintf(fout, \"%s: %.*s\\n\", hdr, len, sp);\n \t\tif (!ep)\n \t\t\tbreak;\n-\t\tdata = ep + 1;\n+\t\tsp = ep + 1;\n \t}\n }\n \n static void handle_info(void)\n {\n-\tchar *sub;\n-\tchar *hdr;\n+\tstruct strbuf *hdr;\n \tint i;\n \n \tfor (i = 0; header[i]; i++) {\n-\n \t\t/* only print inbody headers if we output a patch file */\n \t\tif (patch_lines && s_hdr_data[i])\n \t\t\thdr = s_hdr_data[i];\n@@ -901,20 +849,18 @@ static void handle_info(void)\n \t\t\tcontinue;\n \n \t\tif (!memcmp(header[i], \"Subject\", 7)) {\n-\t\t\tif (keep_subject)\n-\t\t\t\tsub = hdr;\n-\t\t\telse {\n-\t\t\t\tsub = cleanup_subject(hdr);\n-\t\t\t\tcleanup_space(sub);\n+\t\t\tif (!keep_subject) {\n+\t\t\t\tcleanup_subject(hdr);\n+\t\t\t\tcleanup_space(hdr);\n \t\t\t}\n-\t\t\toutput_header_lines(fout, \"Subject\", sub);\n+\t\t\toutput_header_lines(fout, \"Subject\", hdr);\n \t\t} else if (!memcmp(header[i], \"From\", 4)) {\n \t\t\thandle_from(hdr);\n-\t\t\tfprintf(fout, \"Author: %s\\n\", name);\n-\t\t\tfprintf(fout, \"Email: %s\\n\", email);\n+\t\t\tfprintf(fout, \"Author: %s\\n\", name.buf);\n+\t\t\tfprintf(fout, \"Email: %s\\n\", email.buf);\n \t\t} else {\n \t\t\tcleanup_space(hdr);\n-\t\t\tfprintf(fout, \"%s: %s\\n\", header[i], hdr);\n+\t\t\tfprintf(fout, \"%s: %s\\n\", header[i], hdr->buf);\n \t\t}\n \t}\n \tfprintf(fout, \"\\n\");\n@@ -941,8 +887,8 @@ static int mailinfo(FILE *in, FILE *out, int ks, const char *encoding,\n \t\treturn -1;\n \t}\n \n-\tp_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(char *));\n-\ts_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(char *));\n+\tp_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(*p_hdr_data));\n+\ts_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(*s_hdr_data));\n \n \tdo {\n \t\tpeek = fgetc(in);\n@@ -950,8 +896,8 @@ static int mailinfo(FILE *in, FILE *out, int ks, const char *encoding,\n \tungetc(peek, in);\n \n \t/* process the email header */\n-\twhile (read_one_header_line(line, sizeof(line), fin))\n-\t\tcheck_header(line, sizeof(line), p_hdr_data, 1);\n+\twhile (read_one_header_line(&line, fin))\n+\t\tcheck_header(&line, p_hdr_data, 1);\n \n \thandle_body();\n \thandle_info();\n-- \n1.5.4.5\n"},{"id":"83181","messageId":"7v4p6tqxch.fsf@gitster.siamese.dyndns.org","threadId":"14398","inReplyTo":"487A49B4.6000407@etek.chalmers.se","subject":"Re: [PATCH] git-mailinfo: use strbuf's instead of fixed buffers","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-07-13T21:37:50Z","receivedAt":"2008-07-13T21:37:50Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lukas Sandström <lukass@etek.chalmers.se> writes:\n\n> +static void parse_bogus_from(const struct strbuf *line)\n>  {\n> ...\n> +static void handle_from(const struct strbuf *from)\n>  {\n> ...\n>  \tif (!at)\n> -\t\treturn bogus_from(line);\n> +\t\treturn parse_bogus_from(from);\n\nThat's GCCism isn't it?  I'll queue this probably for 'pu' with local\nfixups, so no need to resend for this particular issue, though.\n"},{"id":"83347","messageId":"20080715031356.GQ16127@redhat.com","threadId":"14398","inReplyTo":"7v3amfxx3a.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] git-mailinfo: Fix getting the subject from the body","fromName":"Don Zickus","fromEmail":"dzickus@redhat.com","sentAt":"2008-07-15T03:13:56Z","receivedAt":"2008-07-15T03:13:56Z","isPatch":true,"sender":{"key":"dzickus@redhat.com","avatar":null},"body":"On Sat, Jul 12, 2008 at 02:36:57AM -0700, Junio C Hamano wrote:\n> Another thing I noticed and found puzzling is the handling of \">From \"\n> line that is shown in the context below.  check_header() is supposed to\n> return true when it handled header (i.e. not part of the commit message)\n> and return false when line is not part of the header.  As \">From \" is part\n> of the commit log message, shouldn't it return zero?\n> \n> Don, this part was what you introduced.  Has this codepath ever been\n> exercised in the real life?\n\nHeh.  Most emails I deal with usually wind up causing the code to stop\nlooking for header info (still_looking=0).  So I never ran into that\nscenario.  And I never really tried to rely on inbody stuff.\n\nI thought I was mimicing the original code, guess not.\n\nNow that I think about it, I did run into a situation last year where\ngit-mailinfo parsed the '>From' as an inbody header instead of a commit\nmsg.  I just put a stupid hack in my scripts to work around, thinking it\nwas my scripts.\n\nAnyway if it returns zero, wouldn't it be better to just remove the check\nto begin with?  I kinda forgot why it is there in the first place (my\nchanges just copied it from somewhere else).\n\nCheers,\nDon\n"}]}