{"thread":{"id":"9779","subject":"[RFC] Convert builin-mailinfo.c to use The Better String Library.","startedAt":"2007-09-04T20:50:08Z","lastAt":"2012-05-22T18:30:13Z","messageCount":102,"participants":["Lukas Sandström","Alex Riesen","Pierre Habouzit","Kristian Høgsberg","Matthieu Moy","Miles Bader","Dmitry Kakurin","Shawn O. Pearce","Andreas Ericsson","Junio C Hamano","David Kastrup","Johannes Schindelin","Linus Torvalds","alan","Wincent Colaiuta","Paul Wankadia","Nicolas Pitre","David Symonds","Theodore Tso","Walter Bright","Andy Parkins","Karl Hasselström","John 'Z-Bo' Zabroski","Steven Burns","figo","Bernd Jendrissek","Ian Molton","Jakub Narebski","Dario Rodriguez","Syed M Raihan"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"52514","messageId":"46DDC500.5000606@etek.chalmers.se","threadId":"9779","inReplyTo":null,"subject":"[RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2007-09-04T20:50:08Z","receivedAt":"2007-09-04T20:50:08Z","isPatch":false,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Hi.\n\nThis is an attempt to use \"The Better String Library\"[1] in builtin-mailinfo.c\n\nThe patch doesn't pass all the tests in the testsuit yet, but I thought I'd\nsend it out so people can decide if they like how the code looks.\n\nI'm not sending a patch to add the library files at this time. I'll send\nthat patch when this patch is working.\n\nThe changes required to make it pass the tests shouldn't be very large.\n\n/Lukas\n\n[1] http://bstring.sourceforge.net/\n\n---\n builtin-mailinfo.c |  795 ++++++++++++++++++++++++++--------------------------\n 1 files changed, 392 insertions(+), 403 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex d7cb11d..2ddc15d 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -5,14 +5,14 @@\n #include \"cache.h\"\n #include \"builtin.h\"\n #include \"utf8.h\"\n+#include \"bstring/bstrlib.h\"\n \n static FILE *cmitmsg, *patchfile, *fin, *fout;\n \n static int keep_subject;\n-static const char *metainfo_charset;\n-static char line[1000];\n-static char name[1000];\n-static char email[1000];\n+static bstring metainfo_charset;\n+static bstring name;\n+static bstring email;\n \n static enum  {\n \tTE_DONTCARE, TE_QP, TE_BASE64,\n@@ -21,321 +21,291 @@ static enum  {\n \tTYPE_TEXT, TYPE_OTHER,\n } message_type;\n \n-static char charset[256];\n+static bstring charset;\n static int patch_lines;\n-static char **p_hdr_data, **s_hdr_data;\n+static bstring *p_hdr_data, *s_hdr_data;\n \n #define MAX_HDR_PARSED 10\n #define MAX_BOUNDARIES 5\n \n-static char *sanity_check(char *name, char *email)\n+static bstring sanity_check(bstring name, bstring email)\n {\n-\tint len = strlen(name);\n-\tif (len < 3 || len > 60)\n+\tstatic struct tagbstring email_ind = bsStatic(\"<@>\");\n+\tif (blength(name) < 3 || blength(name) > 60)\n \t\treturn email;\n-\tif (strchr(name, '@') || strchr(name, '<') || strchr(name, '>'))\n+\tif (binchr(name, 0, &email_ind) != BSTR_ERR)\n \t\treturn email;\n \treturn name;\n }\n \n-static int bogus_from(char *line)\n+static int bogus_from(const_bstring line)\n {\n \t/* John Doe <johndoe> */\n-\tchar *bra, *ket, *dst, *cp;\n-\n+\tint bra, ket;\n \t/* This is fallback, so do not bother if we already have an\n \t * e-mail address.\n \t */\n-\tif (*email)\n+\tif (blength(email))\n \t\treturn 0;\n \n-\tbra = strchr(line, '<');\n-\tif (!bra)\n+\tbra = bstrchr(line, '<');\n+\tif (bra == BSTR_ERR)\n \t\treturn 0;\n-\tket = strchr(bra, '>');\n-\tif (!ket)\n+\tket = bstrchrp(line, bra, '>');\n+\tif (ket == BSTR_ERR)\n \t\treturn 0;\n \n-\tfor (dst = email, cp = bra+1; cp < ket; )\n-\t\t*dst++ = *cp++;\n-\t*dst = 0;\n-\tfor (cp = line; isspace(*cp); cp++)\n-\t\t;\n-\tfor (bra--; isspace(*bra); bra--)\n-\t\t*bra = 0;\n-\tcp = sanity_check(cp, email);\n-\tstrcpy(name, cp);\n+\tbdestroy(email);\n+\temail = bmidstr(line, bra + 1, ket - bra - 1);\n+\n+\tname = bmidstr(line, 0, bra);\n+\tbtrimws(name);\n+\tbassign(name, sanity_check(name, email));\n \treturn 1;\n }\n \n-static int handle_from(char *in_line)\n+static int handle_from(const_bstring line)\n {\n-\tchar line[1000];\n-\tchar *at;\n-\tchar *dst;\n+\tint at, es, ee;\n+\tstatic struct tagbstring email_delim = bsStatic(\" \\n\\t\\r\\v\\f<>\");\n \n-\tstrcpy(line, in_line);\n-\tat = strchr(line, '@');\n-\tif (!at)\n+\tat = bstrchr(line, '@');\n+\tif (at == BSTR_ERR)\n \t\treturn bogus_from(line);\n \n \t/*\n \t * If we already have one email, don't take any confusing lines\n \t */\n-\tif (*email && strchr(at+1, '@'))\n+\tif (blength(email) && bstrchrp(line, '@', at + 1) != BSTR_ERR)\n \t\treturn 0;\n \n \t/* Pick up the string around '@', possibly delimited with <>\n-\t * pair; that is the email part.  White them out while copying.\n+\t * pair; that is the email part.\n \t */\n-\twhile (at > line) {\n-\t\tchar c = at[-1];\n-\t\tif (isspace(c))\n-\t\t\tbreak;\n-\t\tif (c == '<') {\n-\t\t\tat[-1] = ' ';\n-\t\t\tbreak;\n-\t\t}\n-\t\tat--;\n-\t}\n-\tdst = email;\n-\tfor (;;) {\n-\t\tunsigned char c = *at;\n-\t\tif (!c || c == '>' || isspace(c)) {\n-\t\t\tif (c == '>')\n-\t\t\t\t*at = ' ';\n-\t\t\tbreak;\n-\t\t}\n-\t\t*at++ = ' ';\n-\t\t*dst++ = c;\n-\t}\n-\t*dst++ = 0;\n+\n+\tes = binchrr(line, at, &email_delim);\n+\tee = binchr(line, at, &email_delim);\n+\tbdestroy(email);\n+\temail = bmidstr(line, es + 1, ee - es - 1);\n \n \t/* The remainder is name.  It could be \"John Doe <john.doe@xz>\"\n \t * or \"john.doe@xz (John Doe)\", but we have whited out the\n \t * email part, so trim from both ends, possibly removing\n \t * the () pair at the end.\n \t */\n-\tat = line + strlen(line);\n-\twhile (at > line) {\n-\t\tunsigned char c = *--at;\n-\t\tif (!isspace(c)) {\n-\t\t\tat[(c == ')') ? 0 : 1] = 0;\n-\t\t\tbreak;\n-\t\t}\n-\t}\n \n-\tat = line;\n-\tfor (;;) {\n-\t\tunsigned char c = *at;\n-\t\tif (!c || !isspace(c)) {\n-\t\t\tif (c == '(')\n-\t\t\t\tat++;\n-\t\t\tbreak;\n-\t\t}\n-\t\tat++;\n-\t}\n-\tat = sanity_check(at, email);\n-\tstrcpy(name, at);\n+\tbdestroy(name);\n+\tname = bstrcpy(line);\n+\tbdelete(name, es, ee - es + 1);\n+\tbtrimws(name);\n+\tif (bchar(name, 0) == '(')\n+\t\tbdelete(name, 0, 1);\n+\tif (bchar(name, blength(name) - 1) == ')')\n+\t\tbtrunc(name, blength(name) - 1);\n+\t\n+\tbassign(name, sanity_check(name, email));\n \treturn 1;\n }\n \n-static int handle_header(char *line, char *data, int ofs)\n-{\n-\tif (!line || !data)\n-\t\treturn 1;\n-\n-\tstrcpy(data, line+ofs);\n-\n-\treturn 0;\n-}\n-\n /* NOTE NOTE NOTE.  We do not claim we do full MIME.  We just attempt\n  * to have enough heuristics to grok MIME encoded patches often found\n  * on our mailing lists.  For example, we do not even treat header lines\n  * case insensitively.\n  */\n \n-static int slurp_attr(const char *line, const char *name, char *attr)\n+static bstring slurp_attr(const_bstring line, const char *name)\n {\n-\tconst char *ends, *ap = strcasestr(line, name);\n-\tsize_t sz;\n+\tint end, start;\n+\tstatic struct tagbstring endchars = bsStatic(\"; \\t\");\n+\tstruct tagbstring bname;\n \n-\tif (!ap) {\n-\t\t*attr = 0;\n-\t\treturn 0;\n-\t}\n-\tap += strlen(name);\n-\tif (*ap == '\"') {\n-\t\tap++;\n-\t\tends = \"\\\"\";\n+\tbtfromcstr(bname, name);\n+\tstart =  binstrcaseless(line, 0, &bname);\n+\tif (start == BSTR_ERR)\n+\t\treturn NULL;\n+\t\n+\tstart += blength(&bname);\n+\tif (blength(line) > start && bchar(line, start) == '\"') {\n+\t\tstart++;\n+\t\tif ((end = bstrchrp(line, start, '\"')) == BSTR_ERR)\n+\t\t\tend = blength(line);\n+\t\treturn bmidstr(line, start, end - start);\n \t}\n-\telse\n-\t\tends = \"; \\t\";\n-\tsz = strcspn(ap, ends);\n-\tmemcpy(attr, ap, sz);\n-\tattr[sz] = 0;\n-\treturn 1;\n+\tif ((end = binchr(line, start, &endchars)) == BSTR_ERR)\n+\t\tend = blength(line);\n+\treturn bmidstr(line, start, end - start);\n }\n \n struct content_type {\n-\tchar *boundary;\n-\tint boundary_len;\n+\tbstring boundary;\n };\n \n static struct content_type content[MAX_BOUNDARIES];\n \n static struct content_type *content_top = content;\n \n-static int handle_content_type(char *line)\n+static int handle_content_type(const_bstring line)\n {\n-\tchar boundary[256];\n-\n-\tif (strcasestr(line, \"text/\") == NULL)\n+\tstatic struct tagbstring cmp_text = bsStatic(\"text/\");\n+\tbstring attr, boundary;\n+\t\n+\tif (binstrcaseless(line, 0, &cmp_text) == BSTR_ERR)\n \t\t message_type = TYPE_OTHER;\n-\tif (slurp_attr(line, \"boundary=\", boundary + 2)) {\n-\t\tmemcpy(boundary, \"--\", 2);\n+\n+\tif ((attr = slurp_attr(line, \"boundary=\"))) {\n+\t\tboundary = bfromcstr(\"--\");\n+\t\tbconcat(boundary, attr);\n+\t\tbdestroy(attr);\n \t\tif (content_top++ >= &content[MAX_BOUNDARIES]) {\n \t\t\tfprintf(stderr, \"Too many boundaries to handle\\n\");\n \t\t\texit(1);\n \t\t}\n-\t\tcontent_top->boundary_len = strlen(boundary);\n-\t\tcontent_top->boundary = xmalloc(content_top->boundary_len+1);\n-\t\tstrcpy(content_top->boundary, boundary);\n+\t\tcontent_top->boundary = boundary;\n+\t\treturn 0;\n \t}\n-\tif (slurp_attr(line, \"charset=\", charset)) {\n-\t\tint i, c;\n-\t\tfor (i = 0; (c = charset[i]) != 0; i++)\n-\t\t\tcharset[i] = tolower(c);\n+\tif ((attr = slurp_attr(line, \"charset=\"))) {\n+\t\tif (btolower(attr) == BSTR_ERR)\n+\t\t\tdie(\"Couldn't convert %s to lowercase.\\n\", attr->data);\n+\t\tcharset = attr;\n \t}\n \treturn 0;\n }\n \n-static int handle_content_transfer_encoding(char *line)\n+static int handle_content_transfer_encoding(const_bstring line)\n {\n-\tif (strcasestr(line, \"base64\"))\n+\tstatic struct tagbstring cmp_base64 = bsStatic(\"base64\");\n+\tstatic struct tagbstring cmp_qp = bsStatic(\"quoted-printable\");\n+\n+\tif (binstrcaseless(line, 0, &cmp_base64) != BSTR_ERR)\n \t\ttransfer_encoding = TE_BASE64;\n-\telse if (strcasestr(line, \"quoted-printable\"))\n+\telse if (binstrcaseless(line, 0, &cmp_qp))\n \t\ttransfer_encoding = TE_QP;\n \telse\n \t\ttransfer_encoding = TE_DONTCARE;\n \treturn 0;\n }\n \n-static int is_multipart_boundary(const char *line)\n-{\n-\treturn (!memcmp(line, content_top->boundary, content_top->boundary_len));\n-}\n-\n-static int eatspace(char *line)\n+static int is_multipart_boundary(const_bstring line)\n {\n-\tint len = strlen(line);\n-\twhile (len > 0 && isspace(line[len-1]))\n-\t\tline[--len] = 0;\n-\treturn len;\n+\treturn !bstrncmp(line, content_top->boundary, blength(content_top->boundary));\n }\n \n-static char *cleanup_subject(char *subject)\n+/*\n+ * Removes (Re:|[ \\[\\t:]|\\[.*\\])* if prefixed and\n+ * trims trailing whitespace.\n+ */\n+static void cleanup_subject(bstring subject)\n {\n-\tfor (;;) {\n-\t\tchar *p;\n-\t\tint len, remove;\n-\t\tswitch (*subject) {\n+\tint pos;\n+\twhile (blength(subject)) {\n+\t\tswitch (bchar(subject, 0)) {\n \t\tcase 'r': case 'R':\n-\t\t\tif (!memcmp(\"e:\", subject+1, 2)) {\n-\t\t\t\tsubject += 3;\n+\t\t\tif (blength(subject) <= 3)\n+\t\t\t\tbreak;\n+\t\t\tif (!memcmp(bdata(subject) + 1, \"e:\", 2)) {\n+\t\t\t\tbdelete(subject, 0, 3);\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tbreak;\n-\t\tcase ' ': case '\\t': case ':':\n-\t\t\tsubject++;\n-\t\t\tcontinue;\n-\n \t\tcase '[':\n-\t\t\tp = strchr(subject, ']');\n-\t\t\tif (!p) {\n-\t\t\t\tsubject++;\n-\t\t\t\tcontinue;\n-\t\t\t}\n-\t\t\tlen = strlen(p);\n-\t\t\tremove = p - subject;\n-\t\t\tif (remove <= len *2) {\n-\t\t\t\tsubject = p+1;\n-\t\t\t\tcontinue;\n+\t\t\tif ((pos = bstrchr(subject, ']')) != BSTR_ERR) {\n+\t\t\t\t/* Don't remove more than a third of the subject. */\n+\t\t\t\tif (pos <= blength(subject)/3) {\n+\t\t\t\t\tbdelete(subject, 0, pos + 1);\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n+\t\t\t\tbreak;\n \t\t\t}\n-\t\t\tbreak;\n+\t\t/* fall through */\n+\t\tcase ' ': case '\\t': case ':':\n+\t\t\tbdelete(subject, 0, 1);\n+\t\t\tcontinue;\n \t\t}\n-\t\teatspace(subject);\n-\t\treturn subject;\n+\n+\t\tbtrimws(subject);\n+\t\treturn;\n \t}\n }\n \n-static void cleanup_space(char *buf)\n+static void cleanup_space(bstring buf)\n {\n-\tunsigned char c;\n-\twhile ((c = *buf) != 0) {\n-\t\tbuf++;\n-\t\tif (isspace(c)) {\n-\t\t\tbuf[-1] = ' ';\n-\t\t\tc = *buf;\n-\t\t\twhile (isspace(c)) {\n-\t\t\t\tint len = strlen(buf);\n-\t\t\t\tmemmove(buf, buf+1, len);\n-\t\t\t\tc = *buf;\n-\t\t\t}\n+\tstruct bstrList *tok;\n+\tstatic struct tagbstring whitespace = bsStatic(\" \\n\\t\\r\\f\\v\");\n+\tint i;\n+\n+\ttok = bsplitstr(buf, &whitespace);\n+\tbtrunc(buf, 0);\n+\tfor (i = 0; i < tok->qty; i++) {\n+\t\tif (blength(tok->entry[i])) {\n+\t\t\tbconcat(buf, tok->entry[i]);\n+\t\t\tbconchar(buf, ' ');\n \t\t}\n \t}\n+\t/* Remove the last ' ' */\n+\tbtrunc(buf, blength(buf) - 1);\n+\tbstrListDestroy(tok);\n+}\n+\n+static int handle_header(bstring line, bstring *data, int ofs)\n+{\n+\tif (!line)\n+\t\treturn 1;\n+\n+\tbdestroy(*data);\n+\t*data = bmidstr(line, ofs, blength(line) - ofs);\n+\n+\treturn 0;\n }\n \n-static void decode_header(char *it, unsigned itsize);\n+static void decode_header(bstring line);\n static char *header[MAX_HDR_PARSED] = {\n \t\"From\",\"Subject\",\"Date\",\n };\n \n-static int check_header(char *line, unsigned linesize, char **hdr_data, int overwrite)\n+static int check_header(bstring line, bstring hdr_data[], int overwrite)\n {\n \tint i;\n \n \t/* search for the interesting parts */\n \tfor (i = 0; header[i]; i++) {\n \t\tint len = strlen(header[i]);\n+\n \t\tif ((!hdr_data[i] || overwrite) &&\n-\t\t    !strncasecmp(line, header[i], len) &&\n-\t\t    line[len] == ':' && isspace(line[len + 1])) {\n+\t\t    bisstemeqcaselessblk(line, header[i], len) &&\n+\t\t    bchar(line, len) == ':' && isspace(bchar(line, len + 1))) {\n \t\t\t/* Unwrap inline B and Q encoding, and optionally\n \t\t\t * normalize the meta information to utf8.\n \t\t\t */\n-\t\t\tdecode_header(line + len + 2, linesize - len - 2);\n-\t\t\thdr_data[i] = xmalloc(1000 * sizeof(char));\n-\t\t\tif (! handle_header(line, hdr_data[i], len + 2)) {\n+\t\t\tdecode_header(line);\n+\t\t\tif (!handle_header(line, &hdr_data[i], len + 2)) {\n \t\t\t\treturn 1;\n \t\t\t}\n \t\t}\n \t}\n \n \t/* Content stuff */\n-\tif (!strncasecmp(line, \"Content-Type\", 12) &&\n-\t\tline[12] == ':' && isspace(line[12 + 1])) {\n-\t\tdecode_header(line + 12 + 2, linesize - 12 - 2);\n+\tif (!bisstemeqcaselessblk(line, bsStaticBlkParms(\"Content-Type\")) &&\n+\t\tbchar(line, 12) == ':' && isspace(bchar(line, 12 + 1))) {\n+\t\tdecode_header(line);\n \t\tif (! handle_content_type(line)) {\n \t\t\treturn 1;\n \t\t}\n \t}\n-\tif (!strncasecmp(line, \"Content-Transfer-Encoding\", 25) &&\n-\t\tline[25] == ':' && isspace(line[25 + 1])) {\n-\t\tdecode_header(line + 25 + 2, linesize - 25 - 2);\n+\tif (!bisstemeqcaselessblk(line, bsStaticBlkParms(\"Content-Transfer-Encoding\")) &&\n+\t\tbchar(line, 25) == ':' && isspace(bchar(line, 25 + 1))) {\n+\t\tdecode_header(line);\n \t\tif (! handle_content_transfer_encoding(line)) {\n \t\t\treturn 1;\n \t\t}\n \t}\n \n \t/* for inbody stuff */\n-\tif (!memcmp(\">From\", line, 5) && isspace(line[5]))\n+\tif (bisstemeqblk(line, bsStaticBlkParms(\">From\")) && isspace(bchar(line, 5)))\n \t\treturn 1;\n-\tif (!memcmp(\"[PATCH]\", line, 7) && isspace(line[7])) {\n+\tif (bisstemeqblk(line, bsStaticBlkParms(\"[PATCH]\")) && isspace(bchar(line, 7))) {\n \t\tfor (i = 0; header[i]; i++) {\n-\t\t\tif (!memcmp(\"Subject: \", header[i], 9)) {\n-\t\t\t\tif (! handle_header(line, hdr_data[i], 0)) {\n+\t\t\tif (!memcmp(\"Subject\", header[i], 7)) {\n+\t\t\t\tif (!handle_header(line, &hdr_data[i], 0)) {\n \t\t\t\t\treturn 1;\n \t\t\t\t}\n \t\t\t}\n@@ -346,7 +316,7 @@ static int check_header(char *line, unsigned linesize, char **hdr_data, int over\n \treturn 0;\n }\n \n-static int is_rfc2822_header(char *line)\n+static int is_rfc2822_header(const_bstring line)\n {\n \t/*\n \t * The section that defines the loosest possible\n@@ -357,15 +327,15 @@ static int is_rfc2822_header(char *line)\n \t * ftext = %d33-57 / %59-126\n \t */\n \tint ch;\n-\tchar *cp = line;\n+\tchar *cp = bdata(line);\n \n \t/* Count mbox From headers as headers */\n-\tif (!memcmp(line, \"From \", 5) || !memcmp(line, \">From \", 6))\n+\tif (blength(line) >= 6 && (!memcmp(cp, \"From \", 5) || !memcmp(cp, \">From \", 6)))\n \t\treturn 1;\n \n \twhile ((ch = *cp++)) {\n \t\tif (ch == ':')\n-\t\t\treturn cp != line;\n+\t\t\treturn cp != bdata(line);\n \t\tif ((33 <= ch && ch <= 57) ||\n \t\t    (59 <= ch && ch <= 126))\n \t\t\tcontinue;\n@@ -375,34 +345,23 @@ static int is_rfc2822_header(char *line)\n }\n \n /*\n- * sz is size of 'line' buffer in bytes.  Must be reasonably\n- * long enough to hold one physical real-world e-mail line.\n+ * 'line' must be a valid bstring\n  */\n-static int read_one_header_line(char *line, int sz, FILE *in)\n+static int read_one_header_line(struct bStream *in, bstring line)\n {\n-\tint len;\n-\n-\t/*\n-\t * We will read at most (sz-1) bytes and then potentially\n-\t * re-add NUL after it.  Accessing line[sz] after this is safe\n-\t * and we can allow len to grow up to and including sz.\n-\t */\n-\tsz--;\n-\n \t/* Get the first part of the line. */\n-\tif (!fgets(line, sz, in))\n-\t\treturn 0;\n+\tif (bsreadln(line, in, '\\n') != BSTR_OK)\n+\t\tgoto unread_line;\n \n \t/*\n \t * Is it an empty line or not a valid rfc2822 header?\n \t * If so, stop here, and return false (\"not a header\")\n \t */\n-\tlen = eatspace(line);\n-\tif (!len || !is_rfc2822_header(line)) {\n+\tbrtrimws(line);\n+\tif (!blength(line) || !is_rfc2822_header(line)) {\n \t\t/* Re-add the newline */\n-\t\tline[len] = '\\n';\n-\t\tline[len + 1] = '\\0';\n-\t\treturn 0;\n+\t\tbconchar(line, '\\n');\n+\t\tgoto unread_line;\n \t}\n \n \t/*\n@@ -410,63 +369,57 @@ static int read_one_header_line(char *line, int sz, FILE *in)\n \t * Yuck, 2822 header \"folding\"\n \t */\n \tfor (;;) {\n-\t\tint peek, addlen;\n-\t\tstatic char continuation[1000];\n+\t\tbstring continuation;\n+\t\tcontinuation = bfromcstr(\"\");\n \n-\t\tpeek = fgetc(in); ungetc(peek, in);\n-\t\tif (peek != ' ' && peek != '\\t')\n+\t\tif (bsreadln(continuation, in, '\\n') != BSTR_OK)\n \t\t\tbreak;\n-\t\tif (!fgets(continuation, sizeof(continuation), in))\n+\t\tif (bchar(continuation, 0) != ' ' && bchar(continuation, 0) != '\\t') {\n+\t\t\tbsunread(in, continuation);\n \t\t\tbreak;\n-\t\taddlen = eatspace(continuation);\n-\t\tif (len < sz - 1) {\n-\t\t\tif (addlen >= sz - len)\n-\t\t\t\taddlen = sz - len - 1;\n-\t\t\tmemcpy(line + len, continuation, addlen);\n-\t\t\tline[len] = '\\n';\n-\t\t\tlen += addlen;\n \t\t}\n+\n+\t\tcontinuation->data[0] = '\\n';\n+\t\tbrtrimws(continuation);\n+\t\tbconcat(line, continuation);\n \t}\n-\tline[len] = 0;\n \n \treturn 1;\n+unread_line:\n+\tbsunread(in, line);\n+\treturn 0;\n }\n \n-static int decode_q_segment(char *in, char *ot, unsigned otsize, char *ep, int rfc2047)\n+static bstring decode_q_segment(bstring line, int rfc2047)\n {\n-\tchar *otend = ot + otsize;\n+\tchar *in = bdata(line);\n \tint c;\n-\twhile ((c = *in++) != 0 && (in <= ep)) {\n-\t\tif (ot == otend) {\n-\t\t\t*--ot = '\\0';\n-\t\t\treturn -1;\n-\t\t}\n+\tbstring out = bfromcstralloc(blength(line), \"\");\n+\n+\twhile ((c = *in++) != 0) {\n \t\tif (c == '=') {\n \t\t\tint d = *in++;\n \t\t\tif (d == '\\n' || !d)\n \t\t\t\tbreak; /* drop trailing newline */\n-\t\t\t*ot++ = ((hexval(d) << 4) | hexval(*in++));\n+\t\t\tbconchar(out, (hexval(d) << 4) | hexval(*in++));\n \t\t\tcontinue;\n \t\t}\n \t\tif (rfc2047 && c == '_') /* rfc2047 4.2 (2) */\n \t\t\tc = 0x20;\n-\t\t*ot++ = c;\n+\t\tbconchar(out, c);\n \t}\n-\t*ot = 0;\n-\treturn 0;\n+\treturn out;\n }\n \n-static int decode_b_segment(char *in, char *ot, unsigned otsize, char *ep)\n+static bstring decode_b_segment(bstring line)\n {\n \t/* Decode in..ep, possibly in-place to ot */\n \tint c, pos = 0, acc = 0;\n-\tchar *otend = ot + otsize;\n+\tchar *in = bdata(line);\n+\tbstring out;\n \n-\twhile ((c = *in++) != 0 && (in <= ep)) {\n-\t\tif (ot == otend) {\n-\t\t\t*--ot = '\\0';\n-\t\t\treturn -1;\n-\t\t}\n+\tout = bfromcstralloc(blength(line), \"\");\n+\twhile ((c = *in++) != 0) {\n \t\tif (c == '+')\n \t\t\tc = 62;\n \t\telse if (c == '/')\n@@ -491,21 +444,20 @@ static int decode_b_segment(char *in, char *ot, unsigned otsize, char *ep)\n \t\t\tacc = (c << 2);\n \t\t\tbreak;\n \t\tcase 1:\n-\t\t\t*ot++ = (acc | (c >> 4));\n+\t\t\tbconchar(out, (acc | (c >> 4)));\n \t\t\tacc = (c & 15) << 4;\n \t\t\tbreak;\n \t\tcase 2:\n-\t\t\t*ot++ = (acc | (c >> 2));\n+\t\t\tbconchar(out, (acc | (c >> 2)));\n \t\t\tacc = (c & 3) << 6;\n \t\t\tbreak;\n \t\tcase 3:\n-\t\t\t*ot++ = (acc | c);\n+\t\t\tbconchar(out, (acc | c));\n \t\t\tacc = pos = 0;\n \t\t\tbreak;\n \t\t}\n \t}\n-\t*ot = 0;\n-\treturn 0;\n+\treturn out;\n }\n \n /*\n@@ -519,147 +471,174 @@ static int decode_b_segment(char *in, char *ot, unsigned otsize, char *ep)\n  * Otherwise, we default to assuming it is Latin1 for historical\n  * reasons.\n  */\n-static const char *guess_charset(const char *line, const char *target_charset)\n+static bstring guess_charset(bstring line, bstring target_charset)\n {\n-\tif (is_encoding_utf8(target_charset)) {\n-\t\tif (is_utf8(line))\n+\t//FIXME: convert utf8.c to bstring\n+\tif (is_encoding_utf8(bdata(target_charset))) {\n+\t\tif (is_utf8(bdata(line)))\n \t\t\treturn NULL;\n \t}\n-\treturn \"latin1\";\n+\treturn bfromcstr(\"latin1\");\n }\n \n-static void convert_to_utf8(char *line, unsigned linesize, const char *charset)\n+static void convert_to_utf8(bstring line, bstring charset)\n {\n-\tchar *out;\n+\tbstring out;\n+\tchar *cout;\n \n-\tif (!charset || !*charset) {\n-\t\tcharset = guess_charset(line, metainfo_charset);\n-\t\tif (!charset)\n+\tif (blength(charset) == 0) {\n+\t\tout = guess_charset(line, metainfo_charset);\n+\t\tif (!out) {\n+\t\t\tbdestroy(charset);\n+\t\t\tcharset = NULL;\n \t\t\treturn;\n+\t\t}\n+\t\tbassign(charset, out);\n+\t\tbdestroy(out);\n \t}\n \n-\tif (!strcmp(metainfo_charset, charset))\n+\tif (!bstrcmp(metainfo_charset, charset))\n \t\treturn;\n-\tout = reencode_string(line, metainfo_charset, charset);\n-\tif (!out)\n+\t//FIXME: convert utf8.c to use bstring\n+\tcout = reencode_string(bdata(line), bdata(metainfo_charset), bdata(charset));\n+\tif (!cout)\n \t\tdie(\"cannot convert from %s to %s\\n\",\n-\t\t    charset, metainfo_charset);\n-\tstrlcpy(line, out, linesize);\n-\tfree(out);\n+\t\t    bdata(charset), bdata(metainfo_charset));\n+\tbassigncstr(line, cout);\n }\n \n-static int decode_header_bq(char *it, unsigned itsize)\n+static int decode_header_bq(bstring line)\n {\n-\tchar *in, *out, *ep, *cp, *sp;\n-\tchar outbuf[1000];\n+\tbstring out;\n+\tbstring decoded = NULL, charset_q = NULL, tmp;\n \tint rfc2047 = 0;\n+\tint in = 0, cp, ep;\n \n-\tin = it;\n-\tout = outbuf;\n-\twhile ((ep = strstr(in, \"=?\")) != NULL) {\n-\t\tint sz, encoding;\n-\t\tchar charset_q[256], piecebuf[256];\n+\tstruct tagbstring cmp_eq_qst = bsStatic(\"=?\");\n+\tstruct tagbstring cmp_qst_eq = bsStatic(\"?=\");\n+\n+\tout = bfromcstralloc(blength(line), \"\");\n+\n+\twhile ((ep = binstr(line, 0, &cmp_eq_qst)) != BSTR_ERR) {\n+\t\tint encoding;\n \t\trfc2047 = 1;\n \n-\t\tif (in != ep) {\n-\t\t\tsz = ep - in;\n-\t\t\tmemcpy(out, in, sz);\n-\t\t\tout += sz;\n-\t\t\tin += sz;\n-\t\t}\n+\t\tbcatblk(out, bdataofs(line, in), ep - in);\n+\t\tin += ep - in + 2;\n \t\t/* E.g.\n \t\t * ep : \"=?iso-2022-jp?B?GyR...?= foo\"\n \t\t * ep : \"=?ISO-8859-1?Q?Foo=FCbar?= baz\"\n \t\t */\n-\t\tep += 2;\n-\t\tcp = strchr(ep, '?');\n-\t\tif (!cp)\n-\t\t\treturn rfc2047; /* no munging */\n-\t\tfor (sp = ep; sp < cp; sp++)\n-\t\t\tcharset_q[sp - ep] = tolower(*sp);\n-\t\tcharset_q[cp - ep] = 0;\n-\t\tencoding = cp[1];\n-\t\tif (!encoding || cp[2] != '?')\n-\t\t\treturn rfc2047; /* no munging */\n-\t\tep = strstr(cp + 3, \"?=\");\n-\t\tif (!ep)\n-\t\t\treturn rfc2047; /* no munging */\n+\n+\t\tcp = bstrchrp(line, in, '?');\n+\t\tif (cp == BSTR_ERR)\n+\t\t\tgoto out0; /* no munging */\n+\n+\t\tcharset_q = bmidstr(line, in, cp - in);\n+\t\tif (charset_q->slen)\n+\t\t\tbtolower(charset_q);\n+\n+\t\tif (line->slen < cp + 2)\n+\t\t\tgoto out1;\n+\t\t\t//die(\"Bad header: %s,\", line->data);\n+\n+\t\tencoding = bchar(line, cp + 1);\n+\t\tif (!encoding || bchar(line, cp + 2) != '?')\n+\t\t\tgoto out1; /* no munging */\n+\t\tep = binstr(line, cp + 3, &cmp_qst_eq);\n+\t\tif (ep == BSTR_ERR)\n+\t\t\tgoto out1; /* no munging */\n \t\tswitch (tolower(encoding)) {\n \t\tdefault:\n-\t\t\treturn rfc2047; /* no munging */\n+\t\t\tgoto out1; /* no munging */\n \t\tcase 'b':\n-\t\t\tsz = decode_b_segment(cp + 3, piecebuf, sizeof(piecebuf), ep);\n+\t\t\t//FIXME: use bmid2tbstr ?\n+\t\t\t// Needs to change the decode function to not look for null\n+\t\t\ttmp = bmidstr(line, cp + 3, ep - cp -3);\n+\t\t\tdecoded = decode_b_segment(tmp);\n \t\t\tbreak;\n \t\tcase 'q':\n-\t\t\tsz = decode_q_segment(cp + 3, piecebuf, sizeof(piecebuf), ep, 1);\n+\t\t\ttmp = bmidstr(line, cp + 3, ep - cp -3);\n+\t\t\tdecoded = decode_q_segment(tmp, 1);\n \t\t\tbreak;\n \t\t}\n-\t\tif (sz < 0)\n-\t\t\treturn rfc2047;\n+\t\tbdestroy(tmp);\n+\t\tif (decoded == NULL)\n+\t\t\tgoto out1;\n \t\tif (metainfo_charset)\n-\t\t\tconvert_to_utf8(piecebuf, sizeof(piecebuf), charset_q);\n+\t\t\tconvert_to_utf8(decoded, charset_q);\n \n-\t\tsz = strlen(piecebuf);\n-\t\tif (outbuf + sizeof(outbuf) <= out + sz)\n-\t\t\treturn rfc2047; /* no munging */\n-\t\tstrcpy(out, piecebuf);\n-\t\tout += sz;\n+\t\tbconcat(out, decoded);\n \t\tin = ep + 2;\n+\n+\t\tbdestroy(decoded);\n+\t\tbdestroy(charset_q);\n \t}\n-\tstrcpy(out, in);\n-\tstrlcpy(it, outbuf, itsize);\n+\t/* Add the remainder of the line. */\n+\tbcatblk(out, bdataofs(line, in), blength(line) - in);\n+\n+\tbassign(line, out);\n+\n+\tbdestroy(decoded);\n+out1:\n+\tbdestroy(charset_q);\n+out0:\n+\tbdestroy(out);\n \treturn rfc2047;\n }\n \n-static void decode_header(char *it, unsigned itsize)\n+static void decode_header(bstring line)\n {\n-\n-\tif (decode_header_bq(it, itsize))\n+\tif (decode_header_bq(line))\n \t\treturn;\n \t/* otherwise \"it\" is a straight copy of the input.\n \t * This can be binary guck but there is no charset specified.\n \t */\n \tif (metainfo_charset)\n-\t\tconvert_to_utf8(it, itsize, \"\");\n+\t\tconvert_to_utf8(line, NULL);\n }\n \n-static void decode_transfer_encoding(char *line, unsigned linesize)\n+static void decode_transfer_encoding(bstring line)\n {\n-\tchar *ep;\n+\tbstring ret = NULL;\n \n \tswitch (transfer_encoding) {\n \tcase TE_QP:\n-\t\tep = line + strlen(line);\n-\t\tdecode_q_segment(line, line, linesize, ep, 0);\n+\t\tret = decode_q_segment(line, 0);\n \t\tbreak;\n \tcase TE_BASE64:\n-\t\tep = line + strlen(line);\n-\t\tdecode_b_segment(line, line, linesize, ep);\n+\t\tret = decode_b_segment(line);\n \t\tbreak;\n \tcase TE_DONTCARE:\n \t\tbreak;\n \t}\n+\tif (ret)\n+\t\tbassign(line, ret);\n+\tbdestroy(ret);\n }\n \n-static int handle_filter(char *line, unsigned linesize);\n+static int handle_filter(bstring line);\n \n-static int find_boundary(void)\n+static int find_boundary(struct bStream *in, bstring line)\n {\n-\twhile(fgets(line, sizeof(line), fin) != NULL) {\n+\twhile(bsreadln(line, in, '\\n') != BSTR_ERR) {\n \t\tif (is_multipart_boundary(line))\n \t\t\treturn 1;\n \t}\n \treturn 0;\n }\n \n-static int handle_boundary(void)\n+static int handle_boundary(struct bStream *in, bstring line)\n {\n-\tchar newline[]=\"\\n\";\n+\tstruct tagbstring newline = bsStatic(\"\\n\");\n+\tchar *c;\n again:\n-\tif (!memcmp(line+content_top->boundary_len, \"--\", 2)) {\n+\tif (blength(line) >= blength(content_top->boundary) + 2 &&\n+\t    (c = bdataofs(line, blength(content_top->boundary))) &&\n+\t    !memcmp(c, \"--\", 2)) {\n \t\t/* we hit an end boundary */\n \t\t/* pop the current boundary off the stack */\n-\t\tfree(content_top->boundary);\n+\t\tbdestroy(content_top->boundary);\n \n \t\t/* technically won't happen as is_multipart_boundary()\n \t\t   will fail first.  But just in case..\n@@ -669,49 +648,52 @@ again:\n \t\t\t\t\t\"can't recover\\n\");\n \t\t\texit(1);\n \t\t}\n-\t\thandle_filter(newline, sizeof(newline));\n+\t\thandle_filter(&newline);\n \n \t\t/* skip to the next boundary */\n-\t\tif (!find_boundary())\n+\t\tif (!find_boundary(in, line)) {\n+\t\t\tbsunread(in, line);\n \t\t\treturn 0;\n+\t\t}\n \t\tgoto again;\n \t}\n \n \t/* set some defaults */\n \ttransfer_encoding = TE_DONTCARE;\n-\tcharset[0] = 0;\n+\tbassigncstr(charset, \"\");\n \tmessage_type = TYPE_TEXT;\n \n \t/* slurp in this section's info */\n-\twhile (read_one_header_line(line, sizeof(line), fin))\n-\t\tcheck_header(line, sizeof(line), p_hdr_data, 0);\n+\twhile (read_one_header_line(in, line))\n+\t\tcheck_header(line, p_hdr_data, 0);\n \n \t/* eat the blank line after section info */\n-\treturn (fgets(line, sizeof(line), fin) != NULL);\n+\treturn (bsreadln(line, in, '\\n') != BSTR_ERR);\n }\n \n-static inline int patchbreak(const char *line)\n+static inline int patchbreak(const_bstring line)\n {\n+\tint i;\n+\n \t/* Beginning of a \"diff -\" header? */\n-\tif (!memcmp(\"diff -\", line, 6))\n+\tif (!bisstemeqblk(line, bsStaticBlkParms(\"diff -\")))\n \t\treturn 1;\n \n \t/* CVS \"Index: \" line? */\n-\tif (!memcmp(\"Index: \", line, 7))\n+\tif (!bisstemeqblk(line, bsStaticBlkParms(\"Index: \")))\n \t\treturn 1;\n \n \t/*\n \t * \"--- <filename>\" starts patches without headers\n \t * \"---<sp>*\" is a manual separator\n \t */\n-\tif (!memcmp(\"---\", line, 3)) {\n-\t\tline += 3;\n+\tif (!bisstemeqblk(line, bsStaticBlkParms(\"---\"))) {\n \t\t/* space followed by a filename? */\n-\t\tif (line[0] == ' ' && !isspace(line[1]))\n+\t\tif (bchar(line, 3) == ' ' && !isspace(bchar(line, 4)))\n \t\t\treturn 1;\n \t\t/* Just whitespace? */\n-\t\tfor (;;) {\n-\t\t\tunsigned char c = *line++;\n+\t\tfor (i = 3; i < blength(line); i++) {\n+\t\t\tunsigned char c = bchar(line, i);\n \t\t\tif (c == '\\n')\n \t\t\t\treturn 1;\n \t\t\tif (!isspace(c))\n@@ -723,31 +705,25 @@ static inline int patchbreak(const char *line)\n }\n \n \n-static int handle_commit_msg(char *line, unsigned linesize)\n+static int handle_commit_msg(bstring line)\n {\n \tstatic int still_looking = 1;\n-\tchar *endline = line + linesize;\n+\tchar *c;\n \n \tif (!cmitmsg)\n \t\treturn 0;\n \n \tif (still_looking) {\n-\t\tchar *cp = line;\n-\t\tif (isspace(*line)) {\n-\t\t\tfor (cp = line + 1; *cp; cp++) {\n-\t\t\t\tif (!isspace(*cp))\n-\t\t\t\t\tbreak;\n-\t\t\t}\n-\t\t\tif (!*cp)\n-\t\t\t\treturn 0;\n-\t\t}\n-\t\tif ((still_looking = check_header(cp, endline - cp, s_hdr_data, 0)) != 0)\n+\t\tbrtrimws(line);\n+\t\tif (blength(line) == 0)\n+\t\t\treturn 0;\n+\t\tif ((still_looking = check_header(line, s_hdr_data, 0)) != 0)\n \t\t\treturn 0;\n \t}\n \n \t/* normalize the log message to UTF-8. */\n \tif (metainfo_charset)\n-\t\tconvert_to_utf8(line, endline - line, charset);\n+\t\tconvert_to_utf8(line, charset);\n \n \tif (patchbreak(line)) {\n \t\tfclose(cmitmsg);\n@@ -755,18 +731,24 @@ static int handle_commit_msg(char *line, unsigned linesize)\n \t\treturn 1;\n \t}\n \n-\tfputs(line, cmitmsg);\n+\tif ((c = bdata(line)) == NULL)\n+\t\tdie(\"Programming error: line had no data\\n\");\n+\tfputs(c, cmitmsg);\n \treturn 0;\n }\n \n-static int handle_patch(char *line)\n+static int handle_patch(const_bstring line)\n {\n-\tfputs(line, patchfile);\n+\tchar *c;\n+\n+\tif ((c = bdata(line)) == NULL)\n+\t\tdie(\"Programming error: patch line had no data\\n\");\n+\tfputs(c, patchfile);\n \tpatch_lines++;\n \treturn 0;\n }\n \n-static int handle_filter(char *line, unsigned linesize)\n+static int handle_filter(bstring line)\n {\n \tstatic int filter = 0;\n \n@@ -775,7 +757,7 @@ static int handle_filter(char *line, unsigned linesize)\n \t */\n \tswitch (filter) {\n \tcase 0:\n-\t\tif (!handle_commit_msg(line, linesize))\n+\t\tif (!handle_commit_msg(line))\n \t\t\tbreak;\n \t\tfilter++;\n \tcase 1:\n@@ -789,16 +771,19 @@ static int handle_filter(char *line, unsigned linesize)\n \treturn 0;\n }\n \n-static void handle_body(void)\n+static void handle_body(struct bStream *in)\n {\n-\tint rc = 0;\n-\tstatic char newline[2000];\n-\tstatic char *np = newline;\n+\t//FIXME: bdestroy line. unread line in more places?\n+\tbstring line;\n+\tint rc = 0, i, end;\n \n+\tline = bfromcstr(\"\");\n \t/* Skip up to the first boundary */\n-\tif (content_top->boundary) {\n-\t\tif (!find_boundary())\n+\tif (content_top->boundary) {//FIXME: ?\n+\t\tif (!find_boundary(in, line)) {\n+\t\t\tbsunread(in, line);\n \t\t\treturn;\n+\t\t}\n \t}\n \n \tdo {\n@@ -806,24 +791,24 @@ static void handle_body(void)\n \t\tif (content_top->boundary && is_multipart_boundary(line)) {\n \t\t\t/* flush any leftover */\n \t\t\tif ((transfer_encoding == TE_BASE64)  &&\n-\t\t\t    (np != newline)) {\n-\t\t\t\thandle_filter(newline, sizeof(newline));\n+\t\t\t    (blength(line))) {\n+\t\t\t\thandle_filter(line);\n \t\t\t}\n-\t\t\tif (!handle_boundary())\n+\t\t\tif (!handle_boundary(in, line))\n \t\t\t\treturn;\n \t\t}\n \n \t\t/* Unwrap transfer encoding */\n-\t\tdecode_transfer_encoding(line, sizeof(line));\n+\t\tdecode_transfer_encoding(line);\n \n \t\tswitch (transfer_encoding) {\n \t\tcase TE_BASE64:\n \t\t{\n-\t\t\tchar *op = line;\n+\t\t\tstruct bstrList *lines;\n \n \t\t\t/* binary data most likely doesn't have newlines */\n \t\t\tif (message_type != TYPE_TEXT) {\n-\t\t\t\trc = handle_filter(line, sizeof(newline));\n+\t\t\t\trc = handle_filter(line);\n \t\t\t\tbreak;\n \t\t\t}\n \n@@ -832,54 +817,55 @@ static void handle_body(void)\n \t\t\t * at a time to handle_filter()\n \t\t\t */\n \n-\t\t\tdo {\n-\t\t\t\twhile (*op != '\\n' && *op != 0)\n-\t\t\t\t\t*np++ = *op++;\n-\t\t\t\t*np = *op;\n-\t\t\t\tif (*np != 0) {\n-\t\t\t\t\t/* should be sitting on a new line */\n-\t\t\t\t\t*(++np) = 0;\n-\t\t\t\t\top++;\n-\t\t\t\t\trc = handle_filter(newline, sizeof(newline));\n-\t\t\t\t\tnp = newline;\n-\t\t\t\t}\n-\t\t\t} while (*op != 0);\n-\t\t\t/* the partial chunk is saved in newline and\n+\t\t\tlines = bsplit(line, '\\n');\n+\t\t\tend = lines->qty - 1;\n+\t\t\t/* the partial chunk is saved in line and\n \t\t\t * will be appended by the next iteration of fgets\n \t\t\t */\n+\t\t\tif (bchar(line, blength(line) - 1) != '\\n') {\n+\t\t\t\tbassign(line, lines->entry[end]);\n+\t\t\t\tend--;\n+\t\t\t} else\n+\t\t\t\tbtrunc(line, 0);\n+\t\t\tfor (i = 0; i <= end; i++)\n+\t\t\t\trc = handle_filter(lines->entry[i]);\n+\n+\t\t\tbstrListDestroy(lines);\n \t\t\tbreak;\n \t\t}\n \t\tdefault:\n-\t\t\trc = handle_filter(line, sizeof(newline));\n+\t\t\trc = handle_filter(line);\n+\t\t\tbtrunc(line, 0);\n \t\t}\n \t\tif (rc)\n \t\t\t/* nothing left to filter */\n \t\t\tbreak;\n-\t} while (fgets(line, sizeof(line), fin));\n+\t} while (bsreadlna(line, in, '\\n') != BSTR_ERR);\n \n \treturn;\n }\n \n-static void output_header_lines(FILE *fout, const char *hdr, char *data)\n+static void output_header_lines(FILE *fout, const char *hdr, const_bstring data)\n {\n+\tchar *sp;\n+\tsp = bdata(data);\n \twhile (1) {\n-\t\tchar *ep = strchr(data, '\\n');\n+\t\tchar *ep = strchr(sp, '\\n');\n \t\tint len;\n \t\tif (!ep)\n-\t\t\tlen = strlen(data);\n+\t\t\tlen = strlen(sp);\n \t\telse\n-\t\t\tlen = ep - data;\n-\t\tfprintf(fout, \"%s: %.*s\\n\", hdr, len, data);\n+\t\t\tlen = ep - sp;\n+\t\tfprintf(fout, \"%s: %.*s\\n\", hdr, len, sp);\n \t\tif (!ep)\n \t\t\tbreak;\n-\t\tdata = ep + 1;\n+\t\tsp = ep + 1;\n \t}\n }\n \n static void handle_info(void)\n {\n-\tchar *sub;\n-\tchar *hdr;\n+\tbstring hdr;\n \tint i;\n \n \tfor (i = 0; header[i]; i++) {\n@@ -893,32 +879,32 @@ static void handle_info(void)\n \t\t\tcontinue;\n \n \t\tif (!memcmp(header[i], \"Subject\", 7)) {\n-\t\t\tif (keep_subject)\n-\t\t\t\tsub = hdr;\n-\t\t\telse {\n-\t\t\t\tsub = cleanup_subject(hdr);\n-\t\t\t\tcleanup_space(sub);\n+\t\t\tif (!keep_subject) {\n+\t\t\t\tcleanup_subject(hdr);\n+\t\t\t\tcleanup_space(hdr);\n \t\t\t}\n-\t\t\toutput_header_lines(fout, \"Subject\", sub);\n+\t\t\toutput_header_lines(fout, \"Subject\", hdr);\n \t\t} else if (!memcmp(header[i], \"From\", 4)) {\n \t\t\thandle_from(hdr);\n-\t\t\tfprintf(fout, \"Author: %s\\n\", name);\n-\t\t\tfprintf(fout, \"Email: %s\\n\", email);\n+\t\t\tfprintf(fout, \"Author: %s\\n\", bdata(name));\n+\t\t\tfprintf(fout, \"Email: %s\\n\", bdata(email));\n \t\t} else {\n \t\t\tcleanup_space(hdr);\n-\t\t\tfprintf(fout, \"%s: %s\\n\", header[i], hdr);\n+\t\t\tfprintf(fout, \"%s: %s\\n\", header[i], bdata(hdr));\n \t\t}\n \t}\n \tfprintf(fout, \"\\n\");\n }\n \n-static int mailinfo(FILE *in, FILE *out, int ks, const char *encoding,\n+static int mailinfo(FILE *in, FILE *out, int ks, const_bstring encoding,\n \t\t    const char *msg, const char *patch)\n {\n \tkeep_subject = ks;\n-\tmetainfo_charset = encoding;\n+\tmetainfo_charset = bstrcpy(encoding);\n \tfin = in;\n \tfout = out;\n+\tbstring line;\n+\tstruct bStream *in_stream = bsopen((bNread) fread, in);\n \n \tcmitmsg = fopen(msg, \"w\");\n \tif (!cmitmsg) {\n@@ -932,14 +918,17 @@ static int mailinfo(FILE *in, FILE *out, int ks, const char *encoding,\n \t\treturn -1;\n \t}\n \n-\tp_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(char *));\n-\ts_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(char *));\n+\tp_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(*p_hdr_data));\n+\ts_hdr_data = xcalloc(MAX_HDR_PARSED, sizeof(*s_hdr_data));\n \n \t/* process the email header */\n-\twhile (read_one_header_line(line, sizeof(line), fin))\n-\t\tcheck_header(line, sizeof(line), p_hdr_data, 1);\n+\tline = bfromcstr(\"\");\n+\twhile (read_one_header_line(in_stream, line))\n+\t\tcheck_header(line, p_hdr_data, 1);\n \n-\thandle_body();\n+\tbsunread(in_stream, line);\n+\t\n+\thandle_body(in_stream);\n \thandle_info();\n \n \treturn 0;\n@@ -958,17 +947,17 @@ int cmd_mailinfo(int argc, const char **argv, const char *prefix)\n \tgit_config(git_default_config);\n \n \tdef_charset = (git_commit_encoding ? git_commit_encoding : \"utf-8\");\n-\tmetainfo_charset = def_charset;\n+\tmetainfo_charset = bfromcstr(def_charset);\n \n \twhile (1 < argc && argv[1][0] == '-') {\n \t\tif (!strcmp(argv[1], \"-k\"))\n \t\t\tkeep_subject = 1;\n \t\telse if (!strcmp(argv[1], \"-u\"))\n-\t\t\tmetainfo_charset = def_charset;\n+\t\t\tbassigncstr(metainfo_charset, def_charset);\n \t\telse if (!strcmp(argv[1], \"-n\"))\n \t\t\tmetainfo_charset = NULL;\n \t\telse if (!prefixcmp(argv[1], \"--encoding=\"))\n-\t\t\tmetainfo_charset = argv[1] + 11;\n+\t\t\tbassigncstr(metainfo_charset, argv[1] + 11);\n \t\telse\n \t\t\tusage(mailinfo_usage);\n \t\targc--; argv++;\n-- \n1.5.3.rc7\n"},{"id":"52522","messageId":"20070904213857.GA21351@steel.home","threadId":"9779","inReplyTo":"46DDC500.5000606@etek.chalmers.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-09-04T21:38:57Z","receivedAt":"2007-09-04T21:38:57Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"Lukas Sandström, Tue, Sep 04, 2007 22:50:08 +0200:\n> Hi.\n> \n> This is an attempt to use \"The Better String Library\"[1] in builtin-mailinfo.c\n> \n> The patch doesn't pass all the tests in the testsuit yet, but I thought I'd\n> send it out so people can decide if they like how the code looks.\n\nIt looks uglier, but what are measurable merits? Object code size,\nperfomance hit/improvement, valgrind logs?\n\n> -static int read_one_header_line(char *line, int sz, FILE *in)\n> +static int read_one_header_line(struct bStream *in, bstring line)\n\nEvery coder has a time in his life when he writes a string library...\nand a stream support for it.\n"},{"id":"52525","messageId":"20070904230117.GA12448@olympe.madism.org","threadId":"9779","inReplyTo":"20070904213857.GA21351@steel.home","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-09-04T23:01:17Z","receivedAt":"2007-09-04T23:01:17Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On mar, sep 04, 2007 at 09:38:57 +0000, Alex Riesen wrote:\n> Lukas Sandström, Tue, Sep 04, 2007 22:50:08 +0200:\n> > Hi.\n> > \n> > This is an attempt to use \"The Better String Library\"[1] in builtin-mailinfo.c\n> > \n> > The patch doesn't pass all the tests in the testsuit yet, but I thought I'd\n> > send it out so people can decide if they like how the code looks.\n> \n> It looks uglier, but what are measurable merits? Object code size,\n> perfomance hit/improvement, valgrind logs?\n\n  Well I honestly believe that putting strbufs/bstrings in mailinfo.c\nadds no value. I was going to give it a try to see how strbufs\nperformed, but it's just useless.\n\n  The main problem mailinfo has, it's according to Junio that it may\nsometimes truncate some things in buffers at 1000 octets, without dying\nloudly. That is bad.\n\n  _but_ there is no point in using arbitrary long string buffers to\nparse a mail. Remember, a mail goes through SMTP, and SMTP is supposed\nto limit its lines at 512 characters (without use of extensions at\nleast). Not to mention that an email address cannot be more than 64+256\nchars long (or sth around that). So using variable lengths buffers is\njust a waste.\n\n  string buffers are not really (IMHO) supposed to help in parsing\ntasks, and when you need to do some serious parsing, either do it by\nhand or use lex, but nothing in between makes sense to me.\n\n  OTOH, string buffers can be used in many places where git has (at\nleast 4 different to my current count, growing) many implementations of\nalways slightly different kind of buffers. I've some more patches\npending here than the one I already sent, and well, here is the\ndiffstat:\n\n$ git diff --stat origin/master.. ^strbuf*\n archive-tar.c         |   67 ++++++++++++------------------------------------\n builtin-apply.c       |   29 ++++++---------------\n builtin-blame.c       |   34 ++++++++-----------------\n builtin-commit-tree.c |   59 +++++++++---------------------------------\n builtin-rerere.c      |   53 +++++++++++---------------------------\n cache-tree.c          |   57 ++++++++++++++---------------------------\n diff.c                |   25 ++++++------------\n fast-import.c         |   38 +++++++++++----------------\n mktree.c              |   26 ++++++-------------\n 9 files changed, 116 insertions(+), 272 deletions(-)\n\n  I mean, there is not even a need to show the diff to understand what\nthe gain is. And that was possible, because strbufs are straightforward,\nand gives you the kind of controls git needs (tweaking how memory will\nbe allocated to avoid reallocs is part of the answer).\n\n\n  A French author once said: “Il semble que la perfection soit atteinte\nnon quand il n'y a plus rien à ajouter, mais quand il n'y a plus rien à\nretrancher.” -- Antoine de St Éxupéry[0]. IMHO git will never need any\nof the bstring splits, streaming functions, tokenization or whatever,\nand supporting those has necessarily led the bstring library to make\nsome choices that may not fit git needs. I don't really like reinventing\nthe wheel, but OTOH buffers and strings are often of the critical path,\nand having a nice fitting buffer API is priceless.\n\n\n  [0] Perfection is achieved, not when there is nothing more to add, but\n      when there is nothing left to take away.\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"52584","messageId":"1189004090.20311.12.camel@hinata.boston.redhat.com","threadId":"9779","inReplyTo":"46DDC500.5000606@etek.chalmers.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Kristian Høgsberg","fromEmail":"krh@redhat.com","sentAt":"2007-09-05T14:54:50Z","receivedAt":"2007-09-05T14:54:50Z","isPatch":false,"sender":{"key":"krh@redhat.com","avatar":"https://gravatar.com/avatar/763dee6f9594ac474f725b137a39565792928e583ddf59b32befc2907409027e?d=mp&s=160"},"body":"On Tue, 2007-09-04 at 22:50 +0200, Lukas Sandström wrote:\n> Hi.\n> \n> This is an attempt to use \"The Better String Library\"[1] in builtin-mailinfo.c\n> \n> The patch doesn't pass all the tests in the testsuit yet, but I thought I'd\n> send it out so people can decide if they like how the code looks.\n> \n> I'm not sending a patch to add the library files at this time. I'll send\n> that patch when this patch is working.\n> \n> The changes required to make it pass the tests shouldn't be very large.\n\nPlease, no.  Let's not pull in a dependency for something as simple as a\nstring library.  How many distros have bstring pcakaged?  \nThe right version?  Does it work on Windows?  We already have strbuf.c,\nlets just consolidate the string manipulation code already in git under\nthat interface.\n\nKristian\n"},{"id":"52586","messageId":"1189006030.20311.14.camel@hinata.boston.redhat.com","threadId":"9779","inReplyTo":"46DDC500.5000606@etek.chalmers.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Kristian Høgsberg","fromEmail":"krh@redhat.com","sentAt":"2007-09-05T15:27:10Z","receivedAt":"2007-09-05T15:27:10Z","isPatch":false,"sender":{"key":"krh@redhat.com","avatar":"https://gravatar.com/avatar/763dee6f9594ac474f725b137a39565792928e583ddf59b32befc2907409027e?d=mp&s=160"},"body":"On Tue, 2007-09-04 at 22:50 +0200, Lukas Sandström wrote:\n> Hi.\n> \n> This is an attempt to use \"The Better String Library\"[1] in builtin-mailinfo.c\n> \n> The patch doesn't pass all the tests in the testsuit yet, but I thought I'd\n> send it out so people can decide if they like how the code looks.\n> \n> I'm not sending a patch to add the library files at this time. I'll send\n> that patch when this patch is working.\n> \n> The changes required to make it pass the tests shouldn't be very large.\n\nPlease, no.  Let's not pull in a dependency for something as simple as a\nstring library.  How many distros have bstring pcakaged?  \nThe right version?  Does it work on Windows?  We already have strbuf.c,\nlets just consolidate the string manipulation code already in git under\nthat interface.\n\nKristian\n"},{"id":"52594","messageId":"vpq642pkoln.fsf@bauges.imag.fr","threadId":"9779","inReplyTo":"1189004090.20311.12.camel@hinata.boston.redhat.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-09-05T17:29:24Z","receivedAt":"2007-09-05T17:29:24Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Kristian Høgsberg <krh@redhat.com> writes:\n\n> On Tue, 2007-09-04 at 22:50 +0200, Lukas Sandström wrote:\n>> Hi.\n>> \n>> This is an attempt to use \"The Better String Library\"[1] in builtin-mailinfo.c\n>> \n>> The patch doesn't pass all the tests in the testsuit yet, but I thought I'd\n>> send it out so people can decide if they like how the code looks.\n>> \n>> I'm not sending a patch to add the library files at this time. I'll send\n>> that patch when this patch is working.\n>> \n>> The changes required to make it pass the tests shouldn't be very large.\n>\n> Please, no.  Let's not pull in a dependency for something as simple as a\n> string library.  How many distros have bstring pcakaged?  \n> The right version?\n\nThat's not a good argument. If dependancy is a problem, bsstring can\neasily be distributed as part of git. It's really small, so it wont\nmake git bloated:\n\n$ wc -l *.c *.h\n    82 bsafe.c\n  3462 bstest.c\n  1134 bstraux.c\n  2964 bstrlib.c\n   358 testaux.c\n    43 bsafe.h\n   112 bstraux.h\n   302 bstrlib.h\n   442 bstrwrap.h\n  8899 total\n\n> Does it work on Windows?\n\nThe library is totally stand alone, portable (known to work with\ngcc/g++, MSVC++, Intel C++, WATCOM C/C++, Turbo C, Borland C++, IBM's\nnative CC compiler on Windows, Linux and Mac OS X)\n\n> We already have strbuf.c, lets just consolidate the string\n> manipulation code already in git under that interface.\n\nThe right question is: what does git need. One way to consolidate\nstrbuf would be to simply\n\n$ rm strbuf.{c,h}\n$ unzip bsstring.zip\n\nand if people decide that git needs a non-trivial string library,\nwritting/testing more code in strbuf.c would probably be more work\nthan just reading what bsstring code does to become familiar enough\nwith it to even be able to maintain it later.\n\nIf people decide that git needs a really trivial string library, then\na few improvements to stbuf.c can be good.\n\nI'd argue in favor of the first option. C strings are horrible, and I\nthink doing something pleasant to use and safe is not completely\ntrivial. But I'm not a big contributor enough to really decide in\nspite of others ;-).\n\n-- \nMatthieu\n"},{"id":"52674","messageId":"buotzq8o78l.fsf@dhapc248.dev.necel.com","threadId":"9779","inReplyTo":"vpq642pkoln.fsf@bauges.imag.fr","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Miles Bader","fromEmail":"miles.bader@necel.com","sentAt":"2007-09-06T02:30:50Z","receivedAt":"2007-09-06T02:30:50Z","isPatch":false,"sender":{"key":"miles.bader@necel.com","avatar":"https://gravatar.com/avatar/be062d4050eb88e04229cbdb60f803e1bd647923a015996c2439e76f23e336a7?d=mp&s=160"},"body":"Matthieu Moy <Matthieu.Moy@imag.fr> writes:\n> and if people decide that git needs a non-trivial string library,\n> writting/testing more code in strbuf.c would probably be more work\n> than just reading what bsstring code does to become familiar enough\n> with it to even be able to maintain it later.\n\n>From what I've seen (by perusing the bstring website), bstring is kind\nof ugly though....\n\n-Miles\n\n-- \n\"Suppose we've chosen the wrong god. Every time we go to church we're\njust making him madder and madder.\" -- Homer Simpson\n"},{"id":"52686","messageId":"4AFD7EAD1AAC4E54A416BA3F6E6A9E52@ntdev.corp.microsoft.com","threadId":"9779","inReplyTo":"vpq642pkoln.fsf@bauges.imag.fr","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Dmitry Kakurin","fromEmail":"dmitry.kakurin@gmail.com","sentAt":"2007-09-06T04:48:47Z","receivedAt":"2007-09-06T04:48:47Z","isPatch":false,"sender":{"key":"dmitry.kakurin@gmail.com","avatar":null},"body":"[ snip ]\n\nWhen I first looked at Git source code two things struck me as odd:\n1. Pure C as opposed to C++. No idea why. Please don't talk about \nportability, it's BS.\n2. Brute-force, direct string manipulation. It's both verbose and \nerror-prone. This makes it hard to follow high-level code logic.\n\n- Dmitry\n"},{"id":"52688","messageId":"20070906045942.GR18160@spearce.org","threadId":"9779","inReplyTo":"4AFD7EAD1AAC4E54A416BA3F6E6A9E52@ntdev.corp.microsoft.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-09-06T04:59:42Z","receivedAt":"2007-09-06T04:59:42Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Dmitry Kakurin <dmitry.kakurin@gmail.com> wrote:\n> When I first looked at Git source code two things struck me as odd:\n> 1. Pure C as opposed to C++. No idea why. Please don't talk about \n> portability, it's BS.\n\nGit's creator (Linus) codes in C, not C++.  He has at various times\nstated reasons why he does not use C++.  I'm sure one can find such\nmessages with a bit of searching on mailing lists that he frequents.\nHe has his reasons.  I also happen to agree with at least some\nof them.  :)\n\nGit evolved from that initial prototype that Linus created.  I'm not\nsure how much code survives from that initial few versions that\nLinus managed before Junio took over, but nobody wanted to rewrite\nthings that already work so it just stayed in C.\n\"If it works, don't fix it.\"\n\nC works.  We (now) have 83,215 lines of it.  Its not going away\nanytime soon in Git.  It is also a relatively simple language that\na large number of open source programmers know.  This makes it easy\nfor them to get involved in the project.  Instead of say Haskell,\nwhich has a smaller community.  Or Tcl/Tk as we recently found out\nin the Git User Survey.  :-\\\n\n-- \nShawn.\n"},{"id":"52690","messageId":"buoir6oo05v.fsf@dhapc248.dev.necel.com","threadId":"9779","inReplyTo":"4AFD7EAD1AAC4E54A416BA3F6E6A9E52@ntdev.corp.microsoft.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Miles Bader","fromEmail":"miles.bader@necel.com","sentAt":"2007-09-06T05:03:40Z","receivedAt":"2007-09-06T05:03:40Z","isPatch":false,"sender":{"key":"miles.bader@necel.com","avatar":"https://gravatar.com/avatar/be062d4050eb88e04229cbdb60f803e1bd647923a015996c2439e76f23e336a7?d=mp&s=160"},"body":"Dmitry Kakurin <dmitry.kakurin@gmail.com> writes:\n> When I first looked at Git source code two things struck me as odd:\n> 1. Pure C as opposed to C++. No idea why. Please don't talk about\n> portability, it's BS.\n\nJust to piss you off.\n\n-Miles\n\n-- \nLove is a snowmobile racing across the tundra.  Suddenly it flips over,\npinning you underneath.  At night the ice weasels come.  --Nietzsche\n"},{"id":"52717","messageId":"46DFC490.3060200@op5.se","threadId":"9779","inReplyTo":"20070906045942.GR18160@spearce.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-06T09:12:48Z","receivedAt":"2007-09-06T09:12:48Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Shawn O. Pearce wrote:\n> Dmitry Kakurin <dmitry.kakurin@gmail.com> wrote:\n>> When I first looked at Git source code two things struck me as odd:\n>> 1. Pure C as opposed to C++. No idea why. Please don't talk about \n>> portability, it's BS.\n> \n> It is also a relatively simple language that\n> a large number of open source programmers know.  This makes it easy\n> for them to get involved in the project.\n\n\nThis is important. Git contains code from more than 300 people. I'm\nguessing you could cut that number by 2/3 if it had been written in C++.\n\nGit is cheating a bit though. Its primary audience was (and is) the\nvarious integrators working on the Linux kernel, all of whom are fairly\ncompetent C programmers.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"52720","messageId":"7vhcm8b0h8.fsf@gitster.siamese.dyndns.org","threadId":"9779","inReplyTo":"46DFC490.3060200@op5.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-09-06T09:35:15Z","receivedAt":"2007-09-06T09:35:15Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Andreas Ericsson <ae@op5.se> writes:\n\n> Git is cheating a bit though. Its primary audience was (and is) the\n> various integrators working on the Linux kernel, all of whom are fairly\n> competent C programmers.\n\nDo we still have a huge overlap with the kernel people?  I had\nan impression that patches from the kernel folks, with notable\nexception from a handful (you know who you are), have petered\nout rapidly after the first several weeks.\n"},{"id":"52724","messageId":"85zm00dsta.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"46DFC490.3060200@op5.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-06T09:52:33Z","receivedAt":"2007-09-06T09:52:33Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Andreas Ericsson <ae@op5.se> writes:\n\n> Shawn O. Pearce wrote:\n>> Dmitry Kakurin <dmitry.kakurin@gmail.com> wrote:\n>>> When I first looked at Git source code two things struck me as odd:\n>>> 1. Pure C as opposed to C++. No idea why. Please don't talk about\n>>> portability, it's BS.\n>>\n>> It is also a relatively simple language that\n>> a large number of open source programmers know.  This makes it easy\n>> for them to get involved in the project.\n>\n>\n> This is important. Git contains code from more than 300 people. I'm\n> guessing you could cut that number by 2/3 if it had been written in\n> C++.\n\nC++ is a language without design discipline.  Its set of features and\nsyntactic elements is incontingent (for example, its templates started\nas a ripoff of Ada generics which would have been ok except for the\ncompletely braindead idea of taking the Ada angle bracket restriction\nsyntax along with it), and it is the task of each programmer to choose\na sane and manageable subset and style, and implement using that.  As\na consequence, every C++ programmer writes his own personal dialect of\nC++, and we have about 20 different incompatible implementations of\nmultidimensional numeric arrays, making a complete mockery of the\n\"code reuse\" mantra: C++ _projects_ can't actually usefully achieve\n\"multiple inheritance\" on a design/meta level: once you start with one\nnon-trivial design, fitting other separately evolved components with a\ndifferent style causes retrofitting nightmares.\n\nSo going to C++ means cutting down the amount of people who find\nthemselves comfortable with the actual design and layout down to maybe\n10% of those who would actually feel ok with the actual _algorithms_\nemployed.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52729","messageId":"46DFD4BA.2000401@op5.se","threadId":"9779","inReplyTo":"7vhcm8b0h8.fsf@gitster.siamese.dyndns.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-06T10:21:46Z","receivedAt":"2007-09-06T10:21:46Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> Andreas Ericsson <ae@op5.se> writes:\n> \n>> Git is cheating a bit though. Its primary audience was (and is) the\n>> various integrators working on the Linux kernel, all of whom are fairly\n>> competent C programmers.\n> \n> Do we still have a huge overlap with the kernel people?  I had\n> an impression that patches from the kernel folks, with notable\n> exception from a handful (you know who you are), have petered\n> out rapidly after the first several weeks.\n\nTrue, but the point I was trying to make is that because git is written\nin C, for an audience who are extremely at home with that particular\nlanguage, it quickly attracted contributors.\n\ngit log --pretty=short | sed -n 's/^Author: \\([^<]*\\)<.*$/\\1/p' | \\\n\tsort | uniq | wc -l\n\nreports 355 unique lines, although some authors are mentioned twice\n(Theodore Tso vs Theodore Ts'o). Cross-matching the kernel authors\nwith the git authors shows that git and linux have 111 developers\nin common, again reporting some of them twice. A quick visual scan\nshows the figure to be 106, assuming no two authors have the same\nname (including email addresses produced more unique contributors as\npeople change email more often than they change name).\n\nIt's not unreasonable to say that git got at least 106 C-programmers\n\"for free\" included in their userbase round about the same second\nLinus went public with his intentions of managing the linux kernel\nin git, all of which are obviously comfortable enough with C to\npoke around in the kernel.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"52743","messageId":"Pine.LNX.4.64.0709061308140.28586@racer.site","threadId":"9779","inReplyTo":"buoir6oo05v.fsf@dhapc248.dev.necel.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-09-06T12:08:53Z","receivedAt":"2007-09-06T12:08:53Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 6 Sep 2007, Miles Bader wrote:\n\n> Dmitry Kakurin <dmitry.kakurin@gmail.com> writes:\n> > When I first looked at Git source code two things struck me as odd:\n> > 1. Pure C as opposed to C++. No idea why. Please don't talk about\n> > portability, it's BS.\n> \n> Just to piss you off.\n\nHehe.\n\nFWIW I strongly disagree that it's BS.  As others have stated, the reasons \nare easily found, and they are no weak arguments.\n\nCiao,\nDscho\n"},{"id":"52789","messageId":"alpine.LFD.0.999.0709061839510.5626@evo.linux-foundation.org","threadId":"9779","inReplyTo":"4AFD7EAD1AAC4E54A416BA3F6E6A9E52@ntdev.corp.microsoft.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-09-06T17:50:28Z","receivedAt":"2007-09-06T17:50:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 5 Sep 2007, Dmitry Kakurin wrote:\n> \n> When I first looked at Git source code two things struck me as odd:\n> 1. Pure C as opposed to C++. No idea why. Please don't talk about portability,\n> it's BS.\n\n*YOU* are full of bullshit.\n\nC++ is a horrible language. It's made more horrible by the fact that a lot \nof substandard programmers use it, to the point where it's much much \neasier to generate total and utter crap with it. Quite frankly, even if \nthe choice of C were to do *nothing* but keep the C++ programmers out, \nthat in itself would be a huge reason to use C.\n\nIn other words: the choice of C is the only sane choice. I know Miles \nBader jokingly said \"to piss you off\", but it's actually true. I've come \nto the conclusion that any programmer that would prefer the project to be \nin C++ over C is likely a programmer that I really *would* prefer to piss \noff, so that he doesn't come and screw up any project I'm involved with.\n\nC++ leads to really really bad design choices. You invariably start using \nthe \"nice\" library features of the language like STL and Boost and other \ntotal and utter crap, that may \"help\" you program, but causes:\n\n - infinite amounts of pain when they don't work (and anybody who tells me \n   that STL and especially Boost are stable and portable is just so full \n   of BS that it's not even funny)\n\n - inefficient abstracted programming models where two years down the road \n   you notice that some abstraction wasn't very efficient, but now all \n   your code depends on all the nice object models around it, and you \n   cannot fix it without rewriting your app.\n\nIn other words, the only way to do good, efficient, and system-level and \nportable C++ ends up to limit yourself to all the things that are \nbasically available in C. And limiting your project to C means that people \ndon't screw that up, and also means that you get a lot of programmers that \ndo actually understand low-level issues and don't screw things up with any \nidiotic \"object model\" crap.\n\nSo I'm sorry, but for something like git, where efficiency was a primary \nobjective, the \"advantages\" of C++ is just a huge mistake. The fact that \nwe also piss off people who cannot see that is just a big additional \nadvantage.\n\nIf you want a VCS that is written in C++, go play with Monotone. Really. \nThey use a \"real database\". They use \"nice object-oriented libraries\". \nThey use \"nice C++ abstractions\". And quite frankly, as a result of all \nthese design decisions that sound so appealing to some CS people, the end \nresult is a horrible and unmaintainable mess.\n\nBut I'm sure you'd like it more than git.\n\n\t\t\tLinus\n"},{"id":"52824","messageId":"a1bbc6950709061721r537b153eu1b0bb3c27fb7bd51@mail.gmail.com","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709061839510.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Dmitry Kakurin","fromEmail":"dmitry.kakurin@gmail.com","sentAt":"2007-09-07T00:21:37Z","receivedAt":"2007-09-07T00:21:37Z","isPatch":false,"sender":{"key":"dmitry.kakurin@gmail.com","avatar":null},"body":"On 9/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> On Wed, 5 Sep 2007, Dmitry Kakurin wrote:\n> >\n> > When I first looked at Git source code two things struck me as odd:\n> > 1. Pure C as opposed to C++. No idea why. Please don't talk about portability,\n> > it's BS.\n>\n> *YOU* are full of bullshit.\n\nnice\n\n> C++ is a horrible language. It's made more horrible by the fact that a lot\n> of substandard programmers use it, to the point where it's much much\n> easier to generate total and utter crap with it. Quite frankly, even if\n> the choice of C were to do *nothing* but keep the C++ programmers out,\n> that in itself would be a huge reason to use C.\n>\n> In other words: the choice of C is the only sane choice. I know Miles\n> Bader jokingly said \"to piss you off\", but it's actually true. I've come\n> to the conclusion that any programmer that would prefer the project to be\n> in C++ over C is likely a programmer that I really *would* prefer to piss\n> off, so that he doesn't come and screw up any project I'm involved with.\n\nAs dinosaurs (who code exclusively in C) are becoming extinct, you\nwill soon find yourself alone with attitude like this.\n\nMeasuring number of people who contributed to Git is incorrect metric.\nObviously C++ developers can contribute C code. But assuming that they\nprefer it that way is wrong.\n\nI was coding in Assembly when there was no C.\nThen in C before C++ was created.\nNow days it's C++ and C#, and I have never looked back.\nBad developers will write bad code in any language. But penalizing\ngood developers for this illusive reason of repealing bad contributors\nis nonsense.\n\nAnyway I don't mean to start a religious C vs. C++ war. It's a matter\nof beliefs and as such pointless.\nI just wanted to get a sense of how many people share this \"Git should\nbe in pure C\" doctrine.\n-- \n- Dmitry\n"},{"id":"52828","messageId":"alpine.LFD.0.999.0709070135361.5626@evo.linux-foundation.org","threadId":"9779","inReplyTo":"a1bbc6950709061721r537b153eu1b0bb3c27fb7bd51@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-09-07T00:38:01Z","receivedAt":"2007-09-07T00:38:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Sep 2007, Dmitry Kakurin wrote:\n> \n> As dinosaurs (who code exclusively in C) are becoming extinct, you\n> will soon find yourself alone with attitude like this.\n\nUnlike you, I actually gave reasons for my dislike of C++, and pointed to \nexamples of the kinds of failures that it leads to.\n\nYou, on the other hand, have given no sane reasons *for* using C++.\n\nThe fact is, git is better than the other SCM's. And good taste (and C) is \none of the reasons for that.\n\nIt has nothing to do with dinosaurs. Good taste doesn't go out of style, \nand comparing C to assembler just shows that you don't have a friggin idea \nabout what you're talking about.\n\n\t\t\tLinus\n"},{"id":"52830","messageId":"a1bbc6950709061808q85cf75co75f2331dc2bdbcbe@mail.gmail.com","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709070135361.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Dmitry Kakurin","fromEmail":"dmitry.kakurin@gmail.com","sentAt":"2007-09-07T01:08:40Z","receivedAt":"2007-09-07T01:08:40Z","isPatch":false,"sender":{"key":"dmitry.kakurin@gmail.com","avatar":null},"body":"On 9/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Thu, 6 Sep 2007, Dmitry Kakurin wrote:\n> >\n> > As dinosaurs (who code exclusively in C) are becoming extinct, you\n> > will soon find yourself alone with attitude like this.\n>\n> Unlike you, I actually gave reasons for my dislike of C++, and pointed to\n> examples of the kinds of failures that it leads to.\n\nAs I said, it's a matter of believes. As such, any reasoning and\narguing will be endless and pointless, as for any other religious\nissue.\n\n> You, on the other hand, have given no sane reasons *for* using C++.\n\nI'll give you reasons why to use C++ for Git (not why C++ is better\nfor any project in general, as that again would be pointless):\n\n1. Good String class will make code much more readable (and\nsignificantly shorter)\n2. Good Buffer class - same reason\n3. Smart pointers and smart handles to manage memory and\nfile/socket/lock handles.\n\nAs it is right now, it's too hard to see the high-level logic thru\nthis endless-busy-work of micro-managing strings and memory.\n\n> The fact is, git is better than the other SCM's. And good taste (and C) is\n> one of the reasons for that.\n\nIMHO Git has a brilliant high-level design (object database, using\nhashes, simple and accessible storage for data and metadata). Kudos to\nyou!\nThe implementation: a mixture of C and shell scripts, command line\ninterface that has evolved bottom-up is so-so.\n\n> and comparing C to assembler just shows that you don't have a friggin idea\n> about what you're talking about.\n\nI don't see myself comparing assembler to C anywhere.\nI was pointing out that I've been programming in different languages\n(many more actually) and observed bad developers writing bad code in\nall of them. So this quality \"bad developer\" is actually\nlanguage-agnostic :-).\n-- \n- Dmitry\n"},{"id":"52831","messageId":"alpine.LFD.0.999.0709070203200.5626@evo.linux-foundation.org","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709070135361.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-09-07T01:12:18Z","receivedAt":"2007-09-07T01:12:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 7 Sep 2007, Linus Torvalds wrote:\n> \n> The fact is, git is better than the other SCM's. And good taste (and C) is \n> one of the reasons for that.\n\nTo be very specific:\n - simple and clear core datastructures, with *very* lean and aggressive \n   code to manage them that takes the whole approach of \"simplicity over \n   fancy\" to the extreme.\n - a willingness to not abstract away the data structures and algorithms, \n   because those are the *whole*point* of core git. \n\nAnd if you want a fancier language, C++ is absolutely the worst one to \nchoose. If you want real high-level, pick one that has true high-level \nfeatures like garbage collection or a good system integration, rather than \nsomething that lacks both the sparseness and straightforwardness of C, \n*and* doesn't even have the high-level bindings to important concepts. \n\nIOW, C++ is in that inconvenient spot where it doesn't help make things \nsimple enough to be truly usable for prototyping or simple GUI \nprogramming, and yet isn't the lean system programming language that C is \nthat actively encourags you to use simple and direct constructs.\n\n\t\t\t\tLinus\n"},{"id":"52832","messageId":"alpine.LFD.0.999.0709070212300.5626@evo.linux-foundation.org","threadId":"9779","inReplyTo":"a1bbc6950709061808q85cf75co75f2331dc2bdbcbe@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-09-07T01:27:46Z","receivedAt":"2007-09-07T01:27:46Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Sep 2007, Dmitry Kakurin wrote:\n> \n> As it is right now, it's too hard to see the high-level logic thru\n> this endless-busy-work of micro-managing strings and memory.\n\nTotal BS. The string/memory management is not at all relevant. Look at the \ncode (I bet you didn't). This isn't the important, or complex part.\n\n> IMHO Git has a brilliant high-level design (object database, using\n> hashes, simple and accessible storage for data and metadata). Kudos to\n> you!\n> The implementation: a mixture of C and shell scripts, command line\n> interface that has evolved bottom-up is so-so.\n\nThe only really important part is the *design*. The fact that some of it \nis in a \"prototyping language\" is exactly because it wasn't the core \nparts, and it's slowly getting replaced. C++ would in *no* way have been \nable to replace the shell scripts or perl parts.\n\nAnd C++ would in no way have made the truly core parts better. \n\n> > and comparing C to assembler just shows that you don't have a friggin idea\n> > about what you're talking about.\n> \n> I don't see myself comparing assembler to C anywhere.\n\nYou made a very clear \"assembler -> C -> C++/C#\" progression nin your \nlife, comparing my staying with C as a \"dinosaur\", as if it was some \ninescapable evolution towards a better/more modern language.\n\nWith zero basis for it, since in many ways C is much superior to C++ (and \neven more so C#) in both its portability and in its availability of \ninterfaces and low-level support.\n\n> I was pointing out that I've been programming in different languages\n> (many more actually) and observed bad developers writing bad code in\n> all of them. So this quality \"bad developer\" is actually\n> language-agnostic :-).\n\nYou can write bad code in any language. However, some languages, and \nespecially some *mental* baggages that go with them are bad.\n\nThe very fact that you come in as a newbie, point to some absolutely \n*trivial* patches, and use that as an argument for a language that the \noriginal author doesn't like, is a sign of you being a person who should \nbe disabused on any idiotic notions as soon as possible.\n\nThe things that actually *matter* for core git code is things like writing \nyour own object allocator to make the footprint be as small as possible in \norder to be able to keep track of object flags for a million objects \nefficiently. It's writing a parser for the tree objects that is basically \nfairly optimal, because there *is* no abstraction. Absolutely all of it is \nat the raw memory byte level.\n\nCan those kinds of things be written in other languages than C? Sure. But \nthey can *not* be written by people who think the \"high-level\" \ncapabilities of C++ string handling somehow matter.\n\nThe fact is, that is *exactly* the kinds of things that C excels at. Not \njust as a language, but as a required *mentality*. One of the great \nstrengths of C is that it doesn't make you think of your program as \nanything high-level. It's what makes you apparently prefer other \nlanguages, but the thing is, from a git standpoint, \"high level\" is \nexactly the wrong thing. \n\n\t\tLinus\n"},{"id":"52833","messageId":"Pine.LNX.4.64.0709061833040.2526@blackbox.fnordora.org","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709070203200.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"alan","fromEmail":"alan@clueserver.org","sentAt":"2007-09-07T01:40:54Z","receivedAt":"2007-09-07T01:40:54Z","isPatch":false,"sender":{"key":"alan@clueserver.org","avatar":null},"body":"On Fri, 7 Sep 2007, Linus Torvalds wrote:\n\n> IOW, C++ is in that inconvenient spot where it doesn't help make things\n> simple enough to be truly usable for prototyping or simple GUI\n> programming, and yet isn't the lean system programming language that C is\n> that actively encourags you to use simple and direct constructs.\n\nNot to mention try finding two C++ compilers that support the same \nlanguage features.  C is a known quantity. C++ depends on whos compiler \nyou use and what class libraries you use.  Trying to make those things \nwork crossplatform is not an easy task.  (Harder than it is in C at \nleast.)\n\nA number of years ago, a programmer who will not be named (and is not me), \ntried to port Perl to C++.  It was a disaster.  He found that every \ncompiler handled something differently.\n\nIf you stuck to one compiler, it might work.  But trying to get GCC to \nwork like MS C++ or Borland C++ or whatever is just asking for pain.\n\n-- \nRefrigerator Rule #1: If you don't remember when you bought it, Don't eat it.\n"},{"id":"52838","messageId":"D7BEA87D-1DCF-4A48-AD5B-0A3FDC973C8A@wincent.com","threadId":"9779","inReplyTo":"a1bbc6950709061721r537b153eu1b0bb3c27fb7bd51@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-09-07T03:06:28Z","receivedAt":"2007-09-07T03:06:28Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 7/9/2007, a las 2:21, Dmitry Kakurin escribió:\n\n> I just wanted to get a sense of how many people share this \"Git should\n> be in pure C\" doctrine.\n\nCount me as one of them. Git is all about speed, and C is the best  \nchoice for speed, especially in context of Git's workload.\n\nCheers,\nWincent\n"},{"id":"52839","messageId":"a1bbc6950709062009x59a41cb7re6051739c11e370c@mail.gmail.com","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709070212300.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Dmitry Kakurin","fromEmail":"dmitry.kakurin@gmail.com","sentAt":"2007-09-07T03:09:23Z","receivedAt":"2007-09-07T03:09:23Z","isPatch":false,"sender":{"key":"dmitry.kakurin@gmail.com","avatar":null},"body":"On 9/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> On Thu, 6 Sep 2007, Dmitry Kakurin wrote:\n> >\n> > As it is right now, it's too hard to see the high-level logic thru\n> > this endless-busy-work of micro-managing strings and memory.\n>\n> Total BS. The string/memory management is not at all relevant. Look at the\n> code (I bet you didn't). This isn't the important, or complex part.\n\nNot only have I looked at the code, I've also debugged it quite a bit.\nGranted most of my problems had to do with handling paths on Windows\n(i.e. string manipulations).\n\nLet me snip \"C is better than C++\" part ...\n> [ snip ]\n... and explain where I'm coming from:\nMy goal is to *use* Git. When something does not work *for me* I want\nto be able to fix it (and contribute the fix) in *shortest time\npossible* and with *minimal efforts*. As for me it's a diversion from\nmy main activities.\nThe fact that Git is written in C does not really contribute to that goal.\nSuggestion to use C++ is the only alternative with existing C codebase.\nSo while C++ may not be the best choice \"academically speaking\" it's\npretty much the only practical choice.\n\n\"Democracy is the worst form of government except for all those others\nthat have been tried.\" - Winston Churchill\n\nNow, I realize that I'm a very infrequent contributor to Git, but I\nwant my opinion to be heard.\nPeople who carry the main weight of developing and maintaining Git\nshould make the call.\n-- \n- Dmitry\n"},{"id":"52842","messageId":"loom.20070907T055946-637@post.gmane.org","threadId":"9779","inReplyTo":"D7BEA87D-1DCF-4A48-AD5B-0A3FDC973C8A@wincent.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Paul Wankadia","fromEmail":"junyer@gmail.com","sentAt":"2007-09-07T04:06:35Z","receivedAt":"2007-09-07T04:06:35Z","isPatch":false,"sender":{"key":"junyer@gmail.com","avatar":null},"body":"Wincent Colaiuta <win <at> wincent.com> writes:\n\n> > I just wanted to get a sense of how many people share this \"Git should\n> > be in pure C\" doctrine.\n> \n> Count me as one of them. Git is all about speed, and C is the best  \n> choice for speed, especially in context of Git's workload.\n\nI concur, but I also feel that D, Clean and OCaml are viable alternatives.\n"},{"id":"52843","messageId":"alpine.LFD.0.9999.0709070020130.21186@xanadu.home","threadId":"9779","inReplyTo":"loom.20070907T055946-637@post.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-09-07T04:30:27Z","receivedAt":"2007-09-07T04:30:27Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 7 Sep 2007, Paul Wankadia wrote:\n\n> Wincent Colaiuta <win <at> wincent.com> writes:\n> \n> > > I just wanted to get a sense of how many people share this \"Git should\n> > > be in pure C\" doctrine.\n> > \n> > Count me as one of them. Git is all about speed, and C is the best  \n> > choice for speed, especially in context of Git's workload.\n> \n> I concur, but I also feel that D, Clean and OCaml are viable alternatives.\n\nI happen to have zero experience with any of those, so if Git \ndevelopment was done with one of them, you'd have to count me out.\n\nC is simply the lingua franca when it comes to programming, and it \nhappens to be the fastest amongst portable languages too.\n\n\nNicolas\n"},{"id":"52865","messageId":"fbqmdu$udg$1@sea.gmane.org","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709070203200.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T05:09:26Z","receivedAt":"2007-09-07T05:09:26Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"Linus Torvalds wrote:\n> And if you want a fancier language, C++ is absolutely the worst one to \n> choose. If you want real high-level, pick one that has true high-level \n> features like garbage collection or a good system integration, rather than \n> something that lacks both the sparseness and straightforwardness of C, \n> *and* doesn't even have the high-level bindings to important concepts. \n> \n> IOW, C++ is in that inconvenient spot where it doesn't help make things \n> simple enough to be truly usable for prototyping or simple GUI \n> programming, and yet isn't the lean system programming language that C is \n> that actively encourags you to use simple and direct constructs.\n\nThe D programming language is a different take than C++ has on growing \nC. I'm curious what your thoughts on that are (D has garbage collection, \nwhile still retaining the ability to directly manage memory). Can you \nenumerate what you feel are the important concepts?\n"},{"id":"52852","messageId":"ee77f5c20709062248oc346ea3p12a410d6d92babfb@mail.gmail.com","threadId":"9779","inReplyTo":"a1bbc6950709062009x59a41cb7re6051739c11e370c@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Symonds","fromEmail":"dsymonds@gmail.com","sentAt":"2007-09-07T05:48:02Z","receivedAt":"2007-09-07T05:48:02Z","isPatch":false,"sender":{"key":"dsymonds@gmail.com","avatar":"https://gravatar.com/avatar/b22f5051cbfc11836e36cf7a690e6cde4e225d835e13295ff98d15c7a9ee3c0f?d=mp&s=160"},"body":"On 07/09/07, Dmitry Kakurin <dmitry.kakurin@gmail.com> wrote:\n> My goal is to *use* Git. When something does not work *for me* I want\n> to be able to fix it (and contribute the fix) in *shortest time\n> possible* and with *minimal efforts*. As for me it's a diversion from\n> my main activities.\n> The fact that Git is written in C does not really contribute to that goal.\n\nThat's just it -- Git's goal isn't to make it as easy as possible for\nGit _users_ to fix it (thought that is a nice thing to have). Git's\ngoal is to be a very good, very fast SCM. Bugs should be found and\nfixed, but that can most effectively be done by the people who are\nalready knowledgeable about Git's codebase (i.e. its developers), not\nits users.\n\n\nDave.\n"},{"id":"52854","messageId":"20070907061554.GB30161@thunk.org","threadId":"9779","inReplyTo":"a1bbc6950709062009x59a41cb7re6051739c11e370c@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-09-07T06:15:54Z","receivedAt":"2007-09-07T06:15:54Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Sep 06, 2007 at 08:09:23PM -0700, Dmitry Kakurin wrote:\n> > Total BS. The string/memory management is not at all relevant. Look at the\n> > code (I bet you didn't). This isn't the important, or complex part.\n> \n> Not only have I looked at the code, I've also debugged it quite a bit.\n> Granted most of my problems had to do with handling paths on Windows\n> (i.e. string manipulations).\n\nI consider string manipulation to be one of the places where C++ is a\ntotal disaster.  It's way to easy for idiots to do something like this:\n\n\ta = b + \"/share/\" + c + serial_num;\n\nwhere you can have absolutely no idea how many memory allocations are\ndone, due to type coercions, overloaded operators (good God, you can\noverload the comma operator in C++!!!), and then when something like\nthat ends up in an inner loop, the result is a disaster from a\nperformance point of view, and it's not even obvious *why*!\n\n> My goal is to *use* Git. When something does not work *for me* I want\n> to be able to fix it (and contribute the fix) in *shortest time\n> possible* and with *minimal efforts*. As for me it's a diversion from\n> my main activities.\n\nYes, and if you contribute something the shortest time possible, and\nit ends up being crap, who gets to rewrite it and fix it?  I've seen\ntoo many C++ programs which get this kind of crap added, and it's not\nnoticed right away (because C++ is really good at hiding such\nperformance killers so they are not visible), and then later on, it's\neven harder to find the performance problems and fix them.\n\n> Now, I realize that I'm a very infrequent contributor to Git, but I\n> want my opinion to be heard.\n\nAnd if git were written in C++, it's precisely the infrequent\ncontributors (who are in a hurry, who only care about the quick hack\nto get them going, and not about the long-term maintainability and\nperformance of the package) that are be in the position to do the\nmost damage...\n\n\t\t\t\t\t\t- Ted\n"},{"id":"52855","messageId":"46E0EEC6.4020004@op5.se","threadId":"9779","inReplyTo":"D7BEA87D-1DCF-4A48-AD5B-0A3FDC973C8A@wincent.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-07T06:25:10Z","receivedAt":"2007-09-07T06:25:10Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Wincent Colaiuta wrote:\n> El 7/9/2007, a las 2:21, Dmitry Kakurin escribió:\n> \n>> I just wanted to get a sense of how many people share this \"Git should\n>> be in pure C\" doctrine.\n> \n> Count me as one of them. Git is all about speed, and C is the best \n> choice for speed, especially in context of Git's workload.\n> \n\nNono, hand-optimized assembly is the best choice for speed. C is just\na little more portable ;-)\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"52857","messageId":"46E0F04D.7040101@op5.se","threadId":"9779","inReplyTo":"a1bbc6950709062009x59a41cb7re6051739c11e370c@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-07T06:31:41Z","receivedAt":"2007-09-07T06:31:41Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Dmitry Kakurin wrote:\n> On 9/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>> On Thu, 6 Sep 2007, Dmitry Kakurin wrote:\n>>> As it is right now, it's too hard to see the high-level logic thru\n>>> this endless-busy-work of micro-managing strings and memory.\n>> Total BS. The string/memory management is not at all relevant. Look at the\n>> code (I bet you didn't). This isn't the important, or complex part.\n> \n> Not only have I looked at the code, I've also debugged it quite a bit.\n> Granted most of my problems had to do with handling paths on Windows\n> (i.e. string manipulations).\n> \n> Let me snip \"C is better than C++\" part ...\n>> [ snip ]\n> ... and explain where I'm coming from:\n> My goal is to *use* Git. When something does not work *for me* I want\n> to be able to fix it (and contribute the fix) in *shortest time\n> possible* and with *minimal efforts*. As for me it's a diversion from\n> my main activities.\n> The fact that Git is written in C does not really contribute to that goal.\n\n\nCoupled with what you said in an earlier mail, namely\n---%<---%<---\n> Obviously C++ developers can contribute C code. But assuming that they\n> prefer it that way is wrong.\n> \n> I was coding in Assembly when there was no C.\n> Then in C before C++ was created.\n> Now days it's C++ and C#, and I have never looked back.\n---%<---%<---\n\nConsidering C appeared in 1972, and C++ appeared in 1985, you have been\nwriting C code for 13 years. And you're telling me that git being written\nin C prevents you from contributing?\n\nIf you want to do something useful in C++ for git, make it easy for C++\nprogrammers to write apps for it.\n\n> \n> Now, I realize that I'm a very infrequent contributor to Git, but I\n> want my opinion to be heard.\n> People who carry the main weight of developing and maintaining Git\n> should make the call.\n\nThey already have, but every now and then someone comes along and suggest\na complete rewrite in some other language. So far we've had Java (there's\nalways one...), Python and now C++.\n\nIt happens to all projects, sooner or later. The funny thing is that all those\npeople that want their favourite software to be rewritten in their favourite\nprogramming language always wants someone else to rewrite it for them.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"52860","messageId":"85ejhb7yzw.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"a1bbc6950709061721r537b153eu1b0bb3c27fb7bd51@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T06:47:47Z","receivedAt":"2007-09-07T06:47:47Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Dmitry Kakurin\" <dmitry.kakurin@gmail.com> writes:\n\n> On 9/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>>\n>> In other words: the choice of C is the only sane choice. I know\n>> Miles Bader jokingly said \"to piss you off\", but it's actually\n>> true. I've come to the conclusion that any programmer that would\n>> prefer the project to be in C++ over C is likely a programmer that\n>> I really *would* prefer to piss off, so that he doesn't come and\n>> screw up any project I'm involved with.\n>\n> As dinosaurs (who code exclusively in C) are becoming extinct, you\n> will soon find yourself alone with attitude like this.\n\nAs long as TeX, Emacs and vi are around, I would not worry too much\nabout dinosaurs in general.  But C++ is a cancerous dinosaur.  It has\ngrowths that just don't belong on a C body.\n\n> I was coding in Assembly when there was no C.  Then in C before C++\n> was created.  Now days it's C++ and C#, and I have never looked\n> back.  Bad developers will write bad code in any language. But\n> penalizing good developers for this illusive reason of repealing bad\n> contributors is nonsense.\n\nThe problem with C++ is that every C++ developer has his own style,\nand reuse is an illusion within that style.  Take a look at classes\nimplementing matrix arithmetic: there are as many around as the day is\nlong, and all of them are incompatible with one another.\n\nWith regard to programming styles, C++ does not support multiple\ninheritance.  For a single project grown from a single start, you can\nget reasonable solutions.  But combining stuff is creating maintenance\nmesses.\n\nWith C, the situation is not dissimilar, but you spent less time\nfighting the illusion that you don't need to reimplement, anyway.\n\n> I just wanted to get a sense of how many people share this \"Git\n> should be in pure C\" doctrine.\n\nWhat nonsense.  Large parts of git already are shell scripts, so\nobviously there is no such doctrine.  Just because C++ is not a sane\nproposition does not mean that others might not work.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52861","messageId":"85abrz7yw4.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"a1bbc6950709061808q85cf75co75f2331dc2bdbcbe@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T06:50:03Z","receivedAt":"2007-09-07T06:50:03Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Dmitry Kakurin\" <dmitry.kakurin@gmail.com> writes:\n\n> I'll give you reasons why to use C++ for Git (not why C++ is better\n> for any project in general, as that again would be pointless):\n>\n> 1. Good String class will make code much more readable (and\n> significantly shorter)\n> 2. Good Buffer class - same reason\n> 3. Smart pointers and smart handles to manage memory and\n> file/socket/lock handles.\n\nBut all of those are incompatible with another and require major\nheadaches and/or interface code to get to run with one another.  And\nthen might use different interface styles, anyway.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52862","messageId":"85642n7yrn.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"a1bbc6950709062009x59a41cb7re6051739c11e370c@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T06:52:44Z","receivedAt":"2007-09-07T06:52:44Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Dmitry Kakurin\" <dmitry.kakurin@gmail.com> writes:\n\n> ... and explain where I'm coming from:\n> My goal is to *use* Git. When something does not work *for me* I want\n> to be able to fix it (and contribute the fix) in *shortest time\n> possible* and with *minimal efforts*. As for me it's a diversion from\n> my main activities.\n> The fact that Git is written in C does not really contribute to that goal.\n> Suggestion to use C++ is the only alternative with existing C codebase.\n> So while C++ may not be the best choice \"academically speaking\" it's\n> pretty much the only practical choice.\n\nSorry, but for fixing things in C, I can look and work locally.  For\nfixing things in C++, I first need to understand the class\nhierarchies used in the project.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52870","messageId":"85k5r27wkv.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"fbqmdu$udg$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T07:40:00Z","receivedAt":"2007-09-07T07:40:00Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Walter Bright <boost@digitalmars.com> writes:\n\n> Linus Torvalds wrote:\n>> And if you want a fancier language, C++ is absolutely the worst one\n>> to choose. If you want real high-level, pick one that has true\n>> high-level features like garbage collection or a good system\n>> integration, rather than something that lacks both the sparseness\n>> and straightforwardness of C, *and* doesn't even have the high-level\n>> bindings to important concepts. \n>>\n>> IOW, C++ is in that inconvenient spot where it doesn't help make\n>> things simple enough to be truly usable for prototyping or simple\n>> GUI programming, and yet isn't the lean system programming language\n>> that C is that actively encourags you to use simple and direct\n>> constructs.\n>\n> The D programming language is a different take than C++ has on growing\n> C. I'm curious what your thoughts on that are (D has garbage\n> collection, while still retaining the ability to directly manage\n> memory). Can you enumerate what you feel are the important concepts?\n\nA design is perfect not when there is no longer anything you can add\nto it, but if there is no longer anything you can take away.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52871","messageId":"200709070841.33057.andyparkins@gmail.com","threadId":"9779","inReplyTo":"85ejhb7yzw.fsf@lola.goethe.zz","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-09-07T07:41:25Z","receivedAt":"2007-09-07T07:41:25Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Friday 2007 September 07, David Kastrup wrote:\n\n(Disclaimer: I'm certainly not joining the \"C++ for git\" chant; this reply is \nmerely to the statements made about C++ in David's message).\n\n> The problem with C++ is that every C++ developer has his own style,\n> and reuse is an illusion within that style.  Take a look at classes\n> implementing matrix arithmetic: there are as many around as the day is\n> long, and all of them are incompatible with one another.\n\nOne could say the same about any API.  \"Take a look at that C library libXYZ - \nit does exactly the same thing as libPQR but all the function calls and \nstructures are different.  Conclusion: C is shit\".  Obviously nonsense.\n\n> With regard to programming styles, C++ does not support multiple\n> inheritance.  For a single project grown from a single start, you can\n\nMultiple inheritance is the spawn of the devil, but C++ _does_ support it.\n\nForgetting about the terrible STL, to me there really is no difference between \nC and C++; you can be object oriented in C.  Take a look at the Linux kernel, \nit should be printed out, rolled up and used to beat the ideas into students \nlearning C++/Java/C#.   Object oriented design is a choice, and if you really \nwanted you could do it in assembly.\n\nI would imagine the reason people often turn up wanting to rewrite Linux and \ngit in C++ is because they are so object oriented in nature already and it's \nnatural to think \"wouldn't this be even better if I wrote it in an object \noriented language\"?  Maybe, maybe not, but why bother?\n\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"52873","messageId":"85bqce7v9j.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"200709070841.33057.andyparkins@gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T08:08:24Z","receivedAt":"2007-09-07T08:08:24Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Andy Parkins <andyparkins@gmail.com> writes:\n\n> On Friday 2007 September 07, David Kastrup wrote:\n>\n> (Disclaimer: I'm certainly not joining the \"C++ for git\" chant; this reply is \n> merely to the statements made about C++ in David's message).\n>\n>> The problem with C++ is that every C++ developer has his own style,\n>> and reuse is an illusion within that style.  Take a look at classes\n>> implementing matrix arithmetic: there are as many around as the day is\n>> long, and all of them are incompatible with one another.\n>\n> One could say the same about any API.  \"Take a look at that C\n> library libXYZ - it does exactly the same thing as libPQR but all\n> the function calls and structures are different.  Conclusion: C is\n> shit\".  Obviously nonsense.\n\nThe difference is that you can pass structures from one library into\nanother with tolerable efficiency.  Because there are only basically 2\nways to lay out a two-dimensional array of floats.\n\n>> With regard to programming styles, C++ does not support multiple\n>> inheritance.  For a single project grown from a single start, you\n>> can\n>\n> Multiple inheritance is the spawn of the devil, but C++ _does_\n> support it.\n\nWhat about \"With regard to programming styles\" did you not understand?\nI was not talking about a technical feature at class level, but about\ncode merging from multiple sources.\n\n> I would imagine the reason people often turn up wanting to rewrite\n> Linux and git in C++ is because they are so object oriented in\n> nature already and it's natural to think \"wouldn't this be even\n> better if I wrote it in an object oriented language\"?  Maybe, maybe\n> not, but why bother?\n\nMaintainability and extensibility certainly are valid arguments for\nrewrites.  But C++ does not really shine in that regard.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52874","messageId":"fbr1a2$qm7$1@sea.gmane.org","threadId":"9779","inReplyTo":"85k5r27wkv.fsf@lola.goethe.zz","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T08:15:06Z","receivedAt":"2007-09-07T08:15:06Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"David Kastrup wrote:\n> Walter Bright <boost@digitalmars.com> writes:\n>> The D programming language is a different take than C++ has on growing\n>> C. I'm curious what your thoughts on that are (D has garbage\n>> collection, while still retaining the ability to directly manage\n>> memory). Can you enumerate what you feel are the important concepts?\n> \n> A design is perfect not when there is no longer anything you can add\n> to it, but if there is no longer anything you can take away.\n\nI like to phrase that a slightly different way: anyone can make \nsomething complicated, but it takes genius to make something simple.\n\nA very big goal for D is to make what should be simple code, simple. It \nturns out that what's simple for a computer is complex for a human. So \nto design a language that is simple for programmers is (unfortunately) a \nrather complex problem. Or perhaps I'm just not smart enough <g>.\n\nA canonical example is that of a loop. Consider a simple C loop over an \narray:\n\nvoid foo(int array[10])\n{\n     for (int i = 0; i < 10; i++)\n     {   int value = array[i];\n         ... do something ...\n     }\n}\n\nIt's simple, but it has a lot of problems:\n\n1) i should be size_t, not int\n2) array is not checked for overflow\n3) 10 may not be the actual array dimension\n4) may be more efficient to step through the array with pointers, rather \nthan indices\n5) type of array may change, but the type of value may not get updated\n6) crashes if array is NULL\n7) only works with arrays and pointers\n\nSince this thread is talking about C++, let's look at the C++ version:\n\nvoid foo(std::vector<int> array)\n{\n   for (std::vector<int>::const_iterator\n        i = array.begin();\n        i != array.end();\n        i++)\n   {\n     int value = *i;\n     ... do something ...\n   }\n}\n\nIt has fewer latent bugs, but still:\n\n1) type of array may change, but the type of value may not get updated\n2) too darned much typing\n3) it's more complicated, not simpler\n\nFrankly, I don't want to write loops that way. I want to write them like \nthis:\n\nvoid foo(int[] array)\n{\n   foreach (value; array)\n   {\n     ... do something ...\n   }\n}\n\nAs a programmer, I'm specifying exactly what I want to happen without \nmuch extra puffery. It's less typing, simpler, and more resistant to bugs.\n\n1) correct loop index type is selected based on the type of array\n2) arrays carry with them their dimension, so foreach is guaranteed to \nstep through the loop the correct number of times\n3) implementation decides if pointers will do a better job than indices, \nbased on the compilation target\n4) type of value is inferred automatically from the type of array, so no \nworries if the type changes\n5) Null arrays have 0 length, so no crashing\n6) works with any collection type\n\n[This example is extracted from a presentation I've made.]\n\n------\nWalter Bright\nhttp://www.digitalmars.com  C, C++, D programming language compilers\nhttp://www.astoriaseminar.com  Extraordinary C++\n"},{"id":"52875","messageId":"851wda7ufz.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"fbr1a2$qm7$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T08:26:08Z","receivedAt":"2007-09-07T08:26:08Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Walter Bright <boost@digitalmars.com> writes:\n\n> A canonical example is that of a loop. Consider a simple C loop over\n> an array:\n>\n> void foo(int array[10])\n> {\n>     for (int i = 0; i < 10; i++)\n>     {   int value = array[i];\n>         ... do something ...\n>     }\n> }\n>\n> It's simple, but it has a lot of problems:\n>\n> 1) i should be size_t, not int\n\nWrong.  size_t is for holding the size of memory objects in bytes, not\nin terms of indices.  For indices, the best variable is of the same\ntype as the declared index maximum size, so here it is typeof(10),\nnamely int.\n\n> 2) array is not checked for overflow\n\nWhy should it?\n\n> 3) 10 may not be the actual array dimension\n\nYour point is?\n\n> 4) may be more efficient to step through the array with pointers,\n> rather than indices\n\nNo.  It is a beginners' and advanced users' mistake to think using\npointers for access is a good idea.  Trivial optimizations are what a\ncompiler is best at, not the user.  Using pointer manipulation will\nmore often than not break loop unrolling, loop reversal, strength\nreduction and other things.\n\n> 5) type of array may change, but the type of value may not get\n> updated\n\nHuh?\n\n> 6) crashes if array is NULL\n\nCertainly.  Your point being?\n\n> 7) only works with arrays and pointers\n\nSince there are only arrays and pointers in C, not really a restriction.\n\n>\n> Since this thread is talking about C++, let's look at the C++ version:\n>\n> void foo(std::vector<int> array)\n> {\n>   for (std::vector<int>::const_iterator\n>        i = array.begin();\n>        i != array.end();\n>        i++)\n>   {\n>     int value = *i;\n>     ... do something ...\n>   }\n> }\n\nWhere is my barf bag?\n\n> Frankly, I don't want to write loops that way. I want to write them\n> like this:\n>\n> void foo(int[] array)\n> {\n>   foreach (value; array)\n>   {\n>     ... do something ...\n>   }\n> }\n>\n> As a programmer, I'm specifying exactly what I want to happen without\n> much extra puffery. It's less typing, simpler, and more resistant to\n> bugs.\n>\n> 1) correct loop index type is selected based on the type of array\n> 2) arrays carry with them their dimension, so foreach is guaranteed to\n> step through the loop the correct number of times\n> 3) implementation decides if pointers will do a better job than\n> indices, based on the compilation target\n> 4) type of value is inferred automatically from the type of array, so\n> no worries if the type changes\n> 5) Null arrays have 0 length, so no crashing\n> 6) works with any collection type\n\nMost of those are toy concerns.  They prevent problems that don't\nactually occur much in practice.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52876","messageId":"fbr2iv$ugg$1@sea.gmane.org","threadId":"9779","inReplyTo":"D7BEA87D-1DCF-4A48-AD5B-0A3FDC973C8A@wincent.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T08:36:56Z","receivedAt":"2007-09-07T08:36:56Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"Wincent Colaiuta wrote:\n> Git is all about speed, and C is the best \n> choice for speed, especially in context of Git's workload.\n\nI can appreciate that. I originally got into writing compilers because \nmy game (Empire) ran too slowly and I thought the existing compilers \ncould be dramatically improved.\n\nAnd technically, yes, you can write code in C that is >= the speed of \nany other language (other than asm). But practically, this isn't \nnecessarily so, for the following reasons:\n\n1) You wind up having to implement the complex, dirty details of things \nyourself. The consequences of this are:\n\n    a) you pick a simpler algorithm (which is likely less efficient - I \nrun across bubble sorts all the time in code)\n\n    b) once you implement, tune, and squeeze all the bugs out of those \ncomplex, dirty details, you're reluctant to change it. You're reluctant \nto try a different algorithm to see if it's faster. I've seen this \neffect a lot in my own code. (I translated a large body of my own C++ \ncode that I'd spent months tuning to D, and quickly managed to get \nsignificantly more speed out of it, because it was much simpler to try \nout different algorithms/data structures.)\n\n2) Garbage collection has an interesting and counterintuitive \nconsequence. If you compare n malloc/free's with n gcnew/collections, \nthe malloc/free will come out faster, and you conclude that gc is slow. \nBut that misses one huge speed advantage of gc - you can do FAR fewer \nallocations! For example, I've done a lot of string manipulating \nprograms in C. The basic problem is keeping track of who owns each \nstring. This is done by, when in doubt, make a copy of the string.\n\nBut if you have gc, you don't worry about who owns the string. You just \nmake another pointer to it. D takes this a step further with the concept \nof array slicing, where one creates windows on existing arrays, or \nwindows on windows on windows, and no allocations are ever done. It's \njust pointer fiddling.\n\n------\nWalter Bright\nhttp://www.digitalmars.com  C, C++, D programming language compilers\nhttp://www.astoriaseminar.com  Extraordinary C++\n"},{"id":"52879","messageId":"fbr4oi$5ko$1@sea.gmane.org","threadId":"9779","inReplyTo":"851wda7ufz.fsf@lola.goethe.zz","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T09:14:01Z","receivedAt":"2007-09-07T09:14:01Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"David Kastrup wrote:\n> Walter Bright <boost@digitalmars.com> writes:\n> \n>> A canonical example is that of a loop. Consider a simple C loop over\n>> an array:\n>>\n>> void foo(int array[10])\n>> {\n>>     for (int i = 0; i < 10; i++)\n>>     {   int value = array[i];\n>>         ... do something ...\n>>     }\n>> }\n>>\n>> It's simple, but it has a lot of problems:\n>>\n>> 1) i should be size_t, not int\n> \n> Wrong.  size_t is for holding the size of memory objects in bytes, not\n> in terms of indices.  For indices, the best variable is of the same\n> type as the declared index maximum size, so here it is typeof(10),\n> namely int.\n\nThe easiest way to show the error is consider the code being ported to a \ntypical 64 bit C compiler. int's are still 32 bits, yet the array can be \nlarger than 32 bits. You're right in that what we want to be able to do \nis typeof(array dimension), but there is no way to do that automatically \nin C, which is my point. If the array dimension changes, you have to \ncarefully check to make sure every loop dependency on the type is \nupdated, too.\n\nsize_t will always work, however, making it a better choice than int, at \nleast for C.\n\n>> 2) array is not checked for overflow\n> \n> Why should it?\n\nBecause the 10 array dimension is not statically checked in C. I could \npass it a pointer to 3 ints without the compiler complaining. This makes \nit a potential maintenance problem. Also, the maintenance programmer may \nchange the array dimension in the function signature, but overlook \nchanging it in the for loop. Again, a maintenance problem.\n\n\n>> 3) 10 may not be the actual array dimension\n> \n> Your point is?\n\nArray buffer overflow errors are commonplace in C, because array \ndimensions are not automatically checked at either compile or run time. \nThis is an expensive problem. Some C APIs try to deal with this by \npassing a second argument for arrays giving the dimension (snprintf, for \nexample), but this tends to be sporadic, not conventional. It being \nextra work for the programmer inevitably means it doesn't get done.\n\n\n>> 4) may be more efficient to step through the array with pointers,\n>> rather than indices\n> \n> No.  It is a beginners' and advanced users' mistake to think using\n> pointers for access is a good idea.  Trivial optimizations are what a\n> compiler is best at, not the user.  Using pointer manipulation will\n> more often than not break loop unrolling, loop reversal, strength\n> reduction and other things.\n\nC compilers vary widely in the optimizations they'll do for simple \nloops. I see often enough attempts by programmers to take such matters \ninto their own hands. I agree with you on that - and suggest the \nlanguage should not tempt the user to do such optimizations.\n\n>> 5) type of array may change, but the type of value may not get\n>> updated\n> \n> Huh?\n\nLet's say our fearless maintenance programmer decides to make it an \narray of longs, not an array of ints. He overlooks changing the type of \nvalue in the loop. Suddenly, things subtly break because of overflows. \nOr maybe he changed the int to an unsigned, now the divides in the loop \ngive different answers. Etc. There really isn't any compiler/language \nhelp in finding these kinds of problems.\n\n\n>> 6) crashes if array is NULL\n> \n> Certainly.  Your point being?\n\nI consider an array that is NULL to have no members, so instead of \ncrashing the loop should execute 0 times.\n\n\n>> 7) only works with arrays and pointers\n> \n> Since there are only arrays and pointers in C, not really a restriction.\n\nC has structs, too, as well as more complicated user defined \ncollections. Essentially, you cannot (simply) write generic algorithms \nin C, because you cannot (simply) generically express iteration. Of \ncourse, you can still express anything in C if you're willing to work \nhard enough to get it. Me, I'm too lazy <g>. It's like why I can't play \nchess - everytime I try to play it instead I think about writing a \nprogram to do the hard work for me.\n\n\n>> As a programmer, I'm specifying exactly what I want to happen without\n>> much extra puffery. It's less typing, simpler, and more resistant to\n>> bugs.\n>>\n>> 1) correct loop index type is selected based on the type of array\n>> 2) arrays carry with them their dimension, so foreach is guaranteed to\n>> step through the loop the correct number of times\n>> 3) implementation decides if pointers will do a better job than\n>> indices, based on the compilation target\n>> 4) type of value is inferred automatically from the type of array, so\n>> no worries if the type changes\n>> 5) Null arrays have 0 length, so no crashing\n>> 6) works with any collection type\n> \n> Most of those are toy concerns.  They prevent problems that don't\n> actually occur much in practice.\n\nI beg to differ - buffer overflow bugs are common and expensive. The \nnice thing about the D loop is it is LESS typing than the C one - you \nget the extra robustness for free.\n\nLet's look at the code gen for the inner loop for C:\n\nL8:             push    [EBX*4][ESI]\n                 call    near ptr _bar\n                 inc     EBX\n                 add     ESP,4\n                 cmp     EBX,0Ah\n                 jb      L8\n\nand for D:\n\nLE:            mov     EAX,[EBX]\n                call    near ptr _D4test3barFiZv\n                add     EBX,4\n                cmp     EBX,ESI\n                jb      LE\n\nI think you can see that performance isn't an impediment.\n"},{"id":"52880","messageId":"774D124B-B37E-40D4-9C0F-B8B2E9C70288@wincent.com","threadId":"9779","inReplyTo":"loom.20070907T055946-637@post.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-09-07T09:19:43Z","receivedAt":"2007-09-07T09:19:43Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 7/9/2007, a las 6:06, Paul Wankadia escribió:\n\n> Wincent Colaiuta <win <at> wincent.com> writes:\n>\n>>> I just wanted to get a sense of how many people share this \"Git  \n>>> should\n>>> be in pure C\" doctrine.\n>>\n>> Count me as one of them. Git is all about speed, and C is the best\n>> choice for speed, especially in context of Git's workload.\n>\n> I concur, but I also feel that D, Clean and OCaml are viable  \n> alternatives.\n\nYes, they have reputation for speed[1], but also a smaller number of  \npeople know them[2].\n\n[1] <http://shootout.alioth.debian.org/gp4/benchmark.php? \ntest=all&lang=all>\n[2] <http://www.tiobe.com/tpci.htm>\n\nCheers,\nWincent\n"},{"id":"52881","messageId":"85wsv26cv8.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"fbr4oi$5ko$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T09:31:07Z","receivedAt":"2007-09-07T09:31:07Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Walter Bright <boost@digitalmars.com> writes:\n\n> David Kastrup wrote:\n>> Walter Bright <boost@digitalmars.com> writes:\n>>\n>>> A canonical example is that of a loop. Consider a simple C loop over\n>>> an array:\n>>>\n>>> void foo(int array[10])\n>>> {\n>>>     for (int i = 0; i < 10; i++)\n>>>     {   int value = array[i];\n>>>         ... do something ...\n>>>     }\n>>> }\n>>>\n>>> It's simple, but it has a lot of problems:\n>>>\n>>> 1) i should be size_t, not int\n>>\n>> Wrong.  size_t is for holding the size of memory objects in bytes, not\n>> in terms of indices.  For indices, the best variable is of the same\n>> type as the declared index maximum size, so here it is typeof(10),\n>> namely int.\n>\n> The easiest way to show the error is consider the code being ported to\n> a typical 64 bit C compiler. int's are still 32 bits, yet the array\n> can be larger than 32 bits.\n\nNot if it is an array declared of size 10.  And if it isn't, you have\nno business stating so in the function prototype.\n\nWillfully obfuscate programming does not prove anything.\n\n>>> 2) array is not checked for overflow\n>>\n>> Why should it?\n>\n> Because the 10 array dimension is not statically checked in C. I\n> could pass it a pointer to 3 ints without the compiler\n> complaining. This makes it a potential maintenance problem.\n\nNonsense.  Again, C won't keep you from shooting yourself in the foot.\n\n>>> 3) 10 may not be the actual array dimension\n>>\n>> Your point is?\n>\n> Array buffer overflow errors are commonplace in C, because array\n> dimensions are not automatically checked at either compile or run\n> time.\n\nNo, because programmers get things wrong.  You can tell C compilers to\ncheck all array accesses, but that is a performance issue.  For gcc,\nwe have\n\n`-fmudflap -fmudflapth -fmudflapir'\n     For front-ends that support it (C and C++), instrument all risky\n     pointer/array dereferencing operations, some standard library\n     string/heap functions, and some other associated constructs with\n     range/validity tests.  Modules so instrumented should be immune to\n     buffer overflows, invalid heap use, and some other classes of C/C++\n     programming errors.  The instrumentation relies on a separate\n     runtime library (`libmudflap'), which will be linked into a\n     program if `-fmudflap' is given at link time.  Run-time behavior\n     of the instrumented program is controlled by the `MUDFLAP_OPTIONS'\n     environment variable.  See `env MUDFLAP_OPTIONS=-help a.out' for\n     its options.\n\nWhy isn't it the default?  Because it is a performance issue.\n\n>>> 5) type of array may change, but the type of value may not get\n>>> updated\n>>\n>> Huh?\n>\n> Let's say our fearless maintenance programmer decides to make it an\n> array of longs, not an array of ints. He overlooks changing the type\n> of value in the loop.\n\nAgain: C does not prevent you from shooting yourself in the foot.\n\n>>> 6) crashes if array is NULL\n>>\n>> Certainly.  Your point being?\n>\n> I consider an array that is NULL to have no members,\n\nNobody else does that.\n\n> so instead of crashing the loop should execute 0 times.\n\nIf the loop count is zero, this is what will happen.\n\n>>> 7) only works with arrays and pointers\n>>\n>> Since there are only arrays and pointers in C, not really a\n>> restriction.\n>\n> C has structs, too, as well as more complicated user defined\n> collections. Essentially, you cannot (simply) write generic\n> algorithms in C, because you cannot (simply) generically express\n> iteration.\n\nOf course you can.  Macros exist.\n\n>> Most of those are toy concerns.  They prevent problems that don't\n>> actually occur much in practice.\n>\n> I beg to differ - buffer overflow bugs are common and expensive.\n\nThen compile your program with appropriate options.  The key word is\n\"option\".  You don't have to take the performance hit if you don't\nwant or need it.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52882","messageId":"20070907094120.GA27754@artemis.corp","threadId":"9779","inReplyTo":"fbqmdu$udg$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-09-07T09:41:20Z","receivedAt":"2007-09-07T09:41:20Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Fri, Sep 07, 2007 at 05:09:26AM +0000, Walter Bright wrote:\n> Linus Torvalds wrote:\n> >And if you want a fancier language, C++ is absolutely the worst one to \n> >choose. If you want real high-level, pick one that has true high-level \n> >features like garbage collection or a good system integration, rather \n> >than something that lacks both the sparseness and straightforwardness of \n> >C, *and* doesn't even have the high-level bindings to important \n> >concepts. IOW, C++ is in that inconvenient spot where it doesn't help \n> >make things simple enough to be truly usable for prototyping or simple \n> >GUI programming, and yet isn't the lean system programming language that \n> >C is that actively encourags you to use simple and direct constructs.\n> \n> The D programming language is a different take than C++ has on growing C. \n> I'm curious what your thoughts on that are (D has garbage collection, \n> while still retaining the ability to directly manage memory). Can you \n> enumerate what you feel are the important concepts?\n\n  Well, to me D has two significant drawbacks to be \"ready to use\". The\nfirst one is that it doesn't has bit-fields. I often deal with bit-fields\non structures that have a _lot_ of instances in my program, and the\nbit-field is chosen for code readability _and_ structure size efficiency.\nI know you pretend that using masks manually often generates better\ncode. But in my case, speed does not matter _that_ much. I mean it does,\nbut not that this micro-level as access to the bit-field is not my\ninner-loop.\n\n  The other second issue I have, is that there is no way to do:\n  import (C) \"foo.h\"\n\n  And this is a big no-go (maybe not for git, but as a general issue)\nbecause it impedes the use of external libraries with a C interface a\n_lot_. E.g. I'd really like to use it to use some GNU libc extensions,\nbut I can't because it has too many dependencies (some async getaddrinfo\ninterface, that need me to import all the signal events and so on\nextensions in the libc, with bitfields, wich send us back to the first\npoint).\n\n\n  I also have a third, but non critical issue, I absolutely don't like\nphobos :) Though I'm obviously free to chose another library. D has\ndefinitely many many many real advances over C (like the .init, .size,\n... and so on fields, known types, and whatever portability nightmare\nthe C impose us). In fact I like to use D like I code in C, using\nmodules and functions, and very few classes, as few as I can. And even\n(under- ?) using D like this, it is a real pleasure to work with. I'm\nreally eager to see gdc be more stable.\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"52883","messageId":"46E11CE1.4030209@op5.se","threadId":"9779","inReplyTo":"fbr2iv$ugg$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-07T09:41:53Z","receivedAt":"2007-09-07T09:41:53Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Walter Bright wrote:\n> Wincent Colaiuta wrote:\n>> Git is all about speed, and C is the best choice for speed, especially \n>> in context of Git's workload.\n> \n> I can appreciate that. I originally got into writing compilers because \n> my game (Empire) ran too slowly and I thought the existing compilers \n> could be dramatically improved.\n> \n> And technically, yes, you can write code in C that is >= the speed of \n> any other language (other than asm). But practically, this isn't \n> necessarily so, for the following reasons:\n> \n> 1) You wind up having to implement the complex, dirty details of things \n> yourself. The consequences of this are:\n> \n>    a) you pick a simpler algorithm (which is likely less efficient - I \n> run across bubble sorts all the time in code)\n> \n>    b) once you implement, tune, and squeeze all the bugs out of those \n> complex, dirty details, you're reluctant to change it. You're reluctant \n> to try a different algorithm to see if it's faster. I've seen this \n> effect a lot in my own code. (I translated a large body of my own C++ \n> code that I'd spent months tuning to D, and quickly managed to get \n> significantly more speed out of it, because it was much simpler to try \n> out different algorithms/data structures.)\n> \n\nI haven't seen this in the development of git, although to be fair, you\ndidn't mention the number of developers that were simultaneously working\non your project. If it was you alone, I can imagine you were reluctant to\nchange it just to see if something is faster.\n\nOpensource projects with many contributors (git, linux) work differently,\nsince one or a few among the plethora of authors will almost always be\na true expert at the problem being solved.\n\nThe current pack-format and how it's read is one such example. It was\ndone once, by the combined efforts of Linus and Junio (this is all off\nthe top of my head and I cba to go looking up the details, so bear with\nme if there are errors). Linus and Junio are both very good C-programmers,\nbut the handling of packfiles was not what you'd call their specialty.\nAlong came Nicolas Pitre, another excellent C programmer, who probably\nhas done some similar work before. He constructed a better algorithm,\neventually resulting in the ultimate performance win with a net gain\nin both time and size (gj, Nicolas).\n\nThe point is that, given enough developers, *someone* is bound to\nfind an algorithm that works so well that it's no longer worth\ninvesting time to even discuss if anything else would work better,\neither because it moves the performance bottleneck to somewhere else\n(where further speedups would no longer produce humanly measurable\nimprovements), or because the action seems instantanous to the user\n(further improvements simply aren't worth it, because no valuable\nresource will be saved from it).\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"52887","messageId":"Pine.LNX.4.64.0709071119510.28586@racer.site","threadId":"9779","inReplyTo":"a1bbc6950709061721r537b153eu1b0bb3c27fb7bd51@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-09-07T10:21:13Z","receivedAt":"2007-09-07T10:21:13Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 6 Sep 2007, Dmitry Kakurin wrote:\n\n> Anyway I don't mean to start a religious C vs. C++ war.\n\nYou have a very strange way of not meaning to start a C vs. C++ war.\n\n> It's a matter of beliefs and as such pointless.\n\nNo, it's not.  As has been shown by some very good _arguments_.  Once you \nhave facts to back up your claims, it is not any belief any longer.\n\nCiao,\nDscho\n"},{"id":"52888","messageId":"Pine.LNX.4.64.0709071124020.28586@racer.site","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709070212300.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-09-07T10:26:39Z","receivedAt":"2007-09-07T10:26:39Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 7 Sep 2007, Linus Torvalds wrote:\n\n> On Thu, 6 Sep 2007, Dmitry Kakurin wrote:\n> \n> > I was pointing out that I've been programming in different languages \n> > (many more actually) and observed bad developers writing bad code in \n> > all of them. So this quality \"bad developer\" is actually \n> > language-agnostic :-).\n> \n> You can write bad code in any language. However, some languages, and \n> especially some *mental* baggages that go with them are bad.\n\nThere is an important additional point: a language like C _holds_ you to a \ncertain degree of diligence.\n\nIn my day-job I have to code in other languages, which make it \"easy\" to \ncode.  As a result, the code I have to work with is sloppy, ugly and \nbuggy.  By applying the same principles I am _forced_ to use in C, with \nGit, I produce better code.\n\nCiao,\nDscho\n"},{"id":"52889","messageId":"Pine.LNX.4.64.0709071126460.28586@racer.site","threadId":"9779","inReplyTo":"a1bbc6950709062009x59a41cb7re6051739c11e370c@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-09-07T10:28:45Z","receivedAt":"2007-09-07T10:28:45Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 6 Sep 2007, Dmitry Kakurin wrote:\n\n> Now, I realize that I'm a very infrequent contributor to Git, but I want \n> my opinion to be heard.\n\nWe are a happy little meritocracy here.  Once you proved that you're not \nfull of shit (some seem to try the opposite, you know who you are), you \ncan go all caps.  Before that, you'll have to show that you earn to be \nheard first.\n\nCiao,\nDscho\n"},{"id":"52892","messageId":"46E12C5D.8000902@etek.chalmers.se","threadId":"9779","inReplyTo":"46DDC500.5000606@etek.chalmers.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2007-09-07T10:47:57Z","receivedAt":"2007-09-07T10:47:57Z","isPatch":false,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Lukas Sandström wrote:\n> Hi.\n> \n> This is an attempt to use \"The Better String Library\"[1] in builtin-mailinfo.c\n> \n> The patch doesn't pass all the tests in the testsuit yet, but I thought I'd\n> send it out so people can decide if they like how the code looks.\n> \n> I'm not sending a patch to add the library files at this time. I'll send\n> that patch when this patch is working.\n> \n> The changes required to make it pass the tests shouldn't be very large.\n> \n> /Lukas\n> \n> [1] http://bstring.sourceforge.net/\n> \n> ---\n>  builtin-mailinfo.c |  795 ++++++++++++++++++++++++++--------------------------\n>  1 files changed, 392 insertions(+), 403 deletions(-)\n\nUnfortunatley, I haven't had any time inte the last few days to code, nor read\nmail. I'm assuming that there is no point in me finishing the patch and that git\nwill go with the strbuf solution?\n\n/Lukas\n"},{"id":"52893","messageId":"Pine.LNX.4.64.0709071155570.28586@racer.site","threadId":"9779","inReplyTo":"46E0EEC6.4020004@op5.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-09-07T10:56:42Z","receivedAt":"2007-09-07T10:56:42Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 7 Sep 2007, Andreas Ericsson wrote:\n\n> Wincent Colaiuta wrote:\n> > El 7/9/2007, a las 2:21, Dmitry Kakurin escribi?:\n> > \n> > > I just wanted to get a sense of how many people share this \"Git should\n> > > be in pure C\" doctrine.\n> > \n> > Count me as one of them. Git is all about speed, and C is the best choice\n> > for speed, especially in context of Git's workload.\n> > \n> \n> Nono, hand-optimized assembly is the best choice for speed. C is just\n> a little more portable ;-)\n\nI have a buck here that says that you cannot hand-optimise assembly (on \nmodern processors at least) as good as even gcc.\n\nCiao,\nDscho\n"},{"id":"52898","messageId":"44DC6433-1BCD-4968-83E6-6631726E0523@wincent.com","threadId":"9779","inReplyTo":"46E0EEC6.4020004@op5.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-09-07T11:30:56Z","receivedAt":"2007-09-07T11:30:56Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 7/9/2007, a las 8:25, Andreas Ericsson escribió:\n\n> Nono, hand-optimized assembly is the best choice for speed. C is just\n> a little more portable ;-)\n\nFunny thing is, GCC almost certainly produces better-optimized  \nassembly than most programmers could... ;-)\n\nCheers,\nWincent\n"},{"id":"52899","messageId":"308098E9-CCAF-44F1-9FEF-F3933218401E@wincent.com","threadId":"9779","inReplyTo":"85k5r27wkv.fsf@lola.goethe.zz","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-09-07T11:36:51Z","receivedAt":"2007-09-07T11:36:51Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 7/9/2007, a las 9:40, David Kastrup escribió:\n\n> A design is perfect not when there is no longer anything you can add\n> to it, but if there is no longer anything you can take away.\n\nIl semble que la perfection soit atteinte non quand il n'y a plus  \nrien à ajouter, mais quand il n'y a plus rien à retrancher.\nPerfection is achieved, not when there is nothing more to add, but  \nwhen there is nothing left to take away.\nCh. III: L'Avion, p. 60\n\n<http://en.wikiquote.org/wiki/Exupery>\n"},{"id":"52901","messageId":"FEA805F3-A6BA-4C76-B2C7-E28C00FDD801@wincent.com","threadId":"9779","inReplyTo":"fbr2iv$ugg$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-09-07T11:52:01Z","receivedAt":"2007-09-07T11:52:01Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 7/9/2007, a las 10:36, Walter Bright escribió:\n\n> Wincent Colaiuta wrote:\n>> Git is all about speed, and C is the best choice for speed,  \n>> especially in context of Git's workload.\n>\n> I can appreciate that. I originally got into writing compilers  \n> because my game (Empire) ran too slowly and I thought the existing  \n> compilers could be dramatically improved.\n>\n> And technically, yes, you can write code in C that is >= the speed  \n> of any other language (other than asm). But practically, this isn't  \n> necessarily so, for the following reasons:\n>\n> 1) You wind up having to implement the complex, dirty details of  \n> things yourself. The consequences of this are:\n>\n>    a) you pick a simpler algorithm (which is likely less efficient  \n> - I run across bubble sorts all the time in code)\n>\n>    b) once you implement, tune, and squeeze all the bugs out of  \n> those complex, dirty details, you're reluctant to change it. You're  \n> reluctant to try a different algorithm to see if it's faster. I've  \n> seen this effect a lot in my own code. (I translated a large body  \n> of my own C++ code that I'd spent months tuning to D, and quickly  \n> managed to get significantly more speed out of it, because it was  \n> much simpler to try out different algorithms/data structures.)\n\nWhile I accept that this is generally true, I think Git is somewhat  \nof a special case. From a design perspective the data structures and  \nalgorithms are remarkably simple -- therein lies its elegance. I  \nthink it's precisely the kind of problem that can be tackled well  \nwith a close-to-the-metal language like C.\n\n> 2) Garbage collection has an interesting and counterintuitive  \n> consequence. If you compare n malloc/free's with n gcnew/ \n> collections, the malloc/free will come out faster, and you conclude  \n> that gc is slow. But that misses one huge speed advantage of gc -  \n> you can do FAR fewer allocations! For example, I've done a lot of  \n> string manipulating programs in C. The basic problem is keeping  \n> track of who owns each string. This is done by, when in doubt, make  \n> a copy of the string.\n>\n> But if you have gc, you don't worry about who owns the string. You  \n> just make another pointer to it. D takes this a step further with  \n> the concept of array slicing, where one creates windows on existing  \n> arrays, or windows on windows on windows, and no allocations are  \n> ever done. It's just pointer fiddling.\n\nThis mirrors my experience in desktop application development.  \nDespite GC being \"slower\" the app actually runs faster and a lot of  \nnasty problems (shared resources, locking etc) just magically go  \naway. Development is easier too.\n\nBut once again I think Git falls into a special category where the  \ndesign makes the \"hassle\" of developing in C worth it.\n\nWincent\n"},{"id":"52902","messageId":"46E13C0F.8040203@op5.se","threadId":"9779","inReplyTo":"Pine.LNX.4.64.0709071155570.28586@racer.site","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-07T11:54:55Z","receivedAt":"2007-09-07T11:54:55Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Johannes Schindelin wrote:\n> Hi,\n> \n> On Fri, 7 Sep 2007, Andreas Ericsson wrote:\n> \n>> Wincent Colaiuta wrote:\n>>> El 7/9/2007, a las 2:21, Dmitry Kakurin escribi?:\n>>>\n>>>> I just wanted to get a sense of how many people share this \"Git should\n>>>> be in pure C\" doctrine.\n>>> Count me as one of them. Git is all about speed, and C is the best choice\n>>> for speed, especially in context of Git's workload.\n>>>\n>> Nono, hand-optimized assembly is the best choice for speed. C is just\n>> a little more portable ;-)\n> \n> I have a buck here that says that you cannot hand-optimise assembly (on \n> modern processors at least) as good as even gcc.\n> \n\n\nhttp://www.gelato.unsw.edu.au/archives/git/0504/1746.html\n\nI win. Donate $1 to FSF next time you get the opportunity ;-)\n\nHand-optimized asm is faster because the optimizer in the compiler is a\ngeneral-purpose one that has to guess and make assumptions about the code\nand its input to make the correct decisions. While it gets things right\nin as many as 80% of the cases, there's still the 20% where it doesn't.\nA human can, with sufficient research and effort, make the same optimizations\nwhere they are correct but avoid the 20% erroneous ones.\n\nIf the compiler gets it wrong inside your innermost loop, it might be worth\nshaving those extra 0.0001 seconds off of each iteration, because in the long\nrun, world-wide, it might save several weeks worth of CPU-time every day.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"52904","messageId":"E4A6490A-ABA9-4383-978E-C7F2E4BC9C23@wincent.com","threadId":"9779","inReplyTo":"46E13C0F.8040203@op5.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-09-07T12:33:42Z","receivedAt":"2007-09-07T12:33:42Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 7/9/2007, a las 13:54, Andreas Ericsson escribió:\n\n> Johannes Schindelin wrote:\n>> Hi,\n>> On Fri, 7 Sep 2007, Andreas Ericsson wrote:\n>>> Wincent Colaiuta wrote:\n>>>> El 7/9/2007, a las 2:21, Dmitry Kakurin escribi?:\n>>>>\n>>>>> I just wanted to get a sense of how many people share this \"Git  \n>>>>> should\n>>>>> be in pure C\" doctrine.\n>>>> Count me as one of them. Git is all about speed, and C is the  \n>>>> best choice\n>>>> for speed, especially in context of Git's workload.\n>>>>\n>>> Nono, hand-optimized assembly is the best choice for speed. C is  \n>>> just\n>>> a little more portable ;-)\n>> I have a buck here that says that you cannot hand-optimise  \n>> assembly (on modern processors at least) as good as even gcc.\n>\n>\n> http://www.gelato.unsw.edu.au/archives/git/0504/1746.html\n>\n> I win. Donate $1 to FSF next time you get the opportunity ;-)\n\nWell, you picked a very specific algorithm amenable to that kind of  \noptimization: small, manageable, with a minimal and well-defined  \nperformance critical section that could be written in assembly. Note  \nhow a good chunk of the implementation was still in C. At most I'd  \ngive you 75 cents for that one. ;-)\n\nWincent\n"},{"id":"52908","messageId":"20070907125501.GA21142@diana.vm.bytemark.co.uk","threadId":"9779","inReplyTo":"E4A6490A-ABA9-4383-978E-C7F2E4BC9C23@wincent.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Karl Hasselström","fromEmail":"kha@treskal.com","sentAt":"2007-09-07T12:55:01Z","receivedAt":"2007-09-07T12:55:01Z","isPatch":false,"sender":{"key":"kha@treskal.com","avatar":"https://gravatar.com/avatar/f0120c734b5279b345075a28521e1ac66acb20c9913ffe9bf6ae97e53f7f3f13?d=mp&s=160"},"body":"On 2007-09-07 14:33:42 +0200, Wincent Colaiuta wrote:\n\n> Well, you picked a very specific algorithm amenable to that kind of\n> optimization: small, manageable, with a minimal and well-defined\n> performance critical section that could be written in assembly. Note\n> how a good chunk of the implementation was still in C.\n\nAnd this is of course exactly the kind of spot where you _would_ use\nassembly in the real world. 99.99% of code is better written in C than\nassembler, but there is that 0.01% where hand-coded assembler is a\nbetter choice.\n\n-- \nKarl Hasselström, kha@treskal.com\n      www.treskal.com/kalle\n"},{"id":"52909","messageId":"46E1590A.4060504@op5.se","threadId":"9779","inReplyTo":"E4A6490A-ABA9-4383-978E-C7F2E4BC9C23@wincent.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-07T13:58:34Z","receivedAt":"2007-09-07T13:58:34Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Wincent Colaiuta wrote:\n> El 7/9/2007, a las 13:54, Andreas Ericsson escribió:\n> \n>> Johannes Schindelin wrote:\n>>> Hi,\n>>> On Fri, 7 Sep 2007, Andreas Ericsson wrote:\n>>>> Wincent Colaiuta wrote:\n>>>>> El 7/9/2007, a las 2:21, Dmitry Kakurin escribi?:\n>>>>>\n>>>>>> I just wanted to get a sense of how many people share this \"Git \n>>>>>> should\n>>>>>> be in pure C\" doctrine.\n>>>>> Count me as one of them. Git is all about speed, and C is the best \n>>>>> choice\n>>>>> for speed, especially in context of Git's workload.\n>>>>>\n>>>> Nono, hand-optimized assembly is the best choice for speed. C is just\n>>>> a little more portable ;-)\n>>> I have a buck here that says that you cannot hand-optimise assembly \n>>> (on modern processors at least) as good as even gcc.\n>>\n>>\n>> http://www.gelato.unsw.edu.au/archives/git/0504/1746.html\n>>\n>> I win. Donate $1 to FSF next time you get the opportunity ;-)\n> \n> Well, you picked a very specific algorithm amenable to that kind of \n> optimization: small, manageable, with a minimal and well-defined \n> performance critical section that could be written in assembly. Note how \n> a good chunk of the implementation was still in C. At most I'd give you \n> 75 cents for that one. ;-)\n> \n\nYes, but that's what I said in the original email as well. C is just so\nmuch more pleasant to write in that the only place you'd (sanely) use\nasm is in exactly these tight loops, where the code is likely to be used\nand reused until the algorithm it describes is no longer a viable option\nfor doing what it was originally designed to do.\n\nIt still proves the point though, as surely as n+1 > n for any value of n:\nHand-optimized assembly is faster than compiler-optimized C code.\n\nIt might be harder to do properly on some architectures than others (RISC\ncomes to mind), but it's still possible.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"52910","messageId":"03C31350-67A5-4B94-A841-5CEB37E78DE8@wincent.com","threadId":"9779","inReplyTo":"46E1590A.4060504@op5.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-09-07T14:13:24Z","receivedAt":"2007-09-07T14:13:24Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 7/9/2007, a las 15:58, Andreas Ericsson escribió:\n\n> Yes, but that's what I said in the original email as well. C is  \n> just so\n> much more pleasant to write in that the only place you'd (sanely) use\n> asm is in exactly these tight loops, where the code is likely to be  \n> used\n> and reused until the algorithm it describes is no longer a viable  \n> option\n> for doing what it was originally designed to do.\n>\n> It still proves the point though, as surely as n+1 > n for any  \n> value of n:\n> Hand-optimized assembly is faster than compiler-optimized C code.\n\nIn a theoretical ideal world, yes; no one would argue that C is  \nfaster than fine-tuned assembly.\n\nBut in the *real world* rewriting Git in assembly would be like  \npainting a house using a single horse hair instead of a paint brush  \nor roller. Your SHA-1 example is a perfect example of where you  \nbenefit from doing a tiny embellished detail using the single hair  \n(assembly) and leave all the rest in C.\n\nIn the real world and not the theoretical ideal world, it's not just  \nabout the diminishing returns you get from writing more and more of a  \ncode base in assembly instead of just the performance-critical  \nbottlenecks; it's that you're more likely to make subtle mistakes or  \neven make things slower. GCC does a remarkable job of optimizing in a  \nhuge number of use cases, and best of all, it does it for free.  \nPersonal opinion, of course, but that's the way I think it is.\n\nCheers,\nWincent\n"},{"id":"52916","messageId":"85ejha5ufw.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"Pine.LNX.4.64.0709071155570.28586@racer.site","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T16:09:07Z","receivedAt":"2007-09-07T16:09:07Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> On Fri, 7 Sep 2007, Andreas Ericsson wrote:\n>\n>> Wincent Colaiuta wrote:\n>> > El 7/9/2007, a las 2:21, Dmitry Kakurin escribi?:\n>> > \n>> > > I just wanted to get a sense of how many people share this \"Git should\n>> > > be in pure C\" doctrine.\n>> > \n>> > Count me as one of them. Git is all about speed, and C is the best choice\n>> > for speed, especially in context of Git's workload.\n>> > \n>> \n>> Nono, hand-optimized assembly is the best choice for speed. C is just\n>> a little more portable ;-)\n>\n> I have a buck here that says that you cannot hand-optimise assembly\n> (on modern processors at least) as good as even gcc.\n\nThat assumes that the original task can even expressed well in C.\nMultiple precision arithmetic, for example, requires access to the\ncarry bit.  You can code around this, for example by writing something\nlike\n\nunsigned a,b,carry;\n\n[...]\n\ncarry = (a+b) < a;\n\nbut the problem is that those are ad-hoc idioms with a variety of\npossibilities, and thus the compilers are not made to recognize them.\nAnother thing is mixed-precision multiplications and divisions: those\nare _natural_ operations on a normal CPU, but have no representation\nin assembly language.\n\nAs a consequence, most high performance multiple-precision packages\ncontain assembly language in some form or other.\n\ngcc's assembly language template are excellent in that they actually\ncooperate nicely with the optimizer, so the optimizer can do all the\naddress calculations and register assignments and opcode reorderings,\nand then the actual operations that are not expressible in C can be\ndone by the programmer.\n\nBut anyway, I have worked as a graphics driver programmer for some\namount of time, and bit-stuffing memory-mapped areas with data was\nstill something where hand assembly was best.\n\nI have also done BIOS terminal emulators, and being able to write\nsomething like\n\nld b,whatever\nmyloop:\npush bc\npush hl\ncall nextchar\npop hl\npop bc\nld (hl),a\ninc hl\ndjnz myloop\n\nin order to suspend the terminal driver until the application comes up\nwith the next `whatever' output characters in an escape sequence is\n_wagonloads_ more maintainable than using a state machine or whatever\nelse for distributing material delivered into the driver.\n\nBut this requires that nextchar can do something like\nnextchar: ld (driverstack),sp\n  ld sp,(appstack)\n  ret\n\nand the entrypoint, in contrast, does\n\noutchar: ld (appstack),sp\n  ld sp,(driverstack)\n  ret\n\nCheap and expedient.  You just need to set up a small stack, and\npresto: coroutines, at absolutely negligible cost.  I know that there\nare some \"portable\" coroutine implementations that use setjmp/longjmp\nin a rather horrific way, but those are way more unnatural.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52927","messageId":"fbs79k$tac$1@sea.gmane.org","threadId":"9779","inReplyTo":"20070907094120.GA27754@artemis.corp","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T19:03:24Z","receivedAt":"2007-09-07T19:03:24Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"Pierre Habouzit wrote:\n>   Well, to me D has two significant drawbacks to be \"ready to use\". The\n> first one is that it doesn't has bit-fields. I often deal with bit-fields\n> on structures that have a _lot_ of instances in my program, and the\n> bit-field is chosen for code readability _and_ structure size efficiency.\n> I know you pretend that using masks manually often generates better\n> code. But in my case, speed does not matter _that_ much. I mean it does,\n> but not that this micro-level as access to the bit-field is not my\n> inner-loop.\n\nI'm surprised this is such an important issue. Others have mentioned it, \nbut regard it as a minor thing. Interestingly, the htod program (which \nconverts C .h files to D import files) will convert bit fields to inline \nfunctions, giving equivalent functionality.\n\n>   The other second issue I have, is that there is no way to do:\n>   import (C) \"foo.h\"\n> \n>   And this is a big no-go (maybe not for git, but as a general issue)\n> because it impedes the use of external libraries with a C interface a\n> _lot_. E.g. I'd really like to use it to use some GNU libc extensions,\n> but I can't because it has too many dependencies (some async getaddrinfo\n> interface, that need me to import all the signal events and so on\n> extensions in the libc, with bitfields, wich send us back to the first\n> point).\n\nD does come with htod, which converts C .h files to D files. It's not \npossible to do a perfect job (because of macros), but it comes pretty \ndarned close. The reason htod gets so close is because it is actually a \nreal C compiler front end, not a perl or regex string processing hack.\n\nBecause it (may) require a little hand tweaking of the results (again, \nbecause C headers may include awful things like:\n\t#define BEGIN {\n\t#define print printf(\n), it's a separate program rather than built-in.\n\n\n>   I also have a third, but non critical issue, I absolutely don't like\n> phobos :)\n\nYou're not the only one <g>. But I'll add that access to the standard C \nruntime library *is* a part of D, so at some level it can't be worse \nthan C. There's also another runtime library available, Tango, which is \nvery popular.\n\n> Though I'm obviously free to chose another library. D has\n> definitely many many many real advances over C (like the .init, .size,\n> ... and so on fields, known types, and whatever portability nightmare\n> the C impose us). In fact I like to use D like I code in C, using\n> modules and functions, and very few classes, as few as I can. And even\n> (under- ?) using D like this, it is a real pleasure to work with. I'm\n> really eager to see gdc be more stable.\n\nThere are a lot of people hard at work on D to make it more stable and \nincrease the breadth and depth of tools available. I am fully aware that \nthere may be non-technical issues to using D in a project like git, like \navailability of other D programmers, tradition, etc., but in this thread \nI'm concerned mainly with technical issues.\n\nP.S. I'm also NOT suggesting that git be converted to D. Translating a \nworking, debugged, 80,000 line codebase from one language to another is \nusually a fool's errand.\n\nThanks for taking the time to post your thoughts.\n\n-----------\nWalter Bright\nhttp://www.digitalmars.com  C, C++, D programming language compilers\nhttp://www.astoriaseminar.com  Extraordinary C++\n"},{"id":"52929","messageId":"fbs8es$1cd$1@sea.gmane.org","threadId":"9779","inReplyTo":"46E11CE1.4030209@op5.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T19:23:16Z","receivedAt":"2007-09-07T19:23:16Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"Andreas Ericsson wrote:\n> Walter Bright wrote:\n>> 1) You wind up having to implement the complex, dirty details of \n>> things yourself. The consequences of this are:\n>>\n>>    a) you pick a simpler algorithm (which is likely less efficient - I \n>> run across bubble sorts all the time in code)\n>>\n>>    b) once you implement, tune, and squeeze all the bugs out of those \n>> complex, dirty details, you're reluctant to change it. You're \n>> reluctant to try a different algorithm to see if it's faster. I've \n>> seen this effect a lot in my own code. (I translated a large body of \n>> my own C++ code that I'd spent months tuning to D, and quickly managed \n>> to get significantly more speed out of it, because it was much simpler \n>> to try out different algorithms/data structures.)\n>>\n> \n> I haven't seen this in the development of git, although to be fair, you\n> didn't mention the number of developers that were simultaneously working\n> on your project.\n\nOn my project, one. But I've seen this problem repeatedly in other \nprojects that had multiple developers. For example, I used to use \nversion 1 of an assembler. It was itself written entirely in assembler. \nIt ran *incredibly* slowly on large asm files. But it was written in \nassembler, which is very fast, so how could that be?\n\nTurns out, the symbol table used internally was a linear one. A linear \nsymbol table is easy to implement, but doesn't scale well at all. A \nlinear symbol table was implemented because it was just harder to do \nmore advanced symbol table algorithms in assembler. In this case, a \nhigher level language re-implementation made the assembler much faster, \neven though that implementation was SLOWER in every detail. It was \nfaster overall, because it was easier to develop faster algorithms.\n\n\n> If it was you alone, I can imagine you were reluctant to\n> change it just to see if something is faster.\n\nMy point was that when I reimplemented it in D, the cost of changing the \nalgorithms got much lower, so I was much more tempted to muck around \ntrying out different ones. The result was I found faster ones.\n\n\n> Opensource projects with many contributors (git, linux) work differently,\n> since one or a few among the plethora of authors will almost always be\n> a true expert at the problem being solved.\n\nThat is a nice advantage. I don't think many projects can rely on having \nthe best in the business working on them, though <g>.\n\n\n> The point is that, given enough developers, *someone* is bound to\n> find an algorithm that works so well that it's no longer worth\n> investing time to even discuss if anything else would work better,\n> either because it moves the performance bottleneck to somewhere else\n> (where further speedups would no longer produce humanly measurable\n> improvements), or because the action seems instantanous to the user\n> (further improvements simply aren't worth it, because no valuable\n> resource will be saved from it).\n\nSure, but I suggest that few projects reach this maxima. Case in point: \nld, the gnu linker. It's terribly slow. To see how slow it is, compare \nit to optlink (the 15 years old one that comes with D for Windows). So I \ndon't believe there is anything inherent about linking that should make \nld so slow. There's some huge leverage possible in speeding up ld \n(spreading out that saved time among all the gnu developers).\n\nSo while git may have reached a maxima in performance, I don't think \nthis principle is applicable in general, even for very widely used open \nsource projects that would profit greatly from improved performance.\n\n------\nWalter Bright\nhttp://www.digitalmars.com  C, C++, D programming language compilers\nhttp://www.astoriaseminar.com  Extraordinary C++\n"},{"id":"52930","messageId":"fbs8jf$1cd$2@sea.gmane.org","threadId":"9779","inReplyTo":"FEA805F3-A6BA-4C76-B2C7-E28C00FDD801@wincent.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T19:25:44Z","receivedAt":"2007-09-07T19:25:44Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"Wincent Colaiuta wrote:\n> But once again I think Git falls into a special category where the \n> design makes the \"hassle\" of developing in C worth it.\n\nThat may very well be true. I've never looked at the source code for \ngit, so I'm not in any position to judge it. Nor do I suggest \ntranslating a debugged, working, 80,000 line project into another language.\n\nMy comments here are in more general terms.\n"},{"id":"52931","messageId":"85wsv22ry8.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"fbs79k$tac$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T19:31:11Z","receivedAt":"2007-09-07T19:31:11Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Walter Bright <boost@digitalmars.com> writes:\n\n> There are a lot of people hard at work on D to make it more stable\n> and increase the breadth and depth of tools available. I am fully\n> aware that there may be non-technical issues to using D in a project\n> like git, like availability of other D programmers, tradition, etc.,\n> but in this thread I'm concerned mainly with technical issues.\n>\n> P.S. I'm also NOT suggesting that git be converted to D. Translating\n> a working, debugged, 80,000 line codebase from one language to\n> another is usually a fool's errand.\n\nIn my opinion there is basically one area which C has botched up\nseriously in order to be useful as a general purpose language, and\nthat is conflating pointers and arrays, and allowing pointer\narithmetic.  The consequences are absolutely awful with regard to\ncompilers being able to optimize, and it is pretty much the primary\nreason that Fortran is still quite in use for numerical work.\n\nC has no usable two-dimensional (never mind higher dimensions) array\nconcept that would allow passing multidimensional arrays of\nruntime-determined size into functions.  Period.\n\nAdd to that the pointer aliasing problems affecting compilers, and C\nis useless for serious portable readable numerical work.\n\nFortran libraries like blas and lapack are ubiquitous after decades\nbecause the language can deal with multiple-dimension arrays sensibly,\nand could do so in the sixties already.\n\nC99 helps a bit.  But messing around with restrict pointers and\nsimilar means that to wring equal performance out of some trivial code\npiece (or permitting the compiler to do so without having to take\naliasing into account) is a lot of work and leads to ugly and\ninscrutable code.\n\nThat's the one thing that has seriously hampered C: the lack of a true\narray type on its own, decoupled from pointers.  It does not need to\ncarry its dimensions with it or other\nhide-the-implementation-from-the-programmer niceties: C is, after all,\na low-level language, and Fortran did not suffer from not having array\ndimensions packed into the arrays as well.\n\nBut that's water down the drawbridge.  This single major deficiency is\nnot anything that would hamper git development.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52932","messageId":"85sl5q2riz.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"fbs8es$1cd$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T19:40:20Z","receivedAt":"2007-09-07T19:40:20Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Walter Bright <boost@digitalmars.com> writes:\n\n> On my project, one. But I've seen this problem repeatedly in other\n> projects that had multiple developers. For example, I used to use\n> version 1 of an assembler. It was itself written entirely in\n> assembler. It ran *incredibly* slowly on large asm files. But it was\n> written in assembler, which is very fast, so how could that be?\n>\n> Turns out, the symbol table used internally was a linear one. A\n> linear symbol table is easy to implement, but doesn't scale well at\n> all.\n\nWell, my first system was a Z80 computer with an editor/assembler in\nROM (4kb).  At one time I tried figuring out the size requirements of\nsymbols.  It was two bytes for each symbol.  Namely the value.  The\n\"symbol table\" was located behind the source code.  Whenever this\nmarvel of technology encountered a label, it searched the source code\nfrom the beginning for the definition of the label, keeping count of\nall label definitions in between.  When it found the definition, the\ncount corresponded to the position in the symbol table.\n\nSo compilation times were O(ns), with n the number of symbol uses and\ns the size of the source code.\n\nImplementing in a higher language would not have helped: memory\nefficiency was what dictated this layout.  Given that the whole\navailable memory was perhaps 50kB, assembly language modules could not\nget so large that scale issues were deadly.  But the assembly times\ndid get annoying sometimes.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52933","messageId":"20070907194115.GA23483@artemis.corp","threadId":"9779","inReplyTo":"fbs79k$tac$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-09-07T19:41:15Z","receivedAt":"2007-09-07T19:41:15Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Fri, Sep 07, 2007 at 07:03:24PM +0000, Walter Bright wrote:\n> Pierre Habouzit wrote:\n> >  Well, to me D has two significant drawbacks to be \"ready to use\". The\n> >first one is that it doesn't has bit-fields. I often deal with \n> >bit-fields\n> >on structures that have a _lot_ of instances in my program, and the\n> >bit-field is chosen for code readability _and_ structure size \n> >efficiency.\n> >I know you pretend that using masks manually often generates better\n> >code. But in my case, speed does not matter _that_ much. I mean it does,\n> >but not that this micro-level as access to the bit-field is not my\n> >inner-loop.\n> \n> I'm surprised this is such an important issue. Others have mentioned it, \n> but regard it as a minor thing. Interestingly, the htod program (which \n> converts C .h files to D import files) will convert bit fields to inline \n> functions, giving equivalent functionality.\n\n  Well htod does that, but it's very impractical to write them from\nscratch. Especially if you want to benefit from the fact that padding\nand integer sizes are very well defined to map e.g. structs onto a raw\nstream, avoiding deserialization and so on. And for that bit-fields are\na really really fast and simple way to describe things.\n\n  I mean, take your classical example of the foreach loop. Your whole\npoint is that it's way shorter, and safer. And now you are saying that\npeople should instead of sth like:\n\n  struct my_struct {\n    unsigned some_field : 2;\n    unsigned has_this_property : 1;\n    unsigned is_in_this_state  : 1;\n    unsigned priority_level    : 2;\n    ...\n  }\n\n  people should write (IIRC it works since ->some_field = 2 calls\n->some_field(2) if the member does not exists, or maybe it's\nset_some_field, it's not very relevant anyway):\n\n  struct my_struct {\n    unsigned some_field() {\n      return this->real_field >> 30;\n    }\n\n    void some_field(unsigned value) {\n      this->real_field |= (value & 3) << 30;\n    }\n\n    ...\n\n  private:\n    unsigned real_field;\n  }\n\n  Please it has to be a joke: there is 42 ways for people to write it\nwrong (wrong shifts, wrong masks, and so on), it's horribly obfuscated,\nhence needs a lot of comments, whereas the bitfield is 90% self\ndocumented, and the syntax is _very_ clear, you cannot beat that. I\nwould be absolutely fine with it being syntactical sugar for some kind\nof template call though.\n\n  Not to mention that the usual C idiom:\n\n  union {\n    unsigned flags;\n    struct {\n      // many bitfields\n    };\n  };\n\n  Would need an explicit copy_flags(const my_struct foo) function to\nwork. Not pretty, not straightforward.\n\n  Really, I feel this is a big lack, for a language that aims at\nsimplicity, conciseness _and_ correctness.\n\n  OK, maybe I'm biased, I work with networks protocols all day long, so\nI often need bitfields, but still, a lot of people deal with network\nprotocols, it's not a niche.\n\n> >  The other second issue I have, is that there is no way to do:\n> >  import (C) \"foo.h\"\n> >  And this is a big no-go (maybe not for git, but as a general issue)\n> >because it impedes the use of external libraries with a C interface a\n> >_lot_. E.g. I'd really like to use it to use some GNU libc extensions,\n> >but I can't because it has too many dependencies (some async getaddrinfo\n> >interface, that need me to import all the signal events and so on\n> >extensions in the libc, with bitfields, wich send us back to the first\n> >point).\n> \n> D does come with htod, which converts C .h files to D files.\n\n  Last time I checked it was only available on windows, and closed\nsource, both are an impediment for many people. It's definitely clear\nthat gcc being opensource and available on so many platforms helped to\nmake C what it is today. Lacking portable and free (as in speech) tools\nare an impediment to the succes of a language. Right now, for D, only\ngdc exists, it lags behind dmd quite a lot afaict, and there is no other\ntoolchain helpers yet.\n\n> It's not possible to do a perfect job (because of macros), but it\n> comes pretty darned close. The reason htod gets so close is because it\n> is actually a real C compiler front end, not a perl or regex string\n> processing hack.\n> \n> Because it (may) require a little hand tweaking of the results (again, \n> because C headers may include awful things like:\n> \t#define BEGIN {\n> \t#define print printf(\n> ), it's a separate program rather than built-in.\n\n  Yeah I'm fine with that, but sadly it's not available everywhere like\nI said.\n\n> >  I also have a third, but non critical issue, I absolutely don't like\n> >phobos :)\n> \n> You're not the only one <g>. But I'll add that access to the standard C \n> runtime library *is* a part of D, so at some level it can't be worse than \n> C. There's also another runtime library available, Tango, which is very \n> popular.\n\n  I completely agree, and I knew about Tango, and anyways, I'm so used\nto C, and D has so few to bring to my code style when I deal with low\nlevel system functions, that I'm totally fine with std.c.* anyways :)\n\n  For the record I wasn't suggesting to rewrite git in D at all. I just\nhappened to see your post, and being very interested in where D is going\nbecause I feel it's an excellent langage, and saw an opportunity to\nmention a few quirks I feel it has, so, well, I answered :)\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"52934","messageId":"85odge2r0w.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"20070907194115.GA23483@artemis.corp","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T19:51:11Z","receivedAt":"2007-09-07T19:51:11Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Pierre Habouzit <madcoder@debian.org> writes:\n\n[bit fields]\n\n>   Really, I feel this is a big lack, for a language that aims at\n> simplicity, conciseness _and_ correctness.\n>\n>   OK, maybe I'm biased, I work with networks protocols all day long, so\n> I often need bitfields, but still, a lot of people deal with network\n> protocols, it's not a niche.\n\nAnd strictly speaking, C bitfields are completely useless for that\npurpose since the compiler is free to use whatever method he wants for\nallocating bit fields.  So if you want to write a portable program,\nyou are back to making the masks yourself.\n\nWhere bit fields work reliably is when you are not interchanging data\nwith other applications, but just laying out your internals.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52935","messageId":"20070907195940.GB23483@artemis.corp","threadId":"9779","inReplyTo":"85odge2r0w.fsf@lola.goethe.zz","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-09-07T19:59:40Z","receivedAt":"2007-09-07T19:59:40Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Fri, Sep 07, 2007 at 07:51:11PM +0000, David Kastrup wrote:\n> Pierre Habouzit <madcoder@debian.org> writes:\n> \n> [bit fields]\n> \n> >   Really, I feel this is a big lack, for a language that aims at\n> > simplicity, conciseness _and_ correctness.\n> >\n> >   OK, maybe I'm biased, I work with networks protocols all day long, so\n> > I often need bitfields, but still, a lot of people deal with network\n> > protocols, it's not a niche.\n> \n> And strictly speaking, C bitfields are completely useless for that\n> purpose since the compiler is free to use whatever method he wants for\n> allocating bit fields.  So if you want to write a portable program,\n> you are back to making the masks yourself.\n\n  The point is (1) D is not C, (2) we all know that linux e.g. does that\nin many places using the fact that it knows how the supported compilers\n(gcc icc tcc maybe some other) do their packing.\n\n  The discussion is about D. D solves the infamous problem with longs\nnot having the same size everywhere, I don't see why it couldn't solve\nthe bitfield issue either.\n\n> Where bit fields work reliably is when you are not interchanging data\n> with other applications, but just laying out your internals.\n\n  Thank you for the _C_ lesson.\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"52936","messageId":"fbsbul$dg0$1@sea.gmane.org","threadId":"9779","inReplyTo":"85wsv26cv8.fsf@lola.goethe.zz","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T20:22:53Z","receivedAt":"2007-09-07T20:22:53Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"David Kastrup wrote:\n> Again, C won't keep you from shooting yourself in the foot.\n\nRight, it won't. A good systems language should do what it can to \nprevent the programmer from *inadvertently* shooting himself in the \nfoot, while allowing him to *deliberately* shoot himself in the foot.\n\nIn the example loop, the cases I pointed out were ones where, if one \nchanges part of the code (such as the array dimension, or the array \ntype, etc.), then there are multiple places in the source that must be \nupdated to reflect it. The reality of human programmers is we update one \nor two places, and overlook the third place, and we now have a bug. \nSaying that one should just be a better programmer and not make such \nmistakes is a pipe dream.\n\nIdeally, each facet of the design of the code should have a single point \nwhere it can be changed, and then all dependencies on that design should \nbe automatically updated. That way, nothing gets overlooked. Doing this \nin C, such as using a #define for the array dimension, involves extra work.\n\nIt's a truism that if it involves extra work, then it often gets omitted.\n\nDoesn't \"A design is perfect not when there is no longer anything you \ncan add to it, but if there is no longer anything you can take away.\" \napply here? Going from:\n\n  void foo(int array[10])\n  {\n     for (int i = 0; i < 10; i++)\n     {   int value = array[i];\n         ... do something ...\n     }\n  }\n\nto:\n\n  void foo(int[] array)\n  {\n    foreach (value; array)\n    {\n      ... do something ...\n    }\n  }\n\ntakes a lot of frankly unnecessary things away, each of which is a \npotential source of error when maintaining the code.\n\n\n> No, because programmers get things wrong.\n\nExactly. That goes back to my point that a good language should help \nprevent inadvertent errors, while still allowing deliberate choices. D \napproaches this by making the correct approach essentially be the one \nwith minimal typing effort. To deliberately shoot yourself in the foot \nusually requires extra typing. For our loop example, we can still write \nthe C style loop (with all its potential problems) in D, but it requires \nextra effort to put in those potential problems. The easier, simpler way \ndoesn't have the problems.\n\n(The issue I have with C++'s fixes to various problems is they require \nextra typing (like the loop example), so guess what, people being people \ntend to not use them. This results in endless attempts to try and push \nC++ programmers into using the more verbose forms, a strategy I suspect \nwill be ultimately futile.)\n\n\n> You can tell C compilers to\n> check all array accesses, but that is a performance issue.\n\nRuntime checking of arrays in D is a performance issue too, so it is \nselectable via a command line switch. But more importantly,\n\n1) Static type checking of fixed size arrays works, so errors can be \ncaught at compile time.\n\n2) For dynamically sized arrays, the dimension of the array is carried \nwith the array, so loops automatically loop the correct number of times. \nNo runtime check is necessary, and it's easier for the code reviewer to \nvisually check the code for correctness.\n\n------\nWalter Bright\nhttp://www.digitalmars.com  C, C++, D programming language compilers\nhttp://www.astoriaseminar.com  Extraordinary C++\n"},{"id":"52938","messageId":"851wda2pbo.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"fbsbul$dg0$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T20:27:55Z","receivedAt":"2007-09-07T20:27:55Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Walter Bright <boost@digitalmars.com> writes:\n\n>  void foo(int array[10])\n>  {\n>     for (int i = 0; i < 10; i++)\n>     {   int value = array[i];\n>         ... do something ...\n>     }\n>  }\n>\n> to:\n>\n>  void foo(int[] array)\n>  {\n>    foreach (value; array)\n>    {\n>      ... do something ...\n>    }\n>  }\n>\n> takes a lot of frankly unnecessary things away, each of which is a\n> potential source of error when maintaining the code.\n\nThe problem is a toy problem: in real applications, you'll need to\naccess several data structures using the same index, and you'll need\nto be able to assign index values to temporary variables and so on.\nSo being able to hide the type of an index in one very specific\napplication (looping through a single array completely) at one place\nis not going to buy you much.\n\nAnyway, D is pretty much irrelevant as a perspective for git, so you\nshould take it to a language advocacy group.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52941","messageId":"fbsd0g$gt6$1@sea.gmane.org","threadId":"9779","inReplyTo":"20070907194115.GA23483@artemis.corp","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T20:40:56Z","receivedAt":"2007-09-07T20:40:56Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"Pierre Habouzit wrote:\n>   Well htod does that, but it's very impractical to write them from\n> scratch.\n\nTrue. I haven't tried yet (nobody else seems to care about it as much as \nyou do!), but I think this could be automated fairly easily with a template.\n\n\n> And for that bit-fields are\n> a really really fast and simple way to describe things.\n\nI should point out that inline functions are inlined, and there is no \nspeed difference in the result.\n\n\n>   Not to mention that the usual C idiom:\n> \n>   union {\n>     unsigned flags;\n>     struct {\n>       // many bitfields\n>     };\n>   };\n> \n>   Would need an explicit copy_flags(const my_struct foo) function to\n> work. Not pretty, not straightforward.\n\nI'm not following this. To copy a union, you just copy it with the \nassignment operator:\n\n\tU a, b;\n\ta = b;\t\t// copies all the bit fields, too!\n\n\n>> D does come with htod, which converts C .h files to D files.\n>   Last time I checked it was only available on windows, and closed\n> source, both are an impediment for many people.\n\nYou're right on both counts. It's because htod is built out of a fork of \nthe Digital Mars C compiler. Something similar could be done with gcc, \nbut I'm not the person to do it. I should also get off my lazy tail and \nport htod to linux.\n\n\n> Right now, for D, only\n> gdc exists, it lags behind dmd quite a lot afaict, and there is no other\n> toolchain helpers yet.\n\nGDC was just released for D 1.020, which is behind D 1.021, but 1.021 \nwas released just a couple days ago <g>.\n\n\n>   For the record I wasn't suggesting to rewrite git in D at all. I just\n> happened to see your post, and being very interested in where D is going\n> because I feel it's an excellent langage, and saw an opportunity to\n> mention a few quirks I feel it has, so, well, I answered :)\n\nAnd it's nice to hear your perspective, which is why I dropped by this \nthread.\n"},{"id":"52943","messageId":"fbsdga$iat$1@sea.gmane.org","threadId":"9779","inReplyTo":"85wsv22ry8.fsf@lola.goethe.zz","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T20:49:23Z","receivedAt":"2007-09-07T20:49:23Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"David Kastrup wrote:\n> In my opinion there is basically one area which C has botched up\n> seriously in order to be useful as a general purpose language, and\n> that is conflating pointers and arrays, and allowing pointer\n> arithmetic.  The consequences are absolutely awful with regard to\n> compilers being able to optimize, and it is pretty much the primary\n> reason that Fortran is still quite in use for numerical work.\n\nI agree. It's one of those things that probably sounded like a good idea \nat the time. The consequences were not foreseen. All languages have a \nfew of these (C++ has the infamous use of < > for template arguments).\n"},{"id":"52945","messageId":"20070907205604.GC23483@artemis.corp","threadId":"9779","inReplyTo":"fbsd0g$gt6$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-09-07T20:56:04Z","receivedAt":"2007-09-07T20:56:04Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Fri, Sep 07, 2007 at 08:40:56PM +0000, Walter Bright wrote:\n> Pierre Habouzit wrote:\n> >And for that bit-fields are\n> >a really really fast and simple way to describe things.\n>\n> I should point out that inline functions are inlined, and there is no \n> speed difference in the result.\n\n  I know that, and that's why I said I was totally fine with the\nbitfield notation to be only syntactic sugar on a template thingy if\nthat's the simplest way to have that it's OKay.\n\n> >  Not to mention that the usual C idiom:\n> >  union {\n> >    unsigned flags;\n> >    struct {\n> >      // many bitfields\n> >    };\n> >  };\n> >  Would need an explicit copy_flags(const my_struct foo) function to\n> >work. Not pretty, not straightforward.\n> \n> I'm not following this. To copy a union, you just copy it with the \n> assignment operator:\n> \n> \tU a, b;\n> \ta = b;\t\t// copies all the bit fields, too!\n\n  That was the point indeed. But if you don't have bitfields, you can't\ndo the union. And if the bitfield is just syntactic sugar, it may be\nunpossible to have such a union. But I may be wrong.\n\n> >Right now, for D, only\n> >gdc exists, it lags behind dmd quite a lot afaict, and there is no other\n> >toolchain helpers yet.\n> \n> GDC was just released for D 1.020, which is behind D 1.021, but 1.021 was \n> released just a couple days ago <g>.\n\n  Sure, but it does not works on amd64 properly (and it's the\narchitecture I care about) and is not ready for the current gcc (4.2,\nonly 4.1 builds) and so on. It's not as stable as DMD is. It does not\nlags too much version-wise, it lags in maturity. But well, youth has a\ncure: time :)\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"52957","messageId":"a1bbc6950709071517n7e7e99ffl3dd351092e7f19d6@mail.gmail.com","threadId":"9779","inReplyTo":"46E0F04D.7040101@op5.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Dmitry Kakurin","fromEmail":"dmitry.kakurin@gmail.com","sentAt":"2007-09-07T22:17:43Z","receivedAt":"2007-09-07T22:17:43Z","isPatch":false,"sender":{"key":"dmitry.kakurin@gmail.com","avatar":null},"body":"On 9/6/07, Andreas Ericsson <ae@op5.se> wrote:\n> They already have, but every now and then someone comes along and suggest\n> a complete rewrite in some other language. So far we've had Java (there's\n> always one...), Python and now C++.\n\nSince this \"complete rewrite\" was mentioned in multiple emails I'd\nlike to rectify that:\nWhat I'm offering (for Git) is to use C++ as a \"better C\".\nDon't change any existing *working* code, but start introducing simple\nC++ constructs in the new code.\nGit is simple enough to not require any high-level abstractions. But\nsome utility classes could make code much simpler.\n\nAnd BTW, I don't even like C++ that much :-), I just like it much\nbetter than C.  I've been saying that C++ is a legacy language for\nquite some time now. But we will use it for many years to come because\nthe size of this legacy code is huge, so there will be plenty of C++\ndevelopers available (to contribute to Git :-).\nAnd C++ is the only way to move with existing C codebase.\n-- \n- Dmitry\n"},{"id":"52958","messageId":"85sl5q1570.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"a1bbc6950709071517n7e7e99ffl3dd351092e7f19d6@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-07T22:28:03Z","receivedAt":"2007-09-07T22:28:03Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Dmitry Kakurin\" <dmitry.kakurin@gmail.com> writes:\n\n> On 9/6/07, Andreas Ericsson <ae@op5.se> wrote:\n>> They already have, but every now and then someone comes along and suggest\n>> a complete rewrite in some other language. So far we've had Java (there's\n>> always one...), Python and now C++.\n>\n> Since this \"complete rewrite\" was mentioned in multiple emails I'd\n> like to rectify that:\n> What I'm offering (for Git) is to use C++ as a \"better C\".\n> Don't change any existing *working* code, but start introducing simple\n> C++ constructs in the new code.\n\nYou are aware that the Linux kernel was kept compilable under g++ for\na while in its history?  You'll need more than vague words to erase\nthe memories from that experiment...\n\nJust compiling under C++, with no source changes, is likely to impact\nperformance and compile time rather badly, not to mention portability\n(you need the C++ runtime, for one thing).\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52960","messageId":"fbskq5$6le$1@sea.gmane.org","threadId":"9779","inReplyTo":"20070907205604.GC23483@artemis.corp","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T22:54:05Z","receivedAt":"2007-09-07T22:54:05Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"Pierre Habouzit wrote:\n>   Sure, but it does not works on amd64 properly (and it's the\n> architecture I care about) and is not ready for the current gcc (4.2,\n> only 4.1 builds) and so on. It's not as stable as DMD is. It does not\n> lags too much version-wise, it lags in maturity. But well, youth has a\n> cure: time :)\n\nYes, and the more people use it, the better it will get. These are all \nenvironmental problems, not technical limitations of the language.\n"},{"id":"52961","messageId":"fbsm43$ave$1@sea.gmane.org","threadId":"9779","inReplyTo":"851wda2pbo.fsf@lola.goethe.zz","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Walter Bright","fromEmail":"boost@digitalmars.com","sentAt":"2007-09-07T23:16:28Z","receivedAt":"2007-09-07T23:16:28Z","isPatch":false,"sender":{"key":"boost@digitalmars.com","avatar":null},"body":"David Kastrup wrote:\n> The problem is a toy problem: in real applications,\n\nNecessarily, to make an example suitable for a n.g. post, I ruthlessly \ncut down the size of it. This can have the inadvertent effect of making \nit appear trivial.\n\n> you'll need to\n> access several data structures using the same index, and you'll need\n> to be able to assign index values to temporary variables and so on.\n\nThe index is available:\n\n\tforeach (index, value; array)\n\t{\n\t\twritefln(\"array[%s] = %s\", index, value);\n\t}\n\nand it isn't necessary to worry about what the correct type for index \nis, as it is inferred.\n\n> So being able to hide the type of an index in one very specific\n> application (looping through a single array completely)\n\n  foreach'ing over a subset (i.e. slice) of an array:\n\n\tforeach (value; array[5 .. $])\n\t\t... loop from 5 to the end ...\n\n> at one place is not going to buy you much.\n\nExperience with foreach in real code shows that the for loop is what \nbecomes a rarity. Simple as it is, foreach is one of the best liked \nimprovements D has. And I speak as one who has written so many for loops \nthat spewing out:\n\n\tfor (int i = 0; i < 10; i++)\n\nis a 'finger' macro for me, i.e. my fingers blit it out without even \nthinking about it.\n\n > Anyway, D is pretty much irrelevant as a perspective for git, so you\n > should take it to a language advocacy group.\n\nI wished to answer your specific comments in this post.\n"},{"id":"52964","messageId":"a1bbc6950709071732s1f15e5ev28bdfc5c1ab5877b@mail.gmail.com","threadId":"9779","inReplyTo":"Pine.LNX.4.64.0709071119510.28586@racer.site","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Dmitry Kakurin","fromEmail":"dmitry.kakurin@gmail.com","sentAt":"2007-09-08T00:32:09Z","receivedAt":"2007-09-08T00:32:09Z","isPatch":false,"sender":{"key":"dmitry.kakurin@gmail.com","avatar":null},"body":"On 9/7/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> On Thu, 6 Sep 2007, Dmitry Kakurin wrote:\n>\n> > Anyway I don't mean to start a religious C vs. C++ war.\n>\n> You have a very strange way of not meaning to start a C vs. C++ war.\n\nI honestly didn't. I didn't even think it's possible. In the\nenvironment of mainstream commercial software development the last war\non this subj was over 8-10 years ago.\nEven wars like \"do we use exceptions/templates/stl\" are pretty much\nover. Now days it's \"do we use Boost\", or \"do we use template\nmetaprogramming\". But even more often it's Java/C# vs. C++.\n\nThat's why I was wondering how come C was chosen for Git.\n\n> > It's a matter of beliefs and as such pointless.\n>\n> No, it's not.  As has been shown by some very good _arguments_.  Once you\n> have facts to back up your claims, it is not any belief any longer.\n\nWell I've heard *opinions* and anecdotal evidence. No facts though.\nAnd it's not surprising. There could be no hard facts in such a\nmatter. It always boils down to \"most of all, I want my software to be\nX\" where X is different for different people (fast,maintainable,quick\nto market, scalable, beautiful, etc ... to name a few).\nWith different values of X any debate is pointless. And X is exactly\nthe matter of believes.\n\nAnyway my curiosity is satisfied (thru the roof so to speak) and I\nthink it's enough on the subj. It has reminded me of good old times\nthough.\n\n-- \n- Dmitry\n"},{"id":"52965","messageId":"a1bbc6950709071737h155c78b7ie6d2b77719239a6a@mail.gmail.com","threadId":"9779","inReplyTo":"85sl5q1570.fsf@lola.goethe.zz","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Dmitry Kakurin","fromEmail":"dmitry.kakurin@gmail.com","sentAt":"2007-09-08T00:37:38Z","receivedAt":"2007-09-08T00:37:38Z","isPatch":false,"sender":{"key":"dmitry.kakurin@gmail.com","avatar":null},"body":"On 9/7/07, David Kastrup <dak@gnu.org> wrote:\n> Just compiling under C++, with no source changes, is likely to impact\n> performance and compile time rather badly\n\nThis in fact is a very specific statement. Would you care to back it\nup with facts?\n\n-- \n- Dmitry\n"},{"id":"52967","messageId":"loom.20070908T025508-792@post.gmane.org","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709070203200.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"John 'Z-Bo' Zabroski","fromEmail":"johnzabroski@yahoo.com","sentAt":"2007-09-08T00:56:54Z","receivedAt":"2007-09-08T00:56:54Z","isPatch":false,"sender":{"key":"johnzabroski@yahoo.com","avatar":null},"body":"Linus Torvalds <torvalds <at> linux-foundation.org> writes:\n\n> \n> And if you want a fancier language, C++ is absolutely the worst one to \n> choose. If you want real high-level, pick one that has true high-level \n> features like garbage collection or a good system integration, rather than \n> something that lacks both the sparseness and straightforwardness of C, \n> *and* doesn't even have the high-level bindings to important concepts. \n> \n> IOW, C++ is in that inconvenient spot where it doesn't help make things \n> simple enough to be truly usable for prototyping or simple GUI \n> programming, and yet isn't the lean system programming language that C is \n> that actively encourags you to use simple and direct constructs.\n> \n> \t\t\t\tLinus\n> \n\nI want code that is Correct, Explicit, Fast, and in that order.\n\nI'm 23 years old and learned C++ when I was 13.  Back then, my compiler didn't\neven support \"bleeding edge\" C++ language features like namespaces.  I'm not a\nC++ expert, and I don't have the ego to call myself a superb programmer.  The\nlargest program I've written is 10K SLOC in C.  Yet, I'd like to participate in\nthis discussion, if that is OKay =)\n\nI do think I am capable of an honest critique of the downside of C++:\n\n_Problems_ _With_ _C++_\n\n *size*\n    On my bookshelf, most recent editions of the canonical C++ _books_:\n        Accelerated C++: Practical Programming by Example (336 pages)\n        The C++ Standard Template Library: A Tutorial and Reference (832 pages)\n        Effective C++: 50 Specific Ways to Improve Your Programs and Design (288\npages)\n        More Effective C++: 35 New Ways to Improve Your Programs and Designs\n(336 pages)\n        Exceptional C++: 47 Engineering Puzzles, Programming Problems, and\nSolutions (240 pages)\n        More Exceptional C++: 40 New Engineering Puzzles, Programming Problems,\nand Solutions (304 pages) \n        The C++ Programming Language (1030 pages)\n        Modern C++ Design: Generic Programming and Design Patterns Applied (352\npages)\n        C++ Templates: The Complete Guide (552 pages)\n\n    Altogether, that is 3918 pages.  K&R, the canonical C _book_, is 272 pages.\n Becoming a C++ language lawyer is much harder than becoming a C language\nlawyer.  Language lawyers know \"how not to hang oneself\" while programming in\nthe language.  I don't know how many of these titles are translated to other\nlanguages, however, I am sure the *effort* required to translate all of them is\nsignificant.  Open source is more successful if there is a lingua franca for\nprogramming, and that is C.  Now, it may move away from C over time, but it will\n*never* be C++ because it's encyclopedic.\n\n *hidden complexity*\n    (1) it's hard to say what code will compile down to.  viz., constructors can\nbe elided, but there is no fitness warranty; profiling your compiler to find out\nwhether it is elided is tedious and \"searching for secrets\" that should be\n_explicit_\n    (2) people don't understand static polymorphism and compile-time dispatch;\npeople are used to objects sending messages dynamically (run-time dispatch)\n    (3) coercion\n    (4) networks of objects are not explicitly laid out, hiding quadratically\ncomplex patterns of communication between objects\n    (5) data structure and data flow come before algorithms.  Sometimes, data\nstructure dictates data flow (ad-hoc networks of objects); sometimes, data flow\ndictates data structure (one of life's most disagreeable tasks - waiting in line\n- is characterized as FIFO).  This, I feel, is the most important point, because\nthe first rule of programming is to figure out what you want to say before you\nfigure out how to say it.  In C++, ad-hoc networks of objects with cyclic\nmessage paths are all too easy to create [see (4)] which means _code_ _is_ _not_\n_explicit_ and as a result _code_ _is_ _not_ _fast_.\n\n *transfer semantics on objects are not robust*\n    this ties into (1) in hidden complexity\n    the code author needs to specify a lot of boilerplate to achieve desired\ntransfer semantics on objects.  Similarly, the code audience, be it reviewer,\nmaintainer or merger, needs to read a lot of boilerplate to understand how\nobjects get moved around in memory.  Moreover, most of these concepts are\nintuitively declarative in nature, such as a parent object/child object relation.\n\n *poor re-use of effort*\n    \"code re-use\" is a misnomer; when programmers speak of code-reuse they mean\nre-use of effort.  There is no benefit to polymorphism if effort cannot be\nconsolidated easily.\n\n *C++ Standard iffy*\n    Some things just disappear quickly for *frantic* reasons (strstream was\nremoved for aesthetics), indicating not enough foresight into what is important.\n I do not want to pick a language where I have to worry about features in it's\n\"standard library\" becoming deprecated mainly for aesthetics.  As Dijkstra\npreached, programming is _not_ supposed to be a frantic exercise.\n\n *usually, better options*\n    See C++??: A Critique of C++ and Programing and Language Trends in the 1990s\n by Ian Joyner http://web.mac.com/joynerian/iWeb/Ian%20Joyner/CPPCritique.pdf\n(Somewhat outdated, but many of the points are intrinsic and will forever be\nrelevant).  You can add to the list of better options D 1.0.\n"},{"id":"52974","messageId":"858x7h1xpj.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"a1bbc6950709071732s1f15e5ev28bdfc5c1ab5877b@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-08T06:24:24Z","receivedAt":"2007-09-08T06:24:24Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Dmitry Kakurin\" <dmitry.kakurin@gmail.com> writes:\n\n> On 9/7/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>\n>> No, it's not.  As has been shown by some very good _arguments_.\n>> Once you have facts to back up your claims, it is not any belief\n>> any longer.\n>\n> Well I've heard *opinions* and anecdotal evidence. No facts though.\n\nAnecdotal evidence _is_ hard facts.  That's what experience is all\nabout.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52975","messageId":"854pi51xoj.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"a1bbc6950709071737h155c78b7ie6d2b77719239a6a@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-08T06:25:00Z","receivedAt":"2007-09-08T06:25:00Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Dmitry Kakurin\" <dmitry.kakurin@gmail.com> writes:\n\n> On 9/7/07, David Kastrup <dak@gnu.org> wrote:\n>> Just compiling under C++, with no source changes, is likely to impact\n>> performance and compile time rather badly\n>\n> This in fact is a very specific statement. Would you care to back it\n> up with facts?\n\nRead up on the Linux kernel history in the archives.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"52976","messageId":"85zlzxzmrn.fsf@lola.goethe.zz","threadId":"9779","inReplyTo":"loom.20070908T025508-792@post.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-08T06:36:44Z","receivedAt":"2007-09-08T06:36:44Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"John 'Z-Bo' Zabroski <johnzabroski@yahoo.com> writes:\n\n> Linus Torvalds <torvalds <at> linux-foundation.org> writes:\n>\n>> IOW, C++ is in that inconvenient spot where it doesn't help make\n>> things simple enough to be truly usable for prototyping or simple\n>> GUI programming, and yet isn't the lean system programming language\n>> that C is that actively encourags you to use simple and direct\n>> constructs.\n>\n> I want code that is Correct, Explicit, Fast, and in that order.\n\nOne beef I have with C++ is its automatic conversion rules.  They were\nobviously designed with two goals:\n\na) behave as C when not using user-defined types.  That's ok.\n\nb) behave like Fortran in mixed-type expressions involving \"complex\"\n   when using C++ (with any arbitrary user-defined type taking the\n   role of \"complex\").\n\nAnd b is just madness.  Not every user-defined arithmetic type is\ncomplex.  I did some work using modular arithmetic (GF(65521) and\nsimilar) and it was some hard work to keep values going through the\nwrong arithmetic conversions.  Basically trial and error and reading\nthe generated assembly code and head scratching and standard-reading.\n\nIn short: the automatic conversions made it hard to express what one\nwanted to get done, both for compiler as well as programmer.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"53015","messageId":"20070908232527.GB2645@steel.home","threadId":"9779","inReplyTo":"a1bbc6950709071732s1f15e5ev28bdfc5c1ab5877b@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-09-08T23:25:27Z","receivedAt":"2007-09-08T23:25:27Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"Dmitry Kakurin, Sat, Sep 08, 2007 02:32:09 +0200:\n> On 9/7/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > Hi,\n> >\n> > On Thu, 6 Sep 2007, Dmitry Kakurin wrote:\n> >\n> > > Anyway I don't mean to start a religious C vs. C++ war.\n> >\n> > You have a very strange way of not meaning to start a C vs. C++ war.\n> \n> I honestly didn't. I didn't even think it's possible. In the\n> environment of mainstream commercial software development the last war\n> on this subj was over 8-10 years ago.\n\nIt is because the \"environment of mainstream commercial software\ndevelopment\" is stuck in \"8-10\" back from now.\n\n> Even wars like \"do we use exceptions/templates/stl\" are pretty much\n> over. Now days it's \"do we use Boost\", or \"do we use template\n> metaprogramming\". But even more often it's Java/C# vs. C++.\n\nNow that's a stupid argument to bring up. Commercial software\ndevelopment is were the most stupid mistakes are done and repeated.\n\n> That's why I was wondering how come C was chosen for Git.\n\n\"Just to annoy mainstream commercial software developers\" would be a\ngood reason.\n"},{"id":"53020","messageId":"46E3354A.7030407@op5.se","threadId":"9779","inReplyTo":"fbsbul$dg0$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-08T23:50:34Z","receivedAt":"2007-09-08T23:50:34Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Walter Bright wrote:\n> David Kastrup wrote:\n>> Again, C won't keep you from shooting yourself in the foot.\n> \n> Right, it won't. A good systems language should do what it can to \n> prevent the programmer from *inadvertently* shooting himself in the \n> foot, while allowing him to *deliberately* shoot himself in the foot.\n> \n\nNo, a good systems language should do exactly what it's told. Supporting\ntools should tell the programmer if he's risking shooting himself in the\nfoot.\n\n> \n>> You can tell C compilers to\n>> check all array accesses, but that is a performance issue.\n> \n> Runtime checking of arrays in D is a performance issue too, so it is \n> selectable via a command line switch.\n\nSame as in C then.\n\n> But more importantly,\n> \n> 2) For dynamically sized arrays, the dimension of the array is carried \n> with the array, so loops automatically loop the correct number of times. \n> No runtime check is necessary, and it's easier for the code reviewer to \n> visually check the code for correctness.\n> \n\nBut this introduces handy but, strictly speaking, unnecessary overhead as\nwell, meaning, in short; 'D is slower than C, but easier to write code in'.\n\nSo in essence, it's a bit like Python, but a teensy bit faster and a lot\neasier to shoot yourself in the foot with.\n\nWhat was the niche you were going for when you thought up D? It can't have\nbeen systems programming, because *any* extra baggage is baggage one would\nlike to get rid of. If it was application programming I fail to see how one\nmore language would help, as there will be portability problems galore and\nit's still considerably slower to develop in than fe Python, while at the\nsame time being considerably easier to mess up in.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"53026","messageId":"46E339BD.1090201@op5.se","threadId":"9779","inReplyTo":"03C31350-67A5-4B94-A841-5CEB37E78DE8@wincent.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-09T00:09:33Z","receivedAt":"2007-09-09T00:09:33Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Wincent Colaiuta wrote:\n> El 7/9/2007, a las 15:58, Andreas Ericsson escribió:\n> \n>> Yes, but that's what I said in the original email as well. C is just so\n>> much more pleasant to write in that the only place you'd (sanely) use\n>> asm is in exactly these tight loops, where the code is likely to be used\n>> and reused until the algorithm it describes is no longer a viable option\n>> for doing what it was originally designed to do.\n>>\n>> It still proves the point though, as surely as n+1 > n for any value \n>> of n:\n>> Hand-optimized assembly is faster than compiler-optimized C code.\n> \n> In a theoretical ideal world, yes; no one would argue that C is faster \n> than fine-tuned assembly.\n> \n> But in the *real world* rewriting Git in assembly would be like painting \n> a house using a single horse hair instead of a paint brush or roller. \n> Your SHA-1 example is a perfect example of where you benefit from doing \n> a tiny embellished detail using the single hair (assembly) and leave all \n> the rest in C.\n> \n> In the real world and not the theoretical ideal world, it's not just \n> about the diminishing returns you get from writing more and more of a \n> code base in assembly instead of just the performance-critical \n> bottlenecks; it's that you're more likely to make subtle mistakes or \n> even make things slower. GCC does a remarkable job of optimizing in a \n> huge number of use cases, and best of all, it does it for free. Personal \n> opinion, of course, but that's the way I think it is.\n> \n\nThe discussion was theoretical from the beginning. Nobody's arguing that\ngit should be rewritten in asm, and you've been preaching to the choir far\ntoo long now. I'll just drop this thread.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"53029","messageId":"46E33D7F.6040501@op5.se","threadId":"9779","inReplyTo":"fbs8es$1cd$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-09T00:25:35Z","receivedAt":"2007-09-09T00:25:35Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Walter Bright wrote:\n> Andreas Ericsson wrote:\n>> Walter Bright wrote:\n>>> 1) You wind up having to implement the complex, dirty details of \n>>> things yourself. The consequences of this are:\n>>>\n>>>    a) you pick a simpler algorithm (which is likely less efficient - \n>>> I run across bubble sorts all the time in code)\n>>>\n>>>    b) once you implement, tune, and squeeze all the bugs out of those \n>>> complex, dirty details, you're reluctant to change it. You're \n>>> reluctant to try a different algorithm to see if it's faster. I've \n>>> seen this effect a lot in my own code. (I translated a large body of \n>>> my own C++ code that I'd spent months tuning to D, and quickly \n>>> managed to get significantly more speed out of it, because it was \n>>> much simpler to try out different algorithms/data structures.)\n>>>\n>>\n>> I haven't seen this in the development of git, although to be fair, you\n>> didn't mention the number of developers that were simultaneously working\n>> on your project.\n> \n> On my project, one. But I've seen this problem repeatedly in other \n> projects that had multiple developers. For example, I used to use \n> version 1 of an assembler. It was itself written entirely in assembler. \n> It ran *incredibly* slowly on large asm files. But it was written in \n> assembler, which is very fast, so how could that be?\n> \n> Turns out, the symbol table used internally was a linear one. A linear \n> symbol table is easy to implement, but doesn't scale well at all. A \n> linear symbol table was implemented because it was just harder to do \n> more advanced symbol table algorithms in assembler. In this case, a \n> higher level language re-implementation made the assembler much faster, \n> even though that implementation was SLOWER in every detail. It was \n> faster overall, because it was easier to develop faster algorithms.\n> \n\nWell, when the ease-of-coding vs the exec-speed of D vs C is that of\nC vs asm, C will be dead fairly soon. However, since C is so ingrained\nin every language designer's head, I find that unlikely to happen any\ntime soon.\n\n> \n>> Opensource projects with many contributors (git, linux) work differently,\n>> since one or a few among the plethora of authors will almost always be\n>> a true expert at the problem being solved.\n> \n> That is a nice advantage. I don't think many projects can rely on having \n> the best in the business working on them, though <g>.\n> \n\nTrue that. I know a fair few projects that could have done with borrowing\none or two proper gurus, but even opensource programmers are selfish in\nthat we usually only work for something that benefits ourselves.\n\n> \n>> The point is that, given enough developers, *someone* is bound to\n>> find an algorithm that works so well that it's no longer worth\n>> investing time to even discuss if anything else would work better,\n>> either because it moves the performance bottleneck to somewhere else\n>> (where further speedups would no longer produce humanly measurable\n>> improvements), or because the action seems instantanous to the user\n>> (further improvements simply aren't worth it, because no valuable\n>> resource will be saved from it).\n> \n> Sure, but I suggest that few projects reach this maxima.\n\nTrue again, but given what I said above holds, it would be madness to\nmove from the lingua franca of oss hacking to a less common one, as it\nwould mean fewer eyes on the code.\n\n> Case in point: \n> ld, the gnu linker. It's terribly slow. To see how slow it is, compare \n> it to optlink (the 15 years old one that comes with D for Windows). So I \n> don't believe there is anything inherent about linking that should make \n> ld so slow. There's some huge leverage possible in speeding up ld \n> (spreading out that saved time among all the gnu developers).\n> \n> So while git may have reached a maxima in performance, I don't think \n> this principle is applicable in general, even for very widely used open \n> source projects that would profit greatly from improved performance.\n> \n\nInteresting. I recently did a spot of work comparing various string-hashing\nalgorithms. Perhaps I should head over to the ld camp and see if I can help.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"53018","messageId":"46E33E74.20105@op5.se","threadId":"9779","inReplyTo":"a1bbc6950709071517n7e7e99ffl3dd351092e7f19d6@mail.gmail.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-09T00:29:40Z","receivedAt":"2007-09-09T00:29:40Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Dmitry Kakurin wrote:\n> On 9/6/07, Andreas Ericsson <ae@op5.se> wrote:\n>> They already have, but every now and then someone comes along and suggest\n>> a complete rewrite in some other language. So far we've had Java (there's\n>> always one...), Python and now C++.\n> \n> Since this \"complete rewrite\" was mentioned in multiple emails I'd\n> like to rectify that:\n> What I'm offering (for Git) is to use C++ as a \"better C\".\n> Don't change any existing *working* code, but start introducing simple\n> C++ constructs in the new code.\n> Git is simple enough to not require any high-level abstractions. But\n> some utility classes could make code much simpler.\n> \n\nThere are far too many highly valuable contributors that have spoken\nagainst C++ for me to believe that C++ and C will ever co-exist in the\nofficial git repo. Good thing utility classes can be developed on top\nof the existing C-code, but in a separate repo, and packed into a\nlibrary. That way, you get some hacking ground for your beloved C++\ncoderswhile the current git contributors can keep contributing in the\nlanguage they like best.\n\n\n> And BTW, I don't even like C++ that much :-), I just like it much\n> better than C.  I've been saying that C++ is a legacy language for\n> quite some time now. But we will use it for many years to come because\n> the size of this legacy code is huge, so there will be plenty of C++\n> developers available (to contribute to Git :-).\n\nThe C code base is a lot larger and C++ will drop dead pretty fast if it's\never removed or left unmaintained. So much for dinosaurs...\n\n> And C++ is the only way to move with existing C codebase.\n\nComplete and utter BS. It can also stay in C, or get language bindings for\nPython/Perl/PHP/LUA(?)/whatever, or both.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"53022","messageId":"20070909003718.GE13385@artemis.corp","threadId":"9779","inReplyTo":"46E3354A.7030407@op5.se","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-09-09T00:37:18Z","receivedAt":"2007-09-09T00:37:18Z","isPatch":false,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Sat, Sep 08, 2007 at 11:50:34PM +0000, Andreas Ericsson wrote:\n> Walter Bright wrote:\n> > David Kastrup wrote:\n> > > Again, C won't keep you from shooting yourself in the foot.\n> > Right, it won't. A good systems language should do what it can to \n> > prevent the programmer from *inadvertently* shooting himself in the \n> > foot, while allowing him to *deliberately* shoot himself in the foot.\n>\n> No, a good systems language should do exactly what it's told.\n> Supporting tools should tell the programmer if he's risking shooting\n> himself in the foot.\n\n  I beg to differ. I mean, knowing enough of D, I think that what Walter\ntries to say is that a good language should provide constructions that\nwhen used prevent the programmer to shoot himself in both foot at the\nsame time.\n\n  D supports most of the C constructions, so when you want to juggle\nwith razor blades, you're free to do so in D. Though, the language\nprovides idioms that prevent you to write stupid mistakes when used. And\nthat is great.\n\n  D is not Java, you have pointers, you can deal with memory\nexplicitely, you can do whatever you can do in C with no or very little\noverhead. Or you can use higher level D, at your own discretion.\n\n> > > You can tell C compilers to\n> > > check all array accesses, but that is a performance issue.\n> > Runtime checking of arrays in D is a performance issue too, so it is \n> > selectable via a command line switch.\n>\n> Same as in C then.\n\n  HAHAHAHAHAHA. Please, who do you try to convince here ? Except in the\nlocal scope, there is few differences between a foo* and a foo[] in C.\n\n> > But more importantly,\n> > 2) For dynamically sized arrays, the dimension of the array is carried\n> > with the array, so loops automatically loop the correct number of times.\n> > No runtime check is necessary, and it's easier for the code reviewer to\n> > visually check the code for correctness.\n>\n> But this introduces handy but, strictly speaking, unnecessary overhead\n> as well, meaning, in short; 'D is slower than C, but easier to write\n> code in'.\n\n  That's BS. See the strbuf API I've been pushing recently ? It has\nsimplified git's code a lot, because each time git had to deal with a\ngrowing string, it had to deal with at least three variables: the buffer\npointer, the current occupied length, and its allocated size. That was\nthree thing to have variable names for, and to pass to functions.\n\n  Now instead, it's just one struct. D gives that gratis. There is no\nperformance loss because you _need_ to do the same. How do you deal with\ndynamic arrays if you dont't store their lenght and size somewhere ? Or\nare you the kind of programmer that write:\n\n  /* 640kb should be enough for everyone… */\n  some_type *array = malloc(640 << 10);\n\n\n> So in essence, it's a bit like Python, but a teensy bit faster and a\n> lot easier to shoot yourself in the foot with.\n\n> What was the niche you were going for when you thought up D? It can't\n> have been systems programming, because *any* extra baggage is baggage\n> one would like to get rid of. If it was application programming I fail\n> to see how one more language would help, as there will be portability\n> problems galore and it's still considerably slower to develop in than\n> fe Python, while at the same time being considerably easier to mess up\n> in.\n\n  Right now I'm just laughing. There is for sure overheads in some\nplaces of D, but the example you take, and what you try to attack in D\nis definitely not where you lose any kind of performance. You could have\nattacked the GC instead (which is after all an easy classical target).\n\n  Just to evaluate the silliness of your arguments:\n  * http://www.digitalmars.com/d/comparison.html so that you can tell\n    what the D features really are,\n  * http://shootout.alioth.debian.org/gp4/benchmark.php?test=all&lang=all\n    so that you can know what the D performance really is about. Of\n    course those are only micro benchmarks, but well, python is \"just\"\n    15 times slower than D, and D seems to be 10% slower. Well then I'm\n    okay with D, I'm ready to buy 10% faster CPUs and avoid a lot of\n    painful debugging time. In my world, 10% faster hardware is cheaper\n    by many orders of magnitude than skilled programmers, but YMMV.\n\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"53013","messageId":"46E34E19.8050402@op5.se","threadId":"9779","inReplyTo":"20070909003718.GE13385@artemis.corp","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-09T01:36:25Z","receivedAt":"2007-09-09T01:36:25Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Pierre Habouzit wrote:\n> On Sat, Sep 08, 2007 at 11:50:34PM +0000, Andreas Ericsson wrote:\n> \n>>>> You can tell C compilers to\n>>>> check all array accesses, but that is a performance issue.\n>>> Runtime checking of arrays in D is a performance issue too, so it is \n>>> selectable via a command line switch.\n>> Same as in C then.\n> \n>   HAHAHAHAHAHA. Please, who do you try to convince here ? Except in the\n> local scope, there is few differences between a foo* and a foo[] in C.\n> \n\n\"Runtime checking of arrays is a performance issue.\" It's true whether it's\ndone manually by the coder or by the compiler. The difference is that in C,\nyou get to choose where it should be done.\n\n\n>>> But more importantly,\n>>> 2) For dynamically sized arrays, the dimension of the array is carried\n>>> with the array, so loops automatically loop the correct number of times.\n>>> No runtime check is necessary, and it's easier for the code reviewer to\n>>> visually check the code for correctness.\n>> But this introduces handy but, strictly speaking, unnecessary overhead\n>> as well, meaning, in short; 'D is slower than C, but easier to write\n>> code in'.\n> \n>   That's BS. See the strbuf API I've been pushing recently ? It has\n> simplified git's code a lot, because each time git had to deal with a\n> growing string, it had to deal with at least three variables: the buffer\n> pointer, the current occupied length, and its allocated size. That was\n> three thing to have variable names for, and to pass to functions.\n> \n\nYup. I applaud your efforts, but it does come with a slight overhead,\nexcept where it replaces faulty code. In practice, it's probably better\nto use the api for all the string-handling, as none of it is performance-\ncritical.\n\n\n>   Now instead, it's just one struct. D gives that gratis. There is no\n> performance loss because you _need_ to do the same. How do you deal with\n> dynamic arrays if you dont't store their lenght and size somewhere ? Or\n> are you the kind of programmer that write:\n> \n>   /* 640kb should be enough for everyone… */\n>   some_type *array = malloc(640 << 10);\n> \n\nNo, but it would depend on what I am to do with it.\n\n> \n>> So in essence, it's a bit like Python, but a teensy bit faster and a\n>> lot easier to shoot yourself in the foot with.\n> \n>> What was the niche you were going for when you thought up D? It can't\n>> have been systems programming, because *any* extra baggage is baggage\n>> one would like to get rid of. If it was application programming I fail\n>> to see how one more language would help, as there will be portability\n>> problems galore and it's still considerably slower to develop in than\n>> fe Python, while at the same time being considerably easier to mess up\n>> in.\n> \n>   Right now I'm just laughing. There is for sure overheads in some\n> places of D, but the example you take, and what you try to attack in D\n> is definitely not where you lose any kind of performance. You could have\n> attacked the GC instead (which is after all an easy classical target).\n> \n\nI was asking what role D was designed to fill. I didn't mean it as an\nattack, but re-reading what I wrote earlier I see it came off a bit harsh.\n\n\n>   Just to evaluate the silliness of your arguments:\n>   * http://www.digitalmars.com/d/comparison.html so that you can tell\n>     what the D features really are,\n\nYou may notice that the feature-list is being provided by the creators\nand marketeers of the D language. Walter Bright certainly seems like a\nnice enough person, but it's possible it's a tad biased.\n\n\n>   * http://shootout.alioth.debian.org/gp4/benchmark.php?test=all&lang=all\n>     so that you can know what the D performance really is about. Of\n>     course those are only micro benchmarks, but well, python is \"just\"\n>     15 times slower than D, and D seems to be 10% slower.\n\n\nI get it to 7.7xC and 1.2xC, respectively, but whatever. It still means\nperformance-critical apps will be written in C, while\ninsert-script-language-of-choice will still be used for prototyping and\nnot-so performance-critical apps.\n\n\n> Well then I'm\n>     okay with D, I'm ready to buy 10% faster CPUs and avoid a lot of\n>     painful debugging time. In my world, 10% faster hardware is cheaper\n>     by many orders of magnitude than skilled programmers, but YMMV.\n> \n\nI'm curious as to how many fewer bugs D developers write compared to C\nprogrammers. I guess it's hard to do a fair test given the comparatively\nshallow pool of D gurus around, but it'd still be interesting to see a\npractical test. 20% increase in runtime is certainly acceptable for\nnever having to see a bug again, but is it acceptable for 10% fewer bugs?\nOr 20% fewer?\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"53609","messageId":"fcrutu$16f$1@sea.gmane.org","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709070203200.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Steven Burns","fromEmail":"royalstream@hotmail.com","sentAt":"2007-09-19T19:56:51Z","receivedAt":"2007-09-19T19:56:51Z","isPatch":false,"sender":{"key":"royalstream@hotmail.com","avatar":null},"body":"To me, the only thing that C++ has that all other mentioned languages lack \nis the power you get from the templates and generic programming.\nSorting will always be faster if you can call the comparison function \ndirectly without using a function pointer, and the only way you can create a \ngeneric sorting algorithm is that way.\n\nThinking about it with a cold head, most things to hate about C++ are not in \nthe language but in its libraries.\nThe only feature I hate from the language itself is the preprocessor \n(macros), which you get in C too.\n\nAnd maybe I also hate the fact that C++ allows for unexperienced programmers \nto create a bunch of classes and hierarchies that make sense to nobody but \nthem. Or even worse, unexperienced programmers start writing their own \nframeworks, wrapping and re-wrapping, the same good old C function one \nthousand times.\n\nI guess that is why most C++ based projects out there have a strict list of \nrules and conventions, you cannot have a stable project without them.\n\nBut, nothing prevents anybody from programming in C++ the way you describe, \nusing simple and clear core structures with some basic methods that \ncomplement them (not obscure them) and make it easier to write the \nalgorithms.\nSadly, once you start using std::string, their overly complicated and fancy \niostreams, and bulky classes that hide too much from you, I have no other \nchoice than to agree and call the whole thing a mess.\n\nSteven Burns\n\n\n\"Linus Torvalds\" <torvalds@linux-foundation.org> wrote in message \nnews:alpine.LFD.0.999.0709070203200.5626@evo.linux-foundation.org...\n>\n>\n> On Fri, 7 Sep 2007, Linus Torvalds wrote:\n>>\n>> The fact is, git is better than the other SCM's. And good taste (and C) \n>> is\n>> one of the reasons for that.\n>\n> To be very specific:\n> - simple and clear core datastructures, with *very* lean and aggressive\n>   code to manage them that takes the whole approach of \"simplicity over\n>   fancy\" to the extreme.\n> - a willingness to not abstract away the data structures and algorithms,\n>   because those are the *whole*point* of core git.\n>\n> And if you want a fancier language, C++ is absolutely the worst one to\n> choose. If you want real high-level, pick one that has true high-level\n> features like garbage collection or a good system integration, rather than\n> something that lacks both the sparseness and straightforwardness of C,\n> *and* doesn't even have the high-level bindings to important concepts.\n>\n> IOW, C++ is in that inconvenient spot where it doesn't help make things\n> simple enough to be truly usable for prototyping or simple GUI\n> programming, and yet isn't the lean system programming language that C is\n> that actively encourags you to use simple and direct constructs.\n>\n> Linus \n"},{"id":"53651","messageId":"fctuo4$t93$1@sea.gmane.org","threadId":"9779","inReplyTo":"20070907061554.GB30161@thunk.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Steven Burns","fromEmail":"royalstream@hotmail.com","sentAt":"2007-09-20T14:06:03Z","receivedAt":"2007-09-20T14:06:03Z","isPatch":false,"sender":{"key":"royalstream@hotmail.com","avatar":null},"body":"> a = b + \"/share/\" + c + serial_num;\n>\n> where you can have absolutely no idea how many memory allocations are\n> done, due to type coercions, overloaded operators\n\nYou are assuming (incorrectly) everybody will use dumb string classes like \nthat.\n\nIt is very possible to create a string class that instead of allocating all\nthose strings simply concatenates tiny temporary objects and performs one\nsingle operation in the end. Not to mention those temporaries are optimized\naway by any decent compiler and you end up with code that runs at the same\nspeed as your C code.\nI've done it, many other programmers have. As a reference, I'd like to\nmention Matthew Wilson's chapter on efficient string concatenation in his\nbook \"Imperfect C++\". He uses expression templates (that's the technique I\njust described) and gets impressive results.\n\nWith that said, your point is valid. 90% of C++ programmers will use string\nclasses that are very inefficient for concatenation, starting with\nstd::string which I hate for that reason (and many other reasons, e.g. you \nhave to\nresort to Boost for mundane things like trimming)\n\nSteven Burns\n\n\"Theodore Tso\" <tytso@mit.edu> wrote in message \nnews:20070907061554.GB30161@thunk.org...\n> On Thu, Sep 06, 2007 at 08:09:23PM -0700, Dmitry Kakurin wrote:\n>> > Total BS. The string/memory management is not at all relevant. Look at \n>> > the\n>> > code (I bet you didn't). This isn't the important, or complex part.\n>>\n>> Not only have I looked at the code, I've also debugged it quite a bit.\n>> Granted most of my problems had to do with handling paths on Windows\n>> (i.e. string manipulations).\n>\n> I consider string manipulation to be one of the places where C++ is a\n> total disaster.  It's way to easy for idiots to do something like this:\n>\n> a = b + \"/share/\" + c + serial_num;\n>\n> where you can have absolutely no idea how many memory allocations are\n> done, due to type coercions, overloaded operators (good God, you can\n> overload the comma operator in C++!!!), and then when something like\n> that ends up in an inner loop, the result is a disaster from a\n> performance point of view, and it's not even obvious *why*!\n>\n>> My goal is to *use* Git. When something does not work *for me* I want\n>> to be able to fix it (and contribute the fix) in *shortest time\n>> possible* and with *minimal efforts*. As for me it's a diversion from\n>> my main activities.\n>\n> Yes, and if you contribute something the shortest time possible, and\n> it ends up being crap, who gets to rewrite it and fix it?  I've seen\n> too many C++ programs which get this kind of crap added, and it's not\n> noticed right away (because C++ is really good at hiding such\n> performance killers so they are not visible), and then later on, it's\n> even harder to find the performance problems and fix them.\n>\n>> Now, I realize that I'm a very infrequent contributor to Git, but I\n>> want my opinion to be heard.\n>\n> And if git were written in C++, it's precisely the infrequent\n> contributors (who are in a hurry, who only care about the quick hack\n> to get them going, and not about the long-term maintainability and\n> performance of the package) that are be in the position to do the\n> most damage...\n>\n> - Ted \n"},{"id":"53653","messageId":"46F28A10.8040803@op5.se","threadId":"9779","inReplyTo":"fctuo4$t93$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-09-20T14:56:16Z","receivedAt":"2007-09-20T14:56:16Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Steven Burns wrote:\n>> a = b + \"/share/\" + c + serial_num;\n>>\n>> where you can have absolutely no idea how many memory allocations are\n>> done, due to type coercions, overloaded operators\n> \n> You are assuming (incorrectly) everybody will use dumb string classes like \n> that.\n> \n\nNot really. He said \"It's way to easy for idiots to do something like this:\"\njust prior to the line you quoted. I wholeheartedly agree, but in no way\ndoes anyone assume that everybody will use dumb string classes.\n\nI'm sure it's perfectly possible to write properly functioning programs in\nC++. I know I use a few of them myself. That doesn't change the fact that\nit's an idiot-friendly language to write code in that's extremely annoying\nfor competent programmers to fix up later.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"53807","messageId":"fd3h7t$59b$1@sea.gmane.org","threadId":"9779","inReplyTo":"fbr2iv$ugg$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Steven Burns","fromEmail":"royalstream@hotmail.com","sentAt":"2007-09-22T16:52:22Z","receivedAt":"2007-09-22T16:52:22Z","isPatch":false,"sender":{"key":"royalstream@hotmail.com","avatar":null},"body":"Another reason GC is sometimes surprisingly faster is not only you end up \nallocating less times like you mention, but because some GC are compacting \ngarbage collectors and that simplyfies allocations dramatically because \nallocating memory is just increasing a pointer. Compare that to the way most \nC++ heaps get implemented.\nI don't know if that's the case with D's GC though.\n\nI completely understand what you say about the strings and who owns it, I've \nran into the same situation a hundred times, not only with strings but with \nvectors, matrixes, lists, etc.\n\nAfter reading your post, I think I will have to revisit D sometime.\nI read about it a few years ago and I got the impression some syntax \ndecisions had been made to ease the writing of the compiler as opposed to \nfavoring the end user/programmer, but it's been a while and maybe I was too \nquick to judge.\n\nSteven\n\n\n\"Walter Bright\" <boost@digitalmars.com> wrote in message \nnews:fbr2iv$ugg$1@sea.gmane.org...\n> Wincent Colaiuta wrote:\n>> Git is all about speed, and C is the best choice for speed, especially in \n>> context of Git's workload.\n>\n> I can appreciate that. I originally got into writing compilers because my \n> game (Empire) ran too slowly and I thought the existing compilers could be \n> dramatically improved.\n>\n> And technically, yes, you can write code in C that is >= the speed of any \n> other language (other than asm). But practically, this isn't necessarily \n> so, for the following reasons:\n>\n> 1) You wind up having to implement the complex, dirty details of things \n> yourself. The consequences of this are:\n>\n>    a) you pick a simpler algorithm (which is likely less efficient - I run \n> across bubble sorts all the time in code)\n>\n>    b) once you implement, tune, and squeeze all the bugs out of those \n> complex, dirty details, you're reluctant to change it. You're reluctant to \n> try a different algorithm to see if it's faster. I've seen this effect a \n> lot in my own code. (I translated a large body of my own C++ code that I'd \n> spent months tuning to D, and quickly managed to get significantly more \n> speed out of it, because it was much simpler to try out different \n> algorithms/data structures.)\n>\n> 2) Garbage collection has an interesting and counterintuitive consequence. \n> If you compare n malloc/free's with n gcnew/collections, the malloc/free \n> will come out faster, and you conclude that gc is slow. But that misses \n> one huge speed advantage of gc - you can do FAR fewer allocations! For \n> example, I've done a lot of string manipulating programs in C. The basic \n> problem is keeping track of who owns each string. This is done by, when in \n> doubt, make a copy of the string.\n>\n> But if you have gc, you don't worry about who owns the string. You just \n> make another pointer to it. D takes this a step further with the concept \n> of array slicing, where one creates windows on existing arrays, or windows \n> on windows on windows, and no allocations are ever done. It's just pointer \n> fiddling.\n>\n> ------\n> Walter Bright\n> http://www.digitalmars.com  C, C++, D programming language compilers\n> http://www.astoriaseminar.com  Extraordinary C++\n> \n"},{"id":"53918","messageId":"loom.20070924T134013-959@post.gmane.org","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709061839510.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"figo","fromEmail":"rcc_dark@hotmail.com","sentAt":"2007-09-24T13:41:17Z","receivedAt":"2007-09-24T13:41:17Z","isPatch":false,"sender":{"key":"rcc_dark@hotmail.com","avatar":null},"body":"http://www.research.att.com/~bs/applications.html\n\njust as Bjarne once wrote in his TC++PL, its hard to teach an old dog new \ntricks. Its even harder to give quality education about how to use something \nto someone who doesnt want to learn.\n\nyou hate high level, then continue programming operative systems, please NEVER \nDO something else. C++ was designed to give programmers high level tools and \nstill being able to take care about performance.\n\nportability wont be possible after a standard is published and some couple of \nyears given to the compiler developers. C++ had its standard in 1998, and add \ntwo or three years for compiler development = 2002. \"Quite recently\", way more \nrecently that your last use of C++ I can bet.\n"},{"id":"53921","messageId":"86odfstbc6.fsf@lola.quinscape.zz","threadId":"9779","inReplyTo":"loom.20070924T134013-959@post.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-24T13:57:45Z","receivedAt":"2007-09-24T13:57:45Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"figo <rcc_dark@hotmail.com> writes:\n\n> http://www.research.att.com/~bs/applications.html\n>\n> just as Bjarne once wrote in his TC++PL, its hard to teach an old dog new \n> tricks. Its even harder to give quality education about how to use something \n> to someone who doesnt want to learn.\n>\n> you hate high level, then continue programming operative systems,\n> please NEVER DO something else. C++ was designed to give programmers\n> high level tools and still being able to take care about\n> performance.\n>\n> portability wont be possible after a standard is published and some\n>couple of years given to the compiler developers. C++ had its\n>standard in 1998, and add two or three years for compiler development\n>= 2002. \"Quite recently\", way more recently that your last use of C++\n>I can bet.\n\nCare to explain why there are still not two numerical C++ libraries\nwith compatible matrix classes?\n\nWhat use is talking about portability and high level when a basic\ninteroperability feature that has been available since the sixties\n(more than 4 decades ago) in Fortran has not yet managed to make it\ninto C++?  C++ by now more or less offers a (somewhat deficient)\nstandardized way to work with complex numbers, but matrices are still\nnot standardized in any manner, and libraries won't interoperate.\n\nSo C++ should get its head wrapped around the _low_ level problems\nfirst.  It is a bloody shame that it still has not caught up with\nFortran IV (or even Fortran II) with regard to usefulness for\nnumerical libraries.\n\nIt is not a matter of \"hating high level\" to see that C++ is mostly\nfocused about addressing the wrong kinds of problems in the wrong\nways.  The pain/gain ratio is just bad.\n\n-- \nDavid Kastrup\n"},{"id":"54025","messageId":"fdbmve$8lj$1@sea.gmane.org","threadId":"9779","inReplyTo":"86odfstbc6.fsf@lola.quinscape.zz","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Steven Burns","fromEmail":"royalstream@hotmail.com","sentAt":"2007-09-25T19:19:19Z","receivedAt":"2007-09-25T19:19:19Z","isPatch":false,"sender":{"key":"royalstream@hotmail.com","avatar":null},"body":"The C++ community in general suffers a lot from the NIH Syndrome.\nMatrixes, Strings, Vectors, everybody creates their own which are always, or \ncourse, superior to what's already available.\n\nAgain, is not the language's fault, a language is just a language.\nIt's the way it has been driven.\n\nMy two cents.\n\n\n\"David Kastrup\" <dak@gnu.org> wrote in message \nnews:86odfstbc6.fsf@lola.quinscape.zz...\n> figo <rcc_dark@hotmail.com> writes:\n>\n>> http://www.research.att.com/~bs/applications.html\n>>\n>> just as Bjarne once wrote in his TC++PL, its hard to teach an old dog new\n>> tricks. Its even harder to give quality education about how to use \n>> something\n>> to someone who doesnt want to learn.\n>>\n>> you hate high level, then continue programming operative systems,\n>> please NEVER DO something else. C++ was designed to give programmers\n>> high level tools and still being able to take care about\n>> performance.\n>>\n>> portability wont be possible after a standard is published and some\n>>couple of years given to the compiler developers. C++ had its\n>>standard in 1998, and add two or three years for compiler development\n>>= 2002. \"Quite recently\", way more recently that your last use of C++\n>>I can bet.\n>\n> Care to explain why there are still not two numerical C++ libraries\n> with compatible matrix classes?\n>\n> What use is talking about portability and high level when a basic\n> interoperability feature that has been available since the sixties\n> (more than 4 decades ago) in Fortran has not yet managed to make it\n> into C++?  C++ by now more or less offers a (somewhat deficient)\n> standardized way to work with complex numbers, but matrices are still\n> not standardized in any manner, and libraries won't interoperate.\n>\n> So C++ should get its head wrapped around the _low_ level problems\n> first.  It is a bloody shame that it still has not caught up with\n> Fortran IV (or even Fortran II) with regard to usefulness for\n> numerical libraries.\n>\n> It is not a matter of \"hating high level\" to see that C++ is mostly\n> focused about addressing the wrong kinds of problems in the wrong\n> ways.  The pain/gain ratio is just bad.\n>\n> -- \n> David Kastrup\n> \n"},{"id":"54035","messageId":"86wsue34gc.fsf@lola.quinscape.zz","threadId":"9779","inReplyTo":"fdbmve$8lj$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-09-25T19:55:31Z","receivedAt":"2007-09-25T19:55:31Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Steven Burns\" <royalstream@hotmail.com> writes:\n\n> The C++ community in general suffers a lot from the NIH Syndrome.\n> Matrixes, Strings, Vectors, everybody creates their own which are always, or \n> course, superior to what's already available.\n>\n> Again, is not the language's fault, a language is just a language.\n> It's the way it has been driven.\n\nHaving loose wires instead of a brake pedal in a car because the user\nmight prefer to brake with his teeth or by wiggling his backside or\nbuilding any other contraption of his own invention is a design\nmistake.  Especially when we are talking about public transportation\nwith changing drivers.\n\nMaking a language huge and bloated in order to be able to use the\nlanguage itself for defining a set of basic data types is just\nmasturbation.  C++ has the most complicated set of implicit\nconversions from any language in the world, and what for?  It is\nmodeled for being able to create a user-defined \"complex\" type which\nbehaves almost as well as Fortran's.  Too bad that this mostly means\neverybody will define his own type (well, at least we have seen two or\nthree different library \"standards\" by now), and that the implicit\nconversion rules and chains are appallingly wrong for a number of\nother possible user-defined arithmetic types.\n\n-- \nDavid Kastrup\n"},{"id":"123458","messageId":"loom.20090917T180857-851@post.gmane.org","threadId":"9779","inReplyTo":"fbs8es$1cd$1@sea.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Bernd Jendrissek","fromEmail":"bernd.jendrissek@gmail.com","sentAt":"2009-09-17T16:23:44Z","receivedAt":"2009-09-17T16:23:44Z","isPatch":false,"sender":{"key":"bernd.jendrissek@gmail.com","avatar":null},"body":"Walter Bright <boost <at> digitalmars.com> writes:\n> Sure, but I suggest that few projects reach this maxima. Case in point: \n> ld, the gnu linker. It's terribly slow. To see how slow it is, compare \n> it to optlink (the 15 years old one that comes with D for Windows). So I \n> don't believe there is anything inherent about linking that should make \n> ld so slow. There's some huge leverage possible in speeding up ld \n> (spreading out that saved time among all the gnu developers).\n\nhttp://en.wikipedia.org/wiki/Gold_(linker)\n\nNote that gold is written in C++; the wikipedia quasi-stub article doesn't make\nthis clear.  Normally that wouldn't be relevant, but in this branch of the\nthread it is.  Its C++-ness seems to be making an argument, but I don't know on\nwhich side!\n"},{"id":"143499","messageId":"loom.20100610T204637-685@post.gmane.org","threadId":"9779","inReplyTo":"4AFD7EAD1AAC4E54A416BA3F6E6A9E52@ntdev.corp.microsoft.com","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Ian Molton","fromEmail":"ian.molton@collabora.co.uk","sentAt":"2010-06-10T19:12:36Z","receivedAt":"2010-06-10T19:12:36Z","isPatch":false,"sender":{"key":"ian.molton@collabora.co.uk","avatar":null},"body":"Dmitry Kakurin <dmitry.kakurin <at> gmail.com> writes:\n\n> \n> [ snip ]\n> \n> When I first looked at Git source code two things struck me as odd:\n> 1. Pure C as opposed to C++. No idea why. Please don't talk about \n> portability, it's BS.\n\nWord to the wise... you effectively just told one of *the* best known\nprogrammers of all time that they are talking BS... nice one. Hope you've got\nsome flameproof undies. Whats that? no? ah well...\n\nI smell a troll, but since everyone else has had a go...\n\nHeres some comments picked out from the thread, in no particular order...\n\n> > You have a very strange way of not meaning to start a C vs. C++ war.\n\n> I honestly didn't. I didn't even think it's possible. In the\n> environment of mainstream commercial software development the last war\n> on this subj was over 8-10 years ago.\n\nReally? I dont know what planet you're from, but this 'war' has been raging for\ndecades, and will probably continue until one side or the other gets round to\nusing tactical nukes.\n\nAnd besides, this *isn't* the commercial (closed) software world - we've moved\non. We no longer depend on closed companies handing out features like orphans in\nthe Victorian times...\n\n>>> [bitfields in D]\n\n>>   Really, I feel this is a big lack, for a language that aims at\n>> simplicity, conciseness _and_ correctness.\n>>\n>>   OK, maybe I'm biased, I work with networks protocols all day long, so\n>> I often need bitfields, but still, a lot of people deal with network\n>> protocols, it's not a niche.\n>\n> And strictly speaking, C bitfields are completely useless for that\n> purpose since the compiler is free to use whatever method he wants for\n> allocating bit fields.  So if you want to write a portable program,\n> you are back to making the masks yourself.\n\nSadly. Thats always been one of the things I found annoying in C. There are\ntimes when you want access to the types the hardware itself uses, and there are\ntimes when you want to know your int is 32 bits long, and there isnt really a\nstandardised way of doing that. Of course, its worked around in practice, but it\nall seems so unnecessary.\n\n> in the *real world* rewriting Git in assembly would be like  \n> painting a house using a single horse hair instead of a paint brush  \n> or roller. Your SHA-1 example is a perfect example of where you  \n> benefit from doing a tiny embellished detail using the single hair  \n> (assembly) and leave all the rest in C.\n\nThe above comment is pure epic win :-)\n\nOn another note, some people talked about code reuse...\n\nIMHO Sourcecode reuse is something of a myth in any language. Sure, some small\nalgorithms get reused, but thats really not a language dependent characteristic.\nAs soon as you build something much bigger than an algorithm, it starts to need\nan interface, and at that point you may as well turn it into a library. Thats\nwhere the REAL code reuse happens. And as it happens at runtime, its good for\nusers - bugfixes help everyone.\n\nOn to language choice...\n\nI have NEVER understood why people seem to think theres some kind of hierarchy\nin either ease of coding or speed. You see it all the time, people think that:\n\nassembler is faster than c is faster than c++ is faster than perl etc.\n\nWHY? I've seen some truely braindamaged assembler that could be outperformed by\nBASIC on a BBC micro. I've seen 'handcrafted' C and C++ that looked like it was\nwritten during a skydive whilst on crack.\n\nlanguages are *tools*. Pick the most appropriate. Use two. Embrace the power of\nand...\n\nLinux make good use of C and assembler, both compiled/assembled seperately and\ninline. Some stuff like accessing weird registers with oddball opcodes is\nactually impossible under C. But (say) write a filesystem in assembler? no\nthanks! (not that it hasn't been done, but for the love of god, why?)\n\nSo, anyway, why do these kind of threads never go away? because opinions are\nlike arseholes. Everyones got one. As you grow older, you learn stuff. You\nhopefully dont repeat the mistakes of ones youth. (theres at least one\nill-conceived C string library out there which I'm embarrased to admit is my\nfault (hopefully it'll never leave the company I was at when I wrote it...).\nThese threads are where the n00bs meet the pros. Usually, the n00bs just need to\nsuck it up and admit it when they've been dumb. Its a very rare day when\nsomething truely radical comes along, and its even rarer when its born of total\ninexperience.\n\nNothing to see here...\n\nAll the best,\n\n-Ian\n\nPS. ironically, in order to post this, gmane required me to enter a word. That\nword was \"restraint\". Gotta love karma.\n"},{"id":"143500","messageId":"m3mxv1pxli.fsf@localhost.localdomain","threadId":"9779","inReplyTo":"loom.20100610T204637-685@post.gmane.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2010-06-11T12:23:55Z","receivedAt":"2010-06-11T12:23:55Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Ian Molton <ian.molton@collabora.co.uk> writes:\n> Dmitry Kakurin <dmitry.kakurin <at> gmail.com> writes:\n>\n> > [ snip ]\n> > \n> > When I first looked at Git source code two things struck me as odd:\n> > 1. Pure C as opposed to C++. No idea why. Please don't talk about \n> > portability, it's BS.\n\nNo gain from C++.\n\nAlso, I don't know when Dmitri written his post, but git uses its own\nstring manipulation mini-library, named strbuf, at least since end of 2007\n(Documentation/technical/api-strbuf.txt was added as stub on 2007-11-24).\n \n> > in the *real world* rewriting Git in assembly would be like  \n> > painting a house using a single horse hair instead of a paint brush  \n> > or roller. Your SHA-1 example is a perfect example of where you  \n> > benefit from doing a tiny embellished detail using the single hair  \n> > (assembly) and leave all the rest in C.\n\nSidenote: block-sha1 implementation is C plus smidgeon of assembly via\n'asm'.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"143508","messageId":"AANLkTikvBg1IEauZIzZP4XINqUjrb3F3ill5vP9zqIpR@mail.gmail.com","threadId":"9779","inReplyTo":"m3mxv1pxli.fsf@localhost.localdomain","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Dario Rodriguez","fromEmail":"soft.d4rio@gmail.com","sentAt":"2010-06-11T13:33:58Z","receivedAt":"2010-06-11T13:33:58Z","isPatch":false,"sender":{"key":"soft.d4rio@gmail.com","avatar":"https://gravatar.com/avatar/9bca8035ad520a6fc0eb66e5126adbe45992c8388c05d1f7952bd4793eb82c22?d=mp&s=160"},"body":"On Fri, Jun 11, 2010 at 9:23 AM, Jakub Narebski <jnareb@gmail.com> wrote:\n> Also, I don't know when Dmitri written his post, but git uses its own\n> string manipulation mini-library, named strbuf, at least since end of 2007\n> (Documentation/technical/api-strbuf.txt was added as stub on 2007-11-24).\n\nAn interesting point, Dmitri written this post on September 2007...\nwhy the flamewar continuation? just curiosity...\n\nNow, as a resume of flames and concepts, the 'better' things in the\nstring library (also the first excuse for C++ apologists, as a\nrepetitive piece of youknowwhat) are pure algorithms and data\nstructures, easy to code in C. The goal isn't in the OO design or\nabstraction itself, or in the language...\n\nPersonally, I use C++ almost every day at work and I found it stupid.\nI love pure C.\n\nCheers,\nDario (argentina)\n"},{"id":"191920","messageId":"loom.20120522T201747-568@post.gmane.org","threadId":"9779","inReplyTo":"alpine.LFD.0.999.0709061839510.5626@evo.linux-foundation.org","subject":"Re: [RFC] Convert builin-mailinfo.c to use The Better String Library.","fromName":"Syed M Raihan","fromEmail":"syed_raihan@yahoo.com","sentAt":"2012-05-22T18:30:13Z","receivedAt":"2012-05-22T18:30:13Z","isPatch":false,"sender":{"key":"syed_raihan@yahoo.com","avatar":null},"body":"Linus Torvalds <torvalds <at> linux-foundation.org> writes:\n\n> \n> \n> On Wed, 5 Sep 2007, Dmitry Kakurin wrote:\n> > \n> > When I first looked at Git source code two things struck me as odd:\n> > 1. Pure C as opposed to C++. No idea why. Please don't talk about port,\n> > it's BS.\n> \n> *YOU* are full of bullshit.\n> \n> C++ is a horrible language. It's made more horrible by the fact that a lot \n> of substandard programmers use it, to the point where it's much much \n> easier to generate total and utter crap with it. Quite frankly, even if \n> the choice of C were to do *nothing* but keep the C++ programmers out, \n> that in itself would be a huge reason to use C.\n> \t\t\tLinus\n> \n\nC++ has one weakness that is ABI compatibility among compilers.\nOther than that Object Model does not make things horrible. \nI have seen 15 years old C++ application library which still \nuses old implementation to implement new enhancement/features \nthat just works seamlessly and even an old dog can learn \nthis old library written in C++.\n\nI have seen C programmers constantly trying how they can mimic \npoly-morphism, inheritance and encapsulation in their C.\nThis is just *BS* - if you dont want C++ then dont use poly-morphism,\ninheritance and encapsulation in your C code!\nOr else just use C++ or Jave or C#.\n\nRegards\nPlease note: I am really a Fan of Linus Torvalds since ever :)\nSyed Raihan\n"}]}