{"thread":{"id":"17092","subject":"[PATCH 0/4] refactor the --color-words to make it more hackable","startedAt":"2009-01-11T19:58:47Z","lastAt":"2009-01-21T19:37:06Z","messageCount":109,"participants":["Johannes Schindelin","Thomas Rast","Junio C Hamano","Jakub Narebski","Santi Béjar","Teemu Likonen","Boyd Stephen Smith Jr.","Markus Heidelberg"],"isPatch":true,"patchVersion":1,"patchTotal":4},"messages":[{"id":"99984","messageId":"alpine.DEB.1.00.0901112057300.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":null,"subject":"[PATCH 0/4] refactor the --color-words to make it more hackable","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-11T19:58:47Z","receivedAt":"2009-01-11T19:58:47Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nSo the total change is pretty large, I have to admit.\n\nBut at least _I_ think it is easy to follow, and it actually makes the code\nmore readable/hackable.  Correct me if I'm wrong.\n\nThe basic idea is to decouple the original text from the text that is\npassed to libxdiff to find the word differences.\n\nTo that end, the words of the pre and post texts are put into two lists that\nare fed to libxdiff.  While the words are extracted, an array is created which\ncontains pointers back to the word boundaries in the original text.\n\nTo make the transition as easy to understand as possible, the code is first\nrefactored without actually changing what makes a word boundary.\n\nJohannes Schindelin (4):\n  Add color_fwrite(), a function coloring each line individually\n  color-words: refactor word splitting and use ALLOC_GROW()\n  color-words: refactor to allow for 0-character word boundaries\n  color-words: take an optional regular expression describing words\n\n color.c |   24 ++++++++\n color.h |    1 +\n diff.c  |  185 +++++++++++++++++++++++++++++++++++++++------------------------\n diff.h  |    1 +\n 4 files changed, 141 insertions(+), 70 deletions(-)\n"},{"id":"99985","messageId":"alpine.DEB.1.00.0901112058570.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901112057300.3586@pacific.mpi-cbg.de","subject":"[PATCH 1/4] Add color_fwrite(), a function coloring each line individually","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-11T19:59:08Z","receivedAt":"2009-01-11T19:59:08Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nWe have to set the color before every line and reset it before every\nnewline.  Add a function color_fwrite() which does that for us.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n color.c |   24 ++++++++++++++++++++++++\n color.h |    1 +\n 2 files changed, 25 insertions(+), 0 deletions(-)\n\ndiff --git a/color.c b/color.c\nindex fc0b72a..bff24ac 100644\n--- a/color.c\n+++ b/color.c\n@@ -191,3 +191,27 @@ int color_fprintf_ln(FILE *fp, const char *color, const char *fmt, ...)\n \tva_end(args);\n \treturn r;\n }\n+\n+/*\n+ * This function splits the buffer by newlines and colors the lines individually.\n+ */\n+void color_fwrite(FILE *f, const char *color, size_t count, const char *buf)\n+{\n+\tif (!*color) {\n+\t\tfwrite(buf, count, 1, f);\n+\t\treturn;\n+\t}\n+\twhile (count) {\n+\t\tchar *p = memchr(buf, '\\n', count);\n+\t\tfputs(color, f);\n+\t\tfwrite(buf, p ? p - buf : count, 1, f);\n+\t\tfputs(COLOR_RESET, f);\n+\t\tif (!p)\n+\t\t\treturn;\n+\t\tfputc('\\n', f);\n+\t\tcount -= p + 1 - buf;\n+\t\tbuf = p + 1;\n+\t}\n+}\n+\n+\ndiff --git a/color.h b/color.h\nindex 6cf5c88..9fb58f5 100644\n--- a/color.h\n+++ b/color.h\n@@ -19,5 +19,6 @@ int git_config_colorbool(const char *var, const char *value, int stdout_is_tty);\n void color_parse(const char *var, const char *value, char *dst);\n int color_fprintf(FILE *fp, const char *color, const char *fmt, ...);\n int color_fprintf_ln(FILE *fp, const char *color, const char *fmt, ...);\n+void color_fwrite(FILE *f, const char *color, size_t count, const char *buf);\n \n #endif /* COLOR_H */\n-- \n1.6.1.186.g48f3bc4\n"},{"id":"99986","messageId":"alpine.DEB.1.00.0901112059160.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901112057300.3586@pacific.mpi-cbg.de","subject":"[PATCH 2/4] color-words: refactor word splitting and use ALLOC_GROW()","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-11T19:59:30Z","receivedAt":"2009-01-11T19:59:30Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nWord splitting is now performed by the function diff_words_fill(),\navoiding having the same code twice.\n\nIn the same spirit, avoid duplicating the code of ALLOC_GROW().\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n diff.c |   40 +++++++++++++++++++---------------------\n 1 files changed, 19 insertions(+), 21 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex f67e0b2..6d87ea5 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -326,10 +326,7 @@ struct diff_words_buffer {\n static void diff_words_append(char *line, unsigned long len,\n \t\tstruct diff_words_buffer *buffer)\n {\n-\tif (buffer->text.size + len > buffer->alloc) {\n-\t\tbuffer->alloc = (buffer->text.size + len) * 3 / 2;\n-\t\tbuffer->text.ptr = xrealloc(buffer->text.ptr, buffer->alloc);\n-\t}\n+\tALLOC_GROW(buffer->text.ptr, buffer->text.size + len, buffer->alloc);\n \tline++;\n \tlen--;\n \tmemcpy(buffer->text.ptr + buffer->text.size, line, len);\n@@ -398,6 +395,22 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \t}\n }\n \n+/*\n+ * This function splits the words in buffer->text, and stores the list with\n+ * newline separator into out.\n+ */\n+static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n+{\n+\tint i;\n+\tout->size = buffer->text.size;\n+\tout->ptr = xmalloc(out->size);\n+\tmemcpy(out->ptr, buffer->text.ptr, out->size);\n+\tfor (i = 0; i < out->size; i++)\n+\t\tif (isspace(out->ptr[i]))\n+\t\t\tout->ptr[i] = '\\n';\n+\tbuffer->current = 0;\n+}\n+\n /* this executes the word diff on the accumulated buffers */\n static void diff_words_show(struct diff_words_data *diff_words)\n {\n@@ -405,26 +418,11 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \txdemitconf_t xecfg;\n \txdemitcb_t ecb;\n \tmmfile_t minus, plus;\n-\tint i;\n \n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n-\tminus.size = diff_words->minus.text.size;\n-\tminus.ptr = xmalloc(minus.size);\n-\tmemcpy(minus.ptr, diff_words->minus.text.ptr, minus.size);\n-\tfor (i = 0; i < minus.size; i++)\n-\t\tif (isspace(minus.ptr[i]))\n-\t\t\tminus.ptr[i] = '\\n';\n-\tdiff_words->minus.current = 0;\n-\n-\tplus.size = diff_words->plus.text.size;\n-\tplus.ptr = xmalloc(plus.size);\n-\tmemcpy(plus.ptr, diff_words->plus.text.ptr, plus.size);\n-\tfor (i = 0; i < plus.size; i++)\n-\t\tif (isspace(plus.ptr[i]))\n-\t\t\tplus.ptr[i] = '\\n';\n-\tdiff_words->plus.current = 0;\n-\n+\tdiff_words_fill(&diff_words->minus, &minus);\n+\tdiff_words_fill(&diff_words->plus, &plus);\n \txpp.flags = XDF_NEED_MINIMAL;\n \txecfg.ctxlen = diff_words->minus.alloc + diff_words->plus.alloc;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n-- \n1.6.1.186.g48f3bc4\n"},{"id":"99987","messageId":"alpine.DEB.1.00.0901112059340.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901112057300.3586@pacific.mpi-cbg.de","subject":"[PATCH 3/4] color-words: refactor to allow for 0-character word boundaries","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-11T19:59:54Z","receivedAt":"2009-01-11T19:59:54Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nUp until now, the color-words code assumed that word boundaries are\nidentical to white space characters.\n\nTherefore, it could get away with a very simple scheme: it copied the\nhunks, substituted newlines for each white space character, called\nlibxdiff with the processed text, but then identified the text to\nprint out by the offsets (which agreed since the original text had the\nsame length).\n\nThis code was ugly, for a number of reasons:\n\n- it was impossible to introduce 0-character word boundaries,\n\n- we had to print everything word by word, and\n\n- the code needed extra special handling of newlines in the removed part.\n\nFix all of these issues by processing the text such that\n\n- we build word lists, separated by newlines,\n\n- we remember the original offsets for every word, and\n\n- after calling libxdiff on the wordlists, we parse the hunk headers, and\n  find the corresponding offsets, and then\n\n- we print the removed/added parts in one go.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n diff.c |  150 +++++++++++++++++++++++++++++++++++-----------------------------\n 1 files changed, 82 insertions(+), 68 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex 6d87ea5..2a3d301 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -319,8 +319,10 @@ static int fill_mmfile(mmfile_t *mf, struct diff_filespec *one)\n struct diff_words_buffer {\n \tmmfile_t text;\n \tlong alloc;\n-\tlong current; /* output pointer */\n-\tint suppressed_newline;\n+\tstruct diff_words_orig {\n+\t\tconst char *begin, *end;\n+\t} *orig;\n+\tint orig_nr, orig_alloc;\n };\n \n static void diff_words_append(char *line, unsigned long len,\n@@ -335,80 +337,79 @@ static void diff_words_append(char *line, unsigned long len,\n \n struct diff_words_data {\n \tstruct diff_words_buffer minus, plus;\n+\tconst char *current_plus;\n \tFILE *file;\n };\n \n-static void print_word(FILE *file, struct diff_words_buffer *buffer, int len, int color,\n-\t\tint suppress_newline)\n+static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n {\n-\tconst char *ptr;\n-\tint eol = 0;\n+\tstruct diff_words_data *diff_words = priv;\n+\tint minus_first, minus_len, plus_first, plus_len;\n+\tconst char *minus_begin, *minus_end, *plus_begin, *plus_end;\n \n-\tif (len == 0)\n+\tif (line[0] != '@' || parse_hunk_header(line, len,\n+\t\t\t&minus_first, &minus_len, &plus_first, &plus_len))\n \t\treturn;\n \n-\tptr  = buffer->text.ptr + buffer->current;\n-\tbuffer->current += len;\n-\n-\tif (ptr[len - 1] == '\\n') {\n-\t\teol = 1;\n-\t\tlen--;\n-\t}\n-\n-\tfputs(diff_get_color(1, color), file);\n-\tfwrite(ptr, len, 1, file);\n-\tfputs(diff_get_color(1, DIFF_RESET), file);\n-\n-\tif (eol) {\n-\t\tif (suppress_newline)\n-\t\t\tbuffer->suppressed_newline = 1;\n-\t\telse\n-\t\t\tputc('\\n', file);\n-\t}\n-}\n-\n-static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n-{\n-\tstruct diff_words_data *diff_words = priv;\n+\tminus_begin = diff_words->minus.orig[minus_first].begin;\n+\tminus_end = minus_len == 0 ? minus_begin :\n+\t\tdiff_words->minus.orig[minus_first + minus_len - 1].end;\n+\tplus_begin = diff_words->plus.orig[plus_first].begin;\n+\tplus_end = plus_len == 0 ? plus_begin :\n+\t\tdiff_words->plus.orig[plus_first + plus_len - 1].end;\n \n-\tif (diff_words->minus.suppressed_newline) {\n-\t\tif (line[0] != '+')\n-\t\t\tputc('\\n', diff_words->file);\n-\t\tdiff_words->minus.suppressed_newline = 0;\n-\t}\n+\tif (diff_words->current_plus != plus_begin)\n+\t\tfwrite(diff_words->current_plus,\n+\t\t\t\tplus_begin - diff_words->current_plus, 1,\n+\t\t\t\tdiff_words->file);\n+\tif (minus_begin != minus_end)\n+\t\tcolor_fwrite(diff_words->file, diff_get_color(1, DIFF_FILE_OLD),\n+\t\t\t\tminus_end - minus_begin, minus_begin);\n+\tif (plus_begin != plus_end)\n+\t\tcolor_fwrite(diff_words->file, diff_get_color(1, DIFF_FILE_NEW),\n+\t\t\t\tplus_end - plus_begin, plus_begin);\n \n-\tlen--;\n-\tswitch (line[0]) {\n-\t\tcase '-':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->minus, len, DIFF_FILE_OLD, 1);\n-\t\t\tbreak;\n-\t\tcase '+':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->plus, len, DIFF_FILE_NEW, 0);\n-\t\t\tbreak;\n-\t\tcase ' ':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->plus, len, DIFF_PLAIN, 0);\n-\t\t\tdiff_words->minus.current += len;\n-\t\t\tbreak;\n-\t}\n+\tdiff_words->current_plus = plus_end;\n }\n \n /*\n- * This function splits the words in buffer->text, and stores the list with\n- * newline separator into out.\n+ * This function splits the words in buffer->text, stores the list with\n+ * newline separator into out, and saves the offsets of the original words\n+ * in buffer->orig.\n  */\n static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n {\n-\tint i;\n-\tout->size = buffer->text.size;\n-\tout->ptr = xmalloc(out->size);\n-\tmemcpy(out->ptr, buffer->text.ptr, out->size);\n-\tfor (i = 0; i < out->size; i++)\n-\t\tif (isspace(out->ptr[i]))\n-\t\t\tout->ptr[i] = '\\n';\n-\tbuffer->current = 0;\n+\tint i, j;\n+\n+\tout->size = 0;\n+\tout->ptr = xmalloc(buffer->text.size);\n+\n+\t/* fake an empty \"0th\" word */\n+\tALLOC_GROW(buffer->orig, 1, buffer->orig_alloc);\n+\tbuffer->orig[0].begin = buffer->orig[0].end = buffer->text.ptr;\n+\tbuffer->orig_nr = 1;\n+\n+\tfor (i = 0; i < buffer->text.size; i++) {\n+\t\tif (isspace(buffer->text.ptr[i]))\n+\t\t\tcontinue;\n+\t\tfor (j = i + 1; j < buffer->text.size &&\n+\t\t\t\t!isspace(buffer->text.ptr[j]); j++)\n+\t\t\t; /* find the end of the word */\n+\n+\t\t/* store original boundaries */\n+\t\tALLOC_GROW(buffer->orig, buffer->orig_nr + 1,\n+\t\t\t\tbuffer->orig_alloc);\n+\t\tbuffer->orig[buffer->orig_nr].begin = buffer->text.ptr + i;\n+\t\tbuffer->orig[buffer->orig_nr].end = buffer->text.ptr + j;\n+\t\tbuffer->orig_nr++;\n+\n+\t\t/* store one word */\n+\t\tmemcpy(out->ptr + out->size, buffer->text.ptr + i, j - i);\n+\t\tout->ptr[out->size + j - i] = '\\n';\n+\t\tout->size += j - i + 1;\n+\n+\t\ti = j - 1;\n+\t}\n }\n \n /* this executes the word diff on the accumulated buffers */\n@@ -419,22 +420,33 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \txdemitcb_t ecb;\n \tmmfile_t minus, plus;\n \n+\t/* special case: only removal */\n+\tif (!diff_words->plus.text.size) {\n+\t\tcolor_fwrite(diff_words->file, diff_get_color(1, DIFF_FILE_OLD),\n+\t\t\tdiff_words->minus.text.size, diff_words->minus.text.ptr);\n+\t\tdiff_words->minus.text.size = 0;\n+\t\treturn;\n+\t}\n+\n+\tdiff_words->current_plus = diff_words->plus.text.ptr;\n+\n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n \tdiff_words_fill(&diff_words->minus, &minus);\n \tdiff_words_fill(&diff_words->plus, &plus);\n \txpp.flags = XDF_NEED_MINIMAL;\n-\txecfg.ctxlen = diff_words->minus.alloc + diff_words->plus.alloc;\n+\txecfg.ctxlen = 0;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n \t\t      &xpp, &xecfg, &ecb);\n \tfree(minus.ptr);\n \tfree(plus.ptr);\n+\tif (diff_words->current_plus != diff_words->plus.text.ptr +\n+\t\t\tdiff_words->plus.text.size)\n+\t\tfwrite(diff_words->current_plus,\n+\t\t\tdiff_words->plus.text.ptr + diff_words->plus.text.size\n+\t\t\t- diff_words->current_plus, 1,\n+\t\t\tdiff_words->file);\n \tdiff_words->minus.text.size = diff_words->plus.text.size = 0;\n-\n-\tif (diff_words->minus.suppressed_newline) {\n-\t\tputc('\\n', diff_words->file);\n-\t\tdiff_words->minus.suppressed_newline = 0;\n-\t}\n }\n \n typedef unsigned long (*sane_truncate_fn)(char *line, unsigned long len);\n@@ -458,7 +470,9 @@ static void free_diff_words_data(struct emit_callback *ecbdata)\n \t\t\tdiff_words_show(ecbdata->diff_words);\n \n \t\tfree (ecbdata->diff_words->minus.text.ptr);\n+\t\tfree (ecbdata->diff_words->minus.orig);\n \t\tfree (ecbdata->diff_words->plus.text.ptr);\n+\t\tfree (ecbdata->diff_words->plus.orig);\n \t\tfree(ecbdata->diff_words);\n \t\tecbdata->diff_words = NULL;\n \t}\n-- \n1.6.1.186.g48f3bc4\n"},{"id":"99988","messageId":"alpine.DEB.1.00.0901112100050.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901112057300.3586@pacific.mpi-cbg.de","subject":"[PATCH 4/4] color-words: take an optional regular expression describing words","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-11T20:00:58Z","receivedAt":"2009-01-11T20:00:58Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nIn some applications, words are not delimited by white space.  To\nallow for that, you can specify a regular expression describing\nwhat makes a word with\n\n\tgit diff --color-words='^[A-Za-z0-9]*'\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n diff.c |   49 +++++++++++++++++++++++++++++++++++++++++--------\n diff.h |    1 +\n 2 files changed, 42 insertions(+), 8 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex 2a3d301..d6bba72 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -333,12 +333,14 @@ static void diff_words_append(char *line, unsigned long len,\n \tlen--;\n \tmemcpy(buffer->text.ptr + buffer->text.size, line, len);\n \tbuffer->text.size += len;\n+\tbuffer->text.ptr[buffer->text.size] = '\\0';\n }\n \n struct diff_words_data {\n \tstruct diff_words_buffer minus, plus;\n \tconst char *current_plus;\n \tFILE *file;\n+\tregex_t *word_regex;\n };\n \n static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n@@ -372,17 +374,36 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \tdiff_words->current_plus = plus_end;\n }\n \n+static int find_word_boundary(mmfile_t *buffer, int i, regex_t *word_regex)\n+{\n+\tif (i >= buffer->size)\n+\t\treturn i;\n+\n+\tif (word_regex) {\n+\t\tregmatch_t match[1];\n+\t\tif (!regexec(word_regex, buffer->ptr + i, 1, match, 0))\n+\t\t\ti += match[0].rm_eo;\n+\t}\n+\telse\n+\t\twhile (i < buffer->size && !isspace(buffer->ptr[i]))\n+\t\t\ti++;\n+\n+\treturn i;\n+}\n+\n /*\n  * This function splits the words in buffer->text, stores the list with\n  * newline separator into out, and saves the offsets of the original words\n  * in buffer->orig.\n  */\n-static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n+static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out,\n+\t\tregex_t *word_regex)\n {\n \tint i, j;\n+\tlong alloc = 0;\n \n \tout->size = 0;\n-\tout->ptr = xmalloc(buffer->text.size);\n+\tout->ptr = NULL;\n \n \t/* fake an empty \"0th\" word */\n \tALLOC_GROW(buffer->orig, 1, buffer->orig_alloc);\n@@ -390,11 +411,9 @@ static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n \tbuffer->orig_nr = 1;\n \n \tfor (i = 0; i < buffer->text.size; i++) {\n-\t\tif (isspace(buffer->text.ptr[i]))\n+\t\tj = find_word_boundary(&buffer->text, i, word_regex);\n+\t\tif (i == j)\n \t\t\tcontinue;\n-\t\tfor (j = i + 1; j < buffer->text.size &&\n-\t\t\t\t!isspace(buffer->text.ptr[j]); j++)\n-\t\t\t; /* find the end of the word */\n \n \t\t/* store original boundaries */\n \t\tALLOC_GROW(buffer->orig, buffer->orig_nr + 1,\n@@ -404,6 +423,7 @@ static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n \t\tbuffer->orig_nr++;\n \n \t\t/* store one word */\n+\t\tALLOC_GROW(out->ptr, out->size + j - i + 1, alloc);\n \t\tmemcpy(out->ptr + out->size, buffer->text.ptr + i, j - i);\n \t\tout->ptr[out->size + j - i] = '\\n';\n \t\tout->size += j - i + 1;\n@@ -432,8 +452,8 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n-\tdiff_words_fill(&diff_words->minus, &minus);\n-\tdiff_words_fill(&diff_words->plus, &plus);\n+\tdiff_words_fill(&diff_words->minus, &minus, diff_words->word_regex);\n+\tdiff_words_fill(&diff_words->plus, &plus, diff_words->word_regex);\n \txpp.flags = XDF_NEED_MINIMAL;\n \txecfg.ctxlen = 0;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n@@ -473,6 +493,7 @@ static void free_diff_words_data(struct emit_callback *ecbdata)\n \t\tfree (ecbdata->diff_words->minus.orig);\n \t\tfree (ecbdata->diff_words->plus.text.ptr);\n \t\tfree (ecbdata->diff_words->plus.orig);\n+\t\tfree(ecbdata->diff_words->word_regex);\n \t\tfree(ecbdata->diff_words);\n \t\tecbdata->diff_words = NULL;\n \t}\n@@ -1495,6 +1516,14 @@ static void builtin_diff(const char *name_a,\n \t\t\tecbdata.diff_words =\n \t\t\t\txcalloc(1, sizeof(struct diff_words_data));\n \t\t\tecbdata.diff_words->file = o->file;\n+\t\t\tif (o->word_regex) {\n+\t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n+\t\t\t\t\txmalloc(sizeof(regex_t));\n+\t\t\t\tif (regcomp(ecbdata.diff_words->word_regex,\n+\t\t\t\t\t\to->word_regex, REG_EXTENDED))\n+\t\t\t\t\tdie (\"Invalid regular expression: %s\",\n+\t\t\t\t\t\t\to->word_regex);\n+\t\t\t}\n \t\t}\n \t\txdi_diff_outf(&mf1, &mf2, fn_out_consume, &ecbdata,\n \t\t\t      &xpp, &xecfg, &ecb);\n@@ -2510,6 +2539,10 @@ int diff_opt_parse(struct diff_options *options, const char **av, int ac)\n \t\tDIFF_OPT_CLR(options, COLOR_DIFF);\n \telse if (!strcmp(arg, \"--color-words\"))\n \t\toptions->flags |= DIFF_OPT_COLOR_DIFF | DIFF_OPT_COLOR_DIFF_WORDS;\n+\telse if (!prefixcmp(arg, \"--color-words=\")) {\n+\t\toptions->flags |= DIFF_OPT_COLOR_DIFF | DIFF_OPT_COLOR_DIFF_WORDS;\n+\t\toptions->word_regex = arg + 14;\n+\t}\n \telse if (!strcmp(arg, \"--exit-code\"))\n \t\tDIFF_OPT_SET(options, EXIT_WITH_STATUS);\n \telse if (!strcmp(arg, \"--quiet\"))\ndiff --git a/diff.h b/diff.h\nindex 4d5a327..23cd90c 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -98,6 +98,7 @@ struct diff_options {\n \n \tint stat_width;\n \tint stat_name_width;\n+\tconst char *word_regex;\n \n \t/* this is set by diffcore for DIFF_FORMAT_PATCH */\n \tint found_changes;\n-- \n1.6.1.186.g48f3bc4\n"},{"id":"100011","messageId":"200901112253.27165.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901112057300.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH 0/4] refactor the --color-words to make it more hackable","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-11T21:53:24Z","receivedAt":"2009-01-11T21:53:24Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> \n> But at least _I_ think it is easy to follow, and it actually makes the code\n> more readable/hackable.  Correct me if I'm wrong.\n\nIt indeed seems a sane approach.  However, the final result segfaults\nand/or prints garbage (on apparently every commit except very small\nchanges) when using the regex '\\S+', which IMHO should give exactly\nthe same result as not using a regex at all.  In git.git:\n\n  $ ./git-show --color-words='\\S+' 7eb5bbdb645\n  Segmentation fault\n  $ ./git-show --color-words='\\S+' d3240d935c4\n  [...garbled output...]\n  Segmentation fault\n\nPlain --color-words is not affected.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n\n"},{"id":"100024","messageId":"7vwsd1o44i.fsf@gitster.siamese.dyndns.org","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901112058570.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH 1/4] Add color_fwrite(), a function coloring each line individually","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-11T22:43:41Z","receivedAt":"2009-01-11T22:43:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> +/*\n> + * This function splits the buffer by newlines and colors the lines individually.\n> + */\n> +void color_fwrite(FILE *f, const char *color, size_t count, const char *buf)\n\nIs it just me that this is grossly misnamed?  It is not about fwrite of\ncount bytes starting at buf in the specified color.  At list it should be\ncalled color_fwrite_lines() or something like that.\n\n> diff --git a/color.h b/color.h\n> index 6cf5c88..9fb58f5 100644\n> --- a/color.h\n> +++ b/color.h\n> @@ -19,5 +19,6 @@ int git_config_colorbool(const char *var, const char *value, int stdout_is_tty);\n>  void color_parse(const char *var, const char *value, char *dst);\n>  int color_fprintf(FILE *fp, const char *color, const char *fmt, ...);\n>  int color_fprintf_ln(FILE *fp, const char *color, const char *fmt, ...);\n> +void color_fwrite(FILE *f, const char *color, size_t count, const char *buf);\n\nAlso if other functions in the family all return int to indicate errors\nand name the FILE * argument fp, I find it a very bad taste not to follow\ntheir patterns without having a good reason (which I do not see).\n"},{"id":"100027","messageId":"alpine.DEB.1.00.0901112351050.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"200901112253.27165.trast@student.ethz.ch","subject":"Re: [PATCH 0/4] refactor the --color-words to make it more hackable","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-11T23:02:00Z","receivedAt":"2009-01-11T23:02:00Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 11 Jan 2009, Thomas Rast wrote:\n\n> Johannes Schindelin wrote:\n> > \n> > But at least _I_ think it is easy to follow, and it actually makes the code\n> > more readable/hackable.  Correct me if I'm wrong.\n> \n> It indeed seems a sane approach.\n\nThanks.\n\n>  However, the final result segfaults and/or prints garbage (on \n> apparently every commit except very small changes) when using the regex \n> '\\S+', which IMHO should give exactly the same result as not using a \n> regex at all.\n\nNo, it should not.  The correct regex is '^\\S+'.\n\nAs it happens, your regex matches _anything_ + non-whitespace.  \nUnfortunately, this includes a newline which utterly confuses the diff, \nand therefore the code that tries to get the true offsets.\n\nConsequently, it crashes.\n\n> Plain --color-words is not affected.\n\nOf course, I did not change anything outside the code path of \n--color-words.\n\nCiao,\nDscho\n\n-- snipsnap --\n[PATCH] color-words: \\n must not be a part of the word.\n\nAllowing \\n as part of a word is a pilot error, but that is not a \nreason for the code to crash.\n\nSigned-off-by: Johannes Schindelin <Johannes.Schindelin@gmx.de>\n---\n diff.c |    6 ++++--\n 1 files changed, 4 insertions(+), 2 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex d6bba72..676eb79 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -381,8 +381,10 @@ static int find_word_boundary(mmfile_t *buffer, int i, regex_t *word_regex)\n \n \tif (word_regex) {\n \t\tregmatch_t match[1];\n-\t\tif (!regexec(word_regex, buffer->ptr + i, 1, match, 0))\n-\t\t\ti += match[0].rm_eo;\n+\t\tif (!regexec(word_regex, buffer->ptr + i, 1, match, 0)) {\n+\t\t\tchar *p = memchr(buffer->ptr + i, '\\n', match[0].rm_eo);\n+\t\t\ti = p ? p - buffer->ptr : match[0].rm_eo + i;\n+\t\t}\n \t}\n \telse\n \t\twhile (i < buffer->size && !isspace(buffer->ptr[i]))\n"},{"id":"100030","messageId":"7viqolo2z9.fsf@gitster.siamese.dyndns.org","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901112059340.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH 3/4] color-words: refactor to allow for 0-character word boundaries","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-11T23:08:26Z","receivedAt":"2009-01-11T23:08:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> This code was ugly, for a number of reasons:\n> ...\n> Fix all of these issues by processing the text such that\n\nLooks much cleaner than the original.  I didn't compare it with Thomas's,\nbut it seems he found some breakages, so I'd expect a second round\nsometime in the future.\n"},{"id":"100032","messageId":"alpine.DEB.1.00.0901120037460.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"7viqolo2z9.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH 3/4] color-words: refactor to allow for 0-character word boundaries","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-11T23:38:00Z","receivedAt":"2009-01-11T23:38:00Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 11 Jan 2009, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > This code was ugly, for a number of reasons:\n> > ...\n> > Fix all of these issues by processing the text such that\n> \n> Looks much cleaner than the original.  I didn't compare it with \n> Thomas's, but it seems he found some breakages, so I'd expect a second \n> round sometime in the future.\n\nCertainly,\nDscho\n"},{"id":"100034","messageId":"alpine.DEB.1.00.0901120048430.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"7vwsd1o44i.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH 1/4] Add color_fwrite(), a function coloring each line individually","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-11T23:49:03Z","receivedAt":"2009-01-11T23:49:03Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 11 Jan 2009, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > +/*\n> > + * This function splits the buffer by newlines and colors the lines individually.\n> > + */\n> > +void color_fwrite(FILE *f, const char *color, size_t count, const char *buf)\n> \n> Is it just me that this is grossly misnamed?  It is not about fwrite of\n> count bytes starting at buf in the specified color.  At list it should be\n> called color_fwrite_lines() or something like that.\n> \n> > diff --git a/color.h b/color.h\n> > index 6cf5c88..9fb58f5 100644\n> > --- a/color.h\n> > +++ b/color.h\n> > @@ -19,5 +19,6 @@ int git_config_colorbool(const char *var, const char *value, int stdout_is_tty);\n> >  void color_parse(const char *var, const char *value, char *dst);\n> >  int color_fprintf(FILE *fp, const char *color, const char *fmt, ...);\n> >  int color_fprintf_ln(FILE *fp, const char *color, const char *fmt, ...);\n> > +void color_fwrite(FILE *f, const char *color, size_t count, const char *buf);\n> \n> Also if other functions in the family all return int to indicate errors\n> and name the FILE * argument fp, I find it a very bad taste not to follow\n> their patterns without having a good reason (which I do not see).\n\nValid points.\n\nSorry,\nDscho\n"},{"id":"100035","messageId":"alpine.DEB.1.00.0901120049190.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901120048430.3586@pacific.mpi-cbg.de","subject":"[PATCH v2 1/4] Add color_fwrite(), a function coloring each line individually","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-11T23:49:32Z","receivedAt":"2009-01-11T23:49:32Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nWe have to set the color before every line and reset it before every\nnewline.  Add a function color_fwrite() which does that for us.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n color.c |   28 ++++++++++++++++++++++++++++\n color.h |    1 +\n 2 files changed, 29 insertions(+), 0 deletions(-)\n\ndiff --git a/color.c b/color.c\nindex fc0b72a..b028880 100644\n--- a/color.c\n+++ b/color.c\n@@ -191,3 +191,31 @@ int color_fprintf_ln(FILE *fp, const char *color, const char *fmt, ...)\n \tva_end(args);\n \treturn r;\n }\n+\n+/*\n+ * This function splits the buffer by newlines and colors the lines individually.\n+ *\n+ * Returns 0 on success.\n+ */\n+int color_fwrite_lines(FILE *fp, const char *color,\n+\t\tsize_t count, const char *buf)\n+{\n+\tif (!*color)\n+\t\treturn fwrite(buf, count, 1, fp) != 1;\n+\twhile (count) {\n+\t\tchar *p = memchr(buf, '\\n', count);\n+\t\tif (fputs(color, fp) < 0 ||\n+\t\t\t\tfwrite(buf, p ? p - buf : count, 1, fp) != 1 ||\n+\t\t\t\tfputs(COLOR_RESET, fp) < 0)\n+\t\t\treturn -1;\n+\t\tif (!p)\n+\t\t\treturn 0;\n+\t\tif (fputc('\\n', fp) < 0)\n+\t\t\treturn -1;\n+\t\tcount -= p + 1 - buf;\n+\t\tbuf = p + 1;\n+\t}\n+\treturn 0;\n+}\n+\n+\ndiff --git a/color.h b/color.h\nindex 6cf5c88..cd5c985 100644\n--- a/color.h\n+++ b/color.h\n@@ -19,5 +19,6 @@ int git_config_colorbool(const char *var, const char *value, int stdout_is_tty);\n void color_parse(const char *var, const char *value, char *dst);\n int color_fprintf(FILE *fp, const char *color, const char *fmt, ...);\n int color_fprintf_ln(FILE *fp, const char *color, const char *fmt, ...);\n+int color_fwrite_lines(FILE *fp, const char *color, size_t count, const char *buf);\n \n #endif /* COLOR_H */\n-- \n1.6.1.223.g50c8f\n"},{"id":"100046","messageId":"gke69p$ms0$1@ger.gmane.org","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901120049190.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH v2 1/4] Add color_fwrite(), a function coloring each line individually","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-01-12T01:27:25Z","receivedAt":"2009-01-12T01:27:25Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Johannes Schindelin wrote:\n\n> We have to set the color before every line and reset it before every\n> newline.  Add a function color_fwrite() which does that for us.\n\ncolor_fwrite_lines(), but I guess Junio can correct this himself.\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"100075","messageId":"200901120725.39463.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901112351050.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH 0/4] refactor the --color-words to make it more hackable","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-12T06:25:28Z","receivedAt":"2009-01-12T06:25:28Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> On Sun, 11 Jan 2009, Thomas Rast wrote:\n> >  However, the final result segfaults and/or prints garbage (on \n> > apparently every commit except very small changes) when using the regex \n> > '\\S+', which IMHO should give exactly the same result as not using a \n> > regex at all.\n> \n> No, it should not.  The correct regex is '^\\S+'.\n> \n> As it happens, your regex matches _anything_ + non-whitespace.\n\nIt definitely doesn't(*).\n\nGiven ' word rest', '^\\S+' would not match at all, and '\\S+' would\nmatch 'word'.  No space there.  However, at a cursory glance your\npatch seems to ignore the rm_so member of match[0], so it'll never\nknow the difference.\n\nWhile it might arguably make sense to enforce that only isspace()\ncharacters are whitespace and !isspace() is at least part of _some_\n(possibly one-character) word, I do not think it is a good idea to\nrequire the anchoring of the user.  If we need it, we must anchor the\nmatch ourselves.\n\n> Unfortunately, this includes a newline which utterly confuses the\n> diff, \n\nI do agree that matching a newline as part of a word is bad because we\nneed it for its diff separator semantics.  Consider passing\nREG_NEWLINE to regcomp() to reduce the risk of matching newlines via\nthings like [^\\[:space:]].\n\n\n\n(*) Well, modulo Junio's objection in the other thread that \\S is\nactually a PCRE extension.  Substitute [^[:space:]] if your local\nflavour doesn't understand it.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n\n\n"},{"id":"100083","messageId":"200901120947.13566.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901112059340.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH 3/4] color-words: refactor to allow for 0-character word boundaries","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-12T08:47:11Z","receivedAt":"2009-01-12T08:47:11Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"As a side remark, this patch makes a good use-case for --patience, and\nis not isomorphic to the other edit-and-move examples; rather it's a\ndelete-and-edit.\n\nJohannes Schindelin wrote:\n> Subject: [PATCH 3/4] color-words: refactor to allow for 0-character word boundaries\n\nI do not think the term \"refactor\" is accurate.  Wikipedia roughly\ndefines it as a code change that preserves all external semantics by\nsome standard method, and lists methods such as variable renaming,\ncommon code extraction, etc.  You are actually completely replacing\nthe algorithm \"under the hood\" with a new one, so no such standard\nmethod applies.\n\nAnd there is also a tiny semantic change: compare\n\n  A: a b  c\n  B: x y  z\n        ^^\n\nThe old version implicitly generated an empty line at the double\nspaces (marked ^^), which subsequently became context and caused the\nwords to be printed as follows, where <..> is old and [..] is new:\n\n  <a b >[x y ] <c>[z]\n\nYour patched version does not generate empty lines for any space\nwhatsoever, not even for newlines.  Thus the result is\n\n  <a b  c>[x y  z]\n\nI think this is actually a good change, since it results in longer\nchunks for \"entirely rewritten\" parts of the diff.  It also answers\nJunio's question in the other thread:\n\nJunio C Hamano wrote:\n>> \n>> What happens if the input \"language\" does not have any inter-word spacing\n>> but its words can still be expressed by regexp patterns?\n>> \n>> ImagineALanguageThatAllowsYouToWriteSomethingLikeThis.  Does the mechanism\n>> help users who want to do word-diff files written in such a language by\n>> outputting:\n>> \n>> \tImagineALanguage<red>That</red><green>Which</green>AllowsYou...\n>> \n>> when '[A-Z][a-z]*' is given by the word pattern?\n\nYour patch handles this as a side-effect *even if the lines are\nindented*, since no sequence of spaces whatsoever is special.  (Mine\nwould have given hard-to-predict results based on the number of\nnewlines between them, and xdiff's decision whether the newlines or\nthe words are more valuable as context.)\n\nSo I think this is actually an improvement, but the commit message\nshould point out the change in semantics.\n\n> +static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n>  {\n> +\tif (line[0] != '@' || parse_hunk_header(line, len,\n> +\t\t\t&minus_first, &minus_len, &plus_first, &plus_len))\n\nIt would be nice to have a comment here that points out that this\nmethod crucially relies on having context length 0 (just as the old\none crucially relied on having the full text in a single hunk).\n\n> +\tfor (i = 0; i < buffer->text.size; i++) {\n> +\t\tif (isspace(buffer->text.ptr[i]))\n> +\t\t\tcontinue;\n\nI think it is this coupling of the loops to find a word, and to find a\nword _beginning_, that comes back to haunt you in 4/4.  If the outer\nloop was strictly about the words, you could use the regex match info\nto find the beginning in the regex case.  This is probably cleaner\nthan attempting to force an anchored match, since at least the 'grep'\non my system takes '^^foo' to mean 'a \"^foo\" at the beginning of a\nline', so you cannot just unconditionally insert a ^.  (Conditionally\ninserting one seems even harder.)\n\n\nThese remarks aside (and the last one is the only one of relevance to\nthe code), this patch would be a vast improvement of the code even if\nwe weren't discussing it in the context of the regex feature.  So FWIW\n\n  Acked-by: Thomas Rast <trast@student.ethz.ch>\n\nup to here.  I hope we can agree on some sane regex semantics for\n4/4...\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n\n"},{"id":"100094","messageId":"7vprisj26i.fsf@gitster.siamese.dyndns.org","threadId":"17092","inReplyTo":"200901120947.13566.trast@student.ethz.ch","subject":"Re: [PATCH 3/4] color-words: refactor to allow for 0-character word boundaries","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-12T09:36:53Z","receivedAt":"2009-01-12T09:36:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Rast <trast@student.ethz.ch> writes:\n\n> These remarks aside (and the last one is the only one of relevance to\n> the code), this patch would be a vast improvement of the code even if\n> we weren't discussing it in the context of the regex feature.  So FWIW\n>\n>   Acked-by: Thomas Rast <trast@student.ethz.ch>\n>\n> up to here.  I hope we can agree on some sane regex semantics for\n> 4/4...\n\nOk, although I've already queued your series to 'pu' for the night, I'll\ndrop and replace it with the one from Dscho.  After a few more iteration\nhopefully we can get it into a reasonable shape.\n\nThanks, both.\n"},{"id":"100402","messageId":"adf1fd3d0901140500j10556a1as6370d40d766f1899@mail.gmail.com","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901112057300.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH 0/4] refactor the --color-words to make it more hackable","fromName":"Santi Béjar","fromEmail":"santi@agolina.net","sentAt":"2009-01-14T13:00:57Z","receivedAt":"2009-01-14T13:00:57Z","isPatch":true,"sender":{"key":"santi@agolina.net","avatar":null},"body":"2009/1/11 Johannes Schindelin <Johannes.Schindelin@gmx.de>:\n>\n\n[...]\n\n> The basic idea is to decouple the original text from the text that is\n> passed to libxdiff to find the word differences.\n>\n> To that end, the words of the pre and post texts are put into two lists that\n> are fed to libxdiff.  While the words are extracted, an array is created which\n> contains pointers back to the word boundaries in the original text.\n>\n\nThanks. With this I will no longer need to add some spurious spaces in\nmy latex files :-)\n\nI've tested and it seems to work, but there are some corner cases that\nit does not handle well. If you have this two files:\n\n---8<--- pre\nh(4)\n\na = b + c\n---8<--- post\nh(4),hh[44]\n\na = b + c\n\naa = a\n\naeff = aeff * ( aaa )\n---8<---\n\nThe \"git diff\" is okay, but not the \"git diff --color-words\", the\naddition of \"aeff = ...\" is not shown.\n\nAdditionally with \"git diff --no-index --color-words='^[A-Za-z0-9]*'\nthe ']' character is not shown as an addition, and instead of the\n\"aeff\" line you get a \")\" in green, as:\n\nh(4),{GREEN}hh[44{ENDGREEN}]\n\na = b + c\n\n{GREEN}aa = a\n ){ENDGREEN}\n\nAlso if the lost text is at the end the next \"diff --git\" line is\nprinted in read:\n\n--8<---\n#!/bin/bash\ngit init\ncat > file <<EOF\na\n\naa\nEOF\ncat > gfile <<EOF\na\nEOF\ngit add .\ngit commit -m \"Initial import\"\ngit rm file\ncat > gfile <<EOF\nb\nEOF\ngit add gfile\ngit commit -m \"changes\"\ngit show --color-words\n---8<---\n\nThanks,\nSanti\n\nP.D.: I've test the version that is in 'pu', it does not have the\npatch to fix the segfault but I've also tested with it.\n"},{"id":"100448","messageId":"alpine.DEB.1.00.0901141840100.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"7vprisj26i.fsf@gitster.siamese.dyndns.org","subject":"[PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T17:49:31Z","receivedAt":"2009-01-14T17:49:31Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nThis series is getting bigger and bigger, unfortunately, just what I tried\nto avoid.\n\nBut at least I am pretty comfortable with the readability of the result,\nand it adds tests -- finally.\n\nChanges relative to the last round: color_fwrite_lines() had problems with\nempty lines, and find_word_boundary() was replaced by find_word_boundaries(),\nwhich finds not only the end of the next word, but the start, too.\n\nThe only \"funny\" thing I realized is that the lines which are output\nby emit_line() add a RESET at the end of the line, and I do not do that\nin color_fwrite_lines().\n\nCan anybody think of undesired behavior as a consequence?\n\nJohannes Schindelin (4):\n  Add color_fwrite_lines(), a function coloring each line individually\n  color-words: refactor word splitting and use ALLOC_GROW()\n  color-words: change algorithm to allow for 0-character word\n    boundaries\n  color-words: take an optional regular expression describing words\n\n Documentation/diff-options.txt |    6 +-\n color.c                        |   28 ++++++\n color.h                        |    1 +\n diff.c                         |  203 ++++++++++++++++++++++++++--------------\n diff.h                         |    1 +\n t/t4034-diff-words.sh          |   86 +++++++++++++++++\n 6 files changed, 253 insertions(+), 72 deletions(-)\n create mode 100755 t/t4034-diff-words.sh\n"},{"id":"100449","messageId":"alpine.DEB.1.00.0901141850050.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901141840100.3586@pacific.mpi-cbg.de","subject":"[PATCH 1/4] Add color_fwrite_lines(), a function coloring each line individually","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T17:50:16Z","receivedAt":"2009-01-14T17:50:16Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nWe have to set the color before every line and reset it before every\nnewline.  Add a function color_fwrite_lines() which does that for us.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n color.c |   28 ++++++++++++++++++++++++++++\n color.h |    1 +\n 2 files changed, 29 insertions(+), 0 deletions(-)\n\ndiff --git a/color.c b/color.c\nindex fc0b72a..d4ae83f 100644\n--- a/color.c\n+++ b/color.c\n@@ -191,3 +191,31 @@ int color_fprintf_ln(FILE *fp, const char *color, const char *fmt, ...)\n \tva_end(args);\n \treturn r;\n }\n+\n+/*\n+ * This function splits the buffer by newlines and colors the lines individually.\n+ *\n+ * Returns 0 on success.\n+ */\n+int color_fwrite_lines(FILE *fp, const char *color,\n+\t\tsize_t count, const char *buf)\n+{\n+\tif (!*color)\n+\t\treturn fwrite(buf, count, 1, fp) != 1;\n+\twhile (count) {\n+\t\tchar *p = memchr(buf, '\\n', count);\n+\t\tif (p != buf && (fputs(color, fp) < 0 ||\n+\t\t\t\tfwrite(buf, p ? p - buf : count, 1, fp) != 1 ||\n+\t\t\t\tfputs(COLOR_RESET, fp) < 0))\n+\t\t\treturn -1;\n+\t\tif (!p)\n+\t\t\treturn 0;\n+\t\tif (fputc('\\n', fp) < 0)\n+\t\t\treturn -1;\n+\t\tcount -= p + 1 - buf;\n+\t\tbuf = p + 1;\n+\t}\n+\treturn 0;\n+}\n+\n+\ndiff --git a/color.h b/color.h\nindex 6cf5c88..cd5c985 100644\n--- a/color.h\n+++ b/color.h\n@@ -19,5 +19,6 @@ int git_config_colorbool(const char *var, const char *value, int stdout_is_tty);\n void color_parse(const char *var, const char *value, char *dst);\n int color_fprintf(FILE *fp, const char *color, const char *fmt, ...);\n int color_fprintf_ln(FILE *fp, const char *color, const char *fmt, ...);\n+int color_fwrite_lines(FILE *fp, const char *color, size_t count, const char *buf);\n \n #endif /* COLOR_H */\n-- \n1.6.1.243.g4c9c5a\n"},{"id":"100451","messageId":"alpine.DEB.1.00.0901141850260.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901141840100.3586@pacific.mpi-cbg.de","subject":"[PATCH 2/4] color-words: refactor word splitting and use ALLOC_GROW()","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T17:50:53Z","receivedAt":"2009-01-14T17:50:53Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nWord splitting is now performed by the function diff_words_fill(),\navoiding having the same code twice.\n\nIn the same spirit, avoid duplicating the code of ALLOC_GROW().\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n\n\tThis has not changed, actually.  Just for your convenience.\n\n diff.c |   40 +++++++++++++++++++---------------------\n 1 files changed, 19 insertions(+), 21 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex f67e0b2..6d87ea5 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -326,10 +326,7 @@ struct diff_words_buffer {\n static void diff_words_append(char *line, unsigned long len,\n \t\tstruct diff_words_buffer *buffer)\n {\n-\tif (buffer->text.size + len > buffer->alloc) {\n-\t\tbuffer->alloc = (buffer->text.size + len) * 3 / 2;\n-\t\tbuffer->text.ptr = xrealloc(buffer->text.ptr, buffer->alloc);\n-\t}\n+\tALLOC_GROW(buffer->text.ptr, buffer->text.size + len, buffer->alloc);\n \tline++;\n \tlen--;\n \tmemcpy(buffer->text.ptr + buffer->text.size, line, len);\n@@ -398,6 +395,22 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \t}\n }\n \n+/*\n+ * This function splits the words in buffer->text, and stores the list with\n+ * newline separator into out.\n+ */\n+static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n+{\n+\tint i;\n+\tout->size = buffer->text.size;\n+\tout->ptr = xmalloc(out->size);\n+\tmemcpy(out->ptr, buffer->text.ptr, out->size);\n+\tfor (i = 0; i < out->size; i++)\n+\t\tif (isspace(out->ptr[i]))\n+\t\t\tout->ptr[i] = '\\n';\n+\tbuffer->current = 0;\n+}\n+\n /* this executes the word diff on the accumulated buffers */\n static void diff_words_show(struct diff_words_data *diff_words)\n {\n@@ -405,26 +418,11 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \txdemitconf_t xecfg;\n \txdemitcb_t ecb;\n \tmmfile_t minus, plus;\n-\tint i;\n \n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n-\tminus.size = diff_words->minus.text.size;\n-\tminus.ptr = xmalloc(minus.size);\n-\tmemcpy(minus.ptr, diff_words->minus.text.ptr, minus.size);\n-\tfor (i = 0; i < minus.size; i++)\n-\t\tif (isspace(minus.ptr[i]))\n-\t\t\tminus.ptr[i] = '\\n';\n-\tdiff_words->minus.current = 0;\n-\n-\tplus.size = diff_words->plus.text.size;\n-\tplus.ptr = xmalloc(plus.size);\n-\tmemcpy(plus.ptr, diff_words->plus.text.ptr, plus.size);\n-\tfor (i = 0; i < plus.size; i++)\n-\t\tif (isspace(plus.ptr[i]))\n-\t\t\tplus.ptr[i] = '\\n';\n-\tdiff_words->plus.current = 0;\n-\n+\tdiff_words_fill(&diff_words->minus, &minus);\n+\tdiff_words_fill(&diff_words->plus, &plus);\n \txpp.flags = XDF_NEED_MINIMAL;\n \txecfg.ctxlen = diff_words->minus.alloc + diff_words->plus.alloc;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n-- \n1.6.1.243.g4c9c5a\n"},{"id":"100450","messageId":"alpine.DEB.1.00.0901141851030.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901141840100.3586@pacific.mpi-cbg.de","subject":"[PATCH 3/4] color-words: change algorithm to allow for 0-character word boundaries","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T17:51:24Z","receivedAt":"2009-01-14T17:51:24Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nUp until now, the color-words code assumed that word boundaries are\nidentical to white space characters.\n\nTherefore, it could get away with a very simple scheme: it copied the\nhunks, substituted newlines for each white space character, called\nlibxdiff with the processed text, and then identified the text to\noutput by the offsets (which agreed since the original text had the\nsame length).\n\nThis code was ugly, for a number of reasons:\n\n- it was impossible to introduce 0-character word boundaries,\n\n- we had to print everything word by word, and\n\n- the code needed extra special handling of newlines in the removed part.\n\nFix all of these issues by processing the text such that\n\n- we build word lists, separated by newlines,\n\n- we remember the original offsets for every word, and\n\n- after calling libxdiff on the wordlists, we parse the hunk headers, and\n  find the corresponding offsets, and then\n\n- we print the removed/added parts in one go.\n\nThe pre and post samples in the test were provided by Santi Béjar.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n diff.c                |  153 +++++++++++++++++++++++++++----------------------\n t/t4034-diff-words.sh |   62 ++++++++++++++++++++\n 2 files changed, 147 insertions(+), 68 deletions(-)\n create mode 100755 t/t4034-diff-words.sh\n\ndiff --git a/diff.c b/diff.c\nindex 6d87ea5..fe8b1f0 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -319,8 +319,10 @@ static int fill_mmfile(mmfile_t *mf, struct diff_filespec *one)\n struct diff_words_buffer {\n \tmmfile_t text;\n \tlong alloc;\n-\tlong current; /* output pointer */\n-\tint suppressed_newline;\n+\tstruct diff_words_orig {\n+\t\tconst char *begin, *end;\n+\t} *orig;\n+\tint orig_nr, orig_alloc;\n };\n \n static void diff_words_append(char *line, unsigned long len,\n@@ -335,80 +337,81 @@ static void diff_words_append(char *line, unsigned long len,\n \n struct diff_words_data {\n \tstruct diff_words_buffer minus, plus;\n+\tconst char *current_plus;\n \tFILE *file;\n };\n \n-static void print_word(FILE *file, struct diff_words_buffer *buffer, int len, int color,\n-\t\tint suppress_newline)\n-{\n-\tconst char *ptr;\n-\tint eol = 0;\n-\n-\tif (len == 0)\n-\t\treturn;\n-\n-\tptr  = buffer->text.ptr + buffer->current;\n-\tbuffer->current += len;\n-\n-\tif (ptr[len - 1] == '\\n') {\n-\t\teol = 1;\n-\t\tlen--;\n-\t}\n-\n-\tfputs(diff_get_color(1, color), file);\n-\tfwrite(ptr, len, 1, file);\n-\tfputs(diff_get_color(1, DIFF_RESET), file);\n-\n-\tif (eol) {\n-\t\tif (suppress_newline)\n-\t\t\tbuffer->suppressed_newline = 1;\n-\t\telse\n-\t\t\tputc('\\n', file);\n-\t}\n-}\n-\n static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n {\n \tstruct diff_words_data *diff_words = priv;\n+\tint minus_first, minus_len, plus_first, plus_len;\n+\tconst char *minus_begin, *minus_end, *plus_begin, *plus_end;\n \n-\tif (diff_words->minus.suppressed_newline) {\n-\t\tif (line[0] != '+')\n-\t\t\tputc('\\n', diff_words->file);\n-\t\tdiff_words->minus.suppressed_newline = 0;\n-\t}\n+\tif (line[0] != '@' || parse_hunk_header(line, len,\n+\t\t\t&minus_first, &minus_len, &plus_first, &plus_len))\n+\t\treturn;\n \n-\tlen--;\n-\tswitch (line[0]) {\n-\t\tcase '-':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->minus, len, DIFF_FILE_OLD, 1);\n-\t\t\tbreak;\n-\t\tcase '+':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->plus, len, DIFF_FILE_NEW, 0);\n-\t\t\tbreak;\n-\t\tcase ' ':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->plus, len, DIFF_PLAIN, 0);\n-\t\t\tdiff_words->minus.current += len;\n-\t\t\tbreak;\n-\t}\n+\tminus_begin = diff_words->minus.orig[minus_first].begin;\n+\tminus_end = minus_len == 0 ? minus_begin :\n+\t\tdiff_words->minus.orig[minus_first + minus_len - 1].end;\n+\tplus_begin = diff_words->plus.orig[plus_first].begin;\n+\tplus_end = plus_len == 0 ? plus_begin :\n+\t\tdiff_words->plus.orig[plus_first + plus_len - 1].end;\n+\n+\tif (diff_words->current_plus != plus_begin)\n+\t\tfwrite(diff_words->current_plus,\n+\t\t\t\tplus_begin - diff_words->current_plus, 1,\n+\t\t\t\tdiff_words->file);\n+\tif (minus_begin != minus_end)\n+\t\tcolor_fwrite_lines(diff_words->file,\n+\t\t\t\tdiff_get_color(1, DIFF_FILE_OLD),\n+\t\t\t\tminus_end - minus_begin, minus_begin);\n+\tif (plus_begin != plus_end)\n+\t\tcolor_fwrite_lines(diff_words->file,\n+\t\t\t\tdiff_get_color(1, DIFF_FILE_NEW),\n+\t\t\t\tplus_end - plus_begin, plus_begin);\n+\n+\tdiff_words->current_plus = plus_end;\n }\n \n /*\n- * This function splits the words in buffer->text, and stores the list with\n- * newline separator into out.\n+ * This function splits the words in buffer->text, stores the list with\n+ * newline separator into out, and saves the offsets of the original words\n+ * in buffer->orig.\n  */\n static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n {\n-\tint i;\n-\tout->size = buffer->text.size;\n-\tout->ptr = xmalloc(out->size);\n-\tmemcpy(out->ptr, buffer->text.ptr, out->size);\n-\tfor (i = 0; i < out->size; i++)\n-\t\tif (isspace(out->ptr[i]))\n-\t\t\tout->ptr[i] = '\\n';\n-\tbuffer->current = 0;\n+\tint i, j;\n+\n+\tout->size = 0;\n+\tout->ptr = xmalloc(buffer->text.size);\n+\n+\t/* fake an empty \"0th\" word */\n+\tALLOC_GROW(buffer->orig, 1, buffer->orig_alloc);\n+\tbuffer->orig[0].begin = buffer->orig[0].end = buffer->text.ptr;\n+\tbuffer->orig_nr = 1;\n+\n+\tfor (i = 0; i < buffer->text.size; i++) {\n+\t\tif (isspace(buffer->text.ptr[i]))\n+\t\t\tcontinue;\n+\t\tfor (j = i + 1; j < buffer->text.size &&\n+\t\t\t\t!isspace(buffer->text.ptr[j]); j++)\n+\t\t\t; /* find the end of the word */\n+\n+\t\t/* store original boundaries */\n+\t\tALLOC_GROW(buffer->orig, buffer->orig_nr + 1,\n+\t\t\t\tbuffer->orig_alloc);\n+\t\tbuffer->orig[buffer->orig_nr].begin = buffer->text.ptr + i;\n+\t\tbuffer->orig[buffer->orig_nr].end = buffer->text.ptr + j;\n+\t\tbuffer->orig_nr++;\n+\n+\t\t/* store one word */\n+\t\tmemcpy(out->ptr + out->size, buffer->text.ptr + i, j - i);\n+\t\tout->ptr[out->size + j - i] = '\\n';\n+\t\tout->size += j - i + 1;\n+\n+\t\ti = j - 1;\n+\t}\n }\n \n /* this executes the word diff on the accumulated buffers */\n@@ -419,22 +422,34 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \txdemitcb_t ecb;\n \tmmfile_t minus, plus;\n \n+\t/* special case: only removal */\n+\tif (!diff_words->plus.text.size) {\n+\t\tcolor_fwrite_lines(diff_words->file,\n+\t\t\tdiff_get_color(1, DIFF_FILE_OLD),\n+\t\t\tdiff_words->minus.text.size, diff_words->minus.text.ptr);\n+\t\tdiff_words->minus.text.size = 0;\n+\t\treturn;\n+\t}\n+\n+\tdiff_words->current_plus = diff_words->plus.text.ptr;\n+\n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n \tdiff_words_fill(&diff_words->minus, &minus);\n \tdiff_words_fill(&diff_words->plus, &plus);\n \txpp.flags = XDF_NEED_MINIMAL;\n-\txecfg.ctxlen = diff_words->minus.alloc + diff_words->plus.alloc;\n+\txecfg.ctxlen = 0;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n \t\t      &xpp, &xecfg, &ecb);\n \tfree(minus.ptr);\n \tfree(plus.ptr);\n+\tif (diff_words->current_plus != diff_words->plus.text.ptr +\n+\t\t\tdiff_words->plus.text.size)\n+\t\tfwrite(diff_words->current_plus,\n+\t\t\tdiff_words->plus.text.ptr + diff_words->plus.text.size\n+\t\t\t- diff_words->current_plus, 1,\n+\t\t\tdiff_words->file);\n \tdiff_words->minus.text.size = diff_words->plus.text.size = 0;\n-\n-\tif (diff_words->minus.suppressed_newline) {\n-\t\tputc('\\n', diff_words->file);\n-\t\tdiff_words->minus.suppressed_newline = 0;\n-\t}\n }\n \n typedef unsigned long (*sane_truncate_fn)(char *line, unsigned long len);\n@@ -458,7 +473,9 @@ static void free_diff_words_data(struct emit_callback *ecbdata)\n \t\t\tdiff_words_show(ecbdata->diff_words);\n \n \t\tfree (ecbdata->diff_words->minus.text.ptr);\n+\t\tfree (ecbdata->diff_words->minus.orig);\n \t\tfree (ecbdata->diff_words->plus.text.ptr);\n+\t\tfree (ecbdata->diff_words->plus.orig);\n \t\tfree(ecbdata->diff_words);\n \t\tecbdata->diff_words = NULL;\n \t}\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nnew file mode 100755\nindex 0000000..b032bd3\n--- /dev/null\n+++ b/t/t4034-diff-words.sh\n@@ -0,0 +1,62 @@\n+#!/bin/sh\n+\n+test_description='word diff colors'\n+\n+. ./test-lib.sh\n+\n+test_expect_success setup '\n+\n+\tgit config diff.color.old red\n+\tgit config diff.color.new green\n+\n+'\n+\n+decrypt_color () {\n+\tsed \\\n+\t\t-e 's/.\\[1m/<WHITE>/g' \\\n+\t\t-e 's/.\\[31m/<RED>/g' \\\n+\t\t-e 's/.\\[32m/<GREEN>/g' \\\n+\t\t-e 's/.\\[36m/<BROWN>/g' \\\n+\t\t-e 's/.\\[m/<RESET>/g'\n+}\n+\n+cat > pre <<\\EOF\n+h(4)\n+\n+a = b + c\n+EOF\n+\n+cat > post <<\\EOF\n+h(4),hh[44]\n+\n+a = b + c\n+\n+aa = a\n+\n+aeff = aeff * ( aaa )\n+EOF\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 330b04f..5ed8eff 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1,3 +1,7 @@<RESET>\n+<RED>h(4)<RESET><GREEN>h(4),hh[44]<RESET>\n+<RESET>\n+a = b + c<RESET>\n+\n+<GREEN>aa = a<RESET>\n+\n+<GREEN>aeff = aeff * ( aaa )<RESET>\n+EOF\n+\n+test_expect_success 'word diff with runs of whitespace' '\n+\n+\ttest_must_fail git diff --no-index --color-words pre post > output &&\n+\tdecrypt_color < output > output.decrypted &&\n+\ttest_cmp expect output.decrypted\n+\n+'\n+\n+test_done\n-- \n1.6.1.243.g4c9c5a\n\n"},{"id":"100452","messageId":"alpine.DEB.1.00.0901141851350.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901141840100.3586@pacific.mpi-cbg.de","subject":"[PATCH 4/4] color-words: take an optional regular expression describing words","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T17:51:47Z","receivedAt":"2009-01-14T17:51:47Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nIn some applications, words are not delimited by white space.  To\nallow for that, you can specify a regular expression describing\nwhat makes a word with\n\n\tgit diff --color-words='[A-Za-z0-9]+'\n\nNote that words cannot contain newline characters.\n\nAs suggested by Thomas Rast, the words are the exact matches of the\nregular expression.\n\nNote that a regular expression beginning with a '^' will match only\na word at the beginning of the hunk, not a word at the beginning of\na line, and is probably not what you want.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/diff-options.txt |    6 +++-\n diff.c                         |   64 ++++++++++++++++++++++++++++++++++-----\n diff.h                         |    1 +\n t/t4034-diff-words.sh          |   24 +++++++++++++++\n 4 files changed, 85 insertions(+), 10 deletions(-)\n\ndiff --git a/Documentation/diff-options.txt b/Documentation/diff-options.txt\nindex 1f8ce97..e546bfa 100644\n--- a/Documentation/diff-options.txt\n+++ b/Documentation/diff-options.txt\n@@ -94,8 +94,12 @@ endif::git-format-patch[]\n \tTurn off colored diff, even when the configuration file\n \tgives the default to color output.\n \n---color-words::\n+--color-words[=regex]::\n \tShow colored word diff, i.e. color words which have changed.\n++\n+Optionally, you can pass a regular expression that tells Git what the\n+words are that you are looking for; The default is to interpret any\n+stretch of non-whitespace as a word.\n \n --no-renames::\n \tTurn off rename detection, even when the configuration\ndiff --git a/diff.c b/diff.c\nindex fe8b1f0..d5d7171 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -333,12 +333,14 @@ static void diff_words_append(char *line, unsigned long len,\n \tlen--;\n \tmemcpy(buffer->text.ptr + buffer->text.size, line, len);\n \tbuffer->text.size += len;\n+\tbuffer->text.ptr[buffer->text.size] = '\\0';\n }\n \n struct diff_words_data {\n \tstruct diff_words_buffer minus, plus;\n \tconst char *current_plus;\n \tFILE *file;\n+\tregex_t *word_regex;\n };\n \n static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n@@ -374,17 +376,49 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \tdiff_words->current_plus = plus_end;\n }\n \n+/* This function starts looking at *begin, and returns 0 iff a word was found. */\n+static int find_word_boundaries(mmfile_t *buffer, regex_t *word_regex,\n+\t\tint *begin, int *end)\n+{\n+\tif (word_regex && *begin < buffer->size) {\n+\t\tregmatch_t match[1];\n+\t\tif (!regexec(word_regex, buffer->ptr + *begin, 1, match, 0)) {\n+\t\t\tchar *p = memchr(buffer->ptr + *begin + match[0].rm_so,\n+\t\t\t\t\t'\\n', match[0].rm_eo);\n+\t\t\t*end = p ? p - buffer->ptr : match[0].rm_eo + *begin;\n+\t\t\t*begin += match[0].rm_so;\n+\t\t\treturn *begin >= *end;\n+\t\t}\n+\t\treturn -1;\n+\t}\n+\n+\t/* find the next word */\n+\twhile (*begin < buffer->size && isspace(buffer->ptr[*begin]))\n+\t\t(*begin)++;\n+\tif (*begin >= buffer->size)\n+\t\treturn -1;\n+\n+\t/* find the end of the word */\n+\t*end = *begin + 1;\n+\twhile (*end < buffer->size && !isspace(buffer->ptr[*end]))\n+\t\t(*end)++;\n+\n+\treturn 0;\n+}\n+\n /*\n  * This function splits the words in buffer->text, stores the list with\n  * newline separator into out, and saves the offsets of the original words\n  * in buffer->orig.\n  */\n-static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n+static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out,\n+\t\tregex_t *word_regex)\n {\n \tint i, j;\n+\tlong alloc = 0;\n \n \tout->size = 0;\n-\tout->ptr = xmalloc(buffer->text.size);\n+\tout->ptr = NULL;\n \n \t/* fake an empty \"0th\" word */\n \tALLOC_GROW(buffer->orig, 1, buffer->orig_alloc);\n@@ -392,11 +426,8 @@ static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n \tbuffer->orig_nr = 1;\n \n \tfor (i = 0; i < buffer->text.size; i++) {\n-\t\tif (isspace(buffer->text.ptr[i]))\n-\t\t\tcontinue;\n-\t\tfor (j = i + 1; j < buffer->text.size &&\n-\t\t\t\t!isspace(buffer->text.ptr[j]); j++)\n-\t\t\t; /* find the end of the word */\n+\t\tif (find_word_boundaries(&buffer->text, word_regex, &i, &j))\n+\t\t\treturn;\n \n \t\t/* store original boundaries */\n \t\tALLOC_GROW(buffer->orig, buffer->orig_nr + 1,\n@@ -406,6 +437,7 @@ static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n \t\tbuffer->orig_nr++;\n \n \t\t/* store one word */\n+\t\tALLOC_GROW(out->ptr, out->size + j - i + 1, alloc);\n \t\tmemcpy(out->ptr + out->size, buffer->text.ptr + i, j - i);\n \t\tout->ptr[out->size + j - i] = '\\n';\n \t\tout->size += j - i + 1;\n@@ -435,9 +467,10 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n-\tdiff_words_fill(&diff_words->minus, &minus);\n-\tdiff_words_fill(&diff_words->plus, &plus);\n+\tdiff_words_fill(&diff_words->minus, &minus, diff_words->word_regex);\n+\tdiff_words_fill(&diff_words->plus, &plus, diff_words->word_regex);\n \txpp.flags = XDF_NEED_MINIMAL;\n+\t/* as only the hunk header will be parsed, we need a 0-context */\n \txecfg.ctxlen = 0;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n \t\t      &xpp, &xecfg, &ecb);\n@@ -476,6 +509,7 @@ static void free_diff_words_data(struct emit_callback *ecbdata)\n \t\tfree (ecbdata->diff_words->minus.orig);\n \t\tfree (ecbdata->diff_words->plus.text.ptr);\n \t\tfree (ecbdata->diff_words->plus.orig);\n+\t\tfree(ecbdata->diff_words->word_regex);\n \t\tfree(ecbdata->diff_words);\n \t\tecbdata->diff_words = NULL;\n \t}\n@@ -1498,6 +1532,14 @@ static void builtin_diff(const char *name_a,\n \t\t\tecbdata.diff_words =\n \t\t\t\txcalloc(1, sizeof(struct diff_words_data));\n \t\t\tecbdata.diff_words->file = o->file;\n+\t\t\tif (o->word_regex) {\n+\t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n+\t\t\t\t\txmalloc(sizeof(regex_t));\n+\t\t\t\tif (regcomp(ecbdata.diff_words->word_regex,\n+\t\t\t\t\t\to->word_regex, REG_EXTENDED))\n+\t\t\t\t\tdie (\"Invalid regular expression: %s\",\n+\t\t\t\t\t\t\to->word_regex);\n+\t\t\t}\n \t\t}\n \t\txdi_diff_outf(&mf1, &mf2, fn_out_consume, &ecbdata,\n \t\t\t      &xpp, &xecfg, &ecb);\n@@ -2513,6 +2555,10 @@ int diff_opt_parse(struct diff_options *options, const char **av, int ac)\n \t\tDIFF_OPT_CLR(options, COLOR_DIFF);\n \telse if (!strcmp(arg, \"--color-words\"))\n \t\toptions->flags |= DIFF_OPT_COLOR_DIFF | DIFF_OPT_COLOR_DIFF_WORDS;\n+\telse if (!prefixcmp(arg, \"--color-words=\")) {\n+\t\toptions->flags |= DIFF_OPT_COLOR_DIFF | DIFF_OPT_COLOR_DIFF_WORDS;\n+\t\toptions->word_regex = arg + 14;\n+\t}\n \telse if (!strcmp(arg, \"--exit-code\"))\n \t\tDIFF_OPT_SET(options, EXIT_WITH_STATUS);\n \telse if (!strcmp(arg, \"--quiet\"))\ndiff --git a/diff.h b/diff.h\nindex 4d5a327..23cd90c 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -98,6 +98,7 @@ struct diff_options {\n \n \tint stat_width;\n \tint stat_name_width;\n+\tconst char *word_regex;\n \n \t/* this is set by diffcore for DIFF_FORMAT_PATCH */\n \tint found_changes;\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex b032bd3..0ed7e53 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -59,4 +59,28 @@ test_expect_success 'word diff with runs of whitespace' '\n \n '\n \n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 330b04f..5ed8eff 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1,3 +1,7 @@<RESET>\n+h(4),<GREEN>hh<RESET>[44]\n+<RESET>\n+a = b + c<RESET>\n+\n+<GREEN>aa = a<RESET>\n+\n+<GREEN>aeff = aeff * ( aaa )<RESET>\n+EOF\n+\n+test_expect_success 'word diff with a regular expression' '\n+\n+\ttest_must_fail git diff --no-index --color-words='[a-z]+' \\\n+\t\tpre post > output &&\n+\tdecrypt_color < output > output.decrypted &&\n+\ttest_cmp expect output.decrypted\n+\n+'\n+\n test_done\n-- \n1.6.1.243.g4c9c5a\n"},{"id":"100455","messageId":"alpine.DEB.1.00.0901141907200.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901141851030.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH 3/4] color-words: change algorithm to allow for 0-character word boundaries","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T18:08:06Z","receivedAt":"2009-01-14T18:08:06Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 14 Jan 2009, Johannes Schindelin wrote:\n\n> +test_expect_success setup '\n> +\n> +\tgit config diff.color.old red\n> +\tgit config diff.color.new green\n> +\n> +'\n\nOops.  This should probably go...\n\nCiao,\nDscho\n"},{"id":"100459","messageId":"87ljtdk9b3.fsf@iki.fi","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901141840100.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Teemu Likonen","fromEmail":"tlikonen@iki.fi","sentAt":"2009-01-14T18:54:24Z","receivedAt":"2009-01-14T18:54:24Z","isPatch":true,"sender":{"key":"tlikonen@iki.fi","avatar":null},"body":"Johannes Schindelin (2009-01-14 18:49 +0100) wrote:\n\n> Can anybody think of undesired behavior as a consequence?\n>\n> Johannes Schindelin (4):\n>   Add color_fwrite_lines(), a function coloring each line individually\n>   color-words: refactor word splitting and use ALLOC_GROW()\n>   color-words: change algorithm to allow for 0-character word\n>     boundaries\n>   color-words: take an optional regular expression describing words\n\nThere is something I don't understand. Maybe it's a bug or maybe it's my\nlimitation. I'd appreciate if you care to explain the reason of the\nfollowing output. Suppose we have two files and the line diff looks like\nthis:\n\n    --- 1/a\n    +++ 2/b\n    @@ -1 +1 @@\n    -aaa (aaa)\n    +aaa (aaa) aaa\n\nWith --color-diff=a+ it looks like \n\n    aaa (aaa)aaa) aaa\n         ^^^^~~~~ ~~~\n\n^ = red, ~ = green\n\nWhy show changes in the \"aaa)\" part when it didn't actually change?\n"},{"id":"100460","messageId":"87d4epk96e.fsf@iki.fi","threadId":"17092","inReplyTo":"87ljtdk9b3.fsf@iki.fi","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Teemu Likonen","fromEmail":"tlikonen@iki.fi","sentAt":"2009-01-14T18:57:13Z","receivedAt":"2009-01-14T18:57:13Z","isPatch":true,"sender":{"key":"tlikonen@iki.fi","avatar":null},"body":"Teemu Likonen (2009-01-14 20:54 +0200) wrote:\n\n> With --color-diff=a+ it looks like \n\nObviously I meant --color-words=a+\n"},{"id":"100464","messageId":"alpine.DEB.1.00.0901142028010.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"87d4epk96e.fsf@iki.fi","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T19:28:33Z","receivedAt":"2009-01-14T19:28:33Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 14 Jan 2009, Teemu Likonen wrote:\n\n> Teemu Likonen (2009-01-14 20:54 +0200) wrote:\n> \n> > With --color-diff=a+ it looks like \n> \n> Obviously I meant --color-words=a+\n\nHeh,  I missed that, even...  Thanks for the report!\n\n-- snipsnap --\n[WILL BE SQUASHED INTO 4/4] Fix find_word_boundaries()\n\nSince newlines cannot be part of words, we have to stop at newlines even\nif the regular expression's match contains one.\n\nOf course, I fscked up the range where to look for the newline when I\nchanged the function from find_word_boundary().\n---\n diff.c                |    2 +-\n t/t4034-diff-words.sh |   20 ++++++++++++++++++++\n 2 files changed, 21 insertions(+), 1 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex d5d7171..1408717 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -384,7 +384,7 @@ static int find_word_boundaries(mmfile_t *buffer, regex_t *word_regex,\n \t\tregmatch_t match[1];\n \t\tif (!regexec(word_regex, buffer->ptr + *begin, 1, match, 0)) {\n \t\t\tchar *p = memchr(buffer->ptr + *begin + match[0].rm_so,\n-\t\t\t\t\t'\\n', match[0].rm_eo);\n+\t\t\t\t\t'\\n', match[0].rm_eo - match[0].rm_so);\n \t\t\t*end = p ? p - buffer->ptr : match[0].rm_eo + *begin;\n \t\t\t*begin += match[0].rm_so;\n \t\t\treturn *begin >= *end;\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 0ed7e53..1137131 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -83,4 +83,24 @@ test_expect_success 'word diff with a regular expression' '\n \n '\n \n+echo 'aaa (aaa)' > pre\n+echo 'aaa (aaa) aaa' > post\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index c29453b..be22f37 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1 +1 @@<RESET>\n+aaa (aaa)<GREEN> aaa<RESET>\n+EOF\n+\n+test_expect_success \"Teemo's example\" '\n+\n+\ttest_must_fail git diff --no-index --color-words='a+' pre post > output &&\n+\tdecrypt_color < output > output.decrypted &&\n+\ttest_cmp expect output.decrypted\n+\n+'\n+\n test_done\n-- \n1.6.1.295.gb16478\n"},{"id":"100465","messageId":"alpine.DEB.1.00.0901142032080.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901142028010.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T19:32:50Z","receivedAt":"2009-01-14T19:32:50Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 14 Jan 2009, Johannes Schindelin wrote:\n\n> +aaa (aaa)<GREEN> aaa<RESET>\n\nOf course, the space must be on the other side of the <GREEN>...  All this \nwill be fixed, and more.\n\nCiao,\nDscho\n"},{"id":"100467","messageId":"1231962401-26974-1-git-send-email-trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901141840100.3586@pacific.mpi-cbg.de","subject":"[PATCH] color-words: make regex configurable via attributes","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T19:46:41Z","receivedAt":"2009-01-14T19:46:41Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Make the --color-words splitting regular expression configurable via\nthe diff driver's 'wordregex' attribute.  The user can then set the\ndriver on a file in .gitattributes.  If a regex is given on the\ncommand line, it overrides the driver's setting.\n\nWe also provide built-in regexes for the languages that already had\nfuncname patterns, and add an appropriate diff driver entry for C/++.\n(The patterns are designed to run UTF-8 sequences into a single chunk\nto make sure they remain readable.)\n\nSigned-off-by: Thomas Rast <trast@student.ethz.ch>\n---\n\nThis is the old 3/4 combined with a test similar to the one it had in\nthe old 4/4, built on top of Dscho's take 3.  I researched the\noperators for each language, but the identifier and number formats may\nbe off in some cases.\n\n\n Documentation/diff-options.txt  |    3 +\n Documentation/gitattributes.txt |   21 ++++++++++\n diff.c                          |   10 +++++\n t/t4034-diff-words.sh           |   40 ++++++++++++++++++++\n userdiff.c                      |   78 +++++++++++++++++++++++++++++++-------\n userdiff.h                      |    1 +\n 6 files changed, 138 insertions(+), 15 deletions(-)\n\ndiff --git a/Documentation/diff-options.txt b/Documentation/diff-options.txt\nindex 2c1fa4b..ef0e2f5 100644\n--- a/Documentation/diff-options.txt\n+++ b/Documentation/diff-options.txt\n@@ -97,6 +97,9 @@ endif::git-format-patch[]\n Optionally, you can pass a regular expression that tells Git what the\n words are that you are looking for; The default is to interpret any\n stretch of non-whitespace as a word.\n+The regex can also be set via a diff driver, see\n+linkgit:gitattributes[1]; giving it explicitly overrides any diff\n+driver setting.\n \n --no-renames::\n \tTurn off rename detection, even when the configuration\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 8af22ec..17707ba 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -317,6 +317,8 @@ patterns are available:\n \n - `bibtex` suitable for files with BibTeX coded references.\n \n+- `cpp` suitable for source code in the C and C++ languages.\n+\n - `html` suitable for HTML/XHTML documents.\n \n - `java` suitable for source code in the Java language.\n@@ -334,6 +336,25 @@ patterns are available:\n - `tex` suitable for source code for LaTeX documents.\n \n \n+Customizing word diff\n+^^^^^^^^^^^^^^^^^^^^^\n+\n+You can customize the rules that `git diff --color-words` uses to\n+split words in a line, by specifying an appropriate regular expression\n+in the \"diff.*.wordregex\" configuration variable.  For example, in TeX\n+a backslash followed by a sequence of letters forms a command, but\n+several such commands can be run together without intervening\n+whitespace.  To separate them, use a regular expression such as\n+\n+------------------------\n+[diff \"tex\"]\n+\twordregex = \"\\\\\\\\[a-zA-Z]+|[{}]|\\\\\\\\.|[^\\\\{}[:space:]]+\"\n+------------------------\n+\n+A built-in pattern is provided for all languages listed in the last\n+section.\n+\n+\n Performing text diffs of binary files\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \ndiff --git a/diff.c b/diff.c\nindex eb67431..08bdc86 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -1372,6 +1372,12 @@ int diff_filespec_is_binary(struct diff_filespec *one)\n \treturn one->driver->funcname.pattern ? &one->driver->funcname : NULL;\n }\n \n+static const char *userdiff_word_regex(struct diff_filespec *one)\n+{\n+\tdiff_filespec_load_driver(one);\n+\treturn one->driver->word_regex;\n+}\n+\n void diff_set_mnemonic_prefix(struct diff_options *options, const char *a, const char *b)\n {\n \tif (!options->a_prefix)\n@@ -1532,6 +1538,10 @@ static void builtin_diff(const char *name_a,\n \t\t\tecbdata.diff_words =\n \t\t\t\txcalloc(1, sizeof(struct diff_words_data));\n \t\t\tecbdata.diff_words->file = o->file;\n+\t\t\tif (!o->word_regex)\n+\t\t\t\to->word_regex = userdiff_word_regex(one);\n+\t\t\tif (!o->word_regex)\n+\t\t\t\to->word_regex = userdiff_word_regex(two);\n \t\t\tif (o->word_regex) {\n \t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n \t\t\t\t\txmalloc(sizeof(regex_t));\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 0ed7e53..d6731d1 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -83,4 +83,44 @@ test_expect_success 'word diff with a regular expression' '\n \n '\n \n+cat > expect-by-chars <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 330b04f..5ed8eff 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1,3 +1,7 @@<RESET>\n+<RED>h(4)<RESET><GREEN>h(4),hh[44]<RESET>\n+<RESET>\n+a = b + c<RESET>\n+\n+<GREEN>aa = a<RESET>\n+\n+<GREEN>aeff = aeff * ( aaa )<RESET>\n+EOF\n+\n+test_expect_success 'set a diff driver' '\n+\tgit config diff.testdriver.wordregex \"[^[:space:]]\" &&\n+\tcat <<EOF > .gitattributes\n+test_* diff=testdriver\n+EOF\n+'\n+\n+test_expect_success 'use default supplied by driver' '\n+\n+\ttest_must_fail git diff --no-index --color-words \\\n+\t\tpre post > output &&\n+\tdecrypt_color < output > output.decrypted &&\n+\ttest_cmp expect-by-chars output.decrypted\n+\n+'\n+\n+test_expect_success 'option overrides default' '\n+\n+\ttest_must_fail git diff --no-index --color-words=\"[a-z]+\" \\\n+\t\tpre post > output &&\n+\tdecrypt_color < output > output.decrypted &&\n+\ttest_cmp expect output.decrypted\n+\n+'\n+\n test_done\ndiff --git a/userdiff.c b/userdiff.c\nindex 3681062..79f9cb9 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -6,14 +6,20 @@\n static int ndrivers;\n static int drivers_alloc;\n \n-#define FUNCNAME(name, pattern) \\\n-\t{ name, NULL, -1, { pattern, REG_EXTENDED } }\n+#define PATTERNS(name, pattern, wordregex)\t\t\t\\\n+\t{ name, NULL, -1, { pattern, REG_EXTENDED }, NULL, wordregex }\n static struct userdiff_driver builtin_drivers[] = {\n-FUNCNAME(\"html\", \"^[ \\t]*(<[Hh][1-6][ \\t].*>.*)$\"),\n-FUNCNAME(\"java\",\n+PATTERNS(\"html\", \"^[ \\t]*(<[Hh][1-6][ \\t].*>.*)$\",\n+\t \"[^<>= \\t]+|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"java\",\n \t \"!^[ \\t]*(catch|do|for|if|instanceof|new|return|switch|throw|while)\\n\"\n-\t \"^[ \\t]*(([ \\t]*[A-Za-z_][A-Za-z_0-9]*){2,}[ \\t]*\\\\([^;]*)$\"),\n-FUNCNAME(\"objc\",\n+\t \"^[ \\t]*(([ \\t]*[A-Za-z_][A-Za-z_0-9]*){2,}[ \\t]*\\\\([^;]*)$\",\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[fFlL]?|0[xXbB]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=\"\n+\t \"|--|\\\\+\\\\+|<<=?|>>>?=?|&&|\\\\|\\\\|\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"objc\",\n \t /* Negate C statements that can look like functions */\n \t \"!^[ \\t]*(do|for|if|else|return|switch|while)\\n\"\n \t /* Objective-C methods */\n@@ -21,20 +27,60 @@\n \t /* C functions */\n \t \"^[ \\t]*(([ \\t]*[A-Za-z_][A-Za-z_0-9]*){2,}[ \\t]*\\\\([^;]*)$\\n\"\n \t /* Objective-C class/protocol definitions */\n-\t \"^(@(implementation|interface|protocol)[ \\t].*)$\"),\n-FUNCNAME(\"pascal\",\n+\t \"^(@(implementation|interface|protocol)[ \\t].*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[fFlL]?|0[xXbB]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=|--|\\\\+\\\\+|<<=?|>>=?|&&|\\\\|\\\\||::|->\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"pascal\",\n \t \"^((procedure|function|constructor|destructor|interface|\"\n \t\t\"implementation|initialization|finalization)[ \\t]*.*)$\"\n \t \"\\n\"\n-\t \"^(.*=[ \\t]*(class|record).*)$\"),\n-FUNCNAME(\"php\", \"^[\\t ]*((function|class).*)\"),\n-FUNCNAME(\"python\", \"^[ \\t]*((class|def)[ \\t].*)$\"),\n-FUNCNAME(\"ruby\", \"^[ \\t]*((class|module|def)[ \\t].*)$\"),\n-FUNCNAME(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\"),\n-FUNCNAME(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\"),\n+\t \"^(.*=[ \\t]*(class|record).*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+|0[xXbB]?[0-9a-fA-F]+\"\n+\t \"|<>|<=|>=|:=|\\\\.\\\\.\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"php\", \"^[\\t ]*((function|class).*)\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+|0[xXbB]?[0-9a-fA-F]+\"\n+\t \"|[-+*/<>%&^|=!.]=|--|\\\\+\\\\+|<<=?|>>=?|===|&&|\\\\|\\\\||::|->\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"python\", \"^[ \\t]*((class|def)[ \\t].*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[jJlL]?|0[xX]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=|//=?|<<=?|>>=?|\\\\*\\\\*=?\"\n+\t \"|[^[:space:]|[\\x80-\\xff]+\"),\n+\t /* -- */\n+PATTERNS(\"ruby\", \"^[ \\t]*((class|module|def)[ \\t].*)$\",\n+\t /* -- */\n+\t \"(@|@@|\\\\$)?[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+|0[xXbB]?[0-9a-fA-F]+|\\\\?(\\\\\\\\C-)?(\\\\\\\\M-)?.\"\n+\t \"|//=?|[-+*/<>%&^|=!]=|<<=?|>>=?|===|\\\\.{1,3}|::|[!=]~\"\n+\t \"|[^[:space:]|[\\x80-\\xff]+\"),\n+PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n+\t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n+PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n+\t \"\\\\\\\\[a-zA-Z@]+|[{}]|\\\\\\\\.|[^\\\\{} \\t]+\"),\n+PATTERNS(\"cpp\",\n+\t /* Jump targets or access declarations */\n+\t \"!^[ \\t]*[A-Za-z_][A-Za-z_0-9]*:.*$\\n\"\n+\t /* C functions at top level */\n+\t \"^([A-Za-z_][A-Za-z_0-9]*([ \\t]+[A-Za-z_][A-Za-z_0-9]*){1,}[ \\t]*\\\\([^;]*)$\\n\"\n+\t /* compound type at top level */\n+\t \"^((struct|class|enum)[^;]*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[fFlL]?|0[xXbB]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=|--|\\\\+\\\\+|<<=?|>>=?|&&|\\\\|\\\\||::|->\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n { \"default\", NULL, -1, { NULL, 0 } },\n };\n-#undef FUNCNAME\n+#undef PATTERNS\n \n static struct userdiff_driver driver_true = {\n \t\"diff=true\",\n@@ -134,6 +180,8 @@ int userdiff_config(const char *k, const char *v)\n \t\treturn parse_string(&drv->external, k, v);\n \tif ((drv = parse_driver(k, v, \"textconv\")))\n \t\treturn parse_string(&drv->textconv, k, v);\n+\tif ((drv = parse_driver(k, v, \"wordregex\")))\n+\t\treturn parse_string(&drv->word_regex, k, v);\n \n \treturn 0;\n }\ndiff --git a/userdiff.h b/userdiff.h\nindex ba29457..2aab13e 100644\n--- a/userdiff.h\n+++ b/userdiff.h\n@@ -12,6 +12,7 @@ struct userdiff_driver {\n \tint binary;\n \tstruct userdiff_funcname funcname;\n \tconst char *textconv;\n+\tconst char *word_regex;\n };\n \n int userdiff_config(const char *k, const char *v);\n-- \n1.6.1.140.ge720e.dirty\n"},{"id":"100468","messageId":"200901142055.27222.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901141851350.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH 4/4] color-words: take an optional regular expression describing words","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T19:55:21Z","receivedAt":"2009-01-14T19:55:21Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> ---color-words::\n> +--color-words[=regex]::\n>  \tShow colored word diff, i.e. color words which have changed.\n> ++\n> +Optionally, you can pass a regular expression that tells Git what the\n> +words are that you are looking for; The default is to interpret any\n> +stretch of non-whitespace as a word.\n\nPerhaps you could resurrect the documentation from my series, adjusted\nfor the different newline rule:\n\n\n--color-words[=<regex>]::\n\tShow colored word diff, i.e., color words which have changed.\n\tBy default, a new word only starts at whitespace, so that a\n\t'word' is defined as a maximal sequence of non-whitespace\n\tcharacters.  The optional argument <regex> can be used to\n\tconfigure this.  It can also be set via a diff driver, see\n\tlinkgit:gitattributes[1]; if a <regex> is given explicitly, it\n\toverrides any diff driver setting.\n+\nThe <regex> must be an (extended) regular expression.  When set, every\nnon-overlapping match of the <regex> is considered a word.  Anything\nbetween these matches is considered whitespace and ignored for the\npurposes of finding differences.  You may want to append\n`|[^[:space:]]` to your regular expression to make sure that it\nmatches all non-whitespace characters.  A match that contains a\nnewline is silently truncated at the newline.\n\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n"},{"id":"100469","messageId":"200901142059.09005.trast@student.ethz.ch","threadId":"17092","inReplyTo":"87ljtdk9b3.fsf@iki.fi","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T19:58:56Z","receivedAt":"2009-01-14T19:58:56Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Teemu Likonen wrote:\n>     -aaa (aaa)\n>     +aaa (aaa) aaa\n\nBug aside, examples like this one make me wonder if we should force a\n\"last resort\" match for `[^[:space:]]`.  For example,\n\n      -aaa [aaa]\n      +aaa (aaa) aaa\n\nwould still give you\n\n      aaa (aaa)<GREEN> aaa<RESET>\n\nwhich may be unexpected.\n\nOf course, when diffing a language where something other than the\n\"usual\" whitespace should be ignored, this behaviour would be useful.\n\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n\n\n"},{"id":"100473","messageId":"200901142104.16134.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901141840100.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T20:04:13Z","receivedAt":"2009-01-14T20:04:13Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> \n> The only \"funny\" thing I realized is that the lines which are output\n> by emit_line() add a RESET at the end of the line, and I do not do that\n> in color_fwrite_lines().\n\nUmm.... but you seem to do?\n\nAck on the new regex semantics, though I'd have implemented it via\ndying on '\\n' instead of silently splitting there (and restarting a\nnew match!).  [I actually _have_ implemented it, but your patch beat\nme to it. :-)]\n\nThus, Ack on 4/4 once the boundary bug is fixed.  Thanks for your\nwork!\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n\n"},{"id":"100475","messageId":"alpine.DEB.1.00.0901142104400.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"1231962401-26974-1-git-send-email-trast@student.ethz.ch","subject":"Re: [PATCH] color-words: make regex configurable via attributes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T20:12:14Z","receivedAt":"2009-01-14T20:12:14Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 14 Jan 2009, Thomas Rast wrote:\n\n> diff --git a/Documentation/diff-options.txt b/Documentation/diff-options.txt\n> index 2c1fa4b..ef0e2f5 100644\n> --- a/Documentation/diff-options.txt\n> +++ b/Documentation/diff-options.txt\n> @@ -97,6 +97,9 @@ endif::git-format-patch[]\n>  Optionally, you can pass a regular expression that tells Git what the\n>  words are that you are looking for; The default is to interpret any\n>  stretch of non-whitespace as a word.\n> +The regex can also be set via a diff driver, see\n> +linkgit:gitattributes[1]; giving it explicitly overrides any diff\n> +driver setting.\n\nHow about making this an extra paragraph?\n\n> diff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\n> index 8af22ec..17707ba 100644\n> --- a/Documentation/gitattributes.txt\n> +++ b/Documentation/gitattributes.txt\n> @@ -317,6 +317,8 @@ patterns are available:\n>  \n>  - `bibtex` suitable for files with BibTeX coded references.\n>  \n> +- `cpp` suitable for source code in the C and C++ languages.\n> +\n\nHow about \"written in C or C++\"?\n\n> +A built-in pattern is provided for all languages listed in the last\n> +section.\n\nWow.  But how about \"previous section\"?\n\n> diff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\n> index 0ed7e53..d6731d1 100755\n> --- a/t/t4034-diff-words.sh\n> +++ b/t/t4034-diff-words.sh\n\nThat was fast!\n\n> +test_expect_success 'use default supplied by driver' '\n> +\n> +\ttest_must_fail git diff --no-index --color-words \\\n> +\t\tpre post > output &&\n> +\tdecrypt_color < output > output.decrypted &&\n> +\ttest_cmp expect-by-chars output.decrypted\n> +\n> +'\n\nI am actually just about to post new revisions of the last two patches \nwhere this would read\n\n\ttest_expect_success 'use default supplied by driver' '\n\n\t\tword_diff --color-words\n\n\t'\n\ninstead...\n\nI don't want to get bitten by stupid mistakes again, though, so I let it \nrun with valgrind while glancing over the code.  Stay tuned.\n\n> +#define PATTERNS(name, pattern, wordregex)\t\t\t\\\n> +\t{ name, NULL, -1, { pattern, REG_EXTENDED }, NULL, wordregex }\n\nYou could get rid of that NULL if...\n\n> diff --git a/userdiff.h b/userdiff.h\n> index ba29457..2aab13e 100644\n> --- a/userdiff.h\n> +++ b/userdiff.h\n> @@ -12,6 +12,7 @@ struct userdiff_driver {\n>  \tint binary;\n>  \tstruct userdiff_funcname funcname;\n>  \tconst char *textconv;\n> +\tconst char *word_regex;\n>  };\n\n... you inserted word_regex before textconv.  In a way, I find this more \nlogical, since both funcname and word_regex have sensible defaults \n(provided by you), whereas textconv is strictly a user's option.\n\nCiao,\nDscho\n"},{"id":"100477","messageId":"200901142118.02041.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901142104400.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH] color-words: make regex configurable via attributes","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T20:17:59Z","receivedAt":"2009-01-14T20:17:59Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> How about making this an extra paragraph?\n\nSure, why not.  Though I'm still in favour of taking some longer\nversion, possibly from my old series.\n\n> On Wed, 14 Jan 2009, Thomas Rast wrote:\n> > +- `cpp` suitable for source code in the C and C++ languages.\n> > +\n> \n> How about \"written in C or C++\"?\n\nI was just trying to be consistent with all other items; all\nprogramming languages are listed as \"Foo language\".\n\n> > +A built-in pattern is provided for all languages listed in the last\n> > +section.\n> \n> Wow.  But how about \"previous section\"?\n\nIndeed, thanks.\n\n> > +#define PATTERNS(name, pattern, wordregex)\t\t\t\\\n> > +\t{ name, NULL, -1, { pattern, REG_EXTENDED }, NULL, wordregex }\n> \n> You could get rid of that NULL if...\n[...]\n> ... you inserted word_regex before textconv.  In a way, I find this more \n> logical, since both funcname and word_regex have sensible defaults \n> (provided by you), whereas textconv is strictly a user's option.\n\nOk, I'll do that.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n"},{"id":"100478","messageId":"alpine.DEB.1.00.0901142142120.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901142104400.3586@pacific.mpi-cbg.de","subject":"[PATCH replacement for take 3 3/4] color-words: change algorithm to allow for 0-character word boundaries","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T20:44:36Z","receivedAt":"2009-01-14T20:44:36Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nUp until now, the color-words code assumed that word boundaries are\nidentical to white space characters.\n\nTherefore, it could get away with a very simple scheme: it copied the\nhunks, substituted newlines for each white space character, called\nlibxdiff with the processed text, and then identified the text to\noutput by the offsets (which agreed since the original text had the\nsame length).\n\nThis code was ugly, for a number of reasons:\n\n- it was impossible to introduce 0-character word boundaries,\n\n- we had to print everything word by word, and\n\n- the code needed extra special handling of newlines in the removed part.\n\nFix all of these issues by processing the text such that\n\n- we build word lists, separated by newlines,\n\n- we remember the original offsets for every word, and\n\n- after calling libxdiff on the wordlists, we parse the hunk headers, and\n  find the corresponding offsets, and then\n\n- we print the removed/added parts in one go.\n\nThe pre and post samples in the test were provided by Santi Béjar.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n\n\tI changed the test script to avoid repeating the same three-command\n\tmantra all the time.\n\n diff.c                |  153 +++++++++++++++++++++++++++----------------------\n t/t4034-diff-words.sh |   66 +++++++++++++++++++++\n 2 files changed, 151 insertions(+), 68 deletions(-)\n create mode 100755 t/t4034-diff-words.sh\n\ndiff --git a/diff.c b/diff.c\nindex 6d87ea5..fe8b1f0 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -319,8 +319,10 @@ static int fill_mmfile(mmfile_t *mf, struct diff_filespec *one)\n struct diff_words_buffer {\n \tmmfile_t text;\n \tlong alloc;\n-\tlong current; /* output pointer */\n-\tint suppressed_newline;\n+\tstruct diff_words_orig {\n+\t\tconst char *begin, *end;\n+\t} *orig;\n+\tint orig_nr, orig_alloc;\n };\n \n static void diff_words_append(char *line, unsigned long len,\n@@ -335,80 +337,81 @@ static void diff_words_append(char *line, unsigned long len,\n \n struct diff_words_data {\n \tstruct diff_words_buffer minus, plus;\n+\tconst char *current_plus;\n \tFILE *file;\n };\n \n-static void print_word(FILE *file, struct diff_words_buffer *buffer, int len, int color,\n-\t\tint suppress_newline)\n-{\n-\tconst char *ptr;\n-\tint eol = 0;\n-\n-\tif (len == 0)\n-\t\treturn;\n-\n-\tptr  = buffer->text.ptr + buffer->current;\n-\tbuffer->current += len;\n-\n-\tif (ptr[len - 1] == '\\n') {\n-\t\teol = 1;\n-\t\tlen--;\n-\t}\n-\n-\tfputs(diff_get_color(1, color), file);\n-\tfwrite(ptr, len, 1, file);\n-\tfputs(diff_get_color(1, DIFF_RESET), file);\n-\n-\tif (eol) {\n-\t\tif (suppress_newline)\n-\t\t\tbuffer->suppressed_newline = 1;\n-\t\telse\n-\t\t\tputc('\\n', file);\n-\t}\n-}\n-\n static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n {\n \tstruct diff_words_data *diff_words = priv;\n+\tint minus_first, minus_len, plus_first, plus_len;\n+\tconst char *minus_begin, *minus_end, *plus_begin, *plus_end;\n \n-\tif (diff_words->minus.suppressed_newline) {\n-\t\tif (line[0] != '+')\n-\t\t\tputc('\\n', diff_words->file);\n-\t\tdiff_words->minus.suppressed_newline = 0;\n-\t}\n+\tif (line[0] != '@' || parse_hunk_header(line, len,\n+\t\t\t&minus_first, &minus_len, &plus_first, &plus_len))\n+\t\treturn;\n \n-\tlen--;\n-\tswitch (line[0]) {\n-\t\tcase '-':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->minus, len, DIFF_FILE_OLD, 1);\n-\t\t\tbreak;\n-\t\tcase '+':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->plus, len, DIFF_FILE_NEW, 0);\n-\t\t\tbreak;\n-\t\tcase ' ':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->plus, len, DIFF_PLAIN, 0);\n-\t\t\tdiff_words->minus.current += len;\n-\t\t\tbreak;\n-\t}\n+\tminus_begin = diff_words->minus.orig[minus_first].begin;\n+\tminus_end = minus_len == 0 ? minus_begin :\n+\t\tdiff_words->minus.orig[minus_first + minus_len - 1].end;\n+\tplus_begin = diff_words->plus.orig[plus_first].begin;\n+\tplus_end = plus_len == 0 ? plus_begin :\n+\t\tdiff_words->plus.orig[plus_first + plus_len - 1].end;\n+\n+\tif (diff_words->current_plus != plus_begin)\n+\t\tfwrite(diff_words->current_plus,\n+\t\t\t\tplus_begin - diff_words->current_plus, 1,\n+\t\t\t\tdiff_words->file);\n+\tif (minus_begin != minus_end)\n+\t\tcolor_fwrite_lines(diff_words->file,\n+\t\t\t\tdiff_get_color(1, DIFF_FILE_OLD),\n+\t\t\t\tminus_end - minus_begin, minus_begin);\n+\tif (plus_begin != plus_end)\n+\t\tcolor_fwrite_lines(diff_words->file,\n+\t\t\t\tdiff_get_color(1, DIFF_FILE_NEW),\n+\t\t\t\tplus_end - plus_begin, plus_begin);\n+\n+\tdiff_words->current_plus = plus_end;\n }\n \n /*\n- * This function splits the words in buffer->text, and stores the list with\n- * newline separator into out.\n+ * This function splits the words in buffer->text, stores the list with\n+ * newline separator into out, and saves the offsets of the original words\n+ * in buffer->orig.\n  */\n static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n {\n-\tint i;\n-\tout->size = buffer->text.size;\n-\tout->ptr = xmalloc(out->size);\n-\tmemcpy(out->ptr, buffer->text.ptr, out->size);\n-\tfor (i = 0; i < out->size; i++)\n-\t\tif (isspace(out->ptr[i]))\n-\t\t\tout->ptr[i] = '\\n';\n-\tbuffer->current = 0;\n+\tint i, j;\n+\n+\tout->size = 0;\n+\tout->ptr = xmalloc(buffer->text.size);\n+\n+\t/* fake an empty \"0th\" word */\n+\tALLOC_GROW(buffer->orig, 1, buffer->orig_alloc);\n+\tbuffer->orig[0].begin = buffer->orig[0].end = buffer->text.ptr;\n+\tbuffer->orig_nr = 1;\n+\n+\tfor (i = 0; i < buffer->text.size; i++) {\n+\t\tif (isspace(buffer->text.ptr[i]))\n+\t\t\tcontinue;\n+\t\tfor (j = i + 1; j < buffer->text.size &&\n+\t\t\t\t!isspace(buffer->text.ptr[j]); j++)\n+\t\t\t; /* find the end of the word */\n+\n+\t\t/* store original boundaries */\n+\t\tALLOC_GROW(buffer->orig, buffer->orig_nr + 1,\n+\t\t\t\tbuffer->orig_alloc);\n+\t\tbuffer->orig[buffer->orig_nr].begin = buffer->text.ptr + i;\n+\t\tbuffer->orig[buffer->orig_nr].end = buffer->text.ptr + j;\n+\t\tbuffer->orig_nr++;\n+\n+\t\t/* store one word */\n+\t\tmemcpy(out->ptr + out->size, buffer->text.ptr + i, j - i);\n+\t\tout->ptr[out->size + j - i] = '\\n';\n+\t\tout->size += j - i + 1;\n+\n+\t\ti = j - 1;\n+\t}\n }\n \n /* this executes the word diff on the accumulated buffers */\n@@ -419,22 +422,34 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \txdemitcb_t ecb;\n \tmmfile_t minus, plus;\n \n+\t/* special case: only removal */\n+\tif (!diff_words->plus.text.size) {\n+\t\tcolor_fwrite_lines(diff_words->file,\n+\t\t\tdiff_get_color(1, DIFF_FILE_OLD),\n+\t\t\tdiff_words->minus.text.size, diff_words->minus.text.ptr);\n+\t\tdiff_words->minus.text.size = 0;\n+\t\treturn;\n+\t}\n+\n+\tdiff_words->current_plus = diff_words->plus.text.ptr;\n+\n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n \tdiff_words_fill(&diff_words->minus, &minus);\n \tdiff_words_fill(&diff_words->plus, &plus);\n \txpp.flags = XDF_NEED_MINIMAL;\n-\txecfg.ctxlen = diff_words->minus.alloc + diff_words->plus.alloc;\n+\txecfg.ctxlen = 0;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n \t\t      &xpp, &xecfg, &ecb);\n \tfree(minus.ptr);\n \tfree(plus.ptr);\n+\tif (diff_words->current_plus != diff_words->plus.text.ptr +\n+\t\t\tdiff_words->plus.text.size)\n+\t\tfwrite(diff_words->current_plus,\n+\t\t\tdiff_words->plus.text.ptr + diff_words->plus.text.size\n+\t\t\t- diff_words->current_plus, 1,\n+\t\t\tdiff_words->file);\n \tdiff_words->minus.text.size = diff_words->plus.text.size = 0;\n-\n-\tif (diff_words->minus.suppressed_newline) {\n-\t\tputc('\\n', diff_words->file);\n-\t\tdiff_words->minus.suppressed_newline = 0;\n-\t}\n }\n \n typedef unsigned long (*sane_truncate_fn)(char *line, unsigned long len);\n@@ -458,7 +473,9 @@ static void free_diff_words_data(struct emit_callback *ecbdata)\n \t\t\tdiff_words_show(ecbdata->diff_words);\n \n \t\tfree (ecbdata->diff_words->minus.text.ptr);\n+\t\tfree (ecbdata->diff_words->minus.orig);\n \t\tfree (ecbdata->diff_words->plus.text.ptr);\n+\t\tfree (ecbdata->diff_words->plus.orig);\n \t\tfree(ecbdata->diff_words);\n \t\tecbdata->diff_words = NULL;\n \t}\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nnew file mode 100755\nindex 0000000..b22195f\n--- /dev/null\n+++ b/t/t4034-diff-words.sh\n@@ -0,0 +1,66 @@\n+#!/bin/sh\n+\n+test_description='word diff colors'\n+\n+. ./test-lib.sh\n+\n+test_expect_success setup '\n+\n+\tgit config diff.color.old red\n+\tgit config diff.color.new green\n+\n+'\n+\n+decrypt_color () {\n+\tsed \\\n+\t\t-e 's/.\\[1m/<WHITE>/g' \\\n+\t\t-e 's/.\\[31m/<RED>/g' \\\n+\t\t-e 's/.\\[32m/<GREEN>/g' \\\n+\t\t-e 's/.\\[36m/<BROWN>/g' \\\n+\t\t-e 's/.\\[m/<RESET>/g'\n+}\n+\n+word_diff () {\n+\ttest_must_fail git diff --no-index \"$@\" pre post > output &&\n+\tdecrypt_color < output > output.decrypted &&\n+\ttest_cmp expect output.decrypted\n+}\n+\n+cat > pre <<\\EOF\n+h(4)\n+\n+a = b + c\n+EOF\n+\n+cat > post <<\\EOF\n+h(4),hh[44]\n+\n+a = b + c\n+\n+aa = a\n+\n+aeff = aeff * ( aaa )\n+EOF\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 330b04f..5ed8eff 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1,3 +1,7 @@<RESET>\n+<RED>h(4)<RESET><GREEN>h(4),hh[44]<RESET>\n+<RESET>\n+a = b + c<RESET>\n+\n+<GREEN>aa = a<RESET>\n+\n+<GREEN>aeff = aeff * ( aaa )<RESET>\n+EOF\n+\n+test_expect_success 'word diff with runs of whitespace' '\n+\n+\tword_diff --color-words\n+\n+'\n+\n+test_done\n-- \n1.6.1.295.g5d331\n\n"},{"id":"100479","messageId":"alpine.DEB.1.00.0901142145200.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901142142120.3586@pacific.mpi-cbg.de","subject":"[PATCH replacement for take 3 4/4] color-words: take an optional regular expression describing words","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T20:46:14Z","receivedAt":"2009-01-14T20:46:14Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nIn some applications, words are not delimited by white space.  To\nallow for that, you can specify a regular expression describing\nwhat makes a word with\n\n\tgit diff --color-words='[A-Za-z0-9]+'\n\nNote that words cannot contain newline characters.\n\nAs suggested by Thomas Rast, the words are the exact matches of the\nregular expression.\n\nNote that a regular expression beginning with a '^' will match only\na word at the beginning of the hunk, not a word at the beginning of\na line, and is probably not what you want.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n\n\tThis basically contains the fix I sent earlier.\n\n\tAs for the documentation, I would not have any issue with your \n\tpatch replacing my documentation in favor of yours.\n\n Documentation/diff-options.txt |    6 +++-\n diff.c                         |   64 ++++++++++++++++++++++++++++++++++-----\n diff.h                         |    1 +\n t/t4034-diff-words.sh          |   39 ++++++++++++++++++++++++\n 4 files changed, 100 insertions(+), 10 deletions(-)\n\ndiff --git a/Documentation/diff-options.txt b/Documentation/diff-options.txt\nindex 1f8ce97..e546bfa 100644\n--- a/Documentation/diff-options.txt\n+++ b/Documentation/diff-options.txt\n@@ -94,8 +94,12 @@ endif::git-format-patch[]\n \tTurn off colored diff, even when the configuration file\n \tgives the default to color output.\n \n---color-words::\n+--color-words[=regex]::\n \tShow colored word diff, i.e. color words which have changed.\n++\n+Optionally, you can pass a regular expression that tells Git what the\n+words are that you are looking for; The default is to interpret any\n+stretch of non-whitespace as a word.\n \n --no-renames::\n \tTurn off rename detection, even when the configuration\ndiff --git a/diff.c b/diff.c\nindex fe8b1f0..1408717 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -333,12 +333,14 @@ static void diff_words_append(char *line, unsigned long len,\n \tlen--;\n \tmemcpy(buffer->text.ptr + buffer->text.size, line, len);\n \tbuffer->text.size += len;\n+\tbuffer->text.ptr[buffer->text.size] = '\\0';\n }\n \n struct diff_words_data {\n \tstruct diff_words_buffer minus, plus;\n \tconst char *current_plus;\n \tFILE *file;\n+\tregex_t *word_regex;\n };\n \n static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n@@ -374,17 +376,49 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \tdiff_words->current_plus = plus_end;\n }\n \n+/* This function starts looking at *begin, and returns 0 iff a word was found. */\n+static int find_word_boundaries(mmfile_t *buffer, regex_t *word_regex,\n+\t\tint *begin, int *end)\n+{\n+\tif (word_regex && *begin < buffer->size) {\n+\t\tregmatch_t match[1];\n+\t\tif (!regexec(word_regex, buffer->ptr + *begin, 1, match, 0)) {\n+\t\t\tchar *p = memchr(buffer->ptr + *begin + match[0].rm_so,\n+\t\t\t\t\t'\\n', match[0].rm_eo - match[0].rm_so);\n+\t\t\t*end = p ? p - buffer->ptr : match[0].rm_eo + *begin;\n+\t\t\t*begin += match[0].rm_so;\n+\t\t\treturn *begin >= *end;\n+\t\t}\n+\t\treturn -1;\n+\t}\n+\n+\t/* find the next word */\n+\twhile (*begin < buffer->size && isspace(buffer->ptr[*begin]))\n+\t\t(*begin)++;\n+\tif (*begin >= buffer->size)\n+\t\treturn -1;\n+\n+\t/* find the end of the word */\n+\t*end = *begin + 1;\n+\twhile (*end < buffer->size && !isspace(buffer->ptr[*end]))\n+\t\t(*end)++;\n+\n+\treturn 0;\n+}\n+\n /*\n  * This function splits the words in buffer->text, stores the list with\n  * newline separator into out, and saves the offsets of the original words\n  * in buffer->orig.\n  */\n-static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n+static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out,\n+\t\tregex_t *word_regex)\n {\n \tint i, j;\n+\tlong alloc = 0;\n \n \tout->size = 0;\n-\tout->ptr = xmalloc(buffer->text.size);\n+\tout->ptr = NULL;\n \n \t/* fake an empty \"0th\" word */\n \tALLOC_GROW(buffer->orig, 1, buffer->orig_alloc);\n@@ -392,11 +426,8 @@ static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n \tbuffer->orig_nr = 1;\n \n \tfor (i = 0; i < buffer->text.size; i++) {\n-\t\tif (isspace(buffer->text.ptr[i]))\n-\t\t\tcontinue;\n-\t\tfor (j = i + 1; j < buffer->text.size &&\n-\t\t\t\t!isspace(buffer->text.ptr[j]); j++)\n-\t\t\t; /* find the end of the word */\n+\t\tif (find_word_boundaries(&buffer->text, word_regex, &i, &j))\n+\t\t\treturn;\n \n \t\t/* store original boundaries */\n \t\tALLOC_GROW(buffer->orig, buffer->orig_nr + 1,\n@@ -406,6 +437,7 @@ static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n \t\tbuffer->orig_nr++;\n \n \t\t/* store one word */\n+\t\tALLOC_GROW(out->ptr, out->size + j - i + 1, alloc);\n \t\tmemcpy(out->ptr + out->size, buffer->text.ptr + i, j - i);\n \t\tout->ptr[out->size + j - i] = '\\n';\n \t\tout->size += j - i + 1;\n@@ -435,9 +467,10 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n-\tdiff_words_fill(&diff_words->minus, &minus);\n-\tdiff_words_fill(&diff_words->plus, &plus);\n+\tdiff_words_fill(&diff_words->minus, &minus, diff_words->word_regex);\n+\tdiff_words_fill(&diff_words->plus, &plus, diff_words->word_regex);\n \txpp.flags = XDF_NEED_MINIMAL;\n+\t/* as only the hunk header will be parsed, we need a 0-context */\n \txecfg.ctxlen = 0;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n \t\t      &xpp, &xecfg, &ecb);\n@@ -476,6 +509,7 @@ static void free_diff_words_data(struct emit_callback *ecbdata)\n \t\tfree (ecbdata->diff_words->minus.orig);\n \t\tfree (ecbdata->diff_words->plus.text.ptr);\n \t\tfree (ecbdata->diff_words->plus.orig);\n+\t\tfree(ecbdata->diff_words->word_regex);\n \t\tfree(ecbdata->diff_words);\n \t\tecbdata->diff_words = NULL;\n \t}\n@@ -1498,6 +1532,14 @@ static void builtin_diff(const char *name_a,\n \t\t\tecbdata.diff_words =\n \t\t\t\txcalloc(1, sizeof(struct diff_words_data));\n \t\t\tecbdata.diff_words->file = o->file;\n+\t\t\tif (o->word_regex) {\n+\t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n+\t\t\t\t\txmalloc(sizeof(regex_t));\n+\t\t\t\tif (regcomp(ecbdata.diff_words->word_regex,\n+\t\t\t\t\t\to->word_regex, REG_EXTENDED))\n+\t\t\t\t\tdie (\"Invalid regular expression: %s\",\n+\t\t\t\t\t\t\to->word_regex);\n+\t\t\t}\n \t\t}\n \t\txdi_diff_outf(&mf1, &mf2, fn_out_consume, &ecbdata,\n \t\t\t      &xpp, &xecfg, &ecb);\n@@ -2513,6 +2555,10 @@ int diff_opt_parse(struct diff_options *options, const char **av, int ac)\n \t\tDIFF_OPT_CLR(options, COLOR_DIFF);\n \telse if (!strcmp(arg, \"--color-words\"))\n \t\toptions->flags |= DIFF_OPT_COLOR_DIFF | DIFF_OPT_COLOR_DIFF_WORDS;\n+\telse if (!prefixcmp(arg, \"--color-words=\")) {\n+\t\toptions->flags |= DIFF_OPT_COLOR_DIFF | DIFF_OPT_COLOR_DIFF_WORDS;\n+\t\toptions->word_regex = arg + 14;\n+\t}\n \telse if (!strcmp(arg, \"--exit-code\"))\n \t\tDIFF_OPT_SET(options, EXIT_WITH_STATUS);\n \telse if (!strcmp(arg, \"--quiet\"))\ndiff --git a/diff.h b/diff.h\nindex 4d5a327..23cd90c 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -98,6 +98,7 @@ struct diff_options {\n \n \tint stat_width;\n \tint stat_name_width;\n+\tconst char *word_regex;\n \n \t/* this is set by diffcore for DIFF_FORMAT_PATCH */\n \tint found_changes;\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex b22195f..f4810e9 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -63,4 +63,43 @@ test_expect_success 'word diff with runs of whitespace' '\n \n '\n \n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 330b04f..5ed8eff 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1,3 +1,7 @@<RESET>\n+h(4),<GREEN>hh<RESET>[44]\n+<RESET>\n+a = b + c<RESET>\n+\n+<GREEN>aa = a<RESET>\n+\n+<GREEN>aeff = aeff * ( aaa<RESET> )\n+EOF\n+\n+test_expect_success 'word diff with a regular expression' '\n+\n+\tword_diff --color-words='[a-z]+'\n+\n+'\n+\n+echo 'aaa (aaa)' > pre\n+echo 'aaa (aaa) aaa' > post\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index c29453b..be22f37 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1 +1 @@<RESET>\n+aaa (aaa) <GREEN>aaa<RESET>\n+EOF\n+\n+test_expect_success \"test parsing words for newline\" '\n+\n+\tword_diff --color-words='a+'\n+\n+'\n+\n test_done\n-- \n1.6.1.295.g5d331\n"},{"id":"100481","messageId":"alpine.DEB.1.00.0901142203190.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"200901142104.16134.trast@student.ethz.ch","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T21:07:48Z","receivedAt":"2009-01-14T21:07:48Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 14 Jan 2009, Thomas Rast wrote:\n\n> Johannes Schindelin wrote:\n> > \n> > The only \"funny\" thing I realized is that the lines which are output\n> > by emit_line() add a RESET at the end of the line, and I do not do that\n> > in color_fwrite_lines().\n> \n> Umm.... but you seem to do?\n\nOh, right!  I think the culprit is in fn_out_diff_words_aux(), which calls \nfwrite() directly for the common words.\n\n> Ack on the new regex semantics, though I'd have implemented it via dying \n> on '\\n' instead of silently splitting there (and restarting a new \n> match!).\n\nHmm.  I'd rather not die() in the middle of it.\n\nMaybe we can even handle newlines correctly by replacing them with NULs \nwhich libxdiff handles just fine?\n\n> Thus, Ack on 4/4 once the boundary bug is fixed.  Thanks for your work!\n\nPhew.  I was almost convinced you would hate me for my criticiscm.\n\nThanks,\nDscho\n"},{"id":"100488","messageId":"alpine.DEB.1.00.0901142258250.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"200901142059.09005.trast@student.ethz.ch","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T22:06:48Z","receivedAt":"2009-01-14T22:06:48Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 14 Jan 2009, Thomas Rast wrote:\n\n> Teemu Likonen wrote:\n> >     -aaa (aaa)\n> >     +aaa (aaa) aaa\n> \n> Bug aside, examples like this one make me wonder if we should force a\n> \"last resort\" match for `[^[:space:]]`.  For example,\n> \n>       -aaa [aaa]\n>       +aaa (aaa) aaa\n> \n> would still give you\n> \n>       aaa (aaa)<GREEN> aaa<RESET>\n> \n> which may be unexpected.\n\nBut why should it be unexpected?  If people say that every length of \"a\" \nmakes a word, and consequently everything else is clutter, then that's \nthat, no?\n\nSo people might be surprised, but then they should have said something \nlike\n\n\t[-.+#@\"'$%^&*([{<>~|]*[A-Za-z][A-Za-z0-9]*[-.+#@\"'$%&*)\\]}>|]*\n\ninstead.\n\nAlthough I have to say that for some applications, it is a pity that \neven POSIX extended regular expressions knows neither lookahead nor \nlookbehind.\n\nWhich reminds me... should we activate REG_EXTENDED by default?\n\nCiao,\nDscho\n"},{"id":"100490","messageId":"200901142311.37342.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901142258250.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T22:11:32Z","receivedAt":"2009-01-14T22:11:32Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> Which reminds me... should we activate REG_EXTENDED by default?\n\nWe (you :-) do, and I think so.  Consider that funcname is not even\ndocumented any more, in favour of xfuncname.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n"},{"id":"100492","messageId":"200901141624.38315.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901142258250.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-14T22:24:32Z","receivedAt":"2009-01-14T22:24:32Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"On Wednesday 2009 January 14 16:06:48 Johannes Schindelin wrote:\n>On Wed, 14 Jan 2009, Thomas Rast wrote:\n>> Bug aside, examples like this one make me wonder if we should force a\n>> \"last resort\" match for `[^[:space:]]`.  For example,\n>>\n>>       -aaa [aaa]\n>>       +aaa (aaa) aaa\n>>\n>> would still give you\n>>\n>>       aaa (aaa)<GREEN> aaa<RESET>\n>>\n>> which may be unexpected.\n>\n>But why should it be unexpected?  If people say that every length of \"a\"\n>makes a word, and consequently everything else is clutter, then that's\n>that, no?\n\nI think some people are going to have problems with the strict dichotomy \nbetween \"part of a word\" and \"ignorable whitespace\" that is being set up.  \nIt makes sense technically, but it could confuse.\n\nImagine with --diff-words=[A-Z][A-Za-z]* and the following change:\n-To be Or Not To be.\n+To ignore Or Not To treat whitespace differently.\n\nI think there is value in being able to ignore anything that's not a word, \nso the documentation that mentions adding '|[^[:space:]]' to your regex \nseems sufficient to me.\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"100493","messageId":"3ff3ccf6e3c1cd6a002d200aee5df88a197a7bf6.1231971446.git.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901142104400.3586@pacific.mpi-cbg.de","subject":"[PATCH 1/4] color-words: fix quoting in t4034","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T22:26:00Z","receivedAt":"2009-01-14T22:26:00Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Since the single quotes match the ones used to quote the test text\nitself, they'd be dropped.  Use double quotes instead.\n---\n\nI'd squash this into Dscho's 4/4, so no SoB.\n\n\n t/t4034-diff-words.sh |    4 ++--\n 1 files changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex f4810e9..6ad1c1f 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -80,7 +80,7 @@ EOF\n \n test_expect_success 'word diff with a regular expression' '\n \n-\tword_diff --color-words='[a-z]+'\n+\tword_diff --color-words=\"[a-z]+\"\n \n '\n \n@@ -98,7 +98,7 @@ EOF\n \n test_expect_success \"test parsing words for newline\" '\n \n-\tword_diff --color-words='a+'\n+\tword_diff --color-words=\"a+\"\n \n '\n \n-- \n1.6.1.142.ge070e\n"},{"id":"100494","messageId":"48504e8a330beca560208ce050d43bc92ac04c90.1231971446.git.trast@student.ethz.ch","threadId":"17092","inReplyTo":"3ff3ccf6e3c1cd6a002d200aee5df88a197a7bf6.1231971446.git.trast@student.ethz.ch","subject":"[PATCH 2/4] color-words: enable REG_NEWLINE to help user","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T22:26:01Z","receivedAt":"2009-01-14T22:26:01Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"We silently truncate a match at the newline, which may lead to\nunexpected behaviour, e.g., when matching \"<[^>]*>\" against\n\n  <foo\n  bar>\n\nsince then \"<foo\" becomes a word (and \"bar>\" doesn't!) even though the\nregex said only angle-bracket-delimited things can be words.\n\nTo alleviate the problem slightly, use REG_NEWLINE so that negated\nclasses can't match a newline.  Of course newlines can still be\nmatched explicitly.\n\nSigned-off-by: Thomas Rast <trast@student.ethz.ch>\n---\n diff.c |    3 ++-\n 1 files changed, 2 insertions(+), 1 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex cc42adf..3f07ac1 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -1536,7 +1536,8 @@ static void builtin_diff(const char *name_a,\n \t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n \t\t\t\t\txmalloc(sizeof(regex_t));\n \t\t\t\tif (regcomp(ecbdata.diff_words->word_regex,\n-\t\t\t\t\t\to->word_regex, REG_EXTENDED))\n+\t\t\t\t\t\to->word_regex,\n+\t\t\t\t\t\tREG_EXTENDED | REG_NEWLINE))\n \t\t\t\t\tdie (\"Invalid regular expression: %s\",\n \t\t\t\t\t\t\to->word_regex);\n \t\t\t}\n-- \n1.6.1.142.ge070e\n"},{"id":"100495","messageId":"b1290f83267e64856e58477e0c19e920dd416c82.1231971446.git.trast@student.ethz.ch","threadId":"17092","inReplyTo":"48504e8a330beca560208ce050d43bc92ac04c90.1231971446.git.trast@student.ethz.ch","subject":"[PATCH 3/4] color-words: expand docs with precise semantics","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T22:26:02Z","receivedAt":"2009-01-14T22:26:02Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Signed-off-by: Thomas Rast <trast@student.ethz.ch>\n---\n Documentation/diff-options.txt |   15 ++++++++++-----\n 1 files changed, 10 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/diff-options.txt b/Documentation/diff-options.txt\nindex 2c1fa4b..8689a92 100644\n--- a/Documentation/diff-options.txt\n+++ b/Documentation/diff-options.txt\n@@ -91,12 +91,17 @@ endif::git-format-patch[]\n \tTurn off colored diff, even when the configuration file\n \tgives the default to color output.\n \n---color-words[=regex]::\n-\tShow colored word diff, i.e. color words which have changed.\n+--color-words[=<regex>]::\n+\tShow colored word diff, i.e., color words which have changed.\n+\tBy default, words are separated by whitespace.\n +\n-Optionally, you can pass a regular expression that tells Git what the\n-words are that you are looking for; The default is to interpret any\n-stretch of non-whitespace as a word.\n+When a <regex> is specified, every non-overlapping match of the\n+<regex> is considered a word.  Anything between these matches is\n+considered whitespace and ignored(!) for the purposes of finding\n+differences.  You may want to append `|[^[:space:]]` to your regular\n+expression to make sure that it matches all non-whitespace characters.\n+A match that contains a newline is silently truncated(!) at the\n+newline.\n \n --no-renames::\n \tTurn off rename detection, even when the configuration\n-- \n1.6.1.142.ge070e\n"},{"id":"100496","messageId":"b404fdfe0f5af535b35d1f239a68f6a7911ede19.1231971446.git.trast@student.ethz.ch","threadId":"17092","inReplyTo":"b1290f83267e64856e58477e0c19e920dd416c82.1231971446.git.trast@student.ethz.ch","subject":"[PATCH 4/4] color-words: make regex configurable via attributes","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T22:26:03Z","receivedAt":"2009-01-14T22:26:03Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Make the --color-words splitting regular expression configurable via\nthe diff driver's 'wordregex' attribute.  The user can then set the\ndriver on a file in .gitattributes.  If a regex is given on the\ncommand line, it overrides the driver's setting.\n\nWe also provide built-in regexes for the languages that already had\nfuncname patterns, and add an appropriate diff driver entry for C/++.\n(The patterns are designed to run UTF-8 sequences into a single chunk\nto make sure they remain readable.)\n\nSigned-off-by: Thomas Rast <trast@student.ethz.ch>\n---\n\nIncorporates the last round of Dscho's suggestions.\n\n Documentation/diff-options.txt  |    4 ++\n Documentation/gitattributes.txt |   21 ++++++++++\n diff.c                          |   10 +++++\n t/t4034-diff-words.sh           |   49 ++++++++++++++++++++++--\n userdiff.c                      |   78 +++++++++++++++++++++++++++++++-------\n userdiff.h                      |    1 +\n 6 files changed, 144 insertions(+), 19 deletions(-)\n\ndiff --git a/Documentation/diff-options.txt b/Documentation/diff-options.txt\nindex 8689a92..1edb82e 100644\n--- a/Documentation/diff-options.txt\n+++ b/Documentation/diff-options.txt\n@@ -102,6 +102,10 @@ differences.  You may want to append `|[^[:space:]]` to your regular\n expression to make sure that it matches all non-whitespace characters.\n A match that contains a newline is silently truncated(!) at the\n newline.\n++\n+The regex can also be set via a diff driver, see\n+linkgit:gitattributes[1]; giving it explicitly overrides any diff\n+driver setting.\n \n --no-renames::\n \tTurn off rename detection, even when the configuration\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 8af22ec..17707ba 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -317,6 +317,8 @@ patterns are available:\n \n - `bibtex` suitable for files with BibTeX coded references.\n \n+- `cpp` suitable for source code in the C and C++ languages.\n+\n - `html` suitable for HTML/XHTML documents.\n \n - `java` suitable for source code in the Java language.\n@@ -334,6 +336,25 @@ patterns are available:\n - `tex` suitable for source code for LaTeX documents.\n \n \n+Customizing word diff\n+^^^^^^^^^^^^^^^^^^^^^\n+\n+You can customize the rules that `git diff --color-words` uses to\n+split words in a line, by specifying an appropriate regular expression\n+in the \"diff.*.wordregex\" configuration variable.  For example, in TeX\n+a backslash followed by a sequence of letters forms a command, but\n+several such commands can be run together without intervening\n+whitespace.  To separate them, use a regular expression such as\n+\n+------------------------\n+[diff \"tex\"]\n+\twordregex = \"\\\\\\\\[a-zA-Z]+|[{}]|\\\\\\\\.|[^\\\\{}[:space:]]+\"\n+------------------------\n+\n+A built-in pattern is provided for all languages listed in the last\n+section.\n+\n+\n Performing text diffs of binary files\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \ndiff --git a/diff.c b/diff.c\nindex 3f07ac1..0e82e18 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -1372,6 +1372,12 @@ int diff_filespec_is_binary(struct diff_filespec *one)\n \treturn one->driver->funcname.pattern ? &one->driver->funcname : NULL;\n }\n \n+static const char *userdiff_word_regex(struct diff_filespec *one)\n+{\n+\tdiff_filespec_load_driver(one);\n+\treturn one->driver->word_regex;\n+}\n+\n void diff_set_mnemonic_prefix(struct diff_options *options, const char *a, const char *b)\n {\n \tif (!options->a_prefix)\n@@ -1532,6 +1538,10 @@ static void builtin_diff(const char *name_a,\n \t\t\tecbdata.diff_words =\n \t\t\t\txcalloc(1, sizeof(struct diff_words_data));\n \t\t\tecbdata.diff_words->file = o->file;\n+\t\t\tif (!o->word_regex)\n+\t\t\t\to->word_regex = userdiff_word_regex(one);\n+\t\t\tif (!o->word_regex)\n+\t\t\t\to->word_regex = userdiff_word_regex(two);\n \t\t\tif (o->word_regex) {\n \t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n \t\t\t\t\txmalloc(sizeof(regex_t));\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 6ad1c1f..631ca44 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -22,8 +22,10 @@ decrypt_color () {\n \n word_diff () {\n \ttest_must_fail git diff --no-index \"$@\" pre post > output &&\n-\tdecrypt_color < output > output.decrypted &&\n-\ttest_cmp expect output.decrypted\n+\tdecrypt_color < output > output.decrypted\n+}\n+word_diff_check () {\n+\ttest_cmp \"$1\" output.decrypted\n }\n \n cat > pre <<\\EOF\n@@ -80,7 +82,45 @@ EOF\n \n test_expect_success 'word diff with a regular expression' '\n \n-\tword_diff --color-words=\"[a-z]+\"\n+\tword_diff --color-words=\"[a-z]+\" &&\n+\tword_diff_check expect\n+\n+'\n+\n+cat > expect-by-chars <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 330b04f..5ed8eff 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1,3 +1,7 @@<RESET>\n+h(4)<GREEN>,hh[44]<RESET>\n+<RESET>\n+a = b + c<RESET>\n+\n+<GREEN>aa = a<RESET>\n+\n+<GREEN>aeff = aeff * ( aaa )<RESET>\n+EOF\n+\n+test_expect_success 'set a diff driver' '\n+\tgit config diff.testdriver.wordregex \"[^[:space:]]\" &&\n+\tcat <<EOF > .gitattributes\n+pre diff=testdriver\n+post diff=testdriver\n+EOF\n+'\n+\n+test_expect_success 'use default supplied by driver' '\n+\n+\tword_diff --color-words &&\n+\tword_diff_check expect-by-chars\n+\n+'\n+\n+test_expect_success 'option overrides default' '\n+\n+\tword_diff --color-words=\"[a-z]+\" &&\n+\tword_diff_check expect\n \n '\n \n@@ -98,7 +138,8 @@ EOF\n \n test_expect_success \"test parsing words for newline\" '\n \n-\tword_diff --color-words=\"a+\"\n+\tword_diff --color-words=\"a+\" &&\n+\tword_diff_check expect\n \n '\n \ndiff --git a/userdiff.c b/userdiff.c\nindex 3681062..dbfda6d 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -6,14 +6,20 @@\n static int ndrivers;\n static int drivers_alloc;\n \n-#define FUNCNAME(name, pattern) \\\n-\t{ name, NULL, -1, { pattern, REG_EXTENDED } }\n+#define PATTERNS(name, pattern, wordregex)\t\t\t\\\n+\t{ name, NULL, -1, { pattern, REG_EXTENDED }, wordregex }\n static struct userdiff_driver builtin_drivers[] = {\n-FUNCNAME(\"html\", \"^[ \\t]*(<[Hh][1-6][ \\t].*>.*)$\"),\n-FUNCNAME(\"java\",\n+PATTERNS(\"html\", \"^[ \\t]*(<[Hh][1-6][ \\t].*>.*)$\",\n+\t \"[^<>= \\t]+|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"java\",\n \t \"!^[ \\t]*(catch|do|for|if|instanceof|new|return|switch|throw|while)\\n\"\n-\t \"^[ \\t]*(([ \\t]*[A-Za-z_][A-Za-z_0-9]*){2,}[ \\t]*\\\\([^;]*)$\"),\n-FUNCNAME(\"objc\",\n+\t \"^[ \\t]*(([ \\t]*[A-Za-z_][A-Za-z_0-9]*){2,}[ \\t]*\\\\([^;]*)$\",\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[fFlL]?|0[xXbB]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=\"\n+\t \"|--|\\\\+\\\\+|<<=?|>>>?=?|&&|\\\\|\\\\|\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"objc\",\n \t /* Negate C statements that can look like functions */\n \t \"!^[ \\t]*(do|for|if|else|return|switch|while)\\n\"\n \t /* Objective-C methods */\n@@ -21,20 +27,60 @@\n \t /* C functions */\n \t \"^[ \\t]*(([ \\t]*[A-Za-z_][A-Za-z_0-9]*){2,}[ \\t]*\\\\([^;]*)$\\n\"\n \t /* Objective-C class/protocol definitions */\n-\t \"^(@(implementation|interface|protocol)[ \\t].*)$\"),\n-FUNCNAME(\"pascal\",\n+\t \"^(@(implementation|interface|protocol)[ \\t].*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[fFlL]?|0[xXbB]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=|--|\\\\+\\\\+|<<=?|>>=?|&&|\\\\|\\\\||::|->\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"pascal\",\n \t \"^((procedure|function|constructor|destructor|interface|\"\n \t\t\"implementation|initialization|finalization)[ \\t]*.*)$\"\n \t \"\\n\"\n-\t \"^(.*=[ \\t]*(class|record).*)$\"),\n-FUNCNAME(\"php\", \"^[\\t ]*((function|class).*)\"),\n-FUNCNAME(\"python\", \"^[ \\t]*((class|def)[ \\t].*)$\"),\n-FUNCNAME(\"ruby\", \"^[ \\t]*((class|module|def)[ \\t].*)$\"),\n-FUNCNAME(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\"),\n-FUNCNAME(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\"),\n+\t \"^(.*=[ \\t]*(class|record).*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+|0[xXbB]?[0-9a-fA-F]+\"\n+\t \"|<>|<=|>=|:=|\\\\.\\\\.\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"php\", \"^[\\t ]*((function|class).*)\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+|0[xXbB]?[0-9a-fA-F]+\"\n+\t \"|[-+*/<>%&^|=!.]=|--|\\\\+\\\\+|<<=?|>>=?|===|&&|\\\\|\\\\||::|->\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"python\", \"^[ \\t]*((class|def)[ \\t].*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[jJlL]?|0[xX]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=|//=?|<<=?|>>=?|\\\\*\\\\*=?\"\n+\t \"|[^[:space:]|[\\x80-\\xff]+\"),\n+\t /* -- */\n+PATTERNS(\"ruby\", \"^[ \\t]*((class|module|def)[ \\t].*)$\",\n+\t /* -- */\n+\t \"(@|@@|\\\\$)?[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+|0[xXbB]?[0-9a-fA-F]+|\\\\?(\\\\\\\\C-)?(\\\\\\\\M-)?.\"\n+\t \"|//=?|[-+*/<>%&^|=!]=|<<=?|>>=?|===|\\\\.{1,3}|::|[!=]~\"\n+\t \"|[^[:space:]|[\\x80-\\xff]+\"),\n+PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n+\t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n+PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n+\t \"\\\\\\\\[a-zA-Z@]+|[{}]|\\\\\\\\.|[^\\\\{} \\t]+\"),\n+PATTERNS(\"cpp\",\n+\t /* Jump targets or access declarations */\n+\t \"!^[ \\t]*[A-Za-z_][A-Za-z_0-9]*:.*$\\n\"\n+\t /* C functions at top level */\n+\t \"^([A-Za-z_][A-Za-z_0-9]*([ \\t]+[A-Za-z_][A-Za-z_0-9]*){1,}[ \\t]*\\\\([^;]*)$\\n\"\n+\t /* compound type at top level */\n+\t \"^((struct|class|enum)[^;]*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[fFlL]?|0[xXbB]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=|--|\\\\+\\\\+|<<=?|>>=?|&&|\\\\|\\\\||::|->\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n { \"default\", NULL, -1, { NULL, 0 } },\n };\n-#undef FUNCNAME\n+#undef PATTERNS\n \n static struct userdiff_driver driver_true = {\n \t\"diff=true\",\n@@ -134,6 +180,8 @@ int userdiff_config(const char *k, const char *v)\n \t\treturn parse_string(&drv->external, k, v);\n \tif ((drv = parse_driver(k, v, \"textconv\")))\n \t\treturn parse_string(&drv->textconv, k, v);\n+\tif ((drv = parse_driver(k, v, \"wordregex\")))\n+\t\treturn parse_string(&drv->word_regex, k, v);\n \n \treturn 0;\n }\ndiff --git a/userdiff.h b/userdiff.h\nindex ba29457..c315159 100644\n--- a/userdiff.h\n+++ b/userdiff.h\n@@ -11,6 +11,7 @@ struct userdiff_driver {\n \tconst char *external;\n \tint binary;\n \tstruct userdiff_funcname funcname;\n+\tconst char *word_regex;\n \tconst char *textconv;\n };\n \n-- \n1.6.1.142.ge070e\n"},{"id":"100499","messageId":"200901142337.33320.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901142203190.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-14T22:37:08Z","receivedAt":"2009-01-14T22:37:08Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> > Ack on the new regex semantics, though I'd have implemented it via dying \n> > on '\\n' instead of silently splitting there (and restarting a new \n> > match!).\n> \n> Hmm.  I'd rather not die() in the middle of it.\n> \n> Maybe we can even handle newlines correctly by replacing them with NULs \n> which libxdiff handles just fine?\n\nI'm not sure it's worth the effort---anyone who wants words to stick\ntogether across newlines probably doesn't put a newline there in the\nfirst place, don't they?  (And it just shifts the problem to another\nspecial character.)\n\n> Phew.  I was almost convinced you would hate me for my criticiscm.\n\nLet's say I wasn't too happy when you asked for two rounds of\nimprovements and _then_ rejected.  But the end result certainly turned\nout better, so the criticism was justified.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n\n"},{"id":"100501","messageId":"alpine.DEB.1.00.0901142341420.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"3ff3ccf6e3c1cd6a002d200aee5df88a197a7bf6.1231971446.git.trast@student.ethz.ch","subject":"Re: [PATCH 1/4] color-words: fix quoting in t4034","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-14T22:41:59Z","receivedAt":"2009-01-14T22:41:59Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 14 Jan 2009, Thomas Rast wrote:\n\n> Since the single quotes match the ones used to quote the test text\n> itself, they'd be dropped.  Use double quotes instead.\n\nSee, I suck with quoting.\n\n> ---\n> \n> I'd squash this into Dscho's 4/4, so no SoB.\n\nSure, done.\n\nThanks,\nDscho\n"},{"id":"100517","messageId":"200901150132.14106.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901142145200.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH replacement for take 3 4/4] color-words: take an optional regular expression describing words","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-15T00:32:11Z","receivedAt":"2009-01-15T00:32:11Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> \tThis basically contains the fix I sent earlier.\n\nUnfortunately I found another case where it breaks.  It even comes\nwith a fairly neat test case:\n\n  $ g diff --no-index test_a test_b\n  diff --git 1/test_a 2/test_b\n  index 289cb9d..2d06f37 100644\n  --- 1/test_a\n  +++ 2/test_b\n  @@ -1 +1 @@\n  -(:\n  +(\n  $ g diff --no-index --color-words='.' test_a test_b\n  diff --git 1/test_a 2/test_b\n  index 289cb9d..2d06f37 100644\n  --- 1/test_a\n  +++ 2/test_b\n  @@ -1 +1 @@\n  :(\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n\n"},{"id":"100529","messageId":"alpine.DEB.1.00.0901150211120.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"200901150132.14106.trast@student.ethz.ch","subject":"Re: [PATCH replacement for take 3 4/4] color-words: take an optional regular expression describing words","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-15T01:12:59Z","receivedAt":"2009-01-15T01:12:59Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 15 Jan 2009, Thomas Rast wrote:\n\n> Johannes Schindelin wrote:\n> > \tThis basically contains the fix I sent earlier.\n> \n> Unfortunately I found another case where it breaks.  It even comes\n> with a fairly neat test case:\n> \n>   $ g diff --no-index test_a test_b\n>   diff --git 1/test_a 2/test_b\n>   index 289cb9d..2d06f37 100644\n>   --- 1/test_a\n>   +++ 2/test_b\n>   @@ -1 +1 @@\n>   -(:\n>   +(\n\nThe diff of the words would look like this:\n\ndiff --git a/a1 b/a2\nindex 8309acb..2d06f37 100644\n--- a/a1\n+++ b/a2\n@@ -2 +1,0 @@\n-:\n\n\nNotice the \"+1,0\"?  I fully expected this to be \"+2,0\", but apparently I \nwas mistaken...\n\nCan anybody explain to me why this is so?\n\nCiao,\nDscho\n"},{"id":"100530","messageId":"alpine.DEB.1.00.0901150233121.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"b404fdfe0f5af535b35d1f239a68f6a7911ede19.1231971446.git.trast@student.ethz.ch","subject":"Re: [PATCH 4/4] color-words: make regex configurable via attributes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-15T01:33:49Z","receivedAt":"2009-01-15T01:33:49Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Thomas,\n\ncould you please squash this in?\n\n-- snipsnap --\n[PATCH to be squashed into the attributes patch] Decomplicate t4034 again\n\n---\n t/t4034-diff-words.sh |   50 ++++++++++++++++++++++--------------------------\n 1 files changed, 23 insertions(+), 27 deletions(-)\n\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 631ca44..07e48d1 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -22,10 +22,8 @@ decrypt_color () {\n \n word_diff () {\n \ttest_must_fail git diff --no-index \"$@\" pre post > output &&\n-\tdecrypt_color < output > output.decrypted\n-}\n-word_diff_check () {\n-\ttest_cmp \"$1\" output.decrypted\n+\tdecrypt_color < output > output.decrypted &&\n+\ttest_cmp expect output.decrypted\n }\n \n cat > pre <<\\EOF\n@@ -82,12 +80,25 @@ EOF\n \n test_expect_success 'word diff with a regular expression' '\n \n-\tword_diff --color-words=\"[a-z]+\" &&\n-\tword_diff_check expect\n+\tword_diff --color-words=\"[a-z]+\"\n \n '\n \n-cat > expect-by-chars <<\\EOF\n+test_expect_success 'set a diff driver' '\n+\tgit config diff.testdriver.wordregex \"[^[:space:]]\" &&\n+\tcat <<EOF > .gitattributes\n+pre diff=testdriver\n+post diff=testdriver\n+EOF\n+'\n+\n+test_expect_success 'option overrides default' '\n+\n+\tword_diff --color-words=\"[a-z]+\"\n+\n+'\n+\n+cat > expect <<\\EOF\n <WHITE>diff --git a/pre b/post<RESET>\n <WHITE>index 330b04f..5ed8eff 100644<RESET>\n <WHITE>--- a/pre<RESET>\n@@ -102,25 +113,9 @@ a = b + c<RESET>\n <GREEN>aeff = aeff * ( aaa )<RESET>\n EOF\n \n-test_expect_success 'set a diff driver' '\n-\tgit config diff.testdriver.wordregex \"[^[:space:]]\" &&\n-\tcat <<EOF > .gitattributes\n-pre diff=testdriver\n-post diff=testdriver\n-EOF\n-'\n-\n test_expect_success 'use default supplied by driver' '\n \n-\tword_diff --color-words &&\n-\tword_diff_check expect-by-chars\n-\n-'\n-\n-test_expect_success 'option overrides default' '\n-\n-\tword_diff --color-words=\"[a-z]+\" &&\n-\tword_diff_check expect\n+\tword_diff --color-words\n \n '\n \n@@ -136,10 +131,11 @@ cat > expect <<\\EOF\n aaa (aaa) <GREEN>aaa<RESET>\n EOF\n \n-test_expect_success \"test parsing words for newline\" '\n+test_expect_success 'test parsing words for newline' '\n+\n+\tword_diff --color-words=\"a+\"\n \n-\tword_diff --color-words=\"a+\" &&\n-\tword_diff_check expect\n+\tword_diff --color-words=.\n \n '\n \n-- \n1.6.1.300.gbc493\n"},{"id":"100532","messageId":"alpine.DEB.1.00.0901150235122.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901150211120.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH replacement for take 3 4/4] color-words: take an optional regular expression describing words","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-15T01:36:37Z","receivedAt":"2009-01-15T01:36:37Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 15 Jan 2009, Johannes Schindelin wrote:\n\n> On Thu, 15 Jan 2009, Thomas Rast wrote:\n> \n> > Johannes Schindelin wrote:\n> > > \tThis basically contains the fix I sent earlier.\n> > \n> > Unfortunately I found another case where it breaks.  It even comes\n> > with a fairly neat test case:\n> > \n> >   $ g diff --no-index test_a test_b\n> >   diff --git 1/test_a 2/test_b\n> >   index 289cb9d..2d06f37 100644\n> >   --- 1/test_a\n> >   +++ 2/test_b\n> >   @@ -1 +1 @@\n> >   -(:\n> >   +(\n> \n> The diff of the words would look like this:\n> \n> diff --git a/a1 b/a2\n> index 8309acb..2d06f37 100644\n> --- a/a1\n> +++ b/a2\n> @@ -2 +1,0 @@\n> -:\n> \n> \n> Notice the \"+1,0\"?  I fully expected this to be \"+2,0\", but apparently I \n> was mistaken...\n> \n> Can anybody explain to me why this is so?\n\n[PATCH to be squashed into the word regex patch] Fix for strange '@@ -2 +1,0 @@' hunk header\n\nIf a hunk header '@@ -2 +1,0 @@' is found that logically should be\n'@@ -2 +2,0 @@', diff_words got confused.\n\nIt would bee squashed into 4/4.\n\nThis might be a libxdiff issue, though.\n\nNot sure yet.\n---\n diff.c                |   18 ++++++++++++++++++\n t/t4034-diff-words.sh |   16 ++++++++++++++++\n 2 files changed, 34 insertions(+), 0 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex c5f7c57..3709651 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -360,6 +360,24 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \tplus_end = plus_len == 0 ? plus_begin :\n \t\tdiff_words->plus.orig[plus_first + plus_len - 1].end;\n \n+\t/*\n+\t * since this is a --unified=0 diff, it can result in a single hunk\n+\t * with a header like this: @@ -2 +1,0 @@\n+\t *\n+\t * This breaks the assumption that minus_first == plus_first.\n+\t *\n+\t * So we have to fix it: whenever we reach the end of pre and post\n+\t * texts, but nothing was added, we need to shift the plus part\n+\t * to the end of the buffer.\n+\t *\n+\t * It is only necessary for the plus part, as we show the common\n+\t * words from that buffer.\n+\t */\n+\tif (plus_len == 0 && minus_first + minus_len\n+\t\t\t== diff_words->minus.orig_nr)\n+\t\tplus_begin = plus_end =\n+\t\t\tdiff_words->plus.orig[diff_words->plus.orig_nr - 1].end;\n+\n \tif (diff_words->current_plus != plus_begin)\n \t\tfwrite(diff_words->current_plus,\n \t\t\t\tplus_begin - diff_words->current_plus, 1,\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 07e48d1..817fba6 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -135,6 +135,22 @@ test_expect_success 'test parsing words for newline' '\n \n \tword_diff --color-words=\"a+\"\n \n+'\n+\n+echo '(:' > pre\n+echo '(' > post\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 289cb9d..2d06f37 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1 +1 @@<RESET>\n+(<RED>:<RESET>\n+EOF\n+\n+test_expect_success 'test when words are only removed at the end' '\n+\n \tword_diff --color-words=.\n \n '\n-- \n1.6.1.300.gbc493\n"},{"id":"100534","messageId":"alpine.DEB.1.00.0901150241210.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901150233121.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH 4/4] color-words: make regex configurable via attributes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-15T01:43:33Z","receivedAt":"2009-01-15T01:43:33Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 15 Jan 2009, Johannes Schindelin wrote:\n\n> @@ -136,10 +131,11 @@ cat > expect <<\\EOF\n>  aaa (aaa) <GREEN>aaa<RESET>\n>  EOF\n>  \n> -test_expect_success \"test parsing words for newline\" '\n> +test_expect_success 'test parsing words for newline' '\n> +\n> +\tword_diff --color-words=\"a+\"\n>  \n> -\tword_diff --color-words=\"a+\" &&\n> -\tword_diff_check expect\n> +\tword_diff --color-words=.\n>  \n>  '\n\nD'oh.  please remove the last word_diff, this comes from my \"fix\" for your \nsmiley issue.\n\nCiao,\nDscho \"off to bed\"\n\n--\n\"The night was so dark that he hardly coulx srr tje keuboarf.\"\n"},{"id":"100544","messageId":"8763khtbfc.fsf@iki.fi","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901142258250.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Teemu Likonen","fromEmail":"tlikonen@iki.fi","sentAt":"2009-01-15T04:56:07Z","receivedAt":"2009-01-15T04:56:07Z","isPatch":true,"sender":{"key":"tlikonen@iki.fi","avatar":null},"body":"Johannes Schindelin (2009-01-14 23:06 +0100) wrote:\n\n> On Wed, 14 Jan 2009, Thomas Rast wrote:\n>>       -aaa [aaa]\n>>       +aaa (aaa) aaa\n>> \n>> would still give you\n>> \n>>       aaa (aaa)<GREEN> aaa<RESET>\n>> \n>> which may be unexpected.\n>\n> But why should it be unexpected?  If people say that every length of \"a\" \n> makes a word, and consequently everything else is clutter, then that's \n> that, no?\n\nIt works logically but I'd very much like to see a some kind of advice\nin the man page. I already faced this (unexpected) situation and wasn't\nable to fix the regexp myself.\n"},{"id":"100551","messageId":"200901150930.38100.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901150235122.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH replacement for take 3 4/4] color-words: take an optional regular expression describing words","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-15T08:30:35Z","receivedAt":"2009-01-15T08:30:35Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> If a hunk header '@@ -2 +1,0 @@' is found that logically should be\n> '@@ -2 +2,0 @@', diff_words got confused.\n[...]\n> This might be a libxdiff issue, though.\n\nLooks like it's just bug-for-bug compatible with diff.  At least my\nGNU diffutils 2.8.7 show the same behaviour.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n\n"},{"id":"100557","messageId":"200901151140.20215.trast@student.ethz.ch","threadId":"17092","inReplyTo":"200901150930.38100.trast@student.ethz.ch","subject":"Re: [PATCH replacement for take 3 4/4] color-words: take an optional regular expression describing words","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-15T10:40:16Z","receivedAt":"2009-01-15T10:40:16Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Thomas Rast wrote:\n> Johannes Schindelin wrote:\n> > If a hunk header '@@ -2 +1,0 @@' is found that logically should be\n> > '@@ -2 +2,0 @@', diff_words got confused.\n> [...]\n> > This might be a libxdiff issue, though.\n> \n> Looks like it's just bug-for-bug compatible with diff.  At least my\n> GNU diffutils 2.8.7 show the same behaviour.\n\nI think the culprit is in\n\n  commit ca557afff9f7dad7a8739cd193ac0730d872e282\n  Author: Davide Libenzi <davidel@xmailserver.org>\n  Date:   Mon Apr 3 18:47:55 2006 -0700\n\n      Clean-up trivially redundant diff.\n\n      Also corrects the line numbers in unified output when using\n      zero lines context.\n[...]\ndiff --git a/xdiff/xutils.c b/xdiff/xutils.c\n[...]\n  @@ -244,7 +257,7 @@ int xdl_emit_hunk_hdr(long s1, long c1, long s2, long c2,\n          memcpy(buf, \"@@ -\", 4);\n          nb += 4;\n\n  -       nb += xdl_num_out(buf + nb, c1 ? s1: 0);\n  +       nb += xdl_num_out(buf + nb, c1 ? s1: s1 - 1);\n\n          if (c1 != 1) {\n                  memcpy(buf + nb, \",\", 1);\n  @@ -256,7 +269,7 @@ int xdl_emit_hunk_hdr(long s1, long c1, long s2, long c2,\n          memcpy(buf + nb, \" +\", 2);\n          nb += 2;\n\n  -       nb += xdl_num_out(buf + nb, c2 ? s2: 0);\n  +       nb += xdl_num_out(buf + nb, c2 ? s2: s2 - 1);\n\n          if (c2 != 1) {\n                  memcpy(buf + nb, \",\", 1);\n\n\nNote how (for some reason I don't quite understand yet) \"correcting\"\nthe offsets involves subtracting 1 if there were no changes on that\nside.\n\nBut skipping ahead to the end doesn't work if there are several such\ninstances where nothing was added.  So I think it must be fixed as\nfollows.\n\n---- 8< ----\ndiff --git a/diff.c b/diff.c\nindex 4174d88..d7bbf74 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -361,8 +361,9 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \t\tdiff_words->plus.orig[plus_first + plus_len - 1].end;\n \n \t/*\n-\t * since this is a --unified=0 diff, it can result in a single hunk\n-\t * with a header like this: @@ -2 +1,0 @@\n+\t * libxdiff subtracts one from the offset if the corresponding\n+\t * length is 0.\t (This can only happen because we use\n+\t * --unified=0.)\n \t *\n \t * This breaks the assumption that minus_first == plus_first.\n \t *\n@@ -373,10 +374,9 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \t * It is only necessary for the plus part, as we show the common\n \t * words from that buffer.\n \t */\n-\tif (plus_len == 0 && minus_first + minus_len\n-\t\t\t== diff_words->minus.orig_nr)\n+\tif (plus_len == 0)\n \t\tplus_begin = plus_end =\n-\t\t\tdiff_words->plus.orig[diff_words->plus.orig_nr - 1].end;\n+\t\t\tdiff_words->plus.orig[plus_first + plus_len].end;\n \n \tif (diff_words->current_plus != plus_begin)\n \t\tfwrite(diff_words->current_plus,\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 744221b..875b464 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -156,4 +156,40 @@ test_expect_success 'test when words are only removed at the end' '\n \n '\n \n+echo 'abcd(Xefghijklmn(YZopqrst' > pre\n+echo 'abcd(efghijklmn(opqrst' > post\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 434ff54..c4bb9f1 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1 +1 @@<RESET>\n+abcd(<RED>X<RESET>efghijklmn(<RED>YZ<RESET>opqrst\n+EOF\n+\n+test_expect_success 'no added words' '\n+\n+\tword_diff --color-words=.\n+\n+'\n+\n+echo 'abcd(efghijklmn(opqrst' > pre\n+echo 'abcd(Xefghijklmn(YZopqrst' > post\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index c4bb9f1..434ff54 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1 +1 @@<RESET>\n+abcd(<GREEN>X<RESET>efghijklmn(<GREEN>YZ<RESET>opqrst\n+EOF\n+\n+test_expect_success 'no removed words' '\n+\n+\tword_diff --color-words=.\n+\n+'\n+\n test_done\n-- \n1.6.1.283.g653b2\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n"},{"id":"100562","messageId":"alpine.DEB.1.00.0901151337080.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"8763khtbfc.fsf@iki.fi","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-15T12:41:37Z","receivedAt":"2009-01-15T12:41:37Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 15 Jan 2009, Teemu Likonen wrote:\n\n> Johannes Schindelin (2009-01-14 23:06 +0100) wrote:\n> \n> > On Wed, 14 Jan 2009, Thomas Rast wrote:\n> >>       -aaa [aaa]\n> >>       +aaa (aaa) aaa\n> >> \n> >> would still give you\n> >> \n> >>       aaa (aaa)<GREEN> aaa<RESET>\n> >> \n> >> which may be unexpected.\n> >\n> > But why should it be unexpected?  If people say that every length of \"a\" \n> > makes a word, and consequently everything else is clutter, then that's \n> > that, no?\n> \n> It works logically but I'd very much like to see a some kind of advice\n> in the man page. I already faced this (unexpected) situation and wasn't\n> able to fix the regexp myself.\n\nExactly because it works logically, I do not want to change it.  This is \nwhat the user said, and for a change, it could be what the user meant.\n\nYou'll have to come up with a method to describe exactly what you want.  \nSo what is it exactly?  What would you want in such a situation?  You \nasked for words that consist solely of the letter 'a'.  Now, the \nsurrounding stuff differs.  What should Git do?\n\nBTW this gets even worse when you compare the following:\n\nbbb aaa\nccc aaa\n\n--color-words=a+ will show\n\nccc aaa\n\n(!!!)\n\nCiao,\nDscho\n"},{"id":"100565","messageId":"alpine.DEB.1.00.0901151342300.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"200901151140.20215.trast@student.ethz.ch","subject":"Re: [PATCH replacement for take 3 4/4] color-words: take an optional regular expression describing words","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-15T12:54:42Z","receivedAt":"2009-01-15T12:54:42Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 15 Jan 2009, Thomas Rast wrote:\n\n> Thomas Rast wrote:\n> > Johannes Schindelin wrote:\n> > > If a hunk header '@@ -2 +1,0 @@' is found that logically should be\n> > > '@@ -2 +2,0 @@', diff_words got confused.\n> > [...]\n> > > This might be a libxdiff issue, though.\n> > \n> > Looks like it's just bug-for-bug compatible with diff.  At least my\n> > GNU diffutils 2.8.7 show the same behaviour.\n> \n> I think the culprit is in\n> \n>   commit ca557afff9f7dad7a8739cd193ac0730d872e282\n>   Author: Davide Libenzi <davidel@xmailserver.org>\n>   Date:   Mon Apr 3 18:47:55 2006 -0700\n> \n>       Clean-up trivially redundant diff.\n> \n>       Also corrects the line numbers in unified output when using\n>       zero lines context.\n> [...]\n> diff --git a/xdiff/xutils.c b/xdiff/xutils.c\n> [...]\n>   @@ -244,7 +257,7 @@ int xdl_emit_hunk_hdr(long s1, long c1, long s2, long c2,\n>           memcpy(buf, \"@@ -\", 4);\n>           nb += 4;\n> \n>   -       nb += xdl_num_out(buf + nb, c1 ? s1: 0);\n>   +       nb += xdl_num_out(buf + nb, c1 ? s1: s1 - 1);\n\nJunio mentioned some POSIX document in which this behavior is actually \nrequired.  So I'll fix my code thusly:\n\n-- snipsnap --\ndiff --git a/diff.c b/diff.c\nindex 3709651..219a242 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -353,30 +353,20 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \t\t\t&minus_first, &minus_len, &plus_first, &plus_len))\n \t\treturn;\n \n-\tminus_begin = diff_words->minus.orig[minus_first].begin;\n-\tminus_end = minus_len == 0 ? minus_begin :\n-\t\tdiff_words->minus.orig[minus_first + minus_len - 1].end;\n-\tplus_begin = diff_words->plus.orig[plus_first].begin;\n-\tplus_end = plus_len == 0 ? plus_begin :\n-\t\tdiff_words->plus.orig[plus_first + plus_len - 1].end;\n+\t/* POSIX requires that first be decremented by one if len == 0... */\n+\tif (minus_len) {\n+\t\tminus_begin = diff_words->minus.orig[minus_first].begin;\n+\t\tminus_end =\n+\t\t\tdiff_words->minus.orig[minus_first + minus_len - 1].end;\n+\t} else\n+\t\tminus_begin = minus_end =\n+\t\t\tdiff_words->minus.orig[minus_first].end;\n \n-\t/*\n-\t * since this is a --unified=0 diff, it can result in a single hunk\n-\t * with a header like this: @@ -2 +1,0 @@\n-\t *\n-\t * This breaks the assumption that minus_first == plus_first.\n-\t *\n-\t * So we have to fix it: whenever we reach the end of pre and post\n-\t * texts, but nothing was added, we need to shift the plus part\n-\t * to the end of the buffer.\n-\t *\n-\t * It is only necessary for the plus part, as we show the common\n-\t * words from that buffer.\n-\t */\n-\tif (plus_len == 0 && minus_first + minus_len\n-\t\t\t== diff_words->minus.orig_nr)\n-\t\tplus_begin = plus_end =\n-\t\t\tdiff_words->plus.orig[diff_words->plus.orig_nr - 1].end;\n+\tif (plus_len) {\n+\t\tplus_begin = diff_words->plus.orig[plus_first].begin;\n+\t\tplus_end = diff_words->plus.orig[plus_first + plus_len - 1].end;\n+\t} else\n+\t\tplus_begin = plus_end = diff_words->plus.orig[plus_first].end;\n \n \tif (diff_words->current_plus != plus_begin)\n \t\tfwrite(diff_words->current_plus,\n-- \n1.6.1.300.gbc493\n"},{"id":"100570","messageId":"871vv4soul.fsf@iki.fi","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901151337080.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Teemu Likonen","fromEmail":"tlikonen@iki.fi","sentAt":"2009-01-15T13:03:46Z","receivedAt":"2009-01-15T13:03:46Z","isPatch":true,"sender":{"key":"tlikonen@iki.fi","avatar":null},"body":"Johannes Schindelin (2009-01-15 13:41 +0100) wrote:\n\n> Exactly because it works logically, I do not want to change it.  This is \n> what the user said, and for a change, it could be what the user meant.\n\nI'm just saying that it would be helpful (to me at least) if the man\npage included this advice. Thomas Rast already suggested this in his\nversion of the man page change:\n\n    You may want to append `|\\S` to your regular expression to make sure\n    that it matches all non-whitespace characters.\n"},{"id":"100576","messageId":"200901151427.38447.trast@student.ethz.ch","threadId":"17092","inReplyTo":"871vv4soul.fsf@iki.fi","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-15T13:27:35Z","receivedAt":"2009-01-15T13:27:35Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Teemu Likonen wrote:\n> Johannes Schindelin (2009-01-15 13:41 +0100) wrote:\n> \n> > Exactly because it works logically, I do not want to change it.  This is \n> > what the user said, and for a change, it could be what the user meant.\n> \n> I'm just saying that it would be helpful (to me at least) if the man\n> page included this advice. Thomas Rast already suggested this in his\n> version of the man page change:\n> \n>     You may want to append `|\\S` to your regular expression to make sure\n>     that it matches all non-whitespace characters.\n\nDscho requested that I put the extended docs in one of my patches, so\nit's currently in\n\n  http://article.gmane.org/gmane.comp.version-control.git/105716\n\nComments welcome of course.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n"},{"id":"100639","messageId":"7vmydstoys.fsf@gitster.siamese.dyndns.org","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901151337080.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-15T18:15:55Z","receivedAt":"2009-01-15T18:15:55Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> BTW this gets even worse when you compare the following:\n>\n> bbb aaa\n> ccc aaa\n>\n> --color-words=a+ will show\n>\n> ccc aaa\n\nNaive question.  What is the expected output?\n\nThe user defines that \"a\", \"aa\", \"aaa\",... are words and everything else\nis the background that the words float on, and asks --color-words to color\ncode where the words differ.  The way to show them is to have the words in\nred (if it comes from preimage) or in green (if it comes from postimage) on\ntop of some background.\n\nIn this case, there is no difference in words, and the only difference is\nthe background.  Should we still see any output?  Shouldn't it behave more\nlike \"diff -w\" that suppresses lines that differ only in whitespace?\n\nI didn't see the semantics of color-words documented in the original\neither, and I think it should be described in a way humans would\nunderstand (in other words, \"here is what we do internally, splitting\nwords into lines, running diff between them and coalescing the result in\nthis and that way, and whatever happens to be output is what you get\" is\nnot the semantics that is explained in a way humans would understand).\n\nThe above \"The way to show them is to have the words in red (if it comes\nfrom preimage) or in green (if it comes from postimage) on top of some\nbackground.\" was my attempt to describe an easier half of the semantics,\nbut I am not sure what definition of \"some background\" the current draft\ncode is designed around; I think the original's definition was \"we discard\nthe background from either preimage or postimage and insert whitespace\noutselves between the words we output; the only exception is the\nend-of-line that appears in the postimage which we try to keep\" or\nsomething like that, but that is not written in the documentation either.\n\nHow should the background computed to draw the result on?  If a\ncorresponding background portion appear in both the preimage and the\npostimage, we use the one from the postimage?  That justifies why bbb is\nnot shown but ccc is, when you compare these two:\n\n  bbb aaa\n  ccc aa\n\nWhat happens if a portion of background is only in the preimage?\nE.g. when these two are compared:\n\n  bbb aaa bb aa b\n  ccc aaa cc\n\nwhat should happen?  We would want to say \"aa\" was removed by showing it\nin red, but on what background should it be displayed?  cc <red>aa</red>\nb?\n"},{"id":"100646","messageId":"alpine.DEB.1.00.0901151940170.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"7vmydstoys.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-15T19:25:49Z","receivedAt":"2009-01-15T19:25:49Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 15 Jan 2009, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> I didn't see the semantics of color-words documented in the original\n> either,\n\nYeah, my bad.  Will try to fix it with this round of patches.\n\nActually, I'll give a quick outline right here:\n\nIdea: the idea of word diff is to show the differences on a word level \ninstead of line level.  To make it easier for humans (albeit we studiously \nexclude color blinds with our defaults), we do not show \"+\" and \"-\" as the \nstandard diff does, but use colors to designate if the words were removed \nor added.\n\nNow, the thing is that the inter-word parts _can_ differ.  The idea here \nis to show the part of the postimage and drop the preimage under the \ntable.\n\nMethod: We use libxdiff as the real workhorse.  First, we let it generate \na line diff.\n\nThen we reconstruct the preimage and postimage for each hunk, process both \ninto new images that have at most one word (in the new code exactly one \nword) per line, and feed the new preimage/postimage pair to libxdiff.\n\n>From the output of libxdiff, we reconstruct which words were actually \nremoved and which were added.  Then -- like the line based diff -- we \ncombine the runs of common words, removed words and added words, and show \nthem.\n\nThe algorithm I implemented in the new patch series is actually much \ncleaner than the old one:\n\n- it feeds images to libxdiff which contain _exactly_ one word per line, \n  decoupling the word offsets in the original image from the offsets in \n  the processed image,\n\n- this decoupling allows for arbitrary word boundaries, even 0-character \n  ones,\n\n- it parses the hunk headers of the libxdiff output instead of the \"-\", \n  \"+\" and \" \" lines, and therefore does not have to play tricks with the \n  newline character in the middle of a run of removed words.\n\n> What happens if a portion of background is only in the preimage?\n\nIf it is in a run of words that were removed, i.e. that are only in the \npreimage, then it is shown in that part.  Otherwise, the background of the \npreimage is never shown.\n\n> E.g. when these two are compared:\n> \n>   bbb aaa bb aa b\n>   ccc aaa cc\n> \n> what should happen?  We would want to say \"aa\" was removed by showing it\n> in red, but on what background should it be displayed?  cc <red>aa</red>\n> b?\n\nIf we are only ever interested in the 'a's, I'd say that the output should \nonly reflect that.  In other words, what the current code does (ccc \naaa<red>aa</red> cc) is okay IMHO.  After all, we said we're interested in \nthe 'a's, so we should not complain that it did not show us the removal of \n'b's.\n\nCiao,\nDscho\n"},{"id":"100674","messageId":"adf1fd3d0901151610p41930ee2gfc7259aee7e15d73@mail.gmail.com","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901151940170.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Santi Béjar","fromEmail":"santi@agolina.net","sentAt":"2009-01-16T00:10:34Z","receivedAt":"2009-01-16T00:10:34Z","isPatch":true,"sender":{"key":"santi@agolina.net","avatar":null},"body":"2009/1/15 Johannes Schindelin <Johannes.Schindelin@gmx.de>:\n> On Thu, 15 Jan 2009, Junio C Hamano wrote:\n>> I didn't see the semantics of color-words documented in the original\n>> either,\n>\n\n[...]\n\n>> E.g. when these two are compared:\n>>\n>>   bbb aaa bb aa b\n>>   ccc aaa cc\n>>\n>> what should happen?  We would want to say \"aa\" was removed by showing it\n>> in red, but on what background should it be displayed?  cc <red>aa</red>\n>> b?\n>\n> If we are only ever interested in the 'a's, I'd say that the output should\n> only reflect that.  In other words, what the current code does (ccc\n> aaa<red>aa</red> cc) is okay IMHO.  After all, we said we're interested in\n> the 'a's, so we should not complain that it did not show us the removal of\n> 'b's.\n\nIt may be ok and logical, but for me it is not what I want. Mmaybe I\ndon't really undestand what I want or is a crazy idea but here it is\nanyway:\n\nTake a simple case with this two lines :\n\nmatrix[a,b,c]\nmatrix{d,b,c}\n\nthere is no space so the standard color-words does not help to\nvisualize that matrix, the b and c are not changed.\n\nWhat I currently do is to add some spaces:\n\nmatrix[ a, b, c ]\nmatrix{ d, b, c }\n\nthen the color-words at least says that \"b, c\" is unchanged.\n\nWhat I would like is that --color-words would act as adding this\nspaces automatically (and even one after \"matrix\").\n\nOr another way to think it could be:\n\na) primary words are those with alphanumerics (or a regex)\nb) secondary \"words\" are the other non-whitespaces characters (in this\ncase \"[]{} and ,\"\nc) whitespaces are cruft.\n\n(having two regexp to specify what is a words but they cannot mix).\n\nIf everything works as I think (it's late night :-) with the above two lines:\n\nmatrix[a,b,c]\nmatrix{d,b,c}\n\nthe word diff would be\n\nmatrix<RED>[<GREEN>{<RED>a<GREEN>d<RESET>,b,c<RED>]<GREEN>}<RED>\n\nSanti\n"},{"id":"100678","messageId":"7vy6xcowsx.fsf@gitster.siamese.dyndns.org","threadId":"17092","inReplyTo":"adf1fd3d0901151610p41930ee2gfc7259aee7e15d73@mail.gmail.com","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-16T01:37:50Z","receivedAt":"2009-01-16T01:37:50Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Santi Béjar <santi@agolina.net> writes:\n\n> It may be ok and logical, but for me it is not what I want. Mmaybe I\n> don't really undestand what I want or is a crazy idea but here it is\n> anyway:\n>\n> Take a simple case with this two lines :\n>\n> matrix[a,b,c]\n> matrix{d,b,c}\n>\n> there is no space so the standard color-words does not help to\n> visualize that matrix, the b and c are not changed.\n>\n> What I currently do is to add some spaces:\n>\n> matrix[ a, b, c ]\n> matrix{ d, b, c }\n>\n> then the color-words at least says that \"b, c\" is unchanged.\n>\n> What I would like is that --color-words would act as adding this\n> spaces automatically (and even one after \"matrix\").\n>\n> Or another way to think it could be:\n>\n> a) primary words are those with alphanumerics (or a regex)\n> b) secondary \"words\" are the other non-whitespaces characters (in this\n> case \"[]{} and ,\"\n> c) whitespaces are cruft.\n\nDscho and Thomas discussed and designed a way to mark \"words look like\nthis\" (and anything that are not words are crufts), and Dscho further\nargues that it is Ok to discard crufts (which I think is fine).\n\nWhat you seem to want in this example is \"there is no cruft other than\nwhitespace, but there are different kinds of words\".  I do not think it is\nincompatible with the way crufts are discarded, but it may be incompatible\nwith the way how words are identified.\n\nI would expect something like:\n\n\t[a-zA-Z0-9]+|[^ a-zA-Z0-9]+\n\nshould define your \"two kinds of words\".  That is, a run of alnums is a\nword, and a run of non-alnums is a word, but \"matrix[a\" is not a word (it\nis a sequence of three words \"matrix\", \"[\" and \"a\").\n"},{"id":"100679","messageId":"200901151942.55742.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"adf1fd3d0901151610p41930ee2gfc7259aee7e15d73@mail.gmail.com","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-16T01:42:55Z","receivedAt":"2009-01-16T01:42:55Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"On Thursday 15 January 2009, Santi Béjar <santi@agolina.net> wrote \nabout 'Re: [PATCH take 3 0/4] color-words improvements':\n>It may be ok and logical, but for me it is not what I want. Mmaybe I\n>don't really undestand what I want or is a crazy idea but here it is\n>anyway:\n\nThe discussion above is mildly theoretical.  I don't imagine someone is \ngoing to intentionally mark 98% of a file as non-words, which is basically \nwhat you are doing with a regex of \"a+\".\n\n>a) primary words are those with alphanumerics (or a regex)\n\nregex: [[:alnum:]]+\n\nexample words: matrix ball I a\nexample non-words: don't haven't\n\n>b) secondary \"words\" are the other non-whitespaces characters (in this\n>case \"[]{} and ,\"\n\nregex: []{}[,]\n\nexample words: [ , }\nexample non-words: [] ball 147\n\n>c) whitespaces are cruft.\n>\n>(having two regexp to specify what is a words but they cannot mix).\n\nCombine regex with '|' to get:\n[[:alnum:]]+|[]{}[,]\n\n>If everything works as I think (it's late night :-) with the above two\n> lines:\n>\n>matrix[a,b,c]\n>matrix{d,b,c}\n>\n>the word diff would be\n>\n>matrix<RED>[<GREEN>{<RED>a<GREEN>d<RESET>,b,c<RED>]<GREEN>}<RED>\n\nFor this specific case, the regex \"[^[:space:]]\" by itself should work, \nalthough it would end up being a character-by-character diff.\n\nThe regex you built from your description \"[[:alnum:]]+|[]}{[,]\" would also \ngive the same diff.  However:\n-dont\n+don't\ngives a word diff of:\ndon't\nnot:\ndon<RED>'<RESET>t\nbecause \"'\" is not recognized as part of any word it is considered \nignorable.\n\nThere was a patch that included documentation that most users should add \n\"|[^[:space:]]\" to the end of their regex, to capture all non-whitespace \ncharacters that are not otherwise part of a word as individual, \nsingle-character \"words\".\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"100682","messageId":"alpine.DEB.1.00.0901160253210.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"adf1fd3d0901151610p41930ee2gfc7259aee7e15d73@mail.gmail.com","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-16T01:55:28Z","receivedAt":"2009-01-16T01:55:28Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 16 Jan 2009, Santi Béjar wrote:\n\n> If everything works as I think (it's late night :-) with the above two lines:\n> \n> matrix[a,b,c]\n> matrix{d,b,c}\n> \n> the word diff would be\n> \n> matrix<RED>[<GREEN>{<RED>a<GREEN>d<RESET>,b,c<RED>]<GREEN>}<RED>\n\nSo I guess that you want something like\n\n\t[A-Za-z0-9]+|[^A-Za-z0-9 \\t]+\n\nNote: I only want to help you finding what you actually want, I am not \ntrying to find it for you.\n\nCiao,\nDscho\n"},{"id":"100707","messageId":"adf1fd3d0901160102y32a08e26q96728495fc0b6fcf@mail.gmail.com","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901160253210.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Santi Béjar","fromEmail":"santi@agolina.net","sentAt":"2009-01-16T09:02:33Z","receivedAt":"2009-01-16T09:02:33Z","isPatch":true,"sender":{"key":"santi@agolina.net","avatar":null},"body":" 2009/1/16 Johannes Schindelin <Johannes.Schindelin@gmx.de>:\n> Hi,\n>\n> On Fri, 16 Jan 2009, Santi Béjar wrote:\n>\n>> If everything works as I think (it's late night :-) with the above two lines:\n>>\n>> matrix[a,b,c]\n>> matrix{d,b,c}\n>>\n>> the word diff would be\n>>\n>> matrix<RED>[<GREEN>{<RED>a<GREEN>d<RESET>,b,c<RED>]<GREEN>}<RED>\n>\n> So I guess that you want something like\n>\n>        [A-Za-z0-9]+|[^A-Za-z0-9 \\t]+\n>\n> Note: I only want to help you finding what you actually want, I am not\n> trying to find it for you.\n>\n\nThanks all for the answers.\n\nSo, I see, it is a matter of finding the right regexp.\n\nBut the only use case for me is of this kind, and I think for the\nothers too. So maybe an easier way to specify it could be worth. But\nI'll write an alias as this is the only regexp I would use, apart from\nthe default word diff.\n\nThanks,\nSanti\n"},{"id":"100721","messageId":"alpine.DEB.1.00.0901161255520.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"adf1fd3d0901160102y32a08e26q96728495fc0b6fcf@mail.gmail.com","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-16T11:57:53Z","receivedAt":"2009-01-16T11:57:53Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 16 Jan 2009, Santi Béjar wrote:\n\n>  2009/1/16 Johannes Schindelin <Johannes.Schindelin@gmx.de>:\n> >\n> > On Fri, 16 Jan 2009, Santi Béjar wrote:\n> >\n> >> If everything works as I think (it's late night :-) with the above \n> >> two lines:\n> >>\n> >> matrix[a,b,c]\n> >> matrix{d,b,c}\n> >>\n> >> the word diff would be\n> >>\n> >> matrix<RED>[<GREEN>{<RED>a<GREEN>d<RESET>,b,c<RED>]<GREEN>}<RED>\n> >\n> > So I guess that you want something like\n> >\n> >        [A-Za-z0-9]+|[^A-Za-z0-9 \\t]+\n> >\n> \n> So, I see, it is a matter of finding the right regexp.\n> \n> But the only use case for me is of this kind, and I think for the\n> others too. So maybe an easier way to specify it could be worth.\n\nSure.  If you can come up with a nice name for it, we could add special \nhandling for something like \"[[:words:]]\" expanding into said regexp.\n\nCiao,\nDscho\n"},{"id":"100722","messageId":"adf1fd3d0901160401s7a363076x1bcd8e90db4f56a1@mail.gmail.com","threadId":"17092","inReplyTo":"adf1fd3d0901160102y32a08e26q96728495fc0b6fcf@mail.gmail.com","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Santi Béjar","fromEmail":"santi@agolina.net","sentAt":"2009-01-16T12:01:35Z","receivedAt":"2009-01-16T12:01:35Z","isPatch":true,"sender":{"key":"santi@agolina.net","avatar":null},"body":"Hi,\n\n  can you both provide a public repository to be able to test the\nlastest version without having to search and apply them?\n\nThanks,\nSanti\n\nP.D.: I know it will be rebased.\n"},{"id":"100725","messageId":"alpine.DEB.1.00.0901161339370.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"adf1fd3d0901160401s7a363076x1bcd8e90db4f56a1@mail.gmail.com","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-16T12:40:06Z","receivedAt":"2009-01-16T12:40:06Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 16 Jan 2009, Santi Béjar wrote:\n\n>   can you both provide a public repository to be able to test the \n> lastest version without having to search and apply them?\n\nYou will always find my latest version in git://repo.or.cz/git/dscho.git.\n\nHth,\nDscho\n"},{"id":"100749","messageId":"200901161011.05614.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"adf1fd3d0901160102y32a08e26q96728495fc0b6fcf@mail.gmail.com","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-16T16:11:01Z","receivedAt":"2009-01-16T16:11:01Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"On Friday 2009 January 16 03:02:33 Santi Béjar wrote:\n> 2009/1/16 Johannes Schindelin <Johannes.Schindelin@gmx.de>:\n>> Hi,\n>>\n>> On Fri, 16 Jan 2009, Santi Béjar wrote:\n>>> If everything works as I think (it's late night :-) with the above two\n>>> lines:\n>>>\n>>> matrix[a,b,c]\n>>> matrix{d,b,c}\n>>>\n>>> the word diff would be\n>>>\n>>> matrix<RED>[<GREEN>{<RED>a<GREEN>d<RESET>,b,c<RED>]<GREEN>}<RED>\n>\n>So, I see, it is a matter of finding the right regexp.\n>\n>But the only use case for me is of this kind, and I think for the\n>others too. So maybe an easier way to specify it could be worth. But\n>I'll write an alias as this is the only regexp I would use, apart from\n>the default word diff.\n\nI think that the C/C++ language word-diff driver would work here, and there \nshould be a shortcut for that.\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"100762","messageId":"200901162004.32557.trast@student.ethz.ch","threadId":"17092","inReplyTo":"adf1fd3d0901160401s7a363076x1bcd8e90db4f56a1@mail.gmail.com","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-16T19:04:29Z","receivedAt":"2009-01-16T19:04:29Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Santi Béjar wrote:\n> Hi,\n> \n>   can you both provide a public repository to be able to test the\n> lastest version without having to search and apply them?\n\nI set up a clone at\n\n  git://repo.or.cz/git/trast.git\n\nThe respective topics are js/word-diff-p1 and tr/word-diff-p2.  For\nyour testing convenience, there are master/next branches that merge\ntr/word-diff-p2 to Junio's master/next.  I *think* I should have\ngathered all squashes from Dscho, too.\n\nThe tip commit has some tweaks to the builtin regexes that aren't in\nany mailed version yet; I'll resend RSN.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n\n"},{"id":"100775","messageId":"alpine.DEB.1.00.0901162208180.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"200901162004.32557.trast@student.ethz.ch","subject":"Re: [PATCH take 3 0/4] color-words improvements","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-16T21:09:56Z","receivedAt":"2009-01-16T21:09:56Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 16 Jan 2009, Thomas Rast wrote:\n\n> The tip commit has some tweaks to the builtin regexes that aren't in any \n> mailed version yet; I'll resend RSN.\n\nNote that I applied the \"better\" fix for the @@ -2 +1,0 @@ issue, but \nhaven't sent out a redone series.\n\nThomas, could you pick up the patches from my 'my-next' branch and \nmaintain an \"official\" topic branch?\n\nCiao,\nDscho\n"},{"id":"100856","messageId":"1232209788-10408-1-git-send-email-trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901162208180.3586@pacific.mpi-cbg.de","subject":"[PATCH v4 0/7] customizable --color-words","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-17T16:29:41Z","receivedAt":"2009-01-17T16:29:41Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> Thomas, could you pick up the patches from my 'my-next' branch and \n> maintain an \"official\" topic branch?\n\nI cherry-picked the three commits you had there, and rebuilt on top.\nI pushed them to\n\n  git://repo.or.cz/git/trast.git tr/word-diff-p2\n\nagain (js/word-diff-p1 again points directly at your half).\n\nThe changes on your side since my last push (hence your last sent\npatches&squashes I collected) were only a pair of quotes changed from\ndouble to single.\n\nOn my side I mainly tweaked the TeX pattern since I noticed it didn't\nmatch many non-alnums such as (), and therefore declare them\nunchanged:\n\n-       \"\\\\\\\\[a-zA-Z@]+|[][{}]|\\\\\\\\.|[a-zA-Z0-9\\x80-\\xff]+\"),\n+       \"\\\\\\\\[a-zA-Z@]+|\\\\\\\\.|[a-zA-Z0-9\\x80-\\xff]+|[^[:space:]]\"),\n\nI also added a clause to the C++ pattern to allow it to match\ndeclarations such as\n\n  int Foo::bar(...)\n\n(it would give up on the :: before).\n\n\nJohannes Schindelin (4):\n  Add color_fwrite_lines(), a function coloring each line individually\n  color-words: refactor word splitting and use ALLOC_GROW()\n  color-words: change algorithm to allow for 0-character word\n    boundaries\n  color-words: take an optional regular expression describing words\n\nThomas Rast (3):\n  color-words: enable REG_NEWLINE to help user\n  color-words: expand docs with precise semantics\n  color-words: make regex configurable via attributes\n\n Documentation/diff-options.txt  |   17 +++-\n Documentation/gitattributes.txt |   21 ++++\n color.c                         |   28 +++++\n color.h                         |    1 +\n diff.c                          |  222 ++++++++++++++++++++++++++-------------\n diff.h                          |    1 +\n t/t4034-diff-words.sh           |  159 ++++++++++++++++++++++++++++\n userdiff.c                      |   78 +++++++++++---\n userdiff.h                      |    1 +\n 9 files changed, 440 insertions(+), 88 deletions(-)\n create mode 100755 t/t4034-diff-words.sh\n"},{"id":"100860","messageId":"1232209788-10408-2-git-send-email-trast@student.ethz.ch","threadId":"17092","inReplyTo":"1232209788-10408-1-git-send-email-trast@student.ethz.ch","subject":"[PATCH v4 1/7] Add color_fwrite_lines(), a function coloring each line individually","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-17T16:29:42Z","receivedAt":"2009-01-17T16:29:42Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWe have to set the color before every line and reset it before every\nnewline.  Add a function color_fwrite_lines() which does that for us.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n color.c |   28 ++++++++++++++++++++++++++++\n color.h |    1 +\n 2 files changed, 29 insertions(+), 0 deletions(-)\n\ndiff --git a/color.c b/color.c\nindex fc0b72a..d4ae83f 100644\n--- a/color.c\n+++ b/color.c\n@@ -191,3 +191,31 @@ int color_fprintf_ln(FILE *fp, const char *color, const char *fmt, ...)\n \tva_end(args);\n \treturn r;\n }\n+\n+/*\n+ * This function splits the buffer by newlines and colors the lines individually.\n+ *\n+ * Returns 0 on success.\n+ */\n+int color_fwrite_lines(FILE *fp, const char *color,\n+\t\tsize_t count, const char *buf)\n+{\n+\tif (!*color)\n+\t\treturn fwrite(buf, count, 1, fp) != 1;\n+\twhile (count) {\n+\t\tchar *p = memchr(buf, '\\n', count);\n+\t\tif (p != buf && (fputs(color, fp) < 0 ||\n+\t\t\t\tfwrite(buf, p ? p - buf : count, 1, fp) != 1 ||\n+\t\t\t\tfputs(COLOR_RESET, fp) < 0))\n+\t\t\treturn -1;\n+\t\tif (!p)\n+\t\t\treturn 0;\n+\t\tif (fputc('\\n', fp) < 0)\n+\t\t\treturn -1;\n+\t\tcount -= p + 1 - buf;\n+\t\tbuf = p + 1;\n+\t}\n+\treturn 0;\n+}\n+\n+\ndiff --git a/color.h b/color.h\nindex 6cf5c88..cd5c985 100644\n--- a/color.h\n+++ b/color.h\n@@ -19,5 +19,6 @@\n void color_parse(const char *var, const char *value, char *dst);\n int color_fprintf(FILE *fp, const char *color, const char *fmt, ...);\n int color_fprintf_ln(FILE *fp, const char *color, const char *fmt, ...);\n+int color_fwrite_lines(FILE *fp, const char *color, size_t count, const char *buf);\n \n #endif /* COLOR_H */\n-- \n1.6.1.315.g92577\n"},{"id":"100857","messageId":"1232209788-10408-3-git-send-email-trast@student.ethz.ch","threadId":"17092","inReplyTo":"1232209788-10408-2-git-send-email-trast@student.ethz.ch","subject":"[PATCH v4 2/7] color-words: refactor word splitting and use ALLOC_GROW()","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-17T16:29:43Z","receivedAt":"2009-01-17T16:29:43Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWord splitting is now performed by the function diff_words_fill(),\navoiding having the same code twice.\n\nIn the same spirit, avoid duplicating the code of ALLOC_GROW().\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n diff.c |   40 +++++++++++++++++++---------------------\n 1 files changed, 19 insertions(+), 21 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex d235482..c111eef 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -326,10 +326,7 @@ struct diff_words_buffer {\n static void diff_words_append(char *line, unsigned long len,\n \t\tstruct diff_words_buffer *buffer)\n {\n-\tif (buffer->text.size + len > buffer->alloc) {\n-\t\tbuffer->alloc = (buffer->text.size + len) * 3 / 2;\n-\t\tbuffer->text.ptr = xrealloc(buffer->text.ptr, buffer->alloc);\n-\t}\n+\tALLOC_GROW(buffer->text.ptr, buffer->text.size + len, buffer->alloc);\n \tline++;\n \tlen--;\n \tmemcpy(buffer->text.ptr + buffer->text.size, line, len);\n@@ -398,6 +395,22 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \t}\n }\n \n+/*\n+ * This function splits the words in buffer->text, and stores the list with\n+ * newline separator into out.\n+ */\n+static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n+{\n+\tint i;\n+\tout->size = buffer->text.size;\n+\tout->ptr = xmalloc(out->size);\n+\tmemcpy(out->ptr, buffer->text.ptr, out->size);\n+\tfor (i = 0; i < out->size; i++)\n+\t\tif (isspace(out->ptr[i]))\n+\t\t\tout->ptr[i] = '\\n';\n+\tbuffer->current = 0;\n+}\n+\n /* this executes the word diff on the accumulated buffers */\n static void diff_words_show(struct diff_words_data *diff_words)\n {\n@@ -405,26 +418,11 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \txdemitconf_t xecfg;\n \txdemitcb_t ecb;\n \tmmfile_t minus, plus;\n-\tint i;\n \n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n-\tminus.size = diff_words->minus.text.size;\n-\tminus.ptr = xmalloc(minus.size);\n-\tmemcpy(minus.ptr, diff_words->minus.text.ptr, minus.size);\n-\tfor (i = 0; i < minus.size; i++)\n-\t\tif (isspace(minus.ptr[i]))\n-\t\t\tminus.ptr[i] = '\\n';\n-\tdiff_words->minus.current = 0;\n-\n-\tplus.size = diff_words->plus.text.size;\n-\tplus.ptr = xmalloc(plus.size);\n-\tmemcpy(plus.ptr, diff_words->plus.text.ptr, plus.size);\n-\tfor (i = 0; i < plus.size; i++)\n-\t\tif (isspace(plus.ptr[i]))\n-\t\t\tplus.ptr[i] = '\\n';\n-\tdiff_words->plus.current = 0;\n-\n+\tdiff_words_fill(&diff_words->minus, &minus);\n+\tdiff_words_fill(&diff_words->plus, &plus);\n \txpp.flags = XDF_NEED_MINIMAL;\n \txecfg.ctxlen = diff_words->minus.alloc + diff_words->plus.alloc;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n-- \n1.6.1.315.g92577\n"},{"id":"100863","messageId":"1232209788-10408-4-git-send-email-trast@student.ethz.ch","threadId":"17092","inReplyTo":"1232209788-10408-3-git-send-email-trast@student.ethz.ch","subject":"[PATCH v4 3/7] color-words: change algorithm to allow for 0-character word boundaries","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-17T16:29:44Z","receivedAt":"2009-01-17T16:29:44Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nUp until now, the color-words code assumed that word boundaries are\nidentical to white space characters.\n\nTherefore, it could get away with a very simple scheme: it copied the\nhunks, substituted newlines for each white space character, called\nlibxdiff with the processed text, and then identified the text to\noutput by the offsets (which agreed since the original text had the\nsame length).\n\nThis code was ugly, for a number of reasons:\n\n- it was impossible to introduce 0-character word boundaries,\n\n- we had to print everything word by word, and\n\n- the code needed extra special handling of newlines in the removed part.\n\nFix all of these issues by processing the text such that\n\n- we build word lists, separated by newlines,\n\n- we remember the original offsets for every word, and\n\n- after calling libxdiff on the wordlists, we parse the hunk headers, and\n  find the corresponding offsets, and then\n\n- we print the removed/added parts in one go.\n\nThe pre and post samples in the test were provided by Santi Béjar.\n\nNote that there is some strange special handling of hunk headers where\none line range is 0 due to POSIX: in this case, the start is one too\nlow.  In other words a hunk header '@@ -1,0 +2 @@' actually means that\nthe line must be added after the _second_ line of the pre text, _not_\nthe first.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n diff.c                |  161 ++++++++++++++++++++++++++++---------------------\n t/t4034-diff-words.sh |   66 ++++++++++++++++++++\n 2 files changed, 159 insertions(+), 68 deletions(-)\n create mode 100755 t/t4034-diff-words.sh\n\ndiff --git a/diff.c b/diff.c\nindex c111eef..37c886a 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -319,8 +319,10 @@ static int fill_mmfile(mmfile_t *mf, struct diff_filespec *one)\n struct diff_words_buffer {\n \tmmfile_t text;\n \tlong alloc;\n-\tlong current; /* output pointer */\n-\tint suppressed_newline;\n+\tstruct diff_words_orig {\n+\t\tconst char *begin, *end;\n+\t} *orig;\n+\tint orig_nr, orig_alloc;\n };\n \n static void diff_words_append(char *line, unsigned long len,\n@@ -335,80 +337,89 @@ static void diff_words_append(char *line, unsigned long len,\n \n struct diff_words_data {\n \tstruct diff_words_buffer minus, plus;\n+\tconst char *current_plus;\n \tFILE *file;\n };\n \n-static void print_word(FILE *file, struct diff_words_buffer *buffer, int len, int color,\n-\t\tint suppress_newline)\n+static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n {\n-\tconst char *ptr;\n-\tint eol = 0;\n+\tstruct diff_words_data *diff_words = priv;\n+\tint minus_first, minus_len, plus_first, plus_len;\n+\tconst char *minus_begin, *minus_end, *plus_begin, *plus_end;\n \n-\tif (len == 0)\n+\tif (line[0] != '@' || parse_hunk_header(line, len,\n+\t\t\t&minus_first, &minus_len, &plus_first, &plus_len))\n \t\treturn;\n \n-\tptr  = buffer->text.ptr + buffer->current;\n-\tbuffer->current += len;\n+\t/* POSIX requires that first be decremented by one if len == 0... */\n+\tif (minus_len) {\n+\t\tminus_begin = diff_words->minus.orig[minus_first].begin;\n+\t\tminus_end =\n+\t\t\tdiff_words->minus.orig[minus_first + minus_len - 1].end;\n+\t} else\n+\t\tminus_begin = minus_end =\n+\t\t\tdiff_words->minus.orig[minus_first].end;\n \n-\tif (ptr[len - 1] == '\\n') {\n-\t\teol = 1;\n-\t\tlen--;\n-\t}\n+\tif (plus_len) {\n+\t\tplus_begin = diff_words->plus.orig[plus_first].begin;\n+\t\tplus_end = diff_words->plus.orig[plus_first + plus_len - 1].end;\n+\t} else\n+\t\tplus_begin = plus_end = diff_words->plus.orig[plus_first].end;\n \n-\tfputs(diff_get_color(1, color), file);\n-\tfwrite(ptr, len, 1, file);\n-\tfputs(diff_get_color(1, DIFF_RESET), file);\n+\tif (diff_words->current_plus != plus_begin)\n+\t\tfwrite(diff_words->current_plus,\n+\t\t\t\tplus_begin - diff_words->current_plus, 1,\n+\t\t\t\tdiff_words->file);\n+\tif (minus_begin != minus_end)\n+\t\tcolor_fwrite_lines(diff_words->file,\n+\t\t\t\tdiff_get_color(1, DIFF_FILE_OLD),\n+\t\t\t\tminus_end - minus_begin, minus_begin);\n+\tif (plus_begin != plus_end)\n+\t\tcolor_fwrite_lines(diff_words->file,\n+\t\t\t\tdiff_get_color(1, DIFF_FILE_NEW),\n+\t\t\t\tplus_end - plus_begin, plus_begin);\n \n-\tif (eol) {\n-\t\tif (suppress_newline)\n-\t\t\tbuffer->suppressed_newline = 1;\n-\t\telse\n-\t\t\tputc('\\n', file);\n-\t}\n-}\n-\n-static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n-{\n-\tstruct diff_words_data *diff_words = priv;\n-\n-\tif (diff_words->minus.suppressed_newline) {\n-\t\tif (line[0] != '+')\n-\t\t\tputc('\\n', diff_words->file);\n-\t\tdiff_words->minus.suppressed_newline = 0;\n-\t}\n-\n-\tlen--;\n-\tswitch (line[0]) {\n-\t\tcase '-':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->minus, len, DIFF_FILE_OLD, 1);\n-\t\t\tbreak;\n-\t\tcase '+':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->plus, len, DIFF_FILE_NEW, 0);\n-\t\t\tbreak;\n-\t\tcase ' ':\n-\t\t\tprint_word(diff_words->file,\n-\t\t\t\t   &diff_words->plus, len, DIFF_PLAIN, 0);\n-\t\t\tdiff_words->minus.current += len;\n-\t\t\tbreak;\n-\t}\n+\tdiff_words->current_plus = plus_end;\n }\n \n /*\n- * This function splits the words in buffer->text, and stores the list with\n- * newline separator into out.\n+ * This function splits the words in buffer->text, stores the list with\n+ * newline separator into out, and saves the offsets of the original words\n+ * in buffer->orig.\n  */\n static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n {\n-\tint i;\n-\tout->size = buffer->text.size;\n-\tout->ptr = xmalloc(out->size);\n-\tmemcpy(out->ptr, buffer->text.ptr, out->size);\n-\tfor (i = 0; i < out->size; i++)\n-\t\tif (isspace(out->ptr[i]))\n-\t\t\tout->ptr[i] = '\\n';\n-\tbuffer->current = 0;\n+\tint i, j;\n+\n+\tout->size = 0;\n+\tout->ptr = xmalloc(buffer->text.size);\n+\n+\t/* fake an empty \"0th\" word */\n+\tALLOC_GROW(buffer->orig, 1, buffer->orig_alloc);\n+\tbuffer->orig[0].begin = buffer->orig[0].end = buffer->text.ptr;\n+\tbuffer->orig_nr = 1;\n+\n+\tfor (i = 0; i < buffer->text.size; i++) {\n+\t\tif (isspace(buffer->text.ptr[i]))\n+\t\t\tcontinue;\n+\t\tfor (j = i + 1; j < buffer->text.size &&\n+\t\t\t\t!isspace(buffer->text.ptr[j]); j++)\n+\t\t\t; /* find the end of the word */\n+\n+\t\t/* store original boundaries */\n+\t\tALLOC_GROW(buffer->orig, buffer->orig_nr + 1,\n+\t\t\t\tbuffer->orig_alloc);\n+\t\tbuffer->orig[buffer->orig_nr].begin = buffer->text.ptr + i;\n+\t\tbuffer->orig[buffer->orig_nr].end = buffer->text.ptr + j;\n+\t\tbuffer->orig_nr++;\n+\n+\t\t/* store one word */\n+\t\tmemcpy(out->ptr + out->size, buffer->text.ptr + i, j - i);\n+\t\tout->ptr[out->size + j - i] = '\\n';\n+\t\tout->size += j - i + 1;\n+\n+\t\ti = j - 1;\n+\t}\n }\n \n /* this executes the word diff on the accumulated buffers */\n@@ -419,22 +430,34 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \txdemitcb_t ecb;\n \tmmfile_t minus, plus;\n \n+\t/* special case: only removal */\n+\tif (!diff_words->plus.text.size) {\n+\t\tcolor_fwrite_lines(diff_words->file,\n+\t\t\tdiff_get_color(1, DIFF_FILE_OLD),\n+\t\t\tdiff_words->minus.text.size, diff_words->minus.text.ptr);\n+\t\tdiff_words->minus.text.size = 0;\n+\t\treturn;\n+\t}\n+\n+\tdiff_words->current_plus = diff_words->plus.text.ptr;\n+\n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n \tdiff_words_fill(&diff_words->minus, &minus);\n \tdiff_words_fill(&diff_words->plus, &plus);\n \txpp.flags = XDF_NEED_MINIMAL;\n-\txecfg.ctxlen = diff_words->minus.alloc + diff_words->plus.alloc;\n+\txecfg.ctxlen = 0;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n \t\t      &xpp, &xecfg, &ecb);\n \tfree(minus.ptr);\n \tfree(plus.ptr);\n+\tif (diff_words->current_plus != diff_words->plus.text.ptr +\n+\t\t\tdiff_words->plus.text.size)\n+\t\tfwrite(diff_words->current_plus,\n+\t\t\tdiff_words->plus.text.ptr + diff_words->plus.text.size\n+\t\t\t- diff_words->current_plus, 1,\n+\t\t\tdiff_words->file);\n \tdiff_words->minus.text.size = diff_words->plus.text.size = 0;\n-\n-\tif (diff_words->minus.suppressed_newline) {\n-\t\tputc('\\n', diff_words->file);\n-\t\tdiff_words->minus.suppressed_newline = 0;\n-\t}\n }\n \n typedef unsigned long (*sane_truncate_fn)(char *line, unsigned long len);\n@@ -458,7 +481,9 @@ static void free_diff_words_data(struct emit_callback *ecbdata)\n \t\t\tdiff_words_show(ecbdata->diff_words);\n \n \t\tfree (ecbdata->diff_words->minus.text.ptr);\n+\t\tfree (ecbdata->diff_words->minus.orig);\n \t\tfree (ecbdata->diff_words->plus.text.ptr);\n+\t\tfree (ecbdata->diff_words->plus.orig);\n \t\tfree(ecbdata->diff_words);\n \t\tecbdata->diff_words = NULL;\n \t}\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nnew file mode 100755\nindex 0000000..b22195f\n--- /dev/null\n+++ b/t/t4034-diff-words.sh\n@@ -0,0 +1,66 @@\n+#!/bin/sh\n+\n+test_description='word diff colors'\n+\n+. ./test-lib.sh\n+\n+test_expect_success setup '\n+\n+\tgit config diff.color.old red\n+\tgit config diff.color.new green\n+\n+'\n+\n+decrypt_color () {\n+\tsed \\\n+\t\t-e 's/.\\[1m/<WHITE>/g' \\\n+\t\t-e 's/.\\[31m/<RED>/g' \\\n+\t\t-e 's/.\\[32m/<GREEN>/g' \\\n+\t\t-e 's/.\\[36m/<BROWN>/g' \\\n+\t\t-e 's/.\\[m/<RESET>/g'\n+}\n+\n+word_diff () {\n+\ttest_must_fail git diff --no-index \"$@\" pre post > output &&\n+\tdecrypt_color < output > output.decrypted &&\n+\ttest_cmp expect output.decrypted\n+}\n+\n+cat > pre <<\\EOF\n+h(4)\n+\n+a = b + c\n+EOF\n+\n+cat > post <<\\EOF\n+h(4),hh[44]\n+\n+a = b + c\n+\n+aa = a\n+\n+aeff = aeff * ( aaa )\n+EOF\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 330b04f..5ed8eff 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1,3 +1,7 @@<RESET>\n+<RED>h(4)<RESET><GREEN>h(4),hh[44]<RESET>\n+<RESET>\n+a = b + c<RESET>\n+\n+<GREEN>aa = a<RESET>\n+\n+<GREEN>aeff = aeff * ( aaa )<RESET>\n+EOF\n+\n+test_expect_success 'word diff with runs of whitespace' '\n+\n+\tword_diff --color-words\n+\n+'\n+\n+test_done\n-- \n1.6.1.315.g92577\n"},{"id":"100859","messageId":"1232209788-10408-5-git-send-email-trast@student.ethz.ch","threadId":"17092","inReplyTo":"1232209788-10408-4-git-send-email-trast@student.ethz.ch","subject":"[PATCH v4 4/7] color-words: take an optional regular expression describing words","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-17T16:29:45Z","receivedAt":"2009-01-17T16:29:45Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nIn some applications, words are not delimited by white space.  To\nallow for that, you can specify a regular expression describing\nwhat makes a word with\n\n\tgit diff --color-words='[A-Za-z0-9]+'\n\nNote that words cannot contain newline characters.\n\nAs suggested by Thomas Rast, the words are the exact matches of the\nregular expression.\n\nNote that a regular expression beginning with a '^' will match only\na word at the beginning of the hunk, not a word at the beginning of\na line, and is probably not what you want.\n\nThis commit contains a quoting fix by Thomas Rast.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/diff-options.txt |    6 +++-\n diff.c                         |   64 ++++++++++++++++++++++++++++++++++-----\n diff.h                         |    1 +\n t/t4034-diff-words.sh          |   57 +++++++++++++++++++++++++++++++++++\n 4 files changed, 118 insertions(+), 10 deletions(-)\n\ndiff --git a/Documentation/diff-options.txt b/Documentation/diff-options.txt\nindex 43793d7..2c1fa4b 100644\n--- a/Documentation/diff-options.txt\n+++ b/Documentation/diff-options.txt\n@@ -91,8 +91,12 @@ endif::git-format-patch[]\n \tTurn off colored diff, even when the configuration file\n \tgives the default to color output.\n \n---color-words::\n+--color-words[=regex]::\n \tShow colored word diff, i.e. color words which have changed.\n++\n+Optionally, you can pass a regular expression that tells Git what the\n+words are that you are looking for; The default is to interpret any\n+stretch of non-whitespace as a word.\n \n --no-renames::\n \tTurn off rename detection, even when the configuration\ndiff --git a/diff.c b/diff.c\nindex 37c886a..9fb3d0d 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -333,12 +333,14 @@ static void diff_words_append(char *line, unsigned long len,\n \tlen--;\n \tmemcpy(buffer->text.ptr + buffer->text.size, line, len);\n \tbuffer->text.size += len;\n+\tbuffer->text.ptr[buffer->text.size] = '\\0';\n }\n \n struct diff_words_data {\n \tstruct diff_words_buffer minus, plus;\n \tconst char *current_plus;\n \tFILE *file;\n+\tregex_t *word_regex;\n };\n \n static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n@@ -382,17 +384,49 @@ static void fn_out_diff_words_aux(void *priv, char *line, unsigned long len)\n \tdiff_words->current_plus = plus_end;\n }\n \n+/* This function starts looking at *begin, and returns 0 iff a word was found. */\n+static int find_word_boundaries(mmfile_t *buffer, regex_t *word_regex,\n+\t\tint *begin, int *end)\n+{\n+\tif (word_regex && *begin < buffer->size) {\n+\t\tregmatch_t match[1];\n+\t\tif (!regexec(word_regex, buffer->ptr + *begin, 1, match, 0)) {\n+\t\t\tchar *p = memchr(buffer->ptr + *begin + match[0].rm_so,\n+\t\t\t\t\t'\\n', match[0].rm_eo - match[0].rm_so);\n+\t\t\t*end = p ? p - buffer->ptr : match[0].rm_eo + *begin;\n+\t\t\t*begin += match[0].rm_so;\n+\t\t\treturn *begin >= *end;\n+\t\t}\n+\t\treturn -1;\n+\t}\n+\n+\t/* find the next word */\n+\twhile (*begin < buffer->size && isspace(buffer->ptr[*begin]))\n+\t\t(*begin)++;\n+\tif (*begin >= buffer->size)\n+\t\treturn -1;\n+\n+\t/* find the end of the word */\n+\t*end = *begin + 1;\n+\twhile (*end < buffer->size && !isspace(buffer->ptr[*end]))\n+\t\t(*end)++;\n+\n+\treturn 0;\n+}\n+\n /*\n  * This function splits the words in buffer->text, stores the list with\n  * newline separator into out, and saves the offsets of the original words\n  * in buffer->orig.\n  */\n-static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n+static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out,\n+\t\tregex_t *word_regex)\n {\n \tint i, j;\n+\tlong alloc = 0;\n \n \tout->size = 0;\n-\tout->ptr = xmalloc(buffer->text.size);\n+\tout->ptr = NULL;\n \n \t/* fake an empty \"0th\" word */\n \tALLOC_GROW(buffer->orig, 1, buffer->orig_alloc);\n@@ -400,11 +434,8 @@ static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n \tbuffer->orig_nr = 1;\n \n \tfor (i = 0; i < buffer->text.size; i++) {\n-\t\tif (isspace(buffer->text.ptr[i]))\n-\t\t\tcontinue;\n-\t\tfor (j = i + 1; j < buffer->text.size &&\n-\t\t\t\t!isspace(buffer->text.ptr[j]); j++)\n-\t\t\t; /* find the end of the word */\n+\t\tif (find_word_boundaries(&buffer->text, word_regex, &i, &j))\n+\t\t\treturn;\n \n \t\t/* store original boundaries */\n \t\tALLOC_GROW(buffer->orig, buffer->orig_nr + 1,\n@@ -414,6 +445,7 @@ static void diff_words_fill(struct diff_words_buffer *buffer, mmfile_t *out)\n \t\tbuffer->orig_nr++;\n \n \t\t/* store one word */\n+\t\tALLOC_GROW(out->ptr, out->size + j - i + 1, alloc);\n \t\tmemcpy(out->ptr + out->size, buffer->text.ptr + i, j - i);\n \t\tout->ptr[out->size + j - i] = '\\n';\n \t\tout->size += j - i + 1;\n@@ -443,9 +475,10 @@ static void diff_words_show(struct diff_words_data *diff_words)\n \n \tmemset(&xpp, 0, sizeof(xpp));\n \tmemset(&xecfg, 0, sizeof(xecfg));\n-\tdiff_words_fill(&diff_words->minus, &minus);\n-\tdiff_words_fill(&diff_words->plus, &plus);\n+\tdiff_words_fill(&diff_words->minus, &minus, diff_words->word_regex);\n+\tdiff_words_fill(&diff_words->plus, &plus, diff_words->word_regex);\n \txpp.flags = XDF_NEED_MINIMAL;\n+\t/* as only the hunk header will be parsed, we need a 0-context */\n \txecfg.ctxlen = 0;\n \txdi_diff_outf(&minus, &plus, fn_out_diff_words_aux, diff_words,\n \t\t      &xpp, &xecfg, &ecb);\n@@ -484,6 +517,7 @@ static void free_diff_words_data(struct emit_callback *ecbdata)\n \t\tfree (ecbdata->diff_words->minus.orig);\n \t\tfree (ecbdata->diff_words->plus.text.ptr);\n \t\tfree (ecbdata->diff_words->plus.orig);\n+\t\tfree(ecbdata->diff_words->word_regex);\n \t\tfree(ecbdata->diff_words);\n \t\tecbdata->diff_words = NULL;\n \t}\n@@ -1506,6 +1540,14 @@ static void builtin_diff(const char *name_a,\n \t\t\tecbdata.diff_words =\n \t\t\t\txcalloc(1, sizeof(struct diff_words_data));\n \t\t\tecbdata.diff_words->file = o->file;\n+\t\t\tif (o->word_regex) {\n+\t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n+\t\t\t\t\txmalloc(sizeof(regex_t));\n+\t\t\t\tif (regcomp(ecbdata.diff_words->word_regex,\n+\t\t\t\t\t\to->word_regex, REG_EXTENDED))\n+\t\t\t\t\tdie (\"Invalid regular expression: %s\",\n+\t\t\t\t\t\t\to->word_regex);\n+\t\t\t}\n \t\t}\n \t\txdi_diff_outf(&mf1, &mf2, fn_out_consume, &ecbdata,\n \t\t\t      &xpp, &xecfg, &ecb);\n@@ -2517,6 +2559,10 @@ int diff_opt_parse(struct diff_options *options, const char **av, int ac)\n \t\tDIFF_OPT_CLR(options, COLOR_DIFF);\n \telse if (!strcmp(arg, \"--color-words\"))\n \t\toptions->flags |= DIFF_OPT_COLOR_DIFF | DIFF_OPT_COLOR_DIFF_WORDS;\n+\telse if (!prefixcmp(arg, \"--color-words=\")) {\n+\t\toptions->flags |= DIFF_OPT_COLOR_DIFF | DIFF_OPT_COLOR_DIFF_WORDS;\n+\t\toptions->word_regex = arg + 14;\n+\t}\n \telse if (!strcmp(arg, \"--exit-code\"))\n \t\tDIFF_OPT_SET(options, EXIT_WITH_STATUS);\n \telse if (!strcmp(arg, \"--quiet\"))\ndiff --git a/diff.h b/diff.h\nindex 4d5a327..23cd90c 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -98,6 +98,7 @@ struct diff_options {\n \n \tint stat_width;\n \tint stat_name_width;\n+\tconst char *word_regex;\n \n \t/* this is set by diffcore for DIFF_FORMAT_PATCH */\n \tint found_changes;\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex b22195f..4873486 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -63,4 +63,61 @@ test_expect_success 'word diff with runs of whitespace' '\n \n '\n \n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 330b04f..5ed8eff 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1,3 +1,7 @@<RESET>\n+h(4),<GREEN>hh<RESET>[44]\n+<RESET>\n+a = b + c<RESET>\n+\n+<GREEN>aa = a<RESET>\n+\n+<GREEN>aeff = aeff * ( aaa<RESET> )\n+EOF\n+\n+test_expect_success 'word diff with a regular expression' '\n+\n+\tword_diff --color-words=\"[a-z]+\"\n+\n+'\n+\n+echo 'aaa (aaa)' > pre\n+echo 'aaa (aaa) aaa' > post\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index c29453b..be22f37 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1 +1 @@<RESET>\n+aaa (aaa) <GREEN>aaa<RESET>\n+EOF\n+\n+test_expect_success 'test parsing words for newline' '\n+\n+\tword_diff --color-words=\"a+\"\n+\n+'\n+\n+echo '(:' > pre\n+echo '(' > post\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 289cb9d..2d06f37 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1 +1 @@<RESET>\n+(<RED>:<RESET>\n+EOF\n+\n+test_expect_success 'test when words are only removed at the end' '\n+\n+\tword_diff --color-words=.\n+\n+'\n+\n test_done\n-- \n1.6.1.315.g92577\n"},{"id":"100862","messageId":"1232209788-10408-6-git-send-email-trast@student.ethz.ch","threadId":"17092","inReplyTo":"1232209788-10408-5-git-send-email-trast@student.ethz.ch","subject":"[PATCH v4 5/7] color-words: enable REG_NEWLINE to help user","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-17T16:29:46Z","receivedAt":"2009-01-17T16:29:46Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"We silently truncate a match at the newline, which may lead to\nunexpected behaviour, e.g., when matching \"<[^>]*>\" against\n\n  <foo\n  bar>\n\nsince then \"<foo\" becomes a word (and \"bar>\" doesn't!) even though the\nregex said only angle-bracket-delimited things can be words.\n\nTo alleviate the problem slightly, use REG_NEWLINE so that negated\nclasses can't match a newline.  Of course newlines can still be\nmatched explicitly.\n\nSigned-off-by: Thomas Rast <trast@student.ethz.ch>\n---\n diff.c |    3 ++-\n 1 files changed, 2 insertions(+), 1 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex 9fb3d0d..00c661f 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -1544,7 +1544,8 @@ static void builtin_diff(const char *name_a,\n \t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n \t\t\t\t\txmalloc(sizeof(regex_t));\n \t\t\t\tif (regcomp(ecbdata.diff_words->word_regex,\n-\t\t\t\t\t\to->word_regex, REG_EXTENDED))\n+\t\t\t\t\t\to->word_regex,\n+\t\t\t\t\t\tREG_EXTENDED | REG_NEWLINE))\n \t\t\t\t\tdie (\"Invalid regular expression: %s\",\n \t\t\t\t\t\t\to->word_regex);\n \t\t\t}\n-- \n1.6.1.315.g92577\n"},{"id":"100861","messageId":"1232209788-10408-7-git-send-email-trast@student.ethz.ch","threadId":"17092","inReplyTo":"1232209788-10408-6-git-send-email-trast@student.ethz.ch","subject":"[PATCH v4 6/7] color-words: expand docs with precise semantics","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-17T16:29:47Z","receivedAt":"2009-01-17T16:29:47Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Signed-off-by: Thomas Rast <trast@student.ethz.ch>\n---\n Documentation/diff-options.txt |   15 ++++++++++-----\n 1 files changed, 10 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/diff-options.txt b/Documentation/diff-options.txt\nindex 2c1fa4b..8689a92 100644\n--- a/Documentation/diff-options.txt\n+++ b/Documentation/diff-options.txt\n@@ -91,12 +91,17 @@ endif::git-format-patch[]\n \tTurn off colored diff, even when the configuration file\n \tgives the default to color output.\n \n---color-words[=regex]::\n-\tShow colored word diff, i.e. color words which have changed.\n+--color-words[=<regex>]::\n+\tShow colored word diff, i.e., color words which have changed.\n+\tBy default, words are separated by whitespace.\n +\n-Optionally, you can pass a regular expression that tells Git what the\n-words are that you are looking for; The default is to interpret any\n-stretch of non-whitespace as a word.\n+When a <regex> is specified, every non-overlapping match of the\n+<regex> is considered a word.  Anything between these matches is\n+considered whitespace and ignored(!) for the purposes of finding\n+differences.  You may want to append `|[^[:space:]]` to your regular\n+expression to make sure that it matches all non-whitespace characters.\n+A match that contains a newline is silently truncated(!) at the\n+newline.\n \n --no-renames::\n \tTurn off rename detection, even when the configuration\n-- \n1.6.1.315.g92577\n"},{"id":"100858","messageId":"1232209788-10408-8-git-send-email-trast@student.ethz.ch","threadId":"17092","inReplyTo":"1232209788-10408-7-git-send-email-trast@student.ethz.ch","subject":"[PATCH v4 7/7] color-words: make regex configurable via attributes","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-17T16:29:48Z","receivedAt":"2009-01-17T16:29:48Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Make the --color-words splitting regular expression configurable via\nthe diff driver's 'wordregex' attribute.  The user can then set the\ndriver on a file in .gitattributes.  If a regex is given on the\ncommand line, it overrides the driver's setting.\n\nWe also provide built-in regexes for the languages that already had\nfuncname patterns, and add an appropriate diff driver entry for C/++.\n(The patterns are designed to run UTF-8 sequences into a single chunk\nto make sure they remain readable.)\n\nSigned-off-by: Thomas Rast <trast@student.ethz.ch>\n---\n Documentation/diff-options.txt  |    4 ++\n Documentation/gitattributes.txt |   21 ++++++++++\n diff.c                          |   10 +++++\n t/t4034-diff-words.sh           |   36 ++++++++++++++++++\n userdiff.c                      |   78 +++++++++++++++++++++++++++++++-------\n userdiff.h                      |    1 +\n 6 files changed, 135 insertions(+), 15 deletions(-)\n\ndiff --git a/Documentation/diff-options.txt b/Documentation/diff-options.txt\nindex 8689a92..1edb82e 100644\n--- a/Documentation/diff-options.txt\n+++ b/Documentation/diff-options.txt\n@@ -102,6 +102,10 @@ differences.  You may want to append `|[^[:space:]]` to your regular\n expression to make sure that it matches all non-whitespace characters.\n A match that contains a newline is silently truncated(!) at the\n newline.\n++\n+The regex can also be set via a diff driver, see\n+linkgit:gitattributes[1]; giving it explicitly overrides any diff\n+driver setting.\n \n --no-renames::\n \tTurn off rename detection, even when the configuration\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 8af22ec..ba3ba12 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -317,6 +317,8 @@ patterns are available:\n \n - `bibtex` suitable for files with BibTeX coded references.\n \n+- `cpp` suitable for source code in the C and C++ languages.\n+\n - `html` suitable for HTML/XHTML documents.\n \n - `java` suitable for source code in the Java language.\n@@ -334,6 +336,25 @@ patterns are available:\n - `tex` suitable for source code for LaTeX documents.\n \n \n+Customizing word diff\n+^^^^^^^^^^^^^^^^^^^^^\n+\n+You can customize the rules that `git diff --color-words` uses to\n+split words in a line, by specifying an appropriate regular expression\n+in the \"diff.*.wordregex\" configuration variable.  For example, in TeX\n+a backslash followed by a sequence of letters forms a command, but\n+several such commands can be run together without intervening\n+whitespace.  To separate them, use a regular expression such as\n+\n+------------------------\n+[diff \"tex\"]\n+\twordregex = \"\\\\\\\\[a-zA-Z]+|[{}]|\\\\\\\\.|[^\\\\{}[:space:]]+\"\n+------------------------\n+\n+A built-in pattern is provided for all languages listed in the\n+previous section.\n+\n+\n Performing text diffs of binary files\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n \ndiff --git a/diff.c b/diff.c\nindex 00c661f..9fcde96 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -1380,6 +1380,12 @@ int diff_filespec_is_binary(struct diff_filespec *one)\n \treturn one->driver->funcname.pattern ? &one->driver->funcname : NULL;\n }\n \n+static const char *userdiff_word_regex(struct diff_filespec *one)\n+{\n+\tdiff_filespec_load_driver(one);\n+\treturn one->driver->word_regex;\n+}\n+\n void diff_set_mnemonic_prefix(struct diff_options *options, const char *a, const char *b)\n {\n \tif (!options->a_prefix)\n@@ -1540,6 +1546,10 @@ static void builtin_diff(const char *name_a,\n \t\t\tecbdata.diff_words =\n \t\t\t\txcalloc(1, sizeof(struct diff_words_data));\n \t\t\tecbdata.diff_words->file = o->file;\n+\t\t\tif (!o->word_regex)\n+\t\t\t\to->word_regex = userdiff_word_regex(one);\n+\t\t\tif (!o->word_regex)\n+\t\t\t\to->word_regex = userdiff_word_regex(two);\n \t\t\tif (o->word_regex) {\n \t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n \t\t\t\t\txmalloc(sizeof(regex_t));\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 4873486..744221b 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -84,6 +84,41 @@ test_expect_success 'word diff with a regular expression' '\n \n '\n \n+test_expect_success 'set a diff driver' '\n+\tgit config diff.testdriver.wordregex \"[^[:space:]]\" &&\n+\tcat <<EOF > .gitattributes\n+pre diff=testdriver\n+post diff=testdriver\n+EOF\n+'\n+\n+test_expect_success 'option overrides default' '\n+\n+\tword_diff --color-words=\"[a-z]+\"\n+\n+'\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 330b04f..5ed8eff 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1,3 +1,7 @@<RESET>\n+h(4)<GREEN>,hh[44]<RESET>\n+<RESET>\n+a = b + c<RESET>\n+\n+<GREEN>aa = a<RESET>\n+\n+<GREEN>aeff = aeff * ( aaa )<RESET>\n+EOF\n+\n+test_expect_success 'use default supplied by driver' '\n+\n+\tword_diff --color-words\n+\n+'\n+\n echo 'aaa (aaa)' > pre\n echo 'aaa (aaa) aaa' > post\n \n@@ -100,6 +135,7 @@ test_expect_success 'test parsing words for newline' '\n \n \tword_diff --color-words=\"a+\"\n \n+\n '\n \n echo '(:' > pre\ndiff --git a/userdiff.c b/userdiff.c\nindex 3681062..2b55509 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -6,14 +6,20 @@\n static int ndrivers;\n static int drivers_alloc;\n \n-#define FUNCNAME(name, pattern) \\\n-\t{ name, NULL, -1, { pattern, REG_EXTENDED } }\n+#define PATTERNS(name, pattern, wordregex)\t\t\t\\\n+\t{ name, NULL, -1, { pattern, REG_EXTENDED }, wordregex }\n static struct userdiff_driver builtin_drivers[] = {\n-FUNCNAME(\"html\", \"^[ \\t]*(<[Hh][1-6][ \\t].*>.*)$\"),\n-FUNCNAME(\"java\",\n+PATTERNS(\"html\", \"^[ \\t]*(<[Hh][1-6][ \\t].*>.*)$\",\n+\t \"[^<>= \\t]+|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"java\",\n \t \"!^[ \\t]*(catch|do|for|if|instanceof|new|return|switch|throw|while)\\n\"\n-\t \"^[ \\t]*(([ \\t]*[A-Za-z_][A-Za-z_0-9]*){2,}[ \\t]*\\\\([^;]*)$\"),\n-FUNCNAME(\"objc\",\n+\t \"^[ \\t]*(([ \\t]*[A-Za-z_][A-Za-z_0-9]*){2,}[ \\t]*\\\\([^;]*)$\",\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[fFlL]?|0[xXbB]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=\"\n+\t \"|--|\\\\+\\\\+|<<=?|>>>?=?|&&|\\\\|\\\\|\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"objc\",\n \t /* Negate C statements that can look like functions */\n \t \"!^[ \\t]*(do|for|if|else|return|switch|while)\\n\"\n \t /* Objective-C methods */\n@@ -21,20 +27,60 @@\n \t /* C functions */\n \t \"^[ \\t]*(([ \\t]*[A-Za-z_][A-Za-z_0-9]*){2,}[ \\t]*\\\\([^;]*)$\\n\"\n \t /* Objective-C class/protocol definitions */\n-\t \"^(@(implementation|interface|protocol)[ \\t].*)$\"),\n-FUNCNAME(\"pascal\",\n+\t \"^(@(implementation|interface|protocol)[ \\t].*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[fFlL]?|0[xXbB]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=|--|\\\\+\\\\+|<<=?|>>=?|&&|\\\\|\\\\||::|->\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"pascal\",\n \t \"^((procedure|function|constructor|destructor|interface|\"\n \t\t\"implementation|initialization|finalization)[ \\t]*.*)$\"\n \t \"\\n\"\n-\t \"^(.*=[ \\t]*(class|record).*)$\"),\n-FUNCNAME(\"php\", \"^[\\t ]*((function|class).*)\"),\n-FUNCNAME(\"python\", \"^[ \\t]*((class|def)[ \\t].*)$\"),\n-FUNCNAME(\"ruby\", \"^[ \\t]*((class|module|def)[ \\t].*)$\"),\n-FUNCNAME(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\"),\n-FUNCNAME(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\"),\n+\t \"^(.*=[ \\t]*(class|record).*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+|0[xXbB]?[0-9a-fA-F]+\"\n+\t \"|<>|<=|>=|:=|\\\\.\\\\.\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"php\", \"^[\\t ]*((function|class).*)\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+|0[xXbB]?[0-9a-fA-F]+\"\n+\t \"|[-+*/<>%&^|=!.]=|--|\\\\+\\\\+|<<=?|>>=?|===|&&|\\\\|\\\\||::|->\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n+PATTERNS(\"python\", \"^[ \\t]*((class|def)[ \\t].*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[jJlL]?|0[xX]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=|//=?|<<=?|>>=?|\\\\*\\\\*=?\"\n+\t \"|[^[:space:]|[\\x80-\\xff]+\"),\n+\t /* -- */\n+PATTERNS(\"ruby\", \"^[ \\t]*((class|module|def)[ \\t].*)$\",\n+\t /* -- */\n+\t \"(@|@@|\\\\$)?[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+|0[xXbB]?[0-9a-fA-F]+|\\\\?(\\\\\\\\C-)?(\\\\\\\\M-)?.\"\n+\t \"|//=?|[-+*/<>%&^|=!]=|<<=?|>>=?|===|\\\\.{1,3}|::|[!=]~\"\n+\t \"|[^[:space:]|[\\x80-\\xff]+\"),\n+PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n+\t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n+PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n+\t \"\\\\\\\\[a-zA-Z@]+|\\\\\\\\.|[a-zA-Z0-9\\x80-\\xff]+|[^[:space:]]\"),\n+PATTERNS(\"cpp\",\n+\t /* Jump targets or access declarations */\n+\t \"!^[ \\t]*[A-Za-z_][A-Za-z_0-9]*:.*$\\n\"\n+\t /* C/++ functions/methods at top level */\n+\t \"^([A-Za-z_][A-Za-z_0-9]*([ \\t]+[A-Za-z_][A-Za-z_0-9]*([ \\t]*::[ \\t]*[^[:space:]]+)?){1,}[ \\t]*\\\\([^;]*)$\\n\"\n+\t /* compound type at top level */\n+\t \"^((struct|class|enum)[^;]*)$\",\n+\t /* -- */\n+\t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n+\t \"|[-+0-9.e]+[fFlL]?|0[xXbB]?[0-9a-fA-F]+[lL]?\"\n+\t \"|[-+*/<>%&^|=!]=|--|\\\\+\\\\+|<<=?|>>=?|&&|\\\\|\\\\||::|->\"\n+\t \"|[^[:space:]]|[\\x80-\\xff]+\"),\n { \"default\", NULL, -1, { NULL, 0 } },\n };\n-#undef FUNCNAME\n+#undef PATTERNS\n \n static struct userdiff_driver driver_true = {\n \t\"diff=true\",\n@@ -134,6 +180,8 @@ int userdiff_config(const char *k, const char *v)\n \t\treturn parse_string(&drv->external, k, v);\n \tif ((drv = parse_driver(k, v, \"textconv\")))\n \t\treturn parse_string(&drv->textconv, k, v);\n+\tif ((drv = parse_driver(k, v, \"wordregex\")))\n+\t\treturn parse_string(&drv->word_regex, k, v);\n \n \treturn 0;\n }\ndiff --git a/userdiff.h b/userdiff.h\nindex ba29457..c315159 100644\n--- a/userdiff.h\n+++ b/userdiff.h\n@@ -11,6 +11,7 @@ struct userdiff_driver {\n \tconst char *external;\n \tint binary;\n \tstruct userdiff_funcname funcname;\n+\tconst char *word_regex;\n \tconst char *textconv;\n };\n \n-- \n1.6.1.315.g92577\n"},{"id":"100987","messageId":"adf1fd3d0901180705s260f0051wb4e3a978601618ec@mail.gmail.com","threadId":"17092","inReplyTo":"1232209788-10408-1-git-send-email-trast@student.ethz.ch","subject":"Re: [PATCH v4 0/7] customizable --color-words","fromName":"Santi Béjar","fromEmail":"santi@agolina.net","sentAt":"2009-01-18T15:05:49Z","receivedAt":"2009-01-18T15:05:49Z","isPatch":true,"sender":{"key":"santi@agolina.net","avatar":null},"body":"2009/1/17 Thomas Rast <trast@student.ethz.ch>:\n> Johannes Schindelin wrote:\n>> Thomas, could you pick up the patches from my 'my-next' branch and\n>> maintain an \"official\" topic branch?\n>\n> I cherry-picked the three commits you had there, and rebuilt on top.\n> I pushed them to\n>\n>  git://repo.or.cz/git/trast.git tr/word-diff-p2\n>\n> again (js/word-diff-p1 again points directly at your half).\n\nI've tested tr/word-diff-p2 and I have not found any issues. I've even\ntested that nothing changed from the tradicional word diff to:\n\ngit log -p --color-words=\"[^[:space:]]+\"\n\nfor the whole git history.\n\nAt the end I've found that a general regex that works best for me is:\n\n\"[[:alpha:]]+|[[:digit:]]+|[^[:alnum:][:space:]]\"\n\nand that is what I tested.\n\nSanti\n"},{"id":"101002","messageId":"adf1fd3d0901180729u3da69108i6140aa7f68bda972@mail.gmail.com","threadId":"17092","inReplyTo":"adf1fd3d0901180705s260f0051wb4e3a978601618ec@mail.gmail.com","subject":"Re: [PATCH v4 0/7] customizable --color-words","fromName":"Santi Béjar","fromEmail":"santi@agolina.net","sentAt":"2009-01-18T15:29:54Z","receivedAt":"2009-01-18T15:29:54Z","isPatch":true,"sender":{"key":"santi@agolina.net","avatar":null},"body":"2009/1/18 Santi Béjar <santi@agolina.net>:\n> 2009/1/17 Thomas Rast <trast@student.ethz.ch>:\n>> Johannes Schindelin wrote:\n>>> Thomas, could you pick up the patches from my 'my-next' branch and\n>>> maintain an \"official\" topic branch?\n>>\n>> I cherry-picked the three commits you had there, and rebuilt on top.\n>> I pushed them to\n>>\n>>  git://repo.or.cz/git/trast.git tr/word-diff-p2\n>>\n>> again (js/word-diff-p1 again points directly at your half).\n>\n> I've tested tr/word-diff-p2 and I have not found any issues. I've even\n> tested that nothing changed from the tradicional word diff to:\n>\n> git log -p --color-words=\"[^[:space:]]+\"\n>\n> for the whole git history.\n>\n\nWhat I tested is that the new code produces the same result for this\ntwo commands:\n\ngit log -p --color-words=\"[^[:space:]]+\"\ngit log -p --color-words\n\nThe old code produced color codes before and after each word, while\nthe new only at the begining of the color and the end of the color. So\nthey cannot produce the same output but equivalent.\n\nSanti\n"},{"id":"101174","messageId":"adf1fd3d0901191447n7fc39dect9cf5afd88a02015b@mail.gmail.com","threadId":"17092","inReplyTo":"1232209788-10408-1-git-send-email-trast@student.ethz.ch","subject":"Re: [PATCH v4 0/7] customizable --color-words","fromName":"Santi Béjar","fromEmail":"santi@agolina.net","sentAt":"2009-01-19T22:47:26Z","receivedAt":"2009-01-19T22:47:26Z","isPatch":true,"sender":{"key":"santi@agolina.net","avatar":null},"body":"2009/1/17 Thomas Rast <trast@student.ethz.ch>:\n> Johannes Schindelin (4):\n>  Add color_fwrite_lines(), a function coloring each line individually\n>  color-words: refactor word splitting and use ALLOC_GROW()\n>  color-words: change algorithm to allow for 0-character word\n>    boundaries\n>  color-words: take an optional regular expression describing words\n>\n> Thomas Rast (3):\n>  color-words: enable REG_NEWLINE to help user\n>  color-words: expand docs with precise semantics\n>  color-words: make regex configurable via attributes\n>\n\nAlso, having a config (diff.color-words?) to set the default regexp\nwould be great. Thanks.\n\nSanti\n"},{"id":"101181","messageId":"alpine.DEB.1.00.0901200031350.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"adf1fd3d0901191447n7fc39dect9cf5afd88a02015b@mail.gmail.com","subject":"Re: [PATCH v4 0/7] customizable --color-words","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-19T23:35:59Z","receivedAt":"2009-01-19T23:35:59Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 19 Jan 2009, Santi Béjar wrote:\n\n> 2009/1/17 Thomas Rast <trast@student.ethz.ch>:\n> > Johannes Schindelin (4):\n> >  Add color_fwrite_lines(), a function coloring each line individually\n> >  color-words: refactor word splitting and use ALLOC_GROW()\n> >  color-words: change algorithm to allow for 0-character word\n> >    boundaries\n> >  color-words: take an optional regular expression describing words\n> >\n> > Thomas Rast (3):\n> >  color-words: enable REG_NEWLINE to help user\n> >  color-words: expand docs with precise semantics\n> >  color-words: make regex configurable via attributes\n> >\n> \n> Also, having a config (diff.color-words?) to set the default regexp\n> would be great. Thanks.\n\n>From \"git log --author==Santi --stat\" it seems that you are quite capable \nof providing that patch.\n\nA few pointers:\n\n- Add a global variable to diff.c, maybe \"char *diff_word_regex\".  \n  (Maybe it should be static instead, as it will be used in diff.c only.)\n\n- Add code to set it in diff.c, function git_diff_ui_config().\n\n- In diff.c, where \"--color-words\" is handled (without \"=\"), add\n\n\tif (diff_words_regex)\n\t\toptions->word_regex = diff_word_regex;\n\n- Add a test to t4034 that tests that the config sets a default, and that \n  the command line can override it.\n\n- Send to this list :-)\n\nCiao,\nDscho\n"},{"id":"101203","messageId":"200901192017.54163.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901200031350.3586@pacific.mpi-cbg.de","subject":"[PATCH] Add tests for diff.color-words configuration option.","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-20T02:17:53Z","receivedAt":"2009-01-20T02:17:53Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"Signed-Off-By: Boyd Stephen Smith Jr. <bss@iguanasuicide.net>\n---\nOn Monday 19 January 2009, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote about 'Re: [PATCH v4 \n0/7] customizable --color-words':\n>On Mon, 19 Jan 2009, Santi Béjar wrote:\n>> Also, having a config (diff.color-words?) to set the default regexp\n>> would be great. Thanks.\n>\n>From \"git log --author==Santi --stat\" it seems that you are quite capable\n>of providing that patch.\n>\n>A few pointers:\n>\n>- Add a test to t4034 that tests that the config sets a default, and that\n>  the command line can override it.\n\nHere's a couple tests to get someone started, adds one \"known breakage\" to\nthe results of the test suite.  This is to be applied on top of\nthe existing patches.\n\nYes, I also think I'll work on the actual implementation, but I'd be glad\nto have someone beat me to it.  I'm not sure why the diff is crazy long.\n\n t/t4034-diff-words.sh |   50 +++++++++++++++++++++++++++++++++++-------------\n 1 files changed, 36 insertions(+), 14 deletions(-)\n\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 744221b..6ebce9d 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -63,7 +63,7 @@ test_expect_success 'word diff with runs of whitespace' '\n \n '\n \n-cat > expect <<\\EOF\n+cat > expect.letter-runs-are-words <<\\EOF\n <WHITE>diff --git a/pre b/post<RESET>\n <WHITE>index 330b04f..5ed8eff 100644<RESET>\n <WHITE>--- a/pre<RESET>\n@@ -77,6 +77,7 @@ a = b + c<RESET>\n \n <GREEN>aeff = aeff * ( aaa<RESET> )\n EOF\n+cp expect.letter-runs-are-words expect\n \n test_expect_success 'word diff with a regular expression' '\n \n@@ -84,21 +85,11 @@ test_expect_success 'word diff with a regular expression' '\n \n '\n \n-test_expect_success 'set a diff driver' '\n-\tgit config diff.testdriver.wordregex \"[^[:space:]]\" &&\n-\tcat <<EOF > .gitattributes\n-pre diff=testdriver\n-post diff=testdriver\n-EOF\n-'\n-\n-test_expect_success 'option overrides default' '\n-\n-\tword_diff --color-words=\"[a-z]+\"\n-\n+test_expect_success 'add configuration for default regex' '\n+\tgit config diff.color-words \"[^[:space:]]\"\n '\n \n-cat > expect <<\\EOF\n+cat > expect.non-whitespace-is-word <<\\EOF\n <WHITE>diff --git a/pre b/post<RESET>\n <WHITE>index 330b04f..5ed8eff 100644<RESET>\n <WHITE>--- a/pre<RESET>\n@@ -112,6 +103,37 @@ a = b + c<RESET>\n \n <GREEN>aeff = aeff * ( aaa )<RESET>\n EOF\n+cp expect.non-whitespace-is-word expect\n+\n+test_expect_failure 'use default supplied by config' '\n+\n+\tword_diff --color-words\n+\n+'\n+\n+cp expect.letter-runs-are-words expect\n+\n+test_expect_success 'option overrides config-default' '\n+\n+\tword_diff --color-words=\"[a-z]+\"\n+\n+'\n+\n+test_expect_success 'set a diff driver' '\n+\tgit config diff.testdriver.wordregex \"[^[:space:]]\" &&\n+\tcat <<EOF > .gitattributes\n+pre diff=testdriver\n+post diff=testdriver\n+EOF\n+'\n+\n+test_expect_success 'option overrides default' '\n+\n+\tword_diff --color-words=\"[a-z]+\"\n+\n+'\n+\n+cp expect.non-whitespace-is-word expect\n \n test_expect_success 'use default supplied by driver' '\n \n-- \n1.5.6.5\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"101208","messageId":"200901192145.21115.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"200901192017.54163.bss@iguanasuicide.net","subject":"[PATCH] diff: Support diff.color-words config option","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-20T03:45:20Z","receivedAt":"2009-01-20T03:45:20Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"When diff is invoked with --color-words (w/o =regex), use the regular\nexpression the user has configured as diff.color-words.\n\ndiff drivers configured via attributes take precedence over the\ndiff.color-words setting.  If the user wants to change them, they have\ntheir own configuration variables.\n\nSigned-off-by: Boyd Stephen Smith Jr <bss@iguanasuicide.net>\n---\nOn Monday 19 January 2009, \"Boyd Stephen Smith Jr.\" <bss@iguanasuicide.net> wrote about '[PATCH] Add \ntests for diff.color-words configuration option.':\n>Yes, I also think I'll work on the actual implementation, but I'd be glad\n>to have someone beat me to it.  I'm not sure why the diff is crazy long.\n\nHere's a patch that makes the added test case succeed, but I think it and\nthe tests themselves should probably be reworked.  Hopefully, this doesn't\nshow up in quoted-printable format (damn you kmail).\n\nWhile it might be a corner-case, we probably need a test of some sort for\nwhen a user/system has a global diff.color-words configuration wants\nto have a single repository (or single run of 'git diff') use the default\nalgorithm. I.e. run as if no regex had been set.\n\n diff.c                |    5 +++++\n t/t4034-diff-words.sh |    2 +-\n 2 files changed, 6 insertions(+), 1 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex 9fcde96..c53e1d1 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -23,6 +23,7 @@ static int diff_detect_rename_default;\n static int diff_rename_limit_default = 200;\n static int diff_suppress_blank_empty;\n int diff_use_color_default = -1;\n+static const char *diff_color_words_cfg = NULL;\n static const char *external_diff_cmd_cfg;\n int diff_auto_refresh_index = 1;\n static int diff_mnemonic_prefix;\n@@ -92,6 +93,8 @@ int git_diff_ui_config(const char *var, const char *value, void *cb)\n \t}\n \tif (!strcmp(var, \"diff.external\"))\n \t\treturn git_config_string(&external_diff_cmd_cfg, var, value);\n+\tif (!strcmp(var, \"diff.color-words\"))\n+\t\treturn git_config_string(&diff_color_words_cfg, var, value);\n \n \treturn git_diff_basic_config(var, value, cb);\n }\n@@ -1550,6 +1553,8 @@ static void builtin_diff(const char *name_a,\n \t\t\t\to->word_regex = userdiff_word_regex(one);\n \t\t\tif (!o->word_regex)\n \t\t\t\to->word_regex = userdiff_word_regex(two);\n+\t\t\tif (!o->word_regex)\n+\t\t\t\to->word_regex = diff_color_words_cfg;\n \t\t\tif (o->word_regex) {\n \t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n \t\t\t\t\txmalloc(sizeof(regex_t));\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 6ebce9d..a207d9e 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -105,7 +105,7 @@ a = b + c<RESET>\n EOF\n cp expect.non-whitespace-is-word expect\n \n-test_expect_failure 'use default supplied by config' '\n+test_expect_success 'use default supplied by config' '\n \n \tword_diff --color-words\n \n-- \n1.5.6.5\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"101223","messageId":"7v1vuympie.fsf@gitster.siamese.dyndns.org","threadId":"17092","inReplyTo":"200901192145.21115.bss@iguanasuicide.net","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-20T06:59:37Z","receivedAt":"2009-01-20T06:59:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Boyd Stephen Smith Jr.\" <bss@iguanasuicide.net> writes:\n\n> When diff is invoked with --color-words (w/o =regex), use the regular\n> expression the user has configured as diff.color-words.\n>\n> diff drivers configured via attributes take precedence over the\n> diff.color-words setting.  If the user wants to change them, they have\n> their own configuration variables.\n\nThis needs an entry in Documentation/config.txt\n\nNone of the existing configuration variables defined use hyphens in\nmulti-word variable names.\n\nOther than that, I think this is a welcome addition to the suite.\n\nThanks.\n"},{"id":"101237","messageId":"alpine.DEB.1.00.0901201057080.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"200901192017.54163.bss@iguanasuicide.net","subject":"Re: [PATCH] Add tests for diff.color-words configuration option.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-20T09:58:37Z","receivedAt":"2009-01-20T09:58:37Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 19 Jan 2009, Boyd Stephen Smith Jr. wrote:\n\n> I'm not sure why the diff is crazy long.\n\nBecause you changed things that need no changing, such as \"cat > expect\" \n-> \"cat > expect.blabla\", and because you inserted your test instead of \nadding it at the end.\n\nCiao,\nDscho\n"},{"id":"101238","messageId":"alpine.DEB.1.00.0901201058520.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"200901192145.21115.bss@iguanasuicide.net","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-20T10:02:00Z","receivedAt":"2009-01-20T10:02:00Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 19 Jan 2009, Boyd Stephen Smith Jr. wrote:\n\n> diff --git a/diff.c b/diff.c\n> index 9fcde96..c53e1d1 100644\n> --- a/diff.c\n> +++ b/diff.c\n> @@ -23,6 +23,7 @@ static int diff_detect_rename_default;\n>  static int diff_rename_limit_default = 200;\n>  static int diff_suppress_blank_empty;\n>  int diff_use_color_default = -1;\n> +static const char *diff_color_words_cfg = NULL;\n>  static const char *external_diff_cmd_cfg;\n\nGuess why external_diff_cmd_cfg is not set to NULL?  All variables \ndefined outside a function are set to all-zero anyway.\n\n> @@ -92,6 +93,8 @@ int git_diff_ui_config(const char *var, const char *value, void *cb)\n>  \t}\n>  \tif (!strcmp(var, \"diff.external\"))\n>  \t\treturn git_config_string(&external_diff_cmd_cfg, var, value);\n> +\tif (!strcmp(var, \"diff.color-words\"))\n\nI'd call it diff.wordregex, because that's what it is.\n\n> @@ -1550,6 +1553,8 @@ static void builtin_diff(const char *name_a,\n>  \t\t\t\to->word_regex = userdiff_word_regex(one);\n>  \t\t\tif (!o->word_regex)\n>  \t\t\t\to->word_regex = userdiff_word_regex(two);\n> +\t\t\tif (!o->word_regex)\n> +\t\t\t\to->word_regex = diff_color_words_cfg;\n\nIMHO this is the wrong order.  config should not override attributes, \nwhich are by definition more specific.\n\n> diff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\n> index 6ebce9d..a207d9e 100755\n> --- a/t/t4034-diff-words.sh\n> +++ b/t/t4034-diff-words.sh\n> @@ -105,7 +105,7 @@ a = b + c<RESET>\n>  EOF\n>  cp expect.non-whitespace-is-word expect\n>  \n> -test_expect_failure 'use default supplied by config' '\n> +test_expect_success 'use default supplied by config' '\n\nLet's squash the two, okay?\n\nThanks,\nDscho\n"},{"id":"101254","messageId":"gl4nkp$4kq$1@ger.gmane.org","threadId":"17092","inReplyTo":"200901192145.21115.bss@iguanasuicide.net","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-01-20T14:38:48Z","receivedAt":"2009-01-20T14:38:48Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Boyd Stephen Smith Jr. wrote:\n\n> Nawiązania: 1 2 3\n> When diff is invoked with --color-words (w/o =regex), use the regular\n> expression the user has configured as diff.color-words.\n> \n> diff drivers configured via attributes take precedence over the\n> diff.color-words setting.  If the user wants to change them, they have\n> their own configuration variables.\n\nJust a nit: all other configuration variables use camelCase or runwords;\nthis would be first configuration variable with '-' as words separator,\nI think.\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"101265","messageId":"200901201034.22478.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901201057080.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH] Add tests for diff.color-words configuration option.","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-20T16:34:18Z","receivedAt":"2009-01-20T16:34:18Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"On Tuesday 2009 January 20 03:58:37 Johannes Schindelin wrote:\n>On Mon, 19 Jan 2009, Boyd Stephen Smith Jr. wrote:\n>> I'm not sure why the diff is crazy long.\n>\n>Because you changed things that need no changing, such as \"cat > expect\"\n>-> \"cat > expect.blabla\",\n\nI suppose I could have gotten away with doing this differently, but I did need \nto save off some of those results to different files because I wanted to \nresuse the results.\n\n>and because you inserted your test instead of \n>adding it at the end.\n\nI put the tests in that order explicitly to test that .gitattributes overrides \nthe configuration option.\n\nI'm going to be reworking both patches anyway, so I should be able to \nrearrange things less, in this file.\n\nThanks for the feedback.\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"101269","messageId":"200901201053.03256.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901201058520.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-20T16:52:58Z","receivedAt":"2009-01-20T16:52:58Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"On Tuesday 2009 January 20 04:02:00 you wrote:\n>On Mon, 19 Jan 2009, Boyd Stephen Smith Jr. wrote:\n>> diff --git a/diff.c b/diff.c\n>> index 9fcde96..c53e1d1 100644\n>> --- a/diff.c\n>> +++ b/diff.c\n>> @@ -23,6 +23,7 @@ static int diff_detect_rename_default;\n>>  static int diff_rename_limit_default = 200;\n>>  static int diff_suppress_blank_empty;\n>>  int diff_use_color_default = -1;\n>> +static const char *diff_color_words_cfg = NULL;\n>>  static const char *external_diff_cmd_cfg;\n>\n>Guess why external_diff_cmd_cfg is not set to NULL?  All variables\n>defined outside a function are set to all-zero anyway.\n\nI suppose I just initialize variables by reflex, having been bitten with too \nmany sometimes-crashes due to variables that were usually-zero.  Assuming C \ndoes guarantee that it is zeroed, I'll drop the \" = NULL\" line noise in the \nnext version.\n\n>> @@ -92,6 +93,8 @@ int git_diff_ui_config(const char *var, const char\n>> *value, void *cb) }\n>>  \tif (!strcmp(var, \"diff.external\"))\n>>  \t\treturn git_config_string(&external_diff_cmd_cfg, var, value);\n>> +\tif (!strcmp(var, \"diff.color-words\"))\n>\n>I'd call it diff.wordregex, because that's what it is.\n\nI don't like runtogetherwords because they are hard to read for me; I tend to \nchoose the wrong word breaks if it is ambiguous.  There are other \nconfiguration values that use camelCaseWords so I will convert over to using \nthat.\n\nI thought \"word regex\" made more sense, but I wanted to match the command-line \noption.  Will change.\n\n>> @@ -1550,6 +1553,8 @@ static void builtin_diff(const char *name_a,\n>>  \t\t\t\to->word_regex = userdiff_word_regex(one);\n>>  \t\t\tif (!o->word_regex)\n>>  \t\t\t\to->word_regex = userdiff_word_regex(two);\n>> +\t\t\tif (!o->word_regex)\n>> +\t\t\t\to->word_regex = diff_color_words_cfg;\n>\n>IMHO this is the wrong order.  config should not override attributes,\n>which are by definition more specific.\n\nYou are up too late Dscho.  This ordering makes the config not override \nattributes.  If one of the files has a diff driver, o->word_regex will be set \nto it (and become non-NULL).  That will prevent execution of the body of the \nadded \"if (!o->word_regex)\" -- preventing the configuration option from being \nused.\n\n>> diff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\n>> index 6ebce9d..a207d9e 100755\n>> --- a/t/t4034-diff-words.sh\n>> +++ b/t/t4034-diff-words.sh\n>> @@ -105,7 +105,7 @@ a = b + c<RESET>\n>>  EOF\n>>  cp expect.non-whitespace-is-word expect\n>>\n>> -test_expect_failure 'use default supplied by config' '\n>> +test_expect_success 'use default supplied by config' '\n>\n>Let's squash the two, okay?\n\nWill do.  I expected the code changes to be larger than the test, and when I \nfinished it was completely the other way.  My next patch will be all-in-one.\n\nThanks for your feedback.\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"101270","messageId":"alpine.DEB.1.00.0901201748230.5159@intel-tinevez-2-302","threadId":"17092","inReplyTo":"200901201034.22478.bss@iguanasuicide.net","subject":"Re: [PATCH] Add tests for diff.color-words configuration option.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-20T16:54:29Z","receivedAt":"2009-01-20T16:54:29Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 20 Jan 2009, Boyd Stephen Smith Jr. wrote:\n\n> On Tuesday 2009 January 20 03:58:37 Johannes Schindelin wrote:\n> >On Mon, 19 Jan 2009, Boyd Stephen Smith Jr. wrote:\n> >> I'm not sure why the diff is crazy long.\n> >\n> >Because you changed things that need no changing, such as \"cat > expect\"\n> >-> \"cat > expect.blabla\",\n> \n> I suppose I could have gotten away with doing this differently, but I \n> did need to save off some of those results to different files because I \n> wanted to resuse the results.\n\nWhy didn't you do that, then?\n\n\tcp expect expect.for-later-use\n\n> >and because you inserted your test instead of adding it at the end.\n> \n> I put the tests in that order explicitly to test that .gitattributes \n> overrides the configuration option.\n\nWhy not just remove the .gitattributes for your second test?\n\nIt would be much clearer that you did not modify any existing tests, then.\n\nCiao,\nDscho\n"},{"id":"101273","messageId":"7vskndkip9.fsf@gitster.siamese.dyndns.org","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901201058520.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-20T17:09:38Z","receivedAt":"2009-01-20T17:09:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> @@ -92,6 +93,8 @@ int git_diff_ui_config(const char *var, const char *value, void *cb)\n>>  \t}\n>>  \tif (!strcmp(var, \"diff.external\"))\n>>  \t\treturn git_config_string(&external_diff_cmd_cfg, var, value);\n>> +\tif (!strcmp(var, \"diff.color-words\"))\n>\n> I'd call it diff.wordregex, because that's what it is.\n\nIf we want to add a new word-oriented option to diff that is not about\ncoloring the word differences, is it safe and sane to reuse the same\ndefinition?  That is, \"git diff --color-words\" would be affected when\ndiff.wordregex is set to some value, so does any new word-oriented\noperation we will add, and the single regex configured would be used as\nthe default value to define how a word would look like.\n\nI think it makes sense; I do not think of a case offhand where you would\nwant to define what a word is for the purpose of coloring diffs in one\nway, and would want to use a different definition for another\nword-oriented operation.\n\n>> @@ -1550,6 +1553,8 @@ static void builtin_diff(const char *name_a,\n>>  \t\t\t\to->word_regex = userdiff_word_regex(one);\n>>  \t\t\tif (!o->word_regex)\n>>  \t\t\t\to->word_regex = userdiff_word_regex(two);\n>> +\t\t\tif (!o->word_regex)\n>> +\t\t\t\to->word_regex = diff_color_words_cfg;\n>\n> IMHO this is the wrong order.  config should not override attributes, \n> which are by definition more specific.\n\nIsn't it merely giving a fallback value when attributes does not give one?\n\nBy the way, wouldn't it make sense to optimize the precontext of that hunk\nby doing _something_ like:\n\n\tif (!o->word_regex && strcmp(one->path, two->path))\n        \to->word_regex = userdiff_word_regex(two);\n\n\"Something like\" comes from special cases like /dev/null for new/deleted\nfiles, etc.\n"},{"id":"101274","messageId":"alpine.DEB.1.00.0901201810170.5159@intel-tinevez-2-302","threadId":"17092","inReplyTo":"200901201053.03256.bss@iguanasuicide.net","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-20T17:14:20Z","receivedAt":"2009-01-20T17:14:20Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 20 Jan 2009, Boyd Stephen Smith Jr. wrote:\n\n> You are up too late Dscho.\n\nYou, sir, are absolutely correct.\n\n> >Let's squash the two, okay?\n> \n> Will do.  I expected the code changes to be larger than the test, and \n> when I finished it was completely the other way.  My next patch will be \n> all-in-one.\n\nFWIW I think it is the correct thing to start with the test script, so \nthat you get a better idea what to look out for.\n\nAnd for patches of which I don't know if they are still necessary, I like \nto \"git checkout <name>^ && make -j50 && git checkout <name> && (cd t && \nsh <test>)\".\n\nBut for submission, I think it makes sense to squash them, except if you \nsubmit a bug report with a test script to show the validity of the report \nfirst, and only later decide that you want to fix it yourself.\n\nCiao,\nDscho\n"},{"id":"101277","messageId":"alpine.DEB.1.00.0901201819490.5159@intel-tinevez-2-302","threadId":"17092","inReplyTo":"7vskndkip9.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-20T17:28:14Z","receivedAt":"2009-01-20T17:28:14Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 20 Jan 2009, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> >> @@ -92,6 +93,8 @@ int git_diff_ui_config(const char *var, const char *value, void *cb)\n> >>  \t}\n> >>  \tif (!strcmp(var, \"diff.external\"))\n> >>  \t\treturn git_config_string(&external_diff_cmd_cfg, var, value);\n> >> +\tif (!strcmp(var, \"diff.color-words\"))\n> >\n> > I'd call it diff.wordregex, because that's what it is.\n> \n> If we want to add a new word-oriented option to diff that is not about \n> coloring the word differences, is it safe and sane to reuse the same \n> definition?  That is, \"git diff --color-words\" would be affected when \n> diff.wordregex is set to some value, so does any new word-oriented \n> operation we will add, and the single regex configured would be used as \n> the default value to define how a word would look like.\n> \n> I think it makes sense; I do not think of a case offhand where you would \n> want to define what a word is for the purpose of coloring diffs in one \n> way, and would want to use a different definition for another \n> word-oriented operation.\n\nWhy not cross that bridge when we're there?  Should we ever feel the need \nfor different word regexes, we would just introduce color.wordregex.\n\n> >> @@ -1550,6 +1553,8 @@ static void builtin_diff(const char *name_a,\n> >>  \t\t\t\to->word_regex = userdiff_word_regex(one);\n> >>  \t\t\tif (!o->word_regex)\n> >>  \t\t\t\to->word_regex = userdiff_word_regex(two);\n> >> +\t\t\tif (!o->word_regex)\n> >> +\t\t\t\to->word_regex = diff_color_words_cfg;\n> >\n> > IMHO this is the wrong order.  config should not override attributes, \n> > which are by definition more specific.\n> \n> Isn't it merely giving a fallback value when attributes does not give one?\n\nYep.  Boyd (or Stephen, as he wants to be called, making it hard to guess \nfrom his email address, but that's all part of the fun, in't it?) already \nrealized that I was up too late and got the order wrong myself.\n\n> By the way, wouldn't it make sense to optimize the precontext of that \n> hunk by doing _something_ like:\n> \n> \tif (!o->word_regex && strcmp(one->path, two->path))\n>         \to->word_regex = userdiff_word_regex(two);\n> \n> \"Something like\" comes from special cases like /dev/null for new/deleted\n> files, etc.\n\nYou mean to avoid the cost of initializing the regex in case one and the \nsame file is diffed against itself?  But that would be better handled \nbefore calling builtin_diff(), don't you think?\n\nI do not know off-hand if diffcore_std() handles that already, so that the \ndiff_flush() ... builtin_diff() cascade is not even called.\n\nBut you raise a valid concern: the regular expression is initialized every \ntime we look at a file.  We probably should have a member \nword_regex_compiled in diff_options, then, and only initialize it the \nfirst time.\n\nCiao,\nDscho \"who does not have the time to work on Git right now\"\n"},{"id":"101279","messageId":"200901201842.24000.markus.heidelberg@web.de","threadId":"17092","inReplyTo":"7v1vuympie.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Markus Heidelberg","fromEmail":"markus.heidelberg@web.de","sentAt":"2009-01-20T17:42:23Z","receivedAt":"2009-01-20T17:42:23Z","isPatch":true,"sender":{"key":"markus.heidelberg@web.de","avatar":"https://avatars.githubusercontent.com/u/6334512?v=4"},"body":"Junio C Hamano, 20.01.2009:\n> \"Boyd Stephen Smith Jr.\" <bss@iguanasuicide.net> writes:\n> \n> > When diff is invoked with --color-words (w/o =regex), use the regular\n> > expression the user has configured as diff.color-words.\n> >\n> > diff drivers configured via attributes take precedence over the\n> > diff.color-words setting.  If the user wants to change them, they have\n> > their own configuration variables.\n> \n> This needs an entry in Documentation/config.txt\n> \n> None of the existing configuration variables defined use hyphens in\n> multi-word variable names.\n\nExcept for diff.suppress-blank-empty\nShould it be converted or is it intention to reflect GNU diff's option?\n\nMarkus\n"},{"id":"101281","messageId":"200901201159.00803.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"200901201842.24000.markus.heidelberg@web.de","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-20T17:58:56Z","receivedAt":"2009-01-20T17:58:56Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"On Tuesday 2009 January 20 11:42:23 Markus Heidelberg wrote:\n>Junio C Hamano, 20.01.2009:\n>> \"Boyd Stephen Smith Jr.\" <bss@iguanasuicide.net> writes:\n>> > When diff is invoked with --color-words (w/o =regex), use the regular\n>> > expression the user has configured as diff.color-words.\n>> >\n>> > diff drivers configured via attributes take precedence over the\n>> > diff.color-words setting.  If the user wants to change them, they have\n>> > their own configuration variables.\n>>\n>> This needs an entry in Documentation/config.txt\n>>\n>> None of the existing configuration variables defined use hyphens in\n>> multi-word variable names.\n>\n>Except for diff.suppress-blank-empty\n>Should it be converted or is it intention to reflect GNU diff's option?\n\nI think best would be to have a project policy, use that for the wordRegex \noption and other options moving forward, then fix the others at some point in \nthe future (1.7?) while having some period of time where both old and \"per \npolicy\" names work.  But, then I'm a big fan of standardization.\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"101296","messageId":"7vk58pk9k5.fsf@gitster.siamese.dyndns.org","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901201819490.5159@intel-tinevez-2-302","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-20T20:27:06Z","receivedAt":"2009-01-20T20:27:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> By the way, wouldn't it make sense to optimize the precontext of that \n>> hunk by doing _something_ like:\n>> \n>> \tif (!o->word_regex && strcmp(one->path, two->path))\n>>         \to->word_regex = userdiff_word_regex(two);\n>> \n>> \"Something like\" comes from special cases like /dev/null for new/deleted\n>> files, etc.\n>\n> You mean to avoid the cost of initializing the regex in case one and the \n> same file is diffed against itself?\n\nNo.\n\nWhat I meant is much simpler than that.\n\nIf one and two are the same filename, and earlier gitattributes lookup for\nthe path already failed to produce any when you checked one, isn't it very\nlikely that the gitattributes lookup for two would fail the same way to\nproduce any result?\n"},{"id":"101300","messageId":"alpine.DEB.1.00.0901202202060.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"7vk58pk9k5.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-20T21:02:16Z","receivedAt":"2009-01-20T21:02:16Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 20 Jan 2009, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> >> By the way, wouldn't it make sense to optimize the precontext of that \n> >> hunk by doing _something_ like:\n> >> \n> >> \tif (!o->word_regex && strcmp(one->path, two->path))\n> >>         \to->word_regex = userdiff_word_regex(two);\n> >> \n> >> \"Something like\" comes from special cases like /dev/null for new/deleted\n> >> files, etc.\n> >\n> > You mean to avoid the cost of initializing the regex in case one and the \n> > same file is diffed against itself?\n> \n> No.\n> \n> What I meant is much simpler than that.\n> \n> If one and two are the same filename, and earlier gitattributes lookup for\n> the path already failed to produce any when you checked one, isn't it very\n> likely that the gitattributes lookup for two would fail the same way to\n> produce any result?\n\nOh, I see!\n\nThanks,\nDscho\n"},{"id":"101302","messageId":"alpine.DEB.1.00.0901202202370.3586@pacific.mpi-cbg.de","threadId":"17092","inReplyTo":"200901201842.24000.markus.heidelberg@web.de","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-20T21:08:33Z","receivedAt":"2009-01-20T21:08:33Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 20 Jan 2009, Markus Heidelberg wrote:\n\n> Junio C Hamano, 20.01.2009:\n> > \"Boyd Stephen Smith Jr.\" <bss@iguanasuicide.net> writes:\n> > \n> > > When diff is invoked with --color-words (w/o =regex), use the regular\n> > > expression the user has configured as diff.color-words.\n> > >\n> > > diff drivers configured via attributes take precedence over the\n> > > diff.color-words setting.  If the user wants to change them, they have\n> > > their own configuration variables.\n> > \n> > This needs an entry in Documentation/config.txt\n> > \n> > None of the existing configuration variables defined use hyphens in\n> > multi-word variable names.\n> \n> Except for diff.suppress-blank-empty\n> Should it be converted or is it intention to reflect GNU diff's option?\n\nGrumble.  It's in v1.6.1-rc1~348, so we cannot just go ahead and fix it.\n\nMy preference would be to convert it _except_ that the old name should \nstill work.  But it should not be advertized.\n\nCiao,\nDscho \"who loves consistency, and knows new users appreciate it, too\"\n\n-- snipsnap --\n[PATCH] Rename diff.suppress-blank-empty to diff.suppressBlankEmpty\n\nAll the other config variables use CamelCase.  This config variable should\nnot be an exception.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Documentation/config.txt       |    2 +-\n diff.c                         |    4 +++-\n t/t4029-diff-trailing-space.sh |    8 ++++----\n 3 files changed, 8 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex c92e7e6..4f0a0b1 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -652,7 +652,7 @@ diff.renames::\n \twill enable basic rename detection.  If set to \"copies\" or\n \t\"copy\", it will detect copies, as well.\n \n-diff.suppress-blank-empty::\n+diff.suppressBlankEmpty::\n \tA boolean to inhibit the standard behavior of printing a space\n \tbefore each empty output line. Defaults to false.\n \ndiff --git a/diff.c b/diff.c\nindex c6a992d..0100b59 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -118,7 +118,9 @@ int git_diff_basic_config(const char *var, const char *value, void *cb)\n \t}\n \n \t/* like GNU diff's --suppress-blank-empty option  */\n-\tif (!strcmp(var, \"diff.suppress-blank-empty\")) {\n+\tif (!strcmp(var, \"diff.suppressblankempty\") ||\n+\t\t\t/* for backwards compatibility */\n+\t\t\t!strcmp(var, \"diff.suppress-blank-empty\")) {\n \t\tdiff_suppress_blank_empty = git_config_bool(var, value);\n \t\treturn 0;\n \t}\ndiff --git a/t/t4029-diff-trailing-space.sh b/t/t4029-diff-trailing-space.sh\nindex 4ca65e0..9ddbbcd 100755\n--- a/t/t4029-diff-trailing-space.sh\n+++ b/t/t4029-diff-trailing-space.sh\n@@ -2,7 +2,7 @@\n #\n # Copyright (c) Jim Meyering\n #\n-test_description='diff honors config option, diff.suppress-blank-empty'\n+test_description='diff honors config option, diff.suppressBlankEmpty'\n \n . ./test-lib.sh\n \n@@ -24,14 +24,14 @@ test_expect_success \\\n      git add f &&\n      git commit -q -m. f &&\n      printf \"\\ny\\n\" > f &&\n-     git config --bool diff.suppress-blank-empty true &&\n+     git config --bool diff.suppressBlankEmpty true &&\n      git diff f > actual &&\n      test_cmp exp actual &&\n      perl -i.bak -p -e \"s/^\\$/ /\" exp &&\n-     git config --bool diff.suppress-blank-empty false &&\n+     git config --bool diff.suppressBlankEmpty false &&\n      git diff f > actual &&\n      test_cmp exp actual &&\n-     git config --bool --unset diff.suppress-blank-empty &&\n+     git config --bool --unset diff.suppressBlankEmpty &&\n      git diff f > actual &&\n      test_cmp exp actual\n      '\n-- \n1.6.1.439.g22f77c\n"},{"id":"101342","messageId":"200901202146.58651.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901201058520.3586@pacific.mpi-cbg.de","subject":"[PATCH] color-words: Support diff.color-words config option","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-21T03:46:57Z","receivedAt":"2009-01-21T03:46:57Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"When diff is invoked with --color-words (w/o =regex), use the regular\nexpression the user has configured as diff.wordregex.\n\ndiff drivers configured via attributes take precedence over the\ndiff.wordregex-words setting.  If the user wants to change them, they have\ntheir own configuration variables.\n\nSigned-off-by: Boyd Stephen Smith Jr <bss@iguanasuicide.net>\n---\nThis version is squashed into one patch and includes documentation and\nrewritten tests.  It was generated against js/diff-color-words~2,\n80c49c3d (color-words: make regex configurable via attributes), replacing\nmy previous 2 patches.  It uses \"diff.wordregex\" for reasons mention by\nDscho and because that was already what the diff drivers were using.\n\nI'm not entirely satisfied with it.  There should probably be some way\nto force the default behavior (which is a bit faster) even if a global\nconfig or diff driver exists.  Also, I think camelCase is better than\nruntogether so I'd prefer to change \"wordregex\" -> \"wordRegex\" across\nthe entire patch set.\n\n Documentation/config.txt       |    6 +++++\n Documentation/diff-options.txt |    7 +++--\n diff.c                         |    5 ++++\n t/t4034-diff-words.sh          |   45 ++++++++++++++++++++++++++++++++++++++-\n 4 files changed, 58 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex 7408bb2..0ca983a 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -639,6 +639,12 @@ diff.suppress-blank-empty::\n \tA boolean to inhibit the standard behavior of printing a space\n \tbefore each empty output line. Defaults to false.\n \n+diff.wordregex::\n+\tA POSIX Extended Regular Expression used to determine what is a \"word\"\n+\twhen performing word-by-word difference calculations.  Character\n+\tsequences that match the regular expression are \"words\", all other\n+\tcharacters are *ignorable* whitespace.\n+\n fetch.unpackLimit::\n \tIf the number of objects fetched over the git native\n \ttransfer is below this\ndiff --git a/Documentation/diff-options.txt b/Documentation/diff-options.txt\nindex 1edb82e..164e2c5 100644\n--- a/Documentation/diff-options.txt\n+++ b/Documentation/diff-options.txt\n@@ -103,9 +103,10 @@ expression to make sure that it matches all non-whitespace characters.\n A match that contains a newline is silently truncated(!) at the\n newline.\n +\n-The regex can also be set via a diff driver, see\n-linkgit:gitattributes[1]; giving it explicitly overrides any diff\n-driver setting.\n+The regex can also be set via a diff driver or configuration option, see\n+linkgit:gitattributes[1] or linkgit:git-config[1].  Giving it explicitly\n+overrides any diff driver or configuration setting.  Diff drivers\n+override configuration settings.\n \n --no-renames::\n \tTurn off rename detection, even when the configuration\ndiff --git a/diff.c b/diff.c\nindex 9fcde96..ed8b83c 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -23,6 +23,7 @@ static int diff_detect_rename_default;\n static int diff_rename_limit_default = 200;\n static int diff_suppress_blank_empty;\n int diff_use_color_default = -1;\n+static const char *diff_word_regex_cfg;\n static const char *external_diff_cmd_cfg;\n int diff_auto_refresh_index = 1;\n static int diff_mnemonic_prefix;\n@@ -92,6 +93,8 @@ int git_diff_ui_config(const char *var, const char *value, void *cb)\n \t}\n \tif (!strcmp(var, \"diff.external\"))\n \t\treturn git_config_string(&external_diff_cmd_cfg, var, value);\n+\tif (!strcmp(var, \"diff.wordregex\"))\n+\t\treturn git_config_string(&diff_word_regex_cfg, var, value);\n \n \treturn git_diff_basic_config(var, value, cb);\n }\n@@ -1550,6 +1553,8 @@ static void builtin_diff(const char *name_a,\n \t\t\t\to->word_regex = userdiff_word_regex(one);\n \t\t\tif (!o->word_regex)\n \t\t\t\to->word_regex = userdiff_word_regex(two);\n+\t\t\tif (!o->word_regex)\n+\t\t\t\to->word_regex = diff_word_regex_cfg;\n \t\t\tif (o->word_regex) {\n \t\t\t\tecbdata.diff_words->word_regex = (regex_t *)\n \t\t\t\t\txmalloc(sizeof(regex_t));\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 744221b..6bcc153 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -77,6 +77,7 @@ a = b + c<RESET>\n \n <GREEN>aeff = aeff * ( aaa<RESET> )\n EOF\n+cp expect expect.letter-runs-are-words\n \n test_expect_success 'word diff with a regular expression' '\n \n@@ -92,7 +93,7 @@ post diff=testdriver\n EOF\n '\n \n-test_expect_success 'option overrides default' '\n+test_expect_success 'option overrides .gitattributes' '\n \n \tword_diff --color-words=\"[a-z]+\"\n \n@@ -112,13 +113,53 @@ a = b + c<RESET>\n \n <GREEN>aeff = aeff * ( aaa )<RESET>\n EOF\n+cp expect expect.non-whitespace-is-word\n \n-test_expect_success 'use default supplied by driver' '\n+test_expect_success 'use regex supplied by driver' '\n \n \tword_diff --color-words\n \n '\n \n+test_expect_success 'set diff.wordregex option' '\n+\tgit config diff.wordregex \"[[:alnum:]]+\"\n+'\n+\n+cp expect.letter-runs-are-words expect\n+\n+test_expect_success 'command-line overrides config' '\n+\tword_diff --color-words=\"[a-z]+\"\n+'\n+\n+cp expect.non-whitespace-is-word expect\n+\n+test_expect_success '.gitattributes override config' '\n+\tword_diff --color-words\n+'\n+\n+test_expect_success 'remove diff driver regex' '\n+\tgit config --unset diff.testdriver.wordregex\n+'\n+\n+cat > expect <<\\EOF\n+<WHITE>diff --git a/pre b/post<RESET>\n+<WHITE>index 330b04f..5ed8eff 100644<RESET>\n+<WHITE>--- a/pre<RESET>\n+<WHITE>+++ b/post<RESET>\n+<BROWN>@@ -1,3 +1,7 @@<RESET>\n+h(4),<GREEN>hh[44<RESET>]\n+<RESET>\n+a = b + c<RESET>\n+\n+<GREEN>aa = a<RESET>\n+\n+<GREEN>aeff = aeff * ( aaa<RESET> )\n+EOF\n+\n+test_expect_success 'use configured regex' '\n+\tword_diff --color-words\n+'\n+\n echo 'aaa (aaa)' > pre\n echo 'aaa (aaa) aaa' > post\n \n-- \n1.5.6.5\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"101344","messageId":"200901202259.54886.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"200901202146.58651.bss@iguanasuicide.net","subject":"[PATCH] Change the spelling of \"wordregex\".","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-21T04:59:54Z","receivedAt":"2009-01-21T04:59:54Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"Use \"wordRegex\" for configuration variable names.  Use \"word_regex\" for C\nlanguage tokens.\n\nSigned-off-by: Boyd Stephen Smith Jr. <bss@iguanasuicide.net>\n---\nOn Tuesday 20 January 2009, \"Boyd Stephen Smith Jr.\" <bss@iguanasuicide.net> wrote about '[PATCH] \ncolor-words: Support diff.color-words config option':\n>I'm not entirely satisfied with it. [...] I think camelCase is better than\n>runtogether so I'd prefer to change \"wordregex\" -> \"wordRegex\" across\n>the entire patch set.\n\nHere's a patch that does something like that, that can be squashed into the\nprevious patch.\n\n Documentation/config.txt        |    2 +-\n Documentation/gitattributes.txt |    4 ++--\n t/t4034-diff-words.sh           |    8 ++++----\n userdiff.c                      |    4 ++--\n 4 files changed, 9 insertions(+), 9 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex 0ca983a..332213e 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -639,7 +639,7 @@ diff.suppress-blank-empty::\n \tA boolean to inhibit the standard behavior of printing a space\n \tbefore each empty output line. Defaults to false.\n \n-diff.wordregex::\n+diff.wordRegex::\n \tA POSIX Extended Regular Expression used to determine what is a \"word\"\n \twhen performing word-by-word difference calculations.  Character\n \tsequences that match the regular expression are \"words\", all other\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex ba3ba12..227934f 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -341,14 +341,14 @@ Customizing word diff\n \n You can customize the rules that `git diff --color-words` uses to\n split words in a line, by specifying an appropriate regular expression\n-in the \"diff.*.wordregex\" configuration variable.  For example, in TeX\n+in the \"diff.*.wordRegex\" configuration variable.  For example, in TeX\n a backslash followed by a sequence of letters forms a command, but\n several such commands can be run together without intervening\n whitespace.  To separate them, use a regular expression such as\n \n ------------------------\n [diff \"tex\"]\n-\twordregex = \"\\\\\\\\[a-zA-Z]+|[{}]|\\\\\\\\.|[^\\\\{}[:space:]]+\"\n+\twordRegex = \"\\\\\\\\[a-zA-Z]+|[{}]|\\\\\\\\.|[^\\\\{}[:space:]]+\"\n ------------------------\n \n A built-in pattern is provided for all languages listed in the\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 6bcc153..4508eff 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -86,7 +86,7 @@ test_expect_success 'word diff with a regular expression' '\n '\n \n test_expect_success 'set a diff driver' '\n-\tgit config diff.testdriver.wordregex \"[^[:space:]]\" &&\n+\tgit config diff.testdriver.wordRegex \"[^[:space:]]\" &&\n \tcat <<EOF > .gitattributes\n pre diff=testdriver\n post diff=testdriver\n@@ -121,8 +121,8 @@ test_expect_success 'use regex supplied by driver' '\n \n '\n \n-test_expect_success 'set diff.wordregex option' '\n-\tgit config diff.wordregex \"[[:alnum:]]+\"\n+test_expect_success 'set diff.wordRegex option' '\n+\tgit config diff.wordRegex \"[[:alnum:]]+\"\n '\n \n cp expect.letter-runs-are-words expect\n@@ -138,7 +138,7 @@ test_expect_success '.gitattributes override config' '\n '\n \n test_expect_success 'remove diff driver regex' '\n-\tgit config --unset diff.testdriver.wordregex\n+\tgit config --unset diff.testdriver.wordRegex\n '\n \n cat > expect <<\\EOF\ndiff --git a/userdiff.c b/userdiff.c\nindex 2b55509..d556da9 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -6,8 +6,8 @@ static struct userdiff_driver *drivers;\n static int ndrivers;\n static int drivers_alloc;\n \n-#define PATTERNS(name, pattern, wordregex)\t\t\t\\\n-\t{ name, NULL, -1, { pattern, REG_EXTENDED }, wordregex }\n+#define PATTERNS(name, pattern, word_regex)\t\t\t\\\n+\t{ name, NULL, -1, { pattern, REG_EXTENDED }, word_regex }\n static struct userdiff_driver builtin_drivers[] = {\n PATTERNS(\"html\", \"^[ \\t]*(<[Hh][1-6][ \\t].*>.*)$\",\n \t \"[^<>= \\t]+|[^[:space:]]|[\\x80-\\xff]+\"),\n-- \n1.5.6.5\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"101353","messageId":"alpine.DEB.1.00.0901210923580.7929@racer","threadId":"17092","inReplyTo":"200901202146.58651.bss@iguanasuicide.net","subject":"Re: [PATCH] color-words: Support diff.color-words config option","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-21T08:25:08Z","receivedAt":"2009-01-21T08:25:08Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 20 Jan 2009, Boyd Stephen Smith Jr. wrote:\n\n> It uses \"diff.wordregex\" for reasons mention by Dscho and because that \n> was already what the diff drivers were using.\n\nTo be fair, Jakub and Junio mentioned it, too.\n\n> I'm not entirely satisfied with it.  There should probably be some way \n> to force the default behavior (which is a bit faster) even if a global \n> config or diff driver exists.  Also, I think camelCase is better than \n> runtogether so I'd prefer to change \"wordregex\" -> \"wordRegex\" across \n> the entire patch set.\n\nWell, the thing is, it _should_ be \"wordRegex\", _except_ in the strcmp() \nbecause the config helpers get a downcased key.\n\nCiao,\nDscho\n"},{"id":"101354","messageId":"alpine.DEB.1.00.0901210925430.7929@racer","threadId":"17092","inReplyTo":"200901202259.54886.bss@iguanasuicide.net","subject":"Re: [PATCH] Change the spelling of \"wordregex\".","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-21T08:26:57Z","receivedAt":"2009-01-21T08:26:57Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 20 Jan 2009, Boyd Stephen Smith Jr. wrote:\n\n> diff --git a/userdiff.c b/userdiff.c\n> index 2b55509..d556da9 100644\n> --- a/userdiff.c\n> +++ b/userdiff.c\n> @@ -6,8 +6,8 @@ static struct userdiff_driver *drivers;\n>  static int ndrivers;\n>  static int drivers_alloc;\n>  \n> -#define PATTERNS(name, pattern, wordregex)\t\t\t\\\n> -\t{ name, NULL, -1, { pattern, REG_EXTENDED }, wordregex }\n> +#define PATTERNS(name, pattern, word_regex)\t\t\t\\\n> +\t{ name, NULL, -1, { pattern, REG_EXTENDED }, word_regex }\n>  static struct userdiff_driver builtin_drivers[] = {\n>  PATTERNS(\"html\", \"^[ \\t]*(<[Hh][1-6][ \\t].*>.*)$\",\n>  \t \"[^<>= \\t]+|[^[:space:]]|[\\x80-\\xff]+\"),\n\nIn general, it is an awesomly good idea to imitate code that is already \nthere.  That literally guarantees consistency (which is Good, as you \nknow).\n\nAnd Thomas just imitated \"xfuncname\", which just so happens to be without \nan \"_\".\n\nCiao,\nDscho\n"},{"id":"101361","messageId":"200901211022.24995.trast@student.ethz.ch","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901210925430.7929@racer","subject":"Re: [PATCH] Change the spelling of \"wordregex\".","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2009-01-21T09:22:22Z","receivedAt":"2009-01-21T09:22:22Z","isPatch":true,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Johannes Schindelin wrote:\n> And Thomas just imitated \"xfuncname\", which just so happens to be without \n> an \"_\".\n\nThen again I ignored the 'x' for \"extended regex\", so it's not\nentirely consistent.\n\n[Mostly because I think the user expects a \"<something>\" whenever\nthere's an \"x<something>\", and \"funcname\" is actually deprecated/not\ndocumented any more, so introducing a basic-regex version seemed\nsilly.]\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"101368","messageId":"7v1vuxeyy9.fsf@gitster.siamese.dyndns.org","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901202202370.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-21T10:27:26Z","receivedAt":"2009-01-21T10:27:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Subject: Rename diff.suppress-blank-empty to diff.suppressBlankEmpty\n>\n> All the other config variables use CamelCase.  This config variable should\n> not be an exception.\n>\n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n\nThanks.\n"},{"id":"101369","messageId":"7vvds9dkdi.fsf@gitster.siamese.dyndns.org","threadId":"17092","inReplyTo":"200901202146.58651.bss@iguanasuicide.net","subject":"Re: [PATCH] color-words: Support diff.wordregex config option","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-21T10:27:37Z","receivedAt":"2009-01-21T10:27:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Boyd Stephen Smith Jr.\" <bss@iguanasuicide.net> writes:\n\n> When diff is invoked with --color-words (w/o =regex), use the regular\n> expression the user has configured as diff.wordregex.\n>\n> diff drivers configured via attributes take precedence over the\n> diff.wordregex-words setting.  If the user wants to change them, they have\n> their own configuration variables.\n>\n> Signed-off-by: Boyd Stephen Smith Jr <bss@iguanasuicide.net>\n> ---\n> This version is squashed into one patch and includes documentation and\n> rewritten tests.  It was generated against js/diff-color-words~2,\n> 80c49c3d (color-words: make regex configurable via attributes), replacing\n> my previous 2 patches.  It uses \"diff.wordregex\" for reasons mention by\n> Dscho and because that was already what the diff drivers were using.\n\nNicely done and very well described.  I fixed the Subject: line, though ;-)\n\nThanks.\n"},{"id":"101393","messageId":"200901210934.03196.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901210925430.7929@racer","subject":"Re: [PATCH] Change the spelling of \"wordregex\".","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-21T15:33:57Z","receivedAt":"2009-01-21T15:33:57Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"On Wednesday 21 January 2009, Johannes Schindelin \n<Johannes.Schindelin@gmx.de> wrote about 'Re: [PATCH] Change the spelling \nof \"wordregex\".':\n>On Tue, 20 Jan 2009, Boyd Stephen Smith Jr. wrote:\n>> diff --git a/userdiff.c b/userdiff.c\n>> index 2b55509..d556da9 100644\n>> --- a/userdiff.c\n>> +++ b/userdiff.c\n>> @@ -6,8 +6,8 @@ static struct userdiff_driver *drivers;\n>>  static int ndrivers;\n>>  static int drivers_alloc;\n>>\n>> -#define PATTERNS(name, pattern, wordregex)\t\t\t\\\n>> -\t{ name, NULL, -1, { pattern, REG_EXTENDED }, wordregex }\n>> +#define PATTERNS(name, pattern, word_regex)\t\t\t\\\n>> +\t{ name, NULL, -1, { pattern, REG_EXTENDED }, word_regex }\n>>  static struct userdiff_driver builtin_drivers[] = {\n>>  PATTERNS(\"html\", \"^[ \\t]*(<[Hh][1-6][ \\t].*>.*)$\",\n>>  \t \"[^<>= \\t]+|[^[:space:]]|[\\x80-\\xff]+\"),\n>\n>In general, it is an awesomly good idea to imitate code that is already\n>there.  That literally guarantees consistency (which is Good, as you\n>know).\n\nAgreed that consistency is good.  However, using \"wordregex\" isn't \nconsistent.  The rest of the time it is used as an identifier in the code, \nit's spelled \"word_regex\" or \"word_regexp\", even before my patch.  \n(Declarations in: userdiff.h, builtin-grep.c, 3x diff.c, and grep.h)\n\nIn particular, the macro is used to initialize \"struct userdiff_driver\"s \nand the relevant member of that struct uses \"word_regex\" before my patch.\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"101399","messageId":"200901211009.36081.bss@iguanasuicide.net","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901210923580.7929@racer","subject":"Re: [PATCH] color-words: Support diff.color-words config option","fromName":"Boyd Stephen Smith Jr.","fromEmail":"bss@iguanasuicide.net","sentAt":"2009-01-21T16:09:35Z","receivedAt":"2009-01-21T16:09:35Z","isPatch":true,"sender":{"key":"bss@iguanasuicide.net","avatar":"https://gravatar.com/avatar/84b95eeff194b816c1568b1339e63e4b229825298664a9037b9f1ec713ead1e3?d=mp&s=160"},"body":"On Wednesday 21 January 2009, Johannes Schindelin \n<Johannes.Schindelin@gmx.de> wrote about 'Re: [PATCH] color-words: Support \ndiff.color-words config option':\n>On Tue, 20 Jan 2009, Boyd Stephen Smith Jr. wrote:\n>> I'm not entirely satisfied with it.  There should probably be some way\n>> to force the default behavior (which is a bit faster) even if a global\n>> config or diff driver exists.  Also, I think camelCase is better than\n>> runtogether so I'd prefer to change \"wordregex\" -> \"wordRegex\" across\n>> the entire patch set.\n>\n>Well, the thing is, it _should_ be \"wordRegex\", _except_ in the strcmp()\n>because the config helpers get a downcased key.\n\nIt would have been nice to know that last night.  I spent far longer than I \nshould have on the \"wordregex\" -> \"wordRegex\" patch.\n-- \nBoyd Stephen Smith Jr.                     ,= ,-_-. =. \nbss@iguanasuicide.net                     ((_/)o o(\\_))\nICQ: 514984 YM/AIM: DaTwinkDaddy           `-'(. .)`-' \nhttp://iguanasuicide.net/                      \\_/     \n"},{"id":"101430","messageId":"200901212037.07014.markus.heidelberg@web.de","threadId":"17092","inReplyTo":"alpine.DEB.1.00.0901202202370.3586@pacific.mpi-cbg.de","subject":"Re: [PATCH] diff: Support diff.color-words config option","fromName":"Markus Heidelberg","fromEmail":"markus.heidelberg@web.de","sentAt":"2009-01-21T19:37:06Z","receivedAt":"2009-01-21T19:37:06Z","isPatch":true,"sender":{"key":"markus.heidelberg@web.de","avatar":"https://avatars.githubusercontent.com/u/6334512?v=4"},"body":"Johannes Schindelin, 20.01.2009:\n> Hi,\n> \n> On Tue, 20 Jan 2009, Markus Heidelberg wrote:\n> \n> > Junio C Hamano, 20.01.2009:\n> > > None of the existing configuration variables defined use hyphens in\n> > > multi-word variable names.\n> > \n> > Except for diff.suppress-blank-empty\n> > Should it be converted or is it intention to reflect GNU diff's option?\n> \n> Grumble.  It's in v1.6.1-rc1~348, so we cannot just go ahead and fix it.\n\nDid I say change it without keeping backward compatibility?\n\n> Ciao,\n> Dscho \"who loves consistency, and knows new users appreciate it, too\"\n\nMe, too.\n\nMarkus\n"}]}