{"thread":{"id":"35947","subject":"[PATCH v2 00/19] Multiparent diff tree-walker + combine-diff speedup","startedAt":"2014-02-24T16:21:32Z","lastAt":"2014-04-10T17:30:32Z","messageCount":64,"participants":["Kirill Smelkov","Duy Nguyen","Thomas Schwinge","Erik Faye-Lund","Junio C Hamano","Johannes Sixt"],"isPatch":true,"patchVersion":2,"patchTotal":19},"messages":[{"id":"235241","messageId":"cover.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":null,"subject":"[PATCH v2 00/19] Multiparent diff tree-walker + combine-diff speedup","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:32Z","receivedAt":"2014-02-24T16:21:32Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Hello up there.\n\nHere go combine-diff speedup patches in form of first reworking diff\ntree-walker to work in general case - when a commit have several parents, not\nonly one - we are traversing all 1+nparent trees in parallel.\n\nThen we are taking advantage of the new diff tree-walker for speeding up\ncombine-diff, which for linux.git results in ~14 times speedup.\n\nThis is the second posting for the whole series - sent here patches should go\ninstead of already-in-pu ks/diff-tree-more and ks/tree-diff-nway into\nks/tree-diff-nway - patches are related and seeing them all at once is more\nlogical to me.\n\nI've tried to do my homework based on review feedback and the changes compared\nto v1 are:\n\n- fixed last-minute thinko/bug last time introduced on my side (sorry) with\n  opt->pathchange manipulation in __diff_tree_sha1() - we were forgetting to\n  restore opt->pathchange, which led to incorrect log -c (merges _and_ plain\n  diff-tree) output;\n\n  This time, I've verified several times, log output stays really the same.\n\n- direct use of alloca() changed to portability wrappers xalloca/xalloca_free\n  which gracefully degrade to xmalloc/free on systems, where alloca is not\n  available (see new patch 17).\n\n- \"i = 0; do { ... } while (++i < nparent)\" is back to usual looping\n  \"for (i = 0; i < nparent; ++)\", as I've re-measured timings and the\n  difference is negligible.\n\n  ( Initially, when I was fighting for every cycle it made sense, but real\n    no-slowdown turned out to be related to avoiding mallocs, load trees in correct\n    order and reducing register pressure. )\n\n- S_IFXMIN_NEQ definition moved out to cache.h, to have all modes registry in one place;\n\n\n- diff_tree() becomes static (new patch 13), as nobody is using it outside\n  tree-diff.c (and is later renamed to __diff_tree_sha1);\n\n- p0 -> first_parent; corrected comments about how emit_diff_first_parent_only\n  behaves;\n\n\nnot changed:\n\n- low-level helpers are still named with \"__\" prefix as, imho, that is the best\n  convention to name such helpers, without sacrificing signal/noise ratio. All\n  of them are now static though.\n\n\nSignoffs were left intact, if a patch was already applied to pu with one, and\nhad not changed.\n\nPlease apply and thanks,\nKirill\n\nP.S. Sorry for the delay - I was very busy.\n\n\nKirill Smelkov (19):\n  combine-diff: move show_log_first logic/action out of paths scanning\n  combine-diff: move changed-paths scanning logic into its own function\n  tree-diff: no need to manually verify that there is no mode change for a path\n  tree-diff: no need to pass match to skip_uninteresting()\n  tree-diff: show_tree() is not needed\n  tree-diff: consolidate code for emitting diffs and recursion in one place\n  tree-diff: don't assume compare_tree_entry() returns -1,0,1\n  tree-diff: move all action-taking code out of compare_tree_entry()\n  tree-diff: rename compare_tree_entry -> tree_entry_pathcmp\n  tree-diff: show_path prototype is not needed anymore\n  tree-diff: simplify tree_entry_pathcmp\n  tree-diff: remove special-case diff-emitting code for empty-tree cases\n  tree-diff: diff_tree() should now be static\n  tree-diff: rework diff_tree interface to be sha1 based\n  tree-diff: no need to call \"full\" diff_tree_sha1 from show_path()\n  tree-diff: reuse base str(buf) memory on sub-tree recursion\n  Portable alloca for Git\n  tree-diff: rework diff_tree() to generate diffs for multiparent cases as well\n  combine-diff: speed it up, by using multiparent diff tree-walker directly\n\n Makefile          |   6 +\n cache.h           |  15 ++\n combine-diff.c    | 170 +++++++++++---\n config.mak.uname  |  10 +-\n configure.ac      |   8 +\n diff.c            |   2 +\n diff.h            |  12 +-\n git-compat-util.h |   8 +\n tree-diff.c       | 666 +++++++++++++++++++++++++++++++++++++++++++-----------\n 9 files changed, 724 insertions(+), 173 deletions(-)\n\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235242","messageId":"bc222e334ac9f12ee946b0ddddc35b13eabb1232.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 01/19] combine-diff: move show_log_first logic/action out of paths scanning","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:33Z","receivedAt":"2014-02-24T16:21:33Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Judging from sample outputs and tests nothing changes in diff -c output,\nand this change will help later patches, when we'll be refactoring paths\nscanning into its own function with several variants - the\nshow_log_first logic / code will stay common to all of them.\n\nNOTE: only now we have to take care to explicitly not show anything if\n    parents array is empty, as in fact there are some clients in Git code,\n    which calls diff_tree_combined() in such a way.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n combine-diff.c | 24 ++++++++++++++----------\n 1 file changed, 14 insertions(+), 10 deletions(-)\n\ndiff --git a/combine-diff.c b/combine-diff.c\nindex 24ca7e2..68d2e53 100644\n--- a/combine-diff.c\n+++ b/combine-diff.c\n@@ -1311,6 +1311,20 @@ void diff_tree_combined(const unsigned char *sha1,\n \tstruct combine_diff_path *p, *paths = NULL;\n \tint i, num_paths, needsep, show_log_first, num_parent = parents->nr;\n \n+\t/* nothing to do, if no parents */\n+\tif (!num_parent)\n+\t\treturn;\n+\n+\tshow_log_first = !!rev->loginfo && !rev->no_commit_id;\n+\tneedsep = 0;\n+\tif (show_log_first) {\n+\t\tshow_log(rev);\n+\n+\t\tif (rev->verbose_header && opt->output_format)\n+\t\t\tprintf(\"%s%c\", diff_line_prefix(opt),\n+\t\t\t       opt->line_termination);\n+\t}\n+\n \tdiffopts = *opt;\n \tcopy_pathspec(&diffopts.pathspec, &opt->pathspec);\n \tdiffopts.output_format = DIFF_FORMAT_NO_OUTPUT;\n@@ -1319,8 +1333,6 @@ void diff_tree_combined(const unsigned char *sha1,\n \t/* tell diff_tree to emit paths in sorted (=tree) order */\n \tdiffopts.orderfile = NULL;\n \n-\tshow_log_first = !!rev->loginfo && !rev->no_commit_id;\n-\tneedsep = 0;\n \t/* find set of paths that everybody touches */\n \tfor (i = 0; i < num_parent; i++) {\n \t\t/* show stat against the first parent even\n@@ -1336,14 +1348,6 @@ void diff_tree_combined(const unsigned char *sha1,\n \t\tdiffcore_std(&diffopts);\n \t\tpaths = intersect_paths(paths, i, num_parent);\n \n-\t\tif (show_log_first && i == 0) {\n-\t\t\tshow_log(rev);\n-\n-\t\t\tif (rev->verbose_header && opt->output_format)\n-\t\t\t\tprintf(\"%s%c\", diff_line_prefix(opt),\n-\t\t\t\t       opt->line_termination);\n-\t\t}\n-\n \t\t/* if showing diff, show it in requested order */\n \t\tif (diffopts.output_format != DIFF_FORMAT_NO_OUTPUT &&\n \t\t    opt->orderfile) {\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235243","messageId":"a1d2bf86a5aacd7be40ddfafaba438ec6e6af41e.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 02/19] combine-diff: move changed-paths scanning logic into its own function","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:34Z","receivedAt":"2014-02-24T16:21:34Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Move code for finding paths for which diff(commit,parent_i) is not-empty\nfor all parents to separate function - at present we have generic (and\nslow) code for this job, which translates 1 n-parent problem to n\n1-parent problems and then intersect results, and will be adding another\nlimited, but faster, paths scanning implementation in the next patch.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n combine-diff.c | 80 ++++++++++++++++++++++++++++++++++++++--------------------\n 1 file changed, 53 insertions(+), 27 deletions(-)\n\ndiff --git a/combine-diff.c b/combine-diff.c\nindex 68d2e53..1732dfd 100644\n--- a/combine-diff.c\n+++ b/combine-diff.c\n@@ -1301,6 +1301,51 @@ static const char *path_path(void *obj)\n \treturn path->path;\n }\n \n+\n+/* find set of paths that every parent touches */\n+static struct combine_diff_path *find_paths(const unsigned char *sha1,\n+\tconst struct sha1_array *parents, struct diff_options *opt)\n+{\n+\tstruct combine_diff_path *paths = NULL;\n+\tint i, num_parent = parents->nr;\n+\n+\tint output_format = opt->output_format;\n+\tconst char *orderfile = opt->orderfile;\n+\n+\topt->output_format = DIFF_FORMAT_NO_OUTPUT;\n+\t/* tell diff_tree to emit paths in sorted (=tree) order */\n+\topt->orderfile = NULL;\n+\n+\tfor (i = 0; i < num_parent; i++) {\n+\t\t/*\n+\t\t * show stat against the first parent even when doing\n+\t\t * combined diff.\n+\t\t */\n+\t\tint stat_opt = (output_format &\n+\t\t\t\t(DIFF_FORMAT_NUMSTAT|DIFF_FORMAT_DIFFSTAT));\n+\t\tif (i == 0 && stat_opt)\n+\t\t\topt->output_format = stat_opt;\n+\t\telse\n+\t\t\topt->output_format = DIFF_FORMAT_NO_OUTPUT;\n+\t\tdiff_tree_sha1(parents->sha1[i], sha1, \"\", opt);\n+\t\tdiffcore_std(opt);\n+\t\tpaths = intersect_paths(paths, i, num_parent);\n+\n+\t\t/* if showing diff, show it in requested order */\n+\t\tif (opt->output_format != DIFF_FORMAT_NO_OUTPUT &&\n+\t\t    orderfile) {\n+\t\t\tdiffcore_order(orderfile);\n+\t\t}\n+\n+\t\tdiff_flush(opt);\n+\t}\n+\n+\topt->output_format = output_format;\n+\topt->orderfile = orderfile;\n+\treturn paths;\n+}\n+\n+\n void diff_tree_combined(const unsigned char *sha1,\n \t\t\tconst struct sha1_array *parents,\n \t\t\tint dense,\n@@ -1308,7 +1353,7 @@ void diff_tree_combined(const unsigned char *sha1,\n {\n \tstruct diff_options *opt = &rev->diffopt;\n \tstruct diff_options diffopts;\n-\tstruct combine_diff_path *p, *paths = NULL;\n+\tstruct combine_diff_path *p, *paths;\n \tint i, num_paths, needsep, show_log_first, num_parent = parents->nr;\n \n \t/* nothing to do, if no parents */\n@@ -1327,35 +1372,16 @@ void diff_tree_combined(const unsigned char *sha1,\n \n \tdiffopts = *opt;\n \tcopy_pathspec(&diffopts.pathspec, &opt->pathspec);\n-\tdiffopts.output_format = DIFF_FORMAT_NO_OUTPUT;\n \tDIFF_OPT_SET(&diffopts, RECURSIVE);\n \tDIFF_OPT_CLR(&diffopts, ALLOW_EXTERNAL);\n-\t/* tell diff_tree to emit paths in sorted (=tree) order */\n-\tdiffopts.orderfile = NULL;\n \n-\t/* find set of paths that everybody touches */\n-\tfor (i = 0; i < num_parent; i++) {\n-\t\t/* show stat against the first parent even\n-\t\t * when doing combined diff.\n-\t\t */\n-\t\tint stat_opt = (opt->output_format &\n-\t\t\t\t(DIFF_FORMAT_NUMSTAT|DIFF_FORMAT_DIFFSTAT));\n-\t\tif (i == 0 && stat_opt)\n-\t\t\tdiffopts.output_format = stat_opt;\n-\t\telse\n-\t\t\tdiffopts.output_format = DIFF_FORMAT_NO_OUTPUT;\n-\t\tdiff_tree_sha1(parents->sha1[i], sha1, \"\", &diffopts);\n-\t\tdiffcore_std(&diffopts);\n-\t\tpaths = intersect_paths(paths, i, num_parent);\n-\n-\t\t/* if showing diff, show it in requested order */\n-\t\tif (diffopts.output_format != DIFF_FORMAT_NO_OUTPUT &&\n-\t\t    opt->orderfile) {\n-\t\t\tdiffcore_order(opt->orderfile);\n-\t\t}\n-\n-\t\tdiff_flush(&diffopts);\n-\t}\n+\t/* find set of paths that everybody touches\n+\t *\n+\t * NOTE find_paths() also handles --stat, as it computes\n+\t * diff(sha1,parent_i) for all i to do the job, specifically\n+\t * for parent0.\n+\t */\n+\tpaths = find_paths(sha1, parents, &diffopts);\n \n \t/* find out number of surviving paths */\n \tfor (num_paths = 0, p = paths; p; p = p->next)\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235245","messageId":"22aebb863fb2a5a556e68d57f3a1095d3c502d4e.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 03/19] tree-diff: no need to manually verify that there is no mode change for a path","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:35Z","receivedAt":"2014-02-24T16:21:35Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Because if there is, such two tree entries would never be compared as\nequal - the code in base_name_compare() explicitly compares modes, if\nthere is a change for dir bit, even for equal paths, entries would\ncompare as different.\n\nThe code I'm removing here is from 2005 April 262e82b4 (Fix diff-tree\nrecursion), which pre-dates base_name_compare() introduction in 958ba6c9\n(Introduce \"base_name_compare()\" helper function) by a month.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n tree-diff.c | 15 +++++----------\n 1 file changed, 5 insertions(+), 10 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 11c3550..5810b00 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -23,6 +23,11 @@ static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2,\n \n \tpathlen1 = tree_entry_len(&t1->entry);\n \tpathlen2 = tree_entry_len(&t2->entry);\n+\n+\t/*\n+\t * NOTE files and directories *always* compare differently,\n+\t * even when having the same name.\n+\t */\n \tcmp = base_name_compare(path1, pathlen1, mode1, path2, pathlen2, mode2);\n \tif (cmp < 0) {\n \t\tshow_entry(opt, \"-\", t1, base);\n@@ -35,16 +40,6 @@ static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2,\n \tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER) && !hashcmp(sha1, sha2) && mode1 == mode2)\n \t\treturn 0;\n \n-\t/*\n-\t * If the filemode has changed to/from a directory from/to a regular\n-\t * file, we need to consider it a remove and an add.\n-\t */\n-\tif (S_ISDIR(mode1) != S_ISDIR(mode2)) {\n-\t\tshow_entry(opt, \"-\", t1, base);\n-\t\tshow_entry(opt, \"+\", t2, base);\n-\t\treturn 0;\n-\t}\n-\n \tstrbuf_add(base, path1, pathlen1);\n \tif (DIFF_OPT_TST(opt, RECURSIVE) && S_ISDIR(mode1)) {\n \t\tif (DIFF_OPT_TST(opt, TREE_IN_RECURSIVE)) {\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235244","messageId":"540b87fe7a353e4f2ac798270991038f3bb89c35.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 04/19] tree-diff: no need to pass match to skip_uninteresting()","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:36Z","receivedAt":"2014-02-24T16:21:36Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"It is neither used there as input, nor the output written through it, is\nused outside.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n tree-diff.c | 17 ++++++++---------\n 1 file changed, 8 insertions(+), 9 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 5810b00..a8c2aec 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -109,13 +109,14 @@ static void show_entry(struct diff_options *opt, const char *prefix,\n }\n \n static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n-\t\t\t       struct diff_options *opt,\n-\t\t\t       enum interesting *match)\n+\t\t\t       struct diff_options *opt)\n {\n+\tenum interesting match;\n+\n \twhile (t->size) {\n-\t\t*match = tree_entry_interesting(&t->entry, base, 0, &opt->pathspec);\n-\t\tif (*match) {\n-\t\t\tif (*match == all_entries_not_interesting)\n+\t\tmatch = tree_entry_interesting(&t->entry, base, 0, &opt->pathspec);\n+\t\tif (match) {\n+\t\t\tif (match == all_entries_not_interesting)\n \t\t\t\tt->size = 0;\n \t\t\tbreak;\n \t\t}\n@@ -128,8 +129,6 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n {\n \tstruct strbuf base;\n \tint baselen = strlen(base_str);\n-\tenum interesting t1_match = entry_not_interesting;\n-\tenum interesting t2_match = entry_not_interesting;\n \n \t/* Enable recursion indefinitely */\n \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n@@ -141,8 +140,8 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \t\tif (opt->pathspec.nr) {\n-\t\t\tskip_uninteresting(t1, &base, opt, &t1_match);\n-\t\t\tskip_uninteresting(t2, &base, opt, &t2_match);\n+\t\t\tskip_uninteresting(t1, &base, opt);\n+\t\t\tskip_uninteresting(t2, &base, opt);\n \t\t}\n \t\tif (!t1->size) {\n \t\t\tif (!t2->size)\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235246","messageId":"4555386618c18b40ee9e06ebaae23e2eb5eb0d1e.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 05/19] tree-diff: show_tree() is not needed","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:37Z","receivedAt":"2014-02-24T16:21:37Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"We don't need special code for showing added/removed subtree, because we\ncan do the same via diff_tree_sha1, just passing NULL for absent tree.\n\nAnd compared to show_tree(), which was calling show_entry() for every\ntree entry, that would lead to the same show_entry() callings:\n\n    show_tree(t):\n        for e in t.entries:\n            show_entry(e)\n\n    diff_tree_sha1(NULL, new):  /* the same applies to (old, NULL) */\n        diff_tree(t1=NULL, t2)\n            ...\n            if (!t1->size)\n                show_entry(t2)\n            ...\n\nand possible overhead is negligible, since after the patch, timing for\n\n    `git log --raw --no-abbrev --no-renames`\n\nfor navy.git and `linux.git v3.10..v3.11` is practically the same.\n\nSo let's say goodbye to show_tree() - it removes some code, but also,\nand what is important, consolidates more code for showing/recursing into\ntrees into one place.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\n( re-posting without change )\n\n tree-diff.c | 35 +++--------------------------------\n 1 file changed, 3 insertions(+), 32 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex a8c2aec..2ad7788 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -55,25 +55,7 @@ static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2,\n \treturn 0;\n }\n \n-/* A whole sub-tree went away or appeared */\n-static void show_tree(struct diff_options *opt, const char *prefix,\n-\t\t      struct tree_desc *desc, struct strbuf *base)\n-{\n-\tenum interesting match = entry_not_interesting;\n-\tfor (; desc->size; update_tree_entry(desc)) {\n-\t\tif (match != all_entries_interesting) {\n-\t\t\tmatch = tree_entry_interesting(&desc->entry, base, 0,\n-\t\t\t\t\t\t       &opt->pathspec);\n-\t\t\tif (match == all_entries_not_interesting)\n-\t\t\t\tbreak;\n-\t\t\tif (match == entry_not_interesting)\n-\t\t\t\tcontinue;\n-\t\t}\n-\t\tshow_entry(opt, prefix, desc, base);\n-\t}\n-}\n-\n-/* A file entry went away or appeared */\n+/* An entry went away or appeared */\n static void show_entry(struct diff_options *opt, const char *prefix,\n \t\t       struct tree_desc *desc, struct strbuf *base)\n {\n@@ -85,23 +67,12 @@ static void show_entry(struct diff_options *opt, const char *prefix,\n \n \tstrbuf_add(base, path, pathlen);\n \tif (DIFF_OPT_TST(opt, RECURSIVE) && S_ISDIR(mode)) {\n-\t\tenum object_type type;\n-\t\tstruct tree_desc inner;\n-\t\tvoid *tree;\n-\t\tunsigned long size;\n-\n-\t\ttree = read_sha1_file(sha1, &type, &size);\n-\t\tif (!tree || type != OBJ_TREE)\n-\t\t\tdie(\"corrupt tree sha %s\", sha1_to_hex(sha1));\n-\n \t\tif (DIFF_OPT_TST(opt, TREE_IN_RECURSIVE))\n \t\t\topt->add_remove(opt, *prefix, mode, sha1, 1, base->buf, 0);\n \n \t\tstrbuf_addch(base, '/');\n-\n-\t\tinit_tree_desc(&inner, tree, size);\n-\t\tshow_tree(opt, prefix, &inner, base);\n-\t\tfree(tree);\n+\t\tdiff_tree_sha1(*prefix == '-' ? sha1 : NULL,\n+\t\t\t       *prefix == '+' ? sha1 : NULL, base->buf, opt);\n \t} else\n \t\topt->add_remove(opt, prefix[0], mode, sha1, 1, base->buf, 0);\n \n-- \n1.9.rc1.181.g641f458\n"},{"id":"235247","messageId":"696b8f0956f5d7699911a3bb56d66602301ee36c.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 06/19] tree-diff: consolidate code for emitting diffs and recursion in one place","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:38Z","receivedAt":"2014-02-24T16:21:38Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Currently both compare_tree_entry() and show_path() invoke opt diff\ncallbacks (opt->add_remove() and opt->change()), and also they both have\ncode which decides whether to recurse into sub-tree, and whether to emit\na tree as separate entry if DIFF_OPT_TREE_IN_RECURSIVE is set.\n\nI.e. we have code duplication and logic scattered on two places.\n\nLet's consolidate it - all diff emmiting code and recurion logic moves\nto show_entry, which is now named as show_path, because it shows diff\nfor a path, based on up to two tree entries, with actual diff emitting\ncode being kept in new helper emit_diff() for clarity.\n\nWhat we have as the result, is that compare_tree_entry is now free from\ncode with logic for diff generation, and also performance is not\naffected as timings for\n\n    `git log --raw --no-abbrev --no-renames`\n\nfor navy.git and `linux.git v3.10..v3.11`, just like in previous patch,\nstay the same.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n tree-diff.c | 115 ++++++++++++++++++++++++++++++++++++++++++++----------------\n 1 file changed, 84 insertions(+), 31 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 2ad7788..a5b9ff9 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -6,8 +6,8 @@\n #include \"diffcore.h\"\n #include \"tree.h\"\n \n-static void show_entry(struct diff_options *opt, const char *prefix,\n-\t\t       struct tree_desc *desc, struct strbuf *base);\n+static void show_path(struct strbuf *base, struct diff_options *opt,\n+\t\t      struct tree_desc *t1, struct tree_desc *t2);\n \n static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2,\n \t\t\t      struct strbuf *base, struct diff_options *opt)\n@@ -16,7 +16,6 @@ static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2,\n \tconst char *path1, *path2;\n \tconst unsigned char *sha1, *sha2;\n \tint cmp, pathlen1, pathlen2;\n-\tint old_baselen = base->len;\n \n \tsha1 = tree_entry_extract(t1, &path1, &mode1);\n \tsha2 = tree_entry_extract(t2, &path2, &mode2);\n@@ -30,51 +29,105 @@ static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2,\n \t */\n \tcmp = base_name_compare(path1, pathlen1, mode1, path2, pathlen2, mode2);\n \tif (cmp < 0) {\n-\t\tshow_entry(opt, \"-\", t1, base);\n+\t\tshow_path(base, opt, t1, /*t2=*/NULL);\n \t\treturn -1;\n \t}\n \tif (cmp > 0) {\n-\t\tshow_entry(opt, \"+\", t2, base);\n+\t\tshow_path(base, opt, /*t1=*/NULL, t2);\n \t\treturn 1;\n \t}\n \tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER) && !hashcmp(sha1, sha2) && mode1 == mode2)\n \t\treturn 0;\n \n-\tstrbuf_add(base, path1, pathlen1);\n-\tif (DIFF_OPT_TST(opt, RECURSIVE) && S_ISDIR(mode1)) {\n-\t\tif (DIFF_OPT_TST(opt, TREE_IN_RECURSIVE)) {\n-\t\t\topt->change(opt, mode1, mode2,\n-\t\t\t\t    sha1, sha2, 1, 1, base->buf, 0, 0);\n-\t\t}\n-\t\tstrbuf_addch(base, '/');\n-\t\tdiff_tree_sha1(sha1, sha2, base->buf, opt);\n-\t} else {\n-\t\topt->change(opt, mode1, mode2, sha1, sha2, 1, 1, base->buf, 0, 0);\n-\t}\n-\tstrbuf_setlen(base, old_baselen);\n+\tshow_path(base, opt, t1, t2);\n \treturn 0;\n }\n \n-/* An entry went away or appeared */\n-static void show_entry(struct diff_options *opt, const char *prefix,\n-\t\t       struct tree_desc *desc, struct strbuf *base)\n+\n+/* convert path, t1/t2 -> opt->diff_*() callbacks */\n+static void emit_diff(struct diff_options *opt, struct strbuf *path,\n+\t\t      struct tree_desc *t1, struct tree_desc *t2)\n+{\n+\tunsigned int mode1 = t1 ? t1->entry.mode : 0;\n+\tunsigned int mode2 = t2 ? t2->entry.mode : 0;\n+\n+\tif (mode1 && mode2) {\n+\t\topt->change(opt, mode1, mode2, t1->entry.sha1, t2->entry.sha1,\n+\t\t\t1, 1, path->buf, 0, 0);\n+\t}\n+\telse {\n+\t\tconst unsigned char *sha1;\n+\t\tunsigned int mode;\n+\t\tint addremove;\n+\n+\t\tif (mode2) {\n+\t\t\taddremove = '+';\n+\t\t\tsha1 = t2->entry.sha1;\n+\t\t\tmode = mode2;\n+\t\t}\n+\t\telse {\n+\t\t\taddremove = '-';\n+\t\t\tsha1 = t1->entry.sha1;\n+\t\t\tmode = mode1;\n+\t\t}\n+\n+\t\topt->add_remove(opt, addremove, mode, sha1, 1, path->buf, 0);\n+\t}\n+}\n+\n+\n+/* new path should be added to diff\n+ *\n+ * 3 cases on how/when it should be called and behaves:\n+ *\n+ *\t!t1,  t2\t-> path added, parent lacks it\n+ *\t t1, !t2\t-> path removed from parent\n+ *\t t1,  t2\t-> path modified\n+ */\n+static void show_path(struct strbuf *base, struct diff_options *opt,\n+\t\t      struct tree_desc *t1, struct tree_desc *t2)\n {\n \tunsigned mode;\n \tconst char *path;\n-\tconst unsigned char *sha1 = tree_entry_extract(desc, &path, &mode);\n-\tint pathlen = tree_entry_len(&desc->entry);\n+\tint pathlen;\n \tint old_baselen = base->len;\n+\tint isdir, recurse = 0, emitthis = 1;\n+\n+\t/* at least something has to be valid */\n+\tassert(t1 || t2);\n+\n+\tif (t2) {\n+\t\t/* path present in resulting tree */\n+\t\ttree_entry_extract(t2, &path, &mode);\n+\t\tpathlen = tree_entry_len(&t2->entry);\n+\t\tisdir = S_ISDIR(mode);\n+\t}\n+\telse {\n+\t\t/* a path was removed - take path from parent. Also take\n+\t\t * mode from parent, to decide on recursion.\n+\t\t */\n+\t\ttree_entry_extract(t1, &path, &mode);\n+\t\tpathlen = tree_entry_len(&t1->entry);\n+\n+\t\tisdir = S_ISDIR(mode);\n+\t\tmode = 0;\n+\t}\n+\n+\tif (DIFF_OPT_TST(opt, RECURSIVE) && isdir) {\n+\t\trecurse = 1;\n+\t\temitthis = DIFF_OPT_TST(opt, TREE_IN_RECURSIVE);\n+\t}\n \n \tstrbuf_add(base, path, pathlen);\n-\tif (DIFF_OPT_TST(opt, RECURSIVE) && S_ISDIR(mode)) {\n-\t\tif (DIFF_OPT_TST(opt, TREE_IN_RECURSIVE))\n-\t\t\topt->add_remove(opt, *prefix, mode, sha1, 1, base->buf, 0);\n \n+\tif (emitthis)\n+\t\temit_diff(opt, base, t1, t2);\n+\n+\tif (recurse) {\n \t\tstrbuf_addch(base, '/');\n-\t\tdiff_tree_sha1(*prefix == '-' ? sha1 : NULL,\n-\t\t\t       *prefix == '+' ? sha1 : NULL, base->buf, opt);\n-\t} else\n-\t\topt->add_remove(opt, prefix[0], mode, sha1, 1, base->buf, 0);\n+\t\tdiff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n+\t\t\t       t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n+\t}\n \n \tstrbuf_setlen(base, old_baselen);\n }\n@@ -117,12 +170,12 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \t\tif (!t1->size) {\n \t\t\tif (!t2->size)\n \t\t\t\tbreak;\n-\t\t\tshow_entry(opt, \"+\", t2, &base);\n+\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n \t\t\tupdate_tree_entry(t2);\n \t\t\tcontinue;\n \t\t}\n \t\tif (!t2->size) {\n-\t\t\tshow_entry(opt, \"-\", t1, &base);\n+\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n \t\t\tupdate_tree_entry(t1);\n \t\t\tcontinue;\n \t\t}\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235248","messageId":"487b0970053a3190da57c30f521c39c23f85dcec.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 07/19] tree-diff: don't assume compare_tree_entry() returns -1,0,1","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:39Z","receivedAt":"2014-02-24T16:21:39Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"It does, but we'll be reworking it in the next patch after it won't, and\nbesides it is better to stick to standard\nstrcmp/memcmp/base_name_compare/etc... convention, where comparison\nfunction returns <0, =0, >0\n\nRegarding performance, comparing for <0, =0, >0 should be a little bit\nfaster, than switch, because it is just 1 test-without-immediate\ninstruction and then up to 3 conditional branches, and in switch you\nhave up to 3 tests with immediate and up to 3 conditional branches.\n\nNo worry, that update_tree_entry(t2) is duplicated for =0 and >0 - it\nwill be good after we'll be adding support for multiparent walker and\nwill stay that way.\n\n=0 case goes first, because it happens more often in real diffs - i.e.\npaths are the same.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n tree-diff.c | 22 ++++++++++++++--------\n 1 file changed, 14 insertions(+), 8 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex a5b9ff9..5f7dbbf 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -179,18 +179,24 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \t\t\tupdate_tree_entry(t1);\n \t\t\tcontinue;\n \t\t}\n-\t\tswitch (compare_tree_entry(t1, t2, &base, opt)) {\n-\t\tcase -1:\n+\n+\t\tcmp = compare_tree_entry(t1, t2, &base, opt);\n+\n+\t\t/* t1 = t2 */\n+\t\tif (cmp == 0) {\n \t\t\tupdate_tree_entry(t1);\n-\t\t\tcontinue;\n-\t\tcase 0:\n+\t\t\tupdate_tree_entry(t2);\n+\t\t}\n+\n+\t\t/* t1 < t2 */\n+\t\telse if (cmp < 0) {\n \t\t\tupdate_tree_entry(t1);\n-\t\t\t/* Fallthrough */\n-\t\tcase 1:\n+\t\t}\n+\n+\t\t/* t1 > t2 */\n+\t\telse {\n \t\t\tupdate_tree_entry(t2);\n-\t\t\tcontinue;\n \t\t}\n-\t\tdie(\"git diff-tree: internal error\");\n \t}\n \n \tstrbuf_release(&base);\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235251","messageId":"d63db800c368a89a9620e07297136d43bf6f9ced.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 08/19] tree-diff: move all action-taking code out of compare_tree_entry()","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:40Z","receivedAt":"2014-02-24T16:21:40Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"- let it do only comparison.\n\nThis way the code is cleaner and more structured - cmp function only\ncompares, and the driver takes action based on comparison result.\n\nThere should be no change in performance, as effectively, we just move\nif series from on place into another, and merge it to was-already-there\nsame switch/if, so the result is maybe a little bit faster.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n tree-diff.c | 28 ++++++++++++----------------\n 1 file changed, 12 insertions(+), 16 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 5f7dbbf..6207372 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -9,8 +9,7 @@\n static void show_path(struct strbuf *base, struct diff_options *opt,\n \t\t      struct tree_desc *t1, struct tree_desc *t2);\n \n-static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2,\n-\t\t\t      struct strbuf *base, struct diff_options *opt)\n+static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2)\n {\n \tunsigned mode1, mode2;\n \tconst char *path1, *path2;\n@@ -28,19 +27,7 @@ static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2,\n \t * even when having the same name.\n \t */\n \tcmp = base_name_compare(path1, pathlen1, mode1, path2, pathlen2, mode2);\n-\tif (cmp < 0) {\n-\t\tshow_path(base, opt, t1, /*t2=*/NULL);\n-\t\treturn -1;\n-\t}\n-\tif (cmp > 0) {\n-\t\tshow_path(base, opt, /*t1=*/NULL, t2);\n-\t\treturn 1;\n-\t}\n-\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER) && !hashcmp(sha1, sha2) && mode1 == mode2)\n-\t\treturn 0;\n-\n-\tshow_path(base, opt, t1, t2);\n-\treturn 0;\n+\treturn cmp;\n }\n \n \n@@ -161,6 +148,8 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \tstrbuf_add(&base, base_str, baselen);\n \n \tfor (;;) {\n+\t\tint cmp;\n+\n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \t\tif (opt->pathspec.nr) {\n@@ -180,21 +169,28 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \t\t\tcontinue;\n \t\t}\n \n-\t\tcmp = compare_tree_entry(t1, t2, &base, opt);\n+\t\tcmp = compare_tree_entry(t1, t2);\n \n \t\t/* t1 = t2 */\n \t\tif (cmp == 0) {\n+\t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n+\t\t\t    hashcmp(t1->entry.sha1, t2->entry.sha1) ||\n+\t\t\t    (t1->entry.mode != t2->entry.mode))\n+\t\t\t\tshow_path(&base, opt, t1, t2);\n+\n \t\t\tupdate_tree_entry(t1);\n \t\t\tupdate_tree_entry(t2);\n \t\t}\n \n \t\t/* t1 < t2 */\n \t\telse if (cmp < 0) {\n+\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n \t\t\tupdate_tree_entry(t1);\n \t\t}\n \n \t\t/* t1 > t2 */\n \t\telse {\n+\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n \t\t\tupdate_tree_entry(t2);\n \t\t}\n \t}\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235252","messageId":"c0c58e97d0ad8cec28704976b1a46dc6a8702f23.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 09/19] tree-diff: rename compare_tree_entry -> tree_entry_pathcmp","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:41Z","receivedAt":"2014-02-24T16:21:41Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Since previous commit, this function does not compare entry hashes, and\nmode are compared fully outside of it. So what it does is compare entry\nnames and DIR bit in modes. Reflect this in its name.\n\nAdd documentation stating the semantics, and move the note about\nfiles/dirs comparison to it.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n tree-diff.c | 15 +++++++++------\n 1 file changed, 9 insertions(+), 6 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 6207372..3345534 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -9,7 +9,14 @@\n static void show_path(struct strbuf *base, struct diff_options *opt,\n \t\t      struct tree_desc *t1, struct tree_desc *t2);\n \n-static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2)\n+/*\n+ * Compare two tree entries, taking into account only path/S_ISDIR(mode),\n+ * but not their sha1's.\n+ *\n+ * NOTE files and directories *always* compare differently, even when having\n+ *      the same name - thanks to base_name_compare().\n+ */\n+static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n {\n \tunsigned mode1, mode2;\n \tconst char *path1, *path2;\n@@ -22,10 +29,6 @@ static int compare_tree_entry(struct tree_desc *t1, struct tree_desc *t2)\n \tpathlen1 = tree_entry_len(&t1->entry);\n \tpathlen2 = tree_entry_len(&t2->entry);\n \n-\t/*\n-\t * NOTE files and directories *always* compare differently,\n-\t * even when having the same name.\n-\t */\n \tcmp = base_name_compare(path1, pathlen1, mode1, path2, pathlen2, mode2);\n \treturn cmp;\n }\n@@ -169,7 +172,7 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \t\t\tcontinue;\n \t\t}\n \n-\t\tcmp = compare_tree_entry(t1, t2);\n+\t\tcmp = tree_entry_pathcmp(t1, t2);\n \n \t\t/* t1 = t2 */\n \t\tif (cmp == 0) {\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235249","messageId":"c35d8878ffbf36601bc8b6b20a4dae709f1cc552.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 10/19] tree-diff: show_path prototype is not needed anymore","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:42Z","receivedAt":"2014-02-24T16:21:42Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"We moved all action-taking code below show_path() in recent HEAD~~\n(tree-diff: move all action-taking code out of compare_tree_entry).\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n tree-diff.c | 3 ---\n 1 file changed, 3 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 3345534..20a4fda 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -6,9 +6,6 @@\n #include \"diffcore.h\"\n #include \"tree.h\"\n \n-static void show_path(struct strbuf *base, struct diff_options *opt,\n-\t\t      struct tree_desc *t1, struct tree_desc *t2);\n-\n /*\n  * Compare two tree entries, taking into account only path/S_ISDIR(mode),\n  * but not their sha1's.\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235250","messageId":"54aeccfe65926ff00147c3045c5bbae1583d68a7.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 11/19] tree-diff: simplify tree_entry_pathcmp","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:43Z","receivedAt":"2014-02-24T16:21:43Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Since an earlier \"Finally switch over tree descriptors to contain a\npre-parsed entry\", we can safely access all tree_desc->entry fields\ndirectly instead of first \"extracting\" them through\ntree_entry_extract.\n\nUse it. The code generated stays the same - only it now visually looks\ncleaner.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n tree-diff.c | 17 ++++++-----------\n 1 file changed, 6 insertions(+), 11 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 20a4fda..cf96ad7 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -15,18 +15,13 @@\n  */\n static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n {\n-\tunsigned mode1, mode2;\n-\tconst char *path1, *path2;\n-\tconst unsigned char *sha1, *sha2;\n-\tint cmp, pathlen1, pathlen2;\n+\tstruct name_entry *e1, *e2;\n+\tint cmp;\n \n-\tsha1 = tree_entry_extract(t1, &path1, &mode1);\n-\tsha2 = tree_entry_extract(t2, &path2, &mode2);\n-\n-\tpathlen1 = tree_entry_len(&t1->entry);\n-\tpathlen2 = tree_entry_len(&t2->entry);\n-\n-\tcmp = base_name_compare(path1, pathlen1, mode1, path2, pathlen2, mode2);\n+\te1 = &t1->entry;\n+\te2 = &t2->entry;\n+\tcmp = base_name_compare(e1->path, tree_entry_len(e1), e1->mode,\n+\t\t\t\te2->path, tree_entry_len(e2), e2->mode);\n \treturn cmp;\n }\n \n-- \n1.9.rc1.181.g641f458\n"},{"id":"235253","messageId":"dad40b2cf785e5951c105cac936d86a7bc6db8a3.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 12/19] tree-diff: remove special-case diff-emitting code for empty-tree cases","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:44Z","receivedAt":"2014-02-24T16:21:44Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"via teaching tree_entry_pathcmp() how to compare empty tree descriptors:\n\nWhile walking trees, we iterate their entries from lowest to highest in\nsort order, so empty tree means all entries were already went over.\n\nIf we artificially assign +infinity value to such tree \"entry\", it will\ngo after all usual entries, and through the usual driver loop we will be\ntaking the same actions, which were hand-coded for special cases, i.e.\n\n    t1 empty, t2 non-empty\n        pathcmp(+∞, t2) -> +1\n        show_path(/*t1=*/NULL, t2);     /* = t1 > t2 case in main loop */\n\n    t1 non-empty, t2-empty\n        pathcmp(t1, +∞) -> -1\n        show_path(t1, /*t2=*/NULL);     /* = t1 < t2 case in main loop */\n\nRight now we never go to when compared tree descriptors are infinity, as\nthis condition is checked in the loop beginning as finishing criteria,\nbut will do in the future, when there will be several parents iterated\nsimultaneously, and some pair of them would run to the end.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n tree-diff.c | 21 +++++++++------------\n 1 file changed, 9 insertions(+), 12 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex cf96ad7..2fd6d0e 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -12,12 +12,19 @@\n  *\n  * NOTE files and directories *always* compare differently, even when having\n  *      the same name - thanks to base_name_compare().\n+ *\n+ * NOTE empty (=invalid) descriptor(s) take part in comparison as +infty.\n  */\n static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n {\n \tstruct name_entry *e1, *e2;\n \tint cmp;\n \n+\tif (!t1->size)\n+\t\treturn t2->size ? +1 /* +∞ > c */  : 0 /* +∞ = +∞ */;\n+\telse if (!t2->size)\n+\t\treturn -1;\t/* c < +∞ */\n+\n \te1 = &t1->entry;\n \te2 = &t2->entry;\n \tcmp = base_name_compare(e1->path, tree_entry_len(e1), e1->mode,\n@@ -151,18 +158,8 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \t\t\tskip_uninteresting(t1, &base, opt);\n \t\t\tskip_uninteresting(t2, &base, opt);\n \t\t}\n-\t\tif (!t1->size) {\n-\t\t\tif (!t2->size)\n-\t\t\t\tbreak;\n-\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n-\t\t\tupdate_tree_entry(t2);\n-\t\t\tcontinue;\n-\t\t}\n-\t\tif (!t2->size) {\n-\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n-\t\t\tupdate_tree_entry(t1);\n-\t\t\tcontinue;\n-\t\t}\n+\t\tif (!t1->size && !t2->size)\n+\t\t\tbreak;\n \n \t\tcmp = tree_entry_pathcmp(t1, t2);\n \n-- \n1.9.rc1.181.g641f458\n"},{"id":"235254","messageId":"dcab79b62b86dea995828195ec8a7693d487d385.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 13/19] tree-diff: diff_tree() should now be static","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:45Z","receivedAt":"2014-02-24T16:21:45Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"We reworked all its users to use the functionality through\ndiff_tree_sha1 variant in recent patches (see \"tree-diff: allow\ndiff_tree_sha1 to accept NULL sha1\" and what comes next).\n\ndiff_tree() is now not used outside tree-diff.c - make it static.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\n ( new patch )\n\n diff.h      | 2 --\n tree-diff.c | 4 ++--\n 2 files changed, 2 insertions(+), 4 deletions(-)\n\ndiff --git a/diff.h b/diff.h\nindex e79f3b3..5d7b9f7 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -189,8 +189,6 @@ const char *diff_line_prefix(struct diff_options *);\n \n extern const char mime_boundary_leader[];\n \n-extern int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n-\t\t     const char *base, struct diff_options *opt);\n extern int diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \t\t\t  const char *base, struct diff_options *opt);\n extern int diff_root_tree_sha1(const unsigned char *new, const char *base,\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 2fd6d0e..b99622c 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -137,8 +137,8 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n \t}\n }\n \n-int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n-\t      const char *base_str, struct diff_options *opt)\n+static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n+\t\t     const char *base_str, struct diff_options *opt)\n {\n \tstruct strbuf base;\n \tint baselen = strlen(base_str);\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235255","messageId":"0b82e2de0edee4a590e7b4165c65938aef7090f5.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:46Z","receivedAt":"2014-02-24T16:21:46Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"In the next commit this will allow to reduce intermediate calls, when\nrecursing into subtrees - at that stage we know only subtree sha1, and\nit is natural for tree walker to start from that phase. For now we do\n\n    diff_tree\n        show_path\n            diff_tree_sha1\n                diff_tree\n                    ...\n\nand the change will allow to reduce it to\n\n    diff_tree\n        show_path\n            diff_tree\n\nAlso, it will allow to omit allocating strbuf for each subtree, and just\nreuse the common strbuf via playing with its len.\n\nThe above-mentioned improvements go in the next 2 patches.\n\nThe downside is that try_to_follow_renames(), if active, we cause\nre-reading of 2 initial trees, which was negligible based on my timings,\nand which is outweighed cogently by the upsides.\n\nNOTE To keep with the current interface and semantics, I needed to\nrename the function from diff_tree() to diff_tree_sha1(). As\ndiff_tree_sha1() was already used, and the function we are talking here\nis its more low-level helper, let's use Linux convention for prefixing\nsuch helpers with double underscore. So the final renaming is\n\n    diff_tree() -> __diff_tree_sha1()\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\nChanges since v1:\n\n - don't need to touch diff.h, as diff_tree() became static.\n\n tree-diff.c | 60 ++++++++++++++++++++++++++++--------------------------------\n 1 file changed, 28 insertions(+), 32 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex b99622c..f90acf5 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -137,12 +137,17 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n \t}\n }\n \n-static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n-\t\t     const char *base_str, struct diff_options *opt)\n+static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n+\t\t\t    const char *base_str, struct diff_options *opt)\n {\n+\tstruct tree_desc t1, t2;\n+\tvoid *t1tree, *t2tree;\n \tstruct strbuf base;\n \tint baselen = strlen(base_str);\n \n+\tt1tree = fill_tree_descriptor(&t1, old);\n+\tt2tree = fill_tree_descriptor(&t2, new);\n+\n \t/* Enable recursion indefinitely */\n \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n \n@@ -155,39 +160,41 @@ static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \t\tif (opt->pathspec.nr) {\n-\t\t\tskip_uninteresting(t1, &base, opt);\n-\t\t\tskip_uninteresting(t2, &base, opt);\n+\t\t\tskip_uninteresting(&t1, &base, opt);\n+\t\t\tskip_uninteresting(&t2, &base, opt);\n \t\t}\n-\t\tif (!t1->size && !t2->size)\n+\t\tif (!t1.size && !t2.size)\n \t\t\tbreak;\n \n-\t\tcmp = tree_entry_pathcmp(t1, t2);\n+\t\tcmp = tree_entry_pathcmp(&t1, &t2);\n \n \t\t/* t1 = t2 */\n \t\tif (cmp == 0) {\n \t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n-\t\t\t    hashcmp(t1->entry.sha1, t2->entry.sha1) ||\n-\t\t\t    (t1->entry.mode != t2->entry.mode))\n-\t\t\t\tshow_path(&base, opt, t1, t2);\n+\t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n+\t\t\t    (t1.entry.mode != t2.entry.mode))\n+\t\t\t\tshow_path(&base, opt, &t1, &t2);\n \n-\t\t\tupdate_tree_entry(t1);\n-\t\t\tupdate_tree_entry(t2);\n+\t\t\tupdate_tree_entry(&t1);\n+\t\t\tupdate_tree_entry(&t2);\n \t\t}\n \n \t\t/* t1 < t2 */\n \t\telse if (cmp < 0) {\n-\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n-\t\t\tupdate_tree_entry(t1);\n+\t\t\tshow_path(&base, opt, &t1, /*t2=*/NULL);\n+\t\t\tupdate_tree_entry(&t1);\n \t\t}\n \n \t\t/* t1 > t2 */\n \t\telse {\n-\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n-\t\t\tupdate_tree_entry(t2);\n+\t\t\tshow_path(&base, opt, /*t1=*/NULL, &t2);\n+\t\t\tupdate_tree_entry(&t2);\n \t\t}\n \t}\n \n \tstrbuf_release(&base);\n+\tfree(t2tree);\n+\tfree(t1tree);\n \treturn 0;\n }\n \n@@ -202,7 +209,7 @@ static inline int diff_might_be_rename(void)\n \t\t!DIFF_FILE_VALID(diff_queued_diff.queue[0]->one);\n }\n \n-static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, const char *base, struct diff_options *opt)\n+static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n {\n \tstruct diff_options diff_opts;\n \tstruct diff_queue_struct *q = &diff_queued_diff;\n@@ -240,7 +247,7 @@ static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, co\n \tdiff_opts.break_opt = opt->break_opt;\n \tdiff_opts.rename_score = opt->rename_score;\n \tdiff_setup_done(&diff_opts);\n-\tdiff_tree(t1, t2, base, &diff_opts);\n+\t__diff_tree_sha1(old, new, base, &diff_opts);\n \tdiffcore_std(&diff_opts);\n \tfree_pathspec(&diff_opts.pathspec);\n \n@@ -301,23 +308,12 @@ static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, co\n \n int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n {\n-\tvoid *tree1, *tree2;\n-\tstruct tree_desc t1, t2;\n-\tunsigned long size1, size2;\n \tint retval;\n \n-\ttree1 = fill_tree_descriptor(&t1, old);\n-\ttree2 = fill_tree_descriptor(&t2, new);\n-\tsize1 = t1.size;\n-\tsize2 = t2.size;\n-\tretval = diff_tree(&t1, &t2, base, opt);\n-\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename()) {\n-\t\tinit_tree_desc(&t1, tree1, size1);\n-\t\tinit_tree_desc(&t2, tree2, size2);\n-\t\ttry_to_follow_renames(&t1, &t2, base, opt);\n-\t}\n-\tfree(tree1);\n-\tfree(tree2);\n+\tretval = __diff_tree_sha1(old, new, base, opt);\n+\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n+\t\ttry_to_follow_renames(old, new, base, opt);\n+\n \treturn retval;\n }\n \n-- \n1.9.rc1.181.g641f458\n"},{"id":"235257","messageId":"7e5e5a381ba4204eac14c5be9e270ffdc0e2be7a.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 15/19] tree-diff: no need to call \"full\" diff_tree_sha1 from show_path()","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:47Z","receivedAt":"2014-02-24T16:21:47Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"As described in previous commit, when recursing into sub-trees, we can\nuse lower-level tree walker, since its interface is now sha1 based.\n\nThe change is ok, because diff_tree_sha1() only invokes\n__diff_tree_sha1(), and also, if base is empty, try_to_follow_renames().\nBut base is not empty here, as we have added a path and '/' before\nrecursing.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n tree-diff.c | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex f90acf5..aea0297 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -114,8 +114,8 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n \n \tif (recurse) {\n \t\tstrbuf_addch(base, '/');\n-\t\tdiff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n-\t\t\t       t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n+\t\t__diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n+\t\t\t\t t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n \t}\n \n \tstrbuf_setlen(base, old_baselen);\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235256","messageId":"301eb2377e0c5f670ffc26bda085d14dbee4f431.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH v2 16/19] tree-diff: reuse base str(buf) memory on sub-tree recursion","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:48Z","receivedAt":"2014-02-24T16:21:48Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"instead of allocating it all the time for every subtree in\n__diff_tree_sha1, let's allocate it once in diff_tree_sha1, and then all\ncallee just use it in stacking style, without memory allocations.\n\nThis should be faster, and for me this change gives the following\nslight speedups for\n\n    git log --raw --no-abbrev --no-renames --format='%H'\n\n                navy.git    linux.git v3.10..v3.11\n\n    before      0.618s      1.903s\n    after       0.611s      1.889s\n    speedup     1.1%        0.7%\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\nChanges since v1:\n\n - don't need to touch diff.h, as the function we are changing became static.\n\n tree-diff.c | 36 ++++++++++++++++++------------------\n 1 file changed, 18 insertions(+), 18 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex aea0297..c76821d 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -115,7 +115,7 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n \tif (recurse) {\n \t\tstrbuf_addch(base, '/');\n \t\t__diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n-\t\t\t\t t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n+\t\t\t\t t2 ? t2->entry.sha1 : NULL, base, opt);\n \t}\n \n \tstrbuf_setlen(base, old_baselen);\n@@ -138,12 +138,10 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n }\n \n static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n-\t\t\t    const char *base_str, struct diff_options *opt)\n+\t\t\t    struct strbuf *base, struct diff_options *opt)\n {\n \tstruct tree_desc t1, t2;\n \tvoid *t1tree, *t2tree;\n-\tstruct strbuf base;\n-\tint baselen = strlen(base_str);\n \n \tt1tree = fill_tree_descriptor(&t1, old);\n \tt2tree = fill_tree_descriptor(&t2, new);\n@@ -151,17 +149,14 @@ static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \t/* Enable recursion indefinitely */\n \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n \n-\tstrbuf_init(&base, PATH_MAX);\n-\tstrbuf_add(&base, base_str, baselen);\n-\n \tfor (;;) {\n \t\tint cmp;\n \n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \t\tif (opt->pathspec.nr) {\n-\t\t\tskip_uninteresting(&t1, &base, opt);\n-\t\t\tskip_uninteresting(&t2, &base, opt);\n+\t\t\tskip_uninteresting(&t1, base, opt);\n+\t\t\tskip_uninteresting(&t2, base, opt);\n \t\t}\n \t\tif (!t1.size && !t2.size)\n \t\t\tbreak;\n@@ -173,7 +168,7 @@ static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n \t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n \t\t\t    (t1.entry.mode != t2.entry.mode))\n-\t\t\t\tshow_path(&base, opt, &t1, &t2);\n+\t\t\t\tshow_path(base, opt, &t1, &t2);\n \n \t\t\tupdate_tree_entry(&t1);\n \t\t\tupdate_tree_entry(&t2);\n@@ -181,18 +176,17 @@ static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \n \t\t/* t1 < t2 */\n \t\telse if (cmp < 0) {\n-\t\t\tshow_path(&base, opt, &t1, /*t2=*/NULL);\n+\t\t\tshow_path(base, opt, &t1, /*t2=*/NULL);\n \t\t\tupdate_tree_entry(&t1);\n \t\t}\n \n \t\t/* t1 > t2 */\n \t\telse {\n-\t\t\tshow_path(&base, opt, /*t1=*/NULL, &t2);\n+\t\t\tshow_path(base, opt, /*t1=*/NULL, &t2);\n \t\t\tupdate_tree_entry(&t2);\n \t\t}\n \t}\n \n-\tstrbuf_release(&base);\n \tfree(t2tree);\n \tfree(t1tree);\n \treturn 0;\n@@ -209,7 +203,7 @@ static inline int diff_might_be_rename(void)\n \t\t!DIFF_FILE_VALID(diff_queued_diff.queue[0]->one);\n }\n \n-static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n+static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, struct strbuf *base, struct diff_options *opt)\n {\n \tstruct diff_options diff_opts;\n \tstruct diff_queue_struct *q = &diff_queued_diff;\n@@ -306,13 +300,19 @@ static void try_to_follow_renames(const unsigned char *old, const unsigned char\n \tq->nr = 1;\n }\n \n-int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n+int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base_str, struct diff_options *opt)\n {\n+\tstruct strbuf base;\n \tint retval;\n \n-\tretval = __diff_tree_sha1(old, new, base, opt);\n-\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n-\t\ttry_to_follow_renames(old, new, base, opt);\n+\tstrbuf_init(&base, PATH_MAX);\n+\tstrbuf_addstr(&base, base_str);\n+\n+\tretval = __diff_tree_sha1(old, new, &base, opt);\n+\tif (!*base_str && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n+\t\ttry_to_follow_renames(old, new, &base, opt);\n+\n+\tstrbuf_release(&base);\n \n \treturn retval;\n }\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235259","messageId":"f08867ee212e27074dbb4cbb06af408b16dba0a1.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 17/19] Portable alloca for Git","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:49Z","receivedAt":"2014-02-24T16:21:49Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"In the next patch we'll have to use alloca() for performance reasons,\nbut since alloca is non-standardized and is not portable, let's have a\ntrick with compatibility wrappers:\n\n1. at configure time, determine, do we have working alloca() through\n   alloca.h, and define\n\n    #define HAVE_ALLOCA_H\n\n   if yes.\n\n2. in code\n\n    #ifdef HAVE_ALLOCA_H\n    # include <alloca.h>\n    # define xalloca(size)      (alloca(size))\n    # define xalloca_free(p)    do {} while(0)\n    #else\n    # define xalloca(size)      (xmalloc(size))\n    # define xalloca_free(p)    (free(p))\n    #endif\n\n   and use it like\n\n   func() {\n       p = xalloca(size);\n       ...\n\n       xalloca_free(p);\n   }\n\nThis way, for systems, where alloca is available, we'll have optimal\non-stack allocations with fast executions. On the other hand, on\nsystems, where alloca is not available, this gracefully fallbacks to\nxmalloc/free.\n\nBoth autoconf and config.mak.uname configurations were updated. For\nautoconf, we are not bothering considering cases, when no alloca.h is\navailable, but alloca() works some other way - its simply alloca.h is\navailable and works or not, everything else is deep legacy.\n\nFor config.mak.uname, I've tried to make my almost-sure guess for where\nalloca() is available, but since I only have access to Linux it is the\nonly change I can be sure about myself, with relevant to other changed\nsystems people Cc'ed.\n\nNOTE\n\nSunOS and Windows had explicit -DHAVE_ALLOCA_H in their configurations.\nI've changed that to now-common HAVE_ALLOCA_H=YesPlease which should be\ncorrect.\n\nCc: Brandon Casey <drafnel@gmail.com>\nCc: Marius Storm-Olsen <mstormo@gmail.com>\nCc: Johannes Sixt <j6t@kdbg.org>\nCc: Johannes Schindelin <Johannes.Schindelin@gmx.de>\nCc: Ramsay Jones <ramsay@ramsay1.demon.co.uk>\nCc: Gerrit Pape <pape@smarden.org>\nCc: Petr Salinger <Petr.Salinger@seznam.cz>\nCc: Jonathan Nieder <jrnieder@gmail.com>\nCc: Thomas Schwinge <tschwinge@gnu.org>\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\n( new patch )\n\n Makefile          |  6 ++++++\n config.mak.uname  | 10 ++++++++--\n configure.ac      |  8 ++++++++\n git-compat-util.h |  8 ++++++++\n 4 files changed, 30 insertions(+), 2 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex dddaf4f..0334806 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -30,6 +30,8 @@ all::\n # Define LIBPCREDIR=/foo/bar if your libpcre header and library files are in\n # /foo/bar/include and /foo/bar/lib directories.\n #\n+# Define HAVE_ALLOCA_H if you have working alloca(3) defined in that header.\n+#\n # Define NO_CURL if you do not have libcurl installed.  git-http-fetch and\n # git-http-push are not built, and you cannot use http:// and https://\n # transports (neither smart nor dumb).\n@@ -1099,6 +1101,10 @@ ifdef USE_LIBPCRE\n \tEXTLIBS += -lpcre\n endif\n \n+ifdef HAVE_ALLOCA_H\n+\tBASIC_CFLAGS += -DHAVE_ALLOCA_H\n+endif\n+\n ifdef NO_CURL\n \tBASIC_CFLAGS += -DNO_CURL\n \tREMOTE_CURL_PRIMARY =\ndiff --git a/config.mak.uname b/config.mak.uname\nindex 7d31fad..71602ee 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -28,6 +28,7 @@ ifeq ($(uname_S),OSF1)\n \tNO_NSEC = YesPlease\n endif\n ifeq ($(uname_S),Linux)\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_STRLCPY = YesPlease\n \tNO_MKSTEMPS = YesPlease\n \tHAVE_PATHS_H = YesPlease\n@@ -35,6 +36,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_DEV_TTY = YesPlease\n endif\n ifeq ($(uname_S),GNU/kFreeBSD)\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_STRLCPY = YesPlease\n \tNO_MKSTEMPS = YesPlease\n \tHAVE_PATHS_H = YesPlease\n@@ -103,6 +105,7 @@ ifeq ($(uname_S),SunOS)\n \tNEEDS_NSL = YesPlease\n \tSHELL_PATH = /bin/bash\n \tSANE_TOOL_PATH = /usr/xpg6/bin:/usr/xpg4/bin\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_STRCASESTR = YesPlease\n \tNO_MEMMEM = YesPlease\n \tNO_MKDTEMP = YesPlease\n@@ -146,7 +149,7 @@ ifeq ($(uname_S),SunOS)\n \tendif\n \tINSTALL = /usr/ucb/install\n \tTAR = gtar\n-\tBASIC_CFLAGS += -D__EXTENSIONS__ -D__sun__ -DHAVE_ALLOCA_H\n+\tBASIC_CFLAGS += -D__EXTENSIONS__ -D__sun__\n endif\n ifeq ($(uname_O),Cygwin)\n \tifeq ($(shell expr \"$(uname_R)\" : '1\\.[1-6]\\.'),4)\n@@ -166,6 +169,7 @@ ifeq ($(uname_O),Cygwin)\n \telse\n \t\tNO_REGEX = UnfortunatelyYes\n \tendif\n+\tHAVE_ALLOCA_H = YesPlease\n \tNEEDS_LIBICONV = YesPlease\n \tNO_FAST_WORKING_DIRECTORY = UnfortunatelyYes\n \tNO_ST_BLOCKS_IN_STRUCT_STAT = YesPlease\n@@ -239,6 +243,7 @@ ifeq ($(uname_S),AIX)\n endif\n ifeq ($(uname_S),GNU)\n \t# GNU/Hurd\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_STRLCPY = YesPlease\n \tNO_MKSTEMPS = YesPlease\n \tHAVE_PATHS_H = YesPlease\n@@ -316,6 +321,7 @@ endif\n ifeq ($(uname_S),Windows)\n \tGIT_VERSION := $(GIT_VERSION).MSVC\n \tpathsep = ;\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_PREAD = YesPlease\n \tNEEDS_CRYPTO_WITH_SSL = YesPlease\n \tNO_LIBGEN_H = YesPlease\n@@ -363,7 +369,7 @@ ifeq ($(uname_S),Windows)\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\n-\tCOMPAT_CFLAGS = -D__USE_MINGW_ACCESS -DNOGDI -DHAVE_STRING_H -DHAVE_ALLOCA_H -Icompat -Icompat/regex -Icompat/win32 -DSTRIP_EXTENSION=\\\".exe\\\"\n+\tCOMPAT_CFLAGS = -D__USE_MINGW_ACCESS -DNOGDI -DHAVE_STRING_H -Icompat -Icompat/regex -Icompat/win32 -DSTRIP_EXTENSION=\\\".exe\\\"\n \tBASIC_LDFLAGS = -IGNORE:4217 -IGNORE:4049 -NOLOGO -SUBSYSTEM:CONSOLE -NODEFAULTLIB:MSVCRT.lib\n \tEXTLIBS = user32.lib advapi32.lib shell32.lib wininet.lib ws2_32.lib\n \tPTHREAD_LIBS =\ndiff --git a/configure.ac b/configure.ac\nindex 2f43393..0eae704 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -272,6 +272,14 @@ AS_HELP_STRING([],           [ARG can be also prefix for libpcre library and hea\n \tGIT_CONF_SUBST([LIBPCREDIR])\n     fi)\n #\n+# Define HAVE_ALLOCA_H if you have working alloca(3) defined in that header.\n+AC_FUNC_ALLOCA\n+case $ac_cv_working_alloca_h in\n+    yes)    HAVE_ALLOCA_H=YesPlease;;\n+    *)      HAVE_ALLOCA_H='';;\n+esac\n+GIT_CONF_SUBST([HAVE_ALLOCA_H])\n+#\n # Define NO_CURL if you do not have curl installed.  git-http-pull and\n # git-http-push are not built, and you cannot use http:// and https://\n # transports.\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex cbd86c3..63b2b3b 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -526,6 +526,14 @@ extern void release_pack_memory(size_t);\n typedef void (*try_to_free_t)(size_t);\n extern try_to_free_t set_try_to_free_routine(try_to_free_t);\n \n+#ifdef HAVE_ALLOCA_H\n+# include <alloca.h>\n+# define xalloca(size)      (alloca(size))\n+# define xalloca_free(p)    do {} while (0)\n+#else\n+# define xalloca(size)      (xmalloc(size))\n+# define xalloca_free(p)    (free(p))\n+#endif\n extern char *xstrdup(const char *str);\n extern void *xmalloc(size_t size);\n extern void *xmallocz(size_t size);\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235261","messageId":"7b307610fe214f47643a46b3e815487558db244e.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH v2 18/19] tree-diff: rework diff_tree() to generate diffs for multiparent cases as well","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:50Z","receivedAt":"2014-02-24T16:21:50Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Previously diff_tree(), which is now named __diff_tree_sha1(), was\ngenerating diff_filepair(s) for two trees t1 and t2, and that was\nusually used for a commit as t1=HEAD~, and t2=HEAD - i.e. to see changes\na commit introduces.\n\nIn Git, however, we have fundamentally built flexibility in that a\ncommit can have many parents - 1 for a plain commit, 2 for a simple merge,\nbut also more than 2 for merging several heads at once.\n\nFor merges there is a so called combine-diff, which shows diff, a merge\nintroduces by itself, omitting changes done by any parent. That works\nthrough first finding paths, that are different to all parents, and then\nshowing generalized diff, with separate columns for +/- for each parent.\nThe code lives in combine-diff.c .\n\nThere is an impedance mismatch, however, in that a commit could\ngenerally have any number of parents, and that while diffing trees, we\ndivide cases for 2-tree diffs and more-than-2-tree diffs. I mean there\nis no special casing for multiple parents commits in e.g.\nrevision-walker .\n\nThat impedance mismatch *hurts* *performance* *badly* for generating\ncombined diffs - in \"combine-diff: optimize combine_diff_path\nsets intersection\" I've already removed some slowness from it, but from\nthe timings provided there, it could be seen, that combined diffs still\ncost more than an order of magnitude more cpu time, compared to diff for\nusual commits, and that would only be an optimistic estimate, if we take\ninto account that for e.g. linux.git there is only one merge for several\ndozens of plain commits.\n\nThat slowness comes from the fact that currently, while generating\ncombined diff, a lot of time is spent computing diff(commit,commit^2)\njust to only then intersect that huge diff to almost small set of files\nfrom diff(commit,commit^1).\n\nThat's because at present, to compute combine-diff, for first finding\npaths, that \"every parent touches\", we use the following combine-diff\nproperty/definition:\n\nD(A,P1...Pn) = D(A,P1) ^ ... ^ D(A,Pn)      (w.r.t. paths)\n\nwhere\n\nD(A,P1...Pn) is combined diff between commit A, and parents Pi\n\nand\n\nD(A,Pi) is usual two-tree diff Pi..A\n\nSo if any of that D(A,Pi) is huge, tracting 1 n-parent combine-diff as n\n1-parent diffs and intersecting results will be slow.\n\nAnd usually, for linux.git and other topic-based workflows, that\nD(A,P2) is huge, because, if merge-base of A and P2, is several dozens\nof merges (from A, via first parent) below, that D(A,P2) will be diffing\nsum of merges from several subsystems to 1 subsystem.\n\nThe solution is to avoid computing n 1-parent diffs, and to find\nchanged-to-all-parents paths via scanning A's and all Pi's trees\nsimultaneously, at each step comparing their entries, and based on that\ncomparison, populate paths result, and deduce we could *skip*\n*recursing* into subdirectories, if at least for 1 parent, sha1 of that\ndir tree is the same as in A. That would save us from doing significant\namount of needless work.\n\nSuch approach is very similar to what diff_tree() does, only there we\ndeal with scanning only 2 trees simultaneously, and for n+1 tree, the\nlogic is a bit more complex:\n\nD(A,X1...Xn) calculation scheme\n-------------------------------\n\nD(A,X1...Xn) = D(A,X1) ^ ... ^ D(A,Xn)       (regarding resulting paths set)\n\n     D(A,Xj)         - diff between A..Xj\n     D(A,X1...Xn)    - combined diff from A to parents X1,...,Xn\n\nWe start from all trees, which are sorted, and compare their entries in\nlock-step:\n\n      A     X1       Xn\n      -     -        -\n     |a|   |x1|     |xn|\n     |-|   |--| ... |--|      i = argmin(x1...xn)\n     | |   |  |     |  |\n     |-|   |--|     |--|\n     |.|   |. |     |. |\n      .     .        .\n      .     .        .\n\nat any time there could be 3 cases:\n\n     1)  a < xi;\n     2)  a > xi;\n     3)  a = xi.\n\nSchematic deduction of what every case means, and what to do, follows:\n\n1)  a < xi  ->  ∀j a ∉ Xj  ->  \"+a\" ∈ D(A,Xj)  ->  D += \"+a\";  a↓\n\n2)  a > xi\n\n    2.1) ∃j: xj > xi  ->  \"-xi\" ∉ D(A,Xj)  ->  D += ø;  ∀ xk=xi  xk↓\n    2.2) ∀j  xj = xi  ->  xj ∉ A  ->  \"-xj\" ∈ D(A,Xj)  ->  D += \"-xi\";  ∀j xj↓\n\n3)  a = xi\n\n    3.1) ∃j: xj > xi  ->  \"+a\" ∈ D(A,Xj)  ->  only xk=xi remains to investigate\n    3.2) xj = xi  ->  investigate δ(a,xj)\n     |\n     |\n     v\n\n    3.1+3.2) looking at δ(a,xk) ∀k: xk=xi - if all != ø  ->\n\n                      ⎧δ(a,xk)  - if xk=xi\n             ->  D += ⎨\n                      ⎩\"+a\"     - if xk>xi\n\n    in any case a↓  ∀ xk=xi  xk↓\n\n~\n\nFor comparison, here is how diff_tree() works:\n\nD(A,B) calculation scheme\n-------------------------\n\n    A     B\n    -     -\n   |a|   |b|    a < b   ->  a ∉ B   ->   D(A,B) +=  +a    a↓\n   |-|   |-|    a > b   ->  b ∉ A   ->   D(A,B) +=  -b    b↓\n   | |   | |    a = b   ->  investigate δ(a,b)            a↓ b↓\n   |-|   |-|\n   |.|   |.|\n    .     .\n    .     .\n\n~~~~~~~~\n\nThis patch generalizes diff tree-walker to work with arbitrary number of\nparents as described above - i.e. now there is a resulting tree t, and\nsome parents trees tp[i] i=[0..nparent). The generalization builds on\nthe fact that usual diff\n\nD(A,B)\n\nis by definition the same as combined diff\n\nD(A,[B]),\n\nso if we could rework the code for common case and make it be not slower\nfor nparent=1 case, usual diff(t1,t2) generation will not be slower, and\nmultiparent diff tree-walker would greatly benefit generating\ncombine-diff.\n\nWhat we do is as follows:\n\n1) diff tree-walker __diff_tree_sha1() is internally reworked to be\n   a paths generator (new name diff_tree_paths()), with each generated path\n   being `struct combine_diff_path` with info for path, new sha1,mode and for\n   every parent which sha1,mode it was in it.\n\n2) From that info, we can still generate usual diff queue with\n   struct diff_filepairs, via \"exporting\" generated\n   combine_diff_path, if we know we run for nparent=1 case.\n   (see emit_diff() which is now named emit_diff_first_parent_only())\n\n3) In order for diff_can_quit_early(), which checks\n\n       DIFF_OPT_TST(opt, HAS_CHANGES))\n\n   to work, that exporting have to be happening not in bulk, but\n   incrementally, one diff path at a time.\n\n   For such consumers, there is a new callback in diff_options\n   introduced:\n\n       ->pathchange(opt, struct combine_diff_path *)\n\n   which, if set to !NULL, is called for every generated path.\n\n   (see new compat __diff_tree_sha1() wrapper around new paths\n    generator for setup)\n\n4) The paths generation itself, is reworked from previous\n   __diff_tree_sha1() code according to \"D(A,X1...Xn) calculation\n   scheme\" provided above:\n\n   On the start we allocate [nparent] arrays in place what was\n   earlier just for one parent tree.\n\n   then we just generalize loops, and comparison according to the\n   algorithm.\n\nSome notes(*):\n\n1) alloca(), for small arrays, is used for \"runs not slower for\n   nparent=1 case than before\" goal - if we change it to xmalloc()/free()\n   the timings get ~1% worse. For alloca() we use just-introduced\n   xalloca/xalloca_free compatibility wrappers, so it should not be a\n   portability problem.\n\n2) For every parent tree, we need to keep a tag, whether entry from that\n   parent equals to entry from minimal parent. For performance reasons I'm\n   keeping that tag in entry's mode field in unused bit - see S_IFXMIN_NEQ.\n   Not doing so, we'd need to alloca another [nparent] array, which hurts\n   performance.\n\n3) For emitted paths, memory could be reused, if we know the path was\n   processed via callback and will not be needed later. We use efficient\n   hand-made realloc-style __path_appendnew(), that saves us from ~1-1.5%\n   of potential additional slowdown.\n\n4) goto(s) are used in several places, as the code executes a little bit\n   faster with lowered register pressure.\n\nAlso\n\n- we should now check for FIND_COPIES_HARDER not only when two entries\n  names are the same, and their hashes are equal, but also for a case,\n  when a path was removed from some of all parents having it.\n\n  The reason is, if we don't, that path won't be emitted at all (see\n  \"a > xi\" case), and we'll just skip it, and FIND_COPIES_HARDER wants\n  all paths - with diff or without - to be emitted, to be later analyzed\n  for being copies sources.\n\n  The new check is only necessary for nparent >1, as for nparent=1 case\n  xmin_eqtotal always =1 =nparent, and a path is always added to diff as\n  removal.\n\n~~~~~~~~\n\nTimings for\n\n    # without -c, i.e. testing only nparent=1 case\n    `git log --raw --no-abbrev --no-renames`\n\nbefore and after the patch are as follows:\n\n                navy.git        linux.git v3.10..v3.11\n\n    before      0.611s          1.889s\n    after       0.619s          1.907s\n    slowdown    1.3%            0.9%\n\nThis timings show we did no harm to usual diff(tree1,tree2) generation.\nFrom the table we can see that we actually did ~1% slowdown, but I think\nI've \"earned\" that 1% in the previous patch (\"tree-diff: reuse base\nstr(buf) memory on sub-tree recursion\", HEAD~~) so for nparent=1 case,\nnet timings stays approximately the same.\n\nThe output also stayed the same.\n\n(*) If we revert 1)-4) to more usual techniques, for nparent=1 case,\n    we'll get ~2-2.5% of additional slowdown, which I've tried to avoid, as\n   \"do no harm for nparent=1 case\" rule.\n\nFor linux.git, combined diff will run an order of magnitude faster and\nappropriate timings will be provided in the next commit, as we'll be\ntaking advantage of the new diff tree-walker for combined-diff\ngeneration there.\n\nP.S. and combined diff is not some exotic/for-play-only stuff - for\nexample for a program I write to represent Git archives as readonly\nfilesystem, there is initial scan with\n\n    `git log --reverse --raw --no-abbrev --no-renames -c`\n\nto extract log of what was created/changed when, as a result building a\nmap\n\n    {}  sha1    ->  in which commit (and date) a content was added\n\nthat `-c` means also show combined diff for merges, and without them, if\na merge is non-trivial (merges changes from two parents with both having\nseparate changes to a file), or an evil one, the map will not be full,\ni.e. some valid sha1 would be absent from it.\n\nThat case was my initial motivation for combined diffs speedup.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\nChanges since v1:\n\n- fixed last-minute thinko/bug last time introduced on my side (sorry) with\n  opt->pathchange manipulation in __diff_tree_sha1() - we were forgetting to\n  restore opt->pathchange, which led to incorrect log -c (merges _and_ plain\n  diff-tree) output;\n\n  This time, I've verified several times, log output stays really the same.\n\n- direct use of alloca() changed to portability wrappers xalloca/xalloca_free\n  which gracefully degrade to xmalloc/free on systems, where alloca is not\n  available (see new patch 17).\n\n- \"i = 0; do { ... } while (++i < nparent)\" is back to usual looping\n  \"for (i = 0; i < nparent; ++)\", as I've re-measured timings and the\n  difference is negligible.\n\n  ( Initially, when I was fighting for every cycle it made sense, but real\n    no-slowdown turned out to be related to avoiding mallocs, load trees in correct\n    order and reducing register pressure. )\n\n- S_IFXMIN_NEQ definition moved out to cache.h, to have all modes registry in one place;\n\n\n- p0 -> first_parent; corrected comments about how emit_diff_first_parent_only\n  behaves;\n\n\nnot changed:\n\n- low-level helpers are still named with \"__\" prefix as, imho, that is the best\n  convention to name such helpers, without sacrificing signal/noise ratio. All\n  of them are now static though.\n\n cache.h     |  15 ++\n diff.c      |   1 +\n diff.h      |  10 ++\n tree-diff.c | 508 ++++++++++++++++++++++++++++++++++++++++++++++++++++--------\n 4 files changed, 471 insertions(+), 63 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex dc040fb..e7f5a0c 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -75,6 +75,21 @@ unsigned long git_deflate_bound(git_zstream *, unsigned long);\n #define S_ISGITLINK(m)\t(((m) & S_IFMT) == S_IFGITLINK)\n \n /*\n+ * Some mode bits are also used internally for computations.\n+ *\n+ * They *must* not overlap with any valid modes, and they *must* not be emitted\n+ * to outside world - i.e. appear on disk or network. In other words, it's just\n+ * temporary fields, which we internally use, but they have to stay in-house.\n+ *\n+ * ( such approach is valid, as standard S_IF* fits into 16 bits, and in Git\n+ *   codebase mode is `unsigned int` which is assumed to be at least 32 bits )\n+ */\n+\n+/* used internally in tree-diff */\n+#define S_DIFFTREE_IFXMIN_NEQ\t0x80000000\n+\n+\n+/*\n  * Intensive research over the course of many years has shown that\n  * port 9418 is totally unused by anything else. Or\n  *\ndiff --git a/diff.c b/diff.c\nindex 8e4a6a9..cda4aa8 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -3216,6 +3216,7 @@ void diff_setup(struct diff_options *options)\n \toptions->context = diff_context_default;\n \tDIFF_OPT_SET(options, RENAME_EMPTY);\n \n+\t/* pathchange left =NULL by default */\n \toptions->change = diff_change;\n \toptions->add_remove = diff_addremove;\n \toptions->use_color = diff_use_color_default;\ndiff --git a/diff.h b/diff.h\nindex 5d7b9f7..732dca7 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -15,6 +15,10 @@ struct diff_filespec;\n struct userdiff_driver;\n struct sha1_array;\n struct commit;\n+struct combine_diff_path;\n+\n+typedef int (*pathchange_fn_t)(struct diff_options *options,\n+\t\t struct combine_diff_path *path);\n \n typedef void (*change_fn_t)(struct diff_options *options,\n \t\t unsigned old_mode, unsigned new_mode,\n@@ -157,6 +161,7 @@ struct diff_options {\n \tint close_file;\n \n \tstruct pathspec pathspec;\n+\tpathchange_fn_t pathchange;\n \tchange_fn_t change;\n \tadd_remove_fn_t add_remove;\n \tdiff_format_fn_t format_callback;\n@@ -189,6 +194,11 @@ const char *diff_line_prefix(struct diff_options *);\n \n extern const char mime_boundary_leader[];\n \n+extern\n+struct combine_diff_path *diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parent_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt);\n extern int diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \t\t\t  const char *base, struct diff_options *opt);\n extern int diff_root_tree_sha1(const unsigned char *new, const char *base,\ndiff --git a/tree-diff.c b/tree-diff.c\nindex c76821d..b682d77 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -7,6 +7,22 @@\n #include \"tree.h\"\n \n /*\n+ * internal mode marker, saying a tree entry != entry of tp[imin]\n+ * (see __diff_tree_paths for what it means there)\n+ *\n+ * we will update/use/emit entry for diff only with it unset.\n+ */\n+#define S_IFXMIN_NEQ\tS_DIFFTREE_IFXMIN_NEQ\n+\n+\n+static struct combine_diff_path *__diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parents_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt);\n+static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n+\t\t\t    struct strbuf *base, struct diff_options *opt);\n+\n+/*\n  * Compare two tree entries, taking into account only path/S_ISDIR(mode),\n  * but not their sha1's.\n  *\n@@ -33,72 +49,153 @@ static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n }\n \n \n-/* convert path, t1/t2 -> opt->diff_*() callbacks */\n-static void emit_diff(struct diff_options *opt, struct strbuf *path,\n-\t\t      struct tree_desc *t1, struct tree_desc *t2)\n+/*\n+ * convert path -> opt->diff_*() callbacks\n+ *\n+ * emits diff to first parent only, and tells diff tree-walker that we are done\n+ * with p and it can be freed.\n+ */\n+static int emit_diff_first_parent_only(struct diff_options *opt, struct combine_diff_path *p)\n {\n-\tunsigned int mode1 = t1 ? t1->entry.mode : 0;\n-\tunsigned int mode2 = t2 ? t2->entry.mode : 0;\n-\n-\tif (mode1 && mode2) {\n-\t\topt->change(opt, mode1, mode2, t1->entry.sha1, t2->entry.sha1,\n-\t\t\t1, 1, path->buf, 0, 0);\n+\tstruct combine_diff_parent *p0 = &p->parent[0];\n+\tif (p->mode && p0->mode) {\n+\t\topt->change(opt, p0->mode, p->mode, p0->sha1, p->sha1,\n+\t\t\t1, 1, p->path, 0, 0);\n \t}\n \telse {\n \t\tconst unsigned char *sha1;\n \t\tunsigned int mode;\n \t\tint addremove;\n \n-\t\tif (mode2) {\n+\t\tif (p->mode) {\n \t\t\taddremove = '+';\n-\t\t\tsha1 = t2->entry.sha1;\n-\t\t\tmode = mode2;\n+\t\t\tsha1 = p->sha1;\n+\t\t\tmode = p->mode;\n \t\t}\n \t\telse {\n \t\t\taddremove = '-';\n-\t\t\tsha1 = t1->entry.sha1;\n-\t\t\tmode = mode1;\n+\t\t\tsha1 = p0->sha1;\n+\t\t\tmode = p0->mode;\n \t\t}\n \n-\t\topt->add_remove(opt, addremove, mode, sha1, 1, path->buf, 0);\n+\t\topt->add_remove(opt, addremove, mode, sha1, 1, p->path, 0);\n \t}\n+\n+\treturn 0;\t/* we are done with p */\n }\n \n \n-/* new path should be added to diff\n+/*\n+ * Make a new combine_diff_path from path/mode/sha1\n+ * and append it to paths list tail.\n+ *\n+ * Memory for created elements could be reused:\n+ *\n+ *\t- if last->next == NULL, the memory is allocated;\n+ *\n+ *\t- if last->next != NULL, it is assumed that p=last->next was returned\n+ *\t  earlier by this function, and p->next was *not* modified.\n+ *\t  The memory is then reused from p.\n+ *\n+ * so for clients,\n+ *\n+ * - if you do need to keep the element\n+ *\n+ *\tp = __path_appendnew(p, ...);\n+ *\tprocess(p);\n+ *\tp->next = NULL;\n+ *\n+ * - if you don't need to keep the element after processing\n+ *\n+ *\tpprev = p;\n+ *\tp = __path_appendnew(p, ...);\n+ *\tprocess(p);\n+ *\tp = pprev;\n+ *\t; don't forget to free tail->next in the end\n+ *\n+ * p->parent[] remains uninitialized.\n+ */\n+static struct combine_diff_path *__path_appendnew(struct combine_diff_path *last,\n+\tint nparent, const struct strbuf *base, const char *path, int pathlen,\n+\tunsigned mode, const unsigned char *sha1)\n+{\n+\tstruct combine_diff_path *p;\n+\tint len = base->len + pathlen;\n+\tint alloclen = combine_diff_path_size(nparent, len);\n+\n+\t/* if last->next is !NULL - it is a pre-allocated memory, we can reuse */\n+\tp = last->next;\n+\tif (p && (alloclen > (intptr_t)p->next)) {\n+\t\tfree(p);\n+\t\tp = NULL;\n+\t}\n+\n+\tif (!p) {\n+\t\tp = xmalloc(alloclen);\n+\n+\t\t/*\n+\t\t * until we go to it next round, .next holds how many bytes we\n+\t\t * allocated (for faster realloc - we don't need copying old data).\n+\t\t */\n+\t\tp->next = (struct combine_diff_path *)(intptr_t)alloclen;\n+\t}\n+\n+\tlast->next = p;\n+\n+\tp->path = (char *)&(p->parent[nparent]);\n+\tmemcpy(p->path, base->buf, base->len);\n+\tmemcpy(p->path + base->len, path, pathlen);\n+\tp->path[len] = 0;\n+\tp->mode = mode;\n+\thashcpy(p->sha1, sha1 ? sha1 : null_sha1);\n+\n+\treturn p;\n+}\n+\n+/*\n+ * new path should be added to combine diff\n  *\n  * 3 cases on how/when it should be called and behaves:\n  *\n- *\t!t1,  t2\t-> path added, parent lacks it\n- *\t t1, !t2\t-> path removed from parent\n- *\t t1,  t2\t-> path modified\n+ *\t t, !tp\t\t-> path added, all parents lack it\n+ *\t!t,  tp\t\t-> path removed from all parents\n+ *\t t,  tp\t\t-> path modified/added\n+ *\t\t\t   (M for tp[i]=tp[imin], A otherwise)\n  */\n-static void show_path(struct strbuf *base, struct diff_options *opt,\n-\t\t      struct tree_desc *t1, struct tree_desc *t2)\n+static struct combine_diff_path *emit_path(struct combine_diff_path *p,\n+\tstruct strbuf *base, struct diff_options *opt, int nparent,\n+\tstruct tree_desc *t, struct tree_desc *tp,\n+\tint imin)\n {\n \tunsigned mode;\n \tconst char *path;\n+\tconst unsigned char *sha1;\n \tint pathlen;\n \tint old_baselen = base->len;\n-\tint isdir, recurse = 0, emitthis = 1;\n+\tint i, isdir, recurse = 0, emitthis = 1;\n \n \t/* at least something has to be valid */\n-\tassert(t1 || t2);\n+\tassert(t || tp);\n \n-\tif (t2) {\n+\tif (t) {\n \t\t/* path present in resulting tree */\n-\t\ttree_entry_extract(t2, &path, &mode);\n-\t\tpathlen = tree_entry_len(&t2->entry);\n+\t\tsha1 = tree_entry_extract(t, &path, &mode);\n+\t\tpathlen = tree_entry_len(&t->entry);\n \t\tisdir = S_ISDIR(mode);\n \t}\n \telse {\n-\t\t/* a path was removed - take path from parent. Also take\n-\t\t * mode from parent, to decide on recursion.\n+\t\t/*\n+\t\t * a path was removed - take path from imin parent. Also take\n+\t\t * mode from that parent, to decide on recursion(1).\n+\t\t *\n+\t\t * 1) all modes for tp[k]=tp[imin] should be the same wrt\n+\t\t *    S_ISDIR, thanks to base_name_compare().\n \t\t */\n-\t\ttree_entry_extract(t1, &path, &mode);\n-\t\tpathlen = tree_entry_len(&t1->entry);\n+\t\ttree_entry_extract(&tp[imin], &path, &mode);\n+\t\tpathlen = tree_entry_len(&tp[imin].entry);\n \n \t\tisdir = S_ISDIR(mode);\n+\t\tsha1 = NULL;\n \t\tmode = 0;\n \t}\n \n@@ -107,18 +204,81 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n \t\temitthis = DIFF_OPT_TST(opt, TREE_IN_RECURSIVE);\n \t}\n \n-\tstrbuf_add(base, path, pathlen);\n+\tif (emitthis) {\n+\t\tint keep;\n+\t\tstruct combine_diff_path *pprev = p;\n+\t\tp = __path_appendnew(p, nparent, base, path, pathlen, mode, sha1);\n \n-\tif (emitthis)\n-\t\temit_diff(opt, base, t1, t2);\n+\t\tfor (i = 0; i < nparent; ++i) {\n+\t\t\t/*\n+\t\t\t * tp[i] is valid, if present and if tp[i]==tp[imin] -\n+\t\t\t * otherwise, we should ignore it.\n+\t\t\t */\n+\t\t\tint tpi_valid = tp && !(tp[i].entry.mode & S_IFXMIN_NEQ);\n+\n+\t\t\tconst unsigned char *sha1_i;\n+\t\t\tunsigned mode_i;\n+\n+\t\t\tp->parent[i].status =\n+\t\t\t\t!t ? DIFF_STATUS_DELETED :\n+\t\t\t\t\ttpi_valid ?\n+\t\t\t\t\t\tDIFF_STATUS_MODIFIED :\n+\t\t\t\t\t\tDIFF_STATUS_ADDED;\n+\n+\t\t\tif (tpi_valid) {\n+\t\t\t\tsha1_i = tp[i].entry.sha1;\n+\t\t\t\tmode_i = tp[i].entry.mode;\n+\t\t\t}\n+\t\t\telse {\n+\t\t\t\tsha1_i = NULL;\n+\t\t\t\tmode_i = 0;\n+\t\t\t}\n+\n+\t\t\tp->parent[i].mode = mode_i;\n+\t\t\thashcpy(p->parent[i].sha1, sha1_i ? sha1_i : null_sha1);\n+\t\t}\n+\n+\t\tkeep = 1;\n+\t\tif (opt->pathchange)\n+\t\t\tkeep = opt->pathchange(opt, p);\n+\n+\t\t/*\n+\t\t * If a path was filtered or consumed - we don't need to add it\n+\t\t * to the list and can reuse its memory, leaving it as\n+\t\t * pre-allocated element on the tail.\n+\t\t *\n+\t\t * On the other hand, if path needs to be kept, we need to\n+\t\t * correct its .next to NULL, as it was pre-initialized to how\n+\t\t * much memory was allocated.\n+\t\t *\n+\t\t * see __path_appendnew() for details.\n+\t\t */\n+\t\tif (!keep)\n+\t\t\tp = pprev;\n+\t\telse\n+\t\t\tp->next = NULL;\n+\t}\n \n \tif (recurse) {\n+\t\tconst unsigned char **parents_sha1;\n+\n+\t\tparents_sha1 = xalloca(nparent * sizeof(parents_sha1[0]));\n+\t\tfor (i = 0; i < nparent; ++i) {\n+\t\t\t/* same rule as in emitthis */\n+\t\t\tint tpi_valid = tp && !(tp[i].entry.mode & S_IFXMIN_NEQ);\n+\n+\t\t\tparents_sha1[i] = tpi_valid ? tp[i].entry.sha1\n+\t\t\t\t\t\t    : NULL;\n+\t\t}\n+\n+\t\tstrbuf_add(base, path, pathlen);\n \t\tstrbuf_addch(base, '/');\n-\t\t__diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n-\t\t\t\t t2 ? t2->entry.sha1 : NULL, base, opt);\n+\t\tp = __diff_tree_paths(p, sha1, parents_sha1, nparent, base, opt);\n+\t\txalloca_free(parents_sha1);\n \t}\n \n \tstrbuf_setlen(base, old_baselen);\n+\treturn p;\n }\n \n static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n@@ -137,59 +297,260 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n \t}\n }\n \n-static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n-\t\t\t    struct strbuf *base, struct diff_options *opt)\n+\n+/*\n+ * generate paths for combined diff D(sha1,parents_sha1[])\n+ *\n+ * Resulting paths are appended to combine_diff_path linked list, and also, are\n+ * emitted on the go via opt->pathchange() callback, so it is possible to\n+ * process the result as batch or incrementally.\n+ *\n+ * The paths are generated scanning new tree and all parents trees\n+ * simultaneously, similarly to what diff_tree() was doing for 2 trees.\n+ * The theory behind such scan is as follows:\n+ *\n+ *\n+ * D(A,X1...Xn) calculation scheme\n+ * -------------------------------\n+ *\n+ * D(A,X1...Xn) = D(A,X1) ^ ... ^ D(A,Xn)\t(regarding resulting paths set)\n+ *\n+ *\tD(A,Xj)\t\t- diff between A..Xj\n+ *\tD(A,X1...Xn)\t- combined diff from A to parents X1,...,Xn\n+ *\n+ *\n+ * We start from all trees, which are sorted, and compare their entries in\n+ * lock-step:\n+ *\n+ *\t A     X1       Xn\n+ *\t -     -        -\n+ *\t|a|   |x1|     |xn|\n+ *\t|-|   |--| ... |--|      i = argmin(x1...xn)\n+ *\t| |   |  |     |  |\n+ *\t|-|   |--|     |--|\n+ *\t|.|   |. |     |. |\n+ *\t .     .        .\n+ *\t .     .        .\n+ *\n+ * at any time there could be 3 cases:\n+ *\n+ *\t1)  a < xi;\n+ *\t2)  a > xi;\n+ *\t3)  a = xi.\n+ *\n+ * Schematic deduction of what every case means, and what to do, follows:\n+ *\n+ * 1)  a < xi  ->  ∀j a ∉ Xj  ->  \"+a\" ∈ D(A,Xj)  ->  D += \"+a\";  a↓\n+ *\n+ * 2)  a > xi\n+ *\n+ *     2.1) ∃j: xj > xi  ->  \"-xi\" ∉ D(A,Xj)  ->  D += ø;  ∀ xk=xi  xk↓\n+ *     2.2) ∀j  xj = xi  ->  xj ∉ A  ->  \"-xj\" ∈ D(A,Xj)  ->  D += \"-xi\";  ∀j xj↓\n+ *\n+ * 3)  a = xi\n+ *\n+ *     3.1) ∃j: xj > xi  ->  \"+a\" ∈ D(A,Xj)  ->  only xk=xi remains to investigate\n+ *     3.2) xj = xi  ->  investigate δ(a,xj)\n+ *      |\n+ *      |\n+ *      v\n+ *\n+ *     3.1+3.2) looking at δ(a,xk) ∀k: xk=xi - if all != ø  ->\n+ *\n+ *                       ⎧δ(a,xk)  - if xk=xi\n+ *              ->  D += ⎨\n+ *                       ⎩\"+a\"     - if xk>xi\n+ *\n+ *\n+ *     in any case a↓  ∀ xk=xi  xk↓\n+ *\n+ *\n+ * ~~~~~~~~\n+ *\n+ * NOTE\n+ *\n+ *\tUsual diff D(A,B) is by definition the same as combined diff D(A,[B]),\n+ *\tso this diff paths generator can, and is used, for plain diffs\n+ *\tgeneration too.\n+ *\n+ *\tPlease keep attention to the common D(A,[B]) case when working on the\n+ *\tcode, in order not to slow it down.\n+ *\n+ * NOTE\n+ *\tnparent must be > 0.\n+ */\n+\n+\n+/* ∀ xk=xi  xk↓ */\n+static inline void update_tp_entries(struct tree_desc *tp, int nparent)\n {\n-\tstruct tree_desc t1, t2;\n-\tvoid *t1tree, *t2tree;\n+\tint i;\n+\tfor (i = 0; i < nparent; ++i)\n+\t\tif (!(tp[i].entry.mode & S_IFXMIN_NEQ))\n+\t\t\tupdate_tree_entry(&tp[i]);\n+}\n+\n+static struct combine_diff_path *__diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parents_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt)\n+{\n+\tstruct tree_desc t, *tp;\n+\tvoid *ttree, **tptree;\n+\tint i;\n+\n+\ttp     = xalloca(nparent * sizeof(tp[0]));\n+\ttptree = xalloca(nparent * sizeof(tptree[0]));\n \n-\tt1tree = fill_tree_descriptor(&t1, old);\n-\tt2tree = fill_tree_descriptor(&t2, new);\n+\t/*\n+\t * load parents first, as they are probably already cached.\n+\t *\n+\t * ( log_tree_diff() parses commit->parent before calling here via\n+\t *   diff_tree_sha1(parent, commit) )\n+\t */\n+\tfor (i = 0; i < nparent; ++i)\n+\t\ttptree[i] = fill_tree_descriptor(&tp[i], parents_sha1[i]);\n+\tttree = fill_tree_descriptor(&t, sha1);\n \n \t/* Enable recursion indefinitely */\n \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n \n \tfor (;;) {\n-\t\tint cmp;\n+\t\tint imin, cmp;\n \n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n+\n \t\tif (opt->pathspec.nr) {\n-\t\t\tskip_uninteresting(&t1, base, opt);\n-\t\t\tskip_uninteresting(&t2, base, opt);\n+\t\t\tskip_uninteresting(&t, base, opt);\n+\t\t\tfor (i = 0; i < nparent; i++)\n+\t\t\t\tskip_uninteresting(&tp[i], base, opt);\n \t\t}\n-\t\tif (!t1.size && !t2.size)\n-\t\t\tbreak;\n \n-\t\tcmp = tree_entry_pathcmp(&t1, &t2);\n+\t\t/* comparing is finished when all trees are done */\n+\t\tif (!t.size) {\n+\t\t\tint done = 1;\n+\t\t\tfor (i = 0; i < nparent; ++i)\n+\t\t\t\tif (tp[i].size) {\n+\t\t\t\t\tdone = 0;\n+\t\t\t\t\tbreak;\n+\t\t\t\t}\n+\t\t\tif (done)\n+\t\t\t\tbreak;\n+\t\t}\n+\n+\t\t/*\n+\t\t * lookup imin = argmin(x1...xn),\n+\t\t * mark entries whether they =tp[imin] along the way\n+\t\t */\n+\t\timin = 0;\n+\t\ttp[0].entry.mode &= ~S_IFXMIN_NEQ;\n+\n+\t\tfor (i = 1; i < nparent; ++i) {\n+\t\t\tcmp = tree_entry_pathcmp(&tp[i], &tp[imin]);\n+\t\t\tif (cmp < 0) {\n+\t\t\t\timin = i;\n+\t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n+\t\t\t}\n+\t\t\telse if (cmp == 0) {\n+\t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n+\t\t\t}\n+\t\t\telse {\n+\t\t\t\ttp[i].entry.mode |= S_IFXMIN_NEQ;\n+\t\t\t}\n+\t\t}\n+\n+\t\t/* fixup markings for entries before imin */\n+\t\tfor (i = 0; i < imin; ++i)\n+\t\t\ttp[i].entry.mode |= S_IFXMIN_NEQ;\t/* x[i] > x[imin] */\n+\n+\n+\n+\t\t/* compare a vs x[imin] */\n+\t\tcmp = tree_entry_pathcmp(&t, &tp[imin]);\n \n-\t\t/* t1 = t2 */\n+\t\t/* a = xi */\n \t\tif (cmp == 0) {\n-\t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n-\t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n-\t\t\t    (t1.entry.mode != t2.entry.mode))\n-\t\t\t\tshow_path(base, opt, &t1, &t2);\n+\t\t\t/* are either xk > xi or diff(a,xk) != ø ? */\n+\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n+\t\t\t\tfor (i = 0; i < nparent; ++i) {\n+\t\t\t\t\t/* x[i] > x[imin] */\n+\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n+\t\t\t\t\t\tcontinue;\n \n-\t\t\tupdate_tree_entry(&t1);\n-\t\t\tupdate_tree_entry(&t2);\n+\t\t\t\t\t/* diff(a,xk) != ø */\n+\t\t\t\t\tif (hashcmp(t.entry.sha1, tp[i].entry.sha1) ||\n+\t\t\t\t\t    (t.entry.mode != tp[i].entry.mode))\n+\t\t\t\t\t\tcontinue;\n+\n+\t\t\t\t\tgoto skip_emit_t_tp;\n+\t\t\t\t}\n+\t\t\t}\n+\n+\t\t\t/* D += {δ(a,xk) if xk=xi;  \"+a\" if xk > xi} */\n+\t\t\tp = emit_path(p, base, opt, nparent,\n+\t\t\t\t\t&t, tp, imin);\n+\n+\t\tskip_emit_t_tp:\n+\t\t\t/* a↓,  ∀ xk=ximin  xk↓ */\n+\t\t\tupdate_tree_entry(&t);\n+\t\t\tupdate_tp_entries(tp, nparent);\n \t\t}\n \n-\t\t/* t1 < t2 */\n+\t\t/* a < xi */\n \t\telse if (cmp < 0) {\n-\t\t\tshow_path(base, opt, &t1, /*t2=*/NULL);\n-\t\t\tupdate_tree_entry(&t1);\n+\t\t\t/* D += \"+a\" */\n+\t\t\tp = emit_path(p, base, opt, nparent,\n+\t\t\t\t\t&t, /*tp=*/NULL, -1);\n+\n+\t\t\t/* a↓ */\n+\t\t\tupdate_tree_entry(&t);\n \t\t}\n \n-\t\t/* t1 > t2 */\n+\t\t/* a > xi */\n \t\telse {\n-\t\t\tshow_path(base, opt, /*t1=*/NULL, &t2);\n-\t\t\tupdate_tree_entry(&t2);\n+\t\t\t/* ∀j xj=ximin -> D += \"-xi\" */\n+\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n+\t\t\t\tfor (i = 0; i < nparent; ++i)\n+\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n+\t\t\t\t\t\tgoto skip_emit_tp;\n+\t\t\t}\n+\n+\t\t\tp = emit_path(p, base, opt, nparent,\n+\t\t\t\t\t/*t=*/NULL, tp, imin);\n+\n+\t\tskip_emit_tp:\n+\t\t\t/* ∀ xk=ximin  xk↓ */\n+\t\t\tupdate_tp_entries(tp, nparent);\n \t\t}\n \t}\n \n-\tfree(t2tree);\n-\tfree(t1tree);\n-\treturn 0;\n+\tfree(ttree);\n+\tfor (i = nparent-1; i >= 0; i--)\n+\t\tfree(tptree[i]);\n+\txalloca_free(tptree);\n+\txalloca_free(tp);\n+\n+\treturn p;\n+}\n+\n+struct combine_diff_path *diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parents_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt)\n+{\n+\tp = __diff_tree_paths(p, sha1, parents_sha1, nparent, base, opt);\n+\n+\t/*\n+\t * free pre-allocated last element, if any\n+\t * (see __path_appendnew() for details about why)\n+\t */\n+\tif (p->next) {\n+\t\tfree(p->next);\n+\t\tp->next = NULL;\n+\t}\n+\n+\treturn p;\n }\n \n /*\n@@ -300,6 +661,27 @@ static void try_to_follow_renames(const unsigned char *old, const unsigned char\n \tq->nr = 1;\n }\n \n+static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n+\t\t\t    struct strbuf *base, struct diff_options *opt)\n+{\n+\tstruct combine_diff_path phead, *p;\n+\tconst unsigned char *parents_sha1[1] = {old};\n+\tpathchange_fn_t pathchange_old = opt->pathchange;\n+\n+\tphead.next = NULL;\n+\topt->pathchange = emit_diff_first_parent_only;\n+\tdiff_tree_paths(&phead, new, parents_sha1, 1, base, opt);\n+\n+\tfor (p = phead.next; p;) {\n+\t\tstruct combine_diff_path *pprev = p;\n+\t\tp = p->next;\n+\t\tfree(pprev);\n+\t}\n+\n+\topt->pathchange = pathchange_old;\n+\treturn 0;\n+}\n+\n int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base_str, struct diff_options *opt)\n {\n \tstruct strbuf base;\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235258","messageId":"7636e51f1efa3ca7e651e9a51ce5810aa9f5693b.1393257006.git.kirr@mns.spb.ru","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"[PATCH 19/19] combine-diff: speed it up, by using multiparent diff tree-walker directly","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-24T16:21:51Z","receivedAt":"2014-02-24T16:21:51Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"As was recently shown in \"combine-diff: optimize\ncombine_diff_path sets intersection\", combine-diff runs very slowly. In\nthat commit we optimized paths sets intersection, but that accounted\nonly for ~ 25% of the slowness, and as my tracing showed, for linux.git\nv3.10..v3.11, for merges a lot of time is spent computing\ndiff(commit,commit^2) just to only then intersect that huge diff to\nalmost small set of files from diff(commit,commit^1).\n\nIn previous commit, we described the problem in more details, and\nreworked the diff tree-walker to be general one - i.e. to work in\nmultiple parent case too. Now is the time to take advantage of it for\nfinding paths for combine diff.\n\nThe implementation is straightforward - if we know, we can get generated\ndiff paths directly, and at present that means no diff filtering or\nrename/copy detection was requested(*), we can call multiparent tree-walker\ndirectly and get ready paths.\n\n(*) because e.g. at present, all diffcore transformations work on\n    diff_filepair queues, but in the future, that limitation can be\n    lifted, if filters would operate directly on combine_diff_paths.\n\nTimings for `git log --raw --no-abbrev --no-renames` without `-c` (\"git log\")\nand with `-c` (\"git log -c\") and with `-c --merges` (\"git log -c --merges\")\nbefore and after the patch are as follows:\n\n                linux.git v3.10..v3.11\n\n            log     log -c     log -c --merges\n\n    before  1.9s    16.4s      15.2s\n    after   1.9s     2.4s       1.1s\n\nThe result stayed the same.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n( re-posting without change )\n\n combine-diff.c | 88 ++++++++++++++++++++++++++++++++++++++++++++++++++++++----\n diff.c         |  1 +\n 2 files changed, 84 insertions(+), 5 deletions(-)\n\ndiff --git a/combine-diff.c b/combine-diff.c\nindex 1732dfd..12764fb 100644\n--- a/combine-diff.c\n+++ b/combine-diff.c\n@@ -1303,7 +1303,7 @@ static const char *path_path(void *obj)\n \n \n /* find set of paths that every parent touches */\n-static struct combine_diff_path *find_paths(const unsigned char *sha1,\n+static struct combine_diff_path *find_paths_generic(const unsigned char *sha1,\n \tconst struct sha1_array *parents, struct diff_options *opt)\n {\n \tstruct combine_diff_path *paths = NULL;\n@@ -1316,6 +1316,7 @@ static struct combine_diff_path *find_paths(const unsigned char *sha1,\n \t/* tell diff_tree to emit paths in sorted (=tree) order */\n \topt->orderfile = NULL;\n \n+\t/* D(A,P1...Pn) = D(A,P1) ^ ... ^ D(A,Pn)  (wrt paths) */\n \tfor (i = 0; i < num_parent; i++) {\n \t\t/*\n \t\t * show stat against the first parent even when doing\n@@ -1346,6 +1347,35 @@ static struct combine_diff_path *find_paths(const unsigned char *sha1,\n }\n \n \n+/*\n+ * find set of paths that everybody touches, assuming diff is run without\n+ * rename/copy detection, etc, comparing all trees simultaneously (= faster).\n+ */\n+static struct combine_diff_path *find_paths_multitree(\n+\tconst unsigned char *sha1, const struct sha1_array *parents,\n+\tstruct diff_options *opt)\n+{\n+\tint i, nparent = parents->nr;\n+\tconst unsigned char **parents_sha1;\n+\tstruct combine_diff_path paths_head;\n+\tstruct strbuf base;\n+\n+\tparents_sha1 = xmalloc(nparent * sizeof(parents_sha1[0]));\n+\tfor (i = 0; i < nparent; i++)\n+\t\tparents_sha1[i] = parents->sha1[i];\n+\n+\t/* fake list head, so worker can assume it is non-NULL */\n+\tpaths_head.next = NULL;\n+\n+\tstrbuf_init(&base, PATH_MAX);\n+\tdiff_tree_paths(&paths_head, sha1, parents_sha1, nparent, &base, opt);\n+\n+\tstrbuf_release(&base);\n+\tfree(parents_sha1);\n+\treturn paths_head.next;\n+}\n+\n+\n void diff_tree_combined(const unsigned char *sha1,\n \t\t\tconst struct sha1_array *parents,\n \t\t\tint dense,\n@@ -1355,6 +1385,7 @@ void diff_tree_combined(const unsigned char *sha1,\n \tstruct diff_options diffopts;\n \tstruct combine_diff_path *p, *paths;\n \tint i, num_paths, needsep, show_log_first, num_parent = parents->nr;\n+\tint need_generic_pathscan;\n \n \t/* nothing to do, if no parents */\n \tif (!num_parent)\n@@ -1377,11 +1408,58 @@ void diff_tree_combined(const unsigned char *sha1,\n \n \t/* find set of paths that everybody touches\n \t *\n-\t * NOTE find_paths() also handles --stat, as it computes\n-\t * diff(sha1,parent_i) for all i to do the job, specifically\n-\t * for parent0.\n+\t * NOTE\n+\t *\n+\t * Diffcore transformations are bound to diff_filespec and logic\n+\t * comparing two entries - i.e. they do not apply directly to combine\n+\t * diff.\n+\t *\n+\t * If some of such transformations is requested - we launch generic\n+\t * path scanning, which works significantly slower compared to\n+\t * simultaneous all-trees-in-one-go scan in find_paths_multitree().\n+\t *\n+\t * TODO some of the filters could be ported to work on\n+\t * combine_diff_paths - i.e. all functionality that skips paths, so in\n+\t * theory, we could end up having only multitree path scanning.\n+\t *\n+\t * NOTE please keep this semantically in sync with diffcore_std()\n \t */\n-\tpaths = find_paths(sha1, parents, &diffopts);\n+\tneed_generic_pathscan = opt->skip_stat_unmatch\t||\n+\t\t\tDIFF_OPT_TST(opt, FOLLOW_RENAMES)\t||\n+\t\t\topt->break_opt != -1\t||\n+\t\t\topt->detect_rename\t||\n+\t\t\topt->pickaxe\t\t||\n+\t\t\topt->filter;\n+\n+\n+\tif (need_generic_pathscan) {\n+\t\t/*\n+\t\t * NOTE generic case also handles --stat, as it computes\n+\t\t * diff(sha1,parent_i) for all i to do the job, specifically\n+\t\t * for parent0.\n+\t\t */\n+\t\tpaths = find_paths_generic(sha1, parents, &diffopts);\n+\t}\n+\telse {\n+\t\tint stat_opt;\n+\t\tpaths = find_paths_multitree(sha1, parents, &diffopts);\n+\n+\t\t/*\n+\t\t * show stat against the first parent even\n+\t\t * when doing combined diff.\n+\t\t */\n+\t\tstat_opt = (opt->output_format &\n+\t\t\t\t(DIFF_FORMAT_NUMSTAT|DIFF_FORMAT_DIFFSTAT));\n+\t\tif (stat_opt) {\n+\t\t\tdiffopts.output_format = stat_opt;\n+\n+\t\t\tdiff_tree_sha1(parents->sha1[0], sha1, \"\", &diffopts);\n+\t\t\tdiffcore_std(&diffopts);\n+\t\t\tif (opt->orderfile)\n+\t\t\t\tdiffcore_order(opt->orderfile);\n+\t\t\tdiff_flush(&diffopts);\n+\t\t}\n+\t}\n \n \t/* find out number of surviving paths */\n \tfor (num_paths = 0, p = paths; p; p = p->next)\ndiff --git a/diff.c b/diff.c\nindex cda4aa8..f2fff46 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -4764,6 +4764,7 @@ void diffcore_fix_diff_index(struct diff_options *options)\n \n void diffcore_std(struct diff_options *options)\n {\n+\t/* NOTE please keep the following in sync with diff_tree_combined() */\n \tif (options->skip_stat_unmatch)\n \t\tdiffcore_skip_stat_unmatch(options);\n \tif (!options->found_follow) {\n-- \n1.9.rc1.181.g641f458\n"},{"id":"235307","messageId":"CACsJy8BXMVNVAyqPEbHTkGxSSEJ6DpYUVwZqthiMQfO7Tj9T8A@mail.gmail.com","threadId":"35947","inReplyTo":"cover.1393257006.git.kirr@mns.spb.ru","subject":"Re: [PATCH v2 00/19] Multiparent diff tree-walker + combine-diff speedup","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2014-02-24T23:43:24Z","receivedAt":"2014-02-24T23:43:24Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Feb 24, 2014 at 11:21 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> Hello up there.\n>\n> Here go combine-diff speedup patches in form of first reworking diff\n> tree-walker to work in general case - when a commit have several parents, not\n> only one - we are traversing all 1+nparent trees in parallel.\n>\n> Then we are taking advantage of the new diff tree-walker for speeding up\n> combine-diff, which for linux.git results in ~14 times speedup.\n\nI think there is another use case for this n-tree walker (but I'm not\nentirely sure yet as I haven't really read the series). In git-log\n(either with pathspec or --patch) we basically do this\n\ndiff HEAD^ HEAD\ndiff HEAD^^ HEAD^\ndiff HEAD^^^ HEAD^^\ndiff HEAD^^^^ HEAD^^^\n...\n\nso except HEAD (and the last commit), all commits' tree will be\nread/diff'd twice. With n-tree walker I think we may be able to diff\nthem in batch to reduce extra processing: commit lists are split into\n16-commit blocks where 16 trees are fed to the new tree walker at the\nsame time. I hope it would make git-log a bit faster (especially for\n-S). Maybe not much.\n-- \nDuy\n"},{"id":"235312","messageId":"20140225103838.GB3844@tugrik.mns.mnsspb.ru","threadId":"35947","inReplyTo":"CACsJy8BXMVNVAyqPEbHTkGxSSEJ6DpYUVwZqthiMQfO7Tj9T8A@mail.gmail.com","subject":"Re: [PATCH v2 00/19] Multiparent diff tree-walker + combine-diff speedup","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-25T10:38:38Z","receivedAt":"2014-02-25T10:38:38Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Tue, Feb 25, 2014 at 06:43:24AM +0700, Duy Nguyen wrote:\n> On Mon, Feb 24, 2014 at 11:21 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> > Hello up there.\n> >\n> > Here go combine-diff speedup patches in form of first reworking diff\n> > tree-walker to work in general case - when a commit have several parents, not\n> > only one - we are traversing all 1+nparent trees in parallel.\n> >\n> > Then we are taking advantage of the new diff tree-walker for speeding up\n> > combine-diff, which for linux.git results in ~14 times speedup.\n> \n> I think there is another use case for this n-tree walker (but I'm not\n> entirely sure yet as I haven't really read the series). In git-log\n> (either with pathspec or --patch) we basically do this\n> \n> diff HEAD^ HEAD\n> diff HEAD^^ HEAD^\n> diff HEAD^^^ HEAD^^\n> diff HEAD^^^^ HEAD^^^\n> ...\n> \n> so except HEAD (and the last commit), all commits' tree will be\n> read/diff'd twice. With n-tree walker I think we may be able to diff\n> them in batch to reduce extra processing: commit lists are split into\n> 16-commit blocks where 16 trees are fed to the new tree walker at the\n> same time. I hope it would make git-log a bit faster (especially for\n> -S). Maybe not much.\n\nThanks for commenting.\n\nUnfortunately, as it is now, no, and I doubt savings will be\nsignificant. The real speedup comes from the fact that for combined\ndiff, we can omit recursing into subdirectories, if we know some diff\nD(commit,parent_i) is empty. Let me quote myself from\n\nhttp://article.gmane.org/gmane.comp.version-control.git/242217\n\nOn Sun, Feb 16, 2014 at 12:08:29PM +0400, Kirill Smelkov wrote:\n> On Fri, Feb 14, 2014 at 09:37:00AM -0800, Junio C Hamano wrote:\n> > I wonder if this machinery can be reused for \"log -m\" as well (or\n> > perhaps you do that already?).  After all, by performing a single\n> > parallel scan, you are gathering all the necessary information to\n> > let you pretend that you did N pairwise diff-tree.\n> \n> Unfortunately, as it is now, no, and let me explain why:\n> \n> The reason that is not true, is that we omit recursing into directories,\n> if we know D(A,some-parent) for that path is empty. That means we don't\n> calculate D(A,any-other-parents) for that path and subpaths.\n> \n> More structured description is that combined diff and \"log -m\", which\n> could be though as all diffs D(A,Pi) are different things:\n> \n>     - the combined diff is D(A,B) generalization based on \"^\" (sets\n>       intersection) operator, and\n> \n>     - log -m, aka \"all diffs\" is D(A,B) generalization based on \"v\"\n>       (sets union) operator.\n> \n> Intersection means, we can omit calculating parts from other sets, if we\n> know some set does not have an element (remember \"don't recurse into\n> subdirectories\"?), and unioning does not have this property.\n> \n> It does so happen, that \"^\" case (combine-diff) is more interesting,\n> because in the end it allows to see new information - the diff a merge\n> itself introduces. \"log -m\" does not have this property and is no more\n> interesting to what plain diff(HEAD,HEAD^n) can provide - in other words\n> it's just a convenience.\n> \n> Now, the diff tree-walker could be generalized once more, to allow\n> clients specify, which diffs combination operator to use - intersection\n> or unioning, but I doubt that for unioning case that would add\n> significant speedup - we can't reduce any diff generation based on\n> another diff and the only saving is that we traverse resulting commit\n> tree once, but for some cases that could be maybe slower, say if result\n> and some parents don't have a path and some parent does, we'll be\n> recursing into that path and do more work compared to plain D(A,Pi) for\n> Pi that lacks the path.\n> \n> In short: it could be generalized more, if needed, but I propose we\n> first establish the ground with generalizing to just combine-diff.\n\nbesides\n\n    D(HEAD~,  HEAD)\n    D(HEAD~2, HEAD~)\n    ...\n    D(HEAD~{n}, HEAD~{n-1})\n\nis different even from \"log -m\" case as now there is no single commit\nwith several parents.\n\nOn a related note, while developing this n-tree walker, I've learned\nthat it is important to load trees in correct order. Quoting patch 18:\n\n-       t1tree = fill_tree_descriptor(&t1, old);\n-       t2tree = fill_tree_descriptor(&t2, new);\n+       /*\n+        * load parents first, as they are probably already cached.\n+        *\n+        * ( log_tree_diff() parses commit->parent before calling here via\n+        *   diff_tree_sha1(parent, commit) )\n+        */\n+       for (i = 0; i < nparent; ++i)\n+               tptree[i] = fill_tree_descriptor(&tp[i], parents_sha1[i]);\n+       ttree = fill_tree_descriptor(&t, sha1);\n\nso it loads parent's tree first. If we change this to be the other way,\ni.e. load commit's tree first, and then parent's tree, there will be up\nto 4% slowdown for whole plain `git log` (without -c).\n\nSo maybe what could be done to speedup plain log is for diff tree-walker\nto populate some form of recently-loaded trees while walking, and drop\ntrees from will not-be used anymore commits - e.g. after doing\nHEAD~..HEAD for next diff for HEAD~~..HEAD~ HEAD~ trees will be there\nand HEAD trees should be dropped (many of them coincides, so reference\ncounting could help).\n\nThis way, we'll leave the walker logic intact, and only there will be\nmore handy trees-pool supported by fill_tree_descriptor, which imho is a\nbetter design. And this way we could indeed add some not-big, but\nnoticeable speedup.\n\nThough it is another topic, and at present my time is limited to only go\nwith combined diff speedup.\n\nThanks,\nKirill\n"},{"id":"235596","messageId":"878usvo64s.fsf@kepler.schwinge.homeip.net","threadId":"35947","inReplyTo":"f08867ee212e27074dbb4cbb06af408b16dba0a1.1393257006.git.kirr@mns.spb.ru","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Thomas Schwinge","fromEmail":"thomas@codesourcery.com","sentAt":"2014-02-28T10:58:59Z","receivedAt":"2014-02-28T10:58:59Z","isPatch":true,"sender":{"key":"thomas@codesourcery.com","avatar":null},"body":"Hi!\n\nOn Mon, 24 Feb 2014 20:21:49 +0400, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> Both autoconf and config.mak.uname configurations were updated. For\n> autoconf, we are not bothering considering cases, when no alloca.h is\n> available, but alloca() works some other way - its simply alloca.h is\n> available and works or not, everything else is deep legacy.\n\nSounds good for GNU Hurd, or any system using glibc (but have not\nexplicitly tested your patch).\n\n> For config.mak.uname, I've tried to make my almost-sure guess for where\n> alloca() is available, but since I only have access to Linux it is the\n> only change I can be sure about myself, with relevant to other changed\n> systems people Cc'ed.\n\n> diff --git a/config.mak.uname b/config.mak.uname\n> index 7d31fad..71602ee 100644\n> --- a/config.mak.uname\n> +++ b/config.mak.uname\n> @@ -239,6 +243,7 @@ ifeq ($(uname_S),AIX)\n>  endif\n>  ifeq ($(uname_S),GNU)\n>  \t# GNU/Hurd\n> +\tHAVE_ALLOCA_H = YesPlease\n>  \tNO_STRLCPY = YesPlease\n>  \tNO_MKSTEMPS = YesPlease\n>  \tHAVE_PATHS_H = YesPlease\n\nAcked-by: Thomas Schwinge <thomas@codesourcery.com> (GNU Hurd changes)\n\n\nGrüße,\n Thomas\n"},{"id":"235607","messageId":"CABPQNSaVQuXBEnSrs6hdHwEbaBKFr-NjKpuBRNnbkM+HtfJ4Ag@mail.gmail.com","threadId":"35947","inReplyTo":"f08867ee212e27074dbb4cbb06af408b16dba0a1.1393257006.git.kirr@mns.spb.ru","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@gmail.com","sentAt":"2014-02-28T13:44:54Z","receivedAt":"2014-02-28T13:44:54Z","isPatch":true,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Mon, Feb 24, 2014 at 5:21 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> diff --git a/Makefile b/Makefile\n> index dddaf4f..0334806 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -316,6 +321,7 @@ endif\n>  ifeq ($(uname_S),Windows)\n>         GIT_VERSION := $(GIT_VERSION).MSVC\n>         pathsep = ;\n> +       HAVE_ALLOCA_H = YesPlease\n>         NO_PREAD = YesPlease\n>         NEEDS_CRYPTO_WITH_SSL = YesPlease\n>         NO_LIBGEN_H = YesPlease\n\nIn MSVC, alloca is defined in in malloc.h, not alloca.h:\n\nhttp://msdn.microsoft.com/en-us/library/wb1s57t5.aspx\n\nIn fact, it has no alloca.h at all. But we don't have malloca.h in\nmingw either, so creating a compat/win32/alloca.h that includes\nmalloc.h is probably sufficient.\n"},{"id":"235608","messageId":"CABPQNSadTGfiue6G+6x7_o10Ri1E7D5vZFU=Cp8rAha+j9jwSA@mail.gmail.com","threadId":"35947","inReplyTo":"CABPQNSaVQuXBEnSrs6hdHwEbaBKFr-NjKpuBRNnbkM+HtfJ4Ag@mail.gmail.com","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@gmail.com","sentAt":"2014-02-28T13:50:04Z","receivedAt":"2014-02-28T13:50:04Z","isPatch":true,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Fri, Feb 28, 2014 at 2:44 PM, Erik Faye-Lund <kusmabite@gmail.com> wrote:\n> On Mon, Feb 24, 2014 at 5:21 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n>> diff --git a/Makefile b/Makefile\n>> index dddaf4f..0334806 100644\n>> --- a/Makefile\n>> +++ b/Makefile\n>> @@ -316,6 +321,7 @@ endif\n>>  ifeq ($(uname_S),Windows)\n>>         GIT_VERSION := $(GIT_VERSION).MSVC\n>>         pathsep = ;\n>> +       HAVE_ALLOCA_H = YesPlease\n>>         NO_PREAD = YesPlease\n>>         NEEDS_CRYPTO_WITH_SSL = YesPlease\n>>         NO_LIBGEN_H = YesPlease\n>\n> In MSVC, alloca is defined in in malloc.h, not alloca.h:\n>\n> http://msdn.microsoft.com/en-us/library/wb1s57t5.aspx\n>\n> In fact, it has no alloca.h at all. But we don't have malloca.h in\n> mingw either, so creating a compat/win32/alloca.h that includes\n> malloc.h is probably sufficient.\n\n\"But we don't have alloca.h in mingw either\", sorry.\n"},{"id":"235638","messageId":"20140228170012.GA5247@tugrik.mns.mnsspb.ru","threadId":"35947","inReplyTo":"CABPQNSadTGfiue6G+6x7_o10Ri1E7D5vZFU=Cp8rAha+j9jwSA@mail.gmail.com","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-02-28T17:00:12Z","receivedAt":"2014-02-28T17:00:12Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Fri, Feb 28, 2014 at 02:50:04PM +0100, Erik Faye-Lund wrote:\n> On Fri, Feb 28, 2014 at 2:44 PM, Erik Faye-Lund <kusmabite@gmail.com> wrote:\n> > On Mon, Feb 24, 2014 at 5:21 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> >> diff --git a/Makefile b/Makefile\n> >> index dddaf4f..0334806 100644\n> >> --- a/Makefile\n> >> +++ b/Makefile\n> >> @@ -316,6 +321,7 @@ endif\n> >>  ifeq ($(uname_S),Windows)\n> >>         GIT_VERSION := $(GIT_VERSION).MSVC\n> >>         pathsep = ;\n> >> +       HAVE_ALLOCA_H = YesPlease\n> >>         NO_PREAD = YesPlease\n> >>         NEEDS_CRYPTO_WITH_SSL = YesPlease\n> >>         NO_LIBGEN_H = YesPlease\n> >\n> > In MSVC, alloca is defined in in malloc.h, not alloca.h:\n> >\n> > http://msdn.microsoft.com/en-us/library/wb1s57t5.aspx\n> >\n> > In fact, it has no alloca.h at all. But we don't have malloca.h in\n> > mingw either, so creating a compat/win32/alloca.h that includes\n> > malloc.h is probably sufficient.\n> \n> \"But we don't have alloca.h in mingw either\", sorry.\n\nDon't we have that for MSVC already in\n\n    compat/vcbuild/include/alloca.h\n\nand\n\n    ifeq ($(uname_S),Windows)\n        ...\n        BASIC_CFLAGS = ... -Icompat/vcbuild/include ...\n\n\nin config.mak.uname ?\n\nAnd as I've not touched MINGW part in config.mak.uname the patch stays\nvalid as it is :) and we can incrementally update what platforms have\nworking alloca with follow-up patches.\n\nIn fact that would be maybe preferred, for maintainers to enable alloca\nwith knowledge and testing, as one person can't have them all at hand.\n\nThanks,\nKirill\n"},{"id":"235641","messageId":"CABPQNSYnDjhxjpyZQkNP_qwect_tnPvJ_nEfGSq9qnYFMpehWg@mail.gmail.com","threadId":"35947","inReplyTo":"20140228170012.GA5247@tugrik.mns.mnsspb.ru","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@gmail.com","sentAt":"2014-02-28T17:19:58Z","receivedAt":"2014-02-28T17:19:58Z","isPatch":true,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Fri, Feb 28, 2014 at 6:00 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> On Fri, Feb 28, 2014 at 02:50:04PM +0100, Erik Faye-Lund wrote:\n>> On Fri, Feb 28, 2014 at 2:44 PM, Erik Faye-Lund <kusmabite@gmail.com> wrote:\n>> > On Mon, Feb 24, 2014 at 5:21 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n>> >> diff --git a/Makefile b/Makefile\n>> >> index dddaf4f..0334806 100644\n>> >> --- a/Makefile\n>> >> +++ b/Makefile\n>> >> @@ -316,6 +321,7 @@ endif\n>> >>  ifeq ($(uname_S),Windows)\n>> >>         GIT_VERSION := $(GIT_VERSION).MSVC\n>> >>         pathsep = ;\n>> >> +       HAVE_ALLOCA_H = YesPlease\n>> >>         NO_PREAD = YesPlease\n>> >>         NEEDS_CRYPTO_WITH_SSL = YesPlease\n>> >>         NO_LIBGEN_H = YesPlease\n>> >\n>> > In MSVC, alloca is defined in in malloc.h, not alloca.h:\n>> >\n>> > http://msdn.microsoft.com/en-us/library/wb1s57t5.aspx\n>> >\n>> > In fact, it has no alloca.h at all. But we don't have malloca.h in\n>> > mingw either, so creating a compat/win32/alloca.h that includes\n>> > malloc.h is probably sufficient.\n>>\n>> \"But we don't have alloca.h in mingw either\", sorry.\n>\n> Don't we have that for MSVC already in\n>\n>     compat/vcbuild/include/alloca.h\n>\n> and\n>\n>     ifeq ($(uname_S),Windows)\n>         ...\n>         BASIC_CFLAGS = ... -Icompat/vcbuild/include ...\n>\n>\n> in config.mak.uname ?\n\nAh, of course. Thanks for setting me straight!\n\n> And as I've not touched MINGW part in config.mak.uname the patch stays\n> valid as it is :) and we can incrementally update what platforms have\n> working alloca with follow-up patches.\n>\n> In fact that would be maybe preferred, for maintainers to enable alloca\n> with knowledge and testing, as one person can't have them all at hand.\n\nYeah, you're probably right.\n"},{"id":"236079","messageId":"20140305093151.GA3994@tugrik.mns.mnsspb.ru","threadId":"35947","inReplyTo":"CABPQNSYnDjhxjpyZQkNP_qwect_tnPvJ_nEfGSq9qnYFMpehWg@mail.gmail.com","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-03-05T09:31:51Z","receivedAt":"2014-03-05T09:31:51Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Fri, Feb 28, 2014 at 06:19:58PM +0100, Erik Faye-Lund wrote:\n> On Fri, Feb 28, 2014 at 6:00 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> > On Fri, Feb 28, 2014 at 02:50:04PM +0100, Erik Faye-Lund wrote:\n> >> On Fri, Feb 28, 2014 at 2:44 PM, Erik Faye-Lund <kusmabite@gmail.com> wrote:\n> >> > On Mon, Feb 24, 2014 at 5:21 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> >> >> diff --git a/Makefile b/Makefile\n> >> >> index dddaf4f..0334806 100644\n> >> >> --- a/Makefile\n> >> >> +++ b/Makefile\n> >> >> @@ -316,6 +321,7 @@ endif\n> >> >>  ifeq ($(uname_S),Windows)\n> >> >>         GIT_VERSION := $(GIT_VERSION).MSVC\n> >> >>         pathsep = ;\n> >> >> +       HAVE_ALLOCA_H = YesPlease\n> >> >>         NO_PREAD = YesPlease\n> >> >>         NEEDS_CRYPTO_WITH_SSL = YesPlease\n> >> >>         NO_LIBGEN_H = YesPlease\n> >> >\n> >> > In MSVC, alloca is defined in in malloc.h, not alloca.h:\n> >> >\n> >> > http://msdn.microsoft.com/en-us/library/wb1s57t5.aspx\n> >> >\n> >> > In fact, it has no alloca.h at all. But we don't have malloca.h in\n> >> > mingw either, so creating a compat/win32/alloca.h that includes\n> >> > malloc.h is probably sufficient.\n> >>\n> >> \"But we don't have alloca.h in mingw either\", sorry.\n> >\n> > Don't we have that for MSVC already in\n> >\n> >     compat/vcbuild/include/alloca.h\n> >\n> > and\n> >\n> >     ifeq ($(uname_S),Windows)\n> >         ...\n> >         BASIC_CFLAGS = ... -Icompat/vcbuild/include ...\n> >\n> >\n> > in config.mak.uname ?\n> \n> Ah, of course. Thanks for setting me straight!\n> \n> > And as I've not touched MINGW part in config.mak.uname the patch stays\n> > valid as it is :) and we can incrementally update what platforms have\n> > working alloca with follow-up patches.\n> >\n> > In fact that would be maybe preferred, for maintainers to enable alloca\n> > with knowledge and testing, as one person can't have them all at hand.\n> \n> Yeah, you're probably right.\n\nErik, the patch has been merged into pu today. Would you please\nfollow-up with tested MINGW change?\n\nThanks beforehand,\nKirill\n"},{"id":"237492","messageId":"xmqqior3pa7h.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"dad40b2cf785e5951c105cac936d86a7bc6db8a3.1393257006.git.kirr@mns.spb.ru","subject":"Re: [PATCH 12/19] tree-diff: remove special-case diff-emitting code for empty-tree cases","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-24T21:18:10Z","receivedAt":"2014-03-24T21:18:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@mns.spb.ru> writes:\n\n> via teaching tree_entry_pathcmp() how to compare empty tree descriptors:\n\nDrop this line, as you explain the \"pretend empty compares bigger\nthan anything else\" idea later anyway?  This early part of the\nproposed log message made me hiccup while reading it.\n\n> While walking trees, we iterate their entries from lowest to highest in\n> sort order, so empty tree means all entries were already went over.\n>\n> If we artificially assign +infinity value to such tree \"entry\", it will\n> go after all usual entries, and through the usual driver loop we will be\n> taking the same actions, which were hand-coded for special cases, i.e.\n>\n>     t1 empty, t2 non-empty\n>         pathcmp(+∞, t2) -> +1\n>         show_path(/*t1=*/NULL, t2);     /* = t1 > t2 case in main loop */\n>\n>     t1 non-empty, t2-empty\n>         pathcmp(t1, +∞) -> -1\n>         show_path(t1, /*t2=*/NULL);     /* = t1 < t2 case in main loop */\n\nSounds good.  I would have phrased a bit differently, though:\n\n    When we have T1 and T2, we return a sign that tells the caller\n    to indicate the \"earlier\" one to be emitted, and by returning\n    the sign that causes the non-empty side to be emitted, we will\n    automatically cause the entries from the remaining side to be\n    emitted, without attempting to touch the empty side at all.  We\n    can teach tree_entry_pathcmp() to pretend that an empty tree has\n    an element that sorts after anything else to achieve this.\n\nwithout saying \"infinity\".\n\n> Right now we never go to when compared tree descriptors are infinity,...\n\nSorry, but I cannot parse this.\n\n> as\n> this condition is checked in the loop beginning as finishing criteria,\n\nWhat condition and which loop?  The loop that immediately surrounds\nthe callsite of tree_entry_pathcmp() is the infinite \"for (;;) {\" loop,\nand after it prepares t1 and t2 by skipping paths outside pathspec,\nwe check if both are empty (i.e. we ran out).  Is that the condition\nyou are referring to?\n\n> but will do in the future, when there will be several parents iterated\n> simultaneously, and some pair of them would run to the end.\n>\n> Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>\n> ( re-posting without change )\n>\n>  tree-diff.c | 21 +++++++++------------\n>  1 file changed, 9 insertions(+), 12 deletions(-)\n>\n> diff --git a/tree-diff.c b/tree-diff.c\n> index cf96ad7..2fd6d0e 100644\n> --- a/tree-diff.c\n> +++ b/tree-diff.c\n> @@ -12,12 +12,19 @@\n>   *\n>   * NOTE files and directories *always* compare differently, even when having\n>   *      the same name - thanks to base_name_compare().\n> + *\n> + * NOTE empty (=invalid) descriptor(s) take part in comparison as +infty.\n\nThe basic idea is very sane.  It is a nice (and obvious---once you\nare told about the trick) and clean restructuring of the code.\n\n>   */\n>  static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n>  {\n>  \tstruct name_entry *e1, *e2;\n>  \tint cmp;\n>  \n> +\tif (!t1->size)\n> +\t\treturn t2->size ? +1 /* +∞ > c */  : 0 /* +∞ = +∞ */;\n> +\telse if (!t2->size)\n> +\t\treturn -1;\t/* c < +∞ */\n\nWhere do these \"c\" come from?  I somehow feel that these comments\nare making it harder to understand what is going on.\n\n>  \te1 = &t1->entry;\n>  \te2 = &t2->entry;\n>  \tcmp = base_name_compare(e1->path, tree_entry_len(e1), e1->mode,\n> @@ -151,18 +158,8 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n>  \t\t\tskip_uninteresting(t1, &base, opt);\n>  \t\t\tskip_uninteresting(t2, &base, opt);\n>  \t\t}\n> -\t\tif (!t1->size) {\n> -\t\t\tif (!t2->size)\n> -\t\t\t\tbreak;\n> -\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n> -\t\t\tupdate_tree_entry(t2);\n> -\t\t\tcontinue;\n> -\t\t}\n> -\t\tif (!t2->size) {\n> -\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n> -\t\t\tupdate_tree_entry(t1);\n> -\t\t\tcontinue;\n> -\t\t}\n> +\t\tif (!t1->size && !t2->size)\n> +\t\t\tbreak;\n>  \n>  \t\tcmp = tree_entry_pathcmp(t1, t2);\n"},{"id":"237496","messageId":"xmqqeh1rp9vz.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"54aeccfe65926ff00147c3045c5bbae1583d68a7.1393257006.git.kirr@mns.spb.ru","subject":"Re: [PATCH 11/19] tree-diff: simplify tree_entry_pathcmp","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-24T21:25:04Z","receivedAt":"2014-03-24T21:25:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@mns.spb.ru> writes:\n\n> Since an earlier \"Finally switch over tree descriptors to contain a\n> pre-parsed entry\", we can safely access all tree_desc->entry fields\n> directly instead of first \"extracting\" them through\n> tree_entry_extract.\n>\n> Use it. The code generated stays the same - only it now visually looks\n> cleaner.\n>\n> Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>\n> ( re-posting without change )\n\nThanks.\n\nHopefully I'll be merging the series up to this point to 'next'\nsoonish.\n\n>\n>  tree-diff.c | 17 ++++++-----------\n>  1 file changed, 6 insertions(+), 11 deletions(-)\n>\n> diff --git a/tree-diff.c b/tree-diff.c\n> index 20a4fda..cf96ad7 100644\n> --- a/tree-diff.c\n> +++ b/tree-diff.c\n> @@ -15,18 +15,13 @@\n>   */\n>  static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n>  {\n> -\tunsigned mode1, mode2;\n> -\tconst char *path1, *path2;\n> -\tconst unsigned char *sha1, *sha2;\n> -\tint cmp, pathlen1, pathlen2;\n> +\tstruct name_entry *e1, *e2;\n> +\tint cmp;\n>  \n> -\tsha1 = tree_entry_extract(t1, &path1, &mode1);\n> -\tsha2 = tree_entry_extract(t2, &path2, &mode2);\n> -\n> -\tpathlen1 = tree_entry_len(&t1->entry);\n> -\tpathlen2 = tree_entry_len(&t2->entry);\n> -\n> -\tcmp = base_name_compare(path1, pathlen1, mode1, path2, pathlen2, mode2);\n> +\te1 = &t1->entry;\n> +\te2 = &t2->entry;\n> +\tcmp = base_name_compare(e1->path, tree_entry_len(e1), e1->mode,\n> +\t\t\t\te2->path, tree_entry_len(e2), e2->mode);\n>  \treturn cmp;\n>  }\n"},{"id":"237499","messageId":"xmqqa9cfp9d5.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"0b82e2de0edee4a590e7b4165c65938aef7090f5.1393257006.git.kirr@mns.spb.ru","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-24T21:36:22Z","receivedAt":"2014-03-24T21:36:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@mns.spb.ru> writes:\n\n> The downside is that try_to_follow_renames(), if active, we cause\n> re-reading of 2 initial trees, which was negligible based on my timings,\n\nThat would depend on how often the codepath triggered in your test\ncase, but is totally understandable.  It fires only when the path we\nhave been following disappears from the parent, and the processing\nof try-to-follow itself is very compute-intensive (it needs to run\nfind-copies-harder logic) that will end up reading many subtrees of\nthe two initial trees; two more reading of tree objects will be\ndwarfed by the actual processing.\n\n> and which is outweighed cogently by the upsides.\n\n> Changes since v1:\n>\n>  - don't need to touch diff.h, as diff_tree() became static.\n\nNice.  I wonder if it is an option to let the function keep its name\ndiff_tree() without renaming it to __diff_tree_whatever(), though.\n\n>  tree-diff.c | 60 ++++++++++++++++++++++++++++--------------------------------\n>  1 file changed, 28 insertions(+), 32 deletions(-)\n>\n> diff --git a/tree-diff.c b/tree-diff.c\n> index b99622c..f90acf5 100644\n> --- a/tree-diff.c\n> +++ b/tree-diff.c\n> @@ -137,12 +137,17 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n>  \t}\n>  }\n>  \n> -static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n> -\t\t     const char *base_str, struct diff_options *opt)\n> +static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n> +\t\t\t    const char *base_str, struct diff_options *opt)\n>  {\n> +\tstruct tree_desc t1, t2;\n> +\tvoid *t1tree, *t2tree;\n>  \tstruct strbuf base;\n>  \tint baselen = strlen(base_str);\n>  \n> +\tt1tree = fill_tree_descriptor(&t1, old);\n> +\tt2tree = fill_tree_descriptor(&t2, new);\n> +\n>  \t/* Enable recursion indefinitely */\n>  \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n>  \n> @@ -155,39 +160,41 @@ static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n>  \t\tif (diff_can_quit_early(opt))\n>  \t\t\tbreak;\n>  \t\tif (opt->pathspec.nr) {\n> -\t\t\tskip_uninteresting(t1, &base, opt);\n> -\t\t\tskip_uninteresting(t2, &base, opt);\n> +\t\t\tskip_uninteresting(&t1, &base, opt);\n> +\t\t\tskip_uninteresting(&t2, &base, opt);\n>  \t\t}\n> -\t\tif (!t1->size && !t2->size)\n> +\t\tif (!t1.size && !t2.size)\n>  \t\t\tbreak;\n>  \n> -\t\tcmp = tree_entry_pathcmp(t1, t2);\n> +\t\tcmp = tree_entry_pathcmp(&t1, &t2);\n>  \n>  \t\t/* t1 = t2 */\n>  \t\tif (cmp == 0) {\n>  \t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n> -\t\t\t    hashcmp(t1->entry.sha1, t2->entry.sha1) ||\n> -\t\t\t    (t1->entry.mode != t2->entry.mode))\n> -\t\t\t\tshow_path(&base, opt, t1, t2);\n> +\t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n> +\t\t\t    (t1.entry.mode != t2.entry.mode))\n> +\t\t\t\tshow_path(&base, opt, &t1, &t2);\n>  \n> -\t\t\tupdate_tree_entry(t1);\n> -\t\t\tupdate_tree_entry(t2);\n> +\t\t\tupdate_tree_entry(&t1);\n> +\t\t\tupdate_tree_entry(&t2);\n>  \t\t}\n>  \n>  \t\t/* t1 < t2 */\n>  \t\telse if (cmp < 0) {\n> -\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n> -\t\t\tupdate_tree_entry(t1);\n> +\t\t\tshow_path(&base, opt, &t1, /*t2=*/NULL);\n> +\t\t\tupdate_tree_entry(&t1);\n>  \t\t}\n>  \n>  \t\t/* t1 > t2 */\n>  \t\telse {\n> -\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n> -\t\t\tupdate_tree_entry(t2);\n> +\t\t\tshow_path(&base, opt, /*t1=*/NULL, &t2);\n> +\t\t\tupdate_tree_entry(&t2);\n>  \t\t}\n>  \t}\n>  \n>  \tstrbuf_release(&base);\n> +\tfree(t2tree);\n> +\tfree(t1tree);\n>  \treturn 0;\n>  }\n>  \n> @@ -202,7 +209,7 @@ static inline int diff_might_be_rename(void)\n>  \t\t!DIFF_FILE_VALID(diff_queued_diff.queue[0]->one);\n>  }\n>  \n> -static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, const char *base, struct diff_options *opt)\n> +static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n>  {\n>  \tstruct diff_options diff_opts;\n>  \tstruct diff_queue_struct *q = &diff_queued_diff;\n> @@ -240,7 +247,7 @@ static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, co\n>  \tdiff_opts.break_opt = opt->break_opt;\n>  \tdiff_opts.rename_score = opt->rename_score;\n>  \tdiff_setup_done(&diff_opts);\n> -\tdiff_tree(t1, t2, base, &diff_opts);\n> +\t__diff_tree_sha1(old, new, base, &diff_opts);\n>  \tdiffcore_std(&diff_opts);\n>  \tfree_pathspec(&diff_opts.pathspec);\n>  \n> @@ -301,23 +308,12 @@ static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, co\n>  \n>  int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n>  {\n> -\tvoid *tree1, *tree2;\n> -\tstruct tree_desc t1, t2;\n> -\tunsigned long size1, size2;\n>  \tint retval;\n>  \n> -\ttree1 = fill_tree_descriptor(&t1, old);\n> -\ttree2 = fill_tree_descriptor(&t2, new);\n> -\tsize1 = t1.size;\n> -\tsize2 = t2.size;\n> -\tretval = diff_tree(&t1, &t2, base, opt);\n> -\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename()) {\n> -\t\tinit_tree_desc(&t1, tree1, size1);\n> -\t\tinit_tree_desc(&t2, tree2, size2);\n> -\t\ttry_to_follow_renames(&t1, &t2, base, opt);\n> -\t}\n> -\tfree(tree1);\n> -\tfree(tree2);\n> +\tretval = __diff_tree_sha1(old, new, base, opt);\n> +\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n> +\t\ttry_to_follow_renames(old, new, base, opt);\n> +\n>  \treturn retval;\n>  }\n"},{"id":"237502","messageId":"xmqq61n3p913.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"301eb2377e0c5f670ffc26bda085d14dbee4f431.1393257006.git.kirr@mns.spb.ru","subject":"Re: [PATCH v2 16/19] tree-diff: reuse base str(buf) memory on sub-tree recursion","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-24T21:43:36Z","receivedAt":"2014-03-24T21:43:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@mns.spb.ru> writes:\n\n> instead of allocating it all the time for every subtree in\n> __diff_tree_sha1, let's allocate it once in diff_tree_sha1, and then all\n> callee just use it in stacking style, without memory allocations.\n>\n> This should be faster, and for me this change gives the following\n> slight speedups for\n>\n>     git log --raw --no-abbrev --no-renames --format='%H'\n>\n>                 navy.git    linux.git v3.10..v3.11\n>\n>     before      0.618s      1.903s\n>     after       0.611s      1.889s\n>     speedup     1.1%        0.7%\n>\n> Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n> ---\n>\n> Changes since v1:\n>\n>  - don't need to touch diff.h, as the function we are changing became static.\n>\n>  tree-diff.c | 36 ++++++++++++++++++------------------\n>  1 file changed, 18 insertions(+), 18 deletions(-)\n>\n> diff --git a/tree-diff.c b/tree-diff.c\n> index aea0297..c76821d 100644\n> --- a/tree-diff.c\n> +++ b/tree-diff.c\n> @@ -115,7 +115,7 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n>  \tif (recurse) {\n>  \t\tstrbuf_addch(base, '/');\n>  \t\t__diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n> -\t\t\t\t t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n> +\t\t\t\t t2 ? t2->entry.sha1 : NULL, base, opt);\n>  \t}\n>  \n>  \tstrbuf_setlen(base, old_baselen);\n\nI was scratching my head for a while, after seeing that there does\nnot seem to be any *new* code added by this patch in order to\nstore-away the original length and restore the singleton base buffer\nto the original length after using addch/addstr to extend it.\n\nBut I see that the code has already been prepared to do this\nconversion.  I wonder why we didn't do this earlier ;-)\n\nLooks good.  Thanks.\n\n> @@ -138,12 +138,10 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n>  }\n>  \n>  static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n> -\t\t\t    const char *base_str, struct diff_options *opt)\n> +\t\t\t    struct strbuf *base, struct diff_options *opt)\n>  {\n>  \tstruct tree_desc t1, t2;\n>  \tvoid *t1tree, *t2tree;\n> -\tstruct strbuf base;\n> -\tint baselen = strlen(base_str);\n>  \n>  \tt1tree = fill_tree_descriptor(&t1, old);\n>  \tt2tree = fill_tree_descriptor(&t2, new);\n> @@ -151,17 +149,14 @@ static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n>  \t/* Enable recursion indefinitely */\n>  \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n>  \n> -\tstrbuf_init(&base, PATH_MAX);\n> -\tstrbuf_add(&base, base_str, baselen);\n> -\n>  \tfor (;;) {\n>  \t\tint cmp;\n>  \n>  \t\tif (diff_can_quit_early(opt))\n>  \t\t\tbreak;\n>  \t\tif (opt->pathspec.nr) {\n> -\t\t\tskip_uninteresting(&t1, &base, opt);\n> -\t\t\tskip_uninteresting(&t2, &base, opt);\n> +\t\t\tskip_uninteresting(&t1, base, opt);\n> +\t\t\tskip_uninteresting(&t2, base, opt);\n>  \t\t}\n>  \t\tif (!t1.size && !t2.size)\n>  \t\t\tbreak;\n> @@ -173,7 +168,7 @@ static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n>  \t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n>  \t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n>  \t\t\t    (t1.entry.mode != t2.entry.mode))\n> -\t\t\t\tshow_path(&base, opt, &t1, &t2);\n> +\t\t\t\tshow_path(base, opt, &t1, &t2);\n>  \n>  \t\t\tupdate_tree_entry(&t1);\n>  \t\t\tupdate_tree_entry(&t2);\n> @@ -181,18 +176,17 @@ static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n>  \n>  \t\t/* t1 < t2 */\n>  \t\telse if (cmp < 0) {\n> -\t\t\tshow_path(&base, opt, &t1, /*t2=*/NULL);\n> +\t\t\tshow_path(base, opt, &t1, /*t2=*/NULL);\n>  \t\t\tupdate_tree_entry(&t1);\n>  \t\t}\n>  \n>  \t\t/* t1 > t2 */\n>  \t\telse {\n> -\t\t\tshow_path(&base, opt, /*t1=*/NULL, &t2);\n> +\t\t\tshow_path(base, opt, /*t1=*/NULL, &t2);\n>  \t\t\tupdate_tree_entry(&t2);\n>  \t\t}\n>  \t}\n>  \n> -\tstrbuf_release(&base);\n>  \tfree(t2tree);\n>  \tfree(t1tree);\n>  \treturn 0;\n> @@ -209,7 +203,7 @@ static inline int diff_might_be_rename(void)\n>  \t\t!DIFF_FILE_VALID(diff_queued_diff.queue[0]->one);\n>  }\n>  \n> -static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n> +static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, struct strbuf *base, struct diff_options *opt)\n>  {\n>  \tstruct diff_options diff_opts;\n>  \tstruct diff_queue_struct *q = &diff_queued_diff;\n> @@ -306,13 +300,19 @@ static void try_to_follow_renames(const unsigned char *old, const unsigned char\n>  \tq->nr = 1;\n>  }\n>  \n> -int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n> +int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base_str, struct diff_options *opt)\n>  {\n> +\tstruct strbuf base;\n>  \tint retval;\n>  \n> -\tretval = __diff_tree_sha1(old, new, base, opt);\n> -\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n> -\t\ttry_to_follow_renames(old, new, base, opt);\n> +\tstrbuf_init(&base, PATH_MAX);\n> +\tstrbuf_addstr(&base, base_str);\n> +\n> +\tretval = __diff_tree_sha1(old, new, &base, opt);\n> +\tif (!*base_str && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n> +\t\ttry_to_follow_renames(old, new, &base, opt);\n> +\n> +\tstrbuf_release(&base);\n>  \n>  \treturn retval;\n>  }\n"},{"id":"237505","messageId":"xmqq1txrp8ur.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"20140305093151.GA3994@tugrik.mns.mnsspb.ru","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-24T21:47:24Z","receivedAt":"2014-03-24T21:47:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@mns.spb.ru> writes:\n\n> On Fri, Feb 28, 2014 at 06:19:58PM +0100, Erik Faye-Lund wrote:\n>> On Fri, Feb 28, 2014 at 6:00 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n>> ...\n>> > In fact that would be maybe preferred, for maintainers to enable alloca\n>> > with knowledge and testing, as one person can't have them all at hand.\n>> \n>> Yeah, you're probably right.\n>\n> Erik, the patch has been merged into pu today. Would you please\n> follow-up with tested MINGW change?\n\nSooo.... I lost track but this discussion seems to have petered out\naround here.  I think the copy we have had for a while on 'pu' is\nbasically sound, and can easily built on by platform folks by adding\nor removing the -DHAVE_ALLOCA_H from the Makefile.\n"},{"id":"237640","messageId":"20140325092040.GA3777@mini.zxlink","threadId":"35947","inReplyTo":"xmqqior3pa7h.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH 12/19] tree-diff: remove special-case diff-emitting code for empty-tree cases","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-25T09:20:40Z","receivedAt":"2014-03-25T09:20:40Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Mon, Mar 24, 2014 at 02:18:10PM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@mns.spb.ru> writes:\n> \n> > via teaching tree_entry_pathcmp() how to compare empty tree descriptors:\n> \n> Drop this line, as you explain the \"pretend empty compares bigger\n> than anything else\" idea later anyway?  This early part of the\n> proposed log message made me hiccup while reading it.\n\nHmm, I was trying to show the big picture first and only then details...\n\n\n> > While walking trees, we iterate their entries from lowest to highest in\n> > sort order, so empty tree means all entries were already went over.\n> >\n> > If we artificially assign +infinity value to such tree \"entry\", it will\n> > go after all usual entries, and through the usual driver loop we will be\n> > taking the same actions, which were hand-coded for special cases, i.e.\n> >\n> >     t1 empty, t2 non-empty\n> >         pathcmp(+∞, t2) -> +1\n> >         show_path(/*t1=*/NULL, t2);     /* = t1 > t2 case in main loop */\n> >\n> >     t1 non-empty, t2-empty\n> >         pathcmp(t1, +∞) -> -1\n> >         show_path(t1, /*t2=*/NULL);     /* = t1 < t2 case in main loop */\n> \n> Sounds good.  I would have phrased a bit differently, though:\n> \n>     When we have T1 and T2, we return a sign that tells the caller\n>     to indicate the \"earlier\" one to be emitted, and by returning\n>     the sign that causes the non-empty side to be emitted, we will\n>     automatically cause the entries from the remaining side to be\n>     emitted, without attempting to touch the empty side at all.  We\n>     can teach tree_entry_pathcmp() to pretend that an empty tree has\n>     an element that sorts after anything else to achieve this.\n> \n> without saying \"infinity\".\n\nDoesn't your description, especially \"an element that sorts after\nanything else\" match what \"infinity\" is pretty exactly? :)\n\nI agree it could read more clearly to those new to the concept, but we\nare basically talking about the same thing and once someone is familiar\nwith infinity and its friends the second description imho is less\nobvious.\n\nLet's maybe as a compromise add your text as \"In other words <textual\ndescription ...>\" ?\n\nThis way, it will hopefully be good both ways...\n\n\n> > Right now we never go to when compared tree descriptors are infinity,...\n> \n> Sorry, but I cannot parse this.\n\nSorry, I've omitted one word here. It should read\n\n    \"Right now we never go to when compared tree descriptors are _both_ infinity,...\"\n\ni.e. right now we never call tree_entry_pathcmp with both t1 and t2\nbeing empty.\n\n\n> > as\n> > this condition is checked in the loop beginning as finishing criteria,\n> \n> What condition and which loop?  The loop that immediately surrounds\n> the callsite of tree_entry_pathcmp() is the infinite \"for (;;) {\" loop,\n> and after it prepares t1 and t2 by skipping paths outside pathspec,\n> we check if both are empty (i.e. we ran out).  Is that the condition\n> you are referring to?\n\nYes exactly. Modulo diff_can_quit_early() logic, we break from loop in\ndiff_tree (the loop in which special-case diff-tree emitting code was)\nwhen both trees were scanned to the end, i.e.\n\n        if (!t1->size && !t2->size)\n                break;\n\nin other words when both t1 and t2 are \"+∞\".\n\nBecause of that, at this stage we will never go into tree_entry_pathcmp\nwith (+∞,+∞) arguments, which could mean (!t1->size && !t2->size) case\ncould be unnecessary in tree_entry_pathcmp and should not be coded at\nall...\n\n> > but will do in the future, when there will be several parents iterated\n> > simultaneously, and some pair of them would run to the end.\n\n... I was trying to say this case will probably be needed later, and that\nit is better to have it for generality.\n\nI hope this should be more clear once that prologue with \"both\" included\nis not confusing.\n\n\n> > Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n> > Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> > ---\n> >\n> > ( re-posting without change )\n> >\n> >  tree-diff.c | 21 +++++++++------------\n> >  1 file changed, 9 insertions(+), 12 deletions(-)\n> >\n> > diff --git a/tree-diff.c b/tree-diff.c\n> > index cf96ad7..2fd6d0e 100644\n> > --- a/tree-diff.c\n> > +++ b/tree-diff.c\n> > @@ -12,12 +12,19 @@\n> >   *\n> >   * NOTE files and directories *always* compare differently, even when having\n> >   *      the same name - thanks to base_name_compare().\n> > + *\n> > + * NOTE empty (=invalid) descriptor(s) take part in comparison as +infty.\n> \n> The basic idea is very sane.  It is a nice (and obvious---once you\n> are told about the trick) and clean restructuring of the code.\n\nThanks. I was surprised it is seen as a trick, as infinity is very handy\nand common concept in many areas and in sorting too.\n\n> \n> >   */\n> >  static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n> >  {\n> >  \tstruct name_entry *e1, *e2;\n> >  \tint cmp;\n> >  \n> > +\tif (!t1->size)\n> > +\t\treturn t2->size ? +1 /* +∞ > c */  : 0 /* +∞ = +∞ */;\n> > +\telse if (!t2->size)\n> > +\t\treturn -1;\t/* c < +∞ */\n> \n> Where do these \"c\" come from?  I somehow feel that these comments\n> are making it harder to understand what is going on.\n\n\"c\" means some finite \"c\"onstant here. When I was studying at school and\nat the university, it was common to denote constants via this letter -\ni.e. in algebra and operators they often show scalar multiplication as\n\n    c·A     (or α·A)\n\netc. I understand it could maybe be confusing (but it came to me as\nsurprise), so would the following be maybe better:\n\n        if (!t1->size)\n        \treturn t2->size ? +1 /* +∞ > const */  : 0 /* +∞ = +∞ */;\n        else if (!t2->size)\n        \treturn -1;\t/* const < +∞ */\n\n?\n\n\nThanks,\nKirill\n\n> >  \te1 = &t1->entry;\n> >  \te2 = &t2->entry;\n> >  \tcmp = base_name_compare(e1->path, tree_entry_len(e1), e1->mode,\n> > @@ -151,18 +158,8 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n> >  \t\t\tskip_uninteresting(t1, &base, opt);\n> >  \t\t\tskip_uninteresting(t2, &base, opt);\n> >  \t\t}\n> > -\t\tif (!t1->size) {\n> > -\t\t\tif (!t2->size)\n> > -\t\t\t\tbreak;\n> > -\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n> > -\t\t\tupdate_tree_entry(t2);\n> > -\t\t\tcontinue;\n> > -\t\t}\n> > -\t\tif (!t2->size) {\n> > -\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n> > -\t\t\tupdate_tree_entry(t1);\n> > -\t\t\tcontinue;\n> > -\t\t}\n> > +\t\tif (!t1->size && !t2->size)\n> > +\t\t\tbreak;\n> >  \n> >  \t\tcmp = tree_entry_pathcmp(t1, t2);\n"},{"id":"237641","messageId":"20140325092215.GB3777@mini.zxlink","threadId":"35947","inReplyTo":"xmqqa9cfp9d5.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-25T09:22:15Z","receivedAt":"2014-03-25T09:22:15Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Mon, Mar 24, 2014 at 02:36:22PM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@mns.spb.ru> writes:\n> \n> > The downside is that try_to_follow_renames(), if active, we cause\n> > re-reading of 2 initial trees, which was negligible based on my timings,\n> \n> That would depend on how often the codepath triggered in your test\n> case, but is totally understandable.  It fires only when the path we\n> have been following disappears from the parent, and the processing\n> of try-to-follow itself is very compute-intensive (it needs to run\n> find-copies-harder logic) that will end up reading many subtrees of\n> the two initial trees; two more reading of tree objects will be\n> dwarfed by the actual processing.\n\nI agree and thanks for the explanation.\n\n> > and which is outweighed cogently by the upsides.\n> \n> > Changes since v1:\n> >\n> >  - don't need to touch diff.h, as diff_tree() became static.\n> \n> Nice.  I wonder if it is an option to let the function keep its name\n> diff_tree() without renaming it to __diff_tree_whatever(), though.\n\nAs I see it, in Git for functions operating on trees, there is convention\nto accept either `struct tree_desc *` and be named simply, or sha1 and\nbe named with _sha1 suffix. From this point of view for new diff_tree()\naccepting sha1's and staying with its old name would be confusing.\n\nBesides, in the end we'll have two function with high-level wrapper, and\nlower-lever worker:\n\n    - diff_tree_sha1(), and\n    - diff_tree_paths().\n\nSo it's not about this only particular case.  Both do some simple\npreparation, call worker, and perform some cleanup.\n\nSo the question is how to name the worker?\n\nIn Linux they use \"__\" prefix. We could also use some other prefix or\nsuffix, e.g. \"_bh\" (for bottom-half), \"_worker\", \"_low\", \"_raw\", etc...\n\n\nTo me, personally, the cleanest is \"__\" prefix, but maybe I'm too used\nto Linux etc... I'm open to other naming scheme, only it should be\nconsistent.\n\nWhat are the downsides of \"__\" prefix by the way?\n\n\n> >  tree-diff.c | 60 ++++++++++++++++++++++++++++--------------------------------\n> >  1 file changed, 28 insertions(+), 32 deletions(-)\n> >\n> > diff --git a/tree-diff.c b/tree-diff.c\n> > index b99622c..f90acf5 100644\n> > --- a/tree-diff.c\n> > +++ b/tree-diff.c\n> > @@ -137,12 +137,17 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n> >  \t}\n> >  }\n> >  \n> > -static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n> > -\t\t     const char *base_str, struct diff_options *opt)\n> > +static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n> > +\t\t\t    const char *base_str, struct diff_options *opt)\n> >  {\n> > +\tstruct tree_desc t1, t2;\n> > +\tvoid *t1tree, *t2tree;\n> >  \tstruct strbuf base;\n> >  \tint baselen = strlen(base_str);\n> >  \n> > +\tt1tree = fill_tree_descriptor(&t1, old);\n> > +\tt2tree = fill_tree_descriptor(&t2, new);\n> > +\n> >  \t/* Enable recursion indefinitely */\n> >  \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n> >  \n> > @@ -155,39 +160,41 @@ static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n> >  \t\tif (diff_can_quit_early(opt))\n> >  \t\t\tbreak;\n> >  \t\tif (opt->pathspec.nr) {\n> > -\t\t\tskip_uninteresting(t1, &base, opt);\n> > -\t\t\tskip_uninteresting(t2, &base, opt);\n> > +\t\t\tskip_uninteresting(&t1, &base, opt);\n> > +\t\t\tskip_uninteresting(&t2, &base, opt);\n> >  \t\t}\n> > -\t\tif (!t1->size && !t2->size)\n> > +\t\tif (!t1.size && !t2.size)\n> >  \t\t\tbreak;\n> >  \n> > -\t\tcmp = tree_entry_pathcmp(t1, t2);\n> > +\t\tcmp = tree_entry_pathcmp(&t1, &t2);\n> >  \n> >  \t\t/* t1 = t2 */\n> >  \t\tif (cmp == 0) {\n> >  \t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n> > -\t\t\t    hashcmp(t1->entry.sha1, t2->entry.sha1) ||\n> > -\t\t\t    (t1->entry.mode != t2->entry.mode))\n> > -\t\t\t\tshow_path(&base, opt, t1, t2);\n> > +\t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n> > +\t\t\t    (t1.entry.mode != t2.entry.mode))\n> > +\t\t\t\tshow_path(&base, opt, &t1, &t2);\n> >  \n> > -\t\t\tupdate_tree_entry(t1);\n> > -\t\t\tupdate_tree_entry(t2);\n> > +\t\t\tupdate_tree_entry(&t1);\n> > +\t\t\tupdate_tree_entry(&t2);\n> >  \t\t}\n> >  \n> >  \t\t/* t1 < t2 */\n> >  \t\telse if (cmp < 0) {\n> > -\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n> > -\t\t\tupdate_tree_entry(t1);\n> > +\t\t\tshow_path(&base, opt, &t1, /*t2=*/NULL);\n> > +\t\t\tupdate_tree_entry(&t1);\n> >  \t\t}\n> >  \n> >  \t\t/* t1 > t2 */\n> >  \t\telse {\n> > -\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n> > -\t\t\tupdate_tree_entry(t2);\n> > +\t\t\tshow_path(&base, opt, /*t1=*/NULL, &t2);\n> > +\t\t\tupdate_tree_entry(&t2);\n> >  \t\t}\n> >  \t}\n> >  \n> >  \tstrbuf_release(&base);\n> > +\tfree(t2tree);\n> > +\tfree(t1tree);\n> >  \treturn 0;\n> >  }\n> >  \n> > @@ -202,7 +209,7 @@ static inline int diff_might_be_rename(void)\n> >  \t\t!DIFF_FILE_VALID(diff_queued_diff.queue[0]->one);\n> >  }\n> >  \n> > -static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, const char *base, struct diff_options *opt)\n> > +static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n> >  {\n> >  \tstruct diff_options diff_opts;\n> >  \tstruct diff_queue_struct *q = &diff_queued_diff;\n> > @@ -240,7 +247,7 @@ static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, co\n> >  \tdiff_opts.break_opt = opt->break_opt;\n> >  \tdiff_opts.rename_score = opt->rename_score;\n> >  \tdiff_setup_done(&diff_opts);\n> > -\tdiff_tree(t1, t2, base, &diff_opts);\n> > +\t__diff_tree_sha1(old, new, base, &diff_opts);\n> >  \tdiffcore_std(&diff_opts);\n> >  \tfree_pathspec(&diff_opts.pathspec);\n> >  \n> > @@ -301,23 +308,12 @@ static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, co\n> >  \n> >  int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n> >  {\n> > -\tvoid *tree1, *tree2;\n> > -\tstruct tree_desc t1, t2;\n> > -\tunsigned long size1, size2;\n> >  \tint retval;\n> >  \n> > -\ttree1 = fill_tree_descriptor(&t1, old);\n> > -\ttree2 = fill_tree_descriptor(&t2, new);\n> > -\tsize1 = t1.size;\n> > -\tsize2 = t2.size;\n> > -\tretval = diff_tree(&t1, &t2, base, opt);\n> > -\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename()) {\n> > -\t\tinit_tree_desc(&t1, tree1, size1);\n> > -\t\tinit_tree_desc(&t2, tree2, size2);\n> > -\t\ttry_to_follow_renames(&t1, &t2, base, opt);\n> > -\t}\n> > -\tfree(tree1);\n> > -\tfree(tree2);\n> > +\tretval = __diff_tree_sha1(old, new, base, opt);\n> > +\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n> > +\t\ttry_to_follow_renames(old, new, base, opt);\n> > +\n> >  \treturn retval;\n> >  }\n"},{"id":"237643","messageId":"20140325092320.GC3777@mini.zxlink","threadId":"35947","inReplyTo":"xmqq61n3p913.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 16/19] tree-diff: reuse base str(buf) memory on sub-tree recursion","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-25T09:23:20Z","receivedAt":"2014-03-25T09:23:20Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Mon, Mar 24, 2014 at 02:43:36PM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@mns.spb.ru> writes:\n> \n> > instead of allocating it all the time for every subtree in\n> > __diff_tree_sha1, let's allocate it once in diff_tree_sha1, and then all\n> > callee just use it in stacking style, without memory allocations.\n> >\n> > This should be faster, and for me this change gives the following\n> > slight speedups for\n> >\n> >     git log --raw --no-abbrev --no-renames --format='%H'\n> >\n> >                 navy.git    linux.git v3.10..v3.11\n> >\n> >     before      0.618s      1.903s\n> >     after       0.611s      1.889s\n> >     speedup     1.1%        0.7%\n> >\n> > Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n> > ---\n> >\n> > Changes since v1:\n> >\n> >  - don't need to touch diff.h, as the function we are changing became static.\n> >\n> >  tree-diff.c | 36 ++++++++++++++++++------------------\n> >  1 file changed, 18 insertions(+), 18 deletions(-)\n> >\n> > diff --git a/tree-diff.c b/tree-diff.c\n> > index aea0297..c76821d 100644\n> > --- a/tree-diff.c\n> > +++ b/tree-diff.c\n> > @@ -115,7 +115,7 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n> >  \tif (recurse) {\n> >  \t\tstrbuf_addch(base, '/');\n> >  \t\t__diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n> > -\t\t\t\t t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n> > +\t\t\t\t t2 ? t2->entry.sha1 : NULL, base, opt);\n> >  \t}\n> >  \n> >  \tstrbuf_setlen(base, old_baselen);\n> \n> I was scratching my head for a while, after seeing that there does\n> not seem to be any *new* code added by this patch in order to\n> store-away the original length and restore the singleton base buffer\n> to the original length after using addch/addstr to extend it.\n> \n> But I see that the code has already been prepared to do this\n> conversion.  I wonder why we didn't do this earlier ;-)\n\nThe conversion to reusing memory started in 48932677 \"diff-tree: convert\nbase+baselen to writable strbuf\" which allowed to avoid \"quite a bit of\nmalloc() and memcpy()\", but for this to work allocation at diff_tree()\nentry had to be there.\n\nIn particular it had to be there, because diff_tree() accepted base as C\nstring, not strbuf, and since diff_tree() was calling itself\nrecursively - oops - new allocation on every subtree.\n\nI've opened the door for avoiding allocations via splitting diff_tree\ninto high-level and low-level parts. The high-level part still accepts\n`char *base`, but low-level function operates on strbuf and recurses\ninto low-level self.\n\nThe high-level diff_tree_sha1() still allocates memory for every\ndiff(tree1,tree2), but that is significantly lower compared to\nallocating memory on every subtree...\n\nThe lesson here is: better use strbuf for api unless there is a reason\nnot to.\n\n\n> Looks good.  Thanks.\n\nThanks.\n\n> > @@ -138,12 +138,10 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n> >  }\n> >  \n> >  static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n> > -\t\t\t    const char *base_str, struct diff_options *opt)\n> > +\t\t\t    struct strbuf *base, struct diff_options *opt)\n> >  {\n> >  \tstruct tree_desc t1, t2;\n> >  \tvoid *t1tree, *t2tree;\n> > -\tstruct strbuf base;\n> > -\tint baselen = strlen(base_str);\n> >  \n> >  \tt1tree = fill_tree_descriptor(&t1, old);\n> >  \tt2tree = fill_tree_descriptor(&t2, new);\n> > @@ -151,17 +149,14 @@ static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n> >  \t/* Enable recursion indefinitely */\n> >  \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n> >  \n> > -\tstrbuf_init(&base, PATH_MAX);\n> > -\tstrbuf_add(&base, base_str, baselen);\n> > -\n> >  \tfor (;;) {\n> >  \t\tint cmp;\n> >  \n> >  \t\tif (diff_can_quit_early(opt))\n> >  \t\t\tbreak;\n> >  \t\tif (opt->pathspec.nr) {\n> > -\t\t\tskip_uninteresting(&t1, &base, opt);\n> > -\t\t\tskip_uninteresting(&t2, &base, opt);\n> > +\t\t\tskip_uninteresting(&t1, base, opt);\n> > +\t\t\tskip_uninteresting(&t2, base, opt);\n> >  \t\t}\n> >  \t\tif (!t1.size && !t2.size)\n> >  \t\t\tbreak;\n> > @@ -173,7 +168,7 @@ static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n> >  \t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n> >  \t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n> >  \t\t\t    (t1.entry.mode != t2.entry.mode))\n> > -\t\t\t\tshow_path(&base, opt, &t1, &t2);\n> > +\t\t\t\tshow_path(base, opt, &t1, &t2);\n> >  \n> >  \t\t\tupdate_tree_entry(&t1);\n> >  \t\t\tupdate_tree_entry(&t2);\n> > @@ -181,18 +176,17 @@ static int __diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n> >  \n> >  \t\t/* t1 < t2 */\n> >  \t\telse if (cmp < 0) {\n> > -\t\t\tshow_path(&base, opt, &t1, /*t2=*/NULL);\n> > +\t\t\tshow_path(base, opt, &t1, /*t2=*/NULL);\n> >  \t\t\tupdate_tree_entry(&t1);\n> >  \t\t}\n> >  \n> >  \t\t/* t1 > t2 */\n> >  \t\telse {\n> > -\t\t\tshow_path(&base, opt, /*t1=*/NULL, &t2);\n> > +\t\t\tshow_path(base, opt, /*t1=*/NULL, &t2);\n> >  \t\t\tupdate_tree_entry(&t2);\n> >  \t\t}\n> >  \t}\n> >  \n> > -\tstrbuf_release(&base);\n> >  \tfree(t2tree);\n> >  \tfree(t1tree);\n> >  \treturn 0;\n> > @@ -209,7 +203,7 @@ static inline int diff_might_be_rename(void)\n> >  \t\t!DIFF_FILE_VALID(diff_queued_diff.queue[0]->one);\n> >  }\n> >  \n> > -static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n> > +static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, struct strbuf *base, struct diff_options *opt)\n> >  {\n> >  \tstruct diff_options diff_opts;\n> >  \tstruct diff_queue_struct *q = &diff_queued_diff;\n> > @@ -306,13 +300,19 @@ static void try_to_follow_renames(const unsigned char *old, const unsigned char\n> >  \tq->nr = 1;\n> >  }\n> >  \n> > -int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n> > +int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base_str, struct diff_options *opt)\n> >  {\n> > +\tstruct strbuf base;\n> >  \tint retval;\n> >  \n> > -\tretval = __diff_tree_sha1(old, new, base, opt);\n> > -\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n> > -\t\ttry_to_follow_renames(old, new, base, opt);\n> > +\tstrbuf_init(&base, PATH_MAX);\n> > +\tstrbuf_addstr(&base, base_str);\n> > +\n> > +\tretval = __diff_tree_sha1(old, new, &base, opt);\n> > +\tif (!*base_str && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n> > +\t\ttry_to_follow_renames(old, new, &base, opt);\n> > +\n> > +\tstrbuf_release(&base);\n> >  \n> >  \treturn retval;\n> >  }\n"},{"id":"237642","messageId":"20140325092336.GD3777@mini.zxlink","threadId":"35947","inReplyTo":"xmqqeh1rp9vz.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH 11/19] tree-diff: simplify tree_entry_pathcmp","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-25T09:23:36Z","receivedAt":"2014-03-25T09:23:36Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Mon, Mar 24, 2014 at 02:25:04PM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@mns.spb.ru> writes:\n> \n> > Since an earlier \"Finally switch over tree descriptors to contain a\n> > pre-parsed entry\", we can safely access all tree_desc->entry fields\n> > directly instead of first \"extracting\" them through\n> > tree_entry_extract.\n> >\n> > Use it. The code generated stays the same - only it now visually looks\n> > cleaner.\n> >\n> > Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n> > Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> > ---\n> >\n> > ( re-posting without change )\n> \n> Thanks.\n> \n> Hopefully I'll be merging the series up to this point to 'next'\n> soonish.\n\nThanks a lot!\n\n\n> >  tree-diff.c | 17 ++++++-----------\n> >  1 file changed, 6 insertions(+), 11 deletions(-)\n> >\n> > diff --git a/tree-diff.c b/tree-diff.c\n> > index 20a4fda..cf96ad7 100644\n> > --- a/tree-diff.c\n> > +++ b/tree-diff.c\n> > @@ -15,18 +15,13 @@\n> >   */\n> >  static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n> >  {\n> > -\tunsigned mode1, mode2;\n> > -\tconst char *path1, *path2;\n> > -\tconst unsigned char *sha1, *sha2;\n> > -\tint cmp, pathlen1, pathlen2;\n> > +\tstruct name_entry *e1, *e2;\n> > +\tint cmp;\n> >  \n> > -\tsha1 = tree_entry_extract(t1, &path1, &mode1);\n> > -\tsha2 = tree_entry_extract(t2, &path2, &mode2);\n> > -\n> > -\tpathlen1 = tree_entry_len(&t1->entry);\n> > -\tpathlen2 = tree_entry_len(&t2->entry);\n> > -\n> > -\tcmp = base_name_compare(path1, pathlen1, mode1, path2, pathlen2, mode2);\n> > +\te1 = &t1->entry;\n> > +\te2 = &t2->entry;\n> > +\tcmp = base_name_compare(e1->path, tree_entry_len(e1), e1->mode,\n> > +\t\t\t\te2->path, tree_entry_len(e2), e2->mode);\n> >  \treturn cmp;\n> >  }\n"},{"id":"237764","messageId":"xmqq8urymaua.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"20140325092040.GA3777@mini.zxlink","subject":"Re: [PATCH 12/19] tree-diff: remove special-case diff-emitting code for empty-tree cases","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-25T17:45:01Z","receivedAt":"2014-03-25T17:45:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@navytux.spb.ru> writes:\n\n>> >  static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n>> >  {\n>> >  \tstruct name_entry *e1, *e2;\n>> >  \tint cmp;\n>> >  \n>> > +\tif (!t1->size)\n>> > +\t\treturn t2->size ? +1 /* +∞ > c */  : 0 /* +∞ = +∞ */;\n>> > +\telse if (!t2->size)\n>> > +\t\treturn -1;\t/* c < +∞ */\n>> \n>> Where do these \"c\" come from?  I somehow feel that these comments\n>> are making it harder to understand what is going on.\n>\n> \"c\" means some finite \"c\"onstant here. When I was studying at school and\n> at the university, it was common to denote constants via this letter -\n> i.e. in algebra and operators they often show scalar multiplication as\n>\n>     c·A     (or α·A)\n>\n> etc. I understand it could maybe be confusing (but it came to me as\n> surprise), so would the following be maybe better:\n>\n>         if (!t1->size)\n>         \treturn t2->size ? +1 /* +∞ > const */  : 0 /* +∞ = +∞ */;\n>         else if (!t2->size)\n>         \treturn -1;\t/* const < +∞ */\n>\n> ?\n\nNot better at all, I am afraid.  A \"const\" in the code usually means\n\"something that does not change, as opposed to a variable\", but what\nyou are saying here is \"t1 does not have an element but t2 still\ndoes. Pretend as if t1 has a virtual/fake element that is larger\nthan any real element t2 may happen to have at the head of its\nqueue\", and you are labeling that \"real element at the head of t2\"\nas \"const\", but as the walker advances, the head element in t1 and\nt2 will change---they are not \"const\" in that sense, and the reader\nis left scratching his head seeing \"const\" there, wondering what the\nauthor of the comment meant.\n\n\"real\" or \"concrete\" might be better a phrasing, but I do not think\nhaving \"/* +inf > concrete */\" there helps the reader understand\nwhat is going on in the first place.  Perhaps:\n\n        /*\n         * When one side is empty, pretend that it has an element\n         * that sorts later than what the other non-empty side has,\n         * so that the caller advances the non-empty side without\n         * touching the empty side.\n         */\n        if (!t1->size)\n                return !t2->size ? 0 : 1;\n        else if (!t2->size)\n                return -1;\n\nor something?\n"},{"id":"237763","messageId":"xmqq4n2mmarr.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"20140325092215.GB3777@mini.zxlink","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-25T17:46:32Z","receivedAt":"2014-03-25T17:46:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@navytux.spb.ru> writes:\n\n> What are the downsides of \"__\" prefix by the way?\n\nAren't these names reserved for compiler/runtime implementations?\n"},{"id":"237776","messageId":"xmqqzjkej5ju.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"20140325092040.GA3777@mini.zxlink","subject":"Re: [PATCH 12/19] tree-diff: remove special-case diff-emitting code for empty-tree cases","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-25T22:07:33Z","receivedAt":"2014-03-25T22:07:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@navytux.spb.ru> writes:\n\n> On Mon, Mar 24, 2014 at 02:18:10PM -0700, Junio C Hamano wrote:\n>> Kirill Smelkov <kirr@mns.spb.ru> writes:\n>> \n>> > via teaching tree_entry_pathcmp() how to compare empty tree descriptors:\n>> \n>> Drop this line, as you explain the \"pretend empty compares bigger\n>> than anything else\" idea later anyway?  This early part of the\n>> proposed log message made me hiccup while reading it.\n>\n> Hmm, I was trying to show the big picture first and only then details...\n\nThe subject should be sufficient for the big picture.  \"OK, we are\nremoving the special casing\" is what we expect the reader to get.\nThen, this\n\n>> > While walking trees, we iterate their entries from lowest to highest in\n>> > sort order, so empty tree means all entries were already went over.\n\nsets the background.  \"OK, the code walks two trees, both have\nsorted elements, in parallel.\" is what we want the reader to\nunderstand.  Then the next part gives the idea of pretending that\nthe empty-side always compare later than the non-empty side while\ndoing that parallel walking (similar to \"merge\").\n\nSo, yes, I think it is a good presentation order to give big picture\npunch-line first on the subject, some background and then the\nsolution.\n"},{"id":"237824","messageId":"20140326183230.GA16002@mini.zxlink","threadId":"35947","inReplyTo":"xmqqzjkej5ju.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH 12/19] tree-diff: remove special-case diff-emitting code for empty-tree cases","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-26T18:32:30Z","receivedAt":"2014-03-26T18:32:30Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Tue, Mar 25, 2014 at 10:45:01AM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n> \n> >> >  static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n> >> >  {\n> >> >  \tstruct name_entry *e1, *e2;\n> >> >  \tint cmp;\n> >> >  \n> >> > +\tif (!t1->size)\n> >> > +\t\treturn t2->size ? +1 /* +∞ > c */  : 0 /* +∞ = +∞ */;\n> >> > +\telse if (!t2->size)\n> >> > +\t\treturn -1;\t/* c < +∞ */\n> >> \n> >> Where do these \"c\" come from?  I somehow feel that these comments\n> >> are making it harder to understand what is going on.\n> >\n> > \"c\" means some finite \"c\"onstant here. When I was studying at school and\n> > at the university, it was common to denote constants via this letter -\n> > i.e. in algebra and operators they often show scalar multiplication as\n> >\n> >     c·A     (or α·A)\n> >\n> > etc. I understand it could maybe be confusing (but it came to me as\n> > surprise), so would the following be maybe better:\n> >\n> >         if (!t1->size)\n> >         \treturn t2->size ? +1 /* +∞ > const */  : 0 /* +∞ = +∞ */;\n> >         else if (!t2->size)\n> >         \treturn -1;\t/* const < +∞ */\n> >\n> > ?\n> \n> Not better at all, I am afraid.  A \"const\" in the code usually means\n> \"something that does not change, as opposed to a variable\", but what\n> you are saying here is \"t1 does not have an element but t2 still\n> does. Pretend as if t1 has a virtual/fake element that is larger\n> than any real element t2 may happen to have at the head of its\n> queue\", and you are labeling that \"real element at the head of t2\"\n> as \"const\", but as the walker advances, the head element in t1 and\n> t2 will change---they are not \"const\" in that sense, and the reader\n> is left scratching his head seeing \"const\" there, wondering what the\n> author of the comment meant.\n\nI agree.\n\n\n> \"real\" or \"concrete\" might be better a phrasing, but I do not think\n> having \"/* +inf > concrete */\" there helps the reader understand\n> what is going on in the first place.  Perhaps:\n> \n>         /*\n>          * When one side is empty, pretend that it has an element\n>          * that sorts later than what the other non-empty side has,\n>          * so that the caller advances the non-empty side without\n>          * touching the empty side.\n>          */\n>         if (!t1->size)\n>                 return !t2->size ? 0 : 1;\n>         else if (!t2->size)\n>                 return -1;\n> \n> or something?\n\nYes, that describe the reasoning without stranger symbols. How about\ntaking it further with\n\n          * NOTE empty (=invalid) descriptor(s) take part in comparison as +infty,\n          *      so that they sort *after* valid tree entries.\n          *\n          *      Due to this convention, if trees are scanned in sorted order, all\n          *      non-empty descriptors will be processed first.\n          */\n         static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n         {\n                struct name_entry *e1, *e2;\n                int cmp;\n         \n                /* empty descriptors sort after valid tree entries */\n                if (!t1->size)\n                        return t2->size ? +1 : 0;\n                else if (!t2->size)\n                        return -1;\n\n?\n\nOn Tue, Mar 25, 2014 at 03:07:33PM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n> \n> > On Mon, Mar 24, 2014 at 02:18:10PM -0700, Junio C Hamano wrote:\n> >> Kirill Smelkov <kirr@mns.spb.ru> writes:\n> >> \n> >> > via teaching tree_entry_pathcmp() how to compare empty tree descriptors:\n> >> \n> >> Drop this line, as you explain the \"pretend empty compares bigger\n> >> than anything else\" idea later anyway?  This early part of the\n> >> proposed log message made me hiccup while reading it.\n> >\n> > Hmm, I was trying to show the big picture first and only then details...\n> \n> The subject should be sufficient for the big picture.  \"OK, we are\n> removing the special casing\" is what we expect the reader to get.\n> Then, this\n> \n> >> > While walking trees, we iterate their entries from lowest to highest in\n> >> > sort order, so empty tree means all entries were already went over.\n> \n> sets the background.  \"OK, the code walks two trees, both have\n> sorted elements, in parallel.\" is what we want the reader to\n> understand.  Then the next part gives the idea of pretending that\n> the empty-side always compare later than the non-empty side while\n> doing that parallel walking (similar to \"merge\").\n> \n> So, yes, I think it is a good presentation order to give big picture\n> punch-line first on the subject, some background and then the\n> solution.\n\nOk, let it be this way and let's drop it.\n\nHere is updated patch:\n(please keep author email)\n\n---- 8< ----\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\nDate: Mon, 24 Feb 2014 20:21:44 +0400\nSubject: [PATCH v2] tree-diff: remove special-case diff-emitting code for empty-tree cases\n\nWhile walking trees, we iterate their entries from lowest to highest in\nsort order, so empty tree means all entries were already went over.\n\nIf we artificially assign +infinity value to such tree \"entry\", it will\ngo after all usual entries, and through the usual driver loop we will be\ntaking the same actions, which were hand-coded for special cases, i.e.\n\n    t1 empty, t2 non-empty\n        pathcmp(+∞, t2) -> +1\n        show_path(/*t1=*/NULL, t2);     /* = t1 > t2 case in main loop */\n\n    t1 non-empty, t2-empty\n        pathcmp(t1, +∞) -> -1\n        show_path(t1, /*t2=*/NULL);     /* = t1 < t2 case in main loop */\n\nIn other words when we have t1 and t2, we return a sign that tells the\ncaller to indicate the \"earlier\" one to be emitted, and by returning the\nsign that causes the non-empty side to be emitted, we will automatically\ncause the entries from the remaining side to be emitted, without\nattempting to touch the empty side at all.  We can teach\ntree_entry_pathcmp() to pretend that an empty tree has an element that\nsorts after anything else to achieve this.\n\nRight now we never go to when compared tree descriptors are both\ninfinity, as this condition is checked in the loop beginning as\nfinishing criteria, but will do so in the future, when there will be\nseveral parents iterated simultaneously, and some pair of them would run\nto the end.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\nChanges since v1:\n\n - reworked commit log and comments as per Junio suggestions.\n\n\n tree-diff.c | 26 ++++++++++++++------------\n 1 file changed, 14 insertions(+), 12 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex f8b2607..6177658 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -12,12 +12,24 @@\n  *\n  * NOTE files and directories *always* compare differently, even when having\n  *      the same name - thanks to base_name_compare().\n+ *\n+ * NOTE empty (=invalid) descriptor(s) take part in comparison as +infty,\n+ *      so that they sort *after* valid tree entries.\n+ *\n+ *      Due to this convention, if trees are scanned in sorted order, all\n+ *      non-empty descriptors will be processed first.\n  */\n static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n {\n \tstruct name_entry *e1, *e2;\n \tint cmp;\n \n+\t/* empty descriptors sort after valid tree entries */\n+\tif (!t1->size)\n+\t\treturn t2->size ? +1 : 0;\n+\telse if (!t2->size)\n+\t\treturn -1;\n+\n \te1 = &t1->entry;\n \te2 = &t2->entry;\n \tcmp = base_name_compare(e1->path, tree_entry_len(e1), e1->mode,\n@@ -150,18 +162,8 @@ int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \t\t\tskip_uninteresting(t1, &base, opt);\n \t\t\tskip_uninteresting(t2, &base, opt);\n \t\t}\n-\t\tif (!t1->size) {\n-\t\t\tif (!t2->size)\n-\t\t\t\tbreak;\n-\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n-\t\t\tupdate_tree_entry(t2);\n-\t\t\tcontinue;\n-\t\t}\n-\t\tif (!t2->size) {\n-\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n-\t\t\tupdate_tree_entry(t1);\n-\t\t\tcontinue;\n-\t\t}\n+\t\tif (!t1->size && !t2->size)\n+\t\t\tbreak;\n \n \t\tcmp = tree_entry_pathcmp(t1, t2);\n \n-- \n1.9.rc0.143.g6fd479e\n"},{"id":"237842","messageId":"20140326195201.GB16002@mini.zxlink","threadId":"35947","inReplyTo":"xmqq4n2mmarr.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-26T19:52:01Z","receivedAt":"2014-03-26T19:52:01Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Tue, Mar 25, 2014 at 10:46:32AM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n> \n> > What are the downsides of \"__\" prefix by the way?\n> \n> Aren't these names reserved for compiler/runtime implementations?\n\nYes, but there are precedents when people don't obey it widely and\nin practice everything works :) Let it be something portable anyway -\nhow about diff_tree_sha1_low() ?\n\nSo corrected patch is below. If such suffixing will be accepted, I will\nsend follow-up patches corrected similiary.\n\n  ( or please pull them from\n    git://repo.or.cz/git/kirr.git y6/tree-diff-walk-multitree )\n\nThanks,\nKirill\n\n---- 8< ----\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\nDate: Mon, 24 Feb 2014 20:21:46 +0400\nSubject: [PATCH v3] tree-diff: rework diff_tree interface to be sha1 based\n\nIn the next commit this will allow to reduce intermediate calls, when\nrecursing into subtrees - at that stage we know only subtree sha1, and\nit is natural for tree walker to start from that phase. For now we do\n\n    diff_tree\n        show_path\n            diff_tree_sha1\n                diff_tree\n                    ...\n\nand the change will allow to reduce it to\n\n    diff_tree\n        show_path\n            diff_tree\n\nAlso, it will allow to omit allocating strbuf for each subtree, and just\nreuse the common strbuf via playing with its len.\n\nThe above-mentioned improvements go in the next 2 patches.\n\nThe downside is that try_to_follow_renames(), if active, we cause\nre-reading of 2 initial trees, which was negligible based on my timings,\nand which is outweighed cogently by the upsides.\n\nNOTE To keep with the current interface and semantics, I needed to\nrename the function from diff_tree() to diff_tree_sha1(). As\ndiff_tree_sha1() was already used, and the function we are talking here\nis its more low-level helper, let's use convention for suffixing\nsuch helpers with \"_low\". So the final renaming is\n\n    diff_tree() -> diff_tree_sha1_low()\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\nChanges since v2:\n\n - renamed __diff_tree_sha1() -> diff_tree_sha1_low() as the former\n   overlaps with reserved-for-implementation identifiers namespace.\n\n\n tree-diff.c | 60 ++++++++++++++++++++++++++++--------------------------------\n 1 file changed, 28 insertions(+), 32 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex f137f39..0277c5c 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -141,12 +141,17 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n \t}\n }\n \n-static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n-\t\t     const char *base_str, struct diff_options *opt)\n+static int diff_tree_sha1_low(const unsigned char *old, const unsigned char *new,\n+\t\t\t      const char *base_str, struct diff_options *opt)\n {\n+\tstruct tree_desc t1, t2;\n+\tvoid *t1tree, *t2tree;\n \tstruct strbuf base;\n \tint baselen = strlen(base_str);\n \n+\tt1tree = fill_tree_descriptor(&t1, old);\n+\tt2tree = fill_tree_descriptor(&t2, new);\n+\n \t/* Enable recursion indefinitely */\n \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n \n@@ -159,39 +164,41 @@ static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \t\tif (opt->pathspec.nr) {\n-\t\t\tskip_uninteresting(t1, &base, opt);\n-\t\t\tskip_uninteresting(t2, &base, opt);\n+\t\t\tskip_uninteresting(&t1, &base, opt);\n+\t\t\tskip_uninteresting(&t2, &base, opt);\n \t\t}\n-\t\tif (!t1->size && !t2->size)\n+\t\tif (!t1.size && !t2.size)\n \t\t\tbreak;\n \n-\t\tcmp = tree_entry_pathcmp(t1, t2);\n+\t\tcmp = tree_entry_pathcmp(&t1, &t2);\n \n \t\t/* t1 = t2 */\n \t\tif (cmp == 0) {\n \t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n-\t\t\t    hashcmp(t1->entry.sha1, t2->entry.sha1) ||\n-\t\t\t    (t1->entry.mode != t2->entry.mode))\n-\t\t\t\tshow_path(&base, opt, t1, t2);\n+\t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n+\t\t\t    (t1.entry.mode != t2.entry.mode))\n+\t\t\t\tshow_path(&base, opt, &t1, &t2);\n \n-\t\t\tupdate_tree_entry(t1);\n-\t\t\tupdate_tree_entry(t2);\n+\t\t\tupdate_tree_entry(&t1);\n+\t\t\tupdate_tree_entry(&t2);\n \t\t}\n \n \t\t/* t1 < t2 */\n \t\telse if (cmp < 0) {\n-\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n-\t\t\tupdate_tree_entry(t1);\n+\t\t\tshow_path(&base, opt, &t1, /*t2=*/NULL);\n+\t\t\tupdate_tree_entry(&t1);\n \t\t}\n \n \t\t/* t1 > t2 */\n \t\telse {\n-\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n-\t\t\tupdate_tree_entry(t2);\n+\t\t\tshow_path(&base, opt, /*t1=*/NULL, &t2);\n+\t\t\tupdate_tree_entry(&t2);\n \t\t}\n \t}\n \n \tstrbuf_release(&base);\n+\tfree(t2tree);\n+\tfree(t1tree);\n \treturn 0;\n }\n \n@@ -206,7 +213,7 @@ static inline int diff_might_be_rename(void)\n \t\t!DIFF_FILE_VALID(diff_queued_diff.queue[0]->one);\n }\n \n-static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, const char *base, struct diff_options *opt)\n+static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n {\n \tstruct diff_options diff_opts;\n \tstruct diff_queue_struct *q = &diff_queued_diff;\n@@ -244,7 +251,7 @@ static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, co\n \tdiff_opts.break_opt = opt->break_opt;\n \tdiff_opts.rename_score = opt->rename_score;\n \tdiff_setup_done(&diff_opts);\n-\tdiff_tree(t1, t2, base, &diff_opts);\n+\tdiff_tree_sha1_low(old, new, base, &diff_opts);\n \tdiffcore_std(&diff_opts);\n \tfree_pathspec(&diff_opts.pathspec);\n \n@@ -305,23 +312,12 @@ static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, co\n \n int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n {\n-\tvoid *tree1, *tree2;\n-\tstruct tree_desc t1, t2;\n-\tunsigned long size1, size2;\n \tint retval;\n \n-\ttree1 = fill_tree_descriptor(&t1, old);\n-\ttree2 = fill_tree_descriptor(&t2, new);\n-\tsize1 = t1.size;\n-\tsize2 = t2.size;\n-\tretval = diff_tree(&t1, &t2, base, opt);\n-\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename()) {\n-\t\tinit_tree_desc(&t1, tree1, size1);\n-\t\tinit_tree_desc(&t2, tree2, size2);\n-\t\ttry_to_follow_renames(&t1, &t2, base, opt);\n-\t}\n-\tfree(tree1);\n-\tfree(tree2);\n+\tretval = diff_tree_sha1_low(old, new, base, opt);\n+\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n+\t\ttry_to_follow_renames(old, new, base, opt);\n+\n \treturn retval;\n }\n \n-- \n1.9.rc0.143.g6fd479e\n"},{"id":"237852","messageId":"xmqq1txoiqzj.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"20140326195201.GB16002@mini.zxlink","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-26T21:34:24Z","receivedAt":"2014-03-26T21:34:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@navytux.spb.ru> writes:\n\n> On Tue, Mar 25, 2014 at 10:46:32AM -0700, Junio C Hamano wrote:\n>> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n>> \n>> > What are the downsides of \"__\" prefix by the way?\n>> \n>> Aren't these names reserved for compiler/runtime implementations?\n>\n> Yes, but there are precedents when people don't obey it widely and\n> in practice everything works :)\n\nI think you are alluding to the practice in the Linux kernel, but\ntheir requirement is vastly different---their product do not even\nlink with libc and they always compile with specific selected\nversions of gcc, no?\n\n> Let it be something portable anyway -\n> how about diff_tree_sha1_low() ?\n\nSure.\n\nAs this is a file-scope static, I do not think the exact naming\nmatters that much.  Just FYI, we seem to use ll_ prefix (standing\nfor low-level) in some places.\n\nThanks.\n"},{"id":"237898","messageId":"20140327142129.GA17333@mini.zxlink","threadId":"35947","inReplyTo":"7e5e5a381ba4204eac14c5be9e270ffdc0e2be7a.1393257006.git.kirr@mns.spb.ru","subject":"Re: [PATCH 15/19] tree-diff: no need to call \"full\" diff_tree_sha1 from show_path()","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-27T14:21:29Z","receivedAt":"2014-03-27T14:21:29Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Mon, Feb 24, 2014 at 08:21:47PM +0400, Kirill Smelkov wrote:\n> As described in previous commit, when recursing into sub-trees, we can\n> use lower-level tree walker, since its interface is now sha1 based.\n> \n> The change is ok, because diff_tree_sha1() only invokes\n> __diff_tree_sha1(), and also, if base is empty, try_to_follow_renames().\n> But base is not empty here, as we have added a path and '/' before\n> recursing.\n> \n> Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n> \n> ( re-posting without change )\n> \n>  tree-diff.c | 4 ++--\n>  1 file changed, 2 insertions(+), 2 deletions(-)\n> \n> diff --git a/tree-diff.c b/tree-diff.c\n> index f90acf5..aea0297 100644\n> --- a/tree-diff.c\n> +++ b/tree-diff.c\n> @@ -114,8 +114,8 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n>  \n>  \tif (recurse) {\n>  \t\tstrbuf_addch(base, '/');\n> -\t\tdiff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n> -\t\t\t       t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n> +\t\t__diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n> +\t\t\t\t t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n>  \t}\n>  \n>  \tstrbuf_setlen(base, old_baselen);\n\nI've found this does not compile as I've forgot to add __diff_tree_sha1\nprototype, and also we are changing naming for __diff_tree_sha1() to\nll_diff_tree_sha1() to follow Git coding style for consistency and\ncorrections to previous patch, so here goes v2:\n\n(please keep author email)\n---- 8< ----\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\nDate: Mon, 24 Feb 2014 20:21:47 +0400\nSubject: [PATCH v2] tree-diff: no need to call \"full\" diff_tree_sha1 from show_path()\n\nAs described in previous commit, when recursing into sub-trees, we can\nuse lower-level tree walker, since its interface is now sha1 based.\n\nThe change is ok, because diff_tree_sha1() only invokes\nll_diff_tree_sha1(), and also, if base is empty, try_to_follow_renames().\nBut base is not empty here, as we have added a path and '/' before\nrecursing.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\nChanges since v1:\n\n - adjust to renaming __diff_tree_sha1 -> ll_diff_tree_sha1;\n - added ll_diff_tree_sha1 prototype as the function is defined below\n   here-introduced call-site.\n\n tree-diff.c | 8 ++++++--\n 1 file changed, 6 insertions(+), 2 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 1d02e43..7fbb022 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -6,6 +6,10 @@\n #include \"diffcore.h\"\n #include \"tree.h\"\n \n+\n+static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n+\t\t\t     const char *base_str, struct diff_options *opt);\n+\n /*\n  * Compare two tree entries, taking into account only path/S_ISDIR(mode),\n  * but not their sha1's.\n@@ -118,8 +122,8 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n \n \tif (recurse) {\n \t\tstrbuf_addch(base, '/');\n-\t\tdiff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n-\t\t\t       t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n+\t\tll_diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n+\t\t\t\t  t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n \t}\n \n \tstrbuf_setlen(base, old_baselen);\n-- \n1.9.rc0.143.g6fd479e\n"},{"id":"237899","messageId":"20140327142207.GB17333@mini.zxlink","threadId":"35947","inReplyTo":"20140325092320.GC3777@mini.zxlink","subject":"Re: [PATCH v2 16/19] tree-diff: reuse base str(buf) memory on sub-tree recursion","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-27T14:22:07Z","receivedAt":"2014-03-27T14:22:07Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Tue, Mar 25, 2014 at 01:23:20PM +0400, Kirill Smelkov wrote:\n> On Mon, Mar 24, 2014 at 02:43:36PM -0700, Junio C Hamano wrote:\n> > Kirill Smelkov <kirr@mns.spb.ru> writes:\n> > \n> > > instead of allocating it all the time for every subtree in\n> > > __diff_tree_sha1, let's allocate it once in diff_tree_sha1, and then all\n> > > callee just use it in stacking style, without memory allocations.\n> > >\n> > > This should be faster, and for me this change gives the following\n> > > slight speedups for\n> > >\n> > >     git log --raw --no-abbrev --no-renames --format='%H'\n> > >\n> > >                 navy.git    linux.git v3.10..v3.11\n> > >\n> > >     before      0.618s      1.903s\n> > >     after       0.611s      1.889s\n> > >     speedup     1.1%        0.7%\n> > >\n> > > Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n> > > ---\n> > >\n> > > Changes since v1:\n> > >\n> > >  - don't need to touch diff.h, as the function we are changing became static.\n> > >\n> > >  tree-diff.c | 36 ++++++++++++++++++------------------\n> > >  1 file changed, 18 insertions(+), 18 deletions(-)\n> > >\n> > > diff --git a/tree-diff.c b/tree-diff.c\n> > > index aea0297..c76821d 100644\n> > > --- a/tree-diff.c\n> > > +++ b/tree-diff.c\n> > > @@ -115,7 +115,7 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n> > >  \tif (recurse) {\n> > >  \t\tstrbuf_addch(base, '/');\n> > >  \t\t__diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n> > > -\t\t\t\t t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n> > > +\t\t\t\t t2 ? t2->entry.sha1 : NULL, base, opt);\n> > >  \t}\n> > >  \n> > >  \tstrbuf_setlen(base, old_baselen);\n> > \n> > I was scratching my head for a while, after seeing that there does\n> > not seem to be any *new* code added by this patch in order to\n> > store-away the original length and restore the singleton base buffer\n> > to the original length after using addch/addstr to extend it.\n> > \n> > But I see that the code has already been prepared to do this\n> > conversion.  I wonder why we didn't do this earlier ;-)\n> \n> The conversion to reusing memory started in 48932677 \"diff-tree: convert\n> base+baselen to writable strbuf\" which allowed to avoid \"quite a bit of\n> malloc() and memcpy()\", but for this to work allocation at diff_tree()\n> entry had to be there.\n> \n> In particular it had to be there, because diff_tree() accepted base as C\n> string, not strbuf, and since diff_tree() was calling itself\n> recursively - oops - new allocation on every subtree.\n> \n> I've opened the door for avoiding allocations via splitting diff_tree\n> into high-level and low-level parts. The high-level part still accepts\n> `char *base`, but low-level function operates on strbuf and recurses\n> into low-level self.\n> \n> The high-level diff_tree_sha1() still allocates memory for every\n> diff(tree1,tree2), but that is significantly lower compared to\n> allocating memory on every subtree...\n> \n> The lesson here is: better use strbuf for api unless there is a reason\n> not to.\n> \n> \n> > Looks good.  Thanks.\n> \n> Thanks.\n\nThanks again. Here it goes adjusted to __diff_tree_sha1 -> ll_diff_tree_sha1 renaming:\n\n(please keep author email)\n---- 8< ----\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\nDate: Mon, 24 Feb 2014 20:21:48 +0400\nSubject: [PATCH v3] tree-diff: reuse base str(buf) memory on sub-tree recursion\n\ninstead of allocating it all the time for every subtree in\nll_diff_tree_sha1, let's allocate it once in diff_tree_sha1, and then all\ncallee just use it in stacking style, without memory allocations.\n\nThis should be faster, and for me this change gives the following\nslight speedups for\n\n    git log --raw --no-abbrev --no-renames --format='%H'\n\n                navy.git    linux.git v3.10..v3.11\n\n    before      0.618s      1.903s\n    after       0.611s      1.889s\n    speedup     1.1%        0.7%\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\nChanges since v2:\n\n - adjust to __diff_tree_sha1 -> ll_diff_tree_sha1 renaming.\n\nChanges since v1:\n\n - don't need to touch diff.h, as the function we are changing became\n   static.\n\n tree-diff.c | 38 +++++++++++++++++++-------------------\n 1 file changed, 19 insertions(+), 19 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 7fbb022..8c8bde6 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -8,7 +8,7 @@\n \n \n static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n-\t\t\t     const char *base_str, struct diff_options *opt);\n+\t\t\t     struct strbuf *base, struct diff_options *opt);\n \n /*\n  * Compare two tree entries, taking into account only path/S_ISDIR(mode),\n@@ -123,7 +123,7 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n \tif (recurse) {\n \t\tstrbuf_addch(base, '/');\n \t\tll_diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n-\t\t\t\t  t2 ? t2->entry.sha1 : NULL, base->buf, opt);\n+\t\t\t\t  t2 ? t2->entry.sha1 : NULL, base, opt);\n \t}\n \n \tstrbuf_setlen(base, old_baselen);\n@@ -146,12 +146,10 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n }\n \n static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n-\t\t\t     const char *base_str, struct diff_options *opt)\n+\t\t\t     struct strbuf *base, struct diff_options *opt)\n {\n \tstruct tree_desc t1, t2;\n \tvoid *t1tree, *t2tree;\n-\tstruct strbuf base;\n-\tint baselen = strlen(base_str);\n \n \tt1tree = fill_tree_descriptor(&t1, old);\n \tt2tree = fill_tree_descriptor(&t2, new);\n@@ -159,17 +157,14 @@ static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \t/* Enable recursion indefinitely */\n \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n \n-\tstrbuf_init(&base, PATH_MAX);\n-\tstrbuf_add(&base, base_str, baselen);\n-\n \tfor (;;) {\n \t\tint cmp;\n \n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \t\tif (opt->pathspec.nr) {\n-\t\t\tskip_uninteresting(&t1, &base, opt);\n-\t\t\tskip_uninteresting(&t2, &base, opt);\n+\t\t\tskip_uninteresting(&t1, base, opt);\n+\t\t\tskip_uninteresting(&t2, base, opt);\n \t\t}\n \t\tif (!t1.size && !t2.size)\n \t\t\tbreak;\n@@ -181,7 +176,7 @@ static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n \t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n \t\t\t    (t1.entry.mode != t2.entry.mode))\n-\t\t\t\tshow_path(&base, opt, &t1, &t2);\n+\t\t\t\tshow_path(base, opt, &t1, &t2);\n \n \t\t\tupdate_tree_entry(&t1);\n \t\t\tupdate_tree_entry(&t2);\n@@ -189,18 +184,17 @@ static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \n \t\t/* t1 < t2 */\n \t\telse if (cmp < 0) {\n-\t\t\tshow_path(&base, opt, &t1, /*t2=*/NULL);\n+\t\t\tshow_path(base, opt, &t1, /*t2=*/NULL);\n \t\t\tupdate_tree_entry(&t1);\n \t\t}\n \n \t\t/* t1 > t2 */\n \t\telse {\n-\t\t\tshow_path(&base, opt, /*t1=*/NULL, &t2);\n+\t\t\tshow_path(base, opt, /*t1=*/NULL, &t2);\n \t\t\tupdate_tree_entry(&t2);\n \t\t}\n \t}\n \n-\tstrbuf_release(&base);\n \tfree(t2tree);\n \tfree(t1tree);\n \treturn 0;\n@@ -217,7 +211,7 @@ static inline int diff_might_be_rename(void)\n \t\t!DIFF_FILE_VALID(diff_queued_diff.queue[0]->one);\n }\n \n-static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n+static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, struct strbuf *base, struct diff_options *opt)\n {\n \tstruct diff_options diff_opts;\n \tstruct diff_queue_struct *q = &diff_queued_diff;\n@@ -314,13 +308,19 @@ static void try_to_follow_renames(const unsigned char *old, const unsigned char\n \tq->nr = 1;\n }\n \n-int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n+int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base_str, struct diff_options *opt)\n {\n+\tstruct strbuf base;\n \tint retval;\n \n-\tretval = ll_diff_tree_sha1(old, new, base, opt);\n-\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n-\t\ttry_to_follow_renames(old, new, base, opt);\n+\tstrbuf_init(&base, PATH_MAX);\n+\tstrbuf_addstr(&base, base_str);\n+\n+\tretval = ll_diff_tree_sha1(old, new, &base, opt);\n+\tif (!*base_str && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n+\t\ttry_to_follow_renames(old, new, &base, opt);\n+\n+\tstrbuf_release(&base);\n \n \treturn retval;\n }\n-- \n1.9.rc0.143.g6fd479e\n"},{"id":"237900","messageId":"20140327142250.GC17333@mini.zxlink","threadId":"35947","inReplyTo":"xmqq1txrp8ur.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-27T14:22:50Z","receivedAt":"2014-03-27T14:22:50Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Mon, Mar 24, 2014 at 02:47:24PM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@mns.spb.ru> writes:\n> \n> > On Fri, Feb 28, 2014 at 06:19:58PM +0100, Erik Faye-Lund wrote:\n> >> On Fri, Feb 28, 2014 at 6:00 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> >> ...\n> >> > In fact that would be maybe preferred, for maintainers to enable alloca\n> >> > with knowledge and testing, as one person can't have them all at hand.\n> >> \n> >> Yeah, you're probably right.\n> >\n> > Erik, the patch has been merged into pu today. Would you please\n> > follow-up with tested MINGW change?\n> \n> Sooo.... I lost track but this discussion seems to have petered out\n> around here.  I think the copy we have had for a while on 'pu' is\n> basically sound, and can easily built on by platform folks by adding\n> or removing the -DHAVE_ALLOCA_H from the Makefile.\n\nYes, that is all correct - that version works and we can improve it in\nthe future with platform-specific follow-up patches, if needed.\n\nPlease pick up the patch with ack from Thomas Schwinge.\n\nThanks,\nKirill\n\n(please keep author email)\n---- 8< ----\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\nDate: Mon, 24 Feb 2014 20:21:49 +0400\nSubject: [PATCH v1a] Portable alloca for Git\n\nIn the next patch we'll have to use alloca() for performance reasons,\nbut since alloca is non-standardized and is not portable, let's have a\ntrick with compatibility wrappers:\n\n1. at configure time, determine, do we have working alloca() through\n   alloca.h, and define\n\n    #define HAVE_ALLOCA_H\n\n   if yes.\n\n2. in code\n\n    #ifdef HAVE_ALLOCA_H\n    # include <alloca.h>\n    # define xalloca(size)      (alloca(size))\n    # define xalloca_free(p)    do {} while(0)\n    #else\n    # define xalloca(size)      (xmalloc(size))\n    # define xalloca_free(p)    (free(p))\n    #endif\n\n   and use it like\n\n   func() {\n       p = xalloca(size);\n       ...\n\n       xalloca_free(p);\n   }\n\nThis way, for systems, where alloca is available, we'll have optimal\non-stack allocations with fast executions. On the other hand, on\nsystems, where alloca is not available, this gracefully fallbacks to\nxmalloc/free.\n\nBoth autoconf and config.mak.uname configurations were updated. For\nautoconf, we are not bothering considering cases, when no alloca.h is\navailable, but alloca() works some other way - its simply alloca.h is\navailable and works or not, everything else is deep legacy.\n\nFor config.mak.uname, I've tried to make my almost-sure guess for where\nalloca() is available, but since I only have access to Linux it is the\nonly change I can be sure about myself, with relevant to other changed\nsystems people Cc'ed.\n\nNOTE\n\nSunOS and Windows had explicit -DHAVE_ALLOCA_H in their configurations.\nI've changed that to now-common HAVE_ALLOCA_H=YesPlease which should be\ncorrect.\n\nCc: Brandon Casey <drafnel@gmail.com>\nCc: Marius Storm-Olsen <mstormo@gmail.com>\nCc: Johannes Sixt <j6t@kdbg.org>\nCc: Johannes Schindelin <Johannes.Schindelin@gmx.de>\nCc: Ramsay Jones <ramsay@ramsay1.demon.co.uk>\nCc: Gerrit Pape <pape@smarden.org>\nCc: Petr Salinger <Petr.Salinger@seznam.cz>\nCc: Jonathan Nieder <jrnieder@gmail.com>\nCc: Thomas Schwinge <tschwinge@gnu.org>\nAcked-by: Thomas Schwinge <thomas@codesourcery.com> (GNU Hurd changes)\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\nChanges since v1:\n\n - added ack for GNU/Hurd.\n\n Makefile          |  6 ++++++\n config.mak.uname  | 10 ++++++++--\n configure.ac      |  8 ++++++++\n git-compat-util.h |  8 ++++++++\n 4 files changed, 30 insertions(+), 2 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex dddaf4f..0334806 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -30,6 +30,8 @@ all::\n # Define LIBPCREDIR=/foo/bar if your libpcre header and library files are in\n # /foo/bar/include and /foo/bar/lib directories.\n #\n+# Define HAVE_ALLOCA_H if you have working alloca(3) defined in that header.\n+#\n # Define NO_CURL if you do not have libcurl installed.  git-http-fetch and\n # git-http-push are not built, and you cannot use http:// and https://\n # transports (neither smart nor dumb).\n@@ -1099,6 +1101,10 @@ ifdef USE_LIBPCRE\n \tEXTLIBS += -lpcre\n endif\n \n+ifdef HAVE_ALLOCA_H\n+\tBASIC_CFLAGS += -DHAVE_ALLOCA_H\n+endif\n+\n ifdef NO_CURL\n \tBASIC_CFLAGS += -DNO_CURL\n \tREMOTE_CURL_PRIMARY =\ndiff --git a/config.mak.uname b/config.mak.uname\nindex 7d31fad..71602ee 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -28,6 +28,7 @@ ifeq ($(uname_S),OSF1)\n \tNO_NSEC = YesPlease\n endif\n ifeq ($(uname_S),Linux)\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_STRLCPY = YesPlease\n \tNO_MKSTEMPS = YesPlease\n \tHAVE_PATHS_H = YesPlease\n@@ -35,6 +36,7 @@ ifeq ($(uname_S),Linux)\n \tHAVE_DEV_TTY = YesPlease\n endif\n ifeq ($(uname_S),GNU/kFreeBSD)\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_STRLCPY = YesPlease\n \tNO_MKSTEMPS = YesPlease\n \tHAVE_PATHS_H = YesPlease\n@@ -103,6 +105,7 @@ ifeq ($(uname_S),SunOS)\n \tNEEDS_NSL = YesPlease\n \tSHELL_PATH = /bin/bash\n \tSANE_TOOL_PATH = /usr/xpg6/bin:/usr/xpg4/bin\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_STRCASESTR = YesPlease\n \tNO_MEMMEM = YesPlease\n \tNO_MKDTEMP = YesPlease\n@@ -146,7 +149,7 @@ ifeq ($(uname_S),SunOS)\n \tendif\n \tINSTALL = /usr/ucb/install\n \tTAR = gtar\n-\tBASIC_CFLAGS += -D__EXTENSIONS__ -D__sun__ -DHAVE_ALLOCA_H\n+\tBASIC_CFLAGS += -D__EXTENSIONS__ -D__sun__\n endif\n ifeq ($(uname_O),Cygwin)\n \tifeq ($(shell expr \"$(uname_R)\" : '1\\.[1-6]\\.'),4)\n@@ -166,6 +169,7 @@ ifeq ($(uname_O),Cygwin)\n \telse\n \t\tNO_REGEX = UnfortunatelyYes\n \tendif\n+\tHAVE_ALLOCA_H = YesPlease\n \tNEEDS_LIBICONV = YesPlease\n \tNO_FAST_WORKING_DIRECTORY = UnfortunatelyYes\n \tNO_ST_BLOCKS_IN_STRUCT_STAT = YesPlease\n@@ -239,6 +243,7 @@ ifeq ($(uname_S),AIX)\n endif\n ifeq ($(uname_S),GNU)\n \t# GNU/Hurd\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_STRLCPY = YesPlease\n \tNO_MKSTEMPS = YesPlease\n \tHAVE_PATHS_H = YesPlease\n@@ -316,6 +321,7 @@ endif\n ifeq ($(uname_S),Windows)\n \tGIT_VERSION := $(GIT_VERSION).MSVC\n \tpathsep = ;\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_PREAD = YesPlease\n \tNEEDS_CRYPTO_WITH_SSL = YesPlease\n \tNO_LIBGEN_H = YesPlease\n@@ -363,7 +369,7 @@ ifeq ($(uname_S),Windows)\n \tCOMPAT_OBJS = compat/msvc.o compat/winansi.o \\\n \t\tcompat/win32/pthread.o compat/win32/syslog.o \\\n \t\tcompat/win32/dirent.o\n-\tCOMPAT_CFLAGS = -D__USE_MINGW_ACCESS -DNOGDI -DHAVE_STRING_H -DHAVE_ALLOCA_H -Icompat -Icompat/regex -Icompat/win32 -DSTRIP_EXTENSION=\\\".exe\\\"\n+\tCOMPAT_CFLAGS = -D__USE_MINGW_ACCESS -DNOGDI -DHAVE_STRING_H -Icompat -Icompat/regex -Icompat/win32 -DSTRIP_EXTENSION=\\\".exe\\\"\n \tBASIC_LDFLAGS = -IGNORE:4217 -IGNORE:4049 -NOLOGO -SUBSYSTEM:CONSOLE -NODEFAULTLIB:MSVCRT.lib\n \tEXTLIBS = user32.lib advapi32.lib shell32.lib wininet.lib ws2_32.lib\n \tPTHREAD_LIBS =\ndiff --git a/configure.ac b/configure.ac\nindex 2f43393..0eae704 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -272,6 +272,14 @@ AS_HELP_STRING([],           [ARG can be also prefix for libpcre library and hea\n \tGIT_CONF_SUBST([LIBPCREDIR])\n     fi)\n #\n+# Define HAVE_ALLOCA_H if you have working alloca(3) defined in that header.\n+AC_FUNC_ALLOCA\n+case $ac_cv_working_alloca_h in\n+    yes)    HAVE_ALLOCA_H=YesPlease;;\n+    *)      HAVE_ALLOCA_H='';;\n+esac\n+GIT_CONF_SUBST([HAVE_ALLOCA_H])\n+#\n # Define NO_CURL if you do not have curl installed.  git-http-pull and\n # git-http-push are not built, and you cannot use http:// and https://\n # transports.\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex cbd86c3..63b2b3b 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -526,6 +526,14 @@ extern void release_pack_memory(size_t);\n typedef void (*try_to_free_t)(size_t);\n extern try_to_free_t set_try_to_free_routine(try_to_free_t);\n \n+#ifdef HAVE_ALLOCA_H\n+# include <alloca.h>\n+# define xalloca(size)      (alloca(size))\n+# define xalloca_free(p)    do {} while (0)\n+#else\n+# define xalloca(size)      (xmalloc(size))\n+# define xalloca_free(p)    (free(p))\n+#endif\n extern char *xstrdup(const char *str);\n extern void *xmalloc(size_t size);\n extern void *xmallocz(size_t size);\n-- \n1.9.rc0.143.g6fd479e\n"},{"id":"237901","messageId":"20140327142354.GD17333@mini.zxlink","threadId":"35947","inReplyTo":"7b307610fe214f47643a46b3e815487558db244e.1393257006.git.kirr@mns.spb.ru","subject":"Re: [PATCH v2 18/19] tree-diff: rework diff_tree() to generate diffs for multiparent cases as well","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-27T14:23:54Z","receivedAt":"2014-03-27T14:23:54Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Mon, Feb 24, 2014 at 08:21:50PM +0400, Kirill Smelkov wrote:\n[...]\n> not changed:\n> \n> - low-level helpers are still named with \"__\" prefix as, imho, that is the best\n>   convention to name such helpers, without sacrificing signal/noise ratio. All\n>   of them are now static though.\n\nPlease find attached corrected version of this patch with\n__diff_tree_sha1() renamed to ll_diff_tree_sha1() and other identifiers\ncorrected similarly for consistency with Git codebase style.\n\nThanks,\nKirill\n\n(please keep author email)\n---- 8< ----\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\nDate: Mon, 24 Feb 2014 20:21:50 +0400\nSubject: [PATCH v3] tree-diff: rework diff_tree() to generate diffs for multiparent cases as well\n\nPreviously diff_tree(), which is now named ll_diff_tree_sha1(), was\ngenerating diff_filepair(s) for two trees t1 and t2, and that was\nusually used for a commit as t1=HEAD~, and t2=HEAD - i.e. to see changes\na commit introduces.\n\nIn Git, however, we have fundamentally built flexibility in that a\ncommit can have many parents - 1 for a plain commit, 2 for a simple merge,\nbut also more than 2 for merging several heads at once.\n\nFor merges there is a so called combine-diff, which shows diff, a merge\nintroduces by itself, omitting changes done by any parent. That works\nthrough first finding paths, that are different to all parents, and then\nshowing generalized diff, with separate columns for +/- for each parent.\nThe code lives in combine-diff.c .\n\nThere is an impedance mismatch, however, in that a commit could\ngenerally have any number of parents, and that while diffing trees, we\ndivide cases for 2-tree diffs and more-than-2-tree diffs. I mean there\nis no special casing for multiple parents commits in e.g.\nrevision-walker .\n\nThat impedance mismatch *hurts* *performance* *badly* for generating\ncombined diffs - in \"combine-diff: optimize combine_diff_path\nsets intersection\" I've already removed some slowness from it, but from\nthe timings provided there, it could be seen, that combined diffs still\ncost more than an order of magnitude more cpu time, compared to diff for\nusual commits, and that would only be an optimistic estimate, if we take\ninto account that for e.g. linux.git there is only one merge for several\ndozens of plain commits.\n\nThat slowness comes from the fact that currently, while generating\ncombined diff, a lot of time is spent computing diff(commit,commit^2)\njust to only then intersect that huge diff to almost small set of files\nfrom diff(commit,commit^1).\n\nThat's because at present, to compute combine-diff, for first finding\npaths, that \"every parent touches\", we use the following combine-diff\nproperty/definition:\n\nD(A,P1...Pn) = D(A,P1) ^ ... ^ D(A,Pn)      (w.r.t. paths)\n\nwhere\n\nD(A,P1...Pn) is combined diff between commit A, and parents Pi\n\nand\n\nD(A,Pi) is usual two-tree diff Pi..A\n\nSo if any of that D(A,Pi) is huge, tracting 1 n-parent combine-diff as n\n1-parent diffs and intersecting results will be slow.\n\nAnd usually, for linux.git and other topic-based workflows, that\nD(A,P2) is huge, because, if merge-base of A and P2, is several dozens\nof merges (from A, via first parent) below, that D(A,P2) will be diffing\nsum of merges from several subsystems to 1 subsystem.\n\nThe solution is to avoid computing n 1-parent diffs, and to find\nchanged-to-all-parents paths via scanning A's and all Pi's trees\nsimultaneously, at each step comparing their entries, and based on that\ncomparison, populate paths result, and deduce we could *skip*\n*recursing* into subdirectories, if at least for 1 parent, sha1 of that\ndir tree is the same as in A. That would save us from doing significant\namount of needless work.\n\nSuch approach is very similar to what diff_tree() does, only there we\ndeal with scanning only 2 trees simultaneously, and for n+1 tree, the\nlogic is a bit more complex:\n\nD(A,X1...Xn) calculation scheme\n-------------------------------\n\nD(A,X1...Xn) = D(A,X1) ^ ... ^ D(A,Xn)       (regarding resulting paths set)\n\n     D(A,Xj)         - diff between A..Xj\n     D(A,X1...Xn)    - combined diff from A to parents X1,...,Xn\n\nWe start from all trees, which are sorted, and compare their entries in\nlock-step:\n\n      A     X1       Xn\n      -     -        -\n     |a|   |x1|     |xn|\n     |-|   |--| ... |--|      i = argmin(x1...xn)\n     | |   |  |     |  |\n     |-|   |--|     |--|\n     |.|   |. |     |. |\n      .     .        .\n      .     .        .\n\nat any time there could be 3 cases:\n\n     1)  a < xi;\n     2)  a > xi;\n     3)  a = xi.\n\nSchematic deduction of what every case means, and what to do, follows:\n\n1)  a < xi  ->  ∀j a ∉ Xj  ->  \"+a\" ∈ D(A,Xj)  ->  D += \"+a\";  a↓\n\n2)  a > xi\n\n    2.1) ∃j: xj > xi  ->  \"-xi\" ∉ D(A,Xj)  ->  D += ø;  ∀ xk=xi  xk↓\n    2.2) ∀j  xj = xi  ->  xj ∉ A  ->  \"-xj\" ∈ D(A,Xj)  ->  D += \"-xi\";  ∀j xj↓\n\n3)  a = xi\n\n    3.1) ∃j: xj > xi  ->  \"+a\" ∈ D(A,Xj)  ->  only xk=xi remains to investigate\n    3.2) xj = xi  ->  investigate δ(a,xj)\n     |\n     |\n     v\n\n    3.1+3.2) looking at δ(a,xk) ∀k: xk=xi - if all != ø  ->\n\n                      ⎧δ(a,xk)  - if xk=xi\n             ->  D += ⎨\n                      ⎩\"+a\"     - if xk>xi\n\n    in any case a↓  ∀ xk=xi  xk↓\n\n~\n\nFor comparison, here is how diff_tree() works:\n\nD(A,B) calculation scheme\n-------------------------\n\n    A     B\n    -     -\n   |a|   |b|    a < b   ->  a ∉ B   ->   D(A,B) +=  +a    a↓\n   |-|   |-|    a > b   ->  b ∉ A   ->   D(A,B) +=  -b    b↓\n   | |   | |    a = b   ->  investigate δ(a,b)            a↓ b↓\n   |-|   |-|\n   |.|   |.|\n    .     .\n    .     .\n\n~~~~~~~~\n\nThis patch generalizes diff tree-walker to work with arbitrary number of\nparents as described above - i.e. now there is a resulting tree t, and\nsome parents trees tp[i] i=[0..nparent). The generalization builds on\nthe fact that usual diff\n\nD(A,B)\n\nis by definition the same as combined diff\n\nD(A,[B]),\n\nso if we could rework the code for common case and make it be not slower\nfor nparent=1 case, usual diff(t1,t2) generation will not be slower, and\nmultiparent diff tree-walker would greatly benefit generating\ncombine-diff.\n\nWhat we do is as follows:\n\n1) diff tree-walker ll_diff_tree_sha1() is internally reworked to be\n   a paths generator (new name diff_tree_paths()), with each generated path\n   being `struct combine_diff_path` with info for path, new sha1,mode and for\n   every parent which sha1,mode it was in it.\n\n2) From that info, we can still generate usual diff queue with\n   struct diff_filepairs, via \"exporting\" generated\n   combine_diff_path, if we know we run for nparent=1 case.\n   (see emit_diff() which is now named emit_diff_first_parent_only())\n\n3) In order for diff_can_quit_early(), which checks\n\n       DIFF_OPT_TST(opt, HAS_CHANGES))\n\n   to work, that exporting have to be happening not in bulk, but\n   incrementally, one diff path at a time.\n\n   For such consumers, there is a new callback in diff_options\n   introduced:\n\n       ->pathchange(opt, struct combine_diff_path *)\n\n   which, if set to !NULL, is called for every generated path.\n\n   (see new compat ll_diff_tree_sha1() wrapper around new paths\n    generator for setup)\n\n4) The paths generation itself, is reworked from previous\n   ll_diff_tree_sha1() code according to \"D(A,X1...Xn) calculation\n   scheme\" provided above:\n\n   On the start we allocate [nparent] arrays in place what was\n   earlier just for one parent tree.\n\n   then we just generalize loops, and comparison according to the\n   algorithm.\n\nSome notes(*):\n\n1) alloca(), for small arrays, is used for \"runs not slower for\n   nparent=1 case than before\" goal - if we change it to xmalloc()/free()\n   the timings get ~1% worse. For alloca() we use just-introduced\n   xalloca/xalloca_free compatibility wrappers, so it should not be a\n   portability problem.\n\n2) For every parent tree, we need to keep a tag, whether entry from that\n   parent equals to entry from minimal parent. For performance reasons I'm\n   keeping that tag in entry's mode field in unused bit - see S_IFXMIN_NEQ.\n   Not doing so, we'd need to alloca another [nparent] array, which hurts\n   performance.\n\n3) For emitted paths, memory could be reused, if we know the path was\n   processed via callback and will not be needed later. We use efficient\n   hand-made realloc-style path_appendnew(), that saves us from ~1-1.5%\n   of potential additional slowdown.\n\n4) goto(s) are used in several places, as the code executes a little bit\n   faster with lowered register pressure.\n\nAlso\n\n- we should now check for FIND_COPIES_HARDER not only when two entries\n  names are the same, and their hashes are equal, but also for a case,\n  when a path was removed from some of all parents having it.\n\n  The reason is, if we don't, that path won't be emitted at all (see\n  \"a > xi\" case), and we'll just skip it, and FIND_COPIES_HARDER wants\n  all paths - with diff or without - to be emitted, to be later analyzed\n  for being copies sources.\n\n  The new check is only necessary for nparent >1, as for nparent=1 case\n  xmin_eqtotal always =1 =nparent, and a path is always added to diff as\n  removal.\n\n~~~~~~~~\n\nTimings for\n\n    # without -c, i.e. testing only nparent=1 case\n    `git log --raw --no-abbrev --no-renames`\n\nbefore and after the patch are as follows:\n\n                navy.git        linux.git v3.10..v3.11\n\n    before      0.611s          1.889s\n    after       0.619s          1.907s\n    slowdown    1.3%            0.9%\n\nThis timings show we did no harm to usual diff(tree1,tree2) generation.\nFrom the table we can see that we actually did ~1% slowdown, but I think\nI've \"earned\" that 1% in the previous patch (\"tree-diff: reuse base\nstr(buf) memory on sub-tree recursion\", HEAD~~) so for nparent=1 case,\nnet timings stays approximately the same.\n\nThe output also stayed the same.\n\n(*) If we revert 1)-4) to more usual techniques, for nparent=1 case,\n    we'll get ~2-2.5% of additional slowdown, which I've tried to avoid, as\n   \"do no harm for nparent=1 case\" rule.\n\nFor linux.git, combined diff will run an order of magnitude faster and\nappropriate timings will be provided in the next commit, as we'll be\ntaking advantage of the new diff tree-walker for combined-diff\ngeneration there.\n\nP.S. and combined diff is not some exotic/for-play-only stuff - for\nexample for a program I write to represent Git archives as readonly\nfilesystem, there is initial scan with\n\n    `git log --reverse --raw --no-abbrev --no-renames -c`\n\nto extract log of what was created/changed when, as a result building a\nmap\n\n    {}  sha1    ->  in which commit (and date) a content was added\n\nthat `-c` means also show combined diff for merges, and without them, if\na merge is non-trivial (merges changes from two parents with both having\nseparate changes to a file), or an evil one, the map will not be full,\ni.e. some valid sha1 would be absent from it.\n\nThat case was my initial motivation for combined diffs speedup.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\nChanges since v2:\n\n- rename:\n    __diff_tree_paths   -> ll_diff_tree_paths\n    __path_append_new   -> path_append_new\n\n  and adjust to previous renaming\n    __diff_tree_sha1() -> ll_diff_tree_sha1()\n\n  to be consistent with Git coding style for naming low-level helpers\n  (ll_ prefix) and not using \"__\" as it overlaps with\n  reserved-for-implementation identifier namespace.\n\n  path_append_new goes without \"ll_\" prefix as it is the only function\n  with no high-level counterpart.\n\nChanges since v1:\n\n- fixed last-minute thinko/bug last time introduced on my side (sorry) with\n  opt->pathchange manipulation in __diff_tree_sha1() - we were forgetting to\n  restore opt->pathchange, which led to incorrect log -c (merges _and_ plain\n  diff-tree) output;\n\n  This time, I've verified several times, log output stays really the same.\n\n- direct use of alloca() changed to portability wrappers xalloca/xalloca_free\n  which gracefully degrade to xmalloc/free on systems, where alloca is not\n  available (see new patch 17).\n\n- \"i = 0; do { ... } while (++i < nparent)\" is back to usual looping\n  \"for (i = 0; i < nparent; ++)\", as I've re-measured timings and the\n  difference is negligible.\n\n  ( Initially, when I was fighting for every cycle it made sense, but real\n    no-slowdown turned out to be related to avoiding mallocs, load trees in correct\n    order and reducing register pressure. )\n\n- S_IFXMIN_NEQ definition moved out to cache.h, to have all modes registry in one place;\n\n\n- p0 -> first_parent; corrected comments about how emit_diff_first_parent_only\n  behaves;\n cache.h     |  15 ++\n diff.c      |   1 +\n diff.h      |  10 ++\n tree-diff.c | 505 ++++++++++++++++++++++++++++++++++++++++++++++++++++--------\n 4 files changed, 467 insertions(+), 64 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex dc040fb..e7f5a0c 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -75,6 +75,21 @@ unsigned long git_deflate_bound(git_zstream *, unsigned long);\n #define S_ISGITLINK(m)\t(((m) & S_IFMT) == S_IFGITLINK)\n \n /*\n+ * Some mode bits are also used internally for computations.\n+ *\n+ * They *must* not overlap with any valid modes, and they *must* not be emitted\n+ * to outside world - i.e. appear on disk or network. In other words, it's just\n+ * temporary fields, which we internally use, but they have to stay in-house.\n+ *\n+ * ( such approach is valid, as standard S_IF* fits into 16 bits, and in Git\n+ *   codebase mode is `unsigned int` which is assumed to be at least 32 bits )\n+ */\n+\n+/* used internally in tree-diff */\n+#define S_DIFFTREE_IFXMIN_NEQ\t0x80000000\n+\n+\n+/*\n  * Intensive research over the course of many years has shown that\n  * port 9418 is totally unused by anything else. Or\n  *\ndiff --git a/diff.c b/diff.c\nindex 8e4a6a9..cda4aa8 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -3216,6 +3216,7 @@ void diff_setup(struct diff_options *options)\n \toptions->context = diff_context_default;\n \tDIFF_OPT_SET(options, RENAME_EMPTY);\n \n+\t/* pathchange left =NULL by default */\n \toptions->change = diff_change;\n \toptions->add_remove = diff_addremove;\n \toptions->use_color = diff_use_color_default;\ndiff --git a/diff.h b/diff.h\nindex 5d7b9f7..732dca7 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -15,6 +15,10 @@ struct diff_filespec;\n struct userdiff_driver;\n struct sha1_array;\n struct commit;\n+struct combine_diff_path;\n+\n+typedef int (*pathchange_fn_t)(struct diff_options *options,\n+\t\t struct combine_diff_path *path);\n \n typedef void (*change_fn_t)(struct diff_options *options,\n \t\t unsigned old_mode, unsigned new_mode,\n@@ -157,6 +161,7 @@ struct diff_options {\n \tint close_file;\n \n \tstruct pathspec pathspec;\n+\tpathchange_fn_t pathchange;\n \tchange_fn_t change;\n \tadd_remove_fn_t add_remove;\n \tdiff_format_fn_t format_callback;\n@@ -189,6 +194,11 @@ const char *diff_line_prefix(struct diff_options *);\n \n extern const char mime_boundary_leader[];\n \n+extern\n+struct combine_diff_path *diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parent_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt);\n extern int diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \t\t\t  const char *base, struct diff_options *opt);\n extern int diff_root_tree_sha1(const unsigned char *new, const char *base,\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 8c8bde6..4a497f7 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -6,7 +6,19 @@\n #include \"diffcore.h\"\n #include \"tree.h\"\n \n+/*\n+ * internal mode marker, saying a tree entry != entry of tp[imin]\n+ * (see ll_diff_tree_paths for what it means there)\n+ *\n+ * we will update/use/emit entry for diff only with it unset.\n+ */\n+#define S_IFXMIN_NEQ\tS_DIFFTREE_IFXMIN_NEQ\n+\n \n+static struct combine_diff_path *ll_diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parents_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt);\n static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \t\t\t     struct strbuf *base, struct diff_options *opt);\n \n@@ -42,71 +54,151 @@ static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n }\n \n \n-/* convert path, t1/t2 -> opt->diff_*() callbacks */\n-static void emit_diff(struct diff_options *opt, struct strbuf *path,\n-\t\t      struct tree_desc *t1, struct tree_desc *t2)\n+/*\n+ * convert path -> opt->diff_*() callbacks\n+ *\n+ * emits diff to first parent only, and tells diff tree-walker that we are done\n+ * with p and it can be freed.\n+ */\n+static int emit_diff_first_parent_only(struct diff_options *opt, struct combine_diff_path *p)\n {\n-\tunsigned int mode1 = t1 ? t1->entry.mode : 0;\n-\tunsigned int mode2 = t2 ? t2->entry.mode : 0;\n-\n-\tif (mode1 && mode2) {\n-\t\topt->change(opt, mode1, mode2, t1->entry.sha1, t2->entry.sha1,\n-\t\t\t1, 1, path->buf, 0, 0);\n+\tstruct combine_diff_parent *p0 = &p->parent[0];\n+\tif (p->mode && p0->mode) {\n+\t\topt->change(opt, p0->mode, p->mode, p0->sha1, p->sha1,\n+\t\t\t1, 1, p->path, 0, 0);\n \t}\n \telse {\n \t\tconst unsigned char *sha1;\n \t\tunsigned int mode;\n \t\tint addremove;\n \n-\t\tif (mode2) {\n+\t\tif (p->mode) {\n \t\t\taddremove = '+';\n-\t\t\tsha1 = t2->entry.sha1;\n-\t\t\tmode = mode2;\n+\t\t\tsha1 = p->sha1;\n+\t\t\tmode = p->mode;\n \t\t} else {\n \t\t\taddremove = '-';\n-\t\t\tsha1 = t1->entry.sha1;\n-\t\t\tmode = mode1;\n+\t\t\tsha1 = p0->sha1;\n+\t\t\tmode = p0->mode;\n \t\t}\n \n-\t\topt->add_remove(opt, addremove, mode, sha1, 1, path->buf, 0);\n+\t\topt->add_remove(opt, addremove, mode, sha1, 1, p->path, 0);\n \t}\n+\n+\treturn 0;\t/* we are done with p */\n }\n \n \n-/* new path should be added to diff\n+/*\n+ * Make a new combine_diff_path from path/mode/sha1\n+ * and append it to paths list tail.\n+ *\n+ * Memory for created elements could be reused:\n+ *\n+ *\t- if last->next == NULL, the memory is allocated;\n+ *\n+ *\t- if last->next != NULL, it is assumed that p=last->next was returned\n+ *\t  earlier by this function, and p->next was *not* modified.\n+ *\t  The memory is then reused from p.\n+ *\n+ * so for clients,\n+ *\n+ * - if you do need to keep the element\n+ *\n+ *\tp = path_appendnew(p, ...);\n+ *\tprocess(p);\n+ *\tp->next = NULL;\n+ *\n+ * - if you don't need to keep the element after processing\n+ *\n+ *\tpprev = p;\n+ *\tp = path_appendnew(p, ...);\n+ *\tprocess(p);\n+ *\tp = pprev;\n+ *\t; don't forget to free tail->next in the end\n+ *\n+ * p->parent[] remains uninitialized.\n+ */\n+static struct combine_diff_path *path_appendnew(struct combine_diff_path *last,\n+\tint nparent, const struct strbuf *base, const char *path, int pathlen,\n+\tunsigned mode, const unsigned char *sha1)\n+{\n+\tstruct combine_diff_path *p;\n+\tint len = base->len + pathlen;\n+\tint alloclen = combine_diff_path_size(nparent, len);\n+\n+\t/* if last->next is !NULL - it is a pre-allocated memory, we can reuse */\n+\tp = last->next;\n+\tif (p && (alloclen > (intptr_t)p->next)) {\n+\t\tfree(p);\n+\t\tp = NULL;\n+\t}\n+\n+\tif (!p) {\n+\t\tp = xmalloc(alloclen);\n+\n+\t\t/*\n+\t\t * until we go to it next round, .next holds how many bytes we\n+\t\t * allocated (for faster realloc - we don't need copying old data).\n+\t\t */\n+\t\tp->next = (struct combine_diff_path *)(intptr_t)alloclen;\n+\t}\n+\n+\tlast->next = p;\n+\n+\tp->path = (char *)&(p->parent[nparent]);\n+\tmemcpy(p->path, base->buf, base->len);\n+\tmemcpy(p->path + base->len, path, pathlen);\n+\tp->path[len] = 0;\n+\tp->mode = mode;\n+\thashcpy(p->sha1, sha1 ? sha1 : null_sha1);\n+\n+\treturn p;\n+}\n+\n+/*\n+ * new path should be added to combine diff\n  *\n  * 3 cases on how/when it should be called and behaves:\n  *\n- *\t!t1,  t2\t-> path added, parent lacks it\n- *\t t1, !t2\t-> path removed from parent\n- *\t t1,  t2\t-> path modified\n+ *\t t, !tp\t\t-> path added, all parents lack it\n+ *\t!t,  tp\t\t-> path removed from all parents\n+ *\t t,  tp\t\t-> path modified/added\n+ *\t\t\t   (M for tp[i]=tp[imin], A otherwise)\n  */\n-static void show_path(struct strbuf *base, struct diff_options *opt,\n-\t\t      struct tree_desc *t1, struct tree_desc *t2)\n+static struct combine_diff_path *emit_path(struct combine_diff_path *p,\n+\tstruct strbuf *base, struct diff_options *opt, int nparent,\n+\tstruct tree_desc *t, struct tree_desc *tp,\n+\tint imin)\n {\n \tunsigned mode;\n \tconst char *path;\n+\tconst unsigned char *sha1;\n \tint pathlen;\n \tint old_baselen = base->len;\n-\tint isdir, recurse = 0, emitthis = 1;\n+\tint i, isdir, recurse = 0, emitthis = 1;\n \n \t/* at least something has to be valid */\n-\tassert(t1 || t2);\n+\tassert(t || tp);\n \n-\tif (t2) {\n+\tif (t) {\n \t\t/* path present in resulting tree */\n-\t\ttree_entry_extract(t2, &path, &mode);\n-\t\tpathlen = tree_entry_len(&t2->entry);\n+\t\tsha1 = tree_entry_extract(t, &path, &mode);\n+\t\tpathlen = tree_entry_len(&t->entry);\n \t\tisdir = S_ISDIR(mode);\n \t} else {\n \t\t/*\n-\t\t * a path was removed - take path from parent. Also take\n-\t\t * mode from parent, to decide on recursion.\n+\t\t * a path was removed - take path from imin parent. Also take\n+\t\t * mode from that parent, to decide on recursion(1).\n+\t\t *\n+\t\t * 1) all modes for tp[k]=tp[imin] should be the same wrt\n+\t\t *    S_ISDIR, thanks to base_name_compare().\n \t\t */\n-\t\ttree_entry_extract(t1, &path, &mode);\n-\t\tpathlen = tree_entry_len(&t1->entry);\n+\t\ttree_entry_extract(&tp[imin], &path, &mode);\n+\t\tpathlen = tree_entry_len(&tp[imin].entry);\n \n \t\tisdir = S_ISDIR(mode);\n+\t\tsha1 = NULL;\n \t\tmode = 0;\n \t}\n \n@@ -115,18 +207,81 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n \t\temitthis = DIFF_OPT_TST(opt, TREE_IN_RECURSIVE);\n \t}\n \n-\tstrbuf_add(base, path, pathlen);\n+\tif (emitthis) {\n+\t\tint keep;\n+\t\tstruct combine_diff_path *pprev = p;\n+\t\tp = path_appendnew(p, nparent, base, path, pathlen, mode, sha1);\n+\n+\t\tfor (i = 0; i < nparent; ++i) {\n+\t\t\t/*\n+\t\t\t * tp[i] is valid, if present and if tp[i]==tp[imin] -\n+\t\t\t * otherwise, we should ignore it.\n+\t\t\t */\n+\t\t\tint tpi_valid = tp && !(tp[i].entry.mode & S_IFXMIN_NEQ);\n+\n+\t\t\tconst unsigned char *sha1_i;\n+\t\t\tunsigned mode_i;\n+\n+\t\t\tp->parent[i].status =\n+\t\t\t\t!t ? DIFF_STATUS_DELETED :\n+\t\t\t\t\ttpi_valid ?\n+\t\t\t\t\t\tDIFF_STATUS_MODIFIED :\n+\t\t\t\t\t\tDIFF_STATUS_ADDED;\n+\n+\t\t\tif (tpi_valid) {\n+\t\t\t\tsha1_i = tp[i].entry.sha1;\n+\t\t\t\tmode_i = tp[i].entry.mode;\n+\t\t\t}\n+\t\t\telse {\n+\t\t\t\tsha1_i = NULL;\n+\t\t\t\tmode_i = 0;\n+\t\t\t}\n+\n+\t\t\tp->parent[i].mode = mode_i;\n+\t\t\thashcpy(p->parent[i].sha1, sha1_i ? sha1_i : null_sha1);\n+\t\t}\n \n-\tif (emitthis)\n-\t\temit_diff(opt, base, t1, t2);\n+\t\tkeep = 1;\n+\t\tif (opt->pathchange)\n+\t\t\tkeep = opt->pathchange(opt, p);\n+\n+\t\t/*\n+\t\t * If a path was filtered or consumed - we don't need to add it\n+\t\t * to the list and can reuse its memory, leaving it as\n+\t\t * pre-allocated element on the tail.\n+\t\t *\n+\t\t * On the other hand, if path needs to be kept, we need to\n+\t\t * correct its .next to NULL, as it was pre-initialized to how\n+\t\t * much memory was allocated.\n+\t\t *\n+\t\t * see path_appendnew() for details.\n+\t\t */\n+\t\tif (!keep)\n+\t\t\tp = pprev;\n+\t\telse\n+\t\t\tp->next = NULL;\n+\t}\n \n \tif (recurse) {\n+\t\tconst unsigned char **parents_sha1;\n+\n+\t\tparents_sha1 = xalloca(nparent * sizeof(parents_sha1[0]));\n+\t\tfor (i = 0; i < nparent; ++i) {\n+\t\t\t/* same rule as in emitthis */\n+\t\t\tint tpi_valid = tp && !(tp[i].entry.mode & S_IFXMIN_NEQ);\n+\n+\t\t\tparents_sha1[i] = tpi_valid ? tp[i].entry.sha1\n+\t\t\t\t\t\t    : NULL;\n+\t\t}\n+\n+\t\tstrbuf_add(base, path, pathlen);\n \t\tstrbuf_addch(base, '/');\n-\t\tll_diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n-\t\t\t\t  t2 ? t2->entry.sha1 : NULL, base, opt);\n+\t\tp = ll_diff_tree_paths(p, sha1, parents_sha1, nparent, base, opt);\n+\t\txalloca_free(parents_sha1);\n \t}\n \n \tstrbuf_setlen(base, old_baselen);\n+\treturn p;\n }\n \n static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n@@ -145,59 +300,260 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n \t}\n }\n \n-static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n-\t\t\t     struct strbuf *base, struct diff_options *opt)\n+\n+/*\n+ * generate paths for combined diff D(sha1,parents_sha1[])\n+ *\n+ * Resulting paths are appended to combine_diff_path linked list, and also, are\n+ * emitted on the go via opt->pathchange() callback, so it is possible to\n+ * process the result as batch or incrementally.\n+ *\n+ * The paths are generated scanning new tree and all parents trees\n+ * simultaneously, similarly to what diff_tree() was doing for 2 trees.\n+ * The theory behind such scan is as follows:\n+ *\n+ *\n+ * D(A,X1...Xn) calculation scheme\n+ * -------------------------------\n+ *\n+ * D(A,X1...Xn) = D(A,X1) ^ ... ^ D(A,Xn)\t(regarding resulting paths set)\n+ *\n+ *\tD(A,Xj)\t\t- diff between A..Xj\n+ *\tD(A,X1...Xn)\t- combined diff from A to parents X1,...,Xn\n+ *\n+ *\n+ * We start from all trees, which are sorted, and compare their entries in\n+ * lock-step:\n+ *\n+ *\t A     X1       Xn\n+ *\t -     -        -\n+ *\t|a|   |x1|     |xn|\n+ *\t|-|   |--| ... |--|      i = argmin(x1...xn)\n+ *\t| |   |  |     |  |\n+ *\t|-|   |--|     |--|\n+ *\t|.|   |. |     |. |\n+ *\t .     .        .\n+ *\t .     .        .\n+ *\n+ * at any time there could be 3 cases:\n+ *\n+ *\t1)  a < xi;\n+ *\t2)  a > xi;\n+ *\t3)  a = xi.\n+ *\n+ * Schematic deduction of what every case means, and what to do, follows:\n+ *\n+ * 1)  a < xi  ->  ∀j a ∉ Xj  ->  \"+a\" ∈ D(A,Xj)  ->  D += \"+a\";  a↓\n+ *\n+ * 2)  a > xi\n+ *\n+ *     2.1) ∃j: xj > xi  ->  \"-xi\" ∉ D(A,Xj)  ->  D += ø;  ∀ xk=xi  xk↓\n+ *     2.2) ∀j  xj = xi  ->  xj ∉ A  ->  \"-xj\" ∈ D(A,Xj)  ->  D += \"-xi\";  ∀j xj↓\n+ *\n+ * 3)  a = xi\n+ *\n+ *     3.1) ∃j: xj > xi  ->  \"+a\" ∈ D(A,Xj)  ->  only xk=xi remains to investigate\n+ *     3.2) xj = xi  ->  investigate δ(a,xj)\n+ *      |\n+ *      |\n+ *      v\n+ *\n+ *     3.1+3.2) looking at δ(a,xk) ∀k: xk=xi - if all != ø  ->\n+ *\n+ *                       ⎧δ(a,xk)  - if xk=xi\n+ *              ->  D += ⎨\n+ *                       ⎩\"+a\"     - if xk>xi\n+ *\n+ *\n+ *     in any case a↓  ∀ xk=xi  xk↓\n+ *\n+ *\n+ * ~~~~~~~~\n+ *\n+ * NOTE\n+ *\n+ *\tUsual diff D(A,B) is by definition the same as combined diff D(A,[B]),\n+ *\tso this diff paths generator can, and is used, for plain diffs\n+ *\tgeneration too.\n+ *\n+ *\tPlease keep attention to the common D(A,[B]) case when working on the\n+ *\tcode, in order not to slow it down.\n+ *\n+ * NOTE\n+ *\tnparent must be > 0.\n+ */\n+\n+\n+/* ∀ xk=xi  xk↓ */\n+static inline void update_tp_entries(struct tree_desc *tp, int nparent)\n {\n-\tstruct tree_desc t1, t2;\n-\tvoid *t1tree, *t2tree;\n+\tint i;\n+\tfor (i = 0; i < nparent; ++i)\n+\t\tif (!(tp[i].entry.mode & S_IFXMIN_NEQ))\n+\t\t\tupdate_tree_entry(&tp[i]);\n+}\n \n-\tt1tree = fill_tree_descriptor(&t1, old);\n-\tt2tree = fill_tree_descriptor(&t2, new);\n+static struct combine_diff_path *ll_diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parents_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt)\n+{\n+\tstruct tree_desc t, *tp;\n+\tvoid *ttree, **tptree;\n+\tint i;\n+\n+\ttp     = xalloca(nparent * sizeof(tp[0]));\n+\ttptree = xalloca(nparent * sizeof(tptree[0]));\n+\n+\t/*\n+\t * load parents first, as they are probably already cached.\n+\t *\n+\t * ( log_tree_diff() parses commit->parent before calling here via\n+\t *   diff_tree_sha1(parent, commit) )\n+\t */\n+\tfor (i = 0; i < nparent; ++i)\n+\t\ttptree[i] = fill_tree_descriptor(&tp[i], parents_sha1[i]);\n+\tttree = fill_tree_descriptor(&t, sha1);\n \n \t/* Enable recursion indefinitely */\n \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n \n \tfor (;;) {\n-\t\tint cmp;\n+\t\tint imin, cmp;\n \n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n+\n \t\tif (opt->pathspec.nr) {\n-\t\t\tskip_uninteresting(&t1, base, opt);\n-\t\t\tskip_uninteresting(&t2, base, opt);\n+\t\t\tskip_uninteresting(&t, base, opt);\n+\t\t\tfor (i = 0; i < nparent; i++)\n+\t\t\t\tskip_uninteresting(&tp[i], base, opt);\n \t\t}\n-\t\tif (!t1.size && !t2.size)\n-\t\t\tbreak;\n \n-\t\tcmp = tree_entry_pathcmp(&t1, &t2);\n+\t\t/* comparing is finished when all trees are done */\n+\t\tif (!t.size) {\n+\t\t\tint done = 1;\n+\t\t\tfor (i = 0; i < nparent; ++i)\n+\t\t\t\tif (tp[i].size) {\n+\t\t\t\t\tdone = 0;\n+\t\t\t\t\tbreak;\n+\t\t\t\t}\n+\t\t\tif (done)\n+\t\t\t\tbreak;\n+\t\t}\n+\n+\t\t/*\n+\t\t * lookup imin = argmin(x1...xn),\n+\t\t * mark entries whether they =tp[imin] along the way\n+\t\t */\n+\t\timin = 0;\n+\t\ttp[0].entry.mode &= ~S_IFXMIN_NEQ;\n+\n+\t\tfor (i = 1; i < nparent; ++i) {\n+\t\t\tcmp = tree_entry_pathcmp(&tp[i], &tp[imin]);\n+\t\t\tif (cmp < 0) {\n+\t\t\t\timin = i;\n+\t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n+\t\t\t}\n+\t\t\telse if (cmp == 0) {\n+\t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n+\t\t\t}\n+\t\t\telse {\n+\t\t\t\ttp[i].entry.mode |= S_IFXMIN_NEQ;\n+\t\t\t}\n+\t\t}\n+\n+\t\t/* fixup markings for entries before imin */\n+\t\tfor (i = 0; i < imin; ++i)\n+\t\t\ttp[i].entry.mode |= S_IFXMIN_NEQ;\t/* x[i] > x[imin] */\n+\n \n-\t\t/* t1 = t2 */\n-\t\tif (cmp == 0) {\n-\t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n-\t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n-\t\t\t    (t1.entry.mode != t2.entry.mode))\n-\t\t\t\tshow_path(base, opt, &t1, &t2);\n \n-\t\t\tupdate_tree_entry(&t1);\n-\t\t\tupdate_tree_entry(&t2);\n+\t\t/* compare a vs x[imin] */\n+\t\tcmp = tree_entry_pathcmp(&t, &tp[imin]);\n+\n+\t\t/* a = xi */\n+\t\tif (cmp == 0) {\n+\t\t\t/* are either xk > xi or diff(a,xk) != ø ? */\n+\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n+\t\t\t\tfor (i = 0; i < nparent; ++i) {\n+\t\t\t\t\t/* x[i] > x[imin] */\n+\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n+\t\t\t\t\t\tcontinue;\n+\n+\t\t\t\t\t/* diff(a,xk) != ø */\n+\t\t\t\t\tif (hashcmp(t.entry.sha1, tp[i].entry.sha1) ||\n+\t\t\t\t\t    (t.entry.mode != tp[i].entry.mode))\n+\t\t\t\t\t\tcontinue;\n+\n+\t\t\t\t\tgoto skip_emit_t_tp;\n+\t\t\t\t}\n+\t\t\t}\n+\n+\t\t\t/* D += {δ(a,xk) if xk=xi;  \"+a\" if xk > xi} */\n+\t\t\tp = emit_path(p, base, opt, nparent,\n+\t\t\t\t\t&t, tp, imin);\n+\n+\t\tskip_emit_t_tp:\n+\t\t\t/* a↓,  ∀ xk=ximin  xk↓ */\n+\t\t\tupdate_tree_entry(&t);\n+\t\t\tupdate_tp_entries(tp, nparent);\n \t\t}\n \n-\t\t/* t1 < t2 */\n+\t\t/* a < xi */\n \t\telse if (cmp < 0) {\n-\t\t\tshow_path(base, opt, &t1, /*t2=*/NULL);\n-\t\t\tupdate_tree_entry(&t1);\n+\t\t\t/* D += \"+a\" */\n+\t\t\tp = emit_path(p, base, opt, nparent,\n+\t\t\t\t\t&t, /*tp=*/NULL, -1);\n+\n+\t\t\t/* a↓ */\n+\t\t\tupdate_tree_entry(&t);\n \t\t}\n \n-\t\t/* t1 > t2 */\n+\t\t/* a > xi */\n \t\telse {\n-\t\t\tshow_path(base, opt, /*t1=*/NULL, &t2);\n-\t\t\tupdate_tree_entry(&t2);\n+\t\t\t/* ∀j xj=ximin -> D += \"-xi\" */\n+\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n+\t\t\t\tfor (i = 0; i < nparent; ++i)\n+\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n+\t\t\t\t\t\tgoto skip_emit_tp;\n+\t\t\t}\n+\n+\t\t\tp = emit_path(p, base, opt, nparent,\n+\t\t\t\t\t/*t=*/NULL, tp, imin);\n+\n+\t\tskip_emit_tp:\n+\t\t\t/* ∀ xk=ximin  xk↓ */\n+\t\t\tupdate_tp_entries(tp, nparent);\n \t\t}\n \t}\n \n-\tfree(t2tree);\n-\tfree(t1tree);\n-\treturn 0;\n+\tfree(ttree);\n+\tfor (i = nparent-1; i >= 0; i--)\n+\t\tfree(tptree[i]);\n+\txalloca_free(tptree);\n+\txalloca_free(tp);\n+\n+\treturn p;\n+}\n+\n+struct combine_diff_path *diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parents_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt)\n+{\n+\tp = ll_diff_tree_paths(p, sha1, parents_sha1, nparent, base, opt);\n+\n+\t/*\n+\t * free pre-allocated last element, if any\n+\t * (see path_appendnew() for details about why)\n+\t */\n+\tif (p->next) {\n+\t\tfree(p->next);\n+\t\tp->next = NULL;\n+\t}\n+\n+\treturn p;\n }\n \n /*\n@@ -308,6 +664,27 @@ static void try_to_follow_renames(const unsigned char *old, const unsigned char\n \tq->nr = 1;\n }\n \n+static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n+\t\t\t     struct strbuf *base, struct diff_options *opt)\n+{\n+\tstruct combine_diff_path phead, *p;\n+\tconst unsigned char *parents_sha1[1] = {old};\n+\tpathchange_fn_t pathchange_old = opt->pathchange;\n+\n+\tphead.next = NULL;\n+\topt->pathchange = emit_diff_first_parent_only;\n+\tdiff_tree_paths(&phead, new, parents_sha1, 1, base, opt);\n+\n+\tfor (p = phead.next; p;) {\n+\t\tstruct combine_diff_path *pprev = p;\n+\t\tp = p->next;\n+\t\tfree(pprev);\n+\t}\n+\n+\topt->pathchange = pathchange_old;\n+\treturn 0;\n+}\n+\n int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base_str, struct diff_options *opt)\n {\n \tstruct strbuf base;\n-- \n1.9.rc0.143.g6fd479e\n"},{"id":"237902","messageId":"20140327142438.GE17333@mini.zxlink","threadId":"35947","inReplyTo":"xmqq1txoiqzj.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-27T14:24:38Z","receivedAt":"2014-03-27T14:24:38Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Wed, Mar 26, 2014 at 02:34:24PM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n> \n> > On Tue, Mar 25, 2014 at 10:46:32AM -0700, Junio C Hamano wrote:\n> >> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n> >> \n> >> > What are the downsides of \"__\" prefix by the way?\n> >> \n> >> Aren't these names reserved for compiler/runtime implementations?\n> >\n> > Yes, but there are precedents when people don't obey it widely and\n> > in practice everything works :)\n> \n> I think you are alluding to the practice in the Linux kernel, but\n> their requirement is vastly different---their product do not even\n> link with libc and they always compile with specific selected\n> versions of gcc, no?\n\nYes, that is correct. Only \"__\" was so visually appealing that there was\na temptation to break the rules, but...\n\n\n> > Let it be something portable anyway -\n> > how about diff_tree_sha1_low() ?\n> \n> Sure.\n> \n> As this is a file-scope static, I do not think the exact naming\n> matters that much.  Just FYI, we seem to use ll_ prefix (standing\n> for low-level) in some places.\n\n... let's then use this \"ll_\" prefix scheme for consistency.\n\nCorrected patch is below, and I've sent corrections to follow-up\npatches as well.\n\nThanks,\nKirill\n\n(please keep author email)\n---- 8< ----\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\nDate: Mon, 24 Feb 2014 20:21:46 +0400\nSubject: [PATCH v3a] tree-diff: rework diff_tree interface to be sha1 based\n\nIn the next commit this will allow to reduce intermediate calls, when\nrecursing into subtrees - at that stage we know only subtree sha1, and\nit is natural for tree walker to start from that phase. For now we do\n\n    diff_tree\n        show_path\n            diff_tree_sha1\n                diff_tree\n                    ...\n\nand the change will allow to reduce it to\n\n    diff_tree\n        show_path\n            diff_tree\n\nAlso, it will allow to omit allocating strbuf for each subtree, and just\nreuse the common strbuf via playing with its len.\n\nThe above-mentioned improvements go in the next 2 patches.\n\nThe downside is that try_to_follow_renames(), if active, we cause\nre-reading of 2 initial trees, which was negligible based on my timings,\nand which is outweighed cogently by the upsides.\n\nNOTE To keep with the current interface and semantics, I needed to\nrename the function from diff_tree() to diff_tree_sha1(). As\ndiff_tree_sha1() was already used, and the function we are talking here\nis its more low-level helper, let's use convention for prefixing\nsuch helpers with \"ll_\". So the final renaming is\n\n    diff_tree() -> ll_diff_tree_sha1()\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n\nChanges since v3:\n\n - further rename diff_tree_sha1_low() -> ll_diff_tree_sha1() to follow Git\n   style for naming low-level helpers.\n\nChanges since v2:\n\n - renamed __diff_tree_sha1() -> diff_tree_sha1_low() as the former\n   overlaps with reserved-for-implementation identifiers namespace.\n\nChanges since v1:\n\n - don't need to touch diff.h, as diff_tree() became static.\n\n\n tree-diff.c | 60 ++++++++++++++++++++++++++++--------------------------------\n 1 file changed, 28 insertions(+), 32 deletions(-)\n\ndiff --git a/tree-diff.c b/tree-diff.c\nindex f137f39..1d02e43 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -141,12 +141,17 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n \t}\n }\n \n-static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n-\t\t     const char *base_str, struct diff_options *opt)\n+static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n+\t\t\t     const char *base_str, struct diff_options *opt)\n {\n+\tstruct tree_desc t1, t2;\n+\tvoid *t1tree, *t2tree;\n \tstruct strbuf base;\n \tint baselen = strlen(base_str);\n \n+\tt1tree = fill_tree_descriptor(&t1, old);\n+\tt2tree = fill_tree_descriptor(&t2, new);\n+\n \t/* Enable recursion indefinitely */\n \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n \n@@ -159,39 +164,41 @@ static int diff_tree(struct tree_desc *t1, struct tree_desc *t2,\n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \t\tif (opt->pathspec.nr) {\n-\t\t\tskip_uninteresting(t1, &base, opt);\n-\t\t\tskip_uninteresting(t2, &base, opt);\n+\t\t\tskip_uninteresting(&t1, &base, opt);\n+\t\t\tskip_uninteresting(&t2, &base, opt);\n \t\t}\n-\t\tif (!t1->size && !t2->size)\n+\t\tif (!t1.size && !t2.size)\n \t\t\tbreak;\n \n-\t\tcmp = tree_entry_pathcmp(t1, t2);\n+\t\tcmp = tree_entry_pathcmp(&t1, &t2);\n \n \t\t/* t1 = t2 */\n \t\tif (cmp == 0) {\n \t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n-\t\t\t    hashcmp(t1->entry.sha1, t2->entry.sha1) ||\n-\t\t\t    (t1->entry.mode != t2->entry.mode))\n-\t\t\t\tshow_path(&base, opt, t1, t2);\n+\t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n+\t\t\t    (t1.entry.mode != t2.entry.mode))\n+\t\t\t\tshow_path(&base, opt, &t1, &t2);\n \n-\t\t\tupdate_tree_entry(t1);\n-\t\t\tupdate_tree_entry(t2);\n+\t\t\tupdate_tree_entry(&t1);\n+\t\t\tupdate_tree_entry(&t2);\n \t\t}\n \n \t\t/* t1 < t2 */\n \t\telse if (cmp < 0) {\n-\t\t\tshow_path(&base, opt, t1, /*t2=*/NULL);\n-\t\t\tupdate_tree_entry(t1);\n+\t\t\tshow_path(&base, opt, &t1, /*t2=*/NULL);\n+\t\t\tupdate_tree_entry(&t1);\n \t\t}\n \n \t\t/* t1 > t2 */\n \t\telse {\n-\t\t\tshow_path(&base, opt, /*t1=*/NULL, t2);\n-\t\t\tupdate_tree_entry(t2);\n+\t\t\tshow_path(&base, opt, /*t1=*/NULL, &t2);\n+\t\t\tupdate_tree_entry(&t2);\n \t\t}\n \t}\n \n \tstrbuf_release(&base);\n+\tfree(t2tree);\n+\tfree(t1tree);\n \treturn 0;\n }\n \n@@ -206,7 +213,7 @@ static inline int diff_might_be_rename(void)\n \t\t!DIFF_FILE_VALID(diff_queued_diff.queue[0]->one);\n }\n \n-static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, const char *base, struct diff_options *opt)\n+static void try_to_follow_renames(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n {\n \tstruct diff_options diff_opts;\n \tstruct diff_queue_struct *q = &diff_queued_diff;\n@@ -244,7 +251,7 @@ static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, co\n \tdiff_opts.break_opt = opt->break_opt;\n \tdiff_opts.rename_score = opt->rename_score;\n \tdiff_setup_done(&diff_opts);\n-\tdiff_tree(t1, t2, base, &diff_opts);\n+\tll_diff_tree_sha1(old, new, base, &diff_opts);\n \tdiffcore_std(&diff_opts);\n \tfree_pathspec(&diff_opts.pathspec);\n \n@@ -305,23 +312,12 @@ static void try_to_follow_renames(struct tree_desc *t1, struct tree_desc *t2, co\n \n int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base, struct diff_options *opt)\n {\n-\tvoid *tree1, *tree2;\n-\tstruct tree_desc t1, t2;\n-\tunsigned long size1, size2;\n \tint retval;\n \n-\ttree1 = fill_tree_descriptor(&t1, old);\n-\ttree2 = fill_tree_descriptor(&t2, new);\n-\tsize1 = t1.size;\n-\tsize2 = t2.size;\n-\tretval = diff_tree(&t1, &t2, base, opt);\n-\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename()) {\n-\t\tinit_tree_desc(&t1, tree1, size1);\n-\t\tinit_tree_desc(&t2, tree2, size2);\n-\t\ttry_to_follow_renames(&t1, &t2, base, opt);\n-\t}\n-\tfree(tree1);\n-\tfree(tree2);\n+\tretval = ll_diff_tree_sha1(old, new, base, opt);\n+\tif (!*base && DIFF_OPT_TST(opt, FOLLOW_RENAMES) && diff_might_be_rename())\n+\t\ttry_to_follow_renames(old, new, base, opt);\n+\n \treturn retval;\n }\n \n-- \n1.9.rc0.143.g6fd479e\n"},{"id":"237926","messageId":"xmqq1txneavo.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"20140327142438.GE17333@mini.zxlink","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-27T18:48:11Z","receivedAt":"2014-03-27T18:48:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@navytux.spb.ru> writes:\n\n> (please keep author email)\n> ---- 8< ----\n> From: Kirill Smelkov <kirr@mns.spb.ru>\n> Date: Mon, 24 Feb 2014 20:21:46 +0400\n> Subject: [PATCH v3a] tree-diff: rework diff_tree interface to be sha1 based\n\n\"git am -c\" will discard everything above the scissors and then\nstart parsing the in-body headers from there, so the above From:\nwill be used.\n\nBut you have a few entries in .mailmap; do you want to update them\nas well?\n\nBy the way, in general I do not appreciate people lying on the Date:\nwith an in-body header in their patches, either in the original or\nin rerolls.\n\nThanks.\n"},{"id":"237942","messageId":"20140327194300.GA5510@mini.zxlink","threadId":"35947","inReplyTo":"xmqq1txneavo.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-03-27T19:43:00Z","receivedAt":"2014-03-27T19:43:00Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"+stefanbeller\n\nOn Thu, Mar 27, 2014 at 11:48:11AM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n> \n> > (please keep author email)\n> > ---- 8< ----\n> > From: Kirill Smelkov <kirr@mns.spb.ru>\n> > Date: Mon, 24 Feb 2014 20:21:46 +0400\n> > Subject: [PATCH v3a] tree-diff: rework diff_tree interface to be sha1 based\n> \n> \"git am -c\" will discard everything above the scissors and then\n> start parsing the in-body headers from there, so the above From:\n> will be used.\n\nThanks.\n\n> But you have a few entries in .mailmap; do you want to update them\n> as well?\n\nWhen Stefan Beller was contacting me on emails, if I recall correctly, I\ntold him all those kirr@... entries are mine, but the one this patch is\nauthored with indicates that something was done at work, and I'd prefer to\nacknowledge that. So maybe\n\n---- 8< ----\nFrom: Kirill Smelkov <kirr@navytux.spb.ru>\nDate: Thu, 27 Mar 2014 23:32:14 +0400\nSubject: [PATCH] .mailmap: Separate Kirill Smelkov personal and work addresses\n\nThe address kirr@mns.spb.ru indicates that a patch was done at work and\nI'd like to acknowledge that.\n\nThe address kirr@navytux.spb.ru is my personal email and indicates that\na contribution is done completely on my own time and resources.\n\nkirr@landau.phys.spbu.ru is old university account which no longer works\n(sigh, to much spam \"because of me\" on the server) and maps to\nkirr@navytux.spb.ru which should be considered as primary.\n\nSigned-off-by: Kirill Smelkov <kirr@navytux.spb.ru>\n---\n .mailmap | 1 -\n 1 file changed, 1 deletion(-)\n\ndiff --git a/.mailmap b/.mailmap\nindex 11057cb..0be5e02 100644\n--- a/.mailmap\n+++ b/.mailmap\n@@ -117,7 +117,6 @@ Keith Cascio <keith@CS.UCLA.EDU> <keith@cs.ucla.edu>\n Kent Engstrom <kent@lysator.liu.se>\n Kevin Leung <kevinlsk@gmail.com>\n Kirill Smelkov <kirr@navytux.spb.ru> <kirr@landau.phys.spbu.ru>\n-Kirill Smelkov <kirr@navytux.spb.ru> <kirr@mns.spb.ru>\n Knut Franke <Knut.Franke@gmx.de> <k.franke@science-computing.de>\n Lars Doelle <lars.doelle@on-line ! de>\n Lars Doelle <lars.doelle@on-line.de>\n-- \n1.9.rc0.143.g6fd479e\n---- 8< ----\n\nOn the other hand, it is still all me, and the main address (navytux) is\nindicated correctly, so I dunno...\n\n> By the way, in general I do not appreciate people lying on the Date:\n> with an in-body header in their patches, either in the original or\n> in rerolls.\n> \n> Thanks.\n\nI see. Somehow it is pity that the date of original work is lost via\nthis approach, as now we are only changing cosmetics etc, and the bulk\nof the work was done earlier.\n\nAnyway, we can drop the date, but please keep the email, as it is used\nfor the acknowledgment.\n\nThanks,\nKirill\n"},{"id":"237983","messageId":"53351C1B.6040609@viscovery.net","threadId":"35947","inReplyTo":"xmqq1txneavo.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2014-03-28T06:52:11Z","receivedAt":"2014-03-28T06:52:11Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 3/27/2014 19:48, schrieb Junio C Hamano:\n>> From: Kirill Smelkov <kirr@mns.spb.ru>\n>> Date: Mon, 24 Feb 2014 20:21:46 +0400\n>> ...\n> \n> By the way, in general I do not appreciate people lying on the Date:\n> with an in-body header in their patches, either in the original or\n> in rerolls.\n\nformat-patch is not very cooperative in this aspect. When I prepare a\npatch series with format-patch, I find myself editing out the Date: line\nfrom all patches it produces again and again. :-(\n\n-- Hannes\n"},{"id":"238016","messageId":"xmqq4n2ickx4.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"53351C1B.6040609@viscovery.net","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-28T17:06:31Z","receivedAt":"2014-03-28T17:06:31Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Sixt <j.sixt@viscovery.net> writes:\n\n> Am 3/27/2014 19:48, schrieb Junio C Hamano:\n>>> From: Kirill Smelkov <kirr@mns.spb.ru>\n>>> Date: Mon, 24 Feb 2014 20:21:46 +0400\n>>> ...\n>> \n>> By the way, in general I do not appreciate people lying on the Date:\n>> with an in-body header in their patches, either in the original or\n>> in rerolls.\n>\n> format-patch is not very cooperative in this aspect. When I prepare a\n> patch series with format-patch, I find myself editing out the Date: line\n> from all patches it produces again and again. :-(\n\nI am not sure what you mean.  If you are pasting the format-patch\noutput into an editor your MUA is using to receive the body of the\nmessage from you, you would remove all the non-body lines, not just\nDate: but Subject: and From:, no?\n"},{"id":"238022","messageId":"5335B57B.4080606@kdbg.org","threadId":"35947","inReplyTo":"xmqq4n2ickx4.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2014-03-28T17:46:35Z","receivedAt":"2014-03-28T17:46:35Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 28.03.2014 18:06, schrieb Junio C Hamano:\n> Johannes Sixt <j.sixt@viscovery.net> writes:\n> \n>> Am 3/27/2014 19:48, schrieb Junio C Hamano:\n>>>> From: Kirill Smelkov <kirr@mns.spb.ru>\n>>>> Date: Mon, 24 Feb 2014 20:21:46 +0400\n>>>> ...\n>>>\n>>> By the way, in general I do not appreciate people lying on the Date:\n>>> with an in-body header in their patches, either in the original or\n>>> in rerolls.\n>>\n>> format-patch is not very cooperative in this aspect. When I prepare a\n>> patch series with format-patch, I find myself editing out the Date: line\n>> from all patches it produces again and again. :-(\n> \n> I am not sure what you mean.  If you are pasting the format-patch\n> output into an editor your MUA is using to receive the body of the\n> message from you, you would remove all the non-body lines, not just\n> Date: but Subject: and From:, no?\n\nCorrect. So I should add that my gripe is about when I want to send a\npatch series with git-send-email that was prepared with git-format-patch.\n\n-- Hannes\n"},{"id":"238028","messageId":"xmqqtxai9nmq.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"5335B57B.4080606@kdbg.org","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-28T18:36:13Z","receivedAt":"2014-03-28T18:36:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Sixt <j6t@kdbg.org> writes:\n\n> Am 28.03.2014 18:06, schrieb Junio C Hamano:\n>> Johannes Sixt <j.sixt@viscovery.net> writes:\n>> \n>>> Am 3/27/2014 19:48, schrieb Junio C Hamano:\n>>>>> From: Kirill Smelkov <kirr@mns.spb.ru>\n>>>>> Date: Mon, 24 Feb 2014 20:21:46 +0400\n>>>>> ...\n>>>>\n>>>> By the way, in general I do not appreciate people lying on the Date:\n>>>> with an in-body header in their patches, either in the original or\n>>>> in rerolls.\n>>>\n>>> format-patch is not very cooperative in this aspect. When I prepare a\n>>> patch series with format-patch, I find myself editing out the Date: line\n>>> from all patches it produces again and again. :-(\n>> \n>> I am not sure what you mean.  If you are pasting the format-patch\n>> output into an editor your MUA is using to receive the body of the\n>> message from you, you would remove all the non-body lines, not just\n>> Date: but Subject: and From:, no?\n>\n> Correct. So I should add that my gripe is about when I want to send a\n> patch series with git-send-email that was prepared with git-format-patch.\n\nHmph.  Don't you get fresh timestamps for your messages in such a\ncase, ignoring whatever is at the beginning of the input files?\n\nMy reading of git-send-email is:\n\n * \"$time = time - scalar $#files\" prepares the initial \"timestamp\",\n   so that running two \"git send-email\" back to back will give\n   timestamps to the series sent out by the first invocation that\n   are older than the ones the second series will get;\n\n * \"sub send_message\" calls \"format_2822_time($time++)\" to send the\n   first message with that initial \"timestamp\", incrementing the\n   timestamps by 1 second intervals (without having to actually wait\n   1 second in between messages) for each patch.\n"},{"id":"238036","messageId":"5335C89A.7040802@kdbg.org","threadId":"35947","inReplyTo":"xmqqtxai9nmq.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2014-03-28T19:08:10Z","receivedAt":"2014-03-28T19:08:10Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 28.03.2014 19:36, schrieb Junio C Hamano:\n> Johannes Sixt <j6t@kdbg.org> writes:\n> \n>> Am 28.03.2014 18:06, schrieb Junio C Hamano:\n>>> Johannes Sixt <j.sixt@viscovery.net> writes:\n>>>\n>>>> Am 3/27/2014 19:48, schrieb Junio C Hamano:\n>>>>>> From: Kirill Smelkov <kirr@mns.spb.ru>\n>>>>>> Date: Mon, 24 Feb 2014 20:21:46 +0400\n>>>>>> ...\n>>>>>\n>>>>> By the way, in general I do not appreciate people lying on the Date:\n>>>>> with an in-body header in their patches, either in the original or\n>>>>> in rerolls.\n>>>>\n>>>> format-patch is not very cooperative in this aspect. When I prepare a\n>>>> patch series with format-patch, I find myself editing out the Date: line\n>>>> from all patches it produces again and again. :-(\n>>>\n>>> I am not sure what you mean.  If you are pasting the format-patch\n>>> output into an editor your MUA is using to receive the body of the\n>>> message from you, you would remove all the non-body lines, not just\n>>> Date: but Subject: and From:, no?\n>>\n>> Correct. So I should add that my gripe is about when I want to send a\n>> patch series with git-send-email that was prepared with git-format-patch.\n> \n> Hmph.  Don't you get fresh timestamps for your messages in such a\n> case, ignoring whatever is at the beginning of the input files?\n> \n> My reading of git-send-email is:\n> \n>  * \"$time = time - scalar $#files\" prepares the initial \"timestamp\",\n>    so that running two \"git send-email\" back to back will give\n>    timestamps to the series sent out by the first invocation that\n>    are older than the ones the second series will get;\n> \n>  * \"sub send_message\" calls \"format_2822_time($time++)\" to send the\n>    first message with that initial \"timestamp\", incrementing the\n>    timestamps by 1 second intervals (without having to actually wait\n>    1 second in between messages) for each patch.\n\nAh, nice! I didn't know that. I never dared to leave an old author date\n(or any date) in the patches, and assumed that it would be kept and\ndisrupt the email time line.\n\nThanks for the hint, and sorry for the noise.\n\n-- Hannes\n"},{"id":"238039","messageId":"xmqqd2h69l9h.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"5335C89A.7040802@kdbg.org","subject":"Re: [PATCH v2 14/19] tree-diff: rework diff_tree interface to be sha1 based","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-03-28T19:27:22Z","receivedAt":"2014-03-28T19:27:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Sixt <j6t@kdbg.org> writes:\n\n>> My reading of git-send-email is:\n>> \n>>  * \"$time = time - scalar $#files\" prepares the initial \"timestamp\",\n>>    so that running two \"git send-email\" back to back will give\n>>    timestamps to the series sent out by the first invocation that\n>>    are older than the ones the second series will get;\n\nA completely irrelevant tangent, but I was being an idiot here.  The\n\"-scaler #$files\" is not about two send-email running back to back.\nA second invocation that sends out a long series will start its\ntimestamp #$files in the past, that will overlap with the timestamp\nof the last one in the first invocation.  And that is not what the\ncode attempts to address.  It wants to merely avoid timestamps from\nthe future.\n"},{"id":"238377","messageId":"xmqqppkxos0w.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"20140327142354.GD17333@mini.zxlink","subject":"Re: [PATCH v2 18/19] tree-diff: rework diff_tree() to generate diffs for multiparent cases as well","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-04-04T18:42:39Z","receivedAt":"2014-04-04T18:42:39Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@navytux.spb.ru> writes:\n\n> +extern\n> +struct combine_diff_path *diff_tree_paths(\n\nThese two on the same line, please.\n\n> +\tstruct combine_diff_path *p, const unsigned char *sha1,\n> +\tconst unsigned char **parent_sha1, int nparent,\n> +\tstruct strbuf *base, struct diff_options *opt);\n>  extern int diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n>  \t\t\t  const char *base, struct diff_options *opt);\n> ...\n> +/*\n> + * convert path -> opt->diff_*() callbacks\n> + *\n> + * emits diff to first parent only, and tells diff tree-walker that we are done\n> + * with p and it can be freed.\n> + */\n> +static int emit_diff_first_parent_only(struct diff_options *opt, struct combine_diff_path *p)\n>  {\n\nVery straight-forward; good.\n\n> +static struct combine_diff_path *path_appendnew(struct combine_diff_path *last,\n> +\tint nparent, const struct strbuf *base, const char *path, int pathlen,\n> +\tunsigned mode, const unsigned char *sha1)\n> +{\n> +\tstruct combine_diff_path *p;\n> +\tint len = base->len + pathlen;\n> +\tint alloclen = combine_diff_path_size(nparent, len);\n> +\n> +\t/* if last->next is !NULL - it is a pre-allocated memory, we can reuse */\n> +\tp = last->next;\n> +\tif (p && (alloclen > (intptr_t)p->next)) {\n> +\t\tfree(p);\n> +\t\tp = NULL;\n> +\t}\n> +\n> +\tif (!p) {\n> +\t\tp = xmalloc(alloclen);\n> +\n> +\t\t/*\n> +\t\t * until we go to it next round, .next holds how many bytes we\n> +\t\t * allocated (for faster realloc - we don't need copying old data).\n> +\t\t */\n> +\t\tp->next = (struct combine_diff_path *)(intptr_t)alloclen;\n\nThis reuse of the .next field is somewhat yucky, but it is very\nlocalized inside a function that has a single callsite to this\nfunction, so let's let it pass.\n\n> +static struct combine_diff_path *emit_path(struct combine_diff_path *p,\n> +\tstruct strbuf *base, struct diff_options *opt, int nparent,\n> +\tstruct tree_desc *t, struct tree_desc *tp,\n> +\tint imin)\n>  {\n\nAgain, fairly straight-forward and good.\n\n> +/*\n> + * generate paths for combined diff D(sha1,parents_sha1[])\n> + ...\n> +static struct combine_diff_path *ll_diff_tree_paths(\n> +\tstruct combine_diff_path *p, const unsigned char *sha1,\n> +\tconst unsigned char **parents_sha1, int nparent,\n> +\tstruct strbuf *base, struct diff_options *opt)\n> +{\n> +\tstruct tree_desc t, *tp;\n> +\tvoid *ttree, **tptree;\n> +\tint i;\n> +\n> +\ttp     = xalloca(nparent * sizeof(tp[0]));\n> +\ttptree = xalloca(nparent * sizeof(tptree[0]));\n> +\n> +\t/*\n> +\t * load parents first, as they are probably already cached.\n> +\t *\n> +\t * ( log_tree_diff() parses commit->parent before calling here via\n> +\t *   diff_tree_sha1(parent, commit) )\n> +\t */\n> +\tfor (i = 0; i < nparent; ++i)\n> +\t\ttptree[i] = fill_tree_descriptor(&tp[i], parents_sha1[i]);\n> +\tttree = fill_tree_descriptor(&t, sha1);\n>  \n>  \t/* Enable recursion indefinitely */\n>  \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n>  \n>  \tfor (;;) {\n> -\t\tint cmp;\n> +\t\tint imin, cmp;\n>  \n>  \t\tif (diff_can_quit_early(opt))\n>  \t\t\tbreak;\n> +\n>  \t\tif (opt->pathspec.nr) {\n> -\t\t\tskip_uninteresting(&t1, base, opt);\n> -\t\t\tskip_uninteresting(&t2, base, opt);\n> +\t\t\tskip_uninteresting(&t, base, opt);\n> +\t\t\tfor (i = 0; i < nparent; i++)\n> +\t\t\t\tskip_uninteresting(&tp[i], base, opt);\n>  \t\t}\n> -\t\tif (!t1.size && !t2.size)\n> -\t\t\tbreak;\n>  \n> -\t\tcmp = tree_entry_pathcmp(&t1, &t2);\n> +\t\t/* comparing is finished when all trees are done */\n> +\t\tif (!t.size) {\n> +\t\t\tint done = 1;\n> +\t\t\tfor (i = 0; i < nparent; ++i)\n> +\t\t\t\tif (tp[i].size) {\n> +\t\t\t\t\tdone = 0;\n> +\t\t\t\t\tbreak;\n> +\t\t\t\t}\n> +\t\t\tif (done)\n> +\t\t\t\tbreak;\n> +\t\t}\n> +\n> +\t\t/*\n> +\t\t * lookup imin = argmin(x1...xn),\n> +\t\t * mark entries whether they =tp[imin] along the way\n> +\t\t */\n> +\t\timin = 0;\n> +\t\ttp[0].entry.mode &= ~S_IFXMIN_NEQ;\n> +\n> +\t\tfor (i = 1; i < nparent; ++i) {\n> +\t\t\tcmp = tree_entry_pathcmp(&tp[i], &tp[imin]);\n> +\t\t\tif (cmp < 0) {\n> +\t\t\t\timin = i;\n> +\t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n> +\t\t\t}\n> +\t\t\telse if (cmp == 0) {\n> +\t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n> +\t\t\t}\n> +\t\t\telse {\n> +\t\t\t\ttp[i].entry.mode |= S_IFXMIN_NEQ;\n> +\t\t\t}\n> +\t\t}\n> +\n> +\t\t/* fixup markings for entries before imin */\n> +\t\tfor (i = 0; i < imin; ++i)\n> +\t\t\ttp[i].entry.mode |= S_IFXMIN_NEQ;\t/* x[i] > x[imin] */\n> +\n\nThese two loop made my reading hiccup for a while.  With these you\nare scanning the tp[] array 1.5 times (and doing the bitwise\nassignment to entry.mode 1.5 * nparent times), but I suspect it may\nhave been a lot easier to read if the first loop only identified the\nimin, and the second loop only did the entry.mode for _all_ nparents.\n\n> +\t\t/* compare a vs x[imin] */\n> +\t\tcmp = tree_entry_pathcmp(&t, &tp[imin]);\n> +\n> +\t\t/* a = xi */\n> +\t\tif (cmp == 0) {\n> +\t\t\t/* are either xk > xi or diff(a,xk) != ø ? */\n> +\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n> +\t\t\t\tfor (i = 0; i < nparent; ++i) {\n> +\t\t\t\t\t/* x[i] > x[imin] */\n> +\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n> +\t\t\t\t\t\tcontinue;\n> +\n> +\t\t\t\t\t/* diff(a,xk) != ø */\n> +\t\t\t\t\tif (hashcmp(t.entry.sha1, tp[i].entry.sha1) ||\n> +\t\t\t\t\t    (t.entry.mode != tp[i].entry.mode))\n> +\t\t\t\t\t\tcontinue;\n> +\n> +\t\t\t\t\tgoto skip_emit_t_tp;\n> +\t\t\t\t}\n> +\t\t\t}\n\nPlease bear with me.  The notation scares me as I am not good at math.\n\nIn short, the above loop is about:\n\n    We are looking at path in 't' and some parents have the same\n    path.  If any of these parents have that path with the contents\n    identical to 't', then do not emit this path.\n\nwhich makes sense to me, but these notation also made my reading\nhiccup, especially because it is hard to guess what \"xk\" refers to\n(e.g. \"any k where 0 <= k < nparent && i != k\"? \"all such k\"?).  I\nstill haven't figured out what you meant to say with \"xk\", but I\nthink I got what the code wants to do.\n\nHow does the \"the (virtual) path from a tree that has ran out of\nentries sorts later than anything else\" comparison rule influence\nthe picture?  A parent that has ran out would have _NEQ bit set and\nwould not count as having the same contents as the path from 't'.\nIf 't' has ran out, the only way t and tp[imin] could compare equal\nis when tp[imin] has also ran out, but that can happen only when all\nthe parents are done with, so we would have broken out of the loop\neven before we try to figure out imin.  So there is no funnies\nthere, which is good.\n\n> +\t\t\t/* D += {δ(a,xk) if xk=xi;  \"+a\" if xk > xi} */\n> +\t\t\tp = emit_path(p, base, opt, nparent,\n> +\t\t\t\t\t&t, tp, imin);\n> +\n> +\t\tskip_emit_t_tp:\n> +\t\t\t/* a↓,  ∀ xk=ximin  xk↓ */\n> +\t\t\tupdate_tree_entry(&t);\n> +\t\t\tupdate_tp_entries(tp, nparent);\n>  \t\t}\n>  \n> -\t\t/* t1 < t2 */\n> +\t\t/* a < xi */\n>  \t\telse if (cmp < 0) {\n> -\t\t\tshow_path(base, opt, &t1, /*t2=*/NULL);\n> -\t\t\tupdate_tree_entry(&t1);\n> +\t\t\t/* D += \"+a\" */\n> +\t\t\tp = emit_path(p, base, opt, nparent,\n> +\t\t\t\t\t&t, /*tp=*/NULL, -1);\n> +\n> +\t\t\t/* a↓ */\n> +\t\t\tupdate_tree_entry(&t);\n\nThis is straight-forward.  No parent has path 't' has, so only the\nentry from 't' is given, and we deal with the next entry in 't'\nwithout touching any of the parents in the next iteration.  Good.\n\n>  \t\t}\n>  \n> -\t\t/* t1 > t2 */\n> +\t\t/* a > xi */\n>  \t\telse {\n> -\t\t\tshow_path(base, opt, /*t1=*/NULL, &t2);\n> -\t\t\tupdate_tree_entry(&t2);\n> +\t\t\t/* ∀j xj=ximin -> D += \"-xi\" */\n\nDid you mean \"-xj\"?\n\n> +\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n> +\t\t\t\tfor (i = 0; i < nparent; ++i)\n> +\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n> +\t\t\t\t\t\tgoto skip_emit_tp;\n> +\t\t\t}\n> +\n> +\t\t\tp = emit_path(p, base, opt, nparent,\n> +\t\t\t\t\t/*t=*/NULL, tp, imin);\n> +\n> +\t\tskip_emit_tp:\n> +\t\t\t/* ∀ xk=ximin  xk↓ */\n> +\t\t\tupdate_tp_entries(tp, nparent);\n\nThere are parents whose path sort earlier than what is in 't'\n(i.e. they were lost in the result---we would want to show\nremoval).  What makes us jump to the skip label?\n\n    We are looking at path in 't', and some parents have paths that\n    sort earlier than that path.  We will not go to skip label if\n    any one of the parent's entry sorts after some other parent (or\n    the parent in question has ran out its entries), which means we\n    show the entry from the parents only when all the parents have\n    that same path, which is missing from 't'.\n\nI am not sure if I am reading this correctly, though.\n\nFor the two-way diff, the above degenerates to \"show all parent\nentries that come before the first entry in 't'\", which is correct.\nFor the combined diff, the current intersect_paths() makes sure that\neach path appears in all the pair-wise diff between t and tp[],\nwhich again means that the above logic match the current behaviour.\n\n\n> +struct combine_diff_path *diff_tree_paths(\n> +\tstruct combine_diff_path *p, const unsigned char *sha1,\n> +\tconst unsigned char **parents_sha1, int nparent,\n> +\tstruct strbuf *base, struct diff_options *opt)\n> +{\n> +\tp = ll_diff_tree_paths(p, sha1, parents_sha1, nparent, base, opt);\n> +\n> +\t/*\n> +\t * free pre-allocated last element, if any\n> +\t * (see path_appendnew() for details about why)\n> +\t */\n> +\tif (p->next) {\n> +\t\tfree(p->next);\n> +\t\tp->next = NULL;\n> +\t}\n> +\n> +\treturn p;\n>  }\n>  \n>  /*\n> @@ -308,6 +664,27 @@ static void try_to_follow_renames(const unsigned char *old, const unsigned char\n>  \tq->nr = 1;\n>  }\n>  \n> +static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n> +\t\t\t     struct strbuf *base, struct diff_options *opt)\n> +{\n> +\tstruct combine_diff_path phead, *p;\n> +\tconst unsigned char *parents_sha1[1] = {old};\n> +\tpathchange_fn_t pathchange_old = opt->pathchange;\n> +\n> +\tphead.next = NULL;\n> +\topt->pathchange = emit_diff_first_parent_only;\n> +\tdiff_tree_paths(&phead, new, parents_sha1, 1, base, opt);\n\nHmph.  I would have expected\n\n\tconst unsigned char **parents_sha1 = &old;\n\nor even\n\n\tdiff_tree_paths(&phead, new, &old, 1, base, opt);\n\nhere.\n\n\nThanks.\n"},{"id":"238400","messageId":"20140406214626.GA3843@mini.zxlink","threadId":"35947","inReplyTo":"xmqqppkxos0w.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 18/19] tree-diff: rework diff_tree() to generate diffs for multiparent cases as well","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-04-06T21:46:26Z","receivedAt":"2014-04-06T21:46:26Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"Junio,\n\nFirst of all thanks a lot for reviewing this patch. I'll reply inline\nwith corrected version attached in the end.\n\nOn Fri, Apr 04, 2014 at 11:42:39AM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n> \n> > +extern\n> > +struct combine_diff_path *diff_tree_paths(\n> \n> These two on the same line, please.\n\nOk\n\n> > +\tstruct combine_diff_path *p, const unsigned char *sha1,\n> > +\tconst unsigned char **parent_sha1, int nparent,\n> > +\tstruct strbuf *base, struct diff_options *opt);\n> >  extern int diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n> >  \t\t\t  const char *base, struct diff_options *opt);\n> > ...\n> > +/*\n> > + * convert path -> opt->diff_*() callbacks\n> > + *\n> > + * emits diff to first parent only, and tells diff tree-walker that we are done\n> > + * with p and it can be freed.\n> > + */\n> > +static int emit_diff_first_parent_only(struct diff_options *opt, struct combine_diff_path *p)\n> >  {\n> \n> Very straight-forward; good.\n\nThanks\n\n> > +static struct combine_diff_path *path_appendnew(struct combine_diff_path *last,\n> > +\tint nparent, const struct strbuf *base, const char *path, int pathlen,\n> > +\tunsigned mode, const unsigned char *sha1)\n> > +{\n> > +\tstruct combine_diff_path *p;\n> > +\tint len = base->len + pathlen;\n> > +\tint alloclen = combine_diff_path_size(nparent, len);\n> > +\n> > +\t/* if last->next is !NULL - it is a pre-allocated memory, we can reuse */\n> > +\tp = last->next;\n> > +\tif (p && (alloclen > (intptr_t)p->next)) {\n> > +\t\tfree(p);\n> > +\t\tp = NULL;\n> > +\t}\n> > +\n> > +\tif (!p) {\n> > +\t\tp = xmalloc(alloclen);\n> > +\n> > +\t\t/*\n> > +\t\t * until we go to it next round, .next holds how many bytes we\n> > +\t\t * allocated (for faster realloc - we don't need copying old data).\n> > +\t\t */\n> > +\t\tp->next = (struct combine_diff_path *)(intptr_t)alloclen;\n> \n> This reuse of the .next field is somewhat yucky, but it is very\n> localized inside a function that has a single callsite to this\n> function, so let's let it pass.\n\nI agree it is not pretty, but it was the best approach I could find\nfor avoiding memory re-allocation without introducing new fields into\n`struct combine_diff_path`. And yes, the trick is localized, so let's\nlet it live.\n\n\n> > +static struct combine_diff_path *emit_path(struct combine_diff_path *p,\n> > +\tstruct strbuf *base, struct diff_options *opt, int nparent,\n> > +\tstruct tree_desc *t, struct tree_desc *tp,\n> > +\tint imin)\n> >  {\n> \n> Again, fairly straight-forward and good.\n\nThanks again.\n\n\n> > +/*\n> > + * generate paths for combined diff D(sha1,parents_sha1[])\n> > + ...\n> > +static struct combine_diff_path *ll_diff_tree_paths(\n> > +\tstruct combine_diff_path *p, const unsigned char *sha1,\n> > +\tconst unsigned char **parents_sha1, int nparent,\n> > +\tstruct strbuf *base, struct diff_options *opt)\n> > +{\n> > +\tstruct tree_desc t, *tp;\n> > +\tvoid *ttree, **tptree;\n> > +\tint i;\n> > +\n> > +\ttp     = xalloca(nparent * sizeof(tp[0]));\n> > +\ttptree = xalloca(nparent * sizeof(tptree[0]));\n> > +\n> > +\t/*\n> > +\t * load parents first, as they are probably already cached.\n> > +\t *\n> > +\t * ( log_tree_diff() parses commit->parent before calling here via\n> > +\t *   diff_tree_sha1(parent, commit) )\n> > +\t */\n> > +\tfor (i = 0; i < nparent; ++i)\n> > +\t\ttptree[i] = fill_tree_descriptor(&tp[i], parents_sha1[i]);\n> > +\tttree = fill_tree_descriptor(&t, sha1);\n> >  \n> >  \t/* Enable recursion indefinitely */\n> >  \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n> >  \n> >  \tfor (;;) {\n> > -\t\tint cmp;\n> > +\t\tint imin, cmp;\n> >  \n> >  \t\tif (diff_can_quit_early(opt))\n> >  \t\t\tbreak;\n> > +\n> >  \t\tif (opt->pathspec.nr) {\n> > -\t\t\tskip_uninteresting(&t1, base, opt);\n> > -\t\t\tskip_uninteresting(&t2, base, opt);\n> > +\t\t\tskip_uninteresting(&t, base, opt);\n> > +\t\t\tfor (i = 0; i < nparent; i++)\n> > +\t\t\t\tskip_uninteresting(&tp[i], base, opt);\n> >  \t\t}\n> > -\t\tif (!t1.size && !t2.size)\n> > -\t\t\tbreak;\n> >  \n> > -\t\tcmp = tree_entry_pathcmp(&t1, &t2);\n> > +\t\t/* comparing is finished when all trees are done */\n> > +\t\tif (!t.size) {\n> > +\t\t\tint done = 1;\n> > +\t\t\tfor (i = 0; i < nparent; ++i)\n> > +\t\t\t\tif (tp[i].size) {\n> > +\t\t\t\t\tdone = 0;\n> > +\t\t\t\t\tbreak;\n> > +\t\t\t\t}\n> > +\t\t\tif (done)\n> > +\t\t\t\tbreak;\n> > +\t\t}\n> > +\n> > +\t\t/*\n> > +\t\t * lookup imin = argmin(x1...xn),\n> > +\t\t * mark entries whether they =tp[imin] along the way\n> > +\t\t */\n> > +\t\timin = 0;\n> > +\t\ttp[0].entry.mode &= ~S_IFXMIN_NEQ;\n> > +\n> > +\t\tfor (i = 1; i < nparent; ++i) {\n> > +\t\t\tcmp = tree_entry_pathcmp(&tp[i], &tp[imin]);\n> > +\t\t\tif (cmp < 0) {\n> > +\t\t\t\timin = i;\n> > +\t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n> > +\t\t\t}\n> > +\t\t\telse if (cmp == 0) {\n> > +\t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n> > +\t\t\t}\n> > +\t\t\telse {\n> > +\t\t\t\ttp[i].entry.mode |= S_IFXMIN_NEQ;\n> > +\t\t\t}\n> > +\t\t}\n> > +\n> > +\t\t/* fixup markings for entries before imin */\n> > +\t\tfor (i = 0; i < imin; ++i)\n> > +\t\t\ttp[i].entry.mode |= S_IFXMIN_NEQ;\t/* x[i] > x[imin] */\n> > +\n> \n> These two loop made my reading hiccup for a while.  With these you\n> are scanning the tp[] array 1.5 times (and doing the bitwise\n> assignment to entry.mode 1.5 * nparent times), but I suspect it may\n> have been a lot easier to read if the first loop only identified the\n> imin, and the second loop only did the entry.mode for _all_ nparents.\n\nHmm, if in the first loop, we identify imin only, then in the second\nloop we would have to call tree_entry_pathcmp(tp[i], tp[imin]) again,\nwhich, in my view, would be not better and even worse - in the original\ncase we are scanning the parents nparent times comparing paths, and only\nthen do simple fixup up-to imin entry.\n\nThe following\n\n---- 8< ---\n                /* lookup imin = argmin(p1...pn) */\n                imin = 0;\n                for (i = 1; i < nparent; ++i) {\n                        cmp = tree_entry_pathcmp(&tp[i], &tp[imin]);\n                        if (cmp < 0) \n                                imin = i;\n                }\n\n                /* mark entries whether they =p[imin] */\n                for (i = 0; i < nparent; ++i) {\n                        cmp = tree_entry_pathcmp(&tp[i], &tp[imin]);\n                        if (cmp)\n                                tp[i].entry.mode |= S_IFXMIN_NEQ;       \n                        else\n                                tp[i].entry.mode &= S_IFXMIN_NEQ;\n                }\n---- 8< ----\n\nmaybe looks a bit simpler, but calls tree_entry_pathcmp twice more times.\n\nBesides for important nparent=1 case we were not calling\ntree_entry_pathcmp at all and here we'll call it once, which would slow\nexecution down a bit, as base_name_compare shows measurable enough in profile.\nTo avoid that we'll need to add 'if (i==imin) continue' and this won't\nbe so simple then. And for general nparent case, as I've said, we'll be\ncalling tree_entry_pathcmp twice more times...\n\nBecause of all that I'd suggest to go with my original version.\n\n\n> > +\t\t/* compare a vs x[imin] */\n> > +\t\tcmp = tree_entry_pathcmp(&t, &tp[imin]);\n> > +\n> > +\t\t/* a = xi */\n> > +\t\tif (cmp == 0) {\n> > +\t\t\t/* are either xk > xi or diff(a,xk) != ø ? */\n> > +\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n> > +\t\t\t\tfor (i = 0; i < nparent; ++i) {\n> > +\t\t\t\t\t/* x[i] > x[imin] */\n> > +\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n> > +\t\t\t\t\t\tcontinue;\n> > +\n> > +\t\t\t\t\t/* diff(a,xk) != ø */\n> > +\t\t\t\t\tif (hashcmp(t.entry.sha1, tp[i].entry.sha1) ||\n> > +\t\t\t\t\t    (t.entry.mode != tp[i].entry.mode))\n> > +\t\t\t\t\t\tcontinue;\n> > +\n> > +\t\t\t\t\tgoto skip_emit_t_tp;\n> > +\t\t\t\t}\n> > +\t\t\t}\n> \n> Please bear with me.  The notation scares me as I am not good at math.\n> \n> In short, the above loop is about:\n> \n>     We are looking at path in 't' and some parents have the same\n>     path.  If any of these parents have that path with the contents\n>     identical to 't', then do not emit this path.\n\nYes, correct.\n\n> which makes sense to me, but these notation also made my reading\n> hiccup, especially because it is hard to guess what \"xk\" refers to\n> (e.g. \"any k where 0 <= k < nparent && i != k\"? \"all such k\"?).  I\n> still haven't figured out what you meant to say with \"xk\", but I\n> think I got what the code wants to do.\n\nSorry about scaring you and about hiccup. After some break on the topic,\nwith a fresh eye I see a lot of confusion goes from the notation I've\nchosen initially (because of how I was reasoning about it on paper, when\nit was in flux) - i.e. xi for x[imin] and also using i as looping\nvariable. And also because xi was already used for x[imin] I've used\nanother letter 'k' denoting all other x'es, which leads to confusion...\n\n\nI propose we do the following renaming to clarify things:\n\n    A/a     ->      T/t     (to match resulting tree t name in the code)\n    X/x     ->      P/p     (to match parents trees tp in the code)\n    i       ->      imin    (so that i would be free for other tasks)\n\nthen the above (with a prologue) would look like\n\n---- 8< ----\n *       T     P1       Pn\n *       -     -        -\n *      |t|   |p1|     |pn|\n *      |-|   |--| ... |--|      imin = argmin(p1...pn)\n *      | |   |  |     |  |\n *      |-|   |--|     |--|\n *      |.|   |. |     |. |\n *       .     .        .\n *       .     .        .\n *\n * at any time there could be 3 cases:\n *\n *      1)  t < p[imin];\n *      2)  t > p[imin];\n *      3)  t = p[imin].\n *\n * Schematic deduction of what every case means, and what to do, follows:\n *\n * 1)  t < p[imin]  ->  ∀j t ∉ Pj  ->  \"+t\" ∈ D(T,Pj)  ->  D += \"+t\";  t↓\n *\n * 2)  t > p[imin]\n *\n *     2.1) ∃j: pj > p[imin]  ->  \"-p[imin]\" ∉ D(T,Pj)  ->  D += ø;  ∀ pi=p[imin]  pi↓\n *     2.2) ∀i  pi = p[imin]  ->  pi ∉ T  ->  \"-pi\" ∈ D(T,Pi)  ->  D += \"-p[imin]\";  ∀i pi↓\n *\n * 3)  t = p[imin]\n *\n *     3.1) ∃j: pj > p[imin]  ->  \"+t\" ∈ D(T,Pj)  ->  only pi=p[imin] remains to investigate\n *     3.2) pi = p[imin]  ->  investigate δ(t,pi)\n *      |\n *      |\n *      v\n *\n *     3.1+3.2) looking at δ(t,pi) ∀i: pi=p[imin] - if all != ø  ->\n *\n *                       ⎧δ(t,pi)  - if pi=p[imin]\n *              ->  D += ⎨\n *                       ⎩\"+t\"     - if pi>p[imin]\n *\n *\n *     in any case t↓  ∀ pi=p[imin]  pi↓\n\n ...\n\n                /* compare t vs p[imin] */\n                cmp = tree_entry_pathcmp(&t, &tp[imin]);\n\n                /* t = p[imin] */\n                if (cmp == 0) {\n                        /* are either pi > p[imin] or diff(t,pi) != ø ? */\n                        if (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n                                for (i = 0; i < nparent; ++i) {\n                                        /* p[i] > p[imin] */\n                                        if (tp[i].entry.mode & S_IFXMIN_NEQ)\n                                                continue;\n\n                                        /* diff(t,pi) != ø */\n                                        if (hashcmp(t.entry.sha1, tp[i].entry.sha1) ||\n                                            (t.entry.mode != tp[i].entry.mode))\n                                                continue;\n\n                                        goto skip_emit_t_tp;\n                                }\n                        }\n\n                        /* D += {δ(t,pi) if pi=p[imin];  \"+a\" if pi > p[imin]} */\n                        p = emit_path(p, base, opt, nparent,\n                                        &t, tp, imin);\n\n                skip_emit_t_tp:\n                        /* t↓,  ∀ pi=p[imin]  pi↓ */\n                        update_tree_entry(&t);\n                        update_tp_entries(tp, nparent);\n                }\n---- 8< ----\n\nnow xk is gone and i matches p[i] (= pi) etc so variable names correlate\nto algorithm description better.\n\nDoes that maybe clarify things?\n\n\n> How does the \"the (virtual) path from a tree that has ran out of\n> entries sorts later than anything else\" comparison rule influence\n> the picture?  A parent that has ran out would have _NEQ bit set and\n> would not count as having the same contents as the path from 't'.\n> If 't' has ran out, the only way t and tp[imin] could compare equal\n> is when tp[imin] has also ran out, but that can happen only when all\n> the parents are done with, so we would have broken out of the loop\n> even before we try to figure out imin.  So there is no funnies\n> there, which is good.\n\nYes, exactly.\n\n\n> > +\t\t\t/* D += {δ(a,xk) if xk=xi;  \"+a\" if xk > xi} */\n> > +\t\t\tp = emit_path(p, base, opt, nparent,\n> > +\t\t\t\t\t&t, tp, imin);\n> > +\n> > +\t\tskip_emit_t_tp:\n> > +\t\t\t/* a↓,  ∀ xk=ximin  xk↓ */\n> > +\t\t\tupdate_tree_entry(&t);\n> > +\t\t\tupdate_tp_entries(tp, nparent);\n> >  \t\t}\n> >  \n> > -\t\t/* t1 < t2 */\n> > +\t\t/* a < xi */\n> >  \t\telse if (cmp < 0) {\n> > -\t\t\tshow_path(base, opt, &t1, /*t2=*/NULL);\n> > -\t\t\tupdate_tree_entry(&t1);\n> > +\t\t\t/* D += \"+a\" */\n> > +\t\t\tp = emit_path(p, base, opt, nparent,\n> > +\t\t\t\t\t&t, /*tp=*/NULL, -1);\n> > +\n> > +\t\t\t/* a↓ */\n> > +\t\t\tupdate_tree_entry(&t);\n> \n> This is straight-forward.  No parent has path 't' has, so only the\n> entry from 't' is given, and we deal with the next entry in 't'\n> without touching any of the parents in the next iteration.  Good.\n\nYes. I hope with the renaming it looks a bit more cleaner:\n\n                /* t < p[imin] */\n                else if (cmp < 0) {\n                        /* D += \"+t\" */\n                        p = emit_path(p, base, opt, nparent,\n                                        &t, /*tp=*/NULL, -1);\n\n                        /* t↓ */\n                        update_tree_entry(&t);\n                }\n\n> \n> >  \t\t}\n> >  \n> > -\t\t/* t1 > t2 */\n> > +\t\t/* a > xi */\n> >  \t\telse {\n> > -\t\t\tshow_path(base, opt, /*t1=*/NULL, &t2);\n> > -\t\t\tupdate_tree_entry(&t2);\n> > +\t\t\t/* ∀j xj=ximin -> D += \"-xi\" */\n> \n> Did you mean \"-xj\"?\n\nNo, ximin, which was denoted in the earlier iterations of patch as xi -\nif all parents current paths are equal and t does not have this path -\nremove the path present in parents. It does not strictly differs between\nximin and xj here, but xj is looping so ximin is better to have as some\ndefined path.\n\nWith the renaming the code looks like this\n\n---- 8< ----\n * 2)  t > p[imin]\n *\n *     2.1) ∃j: pj > p[imin]  ->  \"-p[imin]\" ∉ D(T,Pj)  ->  D += ø;  ∀ pi=p[imin]  pi↓\n *     2.2) ∀i  pi = p[imin]  ->  pi ∉ T  ->  \"-pi\" ∈ D(T,Pi)  ->  D += \"-p[imin]\";  ∀i pi↓\n ...\n                /* t > p[imin] */\n                else {\n                        /* ∀i pi=p[imin] -> D += \"-p[imin]\" */\n                        if (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n                                for (i = 0; i < nparent; ++i)\n                                        if (tp[i].entry.mode & S_IFXMIN_NEQ)\n                                                goto skip_emit_tp;\n                        }\n\n                        p = emit_path(p, base, opt, nparent,\n                                        /*t=*/NULL, tp, imin);\n\n                skip_emit_tp:\n                        /* ∀ pi=p[imin]  pi↓ */\n                        update_tp_entries(tp, nparent);\n                }\n---- 8< ----\n\nThanks for spotting it.\n\n\n> > +\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n> > +\t\t\t\tfor (i = 0; i < nparent; ++i)\n> > +\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n> > +\t\t\t\t\t\tgoto skip_emit_tp;\n> > +\t\t\t}\n> > +\n> > +\t\t\tp = emit_path(p, base, opt, nparent,\n> > +\t\t\t\t\t/*t=*/NULL, tp, imin);\n> > +\n> > +\t\tskip_emit_tp:\n> > +\t\t\t/* ∀ xk=ximin  xk↓ */\n> > +\t\t\tupdate_tp_entries(tp, nparent);\n> \n> There are parents whose path sort earlier than what is in 't'\n> (i.e. they were lost in the result---we would want to show\n> removal).  What makes us jump to the skip label?\n> \n>     We are looking at path in 't', and some parents have paths that\n>     sort earlier than that path.  We will not go to skip label if\n>     any one of the parent's entry sorts after some other parent (or\n>     the parent in question has ran out its entries), which means we\n>     show the entry from the parents only when all the parents have\n>     that same path, which is missing from 't'.\n> \n> I am not sure if I am reading this correctly, though.\n> \n> For the two-way diff, the above degenerates to \"show all parent\n> entries that come before the first entry in 't'\", which is correct.\n> For the combined diff, the current intersect_paths() makes sure that\n> each path appears in all the pair-wise diff between t and tp[],\n> which again means that the above logic match the current behaviour.\n\nYes, correct (modulo we *will* go to skip label if any one of the\nparent's entry sorts after some other parent). By definition of combined\ndiff we show a path only if it shows in every diff D(T,Pi), and if \n\n    2.1) ∃j: pj > p[imin]  ->  \"-p[imin]\" ∉ D(T,Pj)  ->  D += ø;  ∀ pi=p[imin]  pi↓\n\nsome pj sorts after p[imin] that would mean that Pj does not have\np[imin] and since t > p[imin] (which means T does not have p[imin]\neither) diff D(T,Pj) does not have p[imin]. And because of that we know\nthe whole combined-diff will not have p[imin] as, by definition,\ncombined diff is sets intersection and one of the sets does not have\nthat path.\n\n  ( In usual words p[imin] is not changed between Pj..T - it was\n    e.g. removed in Pj~, so merging parents to T does not bring any new\n    information wrt path p[imin] and that is why we do not want to show\n    p[imin] in combined-diff output - no new change about that path )\n\nSo nothing to append to the output, and update minimum tree entries,\npreparing for the next step.\n\n> > +struct combine_diff_path *diff_tree_paths(\n> > +\tstruct combine_diff_path *p, const unsigned char *sha1,\n> > +\tconst unsigned char **parents_sha1, int nparent,\n> > +\tstruct strbuf *base, struct diff_options *opt)\n> > +{\n> > +\tp = ll_diff_tree_paths(p, sha1, parents_sha1, nparent, base, opt);\n> > +\n> > +\t/*\n> > +\t * free pre-allocated last element, if any\n> > +\t * (see path_appendnew() for details about why)\n> > +\t */\n> > +\tif (p->next) {\n> > +\t\tfree(p->next);\n> > +\t\tp->next = NULL;\n> > +\t}\n> > +\n> > +\treturn p;\n> >  }\n> >  \n> >  /*\n> > @@ -308,6 +664,27 @@ static void try_to_follow_renames(const unsigned char *old, const unsigned char\n> >  \tq->nr = 1;\n> >  }\n> >  \n> > +static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n> > +\t\t\t     struct strbuf *base, struct diff_options *opt)\n> > +{\n> > +\tstruct combine_diff_path phead, *p;\n> > +\tconst unsigned char *parents_sha1[1] = {old};\n> > +\tpathchange_fn_t pathchange_old = opt->pathchange;\n> > +\n> > +\tphead.next = NULL;\n> > +\topt->pathchange = emit_diff_first_parent_only;\n> > +\tdiff_tree_paths(&phead, new, parents_sha1, 1, base, opt);\n> \n> Hmph.  I would have expected\n> \n> \tconst unsigned char **parents_sha1 = &old;\n> \n> or even\n> \n> \tdiff_tree_paths(&phead, new, &old, 1, base, opt);\n> \n> here.\n\nI agree, the last one is better - thanks for spotting this.\n\nI'm attaching corrected patch with renaming and fixups with smaller\nissues you've mentioned.\n\nThanks,\nKirill\n\nP.S. Sorry for maybe some crept-in mistakes - I've tried to verify it\nthoroughly, but am too sleepy to be completely sure. On the other hand I\nthink and hope the patch should be ok.\n\n---- 8< ----\nFrom: Kirill Smelkov <kirr@mns.spb.ru>\nSubject: [PATCH] tree-diff: rework diff_tree() to generate diffs for multiparent cases as well\n\nPreviously diff_tree(), which is now named ll_diff_tree_sha1(), was\ngenerating diff_filepair(s) for two trees t1 and t2, and that was\nusually used for a commit as t1=HEAD~, and t2=HEAD - i.e. to see changes\na commit introduces.\n\nIn Git, however, we have fundamentally built flexibility in that a\ncommit can have many parents - 1 for a plain commit, 2 for a simple merge,\nbut also more than 2 for merging several heads at once.\n\nFor merges there is a so called combine-diff, which shows diff, a merge\nintroduces by itself, omitting changes done by any parent. That works\nthrough first finding paths, that are different to all parents, and then\nshowing generalized diff, with separate columns for +/- for each parent.\nThe code lives in combine-diff.c .\n\nThere is an impedance mismatch, however, in that a commit could\ngenerally have any number of parents, and that while diffing trees, we\ndivide cases for 2-tree diffs and more-than-2-tree diffs. I mean there\nis no special casing for multiple parents commits in e.g.\nrevision-walker .\n\nThat impedance mismatch *hurts* *performance* *badly* for generating\ncombined diffs - in \"combine-diff: optimize combine_diff_path\nsets intersection\" I've already removed some slowness from it, but from\nthe timings provided there, it could be seen, that combined diffs still\ncost more than an order of magnitude more cpu time, compared to diff for\nusual commits, and that would only be an optimistic estimate, if we take\ninto account that for e.g. linux.git there is only one merge for several\ndozens of plain commits.\n\nThat slowness comes from the fact that currently, while generating\ncombined diff, a lot of time is spent computing diff(commit,commit^2)\njust to only then intersect that huge diff to almost small set of files\nfrom diff(commit,commit^1).\n\nThat's because at present, to compute combine-diff, for first finding\npaths, that \"every parent touches\", we use the following combine-diff\nproperty/definition:\n\nD(A,P1...Pn) = D(A,P1) ^ ... ^ D(A,Pn)      (w.r.t. paths)\n\nwhere\n\nD(A,P1...Pn) is combined diff between commit A, and parents Pi\n\nand\n\nD(A,Pi) is usual two-tree diff Pi..A\n\nSo if any of that D(A,Pi) is huge, tracting 1 n-parent combine-diff as n\n1-parent diffs and intersecting results will be slow.\n\nAnd usually, for linux.git and other topic-based workflows, that\nD(A,P2) is huge, because, if merge-base of A and P2, is several dozens\nof merges (from A, via first parent) below, that D(A,P2) will be diffing\nsum of merges from several subsystems to 1 subsystem.\n\nThe solution is to avoid computing n 1-parent diffs, and to find\nchanged-to-all-parents paths via scanning A's and all Pi's trees\nsimultaneously, at each step comparing their entries, and based on that\ncomparison, populate paths result, and deduce we could *skip*\n*recursing* into subdirectories, if at least for 1 parent, sha1 of that\ndir tree is the same as in A. That would save us from doing significant\namount of needless work.\n\nSuch approach is very similar to what diff_tree() does, only there we\ndeal with scanning only 2 trees simultaneously, and for n+1 tree, the\nlogic is a bit more complex:\n\nD(T,P1...Pn) calculation scheme\n-------------------------------\n\nD(T,P1...Pn) = D(T,P1) ^ ... ^ D(T,Pn)\t(regarding resulting paths set)\n\n    D(T,Pj)\t\t- diff between T..Pj\n    D(T,P1...Pn)\t- combined diff from T to parents P1,...,Pn\n\nWe start from all trees, which are sorted, and compare their entries in\nlock-step:\n\n     T     P1       Pn\n     -     -        -\n    |t|   |p1|     |pn|\n    |-|   |--| ... |--|      imin = argmin(p1...pn)\n    | |   |  |     |  |\n    |-|   |--|     |--|\n    |.|   |. |     |. |\n     .     .        .\n     .     .        .\n\nat any time there could be 3 cases:\n\n    1)  t < p[imin];\n    2)  t > p[imin];\n    3)  t = p[imin].\n\nSchematic deduction of what every case means, and what to do, follows:\n\n1)  t < p[imin]  ->  ∀j t ∉ Pj  ->  \"+t\" ∈ D(T,Pj)  ->  D += \"+t\";  t↓\n\n2)  t > p[imin]\n\n    2.1) ∃j: pj > p[imin]  ->  \"-p[imin]\" ∉ D(T,Pj)  ->  D += ø;  ∀ pi=p[imin]  pi↓\n    2.2) ∀i  pi = p[imin]  ->  pi ∉ T  ->  \"-pi\" ∈ D(T,Pi)  ->  D += \"-p[imin]\";  ∀i pi↓\n\n3)  t = p[imin]\n\n    3.1) ∃j: pj > p[imin]  ->  \"+t\" ∈ D(T,Pj)  ->  only pi=p[imin] remains to investigate\n    3.2) pi = p[imin]  ->  investigate δ(t,pi)\n     |\n     |\n     v\n\n    3.1+3.2) looking at δ(t,pi) ∀i: pi=p[imin] - if all != ø  ->\n\n                      ⎧δ(t,pi)  - if pi=p[imin]\n             ->  D += ⎨\n                      ⎩\"+t\"     - if pi>p[imin]\n\n    in any case t↓  ∀ pi=p[imin]  pi↓\n\n~\n\nFor comparison, here is how diff_tree() works:\n\nD(A,B) calculation scheme\n-------------------------\n\n    A     B\n    -     -\n   |a|   |b|    a < b   ->  a ∉ B   ->   D(A,B) +=  +a    a↓\n   |-|   |-|    a > b   ->  b ∉ A   ->   D(A,B) +=  -b    b↓\n   | |   | |    a = b   ->  investigate δ(a,b)            a↓ b↓\n   |-|   |-|\n   |.|   |.|\n    .     .\n    .     .\n\n~~~~~~~~\n\nThis patch generalizes diff tree-walker to work with arbitrary number of\nparents as described above - i.e. now there is a resulting tree t, and\nsome parents trees tp[i] i=[0..nparent). The generalization builds on\nthe fact that usual diff\n\nD(A,B)\n\nis by definition the same as combined diff\n\nD(A,[B]),\n\nso if we could rework the code for common case and make it be not slower\nfor nparent=1 case, usual diff(t1,t2) generation will not be slower, and\nmultiparent diff tree-walker would greatly benefit generating\ncombine-diff.\n\nWhat we do is as follows:\n\n1) diff tree-walker ll_diff_tree_sha1() is internally reworked to be\n   a paths generator (new name diff_tree_paths()), with each generated path\n   being `struct combine_diff_path` with info for path, new sha1,mode and for\n   every parent which sha1,mode it was in it.\n\n2) From that info, we can still generate usual diff queue with\n   struct diff_filepairs, via \"exporting\" generated\n   combine_diff_path, if we know we run for nparent=1 case.\n   (see emit_diff() which is now named emit_diff_first_parent_only())\n\n3) In order for diff_can_quit_early(), which checks\n\n       DIFF_OPT_TST(opt, HAS_CHANGES))\n\n   to work, that exporting have to be happening not in bulk, but\n   incrementally, one diff path at a time.\n\n   For such consumers, there is a new callback in diff_options\n   introduced:\n\n       ->pathchange(opt, struct combine_diff_path *)\n\n   which, if set to !NULL, is called for every generated path.\n\n   (see new compat ll_diff_tree_sha1() wrapper around new paths\n    generator for setup)\n\n4) The paths generation itself, is reworked from previous\n   ll_diff_tree_sha1() code according to \"D(A,P1...Pn) calculation\n   scheme\" provided above:\n\n   On the start we allocate [nparent] arrays in place what was\n   earlier just for one parent tree.\n\n   then we just generalize loops, and comparison according to the\n   algorithm.\n\nSome notes(*):\n\n1) alloca(), for small arrays, is used for \"runs not slower for\n   nparent=1 case than before\" goal - if we change it to xmalloc()/free()\n   the timings get ~1% worse. For alloca() we use just-introduced\n   xalloca/xalloca_free compatibility wrappers, so it should not be a\n   portability problem.\n\n2) For every parent tree, we need to keep a tag, whether entry from that\n   parent equals to entry from minimal parent. For performance reasons I'm\n   keeping that tag in entry's mode field in unused bit - see S_IFXMIN_NEQ.\n   Not doing so, we'd need to alloca another [nparent] array, which hurts\n   performance.\n\n3) For emitted paths, memory could be reused, if we know the path was\n   processed via callback and will not be needed later. We use efficient\n   hand-made realloc-style path_appendnew(), that saves us from ~1-1.5%\n   of potential additional slowdown.\n\n4) goto(s) are used in several places, as the code executes a little bit\n   faster with lowered register pressure.\n\nAlso\n\n- we should now check for FIND_COPIES_HARDER not only when two entries\n  names are the same, and their hashes are equal, but also for a case,\n  when a path was removed from some of all parents having it.\n\n  The reason is, if we don't, that path won't be emitted at all (see\n  \"a > xi\" case), and we'll just skip it, and FIND_COPIES_HARDER wants\n  all paths - with diff or without - to be emitted, to be later analyzed\n  for being copies sources.\n\n  The new check is only necessary for nparent >1, as for nparent=1 case\n  xmin_eqtotal always =1 =nparent, and a path is always added to diff as\n  removal.\n\n~~~~~~~~\n\nTimings for\n\n    # without -c, i.e. testing only nparent=1 case\n    `git log --raw --no-abbrev --no-renames`\n\nbefore and after the patch are as follows:\n\n                navy.git        linux.git v3.10..v3.11\n\n    before      0.611s          1.889s\n    after       0.619s          1.907s\n    slowdown    1.3%            0.9%\n\nThis timings show we did no harm to usual diff(tree1,tree2) generation.\nFrom the table we can see that we actually did ~1% slowdown, but I think\nI've \"earned\" that 1% in the previous patch (\"tree-diff: reuse base\nstr(buf) memory on sub-tree recursion\", HEAD~~) so for nparent=1 case,\nnet timings stays approximately the same.\n\nThe output also stayed the same.\n\n(*) If we revert 1)-4) to more usual techniques, for nparent=1 case,\n    we'll get ~2-2.5% of additional slowdown, which I've tried to avoid, as\n   \"do no harm for nparent=1 case\" rule.\n\nFor linux.git, combined diff will run an order of magnitude faster and\nappropriate timings will be provided in the next commit, as we'll be\ntaking advantage of the new diff tree-walker for combined-diff\ngeneration there.\n\nP.S. and combined diff is not some exotic/for-play-only stuff - for\nexample for a program I write to represent Git archives as readonly\nfilesystem, there is initial scan with\n\n    `git log --reverse --raw --no-abbrev --no-renames -c`\n\nto extract log of what was created/changed when, as a result building a\nmap\n\n    {}  sha1    ->  in which commit (and date) a content was added\n\nthat `-c` means also show combined diff for merges, and without them, if\na merge is non-trivial (merges changes from two parents with both having\nseparate changes to a file), or an evil one, the map will not be full,\ni.e. some valid sha1 would be absent from it.\n\nThat case was my initial motivation for combined diffs speedup.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n cache.h     |  15 ++\n diff.c      |   1 +\n diff.h      |   9 ++\n tree-diff.c | 504 ++++++++++++++++++++++++++++++++++++++++++++++++++++--------\n 4 files changed, 465 insertions(+), 64 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex dc040fb..e7f5a0c 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -75,6 +75,21 @@ unsigned long git_deflate_bound(git_zstream *, unsigned long);\n #define S_ISGITLINK(m)\t(((m) & S_IFMT) == S_IFGITLINK)\n \n /*\n+ * Some mode bits are also used internally for computations.\n+ *\n+ * They *must* not overlap with any valid modes, and they *must* not be emitted\n+ * to outside world - i.e. appear on disk or network. In other words, it's just\n+ * temporary fields, which we internally use, but they have to stay in-house.\n+ *\n+ * ( such approach is valid, as standard S_IF* fits into 16 bits, and in Git\n+ *   codebase mode is `unsigned int` which is assumed to be at least 32 bits )\n+ */\n+\n+/* used internally in tree-diff */\n+#define S_DIFFTREE_IFXMIN_NEQ\t0x80000000\n+\n+\n+/*\n  * Intensive research over the course of many years has shown that\n  * port 9418 is totally unused by anything else. Or\n  *\ndiff --git a/diff.c b/diff.c\nindex 8e4a6a9..cda4aa8 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -3216,6 +3216,7 @@ void diff_setup(struct diff_options *options)\n \toptions->context = diff_context_default;\n \tDIFF_OPT_SET(options, RENAME_EMPTY);\n \n+\t/* pathchange left =NULL by default */\n \toptions->change = diff_change;\n \toptions->add_remove = diff_addremove;\n \toptions->use_color = diff_use_color_default;\ndiff --git a/diff.h b/diff.h\nindex 5d7b9f7..0abd735 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -15,6 +15,10 @@ struct diff_filespec;\n struct userdiff_driver;\n struct sha1_array;\n struct commit;\n+struct combine_diff_path;\n+\n+typedef int (*pathchange_fn_t)(struct diff_options *options,\n+\t\t struct combine_diff_path *path);\n \n typedef void (*change_fn_t)(struct diff_options *options,\n \t\t unsigned old_mode, unsigned new_mode,\n@@ -157,6 +161,7 @@ struct diff_options {\n \tint close_file;\n \n \tstruct pathspec pathspec;\n+\tpathchange_fn_t pathchange;\n \tchange_fn_t change;\n \tadd_remove_fn_t add_remove;\n \tdiff_format_fn_t format_callback;\n@@ -189,6 +194,10 @@ const char *diff_line_prefix(struct diff_options *);\n \n extern const char mime_boundary_leader[];\n \n+extern struct combine_diff_path *diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parent_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt);\n extern int diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \t\t\t  const char *base, struct diff_options *opt);\n extern int diff_root_tree_sha1(const unsigned char *new, const char *base,\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 278acc8..e7b378c 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -6,7 +6,19 @@\n #include \"diffcore.h\"\n #include \"tree.h\"\n \n+/*\n+ * internal mode marker, saying a tree entry != entry of tp[imin]\n+ * (see ll_diff_tree_paths for what it means there)\n+ *\n+ * we will update/use/emit entry for diff only with it unset.\n+ */\n+#define S_IFXMIN_NEQ\tS_DIFFTREE_IFXMIN_NEQ\n+\n \n+static struct combine_diff_path *ll_diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parents_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt);\n static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n \t\t\t     struct strbuf *base, struct diff_options *opt);\n \n@@ -42,71 +54,151 @@ static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n }\n \n \n-/* convert path, t1/t2 -> opt->diff_*() callbacks */\n-static void emit_diff(struct diff_options *opt, struct strbuf *path,\n-\t\t      struct tree_desc *t1, struct tree_desc *t2)\n+/*\n+ * convert path -> opt->diff_*() callbacks\n+ *\n+ * emits diff to first parent only, and tells diff tree-walker that we are done\n+ * with p and it can be freed.\n+ */\n+static int emit_diff_first_parent_only(struct diff_options *opt, struct combine_diff_path *p)\n {\n-\tunsigned int mode1 = t1 ? t1->entry.mode : 0;\n-\tunsigned int mode2 = t2 ? t2->entry.mode : 0;\n-\n-\tif (mode1 && mode2) {\n-\t\topt->change(opt, mode1, mode2, t1->entry.sha1, t2->entry.sha1,\n-\t\t\t1, 1, path->buf, 0, 0);\n+\tstruct combine_diff_parent *p0 = &p->parent[0];\n+\tif (p->mode && p0->mode) {\n+\t\topt->change(opt, p0->mode, p->mode, p0->sha1, p->sha1,\n+\t\t\t1, 1, p->path, 0, 0);\n \t}\n \telse {\n \t\tconst unsigned char *sha1;\n \t\tunsigned int mode;\n \t\tint addremove;\n \n-\t\tif (mode2) {\n+\t\tif (p->mode) {\n \t\t\taddremove = '+';\n-\t\t\tsha1 = t2->entry.sha1;\n-\t\t\tmode = mode2;\n+\t\t\tsha1 = p->sha1;\n+\t\t\tmode = p->mode;\n \t\t} else {\n \t\t\taddremove = '-';\n-\t\t\tsha1 = t1->entry.sha1;\n-\t\t\tmode = mode1;\n+\t\t\tsha1 = p0->sha1;\n+\t\t\tmode = p0->mode;\n \t\t}\n \n-\t\topt->add_remove(opt, addremove, mode, sha1, 1, path->buf, 0);\n+\t\topt->add_remove(opt, addremove, mode, sha1, 1, p->path, 0);\n \t}\n+\n+\treturn 0;\t/* we are done with p */\n }\n \n \n-/* new path should be added to diff\n+/*\n+ * Make a new combine_diff_path from path/mode/sha1\n+ * and append it to paths list tail.\n+ *\n+ * Memory for created elements could be reused:\n+ *\n+ *\t- if last->next == NULL, the memory is allocated;\n+ *\n+ *\t- if last->next != NULL, it is assumed that p=last->next was returned\n+ *\t  earlier by this function, and p->next was *not* modified.\n+ *\t  The memory is then reused from p.\n+ *\n+ * so for clients,\n+ *\n+ * - if you do need to keep the element\n+ *\n+ *\tp = path_appendnew(p, ...);\n+ *\tprocess(p);\n+ *\tp->next = NULL;\n+ *\n+ * - if you don't need to keep the element after processing\n+ *\n+ *\tpprev = p;\n+ *\tp = path_appendnew(p, ...);\n+ *\tprocess(p);\n+ *\tp = pprev;\n+ *\t; don't forget to free tail->next in the end\n+ *\n+ * p->parent[] remains uninitialized.\n+ */\n+static struct combine_diff_path *path_appendnew(struct combine_diff_path *last,\n+\tint nparent, const struct strbuf *base, const char *path, int pathlen,\n+\tunsigned mode, const unsigned char *sha1)\n+{\n+\tstruct combine_diff_path *p;\n+\tint len = base->len + pathlen;\n+\tint alloclen = combine_diff_path_size(nparent, len);\n+\n+\t/* if last->next is !NULL - it is a pre-allocated memory, we can reuse */\n+\tp = last->next;\n+\tif (p && (alloclen > (intptr_t)p->next)) {\n+\t\tfree(p);\n+\t\tp = NULL;\n+\t}\n+\n+\tif (!p) {\n+\t\tp = xmalloc(alloclen);\n+\n+\t\t/*\n+\t\t * until we go to it next round, .next holds how many bytes we\n+\t\t * allocated (for faster realloc - we don't need copying old data).\n+\t\t */\n+\t\tp->next = (struct combine_diff_path *)(intptr_t)alloclen;\n+\t}\n+\n+\tlast->next = p;\n+\n+\tp->path = (char *)&(p->parent[nparent]);\n+\tmemcpy(p->path, base->buf, base->len);\n+\tmemcpy(p->path + base->len, path, pathlen);\n+\tp->path[len] = 0;\n+\tp->mode = mode;\n+\thashcpy(p->sha1, sha1 ? sha1 : null_sha1);\n+\n+\treturn p;\n+}\n+\n+/*\n+ * new path should be added to combine diff\n  *\n  * 3 cases on how/when it should be called and behaves:\n  *\n- *\t!t1,  t2\t-> path added, parent lacks it\n- *\t t1, !t2\t-> path removed from parent\n- *\t t1,  t2\t-> path modified\n+ *\t t, !tp\t\t-> path added, all parents lack it\n+ *\t!t,  tp\t\t-> path removed from all parents\n+ *\t t,  tp\t\t-> path modified/added\n+ *\t\t\t   (M for tp[i]=tp[imin], A otherwise)\n  */\n-static void show_path(struct strbuf *base, struct diff_options *opt,\n-\t\t      struct tree_desc *t1, struct tree_desc *t2)\n+static struct combine_diff_path *emit_path(struct combine_diff_path *p,\n+\tstruct strbuf *base, struct diff_options *opt, int nparent,\n+\tstruct tree_desc *t, struct tree_desc *tp,\n+\tint imin)\n {\n \tunsigned mode;\n \tconst char *path;\n+\tconst unsigned char *sha1;\n \tint pathlen;\n \tint old_baselen = base->len;\n-\tint isdir, recurse = 0, emitthis = 1;\n+\tint i, isdir, recurse = 0, emitthis = 1;\n \n \t/* at least something has to be valid */\n-\tassert(t1 || t2);\n+\tassert(t || tp);\n \n-\tif (t2) {\n+\tif (t) {\n \t\t/* path present in resulting tree */\n-\t\ttree_entry_extract(t2, &path, &mode);\n-\t\tpathlen = tree_entry_len(&t2->entry);\n+\t\tsha1 = tree_entry_extract(t, &path, &mode);\n+\t\tpathlen = tree_entry_len(&t->entry);\n \t\tisdir = S_ISDIR(mode);\n \t} else {\n \t\t/*\n-\t\t * a path was removed - take path from parent. Also take\n-\t\t * mode from parent, to decide on recursion.\n+\t\t * a path was removed - take path from imin parent. Also take\n+\t\t * mode from that parent, to decide on recursion(1).\n+\t\t *\n+\t\t * 1) all modes for tp[i]=tp[imin] should be the same wrt\n+\t\t *    S_ISDIR, thanks to base_name_compare().\n \t\t */\n-\t\ttree_entry_extract(t1, &path, &mode);\n-\t\tpathlen = tree_entry_len(&t1->entry);\n+\t\ttree_entry_extract(&tp[imin], &path, &mode);\n+\t\tpathlen = tree_entry_len(&tp[imin].entry);\n \n \t\tisdir = S_ISDIR(mode);\n+\t\tsha1 = NULL;\n \t\tmode = 0;\n \t}\n \n@@ -115,18 +207,81 @@ static void show_path(struct strbuf *base, struct diff_options *opt,\n \t\temitthis = DIFF_OPT_TST(opt, TREE_IN_RECURSIVE);\n \t}\n \n-\tstrbuf_add(base, path, pathlen);\n+\tif (emitthis) {\n+\t\tint keep;\n+\t\tstruct combine_diff_path *pprev = p;\n+\t\tp = path_appendnew(p, nparent, base, path, pathlen, mode, sha1);\n+\n+\t\tfor (i = 0; i < nparent; ++i) {\n+\t\t\t/*\n+\t\t\t * tp[i] is valid, if present and if tp[i]==tp[imin] -\n+\t\t\t * otherwise, we should ignore it.\n+\t\t\t */\n+\t\t\tint tpi_valid = tp && !(tp[i].entry.mode & S_IFXMIN_NEQ);\n+\n+\t\t\tconst unsigned char *sha1_i;\n+\t\t\tunsigned mode_i;\n+\n+\t\t\tp->parent[i].status =\n+\t\t\t\t!t ? DIFF_STATUS_DELETED :\n+\t\t\t\t\ttpi_valid ?\n+\t\t\t\t\t\tDIFF_STATUS_MODIFIED :\n+\t\t\t\t\t\tDIFF_STATUS_ADDED;\n+\n+\t\t\tif (tpi_valid) {\n+\t\t\t\tsha1_i = tp[i].entry.sha1;\n+\t\t\t\tmode_i = tp[i].entry.mode;\n+\t\t\t}\n+\t\t\telse {\n+\t\t\t\tsha1_i = NULL;\n+\t\t\t\tmode_i = 0;\n+\t\t\t}\n+\n+\t\t\tp->parent[i].mode = mode_i;\n+\t\t\thashcpy(p->parent[i].sha1, sha1_i ? sha1_i : null_sha1);\n+\t\t}\n \n-\tif (emitthis)\n-\t\temit_diff(opt, base, t1, t2);\n+\t\tkeep = 1;\n+\t\tif (opt->pathchange)\n+\t\t\tkeep = opt->pathchange(opt, p);\n+\n+\t\t/*\n+\t\t * If a path was filtered or consumed - we don't need to add it\n+\t\t * to the list and can reuse its memory, leaving it as\n+\t\t * pre-allocated element on the tail.\n+\t\t *\n+\t\t * On the other hand, if path needs to be kept, we need to\n+\t\t * correct its .next to NULL, as it was pre-initialized to how\n+\t\t * much memory was allocated.\n+\t\t *\n+\t\t * see path_appendnew() for details.\n+\t\t */\n+\t\tif (!keep)\n+\t\t\tp = pprev;\n+\t\telse\n+\t\t\tp->next = NULL;\n+\t}\n \n \tif (recurse) {\n+\t\tconst unsigned char **parents_sha1;\n+\n+\t\tparents_sha1 = xalloca(nparent * sizeof(parents_sha1[0]));\n+\t\tfor (i = 0; i < nparent; ++i) {\n+\t\t\t/* same rule as in emitthis */\n+\t\t\tint tpi_valid = tp && !(tp[i].entry.mode & S_IFXMIN_NEQ);\n+\n+\t\t\tparents_sha1[i] = tpi_valid ? tp[i].entry.sha1\n+\t\t\t\t\t\t    : NULL;\n+\t\t}\n+\n+\t\tstrbuf_add(base, path, pathlen);\n \t\tstrbuf_addch(base, '/');\n-\t\tll_diff_tree_sha1(t1 ? t1->entry.sha1 : NULL,\n-\t\t\t\t  t2 ? t2->entry.sha1 : NULL, base, opt);\n+\t\tp = ll_diff_tree_paths(p, sha1, parents_sha1, nparent, base, opt);\n+\t\txalloca_free(parents_sha1);\n \t}\n \n \tstrbuf_setlen(base, old_baselen);\n+\treturn p;\n }\n \n static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n@@ -145,59 +300,260 @@ static void skip_uninteresting(struct tree_desc *t, struct strbuf *base,\n \t}\n }\n \n-static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n-\t\t\t     struct strbuf *base, struct diff_options *opt)\n+\n+/*\n+ * generate paths for combined diff D(sha1,parents_sha1[])\n+ *\n+ * Resulting paths are appended to combine_diff_path linked list, and also, are\n+ * emitted on the go via opt->pathchange() callback, so it is possible to\n+ * process the result as batch or incrementally.\n+ *\n+ * The paths are generated scanning new tree and all parents trees\n+ * simultaneously, similarly to what diff_tree() was doing for 2 trees.\n+ * The theory behind such scan is as follows:\n+ *\n+ *\n+ * D(T,P1...Pn) calculation scheme\n+ * -------------------------------\n+ *\n+ * D(T,P1...Pn) = D(T,P1) ^ ... ^ D(T,Pn)\t(regarding resulting paths set)\n+ *\n+ *\tD(T,Pj)\t\t- diff between T..Pj\n+ *\tD(T,P1...Pn)\t- combined diff from T to parents P1,...,Pn\n+ *\n+ *\n+ * We start from all trees, which are sorted, and compare their entries in\n+ * lock-step:\n+ *\n+ *\t T     P1       Pn\n+ *\t -     -        -\n+ *\t|t|   |p1|     |pn|\n+ *\t|-|   |--| ... |--|      imin = argmin(p1...pn)\n+ *\t| |   |  |     |  |\n+ *\t|-|   |--|     |--|\n+ *\t|.|   |. |     |. |\n+ *\t .     .        .\n+ *\t .     .        .\n+ *\n+ * at any time there could be 3 cases:\n+ *\n+ *\t1)  t < p[imin];\n+ *\t2)  t > p[imin];\n+ *\t3)  t = p[imin].\n+ *\n+ * Schematic deduction of what every case means, and what to do, follows:\n+ *\n+ * 1)  t < p[imin]  ->  ∀j t ∉ Pj  ->  \"+t\" ∈ D(T,Pj)  ->  D += \"+t\";  t↓\n+ *\n+ * 2)  t > p[imin]\n+ *\n+ *     2.1) ∃j: pj > p[imin]  ->  \"-p[imin]\" ∉ D(T,Pj)  ->  D += ø;  ∀ pi=p[imin]  pi↓\n+ *     2.2) ∀i  pi = p[imin]  ->  pi ∉ T  ->  \"-pi\" ∈ D(T,Pi)  ->  D += \"-p[imin]\";  ∀i pi↓\n+ *\n+ * 3)  t = p[imin]\n+ *\n+ *     3.1) ∃j: pj > p[imin]  ->  \"+t\" ∈ D(T,Pj)  ->  only pi=p[imin] remains to investigate\n+ *     3.2) pi = p[imin]  ->  investigate δ(t,pi)\n+ *      |\n+ *      |\n+ *      v\n+ *\n+ *     3.1+3.2) looking at δ(t,pi) ∀i: pi=p[imin] - if all != ø  ->\n+ *\n+ *                       ⎧δ(t,pi)  - if pi=p[imin]\n+ *              ->  D += ⎨\n+ *                       ⎩\"+t\"     - if pi>p[imin]\n+ *\n+ *\n+ *     in any case t↓  ∀ pi=p[imin]  pi↓\n+ *\n+ *\n+ * ~~~~~~~~\n+ *\n+ * NOTE\n+ *\n+ *\tUsual diff D(A,B) is by definition the same as combined diff D(A,[B]),\n+ *\tso this diff paths generator can, and is used, for plain diffs\n+ *\tgeneration too.\n+ *\n+ *\tPlease keep attention to the common D(A,[B]) case when working on the\n+ *\tcode, in order not to slow it down.\n+ *\n+ * NOTE\n+ *\tnparent must be > 0.\n+ */\n+\n+\n+/* ∀ pi=p[imin]  pi↓ */\n+static inline void update_tp_entries(struct tree_desc *tp, int nparent)\n {\n-\tstruct tree_desc t1, t2;\n-\tvoid *t1tree, *t2tree;\n+\tint i;\n+\tfor (i = 0; i < nparent; ++i)\n+\t\tif (!(tp[i].entry.mode & S_IFXMIN_NEQ))\n+\t\t\tupdate_tree_entry(&tp[i]);\n+}\n \n-\tt1tree = fill_tree_descriptor(&t1, old);\n-\tt2tree = fill_tree_descriptor(&t2, new);\n+static struct combine_diff_path *ll_diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parents_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt)\n+{\n+\tstruct tree_desc t, *tp;\n+\tvoid *ttree, **tptree;\n+\tint i;\n+\n+\ttp     = xalloca(nparent * sizeof(tp[0]));\n+\ttptree = xalloca(nparent * sizeof(tptree[0]));\n+\n+\t/*\n+\t * load parents first, as they are probably already cached.\n+\t *\n+\t * ( log_tree_diff() parses commit->parent before calling here via\n+\t *   diff_tree_sha1(parent, commit) )\n+\t */\n+\tfor (i = 0; i < nparent; ++i)\n+\t\ttptree[i] = fill_tree_descriptor(&tp[i], parents_sha1[i]);\n+\tttree = fill_tree_descriptor(&t, sha1);\n \n \t/* Enable recursion indefinitely */\n \topt->pathspec.recursive = DIFF_OPT_TST(opt, RECURSIVE);\n \n \tfor (;;) {\n-\t\tint cmp;\n+\t\tint imin, cmp;\n \n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n+\n \t\tif (opt->pathspec.nr) {\n-\t\t\tskip_uninteresting(&t1, base, opt);\n-\t\t\tskip_uninteresting(&t2, base, opt);\n+\t\t\tskip_uninteresting(&t, base, opt);\n+\t\t\tfor (i = 0; i < nparent; i++)\n+\t\t\t\tskip_uninteresting(&tp[i], base, opt);\n+\t\t}\n+\n+\t\t/* comparing is finished when all trees are done */\n+\t\tif (!t.size) {\n+\t\t\tint done = 1;\n+\t\t\tfor (i = 0; i < nparent; ++i)\n+\t\t\t\tif (tp[i].size) {\n+\t\t\t\t\tdone = 0;\n+\t\t\t\t\tbreak;\n+\t\t\t\t}\n+\t\t\tif (done)\n+\t\t\t\tbreak;\n+\t\t}\n+\n+\t\t/*\n+\t\t * lookup imin = argmin(p1...pn),\n+\t\t * mark entries whether they =p[imin] along the way\n+\t\t */\n+\t\timin = 0;\n+\t\ttp[0].entry.mode &= ~S_IFXMIN_NEQ;\n+\n+\t\tfor (i = 1; i < nparent; ++i) {\n+\t\t\tcmp = tree_entry_pathcmp(&tp[i], &tp[imin]);\n+\t\t\tif (cmp < 0) {\n+\t\t\t\timin = i;\n+\t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n+\t\t\t}\n+\t\t\telse if (cmp == 0) {\n+\t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n+\t\t\t}\n+\t\t\telse {\n+\t\t\t\ttp[i].entry.mode |= S_IFXMIN_NEQ;\n+\t\t\t}\n \t\t}\n-\t\tif (!t1.size && !t2.size)\n-\t\t\tbreak;\n \n-\t\tcmp = tree_entry_pathcmp(&t1, &t2);\n+\t\t/* fixup markings for entries before imin */\n+\t\tfor (i = 0; i < imin; ++i)\n+\t\t\ttp[i].entry.mode |= S_IFXMIN_NEQ;\t/* pi > p[imin] */\n \n-\t\t/* t1 = t2 */\n-\t\tif (cmp == 0) {\n-\t\t\tif (DIFF_OPT_TST(opt, FIND_COPIES_HARDER) ||\n-\t\t\t    hashcmp(t1.entry.sha1, t2.entry.sha1) ||\n-\t\t\t    (t1.entry.mode != t2.entry.mode))\n-\t\t\t\tshow_path(base, opt, &t1, &t2);\n \n-\t\t\tupdate_tree_entry(&t1);\n-\t\t\tupdate_tree_entry(&t2);\n+\n+\t\t/* compare t vs p[imin] */\n+\t\tcmp = tree_entry_pathcmp(&t, &tp[imin]);\n+\n+\t\t/* t = p[imin] */\n+\t\tif (cmp == 0) {\n+\t\t\t/* are either pi > p[imin] or diff(t,pi) != ø ? */\n+\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n+\t\t\t\tfor (i = 0; i < nparent; ++i) {\n+\t\t\t\t\t/* p[i] > p[imin] */\n+\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n+\t\t\t\t\t\tcontinue;\n+\n+\t\t\t\t\t/* diff(t,pi) != ø */\n+\t\t\t\t\tif (hashcmp(t.entry.sha1, tp[i].entry.sha1) ||\n+\t\t\t\t\t    (t.entry.mode != tp[i].entry.mode))\n+\t\t\t\t\t\tcontinue;\n+\n+\t\t\t\t\tgoto skip_emit_t_tp;\n+\t\t\t\t}\n+\t\t\t}\n+\n+\t\t\t/* D += {δ(t,pi) if pi=p[imin];  \"+a\" if pi > p[imin]} */\n+\t\t\tp = emit_path(p, base, opt, nparent,\n+\t\t\t\t\t&t, tp, imin);\n+\n+\t\tskip_emit_t_tp:\n+\t\t\t/* t↓,  ∀ pi=p[imin]  pi↓ */\n+\t\t\tupdate_tree_entry(&t);\n+\t\t\tupdate_tp_entries(tp, nparent);\n \t\t}\n \n-\t\t/* t1 < t2 */\n+\t\t/* t < p[imin] */\n \t\telse if (cmp < 0) {\n-\t\t\tshow_path(base, opt, &t1, /*t2=*/NULL);\n-\t\t\tupdate_tree_entry(&t1);\n+\t\t\t/* D += \"+t\" */\n+\t\t\tp = emit_path(p, base, opt, nparent,\n+\t\t\t\t\t&t, /*tp=*/NULL, -1);\n+\n+\t\t\t/* t↓ */\n+\t\t\tupdate_tree_entry(&t);\n \t\t}\n \n-\t\t/* t1 > t2 */\n+\t\t/* t > p[imin] */\n \t\telse {\n-\t\t\tshow_path(base, opt, /*t1=*/NULL, &t2);\n-\t\t\tupdate_tree_entry(&t2);\n+\t\t\t/* ∀i pi=p[imin] -> D += \"-p[imin]\" */\n+\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n+\t\t\t\tfor (i = 0; i < nparent; ++i)\n+\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n+\t\t\t\t\t\tgoto skip_emit_tp;\n+\t\t\t}\n+\n+\t\t\tp = emit_path(p, base, opt, nparent,\n+\t\t\t\t\t/*t=*/NULL, tp, imin);\n+\n+\t\tskip_emit_tp:\n+\t\t\t/* ∀ pi=p[imin]  pi↓ */\n+\t\t\tupdate_tp_entries(tp, nparent);\n \t\t}\n \t}\n \n-\tfree(t2tree);\n-\tfree(t1tree);\n-\treturn 0;\n+\tfree(ttree);\n+\tfor (i = nparent-1; i >= 0; i--)\n+\t\tfree(tptree[i]);\n+\txalloca_free(tptree);\n+\txalloca_free(tp);\n+\n+\treturn p;\n+}\n+\n+struct combine_diff_path *diff_tree_paths(\n+\tstruct combine_diff_path *p, const unsigned char *sha1,\n+\tconst unsigned char **parents_sha1, int nparent,\n+\tstruct strbuf *base, struct diff_options *opt)\n+{\n+\tp = ll_diff_tree_paths(p, sha1, parents_sha1, nparent, base, opt);\n+\n+\t/*\n+\t * free pre-allocated last element, if any\n+\t * (see path_appendnew() for details about why)\n+\t */\n+\tif (p->next) {\n+\t\tfree(p->next);\n+\t\tp->next = NULL;\n+\t}\n+\n+\treturn p;\n }\n \n /*\n@@ -308,6 +664,26 @@ static void try_to_follow_renames(const unsigned char *old, const unsigned char\n \tq->nr = 1;\n }\n \n+static int ll_diff_tree_sha1(const unsigned char *old, const unsigned char *new,\n+\t\t\t     struct strbuf *base, struct diff_options *opt)\n+{\n+\tstruct combine_diff_path phead, *p;\n+\tpathchange_fn_t pathchange_old = opt->pathchange;\n+\n+\tphead.next = NULL;\n+\topt->pathchange = emit_diff_first_parent_only;\n+\tdiff_tree_paths(&phead, new, &old, 1, base, opt);\n+\n+\tfor (p = phead.next; p;) {\n+\t\tstruct combine_diff_path *pprev = p;\n+\t\tp = p->next;\n+\t\tfree(pprev);\n+\t}\n+\n+\topt->pathchange = pathchange_old;\n+\treturn 0;\n+}\n+\n int diff_tree_sha1(const unsigned char *old, const unsigned char *new, const char *base_str, struct diff_options *opt)\n {\n \tstruct strbuf base;\n-- \n1.9.rc0.143.g6fd479e\n"},{"id":"238476","messageId":"xmqqr459m4j9.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"20140406214626.GA3843@mini.zxlink","subject":"Re: [PATCH v2 18/19] tree-diff: rework diff_tree() to generate diffs for multiparent cases as well","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-04-07T17:29:46Z","receivedAt":"2014-04-07T17:29:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@navytux.spb.ru> writes:\n\n> The following\n> ...\n> maybe looks a bit simpler, but calls tree_entry_pathcmp twice more times.\n>\n> Besides for important nparent=1 case we were not calling\n> tree_entry_pathcmp at all and here we'll call it once, which would slow\n> execution down a bit, as base_name_compare shows measurable enough in profile.\n> To avoid that we'll need to add 'if (i==imin) continue' and this won't\n> be so simple then. And for general nparent case, as I've said, we'll be\n> calling tree_entry_pathcmp twice more times...\n>\n> Because of all that I'd suggest to go with my original version.\n\nOK.\n\n> ... After some break on the topic,\n> with a fresh eye I see a lot of confusion goes from the notation I've\n> chosen initially (because of how I was reasoning about it on paper, when\n> it was in flux) - i.e. xi for x[imin] and also using i as looping\n> variable. And also because xi was already used for x[imin] I've used\n> another letter 'k' denoting all other x'es, which leads to confusion...\n>\n>\n> I propose we do the following renaming to clarify things:\n>\n>     A/a     ->      T/t     (to match resulting tree t name in the code)\n>     X/x     ->      P/p     (to match parents trees tp in the code)\n>     i       ->      imin    (so that i would be free for other tasks)\n>\n> then the above (with a prologue) would look like\n>\n> ---- 8< ----\n>  *       T     P1       Pn\n>  *       -     -        -\n>  *      |t|   |p1|     |pn|\n>  *      |-|   |--| ... |--|      imin = argmin(p1...pn)\n>  *      | |   |  |     |  |\n>  *      |-|   |--|     |--|\n>  *      |.|   |. |     |. |\n>  *       .     .        .\n>  *       .     .        .\n>  *\n>  * at any time there could be 3 cases:\n>  *\n>  *      1)  t < p[imin];\n>  *      2)  t > p[imin];\n>  *      3)  t = p[imin].\n>  *\n>  * Schematic deduction of what every case means, and what to do, follows:\n>  *\n>  * 1)  t < p[imin]  ->  ∀j t ∉ Pj  ->  \"+t\" ∈ D(T,Pj)  ->  D += \"+t\";  t↓\n>  *\n>  * 2)  t > p[imin]\n>  *\n>  *     2.1) ∃j: pj > p[imin]  ->  \"-p[imin]\" ∉ D(T,Pj)  ->  D += ø;  ∀ pi=p[imin]  pi↓\n>  *     2.2) ∀i  pi = p[imin]  ->  pi ∉ T  ->  \"-pi\" ∈ D(T,Pi)  ->  D += \"-p[imin]\";  ∀i pi↓\n>  *\n>  * 3)  t = p[imin]\n>  *\n>  *     3.1) ∃j: pj > p[imin]  ->  \"+t\" ∈ D(T,Pj)  ->  only pi=p[imin] remains to investigate\n>  *     3.2) pi = p[imin]  ->  investigate δ(t,pi)\n>  *      |\n>  *      |\n>  *      v\n>  *\n>  *     3.1+3.2) looking at δ(t,pi) ∀i: pi=p[imin] - if all != ø  ->\n>  *\n>  *                       ⎧δ(t,pi)  - if pi=p[imin]\n>  *              ->  D += ⎨\n>  *                       ⎩\"+t\"     - if pi>p[imin]\n>  *\n>  *\n>  *     in any case t↓  ∀ pi=p[imin]  pi↓\n> ...\n> now xk is gone and i matches p[i] (= pi) etc so variable names correlate\n> to algorithm description better.\n>\n> Does that maybe clarify things?\n\nThat sounds more consistent (modulo perhaps s/argmin/min/ at the\nbeginning?).\n\n> P.S. Sorry for maybe some crept-in mistakes - I've tried to verify it\n> thoroughly, but am too sleepy to be completely sure. On the other hand I\n> think and hope the patch should be ok.\n\nThanks and do not be sorry for \"mistakes\"---we have the review\nprocess exactly for catching them.\n"},{"id":"238477","messageId":"xmqqmwfxm2rw.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"20140406214626.GA3843@mini.zxlink","subject":"Re: [PATCH v2 18/19] tree-diff: rework diff_tree() to generate diffs for multiparent cases as well","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-04-07T18:07:47Z","receivedAt":"2014-04-07T18:07:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kirill Smelkov <kirr@navytux.spb.ru> writes:\n\n>> > +\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n>> > +\t\t\t\tfor (i = 0; i < nparent; ++i)\n>> > +\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n>> > +\t\t\t\t\t\tgoto skip_emit_tp;\n>> > +\t\t\t}\n>> > +\n>> > +\t\t\tp = emit_path(p, base, opt, nparent,\n>> > +\t\t\t\t\t/*t=*/NULL, tp, imin);\n>> > +\n>> > +\t\tskip_emit_tp:\n>> > +\t\t\t/* ∀ xk=ximin  xk↓ */\n>> > +\t\t\tupdate_tp_entries(tp, nparent);\n>> \n>> There are parents whose path sort earlier than what is in 't'\n>> (i.e. they were lost in the result---we would want to show\n>> removal).  What makes us jump to the skip label?\n>> \n>>     We are looking at path in 't', and some parents have paths that\n>>     sort earlier than that path.  We will not go to skip label if\n>>     any one of the parent's entry sorts after some other parent (or\n>>     the parent in question has ran out its entries), which means we\n>>     show the entry from the parents only when all the parents have\n>>     that same path, which is missing from 't'.\n>> \n>> I am not sure if I am reading this correctly, though.\n>> \n>> For the two-way diff, the above degenerates to \"show all parent\n>> entries that come before the first entry in 't'\", which is correct.\n>> For the combined diff, the current intersect_paths() makes sure that\n>> each path appears in all the pair-wise diff between t and tp[],\n>> which again means that the above logic match the current behaviour.\n>\n> Yes, correct (modulo we *will* go to skip label if any one of the\n> parent's entry sorts after some other parent). By definition of combined\n> diff we show a path only if it shows in every diff D(T,Pi), and if \n>\n>     2.1) ∃j: pj > p[imin]  ->  \"-p[imin]\" ∉ D(T,Pj)  ->  D += ø;  ∀ pi=p[imin]  pi↓\n>\n> some pj sorts after p[imin] that would mean that Pj does not have\n> p[imin] and since t > p[imin] (which means T does not have p[imin]\n> either) diff D(T,Pj) does not have p[imin]. And because of that we know\n> the whole combined-diff will not have p[imin] as, by definition,\n> combined diff is sets intersection and one of the sets does not have\n> that path.\n>\n>   ( In usual words p[imin] is not changed between Pj..T - it was\n>     e.g. removed in Pj~, so merging parents to T does not bring any new\n>     information wrt path p[imin] and that is why we do not want to show\n>     p[imin] in combined-diff output - no new change about that path )\n>\n> So nothing to append to the output, and update minimum tree entries,\n> preparing for the next step.\n\nThat's all in line with the current and traditional definition of\ncombined diff.\n\nThis is a tangent that is outside the scope of this current topic,\nbut I wonder if you found it disturbing that we treat the result 't'\nthat has a path and the result 't' that does not have a path with\nrespect to a parent that does not have the path in a somewhat\nassymmetric way.\n\nWith a merge M between commits A and B, where they all have the same\npath with different contents, we obviously show that path in the\ncombined diff format.  A merge N that records exactly the same tree\nas M that merges the same commits A and B plus another commit C that\ndoes not have that path still shows the combined diff, with one\nextra column to express \"everything in the result N has been added\nwith respect to C which did not have the path at all\".\n\nHowever, a merge O between the same commits A and B, where A and B\nhave a path and O loses it, shows the path in the combined format.\nA merge P among the same A, B and an extra parent C that does not\nhave that path ceases to show it (this is the assymmetry).\n\nIt is a natural extension of \"Do not show the path when the result\nmatches one of the parent\" rule, and in this case the result P takes\ncontents, \"the path does not exist\", from one parent \"C\", so it is\ninternally consistent, and I originally designed it that way on\npurpose, but somehow it feels a bit strange.\n"},{"id":"238510","messageId":"20140407202616.GA4140@mini.zxlink","threadId":"35947","inReplyTo":"xmqqmwfxm2rw.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH v2 18/19] tree-diff: rework diff_tree() to generate diffs for multiparent cases as well","fromName":"Kirill Smelkov","fromEmail":"kirr@navytux.spb.ru","sentAt":"2014-04-07T20:26:17Z","receivedAt":"2014-04-07T20:26:17Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Mon, Apr 07, 2014 at 10:29:46AM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n> \n> > The following\n> > ...\n> > maybe looks a bit simpler, but calls tree_entry_pathcmp twice more times.\n> >\n> > Besides for important nparent=1 case we were not calling\n> > tree_entry_pathcmp at all and here we'll call it once, which would slow\n> > execution down a bit, as base_name_compare shows measurable enough in profile.\n> > To avoid that we'll need to add 'if (i==imin) continue' and this won't\n> > be so simple then. And for general nparent case, as I've said, we'll be\n> > calling tree_entry_pathcmp twice more times...\n> >\n> > Because of all that I'd suggest to go with my original version.\n> \n> OK.\n\nThanks.\n\n> > ... After some break on the topic,\n> > with a fresh eye I see a lot of confusion goes from the notation I've\n> > chosen initially (because of how I was reasoning about it on paper, when\n> > it was in flux) - i.e. xi for x[imin] and also using i as looping\n> > variable. And also because xi was already used for x[imin] I've used\n> > another letter 'k' denoting all other x'es, which leads to confusion...\n> >\n> >\n> > I propose we do the following renaming to clarify things:\n> >\n> >     A/a     ->      T/t     (to match resulting tree t name in the code)\n> >     X/x     ->      P/p     (to match parents trees tp in the code)\n> >     i       ->      imin    (so that i would be free for other tasks)\n> >\n> > then the above (with a prologue) would look like\n> >\n> > ---- 8< ----\n> >  *       T     P1       Pn\n> >  *       -     -        -\n> >  *      |t|   |p1|     |pn|\n> >  *      |-|   |--| ... |--|      imin = argmin(p1...pn)\n> >  *      | |   |  |     |  |\n> >  *      |-|   |--|     |--|\n> >  *      |.|   |. |     |. |\n> >  *       .     .        .\n> >  *       .     .        .\n> >  *\n> >  * at any time there could be 3 cases:\n> >  *\n> >  *      1)  t < p[imin];\n> >  *      2)  t > p[imin];\n> >  *      3)  t = p[imin].\n> >  *\n> >  * Schematic deduction of what every case means, and what to do, follows:\n> >  *\n> >  * 1)  t < p[imin]  ->  ∀j t ∉ Pj  ->  \"+t\" ∈ D(T,Pj)  ->  D += \"+t\";  t↓\n> >  *\n> >  * 2)  t > p[imin]\n> >  *\n> >  *     2.1) ∃j: pj > p[imin]  ->  \"-p[imin]\" ∉ D(T,Pj)  ->  D += ø;  ∀ pi=p[imin]  pi↓\n> >  *     2.2) ∀i  pi = p[imin]  ->  pi ∉ T  ->  \"-pi\" ∈ D(T,Pi)  ->  D += \"-p[imin]\";  ∀i pi↓\n> >  *\n> >  * 3)  t = p[imin]\n> >  *\n> >  *     3.1) ∃j: pj > p[imin]  ->  \"+t\" ∈ D(T,Pj)  ->  only pi=p[imin] remains to investigate\n> >  *     3.2) pi = p[imin]  ->  investigate δ(t,pi)\n> >  *      |\n> >  *      |\n> >  *      v\n> >  *\n> >  *     3.1+3.2) looking at δ(t,pi) ∀i: pi=p[imin] - if all != ø  ->\n> >  *\n> >  *                       ⎧δ(t,pi)  - if pi=p[imin]\n> >  *              ->  D += ⎨\n> >  *                       ⎩\"+t\"     - if pi>p[imin]\n> >  *\n> >  *\n> >  *     in any case t↓  ∀ pi=p[imin]  pi↓\n> > ...\n> > now xk is gone and i matches p[i] (= pi) etc so variable names correlate\n> > to algorithm description better.\n> >\n> > Does that maybe clarify things?\n> \n> That sounds more consistent (modulo perhaps s/argmin/min/ at the\n> beginning?).\n\nThanks. argmin is there on purpose - min(p1...pn) is the minimal p, and\nargmin(p1...pn) is imin such that p[imin] is minimal. As we are finding\nthe index of the minimal element we should use argmin.\n\n\n> > P.S. Sorry for maybe some crept-in mistakes - I've tried to verify it\n> > thoroughly, but am too sleepy to be completely sure. On the other hand I\n> > think and hope the patch should be ok.\n> \n> Thanks and do not be sorry for \"mistakes\"---we have the review\n> process exactly for catching them.\n\nThanks, I appreciate that.\n\n\nOn Mon, Apr 07, 2014 at 11:07:47AM -0700, Junio C Hamano wrote:\n> Kirill Smelkov <kirr@navytux.spb.ru> writes:\n> \n> >> > +\t\t\tif (!DIFF_OPT_TST(opt, FIND_COPIES_HARDER)) {\n> >> > +\t\t\t\tfor (i = 0; i < nparent; ++i)\n> >> > +\t\t\t\t\tif (tp[i].entry.mode & S_IFXMIN_NEQ)\n> >> > +\t\t\t\t\t\tgoto skip_emit_tp;\n> >> > +\t\t\t}\n> >> > +\n> >> > +\t\t\tp = emit_path(p, base, opt, nparent,\n> >> > +\t\t\t\t\t/*t=*/NULL, tp, imin);\n> >> > +\n> >> > +\t\tskip_emit_tp:\n> >> > +\t\t\t/* ∀ xk=ximin  xk↓ */\n> >> > +\t\t\tupdate_tp_entries(tp, nparent);\n> >> \n> >> There are parents whose path sort earlier than what is in 't'\n> >> (i.e. they were lost in the result---we would want to show\n> >> removal).  What makes us jump to the skip label?\n> >> \n> >>     We are looking at path in 't', and some parents have paths that\n> >>     sort earlier than that path.  We will not go to skip label if\n> >>     any one of the parent's entry sorts after some other parent (or\n> >>     the parent in question has ran out its entries), which means we\n> >>     show the entry from the parents only when all the parents have\n> >>     that same path, which is missing from 't'.\n> >> \n> >> I am not sure if I am reading this correctly, though.\n> >> \n> >> For the two-way diff, the above degenerates to \"show all parent\n> >> entries that come before the first entry in 't'\", which is correct.\n> >> For the combined diff, the current intersect_paths() makes sure that\n> >> each path appears in all the pair-wise diff between t and tp[],\n> >> which again means that the above logic match the current behaviour.\n> >\n> > Yes, correct (modulo we *will* go to skip label if any one of the\n> > parent's entry sorts after some other parent). By definition of combined\n> > diff we show a path only if it shows in every diff D(T,Pi), and if \n> >\n> >     2.1) ∃j: pj > p[imin]  ->  \"-p[imin]\" ∉ D(T,Pj)  ->  D += ø;  ∀ pi=p[imin]  pi↓\n> >\n> > some pj sorts after p[imin] that would mean that Pj does not have\n> > p[imin] and since t > p[imin] (which means T does not have p[imin]\n> > either) diff D(T,Pj) does not have p[imin]. And because of that we know\n> > the whole combined-diff will not have p[imin] as, by definition,\n> > combined diff is sets intersection and one of the sets does not have\n> > that path.\n> >\n> >   ( In usual words p[imin] is not changed between Pj..T - it was\n> >     e.g. removed in Pj~, so merging parents to T does not bring any new\n> >     information wrt path p[imin] and that is why we do not want to show\n> >     p[imin] in combined-diff output - no new change about that path )\n> >\n> > So nothing to append to the output, and update minimum tree entries,\n> > preparing for the next step.\n> \n> That's all in line with the current and traditional definition of\n> combined diff.\n\nOk.\n\n\n> This is a tangent that is outside the scope of this current topic,\n> but I wonder if you found it disturbing that we treat the result 't'\n> that has a path and the result 't' that does not have a path with\n> respect to a parent that does not have the path in a somewhat\n> assymmetric way.\n> \n> With a merge M between commits A and B, where they all have the same\n> path with different contents, we obviously show that path in the\n> combined diff format.  A merge N that records exactly the same tree\n> as M that merges the same commits A and B plus another commit C that\n> does not have that path still shows the combined diff, with one\n> extra column to express \"everything in the result N has been added\n> with respect to C which did not have the path at all\".\n> \n> However, a merge O between the same commits A and B, where A and B\n> have a path and O loses it, shows the path in the combined format.\n> A merge P among the same A, B and an extra parent C that does not\n> have that path ceases to show it (this is the assymmetry).\n\nSymmetry properties are very important and if something does not fit into\nsymmetry - then something somewhere is really wrong for sure, but I\nthink there is no asymmetry here. In your example\n\n      M∆∙    N∆∙\n     / \\ . '''\n    / .'\\  '  '\n    A∆   B∙    Cø\n    \\ ''/. '  '\n     \\ /   '''\n      Oø     Pø\n\nlet's say a path can be:\n\n    ø   - empty\n    ∆   - triangle\n    ∙   - bullet\n    ∆∙  - triangle+bullet\n\nthen some symmetry operation is\n\n    ∆∙  <-> ø\n    ∙   <-> ∙\n    ∆   <-> ∆\n\nso you \"mirror\" ∆∙ to ø (M,N->O,P), and so ø is mirrored back to ∆∙\n(O,P->M,N). But then, when we mirror the whole graph, Cø should be\nmirrored to C'∆∙ and then that would be correct:\n\n    A∆   B∙    C'∆∙\n    \\ ''/. '  '\n     \\ /   '''\n      Oø     Pø\n\nP to A,B,C' would show the path as N to A,B,C(without')\n\n\nIn other words a merge P among the same A, B and extra parent C' with\n\"contains everything\" content is symmetrical to merge N to A,B and C\nwith empty content.\n\n\n> It is a natural extension of \"Do not show the path when the result\n> matches one of the parent\" rule, and in this case the result P takes\n> contents, \"the path does not exist\", from one parent \"C\", so it is\n> internally consistent, and I originally designed it that way on\n> purpose, but somehow it feels a bit strange.\n\nI hope it does not fill strange once the true symmetry is discovered,\nand also P does not add a change compared to C, so the combined diff\nshould be empty as no new information is present in a merge itself.\n\nA bit unusual on the first glance, but not strange, once you are used to\nit and consistent.\n\nThanks,\nKirill\n"},{"id":"238579","messageId":"20140409124827.GA24672@tugrik.mns.mnsspb.ru","threadId":"35947","inReplyTo":"20140327142250.GC17333@mini.zxlink","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Kirill Smelkov","fromEmail":"kirr@mns.spb.ru","sentAt":"2014-04-09T12:48:27Z","receivedAt":"2014-04-09T12:48:27Z","isPatch":true,"sender":{"key":"kirr@navytux.spb.ru","avatar":"https://gravatar.com/avatar/cf3445fdad1849941e17ab25bf1ee7c5ea1be2deb444be25ff5f36b0e50a985f?d=mp&s=160"},"body":"On Thu, Mar 27, 2014 at 06:22:50PM +0400, Kirill Smelkov wrote:\n> On Mon, Mar 24, 2014 at 02:47:24PM -0700, Junio C Hamano wrote:\n> > Kirill Smelkov <kirr@mns.spb.ru> writes:\n> > \n> > > On Fri, Feb 28, 2014 at 06:19:58PM +0100, Erik Faye-Lund wrote:\n> > >> On Fri, Feb 28, 2014 at 6:00 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> > >> ...\n> > >> > In fact that would be maybe preferred, for maintainers to enable alloca\n> > >> > with knowledge and testing, as one person can't have them all at hand.\n> > >> \n> > >> Yeah, you're probably right.\n> > >\n> > > Erik, the patch has been merged into pu today. Would you please\n> > > follow-up with tested MINGW change?\n> > \n> > Sooo.... I lost track but this discussion seems to have petered out\n> > around here.  I think the copy we have had for a while on 'pu' is\n> > basically sound, and can easily built on by platform folks by adding\n> > or removing the -DHAVE_ALLOCA_H from the Makefile.\n> \n> Yes, that is all correct - that version works and we can improve it in\n> the future with platform-specific follow-up patches, if needed.\n\nJunio, thanks for merging this and other diff-tree patches to next.  It\nso happened that I'm wrestling with MSysGit today, so please also find\nalloca-for-mingw patch attached below.\n\nThanks,\nKirill\n\n---- 8< ----\nSubject: [PATCH] mingw: activate alloca\n\nBoth MSVC and MINGW have alloca(3) definitions in malloc.h, so by moving\nwin32-compat alloca.h from compat/vcbuild/include/ to compat/win32/ ,\nwhich is included by both MSVC and MINGW CFLAGS, we can make alloca()\nwork on both those Windows environments.\n\nIn MINGW, malloc.h has explicit check for GNUC and if it is so, defines\nalloca to __builtin_alloca, so it looks like we don't need to add any\ncode to here-shipped alloca.h to get optimum performance.\n\nCompile-tested on Windows in MSysGit.\n\nSigned-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n---\n compat/{vcbuild/include => win32}/alloca.h | 0\n config.mak.uname                           | 1 +\n 2 files changed, 1 insertion(+)\n rename compat/{vcbuild/include => win32}/alloca.h (100%)\n\ndiff --git a/compat/vcbuild/include/alloca.h b/compat/win32/alloca.h\nsimilarity index 100%\nrename from compat/vcbuild/include/alloca.h\nrename to compat/win32/alloca.h\ndiff --git a/config.mak.uname b/config.mak.uname\nindex 17ef893..67bc054 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -480,6 +480,7 @@ ifeq ($(uname_S),NONSTOP_KERNEL)\n endif\n ifneq (,$(findstring MINGW,$(uname_S)))\n \tpathsep = ;\n+\tHAVE_ALLOCA_H = YesPlease\n \tNO_PREAD = YesPlease\n \tNEEDS_CRYPTO_WITH_SSL = YesPlease\n \tNO_LIBGEN_H = YesPlease\n-- \n1.9.0.msysgit.0.31.g74d1b9a.dirty\n"},{"id":"238580","messageId":"CABPQNSbiyzGLC=Y1kiFPOc4WLWnu=ZpPd6aRwSBzBjZYEih5Tw@mail.gmail.com","threadId":"35947","inReplyTo":"20140409124827.GA24672@tugrik.mns.mnsspb.ru","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@gmail.com","sentAt":"2014-04-09T13:01:07Z","receivedAt":"2014-04-09T13:01:07Z","isPatch":true,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Wed, Apr 9, 2014 at 2:48 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n> On Thu, Mar 27, 2014 at 06:22:50PM +0400, Kirill Smelkov wrote:\n>> On Mon, Mar 24, 2014 at 02:47:24PM -0700, Junio C Hamano wrote:\n>> > Kirill Smelkov <kirr@mns.spb.ru> writes:\n>> >\n>> > > On Fri, Feb 28, 2014 at 06:19:58PM +0100, Erik Faye-Lund wrote:\n>> > >> On Fri, Feb 28, 2014 at 6:00 PM, Kirill Smelkov <kirr@mns.spb.ru> wrote:\n>> > >> ...\n>> > >> > In fact that would be maybe preferred, for maintainers to enable alloca\n>> > >> > with knowledge and testing, as one person can't have them all at hand.\n>> > >>\n>> > >> Yeah, you're probably right.\n>> > >\n>> > > Erik, the patch has been merged into pu today. Would you please\n>> > > follow-up with tested MINGW change?\n>> >\n>> > Sooo.... I lost track but this discussion seems to have petered out\n>> > around here.  I think the copy we have had for a while on 'pu' is\n>> > basically sound, and can easily built on by platform folks by adding\n>> > or removing the -DHAVE_ALLOCA_H from the Makefile.\n>>\n>> Yes, that is all correct - that version works and we can improve it in\n>> the future with platform-specific follow-up patches, if needed.\n>\n> Junio, thanks for merging this and other diff-tree patches to next.  It\n> so happened that I'm wrestling with MSysGit today, so please also find\n> alloca-for-mingw patch attached below.\n>\n> Thanks,\n> Kirill\n>\n> ---- 8< ----\n> Subject: [PATCH] mingw: activate alloca\n>\n> Both MSVC and MINGW have alloca(3) definitions in malloc.h, so by moving\n> win32-compat alloca.h from compat/vcbuild/include/ to compat/win32/ ,\n> which is included by both MSVC and MINGW CFLAGS, we can make alloca()\n> work on both those Windows environments.\n>\n> In MINGW, malloc.h has explicit check for GNUC and if it is so, defines\n> alloca to __builtin_alloca, so it looks like we don't need to add any\n> code to here-shipped alloca.h to get optimum performance.\n>\n> Compile-tested on Windows in MSysGit.\n>\n> Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n\nLooks good to me!\n"},{"id":"238625","messageId":"xmqqbnw9xfbb.fsf@gitster.dls.corp.google.com","threadId":"35947","inReplyTo":"CABPQNSbiyzGLC=Y1kiFPOc4WLWnu=ZpPd6aRwSBzBjZYEih5Tw@mail.gmail.com","subject":"Re: [PATCH 17/19] Portable alloca for Git","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-04-10T17:30:32Z","receivedAt":"2014-04-10T17:30:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Erik Faye-Lund <kusmabite@gmail.com> writes:\n\n>> Subject: [PATCH] mingw: activate alloca\n>>\n>> Both MSVC and MINGW have alloca(3) definitions in malloc.h, so by moving\n>> win32-compat alloca.h from compat/vcbuild/include/ to compat/win32/ ,\n>> which is included by both MSVC and MINGW CFLAGS, we can make alloca()\n>> work on both those Windows environments.\n>>\n>> In MINGW, malloc.h has explicit check for GNUC and if it is so, defines\n>> alloca to __builtin_alloca, so it looks like we don't need to add any\n>> code to here-shipped alloca.h to get optimum performance.\n>>\n>> Compile-tested on Windows in MSysGit.\n>>\n>> Signed-off-by: Kirill Smelkov <kirr@mns.spb.ru>\n>\n> Looks good to me!\n\nThanks; queued and pushed out on 'next'.\n"}]}