{"thread":{"id":"48913","subject":"[PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","startedAt":"2018-07-18T20:45:20Z","lastAt":"2018-08-25T13:02:35Z","messageCount":121,"participants":["Ben Peart","Stefan Beller","Jeff King","Junio C Hamano","Duy Nguyen","Nguyễn Thái Ngọc Duy","Elijah Newren","Thomas Adam","Jeff Hostetler","Martin Ågren"],"isPatch":true,"patchVersion":1,"patchTotal":3},"messages":[{"id":"352981","messageId":"20180718204458.20936-1-benpeart@microsoft.com","threadId":"48913","inReplyTo":null,"subject":"[PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Ben Peart","fromEmail":"ben.peart@microsoft.com","sentAt":"2018-07-18T20:45:14Z","receivedAt":"2018-07-18T20:45:20Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"When working directories get big, checkout times start to suffer.  Even with\nGVFS virtualization (which limits git to only having to update those files\nthat have been changed locally) we�re seeing P50 times for checkout of 31\nseconds and the P80 time is 43 seconds.\n\nHere is a checkout command with tracing turned on to demonstrate where the\ntime is spent.  Note, this is somewhat of a �best case� as I�m simply\nchecking out the current commit:\n\nbenpeart@gvfs-perf MINGW64 /f/os/src (official/rs_es_debug_dev)\n$ /usr/src/git/git.exe checkout\n12:31:50.419016 read-cache.c:2006       performance: 1.180966800 s: read cache .git/index\n12:31:51.184636 name-hash.c:605         performance: 0.664575200 s: initialize name hash\n12:31:51.200280 preload-index.c:111     performance: 0.019811600 s: preload index\n12:31:51.294012 read-cache.c:1543       performance: 0.094515600 s: refresh index\n12:32:29.731344 unpack-trees.c:1358     performance: 33.889840200 s: traverse_trees\n12:32:37.512555 read-cache.c:2541       performance: 1.564438300 s: write index, changed mask = 28\n12:32:44.918730 unpack-trees.c:1358     performance: 7.243155600 s: traverse_trees\n12:32:44.965611 diff-lib.c:527          performance: 7.374729200 s: diff-index\nWaiting for GVFS to parse index and update placeholder files...Succeeded\n12:32:46.824986 trace.c:420             performance: 57.715656000 s: git command: 'C:\\git-sdk-64\\usr\\src\\git\\git.exe' checkout\n\nClearly, most of the time (41 seconds) is spent in the traverse_trees() code\nso the question is, how can we significantly speed up that portion of the\ncommand?\n\nI investigated a few options with limited success:\n\nODB cache\n=========\nSince traverse_trees() hits the ODB for each tree object (of which there are\nover 500K in this repo) I wrote and tested having an in-memory ODB cache\nthat cached all tree objects.  This resulted in a > 50% hit ratio (largely\ndue to the fact we traverse the tree twice during checkout) but resulted in\nonly a minimal savings (1.3 seconds).\n\nTree Graph File\n===============\nI also considered storing the commit tree in an alternate structure that is\nfaster to load/parse (ala the Commit graph) but the cache results along with\nthe negligible impact of running checkout back to back (thus ensuring the\nobjects were cached in my file system cache) made me believe this would not\nresult in much savings. MIDX has already helped out here given we end up\nwith a lot of pack files of commits and trees.\n\nSparse tree traversal\n=====================\nWe�ve sped up other parts of git by taking advantage of the existing\nsparse-checkout/excludes logic to limit what files git has to consider to\nthose that have been modified by the user locally.  I haven�t been able to\nthink of a way to take advantage of that with unpack-trees() as when you are\nmerging n commits, a change/conflict can occur in any tree object so they\nmust all be traversed.  If I�m missing something here and there _is_ a way\nto entirely skip large parts of the tree, please let me know!  Please note\nthat we�re already limiting the files that git needs to update in the\nworking directory via sparse-checkout/excludes but the other/merge logic\nstill executes for the entire tree whether there are files to update or not.\n\nMulti-threading unpack_trees()\n==============================\nThe current model of unpack_trees() is that a single thread recursively\ntraverses each tree object as it comes across it.  One thought I had was to\nmulti-thread the traversal so that each tree object could be processed in\nparallel.  To test this idea out, I wrote an unbounded\nMulti-Product-Multi-Consumer queue and then wrote a\ntraverse_trees_parallel() function that would add any new tree objects into\nthe queue where they can be processed by a pool of worker threads.  Each\nthread will wake up when there is work in the queue, remove a tree object,\nprocess it adding any additional tree objects it finds.\n\nMulti-threading anything in git is fraught with challenges as much of the\ncode base is not thread safe.  To make progress, I wrapped mutexes around\ncode paths that were not thread safe.  The end result is that I won�t\ninitially get much parallelization (due to mutexes around all the expensive\nwork) but at least I can test out the idea and resolve any other issues with\nswitching from a serial to a parallel implementation.  If this works out, I\ncan update more of the code paths to be thread safe and/or move to more fine\ngrained mutexes around those paths that are difficult to make thread safe.\n\nFinal thoughts\n==============\n\nThe attached set of patches don�t work!  For some commands they succeed but\nI�m including them only to make it explicit what I�m currently investigating.\nI�d be very interested in design feedback but formatting/spelling/white\nspace errors are less useful at this early stage in the investigation.\n\nWhen I brought up this idea with some other git contributors they mentioned\nthat multi threading unpack_trees() had been discussed a few years ago on\nthe list but that the idea was discarded.  They couldn�t remember exactly\nwhy it was discarded and none of us have been able to find the email threads\nfrom that earlier discussion. As a result, I decided to write up this RFC\nand see if the greater git community has ideas, suggestions, or more\nbackground/history on whether this is a reasonable path to pursue or if\nthere are other/better ideas on how to speed up checkout especially on large\nrepos.\n\n\nBase Ref: master\nWeb-Diff: https://github.com/benpeart/git/commit/a022a91ceb\nCheckout: git fetch https://github.com/benpeart/git unpacktrees-v1 && git checkout a022a91ceb\n\nBen Peart (3):\n  add unbounded Multi-Producer-Multi-Consumer queue\n  add performance tracing around traverse_trees() in unpack_trees()\n  Add initial parallel version of unpack_trees()\n\n Makefile       |   1 +\n cache.h        |   1 +\n config.c       |   5 +\n environment.c  |   1 +\n mpmcqueue.c    |  49 ++++++++\n mpmcqueue.h    |  80 +++++++++++++\n unpack-trees.c | 314 ++++++++++++++++++++++++++++++++++++++++++++++++-\n unpack-trees.h |  30 +++++\n 8 files changed, 480 insertions(+), 1 deletion(-)\n create mode 100644 mpmcqueue.c\n create mode 100644 mpmcqueue.h\n\n\nbase-commit: e3331758f12da22f4103eec7efe1b5304a9be5e9\n-- \n2.17.0.gvfs.1.123.g449c066\n\n\n"},{"id":"352982","messageId":"20180718204458.20936-2-benpeart@microsoft.com","threadId":"48913","inReplyTo":"20180718204458.20936-1-benpeart@microsoft.com","subject":"[PATCH v1 1/3] add unbounded Multi-Producer-Multi-Consumer queue","fromName":"Ben Peart","fromEmail":"ben.peart@microsoft.com","sentAt":"2018-07-18T20:45:15Z","receivedAt":"2018-07-18T20:45:21Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"Signed-off-by: Ben Peart <benpeart@microsoft.com>\n---\n Makefile    |  1 +\n mpmcqueue.c | 49 ++++++++++++++++++++++++++++++++\n mpmcqueue.h | 80 +++++++++++++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 130 insertions(+)\n create mode 100644 mpmcqueue.c\n create mode 100644 mpmcqueue.h\n\ndiff --git a/Makefile b/Makefile\nindex 0cb6590f24..fdaabf0252 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -890,6 +890,7 @@ LIB_OBJS += merge.o\n LIB_OBJS += merge-blobs.o\n LIB_OBJS += merge-recursive.o\n LIB_OBJS += mergesort.o\n+LIB_OBJS += mpmcqueue.o\n LIB_OBJS += name-hash.o\n LIB_OBJS += notes.o\n LIB_OBJS += notes-cache.o\ndiff --git a/mpmcqueue.c b/mpmcqueue.c\nnew file mode 100644\nindex 0000000000..22411af1b0\n--- /dev/null\n+++ b/mpmcqueue.c\n@@ -0,0 +1,49 @@\n+#include \"mpmcqueue.h\"\n+\n+void mpmcq_init(struct mpmcq *queue)\n+{\n+\tqueue->head = NULL;\n+\tqueue->cancel = 0;\n+\tpthread_mutex_init(&queue->mutex, NULL);\n+\tpthread_cond_init(&queue->condition, NULL);\n+}\n+\n+void mpmcq_destroy(struct mpmcq *queue)\n+{\n+\tpthread_mutex_destroy(&queue->mutex);\n+\tpthread_cond_destroy(&queue->condition);\n+}\n+\n+void mpmcq_push(struct mpmcq *queue, struct mpmcq_entry *entry)\n+{\n+\tpthread_mutex_lock(&queue->mutex);\n+\tentry->next = queue->head;\n+\tqueue->head = entry;\n+\tpthread_cond_signal(&queue->condition);\n+\tpthread_mutex_unlock(&queue->mutex);\n+}\n+\n+struct mpmcq_entry *mpmcq_pop(struct mpmcq *queue)\n+{\n+\tstruct mpmcq_entry *entry = NULL;\n+\n+\tpthread_mutex_lock(&queue->mutex);\n+\twhile (!queue->head && !queue->cancel)\n+\t\tpthread_cond_wait(&queue->condition, &queue->mutex);\n+\tif (!queue->cancel) {\n+\t\tentry = queue->head;\n+\t\tqueue->head = entry->next;\n+\t}\n+\tpthread_mutex_unlock(&queue->mutex);\n+\treturn entry;\n+}\n+\n+void mpmcq_cancel(struct mpmcq *queue)\n+{\n+\tstruct mpmcq_entry *entry;\n+\n+\tpthread_mutex_lock(&queue->mutex);\n+\tqueue->cancel = 1;\n+\tpthread_cond_broadcast(&queue->condition);\n+\tpthread_mutex_unlock(&queue->mutex);\n+}\ndiff --git a/mpmcqueue.h b/mpmcqueue.h\nnew file mode 100644\nindex 0000000000..7421e06aad\n--- /dev/null\n+++ b/mpmcqueue.h\n@@ -0,0 +1,80 @@\n+#ifndef MPMCQUEUE_H\n+#define MPMCQUEUE_H\n+\n+#include \"git-compat-util.h\"\n+#include <pthread.h>\n+\n+/*\n+ * Generic implementation of an unbounded Multi-Producer-Multi-Consumer\n+ * queue.\n+ */\n+\n+/*\n+ * struct mpmcq_entry is an opaque structure representing an entry in the\n+ * queue.\n+ */\n+struct mpmcq_entry {\n+\tstruct mpmcq_entry *next;\n+};\n+\n+/*\n+ * struct mpmcq is the concurrent queue structure. Members should not be\n+ * modified directly.\n+ */\n+struct mpmcq {\n+\tstruct mpmcq_entry *head;\n+\tpthread_mutex_t mutex;\n+\tpthread_cond_t condition;\n+\tint cancel;\n+};\n+\n+/*\n+ * Initializes a mpmcq_entry structure.\n+ *\n+ * `entry` points to the entry to initialize.\n+ *\n+ * The mpmcq_entry structure does not hold references to external resources,\n+ * and it is safe to just discard it once you are done with it (i.e. if\n+ * your structure was allocated with xmalloc(), you can just free() it,\n+ * and if it is on stack, you can just let it go out of scope).\n+ */\n+static inline void mpmcq_entry_init(struct mpmcq_entry *entry)\n+{\n+\tentry->next = NULL;\n+}\n+\n+/*\n+ * Initializes a mpmcq structure.\n+ */\n+extern void mpmcq_init(struct mpmcq *queue);\n+\n+/*\n+ * Destroys a mpmcq structure.\n+ */\n+extern void mpmcq_destroy(struct mpmcq *queue);\n+\n+/*\n+ * Pushes an entry on to the queue.\n+ *\n+ * `queue` is the mpmcq structure.\n+ * `entry` is the entry to push.\n+ */\n+extern void mpmcq_push(struct mpmcq *queue, struct mpmcq_entry *entry);\n+\n+/*\n+ * Pops an entry off the queue.\n+ *\n+ * `queue` is the mpmcq structure.\n+ *\n+ * Returns mpmcq_entry on success, NULL on cancel;\n+ */\n+extern struct mpmcq_entry *mpmcq_pop(struct mpmcq *queue);\n+\n+/*\n+ * Cancels any pending pop requests.\n+ *\n+ * `queue` is the mpmcq structure.\n+ */\n+extern void mpmcq_cancel(struct mpmcq *queue);\n+\n+#endif\n-- \n2.17.0.gvfs.1.123.g449c066\n\n"},{"id":"352983","messageId":"20180718204458.20936-3-benpeart@microsoft.com","threadId":"48913","inReplyTo":"20180718204458.20936-1-benpeart@microsoft.com","subject":"[PATCH v1 2/3] add performance tracing around traverse_trees() in unpack_trees()","fromName":"Ben Peart","fromEmail":"ben.peart@microsoft.com","sentAt":"2018-07-18T20:45:16Z","receivedAt":"2018-07-18T20:45:22Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"Signed-off-by: Ben Peart <benpeart@microsoft.com>\n---\n unpack-trees.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 3a85a02a77..1f58efc6bb 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -1326,6 +1326,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \tif (len) {\n \t\tconst char *prefix = o->prefix ? o->prefix : \"\";\n \t\tstruct traverse_info info;\n+\t\tuint64_t start;\n \n \t\tsetup_traverse_info(&info, prefix);\n \t\tinfo.fn = unpack_callback;\n@@ -1350,8 +1351,10 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t\t}\n \t\t}\n \n+\t\tstart = getnanotime();\n \t\tif (traverse_trees(len, t, &info) < 0)\n \t\t\tgoto return_failed;\n+\t\ttrace_performance_since(start, \"traverse_trees\");\n \t}\n \n \t/* Any left-over entries in the index? */\n-- \n2.17.0.gvfs.1.123.g449c066\n\n"},{"id":"352985","messageId":"20180718204458.20936-4-benpeart@microsoft.com","threadId":"48913","inReplyTo":"20180718204458.20936-1-benpeart@microsoft.com","subject":"[PATCH v1 3/3] Add initial parallel version of unpack_trees()","fromName":"Ben Peart","fromEmail":"ben.peart@microsoft.com","sentAt":"2018-07-18T20:45:17Z","receivedAt":"2018-07-18T20:45:25Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"Signed-off-by: Ben Peart <benpeart@microsoft.com>\n---\n cache.h        |   1 +\n config.c       |   5 +\n environment.c  |   1 +\n unpack-trees.c | 313 ++++++++++++++++++++++++++++++++++++++++++++++++-\n unpack-trees.h |  30 +++++\n 5 files changed, 348 insertions(+), 2 deletions(-)\n\ndiff --git a/cache.h b/cache.h\nindex d49092d94d..4bfa35c497 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -815,6 +815,7 @@ extern int fsync_object_files;\n extern int core_preload_index;\n extern int core_commit_graph;\n extern int core_apply_sparse_checkout;\n+extern int core_parallel_unpack_trees;\n extern int precomposed_unicode;\n extern int protect_hfs;\n extern int protect_ntfs;\ndiff --git a/config.c b/config.c\nindex f4a208a166..34d5506588 100644\n--- a/config.c\n+++ b/config.c\n@@ -1346,6 +1346,11 @@ static int git_default_core_config(const char *var, const char *value)\n \t\t\t\t\t var, value);\n \t}\n \n+\tif (!strcmp(var, \"core.parallelunpacktrees\")) {\n+\t\tcore_parallel_unpack_trees = git_config_bool(var, value);\n+\t\treturn 0;\n+\t}\n+\n \t/* Add other config variables here and to Documentation/config.txt. */\n \treturn 0;\n }\ndiff --git a/environment.c b/environment.c\nindex 2a6de2330b..1eb0a05074 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -68,6 +68,7 @@ char *notes_ref_name;\n int grafts_replace_parents = 1;\n int core_commit_graph;\n int core_apply_sparse_checkout;\n+int core_parallel_unpack_trees;\n int merge_log_config = -1;\n int precomposed_unicode = -1; /* see probe_utf8_pathname_composition() */\n unsigned long pack_size_limit_cfg;\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 1f58efc6bb..2333626efd 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -17,6 +17,7 @@\n #include \"submodule-config.h\"\n #include \"fsmonitor.h\"\n #include \"fetch-object.h\"\n+#include \"thread-utils.h\"\n \n /*\n  * Error messages expected by scripts out of plumbing commands such as\n@@ -641,6 +642,98 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n \treturn name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n }\n \n+#ifndef NO_PTHREADS\n+\n+struct traverse_info_parallel {\n+\tstruct mpmcq_entry entry;\n+\tstruct tree_desc t[MAX_UNPACK_TREES];\n+\tvoid *buf[MAX_UNPACK_TREES];\n+\tstruct traverse_info info;\n+\tint n;\n+\tint nr_buf;\n+\tint ret;\n+};\n+\n+static int traverse_trees_parallel(int n, unsigned long dirmask,\n+\t\t\t\t   unsigned long df_conflicts,\n+\t\t\t\t   struct name_entry *names,\n+\t\t\t\t   struct traverse_info *info)\n+{\n+\tint i;\n+\tstruct name_entry *p;\n+\tstruct unpack_trees_options *o = info->data;\n+\tstruct traverse_info_parallel *newinfo;\n+\n+\tp = names;\n+\twhile (!p->mode)\n+\t\tp++;\n+\n+\tnewinfo = xmalloc(sizeof(struct traverse_info_parallel));\n+\tmpmcq_entry_init(&newinfo->entry);\n+\tnewinfo->info = *info;\n+\tnewinfo->info.prev = info;\n+\tnewinfo->info.pathspec = info->pathspec;\n+\tnewinfo->info.name = *p;\n+\tnewinfo->info.pathlen += tree_entry_len(p) + 1;\n+\tnewinfo->info.df_conflicts |= df_conflicts;\n+\tnewinfo->nr_buf = 0;\n+\tnewinfo->n = n;\n+\n+\t/*\n+\t * Fetch the tree from the ODB for each peer directory in the\n+\t * n commits.\n+\t *\n+\t * For 2- and 3-way traversals, we try to avoid hitting the\n+\t * ODB twice for the same OID.  This should yield a nice speed\n+\t * up in checkouts and merges when the commits are similar.\n+\t *\n+\t * We don't bother doing the full O(n^2) search for larger n,\n+\t * because wider traversals don't happen that often and we\n+\t * avoid the search setup.\n+\t *\n+\t * When 2 peer OIDs are the same, we just copy the tree\n+\t * descriptor data.  This implicitly borrows the buffer\n+\t * data from the earlier cell.\n+\t */\n+\tfor (i = 0; i < n; i++, dirmask >>= 1) {\n+\t\tif (i > 0 && are_same_oid(&names[i], &names[i - 1]))\n+\t\t\tnewinfo->t[i] = newinfo->t[i - 1];\n+\t\telse if (i > 1 && are_same_oid(&names[i], &names[i - 2]))\n+\t\t\tnewinfo->t[i] = newinfo->t[i - 2];\n+\t\telse {\n+\t\t\tconst struct object_id *oid = NULL;\n+\t\t\tif (dirmask & 1)\n+\t\t\t\toid = names[i].oid;\n+\n+\t\t\t/*\n+\t\t\t * fill_tree_descriptor() will load the tree from the\n+\t\t\t * ODB. Accessing the ODB is not thread safe so\n+\t\t\t * serialize access using the odb_mutex.\n+\t\t\t */\n+\t\t\tpthread_mutex_lock(&o->odb_mutex);\n+\t\t\tnewinfo->buf[newinfo->nr_buf++] =\n+\t\t\t\tfill_tree_descriptor(newinfo->t + i, oid);\n+\t\t\tpthread_mutex_unlock(&o->odb_mutex);\n+\t\t}\n+\t}\n+\n+\t/*\n+\t * We can't play games with the cache bottom as we are processing\n+\t * the tree objects in parallel.\n+\t * newinfo->bottom = switch_cache_bottom(&newinfo->info);\n+\t */\n+\n+\t/* All I really need here is fetch_and_add() */\n+\tpthread_mutex_lock(&o->work_mutex);\n+\to->remaining_work++;\n+\tpthread_mutex_unlock(&o->work_mutex);\n+\tmpmcq_push(&o->queue, &newinfo->entry);\n+\n+\treturn 0;\n+}\n+\n+#endif\n+\n static int traverse_trees_recursive(int n, unsigned long dirmask,\n \t\t\t\t    unsigned long df_conflicts,\n \t\t\t\t    struct name_entry *names,\n@@ -995,6 +1088,108 @@ static void debug_unpack_callback(int n,\n \t\tdebug_name_entry(i, names + i);\n }\n \n+static int unpack_callback_parallel(int n, unsigned long mask,\n+\t\t\t\t    unsigned long dirmask,\n+\t\t\t\t    struct name_entry *names,\n+\t\t\t\t    struct traverse_info *info)\n+{\n+\tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = {\n+\t\tNULL,\n+\t};\n+\tstruct unpack_trees_options *o = info->data;\n+\tconst struct name_entry *p = names;\n+\n+\t/* Find first entry with a real name (we could use \"mask\" too) */\n+\twhile (!p->mode)\n+\t\tp++;\n+\n+\tif (o->debug_unpack)\n+\t\tdebug_unpack_callback(n, mask, dirmask, names, info);\n+\n+\t/* Are we supposed to look at the index too? */\n+\tif (o->merge) {\n+\t\twhile (1) {\n+\t\t\tint cmp;\n+\t\t\tstruct cache_entry *ce;\n+\n+\t\t\tif (o->diff_index_cached)\n+\t\t\t\tce = next_cache_entry(o);\n+\t\t\telse\n+\t\t\t\tce = find_cache_entry(info, p);\n+\n+\t\t\tif (!ce)\n+\t\t\t\tbreak;\n+\t\t\tcmp = compare_entry(ce, info, p);\n+\t\t\tif (cmp < 0) {\n+\t\t\t\tint ret;\n+\n+\t\t\t\tpthread_mutex_lock(&o->unpack_index_entry_mutex);\n+\t\t\t\tret = unpack_index_entry(ce, o);\n+\t\t\t\tpthread_mutex_unlock(&o->unpack_index_entry_mutex);\n+\t\t\t\tif (ret < 0)\n+\t\t\t\t\treturn unpack_failed(o, NULL);\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!cmp) {\n+\t\t\t\tif (ce_stage(ce)) {\n+\t\t\t\t\t/*\n+\t\t\t\t\t * If we skip unmerged index\n+\t\t\t\t\t * entries, we'll skip this\n+\t\t\t\t\t * entry *and* the tree\n+\t\t\t\t\t * entries associated with it!\n+\t\t\t\t\t */\n+\t\t\t\t\tif (o->skip_unmerged) {\n+\t\t\t\t\t\tadd_same_unmerged(ce, o);\n+\t\t\t\t\t\treturn mask;\n+\t\t\t\t\t}\n+\t\t\t\t}\n+\t\t\t\tsrc[0] = ce;\n+\t\t\t}\n+\t\t\tbreak;\n+\t\t}\n+\t}\n+\n+\tpthread_mutex_lock(&o->unpack_nondirectories_mutex);\n+\tint ret = unpack_nondirectories(n, mask, dirmask, src, names, info);\n+\tpthread_mutex_unlock(&o->unpack_nondirectories_mutex);\n+\tif (ret < 0)\n+\t\treturn -1;\n+\n+\tif (o->merge && src[0]) {\n+\t\tif (ce_stage(src[0]))\n+\t\t\tmark_ce_used_same_name(src[0], o);\n+\t\telse\n+\t\t\tmark_ce_used(src[0], o);\n+\t}\n+\n+\t/* Now handle any directories.. */\n+\tif (dirmask) {\n+\t\t/* special case: \"diff-index --cached\" looking at a tree */\n+\t\tif (o->diff_index_cached && n == 1 && dirmask == 1 &&\n+\t\t    S_ISDIR(names->mode)) {\n+\t\t\tint matches;\n+\t\t\tmatches = cache_tree_matches_traversal(\n+\t\t\t\to->src_index->cache_tree, names, info);\n+\t\t\t/*\n+\t\t\t * Everything under the name matches; skip the\n+\t\t\t * entire hierarchy.  diff_index_cached codepath\n+\t\t\t * special cases D/F conflicts in such a way that\n+\t\t\t * it does not do any look-ahead, so this is safe.\n+\t\t\t */\n+\t\t\tif (matches) {\n+\t\t\t\to->cache_bottom += matches;\n+\t\t\t\treturn mask;\n+\t\t\t}\n+\t\t}\n+\n+\t\tif (traverse_trees_parallel(n, dirmask, mask & ~dirmask, names, info) < 0)\n+\t\t\treturn -1;\n+\t\treturn mask;\n+\t}\n+\n+\treturn mask;\n+}\n+\n static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n@@ -1263,6 +1458,116 @@ static void mark_new_skip_worktree(struct exclude_list *el,\n static int verify_absent(const struct cache_entry *,\n \t\t\t enum unpack_trees_error_types,\n \t\t\t struct unpack_trees_options *);\n+\n+#ifndef NO_PTHREADS\n+static void *traverse_trees_parallel_thread_proc(void *_data)\n+{\n+\tstruct unpack_trees_options *o = _data;\n+\tstruct traverse_info_parallel *info;\n+\tint i;\n+\n+\twhile (1) {\n+\t\tinfo = (struct traverse_info_parallel *)mpmcq_pop(&o->queue);\n+\t\tif (!info)\n+\t\t\tbreak;\n+\n+\t\tinfo->ret = traverse_trees(info->n, info->t, &info->info);\n+\t\t/*\n+\t\t * We can't play games with the cache bottom as we are processing\n+\t\t * the tree objects in parallel.\n+\t\t * restore_cache_bottom(&info->info, info->bottom);\n+\t\t */\n+\n+\t\tfor (i = 0; i < info->nr_buf; i++)\n+\t\t\tfree(info->buf[i]);\n+\t\t/*\n+\t\t * TODO: Can't free \"info\" when thread is done because it can be used\n+\t\t * as ->prev link in child info objects.  Ref count?  Free all at end?\n+\t\tfree(info);\n+\t\t */\n+\n+\t\t/* All I really need here is fetch_and_add() */\n+\t\tpthread_mutex_lock(&o->work_mutex);\n+\t\to->remaining_work--;\n+\t\tif (o->remaining_work == 0)\n+\t\t\tmpmcq_cancel(&o->queue);\n+\t\tpthread_mutex_unlock(&o->work_mutex);\n+\t}\n+\n+\treturn NULL;\n+}\n+\n+static void init_parallel_traverse(struct unpack_trees_options *o,\n+\t\t\t\t   struct traverse_info *info)\n+{\n+\t/*\n+\t * TODO: Add logic to bypass parallel path when not needed.\n+\t *\t\t\t- not enough CPU cores to help\n+\t *\t\t\t- 'git status' is always fast - how to detect?\n+\t *\t\t\t- small trees (may be able to use index size as proxy, small index likely means small commit tree)\n+\t */\n+\tif (core_parallel_unpack_trees) {\n+\t\tint t;\n+\n+\t\tmpmcq_init(&o->queue);\n+\t\to->remaining_work = 0;\n+\t\tpthread_mutex_init(&o->unpack_nondirectories_mutex, NULL);\n+\t\tpthread_mutex_init(&o->unpack_index_entry_mutex, NULL);\n+\t\tpthread_mutex_init(&o->odb_mutex, NULL);\n+\t\tpthread_mutex_init(&o->work_mutex, NULL);\n+\t\to->nr_threads = online_cpus();\n+\t\to->pthreads = xcalloc(o->nr_threads, sizeof(pthread_t));\n+\t\tinfo->fn = unpack_callback_parallel;\n+\n+\t\tfor (t = 0; t < o->nr_threads; t++) {\n+\t\t\tif (pthread_create(&o->pthreads[t], NULL,\n+\t\t\t\t\t   traverse_trees_parallel_thread_proc,\n+\t\t\t\t\t   o))\n+\t\t\t\tdie(\"unable to create traverse_trees_parallel_thread\");\n+\t\t}\n+\t}\n+}\n+\n+static void wait_parallel_traverse(struct unpack_trees_options *o)\n+{\n+\t/*\n+\t * The first tree (root directory) is processed on the main thread.\n+\t * This function is called after it has completed.  If there is no\n+\t * remaining work, we know we are finished.\n+\t */\n+\tif (core_parallel_unpack_trees) {\n+\t\tint t;\n+\n+\t\tpthread_mutex_lock(&o->work_mutex);\n+\t\tif (o->remaining_work == 0)\n+\t\t\tmpmcq_cancel(&o->queue);\n+\t\tpthread_mutex_unlock(&o->work_mutex);\n+\n+\t\tfor (t = 0; t < o->nr_threads; t++) {\n+\t\t\tif (pthread_join(o->pthreads[t], NULL))\n+\t\t\t\tdie(\"unable to join traverse_trees_parallel_thread\");\n+\t\t}\n+\n+\t\tfree(o->pthreads);\n+\t\tpthread_mutex_destroy(&o->work_mutex);\n+\t\tpthread_mutex_destroy(&o->odb_mutex);\n+\t\tpthread_mutex_destroy(&o->unpack_index_entry_mutex);\n+\t\tpthread_mutex_destroy(&o->unpack_nondirectories_mutex);\n+\t\tmpmcq_destroy(&o->queue);\n+\t}\n+}\n+#else\n+static void init_parallel_traverse(struct unpack_trees_options *o)\n+{\n+\treturn;\n+}\n+\n+static void wait_parallel_traverse(struct unpack_trees_options *o)\n+{\n+\treturn;\n+}\n+#endif\n+\n /*\n  * N-way merge \"len\" trees.  Returns 0 on success, -1 on failure to manipulate the\n  * resulting index, -2 on failure to reflect the changes to the work tree.\n@@ -1327,6 +1632,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\tconst char *prefix = o->prefix ? o->prefix : \"\";\n \t\tstruct traverse_info info;\n \t\tuint64_t start;\n+\t\tint ret;\n \n \t\tsetup_traverse_info(&info, prefix);\n \t\tinfo.fn = unpack_callback;\n@@ -1352,9 +1658,12 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t}\n \n \t\tstart = getnanotime();\n-\t\tif (traverse_trees(len, t, &info) < 0)\n-\t\t\tgoto return_failed;\n+\t\tinit_parallel_traverse(o, &info);\n+\t\tret = traverse_trees(len, t, &info);\n+\t\twait_parallel_traverse(o);\n \t\ttrace_performance_since(start, \"traverse_trees\");\n+\t\tif (ret < 0)\n+\t\t\tgoto return_failed;\n \t}\n \n \t/* Any left-over entries in the index? */\ndiff --git a/unpack-trees.h b/unpack-trees.h\nindex c2b434c606..b7140099fa 100644\n--- a/unpack-trees.h\n+++ b/unpack-trees.h\n@@ -3,6 +3,11 @@\n \n #include \"tree-walk.h\"\n #include \"argv-array.h\"\n+#ifndef NO_PTHREADS\n+#include \"git-compat-util.h\"\n+#include <pthread.h>\n+#include \"mpmcqueue.h\"\n+#endif\n \n #define MAX_UNPACK_TREES 8\n \n@@ -80,6 +85,31 @@ struct unpack_trees_options {\n \tstruct index_state result;\n \n \tstruct exclude_list *el; /* for internal use */\n+#ifndef NO_PTHREADS\n+\t/*\n+\t * Speed up the tree traversal by adding all discovered tree objects\n+\t * into a queue and have a pool of worker threads process them in\n+\t * parallel.  Since there is no upper bound on the size of a tree and\n+\t * each worker thread will be adding discovered tree objects to the\n+\t * queue, we need an unbounded multi-producer-multi-consumer queue.\n+\t */\n+\tstruct mpmcq queue;\n+\n+\tint nr_threads;\n+\tpthread_t *pthreads;\n+\n+\t/* need a mutex as we don't have fetch_and_add() */\n+\tint remaining_work;\n+\tpthread_mutex_t work_mutex;\n+\n+\t/* The ODB is not thread safe so we must serialize access to it */\n+\tpthread_mutex_t odb_mutex;\n+\n+\t/* various functions that are not thread safe and must be serialized for now */\n+\tpthread_mutex_t unpack_index_entry_mutex;\n+\tpthread_mutex_t unpack_nondirectories_mutex;\n+\n+#endif\n };\n \n extern int unpack_trees(unsigned n, struct tree_desc *t,\n-- \n2.17.0.gvfs.1.123.g449c066\n\n"},{"id":"352987","messageId":"CAGZ79ka=GkwBLeC1SK33dCH6X91kEeMAFUuDci5GMxyPgO3gXg@mail.gmail.com","threadId":"48913","inReplyTo":"20180718204458.20936-2-benpeart@microsoft.com","subject":"Re: [PATCH v1 1/3] add unbounded Multi-Producer-Multi-Consumer queue","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-07-18T20:57:53Z","receivedAt":"2018-07-18T20:58:07Z","isPatch":true,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Wed, Jul 18, 2018 at 1:45 PM Ben Peart <Ben.Peart@microsoft.com> wrote:\n>\n\nDid you have any further considerations that are worth recording here?\n(memory, performance, CPU execution, threading, would all come to mind)\n\n> Signed-off-by: Ben Peart <benpeart@microsoft.com>\n\n> +/*\n> + * Initializes a mpmcq structure.\n> + */\n\nI'd find the name mpmcq a bit troubling if I were just stumbling upon it\nin the code without the knowledge of this review (and its abbreviation),\nmaybe just 'threadsafe_queue' ?\n\n> +extern void mpmcq_init(struct mpmcq *queue);\n\nWe prefer no extern keyword these days\nc.f. Documentation/CodingGuidelines:\n - Variables and functions local to a given source file should be marked\n   with \"static\". Variables that are visible to other source files\n   must be declared with \"extern\" in header files. However, function\n   declarations should not use \"extern\", as that is already the default.\n"},{"id":"352988","messageId":"CAGZ79kYkoK6R83fQXeXMunzxXYLxFiO8z9+cDY8Ku4fsR+9-TQ@mail.gmail.com","threadId":"48913","inReplyTo":"20180718204458.20936-1-benpeart@microsoft.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-07-18T21:02:19Z","receivedAt":"2018-07-18T21:02:33Z","isPatch":true,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"> don�t\n\nThe encoding seems to be broken somehow (also on)\nhttps://public-inbox.org/git/20180718204458.20936-1-benpeart@microsoft.com/\n\n\n> When I brought up this idea with some other git contributors they mentioned\n> that multi threading unpack_trees() had been discussed a few years ago on\n\nhttps://public-inbox.org/git/CACsJy8A0KUyxK_2NAMh+da9yithZM5d68rhqEVZe3NcMxinAjA@mail.gmail.com/\nhttps://public-inbox.org/git/20160415095139.GA3985@lanh/\n\n\n> the list but that the idea was discarded.  They couldn�t remember exactly\n> why it was discarded and none of us have been able to find the email threads\n> from that earlier discussion. As a result, I decided to write up this RFC\n> and see if the greater git community has ideas, suggestions, or more\n> background/history on whether this is a reasonable path to pursue or if\n> there are other/better ideas on how to speed up checkout especially on large\n> repos.\n\nIf you want more than a bare bones threaded queue, see\nhttps://public-inbox.org/git/1440724495-708-5-git-send-email-sbeller@google.com/\nfor inspiration.\n\nStefan\n"},{"id":"352989","messageId":"20180718213420.GA17291@sigill.intra.peff.net","threadId":"48913","inReplyTo":"20180718204458.20936-1-benpeart@microsoft.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-07-18T21:34:20Z","receivedAt":"2018-07-18T21:34:25Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jul 18, 2018 at 08:45:14PM +0000, Ben Peart wrote:\n\n> When working directories get big, checkout times start to suffer.  Even with\n> GVFS virtualization (which limits git to only having to update those files\n> that have been changed locally) we�re seeing P50 times for checkout of 31\n> seconds and the P80 time is 43 seconds.\n\nFunny aside: all of your apostrophes look like the unicode question\nmark. Looking at raw bytes of your mail, they're actually u+fffd\n(unicode \"replacement character\"). Your headers correctly claim to be\nutf8. So presumably they got munged by whatever converted to unicode and\ndidn't have the original character in its translation table. I wonder if\nthis was send-email (so really perl's encode module), or if your smtp\nserver tried to do an on-the-fly conversion (I know many servers will\nswitch the content-transfer-encoding, but I haven't seen a charset\nconversion before).\n\nAnyway, on to the actual discussion:\n\n> Here is a checkout command with tracing turned on to demonstrate where the\n> time is spent.  Note, this is somewhat of a �best case� as I�m simply\n> checking out the current commit:\n> \n> benpeart@gvfs-perf MINGW64 /f/os/src (official/rs_es_debug_dev)\n> $ /usr/src/git/git.exe checkout\n> 12:31:50.419016 read-cache.c:2006       performance: 1.180966800 s: read cache .git/index\n> 12:31:51.184636 name-hash.c:605         performance: 0.664575200 s: initialize name hash\n> 12:31:51.200280 preload-index.c:111     performance: 0.019811600 s: preload index\n> 12:31:51.294012 read-cache.c:1543       performance: 0.094515600 s: refresh index\n> 12:32:29.731344 unpack-trees.c:1358     performance: 33.889840200 s: traverse_trees\n> 12:32:37.512555 read-cache.c:2541       performance: 1.564438300 s: write index, changed mask = 28\n> 12:32:44.918730 unpack-trees.c:1358     performance: 7.243155600 s: traverse_trees\n> 12:32:44.965611 diff-lib.c:527          performance: 7.374729200 s: diff-index\n> Waiting for GVFS to parse index and update placeholder files...Succeeded\n> 12:32:46.824986 trace.c:420             performance: 57.715656000 s: git command: 'C:\\git-sdk-64\\usr\\src\\git\\git.exe' checkout\n\nWhat's the current state of the index before this checkout? I don't\nrecall offhand how aggressively we prune the tree walk based on the diff\nbetween the index and the tree we're loading. If we're starting from\nscratch, then obviously we do have to walk the whole thing. But in most\ncases we should be able to avoid walking into sub-trees where the index\nhas a matching cache_tree record.\n\nIf we're not doing that, it seems like that's going to be the big\nobvious win, because it reduces the number of trees we have to consider\nin the first place.\n\n> ODB cache\n> =========\n> Since traverse_trees() hits the ODB for each tree object (of which there are\n> over 500K in this repo) I wrote and tested having an in-memory ODB cache\n> that cached all tree objects.  This resulted in a > 50% hit ratio (largely\n> due to the fact we traverse the tree twice during checkout) but resulted in\n> only a minimal savings (1.3 seconds).\n\nIn my experience, one major cost of object access is decompression, both\ndelta and zlib. Trees in particular tend to delta very well across\nversions. We have a cache to try to reuse intermediate delta results,\nbut the default size is probably woefully undersized for your repository\n(I know from past tests it's undersized a bit even for the linux\nkernel).\n\nTry bumping core.deltaBaseCacheLimit to see if that has any impact. It's\n96MB by default.\n\nThere may also be some possible work in making it more aggressive about\nstoring the intermediate results. I seem to recall from past\nexplorations that it doesn't keep everything, and I don't know if its\nheuristics have ever been proven sane.\n\nFor zlib compression, I don't have numbers handy, but previous\nexperiments showed that trees don't actually benefit all that much from\nzlib (presumably because they're mostly random-looking hashes). So one\noption would be to try repacking _just_ the trees with\n\"pack.compression\" set to 0, and see how the result behaves. I suspect\nthat will be pretty painful with your giant multi-pack repo.\n\nIt might be slightly easier if we had an option to set the compression\nlevel on a per-type basis (both to experiment, and then of course if it\nworks to actually tune your repo).\n\nThe numbers above aren't specific enough to know how much time was spent\ndoing zlib stuff, though. And even with more specific probes, it's\ngenerally still hard to tell the difference between what's specific to\nthe compression level, and what's a result of the fact that zlib is\nessentially copying all the bytes from the filesystem into memory.\nStill, my timings with zstd[1] showed something like 10-20% improvement\non object access, so we should be able to get something at least as good\nby moving to no compression.\n\n[1] https://public-inbox.org/git/20161023080552.lma2v6zxmyaiiqz5@sigill.intra.peff.net/\n\n> Tree Graph File\n> ===============\n> I also considered storing the commit tree in an alternate structure that is\n> faster to load/parse (ala the Commit graph) but the cache results along with\n> the negligible impact of running checkout back to back (thus ensuring the\n> objects were cached in my file system cache) made me believe this would not\n> result in much savings. MIDX has already helped out here given we end up\n> with a lot of pack files of commits and trees.\n\nI don't think this will help. Tree objects are actually reasonably\ncompact on disk relative to the information you're getting out of them.\nAs opposed to commit objects, where you expand a kilobyte to get the 40\nbytes of parent pointer.\n\nSo there's probably benefit from storing tree _relationships_ (like the\ndiff between a commit and its parent) if it lets you avoid opening the\ntree, but for a checkout you really are walking the tree entries.\nPossibly with a diff to your current state (as above), but that\nrelationship is much less predictable (your current state is arbitrary,\nnot necessarily the commit parent).\n\n> Sparse tree traversal\n> =====================\n> We�ve sped up other parts of git by taking advantage of the existing\n> sparse-checkout/excludes logic to limit what files git has to consider to\n> those that have been modified by the user locally.  I haven�t been able to\n> think of a way to take advantage of that with unpack-trees() as when you are\n> merging n commits, a change/conflict can occur in any tree object so they\n> must all be traversed.  If I�m missing something here and there _is_ a way\n> to entirely skip large parts of the tree, please let me know!  Please note\n> that we�re already limiting the files that git needs to update in the\n> working directory via sparse-checkout/excludes but the other/merge logic\n> still executes for the entire tree whether there are files to update or not.\n\nFor narrow/partial clones, it seems like we'd ultimately have to\nconsider a merge of unavailable trees to be a conflict anyway. I.e., it\nseems reasonable to me that if we're sparse in path \"subdir\", then in a\nmerge, either:\n\n  1. Neither side touched \"subdir\", and we can ignore it without\n     descending. \n\n  2. One side touched \"subdir\", in which we can take the trivial\n     merge. This would be the common case when you pull somebody else's\n     branch that touched code you have marked as sparse.\n\n  3. Both sides touched it, in which case we must report a conflict.\n     This might happen if you pulled two topics which both touched the\n     same code (even if you didn't). It might even resolve cleanly if we\n     had the actual trees and blobs to look at, but that just means\n     _you_ can't resolve it in your sparse clone.\n\nI'd expect (1) and (2) to be the common cases. I won't be at all\nsurprised if unpack_trees() isn't particularly smart about that, though.\n\n> Multi-threading unpack_trees()\n> ==============================\n> The current model of unpack_trees() is that a single thread recursively\n> traverses each tree object as it comes across it.  One thought I had was to\n> multi-thread the traversal so that each tree object could be processed in\n> parallel.  To test this idea out, I wrote an unbounded\n> Multi-Product-Multi-Consumer queue and then wrote a\n> traverse_trees_parallel() function that would add any new tree objects into\n> the queue where they can be processed by a pool of worker threads.  Each\n> thread will wake up when there is work in the queue, remove a tree object,\n> process it adding any additional tree objects it finds.\n\nI'm generally terrified of multi-threading anything in the core parts of\nGit. There are so many latent bits of non-reentrant or racy code.\n\nI think your queue suggestion may be the sanest approach, though,\nbecause it makes it keeps the responsibilities of the worker threads\npretty clear.\n\n> When I brought up this idea with some other git contributors they mentioned\n> that multi threading unpack_trees() had been discussed a few years ago on\n> the list but that the idea was discarded.  They couldn�t remember exactly\n> why it was discarded and none of us have been able to find the email threads\n> from that earlier discussion. As a result, I decided to write up this RFC\n> and see if the greater git community has ideas, suggestions, or more\n> background/history on whether this is a reasonable path to pursue or if\n> there are other/better ideas on how to speed up checkout especially on large\n> repos.\n\nI don't remember any specific discussion, and didn't dig anything up\nafter a few minutes. But I'd be willing to bet that the primary reason\nit would not be pursued is the general lack of thread safety in the\ncurrent codebase.\n\n-Peff\n"},{"id":"353001","messageId":"xmqqbmb4i25g.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"20180718204458.20936-4-benpeart@microsoft.com","subject":"Re: [PATCH v1 3/3] Add initial parallel version of unpack_trees()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-18T22:56:27Z","receivedAt":"2018-07-18T22:56:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <Ben.Peart@microsoft.com> writes:\n\n> +\t * Fetch the tree from the ODB for each peer directory in the\n> +\t * n commits.\n> +\t *\n> +\t * For 2- and 3-way traversals, we try to avoid hitting the\n> +\t * ODB twice for the same OID.  This should yield a nice speed\n> +\t * up in checkouts and merges when the commits are similar.\n> +\t *\n> +\t * We don't bother doing the full O(n^2) search for larger n,\n> +\t * because wider traversals don't happen that often and we\n> +\t * avoid the search setup.\n\nIt is sensible to optimize for common cases while leaving out the\ncomplexity that is only needed to support rare cases.\n\n> +\t * When 2 peer OIDs are the same, we just copy the tree\n> +\t * descriptor data.  This implicitly borrows the buffer\n> +\t * data from the earlier cell.\n\ncell meaning...?\n\n\n> +\tfor (i = 0; i < n; i++, dirmask >>= 1) {\n> +\t\tif (i > 0 && are_same_oid(&names[i], &names[i - 1]))\n> +\t\t\tnewinfo->t[i] = newinfo->t[i - 1];\n> +\t\telse if (i > 1 && are_same_oid(&names[i], &names[i - 2]))\n> +\t\t\tnewinfo->t[i] = newinfo->t[i - 2];\n> +\t\telse {\n> +\t\t\tconst struct object_id *oid = NULL;\n> +\t\t\tif (dirmask & 1)\n> +\t\t\t\toid = names[i].oid;\n> +\n> +\t\t\t/*\n> +\t\t\t * fill_tree_descriptor() will load the tree from the\n> +\t\t\t * ODB. Accessing the ODB is not thread safe so\n> +\t\t\t * serialize access using the odb_mutex.\n> +\t\t\t */\n> +\t\t\tpthread_mutex_lock(&o->odb_mutex);\n> +\t\t\tnewinfo->buf[newinfo->nr_buf++] =\n> +\t\t\t\tfill_tree_descriptor(newinfo->t + i, oid);\n> +\t\t\tpthread_mutex_unlock(&o->odb_mutex);\n> +\t\t}\n> +\t}\n> +\n> +\t/*\n> +\t * We can't play games with the cache bottom as we are processing\n> +\t * the tree objects in parallel.\n> +\t * newinfo->bottom = switch_cache_bottom(&newinfo->info);\n> +\t */\n\nWould the resulting code match corresponding entries from two/three\ntrees correctly with a tree with entries \"foo.\" (blob), \"foo/\" (has\nsubtree), and \"foo0\" (blob) at the same time, without adjusting the\nbottom?  I am worried because cache_bottom stuff is not about\noptimization but is about correctness.\n\n> +\t/* All I really need here is fetch_and_add() */\n> +\tpthread_mutex_lock(&o->work_mutex);\n> +\to->remaining_work++;\n> +\tpthread_mutex_unlock(&o->work_mutex);\n> +\tmpmcq_push(&o->queue, &newinfo->entry);\n\nNice.  I like the general idea.\n\n"},{"id":"353068","messageId":"xmqqsh4fdos4.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"20180718204458.20936-2-benpeart@microsoft.com","subject":"Re: [PATCH v1 1/3] add unbounded Multi-Producer-Multi-Consumer queue","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-19T19:11:07Z","receivedAt":"2018-07-19T19:11:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <Ben.Peart@microsoft.com> writes:\n\n> +/*\n> + * struct mpmcq_entry is an opaque structure representing an entry in the\n> + * queue.\n> + */\n> +struct mpmcq_entry {\n> +\tstruct mpmcq_entry *next;\n> +};\n> +\n> +/*\n> + * struct mpmcq is the concurrent queue structure. Members should not be\n> + * modified directly.\n> + */\n> +struct mpmcq {\n> +\tstruct mpmcq_entry *head;\n> +\tpthread_mutex_t mutex;\n> +\tpthread_cond_t condition;\n> +\tint cancel;\n> +};\n\nThis calls itself a queue, but a new element goes to the beginning\nof a singly linked list, and the only way to take an element out is\nfrom near the beinning of the linked list, so it looks more like a\nLIFO stack to me.\n\nI do not know how much it matters, as the name mpmcq is totally\nopaque to readers so perhaps readers are not even aware of various\naspects of the service, e.g. how it works, what fairness it gives to\nthe calling code, etc.\n\n"},{"id":"353368","messageId":"a2ad0044-f317-69f7-f2bb-488111c626fb@gmail.com","threadId":"48913","inReplyTo":"20180718213420.GA17291@sigill.intra.peff.net","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-23T15:48:18Z","receivedAt":"2018-07-23T15:48:24Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/18/2018 5:34 PM, Jeff King wrote:\n> On Wed, Jul 18, 2018 at 08:45:14PM +0000, Ben Peart wrote:\n> \n>> When working directories get big, checkout times start to suffer.  Even with\n>> GVFS virtualization (which limits git to only having to update those files\n>> that have been changed locally) we�re seeing P50 times for checkout of 31\n>> seconds and the P80 time is 43 seconds.\n> \n> Funny aside: all of your apostrophes look like the unicode question\n> mark. Looking at raw bytes of your mail, they're actually u+fffd\n> (unicode \"replacement character\"). Your headers correctly claim to be\n> utf8. So presumably they got munged by whatever converted to unicode and\n> didn't have the original character in its translation table. I wonder if\n> this was send-email (so really perl's encode module), or if your smtp\n> server tried to do an on-the-fly conversion (I know many servers will\n> switch the content-transfer-encoding, but I haven't seen a charset\n> conversion before).\n> \n\nThis was my bad.  I wrote the email in Word so I could get spell \nchecking and it has this 'feature' where it converts all straight quotes \nto \"smart quotes.\"  I just forgot to search/replace them back to \nstraight quotes before sending the mail.\n\n> Anyway, on to the actual discussion:\n> \n>> Here is a checkout command with tracing turned on to demonstrate where the\n>> time is spent.  Note, this is somewhat of a �best case� as I�m simply\n>> checking out the current commit:\n>>\n>> benpeart@gvfs-perf MINGW64 /f/os/src (official/rs_es_debug_dev)\n>> $ /usr/src/git/git.exe checkout\n>> 12:31:50.419016 read-cache.c:2006       performance: 1.180966800 s: read cache .git/index\n>> 12:31:51.184636 name-hash.c:605         performance: 0.664575200 s: initialize name hash\n>> 12:31:51.200280 preload-index.c:111     performance: 0.019811600 s: preload index\n>> 12:31:51.294012 read-cache.c:1543       performance: 0.094515600 s: refresh index\n>> 12:32:29.731344 unpack-trees.c:1358     performance: 33.889840200 s: traverse_trees\n>> 12:32:37.512555 read-cache.c:2541       performance: 1.564438300 s: write index, changed mask = 28\n>> 12:32:44.918730 unpack-trees.c:1358     performance: 7.243155600 s: traverse_trees\n>> 12:32:44.965611 diff-lib.c:527          performance: 7.374729200 s: diff-index\n>> Waiting for GVFS to parse index and update placeholder files...Succeeded\n>> 12:32:46.824986 trace.c:420             performance: 57.715656000 s: git command: 'C:\\git-sdk-64\\usr\\src\\git\\git.exe' checkout\n> \n> What's the current state of the index before this checkout? \n\nThis was after running \"git checkout\" multiple times so there was really \nnothing for git to do.\n\n> I don't\n> recall offhand how aggressively we prune the tree walk based on the diff\n> between the index and the tree we're loading. If we're starting from > scratch, then obviously we do have to walk the whole thing. But in most\n> cases we should be able to avoid walking into sub-trees where the index\n> has a matching cache_tree record.\n> \n> If we're not doing that, it seems like that's going to be the big\n> obvious win, because it reduces the number of trees we have to consider\n> in the first place.\n> \n\nI agree this could be a big win.  Especially in large trees, the \npercentage of the tree that changes between two commits is often quite \nsmall.  Saving 100% of that is a much bigger win than actually doing all \nthat work even in parallel. Today, we aren't aggressive at all and do no \npruning.\n\nThis brings up a concern I have with this approach altogether. In an \nearlier patch series, I tried to optimize the \"git checkout -b\" code \npath to not update every file in the working directory but only to \ncreate the new branch and switch to it.  The feedback to that patch was \nthat people rely on the current behavior of rewriting every file so the \npatch was rejected.  This earlier attempt/failure to optimize checkout \nmakes me worried that _any_ effort to prune the tree will be rejected \nfor the same reason.\n\nI'd be interested in how we can prune the tree and only do the work \nrequired without breaking the implied behavior of the current \nimplementation. Would it be acceptable to have two code paths 1) the old \none for back compat that updates every file whether there are changes or \nnot and 2) a new/optimized one that only does the minimum work required? \n  Then we could put which code path executes by default behind by a new \nconfig setting that allows people to opt-in to the new/faster behavior.\n\nAny other ideas or suggestions that don't require coming up with new git \ncommands (ie \"git fast-checkout\") and retraining existing git users?\n\n>> ODB cache\n>> =========\n>> Since traverse_trees() hits the ODB for each tree object (of which there are\n>> over 500K in this repo) I wrote and tested having an in-memory ODB cache\n>> that cached all tree objects.  This resulted in a > 50% hit ratio (largely\n>> due to the fact we traverse the tree twice during checkout) but resulted in\n>> only a minimal savings (1.3 seconds).\n> \n> In my experience, one major cost of object access is decompression, both\n> delta and zlib. Trees in particular tend to delta very well across\n> versions. We have a cache to try to reuse intermediate delta results,\n> but the default size is probably woefully undersized for your repository\n> (I know from past tests it's undersized a bit even for the linux\n> kernel).\n> \n> Try bumping core.deltaBaseCacheLimit to see if that has any impact. It's\n> 96MB by default.\n> \n> There may also be some possible work in making it more aggressive about\n> storing the intermediate results. I seem to recall from past\n> explorations that it doesn't keep everything, and I don't know if its\n> heuristics have ever been proven sane.\n> \n> For zlib compression, I don't have numbers handy, but previous\n> experiments showed that trees don't actually benefit all that much from\n> zlib (presumably because they're mostly random-looking hashes). So one\n> option would be to try repacking _just_ the trees with\n> \"pack.compression\" set to 0, and see how the result behaves. I suspect\n> that will be pretty painful with your giant multi-pack repo.\n> \n> It might be slightly easier if we had an option to set the compression\n> level on a per-type basis (both to experiment, and then of course if it\n> works to actually tune your repo).\n> \n> The numbers above aren't specific enough to know how much time was spent\n> doing zlib stuff, though. And even with more specific probes, it's\n> generally still hard to tell the difference between what's specific to\n> the compression level, and what's a result of the fact that zlib is\n> essentially copying all the bytes from the filesystem into memory.\n> Still, my timings with zstd[1] showed something like 10-20% improvement\n> on object access, so we should be able to get something at least as good\n> by moving to no compression.\n> \n> [1] https://public-inbox.org/git/20161023080552.lma2v6zxmyaiiqz5@sigill.intra.peff.net/\n> \n\nThanks, these are good ideas to pursue.  I've added them to my list of \nthings to look into but believe pruning the tree or traversing it in \nparallel has more performance saving potential so I'll be looking there \nfirst.\n\n<snip>\n\n\n>> Multi-threading unpack_trees()\n>> ==============================\n>> The current model of unpack_trees() is that a single thread recursively\n>> traverses each tree object as it comes across it.  One thought I had was to\n>> multi-thread the traversal so that each tree object could be processed in\n>> parallel.  To test this idea out, I wrote an unbounded\n>> Multi-Product-Multi-Consumer queue and then wrote a\n>> traverse_trees_parallel() function that would add any new tree objects into\n>> the queue where they can be processed by a pool of worker threads.  Each\n>> thread will wake up when there is work in the queue, remove a tree object,\n>> process it adding any additional tree objects it finds.\n> \n> I'm generally terrified of multi-threading anything in the core parts of\n> Git. There are so many latent bits of non-reentrant or racy code.\n> \n> I think your queue suggestion may be the sanest approach, though,\n> because it makes it keeps the responsibilities of the worker threads\n> pretty clear.\n> \n\nI agree the thought of multi-threading unpack_trees() is daunting!  It \nwould be nice if the model of pruning the tree was sufficient to get \nreasonable performance with large repos.  I guess we'll see...\n\n>> When I brought up this idea with some other git contributors they mentioned\n>> that multi threading unpack_trees() had been discussed a few years ago on\n>> the list but that the idea was discarded.  They couldn�t remember exactly\n>> why it was discarded and none of us have been able to find the email threads\n>> from that earlier discussion. As a result, I decided to write up this RFC\n>> and see if the greater git community has ideas, suggestions, or more\n>> background/history on whether this is a reasonable path to pursue or if\n>> there are other/better ideas on how to speed up checkout especially on large\n>> repos.\n> \n> I don't remember any specific discussion, and didn't dig anything up\n> after a few minutes. But I'd be willing to bet that the primary reason\n> it would not be pursued is the general lack of thread safety in the\n> current codebase.\n> \n> -Peff\n> \n"},{"id":"353378","messageId":"CACsJy8D-3sSnoyQZKxeLK-2RmpJSGkziAp5Gf4QpUnxwnhchSQ@mail.gmail.com","threadId":"48913","inReplyTo":"a2ad0044-f317-69f7-f2bb-488111c626fb@gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-23T17:03:16Z","receivedAt":"2018-07-23T17:03:45Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Jul 23, 2018 at 5:50 PM Ben Peart <peartben@gmail.com> wrote:\n> > Anyway, on to the actual discussion:\n> >\n> >> Here is a checkout command with tracing turned on to demonstrate where the\n> >> time is spent.  Note, this is somewhat of a �best case� as I�m simply\n> >> checking out the current commit:\n> >>\n> >> benpeart@gvfs-perf MINGW64 /f/os/src (official/rs_es_debug_dev)\n> >> $ /usr/src/git/git.exe checkout\n> >> 12:31:50.419016 read-cache.c:2006       performance: 1.180966800 s: read cache .git/index\n> >> 12:31:51.184636 name-hash.c:605         performance: 0.664575200 s: initialize name hash\n> >> 12:31:51.200280 preload-index.c:111     performance: 0.019811600 s: preload index\n> >> 12:31:51.294012 read-cache.c:1543       performance: 0.094515600 s: refresh index\n> >> 12:32:29.731344 unpack-trees.c:1358     performance: 33.889840200 s: traverse_trees\n> >> 12:32:37.512555 read-cache.c:2541       performance: 1.564438300 s: write index, changed mask = 28\n> >> 12:32:44.918730 unpack-trees.c:1358     performance: 7.243155600 s: traverse_trees\n> >> 12:32:44.965611 diff-lib.c:527          performance: 7.374729200 s: diff-index\n> >> Waiting for GVFS to parse index and update placeholder files...Succeeded\n> >> 12:32:46.824986 trace.c:420             performance: 57.715656000 s: git command: 'C:\\git-sdk-64\\usr\\src\\git\\git.exe' checkout\n> >\n> > What's the current state of the index before this checkout?\n>\n> This was after running \"git checkout\" multiple times so there was really\n> nothing for git to do.\n\nHmm.. this means cache-tree is fully valid, unless you have changes in\nindex. We're quite aggressive in repairing cache-tree since aecf567cbf\n(cache-tree: create/update cache-tree on checkout - 2014-07-05). If we\nhave very good cache-tree records and still spend 33s on\ntraverse_trees, maybe there's something else.\n\n> >> ODB cache\n> >> =========\n> >> Since traverse_trees() hits the ODB for each tree object (of which there are\n> >> over 500K in this repo) I wrote and tested having an in-memory ODB cache\n> >> that cached all tree objects.  This resulted in a > 50% hit ratio (largely\n> >> due to the fact we traverse the tree twice during checkout) but resulted in\n> >> only a minimal savings (1.3 seconds).\n> >\n> > In my experience, one major cost of object access is decompression, both\n> > delta and zlib. Trees in particular tend to delta very well across\n> > versions. We have a cache to try to reuse intermediate delta results,\n> > but the default size is probably woefully undersized for your repository\n> > (I know from past tests it's undersized a bit even for the linux\n> > kernel).\n> >\n> > Try bumping core.deltaBaseCacheLimit to see if that has any impact. It's\n> > 96MB by default.\n> >\n> > There may also be some possible work in making it more aggressive about\n> > storing the intermediate results. I seem to recall from past\n> > explorations that it doesn't keep everything, and I don't know if its\n> > heuristics have ever been proven sane.\n\nCould we be a bit more flexible about cache size? Say if we know\nthere's 8 GB memory still available, we should be able to use like 1\nGB at least (and that's done automatically without tinkering with\nconfig).\n-- \nDuy\n"},{"id":"353415","messageId":"6ff6fbdc-d9cf-019f-317c-7fdba31105c6@gmail.com","threadId":"48913","inReplyTo":"CACsJy8D-3sSnoyQZKxeLK-2RmpJSGkziAp5Gf4QpUnxwnhchSQ@mail.gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-23T20:51:38Z","receivedAt":"2018-07-23T20:51:42Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/23/2018 1:03 PM, Duy Nguyen wrote:\n> On Mon, Jul 23, 2018 at 5:50 PM Ben Peart <peartben@gmail.com> wrote:\n>>> Anyway, on to the actual discussion:\n>>>\n>>>> Here is a checkout command with tracing turned on to demonstrate where the\n>>>> time is spent.  Note, this is somewhat of a �best case� as I�m simply\n>>>> checking out the current commit:\n>>>>\n>>>> benpeart@gvfs-perf MINGW64 /f/os/src (official/rs_es_debug_dev)\n>>>> $ /usr/src/git/git.exe checkout\n>>>> 12:31:50.419016 read-cache.c:2006       performance: 1.180966800 s: read cache .git/index\n>>>> 12:31:51.184636 name-hash.c:605         performance: 0.664575200 s: initialize name hash\n>>>> 12:31:51.200280 preload-index.c:111     performance: 0.019811600 s: preload index\n>>>> 12:31:51.294012 read-cache.c:1543       performance: 0.094515600 s: refresh index\n>>>> 12:32:29.731344 unpack-trees.c:1358     performance: 33.889840200 s: traverse_trees\n>>>> 12:32:37.512555 read-cache.c:2541       performance: 1.564438300 s: write index, changed mask = 28\n>>>> 12:32:44.918730 unpack-trees.c:1358     performance: 7.243155600 s: traverse_trees\n>>>> 12:32:44.965611 diff-lib.c:527          performance: 7.374729200 s: diff-index\n>>>> Waiting for GVFS to parse index and update placeholder files...Succeeded\n>>>> 12:32:46.824986 trace.c:420             performance: 57.715656000 s: git command: 'C:\\git-sdk-64\\usr\\src\\git\\git.exe' checkout\n>>>\n>>> What's the current state of the index before this checkout?\n>>\n>> This was after running \"git checkout\" multiple times so there was really\n>> nothing for git to do.\n> \n> Hmm.. this means cache-tree is fully valid, unless you have changes in\n> index. We're quite aggressive in repairing cache-tree since aecf567cbf\n> (cache-tree: create/update cache-tree on checkout - 2014-07-05). If we\n> have very good cache-tree records and still spend 33s on\n> traverse_trees, maybe there's something else.\n> \n\nI'm not at all familiar with the cache-tree and couldn't find any \ndocumentation on it other than index-format.txt which says \"it helps \nspeed up tree object generation for a new commit.\"  In this particular \ncase, no new commit is being created so I don't know that the cache-tree \nwould help.\n\nAfter a quick look at the code, the only place I can find that tries to \nuse cache_tree_matches_traversal() is in unpack_callback() and that only \nhappens if n == 1 and in the \"git checkout\" case, n == 2. Am I missing \nsomething?\n"},{"id":"353439","messageId":"20180724042017.GA13248@sigill.intra.peff.net","threadId":"48913","inReplyTo":"6ff6fbdc-d9cf-019f-317c-7fdba31105c6@gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-07-24T04:20:17Z","receivedAt":"2018-07-24T04:20:22Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jul 23, 2018 at 04:51:38PM -0400, Ben Peart wrote:\n\n> > Hmm.. this means cache-tree is fully valid, unless you have changes in\n> > index. We're quite aggressive in repairing cache-tree since aecf567cbf\n> > (cache-tree: create/update cache-tree on checkout - 2014-07-05). If we\n> > have very good cache-tree records and still spend 33s on\n> > traverse_trees, maybe there's something else.\n> \n> I'm not at all familiar with the cache-tree and couldn't find any\n> documentation on it other than index-format.txt which says \"it helps speed\n> up tree object generation for a new commit.\"  In this particular case, no\n> new commit is being created so I don't know that the cache-tree would help.\n\nIt's basically an index extension that mirrors the tree structure within\nthe index, telling you the sha1 of the three that _would_ be generated\nfrom any particular path. So any time you're walking a tree alongside\nthe index, in theory you should be able to say \"the cache-tree for this\nsubset of the index matches the tree\" and skip over a bunch of entries.\n\nAt least that's my view of it. unpack_trees() has always been a\nterrifying beast that I've avoided looking too closely at.\n\n> After a quick look at the code, the only place I can find that tries to use\n> cache_tree_matches_traversal() is in unpack_callback() and that only happens\n> if n == 1 and in the \"git checkout\" case, n == 2. Am I missing something?\n\nLooks like it's trying to special-case \"diff-index --cached\". Which\nkind-of makes sense. In the non-cached case, we're thinking not only\nabout the relationship between the index and the tree, but also whether\nthe on-disk files are up to date.\n\nAnd that would be the same for checkout. We want to know not only\nwhether there are changes to make to the index, but also whether the\non-disk files need to be updated from the index.\n\nBut I assume in your case that we've just refreshed the index quickly\nusing fsmonitor. So I think in the long run what you want is:\n\n  1. fsmonitor tells us which index entries are not clean\n\n  2. based on the unclean list, we invalidate cache-tree entries for\n     those paths\n\n  3. if we have a valid cache-tree entry, we should be able to skip\n     digging into that tree; if not, then we walk the index and tree as\n     normal, adding/deleting index entries and updating (or complaining\n     about) modified on-disk files\n\nI think the \"n\" adds an extra layer of complexity. n==2 means we're\ndoing a \"2-way\" merge. Moving from tree X to tree Y, and dealing with\nthe index as we go. Naively I _think_ we'd be OK to just extend the rule\nto \"if both subtrees match each other _and_ match the valid cache-tree,\nthen we can skip\".\n\nAgain, I'm a little out of my area of expertise here, but cargo-culting\nlike this:\n\ndiff --git a/sha1-file.c b/sha1-file.c\nindex de4839e634..c105af70ce 100644\n--- a/sha1-file.c\n+++ b/sha1-file.c\n@@ -1375,6 +1375,7 @@ static void *read_object(const unsigned char *sha1, enum object_type *type,\n \n \tif (oid_object_info_extended(the_repository, &oid, &oi, 0) < 0)\n \t\treturn NULL;\n+\ttrace_printf(\"reading %s %s\", type_name(*type), sha1_to_hex(sha1));\n \treturn content;\n }\n \ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 66741130ae..cfdad4133d 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -1075,6 +1075,23 @@ static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, str\n \t\t\t\to->cache_bottom += matches;\n \t\t\t\treturn mask;\n \t\t\t}\n+\t\t} else if (n == 2 && S_ISDIR(names->mode) &&\n+\t\t\t   names[0].mode == names[1].mode &&\n+\t\t\t   !strcmp(names[0].path, names[1].path) &&\n+\t\t\t   !oidcmp(names[0].oid, names[1].oid)\n+\t\t\t   /* && somehow account for modified on-disk files */) {\n+\t\t\tint matches;\n+\n+\t\t\t/*\n+\t\t\t * we know that the two trees have the same oid, so we\n+\t\t\t * only need to look at one of them\n+\t\t\t */\n+\t\t\tmatches = cache_tree_matches_traversal(o->src_index->cache_tree,\n+\t\t\t\t\t\t\t       names, info);\n+\t\t\tif (matches) {\n+\t\t\t\to->cache_bottom += matches;\n+\t\t\t\treturn mask;\n+\t\t\t}\n \t\t}\n \n \t\tif (traverse_trees_recursive(n, dirmask, mask & ~dirmask,\n\nseems to avoid the tree reads when running \"GIT_TRACE=1 git checkout\".\nIt also totally empties the index. ;) So clearly we have to do a bit\nmore there. Probably rather than just bumping o->cache_bottom forward,\nwe'd need to actually move those entries into the new index. Or maybe\nit's something else entirely (I did say cargo-culting, right?).\n\n-Peff\n"},{"id":"353440","messageId":"20180724042740.GB13248@sigill.intra.peff.net","threadId":"48913","inReplyTo":"CACsJy8D-3sSnoyQZKxeLK-2RmpJSGkziAp5Gf4QpUnxwnhchSQ@mail.gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-07-24T04:27:40Z","receivedAt":"2018-07-24T04:27:43Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jul 23, 2018 at 07:03:16PM +0200, Duy Nguyen wrote:\n\n> > > Try bumping core.deltaBaseCacheLimit to see if that has any impact. It's\n> > > 96MB by default.\n> > >\n> > > There may also be some possible work in making it more aggressive about\n> > > storing the intermediate results. I seem to recall from past\n> > > explorations that it doesn't keep everything, and I don't know if its\n> > > heuristics have ever been proven sane.\n> \n> Could we be a bit more flexible about cache size? Say if we know\n> there's 8 GB memory still available, we should be able to use like 1\n> GB at least (and that's done automatically without tinkering with\n> config).\n\nI have mixed feelings on that kind of auto-scaling for caches. Git isn't\nalways the only program running (or maybe you even have several git\noperations running at once). So in many cases you'd want a more holistic\nview of the system, and what resources are available.\n\nThe OS already does OK scheduling CPU and the block cache for our mmap'd\nfiles. I don't know if there's a way to communicate with it about this\nkind of cache. I guess asking \"what memory is free\" is one way to do\nthat. But it's not always the best answer (because we might be happy\ntrade off some block cache, etc). On the other hand, that would always\ngive us a conservative value, so if we picked min(96MB, free_mem /\nnr_cpu) or something, that might be an OK rule of thumb.\n\n-Peff\n"},{"id":"353443","messageId":"xmqqefft18m3.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"6ff6fbdc-d9cf-019f-317c-7fdba31105c6@gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-24T05:54:44Z","receivedAt":"2018-07-24T05:54:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <peartben@gmail.com> writes:\n\n>> Hmm.. this means cache-tree is fully valid, unless you have changes in\n>> index. We're quite aggressive in repairing cache-tree since aecf567cbf\n>> (cache-tree: create/update cache-tree on checkout - 2014-07-05). If we\n>> have very good cache-tree records and still spend 33s on\n>> traverse_trees, maybe there's something else.\n>>\n>\n> I'm not at all familiar with the cache-tree and couldn't find any\n> documentation on it other than index-format.txt which says \"it helps\n> speed up tree object generation for a new commit.\"  In this particular\n> case, no new commit is being created so I don't know that the\n> cache-tree would help.\n\ncache-tree is an index extension that records tree object names for\nsubdirectories you see in the index.  Every time you write the\ncontents of the index as a tree object, we need to collect the\nobject name for each top-level paths and write a new top-level tree\nobject out, after doing the same recursively for any modified\nsubdirectory.  Whenever you add, remove or modify a path in the\nindex, the cache-tree entry for enclosing directories are\ninvalidated, so a cache-tree entry that is still valid means that\nall the paths in the index under that directory match the contents\nof the tree object that the cache-tree entry holds.\n\nAnd that property is used by \"diff-index --cached $TREE\" that is run\ninternally.  When we find that the subdirectory \"D\"'s cache-tree\nentry is valid in the index, and the tree object recorded in the\ncache-tree for that subdirectory matches the subtree D in the tree\nobject $TREE, then \"diff-index --cached\" ignores the entire\nsubdirectory D (which saves relatively little in the index as it\nonly needs to scan what is already in the memory forward, but on the\n$TREE traversal side, it does not have to even open a subtree, that\ncan save a lot), and with a well-populated cache-tree, it can save a\nsignificant processing.\n\nI think that is what Duy meant to refer to while looking at the\nnumbers.\n\n> After a quick look at the code, the only place I can find that tries\n> to use cache_tree_matches_traversal() is in unpack_callback() and that\n> only happens if n == 1 and in the \"git checkout\" case, n == 2. Am I\n> missing something?\n"},{"id":"353465","messageId":"20180724151336.GA1957@duynguyen.home","threadId":"48913","inReplyTo":"6ff6fbdc-d9cf-019f-317c-7fdba31105c6@gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-24T15:13:36Z","receivedAt":"2018-07-24T15:13:43Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Jul 23, 2018 at 04:51:38PM -0400, Ben Peart wrote:\n> >>> What's the current state of the index before this checkout?\n> >>\n> >> This was after running \"git checkout\" multiple times so there was really\n> >> nothing for git to do.\n> > \n> > Hmm.. this means cache-tree is fully valid, unless you have changes in\n> > index. We're quite aggressive in repairing cache-tree since aecf567cbf\n> > (cache-tree: create/update cache-tree on checkout - 2014-07-05). If we\n> > have very good cache-tree records and still spend 33s on\n> > traverse_trees, maybe there's something else.\n> > \n> \n> I'm not at all familiar with the cache-tree and couldn't find any \n> documentation on it other than index-format.txt which says \"it helps \n> speed up tree object generation for a new commit.\"\n\nI guess you have the starting points you need after Jeff's and Junio's\nexplanation (and it would be great if cache-tree could actually be for\nfor this two-way merge). But to make it easier for new people in\nfuture, maybe we should add this?\n\nThis is basically a ripoff of Junio's explanation with starting points\n(write-tree and index-format.txt). I wanted to incorporate some pieces\nfrom Jeff's too but I think Junio's already covered it well.\n\n-- 8< --\nSubject: [PATCH] cache-tree.h: more description of what it is and what's it used for\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache-tree.h | 29 +++++++++++++++++++++++++++++\n 1 file changed, 29 insertions(+)\n\ndiff --git a/cache-tree.h b/cache-tree.h\nindex cfd5328cc9..d25a800a72 100644\n--- a/cache-tree.h\n+++ b/cache-tree.h\n@@ -5,6 +5,35 @@\n #include \"tree.h\"\n #include \"tree-walk.h\"\n \n+/*\n+ * cache-tree is an index extension that records tree object names for\n+ * subdirectories you see in the index. It is mainly used for\n+ * generating trees from the index before you create a new commit (see\n+ * builtin/write-tree.c as starting point) but it's also used in \"git\n+ * diff-index --cached $TREE\" as an optimization. See index-format.txt\n+ * for on-disk format.\n+ *\n+ * Every time you write the contents of the index as a tree object, we\n+ * need to collect the object name for each top-level paths and write\n+ * a new top-level tree object out, after doing the same recursively\n+ * for any modified subdirectory. Whenever you add, remove or modify a\n+ * path in the index, the cache-tree entry for enclosing directories\n+ * are invalidated, so a cache-tree entry that is still valid means\n+ * that all the paths in the index under that directory match the\n+ * contents of the tree object that the cache-tree entry holds.\n+ *\n+ * And that property is used by \"diff-index --cached $TREE\" that is\n+ * run internally.  When we find that the subdirectory \"D\"'s\n+ * cache-tree entry is valid in the index, and the tree object\n+ * recorded in the cache-tree for that subdirectory matches the\n+ * subtree D in the tree object $TREE, then \"diff-index --cached\"\n+ * ignores the entire subdirectory D (which saves relatively little in\n+ * the index as it only needs to scan what is already in the memory\n+ * forward, but on the $TREE traversal side, it does not have to even\n+ * open a subtree, that can save a lot), and with a well-populated\n+ * cache-tree, it can save a significant processing.\n+ */\n+\n struct cache_tree;\n struct cache_tree_sub {\n \tstruct cache_tree *cache_tree;\n-- \n2.18.0.656.gda699b98b3\n\n-- 8< --\n"},{"id":"353468","messageId":"CACsJy8Du28jMyfdyhxpVxyw5+Xh+9eX==3x8YJSnmw6GAoRhTA@mail.gmail.com","threadId":"48913","inReplyTo":"20180724042017.GA13248@sigill.intra.peff.net","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-24T15:33:15Z","receivedAt":"2018-07-24T15:33:43Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Jul 24, 2018 at 6:20 AM Jeff King <peff@peff.net> wrote:\n> At least that's my view of it. unpack_trees() has always been a\n> terrifying beast that I've avoided looking too closely at.\n\n/me nods on the terrifying part.\n\n> > After a quick look at the code, the only place I can find that tries to use\n> > cache_tree_matches_traversal() is in unpack_callback() and that only happens\n> > if n == 1 and in the \"git checkout\" case, n == 2. Am I missing something?\n\nSo we do not actually use cache-tree? Big optimization opportunity (if\nwe can make it!).\n\n> Looks like it's trying to special-case \"diff-index --cached\". Which\n> kind-of makes sense. In the non-cached case, we're thinking not only\n> about the relationship between the index and the tree, but also whether\n> the on-disk files are up to date.\n>\n> And that would be the same for checkout. We want to know not only\n> whether there are changes to make to the index, but also whether the\n> on-disk files need to be updated from the index.\n>\n> But I assume in your case that we've just refreshed the index quickly\n> using fsmonitor. So I think in the long run what you want is:\n>\n>   1. fsmonitor tells us which index entries are not clean\n>\n>   2. based on the unclean list, we invalidate cache-tree entries for\n>      those paths\n>\n>   3. if we have a valid cache-tree entry, we should be able to skip\n>      digging into that tree; if not, then we walk the index and tree as\n>      normal, adding/deleting index entries and updating (or complaining\n>      about) modified on-disk files\n\nIf you tie this optimization to twoway_merge specifically (by checking\n\"fn\" field), then I think we can do it even better. Since\ncache_tree_matches_traversal() is one (hopefully not too costly)\nlookup, we can do it without checking with fsmonitor or whatever and\nonly do so when we have found a cache tree.\n\nThen if we write this new special code just for twoway_merge, we need\nto tighten the checks a bit. I think in this case twoway_merge() will\nbe called with \"oldtree\" as same as \"newtree\" (and \"current\" may\ncontains dirty stuff from the index). Then\n\n - o->df_conflict_entry should be NULL (because we handle it slightly\ndifferently in twoway_merge)\n - \"current\" should not have CE_CONFLICTED\n\nthen I believe we will fall into case /* 20 or 21 */ where\nmerged_entry() is suppoed to be called on all entries and it would\nchange nothing in the index since newtree is the same as oldtree, and\nwe could just jump over the whole tree in traverse_trees().\n\n> I think the \"n\" adds an extra layer of complexity. n==2 means we're\n> doing a \"2-way\" merge. Moving from tree X to tree Y, and dealing with\n> the index as we go. Naively I _think_ we'd be OK to just extend the rule\n> to \"if both subtrees match each other _and_ match the valid cache-tree,\n> then we can skip\".\n>\n> Again, I'm a little out of my area of expertise here, but cargo-culting\n> like this:\n>\n> diff --git a/sha1-file.c b/sha1-file.c\n> index de4839e634..c105af70ce 100644\n> --- a/sha1-file.c\n> +++ b/sha1-file.c\n> @@ -1375,6 +1375,7 @@ static void *read_object(const unsigned char *sha1, enum object_type *type,\n>\n>         if (oid_object_info_extended(the_repository, &oid, &oi, 0) < 0)\n>                 return NULL;\n> +       trace_printf(\"reading %s %s\", type_name(*type), sha1_to_hex(sha1));\n>         return content;\n>  }\n>\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index 66741130ae..cfdad4133d 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -1075,6 +1075,23 @@ static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, str\n>                                 o->cache_bottom += matches;\n>                                 return mask;\n>                         }\n> +               } else if (n == 2 && S_ISDIR(names->mode) &&\n> +                          names[0].mode == names[1].mode &&\n> +                          !strcmp(names[0].path, names[1].path) &&\n> +                          !oidcmp(names[0].oid, names[1].oid)\n> +                          /* && somehow account for modified on-disk files */) {\n> +                       int matches;\n> +\n> +                       /*\n> +                        * we know that the two trees have the same oid, so we\n> +                        * only need to look at one of them\n> +                        */\n> +                       matches = cache_tree_matches_traversal(o->src_index->cache_tree,\n> +                                                              names, info);\n> +                       if (matches) {\n> +                               o->cache_bottom += matches;\n> +                               return mask;\n> +                       }\n>                 }\n>\n>                 if (traverse_trees_recursive(n, dirmask, mask & ~dirmask,\n>\n> seems to avoid the tree reads when running \"GIT_TRACE=1 git checkout\".\n> It also totally empties the index. ;) So clearly we have to do a bit\n> more there. Probably rather than just bumping o->cache_bottom forward,\n> we'd need to actually move those entries into the new index. Or maybe\n> it's something else entirely (I did say cargo-culting, right?).\n\nAh this cache_bottom magic. I think this is Junio's alley ;-)\n\n> -Peff\n-- \nDuy\n"},{"id":"353531","messageId":"20180724212126.GB17803@sigill.intra.peff.net","threadId":"48913","inReplyTo":"20180724151336.GA1957@duynguyen.home","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-07-24T21:21:26Z","receivedAt":"2018-07-24T21:21:30Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jul 24, 2018 at 05:13:36PM +0200, Duy Nguyen wrote:\n\n> I guess you have the starting points you need after Jeff's and Junio's\n> explanation (and it would be great if cache-tree could actually be for\n> for this two-way merge). But to make it easier for new people in\n> future, maybe we should add this?\n> \n> This is basically a ripoff of Junio's explanation with starting points\n> (write-tree and index-format.txt). I wanted to incorporate some pieces\n> from Jeff's too but I think Junio's already covered it well.\n> \n> -- 8< --\n> Subject: [PATCH] cache-tree.h: more description of what it is and what's it used for\n\nThere is some discussion of this extension in\nDocumentation/technical/index-format.txt. But it's mostly the mechanical\nbits, not how or why you would use it.\n\nI like the idea of putting this explanation into the repo, though a lot\nof it is pretty specific to \"diff-index --cached\", which could\npotentially grow stale.\n\n-Peff\n"},{"id":"353570","messageId":"93bf2b44-fd05-cb39-cbf2-16a0736f0561@gmail.com","threadId":"48913","inReplyTo":"20180724151336.GA1957@duynguyen.home","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-25T16:09:43Z","receivedAt":"2018-07-25T16:09:49Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/24/2018 11:13 AM, Duy Nguyen wrote:\n> On Mon, Jul 23, 2018 at 04:51:38PM -0400, Ben Peart wrote:\n>>>>> What's the current state of the index before this checkout?\n>>>>\n>>>> This was after running \"git checkout\" multiple times so there was really\n>>>> nothing for git to do.\n>>>\n>>> Hmm.. this means cache-tree is fully valid, unless you have changes in\n>>> index. We're quite aggressive in repairing cache-tree since aecf567cbf\n>>> (cache-tree: create/update cache-tree on checkout - 2014-07-05). If we\n>>> have very good cache-tree records and still spend 33s on\n>>> traverse_trees, maybe there's something else.\n>>>\n>>\n>> I'm not at all familiar with the cache-tree and couldn't find any\n>> documentation on it other than index-format.txt which says \"it helps\n>> speed up tree object generation for a new commit.\"\n> \n> I guess you have the starting points you need after Jeff's and Junio's\n> explanation (and it would be great if cache-tree could actually be for\n> for this two-way merge). But to make it easier for new people in\n> future, maybe we should add this?\n> \n> This is basically a ripoff of Junio's explanation with starting points\n> (write-tree and index-format.txt). I wanted to incorporate some pieces\n> from Jeff's too but I think Junio's already covered it well.\n> \n\nI definitely like capturing this in the code or documentation somewhere. \n  Given I checked the header file for any hints on the design, I think \nthat is a reasonable place to put it.\n\n> -- 8< --\n> Subject: [PATCH] cache-tree.h: more description of what it is and what's it used for\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>   cache-tree.h | 29 +++++++++++++++++++++++++++++\n>   1 file changed, 29 insertions(+)\n> \n> diff --git a/cache-tree.h b/cache-tree.h\n> index cfd5328cc9..d25a800a72 100644\n> --- a/cache-tree.h\n> +++ b/cache-tree.h\n> @@ -5,6 +5,35 @@\n>   #include \"tree.h\"\n>   #include \"tree-walk.h\"\n>   \n> +/*\n> + * cache-tree is an index extension that records tree object names for\n> + * subdirectories you see in the index. It is mainly used for\n> + * generating trees from the index before you create a new commit (see\n> + * builtin/write-tree.c as starting point) but it's also used in \"git\n> + * diff-index --cached $TREE\" as an optimization. See index-format.txt\n> + * for on-disk format.\n> + *\n> + * Every time you write the contents of the index as a tree object, we\n\nI had to read this a couple of times to figure out what was meant by \n\"write the contents of the index as a tree object.\"  Maybe it was just \nme but how about something like:\n\n\"Every time you write a new tree object from the index you need to \ncollect the object name for each top-level path and write a new \ntop-level tree object out and then do the same recursively for any \nsubdirectory.\"\n\n> + * need to collect the object name for each top-level paths and write\n> + * a new top-level tree object out, after doing the same recursively\n> + * for any modified subdirectory. Whenever you add, remove or modify a\n> + * path in the index, the cache-tree entry for enclosing directories\n> + * are invalidated, so a cache-tree entry that is still valid means\n> + * that all the paths in the index under that directory match the\n> + * contents of the tree object that the cache-tree entry holds.\n> + *\n> + * And that property is used by \"diff-index --cached $TREE\" that is\n> + * run internally.  When we find that the subdirectory \"D\"'s\n> + * cache-tree entry is valid in the index, and the tree object\n> + * recorded in the cache-tree for that subdirectory matches the\n> + * subtree D in the tree object $TREE, then \"diff-index --cached\"\n> + * ignores the entire subdirectory D (which saves relatively little in\n> + * the index as it only needs to scan what is already in the memory\n> + * forward, but on the $TREE traversal side, it does not have to even\n> + * open a subtree, that can save a lot), and with a well-populated\n> + * cache-tree, it can save a significant processing.\n> + */\n> +\n>   struct cache_tree;\n>   struct cache_tree_sub {\n>   \tstruct cache_tree *cache_tree;\n> \n"},{"id":"353591","messageId":"0102d204-8be7-618a-69f4-9f924c4e6731@gmail.com","threadId":"48913","inReplyTo":"CACsJy8Du28jMyfdyhxpVxyw5+Xh+9eX==3x8YJSnmw6GAoRhTA@mail.gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-25T20:56:44Z","receivedAt":"2018-07-25T20:56:49Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/24/2018 11:33 AM, Duy Nguyen wrote:\n> On Tue, Jul 24, 2018 at 6:20 AM Jeff King <peff@peff.net> wrote:\n>> At least that's my view of it. unpack_trees() has always been a\n>> terrifying beast that I've avoided looking too closely at.\n> \n> /me nods on the terrifying part.\n> \n>>> After a quick look at the code, the only place I can find that tries to use\n>>> cache_tree_matches_traversal() is in unpack_callback() and that only happens\n>>> if n == 1 and in the \"git checkout\" case, n == 2. Am I missing something?\n> \n> So we do not actually use cache-tree? Big optimization opportunity (if\n> we can make it!).\n> \n\nI agree!  Assuming we can figure out the technical issues around using \nthe cache tree to optimize two way merges, another question I'm trying \nto answer is how we can enable this optimization without causing back \ncompat issues?\n\nWe're discussing detecting that there are no changes for parts of the \ntree between two commits but that isn't the only thing that can trigger \nchanges to be made to the index entries and working directory. Changes \ncan come from other inputs as well.\n\nOne example I am aware of is sparse-checkout.  If you made changes to \nyour sparse checkout settings or $GIT_DIR/info/sparse-checkout file, \nthat could trigger the need to update index entries and files in the \nworking directory.  Since that is a relatively rare occurrence, I can \nsee detecting changes to those settings/file and bypassing the \noptimization if there have been changes.  But are there other cases of \nthings that could cause unexpected changes in behavior?\n\nOne thought I had was to put the optimization behind a config setting so \nthat people had to opt-in to the difference in behavior.  I submitted a \ncanary patch [1] to test out how receptive people would be to that idea. \n  Hopefully I can get some feedback on that aspect of the patch.\n\n[1] \nhttps://public-inbox.org/git/ab8ee481-54fa-a014-69d9-8f621b136766@gmail.com/T/#m2a425a23df5e064a79b0a72537a5dd6ccba3b07b\n\n>> Looks like it's trying to special-case \"diff-index --cached\". Which\n>> kind-of makes sense. In the non-cached case, we're thinking not only\n>> about the relationship between the index and the tree, but also whether\n>> the on-disk files are up to date.\n>>\n>> And that would be the same for checkout. We want to know not only\n>> whether there are changes to make to the index, but also whether the\n>> on-disk files need to be updated from the index.\n>>\n>> But I assume in your case that we've just refreshed the index quickly\n>> using fsmonitor. So I think in the long run what you want is:\n>>\n>>    1. fsmonitor tells us which index entries are not clean\n>>\n>>    2. based on the unclean list, we invalidate cache-tree entries for\n>>       those paths\n>>\n>>    3. if we have a valid cache-tree entry, we should be able to skip\n>>       digging into that tree; if not, then we walk the index and tree as\n>>       normal, adding/deleting index entries and updating (or complaining\n>>       about) modified on-disk files\n> \n> If you tie this optimization to twoway_merge specifically (by checking\n> \"fn\" field), then I think we can do it even better. Since\n> cache_tree_matches_traversal() is one (hopefully not too costly)\n> lookup, we can do it without checking with fsmonitor or whatever and\n> only do so when we have found a cache tree.\n> \n> Then if we write this new special code just for twoway_merge, we need\n> to tighten the checks a bit. I think in this case twoway_merge() will\n> be called with \"oldtree\" as same as \"newtree\" (and \"current\" may\n> contains dirty stuff from the index). Then\n> \n>   - o->df_conflict_entry should be NULL (because we handle it slightly\n> differently in twoway_merge)\n>   - \"current\" should not have CE_CONFLICTED\n> \n> then I believe we will fall into case /* 20 or 21 */ where\n> merged_entry() is suppoed to be called on all entries and it would\n> change nothing in the index since newtree is the same as oldtree, and\n> we could just jump over the whole tree in traverse_trees().\n> \n\nI'm fine with tying specific optimizations to twoway_merge as that is a \nvery common (if not the most common) merge.\n\nI'm still very new to this part of the code so am trying to figure out \nwhat you're suggesting.  I've read your description a few times and what \nI'm getting out of it is that with some additional checks (ie verify \nit's a twoway_merge, df_conflict_entry, not CE_CONFLICTED) that we \nshould be able to skip the whole tree similar to how Peff demonstrated \nbelow without having to invalidate the cache tree to reflect modified \non-disk files.  Is that correct or am I missing something?\n\n>> I think the \"n\" adds an extra layer of complexity. n==2 means we're\n>> doing a \"2-way\" merge. Moving from tree X to tree Y, and dealing with\n>> the index as we go. Naively I _think_ we'd be OK to just extend the rule\n>> to \"if both subtrees match each other _and_ match the valid cache-tree,\n>> then we can skip\".\n>>\n>> Again, I'm a little out of my area of expertise here, but cargo-culting\n>> like this:\n>>\n>> diff --git a/sha1-file.c b/sha1-file.c\n>> index de4839e634..c105af70ce 100644\n>> --- a/sha1-file.c\n>> +++ b/sha1-file.c\n>> @@ -1375,6 +1375,7 @@ static void *read_object(const unsigned char *sha1, enum object_type *type,\n>>\n>>          if (oid_object_info_extended(the_repository, &oid, &oi, 0) < 0)\n>>                  return NULL;\n>> +       trace_printf(\"reading %s %s\", type_name(*type), sha1_to_hex(sha1));\n>>          return content;\n>>   }\n>>\n>> diff --git a/unpack-trees.c b/unpack-trees.c\n>> index 66741130ae..cfdad4133d 100644\n>> --- a/unpack-trees.c\n>> +++ b/unpack-trees.c\n>> @@ -1075,6 +1075,23 @@ static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, str\n>>                                  o->cache_bottom += matches;\n>>                                  return mask;\n>>                          }\n>> +               } else if (n == 2 && S_ISDIR(names->mode) &&\n>> +                          names[0].mode == names[1].mode &&\n>> +                          !strcmp(names[0].path, names[1].path) &&\n>> +                          !oidcmp(names[0].oid, names[1].oid)\n>> +                          /* && somehow account for modified on-disk files */) {\n>> +                       int matches;\n>> +\n>> +                       /*\n>> +                        * we know that the two trees have the same oid, so we\n>> +                        * only need to look at one of them\n>> +                        */\n>> +                       matches = cache_tree_matches_traversal(o->src_index->cache_tree,\n>> +                                                              names, info);\n>> +                       if (matches) {\n>> +                               o->cache_bottom += matches;\n>> +                               return mask;\n>> +                       }\n>>                  }\n>>\n>>                  if (traverse_trees_recursive(n, dirmask, mask & ~dirmask,\n>>\n>> seems to avoid the tree reads when running \"GIT_TRACE=1 git checkout\".\n>> It also totally empties the index. ;) So clearly we have to do a bit\n>> more there. Probably rather than just bumping o->cache_bottom forward,\n>> we'd need to actually move those entries into the new index. Or maybe\n>> it's something else entirely (I did say cargo-culting, right?).\n> \n> Ah this cache_bottom magic. I think this is Junio's alley ;-)\n> \n>> -Peff\n"},{"id":"353612","messageId":"CACsJy8AWcHVYNBZGRUTdcg8FmwOGz3MSUHH+3uVSGrg6MMZMng@mail.gmail.com","threadId":"48913","inReplyTo":"0102d204-8be7-618a-69f4-9f924c4e6731@gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-26T05:30:20Z","receivedAt":"2018-07-26T05:30:50Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Jul 25, 2018 at 10:56 PM Ben Peart <peartben@gmail.com> wrote:\n> I'm still very new to this part of the code so am trying to figure out\n> what you're suggesting.  I've read your description a few times and what\n> I'm getting out of it is that with some additional checks (ie verify\n> it's a twoway_merge, df_conflict_entry, not CE_CONFLICTED) that we\n> should be able to skip the whole tree similar to how Peff demonstrated\n> below without having to invalidate the cache tree to reflect modified\n> on-disk files.  Is that correct or am I missing something?\n\nAnd I didn't give you an easy time because I was not very clear in my\nsuggestion, I think. So let's start again. But first let's start with\na potentially more generic optimization using cache-tree that I\nnoticed just now.\n\nYou now know traverse_trees() is used to walk N trees and the index at\nthe same time. Cache tree is also used to quickly check if a big chunk\nof the index matches some tree object. So what if we try to avoid tree\nobjects if possible (which reduces I/O, object inflation and tree\nparsing cost)? Let's say we're walking two trees X and Y, then we\nnotice through cache-tree that X is the same in the index. Then\ninstead of walking the actual X, you could just get the same entry\nfrom the index and make it \"X\". This way you only need to walk Y and\nthe index (until the shared tree ends of course). If Y happens to\nmatch cache-tree too, all the better!\n\nLet's get back to two-way merge. I suggest you read the two-way merge\nin git-read-tree.txt. That table could give you a pretty good idea\nwhat's going on. twoway_merge() will be given a tuple of three entries\n(I, H, M) of the same path name, for every path. I think what we need\nis determine the condition where the outcome is known in advance, so\nthat we can just skip walking the index for one directory. One of the\nchecks we could do quickly is I==M or I==H (using cache-tree) and H==M\n(using tree hash).\n\nThe first obvious cases that we can optimize are\n\nclean (H==M)\n       ------\n     14 yes                 exists   exists   keep index\n     15 no                  exists   exists   keep index\n\nIn other words if we know H==M, there's no much we need to do since\nwe're keeping the index the same. But you don't really know how many\nentries are in this directory where H==M. You would need cache-tree\nfor that, so in reality it's I==H==M.\n\nThe \"clean\" column is what fsmonitor comes in, though I'm not sure if\nit's actually needed. I haven't checked how '-u' flag works.\n\nThere's two other cases that we can also optimize, though I think it's\nless likely to happen:\n\n        clean I==H  I==M (H!=M)\n       ------------------\n     18 yes   no    yes     exists   exists   keep index\n     19 no    no    yes     exists   exists   keep index\n\nSome other cases where I==H can benefit from the generic tree walk\noptimization above since we can skip parsing H.\n-- \nDuy\n"},{"id":"353655","messageId":"20180726163049.GA15572@duynguyen.home","threadId":"48913","inReplyTo":"CACsJy8AWcHVYNBZGRUTdcg8FmwOGz3MSUHH+3uVSGrg6MMZMng@mail.gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-26T16:30:49Z","receivedAt":"2018-07-26T16:30:58Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Jul 26, 2018 at 07:30:20AM +0200, Duy Nguyen wrote:\n> Let's get back to two-way merge. I suggest you read the two-way merge\n> in git-read-tree.txt. That table could give you a pretty good idea\n> what's going on. twoway_merge() will be given a tuple of three entries\n> (I, H, M) of the same path name, for every path. I think what we need\n> is determine the condition where the outcome is known in advance, so\n> that we can just skip walking the index for one directory. One of the\n> checks we could do quickly is I==M or I==H (using cache-tree) and H==M\n> (using tree hash).\n> \n> The first obvious cases that we can optimize are\n> \n> clean (H==M)\n>        ------\n>      14 yes                 exists   exists   keep index\n>      15 no                  exists   exists   keep index\n> \n> In other words if we know H==M, there's no much we need to do since\n> we're keeping the index the same. But you don't really know how many\n> entries are in this directory where H==M. You would need cache-tree\n> for that, so in reality it's I==H==M.\n> \n> The \"clean\" column is what fsmonitor comes in, though I'm not sure if\n> it's actually needed. I haven't checked how '-u' flag works.\n> \n> There's two other cases that we can also optimize, though I think it's\n> less likely to happen:\n> \n>         clean I==H  I==M (H!=M)\n>        ------------------\n>      18 yes   no    yes     exists   exists   keep index\n>      19 no    no    yes     exists   exists   keep index\n> \n> Some other cases where I==H can benefit from the generic tree walk\n> optimization above since we can skip parsing H.\n\nI'm excited so I decided to try out anyway. This is what I've come up\nwith. Switching trees on git.git shows it could skip plenty entries,\nso promising. It's ugly and it fails at t6020 though, there's still\nwork ahead. But I think it'll stop here.\n\nA few notes after getting my hands dirty\n\n- one big difference between diff --cached and checkout is, diff is a\n  read-only operation while checkout actually creates new index.  One\n  of the side effect is that cache-tree may be destroyed while we're\n  walking the trees, i'm not so sure.\n\n- I don't think we even need to a special twoway_merge_same()\n  here. That function could just call twoway_merge() with the right\n  \"src\" parameter and the outcome should still be the same. Which\n  means it'll work for threeway merge too.\n\n- i'm still scared of that cache_bottom switching to death. no idea\n  how it works or if i broke anything by changing the condition there.\n\n-- 8< --\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex 28627650cd..276712af64 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -515,6 +515,7 @@ static int merge_working_tree(const struct checkout_opts *opts,\n \t\ttopts.gently = opts->merge && old_branch_info->commit;\n \t\ttopts.verbose_update = opts->show_progress;\n \t\ttopts.fn = twoway_merge;\n+\t\ttopts.fn_same = twoway_merge_same;\n \t\tif (opts->overwrite_ignore) {\n \t\t\ttopts.dir = xcalloc(1, sizeof(*topts.dir));\n \t\t\ttopts.dir->flags |= DIR_SHOW_IGNORED;\ndiff --git a/diff-lib.c b/diff-lib.c\nindex a9f38eb5a3..48e6c4ab0d 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -485,6 +485,15 @@ static int oneway_diff(const struct cache_entry * const *src,\n \treturn 0;\n }\n \n+static int oneway_diff_cached(int pos, int nr, struct unpack_trees_options *options)\n+{\n+\t/*\n+\t * Nothing to do. Unpack-trees can safely skip the whole\n+\t * nr_matches cache entries.\n+\t */\n+\treturn 0;\n+}\n+\n static int diff_cache(struct rev_info *revs,\n \t\t      const struct object_id *tree_oid,\n \t\t      const char *tree_name,\n@@ -501,8 +510,8 @@ static int diff_cache(struct rev_info *revs,\n \tmemset(&opts, 0, sizeof(opts));\n \topts.head_idx = 1;\n \topts.index_only = cached;\n-\topts.diff_index_cached = (cached &&\n-\t\t\t\t  !revs->diffopt.flags.find_copies_harder);\n+\tif (cached && !revs->diffopt.flags.find_copies_harder)\n+\t\topts.fn_same = oneway_diff_cached;\n \topts.merge = 1;\n \topts.fn = oneway_diff;\n \topts.unpack_data = revs;\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 66741130ae..01e3f38807 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -615,7 +615,7 @@ static void restore_cache_bottom(struct traverse_info *info, int bottom)\n {\n \tstruct unpack_trees_options *o = info->data;\n \n-\tif (o->diff_index_cached)\n+\tif (o->fn_same)\n \t\treturn;\n \to->cache_bottom = bottom;\n }\n@@ -625,7 +625,7 @@ static int switch_cache_bottom(struct traverse_info *info)\n \tstruct unpack_trees_options *o = info->data;\n \tint ret, pos;\n \n-\tif (o->diff_index_cached)\n+\tif (o->fn_same)\n \t\treturn 0;\n \tret = o->cache_bottom;\n \tpos = find_cache_pos(info->prev, &info->name);\n@@ -996,6 +996,43 @@ static void debug_unpack_callback(int n,\n \t\tdebug_name_entry(i, names + i);\n }\n \n+static int skip_dir(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i, matches;\n+\tint len;\n+\tchar *name;\n+\tint pos;\n+\n+\tif (dirmask != ((1 << n) - 1) || !S_ISDIR(names->mode))\n+\t\treturn 0;\n+\n+\tfor (i = 1; i < n; i++)\n+\t\tif (oidcmp(names[0].oid, names[i].oid))\n+\t\t\treturn 0;\n+\n+\tmatches = cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n+\tif (!matches)\n+\t\treturn 0;\n+\n+\t/*\n+\t * Everything under the name matches; skip the entire\n+\t * hierarchy. fn_same must special cases D/F conflicts in such\n+\t * a way that it does not do any look-ahead, so this is safe.\n+\t */\n+\tlen = traverse_path_len(info, names);\n+\tname = xmalloc(len + 1);\n+\n+\tmake_traverse_path(name, info, names);\n+\tpos = index_name_pos(o->src_index, name, len);\n+\tif (pos >= 0)\n+\t\tdie(\"NOOO\");\n+\ttrace_printf(\"dirmask = %lx, path = %s\\n\", dirmask, name);\n+\tif (o->fn_same(-pos-1, matches, o))\n+\t\tmatches = 0;\n+\tfree(name);\n+\treturn matches;\n+}\n static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n@@ -1015,7 +1052,7 @@ static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, str\n \t\t\tint cmp;\n \t\t\tstruct cache_entry *ce;\n \n-\t\t\tif (o->diff_index_cached)\n+\t\t\tif (o->fn_same)\n \t\t\t\tce = next_cache_entry(o);\n \t\t\telse\n \t\t\t\tce = find_cache_entry(info, p);\n@@ -1059,18 +1096,8 @@ static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, str\n \n \t/* Now handle any directories.. */\n \tif (dirmask) {\n-\t\t/* special case: \"diff-index --cached\" looking at a tree */\n-\t\tif (o->diff_index_cached &&\n-\t\t    n == 1 && dirmask == 1 && S_ISDIR(names->mode)) {\n-\t\t\tint matches;\n-\t\t\tmatches = cache_tree_matches_traversal(o->src_index->cache_tree,\n-\t\t\t\t\t\t\t       names, info);\n-\t\t\t/*\n-\t\t\t * Everything under the name matches; skip the\n-\t\t\t * entire hierarchy.  diff_index_cached codepath\n-\t\t\t * special cases D/F conflicts in such a way that\n-\t\t\t * it does not do any look-ahead, so this is safe.\n-\t\t\t */\n+\t\tif (o->fn_same) {\n+\t\t\tint matches = skip_dir(n, mask, dirmask, names, info);\n \t\t\tif (matches) {\n \t\t\t\to->cache_bottom += matches;\n \t\t\t\treturn mask;\n@@ -1881,6 +1908,7 @@ static int deleted_entry(const struct cache_entry *ce,\n static int keep_entry(const struct cache_entry *ce,\n \t\t      struct unpack_trees_options *o)\n {\n+\ttrace_printf(\"keep_entry(%s)\\n\", ce->name);\n \tadd_entry(o, ce, 0, 0);\n \treturn 1;\n }\n@@ -2132,6 +2160,25 @@ int twoway_merge(const struct cache_entry * const *src,\n \treturn deleted_entry(oldtree, current, o);\n }\n \n+int twoway_merge_same(int pos, int nr, struct unpack_trees_options *o)\n+{\n+\tint i;\n+\n+\t/*\n+\t * Since cache-tree at \"src\" exists, it means there's no\n+\t * staged entries here (they would have invalidated cache-tree\n+\t * otherwise). So no CE_CONFLICTED.\n+\t *\n+\t * And because I==H==M, we can't run into d/f conflicts\n+\t * either: for every path name, we will always find a _file_\n+\t * in the index as well as the two other trees.\n+\t */\n+\ttrace_printf(\"Skipping %d entries\\n\", nr);\n+\tfor (i = 0; i < nr; i++)\n+\t\tkeep_entry(o->src_index->cache[pos + i], o);\n+\treturn 0;\n+}\n+\n /*\n  * Bind merge.\n  *\ndiff --git a/unpack-trees.h b/unpack-trees.h\nindex c2b434c606..45c69e2ed0 100644\n--- a/unpack-trees.h\n+++ b/unpack-trees.h\n@@ -12,6 +12,9 @@ struct exclude_list;\n typedef int (*merge_fn_t)(const struct cache_entry * const *src,\n \t\tstruct unpack_trees_options *options);\n \n+typedef int (*merge_same_fn_t)(int pos, int nr,\n+\t\t\t       struct unpack_trees_options *options);\n+\n enum unpack_trees_error_types {\n \tERROR_WOULD_OVERWRITE = 0,\n \tERROR_NOT_UPTODATE_FILE,\n@@ -49,7 +52,6 @@ struct unpack_trees_options {\n \t\t     aggressive,\n \t\t     skip_unmerged,\n \t\t     initial_checkout,\n-\t\t     diff_index_cached,\n \t\t     debug_unpack,\n \t\t     skip_sparse_checkout,\n \t\t     gently,\n@@ -61,6 +63,7 @@ struct unpack_trees_options {\n \tstruct dir_struct *dir;\n \tstruct pathspec *pathspec;\n \tmerge_fn_t fn;\n+\tmerge_same_fn_t fn_same;\n \tconst char *msgs[NB_UNPACK_TREES_ERROR_TYPES];\n \tstruct argv_array msgs_to_free;\n \t/*\n@@ -92,6 +95,8 @@ int threeway_merge(const struct cache_entry * const *stages,\n \t\t   struct unpack_trees_options *o);\n int twoway_merge(const struct cache_entry * const *src,\n \t\t struct unpack_trees_options *o);\n+int twoway_merge_same(int pos, int nr,\n+\t\t      struct unpack_trees_options *o);\n int bind_merge(const struct cache_entry * const *src,\n \t       struct unpack_trees_options *o);\n int oneway_merge(const struct cache_entry * const *src,\n-- 8< --\n"},{"id":"353656","messageId":"CACsJy8Af4K=XOcGaZxhbWOD7OYgtW7bXLved_Xv-XgVM_6AQwQ@mail.gmail.com","threadId":"48913","inReplyTo":"0102d204-8be7-618a-69f4-9f924c4e6731@gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-26T16:35:19Z","receivedAt":"2018-07-26T16:35:47Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Jul 25, 2018 at 10:56 PM Ben Peart <peartben@gmail.com> wrote:\n>\n>\n>\n> On 7/24/2018 11:33 AM, Duy Nguyen wrote:\n> > On Tue, Jul 24, 2018 at 6:20 AM Jeff King <peff@peff.net> wrote:\n> >> At least that's my view of it. unpack_trees() has always been a\n> >> terrifying beast that I've avoided looking too closely at.\n> >\n> > /me nods on the terrifying part.\n> >\n> >>> After a quick look at the code, the only place I can find that tries to use\n> >>> cache_tree_matches_traversal() is in unpack_callback() and that only happens\n> >>> if n == 1 and in the \"git checkout\" case, n == 2. Am I missing something?\n> >\n> > So we do not actually use cache-tree? Big optimization opportunity (if\n> > we can make it!).\n> >\n>\n> I agree!  Assuming we can figure out the technical issues around using\n> the cache tree to optimize two way merges, another question I'm trying\n> to answer is how we can enable this optimization without causing back\n> compat issues?\n\nIf it works as I expect, then there's no compat issues at all (exactly\nlike the diff_index_cached optimization we already have). We simply\nfind a safe shortcut that does not add any side effects.\n-- \nDuy\n"},{"id":"353675","messageId":"xmqqd0v9pyzu.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"20180726163049.GA15572@duynguyen.home","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-26T19:40:05Z","receivedAt":"2018-07-26T19:40:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> I'm excited so I decided to try out anyway. This is what I've come up\n> with. Switching trees on git.git shows it could skip plenty entries,\n> so promising. It's ugly and it fails at t6020 though, there's still\n> work ahead. But I think it'll stop here.\n\nWe are extremely shallow compared to projects like the kernel and\nstuff from java land, so that is quite an interesting find.\n\n"},{"id":"353720","messageId":"20180727154241.GA21288@duynguyen.home","threadId":"48913","inReplyTo":"xmqqd0v9pyzu.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-27T15:42:41Z","receivedAt":"2018-07-27T15:42:49Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Jul 26, 2018 at 12:40:05PM -0700, Junio C Hamano wrote:\n> Duy Nguyen <pclouds@gmail.com> writes:\n> \n> > I'm excited so I decided to try out anyway. This is what I've come up\n> > with. Switching trees on git.git shows it could skip plenty entries,\n> > so promising. It's ugly and it fails at t6020 though, there's still\n> > work ahead. But I think it'll stop here.\n> \n> We are extremely shallow compared to projects like the kernel and\n> stuff from java land, so that is quite an interesting find.\n> \n\nYeah. I've got a more or less complete patch now with full test suite\npassed and even with linux.git, the numbers look pretty good.\n\nBen, is it possible for you to try this one out? I don't suppose it\nwill be that good on a real big repo. But I'm curious how much faster\ncould this patch does.\n\nI'm quite happy that I don't have to make specific code for twoway\nmerge, which means this patch would also help real merges (3way)\ntoo. Interestingly this also helps reduce traverse_trees() when\ndiff_index_cached optimization is on. I have no idea how but\nwell.. can't complain.\n\n-- 8< --\nSubject: [PATCH] unpack-trees: optimize walking same trees with cache-tree\n\nIn order to merge one or many trees with the index, unpack-trees code\nwalk multiple trees in parallel with the index and perform n-way\nmerge. If we find out at start of a directory that all trees are the\nsame (by comparing OID) and cache-tree happens to be available for\nthat directory as well, we could avoid walking the trees.\n\nOne nice attribute of cache-tree (and the index) is that the tree is\nonly flattened (and it's called \"the index\") and we know how many\nfiles that directory has. With this information, we could avoid\naccessing object database to walk tree objects and just take the\nentries from the index instead.\n\nThe upside is of course a lot less I/O since we can potentially skip\nlots of trees (think subtrees). We also save CPU because we don't have\nto inflate and the apply deltas. The downside is of course more\nfragile code since the logic in some functions are now duplicated\nelsewhere.\n\nWIth this patch, switching between two trees on linux.git where\nthere's only one file changed (toplevel Makefile) seems sped up pretty\ngood. Total checkout time goes down from 0.543 to 0.352 (35%).\ntraverse_trees() one twoway merge (the big one in unpack_trees()) goes\nfrom 0.157s to 0.036 (70%).\n\nNote that compared to diff_index_cached optimization (which is very\nsimilar to this) we do more work here. This is because diff_index_cached\nonly cares about side effect, it does not modify the index, so we can\nquickly jump through a big chunk of cache entries. For n-way merge, we\nneed to add entries and verify stuff, so more CPU cycles.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n unpack-trees.c | 125 +++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 125 insertions(+)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 66741130ae..9c791b55b2 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -642,6 +642,110 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n \treturn name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n }\n \n+static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n+\t\t\t\t\tstruct name_entry *names,\n+\t\t\t\t\tstruct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i;\n+\n+\tif (dirmask != ((1 << n) - 1) || !S_ISDIR(names->mode) || !o->merge)\n+\t\treturn 0;\n+\n+\tfor (i = 1; i < n; i++)\n+\t\tif (!are_same_oid(names, names + i))\n+\t\t\treturn 0;\n+\n+\treturn cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n+}\n+\n+/*\n+ * Fast path if we detect that all trees are the same as cache-tree at this\n+ * path. We'll walk these trees recursively using cache-tree/index instead of\n+ * ODB since already know what these trees contain.\n+ */\n+static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n+\t\t\t\t  struct name_entry *names,\n+\t\t\t\t  struct traverse_info *info)\n+{\n+\tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i, d;\n+\n+\t/*\n+\t * Do what unpack_callback() and unpack_nondirectories() normally\n+\t * do. But we do it in one function call (for even nested trees)\n+\t * instead.\n+\t *\n+\t * D/F conflicts and staged entries are not a concern because cache-tree\n+\t * would be invalidated and we would never get here in the first place.\n+\t */\n+\tfor (i = 0; i < nr_entries; i++) {\n+\t\tstruct cache_entry *tree_ce;\n+\t\tint len, rc;\n+\n+\t\tsrc[0] = o->src_index->cache[pos + i];\n+\n+\t\t/* Do what unpack_nondirectories() normally does */\n+\t\tlen = ce_namelen(src[0]);\n+\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n+\n+\t\ttree_ce->ce_mode = src[0]->ce_mode;\n+\t\ttree_ce->ce_flags = create_ce_flags(0);\n+\t\ttree_ce->ce_namelen = len;\n+\t\toidcpy(&tree_ce->oid, &src[0]->oid);\n+\t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n+\n+\t\tfor (d = 1; d <= nr_names; d++)\n+\t\t\tsrc[d] = tree_ce;\n+\n+\t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n+\t\tfree(tree_ce);\n+\t\tif (rc < 0)\n+\t\t\treturn rc;\n+\n+\t\tmark_ce_used(src[0], o);\n+\t}\n+\ttrace_printf(\"Quick traverse over %d entries from %s to %s\\n\",\n+\t\t     nr_entries,\n+\t\t     o->src_index->cache[pos]->name,\n+\t\t     o->src_index->cache[pos + nr_entries - 1]->name);\n+\treturn 0;\n+}\n+\n+static int index_pos_by_traverse_info(struct name_entry *names,\n+\t\t\t\t      struct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint len = traverse_path_len(info, names);\n+\tchar *name = xmalloc(len + 1);\n+\tint pos;\n+\n+\tmake_traverse_path(name, info, names);\n+\tpos = index_name_pos(o->src_index, name, len);\n+\tif (pos >= 0)\n+\t\tBUG(\"This is so wrong. This is a directory and should not exist in index\");\n+\tpos = -pos - 1;\n+\t/*\n+\t * There's no guarantee that pos points to the first entry of the\n+\t * directory. If the directory name is \"letters\" and there's another\n+\t * file named \"letters.txt\" in the index, pos will point to that file\n+\t * instead.\n+\t */\n+\twhile (pos < o->src_index->cache_nr) {\n+\t\tconst struct cache_entry *ce = o->src_index->cache[pos];\n+\t\tif (ce_namelen(ce) > len &&\n+\t\t    ce->name[len] == '/' &&\n+\t\t    !memcmp(ce->name, name, len))\n+\t\t\tbreak;\n+\t\tpos++;\n+\t}\n+\tif (pos == o->src_index->cache_nr)\n+\t\tBUG(\"This is still wrong\");\n+\tfree(name);\n+\treturn pos;\n+}\n+\n static int traverse_trees_recursive(int n, unsigned long dirmask,\n \t\t\t\t    unsigned long df_conflicts,\n \t\t\t\t    struct name_entry *names,\n@@ -653,6 +757,17 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n \tvoid *buf[MAX_UNPACK_TREES];\n \tstruct traverse_info newinfo;\n \tstruct name_entry *p;\n+\tint nr_entries;\n+\n+\tnr_entries = all_trees_same_as_cache_tree(n, dirmask, names, info);\n+\tif (nr_entries > 0) {\n+\t\tstruct unpack_trees_options *o = info->data;\n+\t\tint pos = index_pos_by_traverse_info(names, info);\n+\n+\t\tif (!o->merge || df_conflicts)\n+\t\t\tBUG(\"Wrong condition to get here buddy\");\n+\t\treturn traverse_by_cache_tree(pos, nr_entries, n, names, info);\n+\t}\n \n \tp = names;\n \twhile (!p->mode)\n@@ -812,6 +927,11 @@ static struct cache_entry *create_ce_entry(const struct traverse_info *info, con\n \treturn ce;\n }\n \n+/*\n+ * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n+ * without actually calling it. If you change the logic here you may need to\n+ * check and change there as well.\n+ */\n static int unpack_nondirectories(int n, unsigned long mask,\n \t\t\t\t unsigned long dirmask,\n \t\t\t\t struct cache_entry **src,\n@@ -996,6 +1116,11 @@ static void debug_unpack_callback(int n,\n \t\tdebug_name_entry(i, names + i);\n }\n \n+/*\n+ * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n+ * without actually calling it. If you change the logic here you may need to\n+ * check and change there as well.\n+ */\n static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n-- \n2.18.0.656.gda699b98b3\n\n-- 8< --\n--\nDuy\n"},{"id":"353723","messageId":"ab7338a4-4a63-c722-35b1-a5b0784d66bf@gmail.com","threadId":"48913","inReplyTo":"xmqqd0v9pyzu.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-27T15:50:56Z","receivedAt":"2018-07-27T15:51:01Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/26/2018 3:40 PM, Junio C Hamano wrote:\n> Duy Nguyen <pclouds@gmail.com> writes:\n> \n>> I'm excited so I decided to try out anyway. This is what I've come up\n>> with. Switching trees on git.git shows it could skip plenty entries,\n>> so promising. It's ugly and it fails at t6020 though, there's still\n>> work ahead. But I think it'll stop here.\n> \n> We are extremely shallow compared to projects like the kernel and\n> stuff from java land, so that is quite an interesting find.\n> \n\nI had a few minutes so applied this patch to the latest git for windows \nand ran the p0006-read-tree-checkout.sh perf test on the git repo as \nwell as a synthetic large repo.  The results look quite promising - up \nto a 28.8% savings!\n\nI'm out of time this week but am _very_ interested in seeing if this can \nbe completed successfully.\n\nBen\n\n\ngit repo results\n================\n\nTest                                                            this \ntree          gfw\n---------------------------------------------------------------------------------------------------------\n0006.2: read-tree br_base br_ballast (1000001) \n1.37(0.04+0.09)    1.34(0.03+0.09) -2.2%\n0006.3: switch between br_base br_ballast (1000001) \n50.21(0.07+0.09)   50.22(0.03+0.09) +0.0%\n0006.4: switch between br_ballast br_ballast_plus_1 (1000001) \n3.58(0.03+0.09)    4.61(0.03+0.10) +28.8%\n0006.5: switch between aliases (1000001) \n3.67(0.03+0.07)    4.56(0.01+0.07) +24.3%\n\n\nlarge synthetic repo results\n============================\n\nTest                                                            this \ntree          gfw\n---------------------------------------------------------------------------------------------------------\n0006.2: read-tree br_base br_ballast (1000001) \n1.33(0.04+0.04)    1.33(0.04+0.06) +0.0%\n0006.3: switch between br_base br_ballast (1000001) \n48.96(0.03+0.12)   50.76(0.03+0.07) +3.7%\n0006.4: switch between br_ballast br_ballast_plus_1 (1000001) \n3.64(0.01+0.09)    4.59(0.06+0.07) +26.1%\n0006.5: switch between aliases (1000001) \n3.68(0.03+0.07)    4.66(0.04+0.06) +26.6%\n"},{"id":"353728","messageId":"434074a8-1045-8c8f-da0c-873436acf40e@gmail.com","threadId":"48913","inReplyTo":"20180727154241.GA21288@duynguyen.home","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-27T16:22:18Z","receivedAt":"2018-07-27T16:22:22Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/27/2018 11:42 AM, Duy Nguyen wrote:\n> On Thu, Jul 26, 2018 at 12:40:05PM -0700, Junio C Hamano wrote:\n>> Duy Nguyen <pclouds@gmail.com> writes:\n>>\n>>> I'm excited so I decided to try out anyway. This is what I've come up\n>>> with. Switching trees on git.git shows it could skip plenty entries,\n>>> so promising. It's ugly and it fails at t6020 though, there's still\n>>> work ahead. But I think it'll stop here.\n>>\n>> We are extremely shallow compared to projects like the kernel and\n>> stuff from java land, so that is quite an interesting find.\n>>\n> \n> Yeah. I've got a more or less complete patch now with full test suite\n> passed and even with linux.git, the numbers look pretty good.\n> \n> Ben, is it possible for you to try this one out? I don't suppose it\n> will be that good on a real big repo. But I'm curious how much faster\n> could this patch does.\n> \n\nThanks Duy.  I'm super excited about this so did a quick and dirty \nmanual perf test.\n\nI ran \"git checkout\" 5 times, discarded the first 2 runs and averaged \nthe last 3 with and without this patch on top of VFSForGit in a large repo.\n\nWithout this patch average times were 16.97\nWith this patch average times were 10.55\n\nThat is a significant improvement!\n\nI really have to run but I'll be back next week to dig in more.\n\nBen\n"},{"id":"353729","messageId":"xmqqpnz8ob2x.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"20180727154241.GA21288@duynguyen.home","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-27T17:14:14Z","receivedAt":"2018-07-27T17:14:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index 66741130ae..9c791b55b2 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -642,6 +642,110 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n>  \treturn name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n>  }\n>  \n> +static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n> +\t\t\t\t\tstruct name_entry *names,\n> +\t\t\t\t\tstruct traverse_info *info)\n> +{\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint i;\n> +\n> +\tif (dirmask != ((1 << n) - 1) || !S_ISDIR(names->mode) || !o->merge)\n> +\t\treturn 0;\n\nIn other words, punt if (1) not all are directories, (2) the first\nname entry given by the caller in names[] is not ISDIR(), or (3) we\nare not merging i.e. not \"Are we supposed to look at the index too?\"\nin unpack_callback().\n\nI am not sure if the second one is doing us any good.  When\nS_ISDIR(names->mode) is not true, then the bit in dirmask that\ncorresponds to the one in the entry[] traverse_trees() filled and\npassed to us must be zero, so the dirmask check would reject such a\ncase anyway, no?\n\nI would have moved !o->merge to the front, not for performance\nreasons but to make it clear that this function helps an\noptimization that matters only when we are walking tree(s) together\nwith the index.\n\n> +\tfor (i = 1; i < n; i++)\n> +\t\tif (!are_same_oid(names, names + i))\n> +\t\t\treturn 0;\n> +\n> +\treturn cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n> +}\n> +\n> +/*\n> + * Fast path if we detect that all trees are the same as cache-tree at this\n> + * path. We'll walk these trees recursively using cache-tree/index instead of\n> + * ODB since already know what these trees contain.\n> + */\n> +static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n> +\t\t\t\t  struct name_entry *names,\n> +\t\t\t\t  struct traverse_info *info)\n> +{\n> +\tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint i, d;\n> +\n> +\t/*\n> +\t * Do what unpack_callback() and unpack_nondirectories() normally\n> +\t * do. But we do it in one function call (for even nested trees)\n> +\t * instead.\n> +\t *\n> +\t * D/F conflicts and staged entries are not a concern because cache-tree\n> +\t * would be invalidated and we would never get here in the first place.\n> +\t */\n\nWe want to at least have\n\n\tif (!o->merge || ARRAY_SIZE(src) <= nr_names)\n\t\tBUG(\"\");\n\nhere, I'd think.\n\n> +\tfor (i = 0; i < nr_entries; i++) {\n> +\t\tstruct cache_entry *tree_ce;\n> +\t\tint len, rc;\n> +\n> +\t\tsrc[0] = o->src_index->cache[pos + i];\n> +\n> +\t\t/* Do what unpack_nondirectories() normally does */\n> +\t\tlen = ce_namelen(src[0]);\n> +\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n\nunpack_nondirectories() uses create_ce_entry() here.  Any reason why\nwe shouldn't use it and tell it to make a transient one?\n\n> +\t\ttree_ce->ce_mode = src[0]->ce_mode;\n> +\t\ttree_ce->ce_flags = create_ce_flags(0);\n> +\t\ttree_ce->ce_namelen = len;\n> +\t\toidcpy(&tree_ce->oid, &src[0]->oid);\n> +\t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n> +\n> +\t\tfor (d = 1; d <= nr_names; d++)\n> +\t\t\tsrc[d] = tree_ce;\n> +\n> +\t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n> +\t\tfree(tree_ce);\n> +\t\tif (rc < 0)\n> +\t\t\treturn rc;\n> +\n> +\t\tmark_ce_used(src[0], o);\n> +\t}\n> +\ttrace_printf(\"Quick traverse over %d entries from %s to %s\\n\",\n> +\t\t     nr_entries,\n> +\t\t     o->src_index->cache[pos]->name,\n> +\t\t     o->src_index->cache[pos + nr_entries - 1]->name);\n> +\treturn 0;\n> +}\n\nWhen I invented the cache-tree originally, primarily to speed up\nwriting of deeply nested trees, I had the \"diff-index --cached\"\noptimization where a subtree with contents known to be the same as\nthe corresponding span in the index is entirely skipped without\ngetting even looked at.  I didn't realize this (now obvious)\noptimization that scanning the index is faster than opening and\ntraversing trees (I was more focused on not even scanning, which\nis what \"diff-index --cached\" optimization was about).\n\nNice.\n\n\n> +static int index_pos_by_traverse_info(struct name_entry *names,\n> +\t\t\t\t      struct traverse_info *info)\n> +{\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint len = traverse_path_len(info, names);\n> +\tchar *name = xmalloc(len + 1);\n> +\tint pos;\n> +\n> +\tmake_traverse_path(name, info, names);\n> +\tpos = index_name_pos(o->src_index, name, len);\n> +\tif (pos >= 0)\n> +\t\tBUG(\"This is so wrong. This is a directory and should not exist in index\");\n> +\tpos = -pos - 1;\n> +\t/*\n> +\t * There's no guarantee that pos points to the first entry of the\n> +\t * directory. If the directory name is \"letters\" and there's another\n> +\t * file named \"letters.txt\" in the index, pos will point to that file\n> +\t * instead.\n> +\t */\n\nIs this trying to address the issue o->cache_bottom,\nnext_cache_entry(), etc. are trying to address?  i.e. an entry\n\"letters\" appears at a different place relative to other entries in\na tree, depending on the type of the entry itself, so linear and\nparallel scan of the index and the trees may miss matching entries\nwithout backtracking?  If so, I am not sure if the loop below is\nsufficient.\n\n> +\twhile (pos < o->src_index->cache_nr) {\n> +\t\tconst struct cache_entry *ce = o->src_index->cache[pos];\n> +\t\tif (ce_namelen(ce) > len &&\n> +\t\t    ce->name[len] == '/' &&\n> +\t\t    !memcmp(ce->name, name, len))\n> +\t\t\tbreak;\n> +\t\tpos++;\n> +\t}\n> +\tif (pos == o->src_index->cache_nr)\n> +\t\tBUG(\"This is still wrong\");\n> +\tfree(name);\n> +\treturn pos;\n> +}\n> +\n\nIn anycase, nice progress.\n"},{"id":"353737","messageId":"CACsJy8CeF53pA8jfVcY+50-Y_HLm0KkzWvDcTgGV0692hsTHZA@mail.gmail.com","threadId":"48913","inReplyTo":"xmqqpnz8ob2x.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-27T17:52:33Z","receivedAt":"2018-07-27T17:53:04Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Jul 27, 2018 at 7:14 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Duy Nguyen <pclouds@gmail.com> writes:\n>\n> > diff --git a/unpack-trees.c b/unpack-trees.c\n> > index 66741130ae..9c791b55b2 100644\n> > --- a/unpack-trees.c\n> > +++ b/unpack-trees.c\n> > @@ -642,6 +642,110 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n> >       return name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n> >  }\n> >\n> > +static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n> > +                                     struct name_entry *names,\n> > +                                     struct traverse_info *info)\n> > +{\n> > +     struct unpack_trees_options *o = info->data;\n> > +     int i;\n> > +\n> > +     if (dirmask != ((1 << n) - 1) || !S_ISDIR(names->mode) || !o->merge)\n> > +             return 0;\n>\n> In other words, punt if (1) not all are directories, (2) the first\n> name entry given by the caller in names[] is not ISDIR(), or (3) we\n> are not merging i.e. not \"Are we supposed to look at the index too?\"\n> in unpack_callback().\n>\n> I am not sure if the second one is doing us any good.  When\n> S_ISDIR(names->mode) is not true, then the bit in dirmask that\n> corresponds to the one in the entry[] traverse_trees() filled and\n> passed to us must be zero, so the dirmask check would reject such a\n> case anyway, no?\n\nYou're right. This code kinda evolved from the diff_index_cached and I\nforgot about this.\n\n> > +     for (i = 0; i < nr_entries; i++) {\n> > +             struct cache_entry *tree_ce;\n> > +             int len, rc;\n> > +\n> > +             src[0] = o->src_index->cache[pos + i];\n> > +\n> > +             /* Do what unpack_nondirectories() normally does */\n> > +             len = ce_namelen(src[0]);\n> > +             tree_ce = xcalloc(1, cache_entry_size(len));\n>\n> unpack_nondirectories() uses create_ce_entry() here.  Any reason why\n> we shouldn't use it and tell it to make a transient one?\n\nThat one takes a struct name_entry to recreate the path, which will\nnot be correct since we will go deep in subdirs in this loop as well.\n\nSide note. I notice that I allocate/free (and memcpy even) more than I\nshould. The directory part in ce->name for example will never change.\nAnd if the old tree_ce is large enough, we could avoid reallocation\ntoo.\n\n> > +             tree_ce->ce_mode = src[0]->ce_mode;\n> > +             tree_ce->ce_flags = create_ce_flags(0);\n> > +             tree_ce->ce_namelen = len;\n> > +             oidcpy(&tree_ce->oid, &src[0]->oid);\n> > +             memcpy(tree_ce->name, src[0]->name, len + 1);\n> > +\n> > +             for (d = 1; d <= nr_names; d++)\n> > +                     src[d] = tree_ce;\n> > +\n> > +             rc = call_unpack_fn((const struct cache_entry * const *)src, o);\n> > +             free(tree_ce);\n> > +             if (rc < 0)\n> > +                     return rc;\n> > +\n> > +             mark_ce_used(src[0], o);\n> > +     }\n> > +     trace_printf(\"Quick traverse over %d entries from %s to %s\\n\",\n> > +                  nr_entries,\n> > +                  o->src_index->cache[pos]->name,\n> > +                  o->src_index->cache[pos + nr_entries - 1]->name);\n> > +     return 0;\n> > +}\n>\n> When I invented the cache-tree originally, primarily to speed up\n> writing of deeply nested trees, I had the \"diff-index --cached\"\n> optimization where a subtree with contents known to be the same as\n> the corresponding span in the index is entirely skipped without\n> getting even looked at.  I didn't realize this (now obvious)\n> optimization that scanning the index is faster than opening and\n> traversing trees (I was more focused on not even scanning, which\n> is what \"diff-index --cached\" optimization was about).\n>\n> Nice.\n\nI would still love to take this further. We should have cache-tree for\nlike 90% of HEAD, and even if we do 2 or 3 merge where the other trees\nare very different, we should be able to just \"recreate\" HEAD from the\nindex by using cache-tree.\n\nThis is hard though, much trickier than dealing with this case. And I\nguess that the benefit will be much smaller so probably not worth the\ncomplexity.\n\n> > +static int index_pos_by_traverse_info(struct name_entry *names,\n> > +                                   struct traverse_info *info)\n> > +{\n> > +     struct unpack_trees_options *o = info->data;\n> > +     int len = traverse_path_len(info, names);\n> > +     char *name = xmalloc(len + 1);\n> > +     int pos;\n> > +\n> > +     make_traverse_path(name, info, names);\n> > +     pos = index_name_pos(o->src_index, name, len);\n> > +     if (pos >= 0)\n> > +             BUG(\"This is so wrong. This is a directory and should not exist in index\");\n> > +     pos = -pos - 1;\n> > +     /*\n> > +      * There's no guarantee that pos points to the first entry of the\n> > +      * directory. If the directory name is \"letters\" and there's another\n> > +      * file named \"letters.txt\" in the index, pos will point to that file\n> > +      * instead.\n> > +      */\n>\n> Is this trying to address the issue o->cache_bottom,\n> next_cache_entry(), etc. are trying to address?  i.e. an entry\n> \"letters\" appears at a different place relative to other entries in\n> a tree, depending on the type of the entry itself, so linear and\n> parallel scan of the index and the trees may miss matching entries\n> without backtracking?  If so, I am not sure if the loop below is\n> sufficient.\n\nNo it's because index_name_pos does not necessarily give us the right\nstarting point. This is why t6020 fails, where the index has \"letters\"\nand \"letters/foo\" when the cache-tree for \"letters\" is valid. -pos-1\nwould give me the position of \"letters\", not \"letters/foo\". Ideally we\nshould be able to get this starting index from cache-tree code since\nwe're searching for it in there anyway. Then this code could be gone.\n\nThe cache_bottom stuff still scares me though. I reuse mark_ce_used()\nwith hope that it deals with cache_bottom correctly. And as you note,\nthe lookahead code to deal with D/F conflicts could probably mess up\nhere too. You're probably the best one to check this ;-)\n\n> > +     while (pos < o->src_index->cache_nr) {\n> > +             const struct cache_entry *ce = o->src_index->cache[pos];\n> > +             if (ce_namelen(ce) > len &&\n> > +                 ce->name[len] == '/' &&\n> > +                 !memcmp(ce->name, name, len))\n> > +                     break;\n> > +             pos++;\n> > +     }\n> > +     if (pos == o->src_index->cache_nr)\n> > +             BUG(\"This is still wrong\");\n> > +     free(name);\n> > +     return pos;\n> > +}\n> > +\n>\n> In anycase, nice progress.\n\nJust FYI I'm still trying to reduce execution time further and this\nchange happens to half traverse_trees() time (which is a huge deal)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex f0be9f298d..a2e63ad5bf 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -201,7 +201,7 @@ static int do_add_entry(struct\nunpack_trees_options *o, struct cache_entry *ce,\n\n        ce->ce_flags = (ce->ce_flags & ~clear) | set;\n        return add_index_entry(&o->result, ce,\n-                              ADD_CACHE_OK_TO_ADD | ADD_CACHE_OK_TO_REPLACE);\n+                              ADD_CACHE_JUST_APPEND |\nADD_CACHE_OK_TO_ADD | ADD_CACHE_OK_TO_REPLACE);\n }\n\n static struct cache_entry *dup_entry(const struct cache_entry *ce)\n\nIt's probably not the right thing to do of course. But perhaps we\ncould do something in that direction (e.g. validate everything at the\nend of traverse_by_cache_tree...)\n-- \nDuy\n"},{"id":"353738","messageId":"CACsJy8DOhfjMWAb4hP6aoBS6i6DyPuJqj7w2qC3hndo=gy5=zg@mail.gmail.com","threadId":"48913","inReplyTo":"434074a8-1045-8c8f-da0c-873436acf40e@gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-27T18:00:31Z","receivedAt":"2018-07-27T18:01:00Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Jul 27, 2018 at 6:22 PM Ben Peart <peartben@gmail.com> wrote:\n>\n>\n>\n> On 7/27/2018 11:42 AM, Duy Nguyen wrote:\n> > On Thu, Jul 26, 2018 at 12:40:05PM -0700, Junio C Hamano wrote:\n> >> Duy Nguyen <pclouds@gmail.com> writes:\n> >>\n> >>> I'm excited so I decided to try out anyway. This is what I've come up\n> >>> with. Switching trees on git.git shows it could skip plenty entries,\n> >>> so promising. It's ugly and it fails at t6020 though, there's still\n> >>> work ahead. But I think it'll stop here.\n> >>\n> >> We are extremely shallow compared to projects like the kernel and\n> >> stuff from java land, so that is quite an interesting find.\n> >>\n> >\n> > Yeah. I've got a more or less complete patch now with full test suite\n> > passed and even with linux.git, the numbers look pretty good.\n> >\n> > Ben, is it possible for you to try this one out? I don't suppose it\n> > will be that good on a real big repo. But I'm curious how much faster\n> > could this patch does.\n> >\n>\n> Thanks Duy.  I'm super excited about this so did a quick and dirty\n> manual perf test.\n>\n> I ran \"git checkout\" 5 times, discarded the first 2 runs and averaged\n> the last 3 with and without this patch on top of VFSForGit in a large repo.\n>\n> Without this patch average times were 16.97\n> With this patch average times were 10.55\n>\n> That is a significant improvement!\n\nMeh! Junio cut down time to like 1/5th in b65982b608 (Optimize\n\"diff-index --cached\" using cache-tree - 2009-05-20). This is not\nenough!\n\nOK i'm kidding :) I'd like to see you measure traverse_trees like in\nyour first mail though. Total checkout number is nice and all but I\nstill like to see exactly how much time is reduced in traverse_trees()\nalone (or unpack_trees() to be precise). That would give me a much\nbetter picture of this unpacking business.\n-- \nDuy\n"},{"id":"353813","messageId":"20180729062424.GA22870@duynguyen.home","threadId":"48913","inReplyTo":"CACsJy8CeF53pA8jfVcY+50-Y_HLm0KkzWvDcTgGV0692hsTHZA@mail.gmail.com","subject":"Re: [PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-29T06:24:24Z","receivedAt":"2018-07-29T06:24:31Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Jul 27, 2018 at 07:52:33PM +0200, Duy Nguyen wrote:\n> Just FYI I'm still trying to reduce execution time further and this\n> change happens to half traverse_trees() time (which is a huge deal)\n> \n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index f0be9f298d..a2e63ad5bf 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -201,7 +201,7 @@ static int do_add_entry(struct\n> unpack_trees_options *o, struct cache_entry *ce,\n> \n>         ce->ce_flags = (ce->ce_flags & ~clear) | set;\n>         return add_index_entry(&o->result, ce,\n> -                              ADD_CACHE_OK_TO_ADD | ADD_CACHE_OK_TO_REPLACE);\n> +                              ADD_CACHE_JUST_APPEND |\n> ADD_CACHE_OK_TO_ADD | ADD_CACHE_OK_TO_REPLACE);\n>  }\n> \n>  static struct cache_entry *dup_entry(const struct cache_entry *ce)\n> \n> It's probably not the right thing to do of course. But perhaps we\n> could do something in that direction (e.g. validate everything at the\n> end of traverse_by_cache_tree...)\n\nIt's just too much computation that could be reduced. The following\npatch gives more or less the same performance gain as adding\nADD_CACHE_JUST_APPEND (traverse_trees() time cut down by half).\n\nOf these, the walking cache-tree inside add_index_entry_with_check()\nis most expensive and we probably could just walk the cache-tree in\ntraverse_by_cache_tree() loop and do the invalidation there instead.\n\n-- 8< --\ndiff --git a/cache.h b/cache.h\nindex 8b447652a7..e6f7ee4b64 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -673,6 +673,7 @@ extern int index_name_pos(const struct index_state *, const char *name, int name\n #define ADD_CACHE_JUST_APPEND 8\t\t/* Append only; tree.c::read_tree() */\n #define ADD_CACHE_NEW_ONLY 16\t\t/* Do not replace existing ones */\n #define ADD_CACHE_KEEP_CACHE_TREE 32\t/* Do not invalidate cache-tree */\n+#define ADD_CACHE_SKIP_VERIFY_PATH 64\t/* Do not verify path */\n extern int add_index_entry(struct index_state *, struct cache_entry *ce, int option);\n extern void rename_index_entry_at(struct index_state *, int pos, const char *new_name);\n \ndiff --git a/read-cache.c b/read-cache.c\nindex e865254bea..b0b5df5de7 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1170,6 +1170,7 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n \tint ok_to_add = option & ADD_CACHE_OK_TO_ADD;\n \tint ok_to_replace = option & ADD_CACHE_OK_TO_REPLACE;\n \tint skip_df_check = option & ADD_CACHE_SKIP_DFCHECK;\n+\tint skip_verify_path = option & ADD_CACHE_SKIP_VERIFY_PATH;\n \tint new_only = option & ADD_CACHE_NEW_ONLY;\n \n \tif (!(option & ADD_CACHE_KEEP_CACHE_TREE))\n@@ -1210,7 +1211,7 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n \n \tif (!ok_to_add)\n \t\treturn -1;\n-\tif (!verify_path(ce->name, ce->ce_mode))\n+\tif (!skip_verify_path && !verify_path(ce->name, ce->ce_mode))\n \t\treturn error(\"Invalid path '%s'\", ce->name);\n \n \tif (!skip_df_check &&\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex f2a2db6ab8..ff6a0f2bd3 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -201,6 +201,7 @@ static int do_add_entry(struct unpack_trees_options *o, struct cache_entry *ce,\n \n \tce->ce_flags = (ce->ce_flags & ~clear) | set;\n \treturn add_index_entry(&o->result, ce,\n+\t\t\t       o->extra_add_index_flags |\n \t\t\t       ADD_CACHE_OK_TO_ADD | ADD_CACHE_OK_TO_REPLACE);\n }\n \n@@ -678,6 +679,25 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \tconst char *first_name = o->src_index->cache[pos]->name;\n \tint dirlen = (strrchr(first_name, '/') - first_name)+1;\n \n+\t/*\n+\t * Try to keep add_index_entry() as fast as possible since\n+\t * we're going to do a lot of them.\n+\t *\n+\t * Skipping verify_path() should totally be safe because these\n+\t * paths are from the source index, which must have been\n+\t * verified.\n+\t *\n+\t * Skipping D/F and cache-tree validation checks is trickier\n+\t * because it assumes what n-merge code would do when all\n+\t * trees and the index are the same. We probably could just\n+\t * optimize those code instead (e.g. we don't invalidate that\n+\t * many cache-tree, but the searching for them is very\n+\t * expensive).\n+\t */\n+\to->extra_add_index_flags = ADD_CACHE_SKIP_DFCHECK;\n+\to->extra_add_index_flags |= ADD_CACHE_KEEP_CACHE_TREE;\n+\to->extra_add_index_flags |= ADD_CACHE_SKIP_VERIFY_PATH;\n+\n \t/*\n \t * Do what unpack_callback() and unpack_nondirectories() normally\n \t * do. But we do it in one function call (for even nested trees)\n@@ -721,6 +741,7 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \n \t\tmark_ce_used(src[0], o);\n \t}\n+\to->extra_add_index_flags = 0;\n \tfree(tree_ce);\n \ttrace_printf(\"Quick traverse over %d entries from %s to %s\\n\",\n \t\t     nr_entries,\ndiff --git a/unpack-trees.h b/unpack-trees.h\nindex c2b434c606..94e1b14078 100644\n--- a/unpack-trees.h\n+++ b/unpack-trees.h\n@@ -80,6 +80,7 @@ struct unpack_trees_options {\n \tstruct index_state result;\n \n \tstruct exclude_list *el; /* for internal use */\n+\tunsigned int extra_add_index_flags;\n };\n \n extern int unpack_trees(unsigned n, struct tree_desc *t,\n-- 8< --\n"},{"id":"353815","messageId":"20180729103306.16403-1-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180727154241.GA21288@duynguyen.home","subject":"[PATCH v2 0/4] Speed up unpack_trees()","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-29T10:33:02Z","receivedAt":"2018-07-29T10:33:13Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This series speeds up unpack_trees() a bit by using cache-tree.\nunpack-trees could bit split in three big parts\n\n- the actual tree unpacking and running n-way merging\n- update worktree, which could be expensive depending on how much I/O\n  is involved\n- repair cache-tree\n\nThis series focuses on the first part alone and could give 700%\nspeedup (best case possible scenario, real life ones probably not that\nimpressive).\n\nIt also shows that the reparing cache-tree is kinda expensive. I have\nan idea of reusing cache-tree from the original index, but I'll leave\nthat to Ben or others to try out and see if it helps at all.\n\nv2 fixes the comments from Junio, adds more performance tracing and\nreduces the cost of adding index entries.\n\nNguyễn Thái Ngọc Duy (4):\n  unpack-trees.c: add performance tracing\n  unpack-trees: optimize walking same trees with cache-tree\n  unpack-trees: reduce malloc in cache-tree walk\n  unpack-trees: cheaper index update when walking by cache-tree\n\n cache-tree.c   |   2 +\n cache.h        |   1 +\n read-cache.c   |   3 +-\n unpack-trees.c | 161 ++++++++++++++++++++++++++++++++++++++++++++++++-\n unpack-trees.h |   1 +\n 5 files changed, 166 insertions(+), 2 deletions(-)\n\n-- \n2.18.0.656.gda699b98b3\n\n"},{"id":"353816","messageId":"20180729103306.16403-2-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-1-pclouds@gmail.com","subject":"[PATCH v2 1/4] unpack-trees.c: add performance tracing","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-29T10:33:03Z","receivedAt":"2018-07-29T10:33:14Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"We're going to optimize unpack_trees() a bit in the following\npatches. Let's add some tracing to measure how long it takes before\nand after. This is the baseline (\"git checkout -\" on gcc.git, 80k\nfiles on worktree)\n\n    0.018239226 s: read cache .git/index\n    0.052541655 s: preload index\n    0.001537598 s: refresh index\n    0.168167768 s: unpack trees\n    0.002897186 s: update worktree after a merge\n    0.131661745 s: repair cache-tree\n    0.075389117 s: write index, changed mask = 2a\n    0.111702023 s: unpack trees\n    0.000023245 s: update worktree after a merge\n    0.111793866 s: diff-index\n    0.587933288 s: git command: /home/pclouds/w/git/git checkout -\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache-tree.c   | 2 ++\n unpack-trees.c | 4 ++++\n 2 files changed, 6 insertions(+)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 6b46711996..0dbe10fc85 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -426,6 +426,7 @@ static int update_one(struct cache_tree *it,\n \n int cache_tree_update(struct index_state *istate, int flags)\n {\n+\tuint64_t start = getnanotime();\n \tstruct cache_tree *it = istate->cache_tree;\n \tstruct cache_entry **cache = istate->cache;\n \tint entries = istate->cache_nr;\n@@ -437,6 +438,7 @@ int cache_tree_update(struct index_state *istate, int flags)\n \tif (i < 0)\n \t\treturn i;\n \tistate->cache_changed |= CACHE_TREE_CHANGED;\n+\ttrace_performance_since(start, \"repair cache-tree\");\n \treturn 0;\n }\n \ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 66741130ae..dc58d1f5ae 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -352,6 +352,7 @@ static int check_updates(struct unpack_trees_options *o)\n \tstruct progress *progress = NULL;\n \tstruct index_state *index = &o->result;\n \tstruct checkout state = CHECKOUT_INIT;\n+\tuint64_t start = getnanotime();\n \tint i;\n \n \tstate.force = 1;\n@@ -423,6 +424,7 @@ static int check_updates(struct unpack_trees_options *o)\n \terrs |= finish_delayed_checkout(&state);\n \tif (o->update)\n \t\tgit_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n+\ttrace_performance_since(start, \"update worktree after a merge\");\n \treturn errs != 0;\n }\n \n@@ -1275,6 +1277,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \tint i, ret;\n \tstatic struct cache_entry *dfc;\n \tstruct exclude_list el;\n+\tuint64_t start = getnanotime();\n \n \tif (len > MAX_UNPACK_TREES)\n \t\tdie(\"unpack_trees takes at most %d trees\", MAX_UNPACK_TREES);\n@@ -1423,6 +1426,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t\tgoto done;\n \t\t}\n \t}\n+\ttrace_performance_since(start, \"unpack trees\");\n \n \tret = check_updates(o) ? (-2) : 0;\n \tif (o->dst_index) {\n-- \n2.18.0.656.gda699b98b3\n\n"},{"id":"353817","messageId":"20180729103306.16403-3-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-1-pclouds@gmail.com","subject":"[PATCH v2 2/4] unpack-trees: optimize walking same trees with cache-tree","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-29T10:33:04Z","receivedAt":"2018-07-29T10:33:15Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"From: Duy Nguyen <pclouds@gmail.com>\n\nIn order to merge one or many trees with the index, unpack-trees code\nwalks multiple trees in parallel with the index and performs n-way\nmerge. If we find out at start of a directory that all trees are the\nsame (by comparing OID) and cache-tree happens to be available for\nthat directory as well, we could avoid walking the trees because we\nalready know what these trees contain: it's flattened in what's called\n\"the index\".\n\nThe upside is of course a lot less I/O since we can potentially skip\nlots of trees (think subtrees). We also save CPU because we don't have\nto inflate and the apply deltas. The downside is of course more\nfragile code since the logic in some functions are now duplicated\nelsewhere.\n\n\"checkout -\" with this patch on gcc.git:\n\n    baseline      new\n  --------------------------------------------------------------------\n    0.018239226   0.019365414 s: read cache .git/index\n    0.052541655   0.049605548 s: preload index\n    0.001537598   0.001571695 s: refresh index\n    0.168167768   0.049677212 s: unpack trees\n    0.002897186   0.002845256 s: update worktree after a merge\n    0.131661745   0.136597522 s: repair cache-tree\n    0.075389117   0.075422517 s: write index, changed mask = 2a\n    0.111702023   0.032813253 s: unpack trees\n    0.000023245   0.000022002 s: update worktree after a merge\n    0.111793866   0.032933140 s: diff-index\n    0.587933288   0.398924370 s: git command: /home/pclouds/w/git/git\n\nThis command calls unpack_trees() twice, the first time on 2way merge\nand the second 1way merge. In both times, \"unpack trees\" time is\nreduced to one third. Overall time reduction is not that impressive of\ncourse because index operations take a big chunk. And there's that\nrepair cache-tree line.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n unpack-trees.c | 119 ++++++++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 118 insertions(+), 1 deletion(-)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex dc58d1f5ae..39566b28fb 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -644,6 +644,102 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n \treturn name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n }\n \n+static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n+\t\t\t\t\tstruct name_entry *names,\n+\t\t\t\t\tstruct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i;\n+\n+\tif (!o->merge || dirmask != ((1 << n) - 1))\n+\t\treturn 0;\n+\n+\tfor (i = 1; i < n; i++)\n+\t\tif (!are_same_oid(names, names + i))\n+\t\t\treturn 0;\n+\n+\treturn cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n+}\n+\n+static int index_pos_by_traverse_info(struct name_entry *names,\n+\t\t\t\t      struct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint len = traverse_path_len(info, names);\n+\tchar *name = xmalloc(len + 1 /* slash */ + 1 /* NUL */);\n+\tint pos;\n+\n+\tmake_traverse_path(name, info, names);\n+\tname[len++] = '/';\n+\tname[len] = '\\0';\n+\tpos = index_name_pos(o->src_index, name, len);\n+\tif (pos >= 0)\n+\t\tBUG(\"This is a directory and should not exist in index\");\n+\tpos = -pos - 1;\n+\tif (!starts_with(o->src_index->cache[pos]->name, name) ||\n+\t    (pos > 0 && starts_with(o->src_index->cache[pos-1]->name, name)))\n+\t\tBUG(\"pos must point at the first entry in this directory\");\n+\tfree(name);\n+\treturn pos;\n+}\n+\n+/*\n+ * Fast path if we detect that all trees are the same as cache-tree at this\n+ * path. We'll walk these trees recursively using cache-tree/index instead of\n+ * ODB since already know what these trees contain.\n+ */\n+static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n+\t\t\t\t  struct name_entry *names,\n+\t\t\t\t  struct traverse_info *info)\n+{\n+\tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i, d;\n+\n+\tif (!o->merge)\n+\t\tBUG(\"We need cache-tree to do this optimization\");\n+\n+\t/*\n+\t * Do what unpack_callback() and unpack_nondirectories() normally\n+\t * do. But we walk all paths recursively in just one loop instead.\n+\t *\n+\t * D/F conflicts and staged entries are not a concern because\n+\t * cache-tree would be invalidated and we would never get here\n+\t * in the first place.\n+\t */\n+\tfor (i = 0; i < nr_entries; i++) {\n+\t\tstruct cache_entry *tree_ce;\n+\t\tint len, rc;\n+\n+\t\tsrc[0] = o->src_index->cache[pos + i];\n+\n+\t\tlen = ce_namelen(src[0]);\n+\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n+\n+\t\ttree_ce->ce_mode = src[0]->ce_mode;\n+\t\ttree_ce->ce_flags = create_ce_flags(0);\n+\t\ttree_ce->ce_namelen = len;\n+\t\toidcpy(&tree_ce->oid, &src[0]->oid);\n+\t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n+\n+\t\tfor (d = 1; d <= nr_names; d++)\n+\t\t\tsrc[d] = tree_ce;\n+\n+\t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n+\t\tfree(tree_ce);\n+\t\tif (rc < 0)\n+\t\t\treturn rc;\n+\n+\t\tmark_ce_used(src[0], o);\n+\t}\n+\tif (o->debug_unpack)\n+\t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n+\t\t       nr_entries,\n+\t\t       o->src_index->cache[pos]->name,\n+\t\t       o->src_index->cache[pos + nr_entries - 1]->name);\n+\treturn 0;\n+}\n+\n static int traverse_trees_recursive(int n, unsigned long dirmask,\n \t\t\t\t    unsigned long df_conflicts,\n \t\t\t\t    struct name_entry *names,\n@@ -655,6 +751,17 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n \tvoid *buf[MAX_UNPACK_TREES];\n \tstruct traverse_info newinfo;\n \tstruct name_entry *p;\n+\tint nr_entries;\n+\n+\tnr_entries = all_trees_same_as_cache_tree(n, dirmask, names, info);\n+\tif (nr_entries > 0) {\n+\t\tstruct unpack_trees_options *o = info->data;\n+\t\tint pos = index_pos_by_traverse_info(names, info);\n+\n+\t\tif (!o->merge || df_conflicts)\n+\t\t\tBUG(\"Wrong condition to get here buddy\");\n+\t\treturn traverse_by_cache_tree(pos, nr_entries, n, names, info);\n+\t}\n \n \tp = names;\n \twhile (!p->mode)\n@@ -814,6 +921,11 @@ static struct cache_entry *create_ce_entry(const struct traverse_info *info, con\n \treturn ce;\n }\n \n+/*\n+ * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n+ * without actually calling it. If you change the logic here you may need to\n+ * check and change there as well.\n+ */\n static int unpack_nondirectories(int n, unsigned long mask,\n \t\t\t\t unsigned long dirmask,\n \t\t\t\t struct cache_entry **src,\n@@ -998,6 +1110,11 @@ static void debug_unpack_callback(int n,\n \t\tdebug_name_entry(i, names + i);\n }\n \n+/*\n+ * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n+ * without actually calling it. If you change the logic here you may need to\n+ * check and change there as well.\n+ */\n static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n@@ -1280,7 +1397,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \tuint64_t start = getnanotime();\n \n \tif (len > MAX_UNPACK_TREES)\n-\t\tdie(\"unpack_trees takes at most %d trees\", MAX_UNPACK_TREES);\n+\t\tdie(_(\"unpack_trees takes at most %d trees\"), MAX_UNPACK_TREES);\n \n \tmemset(&el, 0, sizeof(el));\n \tif (!core_apply_sparse_checkout || !o->update)\n-- \n2.18.0.656.gda699b98b3\n\n"},{"id":"353818","messageId":"20180729103306.16403-4-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-1-pclouds@gmail.com","subject":"[PATCH v2 3/4] unpack-trees: reduce malloc in cache-tree walk","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-29T10:33:05Z","receivedAt":"2018-07-29T10:33:16Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This is a micro optimization that probably only shines on repos with\ndeep directory structure. Instead of allocating and freeing a new\ncache_entry in every iteration, we reuse the last one and only update\nthe parts that are new each iteration.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n unpack-trees.c | 29 ++++++++++++++++++++---------\n 1 file changed, 20 insertions(+), 9 deletions(-)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 39566b28fb..c33ebaf001 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -694,6 +694,8 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n \tstruct unpack_trees_options *o = info->data;\n+\tstruct cache_entry *tree_ce = NULL;\n+\tint ce_len = 0;\n \tint i, d;\n \n \tif (!o->merge)\n@@ -708,30 +710,39 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \t * in the first place.\n \t */\n \tfor (i = 0; i < nr_entries; i++) {\n-\t\tstruct cache_entry *tree_ce;\n-\t\tint len, rc;\n+\t\tint new_ce_len, len, rc;\n \n \t\tsrc[0] = o->src_index->cache[pos + i];\n \n \t\tlen = ce_namelen(src[0]);\n-\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n+\t\tnew_ce_len = cache_entry_size(len);\n+\n+\t\tif (new_ce_len > ce_len) {\n+\t\t\tnew_ce_len <<= 1;\n+\t\t\ttree_ce = xrealloc(tree_ce, new_ce_len);\n+\t\t\tmemset(tree_ce, 0, new_ce_len);\n+\t\t\tce_len = new_ce_len;\n+\n+\t\t\ttree_ce->ce_flags = create_ce_flags(0);\n+\n+\t\t\tfor (d = 1; d <= nr_names; d++)\n+\t\t\t\tsrc[d] = tree_ce;\n+\t\t}\n \n \t\ttree_ce->ce_mode = src[0]->ce_mode;\n-\t\ttree_ce->ce_flags = create_ce_flags(0);\n \t\ttree_ce->ce_namelen = len;\n \t\toidcpy(&tree_ce->oid, &src[0]->oid);\n \t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n \n-\t\tfor (d = 1; d <= nr_names; d++)\n-\t\t\tsrc[d] = tree_ce;\n-\n \t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n-\t\tfree(tree_ce);\n-\t\tif (rc < 0)\n+\t\tif (rc < 0) {\n+\t\t\tfree(tree_ce);\n \t\t\treturn rc;\n+\t\t}\n \n \t\tmark_ce_used(src[0], o);\n \t}\n+\tfree(tree_ce);\n \tif (o->debug_unpack)\n \t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n \t\t       nr_entries,\n-- \n2.18.0.656.gda699b98b3\n\n"},{"id":"353819","messageId":"20180729103306.16403-5-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-1-pclouds@gmail.com","subject":"[PATCH v2 4/4] unpack-trees: cheaper index update when walking by cache-tree","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-29T10:33:06Z","receivedAt":"2018-07-29T10:33:18Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"With the new cache-tree, we could mostly avoid I/O (due to odb access)\nthe code mostly becomes a loop of \"check this, check that, add the\nentry to the index\". We could skip a couple checks in this giant loop\nto go faster:\n\n- We know here that we're copying entries from the source index to the\n  result one. All paths in the source index must have been validated\n  at load time already (and we're not taking strange paths from tree\n  objects) which means we can skip verify_path() without compromise.\n\n- We also know that D/F conflicts can't happen for all these entries\n  (since cache-tree and all the trees are the same) so we can skip\n  that as well.\n\nThis gives rather nice speedups for \"unpack trees\" rows where \"unpack\ntrees\" time is now cut in half compared to when\ntraverse_by_cache_tree() is added, or 1/7 of the original \"unpack\ntrees\" time.\n\n   baseline      cache-tree    this patch\n --------------------------------------------------------------------\n   0.018239226   0.019365414   0.020519621 s: read cache .git/index\n   0.052541655   0.049605548   0.048814384 s: preload index\n   0.001537598   0.001571695   0.001575382 s: refresh index\n   0.168167768   0.049677212   0.024719308 s: unpack trees\n   0.002897186   0.002845256   0.002805555 s: update worktree after a merge\n   0.131661745   0.136597522   0.134891617 s: repair cache-tree\n   0.075389117   0.075422517   0.074832291 s: write index, changed mask = 2a\n   0.111702023   0.032813253   0.008616479 s: unpack trees\n   0.000023245   0.000022002   0.000026630 s: update worktree after a merge\n   0.111793866   0.032933140   0.008714071 s: diff-index\n   0.587933288   0.398924370   0.380452871 s: git command: /home/pclouds/w/git/git\n\nTotal saving of this new patch looks even less impressive, now that\ntime spent in unpacking trees is so small. Which is why the next\nattempt should be on that \"repair cache-tree\" line.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache.h        |  1 +\n read-cache.c   |  3 ++-\n unpack-trees.c | 27 +++++++++++++++++++++++++++\n unpack-trees.h |  1 +\n 4 files changed, 31 insertions(+), 1 deletion(-)\n\ndiff --git a/cache.h b/cache.h\nindex 8b447652a7..e6f7ee4b64 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -673,6 +673,7 @@ extern int index_name_pos(const struct index_state *, const char *name, int name\n #define ADD_CACHE_JUST_APPEND 8\t\t/* Append only; tree.c::read_tree() */\n #define ADD_CACHE_NEW_ONLY 16\t\t/* Do not replace existing ones */\n #define ADD_CACHE_KEEP_CACHE_TREE 32\t/* Do not invalidate cache-tree */\n+#define ADD_CACHE_SKIP_VERIFY_PATH 64\t/* Do not verify path */\n extern int add_index_entry(struct index_state *, struct cache_entry *ce, int option);\n extern void rename_index_entry_at(struct index_state *, int pos, const char *new_name);\n \ndiff --git a/read-cache.c b/read-cache.c\nindex e865254bea..b0b5df5de7 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1170,6 +1170,7 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n \tint ok_to_add = option & ADD_CACHE_OK_TO_ADD;\n \tint ok_to_replace = option & ADD_CACHE_OK_TO_REPLACE;\n \tint skip_df_check = option & ADD_CACHE_SKIP_DFCHECK;\n+\tint skip_verify_path = option & ADD_CACHE_SKIP_VERIFY_PATH;\n \tint new_only = option & ADD_CACHE_NEW_ONLY;\n \n \tif (!(option & ADD_CACHE_KEEP_CACHE_TREE))\n@@ -1210,7 +1211,7 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n \n \tif (!ok_to_add)\n \t\treturn -1;\n-\tif (!verify_path(ce->name, ce->ce_mode))\n+\tif (!skip_verify_path && !verify_path(ce->name, ce->ce_mode))\n \t\treturn error(\"Invalid path '%s'\", ce->name);\n \n \tif (!skip_df_check &&\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex c33ebaf001..dc62afd968 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -201,6 +201,7 @@ static int do_add_entry(struct unpack_trees_options *o, struct cache_entry *ce,\n \n \tce->ce_flags = (ce->ce_flags & ~clear) | set;\n \treturn add_index_entry(&o->result, ce,\n+\t\t\t       o->extra_add_index_flags |\n \t\t\t       ADD_CACHE_OK_TO_ADD | ADD_CACHE_OK_TO_REPLACE);\n }\n \n@@ -701,6 +702,24 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \tif (!o->merge)\n \t\tBUG(\"We need cache-tree to do this optimization\");\n \n+\t/*\n+\t * Try to keep add_index_entry() as fast as possible since\n+\t * we're going to do a lot of them.\n+\t *\n+\t * Skipping verify_path() should totally be safe because these\n+\t * paths are from the source index, which must have been\n+\t * verified.\n+\t *\n+\t * Skipping D/F and cache-tree validation checks is trickier\n+\t * because it assumes what n-merge code would do when all\n+\t * trees and the index are the same. We probably could just\n+\t * optimize those code instead (e.g. we don't invalidate that\n+\t * many cache-tree, but the searching for them is very\n+\t * expensive).\n+\t */\n+\to->extra_add_index_flags = ADD_CACHE_SKIP_DFCHECK;\n+\to->extra_add_index_flags |= ADD_CACHE_SKIP_VERIFY_PATH;\n+\n \t/*\n \t * Do what unpack_callback() and unpack_nondirectories() normally\n \t * do. But we walk all paths recursively in just one loop instead.\n@@ -742,6 +761,7 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \n \t\tmark_ce_used(src[0], o);\n \t}\n+\to->extra_add_index_flags = 0;\n \tfree(tree_ce);\n \tif (o->debug_unpack)\n \t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n@@ -1561,6 +1581,13 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\tif (!ret) {\n \t\t\tif (!o->result.cache_tree)\n \t\t\t\to->result.cache_tree = cache_tree();\n+\t\t\t/*\n+\t\t\t * TODO: Walk o.src_index->cache_tree, quickly check\n+\t\t\t * if o->result.cache has the exact same content for\n+\t\t\t * any valid cache-tree in o.src_index, then we can\n+\t\t\t * just copy the cache-tree over instead of hashing a\n+\t\t\t * new tree object.\n+\t\t\t */\n \t\t\tif (!cache_tree_fully_valid(o->result.cache_tree))\n \t\t\t\tcache_tree_update(&o->result,\n \t\t\t\t\t\t  WRITE_TREE_SILENT |\ndiff --git a/unpack-trees.h b/unpack-trees.h\nindex c2b434c606..94e1b14078 100644\n--- a/unpack-trees.h\n+++ b/unpack-trees.h\n@@ -80,6 +80,7 @@ struct unpack_trees_options {\n \tstruct index_state result;\n \n \tstruct exclude_list *el; /* for internal use */\n+\tunsigned int extra_add_index_flags;\n };\n \n extern int unpack_trees(unsigned n, struct tree_desc *t,\n-- \n2.18.0.656.gda699b98b3\n\n"},{"id":"353925","messageId":"9a9a309c-7143-e642-cfd8-6df76e77995a@gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-1-pclouds@gmail.com","subject":"Re: [PATCH v2 0/4] Speed up unpack_trees()","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-30T18:10:34Z","receivedAt":"2018-07-30T18:10:39Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/29/2018 6:33 AM, Nguyễn Thái Ngọc Duy wrote:\n> This series speeds up unpack_trees() a bit by using cache-tree.\n> unpack-trees could bit split in three big parts\n> \n> - the actual tree unpacking and running n-way merging\n> - update worktree, which could be expensive depending on how much I/O\n>    is involved\n> - repair cache-tree\n> \n> This series focuses on the first part alone and could give 700%\n> speedup (best case possible scenario, real life ones probably not that\n> impressive).\n> \n> It also shows that the reparing cache-tree is kinda expensive. I have\n> an idea of reusing cache-tree from the original index, but I'll leave\n> that to Ben or others to try out and see if it helps at all.\n> \n> v2 fixes the comments from Junio, adds more performance tracing and\n> reduces the cost of adding index entries.\n> \n> Nguyễn Thái Ngọc Duy (4):\n>    unpack-trees.c: add performance tracing\n>    unpack-trees: optimize walking same trees with cache-tree\n>    unpack-trees: reduce malloc in cache-tree walk\n>    unpack-trees: cheaper index update when walking by cache-tree\n> \n>   cache-tree.c   |   2 +\n>   cache.h        |   1 +\n>   read-cache.c   |   3 +-\n>   unpack-trees.c | 161 ++++++++++++++++++++++++++++++++++++++++++++++++-\n>   unpack-trees.h |   1 +\n>   5 files changed, 166 insertions(+), 2 deletions(-)\n> \n\nI ran \"git checkout\" on a large repo and averaged the results of 3 runs. \n  This clearly demonstrates the benefit of the optimized unpack_trees() \nas even the final \"diff-index\" is essentially a 3rd call to unpack_trees().\n\nbaseline\tnew\t\n----------------------------------------------------------------------\n0.535510167\t0.556558733\ts: read cache .git/index\n0.3057373\t0.3147105\ts: initialize name hash\n0.0184082\t0.023558433\ts: preload index\n0.086910967\t0.089085967\ts: refresh index\n7.889590767\t2.191554433\ts: unpack trees\n0.120760833\t0.131941267\ts: update worktree after a merge\n2.2583504\t2.572663167\ts: repair cache-tree\n0.8916137\t0.959495233\ts: write index, changed mask = 28\n3.405199233\t0.2710663\ts: unpack trees\n0.000999667\t0.0021554\ts: update worktree after a merge\n3.4063306\t0.273318333\ts: diff-index\n16.9524923\t9.462943133\ts: git command: \n'c:\\git-sdk-64\\usr\\src\\git\\git.exe' checkout\n\nThe first call to unpack_trees() saves 72%\nThe 2nd and 3rd call save 92%\nTotal time savings for the entire command was 44%\n\nIn the performance game of whack-a-mole, that call to repair cache-tree \nis now looking quite expensive...\n\nBen\n"},{"id":"353947","messageId":"cae996fc-38d7-a691-3b66-b0b504513f75@gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-2-pclouds@gmail.com","subject":"Re: [PATCH v2 1/4] unpack-trees.c: add performance tracing","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-30T20:16:03Z","receivedAt":"2018-07-30T20:16:09Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/29/2018 6:33 AM, Nguyễn Thái Ngọc Duy wrote:\n> We're going to optimize unpack_trees() a bit in the following\n> patches. Let's add some tracing to measure how long it takes before\n> and after. This is the baseline (\"git checkout -\" on gcc.git, 80k\n> files on worktree)\n> \n>      0.018239226 s: read cache .git/index\n>      0.052541655 s: preload index\n>      0.001537598 s: refresh index\n>      0.168167768 s: unpack trees\n>      0.002897186 s: update worktree after a merge\n>      0.131661745 s: repair cache-tree\n>      0.075389117 s: write index, changed mask = 2a\n>      0.111702023 s: unpack trees\n>      0.000023245 s: update worktree after a merge\n>      0.111793866 s: diff-index\n>      0.587933288 s: git command: /home/pclouds/w/git/git checkout -\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\nI've reviewed this patch and it looks good to me.  Nice to see the \nadditional breakdown on where time is being spent.\n\n"},{"id":"353956","messageId":"2b97c84e-0634-3879-18a1-8f96a7b4a30d@gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-3-pclouds@gmail.com","subject":"Re: [PATCH v2 2/4] unpack-trees: optimize walking same trees with cache-tree","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-30T20:52:54Z","receivedAt":"2018-07-30T20:53:00Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/29/2018 6:33 AM, Nguyễn Thái Ngọc Duy wrote:\n> From: Duy Nguyen <pclouds@gmail.com>\n> \n> In order to merge one or many trees with the index, unpack-trees code\n> walks multiple trees in parallel with the index and performs n-way\n> merge. If we find out at start of a directory that all trees are the\n> same (by comparing OID) and cache-tree happens to be available for\n> that directory as well, we could avoid walking the trees because we\n> already know what these trees contain: it's flattened in what's called\n> \"the index\".\n> \n> The upside is of course a lot less I/O since we can potentially skip\n> lots of trees (think subtrees). We also save CPU because we don't have\n> to inflate and the apply deltas. The downside is of course more\n> fragile code since the logic in some functions are now duplicated\n> elsewhere.\n> \n> \"checkout -\" with this patch on gcc.git:\n> \n>      baseline      new\n>    --------------------------------------------------------------------\n>      0.018239226   0.019365414 s: read cache .git/index\n>      0.052541655   0.049605548 s: preload index\n>      0.001537598   0.001571695 s: refresh index\n>      0.168167768   0.049677212 s: unpack trees\n>      0.002897186   0.002845256 s: update worktree after a merge\n>      0.131661745   0.136597522 s: repair cache-tree\n>      0.075389117   0.075422517 s: write index, changed mask = 2a\n>      0.111702023   0.032813253 s: unpack trees\n>      0.000023245   0.000022002 s: update worktree after a merge\n>      0.111793866   0.032933140 s: diff-index\n>      0.587933288   0.398924370 s: git command: /home/pclouds/w/git/git\n> \n> This command calls unpack_trees() twice, the first time on 2way merge\n> and the second 1way merge. In both times, \"unpack trees\" time is\n> reduced to one third. Overall time reduction is not that impressive of\n> course because index operations take a big chunk. And there's that\n> repair cache-tree line.\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>   unpack-trees.c | 119 ++++++++++++++++++++++++++++++++++++++++++++++++-\n>   1 file changed, 118 insertions(+), 1 deletion(-)\n> \n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index dc58d1f5ae..39566b28fb 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -644,6 +644,102 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n>   \treturn name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n>   }\n>   \n> +static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n> +\t\t\t\t\tstruct name_entry *names,\n> +\t\t\t\t\tstruct traverse_info *info)\n> +{\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint i;\n> +\n> +\tif (!o->merge || dirmask != ((1 << n) - 1))\n> +\t\treturn 0;\n> +\n> +\tfor (i = 1; i < n; i++)\n> +\t\tif (!are_same_oid(names, names + i))\n> +\t\t\treturn 0;\n> +\n> +\treturn cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n> +}\n> +\n> +static int index_pos_by_traverse_info(struct name_entry *names,\n> +\t\t\t\t      struct traverse_info *info)\n> +{\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint len = traverse_path_len(info, names);\n> +\tchar *name = xmalloc(len + 1 /* slash */ + 1 /* NUL */);\n> +\tint pos;\n> +\n> +\tmake_traverse_path(name, info, names);\n> +\tname[len++] = '/';\n> +\tname[len] = '\\0';\n> +\tpos = index_name_pos(o->src_index, name, len);\n> +\tif (pos >= 0)\n> +\t\tBUG(\"This is a directory and should not exist in index\");\n> +\tpos = -pos - 1;\n> +\tif (!starts_with(o->src_index->cache[pos]->name, name) ||\n> +\t    (pos > 0 && starts_with(o->src_index->cache[pos-1]->name, name)))\n> +\t\tBUG(\"pos must point at the first entry in this directory\");\n> +\tfree(name);\n> +\treturn pos;\n> +}\n> +\n> +/*\n> + * Fast path if we detect that all trees are the same as cache-tree at this\n> + * path. We'll walk these trees recursively using cache-tree/index instead of\n> + * ODB since already know what these trees contain.\n> + */\n> +static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n> +\t\t\t\t  struct name_entry *names,\n> +\t\t\t\t  struct traverse_info *info)\n> +{\n> +\tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint i, d;\n> +\n> +\tif (!o->merge)\n> +\t\tBUG(\"We need cache-tree to do this optimization\");\n> +\n> +\t/*\n> +\t * Do what unpack_callback() and unpack_nondirectories() normally\n> +\t * do. But we walk all paths recursively in just one loop instead.\n> +\t *\n> +\t * D/F conflicts and staged entries are not a concern because\n> +\t * cache-tree would be invalidated and we would never get here\n> +\t * in the first place.\n> +\t */\n> +\tfor (i = 0; i < nr_entries; i++) {\n> +\t\tstruct cache_entry *tree_ce;\n> +\t\tint len, rc;\n> +\n> +\t\tsrc[0] = o->src_index->cache[pos + i];\n> +\n> +\t\tlen = ce_namelen(src[0]);\n> +\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n> +\n> +\t\ttree_ce->ce_mode = src[0]->ce_mode;\n> +\t\ttree_ce->ce_flags = create_ce_flags(0);\n> +\t\ttree_ce->ce_namelen = len;\n> +\t\toidcpy(&tree_ce->oid, &src[0]->oid);\n> +\t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n> +\n\nI don't like the overhead of having to create an entirely new cache \nentry here but I see you clean this up in the next patch.\n\n> +\t\tfor (d = 1; d <= nr_names; d++)\n> +\t\t\tsrc[d] = tree_ce;\n> +\n> +\t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n> +\t\tfree(tree_ce);\n> +\t\tif (rc < 0)\n> +\t\t\treturn rc;\n> +\n> +\t\tmark_ce_used(src[0], o);\n> +\t}\n> +\tif (o->debug_unpack)\n> +\t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n> +\t\t       nr_entries,\n> +\t\t       o->src_index->cache[pos]->name,\n> +\t\t       o->src_index->cache[pos + nr_entries - 1]->name);\n> +\treturn 0;\n> +}\n> +\n>   static int traverse_trees_recursive(int n, unsigned long dirmask,\n>   \t\t\t\t    unsigned long df_conflicts,\n>   \t\t\t\t    struct name_entry *names,\n> @@ -655,6 +751,17 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n>   \tvoid *buf[MAX_UNPACK_TREES];\n>   \tstruct traverse_info newinfo;\n>   \tstruct name_entry *p;\n> +\tint nr_entries;\n> +\n> +\tnr_entries = all_trees_same_as_cache_tree(n, dirmask, names, info);\n> +\tif (nr_entries > 0) {\n> +\t\tstruct unpack_trees_options *o = info->data;\n> +\t\tint pos = index_pos_by_traverse_info(names, info);\n> +\n> +\t\tif (!o->merge || df_conflicts)\n> +\t\t\tBUG(\"Wrong condition to get here buddy\");\n> +\t\treturn traverse_by_cache_tree(pos, nr_entries, n, names, info);\n> +\t}\n>   \n>   \tp = names;\n>   \twhile (!p->mode)\n> @@ -814,6 +921,11 @@ static struct cache_entry *create_ce_entry(const struct traverse_info *info, con\n>   \treturn ce;\n>   }\n>   \n> +/*\n> + * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n\ns/funciton/function\n\n> + * without actually calling it. If you change the logic here you may need to\n> + * check and change there as well.\n> + */\n>   static int unpack_nondirectories(int n, unsigned long mask,\n>   \t\t\t\t unsigned long dirmask,\n>   \t\t\t\t struct cache_entry **src,\n> @@ -998,6 +1110,11 @@ static void debug_unpack_callback(int n,\n>   \t\tdebug_name_entry(i, names + i);\n>   }\n>   \n> +/*\n> + * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n> + * without actually calling it. If you change the logic here you may need to\n> + * check and change there as well.\n> + */\n>   static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n>   {\n>   \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n> @@ -1280,7 +1397,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>   \tuint64_t start = getnanotime();\n>   \n>   \tif (len > MAX_UNPACK_TREES)\n> -\t\tdie(\"unpack_trees takes at most %d trees\", MAX_UNPACK_TREES);\n> +\t\tdie(_(\"unpack_trees takes at most %d trees\"), MAX_UNPACK_TREES);\n>   \n\nI'd like to see this get in independently of this patch series.\n\n>   \tmemset(&el, 0, sizeof(el));\n>   \tif (!core_apply_sparse_checkout || !o->update)\n> \n"},{"id":"353957","messageId":"0cbba28a-1820-63bf-b31c-cb3774d7d5bd@gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-4-pclouds@gmail.com","subject":"Re: [PATCH v2 3/4] unpack-trees: reduce malloc in cache-tree walk","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-30T20:58:28Z","receivedAt":"2018-07-30T20:58:33Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/29/2018 6:33 AM, Nguyễn Thái Ngọc Duy wrote:\n> This is a micro optimization that probably only shines on repos with\n> deep directory structure. Instead of allocating and freeing a new\n> cache_entry in every iteration, we reuse the last one and only update\n> the parts that are new each iteration.\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>   unpack-trees.c | 29 ++++++++++++++++++++---------\n>   1 file changed, 20 insertions(+), 9 deletions(-)\n> \n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index 39566b28fb..c33ebaf001 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -694,6 +694,8 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n>   {\n>   \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n>   \tstruct unpack_trees_options *o = info->data;\n> +\tstruct cache_entry *tree_ce = NULL;\n> +\tint ce_len = 0;\n>   \tint i, d;\n>   \n>   \tif (!o->merge)\n> @@ -708,30 +710,39 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n>   \t * in the first place.\n>   \t */\n>   \tfor (i = 0; i < nr_entries; i++) {\n> -\t\tstruct cache_entry *tree_ce;\n> -\t\tint len, rc;\n> +\t\tint new_ce_len, len, rc;\n>   \n>   \t\tsrc[0] = o->src_index->cache[pos + i];\n>   \n>   \t\tlen = ce_namelen(src[0]);\n> -\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n> +\t\tnew_ce_len = cache_entry_size(len);\n> +\n> +\t\tif (new_ce_len > ce_len) {\n> +\t\t\tnew_ce_len <<= 1;\n> +\t\t\ttree_ce = xrealloc(tree_ce, new_ce_len);\n> +\t\t\tmemset(tree_ce, 0, new_ce_len);\n> +\t\t\tce_len = new_ce_len;\n> +\n> +\t\t\ttree_ce->ce_flags = create_ce_flags(0);\n> +\n> +\t\t\tfor (d = 1; d <= nr_names; d++)\n> +\t\t\t\tsrc[d] = tree_ce;\n> +\t\t}\n\nNice optimization - especially when there are a lot of cache entries and \nlarge trees.\n\n>   \n>   \t\ttree_ce->ce_mode = src[0]->ce_mode;\n> -\t\ttree_ce->ce_flags = create_ce_flags(0);\n>   \t\ttree_ce->ce_namelen = len;\n>   \t\toidcpy(&tree_ce->oid, &src[0]->oid);\n>   \t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n>   \n> -\t\tfor (d = 1; d <= nr_names; d++)\n> -\t\t\tsrc[d] = tree_ce;\n> -\n>   \t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n> -\t\tfree(tree_ce);\n> -\t\tif (rc < 0)\n> +\t\tif (rc < 0) {\n> +\t\t\tfree(tree_ce);\n>   \t\t\treturn rc;\n> +\t\t}\n>   \n>   \t\tmark_ce_used(src[0], o);\n>   \t}\n> +\tfree(tree_ce);\n>   \tif (o->debug_unpack)\n>   \t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n>   \t\t       nr_entries,\n> \n"},{"id":"353960","messageId":"dfe67a11-8b0b-682e-ff6c-4341808339bc@gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-1-pclouds@gmail.com","subject":"Re: [PATCH v2 0/4] Speed up unpack_trees()","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-30T21:04:10Z","receivedAt":"2018-07-30T21:04:14Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/29/2018 6:33 AM, Nguyễn Thái Ngọc Duy wrote:\n> This series speeds up unpack_trees() a bit by using cache-tree.\n> unpack-trees could bit split in three big parts\n> \n> - the actual tree unpacking and running n-way merging\n> - update worktree, which could be expensive depending on how much I/O\n>    is involved\n> - repair cache-tree\n> \n> This series focuses on the first part alone and could give 700%\n> speedup (best case possible scenario, real life ones probably not that\n> impressive).\n> \n> It also shows that the reparing cache-tree is kinda expensive. I have\n> an idea of reusing cache-tree from the original index, but I'll leave\n> that to Ben or others to try out and see if it helps at all.\n> \n> v2 fixes the comments from Junio, adds more performance tracing and\n> reduces the cost of adding index entries.\n> \n> Nguyễn Thái Ngọc Duy (4):\n>    unpack-trees.c: add performance tracing\n>    unpack-trees: optimize walking same trees with cache-tree\n>    unpack-trees: reduce malloc in cache-tree walk\n>    unpack-trees: cheaper index update when walking by cache-tree\n> \n>   cache-tree.c   |   2 +\n>   cache.h        |   1 +\n>   read-cache.c   |   3 +-\n>   unpack-trees.c | 161 ++++++++++++++++++++++++++++++++++++++++++++++++-\n>   unpack-trees.h |   1 +\n>   5 files changed, 166 insertions(+), 2 deletions(-)\n> \n\nI have a limited understanding of this code path so I'm not the best \nperson to review this but I didn't see any issues that concerned me.  I \nalso was able to run our internal functional and performance tests in \naddition to the git tests and the results were positive.\n\nBen\n"},{"id":"354045","messageId":"CACsJy8BUBjPngHz=icHomor-LJOkMLwZ9bQ6YJDxnoXGg++vjg@mail.gmail.com","threadId":"48913","inReplyTo":"9a9a309c-7143-e642-cfd8-6df76e77995a@gmail.com","subject":"Re: [PATCH v2 0/4] Speed up unpack_trees()","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-07-31T15:31:45Z","receivedAt":"2018-07-31T15:32:14Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Jul 30, 2018 at 8:10 PM Ben Peart <peartben@gmail.com> wrote:\n> I ran \"git checkout\" on a large repo and averaged the results of 3 runs.\n>   This clearly demonstrates the benefit of the optimized unpack_trees()\n> as even the final \"diff-index\" is essentially a 3rd call to unpack_trees().\n>\n> baseline        new\n> ----------------------------------------------------------------------\n> 0.535510167     0.556558733     s: read cache .git/index\n> 0.3057373       0.3147105       s: initialize name hash\n> 0.0184082       0.023558433     s: preload index\n> 0.086910967     0.089085967     s: refresh index\n> 7.889590767     2.191554433     s: unpack trees\n> 0.120760833     0.131941267     s: update worktree after a merge\n> 2.2583504       2.572663167     s: repair cache-tree\n> 0.8916137       0.959495233     s: write index, changed mask = 28\n> 3.405199233     0.2710663       s: unpack trees\n> 0.000999667     0.0021554       s: update worktree after a merge\n> 3.4063306       0.273318333     s: diff-index\n> 16.9524923      9.462943133     s: git command:\n> 'c:\\git-sdk-64\\usr\\src\\git\\git.exe' checkout\n>\n> The first call to unpack_trees() saves 72%\n> The 2nd and 3rd call save 92%\n\nBy the 3rd I guess you meant \"diff-index\" line. I think it's the same\nwith the second call. diff-index triggers the second unpack-trees but\nthere's no indent here and it's misleading to read this as diff-index\nand unpack-trees execute one after the other.\n\n> Total time savings for the entire command was 44%\n\nWow.. I guess you have more trees since I could only save 30% on gcc.git.\n\n> In the performance game of whack-a-mole, that call to repair cache-tree\n> is now looking quite expensive...\n\nYeah and I think we can whack that mole too. I did some measurement.\nBest case possible, we just need to scan through two indexes (one with\nmany good cache-tree, one with no cache-tree), compare and copy\ncache-tree over. The scanning takes like 1% time of current repair\nstep and I suspect it's the hashing that takes most of the time. Of\ncourse real world won't have such nice numbers, but I guess we could\nmaybe half cache-tree update/repair time.\n-- \nDuy\n"},{"id":"354058","messageId":"cc3c4dbb-d545-6a6c-b20e-6a8ca66fc210@gmail.com","threadId":"48913","inReplyTo":"CACsJy8BUBjPngHz=icHomor-LJOkMLwZ9bQ6YJDxnoXGg++vjg@mail.gmail.com","subject":"Re: [PATCH v2 0/4] Speed up unpack_trees()","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-31T16:50:48Z","receivedAt":"2018-07-31T16:50:54Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/31/2018 11:31 AM, Duy Nguyen wrote:\n> On Mon, Jul 30, 2018 at 8:10 PM Ben Peart <peartben@gmail.com> wrote:\n>> I ran \"git checkout\" on a large repo and averaged the results of 3 runs.\n>>    This clearly demonstrates the benefit of the optimized unpack_trees()\n>> as even the final \"diff-index\" is essentially a 3rd call to unpack_trees().\n>>\n>> baseline        new\n>> ----------------------------------------------------------------------\n>> 0.535510167     0.556558733     s: read cache .git/index\n>> 0.3057373       0.3147105       s: initialize name hash\n>> 0.0184082       0.023558433     s: preload index\n>> 0.086910967     0.089085967     s: refresh index\n>> 7.889590767     2.191554433     s: unpack trees\n>> 0.120760833     0.131941267     s: update worktree after a merge\n>> 2.2583504       2.572663167     s: repair cache-tree\n>> 0.8916137       0.959495233     s: write index, changed mask = 28\n>> 3.405199233     0.2710663       s: unpack trees\n>> 0.000999667     0.0021554       s: update worktree after a merge\n>> 3.4063306       0.273318333     s: diff-index\n>> 16.9524923      9.462943133     s: git command:\n>> 'c:\\git-sdk-64\\usr\\src\\git\\git.exe' checkout\n>>\n>> The first call to unpack_trees() saves 72%\n>> The 2nd and 3rd call save 92%\n> \n> By the 3rd I guess you meant \"diff-index\" line. I think it's the same\n> with the second call. diff-index triggers the second unpack-trees but\n> there's no indent here and it's misleading to read this as diff-index\n> and unpack-trees execute one after the other.\n> \n>> Total time savings for the entire command was 44%\n> \n> Wow.. I guess you have more trees since I could only save 30% on gcc.git.\n\nYes, with over 500K trees, this optimization really pays off for us.  I \ncan't wait to see how this works out in the wild (vs my \"lab\" based \nperformance testing).\n\nThank you!  I definitely owe you lunch. :)\n\n> \n>> In the performance game of whack-a-mole, that call to repair cache-tree\n>> is now looking quite expensive...\n> \n> Yeah and I think we can whack that mole too. I did some measurement.\n> Best case possible, we just need to scan through two indexes (one with\n> many good cache-tree, one with no cache-tree), compare and copy\n> cache-tree over. The scanning takes like 1% time of current repair\n> step and I suspect it's the hashing that takes most of the time. Of\n> course real world won't have such nice numbers, but I guess we could\n> maybe half cache-tree update/repair time.\n> \n\nI have some great profiling tools available so will take a look at this \nnext and see exactly where the time is being spent.\n"},{"id":"354068","messageId":"57d146a2-9bf8-66c9-9cb4-c05f93b63319@gmail.com","threadId":"48913","inReplyTo":"cc3c4dbb-d545-6a6c-b20e-6a8ca66fc210@gmail.com","subject":"Re: [PATCH v2 0/4] Speed up unpack_trees()","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-07-31T17:31:31Z","receivedAt":"2018-07-31T17:31:36Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 7/31/2018 12:50 PM, Ben Peart wrote:\n> \n> \n> On 7/31/2018 11:31 AM, Duy Nguyen wrote:\n\n>>\n>>> In the performance game of whack-a-mole, that call to repair cache-tree\n>>> is now looking quite expensive...\n>>\n>> Yeah and I think we can whack that mole too. I did some measurement.\n>> Best case possible, we just need to scan through two indexes (one with\n>> many good cache-tree, one with no cache-tree), compare and copy\n>> cache-tree over. The scanning takes like 1% time of current repair\n>> step and I suspect it's the hashing that takes most of the time. Of\n>> course real world won't have such nice numbers, but I guess we could\n>> maybe half cache-tree update/repair time.\n>>\n> \n> I have some great profiling tools available so will take a look at this \n> next and see exactly where the time is being spent.\n\nGood instincts.  In cache_tree_update, the heavy hitter is definitely \nhash_object_file followed by has_object_file.\n\nName                               \tInc %\t     Inc\n+ git!cache_tree_update            \t 12.4\t   4,935\n|+ git!update_one                  \t 11.8\t   4,706\n| + git!update_one                 \t 11.8\t   4,706\n|  + git!hash_object_file          \t  6.1\t   2,406\n|  + git!has_object_file           \t  2.0\t     813\n|  + OTHER <<vcruntime140d!strchr>>\t  0.5\t     203\n|  + git!strbuf_addf               \t  0.4\t     155\n|  + git!strbuf_release            \t  0.4\t     143\n|  + git!strbuf_add                \t  0.3\t     121\n|  + OTHER <<vcruntime140d!memcmp>>\t  0.2\t      93\n|  + git!strbuf_grow               \t  0.1\t      25\n"},{"id":"354168","messageId":"20180801163830.GA31968@duynguyen.home","threadId":"48913","inReplyTo":"57d146a2-9bf8-66c9-9cb4-c05f93b63319@gmail.com","subject":"Re: [PATCH v2 0/4] Speed up unpack_trees()","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-01T16:38:30Z","receivedAt":"2018-08-01T16:38:36Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Jul 31, 2018 at 01:31:31PM -0400, Ben Peart wrote:\n> \n> \n> On 7/31/2018 12:50 PM, Ben Peart wrote:\n> > \n> > \n> > On 7/31/2018 11:31 AM, Duy Nguyen wrote:\n> \n> >>\n> >>> In the performance game of whack-a-mole, that call to repair cache-tree\n> >>> is now looking quite expensive...\n> >>\n> >> Yeah and I think we can whack that mole too. I did some measurement.\n> >> Best case possible, we just need to scan through two indexes (one with\n> >> many good cache-tree, one with no cache-tree), compare and copy\n> >> cache-tree over. The scanning takes like 1% time of current repair\n> >> step and I suspect it's the hashing that takes most of the time. Of\n> >> course real world won't have such nice numbers, but I guess we could\n> >> maybe half cache-tree update/repair time.\n> >>\n> > \n> > I have some great profiling tools available so will take a look at this \n> > next and see exactly where the time is being spent.\n> \n> Good instincts.  In cache_tree_update, the heavy hitter is definitely \n> hash_object_file followed by has_object_file.\n> \n> Name                               \tInc %\t     Inc\n> + git!cache_tree_update            \t 12.4\t   4,935\n> |+ git!update_one                  \t 11.8\t   4,706\n> | + git!update_one                 \t 11.8\t   4,706\n> |  + git!hash_object_file          \t  6.1\t   2,406\n> |  + git!has_object_file           \t  2.0\t     813\n> |  + OTHER <<vcruntime140d!strchr>>\t  0.5\t     203\n> |  + git!strbuf_addf               \t  0.4\t     155\n> |  + git!strbuf_release            \t  0.4\t     143\n> |  + git!strbuf_add                \t  0.3\t     121\n> |  + OTHER <<vcruntime140d!memcmp>>\t  0.2\t      93\n> |  + git!strbuf_grow               \t  0.1\t      25\n\nBen, if you work on this, this could be a good starting point. I will\nnot work on this because I still have some other things to catch up\nand follow through. You can have my sign off if you reuse something\nfrom this patch\n\nEven if it's a naive implementation, the initial numbers look pretty\ngood. Without the patch we have\n\n18:31:05.970621 unpack-trees.c:1437     performance: 0.000001029 s: copy\n18:31:05.975729 unpack-trees.c:1444     performance: 0.005082004 s: update\n\nAnd with the patch\n\n18:31:13.295655 unpack-trees.c:1437     performance: 0.000198017 s: copy\n18:31:13.296757 unpack-trees.c:1444     performance: 0.001075935 s: update\n\nTime saving is about 80% by the look of this (best possible case\nbecause only the top tree needs to be hashed and written out).\n\n-- 8< --\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 6b46711996..67a4a93100 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -440,6 +440,147 @@ int cache_tree_update(struct index_state *istate, int flags)\n \treturn 0;\n }\n \n+static int same(const struct cache_entry *a, const struct cache_entry *b)\n+{\n+\tif (ce_stage(a) || ce_stage(b))\n+\t\treturn 0;\n+\tif ((a->ce_flags | b->ce_flags) & CE_CONFLICTED)\n+\t\treturn 0;\n+\treturn a->ce_mode == b->ce_mode &&\n+\t       !oidcmp(&a->oid, &b->oid);\n+}\n+\n+static int cache_tree_name_pos(const struct index_state *istate,\n+\t\t\t       const struct strbuf *path)\n+{\n+\tint pos;\n+\n+\tif (!path->len)\n+\t\treturn 0;\n+\n+\tpos = index_name_pos(istate, path->buf, path->len);\n+\tif (pos >= 0)\n+\t\tBUG(\"No no no, directory path must not exist in index\");\n+\treturn -pos - 1;\n+}\n+\n+/*\n+ * Locate the same cache-tree in two separate indexes. Check the\n+ * cache-tree is still valid for the \"to\" index (i.e. it contains the\n+ * same set of entries in the \"from\" index).\n+ */\n+static int verify_one_cache_tree(const struct index_state *to,\n+\t\t\t\t const struct index_state *from,\n+\t\t\t\t const struct cache_tree *it,\n+\t\t\t\t const struct strbuf *path)\n+{\n+\tint i, spos, dpos;\n+\n+\tspos = cache_tree_name_pos(from, path);\n+\tif (spos + it->entry_count > from->cache_nr)\n+\t\treturn -1;\n+\n+\tdpos = cache_tree_name_pos(to, path);\n+\tif (dpos + it->entry_count > to->cache_nr)\n+\t\treturn -1;\n+\n+\t/* Can we quickly check head and tail and bail out early */\n+\tif (!same(from->cache[spos], to->cache[spos]) ||\n+\t    !same(from->cache[spos + it->entry_count - 1],\n+\t\t  to->cache[spos + it->entry_count - 1]))\n+\t\treturn -1;\n+\n+\tfor (i = 1; i < it->entry_count - 1; i++)\n+\t\tif (!same(from->cache[spos + i],\n+\t\t\t  to->cache[dpos + i]))\n+\t\t\treturn -1;\n+\n+\treturn 0;\n+}\n+\n+static int verify_and_invalidate(struct index_state *to,\n+\t\t\t\t const struct index_state *from,\n+\t\t\t\t struct cache_tree *it,\n+\t\t\t\t struct strbuf *path)\n+{\n+\t/*\n+\t * Optimistically verify the current tree first. Alternatively\n+\t * we could verify all the subtrees first then do this\n+\t * last. Any invalid subtree would also invalidates its\n+\t * ancestors.\n+\t */\n+\tif (it->entry_count != -1 &&\n+\t    verify_one_cache_tree(to, from, it, path))\n+\t\tit->entry_count = -1;\n+\n+\t/*\n+\t * If the current tree is valid, don't bother checking\n+\t * inside. All subtrees _should_ also be valid\n+\t */\n+\tif (it->entry_count == -1) {\n+\t\tint i, len = path->len;\n+\n+\t\tfor (i = 0; i < it->subtree_nr; i++) {\n+\t\t\tstruct cache_tree_sub *down = it->down[i];\n+\n+\t\t\tif (!down || !down->cache_tree)\n+\t\t\t\tcontinue;\n+\n+\t\t\tstrbuf_setlen(path, len);\n+\t\t\tstrbuf_add(path, down->name, down->namelen);\n+\t\t\tstrbuf_addch(path, '/');\n+\t\t\tif (verify_and_invalidate(to, from,\n+\t\t\t\t\t\t  down->cache_tree, path))\n+\t\t\t\treturn -1;\n+\t\t}\n+\t\tstrbuf_setlen(path, len);\n+\t}\n+\treturn 0;\n+}\n+\n+static struct cache_tree *duplicate_cache_tree(const struct cache_tree *src)\n+{\n+\tstruct cache_tree *dst;\n+\tint i;\n+\n+\tif (!src)\n+\t\treturn NULL;\n+\n+\tdst = xmalloc(sizeof(*dst));\n+\tdst->entry_count = src->entry_count;\n+\toidcpy(&dst->oid, &src->oid);\n+\tdst->subtree_nr = src->subtree_nr;\n+\tdst->subtree_alloc = dst->subtree_nr;\n+\tALLOC_ARRAY(dst->down, dst->subtree_alloc);\n+\tfor (i = 0; i < src->subtree_nr; i++) {\n+\t\tstruct cache_tree_sub *dsrc = src->down[i];\n+\t\tstruct cache_tree_sub *down;\n+\n+\t\tFLEX_ALLOC_MEM(down, name, dsrc->name, dsrc->namelen);\n+\t\tdown->count = dsrc->count;\n+\t\tdown->namelen = dsrc->namelen;\n+\t\tdown->used = dsrc->used;\n+\t\tdown->cache_tree = duplicate_cache_tree(dsrc->cache_tree);\n+\t\tdst->down[i] = down;\n+\t}\n+\treturn dst;\n+}\n+\n+int cache_tree_copy(struct index_state *to, const struct index_state *from)\n+{\n+\tstruct cache_tree *it = duplicate_cache_tree(from->cache_tree);\n+\tstruct strbuf path = STRBUF_INIT;\n+\tint ret;\n+\n+\tif (to->cache_tree)\n+\t\tBUG(\"Sorry merging cache-tree is not supported yet\");\n+\tret = verify_and_invalidate(to, from, it, &path);\n+\tto->cache_tree = it;\n+\tto->cache_changed |= CACHE_TREE_CHANGED;\n+\tstrbuf_release(&path);\n+\treturn ret;\n+}\n+\n static void write_one(struct strbuf *buffer, struct cache_tree *it,\n                       const char *path, int pathlen)\n {\ndiff --git a/cache-tree.h b/cache-tree.h\nindex cfd5328cc9..6981da8e0d 100644\n--- a/cache-tree.h\n+++ b/cache-tree.h\n@@ -53,4 +53,6 @@ void prime_cache_tree(struct index_state *, struct tree *);\n \n extern int cache_tree_matches_traversal(struct cache_tree *, struct name_entry *ent, struct traverse_info *info);\n \n+int cache_tree_copy(struct index_state *to, const struct index_state *from);\n+\n #endif\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex cd0680f11e..cb3fdd42a6 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -1427,12 +1427,22 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \tret = check_updates(o) ? (-2) : 0;\n \tif (o->dst_index) {\n \t\tif (!ret) {\n-\t\t\tif (!o->result.cache_tree)\n+\t\t\tif (!o->result.cache_tree) {\n+\t\t\t\tuint64_t start = getnanotime();\n+#if 0\n \t\t\t\to->result.cache_tree = cache_tree();\n-\t\t\tif (!cache_tree_fully_valid(o->result.cache_tree))\n+#else\n+\t\t\t\tcache_tree_copy(&o->result, o->src_index);\n+#endif\n+\t\t\t\ttrace_performance_since(start, \"copy\");\n+\t\t\t}\n+\t\t\tif (!cache_tree_fully_valid(o->result.cache_tree)) {\n+\t\t\t\tuint64_t start = getnanotime();\n \t\t\t\tcache_tree_update(&o->result,\n \t\t\t\t\t\t  WRITE_TREE_SILENT |\n \t\t\t\t\t\t  WRITE_TREE_REPAIR);\n+\t\t\t\ttrace_performance_since(start, \"update\");\n+\t\t\t}\n \t\t}\n \t\tmove_index_extensions(&o->result, o->src_index);\n \t\tdiscard_index(o->dst_index);\n-- 8< --\n"},{"id":"354474","messageId":"20180804053723.4695-1-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-1-pclouds@gmail.com","subject":"[PATCH v3 0/4] Speed up unpack_trees()","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-04T05:37:19Z","receivedAt":"2018-08-04T05:37:35Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This is a minor update to address Ben's comments and add his\nmeasurements in the commit message of 2/4 for the record.\n\nI've also checked about the lookahead thing in unpack_trees() to see\nif we accidentally break something there, which is my biggest worry.\nSee [1] and [2] for context, but I believe since we can't have D/F\nconflicts, the situation where lookahead is needed will not occur. So\nwe should be safe.\n\n[1] da165f470e (unpack-trees.c: prepare for looking ahead in the index - 2010-01-07)\n[2] 730f72840c (unpack-trees.c: look ahead in the index - 2009-09-20)\n\nrange-diff:\n\n1:  789f7e2872 ! 1:  05eb762d2d unpack-trees.c: add performance tracing\n    @@ -1,6 +1,6 @@\n     Author: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n     \n    -    unpack-trees.c: add performance tracing\n    +    unpack-trees: add performance tracing\n     \n         We're going to optimize unpack_trees() a bit in the following\n         patches. Let's add some tracing to measure how long it takes before\n2:  589bed1366 ! 2:  02286ad123 unpack-trees: optimize walking same trees with cache-tree\n    @@ -32,6 +32,24 @@\n             0.111793866   0.032933140 s: diff-index\n             0.587933288   0.398924370 s: git command: /home/pclouds/w/git/git\n     \n    +    Another measurement from Ben's running \"git checkout\" with over 500k\n    +    trees (on the whole series):\n    +\n    +        baseline        new\n    +      ----------------------------------------------------------------------\n    +        0.535510167     0.556558733     s: read cache .git/index\n    +        0.3057373       0.3147105       s: initialize name hash\n    +        0.0184082       0.023558433     s: preload index\n    +        0.086910967     0.089085967     s: refresh index\n    +        7.889590767     2.191554433     s: unpack trees\n    +        0.120760833     0.131941267     s: update worktree after a merge\n    +        2.2583504       2.572663167     s: repair cache-tree\n    +        0.8916137       0.959495233     s: write index, changed mask = 28\n    +        3.405199233     0.2710663       s: unpack trees\n    +        0.000999667     0.0021554       s: update worktree after a merge\n    +        3.4063306       0.273318333     s: diff-index\n    +        16.9524923      9.462943133     s: git command: git.exe checkout\n    +\n         This command calls unpack_trees() twice, the first time on 2way merge\n         and the second 1way merge. In both times, \"unpack trees\" time is\n         reduced to one third. Overall time reduction is not that impressive of\n    @@ -39,7 +57,6 @@\n         repair cache-tree line.\n     \n         Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n    -    Signed-off-by: Junio C Hamano <gitster@pobox.com>\n     \n     diff --git a/unpack-trees.c b/unpack-trees.c\n     --- a/unpack-trees.c\n    @@ -170,7 +187,7 @@\n      }\n      \n     +/*\n    -+ * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n    ++ * Note that traverse_by_cache_tree() duplicates some logic in this function\n     + * without actually calling it. If you change the logic here you may need to\n     + * check and change there as well.\n     + */\n    @@ -189,12 +206,3 @@\n      static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n      {\n      \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n    -@@\n    - \tuint64_t start = getnanotime();\n    - \n    - \tif (len > MAX_UNPACK_TREES)\n    --\t\tdie(\"unpack_trees takes at most %d trees\", MAX_UNPACK_TREES);\n    -+\t\tdie(_(\"unpack_trees takes at most %d trees\"), MAX_UNPACK_TREES);\n    - \n    - \tmemset(&el, 0, sizeof(el));\n    - \tif (!core_apply_sparse_checkout || !o->update)\n3:  7c6f863fc0 = 3:  c87b82ffee unpack-trees: reduce malloc in cache-tree walk\n4:  6ca17b1138 ! 4:  e791cdfc82 unpack-trees: cheaper index update when walking by cache-tree\n    @@ -40,7 +40,6 @@\n         attempt should be on that \"repair cache-tree\" line.\n     \n         Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n    -    Signed-off-by: Junio C Hamano <gitster@pobox.com>\n     \n     diff --git a/cache.h b/cache.h\n     --- a/cache.h\n    @@ -119,20 +118,6 @@\n      \tfree(tree_ce);\n      \tif (o->debug_unpack)\n      \t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n    -@@\n    - \t\tif (!ret) {\n    - \t\t\tif (!o->result.cache_tree)\n    - \t\t\t\to->result.cache_tree = cache_tree();\n    -+\t\t\t/*\n    -+\t\t\t * TODO: Walk o.src_index->cache_tree, quickly check\n    -+\t\t\t * if o->result.cache has the exact same content for\n    -+\t\t\t * any valid cache-tree in o.src_index, then we can\n    -+\t\t\t * just copy the cache-tree over instead of hashing a\n    -+\t\t\t * new tree object.\n    -+\t\t\t */\n    - \t\t\tif (!cache_tree_fully_valid(o->result.cache_tree))\n    - \t\t\t\tcache_tree_update(&o->result,\n    - \t\t\t\t\t\t  WRITE_TREE_SILENT |\n     \n     diff --git a/unpack-trees.h b/unpack-trees.h\n     --- a/unpack-trees.h\n\nNguyễn Thái Ngọc Duy (4):\n  unpack-trees: add performance tracing\n  unpack-trees: optimize walking same trees with cache-tree\n  unpack-trees: reduce malloc in cache-tree walk\n  unpack-trees: cheaper index update when walking by cache-tree\n\n cache-tree.c   |   2 +\n cache.h        |   1 +\n read-cache.c   |   3 +-\n unpack-trees.c | 152 +++++++++++++++++++++++++++++++++++++++++++++++++\n unpack-trees.h |   1 +\n 5 files changed, 158 insertions(+), 1 deletion(-)\n\n-- \n2.18.0.656.gda699b98b3\n"},{"id":"354475","messageId":"20180804053723.4695-2-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180804053723.4695-1-pclouds@gmail.com","subject":"[PATCH v3 1/4] unpack-trees: add performance tracing","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-04T05:37:20Z","receivedAt":"2018-08-04T05:37:38Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"We're going to optimize unpack_trees() a bit in the following\npatches. Let's add some tracing to measure how long it takes before\nand after. This is the baseline (\"git checkout -\" on gcc.git, 80k\nfiles on worktree)\n\n    0.018239226 s: read cache .git/index\n    0.052541655 s: preload index\n    0.001537598 s: refresh index\n    0.168167768 s: unpack trees\n    0.002897186 s: update worktree after a merge\n    0.131661745 s: repair cache-tree\n    0.075389117 s: write index, changed mask = 2a\n    0.111702023 s: unpack trees\n    0.000023245 s: update worktree after a merge\n    0.111793866 s: diff-index\n    0.587933288 s: git command: /home/pclouds/w/git/git checkout -\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n cache-tree.c   | 2 ++\n unpack-trees.c | 4 ++++\n 2 files changed, 6 insertions(+)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 6b46711996..0dbe10fc85 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -426,6 +426,7 @@ static int update_one(struct cache_tree *it,\n \n int cache_tree_update(struct index_state *istate, int flags)\n {\n+\tuint64_t start = getnanotime();\n \tstruct cache_tree *it = istate->cache_tree;\n \tstruct cache_entry **cache = istate->cache;\n \tint entries = istate->cache_nr;\n@@ -437,6 +438,7 @@ int cache_tree_update(struct index_state *istate, int flags)\n \tif (i < 0)\n \t\treturn i;\n \tistate->cache_changed |= CACHE_TREE_CHANGED;\n+\ttrace_performance_since(start, \"repair cache-tree\");\n \treturn 0;\n }\n \ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex cd0680f11e..a32ddee159 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -352,6 +352,7 @@ static int check_updates(struct unpack_trees_options *o)\n \tstruct progress *progress = NULL;\n \tstruct index_state *index = &o->result;\n \tstruct checkout state = CHECKOUT_INIT;\n+\tuint64_t start = getnanotime();\n \tint i;\n \n \tstate.force = 1;\n@@ -423,6 +424,7 @@ static int check_updates(struct unpack_trees_options *o)\n \terrs |= finish_delayed_checkout(&state);\n \tif (o->update)\n \t\tgit_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n+\ttrace_performance_since(start, \"update worktree after a merge\");\n \treturn errs != 0;\n }\n \n@@ -1275,6 +1277,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \tint i, ret;\n \tstatic struct cache_entry *dfc;\n \tstruct exclude_list el;\n+\tuint64_t start = getnanotime();\n \n \tif (len > MAX_UNPACK_TREES)\n \t\tdie(\"unpack_trees takes at most %d trees\", MAX_UNPACK_TREES);\n@@ -1423,6 +1426,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t\tgoto done;\n \t\t}\n \t}\n+\ttrace_performance_since(start, \"unpack trees\");\n \n \tret = check_updates(o) ? (-2) : 0;\n \tif (o->dst_index) {\n-- \n2.18.0.656.gda699b98b3\n\n"},{"id":"354476","messageId":"20180804053723.4695-3-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180804053723.4695-1-pclouds@gmail.com","subject":"[PATCH v3 2/4] unpack-trees: optimize walking same trees with cache-tree","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-04T05:37:21Z","receivedAt":"2018-08-04T05:37:38Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"From: Duy Nguyen <pclouds@gmail.com>\n\nIn order to merge one or many trees with the index, unpack-trees code\nwalks multiple trees in parallel with the index and performs n-way\nmerge. If we find out at start of a directory that all trees are the\nsame (by comparing OID) and cache-tree happens to be available for\nthat directory as well, we could avoid walking the trees because we\nalready know what these trees contain: it's flattened in what's called\n\"the index\".\n\nThe upside is of course a lot less I/O since we can potentially skip\nlots of trees (think subtrees). We also save CPU because we don't have\nto inflate and the apply deltas. The downside is of course more\nfragile code since the logic in some functions are now duplicated\nelsewhere.\n\n\"checkout -\" with this patch on gcc.git:\n\n    baseline      new\n  --------------------------------------------------------------------\n    0.018239226   0.019365414 s: read cache .git/index\n    0.052541655   0.049605548 s: preload index\n    0.001537598   0.001571695 s: refresh index\n    0.168167768   0.049677212 s: unpack trees\n    0.002897186   0.002845256 s: update worktree after a merge\n    0.131661745   0.136597522 s: repair cache-tree\n    0.075389117   0.075422517 s: write index, changed mask = 2a\n    0.111702023   0.032813253 s: unpack trees\n    0.000023245   0.000022002 s: update worktree after a merge\n    0.111793866   0.032933140 s: diff-index\n    0.587933288   0.398924370 s: git command: /home/pclouds/w/git/git\n\nAnother measurement from Ben's running \"git checkout\" with over 500k\ntrees (on the whole series):\n\n    baseline        new\n  ----------------------------------------------------------------------\n    0.535510167     0.556558733     s: read cache .git/index\n    0.3057373       0.3147105       s: initialize name hash\n    0.0184082       0.023558433     s: preload index\n    0.086910967     0.089085967     s: refresh index\n    7.889590767     2.191554433     s: unpack trees\n    0.120760833     0.131941267     s: update worktree after a merge\n    2.2583504       2.572663167     s: repair cache-tree\n    0.8916137       0.959495233     s: write index, changed mask = 28\n    3.405199233     0.2710663       s: unpack trees\n    0.000999667     0.0021554       s: update worktree after a merge\n    3.4063306       0.273318333     s: diff-index\n    16.9524923      9.462943133     s: git command: git.exe checkout\n\nThis command calls unpack_trees() twice, the first time on 2way merge\nand the second 1way merge. In both times, \"unpack trees\" time is\nreduced to one third. Overall time reduction is not that impressive of\ncourse because index operations take a big chunk. And there's that\nrepair cache-tree line.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n unpack-trees.c | 117 +++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 117 insertions(+)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex a32ddee159..ba3d2e947e 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -644,6 +644,102 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n \treturn name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n }\n \n+static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n+\t\t\t\t\tstruct name_entry *names,\n+\t\t\t\t\tstruct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i;\n+\n+\tif (!o->merge || dirmask != ((1 << n) - 1))\n+\t\treturn 0;\n+\n+\tfor (i = 1; i < n; i++)\n+\t\tif (!are_same_oid(names, names + i))\n+\t\t\treturn 0;\n+\n+\treturn cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n+}\n+\n+static int index_pos_by_traverse_info(struct name_entry *names,\n+\t\t\t\t      struct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint len = traverse_path_len(info, names);\n+\tchar *name = xmalloc(len + 1 /* slash */ + 1 /* NUL */);\n+\tint pos;\n+\n+\tmake_traverse_path(name, info, names);\n+\tname[len++] = '/';\n+\tname[len] = '\\0';\n+\tpos = index_name_pos(o->src_index, name, len);\n+\tif (pos >= 0)\n+\t\tBUG(\"This is a directory and should not exist in index\");\n+\tpos = -pos - 1;\n+\tif (!starts_with(o->src_index->cache[pos]->name, name) ||\n+\t    (pos > 0 && starts_with(o->src_index->cache[pos-1]->name, name)))\n+\t\tBUG(\"pos must point at the first entry in this directory\");\n+\tfree(name);\n+\treturn pos;\n+}\n+\n+/*\n+ * Fast path if we detect that all trees are the same as cache-tree at this\n+ * path. We'll walk these trees recursively using cache-tree/index instead of\n+ * ODB since already know what these trees contain.\n+ */\n+static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n+\t\t\t\t  struct name_entry *names,\n+\t\t\t\t  struct traverse_info *info)\n+{\n+\tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i, d;\n+\n+\tif (!o->merge)\n+\t\tBUG(\"We need cache-tree to do this optimization\");\n+\n+\t/*\n+\t * Do what unpack_callback() and unpack_nondirectories() normally\n+\t * do. But we walk all paths recursively in just one loop instead.\n+\t *\n+\t * D/F conflicts and staged entries are not a concern because\n+\t * cache-tree would be invalidated and we would never get here\n+\t * in the first place.\n+\t */\n+\tfor (i = 0; i < nr_entries; i++) {\n+\t\tstruct cache_entry *tree_ce;\n+\t\tint len, rc;\n+\n+\t\tsrc[0] = o->src_index->cache[pos + i];\n+\n+\t\tlen = ce_namelen(src[0]);\n+\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n+\n+\t\ttree_ce->ce_mode = src[0]->ce_mode;\n+\t\ttree_ce->ce_flags = create_ce_flags(0);\n+\t\ttree_ce->ce_namelen = len;\n+\t\toidcpy(&tree_ce->oid, &src[0]->oid);\n+\t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n+\n+\t\tfor (d = 1; d <= nr_names; d++)\n+\t\t\tsrc[d] = tree_ce;\n+\n+\t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n+\t\tfree(tree_ce);\n+\t\tif (rc < 0)\n+\t\t\treturn rc;\n+\n+\t\tmark_ce_used(src[0], o);\n+\t}\n+\tif (o->debug_unpack)\n+\t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n+\t\t       nr_entries,\n+\t\t       o->src_index->cache[pos]->name,\n+\t\t       o->src_index->cache[pos + nr_entries - 1]->name);\n+\treturn 0;\n+}\n+\n static int traverse_trees_recursive(int n, unsigned long dirmask,\n \t\t\t\t    unsigned long df_conflicts,\n \t\t\t\t    struct name_entry *names,\n@@ -655,6 +751,17 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n \tvoid *buf[MAX_UNPACK_TREES];\n \tstruct traverse_info newinfo;\n \tstruct name_entry *p;\n+\tint nr_entries;\n+\n+\tnr_entries = all_trees_same_as_cache_tree(n, dirmask, names, info);\n+\tif (nr_entries > 0) {\n+\t\tstruct unpack_trees_options *o = info->data;\n+\t\tint pos = index_pos_by_traverse_info(names, info);\n+\n+\t\tif (!o->merge || df_conflicts)\n+\t\t\tBUG(\"Wrong condition to get here buddy\");\n+\t\treturn traverse_by_cache_tree(pos, nr_entries, n, names, info);\n+\t}\n \n \tp = names;\n \twhile (!p->mode)\n@@ -814,6 +921,11 @@ static struct cache_entry *create_ce_entry(const struct traverse_info *info, con\n \treturn ce;\n }\n \n+/*\n+ * Note that traverse_by_cache_tree() duplicates some logic in this function\n+ * without actually calling it. If you change the logic here you may need to\n+ * check and change there as well.\n+ */\n static int unpack_nondirectories(int n, unsigned long mask,\n \t\t\t\t unsigned long dirmask,\n \t\t\t\t struct cache_entry **src,\n@@ -998,6 +1110,11 @@ static void debug_unpack_callback(int n,\n \t\tdebug_name_entry(i, names + i);\n }\n \n+/*\n+ * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n+ * without actually calling it. If you change the logic here you may need to\n+ * check and change there as well.\n+ */\n static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n-- \n2.18.0.656.gda699b98b3\n\n"},{"id":"354477","messageId":"20180804053723.4695-4-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180804053723.4695-1-pclouds@gmail.com","subject":"[PATCH v3 3/4] unpack-trees: reduce malloc in cache-tree walk","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-04T05:37:22Z","receivedAt":"2018-08-04T05:37:39Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This is a micro optimization that probably only shines on repos with\ndeep directory structure. Instead of allocating and freeing a new\ncache_entry in every iteration, we reuse the last one and only update\nthe parts that are new each iteration.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n unpack-trees.c | 29 ++++++++++++++++++++---------\n 1 file changed, 20 insertions(+), 9 deletions(-)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex ba3d2e947e..c8defc2015 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -694,6 +694,8 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n \tstruct unpack_trees_options *o = info->data;\n+\tstruct cache_entry *tree_ce = NULL;\n+\tint ce_len = 0;\n \tint i, d;\n \n \tif (!o->merge)\n@@ -708,30 +710,39 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \t * in the first place.\n \t */\n \tfor (i = 0; i < nr_entries; i++) {\n-\t\tstruct cache_entry *tree_ce;\n-\t\tint len, rc;\n+\t\tint new_ce_len, len, rc;\n \n \t\tsrc[0] = o->src_index->cache[pos + i];\n \n \t\tlen = ce_namelen(src[0]);\n-\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n+\t\tnew_ce_len = cache_entry_size(len);\n+\n+\t\tif (new_ce_len > ce_len) {\n+\t\t\tnew_ce_len <<= 1;\n+\t\t\ttree_ce = xrealloc(tree_ce, new_ce_len);\n+\t\t\tmemset(tree_ce, 0, new_ce_len);\n+\t\t\tce_len = new_ce_len;\n+\n+\t\t\ttree_ce->ce_flags = create_ce_flags(0);\n+\n+\t\t\tfor (d = 1; d <= nr_names; d++)\n+\t\t\t\tsrc[d] = tree_ce;\n+\t\t}\n \n \t\ttree_ce->ce_mode = src[0]->ce_mode;\n-\t\ttree_ce->ce_flags = create_ce_flags(0);\n \t\ttree_ce->ce_namelen = len;\n \t\toidcpy(&tree_ce->oid, &src[0]->oid);\n \t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n \n-\t\tfor (d = 1; d <= nr_names; d++)\n-\t\t\tsrc[d] = tree_ce;\n-\n \t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n-\t\tfree(tree_ce);\n-\t\tif (rc < 0)\n+\t\tif (rc < 0) {\n+\t\t\tfree(tree_ce);\n \t\t\treturn rc;\n+\t\t}\n \n \t\tmark_ce_used(src[0], o);\n \t}\n+\tfree(tree_ce);\n \tif (o->debug_unpack)\n \t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n \t\t       nr_entries,\n-- \n2.18.0.656.gda699b98b3\n\n"},{"id":"354478","messageId":"20180804053723.4695-5-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180804053723.4695-1-pclouds@gmail.com","subject":"[PATCH v3 4/4] unpack-trees: cheaper index update when walking by cache-tree","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-04T05:37:23Z","receivedAt":"2018-08-04T05:37:41Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"With the new cache-tree, we could mostly avoid I/O (due to odb access)\nthe code mostly becomes a loop of \"check this, check that, add the\nentry to the index\". We could skip a couple checks in this giant loop\nto go faster:\n\n- We know here that we're copying entries from the source index to the\n  result one. All paths in the source index must have been validated\n  at load time already (and we're not taking strange paths from tree\n  objects) which means we can skip verify_path() without compromise.\n\n- We also know that D/F conflicts can't happen for all these entries\n  (since cache-tree and all the trees are the same) so we can skip\n  that as well.\n\nThis gives rather nice speedups for \"unpack trees\" rows where \"unpack\ntrees\" time is now cut in half compared to when\ntraverse_by_cache_tree() is added, or 1/7 of the original \"unpack\ntrees\" time.\n\n   baseline      cache-tree    this patch\n --------------------------------------------------------------------\n   0.018239226   0.019365414   0.020519621 s: read cache .git/index\n   0.052541655   0.049605548   0.048814384 s: preload index\n   0.001537598   0.001571695   0.001575382 s: refresh index\n   0.168167768   0.049677212   0.024719308 s: unpack trees\n   0.002897186   0.002845256   0.002805555 s: update worktree after a merge\n   0.131661745   0.136597522   0.134891617 s: repair cache-tree\n   0.075389117   0.075422517   0.074832291 s: write index, changed mask = 2a\n   0.111702023   0.032813253   0.008616479 s: unpack trees\n   0.000023245   0.000022002   0.000026630 s: update worktree after a merge\n   0.111793866   0.032933140   0.008714071 s: diff-index\n   0.587933288   0.398924370   0.380452871 s: git command: /home/pclouds/w/git/git\n\nTotal saving of this new patch looks even less impressive, now that\ntime spent in unpacking trees is so small. Which is why the next\nattempt should be on that \"repair cache-tree\" line.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache.h        |  1 +\n read-cache.c   |  3 ++-\n unpack-trees.c | 20 ++++++++++++++++++++\n unpack-trees.h |  1 +\n 4 files changed, 24 insertions(+), 1 deletion(-)\n\ndiff --git a/cache.h b/cache.h\nindex 8b447652a7..e6f7ee4b64 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -673,6 +673,7 @@ extern int index_name_pos(const struct index_state *, const char *name, int name\n #define ADD_CACHE_JUST_APPEND 8\t\t/* Append only; tree.c::read_tree() */\n #define ADD_CACHE_NEW_ONLY 16\t\t/* Do not replace existing ones */\n #define ADD_CACHE_KEEP_CACHE_TREE 32\t/* Do not invalidate cache-tree */\n+#define ADD_CACHE_SKIP_VERIFY_PATH 64\t/* Do not verify path */\n extern int add_index_entry(struct index_state *, struct cache_entry *ce, int option);\n extern void rename_index_entry_at(struct index_state *, int pos, const char *new_name);\n \ndiff --git a/read-cache.c b/read-cache.c\nindex e865254bea..b0b5df5de7 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1170,6 +1170,7 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n \tint ok_to_add = option & ADD_CACHE_OK_TO_ADD;\n \tint ok_to_replace = option & ADD_CACHE_OK_TO_REPLACE;\n \tint skip_df_check = option & ADD_CACHE_SKIP_DFCHECK;\n+\tint skip_verify_path = option & ADD_CACHE_SKIP_VERIFY_PATH;\n \tint new_only = option & ADD_CACHE_NEW_ONLY;\n \n \tif (!(option & ADD_CACHE_KEEP_CACHE_TREE))\n@@ -1210,7 +1211,7 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n \n \tif (!ok_to_add)\n \t\treturn -1;\n-\tif (!verify_path(ce->name, ce->ce_mode))\n+\tif (!skip_verify_path && !verify_path(ce->name, ce->ce_mode))\n \t\treturn error(\"Invalid path '%s'\", ce->name);\n \n \tif (!skip_df_check &&\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex c8defc2015..1438ee1555 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -201,6 +201,7 @@ static int do_add_entry(struct unpack_trees_options *o, struct cache_entry *ce,\n \n \tce->ce_flags = (ce->ce_flags & ~clear) | set;\n \treturn add_index_entry(&o->result, ce,\n+\t\t\t       o->extra_add_index_flags |\n \t\t\t       ADD_CACHE_OK_TO_ADD | ADD_CACHE_OK_TO_REPLACE);\n }\n \n@@ -701,6 +702,24 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \tif (!o->merge)\n \t\tBUG(\"We need cache-tree to do this optimization\");\n \n+\t/*\n+\t * Try to keep add_index_entry() as fast as possible since\n+\t * we're going to do a lot of them.\n+\t *\n+\t * Skipping verify_path() should totally be safe because these\n+\t * paths are from the source index, which must have been\n+\t * verified.\n+\t *\n+\t * Skipping D/F and cache-tree validation checks is trickier\n+\t * because it assumes what n-merge code would do when all\n+\t * trees and the index are the same. We probably could just\n+\t * optimize those code instead (e.g. we don't invalidate that\n+\t * many cache-tree, but the searching for them is very\n+\t * expensive).\n+\t */\n+\to->extra_add_index_flags = ADD_CACHE_SKIP_DFCHECK;\n+\to->extra_add_index_flags |= ADD_CACHE_SKIP_VERIFY_PATH;\n+\n \t/*\n \t * Do what unpack_callback() and unpack_nondirectories() normally\n \t * do. But we walk all paths recursively in just one loop instead.\n@@ -742,6 +761,7 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \n \t\tmark_ce_used(src[0], o);\n \t}\n+\to->extra_add_index_flags = 0;\n \tfree(tree_ce);\n \tif (o->debug_unpack)\n \t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\ndiff --git a/unpack-trees.h b/unpack-trees.h\nindex c2b434c606..94e1b14078 100644\n--- a/unpack-trees.h\n+++ b/unpack-trees.h\n@@ -80,6 +80,7 @@ struct unpack_trees_options {\n \tstruct index_state result;\n \n \tstruct exclude_list *el; /* for internal use */\n+\tunsigned int extra_add_index_flags;\n };\n \n extern int unpack_trees(unsigned n, struct tree_desc *t,\n-- \n2.18.0.656.gda699b98b3\n\n"},{"id":"354631","messageId":"xmqq7el3qywq.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"20180804053723.4695-1-pclouds@gmail.com","subject":"Re: [PATCH v3 0/4] Speed up unpack_trees()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-06T15:48:21Z","receivedAt":"2018-08-06T15:48:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> This is a minor update to address Ben's comments and add his\n> measurements in the commit message of 2/4 for the record.\n\nYay.\n\n> I've also checked about the lookahead thing in unpack_trees() to see\n> if we accidentally break something there, which is my biggest worry.\n> See [1] and [2] for context, but I believe since we can't have D/F\n> conflicts, the situation where lookahead is needed will not occur. So\n> we should be safe.\n\nIsn't this about branch switching, where the currently checked out\nbranch may have a regular file 't' and checking out another branch\nthat has directory 't' in it (or vice versa, possibly with the index\nhaving either a regular file 't' or requiring 't' to be a diretory\nby having a blob 't/1' in it)?  The log messge of [1] talks about\nwalking three trees together with the index, but even if we limit us\nto two-tree walk, I do not think that the picture fundamentally\nchanges.  So I am not sure how we can confidently say \"we can't have\nD/F\".  I'd need to block a solid time to take a look at the patches.\n\n> [1] da165f470e (unpack-trees.c: prepare for looking ahead in the index - 2010-01-07)\n> [2] 730f72840c (unpack-trees.c: look ahead in the index - 2009-09-20)\n\nThanks.\n"},{"id":"354633","messageId":"CACsJy8CzuxjjLyf637dtTHc1wK-UFVnNjwa0O300kYOWehz1vA@mail.gmail.com","threadId":"48913","inReplyTo":"xmqq7el3qywq.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 0/4] Speed up unpack_trees()","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-06T15:59:28Z","receivedAt":"2018-08-06T15:59:57Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 6, 2018 at 5:48 PM Junio C Hamano <gitster@pobox.com> wrote:\n> > I've also checked about the lookahead thing in unpack_trees() to see\n> > if we accidentally break something there, which is my biggest worry.\n> > See [1] and [2] for context, but I believe since we can't have D/F\n> > conflicts, the situation where lookahead is needed will not occur. So\n> > we should be safe.\n>\n> Isn't this about branch switching, where the currently checked out\n> branch may have a regular file 't' and checking out another branch\n> that has directory 't' in it (or vice versa, possibly with the index\n> having either a regular file 't' or requiring 't' to be a diretory\n> by having a blob 't/1' in it)?\n\nWe require the unpacked entry from all input trees to be a tree\nobjects (the dirmask thing), so if one tree has 't' as a file,\nall_trees_same_as_cache_tree() should return false and not trigger\nthis optimization. Same thing for the index, if it has the file 't',\nthen we should not have the cache-tree at path 't' and the\noptimization is skipped as well.\n\nSo yes branch switching definitely can have d/f conflicts, but we\nshould never ever accidentally run this new optimization when that\nhappens.\n\n> The log messge of [1] talks about\n> walking three trees together with the index, but even if we limit us\n> to two-tree walk, I do not think that the picture fundamentally\n> changes.  So I am not sure how we can confidently say \"we can't have\n> D/F\".  I'd need to block a solid time to take a look at the patches.\n\nYes please :)\n-- \nDuy\n"},{"id":"354670","messageId":"xmqqr2jbnwy2.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"CACsJy8CzuxjjLyf637dtTHc1wK-UFVnNjwa0O300kYOWehz1vA@mail.gmail.com","subject":"Re: [PATCH v3 0/4] Speed up unpack_trees()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-06T18:59:01Z","receivedAt":"2018-08-06T18:59:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> We require the unpacked entry from all input trees to be a tree\n> objects (the dirmask thing), so if one tree has 't' as a file,\n\nAh, OK, this is still part of that \"all the trees match cache tree\nso we walk the index instead\" optimization.  I forgot about that.\n\n"},{"id":"354871","messageId":"a8be62e2-528b-72cc-56c7-55291af8dd66@gmail.com","threadId":"48913","inReplyTo":"xmqqr2jbnwy2.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 0/4] Speed up unpack_trees()","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-08T17:00:26Z","receivedAt":"2018-08-08T17:00:31Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/6/2018 2:59 PM, Junio C Hamano wrote:\n> Duy Nguyen <pclouds@gmail.com> writes:\n> \n>> We require the unpacked entry from all input trees to be a tree\n>> objects (the dirmask thing), so if one tree has 't' as a file,\n> \n> Ah, OK, this is still part of that \"all the trees match cache tree\n> so we walk the index instead\" optimization.  I forgot about that.\n> \n\nI ran this set of patches through the VFS For Git set of functional \ntests as well as our performance test suite (my earlier perf numbers \nwere from manual testing).  All the functional tests pass and the \nperformance tests are looking _very_ promising.\n\nCheckout times are impacted most and on average drop from 20.96\tseconds \nto 11.63 seconds for a 45% savings.\n\nMerge times drop from 19.44 seconds to 12.88 for a 34% savings.\n\nRebase times drop from 26.78 seconds to 20.72 for a 23% savings.\n\nOverall, I'm looking forward to a good review of the patches and seeing \nthem get merged as soon as they are ready.\n\n"},{"id":"354873","messageId":"xmqqpnyshhtt.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"CACsJy8CzuxjjLyf637dtTHc1wK-UFVnNjwa0O300kYOWehz1vA@mail.gmail.com","subject":"Re: [PATCH v3 0/4] Speed up unpack_trees()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-08T17:46:38Z","receivedAt":"2018-08-08T17:46:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> On Mon, Aug 6, 2018 at 5:48 PM Junio C Hamano <gitster@pobox.com> wrote:\n>> > I've also checked about the lookahead thing in unpack_trees() to see\n>> > if we accidentally break something there, which is my biggest worry.\n>> > See [1] and [2] for context, but I believe since we can't have D/F\n>> > conflicts, the situation where lookahead is needed will not occur. So\n>> > we should be safe.\n\nI think you would want the same \"switch cache-bottom before\ndescending into a subdirectory, and then restore cache-bottom after\ntraversal comes back\" dance that is done for the normal tree\ntraversal case to happen.\n\n\tbottom = switch_cache_bottom(&newinfo);\n\tret = traverse_trees(n, t, &newinfo);\n\trestore_cache_bottom(&newinfo, bottom);\n\nDuring your walk of the index and the trees that are known to be in\nsync, there is little reason to worry about the cache_bottom, which\nis advanced by calling mark_ce_used() in traverse_by_cache_tree().\nWhere it matters is what happens after the traversal comes back out\nof the subtree.  find_cache_pos() uses the bottom pointer so that it\ndoes not have to go back to far to find an index entry that has not\nbeen used to match with the entries from the trees (which are not\nsorted exactly the same way as the index, unfortunately), so\nforgetting to advance the bottom pointer while correctly marking a\nce as \"used\" is OK (i.e. hurts performance but not correctness), but\nadvancing the bottom pointer too much and leaving entries that are\nnot used behind is *not* OK.  And lack of restoring the bottom in\nthe new codepath makes me suspect exactly such a bug _after_ the\ntraversal exits the subtree we are using this new optimization in\nand moves on.\n\nImagine we are iterating over the top-level of the trees, and found\na subtree in them.  There may be some index entries before the first\npath in this subtree that are not yet marked as \"used\", the earliest\nof which is pointed at by the cache_bottom pointer.\n\nBefore descending into the subtree (and start consuming the entry\nwith the first in this subtree from the index), we stash away the\ncurrent cache_bottom, and then start walking the subtree.  While we\nare in that subtree, the cache_bottom starts from the first name in\nthe subtree and increments, as we _know_ the entry at the old\ncache_bottom is outside this subtree and will not match any entry\nform the subtree.  Then when the traversal returns, all index\nentries within the \"subtree/\" path will be marked \"used\".  At that\npoint, when we continue to scan the top-level of the trees, we need\nto restore the cache_bottom, so that we do not forget entries that\nwe knew we needed to scan eventually, if there was any.\n"},{"id":"354875","messageId":"xmqqin4khgm5.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"xmqqpnyshhtt.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 0/4] Speed up unpack_trees()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-08T18:12:50Z","receivedAt":"2018-08-08T18:12:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> not used behind is *not* OK.  And lack of restoring the bottom in\n> the new codepath makes me suspect exactly such a bug _after_ the\n> traversal exits the subtree we are using this new optimization in\n> and moves on.\n\nHmph, thinking about this further, I cannot convince myself that\nlack of bottom adjustment can lead to a triggerable bug.  The only\ncase that a subtree traversal need to skip some unpacked entries in\nthe index and then revisit them by rewinding, e.g. entries \"t-i\" and\n\"t-j\" that are left unprocessed while entries \"t/1\", \"t/2\", etc. are\nprocessed, in the illustration of da165f47 (\"unpack-trees.c: prepare\nfor looking ahead in the index\", 2010-01-07), is when one of the\ntrees have a non-tree with the same name as the subtree we are\ntrying to descend into, and as long as we know all trees have the\nthing as a tree, I do not think of a case where such ordering\ninversion would get in the way.\n\nThat was the only thing I found questionable in 2/4, which is the\nmost important piece in the series, so we probably are OK.\n\nThanks for working on this one.\n"},{"id":"354876","messageId":"CABPp-BGcPV0RA624_1UOXYkvaNhW4yR2ifhV_MVFZQOgBb_Ydg@mail.gmail.com","threadId":"48913","inReplyTo":"20180804053723.4695-3-pclouds@gmail.com","subject":"Re: [PATCH v3 2/4] unpack-trees: optimize walking same trees with cache-tree","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-08-08T18:23:05Z","receivedAt":"2018-08-08T18:23:22Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Fri, Aug 3, 2018 at 10:39 PM Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n> From: Duy Nguyen <pclouds@gmail.com>\n>\n> In order to merge one or many trees with the index, unpack-trees code\n> walks multiple trees in parallel with the index and performs n-way\n> merge. If we find out at start of a directory that all trees are the\n> same (by comparing OID) and cache-tree happens to be available for\n> that directory as well, we could avoid walking the trees because we\n> already know what these trees contain: it's flattened in what's called\n> \"the index\".\n\nThis is cool.\n\n> The upside is of course a lot less I/O since we can potentially skip\n> lots of trees (think subtrees). We also save CPU because we don't have\n> to inflate and the apply deltas. The downside is of course more\n\ns/and the apply/and apply the/\n\n> fragile code since the logic in some functions are now duplicated\n> elsewhere.\n>\n> \"checkout -\" with this patch on gcc.git:\n>\n>     baseline      new\n>   --------------------------------------------------------------------\n>     0.018239226   0.019365414 s: read cache .git/index\n>     0.052541655   0.049605548 s: preload index\n>     0.001537598   0.001571695 s: refresh index\n>     0.168167768   0.049677212 s: unpack trees\n>     0.002897186   0.002845256 s: update worktree after a merge\n>     0.131661745   0.136597522 s: repair cache-tree\n>     0.075389117   0.075422517 s: write index, changed mask = 2a\n>     0.111702023   0.032813253 s: unpack trees\n>     0.000023245   0.000022002 s: update worktree after a merge\n>     0.111793866   0.032933140 s: diff-index\n>     0.587933288   0.398924370 s: git command: /home/pclouds/w/git/git\n>\n> Another measurement from Ben's running \"git checkout\" with over 500k\n> trees (on the whole series):\n>\n>     baseline        new\n>   ----------------------------------------------------------------------\n>     0.535510167     0.556558733     s: read cache .git/index\n>     0.3057373       0.3147105       s: initialize name hash\n>     0.0184082       0.023558433     s: preload index\n>     0.086910967     0.089085967     s: refresh index\n>     7.889590767     2.191554433     s: unpack trees\n>     0.120760833     0.131941267     s: update worktree after a merge\n>     2.2583504       2.572663167     s: repair cache-tree\n>     0.8916137       0.959495233     s: write index, changed mask = 28\n>     3.405199233     0.2710663       s: unpack trees\n>     0.000999667     0.0021554       s: update worktree after a merge\n>     3.4063306       0.273318333     s: diff-index\n>     16.9524923      9.462943133     s: git command: git.exe checkout\n>\n> This command calls unpack_trees() twice, the first time on 2way merge\n> and the second 1way merge. In both times, \"unpack trees\" time is\n> reduced to one third. Overall time reduction is not that impressive of\n> course because index operations take a big chunk. And there's that\n> repair cache-tree line.\n>\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  unpack-trees.c | 117 +++++++++++++++++++++++++++++++++++++++++++++++++\n>  1 file changed, 117 insertions(+)\n>\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index a32ddee159..ba3d2e947e 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -644,6 +644,102 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n>         return name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n>  }\n>\n> +static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n> +                                       struct name_entry *names,\n> +                                       struct traverse_info *info)\n> +{\n> +       struct unpack_trees_options *o = info->data;\n> +       int i;\n> +\n> +       if (!o->merge || dirmask != ((1 << n) - 1))\n> +               return 0;\n> +\n> +       for (i = 1; i < n; i++)\n> +               if (!are_same_oid(names, names + i))\n> +                       return 0;\n> +\n> +       return cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n> +}\n\nI was curious whether this could also be extended in the case of a\nmerge; as long as HEAD and MERGE have the same tree, even if the base\ncommit doesn't match, we can still just use the tree from HEAD which\nshould be in the current index/cache_tree.  However, it'd be a\nsomewhat odd history for HEAD and MERGE to match on some significantly\nsized tree when the base commit doesn't also match.\n\n> +\n> +static int index_pos_by_traverse_info(struct name_entry *names,\n> +                                     struct traverse_info *info)\n> +{\n> +       struct unpack_trees_options *o = info->data;\n> +       int len = traverse_path_len(info, names);\n> +       char *name = xmalloc(len + 1 /* slash */ + 1 /* NUL */);\n> +       int pos;\n> +\n> +       make_traverse_path(name, info, names);\n> +       name[len++] = '/';\n> +       name[len] = '\\0';\n> +       pos = index_name_pos(o->src_index, name, len);\n> +       if (pos >= 0)\n> +               BUG(\"This is a directory and should not exist in index\");\n> +       pos = -pos - 1;\n> +       if (!starts_with(o->src_index->cache[pos]->name, name) ||\n> +           (pos > 0 && starts_with(o->src_index->cache[pos-1]->name, name)))\n> +               BUG(\"pos must point at the first entry in this directory\");\n> +       free(name);\n> +       return pos;\n> +}\n> +\n> +/*\n> + * Fast path if we detect that all trees are the same as cache-tree at this\n> + * path. We'll walk these trees recursively using cache-tree/index instead of\n> + * ODB since already know what these trees contain.\n> + */\n> +static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n> +                                 struct name_entry *names,\n> +                                 struct traverse_info *info)\n> +{\n> +       struct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n> +       struct unpack_trees_options *o = info->data;\n> +       int i, d;\n> +\n> +       if (!o->merge)\n> +               BUG(\"We need cache-tree to do this optimization\");\n> +\n> +       /*\n> +        * Do what unpack_callback() and unpack_nondirectories() normally\n> +        * do. But we walk all paths recursively in just one loop instead.\n> +        *\n> +        * D/F conflicts and staged entries are not a concern because\n\n\"staged entries\"?  Do you mean \"higher stage entries\"?  I'm not sure\nthe correct terminology here, but the former makes me think of changes\nthe user has staged but not committed (i.e. stuff found at stage #0 in\nthe index, but which isn't found in any tree yet) vs. the latter which\nI'd use to refer to entries at stages 1 or higher.\n\n> +        * cache-tree would be invalidated and we would never get here\n> +        * in the first place.\n> +        */\n> +       for (i = 0; i < nr_entries; i++) {\n> +               struct cache_entry *tree_ce;\n> +               int len, rc;\n> +\n> +               src[0] = o->src_index->cache[pos + i];\n> +\n> +               len = ce_namelen(src[0]);\n> +               tree_ce = xcalloc(1, cache_entry_size(len));\n> +\n> +               tree_ce->ce_mode = src[0]->ce_mode;\n> +               tree_ce->ce_flags = create_ce_flags(0);\n> +               tree_ce->ce_namelen = len;\n> +               oidcpy(&tree_ce->oid, &src[0]->oid);\n> +               memcpy(tree_ce->name, src[0]->name, len + 1);\n\nWe do a bunch of work to setup tree_ce...\n\n> +               for (d = 1; d <= nr_names; d++)\n> +                       src[d] = tree_ce;\n\n...then we make nr_names copies of tree_ce (so that *way_merge or\nbind_merge or oneway_diff or whatever will have the expected number of\nentries).\n\n> +               rc = call_unpack_fn((const struct cache_entry * const *)src, o);\n\n...then we call o->fn (via call_unpack_fn) to do various complicated\nlogic to figure out which tree_ce to use??  Isn't that just an\nexpensive way to recompute that what we currently have in the index is\nwhat we want to keep there?\n\nGranted, a caller of this may have set o->fn to something other than\n{one,two,three}way_merge (or bind_merge), and that function might have\nimportant side effects...but it just seems annoying to have to do so\nmuch work when for most uses we already know the entry in the index is\nthe one we already want.  In fact, the only other thing in the\ncodebase that o->fn is now set to is oneway_diff, which I think is a\nno-op when the two trees match.\n\nWould be nice if we could avoid all this, at least in the common cases\nwhere o->fn is a function known to not have side effects.  Or did I\nnot read those functions closely enough and they do have important\nside effects?\n\n> +               free(tree_ce);\n> +               if (rc < 0)\n> +                       return rc;\n> +\n> +               mark_ce_used(src[0], o);\n> +       }\n> +       if (o->debug_unpack)\n> +               printf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n> +                      nr_entries,\n> +                      o->src_index->cache[pos]->name,\n> +                      o->src_index->cache[pos + nr_entries - 1]->name);\n> +       return 0;\n> +}\n> +\n>  static int traverse_trees_recursive(int n, unsigned long dirmask,\n>                                     unsigned long df_conflicts,\n>                                     struct name_entry *names,\n> @@ -655,6 +751,17 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n>         void *buf[MAX_UNPACK_TREES];\n>         struct traverse_info newinfo;\n>         struct name_entry *p;\n> +       int nr_entries;\n> +\n> +       nr_entries = all_trees_same_as_cache_tree(n, dirmask, names, info);\n> +       if (nr_entries > 0) {\n> +               struct unpack_trees_options *o = info->data;\n> +               int pos = index_pos_by_traverse_info(names, info);\n> +\n> +               if (!o->merge || df_conflicts)\n> +                       BUG(\"Wrong condition to get here buddy\");\n\nheh.  :)\n\n> +               return traverse_by_cache_tree(pos, nr_entries, n, names, info);\n> +       }\n>\n>         p = names;\n>         while (!p->mode)\n> @@ -814,6 +921,11 @@ static struct cache_entry *create_ce_entry(const struct traverse_info *info, con\n>         return ce;\n>  }\n>\n> +/*\n> + * Note that traverse_by_cache_tree() duplicates some logic in this function\n> + * without actually calling it. If you change the logic here you may need to\n> + * check and change there as well.\n> + */\n>  static int unpack_nondirectories(int n, unsigned long mask,\n>                                  unsigned long dirmask,\n>                                  struct cache_entry **src,\n> @@ -998,6 +1110,11 @@ static void debug_unpack_callback(int n,\n>                 debug_name_entry(i, names + i);\n>  }\n>\n> +/*\n> + * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n\ns/funciton/function/\n\n> + * without actually calling it. If you change the logic here you may need to\n> + * check and change there as well.\n> + */\n>  static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n>  {\n>         struct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n> --\n> 2.18.0.656.gda699b98b3\n"},{"id":"354878","messageId":"CABPp-BF7kfeTc7rdg8QFMq0MkP+GOQK23KBPW=A6FrD0SymjZg@mail.gmail.com","threadId":"48913","inReplyTo":"20180804053723.4695-4-pclouds@gmail.com","subject":"Re: [PATCH v3 3/4] unpack-trees: reduce malloc in cache-tree walk","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-08-08T18:30:20Z","receivedAt":"2018-08-08T18:30:34Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Fri, Aug 3, 2018 at 10:39 PM Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n>\n> This is a micro optimization that probably only shines on repos with\n> deep directory structure. Instead of allocating and freeing a new\n> cache_entry in every iteration, we reuse the last one and only update\n> the parts that are new each iteration.\n>\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>  unpack-trees.c | 29 ++++++++++++++++++++---------\n>  1 file changed, 20 insertions(+), 9 deletions(-)\n>\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index ba3d2e947e..c8defc2015 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -694,6 +694,8 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n>  {\n>         struct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n>         struct unpack_trees_options *o = info->data;\n> +       struct cache_entry *tree_ce = NULL;\n> +       int ce_len = 0;\n>         int i, d;\n>\n>         if (!o->merge)\n> @@ -708,30 +710,39 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n>          * in the first place.\n>          */\n>         for (i = 0; i < nr_entries; i++) {\n> -               struct cache_entry *tree_ce;\n> -               int len, rc;\n> +               int new_ce_len, len, rc;\n>\n>                 src[0] = o->src_index->cache[pos + i];\n>\n>                 len = ce_namelen(src[0]);\n> -               tree_ce = xcalloc(1, cache_entry_size(len));\n> +               new_ce_len = cache_entry_size(len);\n> +\n> +               if (new_ce_len > ce_len) {\n> +                       new_ce_len <<= 1;\n> +                       tree_ce = xrealloc(tree_ce, new_ce_len);\n> +                       memset(tree_ce, 0, new_ce_len);\n> +                       ce_len = new_ce_len;\n> +\n> +                       tree_ce->ce_flags = create_ce_flags(0);\n> +\n> +                       for (d = 1; d <= nr_names; d++)\n> +                               src[d] = tree_ce;\n> +               }\n>\n>                 tree_ce->ce_mode = src[0]->ce_mode;\n> -               tree_ce->ce_flags = create_ce_flags(0);\n>                 tree_ce->ce_namelen = len;\n>                 oidcpy(&tree_ce->oid, &src[0]->oid);\n>                 memcpy(tree_ce->name, src[0]->name, len + 1);\n>\n> -               for (d = 1; d <= nr_names; d++)\n> -                       src[d] = tree_ce;\n> -\n>                 rc = call_unpack_fn((const struct cache_entry * const *)src, o);\n> -               free(tree_ce);\n> -               if (rc < 0)\n> +               if (rc < 0) {\n> +                       free(tree_ce);\n>                         return rc;\n> +               }\n>\n>                 mark_ce_used(src[0], o);\n>         }\n> +       free(tree_ce);\n>         if (o->debug_unpack)\n>                 printf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n>                        nr_entries,\n> --\n> 2.18.0.656.gda699b98b3\n\nSeems reasonable, when we really do have to invoke call_unpack_fn.\nI'm still curious if there are reasons why we couldn't just skip that\ncall (at least when o->fn is one of {oneway_merge, twoway_merge,\nthreeway_merge, bind_merge}), but I already brought that up in my\ncomments on patch 2.\n"},{"id":"354880","messageId":"xmqqeff8hfdg.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"xmqqin4khgm5.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 0/4] Speed up unpack_trees()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-08T18:39:39Z","receivedAt":"2018-08-08T18:39:45Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> not used behind is *not* OK.  And lack of restoring the bottom in\n>> the new codepath makes me suspect exactly such a bug _after_ the\n>> traversal exits the subtree we are using this new optimization in\n>> and moves on.\n>\n> Hmph, thinking about this further, I cannot convince myself that\n> lack of bottom adjustment can lead to a triggerable bug.  The only\n> case that a subtree traversal need to skip some unpacked entries in\n> the index and then revisit them by rewinding, e.g. entries \"t-i\" and\n> \"t-j\" that are left unprocessed while entries \"t/1\", \"t/2\", etc. are\n> processed, in the illustration of da165f47 (\"unpack-trees.c: prepare\n> for looking ahead in the index\", 2010-01-07), is when one of the\n> trees have a non-tree with the same name as the subtree we are\n> trying to descend into, and as long as we know all trees have the\n> thing as a tree, I do not think of a case where such ordering\n> inversion would get in the way.\n\nOne more, and hopefully the final, note.\n\nA paranoid may be soothed by a simple \"cache_bottom must match pos\nat this point\" at the beginning of the optimized traversal.  Just\nlike you already have an assert to ensure that pos points at the\nfirst entry in the directory in index_pos_by_traverse_info(), there\nshould not be any unused entry in the index before that entry and\nthe bottom pointer must be pointing at it.  It is a cheap check, and\nif violated, would indicate that the above \"I do not think of a\ncase ...\" was incomplete.\n\n> That was the only thing I found questionable in 2/4, which is the\n> most important piece in the series, so we probably are OK.\n>\n> Thanks for working on this one.\n"},{"id":"354883","messageId":"CABPp-BGF+GZjm-DiveLjFOESKwPz2F0Y7X4_kXyem2xFo2odUw@mail.gmail.com","threadId":"48913","inReplyTo":"20180729103306.16403-5-pclouds@gmail.com","subject":"Re: [PATCH v2 4/4] unpack-trees: cheaper index update when walking by cache-tree","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-08-08T18:46:24Z","receivedAt":"2018-08-08T18:46:39Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sun, Jul 29, 2018 at 3:36 AM Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n>\n> With the new cache-tree, we could mostly avoid I/O (due to odb access)\n> the code mostly becomes a loop of \"check this, check that, add the\n> entry to the index\". We could skip a couple checks in this giant loop\n> to go faster:\n>\n> - We know here that we're copying entries from the source index to the\n>   result one. All paths in the source index must have been validated\n>   at load time already (and we're not taking strange paths from tree\n>   objects) which means we can skip verify_path() without compromise.\n>\n> - We also know that D/F conflicts can't happen for all these entries\n>   (since cache-tree and all the trees are the same) so we can skip\n>   that as well.\n>\n> This gives rather nice speedups for \"unpack trees\" rows where \"unpack\n> trees\" time is now cut in half compared to when\n> traverse_by_cache_tree() is added, or 1/7 of the original \"unpack\n> trees\" time.\n>\n>    baseline      cache-tree    this patch\n>  --------------------------------------------------------------------\n>    0.018239226   0.019365414   0.020519621 s: read cache .git/index\n>    0.052541655   0.049605548   0.048814384 s: preload index\n>    0.001537598   0.001571695   0.001575382 s: refresh index\n>    0.168167768   0.049677212   0.024719308 s: unpack trees\n>    0.002897186   0.002845256   0.002805555 s: update worktree after a merge\n>    0.131661745   0.136597522   0.134891617 s: repair cache-tree\n>    0.075389117   0.075422517   0.074832291 s: write index, changed mask = 2a\n>    0.111702023   0.032813253   0.008616479 s: unpack trees\n>    0.000023245   0.000022002   0.000026630 s: update worktree after a merge\n>    0.111793866   0.032933140   0.008714071 s: diff-index\n>    0.587933288   0.398924370   0.380452871 s: git command: /home/pclouds/w/git/git\n>\n> Total saving of this new patch looks even less impressive, now that\n> time spent in unpacking trees is so small. Which is why the next\n> attempt should be on that \"repair cache-tree\" line.\n>\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  cache.h        |  1 +\n>  read-cache.c   |  3 ++-\n>  unpack-trees.c | 27 +++++++++++++++++++++++++++\n>  unpack-trees.h |  1 +\n>  4 files changed, 31 insertions(+), 1 deletion(-)\n>\n> diff --git a/cache.h b/cache.h\n> index 8b447652a7..e6f7ee4b64 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -673,6 +673,7 @@ extern int index_name_pos(const struct index_state *, const char *name, int name\n>  #define ADD_CACHE_JUST_APPEND 8                /* Append only; tree.c::read_tree() */\n>  #define ADD_CACHE_NEW_ONLY 16          /* Do not replace existing ones */\n>  #define ADD_CACHE_KEEP_CACHE_TREE 32   /* Do not invalidate cache-tree */\n> +#define ADD_CACHE_SKIP_VERIFY_PATH 64  /* Do not verify path */\n>  extern int add_index_entry(struct index_state *, struct cache_entry *ce, int option);\n>  extern void rename_index_entry_at(struct index_state *, int pos, const char *new_name);\n>\n> diff --git a/read-cache.c b/read-cache.c\n> index e865254bea..b0b5df5de7 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1170,6 +1170,7 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n>         int ok_to_add = option & ADD_CACHE_OK_TO_ADD;\n>         int ok_to_replace = option & ADD_CACHE_OK_TO_REPLACE;\n>         int skip_df_check = option & ADD_CACHE_SKIP_DFCHECK;\n> +       int skip_verify_path = option & ADD_CACHE_SKIP_VERIFY_PATH;\n>         int new_only = option & ADD_CACHE_NEW_ONLY;\n>\n>         if (!(option & ADD_CACHE_KEEP_CACHE_TREE))\n> @@ -1210,7 +1211,7 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n>\n>         if (!ok_to_add)\n>                 return -1;\n> -       if (!verify_path(ce->name, ce->ce_mode))\n> +       if (!skip_verify_path && !verify_path(ce->name, ce->ce_mode))\n>                 return error(\"Invalid path '%s'\", ce->name);\n>\n>         if (!skip_df_check &&\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index c33ebaf001..dc62afd968 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -201,6 +201,7 @@ static int do_add_entry(struct unpack_trees_options *o, struct cache_entry *ce,\n>\n>         ce->ce_flags = (ce->ce_flags & ~clear) | set;\n>         return add_index_entry(&o->result, ce,\n> +                              o->extra_add_index_flags |\n>                                ADD_CACHE_OK_TO_ADD | ADD_CACHE_OK_TO_REPLACE);\n>  }\n>\n> @@ -701,6 +702,24 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n>         if (!o->merge)\n>                 BUG(\"We need cache-tree to do this optimization\");\n>\n> +       /*\n> +        * Try to keep add_index_entry() as fast as possible since\n> +        * we're going to do a lot of them.\n> +        *\n> +        * Skipping verify_path() should totally be safe because these\n> +        * paths are from the source index, which must have been\n> +        * verified.\n> +        *\n> +        * Skipping D/F and cache-tree validation checks is trickier\n> +        * because it assumes what n-merge code would do when all\n> +        * trees and the index are the same. We probably could just\n> +        * optimize those code instead (e.g. we don't invalidate that\n> +        * many cache-tree, but the searching for them is very\n> +        * expensive).\n> +        */\n> +       o->extra_add_index_flags = ADD_CACHE_SKIP_DFCHECK;\n> +       o->extra_add_index_flags |= ADD_CACHE_SKIP_VERIFY_PATH;\n> +\n\nIn sum of this whole patch, you notice that the Nway_merge functions\nare still a bit of a bottleneck, but you know you have a special case\nwhere you want them to put an entry in the index that matches what is\nalready there, so you try to set some extra flags to short-circuit\npart of their logic and get to what you know is the correct result.\n\nThis seems a little scary to me.  I think it's probably safe as long\nas o->fn is one of {oneway_merge, twoway_merge, threeway_merge,\nbind_merge} (the cases you have in mind and which the current code\nuses), but the caller isn't limited to those.  Right now in\ndiff-lib.c, there's a caller that has their own function, oneway_diff.\nMore could be added in the future.\n\nIf we're going to go this route, I think we should first check that\no->fn is one of those known safe functions.  And if we're going that\nroute, the comments I bring up on patch 2 about possibly avoiding\ncall_unpack_fn() altogether might even obviate this patch while\nspeeding things up more.\n\n>         /*\n>          * Do what unpack_callback() and unpack_nondirectories() normally\n>          * do. But we walk all paths recursively in just one loop instead.\n> @@ -742,6 +761,7 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n>\n>                 mark_ce_used(src[0], o);\n>         }\n> +       o->extra_add_index_flags = 0;\n>         free(tree_ce);\n>         if (o->debug_unpack)\n>                 printf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n> @@ -1561,6 +1581,13 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>                 if (!ret) {\n>                         if (!o->result.cache_tree)\n>                                 o->result.cache_tree = cache_tree();\n> +                       /*\n> +                        * TODO: Walk o.src_index->cache_tree, quickly check\n> +                        * if o->result.cache has the exact same content for\n> +                        * any valid cache-tree in o.src_index, then we can\n> +                        * just copy the cache-tree over instead of hashing a\n> +                        * new tree object.\n> +                        */\n\nInteresting.  I really don't know how cache_tree works...but if we\navoided calling call_unpack_fn, and thus left the original index entry\nin place instead of replacing it with an equal one, would that as a\nside effect speed up the cache_tree_valid/cache_tree_update calls for\nus?  Or is there still work here?\n\n>                         if (!cache_tree_fully_valid(o->result.cache_tree))\n>                                 cache_tree_update(&o->result,\n>                                                   WRITE_TREE_SILENT |\n> diff --git a/unpack-trees.h b/unpack-trees.h\n> index c2b434c606..94e1b14078 100644\n> --- a/unpack-trees.h\n> +++ b/unpack-trees.h\n> @@ -80,6 +80,7 @@ struct unpack_trees_options {\n>         struct index_state result;\n>\n>         struct exclude_list *el; /* for internal use */\n> +       unsigned int extra_add_index_flags;\n>  };\n>\n>  extern int unpack_trees(unsigned n, struct tree_desc *t,\n> --\n> 2.18.0.656.gda699b98b3\n"},{"id":"354928","messageId":"ccec34c9-b81a-bcb4-7d05-48dccc059cc8@gmail.com","threadId":"48913","inReplyTo":"20180801163830.GA31968@duynguyen.home","subject":"Re: [PATCH v2 0/4] Speed up unpack_trees()","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-08T20:53:09Z","receivedAt":"2018-08-08T20:53:14Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/1/2018 12:38 PM, Duy Nguyen wrote:\n> On Tue, Jul 31, 2018 at 01:31:31PM -0400, Ben Peart wrote:\n>>\n>>\n>> On 7/31/2018 12:50 PM, Ben Peart wrote:\n>>>\n>>>\n>>> On 7/31/2018 11:31 AM, Duy Nguyen wrote:\n>>\n>>>>\n>>>>> In the performance game of whack-a-mole, that call to repair cache-tree\n>>>>> is now looking quite expensive...\n>>>>\n>>>> Yeah and I think we can whack that mole too. I did some measurement.\n>>>> Best case possible, we just need to scan through two indexes (one with\n>>>> many good cache-tree, one with no cache-tree), compare and copy\n>>>> cache-tree over. The scanning takes like 1% time of current repair\n>>>> step and I suspect it's the hashing that takes most of the time. Of\n>>>> course real world won't have such nice numbers, but I guess we could\n>>>> maybe half cache-tree update/repair time.\n>>>>\n>>>\n>>> I have some great profiling tools available so will take a look at this\n>>> next and see exactly where the time is being spent.\n>>\n>> Good instincts.  In cache_tree_update, the heavy hitter is definitely\n>> hash_object_file followed by has_object_file.\n>>\n>> Name                               \tInc %\t     Inc\n>> + git!cache_tree_update            \t 12.4\t   4,935\n>> |+ git!update_one                  \t 11.8\t   4,706\n>> | + git!update_one                 \t 11.8\t   4,706\n>> |  + git!hash_object_file          \t  6.1\t   2,406\n>> |  + git!has_object_file           \t  2.0\t     813\n>> |  + OTHER <<vcruntime140d!strchr>>\t  0.5\t     203\n>> |  + git!strbuf_addf               \t  0.4\t     155\n>> |  + git!strbuf_release            \t  0.4\t     143\n>> |  + git!strbuf_add                \t  0.3\t     121\n>> |  + OTHER <<vcruntime140d!memcmp>>\t  0.2\t      93\n>> |  + git!strbuf_grow               \t  0.1\t      25\n> \n> Ben, if you work on this, this could be a good starting point. I will\n> not work on this because I still have some other things to catch up\n> and follow through. You can have my sign off if you reuse something\n> from this patch\n> \n> Even if it's a naive implementation, the initial numbers look pretty\n> good. Without the patch we have\n> \n> 18:31:05.970621 unpack-trees.c:1437     performance: 0.000001029 s: copy\n> 18:31:05.975729 unpack-trees.c:1444     performance: 0.005082004 s: update\n> \n> And with the patch\n> \n> 18:31:13.295655 unpack-trees.c:1437     performance: 0.000198017 s: copy\n> 18:31:13.296757 unpack-trees.c:1444     performance: 0.001075935 s: update\n> \n> Time saving is about 80% by the look of this (best possible case\n> because only the top tree needs to be hashed and written out).\n> \n> -- 8< --\n> diff --git a/cache-tree.c b/cache-tree.c\n> index 6b46711996..67a4a93100 100644\n> --- a/cache-tree.c\n> +++ b/cache-tree.c\n> @@ -440,6 +440,147 @@ int cache_tree_update(struct index_state *istate, int flags)\n>   \treturn 0;\n>   }\n>   \n> +static int same(const struct cache_entry *a, const struct cache_entry *b)\n> +{\n> +\tif (ce_stage(a) || ce_stage(b))\n> +\t\treturn 0;\n> +\tif ((a->ce_flags | b->ce_flags) & CE_CONFLICTED)\n> +\t\treturn 0;\n> +\treturn a->ce_mode == b->ce_mode &&\n> +\t       !oidcmp(&a->oid, &b->oid);\n> +}\n> +\n> +static int cache_tree_name_pos(const struct index_state *istate,\n> +\t\t\t       const struct strbuf *path)\n> +{\n> +\tint pos;\n> +\n> +\tif (!path->len)\n> +\t\treturn 0;\n> +\n> +\tpos = index_name_pos(istate, path->buf, path->len);\n> +\tif (pos >= 0)\n> +\t\tBUG(\"No no no, directory path must not exist in index\");\n> +\treturn -pos - 1;\n> +}\n> +\n> +/*\n> + * Locate the same cache-tree in two separate indexes. Check the\n> + * cache-tree is still valid for the \"to\" index (i.e. it contains the\n> + * same set of entries in the \"from\" index).\n> + */\n> +static int verify_one_cache_tree(const struct index_state *to,\n> +\t\t\t\t const struct index_state *from,\n> +\t\t\t\t const struct cache_tree *it,\n> +\t\t\t\t const struct strbuf *path)\n> +{\n> +\tint i, spos, dpos;\n> +\n> +\tspos = cache_tree_name_pos(from, path);\n> +\tif (spos + it->entry_count > from->cache_nr)\n> +\t\treturn -1;\n> +\n> +\tdpos = cache_tree_name_pos(to, path);\n> +\tif (dpos + it->entry_count > to->cache_nr)\n> +\t\treturn -1;\n> +\n> +\t/* Can we quickly check head and tail and bail out early */\n> +\tif (!same(from->cache[spos], to->cache[spos]) ||\n> +\t    !same(from->cache[spos + it->entry_count - 1],\n> +\t\t  to->cache[spos + it->entry_count - 1]))\n> +\t\treturn -1;\n> +\n> +\tfor (i = 1; i < it->entry_count - 1; i++)\n> +\t\tif (!same(from->cache[spos + i],\n> +\t\t\t  to->cache[dpos + i]))\n> +\t\t\treturn -1;\n> +\n> +\treturn 0;\n> +}\n> +\n> +static int verify_and_invalidate(struct index_state *to,\n> +\t\t\t\t const struct index_state *from,\n> +\t\t\t\t struct cache_tree *it,\n> +\t\t\t\t struct strbuf *path)\n> +{\n> +\t/*\n> +\t * Optimistically verify the current tree first. Alternatively\n> +\t * we could verify all the subtrees first then do this\n> +\t * last. Any invalid subtree would also invalidates its\n> +\t * ancestors.\n> +\t */\n> +\tif (it->entry_count != -1 &&\n> +\t    verify_one_cache_tree(to, from, it, path))\n> +\t\tit->entry_count = -1;\n> +\n> +\t/*\n> +\t * If the current tree is valid, don't bother checking\n> +\t * inside. All subtrees _should_ also be valid\n> +\t */\n> +\tif (it->entry_count == -1) {\n> +\t\tint i, len = path->len;\n> +\n> +\t\tfor (i = 0; i < it->subtree_nr; i++) {\n> +\t\t\tstruct cache_tree_sub *down = it->down[i];\n> +\n> +\t\t\tif (!down || !down->cache_tree)\n> +\t\t\t\tcontinue;\n> +\n> +\t\t\tstrbuf_setlen(path, len);\n> +\t\t\tstrbuf_add(path, down->name, down->namelen);\n> +\t\t\tstrbuf_addch(path, '/');\n> +\t\t\tif (verify_and_invalidate(to, from,\n> +\t\t\t\t\t\t  down->cache_tree, path))\n> +\t\t\t\treturn -1;\n> +\t\t}\n> +\t\tstrbuf_setlen(path, len);\n> +\t}\n> +\treturn 0;\n> +}\n> +\n> +static struct cache_tree *duplicate_cache_tree(const struct cache_tree *src)\n> +{\n> +\tstruct cache_tree *dst;\n> +\tint i;\n> +\n> +\tif (!src)\n> +\t\treturn NULL;\n> +\n> +\tdst = xmalloc(sizeof(*dst));\n> +\tdst->entry_count = src->entry_count;\n> +\toidcpy(&dst->oid, &src->oid);\n> +\tdst->subtree_nr = src->subtree_nr;\n> +\tdst->subtree_alloc = dst->subtree_nr;\n> +\tALLOC_ARRAY(dst->down, dst->subtree_alloc);\n> +\tfor (i = 0; i < src->subtree_nr; i++) {\n> +\t\tstruct cache_tree_sub *dsrc = src->down[i];\n> +\t\tstruct cache_tree_sub *down;\n> +\n> +\t\tFLEX_ALLOC_MEM(down, name, dsrc->name, dsrc->namelen);\n> +\t\tdown->count = dsrc->count;\n> +\t\tdown->namelen = dsrc->namelen;\n> +\t\tdown->used = dsrc->used;\n> +\t\tdown->cache_tree = duplicate_cache_tree(dsrc->cache_tree);\n> +\t\tdst->down[i] = down;\n> +\t}\n> +\treturn dst;\n> +}\n> +\n> +int cache_tree_copy(struct index_state *to, const struct index_state *from)\n> +{\n> +\tstruct cache_tree *it = duplicate_cache_tree(from->cache_tree);\n> +\tstruct strbuf path = STRBUF_INIT;\n> +\tint ret;\n> +\n> +\tif (to->cache_tree)\n> +\t\tBUG(\"Sorry merging cache-tree is not supported yet\");\n> +\tret = verify_and_invalidate(to, from, it, &path);\n> +\tto->cache_tree = it;\n> +\tto->cache_changed |= CACHE_TREE_CHANGED;\n> +\tstrbuf_release(&path);\n> +\treturn ret;\n> +}\n> +\n>   static void write_one(struct strbuf *buffer, struct cache_tree *it,\n>                         const char *path, int pathlen)\n>   {\n> diff --git a/cache-tree.h b/cache-tree.h\n> index cfd5328cc9..6981da8e0d 100644\n> --- a/cache-tree.h\n> +++ b/cache-tree.h\n> @@ -53,4 +53,6 @@ void prime_cache_tree(struct index_state *, struct tree *);\n>   \n>   extern int cache_tree_matches_traversal(struct cache_tree *, struct name_entry *ent, struct traverse_info *info);\n>   \n> +int cache_tree_copy(struct index_state *to, const struct index_state *from);\n> +\n>   #endif\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index cd0680f11e..cb3fdd42a6 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -1427,12 +1427,22 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>   \tret = check_updates(o) ? (-2) : 0;\n>   \tif (o->dst_index) {\n>   \t\tif (!ret) {\n> -\t\t\tif (!o->result.cache_tree)\n> +\t\t\tif (!o->result.cache_tree) {\n> +\t\t\t\tuint64_t start = getnanotime();\n> +#if 0\n>   \t\t\t\to->result.cache_tree = cache_tree();\n> -\t\t\tif (!cache_tree_fully_valid(o->result.cache_tree))\n> +#else\n> +\t\t\t\tcache_tree_copy(&o->result, o->src_index);\n> +#endif\n> +\t\t\t\ttrace_performance_since(start, \"copy\");\n> +\t\t\t}\n> +\t\t\tif (!cache_tree_fully_valid(o->result.cache_tree)) {\n> +\t\t\t\tuint64_t start = getnanotime();\n>   \t\t\t\tcache_tree_update(&o->result,\n>   \t\t\t\t\t\t  WRITE_TREE_SILENT |\n>   \t\t\t\t\t\t  WRITE_TREE_REPAIR);\n> +\t\t\t\ttrace_performance_since(start, \"update\");\n> +\t\t\t}\n>   \t\t}\n>   \t\tmove_index_extensions(&o->result, o->src_index);\n>   \t\tdiscard_index(o->dst_index);\n> -- 8< --\n> \n\nI like the idea (and the perf win!) but it seems like there is an \nimportant piece missing.  If I'm reading this correctly, unpack_trees() \nwill copy the source cache tree (instead of creating a new one) and then \nverify_and_invalidate() will walk the cache tree and for any tree that \nis dirty, it will flag its ancestors as dirty as well.\n\nWhat I don't understand is how any cache tree entries that became \ninvalid as a result of the merge of the n-trees are marked as invalid. \nIt seems like something needs to walk the cache tree and call \ncache_tree_invalidate_path() for all entries that changed as a result of \nthe merge before the call to verify_and_invalidate().\n\nI thought at first cache_tree_fully_valid() might do that but it only \nlooks for entries that are already marked as invalid (or are missing \ntheir corresponding object in the object store).  It assumes something \nelse has marked the invalid paths already.\n"},{"id":"354972","messageId":"eb39eecf-81b0-e937-d686-47b7565d6511@gmail.com","threadId":"48913","inReplyTo":"ccec34c9-b81a-bcb4-7d05-48dccc059cc8@gmail.com","subject":"Re: [PATCH v2 0/4] Speed up unpack_trees()","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-09T08:16:03Z","receivedAt":"2018-08-09T08:16:08Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/8/2018 4:53 PM, Ben Peart wrote:\n> \n> \n> On 8/1/2018 12:38 PM, Duy Nguyen wrote:\n>> On Tue, Jul 31, 2018 at 01:31:31PM -0400, Ben Peart wrote:\n>>>\n>>>\n>>> On 7/31/2018 12:50 PM, Ben Peart wrote:\n>>>>\n>>>>\n>>>> On 7/31/2018 11:31 AM, Duy Nguyen wrote:\n>>>\n>>>>>\n>>>>>> In the performance game of whack-a-mole, that call to repair \n>>>>>> cache-tree\n>>>>>> is now looking quite expensive...\n>>>>>\n>>>>> Yeah and I think we can whack that mole too. I did some measurement.\n>>>>> Best case possible, we just need to scan through two indexes (one with\n>>>>> many good cache-tree, one with no cache-tree), compare and copy\n>>>>> cache-tree over. The scanning takes like 1% time of current repair\n>>>>> step and I suspect it's the hashing that takes most of the time. Of\n>>>>> course real world won't have such nice numbers, but I guess we could\n>>>>> maybe half cache-tree update/repair time.\n>>>>>\n>>>>\n>>>> I have some great profiling tools available so will take a look at this\n>>>> next and see exactly where the time is being spent.\n>>>\n>>> Good instincts.  In cache_tree_update, the heavy hitter is definitely\n>>> hash_object_file followed by has_object_file.\n>>>\n>>> Name                                   Inc %         Inc\n>>> + git!cache_tree_update                 12.4       4,935\n>>> |+ git!update_one                       11.8       4,706\n>>> | + git!update_one                      11.8       4,706\n>>> |  + git!hash_object_file                6.1       2,406\n>>> |  + git!has_object_file                 2.0         813\n>>> |  + OTHER <<vcruntime140d!strchr>>      0.5         203\n>>> |  + git!strbuf_addf                     0.4         155\n>>> |  + git!strbuf_release                  0.4         143\n>>> |  + git!strbuf_add                      0.3         121\n>>> |  + OTHER <<vcruntime140d!memcmp>>      0.2          93\n>>> |  + git!strbuf_grow                     0.1          25\n>>\n>> Ben, if you work on this, this could be a good starting point. I will\n>> not work on this because I still have some other things to catch up\n>> and follow through. You can have my sign off if you reuse something\n>> from this patch\n>>\n>> Even if it's a naive implementation, the initial numbers look pretty\n>> good. Without the patch we have\n>>\n>> 18:31:05.970621 unpack-trees.c:1437     performance: 0.000001029 s: copy\n>> 18:31:05.975729 unpack-trees.c:1444     performance: 0.005082004 s: \n>> update\n>>\n>> And with the patch\n>>\n>> 18:31:13.295655 unpack-trees.c:1437     performance: 0.000198017 s: copy\n>> 18:31:13.296757 unpack-trees.c:1444     performance: 0.001075935 s: \n>> update\n>>\n>> Time saving is about 80% by the look of this (best possible case\n>> because only the top tree needs to be hashed and written out).\n>>\n>> -- 8< --\n>> diff --git a/cache-tree.c b/cache-tree.c\n>> index 6b46711996..67a4a93100 100644\n>> --- a/cache-tree.c\n>> +++ b/cache-tree.c\n>> @@ -440,6 +440,147 @@ int cache_tree_update(struct index_state \n>> *istate, int flags)\n>>       return 0;\n>>   }\n>> +static int same(const struct cache_entry *a, const struct cache_entry \n>> *b)\n>> +{\n>> +    if (ce_stage(a) || ce_stage(b))\n>> +        return 0;\n>> +    if ((a->ce_flags | b->ce_flags) & CE_CONFLICTED)\n>> +        return 0;\n>> +    return a->ce_mode == b->ce_mode &&\n>> +           !oidcmp(&a->oid, &b->oid);\n>> +}\n>> +\n>> +static int cache_tree_name_pos(const struct index_state *istate,\n>> +                   const struct strbuf *path)\n>> +{\n>> +    int pos;\n>> +\n>> +    if (!path->len)\n>> +        return 0;\n>> +\n>> +    pos = index_name_pos(istate, path->buf, path->len);\n>> +    if (pos >= 0)\n>> +        BUG(\"No no no, directory path must not exist in index\");\n>> +    return -pos - 1;\n>> +}\n>> +\n>> +/*\n>> + * Locate the same cache-tree in two separate indexes. Check the\n>> + * cache-tree is still valid for the \"to\" index (i.e. it contains the\n>> + * same set of entries in the \"from\" index).\n>> + */\n>> +static int verify_one_cache_tree(const struct index_state *to,\n>> +                 const struct index_state *from,\n>> +                 const struct cache_tree *it,\n>> +                 const struct strbuf *path)\n>> +{\n>> +    int i, spos, dpos;\n>> +\n>> +    spos = cache_tree_name_pos(from, path);\n>> +    if (spos + it->entry_count > from->cache_nr)\n>> +        return -1;\n>> +\n>> +    dpos = cache_tree_name_pos(to, path);\n>> +    if (dpos + it->entry_count > to->cache_nr)\n>> +        return -1;\n>> +\n>> +    /* Can we quickly check head and tail and bail out early */\n>> +    if (!same(from->cache[spos], to->cache[spos]) ||\n>> +        !same(from->cache[spos + it->entry_count - 1],\n>> +          to->cache[spos + it->entry_count - 1]))\n>> +        return -1;\n>> +\n>> +    for (i = 1; i < it->entry_count - 1; i++)\n>> +        if (!same(from->cache[spos + i],\n>> +              to->cache[dpos + i]))\n>> +            return -1;\n>> +\n>> +    return 0;\n>> +}\n>> +\n>> +static int verify_and_invalidate(struct index_state *to,\n>> +                 const struct index_state *from,\n>> +                 struct cache_tree *it,\n>> +                 struct strbuf *path)\n>> +{\n>> +    /*\n>> +     * Optimistically verify the current tree first. Alternatively\n>> +     * we could verify all the subtrees first then do this\n>> +     * last. Any invalid subtree would also invalidates its\n>> +     * ancestors.\n>> +     */\n>> +    if (it->entry_count != -1 &&\n>> +        verify_one_cache_tree(to, from, it, path))\n>> +        it->entry_count = -1;\n>> +\n>> +    /*\n>> +     * If the current tree is valid, don't bother checking\n>> +     * inside. All subtrees _should_ also be valid\n>> +     */\n>> +    if (it->entry_count == -1) {\n>> +        int i, len = path->len;\n>> +\n>> +        for (i = 0; i < it->subtree_nr; i++) {\n>> +            struct cache_tree_sub *down = it->down[i];\n>> +\n>> +            if (!down || !down->cache_tree)\n>> +                continue;\n>> +\n>> +            strbuf_setlen(path, len);\n>> +            strbuf_add(path, down->name, down->namelen);\n>> +            strbuf_addch(path, '/');\n>> +            if (verify_and_invalidate(to, from,\n>> +                          down->cache_tree, path))\n>> +                return -1;\n>> +        }\n>> +        strbuf_setlen(path, len);\n>> +    }\n>> +    return 0;\n>> +}\n>> +\n>> +static struct cache_tree *duplicate_cache_tree(const struct \n>> cache_tree *src)\n>> +{\n>> +    struct cache_tree *dst;\n>> +    int i;\n>> +\n>> +    if (!src)\n>> +        return NULL;\n>> +\n>> +    dst = xmalloc(sizeof(*dst));\n>> +    dst->entry_count = src->entry_count;\n>> +    oidcpy(&dst->oid, &src->oid);\n>> +    dst->subtree_nr = src->subtree_nr;\n>> +    dst->subtree_alloc = dst->subtree_nr;\n>> +    ALLOC_ARRAY(dst->down, dst->subtree_alloc);\n>> +    for (i = 0; i < src->subtree_nr; i++) {\n>> +        struct cache_tree_sub *dsrc = src->down[i];\n>> +        struct cache_tree_sub *down;\n>> +\n>> +        FLEX_ALLOC_MEM(down, name, dsrc->name, dsrc->namelen);\n>> +        down->count = dsrc->count;\n>> +        down->namelen = dsrc->namelen;\n>> +        down->used = dsrc->used;\n>> +        down->cache_tree = duplicate_cache_tree(dsrc->cache_tree);\n>> +        dst->down[i] = down;\n>> +    }\n>> +    return dst;\n>> +}\n>> +\n>> +int cache_tree_copy(struct index_state *to, const struct index_state \n>> *from)\n>> +{\n>> +    struct cache_tree *it = duplicate_cache_tree(from->cache_tree);\n>> +    struct strbuf path = STRBUF_INIT;\n>> +    int ret;\n>> +\n>> +    if (to->cache_tree)\n>> +        BUG(\"Sorry merging cache-tree is not supported yet\");\n>> +    ret = verify_and_invalidate(to, from, it, &path);\n>> +    to->cache_tree = it;\n>> +    to->cache_changed |= CACHE_TREE_CHANGED;\n>> +    strbuf_release(&path);\n>> +    return ret;\n>> +}\n>> +\n>>   static void write_one(struct strbuf *buffer, struct cache_tree *it,\n>>                         const char *path, int pathlen)\n>>   {\n>> diff --git a/cache-tree.h b/cache-tree.h\n>> index cfd5328cc9..6981da8e0d 100644\n>> --- a/cache-tree.h\n>> +++ b/cache-tree.h\n>> @@ -53,4 +53,6 @@ void prime_cache_tree(struct index_state *, struct \n>> tree *);\n>>   extern int cache_tree_matches_traversal(struct cache_tree *, struct \n>> name_entry *ent, struct traverse_info *info);\n>> +int cache_tree_copy(struct index_state *to, const struct index_state \n>> *from);\n>> +\n>>   #endif\n>> diff --git a/unpack-trees.c b/unpack-trees.c\n>> index cd0680f11e..cb3fdd42a6 100644\n>> --- a/unpack-trees.c\n>> +++ b/unpack-trees.c\n>> @@ -1427,12 +1427,22 @@ int unpack_trees(unsigned len, struct \n>> tree_desc *t, struct unpack_trees_options\n>>       ret = check_updates(o) ? (-2) : 0;\n>>       if (o->dst_index) {\n>>           if (!ret) {\n>> -            if (!o->result.cache_tree)\n>> +            if (!o->result.cache_tree) {\n>> +                uint64_t start = getnanotime();\n>> +#if 0\n>>                   o->result.cache_tree = cache_tree();\n>> -            if (!cache_tree_fully_valid(o->result.cache_tree))\n>> +#else\n>> +                cache_tree_copy(&o->result, o->src_index);\n>> +#endif\n>> +                trace_performance_since(start, \"copy\");\n>> +            }\n>> +            if (!cache_tree_fully_valid(o->result.cache_tree)) {\n>> +                uint64_t start = getnanotime();\n>>                   cache_tree_update(&o->result,\n>>                             WRITE_TREE_SILENT |\n>>                             WRITE_TREE_REPAIR);\n>> +                trace_performance_since(start, \"update\");\n>> +            }\n>>           }\n>>           move_index_extensions(&o->result, o->src_index);\n>>           discard_index(o->dst_index);\n>> -- 8< --\n>>\n> \n> I like the idea (and the perf win!) but it seems like there is an \n> important piece missing.  If I'm reading this correctly, unpack_trees() \n> will copy the source cache tree (instead of creating a new one) and then \n> verify_and_invalidate() will walk the cache tree and for any tree that \n> is dirty, it will flag its ancestors as dirty as well.\n> \n> What I don't understand is how any cache tree entries that became \n> invalid as a result of the merge of the n-trees are marked as invalid. \n> It seems like something needs to walk the cache tree and call \n> cache_tree_invalidate_path() for all entries that changed as a result of \n> the merge before the call to verify_and_invalidate().\n> \n> I thought at first cache_tree_fully_valid() might do that but it only \n> looks for entries that are already marked as invalid (or are missing \n> their corresponding object in the object store).  It assumes something \n> else has marked the invalid paths already.\n\nIn fact, in the other [1] patch series, we're detecting the number of \ncache entries that are the same as the cache tree and using that to \ntraverse_by_cache_tree().  At that point, couldn't we copy the \ncorresponding cache tree entries over to the destination so that those \ndon't have to get recreated in the later call to cache_tree_update()?\n\n[1] \nhttps://public-inbox.org/git/20180727154241.GA21288@duynguyen.home/T/#mad6b94733dcf16c29350cbad4beccd9ca93beaed\n"},{"id":"355083","messageId":"CACsJy8B9=HL3mnBVuPEmAR=ukCysxGtpAbr3KY-duL7cQ4D=CQ@mail.gmail.com","threadId":"48913","inReplyTo":"ccec34c9-b81a-bcb4-7d05-48dccc059cc8@gmail.com","subject":"Re: [PATCH v2 0/4] Speed up unpack_trees()","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-10T15:51:43Z","receivedAt":"2018-08-10T15:52:12Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Aug 8, 2018 at 10:53 PM Ben Peart <peartben@gmail.com> wrote:\n>\n>\n>\n> On 8/1/2018 12:38 PM, Duy Nguyen wrote:\n> > On Tue, Jul 31, 2018 at 01:31:31PM -0400, Ben Peart wrote:\n> >>\n> >>\n> >> On 7/31/2018 12:50 PM, Ben Peart wrote:\n> >>>\n> >>>\n> >>> On 7/31/2018 11:31 AM, Duy Nguyen wrote:\n> >>\n> >>>>\n> >>>>> In the performance game of whack-a-mole, that call to repair cache-tree\n> >>>>> is now looking quite expensive...\n> >>>>\n> >>>> Yeah and I think we can whack that mole too. I did some measurement.\n> >>>> Best case possible, we just need to scan through two indexes (one with\n> >>>> many good cache-tree, one with no cache-tree), compare and copy\n> >>>> cache-tree over. The scanning takes like 1% time of current repair\n> >>>> step and I suspect it's the hashing that takes most of the time. Of\n> >>>> course real world won't have such nice numbers, but I guess we could\n> >>>> maybe half cache-tree update/repair time.\n> >>>>\n> >>>\n> >>> I have some great profiling tools available so will take a look at this\n> >>> next and see exactly where the time is being spent.\n> >>\n> >> Good instincts.  In cache_tree_update, the heavy hitter is definitely\n> >> hash_object_file followed by has_object_file.\n> >>\n> >> Name                                 Inc %        Inc\n> >> + git!cache_tree_update               12.4      4,935\n> >> |+ git!update_one                     11.8      4,706\n> >> | + git!update_one                    11.8      4,706\n> >> |  + git!hash_object_file              6.1      2,406\n> >> |  + git!has_object_file               2.0        813\n> >> |  + OTHER <<vcruntime140d!strchr>>    0.5        203\n> >> |  + git!strbuf_addf                   0.4        155\n> >> |  + git!strbuf_release                0.4        143\n> >> |  + git!strbuf_add                    0.3        121\n> >> |  + OTHER <<vcruntime140d!memcmp>>    0.2         93\n> >> |  + git!strbuf_grow                   0.1         25\n> >\n> > Ben, if you work on this, this could be a good starting point. I will\n> > not work on this because I still have some other things to catch up\n> > and follow through. You can have my sign off if you reuse something\n> > from this patch\n> >\n> > Even if it's a naive implementation, the initial numbers look pretty\n> > good. Without the patch we have\n> >\n> > 18:31:05.970621 unpack-trees.c:1437     performance: 0.000001029 s: copy\n> > 18:31:05.975729 unpack-trees.c:1444     performance: 0.005082004 s: update\n> >\n> > And with the patch\n> >\n> > 18:31:13.295655 unpack-trees.c:1437     performance: 0.000198017 s: copy\n> > 18:31:13.296757 unpack-trees.c:1444     performance: 0.001075935 s: update\n> >\n> > Time saving is about 80% by the look of this (best possible case\n> > because only the top tree needs to be hashed and written out).\n> >\n> > -- 8< --\n> > diff --git a/cache-tree.c b/cache-tree.c\n> > index 6b46711996..67a4a93100 100644\n> > --- a/cache-tree.c\n> > +++ b/cache-tree.c\n> > @@ -440,6 +440,147 @@ int cache_tree_update(struct index_state *istate, int flags)\n> >       return 0;\n> >   }\n> >\n> > +static int same(const struct cache_entry *a, const struct cache_entry *b)\n> > +{\n> > +     if (ce_stage(a) || ce_stage(b))\n> > +             return 0;\n> > +     if ((a->ce_flags | b->ce_flags) & CE_CONFLICTED)\n> > +             return 0;\n> > +     return a->ce_mode == b->ce_mode &&\n> > +            !oidcmp(&a->oid, &b->oid);\n> > +}\n> > +\n> > +static int cache_tree_name_pos(const struct index_state *istate,\n> > +                            const struct strbuf *path)\n> > +{\n> > +     int pos;\n> > +\n> > +     if (!path->len)\n> > +             return 0;\n> > +\n> > +     pos = index_name_pos(istate, path->buf, path->len);\n> > +     if (pos >= 0)\n> > +             BUG(\"No no no, directory path must not exist in index\");\n> > +     return -pos - 1;\n> > +}\n> > +\n> > +/*\n> > + * Locate the same cache-tree in two separate indexes. Check the\n> > + * cache-tree is still valid for the \"to\" index (i.e. it contains the\n> > + * same set of entries in the \"from\" index).\n> > + */\n> > +static int verify_one_cache_tree(const struct index_state *to,\n> > +                              const struct index_state *from,\n> > +                              const struct cache_tree *it,\n> > +                              const struct strbuf *path)\n> > +{\n> > +     int i, spos, dpos;\n> > +\n> > +     spos = cache_tree_name_pos(from, path);\n> > +     if (spos + it->entry_count > from->cache_nr)\n> > +             return -1;\n> > +\n> > +     dpos = cache_tree_name_pos(to, path);\n> > +     if (dpos + it->entry_count > to->cache_nr)\n> > +             return -1;\n> > +\n> > +     /* Can we quickly check head and tail and bail out early */\n> > +     if (!same(from->cache[spos], to->cache[spos]) ||\n> > +         !same(from->cache[spos + it->entry_count - 1],\n> > +               to->cache[spos + it->entry_count - 1]))\n> > +             return -1;\n> > +\n> > +     for (i = 1; i < it->entry_count - 1; i++)\n> > +             if (!same(from->cache[spos + i],\n> > +                       to->cache[dpos + i]))\n> > +                     return -1;\n> > +\n> > +     return 0;\n> > +}\n> > +\n> > +static int verify_and_invalidate(struct index_state *to,\n> > +                              const struct index_state *from,\n> > +                              struct cache_tree *it,\n> > +                              struct strbuf *path)\n> > +{\n> > +     /*\n> > +      * Optimistically verify the current tree first. Alternatively\n> > +      * we could verify all the subtrees first then do this\n> > +      * last. Any invalid subtree would also invalidates its\n> > +      * ancestors.\n> > +      */\n> > +     if (it->entry_count != -1 &&\n> > +         verify_one_cache_tree(to, from, it, path))\n> > +             it->entry_count = -1;\n> > +\n> > +     /*\n> > +      * If the current tree is valid, don't bother checking\n> > +      * inside. All subtrees _should_ also be valid\n> > +      */\n> > +     if (it->entry_count == -1) {\n> > +             int i, len = path->len;\n> > +\n> > +             for (i = 0; i < it->subtree_nr; i++) {\n> > +                     struct cache_tree_sub *down = it->down[i];\n> > +\n> > +                     if (!down || !down->cache_tree)\n> > +                             continue;\n> > +\n> > +                     strbuf_setlen(path, len);\n> > +                     strbuf_add(path, down->name, down->namelen);\n> > +                     strbuf_addch(path, '/');\n> > +                     if (verify_and_invalidate(to, from,\n> > +                                               down->cache_tree, path))\n> > +                             return -1;\n> > +             }\n> > +             strbuf_setlen(path, len);\n> > +     }\n> > +     return 0;\n> > +}\n> > +\n> > +static struct cache_tree *duplicate_cache_tree(const struct cache_tree *src)\n> > +{\n> > +     struct cache_tree *dst;\n> > +     int i;\n> > +\n> > +     if (!src)\n> > +             return NULL;\n> > +\n> > +     dst = xmalloc(sizeof(*dst));\n> > +     dst->entry_count = src->entry_count;\n> > +     oidcpy(&dst->oid, &src->oid);\n> > +     dst->subtree_nr = src->subtree_nr;\n> > +     dst->subtree_alloc = dst->subtree_nr;\n> > +     ALLOC_ARRAY(dst->down, dst->subtree_alloc);\n> > +     for (i = 0; i < src->subtree_nr; i++) {\n> > +             struct cache_tree_sub *dsrc = src->down[i];\n> > +             struct cache_tree_sub *down;\n> > +\n> > +             FLEX_ALLOC_MEM(down, name, dsrc->name, dsrc->namelen);\n> > +             down->count = dsrc->count;\n> > +             down->namelen = dsrc->namelen;\n> > +             down->used = dsrc->used;\n> > +             down->cache_tree = duplicate_cache_tree(dsrc->cache_tree);\n> > +             dst->down[i] = down;\n> > +     }\n> > +     return dst;\n> > +}\n> > +\n> > +int cache_tree_copy(struct index_state *to, const struct index_state *from)\n> > +{\n> > +     struct cache_tree *it = duplicate_cache_tree(from->cache_tree);\n> > +     struct strbuf path = STRBUF_INIT;\n> > +     int ret;\n> > +\n> > +     if (to->cache_tree)\n> > +             BUG(\"Sorry merging cache-tree is not supported yet\");\n> > +     ret = verify_and_invalidate(to, from, it, &path);\n> > +     to->cache_tree = it;\n> > +     to->cache_changed |= CACHE_TREE_CHANGED;\n> > +     strbuf_release(&path);\n> > +     return ret;\n> > +}\n> > +\n> >   static void write_one(struct strbuf *buffer, struct cache_tree *it,\n> >                         const char *path, int pathlen)\n> >   {\n> > diff --git a/cache-tree.h b/cache-tree.h\n> > index cfd5328cc9..6981da8e0d 100644\n> > --- a/cache-tree.h\n> > +++ b/cache-tree.h\n> > @@ -53,4 +53,6 @@ void prime_cache_tree(struct index_state *, struct tree *);\n> >\n> >   extern int cache_tree_matches_traversal(struct cache_tree *, struct name_entry *ent, struct traverse_info *info);\n> >\n> > +int cache_tree_copy(struct index_state *to, const struct index_state *from);\n> > +\n> >   #endif\n> > diff --git a/unpack-trees.c b/unpack-trees.c\n> > index cd0680f11e..cb3fdd42a6 100644\n> > --- a/unpack-trees.c\n> > +++ b/unpack-trees.c\n> > @@ -1427,12 +1427,22 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n> >       ret = check_updates(o) ? (-2) : 0;\n> >       if (o->dst_index) {\n> >               if (!ret) {\n> > -                     if (!o->result.cache_tree)\n> > +                     if (!o->result.cache_tree) {\n> > +                             uint64_t start = getnanotime();\n> > +#if 0\n> >                               o->result.cache_tree = cache_tree();\n> > -                     if (!cache_tree_fully_valid(o->result.cache_tree))\n> > +#else\n> > +                             cache_tree_copy(&o->result, o->src_index);\n> > +#endif\n> > +                             trace_performance_since(start, \"copy\");\n> > +                     }\n> > +                     if (!cache_tree_fully_valid(o->result.cache_tree)) {\n> > +                             uint64_t start = getnanotime();\n> >                               cache_tree_update(&o->result,\n> >                                                 WRITE_TREE_SILENT |\n> >                                                 WRITE_TREE_REPAIR);\n> > +                             trace_performance_since(start, \"update\");\n> > +                     }\n> >               }\n> >               move_index_extensions(&o->result, o->src_index);\n> >               discard_index(o->dst_index);\n> > -- 8< --\n> >\n>\n> I like the idea (and the perf win!) but it seems like there is an\n> important piece missing.  If I'm reading this correctly, unpack_trees()\n> will copy the source cache tree (instead of creating a new one) and then\n> verify_and_invalidate() will walk the cache tree and for any tree that\n> is dirty, it will flag its ancestors as dirty as well.\n\nThat, and the verification part. The for loop at the bottom of\nverify_one_cache_tree() makes sure that the cache-tree is valid. That\nis, if we recreate cache-tree from scratch in the destination index,\nit should produce the same OID as the cache-tree we copy over.\n\nBut I think I'm a bit loose in that check. Suppose in the source index we have\n\nabc\nfoo/abc\nfoo/def\nxyz\n\nthe for loop tries to make sure that for cache-tree of 'foo', the\ndestination index must have foo/abc and foo/def (with same mode,\noid....) but it fails to catch this\n\nabc\nfoo/abc\nfoo/def\nfoo/xyz\nxyz\n\nIf we recreate cache-tree from scratch, the cache-tree for 'foo'\nshould cover three items and have different oid than one we copied\nfrom the source index. Same problem could happen if we have something\nin foo, but before foo/abc.\n\n> What I don't understand is how any cache tree entries that became\n> invalid as a result of the merge of the n-trees are marked as invalid.\n> It seems like something needs to walk the cache tree and call\n> cache_tree_invalidate_path() for all entries that changed as a result of\n> the merge before the call to verify_and_invalidate().\n\nI'm not sure I understand but anyway the way I understand it, when we\nmerge from o->src_index to o->result, we start o->result with empty\ncache-tree. There's nothing in there to invalidate, even though we do\ncall cache_tree_invalidate_path() (from invalidate_ce_path() in\nunpack-trees.c)\n\nI don't think the merge operation is related to this at all. This\nproblem can be stated as \"I have a set of good cache-trees that are\nassociated with index 'A' and a new index 'B' with no cache-tree at\nall. Can I (cheaply) reuse some  cache-tree from 'A'?\". My answer in\nthis patch is yes, for each cache-tree in A, make sure that the list\nof cache-entries associated with it is present in B (which is almost\ncorrect).\n-- \nDuy\n"},{"id":"355086","messageId":"CACsJy8AOV+73RLx2GyWTvcoKxop2t_9x0mFjN9raOsGjDdQ2bg@mail.gmail.com","threadId":"48913","inReplyTo":"eb39eecf-81b0-e937-d686-47b7565d6511@gmail.com","subject":"Re: [PATCH v2 0/4] Speed up unpack_trees()","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-10T16:08:55Z","receivedAt":"2018-08-10T16:09:24Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Aug 9, 2018 at 10:16 AM Ben Peart <peartben@gmail.com> wrote:\n> In fact, in the other [1] patch series, we're detecting the number of\n> cache entries that are the same as the cache tree and using that to\n> traverse_by_cache_tree().  At that point, couldn't we copy the\n> corresponding cache tree entries over to the destination so that those\n> don't have to get recreated in the later call to cache_tree_update()?\n\nWe could. But as I stated in another mail, I saw this cache-tree\noptimization as a separate problem and didn't want to mix them up.\nThat way cache_tree_copy() could be used elsewhere if the opportunity\nshows up.\n\nMixing them up could also complicate the problem. You you merge stuff,\nyou add new cache-entries to o->result with add_index_entry() which\ntries to invalidate those paths in o->result's cache-tree. Right now\nthe cache-tree is empty so it's really no-op. But if you copy\ncache-tree over while merging, that invalidation might either\ninvalidate your newly copied cache-tree, or get slowed down because\nnon-empty o->result's cache-tree means you start to need to walk it to\nfind if there's any path to invalidate.\n\nPS. This code keeps messing me up. invalidate_ce_path() may also\ninvalidate cache-tree in the _source_ index. For this optimization to\nreally shine, you better keep the the original cache-tree intact (so\nthat you can reuse as much as possible).\n\nI don't see the purpose of this source cache tree invalidation at all.\nMy guess at this point is Linus actually made a mistake in 34110cd4e3\n(Make 'unpack_trees()' have a separate source and destination index -\n2008-03-06) and he should have invalidated _destination_ index instead\nof the source one. I'm going to dig in some more and probably will\nsend a patch to remove this invalidation.\n-- \nDuy\n"},{"id":"355089","messageId":"CACsJy8BtgMSYqkD1EaFQ=S49BA-veyTO1qU0FaPMkHY-KeggfA@mail.gmail.com","threadId":"48913","inReplyTo":"CABPp-BGcPV0RA624_1UOXYkvaNhW4yR2ifhV_MVFZQOgBb_Ydg@mail.gmail.com","subject":"Re: [PATCH v3 2/4] unpack-trees: optimize walking same trees with cache-tree","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-10T16:29:23Z","receivedAt":"2018-08-10T16:29:55Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Aug 8, 2018 at 8:23 PM Elijah Newren <newren@gmail.com> wrote:\n> > diff --git a/unpack-trees.c b/unpack-trees.c\n> > index a32ddee159..ba3d2e947e 100644\n> > --- a/unpack-trees.c\n> > +++ b/unpack-trees.c\n> > @@ -644,6 +644,102 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n> >         return name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n> >  }\n> >\n> > +static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n> > +                                       struct name_entry *names,\n> > +                                       struct traverse_info *info)\n> > +{\n> > +       struct unpack_trees_options *o = info->data;\n> > +       int i;\n> > +\n> > +       if (!o->merge || dirmask != ((1 << n) - 1))\n> > +               return 0;\n> > +\n> > +       for (i = 1; i < n; i++)\n> > +               if (!are_same_oid(names, names + i))\n> > +                       return 0;\n> > +\n> > +       return cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n> > +}\n>\n> I was curious whether this could also be extended in the case of a\n> merge; as long as HEAD and MERGE have the same tree, even if the base\n> commit doesn't match, we can still just use the tree from HEAD which\n> should be in the current index/cache_tree.  However, it'd be a\n> somewhat odd history for HEAD and MERGE to match on some significantly\n> sized tree when the base commit doesn't also match.\n\nI did have 3-way merge in mind when I wrote this patch. Yes it's\nunlikely except one case (I think). Consider a large \"mono repo\" that\ncontains stuff from many teams. When you branch out for your own team,\nthen most of your changes will be in a few directories, the rest of\nthe code base untouched. In that case we could have a lot of same\ntrees in subdirectories outside the stuff your team touches. This of\ncourse assumes that your team keeps the same base static for some\ntime, not constantly rebasing/merging on top of 'master'.\n\n> > +       /*\n> > +        * Do what unpack_callback() and unpack_nondirectories() normally\n> > +        * do. But we walk all paths recursively in just one loop instead.\n> > +        *\n> > +        * D/F conflicts and staged entries are not a concern because\n>\n> \"staged entries\"?  Do you mean \"higher stage entries\"?  I'm not sure\n> the correct terminology here, but the former makes me think of changes\n> the user has staged but not committed (i.e. stuff found at stage #0 in\n> the index, but which isn't found in any tree yet) vs. the latter which\n> I'd use to refer to entries at stages 1 or higher.\n\nYep stage 1 or higher (I was thinking ce_stage() when I wrote this).\nWill clarify.\n\n\n> > +        * cache-tree would be invalidated and we would never get here\n> > +        * in the first place.\n> > +        */\n> > +       for (i = 0; i < nr_entries; i++) {\n> > +               struct cache_entry *tree_ce;\n> > +               int len, rc;\n> > +\n> > +               src[0] = o->src_index->cache[pos + i];\n> > +\n> > +               len = ce_namelen(src[0]);\n> > +               tree_ce = xcalloc(1, cache_entry_size(len));\n> > +\n> > +               tree_ce->ce_mode = src[0]->ce_mode;\n> > +               tree_ce->ce_flags = create_ce_flags(0);\n> > +               tree_ce->ce_namelen = len;\n> > +               oidcpy(&tree_ce->oid, &src[0]->oid);\n> > +               memcpy(tree_ce->name, src[0]->name, len + 1);\n>\n> We do a bunch of work to setup tree_ce...\n>\n> > +               for (d = 1; d <= nr_names; d++)\n> > +                       src[d] = tree_ce;\n>\n> ...then we make nr_names copies of tree_ce (so that *way_merge or\n> bind_merge or oneway_diff or whatever will have the expected number of\n> entries).\n>\n> > +               rc = call_unpack_fn((const struct cache_entry * const *)src, o);\n>\n> ...then we call o->fn (via call_unpack_fn) to do various complicated\n> logic to figure out which tree_ce to use??  Isn't that just an\n> expensive way to recompute that what we currently have in the index is\n> what we want to keep there?\n>\n> Granted, a caller of this may have set o->fn to something other than\n> {one,two,three}way_merge (or bind_merge), and that function might have\n> important side effects...but it just seems annoying to have to do so\n> much work when for most uses we already know the entry in the index is\n> the one we already want.\n\nI'm not so sure about that. Which is why I keep it generic.\n\n> In fact, the only other thing in the\n> codebase that o->fn is now set to is oneway_diff, which I think is a\n> no-op when the two trees match.\n>\n> Would be nice if we could avoid all this, at least in the common cases\n> where o->fn is a function known to not have side effects.  Or did I\n> not read those functions closely enough and they do have important\n> side effects?\n\nIn one of my earlier \"how about this\" attempts, I introduced fn_same\n[1] that can help achieve this without carving \"known not to have side\neffects\" in common code. Which I think is still a good direction to go\nif we want to optimize more aggressively. We could have something like\nthis\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 1f11991a51..01b80389e0 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -699,6 +699,9 @@ static int traverse_by_cache_tree(int pos, int\nnr_entries, int nr_names,\n        int ce_len = 0;\n        int i, d;\n\n+       if (o->fn_cache_tree)\n+               return o->fn_cache_tree(pos, nr_entries, nr_names, names, info);\n+\n        if (!o->merge)\n                BUG(\"We need cache-tree to do this optimization\");\n\nthen you can add, say threeway_cache_tree_merge(), that does what\ntraverse_by_cache_tree() does but more efficient. This involves a lot\nmore work (mostly staring and those n-merge functions and making sure\nyou don't set the right conditions before going the fast path).\n\nI didn't do it because.. well.. it's more work and also riskier. I\nthink we can leave that for later, unless you think we should do it\nnow.\n\n[1] https://public-inbox.org/git/20180726163049.GA15572@duynguyen.home/\n-- \nDuy\n"},{"id":"355090","messageId":"CACsJy8DF5XLf-RF3SwTpRynYALJUPO_VTK=fpx1oabwB80ZpPw@mail.gmail.com","threadId":"48913","inReplyTo":"CABPp-BGF+GZjm-DiveLjFOESKwPz2F0Y7X4_kXyem2xFo2odUw@mail.gmail.com","subject":"Re: [PATCH v2 4/4] unpack-trees: cheaper index update when walking by cache-tree","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-10T16:39:01Z","receivedAt":"2018-08-10T16:39:30Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Aug 8, 2018 at 8:46 PM Elijah Newren <newren@gmail.com> wrote:\n> > @@ -701,6 +702,24 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n> >         if (!o->merge)\n> >                 BUG(\"We need cache-tree to do this optimization\");\n> >\n> > +       /*\n> > +        * Try to keep add_index_entry() as fast as possible since\n> > +        * we're going to do a lot of them.\n> > +        *\n> > +        * Skipping verify_path() should totally be safe because these\n> > +        * paths are from the source index, which must have been\n> > +        * verified.\n> > +        *\n> > +        * Skipping D/F and cache-tree validation checks is trickier\n> > +        * because it assumes what n-merge code would do when all\n> > +        * trees and the index are the same. We probably could just\n> > +        * optimize those code instead (e.g. we don't invalidate that\n> > +        * many cache-tree, but the searching for them is very\n> > +        * expensive).\n> > +        */\n> > +       o->extra_add_index_flags = ADD_CACHE_SKIP_DFCHECK;\n> > +       o->extra_add_index_flags |= ADD_CACHE_SKIP_VERIFY_PATH;\n> > +\n>\n> In sum of this whole patch, you notice that the Nway_merge functions\n> are still a bit of a bottleneck, but you know you have a special case\n> where you want them to put an entry in the index that matches what is\n> already there, so you try to set some extra flags to short-circuit\n> part of their logic and get to what you know is the correct result.\n>\n> This seems a little scary to me.  I think it's probably safe as long\n> as o->fn is one of {oneway_merge, twoway_merge, threeway_merge,\n> bind_merge} (the cases you have in mind and which the current code\n> uses), but the caller isn't limited to those.  Right now in\n> diff-lib.c, there's a caller that has their own function, oneway_diff.\n> More could be added in the future.\n>\n> If we're going to go this route, I think we should first check that\n> o->fn is one of those known safe functions.  And if we're going that\n> route, the comments I bring up on patch 2 about possibly avoiding\n> call_unpack_fn() altogether might even obviate this patch while\n> speeding things up more.\n\nYes I do need to check o->fn. I might have to think more about\navoiding call_unpack_fn(). Even if we avoid it though, we still go\nthrough add_index_entry() and suffer the same checks every time unless\nwe do somethine like this (but then of course it's safer because\nyou're doing it in a specific x-way merge, not generic code like\nthis).\n\n> > @@ -1561,6 +1581,13 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n> >                 if (!ret) {\n> >                         if (!o->result.cache_tree)\n> >                                 o->result.cache_tree = cache_tree();\n> > +                       /*\n> > +                        * TODO: Walk o.src_index->cache_tree, quickly check\n> > +                        * if o->result.cache has the exact same content for\n> > +                        * any valid cache-tree in o.src_index, then we can\n> > +                        * just copy the cache-tree over instead of hashing a\n> > +                        * new tree object.\n> > +                        */\n>\n> Interesting.  I really don't know how cache_tree works...but if we\n> avoided calling call_unpack_fn, and thus left the original index entry\n> in place instead of replacing it with an equal one, would that as a\n> side effect speed up the cache_tree_valid/cache_tree_update calls for\n> us?  Or is there still work here?\n\nNaah. Notice that we don't care at all about the source's cache-tree\nwhen we update o->result one (and we never ever do anything about\no->result's cache-tree during the merge). Whether you invalidate or\nnot, o->result's cache-tree is always empty and you still have to\nrecreate all cache-tree in o->result. You essentially play full cost\nof \"git write-tree\" here if I'm not mistaken.\n-- \nDuy\n"},{"id":"355115","messageId":"CACsJy8Ag3N6-A7YLOxBkE-aEUwB+PwNQr63GhiYoxjAsuRmQ5w@mail.gmail.com","threadId":"48913","inReplyTo":"xmqqeff8hfdg.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 0/4] Speed up unpack_trees()","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-10T16:53:43Z","receivedAt":"2018-08-10T16:54:12Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Aug 8, 2018 at 8:39 PM Junio C Hamano <gitster@pobox.com> wrote:\n> One more, and hopefully the final, note.\n>\n> ..\n\nMuch appreciated. I don't think I could figure all this out no matter\nhow long I stare at those commits and current code.\n-- \nDuy\n"},{"id":"355128","messageId":"CABPp-BGU6QnUwQgkhwx6vLBc3ozoEScQ4DaZd-9ZZfQhXfxPww@mail.gmail.com","threadId":"48913","inReplyTo":"CACsJy8DF5XLf-RF3SwTpRynYALJUPO_VTK=fpx1oabwB80ZpPw@mail.gmail.com","subject":"Re: [PATCH v2 4/4] unpack-trees: cheaper index update when walking by cache-tree","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-08-10T18:39:45Z","receivedAt":"2018-08-10T18:40:00Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Fri, Aug 10, 2018 at 9:39 AM Duy Nguyen <pclouds@gmail.com> wrote:\n>\n> On Wed, Aug 8, 2018 at 8:46 PM Elijah Newren <newren@gmail.com> wrote:\n> > > @@ -701,6 +702,24 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n\n> > If we're going to go this route, I think we should first check that\n> > o->fn is one of those known safe functions.  And if we're going that\n> > route, the comments I bring up on patch 2 about possibly avoiding\n> > call_unpack_fn() altogether might even obviate this patch while\n> > speeding things up more.\n>\n> Yes I do need to check o->fn. I might have to think more about\n> avoiding call_unpack_fn(). Even if we avoid it though, we still go\n> through add_index_entry() and suffer the same checks every time unless\n> we do somethine like this (but then of course it's safer because\n> you're doing it in a specific x-way merge, not generic code like\n> this).\n\nWhy do we still need to go through add_index_entry()?  I thought that\nthe whole point was that you already checked that at the current path,\nthe trees being unpacked were all equal and matched both the index and\nthe cache_tree.  If so, why is there any need for an update at all?\n(Did I read your all_trees_same_as_cache_tree() function wrong, and\nyou don't actually know these all match in some important way?)\n\n\n> > > @@ -1561,6 +1581,13 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n> > >                 if (!ret) {\n> > >                         if (!o->result.cache_tree)\n> > >                                 o->result.cache_tree = cache_tree();\n> > > +                       /*\n> > > +                        * TODO: Walk o.src_index->cache_tree, quickly check\n> > > +                        * if o->result.cache has the exact same content for\n> > > +                        * any valid cache-tree in o.src_index, then we can\n> > > +                        * just copy the cache-tree over instead of hashing a\n> > > +                        * new tree object.\n> > > +                        */\n> >\n> > Interesting.  I really don't know how cache_tree works...but if we\n> > avoided calling call_unpack_fn, and thus left the original index entry\n> > in place instead of replacing it with an equal one, would that as a\n> > side effect speed up the cache_tree_valid/cache_tree_update calls for\n> > us?  Or is there still work here?\n>\n> Naah. Notice that we don't care at all about the source's cache-tree\n> when we update o->result one (and we never ever do anything about\n> o->result's cache-tree during the merge). Whether you invalidate or\n> not, o->result's cache-tree is always empty and you still have to\n> recreate all cache-tree in o->result. You essentially play full cost\n> of \"git write-tree\" here if I'm not mistaken.\n\nOh...perhaps that answers my question above.  So we have to call\nadd_index_entry() for the side effect of populating the new\ncache_tree?\n"},{"id":"355131","messageId":"CABPp-BEOFSU2k+DKuTQZtz+c6eboiUo3RDzBZHkB6=V3SFAigQ@mail.gmail.com","threadId":"48913","inReplyTo":"CACsJy8BtgMSYqkD1EaFQ=S49BA-veyTO1qU0FaPMkHY-KeggfA@mail.gmail.com","subject":"Re: [PATCH v3 2/4] unpack-trees: optimize walking same trees with cache-tree","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-08-10T18:48:33Z","receivedAt":"2018-08-10T18:48:47Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Fri, Aug 10, 2018 at 9:29 AM Duy Nguyen <pclouds@gmail.com> wrote:\n> On Wed, Aug 8, 2018 at 8:23 PM Elijah Newren <newren@gmail.com> wrote:\n\n> > > +        * cache-tree would be invalidated and we would never get here\n> > > +        * in the first place.\n> > > +        */\n> > > +       for (i = 0; i < nr_entries; i++) {\n> > > +               struct cache_entry *tree_ce;\n> > > +               int len, rc;\n> > > +\n> > > +               src[0] = o->src_index->cache[pos + i];\n> > > +\n> > > +               len = ce_namelen(src[0]);\n> > > +               tree_ce = xcalloc(1, cache_entry_size(len));\n> > > +\n> > > +               tree_ce->ce_mode = src[0]->ce_mode;\n> > > +               tree_ce->ce_flags = create_ce_flags(0);\n> > > +               tree_ce->ce_namelen = len;\n> > > +               oidcpy(&tree_ce->oid, &src[0]->oid);\n> > > +               memcpy(tree_ce->name, src[0]->name, len + 1);\n> >\n> > We do a bunch of work to setup tree_ce...\n> >\n> > > +               for (d = 1; d <= nr_names; d++)\n> > > +                       src[d] = tree_ce;\n> >\n> > ...then we make nr_names copies of tree_ce (so that *way_merge or\n> > bind_merge or oneway_diff or whatever will have the expected number of\n> > entries).\n> >\n> > > +               rc = call_unpack_fn((const struct cache_entry * const *)src, o);\n> >\n> > ...then we call o->fn (via call_unpack_fn) to do various complicated\n> > logic to figure out which tree_ce to use??  Isn't that just an\n> > expensive way to recompute that what we currently have in the index is\n> > what we want to keep there?\n> >\n> > Granted, a caller of this may have set o->fn to something other than\n> > {one,two,three}way_merge (or bind_merge), and that function might have\n> > important side effects...but it just seems annoying to have to do so\n> > much work when for most uses we already know the entry in the index is\n> > the one we already want.\n>\n> I'm not so sure about that. Which is why I keep it generic.\n>\n> > In fact, the only other thing in the\n> > codebase that o->fn is now set to is oneway_diff, which I think is a\n> > no-op when the two trees match.\n> >\n> > Would be nice if we could avoid all this, at least in the common cases\n> > where o->fn is a function known to not have side effects.  Or did I\n> > not read those functions closely enough and they do have important\n> > side effects?\n>\n> In one of my earlier \"how about this\" attempts, I introduced fn_same\n> [1] that can help achieve this without carving \"known not to have side\n> effects\" in common code. Which I think is still a good direction to go\n> if we want to optimize more aggressively. We could have something like\n> this\n>\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index 1f11991a51..01b80389e0 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -699,6 +699,9 @@ static int traverse_by_cache_tree(int pos, int\n> nr_entries, int nr_names,\n>         int ce_len = 0;\n>         int i, d;\n>\n> +       if (o->fn_cache_tree)\n> +               return o->fn_cache_tree(pos, nr_entries, nr_names, names, info);\n> +\n>         if (!o->merge)\n>                 BUG(\"We need cache-tree to do this optimization\");\n>\n> then you can add, say threeway_cache_tree_merge(), that does what\n> traverse_by_cache_tree() does but more efficient. This involves a lot\n> more work (mostly staring and those n-merge functions and making sure\n> you don't set the right conditions before going the fast path).\n>\n> I didn't do it because.. well.. it's more work and also riskier. I\n> think we can leave that for later, unless you think we should do it\n> now.\n>\n> [1] https://public-inbox.org/git/20180726163049.GA15572@duynguyen.home/\n\nYeah, from your other thread, I think I was missing some of the\nintracacies of how the cache-tree works and the extra work that'd be\nneeded to bring it along. Deferring until later makes sense.\n"},{"id":"355137","messageId":"CACsJy8AeptcqwRC+DOrdhvk69kEQT6+S6M=0OGWBFOE5gihGzA@mail.gmail.com","threadId":"48913","inReplyTo":"CABPp-BGU6QnUwQgkhwx6vLBc3ozoEScQ4DaZd-9ZZfQhXfxPww@mail.gmail.com","subject":"Re: [PATCH v2 4/4] unpack-trees: cheaper index update when walking by cache-tree","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-10T19:30:05Z","receivedAt":"2018-08-10T19:30:34Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Aug 10, 2018 at 8:39 PM Elijah Newren <newren@gmail.com> wrote:\n>\n> On Fri, Aug 10, 2018 at 9:39 AM Duy Nguyen <pclouds@gmail.com> wrote:\n> >\n> > On Wed, Aug 8, 2018 at 8:46 PM Elijah Newren <newren@gmail.com> wrote:\n> > > > @@ -701,6 +702,24 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n>\n> > > If we're going to go this route, I think we should first check that\n> > > o->fn is one of those known safe functions.  And if we're going that\n> > > route, the comments I bring up on patch 2 about possibly avoiding\n> > > call_unpack_fn() altogether might even obviate this patch while\n> > > speeding things up more.\n> >\n> > Yes I do need to check o->fn. I might have to think more about\n> > avoiding call_unpack_fn(). Even if we avoid it though, we still go\n> > through add_index_entry() and suffer the same checks every time unless\n> > we do somethine like this (but then of course it's safer because\n> > you're doing it in a specific x-way merge, not generic code like\n> > this).\n>\n> Why do we still need to go through add_index_entry()?  I thought that\n> the whole point was that you already checked that at the current path,\n> the trees being unpacked were all equal and matched both the index and\n> the cache_tree.  If so, why is there any need for an update at all?\n> (Did I read your all_trees_same_as_cache_tree() function wrong, and\n> you don't actually know these all match in some important way?)\n\nUnless fn is oneway_diff, we have to create a new index (in o->result)\nbased on o->src_index and some other trees. So we have to add entries\nto o->result and add_index_entry() is the way to do that (granted if\nwe feel confident we could add ADD_CACHE_JUST_APPEND which makes it\nsuper cheap). This is the outcome of n-way merge,\n\nall_trees_same_as_cache_tree() only gurantees the input condition (all\ntrees the same, index also the same) but it can't affect what fn does.\nI don't think we can just simply skip and not update anything (like\no->diff_index_cached case) because o->result would be empty in the\nend. And we need to create (temporary) o->result before we can swap it\nto o->dst_index as the result of a merge operation.\n\n> > > > @@ -1561,6 +1581,13 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n> > > >                 if (!ret) {\n> > > >                         if (!o->result.cache_tree)\n> > > >                                 o->result.cache_tree = cache_tree();\n> > > > +                       /*\n> > > > +                        * TODO: Walk o.src_index->cache_tree, quickly check\n> > > > +                        * if o->result.cache has the exact same content for\n> > > > +                        * any valid cache-tree in o.src_index, then we can\n> > > > +                        * just copy the cache-tree over instead of hashing a\n> > > > +                        * new tree object.\n> > > > +                        */\n> > >\n> > > Interesting.  I really don't know how cache_tree works...but if we\n> > > avoided calling call_unpack_fn, and thus left the original index entry\n> > > in place instead of replacing it with an equal one, would that as a\n> > > side effect speed up the cache_tree_valid/cache_tree_update calls for\n> > > us?  Or is there still work here?\n> >\n> > Naah. Notice that we don't care at all about the source's cache-tree\n> > when we update o->result one (and we never ever do anything about\n> > o->result's cache-tree during the merge). Whether you invalidate or\n> > not, o->result's cache-tree is always empty and you still have to\n> > recreate all cache-tree in o->result. You essentially play full cost\n> > of \"git write-tree\" here if I'm not mistaken.\n>\n> Oh...perhaps that answers my question above.  So we have to call\n> add_index_entry() for the side effect of populating the new\n> cache_tree?\n\nI have a feeling that you're thinking we can swap o->src_index to\no->dst_index at the end? That might explain your confusion about\no->result (or I misread your replies horribly) and the original\nindex...\n-- \nDuy\n"},{"id":"355140","messageId":"CABPp-BHMC8k3t2_9KzdJvg80e-nqwsbLUceTLNjQ=ST=9XthEA@mail.gmail.com","threadId":"48913","inReplyTo":"CACsJy8AeptcqwRC+DOrdhvk69kEQT6+S6M=0OGWBFOE5gihGzA@mail.gmail.com","subject":"Re: [PATCH v2 4/4] unpack-trees: cheaper index update when walking by cache-tree","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-08-10T19:40:19Z","receivedAt":"2018-08-10T19:40:33Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Fri, Aug 10, 2018 at 12:30 PM Duy Nguyen <pclouds@gmail.com> wrote:\n> On Fri, Aug 10, 2018 at 8:39 PM Elijah Newren <newren@gmail.com> wrote:\n...\n> > Why do we still need to go through add_index_entry()?  I thought that\n> > the whole point was that you already checked that at the current path,\n> > the trees being unpacked were all equal and matched both the index and\n> > the cache_tree.  If so, why is there any need for an update at all?\n> > (Did I read your all_trees_same_as_cache_tree() function wrong, and\n> > you don't actually know these all match in some important way?)\n>\n> Unless fn is oneway_diff, we have to create a new index (in o->result)\n> based on o->src_index and some other trees. So we have to add entries\n\nOh, right, o->src_index may not equal o->dst_index (because of people\nlike me who call it that way from merge-recursive.c) and even if it\ndoes, we still have the temporary o->result in the mean time.  I\nshould have remembered that; just didn't.\n\n> to o->result and add_index_entry() is the way to do that (granted if\n> we feel confident we could add ADD_CACHE_JUST_APPEND which makes it\n> super cheap). This is the outcome of n-way merge,\n>\n> all_trees_same_as_cache_tree() only gurantees the input condition (all\n> trees the same, index also the same) but it can't affect what fn does.\n> I don't think we can just simply skip and not update anything (like\n> o->diff_index_cached case) because o->result would be empty in the\n> end. And we need to create (temporary) o->result before we can swap it\n> to o->dst_index as the result of a merge operation.\n>\n\n...\n\n> I have a feeling that you're thinking we can swap o->src_index to\n> o->dst_index at the end? That might explain your confusion about\n> o->result (or I misread your replies horribly) and the original\n> index...\n\nYeah, thanks for figuring out my confusion and jogging my memory.\n"},{"id":"355146","messageId":"CACsJy8D+UcmokEVn-=PRp7cZMK9fY0H+epKJpR8ytLSJdjWHcg@mail.gmail.com","threadId":"48913","inReplyTo":"CABPp-BHMC8k3t2_9KzdJvg80e-nqwsbLUceTLNjQ=ST=9XthEA@mail.gmail.com","subject":"Re: [PATCH v2 4/4] unpack-trees: cheaper index update when walking by cache-tree","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-10T19:48:30Z","receivedAt":"2018-08-10T19:48:59Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Aug 10, 2018 at 9:40 PM Elijah Newren <newren@gmail.com> wrote:\n>\n> On Fri, Aug 10, 2018 at 12:30 PM Duy Nguyen <pclouds@gmail.com> wrote:\n> > On Fri, Aug 10, 2018 at 8:39 PM Elijah Newren <newren@gmail.com> wrote:\n> ...\n> > > Why do we still need to go through add_index_entry()?  I thought that\n> > > the whole point was that you already checked that at the current path,\n> > > the trees being unpacked were all equal and matched both the index and\n> > > the cache_tree.  If so, why is there any need for an update at all?\n> > > (Did I read your all_trees_same_as_cache_tree() function wrong, and\n> > > you don't actually know these all match in some important way?)\n> >\n> > Unless fn is oneway_diff, we have to create a new index (in o->result)\n> > based on o->src_index and some other trees. So we have to add entries\n>\n> Oh, right, o->src_index may not equal o->dst_index (because of people\n> like me who call it that way from merge-recursive.c) and even if it\n> does, we still have the temporary o->result in the mean time.  I\n> should have remembered that; just didn't.\n\nYour forgetting about this actually helps. I think the idea of\navoiding add_index_entry() may be worth considering.\n\nWe know that 90% of cases of unpack_trees() is from the_index to\nthe_index. So if instead of creating a full temporary index, where 90%\nof it might be the same as source index, if we just mark in the source\nindex (e.g. in ce_flags) the entries that should be copied to\no->result and _not_ create them in o->result. When it's time to create\no->dst_index (which is the_index) from o->result, we could just do\nlittle manipulation to delete stuff that the_index has but o->result\ndoes not and add a bit more things. It is something that at least\nsounds nice in my head, but I'm not sure if it works out...\n-- \nDuy\n"},{"id":"355291","messageId":"20180812081551.27927-1-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180804053723.4695-1-pclouds@gmail.com","subject":"[PATCH v4 0/5] Speed up unpack_trees()","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-12T08:15:46Z","receivedAt":"2018-08-12T08:16:02Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"v4 has a bunch of changes\n\n- 1/5 is a new one to show indented tracing. This way it's less\n  misleading to read nested time measurements\n- 3/5 now has the switch/restore cache_bottom logic. Junio suggested a\n  check instead in his final note, but I think this is safer (yeah I'm\n  scared too)\n- the old 4/4 is dropped because\n  - it assumes n-way logic\n  - the visible time saving is not worth the tradeoff\n  - Elijah gave me an idea to avoid add_index_entry() that I think\n    does not have n-way logic assumptions and gives better saving.\n    But it requires some more changes so I'm going to do it later\n- 5/5 is also new and should help reduce cache_tree_update() cost.\n  I wrote somewhere I was not going to work on this part, but it turns\n  out just a couple lines, might as well do it now.\n\nInterdiff\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 0dbe10fc85..105f13806f 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -426,7 +426,6 @@ static int update_one(struct cache_tree *it,\n \n int cache_tree_update(struct index_state *istate, int flags)\n {\n-\tuint64_t start = getnanotime();\n \tstruct cache_tree *it = istate->cache_tree;\n \tstruct cache_entry **cache = istate->cache;\n \tint entries = istate->cache_nr;\n@@ -434,11 +433,12 @@ int cache_tree_update(struct index_state *istate, int flags)\n \n \tif (i)\n \t\treturn i;\n+\ttrace_performance_enter();\n \ti = update_one(it, cache, entries, \"\", 0, &skip, flags);\n+\ttrace_performance_leave(\"cache_tree_update\");\n \tif (i < 0)\n \t\treturn i;\n \tistate->cache_changed |= CACHE_TREE_CHANGED;\n-\ttrace_performance_since(start, \"repair cache-tree\");\n \treturn 0;\n }\n \ndiff --git a/cache.h b/cache.h\nindex e6f7ee4b64..8b447652a7 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -673,7 +673,6 @@ extern int index_name_pos(const struct index_state *, const char *name, int name\n #define ADD_CACHE_JUST_APPEND 8\t\t/* Append only; tree.c::read_tree() */\n #define ADD_CACHE_NEW_ONLY 16\t\t/* Do not replace existing ones */\n #define ADD_CACHE_KEEP_CACHE_TREE 32\t/* Do not invalidate cache-tree */\n-#define ADD_CACHE_SKIP_VERIFY_PATH 64\t/* Do not verify path */\n extern int add_index_entry(struct index_state *, struct cache_entry *ce, int option);\n extern void rename_index_entry_at(struct index_state *, int pos, const char *new_name);\n \ndiff --git a/diff-lib.c b/diff-lib.c\nindex a9f38eb5a3..1ffa22c882 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -518,8 +518,8 @@ static int diff_cache(struct rev_info *revs,\n int run_diff_index(struct rev_info *revs, int cached)\n {\n \tstruct object_array_entry *ent;\n-\tuint64_t start = getnanotime();\n \n+\ttrace_performance_enter();\n \tent = revs->pending.objects;\n \tif (diff_cache(revs, &ent->item->oid, ent->name, cached))\n \t\texit(128);\n@@ -528,7 +528,7 @@ int run_diff_index(struct rev_info *revs, int cached)\n \tdiffcore_fix_diff_index(&revs->diffopt);\n \tdiffcore_std(&revs->diffopt);\n \tdiff_flush(&revs->diffopt);\n-\ttrace_performance_since(start, \"diff-index\");\n+\ttrace_performance_leave(\"diff-index\");\n \treturn 0;\n }\n \ndiff --git a/dir.c b/dir.c\nindex 21e6f2520a..c5e9fc8cea 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -2263,11 +2263,11 @@ int read_directory(struct dir_struct *dir, struct index_state *istate,\n \t\t   const char *path, int len, const struct pathspec *pathspec)\n {\n \tstruct untracked_cache_dir *untracked;\n-\tuint64_t start = getnanotime();\n \n \tif (has_symlink_leading_path(path, len))\n \t\treturn dir->nr;\n \n+\ttrace_performance_enter();\n \tuntracked = validate_untracked_cache(dir, len, pathspec);\n \tif (!untracked)\n \t\t/*\n@@ -2302,7 +2302,7 @@ int read_directory(struct dir_struct *dir, struct index_state *istate,\n \t\tdir->nr = i;\n \t}\n \n-\ttrace_performance_since(start, \"read directory %.*s\", len, path);\n+\ttrace_performance_leave(\"read directory %.*s\", len, path);\n \tif (dir->untracked) {\n \t\tstatic int force_untracked_cache = -1;\n \t\tstatic struct trace_key trace_untracked_stats = TRACE_KEY_INIT(UNTRACKED_STATS);\ndiff --git a/name-hash.c b/name-hash.c\nindex 163849831c..1fcda73cb3 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -578,10 +578,10 @@ static void threaded_lazy_init_name_hash(\n \n static void lazy_init_name_hash(struct index_state *istate)\n {\n-\tuint64_t start = getnanotime();\n \n \tif (istate->name_hash_initialized)\n \t\treturn;\n+\ttrace_performance_enter();\n \thashmap_init(&istate->name_hash, cache_entry_cmp, NULL, istate->cache_nr);\n \thashmap_init(&istate->dir_hash, dir_entry_cmp, NULL, istate->cache_nr);\n \n@@ -602,7 +602,7 @@ static void lazy_init_name_hash(struct index_state *istate)\n \t}\n \n \tistate->name_hash_initialized = 1;\n-\ttrace_performance_since(start, \"initialize name hash\");\n+\ttrace_performance_leave(\"initialize name hash\");\n }\n \n /*\ndiff --git a/preload-index.c b/preload-index.c\nindex 4d08d44874..d7f7919ba2 100644\n--- a/preload-index.c\n+++ b/preload-index.c\n@@ -78,7 +78,6 @@ static void preload_index(struct index_state *index,\n {\n \tint threads, i, work, offset;\n \tstruct thread_data data[MAX_PARALLEL];\n-\tuint64_t start = getnanotime();\n \n \tif (!core_preload_index)\n \t\treturn;\n@@ -88,6 +87,7 @@ static void preload_index(struct index_state *index,\n \t\tthreads = 2;\n \tif (threads < 2)\n \t\treturn;\n+\ttrace_performance_enter();\n \tif (threads > MAX_PARALLEL)\n \t\tthreads = MAX_PARALLEL;\n \toffset = 0;\n@@ -109,7 +109,7 @@ static void preload_index(struct index_state *index,\n \t\tif (pthread_join(p->pthread, NULL))\n \t\t\tdie(\"unable to join threaded lstat\");\n \t}\n-\ttrace_performance_since(start, \"preload index\");\n+\ttrace_performance_leave(\"preload index\");\n }\n #endif\n \ndiff --git a/read-cache.c b/read-cache.c\nindex b0b5df5de7..2b5646ef26 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1170,7 +1170,6 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n \tint ok_to_add = option & ADD_CACHE_OK_TO_ADD;\n \tint ok_to_replace = option & ADD_CACHE_OK_TO_REPLACE;\n \tint skip_df_check = option & ADD_CACHE_SKIP_DFCHECK;\n-\tint skip_verify_path = option & ADD_CACHE_SKIP_VERIFY_PATH;\n \tint new_only = option & ADD_CACHE_NEW_ONLY;\n \n \tif (!(option & ADD_CACHE_KEEP_CACHE_TREE))\n@@ -1211,7 +1210,7 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n \n \tif (!ok_to_add)\n \t\treturn -1;\n-\tif (!skip_verify_path && !verify_path(ce->name, ce->ce_mode))\n+\tif (!verify_path(ce->name, ce->ce_mode))\n \t\treturn error(\"Invalid path '%s'\", ce->name);\n \n \tif (!skip_df_check &&\n@@ -1400,8 +1399,8 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n \tconst char *typechange_fmt;\n \tconst char *added_fmt;\n \tconst char *unmerged_fmt;\n-\tuint64_t start = getnanotime();\n \n+\ttrace_performance_enter();\n \tmodified_fmt = (in_porcelain ? \"M\\t%s\\n\" : \"%s: needs update\\n\");\n \tdeleted_fmt = (in_porcelain ? \"D\\t%s\\n\" : \"%s: needs update\\n\");\n \ttypechange_fmt = (in_porcelain ? \"T\\t%s\\n\" : \"%s needs update\\n\");\n@@ -1471,7 +1470,7 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n \n \t\treplace_index_entry(istate, i, new_entry);\n \t}\n-\ttrace_performance_since(start, \"refresh index\");\n+\ttrace_performance_leave(\"refresh index\");\n \treturn has_errors;\n }\n \n@@ -1902,7 +1901,6 @@ static void freshen_shared_index(const char *shared_index, int warn)\n int read_index_from(struct index_state *istate, const char *path,\n \t\t    const char *gitdir)\n {\n-\tuint64_t start = getnanotime();\n \tstruct split_index *split_index;\n \tint ret;\n \tchar *base_oid_hex;\n@@ -1912,8 +1910,9 @@ int read_index_from(struct index_state *istate, const char *path,\n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n \n+\ttrace_performance_enter();\n \tret = do_read_index(istate, path, 0);\n-\ttrace_performance_since(start, \"read cache %s\", path);\n+\ttrace_performance_leave(\"read cache %s\", path);\n \n \tsplit_index = istate->split_index;\n \tif (!split_index || is_null_oid(&split_index->base_oid)) {\n@@ -1921,6 +1920,7 @@ int read_index_from(struct index_state *istate, const char *path,\n \t\treturn ret;\n \t}\n \n+\ttrace_performance_enter();\n \tif (split_index->base)\n \t\tdiscard_index(split_index->base);\n \telse\n@@ -1937,8 +1937,8 @@ int read_index_from(struct index_state *istate, const char *path,\n \tfreshen_shared_index(base_path, 0);\n \tmerge_base_index(istate);\n \tpost_read_index_from(istate);\n-\ttrace_performance_since(start, \"read cache %s\", base_path);\n \tfree(base_path);\n+\ttrace_performance_leave(\"read cache %s\", base_path);\n \treturn ret;\n }\n \n@@ -2763,4 +2763,6 @@ void move_index_extensions(struct index_state *dst, struct index_state *src)\n {\n \tdst->untracked = src->untracked;\n \tsrc->untracked = NULL;\n+\tdst->cache_tree = src->cache_tree;\n+\tsrc->cache_tree = NULL;\n }\ndiff --git a/trace.c b/trace.c\nindex fc623e91fd..fa4a2e7120 100644\n--- a/trace.c\n+++ b/trace.c\n@@ -176,10 +176,30 @@ void trace_strbuf_fl(const char *file, int line, struct trace_key *key,\n \tstrbuf_release(&buf);\n }\n \n+static uint64_t perf_start_times[10];\n+static int perf_indent;\n+\n+uint64_t trace_performance_enter(void)\n+{\n+\tuint64_t now;\n+\n+\tif (!trace_want(&trace_perf_key))\n+\t\treturn 0;\n+\n+\tnow = getnanotime();\n+\tperf_start_times[perf_indent] = now;\n+\tif (perf_indent + 1 < ARRAY_SIZE(perf_start_times))\n+\t\tperf_indent++;\n+\telse\n+\t\tBUG(\"Too deep indentation\");\n+\treturn now;\n+}\n+\n static void trace_performance_vprintf_fl(const char *file, int line,\n \t\t\t\t\t uint64_t nanos, const char *format,\n \t\t\t\t\t va_list ap)\n {\n+\tstatic const char space[] = \"          \";\n \tstruct strbuf buf = STRBUF_INIT;\n \n \tif (!prepare_trace_line(file, line, &trace_perf_key, &buf))\n@@ -188,7 +208,10 @@ static void trace_performance_vprintf_fl(const char *file, int line,\n \tstrbuf_addf(&buf, \"performance: %.9f s\", (double) nanos / 1000000000);\n \n \tif (format && *format) {\n-\t\tstrbuf_addstr(&buf, \": \");\n+\t\tif (perf_indent >= strlen(space))\n+\t\t\tBUG(\"Too deep indentation\");\n+\n+\t\tstrbuf_addf(&buf, \":%.*s \", perf_indent, space);\n \t\tstrbuf_vaddf(&buf, format, ap);\n \t}\n \n@@ -244,6 +267,24 @@ void trace_performance_since(uint64_t start, const char *format, ...)\n \tva_end(ap);\n }\n \n+void trace_performance_leave(const char *format, ...)\n+{\n+\tva_list ap;\n+\tuint64_t since;\n+\n+\tif (perf_indent)\n+\t\tperf_indent--;\n+\n+\tif (!format) /* Allow callers to leave without tracing anything */\n+\t\treturn;\n+\n+\tsince = perf_start_times[perf_indent];\n+\tva_start(ap, format);\n+\ttrace_performance_vprintf_fl(NULL, 0, getnanotime() - since,\n+\t\t\t\t     format, ap);\n+\tva_end(ap);\n+}\n+\n #else\n \n void trace_printf_key_fl(const char *file, int line, struct trace_key *key,\n@@ -273,6 +314,24 @@ void trace_performance_fl(const char *file, int line, uint64_t nanos,\n \tva_end(ap);\n }\n \n+void trace_performance_leave_fl(const char *file, int line,\n+\t\t\t\tuint64_t nanos, const char *format, ...)\n+{\n+\tva_list ap;\n+\tuint64_t since;\n+\n+\tif (perf_indent)\n+\t\tperf_indent--;\n+\n+\tif (!format) /* Allow callers to leave without tracing anything */\n+\t\treturn;\n+\n+\tsince = perf_start_times[perf_indent];\n+\tva_start(ap, format);\n+\ttrace_performance_vprintf_fl(file, line, nanos - since, format, ap);\n+\tva_end(ap);\n+}\n+\n #endif /* HAVE_VARIADIC_MACROS */\n \n \n@@ -411,13 +470,11 @@ uint64_t getnanotime(void)\n \t}\n }\n \n-static uint64_t command_start_time;\n static struct strbuf command_line = STRBUF_INIT;\n \n static void print_command_performance_atexit(void)\n {\n-\ttrace_performance_since(command_start_time, \"git command:%s\",\n-\t\t\t\tcommand_line.buf);\n+\ttrace_performance_leave(\"git command:%s\", command_line.buf);\n }\n \n void trace_command_performance(const char **argv)\n@@ -425,10 +482,10 @@ void trace_command_performance(const char **argv)\n \tif (!trace_want(&trace_perf_key))\n \t\treturn;\n \n-\tif (!command_start_time)\n+\tif (!command_line.len)\n \t\tatexit(print_command_performance_atexit);\n \n \tstrbuf_reset(&command_line);\n \tsq_quote_argv_pretty(&command_line, argv);\n-\tcommand_start_time = getnanotime();\n+\ttrace_performance_enter();\n }\ndiff --git a/trace.h b/trace.h\nindex 2b6a1bc17c..171b256d26 100644\n--- a/trace.h\n+++ b/trace.h\n@@ -23,6 +23,7 @@ extern void trace_disable(struct trace_key *key);\n extern uint64_t getnanotime(void);\n extern void trace_command_performance(const char **argv);\n extern void trace_verbatim(struct trace_key *key, const void *buf, unsigned len);\n+uint64_t trace_performance_enter(void);\n \n #ifndef HAVE_VARIADIC_MACROS\n \n@@ -45,6 +46,9 @@ extern void trace_performance(uint64_t nanos, const char *format, ...);\n __attribute__((format (printf, 2, 3)))\n extern void trace_performance_since(uint64_t start, const char *format, ...);\n \n+__attribute__((format (printf, 1, 2)))\n+void trace_performance_leave(const char *format, ...);\n+\n #else\n \n /*\n@@ -118,6 +122,14 @@ extern void trace_performance_since(uint64_t start, const char *format, ...);\n \t\t\t\t\t     __VA_ARGS__);\t\t    \\\n \t} while (0)\n \n+#define trace_performance_leave(...)\t\t\t\t\t    \\\n+\tdo {\t\t\t\t\t\t\t\t    \\\n+\t\tif (trace_pass_fl(&trace_perf_key))\t\t\t    \\\n+\t\t\ttrace_performance_leave_fl(TRACE_CONTEXT, __LINE__, \\\n+\t\t\t\t\t\t   getnanotime(),\t    \\\n+\t\t\t\t\t\t   __VA_ARGS__);\t    \\\n+\t} while (0)\n+\n /* backend functions, use non-*fl macros instead */\n __attribute__((format (printf, 4, 5)))\n extern void trace_printf_key_fl(const char *file, int line, struct trace_key *key,\n@@ -130,6 +142,9 @@ extern void trace_strbuf_fl(const char *file, int line, struct trace_key *key,\n __attribute__((format (printf, 4, 5)))\n extern void trace_performance_fl(const char *file, int line,\n \t\t\t\t uint64_t nanos, const char *fmt, ...);\n+__attribute__((format (printf, 4, 5)))\n+extern void trace_performance_leave_fl(const char *file, int line,\n+\t\t\t\t       uint64_t nanos, const char *fmt, ...);\n static inline int trace_pass_fl(struct trace_key *key)\n {\n \treturn key->fd || !key->initialized;\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 1438ee1555..d822662c75 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -201,7 +201,6 @@ static int do_add_entry(struct unpack_trees_options *o, struct cache_entry *ce,\n \n \tce->ce_flags = (ce->ce_flags & ~clear) | set;\n \treturn add_index_entry(&o->result, ce,\n-\t\t\t       o->extra_add_index_flags |\n \t\t\t       ADD_CACHE_OK_TO_ADD | ADD_CACHE_OK_TO_REPLACE);\n }\n \n@@ -353,9 +352,9 @@ static int check_updates(struct unpack_trees_options *o)\n \tstruct progress *progress = NULL;\n \tstruct index_state *index = &o->result;\n \tstruct checkout state = CHECKOUT_INIT;\n-\tuint64_t start = getnanotime();\n \tint i;\n \n+\ttrace_performance_enter();\n \tstate.force = 1;\n \tstate.quiet = 1;\n \tstate.refresh_cache = 1;\n@@ -425,7 +424,7 @@ static int check_updates(struct unpack_trees_options *o)\n \terrs |= finish_delayed_checkout(&state);\n \tif (o->update)\n \t\tgit_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n-\ttrace_performance_since(start, \"update worktree after a merge\");\n+\ttrace_performance_leave(\"check_updates\");\n \treturn errs != 0;\n }\n \n@@ -702,31 +701,13 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \tif (!o->merge)\n \t\tBUG(\"We need cache-tree to do this optimization\");\n \n-\t/*\n-\t * Try to keep add_index_entry() as fast as possible since\n-\t * we're going to do a lot of them.\n-\t *\n-\t * Skipping verify_path() should totally be safe because these\n-\t * paths are from the source index, which must have been\n-\t * verified.\n-\t *\n-\t * Skipping D/F and cache-tree validation checks is trickier\n-\t * because it assumes what n-merge code would do when all\n-\t * trees and the index are the same. We probably could just\n-\t * optimize those code instead (e.g. we don't invalidate that\n-\t * many cache-tree, but the searching for them is very\n-\t * expensive).\n-\t */\n-\to->extra_add_index_flags = ADD_CACHE_SKIP_DFCHECK;\n-\to->extra_add_index_flags |= ADD_CACHE_SKIP_VERIFY_PATH;\n-\n \t/*\n \t * Do what unpack_callback() and unpack_nondirectories() normally\n \t * do. But we walk all paths recursively in just one loop instead.\n \t *\n-\t * D/F conflicts and staged entries are not a concern because\n-\t * cache-tree would be invalidated and we would never get here\n-\t * in the first place.\n+\t * D/F conflicts and higher stage entries are not a concern\n+\t * because cache-tree would be invalidated and we would never\n+\t * get here in the first place.\n \t */\n \tfor (i = 0; i < nr_entries; i++) {\n \t\tint new_ce_len, len, rc;\n@@ -761,7 +742,6 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \n \t\tmark_ce_used(src[0], o);\n \t}\n-\to->extra_add_index_flags = 0;\n \tfree(tree_ce);\n \tif (o->debug_unpack)\n \t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n@@ -791,7 +771,17 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n \n \t\tif (!o->merge || df_conflicts)\n \t\t\tBUG(\"Wrong condition to get here buddy\");\n-\t\treturn traverse_by_cache_tree(pos, nr_entries, n, names, info);\n+\n+\t\t/*\n+\t\t * All entries up to 'pos' must have been processed\n+\t\t * (i.e. marked CE_UNPACKED) at this point. But to be safe,\n+\t\t * save and restore cache_bottom anyway to not miss\n+\t\t * unprocessed entries before 'pos'.\n+\t\t */\n+\t\tbottom = o->cache_bottom;\n+\t\tret = traverse_by_cache_tree(pos, nr_entries, n, names, info);\n+\t\to->cache_bottom = bottom;\n+\t\treturn ret;\n \t}\n \n \tp = names;\n@@ -1142,7 +1132,7 @@ static void debug_unpack_callback(int n,\n }\n \n /*\n- * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n+ * Note that traverse_by_cache_tree() duplicates some logic in this function\n  * without actually calling it. If you change the logic here you may need to\n  * check and change there as well.\n  */\n@@ -1425,11 +1415,11 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \tint i, ret;\n \tstatic struct cache_entry *dfc;\n \tstruct exclude_list el;\n-\tuint64_t start = getnanotime();\n \n \tif (len > MAX_UNPACK_TREES)\n \t\tdie(\"unpack_trees takes at most %d trees\", MAX_UNPACK_TREES);\n \n+\ttrace_performance_enter();\n \tmemset(&el, 0, sizeof(el));\n \tif (!core_apply_sparse_checkout || !o->update)\n \t\to->skip_sparse_checkout = 1;\n@@ -1502,7 +1492,10 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t\t}\n \t\t}\n \n-\t\tif (traverse_trees(len, t, &info) < 0)\n+\t\ttrace_performance_enter();\n+\t\tret = traverse_trees(len, t, &info);\n+\t\ttrace_performance_leave(\"traverse_trees\");\n+\t\tif (ret < 0)\n \t\t\tgoto return_failed;\n \t}\n \n@@ -1574,10 +1567,10 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t\tgoto done;\n \t\t}\n \t}\n-\ttrace_performance_since(start, \"unpack trees\");\n \n \tret = check_updates(o) ? (-2) : 0;\n \tif (o->dst_index) {\n+\t\tmove_index_extensions(&o->result, o->src_index);\n \t\tif (!ret) {\n \t\t\tif (!o->result.cache_tree)\n \t\t\t\to->result.cache_tree = cache_tree();\n@@ -1586,7 +1579,6 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t\t\t\t\t  WRITE_TREE_SILENT |\n \t\t\t\t\t\t  WRITE_TREE_REPAIR);\n \t\t}\n-\t\tmove_index_extensions(&o->result, o->src_index);\n \t\tdiscard_index(o->dst_index);\n \t\t*o->dst_index = o->result;\n \t} else {\n@@ -1595,6 +1587,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \to->src_index = NULL;\n \n done:\n+\ttrace_performance_leave(\"unpack_trees\");\n \tclear_exclude_list(&el);\n \treturn ret;\n \ndiff --git a/unpack-trees.h b/unpack-trees.h\nindex 94e1b14078..c2b434c606 100644\n--- a/unpack-trees.h\n+++ b/unpack-trees.h\n@@ -80,7 +80,6 @@ struct unpack_trees_options {\n \tstruct index_state result;\n \n \tstruct exclude_list *el; /* for internal use */\n-\tunsigned int extra_add_index_flags;\n };\n \n extern int unpack_trees(unsigned n, struct tree_desc *t,\n\nNguyễn Thái Ngọc Duy (5):\n  trace.h: support nested performance tracing\n  unpack-trees: add performance tracing\n  unpack-trees: optimize walking same trees with cache-tree\n  unpack-trees: reduce malloc in cache-tree walk\n  unpack-trees: reuse (still valid) cache-tree from src_index\n\n cache-tree.c    |   2 +\n diff-lib.c      |   4 +-\n dir.c           |   4 +-\n name-hash.c     |   4 +-\n preload-index.c |   4 +-\n read-cache.c    |  13 +++--\n trace.c         |  69 ++++++++++++++++++++--\n trace.h         |  15 +++++\n unpack-trees.c  | 149 +++++++++++++++++++++++++++++++++++++++++++++++-\n 9 files changed, 243 insertions(+), 21 deletions(-)\n\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355293","messageId":"20180812081551.27927-4-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-1-pclouds@gmail.com","subject":"[PATCH v4 3/5] unpack-trees: optimize walking same trees with cache-tree","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-12T08:15:49Z","receivedAt":"2018-08-12T08:16:02Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"In order to merge one or many trees with the index, unpack-trees code\nwalks multiple trees in parallel with the index and performs n-way\nmerge. If we find out at start of a directory that all trees are the\nsame (by comparing OID) and cache-tree happens to be available for\nthat directory as well, we could avoid walking the trees because we\nalready know what these trees contain: it's flattened in what's called\n\"the index\".\n\nThe upside is of course a lot less I/O since we can potentially skip\nlots of trees (think subtrees). We also save CPU because we don't have\nto inflate and apply the deltas. The downside is of course more\nfragile code since the logic in some functions are now duplicated\nelsewhere.\n\n\"checkout -\" with this patch on webkit.git (275k files):\n\n    baseline      new\n  --------------------------------------------------------------------\n    0.056651714   0.080394752 s:  read cache .git/index\n    0.183101080   0.216010838 s:  preload index\n    0.008584433   0.008534301 s:  refresh index\n    0.633767589   0.251992198 s:   traverse_trees\n    0.340265448   0.377031383 s:   check_updates\n    0.381884638   0.372768105 s:   cache_tree_update\n    1.401562947   1.045887251 s:  unpack_trees\n    0.338687914   0.314983512 s:  write index, changed mask = 2e\n    0.411927922   0.062572653 s:    traverse_trees\n    0.000023335   0.000022544 s:    check_updates\n    0.423697246   0.073795585 s:   unpack_trees\n    0.423708360   0.073807557 s:  diff-index\n    2.559524127   1.938191592 s: git command: git checkout -\n\nAnother measurement from Ben's running \"git checkout\" with over 500k\ntrees (on the whole series):\n\n    baseline        new\n  ----------------------------------------------------------------------\n    0.535510167     0.556558733     s: read cache .git/index\n    0.3057373       0.3147105       s: initialize name hash\n    0.0184082       0.023558433     s: preload index\n    0.086910967     0.089085967     s: refresh index\n    7.889590767     2.191554433     s: unpack trees\n    0.120760833     0.131941267     s: update worktree after a merge\n    2.2583504       2.572663167     s: repair cache-tree\n    0.8916137       0.959495233     s: write index, changed mask = 28\n    3.405199233     0.2710663       s: unpack trees\n    0.000999667     0.0021554       s: update worktree after a merge\n    3.4063306       0.273318333     s: diff-index\n    16.9524923      9.462943133     s: git command: git.exe checkout\n\nThis command calls unpack_trees() twice, the first time on 2way merge\nand the second 1way merge. In both times, \"unpack trees\" time is\nreduced to one third. Overall time reduction is not that impressive of\ncourse because index operations take a big chunk. And there's that\nrepair cache-tree line.\n\nPS. A note about cache-tree invalidation and the use of it in this\ncode.\n\nWe do invalidate cache-tree in _source_ index when we add new entries\nto the (temporary) \"result\" index. But we also use the cache-tree from\nsource index in this optimization. Does this mean we end up having no\ncache-tree in the source index to activate this optimization?\n\nThe answer is twisted: the order of finding a good cache-tree and\ninvalidating it matters. In this case we check for a good cache-tree\nfirst in all_trees_same_as_cache_tree(), then we start to merge things\nand potentially invalidate that same cache-tree in the process. Since\ncache-tree invalidation happens after the optimization kicks in, we're\nstill good. But we may lose that cache-tree at the very first\ncall_unpack_fn() call in traverse_by_cache_tree().\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n unpack-trees.c | 127 +++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 127 insertions(+)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex b237eaa0f2..07456d0fb2 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -644,6 +644,102 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n \treturn name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n }\n \n+static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n+\t\t\t\t\tstruct name_entry *names,\n+\t\t\t\t\tstruct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i;\n+\n+\tif (!o->merge || dirmask != ((1 << n) - 1))\n+\t\treturn 0;\n+\n+\tfor (i = 1; i < n; i++)\n+\t\tif (!are_same_oid(names, names + i))\n+\t\t\treturn 0;\n+\n+\treturn cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n+}\n+\n+static int index_pos_by_traverse_info(struct name_entry *names,\n+\t\t\t\t      struct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint len = traverse_path_len(info, names);\n+\tchar *name = xmalloc(len + 1 /* slash */ + 1 /* NUL */);\n+\tint pos;\n+\n+\tmake_traverse_path(name, info, names);\n+\tname[len++] = '/';\n+\tname[len] = '\\0';\n+\tpos = index_name_pos(o->src_index, name, len);\n+\tif (pos >= 0)\n+\t\tBUG(\"This is a directory and should not exist in index\");\n+\tpos = -pos - 1;\n+\tif (!starts_with(o->src_index->cache[pos]->name, name) ||\n+\t    (pos > 0 && starts_with(o->src_index->cache[pos-1]->name, name)))\n+\t\tBUG(\"pos must point at the first entry in this directory\");\n+\tfree(name);\n+\treturn pos;\n+}\n+\n+/*\n+ * Fast path if we detect that all trees are the same as cache-tree at this\n+ * path. We'll walk these trees recursively using cache-tree/index instead of\n+ * ODB since already know what these trees contain.\n+ */\n+static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n+\t\t\t\t  struct name_entry *names,\n+\t\t\t\t  struct traverse_info *info)\n+{\n+\tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i, d;\n+\n+\tif (!o->merge)\n+\t\tBUG(\"We need cache-tree to do this optimization\");\n+\n+\t/*\n+\t * Do what unpack_callback() and unpack_nondirectories() normally\n+\t * do. But we walk all paths recursively in just one loop instead.\n+\t *\n+\t * D/F conflicts and higher stage entries are not a concern\n+\t * because cache-tree would be invalidated and we would never\n+\t * get here in the first place.\n+\t */\n+\tfor (i = 0; i < nr_entries; i++) {\n+\t\tstruct cache_entry *tree_ce;\n+\t\tint len, rc;\n+\n+\t\tsrc[0] = o->src_index->cache[pos + i];\n+\n+\t\tlen = ce_namelen(src[0]);\n+\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n+\n+\t\ttree_ce->ce_mode = src[0]->ce_mode;\n+\t\ttree_ce->ce_flags = create_ce_flags(0);\n+\t\ttree_ce->ce_namelen = len;\n+\t\toidcpy(&tree_ce->oid, &src[0]->oid);\n+\t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n+\n+\t\tfor (d = 1; d <= nr_names; d++)\n+\t\t\tsrc[d] = tree_ce;\n+\n+\t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n+\t\tfree(tree_ce);\n+\t\tif (rc < 0)\n+\t\t\treturn rc;\n+\n+\t\tmark_ce_used(src[0], o);\n+\t}\n+\tif (o->debug_unpack)\n+\t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n+\t\t       nr_entries,\n+\t\t       o->src_index->cache[pos]->name,\n+\t\t       o->src_index->cache[pos + nr_entries - 1]->name);\n+\treturn 0;\n+}\n+\n static int traverse_trees_recursive(int n, unsigned long dirmask,\n \t\t\t\t    unsigned long df_conflicts,\n \t\t\t\t    struct name_entry *names,\n@@ -655,6 +751,27 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n \tvoid *buf[MAX_UNPACK_TREES];\n \tstruct traverse_info newinfo;\n \tstruct name_entry *p;\n+\tint nr_entries;\n+\n+\tnr_entries = all_trees_same_as_cache_tree(n, dirmask, names, info);\n+\tif (nr_entries > 0) {\n+\t\tstruct unpack_trees_options *o = info->data;\n+\t\tint pos = index_pos_by_traverse_info(names, info);\n+\n+\t\tif (!o->merge || df_conflicts)\n+\t\t\tBUG(\"Wrong condition to get here buddy\");\n+\n+\t\t/*\n+\t\t * All entries up to 'pos' must have been processed\n+\t\t * (i.e. marked CE_UNPACKED) at this point. But to be safe,\n+\t\t * save and restore cache_bottom anyway to not miss\n+\t\t * unprocessed entries before 'pos'.\n+\t\t */\n+\t\tbottom = o->cache_bottom;\n+\t\tret = traverse_by_cache_tree(pos, nr_entries, n, names, info);\n+\t\to->cache_bottom = bottom;\n+\t\treturn ret;\n+\t}\n \n \tp = names;\n \twhile (!p->mode)\n@@ -814,6 +931,11 @@ static struct cache_entry *create_ce_entry(const struct traverse_info *info, con\n \treturn ce;\n }\n \n+/*\n+ * Note that traverse_by_cache_tree() duplicates some logic in this function\n+ * without actually calling it. If you change the logic here you may need to\n+ * check and change there as well.\n+ */\n static int unpack_nondirectories(int n, unsigned long mask,\n \t\t\t\t unsigned long dirmask,\n \t\t\t\t struct cache_entry **src,\n@@ -998,6 +1120,11 @@ static void debug_unpack_callback(int n,\n \t\tdebug_name_entry(i, names + i);\n }\n \n+/*\n+ * Note that traverse_by_cache_tree() duplicates some logic in this function\n+ * without actually calling it. If you change the logic here you may need to\n+ * check and change there as well.\n+ */\n static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355294","messageId":"20180812081551.27927-2-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-1-pclouds@gmail.com","subject":"[PATCH v4 1/5] trace.h: support nested performance tracing","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-12T08:15:47Z","receivedAt":"2018-08-12T08:16:02Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Performance measurements are listed right now as a flat list, which is\nfine when we measure big blocks. But when we start adding more and\nmore measurements, some of them could be just part of a bigger\nmeasurement and a flat list gives a wrong impression that they are\nexecuted at the same level instead of nested.\n\nAdd trace_performance_enter() and trace_performance_leave() to allow\nindent these nested measurements. For now it does not help much\nbecause the only nested thing is (lazy) name hash initialization\n(e.g. called in diff-index from \"git status\"). This will help more\nbecause I'm going to add some more tracing that's actually nested.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n diff-lib.c      |  4 +--\n dir.c           |  4 +--\n name-hash.c     |  4 +--\n preload-index.c |  4 +--\n read-cache.c    | 11 ++++----\n trace.c         | 69 ++++++++++++++++++++++++++++++++++++++++++++-----\n trace.h         | 15 +++++++++++\n 7 files changed, 92 insertions(+), 19 deletions(-)\n\ndiff --git a/diff-lib.c b/diff-lib.c\nindex a9f38eb5a3..1ffa22c882 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -518,8 +518,8 @@ static int diff_cache(struct rev_info *revs,\n int run_diff_index(struct rev_info *revs, int cached)\n {\n \tstruct object_array_entry *ent;\n-\tuint64_t start = getnanotime();\n \n+\ttrace_performance_enter();\n \tent = revs->pending.objects;\n \tif (diff_cache(revs, &ent->item->oid, ent->name, cached))\n \t\texit(128);\n@@ -528,7 +528,7 @@ int run_diff_index(struct rev_info *revs, int cached)\n \tdiffcore_fix_diff_index(&revs->diffopt);\n \tdiffcore_std(&revs->diffopt);\n \tdiff_flush(&revs->diffopt);\n-\ttrace_performance_since(start, \"diff-index\");\n+\ttrace_performance_leave(\"diff-index\");\n \treturn 0;\n }\n \ndiff --git a/dir.c b/dir.c\nindex 21e6f2520a..c5e9fc8cea 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -2263,11 +2263,11 @@ int read_directory(struct dir_struct *dir, struct index_state *istate,\n \t\t   const char *path, int len, const struct pathspec *pathspec)\n {\n \tstruct untracked_cache_dir *untracked;\n-\tuint64_t start = getnanotime();\n \n \tif (has_symlink_leading_path(path, len))\n \t\treturn dir->nr;\n \n+\ttrace_performance_enter();\n \tuntracked = validate_untracked_cache(dir, len, pathspec);\n \tif (!untracked)\n \t\t/*\n@@ -2302,7 +2302,7 @@ int read_directory(struct dir_struct *dir, struct index_state *istate,\n \t\tdir->nr = i;\n \t}\n \n-\ttrace_performance_since(start, \"read directory %.*s\", len, path);\n+\ttrace_performance_leave(\"read directory %.*s\", len, path);\n \tif (dir->untracked) {\n \t\tstatic int force_untracked_cache = -1;\n \t\tstatic struct trace_key trace_untracked_stats = TRACE_KEY_INIT(UNTRACKED_STATS);\ndiff --git a/name-hash.c b/name-hash.c\nindex 163849831c..1fcda73cb3 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -578,10 +578,10 @@ static void threaded_lazy_init_name_hash(\n \n static void lazy_init_name_hash(struct index_state *istate)\n {\n-\tuint64_t start = getnanotime();\n \n \tif (istate->name_hash_initialized)\n \t\treturn;\n+\ttrace_performance_enter();\n \thashmap_init(&istate->name_hash, cache_entry_cmp, NULL, istate->cache_nr);\n \thashmap_init(&istate->dir_hash, dir_entry_cmp, NULL, istate->cache_nr);\n \n@@ -602,7 +602,7 @@ static void lazy_init_name_hash(struct index_state *istate)\n \t}\n \n \tistate->name_hash_initialized = 1;\n-\ttrace_performance_since(start, \"initialize name hash\");\n+\ttrace_performance_leave(\"initialize name hash\");\n }\n \n /*\ndiff --git a/preload-index.c b/preload-index.c\nindex 4d08d44874..d7f7919ba2 100644\n--- a/preload-index.c\n+++ b/preload-index.c\n@@ -78,7 +78,6 @@ static void preload_index(struct index_state *index,\n {\n \tint threads, i, work, offset;\n \tstruct thread_data data[MAX_PARALLEL];\n-\tuint64_t start = getnanotime();\n \n \tif (!core_preload_index)\n \t\treturn;\n@@ -88,6 +87,7 @@ static void preload_index(struct index_state *index,\n \t\tthreads = 2;\n \tif (threads < 2)\n \t\treturn;\n+\ttrace_performance_enter();\n \tif (threads > MAX_PARALLEL)\n \t\tthreads = MAX_PARALLEL;\n \toffset = 0;\n@@ -109,7 +109,7 @@ static void preload_index(struct index_state *index,\n \t\tif (pthread_join(p->pthread, NULL))\n \t\t\tdie(\"unable to join threaded lstat\");\n \t}\n-\ttrace_performance_since(start, \"preload index\");\n+\ttrace_performance_leave(\"preload index\");\n }\n #endif\n \ndiff --git a/read-cache.c b/read-cache.c\nindex e865254bea..4fd35f4f37 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1399,8 +1399,8 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n \tconst char *typechange_fmt;\n \tconst char *added_fmt;\n \tconst char *unmerged_fmt;\n-\tuint64_t start = getnanotime();\n \n+\ttrace_performance_enter();\n \tmodified_fmt = (in_porcelain ? \"M\\t%s\\n\" : \"%s: needs update\\n\");\n \tdeleted_fmt = (in_porcelain ? \"D\\t%s\\n\" : \"%s: needs update\\n\");\n \ttypechange_fmt = (in_porcelain ? \"T\\t%s\\n\" : \"%s needs update\\n\");\n@@ -1470,7 +1470,7 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n \n \t\treplace_index_entry(istate, i, new_entry);\n \t}\n-\ttrace_performance_since(start, \"refresh index\");\n+\ttrace_performance_leave(\"refresh index\");\n \treturn has_errors;\n }\n \n@@ -1901,7 +1901,6 @@ static void freshen_shared_index(const char *shared_index, int warn)\n int read_index_from(struct index_state *istate, const char *path,\n \t\t    const char *gitdir)\n {\n-\tuint64_t start = getnanotime();\n \tstruct split_index *split_index;\n \tint ret;\n \tchar *base_oid_hex;\n@@ -1911,8 +1910,9 @@ int read_index_from(struct index_state *istate, const char *path,\n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n \n+\ttrace_performance_enter();\n \tret = do_read_index(istate, path, 0);\n-\ttrace_performance_since(start, \"read cache %s\", path);\n+\ttrace_performance_leave(\"read cache %s\", path);\n \n \tsplit_index = istate->split_index;\n \tif (!split_index || is_null_oid(&split_index->base_oid)) {\n@@ -1920,6 +1920,7 @@ int read_index_from(struct index_state *istate, const char *path,\n \t\treturn ret;\n \t}\n \n+\ttrace_performance_enter();\n \tif (split_index->base)\n \t\tdiscard_index(split_index->base);\n \telse\n@@ -1936,8 +1937,8 @@ int read_index_from(struct index_state *istate, const char *path,\n \tfreshen_shared_index(base_path, 0);\n \tmerge_base_index(istate);\n \tpost_read_index_from(istate);\n-\ttrace_performance_since(start, \"read cache %s\", base_path);\n \tfree(base_path);\n+\ttrace_performance_leave(\"read cache %s\", base_path);\n \treturn ret;\n }\n \ndiff --git a/trace.c b/trace.c\nindex fc623e91fd..fa4a2e7120 100644\n--- a/trace.c\n+++ b/trace.c\n@@ -176,10 +176,30 @@ void trace_strbuf_fl(const char *file, int line, struct trace_key *key,\n \tstrbuf_release(&buf);\n }\n \n+static uint64_t perf_start_times[10];\n+static int perf_indent;\n+\n+uint64_t trace_performance_enter(void)\n+{\n+\tuint64_t now;\n+\n+\tif (!trace_want(&trace_perf_key))\n+\t\treturn 0;\n+\n+\tnow = getnanotime();\n+\tperf_start_times[perf_indent] = now;\n+\tif (perf_indent + 1 < ARRAY_SIZE(perf_start_times))\n+\t\tperf_indent++;\n+\telse\n+\t\tBUG(\"Too deep indentation\");\n+\treturn now;\n+}\n+\n static void trace_performance_vprintf_fl(const char *file, int line,\n \t\t\t\t\t uint64_t nanos, const char *format,\n \t\t\t\t\t va_list ap)\n {\n+\tstatic const char space[] = \"          \";\n \tstruct strbuf buf = STRBUF_INIT;\n \n \tif (!prepare_trace_line(file, line, &trace_perf_key, &buf))\n@@ -188,7 +208,10 @@ static void trace_performance_vprintf_fl(const char *file, int line,\n \tstrbuf_addf(&buf, \"performance: %.9f s\", (double) nanos / 1000000000);\n \n \tif (format && *format) {\n-\t\tstrbuf_addstr(&buf, \": \");\n+\t\tif (perf_indent >= strlen(space))\n+\t\t\tBUG(\"Too deep indentation\");\n+\n+\t\tstrbuf_addf(&buf, \":%.*s \", perf_indent, space);\n \t\tstrbuf_vaddf(&buf, format, ap);\n \t}\n \n@@ -244,6 +267,24 @@ void trace_performance_since(uint64_t start, const char *format, ...)\n \tva_end(ap);\n }\n \n+void trace_performance_leave(const char *format, ...)\n+{\n+\tva_list ap;\n+\tuint64_t since;\n+\n+\tif (perf_indent)\n+\t\tperf_indent--;\n+\n+\tif (!format) /* Allow callers to leave without tracing anything */\n+\t\treturn;\n+\n+\tsince = perf_start_times[perf_indent];\n+\tva_start(ap, format);\n+\ttrace_performance_vprintf_fl(NULL, 0, getnanotime() - since,\n+\t\t\t\t     format, ap);\n+\tva_end(ap);\n+}\n+\n #else\n \n void trace_printf_key_fl(const char *file, int line, struct trace_key *key,\n@@ -273,6 +314,24 @@ void trace_performance_fl(const char *file, int line, uint64_t nanos,\n \tva_end(ap);\n }\n \n+void trace_performance_leave_fl(const char *file, int line,\n+\t\t\t\tuint64_t nanos, const char *format, ...)\n+{\n+\tva_list ap;\n+\tuint64_t since;\n+\n+\tif (perf_indent)\n+\t\tperf_indent--;\n+\n+\tif (!format) /* Allow callers to leave without tracing anything */\n+\t\treturn;\n+\n+\tsince = perf_start_times[perf_indent];\n+\tva_start(ap, format);\n+\ttrace_performance_vprintf_fl(file, line, nanos - since, format, ap);\n+\tva_end(ap);\n+}\n+\n #endif /* HAVE_VARIADIC_MACROS */\n \n \n@@ -411,13 +470,11 @@ uint64_t getnanotime(void)\n \t}\n }\n \n-static uint64_t command_start_time;\n static struct strbuf command_line = STRBUF_INIT;\n \n static void print_command_performance_atexit(void)\n {\n-\ttrace_performance_since(command_start_time, \"git command:%s\",\n-\t\t\t\tcommand_line.buf);\n+\ttrace_performance_leave(\"git command:%s\", command_line.buf);\n }\n \n void trace_command_performance(const char **argv)\n@@ -425,10 +482,10 @@ void trace_command_performance(const char **argv)\n \tif (!trace_want(&trace_perf_key))\n \t\treturn;\n \n-\tif (!command_start_time)\n+\tif (!command_line.len)\n \t\tatexit(print_command_performance_atexit);\n \n \tstrbuf_reset(&command_line);\n \tsq_quote_argv_pretty(&command_line, argv);\n-\tcommand_start_time = getnanotime();\n+\ttrace_performance_enter();\n }\ndiff --git a/trace.h b/trace.h\nindex 2b6a1bc17c..171b256d26 100644\n--- a/trace.h\n+++ b/trace.h\n@@ -23,6 +23,7 @@ extern void trace_disable(struct trace_key *key);\n extern uint64_t getnanotime(void);\n extern void trace_command_performance(const char **argv);\n extern void trace_verbatim(struct trace_key *key, const void *buf, unsigned len);\n+uint64_t trace_performance_enter(void);\n \n #ifndef HAVE_VARIADIC_MACROS\n \n@@ -45,6 +46,9 @@ extern void trace_performance(uint64_t nanos, const char *format, ...);\n __attribute__((format (printf, 2, 3)))\n extern void trace_performance_since(uint64_t start, const char *format, ...);\n \n+__attribute__((format (printf, 1, 2)))\n+void trace_performance_leave(const char *format, ...);\n+\n #else\n \n /*\n@@ -118,6 +122,14 @@ extern void trace_performance_since(uint64_t start, const char *format, ...);\n \t\t\t\t\t     __VA_ARGS__);\t\t    \\\n \t} while (0)\n \n+#define trace_performance_leave(...)\t\t\t\t\t    \\\n+\tdo {\t\t\t\t\t\t\t\t    \\\n+\t\tif (trace_pass_fl(&trace_perf_key))\t\t\t    \\\n+\t\t\ttrace_performance_leave_fl(TRACE_CONTEXT, __LINE__, \\\n+\t\t\t\t\t\t   getnanotime(),\t    \\\n+\t\t\t\t\t\t   __VA_ARGS__);\t    \\\n+\t} while (0)\n+\n /* backend functions, use non-*fl macros instead */\n __attribute__((format (printf, 4, 5)))\n extern void trace_printf_key_fl(const char *file, int line, struct trace_key *key,\n@@ -130,6 +142,9 @@ extern void trace_strbuf_fl(const char *file, int line, struct trace_key *key,\n __attribute__((format (printf, 4, 5)))\n extern void trace_performance_fl(const char *file, int line,\n \t\t\t\t uint64_t nanos, const char *fmt, ...);\n+__attribute__((format (printf, 4, 5)))\n+extern void trace_performance_leave_fl(const char *file, int line,\n+\t\t\t\t       uint64_t nanos, const char *fmt, ...);\n static inline int trace_pass_fl(struct trace_key *key)\n {\n \treturn key->fd || !key->initialized;\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355295","messageId":"20180812081551.27927-3-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-1-pclouds@gmail.com","subject":"[PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-12T08:15:48Z","receivedAt":"2018-08-12T08:16:02Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"We're going to optimize unpack_trees() a bit in the following\npatches. Let's add some tracing to measure how long it takes before\nand after. This is the baseline (\"git checkout -\" on webkit.git, 275k\nfiles on worktree)\n\n    performance: 0.056651714 s:  read cache .git/index\n    performance: 0.183101080 s:  preload index\n    performance: 0.008584433 s:  refresh index\n    performance: 0.633767589 s:   traverse_trees\n    performance: 0.340265448 s:   check_updates\n    performance: 0.381884638 s:   cache_tree_update\n    performance: 1.401562947 s:  unpack_trees\n    performance: 0.338687914 s:  write index, changed mask = 2e\n    performance: 0.411927922 s:    traverse_trees\n    performance: 0.000023335 s:    check_updates\n    performance: 0.423697246 s:   unpack_trees\n    performance: 0.423708360 s:  diff-index\n    performance: 2.559524127 s: git command: git checkout -\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache-tree.c   | 2 ++\n unpack-trees.c | 9 ++++++++-\n 2 files changed, 10 insertions(+), 1 deletion(-)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 6b46711996..105f13806f 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -433,7 +433,9 @@ int cache_tree_update(struct index_state *istate, int flags)\n \n \tif (i)\n \t\treturn i;\n+\ttrace_performance_enter();\n \ti = update_one(it, cache, entries, \"\", 0, &skip, flags);\n+\ttrace_performance_leave(\"cache_tree_update\");\n \tif (i < 0)\n \t\treturn i;\n \tistate->cache_changed |= CACHE_TREE_CHANGED;\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex cd0680f11e..b237eaa0f2 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -354,6 +354,7 @@ static int check_updates(struct unpack_trees_options *o)\n \tstruct checkout state = CHECKOUT_INIT;\n \tint i;\n \n+\ttrace_performance_enter();\n \tstate.force = 1;\n \tstate.quiet = 1;\n \tstate.refresh_cache = 1;\n@@ -423,6 +424,7 @@ static int check_updates(struct unpack_trees_options *o)\n \terrs |= finish_delayed_checkout(&state);\n \tif (o->update)\n \t\tgit_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n+\ttrace_performance_leave(\"check_updates\");\n \treturn errs != 0;\n }\n \n@@ -1279,6 +1281,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \tif (len > MAX_UNPACK_TREES)\n \t\tdie(\"unpack_trees takes at most %d trees\", MAX_UNPACK_TREES);\n \n+\ttrace_performance_enter();\n \tmemset(&el, 0, sizeof(el));\n \tif (!core_apply_sparse_checkout || !o->update)\n \t\to->skip_sparse_checkout = 1;\n@@ -1351,7 +1354,10 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t\t}\n \t\t}\n \n-\t\tif (traverse_trees(len, t, &info) < 0)\n+\t\ttrace_performance_enter();\n+\t\tret = traverse_trees(len, t, &info);\n+\t\ttrace_performance_leave(\"traverse_trees\");\n+\t\tif (ret < 0)\n \t\t\tgoto return_failed;\n \t}\n \n@@ -1443,6 +1449,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \to->src_index = NULL;\n \n done:\n+\ttrace_performance_leave(\"unpack_trees\");\n \tclear_exclude_list(&el);\n \treturn ret;\n \n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355292","messageId":"20180812081551.27927-5-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-1-pclouds@gmail.com","subject":"[PATCH v4 4/5] unpack-trees: reduce malloc in cache-tree walk","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-12T08:15:50Z","receivedAt":"2018-08-12T08:16:03Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This is a micro optimization that probably only shines on repos with\ndeep directory structure. Instead of allocating and freeing a new\ncache_entry in every iteration, we reuse the last one and only update\nthe parts that are new each iteration.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n unpack-trees.c | 29 ++++++++++++++++++++---------\n 1 file changed, 20 insertions(+), 9 deletions(-)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 07456d0fb2..6deb04c163 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -694,6 +694,8 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n \tstruct unpack_trees_options *o = info->data;\n+\tstruct cache_entry *tree_ce = NULL;\n+\tint ce_len = 0;\n \tint i, d;\n \n \tif (!o->merge)\n@@ -708,30 +710,39 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \t * get here in the first place.\n \t */\n \tfor (i = 0; i < nr_entries; i++) {\n-\t\tstruct cache_entry *tree_ce;\n-\t\tint len, rc;\n+\t\tint new_ce_len, len, rc;\n \n \t\tsrc[0] = o->src_index->cache[pos + i];\n \n \t\tlen = ce_namelen(src[0]);\n-\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n+\t\tnew_ce_len = cache_entry_size(len);\n+\n+\t\tif (new_ce_len > ce_len) {\n+\t\t\tnew_ce_len <<= 1;\n+\t\t\ttree_ce = xrealloc(tree_ce, new_ce_len);\n+\t\t\tmemset(tree_ce, 0, new_ce_len);\n+\t\t\tce_len = new_ce_len;\n+\n+\t\t\ttree_ce->ce_flags = create_ce_flags(0);\n+\n+\t\t\tfor (d = 1; d <= nr_names; d++)\n+\t\t\t\tsrc[d] = tree_ce;\n+\t\t}\n \n \t\ttree_ce->ce_mode = src[0]->ce_mode;\n-\t\ttree_ce->ce_flags = create_ce_flags(0);\n \t\ttree_ce->ce_namelen = len;\n \t\toidcpy(&tree_ce->oid, &src[0]->oid);\n \t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n \n-\t\tfor (d = 1; d <= nr_names; d++)\n-\t\t\tsrc[d] = tree_ce;\n-\n \t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n-\t\tfree(tree_ce);\n-\t\tif (rc < 0)\n+\t\tif (rc < 0) {\n+\t\t\tfree(tree_ce);\n \t\t\treturn rc;\n+\t\t}\n \n \t\tmark_ce_used(src[0], o);\n \t}\n+\tfree(tree_ce);\n \tif (o->debug_unpack)\n \t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n \t\t       nr_entries,\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355296","messageId":"20180812081551.27927-6-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-1-pclouds@gmail.com","subject":"[PATCH v4 5/5] unpack-trees: reuse (still valid) cache-tree from src_index","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-12T08:15:51Z","receivedAt":"2018-08-12T08:16:04Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"We do n-way merge by walking the source index and n trees at the same\ntime and add merge results to a new temporary index called o->result.\nThe merge result for any given path could be either\n\n- keep_entry(): same old index entry in o->src_index is reused\n- merged_entry(): either a new entry is added, or an existing one updated\n- deleted_entry(): one entry from o->src_index is removed\n\nFor some reason [1] we keep making sure that the source index's\ncache-tree is still valid if used by o->result: for all those\nmerged/deleted entries, we invalidate the same path in o->src_index,\nso only cache-trees covering the \"keep_entry\" parts remain good.\n\nBecause of this, the cache-tree from o->src_index can be perfectly\nreused in o->result. And in fact we already rely on this logic to\nreuse untracked cache in edf3b90553 (unpack-trees: preserve index\nextensions - 2017-05-08). Move the cache-tree to o->result before\ndoing cache_tree_update() to reduce hashing cost.\n\nSince cache_tree_update() has risen up as one of the most expensive\nparts in unpack_trees() after the last few patches. This does help\nreduce unpack_trees() time significantly (on webkit.git):\n\n    before       after\n  --------------------------------------------------------------------\n    0.080394752  0.051258167 s:  read cache .git/index\n    0.216010838  0.212106298 s:  preload index\n    0.008534301  0.280521764 s:  refresh index\n    0.251992198  0.218160442 s:   traverse_trees\n    0.377031383  0.374948191 s:   check_updates\n    0.372768105  0.037040114 s:   cache_tree_update\n    1.045887251  0.672031609 s:  unpack_trees\n    0.314983512  0.317456290 s:  write index, changed mask = 2e\n    0.062572653  0.038382654 s:    traverse_trees\n    0.000022544  0.000042731 s:    check_updates\n    0.073795585  0.050930053 s:   unpack_trees\n    0.073807557  0.051099735 s:  diff-index\n    1.938191592  1.614241153 s: git command: git checkout -\n\n[1] I'm pretty sure the reason is an oversight in 34110cd4e3 (Make\n    'unpack_trees()' have a separate source and destination index -\n    2008-03-06). That patch aims to _not_ update the source index at\n    all. The invalidation should have been done on o->result in that\n    patch. But then there was no cache-tree on o->result even then so\n    it's pointless to do so.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c   | 2 ++\n unpack-trees.c | 2 +-\n 2 files changed, 3 insertions(+), 1 deletion(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 4fd35f4f37..2b5646ef26 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2763,4 +2763,6 @@ void move_index_extensions(struct index_state *dst, struct index_state *src)\n {\n \tdst->untracked = src->untracked;\n \tsrc->untracked = NULL;\n+\tdst->cache_tree = src->cache_tree;\n+\tsrc->cache_tree = NULL;\n }\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 6deb04c163..d822662c75 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -1570,6 +1570,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \n \tret = check_updates(o) ? (-2) : 0;\n \tif (o->dst_index) {\n+\t\tmove_index_extensions(&o->result, o->src_index);\n \t\tif (!ret) {\n \t\t\tif (!o->result.cache_tree)\n \t\t\t\to->result.cache_tree = cache_tree();\n@@ -1578,7 +1579,6 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t\t\t\t\t  WRITE_TREE_SILENT |\n \t\t\t\t\t\t  WRITE_TREE_REPAIR);\n \t\t}\n-\t\tmove_index_extensions(&o->result, o->src_index);\n \t\tdiscard_index(o->dst_index);\n \t\t*o->dst_index = o->result;\n \t} else {\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355299","messageId":"CAOhcEPZphaKASyMAmZ5erdn-fygdVrvtPScTL_zZmAAgCYYKqQ@mail.gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-3-pclouds@gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Thomas Adam","fromEmail":"thomas@xteddy.org","sentAt":"2018-08-12T10:05:37Z","receivedAt":"2018-08-12T10:05:54Z","isPatch":true,"sender":{"key":"thomas@xteddy.org","avatar":"https://gravatar.com/avatar/e7256db4738e501e5d2e84f00bb0bd99503165729573848a03330301fc2adc4a?d=mp&s=160"},"body":"On Sun, 12 Aug 2018 at 09:19, Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n\nHi,\n\n> +       trace_performance_leave(\"cache_tree_update\");\n\nI would suggest trace_performance_leave() calls use __func__ instead.\nThat way, there's no ambiguity if the function name ever changes.\n\nKindly,\nThomas\n"},{"id":"355365","messageId":"CABPp-BEDQfzyZjD0CuZKhvj3iUi0H6Ar0Fgm2UhehjP1pnWKgA@mail.gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-6-pclouds@gmail.com","subject":"Re: [PATCH v4 5/5] unpack-trees: reuse (still valid) cache-tree from src_index","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-08-13T15:48:45Z","receivedAt":"2018-08-13T15:49:00Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sun, Aug 12, 2018 at 1:16 AM Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n>\n> We do n-way merge by walking the source index and n trees at the same\n> time and add merge results to a new temporary index called o->result.\n> The merge result for any given path could be either\n>\n> - keep_entry(): same old index entry in o->src_index is reused\n> - merged_entry(): either a new entry is added, or an existing one updated\n> - deleted_entry(): one entry from o->src_index is removed\n>\n> For some reason [1] we keep making sure that the source index's\n> cache-tree is still valid if used by o->result: for all those\n> merged/deleted entries, we invalidate the same path in o->src_index,\n> so only cache-trees covering the \"keep_entry\" parts remain good.\n>\n> Because of this, the cache-tree from o->src_index can be perfectly\n> reused in o->result. And in fact we already rely on this logic to\n> reuse untracked cache in edf3b90553 (unpack-trees: preserve index\n> extensions - 2017-05-08). Move the cache-tree to o->result before\n> doing cache_tree_update() to reduce hashing cost.\n>\n> Since cache_tree_update() has risen up as one of the most expensive\n> parts in unpack_trees() after the last few patches. This does help\n> reduce unpack_trees() time significantly (on webkit.git):\n>\n>     before       after\n>   --------------------------------------------------------------------\n>     0.080394752  0.051258167 s:  read cache .git/index\n>     0.216010838  0.212106298 s:  preload index\n>     0.008534301  0.280521764 s:  refresh index\n>     0.251992198  0.218160442 s:   traverse_trees\n>     0.377031383  0.374948191 s:   check_updates\n>     0.372768105  0.037040114 s:   cache_tree_update\n>     1.045887251  0.672031609 s:  unpack_trees\n\nCool, nice drop in both cache_tree_update() and unpack_trees().  But\nwhy did refresh_index() go up so much?  That should have been\nunaffected by this patch to, so it seems like something odd is going\non.  Any ideas?\n"},{"id":"355366","messageId":"CACsJy8D+VMBO7oNH3DJ4SspLj7OBq78rkaBHB35BTLwuVsLekQ@mail.gmail.com","threadId":"48913","inReplyTo":"CABPp-BEDQfzyZjD0CuZKhvj3iUi0H6Ar0Fgm2UhehjP1pnWKgA@mail.gmail.com","subject":"Re: [PATCH v4 5/5] unpack-trees: reuse (still valid) cache-tree from src_index","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-13T15:57:14Z","receivedAt":"2018-08-13T15:57:43Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 13, 2018 at 5:48 PM Elijah Newren <newren@gmail.com> wrote:\n>\n> On Sun, Aug 12, 2018 at 1:16 AM Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n> >\n> > We do n-way merge by walking the source index and n trees at the same\n> > time and add merge results to a new temporary index called o->result.\n> > The merge result for any given path could be either\n> >\n> > - keep_entry(): same old index entry in o->src_index is reused\n> > - merged_entry(): either a new entry is added, or an existing one updated\n> > - deleted_entry(): one entry from o->src_index is removed\n> >\n> > For some reason [1] we keep making sure that the source index's\n> > cache-tree is still valid if used by o->result: for all those\n> > merged/deleted entries, we invalidate the same path in o->src_index,\n> > so only cache-trees covering the \"keep_entry\" parts remain good.\n> >\n> > Because of this, the cache-tree from o->src_index can be perfectly\n> > reused in o->result. And in fact we already rely on this logic to\n> > reuse untracked cache in edf3b90553 (unpack-trees: preserve index\n> > extensions - 2017-05-08). Move the cache-tree to o->result before\n> > doing cache_tree_update() to reduce hashing cost.\n> >\n> > Since cache_tree_update() has risen up as one of the most expensive\n> > parts in unpack_trees() after the last few patches. This does help\n> > reduce unpack_trees() time significantly (on webkit.git):\n> >\n> >     before       after\n> >   --------------------------------------------------------------------\n> >     0.080394752  0.051258167 s:  read cache .git/index\n> >     0.216010838  0.212106298 s:  preload index\n> >     0.008534301  0.280521764 s:  refresh index\n> >     0.251992198  0.218160442 s:   traverse_trees\n> >     0.377031383  0.374948191 s:   check_updates\n> >     0.372768105  0.037040114 s:   cache_tree_update\n> >     1.045887251  0.672031609 s:  unpack_trees\n>\n> Cool, nice drop in both cache_tree_update() and unpack_trees().  But\n> why did refresh_index() go up so much?  That should have been\n> unaffected by this patch to, so it seems like something odd is going\n> on.  Any ideas?\n\nProbably fs cache and stuff. This is a laptop with just 4GB RAM and a\nvery slow disk so if something triggers in the background and evicts\nsome webkit.git's stat info, refresh_index will get hot fast (and with\n275k files, webkit.git needs quite a bit of ram to make sure stat()\ncalls don't hit the disk).\n-- \nDuy\n"},{"id":"355367","messageId":"a8543f33-eb1a-4122-87a4-f8d888af7381@gmail.com","threadId":"48913","inReplyTo":"CABPp-BEDQfzyZjD0CuZKhvj3iUi0H6Ar0Fgm2UhehjP1pnWKgA@mail.gmail.com","subject":"Re: [PATCH v4 5/5] unpack-trees: reuse (still valid) cache-tree from src_index","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-13T16:05:06Z","receivedAt":"2018-08-13T16:05:11Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/13/2018 11:48 AM, Elijah Newren wrote:\n> On Sun, Aug 12, 2018 at 1:16 AM Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n>>\n>> We do n-way merge by walking the source index and n trees at the same\n>> time and add merge results to a new temporary index called o->result.\n>> The merge result for any given path could be either\n>>\n>> - keep_entry(): same old index entry in o->src_index is reused\n>> - merged_entry(): either a new entry is added, or an existing one updated\n>> - deleted_entry(): one entry from o->src_index is removed\n>>\n>> For some reason [1] we keep making sure that the source index's\n>> cache-tree is still valid if used by o->result: for all those\n>> merged/deleted entries, we invalidate the same path in o->src_index,\n>> so only cache-trees covering the \"keep_entry\" parts remain good.\n>>\n>> Because of this, the cache-tree from o->src_index can be perfectly\n>> reused in o->result. And in fact we already rely on this logic to\n>> reuse untracked cache in edf3b90553 (unpack-trees: preserve index\n>> extensions - 2017-05-08). Move the cache-tree to o->result before\n>> doing cache_tree_update() to reduce hashing cost.\n>>\n>> Since cache_tree_update() has risen up as one of the most expensive\n>> parts in unpack_trees() after the last few patches. This does help\n>> reduce unpack_trees() time significantly (on webkit.git):\n>>\n>>      before       after\n>>    --------------------------------------------------------------------\n>>      0.080394752  0.051258167 s:  read cache .git/index\n>>      0.216010838  0.212106298 s:  preload index\n>>      0.008534301  0.280521764 s:  refresh index\n>>      0.251992198  0.218160442 s:   traverse_trees\n>>      0.377031383  0.374948191 s:   check_updates\n>>      0.372768105  0.037040114 s:   cache_tree_update\n>>      1.045887251  0.672031609 s:  unpack_trees\n> \n> Cool, nice drop in both cache_tree_update() and unpack_trees().  But\n> why did refresh_index() go up so much?  That should have been\n> unaffected by this patch to, so it seems like something odd is going\n> on.  Any ideas?\n> \n\nI was part way through writing a patch that would copy the valid parts \nof the cache-tree from the source index to the dest index but the \nobservation that the source index cache tree was already being \ninvalidated properly which allows the simple pointer \"copy\" is much better!\n\nI run some tests on a large repo and the results look very promising.\n\nbase\tnew\tdiff\t% saved\t\n0.55\t0.52\t0.02\t4.32%\ts:  read cache .git/index\n0.31\t0.30\t0.01\t2.98%\ts:  initialize name hash\n0.03\t0.02\t0.00\t9.98%\ts:  preload index\n0.09\t0.09\t0.00\t4.86%\ts:  refresh index\n5.93\t1.19\t4.74\t79.95%\ts:   traverse_trees\n0.12\t0.13\t-0.01\t-4.15%\ts:   check_updates\n2.14\t0.00\t2.14\t100.00%\ts:   cache_tree_update\n10.63\t4.29\t6.33\t59.59%\ts:  unpack_trees\n0.97\t0.91\t0.06\t6.41%\ts:  write index, changed mask = 28\n3.49\t0.18\t3.31\t94.91%\ts:    traverse_trees\n0.00\t0.00\t0.00\t17.53%\ts:    check_updates\n3.61\t0.30\t3.31\t91.77%\ts:   unpack_trees\n3.61\t0.30\t3.31\t91.77%\ts:  diff-index\n17.28\t8.36\t8.92\t51.62%\ts: git command: c:git.exe checkout\n\nSame methodology as before, I ran \"git checkout\" 5 times, threw away the \nfirst 2 runs and averaged the last 3.  I entered 0 for the \"new\" \ncache_tree_update line as it no longer reports anything.\n"},{"id":"355394","messageId":"CACsJy8A2L-WX_RLP37fF-f=YBLkCY7+iZLo+Uz=H9EFOWMVnoA@mail.gmail.com","threadId":"48913","inReplyTo":"a8543f33-eb1a-4122-87a4-f8d888af7381@gmail.com","subject":"Re: [PATCH v4 5/5] unpack-trees: reuse (still valid) cache-tree from src_index","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-13T16:25:36Z","receivedAt":"2018-08-13T16:26:05Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 13, 2018 at 6:05 PM Ben Peart <peartben@gmail.com> wrote:\n> I was part way through writing a patch that would copy the valid parts\n> of the cache-tree from the source index to the dest index\n\nYeah sorry about that. I make bad judgements all the time, unfortunately.\n\nIf it's sort of working though, please post to the list anyway to\narchive it. Who knows, some time down the road we might actually need\nit again.\n\n> I run some tests on a large repo and the results look very promising.\n>\n> base    new     diff    % saved\n> 0.55    0.52    0.02    4.32%   s:  read cache .git/index\n> 0.31    0.30    0.01    2.98%   s:  initialize name hash\n> 0.03    0.02    0.00    9.98%   s:  preload index\n> 0.09    0.09    0.00    4.86%   s:  refresh index\n> 5.93    1.19    4.74    79.95%  s:   traverse_trees\n> 0.12    0.13    -0.01   -4.15%  s:   check_updates\n> 2.14    0.00    2.14    100.00% s:   cache_tree_update\n> 10.63   4.29    6.33    59.59%  s:  unpack_trees\n\nThere's a big gap here, I think. unpack_trees() takes 4s but the sum\nof traverse_trees, check_updates and cache_tree_update is 1.5s top. I\nguess that's sparse checkout and stuff? It's either that or there's\nanother big hidden thing we should pay attention to ;-)\n\n> 0.97    0.91    0.06    6.41%   s:  write index, changed mask = 28\n> 3.49    0.18    3.31    94.91%  s:    traverse_trees\n> 0.00    0.00    0.00    17.53%  s:    check_updates\n> 3.61    0.30    3.31    91.77%  s:   unpack_trees\n> 3.61    0.30    3.31    91.77%  s:  diff-index\n> 17.28   8.36    8.92    51.62%  s: git command: c:git.exe checkout\n>\n> Same methodology as before, I ran \"git checkout\" 5 times, threw away the\n> first 2 runs and averaged the last 3.  I entered 0 for the \"new\"\n> cache_tree_update line as it no longer reports anything.\n-- \nDuy\n"},{"id":"355406","messageId":"d3faf3e0-5526-8613-9c7e-796f627eaef3@gmail.com","threadId":"48913","inReplyTo":"CACsJy8A2L-WX_RLP37fF-f=YBLkCY7+iZLo+Uz=H9EFOWMVnoA@mail.gmail.com","subject":"Re: [PATCH v4 5/5] unpack-trees: reuse (still valid) cache-tree from src_index","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-13T17:15:58Z","receivedAt":"2018-08-13T17:16:02Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/13/2018 12:25 PM, Duy Nguyen wrote:\n> On Mon, Aug 13, 2018 at 6:05 PM Ben Peart <peartben@gmail.com> wrote:\n>> I was part way through writing a patch that would copy the valid parts\n>> of the cache-tree from the source index to the dest index\n> \n> Yeah sorry about that. I make bad judgements all the time, unfortunately.\n> \n> If it's sort of working though, please post to the list anyway to\n> archive it. Who knows, some time down the road we might actually need\n> it again.\n> \n>> I run some tests on a large repo and the results look very promising.\n>>\n>> base    new     diff    % saved\n>> 0.55    0.52    0.02    4.32%   s:  read cache .git/index\n>> 0.31    0.30    0.01    2.98%   s:  initialize name hash\n>> 0.03    0.02    0.00    9.98%   s:  preload index\n>> 0.09    0.09    0.00    4.86%   s:  refresh index\n>> 5.93    1.19    4.74    79.95%  s:   traverse_trees\n>> 0.12    0.13    -0.01   -4.15%  s:   check_updates\n>> 2.14    0.00    2.14    100.00% s:   cache_tree_update\n>> 10.63   4.29    6.33    59.59%  s:  unpack_trees\n> \n> There's a big gap here, I think. unpack_trees() takes 4s but the sum\n> of traverse_trees, check_updates and cache_tree_update is 1.5s top. I\n> guess that's sparse checkout and stuff? It's either that or there's\n> another big hidden thing we should pay attention to ;-)\n> \n\nYes, there are additional costs associated with the sparse-checkout and \nexcludes logic.  We've sped that up significantly by converting it to a \nhashmap but it still has measurable cost as we have to compute the hash \nof the cache entry name before looking it up.\n\nName                                            Inc %\t     Inc\n+ git!unpack_trees                      \t 50.9\t   4,575\n|+ git!clear_ce_flags_1                 \t 16.5\t   1,479\n||+ git!is_included_in_virtualfilesystem\t 15.9\t   1,430\n|| + git!check_includes_hashmap         \t 15.8\t   1,418\n|+ git!traverse_trees                   \t 16.3\t   1,468\n|+ git!cache_tree_fully_valid           \t 11.7\t   1,055\n|+ git!check_updates                    \t  1.9\t     169\n|+ git!discard_index                    \t  1.8\t     162\n|+ git!apply_sparse_checkout            \t  0.2\t      15\n|+ git!next_cache_entry                 \t  0.0\t       3\n\n"},{"id":"355439","messageId":"5361e977-c476-0efe-5a2a-7d377dd51bbe@gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-2-pclouds@gmail.com","subject":"Re: [PATCH v4 1/5] trace.h: support nested performance tracing","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-13T18:39:16Z","receivedAt":"2018-08-13T18:39:21Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/12/2018 4:15 AM, Nguyễn Thái Ngọc Duy wrote:\n> Performance measurements are listed right now as a flat list, which is\n> fine when we measure big blocks. But when we start adding more and\n> more measurements, some of them could be just part of a bigger\n> measurement and a flat list gives a wrong impression that they are\n> executed at the same level instead of nested.\n> \n> Add trace_performance_enter() and trace_performance_leave() to allow\n> indent these nested measurements. For now it does not help much\n> because the only nested thing is (lazy) name hash initialization\n> (e.g. called in diff-index from \"git status\"). This will help more\n> because I'm going to add some more tracing that's actually nested.\n> \n\nI reviewed this and it looks reasonable to me.\n\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>   diff-lib.c      |  4 +--\n>   dir.c           |  4 +--\n>   name-hash.c     |  4 +--\n>   preload-index.c |  4 +--\n>   read-cache.c    | 11 ++++----\n>   trace.c         | 69 ++++++++++++++++++++++++++++++++++++++++++++-----\n>   trace.h         | 15 +++++++++++\n>   7 files changed, 92 insertions(+), 19 deletions(-)\n> \n> diff --git a/diff-lib.c b/diff-lib.c\n> index a9f38eb5a3..1ffa22c882 100644\n> --- a/diff-lib.c\n> +++ b/diff-lib.c\n> @@ -518,8 +518,8 @@ static int diff_cache(struct rev_info *revs,\n>   int run_diff_index(struct rev_info *revs, int cached)\n>   {\n>   \tstruct object_array_entry *ent;\n> -\tuint64_t start = getnanotime();\n>   \n> +\ttrace_performance_enter();\n>   \tent = revs->pending.objects;\n>   \tif (diff_cache(revs, &ent->item->oid, ent->name, cached))\n>   \t\texit(128);\n> @@ -528,7 +528,7 @@ int run_diff_index(struct rev_info *revs, int cached)\n>   \tdiffcore_fix_diff_index(&revs->diffopt);\n>   \tdiffcore_std(&revs->diffopt);\n>   \tdiff_flush(&revs->diffopt);\n> -\ttrace_performance_since(start, \"diff-index\");\n> +\ttrace_performance_leave(\"diff-index\");\n>   \treturn 0;\n>   }\n>   \n> diff --git a/dir.c b/dir.c\n> index 21e6f2520a..c5e9fc8cea 100644\n> --- a/dir.c\n> +++ b/dir.c\n> @@ -2263,11 +2263,11 @@ int read_directory(struct dir_struct *dir, struct index_state *istate,\n>   \t\t   const char *path, int len, const struct pathspec *pathspec)\n>   {\n>   \tstruct untracked_cache_dir *untracked;\n> -\tuint64_t start = getnanotime();\n>   \n\nI think removing the cost of has_symlink_leading_path() from this perf \ntrace is probably OK to simplify the enter/leave logic.\n\n>   \tif (has_symlink_leading_path(path, len))\n>   \t\treturn dir->nr;\n>   \n> +\ttrace_performance_enter();\n>   \tuntracked = validate_untracked_cache(dir, len, pathspec);\n>   \tif (!untracked)\n>   \t\t/*\n> @@ -2302,7 +2302,7 @@ int read_directory(struct dir_struct *dir, struct index_state *istate,\n>   \t\tdir->nr = i;\n>   \t}\n>   \n> -\ttrace_performance_since(start, \"read directory %.*s\", len, path);\n> +\ttrace_performance_leave(\"read directory %.*s\", len, path);\n>   \tif (dir->untracked) {\n>   \t\tstatic int force_untracked_cache = -1;\n>   \t\tstatic struct trace_key trace_untracked_stats = TRACE_KEY_INIT(UNTRACKED_STATS);\n> diff --git a/name-hash.c b/name-hash.c\n> index 163849831c..1fcda73cb3 100644\n> --- a/name-hash.c\n> +++ b/name-hash.c\n> @@ -578,10 +578,10 @@ static void threaded_lazy_init_name_hash(\n>   \n>   static void lazy_init_name_hash(struct index_state *istate)\n>   {\n> -\tuint64_t start = getnanotime();\n>   \n>   \tif (istate->name_hash_initialized)\n>   \t\treturn;\n> +\ttrace_performance_enter();\n>   \thashmap_init(&istate->name_hash, cache_entry_cmp, NULL, istate->cache_nr);\n>   \thashmap_init(&istate->dir_hash, dir_entry_cmp, NULL, istate->cache_nr);\n>   \n> @@ -602,7 +602,7 @@ static void lazy_init_name_hash(struct index_state *istate)\n>   \t}\n>   \n>   \tistate->name_hash_initialized = 1;\n> -\ttrace_performance_since(start, \"initialize name hash\");\n> +\ttrace_performance_leave(\"initialize name hash\");\n>   }\n>   \n>   /*\n> diff --git a/preload-index.c b/preload-index.c\n> index 4d08d44874..d7f7919ba2 100644\n> --- a/preload-index.c\n> +++ b/preload-index.c\n> @@ -78,7 +78,6 @@ static void preload_index(struct index_state *index,\n>   {\n>   \tint threads, i, work, offset;\n>   \tstruct thread_data data[MAX_PARALLEL];\n> -\tuint64_t start = getnanotime();\n>   \n>   \tif (!core_preload_index)\n>   \t\treturn;\n> @@ -88,6 +87,7 @@ static void preload_index(struct index_state *index,\n>   \t\tthreads = 2;\n>   \tif (threads < 2)\n>   \t\treturn;\n> +\ttrace_performance_enter();\n>   \tif (threads > MAX_PARALLEL)\n>   \t\tthreads = MAX_PARALLEL;\n>   \toffset = 0;\n> @@ -109,7 +109,7 @@ static void preload_index(struct index_state *index,\n>   \t\tif (pthread_join(p->pthread, NULL))\n>   \t\t\tdie(\"unable to join threaded lstat\");\n>   \t}\n> -\ttrace_performance_since(start, \"preload index\");\n> +\ttrace_performance_leave(\"preload index\");\n>   }\n>   #endif\n>   \n> diff --git a/read-cache.c b/read-cache.c\n> index e865254bea..4fd35f4f37 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1399,8 +1399,8 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n>   \tconst char *typechange_fmt;\n>   \tconst char *added_fmt;\n>   \tconst char *unmerged_fmt;\n> -\tuint64_t start = getnanotime();\n>   \n> +\ttrace_performance_enter();\n>   \tmodified_fmt = (in_porcelain ? \"M\\t%s\\n\" : \"%s: needs update\\n\");\n>   \tdeleted_fmt = (in_porcelain ? \"D\\t%s\\n\" : \"%s: needs update\\n\");\n>   \ttypechange_fmt = (in_porcelain ? \"T\\t%s\\n\" : \"%s needs update\\n\");\n> @@ -1470,7 +1470,7 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n>   \n>   \t\treplace_index_entry(istate, i, new_entry);\n>   \t}\n> -\ttrace_performance_since(start, \"refresh index\");\n> +\ttrace_performance_leave(\"refresh index\");\n>   \treturn has_errors;\n>   }\n>   \n> @@ -1901,7 +1901,6 @@ static void freshen_shared_index(const char *shared_index, int warn)\n>   int read_index_from(struct index_state *istate, const char *path,\n>   \t\t    const char *gitdir)\n>   {\n> -\tuint64_t start = getnanotime();\n>   \tstruct split_index *split_index;\n>   \tint ret;\n>   \tchar *base_oid_hex;\n> @@ -1911,8 +1910,9 @@ int read_index_from(struct index_state *istate, const char *path,\n>   \tif (istate->initialized)\n>   \t\treturn istate->cache_nr;\n>   \n> +\ttrace_performance_enter();\n>   \tret = do_read_index(istate, path, 0);\n> -\ttrace_performance_since(start, \"read cache %s\", path);\n> +\ttrace_performance_leave(\"read cache %s\", path);\n>   \n>   \tsplit_index = istate->split_index;\n>   \tif (!split_index || is_null_oid(&split_index->base_oid)) {\n> @@ -1920,6 +1920,7 @@ int read_index_from(struct index_state *istate, const char *path,\n>   \t\treturn ret;\n>   \t}\n>   \n\nThis one is kind of odd how it's splitting up the index read from the \nsplit index but it's no more odd than it was before.\n\n> +\ttrace_performance_enter();\n>   \tif (split_index->base)\n>   \t\tdiscard_index(split_index->base);\n>   \telse\n> @@ -1936,8 +1937,8 @@ int read_index_from(struct index_state *istate, const char *path,\n>   \tfreshen_shared_index(base_path, 0);\n>   \tmerge_base_index(istate);\n>   \tpost_read_index_from(istate);\n> -\ttrace_performance_since(start, \"read cache %s\", base_path);\n>   \tfree(base_path);\n> +\ttrace_performance_leave(\"read cache %s\", base_path);\n>   \treturn ret;\n>   }\n>   \n> diff --git a/trace.c b/trace.c\n> index fc623e91fd..fa4a2e7120 100644\n> --- a/trace.c\n> +++ b/trace.c\n> @@ -176,10 +176,30 @@ void trace_strbuf_fl(const char *file, int line, struct trace_key *key,\n>   \tstrbuf_release(&buf);\n>   }\n>   \n> +static uint64_t perf_start_times[10];\n> +static int perf_indent;\n> +\n> +uint64_t trace_performance_enter(void)\n> +{\n> +\tuint64_t now;\n> +\n> +\tif (!trace_want(&trace_perf_key))\n> +\t\treturn 0;\n> +\n> +\tnow = getnanotime();\n> +\tperf_start_times[perf_indent] = now;\n> +\tif (perf_indent + 1 < ARRAY_SIZE(perf_start_times))\n> +\t\tperf_indent++;\n> +\telse\n> +\t\tBUG(\"Too deep indentation\");\n> +\treturn now;\n> +}\n> +\n>   static void trace_performance_vprintf_fl(const char *file, int line,\n>   \t\t\t\t\t uint64_t nanos, const char *format,\n>   \t\t\t\t\t va_list ap)\n>   {\n> +\tstatic const char space[] = \"          \";\n>   \tstruct strbuf buf = STRBUF_INIT;\n>   \n>   \tif (!prepare_trace_line(file, line, &trace_perf_key, &buf))\n> @@ -188,7 +208,10 @@ static void trace_performance_vprintf_fl(const char *file, int line,\n>   \tstrbuf_addf(&buf, \"performance: %.9f s\", (double) nanos / 1000000000);\n>   \n>   \tif (format && *format) {\n> -\t\tstrbuf_addstr(&buf, \": \");\n> +\t\tif (perf_indent >= strlen(space))\n> +\t\t\tBUG(\"Too deep indentation\");\n> +\n> +\t\tstrbuf_addf(&buf, \":%.*s \", perf_indent, space);\n>   \t\tstrbuf_vaddf(&buf, format, ap);\n>   \t}\n>   \n> @@ -244,6 +267,24 @@ void trace_performance_since(uint64_t start, const char *format, ...)\n>   \tva_end(ap);\n>   }\n>   \n> +void trace_performance_leave(const char *format, ...)\n> +{\n> +\tva_list ap;\n> +\tuint64_t since;\n> +\n> +\tif (perf_indent)\n> +\t\tperf_indent--;\n> +\n> +\tif (!format) /* Allow callers to leave without tracing anything */\n> +\t\treturn;\n> +\n> +\tsince = perf_start_times[perf_indent];\n> +\tva_start(ap, format);\n> +\ttrace_performance_vprintf_fl(NULL, 0, getnanotime() - since,\n> +\t\t\t\t     format, ap);\n> +\tva_end(ap);\n> +}\n> +\n>   #else\n>   \n>   void trace_printf_key_fl(const char *file, int line, struct trace_key *key,\n> @@ -273,6 +314,24 @@ void trace_performance_fl(const char *file, int line, uint64_t nanos,\n>   \tva_end(ap);\n>   }\n>   \n> +void trace_performance_leave_fl(const char *file, int line,\n> +\t\t\t\tuint64_t nanos, const char *format, ...)\n> +{\n> +\tva_list ap;\n> +\tuint64_t since;\n> +\n> +\tif (perf_indent)\n> +\t\tperf_indent--;\n> +\n> +\tif (!format) /* Allow callers to leave without tracing anything */\n> +\t\treturn;\n> +\n> +\tsince = perf_start_times[perf_indent];\n> +\tva_start(ap, format);\n> +\ttrace_performance_vprintf_fl(file, line, nanos - since, format, ap);\n> +\tva_end(ap);\n> +}\n> +\n>   #endif /* HAVE_VARIADIC_MACROS */\n>   \n>   \n> @@ -411,13 +470,11 @@ uint64_t getnanotime(void)\n>   \t}\n>   }\n>   \n> -static uint64_t command_start_time;\n>   static struct strbuf command_line = STRBUF_INIT;\n>   \n>   static void print_command_performance_atexit(void)\n>   {\n> -\ttrace_performance_since(command_start_time, \"git command:%s\",\n> -\t\t\t\tcommand_line.buf);\n> +\ttrace_performance_leave(\"git command:%s\", command_line.buf);\n>   }\n>   \n>   void trace_command_performance(const char **argv)\n> @@ -425,10 +482,10 @@ void trace_command_performance(const char **argv)\n>   \tif (!trace_want(&trace_perf_key))\n>   \t\treturn;\n>   \n> -\tif (!command_start_time)\n> +\tif (!command_line.len)\n>   \t\tatexit(print_command_performance_atexit);\n>   \n>   \tstrbuf_reset(&command_line);\n>   \tsq_quote_argv_pretty(&command_line, argv);\n> -\tcommand_start_time = getnanotime();\n> +\ttrace_performance_enter();\n>   }\n> diff --git a/trace.h b/trace.h\n> index 2b6a1bc17c..171b256d26 100644\n> --- a/trace.h\n> +++ b/trace.h\n> @@ -23,6 +23,7 @@ extern void trace_disable(struct trace_key *key);\n>   extern uint64_t getnanotime(void);\n>   extern void trace_command_performance(const char **argv);\n>   extern void trace_verbatim(struct trace_key *key, const void *buf, unsigned len);\n> +uint64_t trace_performance_enter(void);\n>   \n>   #ifndef HAVE_VARIADIC_MACROS\n>   \n> @@ -45,6 +46,9 @@ extern void trace_performance(uint64_t nanos, const char *format, ...);\n>   __attribute__((format (printf, 2, 3)))\n>   extern void trace_performance_since(uint64_t start, const char *format, ...);\n>   \n> +__attribute__((format (printf, 1, 2)))\n> +void trace_performance_leave(const char *format, ...);\n> +\n>   #else\n>   \n>   /*\n> @@ -118,6 +122,14 @@ extern void trace_performance_since(uint64_t start, const char *format, ...);\n>   \t\t\t\t\t     __VA_ARGS__);\t\t    \\\n>   \t} while (0)\n>   \n> +#define trace_performance_leave(...)\t\t\t\t\t    \\\n> +\tdo {\t\t\t\t\t\t\t\t    \\\n> +\t\tif (trace_pass_fl(&trace_perf_key))\t\t\t    \\\n> +\t\t\ttrace_performance_leave_fl(TRACE_CONTEXT, __LINE__, \\\n> +\t\t\t\t\t\t   getnanotime(),\t    \\\n> +\t\t\t\t\t\t   __VA_ARGS__);\t    \\\n> +\t} while (0)\n> +\n>   /* backend functions, use non-*fl macros instead */\n>   __attribute__((format (printf, 4, 5)))\n>   extern void trace_printf_key_fl(const char *file, int line, struct trace_key *key,\n> @@ -130,6 +142,9 @@ extern void trace_strbuf_fl(const char *file, int line, struct trace_key *key,\n>   __attribute__((format (printf, 4, 5)))\n>   extern void trace_performance_fl(const char *file, int line,\n>   \t\t\t\t uint64_t nanos, const char *fmt, ...);\n> +__attribute__((format (printf, 4, 5)))\n> +extern void trace_performance_leave_fl(const char *file, int line,\n> +\t\t\t\t       uint64_t nanos, const char *fmt, ...);\n>   static inline int trace_pass_fl(struct trace_key *key)\n>   {\n>   \treturn key->fd || !key->initialized;\n> \n"},{"id":"355442","messageId":"34bbad06-4b19-6558-3004-58b0498a5166@gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-3-pclouds@gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-13T18:44:41Z","receivedAt":"2018-08-13T18:44:45Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/12/2018 4:15 AM, Nguyễn Thái Ngọc Duy wrote:\n> We're going to optimize unpack_trees() a bit in the following\n> patches. Let's add some tracing to measure how long it takes before\n> and after. This is the baseline (\"git checkout -\" on webkit.git, 275k\n> files on worktree)\n> \n>      performance: 0.056651714 s:  read cache .git/index\n>      performance: 0.183101080 s:  preload index\n>      performance: 0.008584433 s:  refresh index\n>      performance: 0.633767589 s:   traverse_trees\n>      performance: 0.340265448 s:   check_updates\n>      performance: 0.381884638 s:   cache_tree_update\n>      performance: 1.401562947 s:  unpack_trees\n>      performance: 0.338687914 s:  write index, changed mask = 2e\n>      performance: 0.411927922 s:    traverse_trees\n>      performance: 0.000023335 s:    check_updates\n>      performance: 0.423697246 s:   unpack_trees\n>      performance: 0.423708360 s:  diff-index\n>      performance: 2.559524127 s: git command: git checkout -\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>   cache-tree.c   | 2 ++\n>   unpack-trees.c | 9 ++++++++-\n>   2 files changed, 10 insertions(+), 1 deletion(-)\n> \n> diff --git a/cache-tree.c b/cache-tree.c\n> index 6b46711996..105f13806f 100644\n> --- a/cache-tree.c\n> +++ b/cache-tree.c\n> @@ -433,7 +433,9 @@ int cache_tree_update(struct index_state *istate, int flags)\n>   \n>   \tif (i)\n>   \t\treturn i;\n> +\ttrace_performance_enter();\n\nThis one is a little odd to me.  I think the either the \ntrace_performance_enter() call should move up to include the \nverify_cache() call or the enter/leave should move into the update_one() \ncall as that is all it is measuring/reporting on.\n\n>   \ti = update_one(it, cache, entries, \"\", 0, &skip, flags);\n> +\ttrace_performance_leave(\"cache_tree_update\");\n>   \tif (i < 0)\n>   \t\treturn i;\n>   \tistate->cache_changed |= CACHE_TREE_CHANGED;\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index cd0680f11e..b237eaa0f2 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -354,6 +354,7 @@ static int check_updates(struct unpack_trees_options *o)\n>   \tstruct checkout state = CHECKOUT_INIT;\n>   \tint i;\n>   \n> +\ttrace_performance_enter();\n>   \tstate.force = 1;\n>   \tstate.quiet = 1;\n>   \tstate.refresh_cache = 1;\n> @@ -423,6 +424,7 @@ static int check_updates(struct unpack_trees_options *o)\n>   \terrs |= finish_delayed_checkout(&state);\n>   \tif (o->update)\n>   \t\tgit_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n> +\ttrace_performance_leave(\"check_updates\");\n>   \treturn errs != 0;\n>   }\n>   \n> @@ -1279,6 +1281,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>   \tif (len > MAX_UNPACK_TREES)\n>   \t\tdie(\"unpack_trees takes at most %d trees\", MAX_UNPACK_TREES);\n>   \n> +\ttrace_performance_enter();\n>   \tmemset(&el, 0, sizeof(el));\n>   \tif (!core_apply_sparse_checkout || !o->update)\n>   \t\to->skip_sparse_checkout = 1;\n> @@ -1351,7 +1354,10 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>   \t\t\t}\n>   \t\t}\n>   \n> -\t\tif (traverse_trees(len, t, &info) < 0)\n> +\t\ttrace_performance_enter();\n> +\t\tret = traverse_trees(len, t, &info);\n> +\t\ttrace_performance_leave(\"traverse_trees\");\n\nWhy not move this enter/leave pair into the traverse_trees() function \nitself?\n\n> +\t\tif (ret < 0)\n>   \t\t\tgoto return_failed;\n>   \t}\n>   \n> @@ -1443,6 +1449,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>   \to->src_index = NULL;\n>   \n>   done:\n> +\ttrace_performance_leave(\"unpack_trees\");\n>   \tclear_exclude_list(&el);\n>   \treturn ret;\n>   \n> \n"},{"id":"355445","messageId":"xmqqin4e5ced.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"CAOhcEPZphaKASyMAmZ5erdn-fygdVrvtPScTL_zZmAAgCYYKqQ@mail.gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-13T18:50:34Z","receivedAt":"2018-08-13T18:50:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Adam <thomas@xteddy.org> writes:\n\n> On Sun, 12 Aug 2018 at 09:19, Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n>\n> Hi,\n>\n>> +       trace_performance_leave(\"cache_tree_update\");\n>\n> I would suggest trace_performance_leave() calls use __func__ instead.\n> That way, there's no ambiguity if the function name ever changes.\n\nPlease don't, unless you are certain that everybody has __func__ in\nthe first place.\n"},{"id":"355447","messageId":"f3403347-607b-b67c-297c-eeb9190a7de7@gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-4-pclouds@gmail.com","subject":"Re: [PATCH v4 3/5] unpack-trees: optimize walking same trees with cache-tree","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-13T18:58:24Z","receivedAt":"2018-08-13T18:58:29Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/12/2018 4:15 AM, Nguyễn Thái Ngọc Duy wrote:\n> In order to merge one or many trees with the index, unpack-trees code\n> walks multiple trees in parallel with the index and performs n-way\n> merge. If we find out at start of a directory that all trees are the\n> same (by comparing OID) and cache-tree happens to be available for\n> that directory as well, we could avoid walking the trees because we\n> already know what these trees contain: it's flattened in what's called\n> \"the index\".\n> \n> The upside is of course a lot less I/O since we can potentially skip\n> lots of trees (think subtrees). We also save CPU because we don't have\n> to inflate and apply the deltas. The downside is of course more\n> fragile code since the logic in some functions are now duplicated\n> elsewhere.\n> \n> \"checkout -\" with this patch on webkit.git (275k files):\n> \n>      baseline      new\n>    --------------------------------------------------------------------\n>      0.056651714   0.080394752 s:  read cache .git/index\n>      0.183101080   0.216010838 s:  preload index\n>      0.008584433   0.008534301 s:  refresh index\n>      0.633767589   0.251992198 s:   traverse_trees\n>      0.340265448   0.377031383 s:   check_updates\n>      0.381884638   0.372768105 s:   cache_tree_update\n>      1.401562947   1.045887251 s:  unpack_trees\n>      0.338687914   0.314983512 s:  write index, changed mask = 2e\n>      0.411927922   0.062572653 s:    traverse_trees\n>      0.000023335   0.000022544 s:    check_updates\n>      0.423697246   0.073795585 s:   unpack_trees\n>      0.423708360   0.073807557 s:  diff-index\n>      2.559524127   1.938191592 s: git command: git checkout -\n> \n> Another measurement from Ben's running \"git checkout\" with over 500k\n> trees (on the whole series):\n> \n>      baseline        new\n>    ----------------------------------------------------------------------\n>      0.535510167     0.556558733     s: read cache .git/index\n>      0.3057373       0.3147105       s: initialize name hash\n>      0.0184082       0.023558433     s: preload index\n>      0.086910967     0.089085967     s: refresh index\n>      7.889590767     2.191554433     s: unpack trees\n>      0.120760833     0.131941267     s: update worktree after a merge\n>      2.2583504       2.572663167     s: repair cache-tree\n>      0.8916137       0.959495233     s: write index, changed mask = 28\n>      3.405199233     0.2710663       s: unpack trees\n>      0.000999667     0.0021554       s: update worktree after a merge\n>      3.4063306       0.273318333     s: diff-index\n>      16.9524923      9.462943133     s: git command: git.exe checkout\n> \n> This command calls unpack_trees() twice, the first time on 2way merge\n> and the second 1way merge. In both times, \"unpack trees\" time is\n> reduced to one third. Overall time reduction is not that impressive of\n> course because index operations take a big chunk. And there's that\n> repair cache-tree line.\n> \n> PS. A note about cache-tree invalidation and the use of it in this\n> code.\n> \n> We do invalidate cache-tree in _source_ index when we add new entries\n> to the (temporary) \"result\" index. But we also use the cache-tree from\n> source index in this optimization. Does this mean we end up having no\n> cache-tree in the source index to activate this optimization?\n> \n> The answer is twisted: the order of finding a good cache-tree and\n> invalidating it matters. In this case we check for a good cache-tree\n> first in all_trees_same_as_cache_tree(), then we start to merge things\n> and potentially invalidate that same cache-tree in the process. Since\n> cache-tree invalidation happens after the optimization kicks in, we're\n> still good. But we may lose that cache-tree at the very first\n> call_unpack_fn() call in traverse_by_cache_tree().\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>   unpack-trees.c | 127 +++++++++++++++++++++++++++++++++++++++++++++++++\n>   1 file changed, 127 insertions(+)\n> \n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index b237eaa0f2..07456d0fb2 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -644,6 +644,102 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n>   \treturn name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n>   }\n>   \n> +static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n> +\t\t\t\t\tstruct name_entry *names,\n> +\t\t\t\t\tstruct traverse_info *info)\n> +{\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint i;\n> +\n> +\tif (!o->merge || dirmask != ((1 << n) - 1))\n> +\t\treturn 0;\n> +\n> +\tfor (i = 1; i < n; i++)\n> +\t\tif (!are_same_oid(names, names + i))\n> +\t\t\treturn 0;\n> +\n> +\treturn cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n> +}\n> +\n> +static int index_pos_by_traverse_info(struct name_entry *names,\n> +\t\t\t\t      struct traverse_info *info)\n> +{\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint len = traverse_path_len(info, names);\n> +\tchar *name = xmalloc(len + 1 /* slash */ + 1 /* NUL */);\n> +\tint pos;\n> +\n> +\tmake_traverse_path(name, info, names);\n> +\tname[len++] = '/';\n> +\tname[len] = '\\0';\n> +\tpos = index_name_pos(o->src_index, name, len);\n> +\tif (pos >= 0)\n> +\t\tBUG(\"This is a directory and should not exist in index\");\n> +\tpos = -pos - 1;\n> +\tif (!starts_with(o->src_index->cache[pos]->name, name) ||\n> +\t    (pos > 0 && starts_with(o->src_index->cache[pos-1]->name, name)))\n> +\t\tBUG(\"pos must point at the first entry in this directory\");\n> +\tfree(name);\n> +\treturn pos;\n> +}\n> +\n> +/*\n> + * Fast path if we detect that all trees are the same as cache-tree at this\n> + * path. We'll walk these trees recursively using cache-tree/index instead of\n> + * ODB since already know what these trees contain.\n> + */\n> +static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n> +\t\t\t\t  struct name_entry *names,\n> +\t\t\t\t  struct traverse_info *info)\n> +{\n> +\tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint i, d;\n> +\n> +\tif (!o->merge)\n> +\t\tBUG(\"We need cache-tree to do this optimization\");\n> +\n> +\t/*\n> +\t * Do what unpack_callback() and unpack_nondirectories() normally\n> +\t * do. But we walk all paths recursively in just one loop instead.\n\nThis comment threw me for a second.  Instead of walking the paths \nrecursively (like the old code path does), this code path actually does \nit in an iterative loop.  How about:\n\n\"But we walk all paths in an iterative loop instead.\"\n\n> +\t *\n> +\t * D/F conflicts and higher stage entries are not a concern\n> +\t * because cache-tree would be invalidated and we would never\n> +\t * get here in the first place.\n> +\t */\n> +\tfor (i = 0; i < nr_entries; i++) {\n> +\t\tstruct cache_entry *tree_ce;\n> +\t\tint len, rc;\n> +\n> +\t\tsrc[0] = o->src_index->cache[pos + i];\n> +\n> +\t\tlen = ce_namelen(src[0]);\n> +\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n> +\n> +\t\ttree_ce->ce_mode = src[0]->ce_mode;\n> +\t\ttree_ce->ce_flags = create_ce_flags(0);\n> +\t\ttree_ce->ce_namelen = len;\n> +\t\toidcpy(&tree_ce->oid, &src[0]->oid);\n> +\t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n> +\n> +\t\tfor (d = 1; d <= nr_names; d++)\n> +\t\t\tsrc[d] = tree_ce;\n> +\n> +\t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n\nI don't fully understand why this is still necessary since \"we detect \nthat all trees are the same as cache-tree at this path.\"  I do know \n(because I tried it :)) that if we don't actually call the unpack \nfunction the patch fails a bunch of tests so clearly something important \nis being missed.\n\n> +\t\tfree(tree_ce);\n> +\t\tif (rc < 0)\n> +\t\t\treturn rc;\n> +\n> +\t\tmark_ce_used(src[0], o);\n> +\t}\n> +\tif (o->debug_unpack)\n> +\t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n> +\t\t       nr_entries,\n> +\t\t       o->src_index->cache[pos]->name,\n> +\t\t       o->src_index->cache[pos + nr_entries - 1]->name);\n> +\treturn 0;\n> +}\n> +\n>   static int traverse_trees_recursive(int n, unsigned long dirmask,\n>   \t\t\t\t    unsigned long df_conflicts,\n>   \t\t\t\t    struct name_entry *names,\n> @@ -655,6 +751,27 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n>   \tvoid *buf[MAX_UNPACK_TREES];\n>   \tstruct traverse_info newinfo;\n>   \tstruct name_entry *p;\n> +\tint nr_entries;\n> +\n> +\tnr_entries = all_trees_same_as_cache_tree(n, dirmask, names, info);\n> +\tif (nr_entries > 0) {\n> +\t\tstruct unpack_trees_options *o = info->data;\n> +\t\tint pos = index_pos_by_traverse_info(names, info);\n> +\n> +\t\tif (!o->merge || df_conflicts)\n> +\t\t\tBUG(\"Wrong condition to get here buddy\");\n> +\n> +\t\t/*\n> +\t\t * All entries up to 'pos' must have been processed\n> +\t\t * (i.e. marked CE_UNPACKED) at this point. But to be safe,\n> +\t\t * save and restore cache_bottom anyway to not miss\n> +\t\t * unprocessed entries before 'pos'.\n> +\t\t */\n> +\t\tbottom = o->cache_bottom;\n> +\t\tret = traverse_by_cache_tree(pos, nr_entries, n, names, info);\n> +\t\to->cache_bottom = bottom;\n\nI agree with adding this back in - very low cost to provide some \nconsistency and additional safety.\n\n> +\t\treturn ret;\n> +\t}\n>   \n>   \tp = names;\n>   \twhile (!p->mode)\n> @@ -814,6 +931,11 @@ static struct cache_entry *create_ce_entry(const struct traverse_info *info, con\n>   \treturn ce;\n>   }\n>   \n> +/*\n> + * Note that traverse_by_cache_tree() duplicates some logic in this function\n> + * without actually calling it. If you change the logic here you may need to\n> + * check and change there as well.\n> + */\n>   static int unpack_nondirectories(int n, unsigned long mask,\n>   \t\t\t\t unsigned long dirmask,\n>   \t\t\t\t struct cache_entry **src,\n> @@ -998,6 +1120,11 @@ static void debug_unpack_callback(int n,\n>   \t\tdebug_name_entry(i, names + i);\n>   }\n>   \n> +/*\n> + * Note that traverse_by_cache_tree() duplicates some logic in this function\n> + * without actually calling it. If you change the logic here you may need to\n> + * check and change there as well.\n> + */\n>   static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n>   {\n>   \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n> \n"},{"id":"355450","messageId":"xmqqeff25bwb.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"20180812081551.27927-1-pclouds@gmail.com","subject":"Re: [PATCH v4 0/5] Speed up unpack_trees()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-13T19:01:24Z","receivedAt":"2018-08-13T19:01:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> v4 has a bunch of changes\n>\n> - 1/5 is a new one to show indented tracing. This way it's less\n>   misleading to read nested time measurements\n> - 3/5 now has the switch/restore cache_bottom logic. Junio suggested a\n>   check instead in his final note, but I think this is safer (yeah I'm\n>   scared too)\n> - the old 4/4 is dropped because\n>   - it assumes n-way logic\n>   - the visible time saving is not worth the tradeoff\n>   - Elijah gave me an idea to avoid add_index_entry() that I think\n>     does not have n-way logic assumptions and gives better saving.\n>     But it requires some more changes so I'm going to do it later\n> - 5/5 is also new and should help reduce cache_tree_update() cost.\n>   I wrote somewhere I was not going to work on this part, but it turns\n>   out just a couple lines, might as well do it now.\n\nThe last step feels a bit scary, but other than that I did not spot\nanything iffy in the series.  Nicely done.\n\nThanks.\n"},{"id":"355463","messageId":"20180813192526.GC10013@sigill.intra.peff.net","threadId":"48913","inReplyTo":"20180812081551.27927-3-pclouds@gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-08-13T19:25:26Z","receivedAt":"2018-08-13T19:25:30Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Aug 12, 2018 at 10:15:48AM +0200, Nguyễn Thái Ngọc Duy wrote:\n\n> We're going to optimize unpack_trees() a bit in the following\n> patches. Let's add some tracing to measure how long it takes before\n> and after. This is the baseline (\"git checkout -\" on webkit.git, 275k\n> files on worktree)\n> \n>     performance: 0.056651714 s:  read cache .git/index\n>     performance: 0.183101080 s:  preload index\n>     performance: 0.008584433 s:  refresh index\n>     performance: 0.633767589 s:   traverse_trees\n>     performance: 0.340265448 s:   check_updates\n>     performance: 0.381884638 s:   cache_tree_update\n>     performance: 1.401562947 s:  unpack_trees\n>     performance: 0.338687914 s:  write index, changed mask = 2e\n>     performance: 0.411927922 s:    traverse_trees\n>     performance: 0.000023335 s:    check_updates\n>     performance: 0.423697246 s:   unpack_trees\n>     performance: 0.423708360 s:  diff-index\n>     performance: 2.559524127 s: git command: git checkout -\n\nAm I the only one who feels a little funny about us sprinkling these\nperformance probes through the code base?\n\nOn Linux, \"perf\" already does a great job of this without having to\nmodify the source, and there are tools like:\n\n  http://www.brendangregg.com/FlameGraphs/cpuflamegraphs.html\n\nthat help make sense of the results.\n\nI know that's not going to help on Windows, but presumably there are\nhardware-counter based perf tools there, too.\n\nI can buy the argument that it's nice to have some form of profiling\nthat works everywhere, even if it's lowest-common-denominator. I just\nwonder if we could be investing effort into tooling around existing\nsolutions that will end up more powerful and flexible in the long run.\n\n-Peff\n"},{"id":"355465","messageId":"CAGZ79kYLpCZ=JNO7eun3mrhui+rymzr3i-VMLyFvKCxOaFNe-A@mail.gmail.com","threadId":"48913","inReplyTo":"20180813192526.GC10013@sigill.intra.peff.net","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-08-13T19:36:21Z","receivedAt":"2018-08-13T19:36:35Z","isPatch":true,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Mon, Aug 13, 2018 at 12:25 PM Jeff King <peff@peff.net> wrote:\n\n> I can buy the argument that it's nice to have some form of profiling\n> that works everywhere, even if it's lowest-common-denominator. I just\n> wonder if we could be investing effort into tooling around existing\n> solutions that will end up more powerful and flexible in the long run.\n\nThe issue AFAICT is that running perf is done by $YOU, the specialist,\nwhereas the performance framework put into place here can be\n\"turned on for the whole fleet\" and the ability to collect data from\nnon-specialists is there. (Note: At GitHub you do the serving side,\nwhereas Google, MS also control the shipped binary on the client\nside; asking a random engineer to run perf on their Git thing only\nhelps their special case and is unstructured; what helps is colorful\ndashboards aggregating all the results from all the people).\n\nSo it really is \"works everywhere,\" but not as you envisioned\n(cross platform vs more machines) ;-)\n"},{"id":"355467","messageId":"CACsJy8Cxp9+xiMB6C71Kr63EWyAni-K0ZwVBbpBjUieDbZ+6AA@mail.gmail.com","threadId":"48913","inReplyTo":"20180813192526.GC10013@sigill.intra.peff.net","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-13T19:52:41Z","receivedAt":"2018-08-13T19:53:09Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 13, 2018 at 9:25 PM Jeff King <peff@peff.net> wrote:\n> Am I the only one who feels a little funny about us sprinkling these\n> performance probes through the code base?\n>\n> On Linux, \"perf\" already does a great job of this without having to\n> modify the source, and there are tools like:\n>\n>   http://www.brendangregg.com/FlameGraphs/cpuflamegraphs.html\n>\n> that help make sense of the results.\n\nI don't think I have really fully mastered 'perf'. In this case for\nexample, I don't think the default event 'cycles' is the right one\nbecause we are hit hard by I/O as well. I think at least I now have an\nexcuse to try that famous flamegraph out ;-) but if you have time to\nrun a quick analysis of this unpack-trees with 'perf', I'd love to\nlearn a trick or two from you.\n\n> I know that's not going to help on Windows, but presumably there are\n> hardware-counter based perf tools there, too.\n>\n> I can buy the argument that it's nice to have some form of profiling\n> that works everywhere, even if it's lowest-common-denominator. I just\n> wonder if we could be investing effort into tooling around existing\n> solutions that will end up more powerful and flexible in the long run.\n\nI think part of this sprinkling is to highlight the performance\nsensitive spots in the code. And it would be helpful to ask a user to\nenable GIT_TRACE_PERFORMANCE to have a quick breakdown when something\nis reported slow. I don't care that much about other platforms to be\nhonest, but perf being largely restricted to root does prevent it from\nreplacing GIT_TRACE_PERFORMANCE in this case.\n-- \nDuy\n"},{"id":"355468","messageId":"33dff7ef-e68b-91aa-22b8-e35947e4f7b8@gmail.com","threadId":"48913","inReplyTo":"CAGZ79kYLpCZ=JNO7eun3mrhui+rymzr3i-VMLyFvKCxOaFNe-A@mail.gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-13T20:11:14Z","receivedAt":"2018-08-13T20:11:18Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/13/2018 3:36 PM, Stefan Beller wrote:\n> On Mon, Aug 13, 2018 at 12:25 PM Jeff King <peff@peff.net> wrote:\n> \n>> I can buy the argument that it's nice to have some form of profiling\n>> that works everywhere, even if it's lowest-common-denominator. I just\n>> wonder if we could be investing effort into tooling around existing\n>> solutions that will end up more powerful and flexible in the long run.\n> \n> The issue AFAICT is that running perf is done by $YOU, the specialist,\n> whereas the performance framework put into place here can be\n> \"turned on for the whole fleet\" and the ability to collect data from\n> non-specialists is there. (Note: At GitHub you do the serving side,\n> whereas Google, MS also control the shipped binary on the client\n> side; asking a random engineer to run perf on their Git thing only\n> helps their special case and is unstructured; what helps is colorful\n> dashboards aggregating all the results from all the people).\n> \n> So it really is \"works everywhere,\" but not as you envisioned\n> (cross platform vs more machines) ;-)\n> \n\nI currently use GIT_TRACE_PERFORMANCE primarily to communicate \nperformance measurements on the mailing list.  While it is convenient \noccasionally to run with it turned on locally, primarily it gives me a \ncommon reference when communicating with others on the list about \nperformance.\n\nWe have several excellent profiling tools available on Windows \n(perfview, VS, wpa, etc) so for any detailed investigations, I use \nthose.  They obviously don't require any instrumenting in the code.\n\nFor our internal end user performance data, we'll use structured logging \nand our custom telemetry solution rather than the GIT_TRACE_PERFORMANCE \nmechanism.  We never ask end users to turn on GIT_TRACE_PERFORMANCE.  If \nwe need more than what we can gather using telemetry, we ask them to \ncapture a perfview along with other diagnostic data and send it to us \nfor evaluation.\n\n"},{"id":"355490","messageId":"20180813214702.GA16006@sigill.intra.peff.net","threadId":"48913","inReplyTo":"CACsJy8Cxp9+xiMB6C71Kr63EWyAni-K0ZwVBbpBjUieDbZ+6AA@mail.gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-08-13T21:47:03Z","receivedAt":"2018-08-13T21:47:06Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Aug 13, 2018 at 09:52:41PM +0200, Duy Nguyen wrote:\n\n> I don't think I have really fully mastered 'perf'. In this case for\n> example, I don't think the default event 'cycles' is the right one\n> because we are hit hard by I/O as well. I think at least I now have an\n> excuse to try that famous flamegraph out ;-) but if you have time to\n> run a quick analysis of this unpack-trees with 'perf', I'd love to\n> learn a trick or two from you.\n\nTo be honest, I don't feel like I know how to use perf either. ;) But\nI'll try to contribute what I know.\n\nUsually I'd just use perf to get a callgraph with hot-spots, like:\n\n  perf record -g git ...\n  perf report --call-graph=fractal,0.05,caller\n\nBut that's not going to show you absolute times, which makes it lousy\nfor comparing run-to-run (if you speed something up, its percentage gets\nsmaller, but it's hard to tell _how much_ you've sped it up). And as you\nnote, it's measuring CPU cycles, not wall-clock.\n\nTo get output most similar to what you've shown, I think you'd define\nsome probes at functions of interest:\n\n  for i in unpack_trees cache_tree_update; do\n    # Cover both function entrance and return.\n    perf probe -x $(which git) $i\n    perf probe -x $(which git) ${i}%return\n  done\n\nand then record a run looking for those events:\n\n  perf record -e 'probe_git:*' git ...\n\nand then dump the result:\n\n  perf script -F time,event\n\nwhich gives you the times for each event. If you want elapsed times, you\nhave to compute them yourself:\n\n  perf script -F time,event |\n  perl -ne '\n    /([0-9.]+):\\s+probe_git:(.*):/ or die \"confusing: $_\";\n    my ($t, $func) = ($1, $2);\n    if ($func =~ s/__return$//) {\n      my $start = pop @stack;\n      printf \"%0.9f\", $t - $start;\n      print \" s: \";\n      print \"  \" for (0..@stack-1);\n      print $func, \"\\n\";\n    } else {\n      push @stack, $t;\n    }\n  '\n\nwhich gives a similar inverted-graph elapsed-time output that your trace\noutput does. One annoying downside is that you have to be root to create\nor use the dynamic probes. I don't know if there's an easy way around\nthat. Or if there's a perf command which already handles this kind of\nelapsed stuff (there's a \"perf trace\" which seems really close, but I\ncouldn't convince it to look at elapsed time for non-syscalls).\n\n> > I can buy the argument that it's nice to have some form of profiling\n> > that works everywhere, even if it's lowest-common-denominator. I just\n> > wonder if we could be investing effort into tooling around existing\n> > solutions that will end up more powerful and flexible in the long run.\n> \n> I think part of this sprinkling is to highlight the performance\n> sensitive spots in the code. And it would be helpful to ask a user to\n> enable GIT_TRACE_PERFORMANCE to have a quick breakdown when something\n> is reported slow. I don't care that much about other platforms to be\n> honest, but perf being largely restricted to root does prevent it from\n> replacing GIT_TRACE_PERFORMANCE in this case.\n\nYeah, this line of reasoning (which is similar to what Stefan said) is\ncompelling to me. GIT_TRACE_* is _most_ useful when we can ask ordinary\nusers to give us output. Even if we scripted the complexity I showed\nabove, it's not guaranteed that perf is even available or that the user\nhas permissions to use it.\n\n-Peff\n"},{"id":"355494","messageId":"xmqqk1ot3n4h.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"20180813192526.GC10013@sigill.intra.peff.net","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-13T22:41:50Z","receivedAt":"2018-08-13T22:41:55Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> I can buy the argument that it's nice to have some form of profiling\n> that works everywhere, even if it's lowest-common-denominator. I just\n> wonder if we could be investing effort into tooling around existing\n> solutions that will end up more powerful and flexible in the long run.\n\nAnother thing I noticed is that the codepaths we would find\ninteresting to annotate with trace_performance_* stuff often\noverlaps with the \"slog\" thing.  If the latter aims to eventually\nreplace GIT_TRACE (and if not, I suspect there is not much point\nadding it in the first place), perhaps we can extend it to also\ncover the need of these trace_performance_* calls, so that we do not\nhave to carry three different tracing mechanisms.\n\n"},{"id":"355580","messageId":"90d1bbf7-91a3-74ac-de65-1eb8405dc1f7@jeffhostetler.com","threadId":"48913","inReplyTo":"xmqqk1ot3n4h.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Jeff Hostetler","fromEmail":"git@jeffhostetler.com","sentAt":"2018-08-14T18:19:31Z","receivedAt":"2018-08-14T18:19:35Z","isPatch":true,"sender":{"key":"git@jeffhostetler.com","avatar":null},"body":"\n\nOn 8/13/2018 6:41 PM, Junio C Hamano wrote:\n> Jeff King <peff@peff.net> writes:\n> \n>> I can buy the argument that it's nice to have some form of profiling\n>> that works everywhere, even if it's lowest-common-denominator. I just\n>> wonder if we could be investing effort into tooling around existing\n>> solutions that will end up more powerful and flexible in the long run.\n> \n> Another thing I noticed is that the codepaths we would find\n> interesting to annotate with trace_performance_* stuff often\n> overlaps with the \"slog\" thing.  If the latter aims to eventually\n> replace GIT_TRACE (and if not, I suspect there is not much point\n> adding it in the first place), perhaps we can extend it to also\n> cover the need of these trace_performance_* calls, so that we do not\n> have to carry three different tracing mechanisms.\n> \n\nI'm looking at adding code to my SLOG (better name suggestions welcome)\npatch series to eventually replace the existing git_trace facility.\nAnd I would like to have a set of nested messages like Duy has proposed\nbe a part of that.\n\nIn an independent effort I've found the nested messages being very\nhelpful in certain contexts.  They are not a replacement for the\nvarious platform tools, like PerfView and friends as discussed earlier\non this thread, but then again I can ask a customer to turn a knob and\nrun it again and send me the output and hopefully get a rough idea of\nthe problem -- without having them install a bunch of perf tools.\n\nJeff\n\n"},{"id":"355586","messageId":"CACsJy8DQmOCD2a5QFUiyPuoPZLq-QEejLhWACKpsJLvK5ERAMg@mail.gmail.com","threadId":"48913","inReplyTo":"90d1bbf7-91a3-74ac-de65-1eb8405dc1f7@jeffhostetler.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-14T18:32:11Z","receivedAt":"2018-08-14T18:32:40Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Aug 14, 2018 at 8:19 PM Jeff Hostetler <git@jeffhostetler.com> wrote:\n> I'm looking at adding code to my SLOG (better name suggestions welcome)\n> patch series to eventually replace the existing git_trace facility.\n\nComplement maybe. Replace, please no. I'd rather not stare at json messages.\n-- \nDuy\n"},{"id":"355588","messageId":"CAGZ79kZwVpCBMkBKuYpwZFgAN50wZub_fyzWrAsE=ksuc-aCgQ@mail.gmail.com","threadId":"48913","inReplyTo":"CACsJy8DQmOCD2a5QFUiyPuoPZLq-QEejLhWACKpsJLvK5ERAMg@mail.gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-08-14T18:44:09Z","receivedAt":"2018-08-14T18:44:23Z","isPatch":true,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Tue, Aug 14, 2018 at 11:32 AM Duy Nguyen <pclouds@gmail.com> wrote:\n>\n> On Tue, Aug 14, 2018 at 8:19 PM Jeff Hostetler <git@jeffhostetler.com> wrote:\n> > I'm looking at adding code to my SLOG (better name suggestions welcome)\n> > patch series to eventually replace the existing git_trace facility.\n>\n> Complement maybe. Replace, please no. I'd rather not stare at json messages.\n\nFrom the sidelines: We'd only need one logging infrastructure in place, as the\nformatting would be done as a later step? For local operations we'd certainly\nfind better formatting than json, and we figured that we might end up desiring\nProtocolBuffers[1] instead of JSon, so if it would be easy to change\nthe output of\nthe structured logging easily that would be great.\n\nBut AFAICT these series are all about putting the sampling points into the\ncode base, so formatting would be orthogonal to it?\n\nStefan\n\n[1] https://developers.google.com/protocol-buffers/\n"},{"id":"355590","messageId":"CACsJy8CTNeR8Bchj37yNL+mWp1Y5rhD6QV2Gf06CPLHVXd8TDQ@mail.gmail.com","threadId":"48913","inReplyTo":"CAGZ79kZwVpCBMkBKuYpwZFgAN50wZub_fyzWrAsE=ksuc-aCgQ@mail.gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-14T18:51:41Z","receivedAt":"2018-08-14T18:52:10Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Aug 14, 2018 at 8:44 PM Stefan Beller <sbeller@google.com> wrote:\n>\n> On Tue, Aug 14, 2018 at 11:32 AM Duy Nguyen <pclouds@gmail.com> wrote:\n> >\n> > On Tue, Aug 14, 2018 at 8:19 PM Jeff Hostetler <git@jeffhostetler.com> wrote:\n> > > I'm looking at adding code to my SLOG (better name suggestions welcome)\n> > > patch series to eventually replace the existing git_trace facility.\n> >\n> > Complement maybe. Replace, please no. I'd rather not stare at json messages.\n>\n> From the sidelines: We'd only need one logging infrastructure in place, as the\n> formatting would be done as a later step? For local operations we'd certainly\n> find better formatting than json, and we figured that we might end up desiring\n> ProtocolBuffers[1] instead of JSon, so if it would be easy to change\n> the output of\n> the structured logging easily that would be great.\n\nThese trace messages are made for human consumption. Granted\noccasionally we need some processing but I find one liners mostly\nsuffice. Now we turn these into something made for machines, turning\npeople to second citizens. I've read these messages reformatted for\nhuman, it's usually too verbose even if it's reformatted.\n\n> But AFAICT these series are all about putting the sampling points into the\n> code base, so formatting would be orthogonal to it?\n\nIt's not just sampling points. There's things like index id being\nshown in the message for example. I prefer to keep free style format\nto help me read. There's also things like indentation I do here to\nhelp me read. Granted you could do all that with scripts and stuff,\nbut will we pass around in mail  dumps of json messages to be decoded\nlocally?\n\n> Stefan\n>\n> [1] https://developers.google.com/protocol-buffers/\n\n\n\n-- \nDuy\n"},{"id":"355598","messageId":"a0d4142a-795d-ddd3-0c9b-b54131591b7c@gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-1-pclouds@gmail.com","subject":"Re: [PATCH v4 0/5] Speed up unpack_trees()","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-14T19:19:42Z","receivedAt":"2018-08-14T19:19:47Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/12/2018 4:15 AM, Nguyễn Thái Ngọc Duy wrote:\n> v4 has a bunch of changes\n> \n> - 1/5 is a new one to show indented tracing. This way it's less\n>    misleading to read nested time measurements\n> - 3/5 now has the switch/restore cache_bottom logic. Junio suggested a\n>    check instead in his final note, but I think this is safer (yeah I'm\n>    scared too)\n> - the old 4/4 is dropped because\n>    - it assumes n-way logic\n>    - the visible time saving is not worth the tradeoff\n>    - Elijah gave me an idea to avoid add_index_entry() that I think\n>      does not have n-way logic assumptions and gives better saving.\n>      But it requires some more changes so I'm going to do it later\n> - 5/5 is also new and should help reduce cache_tree_update() cost.\n>    I wrote somewhere I was not going to work on this part, but it turns\n>    out just a couple lines, might as well do it now.\n> \n> Interdiff\n> \n\nI've now had a chance to run the git tests, as well as our own unit and \nfunctional tests with this patch series and all passed.\n\nI reviewed the tests in t0090-cache-tree.h and verified that there are \ntests that validate the cache tree is correct after doing a checkout and \nmerge (both of which exercise the new cache tree optimization in patch 5).\n\nI've also run our perf test suite and the results are outstanding:\n\nCheckout saves 51% on average\nMerge saves 44%\nPull saves 30%\nRebase saves 26%\n\nFor perspective, that means these commands are going from ~20 seconds to \n~10 seconds.\n\nI don't feel that any of my comments to the individual patches deserve a \nre-roll.  Given the ongoing discussion about the additional tracing - \nI'm happy to leave out the first 2 patches so that the rest can go in \nsooner rather than later.\n\nLooks good!\n\n> diff --git a/cache-tree.c b/cache-tree.c\n> index 0dbe10fc85..105f13806f 100644\n> --- a/cache-tree.c\n> +++ b/cache-tree.c\n> @@ -426,7 +426,6 @@ static int update_one(struct cache_tree *it,\n>   \n>   int cache_tree_update(struct index_state *istate, int flags)\n>   {\n> -\tuint64_t start = getnanotime();\n>   \tstruct cache_tree *it = istate->cache_tree;\n>   \tstruct cache_entry **cache = istate->cache;\n>   \tint entries = istate->cache_nr;\n> @@ -434,11 +433,12 @@ int cache_tree_update(struct index_state *istate, int flags)\n>   \n>   \tif (i)\n>   \t\treturn i;\n> +\ttrace_performance_enter();\n>   \ti = update_one(it, cache, entries, \"\", 0, &skip, flags);\n> +\ttrace_performance_leave(\"cache_tree_update\");\n>   \tif (i < 0)\n>   \t\treturn i;\n>   \tistate->cache_changed |= CACHE_TREE_CHANGED;\n> -\ttrace_performance_since(start, \"repair cache-tree\");\n>   \treturn 0;\n>   }\n>   \n> diff --git a/cache.h b/cache.h\n> index e6f7ee4b64..8b447652a7 100644\n> --- a/cache.h\n> +++ b/cache.h\n> @@ -673,7 +673,6 @@ extern int index_name_pos(const struct index_state *, const char *name, int name\n>   #define ADD_CACHE_JUST_APPEND 8\t\t/* Append only; tree.c::read_tree() */\n>   #define ADD_CACHE_NEW_ONLY 16\t\t/* Do not replace existing ones */\n>   #define ADD_CACHE_KEEP_CACHE_TREE 32\t/* Do not invalidate cache-tree */\n> -#define ADD_CACHE_SKIP_VERIFY_PATH 64\t/* Do not verify path */\n>   extern int add_index_entry(struct index_state *, struct cache_entry *ce, int option);\n>   extern void rename_index_entry_at(struct index_state *, int pos, const char *new_name);\n>   \n> diff --git a/diff-lib.c b/diff-lib.c\n> index a9f38eb5a3..1ffa22c882 100644\n> --- a/diff-lib.c\n> +++ b/diff-lib.c\n> @@ -518,8 +518,8 @@ static int diff_cache(struct rev_info *revs,\n>   int run_diff_index(struct rev_info *revs, int cached)\n>   {\n>   \tstruct object_array_entry *ent;\n> -\tuint64_t start = getnanotime();\n>   \n> +\ttrace_performance_enter();\n>   \tent = revs->pending.objects;\n>   \tif (diff_cache(revs, &ent->item->oid, ent->name, cached))\n>   \t\texit(128);\n> @@ -528,7 +528,7 @@ int run_diff_index(struct rev_info *revs, int cached)\n>   \tdiffcore_fix_diff_index(&revs->diffopt);\n>   \tdiffcore_std(&revs->diffopt);\n>   \tdiff_flush(&revs->diffopt);\n> -\ttrace_performance_since(start, \"diff-index\");\n> +\ttrace_performance_leave(\"diff-index\");\n>   \treturn 0;\n>   }\n>   \n> diff --git a/dir.c b/dir.c\n> index 21e6f2520a..c5e9fc8cea 100644\n> --- a/dir.c\n> +++ b/dir.c\n> @@ -2263,11 +2263,11 @@ int read_directory(struct dir_struct *dir, struct index_state *istate,\n>   \t\t   const char *path, int len, const struct pathspec *pathspec)\n>   {\n>   \tstruct untracked_cache_dir *untracked;\n> -\tuint64_t start = getnanotime();\n>   \n>   \tif (has_symlink_leading_path(path, len))\n>   \t\treturn dir->nr;\n>   \n> +\ttrace_performance_enter();\n>   \tuntracked = validate_untracked_cache(dir, len, pathspec);\n>   \tif (!untracked)\n>   \t\t/*\n> @@ -2302,7 +2302,7 @@ int read_directory(struct dir_struct *dir, struct index_state *istate,\n>   \t\tdir->nr = i;\n>   \t}\n>   \n> -\ttrace_performance_since(start, \"read directory %.*s\", len, path);\n> +\ttrace_performance_leave(\"read directory %.*s\", len, path);\n>   \tif (dir->untracked) {\n>   \t\tstatic int force_untracked_cache = -1;\n>   \t\tstatic struct trace_key trace_untracked_stats = TRACE_KEY_INIT(UNTRACKED_STATS);\n> diff --git a/name-hash.c b/name-hash.c\n> index 163849831c..1fcda73cb3 100644\n> --- a/name-hash.c\n> +++ b/name-hash.c\n> @@ -578,10 +578,10 @@ static void threaded_lazy_init_name_hash(\n>   \n>   static void lazy_init_name_hash(struct index_state *istate)\n>   {\n> -\tuint64_t start = getnanotime();\n>   \n>   \tif (istate->name_hash_initialized)\n>   \t\treturn;\n> +\ttrace_performance_enter();\n>   \thashmap_init(&istate->name_hash, cache_entry_cmp, NULL, istate->cache_nr);\n>   \thashmap_init(&istate->dir_hash, dir_entry_cmp, NULL, istate->cache_nr);\n>   \n> @@ -602,7 +602,7 @@ static void lazy_init_name_hash(struct index_state *istate)\n>   \t}\n>   \n>   \tistate->name_hash_initialized = 1;\n> -\ttrace_performance_since(start, \"initialize name hash\");\n> +\ttrace_performance_leave(\"initialize name hash\");\n>   }\n>   \n>   /*\n> diff --git a/preload-index.c b/preload-index.c\n> index 4d08d44874..d7f7919ba2 100644\n> --- a/preload-index.c\n> +++ b/preload-index.c\n> @@ -78,7 +78,6 @@ static void preload_index(struct index_state *index,\n>   {\n>   \tint threads, i, work, offset;\n>   \tstruct thread_data data[MAX_PARALLEL];\n> -\tuint64_t start = getnanotime();\n>   \n>   \tif (!core_preload_index)\n>   \t\treturn;\n> @@ -88,6 +87,7 @@ static void preload_index(struct index_state *index,\n>   \t\tthreads = 2;\n>   \tif (threads < 2)\n>   \t\treturn;\n> +\ttrace_performance_enter();\n>   \tif (threads > MAX_PARALLEL)\n>   \t\tthreads = MAX_PARALLEL;\n>   \toffset = 0;\n> @@ -109,7 +109,7 @@ static void preload_index(struct index_state *index,\n>   \t\tif (pthread_join(p->pthread, NULL))\n>   \t\t\tdie(\"unable to join threaded lstat\");\n>   \t}\n> -\ttrace_performance_since(start, \"preload index\");\n> +\ttrace_performance_leave(\"preload index\");\n>   }\n>   #endif\n>   \n> diff --git a/read-cache.c b/read-cache.c\n> index b0b5df5de7..2b5646ef26 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1170,7 +1170,6 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n>   \tint ok_to_add = option & ADD_CACHE_OK_TO_ADD;\n>   \tint ok_to_replace = option & ADD_CACHE_OK_TO_REPLACE;\n>   \tint skip_df_check = option & ADD_CACHE_SKIP_DFCHECK;\n> -\tint skip_verify_path = option & ADD_CACHE_SKIP_VERIFY_PATH;\n>   \tint new_only = option & ADD_CACHE_NEW_ONLY;\n>   \n>   \tif (!(option & ADD_CACHE_KEEP_CACHE_TREE))\n> @@ -1211,7 +1210,7 @@ static int add_index_entry_with_check(struct index_state *istate, struct cache_e\n>   \n>   \tif (!ok_to_add)\n>   \t\treturn -1;\n> -\tif (!skip_verify_path && !verify_path(ce->name, ce->ce_mode))\n> +\tif (!verify_path(ce->name, ce->ce_mode))\n>   \t\treturn error(\"Invalid path '%s'\", ce->name);\n>   \n>   \tif (!skip_df_check &&\n> @@ -1400,8 +1399,8 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n>   \tconst char *typechange_fmt;\n>   \tconst char *added_fmt;\n>   \tconst char *unmerged_fmt;\n> -\tuint64_t start = getnanotime();\n>   \n> +\ttrace_performance_enter();\n>   \tmodified_fmt = (in_porcelain ? \"M\\t%s\\n\" : \"%s: needs update\\n\");\n>   \tdeleted_fmt = (in_porcelain ? \"D\\t%s\\n\" : \"%s: needs update\\n\");\n>   \ttypechange_fmt = (in_porcelain ? \"T\\t%s\\n\" : \"%s needs update\\n\");\n> @@ -1471,7 +1470,7 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n>   \n>   \t\treplace_index_entry(istate, i, new_entry);\n>   \t}\n> -\ttrace_performance_since(start, \"refresh index\");\n> +\ttrace_performance_leave(\"refresh index\");\n>   \treturn has_errors;\n>   }\n>   \n> @@ -1902,7 +1901,6 @@ static void freshen_shared_index(const char *shared_index, int warn)\n>   int read_index_from(struct index_state *istate, const char *path,\n>   \t\t    const char *gitdir)\n>   {\n> -\tuint64_t start = getnanotime();\n>   \tstruct split_index *split_index;\n>   \tint ret;\n>   \tchar *base_oid_hex;\n> @@ -1912,8 +1910,9 @@ int read_index_from(struct index_state *istate, const char *path,\n>   \tif (istate->initialized)\n>   \t\treturn istate->cache_nr;\n>   \n> +\ttrace_performance_enter();\n>   \tret = do_read_index(istate, path, 0);\n> -\ttrace_performance_since(start, \"read cache %s\", path);\n> +\ttrace_performance_leave(\"read cache %s\", path);\n>   \n>   \tsplit_index = istate->split_index;\n>   \tif (!split_index || is_null_oid(&split_index->base_oid)) {\n> @@ -1921,6 +1920,7 @@ int read_index_from(struct index_state *istate, const char *path,\n>   \t\treturn ret;\n>   \t}\n>   \n> +\ttrace_performance_enter();\n>   \tif (split_index->base)\n>   \t\tdiscard_index(split_index->base);\n>   \telse\n> @@ -1937,8 +1937,8 @@ int read_index_from(struct index_state *istate, const char *path,\n>   \tfreshen_shared_index(base_path, 0);\n>   \tmerge_base_index(istate);\n>   \tpost_read_index_from(istate);\n> -\ttrace_performance_since(start, \"read cache %s\", base_path);\n>   \tfree(base_path);\n> +\ttrace_performance_leave(\"read cache %s\", base_path);\n>   \treturn ret;\n>   }\n>   \n> @@ -2763,4 +2763,6 @@ void move_index_extensions(struct index_state *dst, struct index_state *src)\n>   {\n>   \tdst->untracked = src->untracked;\n>   \tsrc->untracked = NULL;\n> +\tdst->cache_tree = src->cache_tree;\n> +\tsrc->cache_tree = NULL;\n>   }\n> diff --git a/trace.c b/trace.c\n> index fc623e91fd..fa4a2e7120 100644\n> --- a/trace.c\n> +++ b/trace.c\n> @@ -176,10 +176,30 @@ void trace_strbuf_fl(const char *file, int line, struct trace_key *key,\n>   \tstrbuf_release(&buf);\n>   }\n>   \n> +static uint64_t perf_start_times[10];\n> +static int perf_indent;\n> +\n> +uint64_t trace_performance_enter(void)\n> +{\n> +\tuint64_t now;\n> +\n> +\tif (!trace_want(&trace_perf_key))\n> +\t\treturn 0;\n> +\n> +\tnow = getnanotime();\n> +\tperf_start_times[perf_indent] = now;\n> +\tif (perf_indent + 1 < ARRAY_SIZE(perf_start_times))\n> +\t\tperf_indent++;\n> +\telse\n> +\t\tBUG(\"Too deep indentation\");\n> +\treturn now;\n> +}\n> +\n>   static void trace_performance_vprintf_fl(const char *file, int line,\n>   \t\t\t\t\t uint64_t nanos, const char *format,\n>   \t\t\t\t\t va_list ap)\n>   {\n> +\tstatic const char space[] = \"          \";\n>   \tstruct strbuf buf = STRBUF_INIT;\n>   \n>   \tif (!prepare_trace_line(file, line, &trace_perf_key, &buf))\n> @@ -188,7 +208,10 @@ static void trace_performance_vprintf_fl(const char *file, int line,\n>   \tstrbuf_addf(&buf, \"performance: %.9f s\", (double) nanos / 1000000000);\n>   \n>   \tif (format && *format) {\n> -\t\tstrbuf_addstr(&buf, \": \");\n> +\t\tif (perf_indent >= strlen(space))\n> +\t\t\tBUG(\"Too deep indentation\");\n> +\n> +\t\tstrbuf_addf(&buf, \":%.*s \", perf_indent, space);\n>   \t\tstrbuf_vaddf(&buf, format, ap);\n>   \t}\n>   \n> @@ -244,6 +267,24 @@ void trace_performance_since(uint64_t start, const char *format, ...)\n>   \tva_end(ap);\n>   }\n>   \n> +void trace_performance_leave(const char *format, ...)\n> +{\n> +\tva_list ap;\n> +\tuint64_t since;\n> +\n> +\tif (perf_indent)\n> +\t\tperf_indent--;\n> +\n> +\tif (!format) /* Allow callers to leave without tracing anything */\n> +\t\treturn;\n> +\n> +\tsince = perf_start_times[perf_indent];\n> +\tva_start(ap, format);\n> +\ttrace_performance_vprintf_fl(NULL, 0, getnanotime() - since,\n> +\t\t\t\t     format, ap);\n> +\tva_end(ap);\n> +}\n> +\n>   #else\n>   \n>   void trace_printf_key_fl(const char *file, int line, struct trace_key *key,\n> @@ -273,6 +314,24 @@ void trace_performance_fl(const char *file, int line, uint64_t nanos,\n>   \tva_end(ap);\n>   }\n>   \n> +void trace_performance_leave_fl(const char *file, int line,\n> +\t\t\t\tuint64_t nanos, const char *format, ...)\n> +{\n> +\tva_list ap;\n> +\tuint64_t since;\n> +\n> +\tif (perf_indent)\n> +\t\tperf_indent--;\n> +\n> +\tif (!format) /* Allow callers to leave without tracing anything */\n> +\t\treturn;\n> +\n> +\tsince = perf_start_times[perf_indent];\n> +\tva_start(ap, format);\n> +\ttrace_performance_vprintf_fl(file, line, nanos - since, format, ap);\n> +\tva_end(ap);\n> +}\n> +\n>   #endif /* HAVE_VARIADIC_MACROS */\n>   \n>   \n> @@ -411,13 +470,11 @@ uint64_t getnanotime(void)\n>   \t}\n>   }\n>   \n> -static uint64_t command_start_time;\n>   static struct strbuf command_line = STRBUF_INIT;\n>   \n>   static void print_command_performance_atexit(void)\n>   {\n> -\ttrace_performance_since(command_start_time, \"git command:%s\",\n> -\t\t\t\tcommand_line.buf);\n> +\ttrace_performance_leave(\"git command:%s\", command_line.buf);\n>   }\n>   \n>   void trace_command_performance(const char **argv)\n> @@ -425,10 +482,10 @@ void trace_command_performance(const char **argv)\n>   \tif (!trace_want(&trace_perf_key))\n>   \t\treturn;\n>   \n> -\tif (!command_start_time)\n> +\tif (!command_line.len)\n>   \t\tatexit(print_command_performance_atexit);\n>   \n>   \tstrbuf_reset(&command_line);\n>   \tsq_quote_argv_pretty(&command_line, argv);\n> -\tcommand_start_time = getnanotime();\n> +\ttrace_performance_enter();\n>   }\n> diff --git a/trace.h b/trace.h\n> index 2b6a1bc17c..171b256d26 100644\n> --- a/trace.h\n> +++ b/trace.h\n> @@ -23,6 +23,7 @@ extern void trace_disable(struct trace_key *key);\n>   extern uint64_t getnanotime(void);\n>   extern void trace_command_performance(const char **argv);\n>   extern void trace_verbatim(struct trace_key *key, const void *buf, unsigned len);\n> +uint64_t trace_performance_enter(void);\n>   \n>   #ifndef HAVE_VARIADIC_MACROS\n>   \n> @@ -45,6 +46,9 @@ extern void trace_performance(uint64_t nanos, const char *format, ...);\n>   __attribute__((format (printf, 2, 3)))\n>   extern void trace_performance_since(uint64_t start, const char *format, ...);\n>   \n> +__attribute__((format (printf, 1, 2)))\n> +void trace_performance_leave(const char *format, ...);\n> +\n>   #else\n>   \n>   /*\n> @@ -118,6 +122,14 @@ extern void trace_performance_since(uint64_t start, const char *format, ...);\n>   \t\t\t\t\t     __VA_ARGS__);\t\t    \\\n>   \t} while (0)\n>   \n> +#define trace_performance_leave(...)\t\t\t\t\t    \\\n> +\tdo {\t\t\t\t\t\t\t\t    \\\n> +\t\tif (trace_pass_fl(&trace_perf_key))\t\t\t    \\\n> +\t\t\ttrace_performance_leave_fl(TRACE_CONTEXT, __LINE__, \\\n> +\t\t\t\t\t\t   getnanotime(),\t    \\\n> +\t\t\t\t\t\t   __VA_ARGS__);\t    \\\n> +\t} while (0)\n> +\n>   /* backend functions, use non-*fl macros instead */\n>   __attribute__((format (printf, 4, 5)))\n>   extern void trace_printf_key_fl(const char *file, int line, struct trace_key *key,\n> @@ -130,6 +142,9 @@ extern void trace_strbuf_fl(const char *file, int line, struct trace_key *key,\n>   __attribute__((format (printf, 4, 5)))\n>   extern void trace_performance_fl(const char *file, int line,\n>   \t\t\t\t uint64_t nanos, const char *fmt, ...);\n> +__attribute__((format (printf, 4, 5)))\n> +extern void trace_performance_leave_fl(const char *file, int line,\n> +\t\t\t\t       uint64_t nanos, const char *fmt, ...);\n>   static inline int trace_pass_fl(struct trace_key *key)\n>   {\n>   \treturn key->fd || !key->initialized;\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index 1438ee1555..d822662c75 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -201,7 +201,6 @@ static int do_add_entry(struct unpack_trees_options *o, struct cache_entry *ce,\n>   \n>   \tce->ce_flags = (ce->ce_flags & ~clear) | set;\n>   \treturn add_index_entry(&o->result, ce,\n> -\t\t\t       o->extra_add_index_flags |\n>   \t\t\t       ADD_CACHE_OK_TO_ADD | ADD_CACHE_OK_TO_REPLACE);\n>   }\n>   \n> @@ -353,9 +352,9 @@ static int check_updates(struct unpack_trees_options *o)\n>   \tstruct progress *progress = NULL;\n>   \tstruct index_state *index = &o->result;\n>   \tstruct checkout state = CHECKOUT_INIT;\n> -\tuint64_t start = getnanotime();\n>   \tint i;\n>   \n> +\ttrace_performance_enter();\n>   \tstate.force = 1;\n>   \tstate.quiet = 1;\n>   \tstate.refresh_cache = 1;\n> @@ -425,7 +424,7 @@ static int check_updates(struct unpack_trees_options *o)\n>   \terrs |= finish_delayed_checkout(&state);\n>   \tif (o->update)\n>   \t\tgit_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n> -\ttrace_performance_since(start, \"update worktree after a merge\");\n> +\ttrace_performance_leave(\"check_updates\");\n>   \treturn errs != 0;\n>   }\n>   \n> @@ -702,31 +701,13 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n>   \tif (!o->merge)\n>   \t\tBUG(\"We need cache-tree to do this optimization\");\n>   \n> -\t/*\n> -\t * Try to keep add_index_entry() as fast as possible since\n> -\t * we're going to do a lot of them.\n> -\t *\n> -\t * Skipping verify_path() should totally be safe because these\n> -\t * paths are from the source index, which must have been\n> -\t * verified.\n> -\t *\n> -\t * Skipping D/F and cache-tree validation checks is trickier\n> -\t * because it assumes what n-merge code would do when all\n> -\t * trees and the index are the same. We probably could just\n> -\t * optimize those code instead (e.g. we don't invalidate that\n> -\t * many cache-tree, but the searching for them is very\n> -\t * expensive).\n> -\t */\n> -\to->extra_add_index_flags = ADD_CACHE_SKIP_DFCHECK;\n> -\to->extra_add_index_flags |= ADD_CACHE_SKIP_VERIFY_PATH;\n> -\n>   \t/*\n>   \t * Do what unpack_callback() and unpack_nondirectories() normally\n>   \t * do. But we walk all paths recursively in just one loop instead.\n>   \t *\n> -\t * D/F conflicts and staged entries are not a concern because\n> -\t * cache-tree would be invalidated and we would never get here\n> -\t * in the first place.\n> +\t * D/F conflicts and higher stage entries are not a concern\n> +\t * because cache-tree would be invalidated and we would never\n> +\t * get here in the first place.\n>   \t */\n>   \tfor (i = 0; i < nr_entries; i++) {\n>   \t\tint new_ce_len, len, rc;\n> @@ -761,7 +742,6 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n>   \n>   \t\tmark_ce_used(src[0], o);\n>   \t}\n> -\to->extra_add_index_flags = 0;\n>   \tfree(tree_ce);\n>   \tif (o->debug_unpack)\n>   \t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n> @@ -791,7 +771,17 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n>   \n>   \t\tif (!o->merge || df_conflicts)\n>   \t\t\tBUG(\"Wrong condition to get here buddy\");\n> -\t\treturn traverse_by_cache_tree(pos, nr_entries, n, names, info);\n> +\n> +\t\t/*\n> +\t\t * All entries up to 'pos' must have been processed\n> +\t\t * (i.e. marked CE_UNPACKED) at this point. But to be safe,\n> +\t\t * save and restore cache_bottom anyway to not miss\n> +\t\t * unprocessed entries before 'pos'.\n> +\t\t */\n> +\t\tbottom = o->cache_bottom;\n> +\t\tret = traverse_by_cache_tree(pos, nr_entries, n, names, info);\n> +\t\to->cache_bottom = bottom;\n> +\t\treturn ret;\n>   \t}\n>   \n>   \tp = names;\n> @@ -1142,7 +1132,7 @@ static void debug_unpack_callback(int n,\n>   }\n>   \n>   /*\n> - * Note that traverse_by_cache_tree() duplicates some logic in this funciton\n> + * Note that traverse_by_cache_tree() duplicates some logic in this function\n>    * without actually calling it. If you change the logic here you may need to\n>    * check and change there as well.\n>    */\n> @@ -1425,11 +1415,11 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>   \tint i, ret;\n>   \tstatic struct cache_entry *dfc;\n>   \tstruct exclude_list el;\n> -\tuint64_t start = getnanotime();\n>   \n>   \tif (len > MAX_UNPACK_TREES)\n>   \t\tdie(\"unpack_trees takes at most %d trees\", MAX_UNPACK_TREES);\n>   \n> +\ttrace_performance_enter();\n>   \tmemset(&el, 0, sizeof(el));\n>   \tif (!core_apply_sparse_checkout || !o->update)\n>   \t\to->skip_sparse_checkout = 1;\n> @@ -1502,7 +1492,10 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>   \t\t\t}\n>   \t\t}\n>   \n> -\t\tif (traverse_trees(len, t, &info) < 0)\n> +\t\ttrace_performance_enter();\n> +\t\tret = traverse_trees(len, t, &info);\n> +\t\ttrace_performance_leave(\"traverse_trees\");\n> +\t\tif (ret < 0)\n>   \t\t\tgoto return_failed;\n>   \t}\n>   \n> @@ -1574,10 +1567,10 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>   \t\t\tgoto done;\n>   \t\t}\n>   \t}\n> -\ttrace_performance_since(start, \"unpack trees\");\n>   \n>   \tret = check_updates(o) ? (-2) : 0;\n>   \tif (o->dst_index) {\n> +\t\tmove_index_extensions(&o->result, o->src_index);\n>   \t\tif (!ret) {\n>   \t\t\tif (!o->result.cache_tree)\n>   \t\t\t\to->result.cache_tree = cache_tree();\n> @@ -1586,7 +1579,6 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>   \t\t\t\t\t\t  WRITE_TREE_SILENT |\n>   \t\t\t\t\t\t  WRITE_TREE_REPAIR);\n>   \t\t}\n> -\t\tmove_index_extensions(&o->result, o->src_index);\n>   \t\tdiscard_index(o->dst_index);\n>   \t\t*o->dst_index = o->result;\n>   \t} else {\n> @@ -1595,6 +1587,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>   \to->src_index = NULL;\n>   \n>   done:\n> +\ttrace_performance_leave(\"unpack_trees\");\n>   \tclear_exclude_list(&el);\n>   \treturn ret;\n>   \n> diff --git a/unpack-trees.h b/unpack-trees.h\n> index 94e1b14078..c2b434c606 100644\n> --- a/unpack-trees.h\n> +++ b/unpack-trees.h\n> @@ -80,7 +80,6 @@ struct unpack_trees_options {\n>   \tstruct index_state result;\n>   \n>   \tstruct exclude_list *el; /* for internal use */\n> -\tunsigned int extra_add_index_flags;\n>   };\n>   \n>   extern int unpack_trees(unsigned n, struct tree_desc *t,\n> \n> Nguyễn Thái Ngọc Duy (5):\n>    trace.h: support nested performance tracing\n>    unpack-trees: add performance tracing\n>    unpack-trees: optimize walking same trees with cache-tree\n>    unpack-trees: reduce malloc in cache-tree walk\n>    unpack-trees: reuse (still valid) cache-tree from src_index\n> \n>   cache-tree.c    |   2 +\n>   diff-lib.c      |   4 +-\n>   dir.c           |   4 +-\n>   name-hash.c     |   4 +-\n>   preload-index.c |   4 +-\n>   read-cache.c    |  13 +++--\n>   trace.c         |  69 ++++++++++++++++++++--\n>   trace.h         |  15 +++++\n>   unpack-trees.c  | 149 +++++++++++++++++++++++++++++++++++++++++++++++-\n>   9 files changed, 243 insertions(+), 21 deletions(-)\n> \n"},{"id":"355605","messageId":"20180814195456.GE28452@sigill.intra.peff.net","threadId":"48913","inReplyTo":"CACsJy8CTNeR8Bchj37yNL+mWp1Y5rhD6QV2Gf06CPLHVXd8TDQ@mail.gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-08-14T19:54:56Z","receivedAt":"2018-08-14T19:54:59Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Aug 14, 2018 at 08:51:41PM +0200, Duy Nguyen wrote:\n\n> > But AFAICT these series are all about putting the sampling points into the\n> > code base, so formatting would be orthogonal to it?\n> \n> It's not just sampling points. There's things like index id being\n> shown in the message for example. I prefer to keep free style format\n> to help me read. There's also things like indentation I do here to\n> help me read. Granted you could do all that with scripts and stuff,\n> but will we pass around in mail  dumps of json messages to be decoded\n> locally?\n\nI think you could have both forms using the same entry points sprinkled\nthrough the code.\n\nAt GitHub we have a similar telemetry-ish thing, where we collect some\ndata points and then the resulting JSON is stored for every operation\n(for a few weeks for read ops, and indefinitely attached to every ref\nwrite).\n\nAnd I've found that the storage and the trace-style \"just show a\nhuman-readable message to stderr\" interface complement each other in\nboth directions:\n\n - you can output a human readable message that is sent immediately to\n   the trace mechanism but _also_ becomes part of the telemetry. E.g.,\n   imagine that one item in the json blob is \"this is the last message\n   from GIT_TRACE_FOO\". Now you can push tracing messages into whatever\n   plan you're using to store SLOG. We do this less with TRACE, and much\n   more with error() and die() messages.\n\n - when a structured telemetry item is updated, we can still output a\n   human-readable trace message with just that item. E.g., with:\n\n     trace_performance(n, \"foo\");\n\n   we could either store a json key (perf.foo=n) or output a nicely\n   formatted string like we do now, depending on what the user has\n   configured (or even both, of course).\n\nIt helps if the sampling points give enough information to cover both\ncases (as in the trace_performance example), but you can generally\nshoe-horn unstructured data into the structured log, and pretty-print\nstructured data.\n\n-Peff\n"},{"id":"355610","messageId":"603037bc-57a4-92a6-9c13-ae5b253d3ba3@jeffhostetler.com","threadId":"48913","inReplyTo":"CAGZ79kZwVpCBMkBKuYpwZFgAN50wZub_fyzWrAsE=ksuc-aCgQ@mail.gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Jeff Hostetler","fromEmail":"git@jeffhostetler.com","sentAt":"2018-08-14T20:14:00Z","receivedAt":"2018-08-14T20:14:04Z","isPatch":true,"sender":{"key":"git@jeffhostetler.com","avatar":null},"body":"\n\nOn 8/14/2018 2:44 PM, Stefan Beller wrote:\n> On Tue, Aug 14, 2018 at 11:32 AM Duy Nguyen <pclouds@gmail.com> wrote:\n>>\n>> On Tue, Aug 14, 2018 at 8:19 PM Jeff Hostetler <git@jeffhostetler.com> wrote:\n>>> I'm looking at adding code to my SLOG (better name suggestions welcome)\n>>> patch series to eventually replace the existing git_trace facility.\n>>\n>> Complement maybe. Replace, please no. I'd rather not stare at json messages.\n> \n>  From the sidelines: We'd only need one logging infrastructure in place, as the\n> formatting would be done as a later step? For local operations we'd certainly\n> find better formatting than json, and we figured that we might end up desiring\n> ProtocolBuffers[1] instead of JSon, so if it would be easy to change\n> the output of\n> the structured logging easily that would be great.\n> \n> But AFAICT these series are all about putting the sampling points into the\n> code base, so formatting would be orthogonal to it?\n> \n> Stefan\n> \n> [1] https://developers.google.com/protocol-buffers/\n> \n\nLast time I checked, protocol-buffers has a C++ binding but not\na C binding.\n\nI've not had a chance to use pbuffers, so I have to ask what advantages\nwould they have over JSON or some other similar self-describing format?\nAnd/or would it be possible for you to tail the json log file and\nconvert it to whatever format you preferred?\n\nIt seems like the important thing is to capture structured data\n(whatever the format) to disk first.\n\nJeff\n"},{"id":"355615","messageId":"xmqqeff0zn53.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"CACsJy8CTNeR8Bchj37yNL+mWp1Y5rhD6QV2Gf06CPLHVXd8TDQ@mail.gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-14T20:52:40Z","receivedAt":"2018-08-14T20:52:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> These trace messages are made for human consumption. Granted\n> occasionally we need some processing but I find one liners mostly\n> suffice. Now we turn these into something made for machines, turning\n> people to second citizens. I've read these messages reformatted for\n> human, it's usually too verbose even if it's reformatted.\n\nI actually actively hate the aspect of the slog thing that exposes\nthe fact that it wants to take and show JSON too much in its API,\nbut if you look at these \"jw_object_*()\" thing as _only_ filling\nparameters to be emitted, there is no reason to think we cannot\nenhance/extend slog_emit_*() thing to take a format string (perhaps\ninside the jw structure) so that the formatter does not have to\ngenerate JSON at all.  Envisioning that kind of future, json_writer\nis a misnomer that too narrowly defines what it is---it is merely a\ngeneric data container that the codepath being traced can use to\ncommunicate what needs to be logged to the outside world.\nslog_emit* can (and when enhanced, should) be capable of paying\nattention to an external input (e.g. environment variable) to switch\nthe output format, and JSON could be just one of the choices.\n\n> It's not just sampling points. There's things like index id being\n> shown in the message for example. I prefer to keep free style format\n> to help me read. There's also things like indentation I do here to\n> help me read.\n\nYup, I do not think that contradicts with the approach to have a\nsingle unified \"data collection\" API; you should also be able to\nspecify how that collection of data is to be presented in the trace\nmessages meant for humans, which would be discarded when emitting\njson but would be used when showing human-readble trace, no?\n"},{"id":"355702","messageId":"CACsJy8C5xPOa26q_dvGgrmkV+C-k2kmc8_nQbwzcDVNue4ehYw@mail.gmail.com","threadId":"48913","inReplyTo":"xmqqeff0zn53.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-15T16:32:35Z","receivedAt":"2018-08-15T16:33:03Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Aug 14, 2018 at 10:52 PM Junio C Hamano <gitster@pobox.com> wrote:\n> > It's not just sampling points. There's things like index id being\n> > shown in the message for example. I prefer to keep free style format\n> > to help me read. There's also things like indentation I do here to\n> > help me read.\n>\n> Yup, I do not think that contradicts with the approach to have a\n> single unified \"data collection\" API; you should also be able to\n> specify how that collection of data is to be presented in the trace\n> messages meant for humans, which would be discarded when emitting\n> json but would be used when showing human-readble trace, no?\n\nYes. As Peff also pointed out in another mail, as long as this\nstructured logging stuff does not stop me from manual trace messages\nand don't force more work on me when I add new traces, I don't care if\nit exists.\n-- \nDuy\n"},{"id":"355704","messageId":"CACsJy8BFKzQurH6v-geY_ZJjtwF+S2YnZQA=Q6FajLm=GvHjiw@mail.gmail.com","threadId":"48913","inReplyTo":"f3403347-607b-b67c-297c-eeb9190a7de7@gmail.com","subject":"Re: [PATCH v4 3/5] unpack-trees: optimize walking same trees with cache-tree","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-15T16:38:03Z","receivedAt":"2018-08-15T16:38:31Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 13, 2018 at 8:58 PM Ben Peart <peartben@gmail.com> wrote:\n> > +      *\n> > +      * D/F conflicts and higher stage entries are not a concern\n> > +      * because cache-tree would be invalidated and we would never\n> > +      * get here in the first place.\n> > +      */\n> > +     for (i = 0; i < nr_entries; i++) {\n> > +             struct cache_entry *tree_ce;\n> > +             int len, rc;\n> > +\n> > +             src[0] = o->src_index->cache[pos + i];\n> > +\n> > +             len = ce_namelen(src[0]);\n> > +             tree_ce = xcalloc(1, cache_entry_size(len));\n> > +\n> > +             tree_ce->ce_mode = src[0]->ce_mode;\n> > +             tree_ce->ce_flags = create_ce_flags(0);\n> > +             tree_ce->ce_namelen = len;\n> > +             oidcpy(&tree_ce->oid, &src[0]->oid);\n> > +             memcpy(tree_ce->name, src[0]->name, len + 1);\n> > +\n> > +             for (d = 1; d <= nr_names; d++)\n> > +                     src[d] = tree_ce;\n> > +\n> > +             rc = call_unpack_fn((const struct cache_entry * const *)src, o);\n>\n> I don't fully understand why this is still necessary since \"we detect\n> that all trees are the same as cache-tree at this path.\"  I do know\n> (because I tried it :)) that if we don't actually call the unpack\n> function the patch fails a bunch of tests so clearly something important\n> is being missed.\n\nYeah because removing this line assumes n-way logic, which most likely\nmeans \"use the index version if all trees are the same as the index\"\nbut it's not necessarily true. There could be flags that make n-way\nbehave differently. And even if we make that assumption, we need to\ncopy src[0] to o->result (heh I tried that \"skip call_unpack_fn\" thing\ntoo when I thought this would be the same as the diff-index --cached\noptimization path, and only realized copying to o->result was needed\nafterwards).\n-- \nDuy\n"},{"id":"355733","messageId":"xmqq36vfiiws.fsf@gitster-ct.c.googlers.com","threadId":"48913","inReplyTo":"CACsJy8C5xPOa26q_dvGgrmkV+C-k2kmc8_nQbwzcDVNue4ehYw@mail.gmail.com","subject":"Re: [PATCH v4 2/5] unpack-trees: add performance tracing","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-15T18:28:19Z","receivedAt":"2018-08-15T18:30:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> On Tue, Aug 14, 2018 at 10:52 PM Junio C Hamano <gitster@pobox.com> wrote:\n>> > It's not just sampling points. There's things like index id being\n>> > shown in the message for example. I prefer to keep free style format\n>> > to help me read. There's also things like indentation I do here to\n>> > help me read.\n>>\n>> Yup, I do not think that contradicts with the approach to have a\n>> single unified \"data collection\" API; you should also be able to\n>> specify how that collection of data is to be presented in the trace\n>> messages meant for humans, which would be discarded when emitting\n>> json but would be used when showing human-readble trace, no?\n>\n> Yes. As Peff also pointed out in another mail, as long as this\n> structured logging stuff does not stop me from manual trace messages\n> and don't force more work on me when I add new traces, I don't care if\n> it exists.\n\nI am hoping that we are on the same page, but just to make sure,\nwhat I think we would want is to have just a single set of\nannotations in the codepath, instead of \"we can add annotations from\nthese two separate sets, and they do not interfere each other so I\ndo not care about what the other guy is doing\".\n\nIOW, I found it highly annoying having to resolve merges like\n7234f27b (\"Merge branch 'nd/unpack-trees-with-cache-tree' into pu\",\n2018-08-14), taking two topics that try to use different tracing\nmechanisms in the same codepath.\n"},{"id":"355978","messageId":"20180818144128.19361-1-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180812081551.27927-1-pclouds@gmail.com","subject":"[PATCH v5 0/7] Speed up unpack_trees()","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-18T14:41:21Z","receivedAt":"2018-08-18T14:41:41Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"v5 fixes some minor comments from round 4 and a big mistake in 5/5.\nJunio's scary feeling turns out true. There is a missing invalidation\nin keep_entry() which is not added in 6/7. 7/7 makes sure that similar\nproblems will not slip through.\n\nI had to rebase this series on top of 'master' because 7/7 caught a\nbad cache-tree situation that has been fixed by Elijah in ad3762042a\n(read-cache: fix directory/file conflict handling in\nread_index_unmerged() - 2018-07-31). I believe the issue was we prime\ncache-tree in 'git reset --hard' even though the index has conflicts.\n\nRange-diff (before the rebase):\n\n1:  a192faf79e ! 1:  ed8763726b trace.h: support nested performance tracing\n    @@ -49,13 +49,16 @@\n      \tstruct untracked_cache_dir *untracked;\n     -\tuint64_t start = getnanotime();\n      \n    - \tif (has_symlink_leading_path(path, len))\n    +-\tif (has_symlink_leading_path(path, len))\n    ++\ttrace_performance_enter();\n    ++\n    ++\tif (has_symlink_leading_path(path, len)) {\n    ++\t\ttrace_performance_leave(\"read directory %.*s\", len, path);\n      \t\treturn dir->nr;\n    ++\t}\n      \n    -+\ttrace_performance_enter();\n      \tuntracked = validate_untracked_cache(dir, len, pathspec);\n      \tif (!untracked)\n    - \t\t/*\n     @@\n      \t\tdir->nr = i;\n      \t}\n2:  9afe7c488a = 2:  9b70652fa2 unpack-trees: add performance tracing\n3:  74101edb60 ! 3:  8b3cfea623 unpack-trees: optimize walking same trees with cache-tree\n    @@ -141,7 +141,7 @@\n     +\n     +\t/*\n     +\t * Do what unpack_callback() and unpack_nondirectories() normally\n    -+\t * do. But we walk all paths recursively in just one loop instead.\n    ++\t * do. But we walk all paths in an iterative loop instead.\n     +\t *\n     +\t * D/F conflicts and higher stage entries are not a concern\n     +\t * because cache-tree would be invalidated and we would never\n4:  9261c5920e = 4:  5af28d44ca unpack-trees: reduce malloc in cache-tree walk\n5:  43fac1154f = 5:  5657c92fe9 unpack-trees: reuse (still valid) cache-tree from src_index\n-:  ---------- > 6:  3b91783afc unpack-trees: add missing cache invalidation\n-:  ---------- > 7:  0d5464c0dc cache-tree: verify valid cache-tree in the test suite\n\nNguyễn Thái Ngọc Duy (7):\n  trace.h: support nested performance tracing\n  unpack-trees: add performance tracing\n  unpack-trees: optimize walking same trees with cache-tree\n  unpack-trees: reduce malloc in cache-tree walk\n  unpack-trees: reuse (still valid) cache-tree from src_index\n  unpack-trees: add missing cache invalidation\n  cache-tree: verify valid cache-tree in the test suite\n\n cache-tree.c    |  80 +++++++++++++++++++++++++\n cache-tree.h    |   1 +\n diff-lib.c      |   4 +-\n dir.c           |   9 ++-\n name-hash.c     |   4 +-\n preload-index.c |   4 +-\n read-cache.c    |  16 +++--\n t/test-lib.sh   |   6 ++\n trace.c         |  69 ++++++++++++++++++++--\n trace.h         |  15 +++++\n unpack-trees.c  | 154 +++++++++++++++++++++++++++++++++++++++++++++++-\n 11 files changed, 340 insertions(+), 22 deletions(-)\n\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355979","messageId":"20180818144128.19361-2-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-1-pclouds@gmail.com","subject":"[PATCH v5 1/7] trace.h: support nested performance tracing","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-18T14:41:22Z","receivedAt":"2018-08-18T14:41:42Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Performance measurements are listed right now as a flat list, which is\nfine when we measure big blocks. But when we start adding more and\nmore measurements, some of them could be just part of a bigger\nmeasurement and a flat list gives a wrong impression that they are\nexecuted at the same level instead of nested.\n\nAdd trace_performance_enter() and trace_performance_leave() to allow\nindent these nested measurements. For now it does not help much\nbecause the only nested thing is (lazy) name hash initialization\n(e.g. called in diff-index from \"git status\"). This will help more\nbecause I'm going to add some more tracing that's actually nested.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n diff-lib.c      |  4 +--\n dir.c           |  9 ++++---\n name-hash.c     |  4 +--\n preload-index.c |  4 +--\n read-cache.c    | 11 ++++----\n trace.c         | 69 ++++++++++++++++++++++++++++++++++++++++++++-----\n trace.h         | 15 +++++++++++\n 7 files changed, 96 insertions(+), 20 deletions(-)\n\ndiff --git a/diff-lib.c b/diff-lib.c\nindex 732f684a49..d5bbb7ea50 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -518,11 +518,11 @@ static int diff_cache(struct rev_info *revs,\n int run_diff_index(struct rev_info *revs, int cached)\n {\n \tstruct object_array_entry *ent;\n-\tuint64_t start = getnanotime();\n \n \tif (revs->pending.nr != 1)\n \t\tBUG(\"run_diff_index must be passed exactly one tree\");\n \n+\ttrace_performance_enter();\n \tent = revs->pending.objects;\n \tif (diff_cache(revs, &ent->item->oid, ent->name, cached))\n \t\texit(128);\n@@ -531,7 +531,7 @@ int run_diff_index(struct rev_info *revs, int cached)\n \tdiffcore_fix_diff_index(&revs->diffopt);\n \tdiffcore_std(&revs->diffopt);\n \tdiff_flush(&revs->diffopt);\n-\ttrace_performance_since(start, \"diff-index\");\n+\ttrace_performance_leave(\"diff-index\");\n \treturn 0;\n }\n \ndiff --git a/dir.c b/dir.c\nindex 32f5f72759..18b57b94cc 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -2263,10 +2263,13 @@ int read_directory(struct dir_struct *dir, struct index_state *istate,\n \t\t   const char *path, int len, const struct pathspec *pathspec)\n {\n \tstruct untracked_cache_dir *untracked;\n-\tuint64_t start = getnanotime();\n \n-\tif (has_symlink_leading_path(path, len))\n+\ttrace_performance_enter();\n+\n+\tif (has_symlink_leading_path(path, len)) {\n+\t\ttrace_performance_leave(\"read directory %.*s\", len, path);\n \t\treturn dir->nr;\n+\t}\n \n \tuntracked = validate_untracked_cache(dir, len, pathspec);\n \tif (!untracked)\n@@ -2302,7 +2305,7 @@ int read_directory(struct dir_struct *dir, struct index_state *istate,\n \t\tdir->nr = i;\n \t}\n \n-\ttrace_performance_since(start, \"read directory %.*s\", len, path);\n+\ttrace_performance_leave(\"read directory %.*s\", len, path);\n \tif (dir->untracked) {\n \t\tstatic int force_untracked_cache = -1;\n \t\tstatic struct trace_key trace_untracked_stats = TRACE_KEY_INIT(UNTRACKED_STATS);\ndiff --git a/name-hash.c b/name-hash.c\nindex 163849831c..1fcda73cb3 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -578,10 +578,10 @@ static void threaded_lazy_init_name_hash(\n \n static void lazy_init_name_hash(struct index_state *istate)\n {\n-\tuint64_t start = getnanotime();\n \n \tif (istate->name_hash_initialized)\n \t\treturn;\n+\ttrace_performance_enter();\n \thashmap_init(&istate->name_hash, cache_entry_cmp, NULL, istate->cache_nr);\n \thashmap_init(&istate->dir_hash, dir_entry_cmp, NULL, istate->cache_nr);\n \n@@ -602,7 +602,7 @@ static void lazy_init_name_hash(struct index_state *istate)\n \t}\n \n \tistate->name_hash_initialized = 1;\n-\ttrace_performance_since(start, \"initialize name hash\");\n+\ttrace_performance_leave(\"initialize name hash\");\n }\n \n /*\ndiff --git a/preload-index.c b/preload-index.c\nindex 4d08d44874..d7f7919ba2 100644\n--- a/preload-index.c\n+++ b/preload-index.c\n@@ -78,7 +78,6 @@ static void preload_index(struct index_state *index,\n {\n \tint threads, i, work, offset;\n \tstruct thread_data data[MAX_PARALLEL];\n-\tuint64_t start = getnanotime();\n \n \tif (!core_preload_index)\n \t\treturn;\n@@ -88,6 +87,7 @@ static void preload_index(struct index_state *index,\n \t\tthreads = 2;\n \tif (threads < 2)\n \t\treturn;\n+\ttrace_performance_enter();\n \tif (threads > MAX_PARALLEL)\n \t\tthreads = MAX_PARALLEL;\n \toffset = 0;\n@@ -109,7 +109,7 @@ static void preload_index(struct index_state *index,\n \t\tif (pthread_join(p->pthread, NULL))\n \t\t\tdie(\"unable to join threaded lstat\");\n \t}\n-\ttrace_performance_since(start, \"preload index\");\n+\ttrace_performance_leave(\"preload index\");\n }\n #endif\n \ndiff --git a/read-cache.c b/read-cache.c\nindex c5fabc844a..1c9c88c130 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1476,8 +1476,8 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n \tconst char *typechange_fmt;\n \tconst char *added_fmt;\n \tconst char *unmerged_fmt;\n-\tuint64_t start = getnanotime();\n \n+\ttrace_performance_enter();\n \tmodified_fmt = (in_porcelain ? \"M\\t%s\\n\" : \"%s: needs update\\n\");\n \tdeleted_fmt = (in_porcelain ? \"D\\t%s\\n\" : \"%s: needs update\\n\");\n \ttypechange_fmt = (in_porcelain ? \"T\\t%s\\n\" : \"%s needs update\\n\");\n@@ -1547,7 +1547,7 @@ int refresh_index(struct index_state *istate, unsigned int flags,\n \n \t\treplace_index_entry(istate, i, new_entry);\n \t}\n-\ttrace_performance_since(start, \"refresh index\");\n+\ttrace_performance_leave(\"refresh index\");\n \treturn has_errors;\n }\n \n@@ -2002,7 +2002,6 @@ static void freshen_shared_index(const char *shared_index, int warn)\n int read_index_from(struct index_state *istate, const char *path,\n \t\t    const char *gitdir)\n {\n-\tuint64_t start = getnanotime();\n \tstruct split_index *split_index;\n \tint ret;\n \tchar *base_oid_hex;\n@@ -2012,8 +2011,9 @@ int read_index_from(struct index_state *istate, const char *path,\n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n \n+\ttrace_performance_enter();\n \tret = do_read_index(istate, path, 0);\n-\ttrace_performance_since(start, \"read cache %s\", path);\n+\ttrace_performance_leave(\"read cache %s\", path);\n \n \tsplit_index = istate->split_index;\n \tif (!split_index || is_null_oid(&split_index->base_oid)) {\n@@ -2021,6 +2021,7 @@ int read_index_from(struct index_state *istate, const char *path,\n \t\treturn ret;\n \t}\n \n+\ttrace_performance_enter();\n \tif (split_index->base)\n \t\tdiscard_index(split_index->base);\n \telse\n@@ -2037,8 +2038,8 @@ int read_index_from(struct index_state *istate, const char *path,\n \tfreshen_shared_index(base_path, 0);\n \tmerge_base_index(istate);\n \tpost_read_index_from(istate);\n-\ttrace_performance_since(start, \"read cache %s\", base_path);\n \tfree(base_path);\n+\ttrace_performance_leave(\"read cache %s\", base_path);\n \treturn ret;\n }\n \ndiff --git a/trace.c b/trace.c\nindex fc623e91fd..fa4a2e7120 100644\n--- a/trace.c\n+++ b/trace.c\n@@ -176,10 +176,30 @@ void trace_strbuf_fl(const char *file, int line, struct trace_key *key,\n \tstrbuf_release(&buf);\n }\n \n+static uint64_t perf_start_times[10];\n+static int perf_indent;\n+\n+uint64_t trace_performance_enter(void)\n+{\n+\tuint64_t now;\n+\n+\tif (!trace_want(&trace_perf_key))\n+\t\treturn 0;\n+\n+\tnow = getnanotime();\n+\tperf_start_times[perf_indent] = now;\n+\tif (perf_indent + 1 < ARRAY_SIZE(perf_start_times))\n+\t\tperf_indent++;\n+\telse\n+\t\tBUG(\"Too deep indentation\");\n+\treturn now;\n+}\n+\n static void trace_performance_vprintf_fl(const char *file, int line,\n \t\t\t\t\t uint64_t nanos, const char *format,\n \t\t\t\t\t va_list ap)\n {\n+\tstatic const char space[] = \"          \";\n \tstruct strbuf buf = STRBUF_INIT;\n \n \tif (!prepare_trace_line(file, line, &trace_perf_key, &buf))\n@@ -188,7 +208,10 @@ static void trace_performance_vprintf_fl(const char *file, int line,\n \tstrbuf_addf(&buf, \"performance: %.9f s\", (double) nanos / 1000000000);\n \n \tif (format && *format) {\n-\t\tstrbuf_addstr(&buf, \": \");\n+\t\tif (perf_indent >= strlen(space))\n+\t\t\tBUG(\"Too deep indentation\");\n+\n+\t\tstrbuf_addf(&buf, \":%.*s \", perf_indent, space);\n \t\tstrbuf_vaddf(&buf, format, ap);\n \t}\n \n@@ -244,6 +267,24 @@ void trace_performance_since(uint64_t start, const char *format, ...)\n \tva_end(ap);\n }\n \n+void trace_performance_leave(const char *format, ...)\n+{\n+\tva_list ap;\n+\tuint64_t since;\n+\n+\tif (perf_indent)\n+\t\tperf_indent--;\n+\n+\tif (!format) /* Allow callers to leave without tracing anything */\n+\t\treturn;\n+\n+\tsince = perf_start_times[perf_indent];\n+\tva_start(ap, format);\n+\ttrace_performance_vprintf_fl(NULL, 0, getnanotime() - since,\n+\t\t\t\t     format, ap);\n+\tva_end(ap);\n+}\n+\n #else\n \n void trace_printf_key_fl(const char *file, int line, struct trace_key *key,\n@@ -273,6 +314,24 @@ void trace_performance_fl(const char *file, int line, uint64_t nanos,\n \tva_end(ap);\n }\n \n+void trace_performance_leave_fl(const char *file, int line,\n+\t\t\t\tuint64_t nanos, const char *format, ...)\n+{\n+\tva_list ap;\n+\tuint64_t since;\n+\n+\tif (perf_indent)\n+\t\tperf_indent--;\n+\n+\tif (!format) /* Allow callers to leave without tracing anything */\n+\t\treturn;\n+\n+\tsince = perf_start_times[perf_indent];\n+\tva_start(ap, format);\n+\ttrace_performance_vprintf_fl(file, line, nanos - since, format, ap);\n+\tva_end(ap);\n+}\n+\n #endif /* HAVE_VARIADIC_MACROS */\n \n \n@@ -411,13 +470,11 @@ uint64_t getnanotime(void)\n \t}\n }\n \n-static uint64_t command_start_time;\n static struct strbuf command_line = STRBUF_INIT;\n \n static void print_command_performance_atexit(void)\n {\n-\ttrace_performance_since(command_start_time, \"git command:%s\",\n-\t\t\t\tcommand_line.buf);\n+\ttrace_performance_leave(\"git command:%s\", command_line.buf);\n }\n \n void trace_command_performance(const char **argv)\n@@ -425,10 +482,10 @@ void trace_command_performance(const char **argv)\n \tif (!trace_want(&trace_perf_key))\n \t\treturn;\n \n-\tif (!command_start_time)\n+\tif (!command_line.len)\n \t\tatexit(print_command_performance_atexit);\n \n \tstrbuf_reset(&command_line);\n \tsq_quote_argv_pretty(&command_line, argv);\n-\tcommand_start_time = getnanotime();\n+\ttrace_performance_enter();\n }\ndiff --git a/trace.h b/trace.h\nindex 2b6a1bc17c..171b256d26 100644\n--- a/trace.h\n+++ b/trace.h\n@@ -23,6 +23,7 @@ extern void trace_disable(struct trace_key *key);\n extern uint64_t getnanotime(void);\n extern void trace_command_performance(const char **argv);\n extern void trace_verbatim(struct trace_key *key, const void *buf, unsigned len);\n+uint64_t trace_performance_enter(void);\n \n #ifndef HAVE_VARIADIC_MACROS\n \n@@ -45,6 +46,9 @@ extern void trace_performance(uint64_t nanos, const char *format, ...);\n __attribute__((format (printf, 2, 3)))\n extern void trace_performance_since(uint64_t start, const char *format, ...);\n \n+__attribute__((format (printf, 1, 2)))\n+void trace_performance_leave(const char *format, ...);\n+\n #else\n \n /*\n@@ -118,6 +122,14 @@ extern void trace_performance_since(uint64_t start, const char *format, ...);\n \t\t\t\t\t     __VA_ARGS__);\t\t    \\\n \t} while (0)\n \n+#define trace_performance_leave(...)\t\t\t\t\t    \\\n+\tdo {\t\t\t\t\t\t\t\t    \\\n+\t\tif (trace_pass_fl(&trace_perf_key))\t\t\t    \\\n+\t\t\ttrace_performance_leave_fl(TRACE_CONTEXT, __LINE__, \\\n+\t\t\t\t\t\t   getnanotime(),\t    \\\n+\t\t\t\t\t\t   __VA_ARGS__);\t    \\\n+\t} while (0)\n+\n /* backend functions, use non-*fl macros instead */\n __attribute__((format (printf, 4, 5)))\n extern void trace_printf_key_fl(const char *file, int line, struct trace_key *key,\n@@ -130,6 +142,9 @@ extern void trace_strbuf_fl(const char *file, int line, struct trace_key *key,\n __attribute__((format (printf, 4, 5)))\n extern void trace_performance_fl(const char *file, int line,\n \t\t\t\t uint64_t nanos, const char *fmt, ...);\n+__attribute__((format (printf, 4, 5)))\n+extern void trace_performance_leave_fl(const char *file, int line,\n+\t\t\t\t       uint64_t nanos, const char *fmt, ...);\n static inline int trace_pass_fl(struct trace_key *key)\n {\n \treturn key->fd || !key->initialized;\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355980","messageId":"20180818144128.19361-3-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-1-pclouds@gmail.com","subject":"[PATCH v5 2/7] unpack-trees: add performance tracing","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-18T14:41:23Z","receivedAt":"2018-08-18T14:41:43Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"We're going to optimize unpack_trees() a bit in the following\npatches. Let's add some tracing to measure how long it takes before\nand after. This is the baseline (\"git checkout -\" on webkit.git, 275k\nfiles on worktree)\n\n    performance: 0.056651714 s:  read cache .git/index\n    performance: 0.183101080 s:  preload index\n    performance: 0.008584433 s:  refresh index\n    performance: 0.633767589 s:   traverse_trees\n    performance: 0.340265448 s:   check_updates\n    performance: 0.381884638 s:   cache_tree_update\n    performance: 1.401562947 s:  unpack_trees\n    performance: 0.338687914 s:  write index, changed mask = 2e\n    performance: 0.411927922 s:    traverse_trees\n    performance: 0.000023335 s:    check_updates\n    performance: 0.423697246 s:   unpack_trees\n    performance: 0.423708360 s:  diff-index\n    performance: 2.559524127 s: git command: git checkout -\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n cache-tree.c   | 2 ++\n unpack-trees.c | 9 ++++++++-\n 2 files changed, 10 insertions(+), 1 deletion(-)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 181d5919f0..caafbff2ff 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -433,7 +433,9 @@ int cache_tree_update(struct index_state *istate, int flags)\n \n \tif (i)\n \t\treturn i;\n+\ttrace_performance_enter();\n \ti = update_one(it, cache, entries, \"\", 0, &skip, flags);\n+\ttrace_performance_leave(\"cache_tree_update\");\n \tif (i < 0)\n \t\treturn i;\n \tistate->cache_changed |= CACHE_TREE_CHANGED;\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex f9efee0836..6d9f692ea6 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -345,6 +345,7 @@ static int check_updates(struct unpack_trees_options *o)\n \tstruct checkout state = CHECKOUT_INIT;\n \tint i;\n \n+\ttrace_performance_enter();\n \tstate.force = 1;\n \tstate.quiet = 1;\n \tstate.refresh_cache = 1;\n@@ -414,6 +415,7 @@ static int check_updates(struct unpack_trees_options *o)\n \terrs |= finish_delayed_checkout(&state);\n \tif (o->update)\n \t\tgit_attr_set_direction(GIT_ATTR_CHECKIN, NULL);\n+\ttrace_performance_leave(\"check_updates\");\n \treturn errs != 0;\n }\n \n@@ -1285,6 +1287,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \tif (len > MAX_UNPACK_TREES)\n \t\tdie(\"unpack_trees takes at most %d trees\", MAX_UNPACK_TREES);\n \n+\ttrace_performance_enter();\n \tmemset(&el, 0, sizeof(el));\n \tif (!core_apply_sparse_checkout || !o->update)\n \t\to->skip_sparse_checkout = 1;\n@@ -1357,7 +1360,10 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t\t}\n \t\t}\n \n-\t\tif (traverse_trees(len, t, &info) < 0)\n+\t\ttrace_performance_enter();\n+\t\tret = traverse_trees(len, t, &info);\n+\t\ttrace_performance_leave(\"traverse_trees\");\n+\t\tif (ret < 0)\n \t\t\tgoto return_failed;\n \t}\n \n@@ -1449,6 +1455,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \to->src_index = NULL;\n \n done:\n+\ttrace_performance_leave(\"unpack_trees\");\n \tclear_exclude_list(&el);\n \treturn ret;\n \n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355982","messageId":"20180818144128.19361-4-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-1-pclouds@gmail.com","subject":"[PATCH v5 3/7] unpack-trees: optimize walking same trees with cache-tree","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-18T14:41:24Z","receivedAt":"2018-08-18T14:41:46Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"In order to merge one or many trees with the index, unpack-trees code\nwalks multiple trees in parallel with the index and performs n-way\nmerge. If we find out at start of a directory that all trees are the\nsame (by comparing OID) and cache-tree happens to be available for\nthat directory as well, we could avoid walking the trees because we\nalready know what these trees contain: it's flattened in what's called\n\"the index\".\n\nThe upside is of course a lot less I/O since we can potentially skip\nlots of trees (think subtrees). We also save CPU because we don't have\nto inflate and apply the deltas. The downside is of course more\nfragile code since the logic in some functions are now duplicated\nelsewhere.\n\n\"checkout -\" with this patch on webkit.git (275k files):\n\n    baseline      new\n  --------------------------------------------------------------------\n    0.056651714   0.080394752 s:  read cache .git/index\n    0.183101080   0.216010838 s:  preload index\n    0.008584433   0.008534301 s:  refresh index\n    0.633767589   0.251992198 s:   traverse_trees\n    0.340265448   0.377031383 s:   check_updates\n    0.381884638   0.372768105 s:   cache_tree_update\n    1.401562947   1.045887251 s:  unpack_trees\n    0.338687914   0.314983512 s:  write index, changed mask = 2e\n    0.411927922   0.062572653 s:    traverse_trees\n    0.000023335   0.000022544 s:    check_updates\n    0.423697246   0.073795585 s:   unpack_trees\n    0.423708360   0.073807557 s:  diff-index\n    2.559524127   1.938191592 s: git command: git checkout -\n\nAnother measurement from Ben's running \"git checkout\" with over 500k\ntrees (on the whole series):\n\n    baseline        new\n  ----------------------------------------------------------------------\n    0.535510167     0.556558733     s: read cache .git/index\n    0.3057373       0.3147105       s: initialize name hash\n    0.0184082       0.023558433     s: preload index\n    0.086910967     0.089085967     s: refresh index\n    7.889590767     2.191554433     s: unpack trees\n    0.120760833     0.131941267     s: update worktree after a merge\n    2.2583504       2.572663167     s: repair cache-tree\n    0.8916137       0.959495233     s: write index, changed mask = 28\n    3.405199233     0.2710663       s: unpack trees\n    0.000999667     0.0021554       s: update worktree after a merge\n    3.4063306       0.273318333     s: diff-index\n    16.9524923      9.462943133     s: git command: git.exe checkout\n\nThis command calls unpack_trees() twice, the first time on 2way merge\nand the second 1way merge. In both times, \"unpack trees\" time is\nreduced to one third. Overall time reduction is not that impressive of\ncourse because index operations take a big chunk. And there's that\nrepair cache-tree line.\n\nPS. A note about cache-tree invalidation and the use of it in this\ncode.\n\nWe do invalidate cache-tree in _source_ index when we add new entries\nto the (temporary) \"result\" index. But we also use the cache-tree from\nsource index in this optimization. Does this mean we end up having no\ncache-tree in the source index to activate this optimization?\n\nThe answer is twisted: the order of finding a good cache-tree and\ninvalidating it matters. In this case we check for a good cache-tree\nfirst in all_trees_same_as_cache_tree(), then we start to merge things\nand potentially invalidate that same cache-tree in the process. Since\ncache-tree invalidation happens after the optimization kicks in, we're\nstill good. But we may lose that cache-tree at the very first\ncall_unpack_fn() call in traverse_by_cache_tree().\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n unpack-trees.c | 127 +++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 127 insertions(+)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 6d9f692ea6..8376663b59 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -635,6 +635,102 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n \treturn name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n }\n \n+static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n+\t\t\t\t\tstruct name_entry *names,\n+\t\t\t\t\tstruct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i;\n+\n+\tif (!o->merge || dirmask != ((1 << n) - 1))\n+\t\treturn 0;\n+\n+\tfor (i = 1; i < n; i++)\n+\t\tif (!are_same_oid(names, names + i))\n+\t\t\treturn 0;\n+\n+\treturn cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n+}\n+\n+static int index_pos_by_traverse_info(struct name_entry *names,\n+\t\t\t\t      struct traverse_info *info)\n+{\n+\tstruct unpack_trees_options *o = info->data;\n+\tint len = traverse_path_len(info, names);\n+\tchar *name = xmalloc(len + 1 /* slash */ + 1 /* NUL */);\n+\tint pos;\n+\n+\tmake_traverse_path(name, info, names);\n+\tname[len++] = '/';\n+\tname[len] = '\\0';\n+\tpos = index_name_pos(o->src_index, name, len);\n+\tif (pos >= 0)\n+\t\tBUG(\"This is a directory and should not exist in index\");\n+\tpos = -pos - 1;\n+\tif (!starts_with(o->src_index->cache[pos]->name, name) ||\n+\t    (pos > 0 && starts_with(o->src_index->cache[pos-1]->name, name)))\n+\t\tBUG(\"pos must point at the first entry in this directory\");\n+\tfree(name);\n+\treturn pos;\n+}\n+\n+/*\n+ * Fast path if we detect that all trees are the same as cache-tree at this\n+ * path. We'll walk these trees recursively using cache-tree/index instead of\n+ * ODB since already know what these trees contain.\n+ */\n+static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n+\t\t\t\t  struct name_entry *names,\n+\t\t\t\t  struct traverse_info *info)\n+{\n+\tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n+\tstruct unpack_trees_options *o = info->data;\n+\tint i, d;\n+\n+\tif (!o->merge)\n+\t\tBUG(\"We need cache-tree to do this optimization\");\n+\n+\t/*\n+\t * Do what unpack_callback() and unpack_nondirectories() normally\n+\t * do. But we walk all paths in an iterative loop instead.\n+\t *\n+\t * D/F conflicts and higher stage entries are not a concern\n+\t * because cache-tree would be invalidated and we would never\n+\t * get here in the first place.\n+\t */\n+\tfor (i = 0; i < nr_entries; i++) {\n+\t\tstruct cache_entry *tree_ce;\n+\t\tint len, rc;\n+\n+\t\tsrc[0] = o->src_index->cache[pos + i];\n+\n+\t\tlen = ce_namelen(src[0]);\n+\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n+\n+\t\ttree_ce->ce_mode = src[0]->ce_mode;\n+\t\ttree_ce->ce_flags = create_ce_flags(0);\n+\t\ttree_ce->ce_namelen = len;\n+\t\toidcpy(&tree_ce->oid, &src[0]->oid);\n+\t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n+\n+\t\tfor (d = 1; d <= nr_names; d++)\n+\t\t\tsrc[d] = tree_ce;\n+\n+\t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n+\t\tfree(tree_ce);\n+\t\tif (rc < 0)\n+\t\t\treturn rc;\n+\n+\t\tmark_ce_used(src[0], o);\n+\t}\n+\tif (o->debug_unpack)\n+\t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n+\t\t       nr_entries,\n+\t\t       o->src_index->cache[pos]->name,\n+\t\t       o->src_index->cache[pos + nr_entries - 1]->name);\n+\treturn 0;\n+}\n+\n static int traverse_trees_recursive(int n, unsigned long dirmask,\n \t\t\t\t    unsigned long df_conflicts,\n \t\t\t\t    struct name_entry *names,\n@@ -646,6 +742,27 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n \tvoid *buf[MAX_UNPACK_TREES];\n \tstruct traverse_info newinfo;\n \tstruct name_entry *p;\n+\tint nr_entries;\n+\n+\tnr_entries = all_trees_same_as_cache_tree(n, dirmask, names, info);\n+\tif (nr_entries > 0) {\n+\t\tstruct unpack_trees_options *o = info->data;\n+\t\tint pos = index_pos_by_traverse_info(names, info);\n+\n+\t\tif (!o->merge || df_conflicts)\n+\t\t\tBUG(\"Wrong condition to get here buddy\");\n+\n+\t\t/*\n+\t\t * All entries up to 'pos' must have been processed\n+\t\t * (i.e. marked CE_UNPACKED) at this point. But to be safe,\n+\t\t * save and restore cache_bottom anyway to not miss\n+\t\t * unprocessed entries before 'pos'.\n+\t\t */\n+\t\tbottom = o->cache_bottom;\n+\t\tret = traverse_by_cache_tree(pos, nr_entries, n, names, info);\n+\t\to->cache_bottom = bottom;\n+\t\treturn ret;\n+\t}\n \n \tp = names;\n \twhile (!p->mode)\n@@ -812,6 +929,11 @@ static struct cache_entry *create_ce_entry(const struct traverse_info *info,\n \treturn ce;\n }\n \n+/*\n+ * Note that traverse_by_cache_tree() duplicates some logic in this function\n+ * without actually calling it. If you change the logic here you may need to\n+ * check and change there as well.\n+ */\n static int unpack_nondirectories(int n, unsigned long mask,\n \t\t\t\t unsigned long dirmask,\n \t\t\t\t struct cache_entry **src,\n@@ -1004,6 +1126,11 @@ static void debug_unpack_callback(int n,\n \t\tdebug_name_entry(i, names + i);\n }\n \n+/*\n+ * Note that traverse_by_cache_tree() duplicates some logic in this function\n+ * without actually calling it. If you change the logic here you may need to\n+ * check and change there as well.\n+ */\n static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355981","messageId":"20180818144128.19361-6-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-1-pclouds@gmail.com","subject":"[PATCH v5 5/7] unpack-trees: reuse (still valid) cache-tree from src_index","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-18T14:41:26Z","receivedAt":"2018-08-18T14:41:47Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"We do n-way merge by walking the source index and n trees at the same\ntime and add merge results to a new temporary index called o->result.\nThe merge result for any given path could be either\n\n- keep_entry(): same old index entry in o->src_index is reused\n- merged_entry(): either a new entry is added, or an existing one updated\n- deleted_entry(): one entry from o->src_index is removed\n\nFor some reason [1] we keep making sure that the source index's\ncache-tree is still valid if used by o->result: for all those\nmerged/deleted entries, we invalidate the same path in o->src_index,\nso only cache-trees covering the \"keep_entry\" parts remain good.\n\nBecause of this, the cache-tree from o->src_index can be perfectly\nreused in o->result. And in fact we already rely on this logic to\nreuse untracked cache in edf3b90553 (unpack-trees: preserve index\nextensions - 2017-05-08). Move the cache-tree to o->result before\ndoing cache_tree_update() to reduce hashing cost.\n\nSince cache_tree_update() has risen up as one of the most expensive\nparts in unpack_trees() after the last few patches. This does help\nreduce unpack_trees() time significantly (on webkit.git):\n\n    before       after\n  --------------------------------------------------------------------\n    0.080394752  0.051258167 s:  read cache .git/index\n    0.216010838  0.212106298 s:  preload index\n    0.008534301  0.280521764 s:  refresh index\n    0.251992198  0.218160442 s:   traverse_trees\n    0.377031383  0.374948191 s:   check_updates\n    0.372768105  0.037040114 s:   cache_tree_update\n    1.045887251  0.672031609 s:  unpack_trees\n    0.314983512  0.317456290 s:  write index, changed mask = 2e\n    0.062572653  0.038382654 s:    traverse_trees\n    0.000022544  0.000042731 s:    check_updates\n    0.073795585  0.050930053 s:   unpack_trees\n    0.073807557  0.051099735 s:  diff-index\n    1.938191592  1.614241153 s: git command: git checkout -\n\n[1] I'm pretty sure the reason is an oversight in 34110cd4e3 (Make\n    'unpack_trees()' have a separate source and destination index -\n    2008-03-06). That patch aims to _not_ update the source index at\n    all. The invalidation should have been done on o->result in that\n    patch. But then there was no cache-tree on o->result even then so\n    it's pointless to do so.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n read-cache.c   | 2 ++\n unpack-trees.c | 2 +-\n 2 files changed, 3 insertions(+), 1 deletion(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 1c9c88c130..5ce40f39b3 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2940,6 +2940,8 @@ void move_index_extensions(struct index_state *dst, struct index_state *src)\n {\n \tdst->untracked = src->untracked;\n \tsrc->untracked = NULL;\n+\tdst->cache_tree = src->cache_tree;\n+\tsrc->cache_tree = NULL;\n }\n \n struct cache_entry *dup_cache_entry(const struct cache_entry *ce,\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex dbef6e1b8a..aa80b65ee1 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -1576,6 +1576,7 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \n \tret = check_updates(o) ? (-2) : 0;\n \tif (o->dst_index) {\n+\t\tmove_index_extensions(&o->result, o->src_index);\n \t\tif (!ret) {\n \t\t\tif (!o->result.cache_tree)\n \t\t\t\to->result.cache_tree = cache_tree();\n@@ -1584,7 +1585,6 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \t\t\t\t\t\t  WRITE_TREE_SILENT |\n \t\t\t\t\t\t  WRITE_TREE_REPAIR);\n \t\t}\n-\t\tmove_index_extensions(&o->result, o->src_index);\n \t\tdiscard_index(o->dst_index);\n \t\t*o->dst_index = o->result;\n \t} else {\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355983","messageId":"20180818144128.19361-7-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-1-pclouds@gmail.com","subject":"[PATCH v5 6/7] unpack-trees: add missing cache invalidation","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-18T14:41:27Z","receivedAt":"2018-08-18T14:41:48Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Any changes to the output index should be (confusingly) marked in the\nsource index with invalidate_ce_path(). This is used to make sure we\nstill have valid untracked cache and cache-tree extensions in the end.\n\nWe do a pretty good job of invalidating except in two places.\nverify_clean_subdirectory() is part of verify_absent() and\nverify_absent_sparse(). The former is usually called by merged_entry()\nor directly in threeway_merge(). The latter is obviously used by\nsparse checkout.\n\nIn these three call sites, only merged_entry() follows up with\ninvalidate_ce_path(). The other two don't, but they should not trigger\nthis ce removal because this is about D/F conflicts [1]. But let's be\nsafe and invalidate_ce_path() here as well.\n\nThe second place is keep_entry() which is also used by threeway_merge()\nto keep higher stage entries. In order to reuse cache-tree we need to\ninvalidate these paths as well. It's not a problem in the past because\nwhenever a higher stage entry is present, cache-tree will not be\ncreated [2]. Now we salvage cache-tree even when higher stage entries\nare present, we need more invalidation.\n\n[1] c81935348b (Fix switching to a branch with D/F when current branch\n    has file D. - 2007-03-15)\n\n[2] This is probably too strict. We should be able to create and save\n    cache-tree for the directories that do not have conflict entries\n    in cache_tree_update(). And this becomes more important when\n    cache-tree plays bigger role in terms of performance.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n unpack-trees.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex aa80b65ee1..bc43922922 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -1774,6 +1774,7 @@ static int verify_clean_subdirectory(const struct cache_entry *ce,\n \t\t\tif (verify_uptodate(ce2, o))\n \t\t\t\treturn -1;\n \t\t\tadd_entry(o, ce2, CE_REMOVE, 0);\n+\t\t\tinvalidate_ce_path(ce, o);\n \t\t\tmark_ce_used(ce2, o);\n \t\t}\n \t\tcnt++;\n@@ -2033,6 +2034,8 @@ static int keep_entry(const struct cache_entry *ce,\n \t\t      struct unpack_trees_options *o)\n {\n \tadd_entry(o, ce, 0, 0);\n+\tif (ce_stage(ce))\n+\t\tinvalidate_ce_path(ce, o);\n \treturn 1;\n }\n \n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355985","messageId":"20180818144128.19361-5-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-1-pclouds@gmail.com","subject":"[PATCH v5 4/7] unpack-trees: reduce malloc in cache-tree walk","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-18T14:41:25Z","receivedAt":"2018-08-18T14:41:49Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This is a micro optimization that probably only shines on repos with\ndeep directory structure. Instead of allocating and freeing a new\ncache_entry in every iteration, we reuse the last one and only update\nthe parts that are new each iteration.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n unpack-trees.c | 29 ++++++++++++++++++++---------\n 1 file changed, 20 insertions(+), 9 deletions(-)\n\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 8376663b59..dbef6e1b8a 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -685,6 +685,8 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n {\n \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n \tstruct unpack_trees_options *o = info->data;\n+\tstruct cache_entry *tree_ce = NULL;\n+\tint ce_len = 0;\n \tint i, d;\n \n \tif (!o->merge)\n@@ -699,30 +701,39 @@ static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \t * get here in the first place.\n \t */\n \tfor (i = 0; i < nr_entries; i++) {\n-\t\tstruct cache_entry *tree_ce;\n-\t\tint len, rc;\n+\t\tint new_ce_len, len, rc;\n \n \t\tsrc[0] = o->src_index->cache[pos + i];\n \n \t\tlen = ce_namelen(src[0]);\n-\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n+\t\tnew_ce_len = cache_entry_size(len);\n+\n+\t\tif (new_ce_len > ce_len) {\n+\t\t\tnew_ce_len <<= 1;\n+\t\t\ttree_ce = xrealloc(tree_ce, new_ce_len);\n+\t\t\tmemset(tree_ce, 0, new_ce_len);\n+\t\t\tce_len = new_ce_len;\n+\n+\t\t\ttree_ce->ce_flags = create_ce_flags(0);\n+\n+\t\t\tfor (d = 1; d <= nr_names; d++)\n+\t\t\t\tsrc[d] = tree_ce;\n+\t\t}\n \n \t\ttree_ce->ce_mode = src[0]->ce_mode;\n-\t\ttree_ce->ce_flags = create_ce_flags(0);\n \t\ttree_ce->ce_namelen = len;\n \t\toidcpy(&tree_ce->oid, &src[0]->oid);\n \t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n \n-\t\tfor (d = 1; d <= nr_names; d++)\n-\t\t\tsrc[d] = tree_ce;\n-\n \t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n-\t\tfree(tree_ce);\n-\t\tif (rc < 0)\n+\t\tif (rc < 0) {\n+\t\t\tfree(tree_ce);\n \t\t\treturn rc;\n+\t\t}\n \n \t\tmark_ce_used(src[0], o);\n \t}\n+\tfree(tree_ce);\n \tif (o->debug_unpack)\n \t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n \t\t       nr_entries,\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"355984","messageId":"20180818144128.19361-8-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-1-pclouds@gmail.com","subject":"[PATCH v5 7/7] cache-tree: verify valid cache-tree in the test suite","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-18T14:41:28Z","receivedAt":"2018-08-18T14:41:50Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This makes sure that cache-tree is consistent with the index. The main\npurpose is to catch potential problems by saving the index in\nunpack_trees() but the line in write_index() would also help spot\nmissing invalidation in other code.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n cache-tree.c   | 78 ++++++++++++++++++++++++++++++++++++++++++++++++++\n cache-tree.h   |  1 +\n read-cache.c   |  3 ++\n t/test-lib.sh  |  6 ++++\n unpack-trees.c |  2 ++\n 5 files changed, 90 insertions(+)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex caafbff2ff..c3c206427c 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -4,6 +4,7 @@\n #include \"tree-walk.h\"\n #include \"cache-tree.h\"\n #include \"object-store.h\"\n+#include \"replace-object.h\"\n \n #ifndef DEBUG\n #define DEBUG 0\n@@ -732,3 +733,80 @@ int update_main_cache_tree(int flags)\n \t\tthe_index.cache_tree = cache_tree();\n \treturn cache_tree_update(&the_index, flags);\n }\n+\n+static void verify_one(struct index_state *istate,\n+\t\t       struct cache_tree *it,\n+\t\t       struct strbuf *path)\n+{\n+\tint i, pos, len = path->len;\n+\tstruct strbuf tree_buf = STRBUF_INIT;\n+\tstruct object_id new_oid;\n+\n+\tfor (i = 0; i < it->subtree_nr; i++) {\n+\t\tstrbuf_addf(path, \"%s/\", it->down[i]->name);\n+\t\tverify_one(istate, it->down[i]->cache_tree, path);\n+\t\tstrbuf_setlen(path, len);\n+\t}\n+\n+\tif (it->entry_count < 0 ||\n+\t    /* no verification on tests (t7003) that replace trees */\n+\t    lookup_replace_object(the_repository, &it->oid) != &it->oid)\n+\t\treturn;\n+\n+\tif (path->len) {\n+\t\tpos = index_name_pos(istate, path->buf, path->len);\n+\t\tpos = -pos - 1;\n+\t} else {\n+\t\tpos = 0;\n+\t}\n+\n+\ti = 0;\n+\twhile (i < it->entry_count) {\n+\t\tstruct cache_entry *ce = istate->cache[pos + i];\n+\t\tconst char *slash;\n+\t\tstruct cache_tree_sub *sub = NULL;\n+\t\tconst struct object_id *oid;\n+\t\tconst char *name;\n+\t\tunsigned mode;\n+\t\tint entlen;\n+\n+\t\tif (ce->ce_flags & (CE_STAGEMASK | CE_INTENT_TO_ADD | CE_REMOVE))\n+\t\t\tBUG(\"%s with flags 0x%x should not be in cache-tree\",\n+\t\t\t    ce->name, ce->ce_flags);\n+\t\tname = ce->name + path->len;\n+\t\tslash = strchr(name, '/');\n+\t\tif (slash) {\n+\t\t\tentlen = slash - name;\n+\t\t\tsub = find_subtree(it, ce->name + path->len, entlen, 0);\n+\t\t\tif (!sub || sub->cache_tree->entry_count < 0)\n+\t\t\t\tBUG(\"bad subtree '%.*s'\", entlen, name);\n+\t\t\toid = &sub->cache_tree->oid;\n+\t\t\tmode = S_IFDIR;\n+\t\t\ti += sub->cache_tree->entry_count;\n+\t\t} else {\n+\t\t\toid = &ce->oid;\n+\t\t\tmode = ce->ce_mode;\n+\t\t\tentlen = ce_namelen(ce) - path->len;\n+\t\t\ti++;\n+\t\t}\n+\t\tstrbuf_addf(&tree_buf, \"%o %.*s%c\", mode, entlen, name, '\\0');\n+\t\tstrbuf_add(&tree_buf, oid->hash, the_hash_algo->rawsz);\n+\t}\n+\thash_object_file(tree_buf.buf, tree_buf.len, tree_type, &new_oid);\n+\tif (oidcmp(&new_oid, &it->oid))\n+\t\tBUG(\"cache-tree for path %.*s does not match. \"\n+\t\t    \"Expected %s got %s\", len, path->buf,\n+\t\t    oid_to_hex(&new_oid), oid_to_hex(&it->oid));\n+\tstrbuf_setlen(path, len);\n+\tstrbuf_release(&tree_buf);\n+}\n+\n+void cache_tree_verify(struct index_state *istate)\n+{\n+\tstruct strbuf path = STRBUF_INIT;\n+\n+\tif (!istate->cache_tree)\n+\t\treturn;\n+\tverify_one(istate, istate->cache_tree, &path);\n+\tstrbuf_release(&path);\n+}\ndiff --git a/cache-tree.h b/cache-tree.h\nindex 9799e894f7..c1fde531f9 100644\n--- a/cache-tree.h\n+++ b/cache-tree.h\n@@ -32,6 +32,7 @@ struct cache_tree *cache_tree_read(const char *buffer, unsigned long size);\n \n int cache_tree_fully_valid(struct cache_tree *);\n int cache_tree_update(struct index_state *, int);\n+void cache_tree_verify(struct index_state *);\n \n int update_main_cache_tree(int);\n \ndiff --git a/read-cache.c b/read-cache.c\nindex 5ce40f39b3..41f313bc9e 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2744,6 +2744,9 @@ int write_locked_index(struct index_state *istate, struct lock_file *lock,\n \tint new_shared_index, ret;\n \tstruct split_index *si = istate->split_index;\n \n+\tif (git_env_bool(\"GIT_TEST_CHECK_CACHE_TREE\", 0))\n+\t\tcache_tree_verify(istate);\n+\n \tif ((flags & SKIP_IF_UNCHANGED) && !istate->cache_changed) {\n \t\tif (flags & COMMIT_LOCK)\n \t\t\trollback_lock_file(lock);\ndiff --git a/t/test-lib.sh b/t/test-lib.sh\nindex 78f7097746..5b50f6e2e6 100644\n--- a/t/test-lib.sh\n+++ b/t/test-lib.sh\n@@ -1083,6 +1083,12 @@ else\n \ttest_set_prereq C_LOCALE_OUTPUT\n fi\n \n+if test -z \"$GIT_TEST_CHECK_CACHE_TREE\"\n+then\n+\tGIT_TEST_CHECK_CACHE_TREE=true\n+\texport GIT_TEST_CHECK_CACHE_TREE\n+fi\n+\n test_lazy_prereq PIPE '\n \t# test whether the filesystem supports FIFOs\n \ttest_have_prereq !MINGW,!CYGWIN &&\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex bc43922922..3394540842 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -1578,6 +1578,8 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n \tif (o->dst_index) {\n \t\tmove_index_extensions(&o->result, o->src_index);\n \t\tif (!ret) {\n+\t\t\tif (git_env_bool(\"GIT_TEST_CHECK_CACHE_TREE\", 0))\n+\t\t\t\tcache_tree_verify(&o->result);\n \t\t\tif (!o->result.cache_tree)\n \t\t\t\to->result.cache_tree = cache_tree();\n \t\t\tif (!cache_tree_fully_valid(o->result.cache_tree))\n-- \n2.18.0.1004.g6639190530\n\n"},{"id":"356008","messageId":"CABPp-BGn_BauwEGuHBR92CBQ-sOuS_tDwW5uDdmjrqpF2jxwxA@mail.gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-8-pclouds@gmail.com","subject":"Re: [PATCH v5 7/7] cache-tree: verify valid cache-tree in the test suite","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-08-18T21:45:00Z","receivedAt":"2018-08-18T21:45:20Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Aug 18, 2018 at 7:41 AM Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n...\n> diff --git a/read-cache.c b/read-cache.c\n> index 5ce40f39b3..41f313bc9e 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -2744,6 +2744,9 @@ int write_locked_index(struct index_state *istate, struct lock_file *lock,\n>         int new_shared_index, ret;\n>         struct split_index *si = istate->split_index;\n>\n> +       if (git_env_bool(\"GIT_TEST_CHECK_CACHE_TREE\", 0))\n> +               cache_tree_verify(istate);\n> +\n>         if ((flags & SKIP_IF_UNCHANGED) && !istate->cache_changed) {\n>                 if (flags & COMMIT_LOCK)\n>                         rollback_lock_file(lock);\n> diff --git a/t/test-lib.sh b/t/test-lib.sh\n> index 78f7097746..5b50f6e2e6 100644\n> --- a/t/test-lib.sh\n> +++ b/t/test-lib.sh\n> @@ -1083,6 +1083,12 @@ else\n>         test_set_prereq C_LOCALE_OUTPUT\n>  fi\n>\n> +if test -z \"$GIT_TEST_CHECK_CACHE_TREE\"\n> +then\n> +       GIT_TEST_CHECK_CACHE_TREE=true\n> +       export GIT_TEST_CHECK_CACHE_TREE\n> +fi\n> +\n>  test_lazy_prereq PIPE '\n>         # test whether the filesystem supports FIFOs\n>         test_have_prereq !MINGW,!CYGWIN &&\n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index bc43922922..3394540842 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -1578,6 +1578,8 @@ int unpack_trees(unsigned len, struct tree_desc *t, struct unpack_trees_options\n>         if (o->dst_index) {\n>                 move_index_extensions(&o->result, o->src_index);\n>                 if (!ret) {\n> +                       if (git_env_bool(\"GIT_TEST_CHECK_CACHE_TREE\", 0))\n> +                               cache_tree_verify(&o->result);\n>                         if (!o->result.cache_tree)\n>                                 o->result.cache_tree = cache_tree();\n>                         if (!cache_tree_fully_valid(o->result.cache_tree))\n> --\n> 2.18.0.1004.g6639190530\n\nShould documentation of GIT_TEST_CHECK_CACHE_TREE be added in\nt/README, int the \"Running tests with special setups\" section?\n"},{"id":"356010","messageId":"CABPp-BEK-oFWBbjgZBCDaixtmnxrTYtvHnPeT5enHBr9XJ8fGg@mail.gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-1-pclouds@gmail.com","subject":"Re: [PATCH v5 0/7] Speed up unpack_trees()","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2018-08-18T22:01:27Z","receivedAt":"2018-08-18T22:01:40Z","isPatch":true,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sat, Aug 18, 2018 at 7:41 AM Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n>\n> v5 fixes some minor comments from round 4 and a big mistake in 5/5.\n> Junio's scary feeling turns out true. There is a missing invalidation\n> in keep_entry() which is not added in 6/7. 7/7 makes sure that similar\n\nI'm having trouble parsing this.  Did you mean \"...which is now\nadded...\"?  Also, if 6/7 represents a fix to the \"big mistake in 5/5\",\nwhy is 6/7 separate from 5/7 instead of squashed in?\n\n> problems will not slip through.\n>\n> I had to rebase this series on top of 'master' because 7/7 caught a\n> bad cache-tree situation that has been fixed by Elijah in ad3762042a\n\nCool, glad that helped.\n\n...\n> Nguyễn Thái Ngọc Duy (7):\n>   trace.h: support nested performance tracing\n>   unpack-trees: add performance tracing\n>   unpack-trees: optimize walking same trees with cache-tree\n>   unpack-trees: reduce malloc in cache-tree walk\n>   unpack-trees: reuse (still valid) cache-tree from src_index\n>   unpack-trees: add missing cache invalidation\n>   cache-tree: verify valid cache-tree in the test suite\n\nI read through the new series and only had one small comment.  I'm not\nup to speed on cache-tree stuff, still, so don't feel qualified to\ngive an Ack on it.\n"},{"id":"356028","messageId":"CACsJy8AmX48=2N-MsXcnaFrCybArj8YaCpc7+LvahUQQBvSXAQ@mail.gmail.com","threadId":"48913","inReplyTo":"CABPp-BEK-oFWBbjgZBCDaixtmnxrTYtvHnPeT5enHBr9XJ8fGg@mail.gmail.com","subject":"Re: [PATCH v5 0/7] Speed up unpack_trees()","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-19T05:09:53Z","receivedAt":"2018-08-19T05:12:47Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Aug 19, 2018 at 12:01 AM Elijah Newren <newren@gmail.com> wrote:\n>\n> On Sat, Aug 18, 2018 at 7:41 AM Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n> >\n> > v5 fixes some minor comments from round 4 and a big mistake in 5/5.\n> > Junio's scary feeling turns out true. There is a missing invalidation\n> > in keep_entry() which is not added in 6/7. 7/7 makes sure that similar\n>\n> I'm having trouble parsing this.  Did you mean \"...which is now\n> added...\"?\n\nOops. Yes.\n\n>  Also, if 6/7 represents a fix to the \"big mistake in 5/5\",\n> why is 6/7 separate from 5/7 instead of squashed in?\n\nI felt that was cramming up too much in the commit message. But if\nit's the right thing to do, I'll reroll and combine 5/7 and 6/7 .\n-- \nDuy\n"},{"id":"356083","messageId":"b9c78c72-5d3a-9bdd-f3eb-b383ca6676a7@gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-4-pclouds@gmail.com","subject":"Re: [PATCH v5 3/7] unpack-trees: optimize walking same trees with cache-tree","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-20T12:43:22Z","receivedAt":"2018-08-20T12:43:28Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/18/2018 10:41 AM, Nguyễn Thái Ngọc Duy wrote:\n> In order to merge one or many trees with the index, unpack-trees code\n> walks multiple trees in parallel with the index and performs n-way\n> merge. If we find out at start of a directory that all trees are the\n> same (by comparing OID) and cache-tree happens to be available for\n> that directory as well, we could avoid walking the trees because we\n> already know what these trees contain: it's flattened in what's called\n> \"the index\".\n> \n> The upside is of course a lot less I/O since we can potentially skip\n> lots of trees (think subtrees). We also save CPU because we don't have\n> to inflate and apply the deltas. The downside is of course more\n> fragile code since the logic in some functions are now duplicated\n> elsewhere.\n> \n> \"checkout -\" with this patch on webkit.git (275k files):\n> \n>      baseline      new\n>    --------------------------------------------------------------------\n>      0.056651714   0.080394752 s:  read cache .git/index\n>      0.183101080   0.216010838 s:  preload index\n>      0.008584433   0.008534301 s:  refresh index\n>      0.633767589   0.251992198 s:   traverse_trees\n>      0.340265448   0.377031383 s:   check_updates\n>      0.381884638   0.372768105 s:   cache_tree_update\n>      1.401562947   1.045887251 s:  unpack_trees\n>      0.338687914   0.314983512 s:  write index, changed mask = 2e\n>      0.411927922   0.062572653 s:    traverse_trees\n>      0.000023335   0.000022544 s:    check_updates\n>      0.423697246   0.073795585 s:   unpack_trees\n>      0.423708360   0.073807557 s:  diff-index\n>      2.559524127   1.938191592 s: git command: git checkout -\n> \n> Another measurement from Ben's running \"git checkout\" with over 500k\n> trees (on the whole series):\n> \n>      baseline        new\n>    ----------------------------------------------------------------------\n>      0.535510167     0.556558733     s: read cache .git/index\n>      0.3057373       0.3147105       s: initialize name hash\n>      0.0184082       0.023558433     s: preload index\n>      0.086910967     0.089085967     s: refresh index\n>      7.889590767     2.191554433     s: unpack trees\n>      0.120760833     0.131941267     s: update worktree after a merge\n>      2.2583504       2.572663167     s: repair cache-tree\n>      0.8916137       0.959495233     s: write index, changed mask = 28\n>      3.405199233     0.2710663       s: unpack trees\n>      0.000999667     0.0021554       s: update worktree after a merge\n>      3.4063306       0.273318333     s: diff-index\n>      16.9524923      9.462943133     s: git command: git.exe checkout\n> \n> This command calls unpack_trees() twice, the first time on 2way merge\n> and the second 1way merge. In both times, \"unpack trees\" time is\n> reduced to one third. Overall time reduction is not that impressive of\n> course because index operations take a big chunk. And there's that\n> repair cache-tree line.\n> \n> PS. A note about cache-tree invalidation and the use of it in this\n> code.\n> \n> We do invalidate cache-tree in _source_ index when we add new entries\n> to the (temporary) \"result\" index. But we also use the cache-tree from\n> source index in this optimization. Does this mean we end up having no\n> cache-tree in the source index to activate this optimization?\n> \n> The answer is twisted: the order of finding a good cache-tree and\n> invalidating it matters. In this case we check for a good cache-tree\n> first in all_trees_same_as_cache_tree(), then we start to merge things\n> and potentially invalidate that same cache-tree in the process. Since\n> cache-tree invalidation happens after the optimization kicks in, we're\n> still good. But we may lose that cache-tree at the very first\n> call_unpack_fn() call in traverse_by_cache_tree().\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>   unpack-trees.c | 127 +++++++++++++++++++++++++++++++++++++++++++++++++\n>   1 file changed, 127 insertions(+)\n> \n> diff --git a/unpack-trees.c b/unpack-trees.c\n> index 6d9f692ea6..8376663b59 100644\n> --- a/unpack-trees.c\n> +++ b/unpack-trees.c\n> @@ -635,6 +635,102 @@ static inline int are_same_oid(struct name_entry *name_j, struct name_entry *nam\n>   \treturn name_j->oid && name_k->oid && !oidcmp(name_j->oid, name_k->oid);\n>   }\n>   \n> +static int all_trees_same_as_cache_tree(int n, unsigned long dirmask,\n> +\t\t\t\t\tstruct name_entry *names,\n> +\t\t\t\t\tstruct traverse_info *info)\n> +{\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint i;\n> +\n> +\tif (!o->merge || dirmask != ((1 << n) - 1))\n> +\t\treturn 0;\n> +\n> +\tfor (i = 1; i < n; i++)\n> +\t\tif (!are_same_oid(names, names + i))\n> +\t\t\treturn 0;\n> +\n> +\treturn cache_tree_matches_traversal(o->src_index->cache_tree, names, info);\n> +}\n> +\n> +static int index_pos_by_traverse_info(struct name_entry *names,\n> +\t\t\t\t      struct traverse_info *info)\n> +{\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint len = traverse_path_len(info, names);\n> +\tchar *name = xmalloc(len + 1 /* slash */ + 1 /* NUL */);\n> +\tint pos;\n> +\n> +\tmake_traverse_path(name, info, names);\n> +\tname[len++] = '/';\n> +\tname[len] = '\\0';\n> +\tpos = index_name_pos(o->src_index, name, len);\n> +\tif (pos >= 0)\n> +\t\tBUG(\"This is a directory and should not exist in index\");\n> +\tpos = -pos - 1;\n> +\tif (!starts_with(o->src_index->cache[pos]->name, name) ||\n> +\t    (pos > 0 && starts_with(o->src_index->cache[pos-1]->name, name)))\n> +\t\tBUG(\"pos must point at the first entry in this directory\");\n> +\tfree(name);\n> +\treturn pos;\n> +}\n> +\n> +/*\n> + * Fast path if we detect that all trees are the same as cache-tree at this\n> + * path. We'll walk these trees recursively using cache-tree/index instead of\n\nnit, not worth a re-roll\n\n\"We'll walk these trees in an iterative loop using cache-tree/index...\"\n\n> + * ODB since already know what these trees contain.\n> + */\n> +static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n> +\t\t\t\t  struct name_entry *names,\n> +\t\t\t\t  struct traverse_info *info)\n> +{\n> +\tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n> +\tstruct unpack_trees_options *o = info->data;\n> +\tint i, d;\n> +\n> +\tif (!o->merge)\n> +\t\tBUG(\"We need cache-tree to do this optimization\");\n> +\n> +\t/*\n> +\t * Do what unpack_callback() and unpack_nondirectories() normally\n> +\t * do. But we walk all paths in an iterative loop instead.\n> +\t *\n> +\t * D/F conflicts and higher stage entries are not a concern\n> +\t * because cache-tree would be invalidated and we would never\n> +\t * get here in the first place.\n> +\t */\n> +\tfor (i = 0; i < nr_entries; i++) {\n> +\t\tstruct cache_entry *tree_ce;\n> +\t\tint len, rc;\n> +\n> +\t\tsrc[0] = o->src_index->cache[pos + i];\n> +\n> +\t\tlen = ce_namelen(src[0]);\n> +\t\ttree_ce = xcalloc(1, cache_entry_size(len));\n> +\n> +\t\ttree_ce->ce_mode = src[0]->ce_mode;\n> +\t\ttree_ce->ce_flags = create_ce_flags(0);\n> +\t\ttree_ce->ce_namelen = len;\n> +\t\toidcpy(&tree_ce->oid, &src[0]->oid);\n> +\t\tmemcpy(tree_ce->name, src[0]->name, len + 1);\n> +\n> +\t\tfor (d = 1; d <= nr_names; d++)\n> +\t\t\tsrc[d] = tree_ce;\n> +\n> +\t\trc = call_unpack_fn((const struct cache_entry * const *)src, o);\n> +\t\tfree(tree_ce);\n> +\t\tif (rc < 0)\n> +\t\t\treturn rc;\n> +\n> +\t\tmark_ce_used(src[0], o);\n> +\t}\n> +\tif (o->debug_unpack)\n> +\t\tprintf(\"Unpacked %d entries from %s to %s using cache-tree\\n\",\n> +\t\t       nr_entries,\n> +\t\t       o->src_index->cache[pos]->name,\n> +\t\t       o->src_index->cache[pos + nr_entries - 1]->name);\n> +\treturn 0;\n> +}\n> +\n>   static int traverse_trees_recursive(int n, unsigned long dirmask,\n>   \t\t\t\t    unsigned long df_conflicts,\n>   \t\t\t\t    struct name_entry *names,\n> @@ -646,6 +742,27 @@ static int traverse_trees_recursive(int n, unsigned long dirmask,\n>   \tvoid *buf[MAX_UNPACK_TREES];\n>   \tstruct traverse_info newinfo;\n>   \tstruct name_entry *p;\n> +\tint nr_entries;\n> +\n> +\tnr_entries = all_trees_same_as_cache_tree(n, dirmask, names, info);\n> +\tif (nr_entries > 0) {\n> +\t\tstruct unpack_trees_options *o = info->data;\n> +\t\tint pos = index_pos_by_traverse_info(names, info);\n> +\n> +\t\tif (!o->merge || df_conflicts)\n> +\t\t\tBUG(\"Wrong condition to get here buddy\");\n> +\n> +\t\t/*\n> +\t\t * All entries up to 'pos' must have been processed\n> +\t\t * (i.e. marked CE_UNPACKED) at this point. But to be safe,\n> +\t\t * save and restore cache_bottom anyway to not miss\n> +\t\t * unprocessed entries before 'pos'.\n> +\t\t */\n> +\t\tbottom = o->cache_bottom;\n> +\t\tret = traverse_by_cache_tree(pos, nr_entries, n, names, info);\n> +\t\to->cache_bottom = bottom;\n> +\t\treturn ret;\n> +\t}\n>   \n>   \tp = names;\n>   \twhile (!p->mode)\n> @@ -812,6 +929,11 @@ static struct cache_entry *create_ce_entry(const struct traverse_info *info,\n>   \treturn ce;\n>   }\n>   \n> +/*\n> + * Note that traverse_by_cache_tree() duplicates some logic in this function\n> + * without actually calling it. If you change the logic here you may need to\n> + * check and change there as well.\n> + */\n>   static int unpack_nondirectories(int n, unsigned long mask,\n>   \t\t\t\t unsigned long dirmask,\n>   \t\t\t\t struct cache_entry **src,\n> @@ -1004,6 +1126,11 @@ static void debug_unpack_callback(int n,\n>   \t\tdebug_name_entry(i, names + i);\n>   }\n>   \n> +/*\n> + * Note that traverse_by_cache_tree() duplicates some logic in this function\n> + * without actually calling it. If you change the logic here you may need to\n> + * check and change there as well.\n> + */\n>   static int unpack_callback(int n, unsigned long mask, unsigned long dirmask, struct name_entry *names, struct traverse_info *info)\n>   {\n>   \tstruct cache_entry *src[MAX_UNPACK_TREES + 1] = { NULL, };\n> \n"},{"id":"356517","messageId":"20180825121848.11606-1-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180818144128.19361-1-pclouds@gmail.com","subject":"[PATCH] Document update for nd/unpack-trees-with-cache-tree","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-25T12:18:48Z","receivedAt":"2018-08-25T12:18:57Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Fix an incorrect comment in the new code added in b4da37380b\n(unpack-trees: optimize walking same trees with cache-tree -\n2018-08-18) and document about the new test variable that is enabled\nby default in test-lib.sh in 4592e6080f (cache-tree: verify valid\ncache-tree in the test suite - 2018-08-18)\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n On top of nd/unpack-trees-with-cache-tree. Incremental update since\n this topic has entered 'next'\n\n t/README       | 4 ++++\n unpack-trees.c | 4 ++--\n 2 files changed, 6 insertions(+), 2 deletions(-)\n\ndiff --git a/t/README b/t/README\nindex 8373a27fea..0e7cc23734 100644\n--- a/t/README\n+++ b/t/README\n@@ -315,6 +315,10 @@ packs on demand. This normally only happens when the object size is\n over 2GB. This variable forces the code path on any object larger than\n <n> bytes.\n \n+GIT_TEST_VALIDATE_INDEX_CACHE_ENTRIES=<boolean> checks that cache-tree\n+records are valid when the index is written out or after a merge. This\n+is mostly to catch missing invalidation. Default is true.\n+\n Naming Tests\n ------------\n \ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 3394540842..5a18f36143 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -676,8 +676,8 @@ static int index_pos_by_traverse_info(struct name_entry *names,\n \n /*\n  * Fast path if we detect that all trees are the same as cache-tree at this\n- * path. We'll walk these trees recursively using cache-tree/index instead of\n- * ODB since already know what these trees contain.\n+ * path. We'll walk these trees in an iteractive loop using cache-tree/index\n+ * instead of ODB since already know what these trees contain.\n  */\n static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \t\t\t\t  struct name_entry *names,\n-- \n2.19.0.rc0.337.ge906d732e7\n\n"},{"id":"356518","messageId":"CAN0heSqcNBJ6-6YBFixd1+h1fg3o=2fJZXhmDehNLLYCN=RFqg@mail.gmail.com","threadId":"48913","inReplyTo":"20180825121848.11606-1-pclouds@gmail.com","subject":"Re: [PATCH] Document update for nd/unpack-trees-with-cache-tree","fromName":"Martin Ågren","fromEmail":"martin.agren@gmail.com","sentAt":"2018-08-25T12:31:09Z","receivedAt":"2018-08-25T12:39:55Z","isPatch":true,"sender":{"key":"martin.agren@gmail.com","avatar":null},"body":"On Sat, 25 Aug 2018 at 14:22, Nguyễn Thái Ngọc Duy <pclouds@gmail.com> wrote:\n>   * Fast path if we detect that all trees are the same as cache-tree at this\n> - * path. We'll walk these trees recursively using cache-tree/index instead of\n> - * ODB since already know what these trees contain.\n> + * path. We'll walk these trees in an iteractive loop using cache-tree/index\n> + * instead of ODB since already know what these trees contain.\n\ns/iteractive/iterative/ (i.e., drop \"c\")\n\nNot new, but still: s/already/we already/\n\nMartin\n"},{"id":"356519","messageId":"20180825130209.31231-1-pclouds@gmail.com","threadId":"48913","inReplyTo":"20180825121848.11606-1-pclouds@gmail.com","subject":"[PATCH v2] Document update for nd/unpack-trees-with-cache-tree","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-25T13:02:09Z","receivedAt":"2018-08-25T13:02:35Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Fix an incorrect comment in the new code added in b4da37380b\n(unpack-trees: optimize walking same trees with cache-tree -\n2018-08-18) and document about the new test variable that is enabled\nby default in test-lib.sh in 4592e6080f (cache-tree: verify valid\ncache-tree in the test suite - 2018-08-18)\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Some more typo fix, found by Martin.\n\n t/README       | 4 ++++\n unpack-trees.c | 4 ++--\n 2 files changed, 6 insertions(+), 2 deletions(-)\n\ndiff --git a/t/README b/t/README\nindex 8373a27fea..0e7cc23734 100644\n--- a/t/README\n+++ b/t/README\n@@ -315,6 +315,10 @@ packs on demand. This normally only happens when the object size is\n over 2GB. This variable forces the code path on any object larger than\n <n> bytes.\n \n+GIT_TEST_VALIDATE_INDEX_CACHE_ENTRIES=<boolean> checks that cache-tree\n+records are valid when the index is written out or after a merge. This\n+is mostly to catch missing invalidation. Default is true.\n+\n Naming Tests\n ------------\n \ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex 3394540842..515c374373 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -676,8 +676,8 @@ static int index_pos_by_traverse_info(struct name_entry *names,\n \n /*\n  * Fast path if we detect that all trees are the same as cache-tree at this\n- * path. We'll walk these trees recursively using cache-tree/index instead of\n- * ODB since already know what these trees contain.\n+ * path. We'll walk these trees in an iterative loop using cache-tree/index\n+ * instead of ODB since we already know what these trees contain.\n  */\n static int traverse_by_cache_tree(int pos, int nr_entries, int nr_names,\n \t\t\t\t  struct name_entry *names,\n-- \n2.19.0.rc0.337.ge906d732e7\n\n"}]}