{"thread":{"id":"49204","subject":"[PATCH v1] read-cache: speed up index load through parallelization","startedAt":"2018-08-23T15:45:20Z","lastAt":"2018-11-27T00:50:17Z","messageCount":199,"participants":["Ben Peart","Stefan Beller","Junio C Hamano","Duy Nguyen","Nguyễn Thái Ngọc Duy","Torsten Bögershausen","Martin Ågren","SZEDER Gábor","Ramsay Jones","Jeff King","Jonathan Nieder","Jonathan Tan","Ævar Arnfjörð Bjarmason"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"356375","messageId":"20180823154053.20212-1-benpeart@microsoft.com","threadId":"49204","inReplyTo":null,"subject":"[PATCH v1] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"ben.peart@microsoft.com","sentAt":"2018-08-23T15:41:12Z","receivedAt":"2018-08-23T15:45:20Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by creating\nmultiple threads to divide the work of loading and converting the cache\nentries across all available CPU cores.\n\nIt accomplishes this by having the primary thread loop across the index file\ntracking the offset and (for V4 indexes) expanding the name. It creates a\nthread to process each block of entries as it comes to them. Once the\nthreads are complete and the cache entries are loaded, the rest of the\nextensions can be loaded and processed normally on the primary thread.\n\nPerformance impact:\n\nread cache .git/index times on a synthetic repo with:\n\n100,000 entries\nFALSE       TRUE        Savings     %Savings\n0.014798767 0.009580433 0.005218333 35.26%\n\n1,000,000 entries\nFALSE       TRUE        Savings     %Savings\n0.240896533 0.1751243   0.065772233 27.30%\n\nread cache .git/index times on an actual repo with:\n\n~3M entries\nFALSE       TRUE        Savings     %Savings\n0.59898098  0.4513169   0.14766408  24.65%\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n\nNotes:\n    Base Ref: master\n    Web-Diff: https://github.com/benpeart/git/commit/67a700419b\n    Checkout: git fetch https://github.com/benpeart/git read-index-multithread-v1 && git checkout 67a700419b\n\n Documentation/config.txt |   8 ++\n config.c                 |  13 +++\n config.h                 |   1 +\n read-cache.c             | 218 ++++++++++++++++++++++++++++++++++-----\n 4 files changed, 216 insertions(+), 24 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex 1c42364988..3344685cc4 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -899,6 +899,14 @@ relatively high IO latencies.  When enabled, Git will do the\n index comparison to the filesystem data in parallel, allowing\n overlapping IO's.  Defaults to true.\n \n+core.fastIndex::\n+       Enable parallel index loading\n++\n+This can speed up operations like 'git diff' and 'git status' especially\n+when the index is very large.  When enabled, Git will do the index\n+loading from the on disk format to the in-memory format in parallel.\n+Defaults to true.\n+\n core.createObject::\n \tYou can set this to 'link', in which case a hardlink followed by\n \ta delete of the source are used to make sure that object creation\ndiff --git a/config.c b/config.c\nindex 9a0b10d4bc..883092fdd3 100644\n--- a/config.c\n+++ b/config.c\n@@ -2289,6 +2289,19 @@ int git_config_get_fsmonitor(void)\n \treturn 0;\n }\n \n+int git_config_get_fast_index(void)\n+{\n+\tint val;\n+\n+\tif (!git_config_get_maybe_bool(\"core.fastindex\", &val))\n+\t\treturn val;\n+\n+\tif (getenv(\"GIT_FASTINDEX_TEST\"))\n+\t\treturn 1;\n+\n+\treturn -1; /* default value */\n+}\n+\n NORETURN\n void git_die_config_linenr(const char *key, const char *filename, int linenr)\n {\ndiff --git a/config.h b/config.h\nindex ab46e0165d..74ca4e7db5 100644\n--- a/config.h\n+++ b/config.h\n@@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n extern int git_config_get_split_index(void);\n extern int git_config_get_max_percent_split_change(void);\n extern int git_config_get_fsmonitor(void);\n+extern int git_config_get_fast_index(void);\n \n /* This dies if the configured or default date is in the future */\n extern int git_config_get_expiry(const char *key, const char **output);\ndiff --git a/read-cache.c b/read-cache.c\nindex 7b1354d759..0fa7e1a04c 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -24,6 +24,10 @@\n #include \"utf8.h\"\n #include \"fsmonitor.h\"\n \n+#ifndef min\n+#define min(a,b) (((a) < (b)) ? (a) : (b))\n+#endif\n+\n /* Mask for the name length in ce_flags in the on-disk index */\n \n #define CE_NAMEMASK  (0x0fff)\n@@ -1889,16 +1893,203 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+static unsigned long load_cache_entry_block(struct index_state *istate, struct mem_pool *ce_mem_pool, int offset, int nr, void *mmap, unsigned long start_offset, struct strbuf *previous_name)\n+{\n+\tint i;\n+\tunsigned long src_offset = start_offset;\n+\n+\tfor (i = offset; i < offset + nr; i++) {\n+\t\tstruct ondisk_cache_entry *disk_ce;\n+\t\tstruct cache_entry *ce;\n+\t\tunsigned long consumed;\n+\n+\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tset_index_entry(istate, i, ce);\n+\n+\t\tsrc_offset += consumed;\n+\t}\n+\treturn src_offset - start_offset;\n+}\n+\n+static unsigned long load_all_cache_entries(struct index_state *istate, void *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tunsigned long consumed;\n+\n+\tif (istate->version == 4) {\n+\t\tprevious_name = &previous_name_buf;\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n+\t} else {\n+\t\tprevious_name = NULL;\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n+\t}\n+\n+\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool, 0, istate->cache_nr, mmap, src_offset, previous_name);\n+\tstrbuf_release(&previous_name_buf);\n+\treturn consumed;\n+}\n+\n+#ifdef NO_PTHREADS\n+\n+#define load_cache_entries load_all_cache_entries\n+\n+#else\n+\n+#include \"thread-utils.h\"\n+\n+/*\n+* Mostly randomly chosen maximum thread counts: we\n+* cap the parallelism to online_cpus() threads, and we want\n+* to have at least 7500 cache entries per thread for it to\n+* be worth starting a thread.\n+*/\n+#define THREAD_COST\t\t(7500)\n+\n+struct load_cache_entries_thread_data\n+{\n+\tpthread_t pthread;\n+\tstruct index_state *istate;\n+\tstruct mem_pool *ce_mem_pool;\n+\tint offset, nr;\n+\tvoid *mmap;\n+\tunsigned long start_offset;\n+\tstruct strbuf previous_name_buf;\n+\tstruct strbuf *previous_name;\n+\tunsigned long consumed;\t/* return # of bytes in index file processed */\n+};\n+\n+/*\n+* A thread proc to run the load_cache_entries() computation\n+* across multiple background threads.\n+*/\n+static void *load_cache_entries_thread(void *_data)\n+{\n+\tstruct load_cache_entries_thread_data *p = _data;\n+\n+\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool, p->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n+\treturn NULL;\n+}\n+\n+static unsigned long load_cache_entries(struct index_state *istate, void *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tstruct load_cache_entries_thread_data *data;\n+\tint threads, cpus, thread_nr;\n+\tunsigned long consumed;\n+\tint i, thread;\n+\n+\tcpus = online_cpus();\n+\tthreads = istate->cache_nr / THREAD_COST;\n+\tif (threads > cpus)\n+\t\tthreads = cpus;\n+\n+\t/* enable testing with fewer than default minimum of entries */\n+\tif ((istate->cache_nr > 1) && (threads < 2) && getenv(\"GIT_FASTINDEX_TEST\"))\n+\t\tthreads = 2;\n+\n+\tif (threads < 2 || !git_config_get_fast_index())\n+\t\treturn load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+\n+\tmem_pool_init(&istate->ce_mem_pool, 0);\n+\tif (istate->version == 4)\n+\t\tprevious_name = &previous_name_buf;\n+\telse\n+\t\tprevious_name = NULL;\n+\n+\tthread_nr = (istate->cache_nr + threads - 1) / threads;\n+\tdata = xcalloc(threads, sizeof(struct load_cache_entries_thread_data));\n+\n+\t/* loop through index entries starting a thread for every thread_nr entries */\n+\tconsumed = thread = 0;\n+\tfor (i = 0; ; i++) {\n+\t\tstruct ondisk_cache_entry *ondisk;\n+\t\tconst char *name;\n+\t\tunsigned int flags;\n+\n+\t\t/* we've reached the begining of a block of cache entries, kick off a thread to process them */\n+\t\tif (0 == i % thread_nr) {\n+\t\t\tstruct load_cache_entries_thread_data *p = &data[thread];\n+\n+\t\t\tp->istate = istate;\n+\t\t\tp->offset = i;\n+\t\t\tp->nr = min(thread_nr, istate->cache_nr - i);\n+\n+\t\t\t/* create a mem_pool for each thread */\n+\t\t\tif (istate->version == 4)\n+\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\t\t\t  estimate_cache_size_from_compressed(p->nr));\n+\t\t\telse\n+\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\t\t\t  estimate_cache_size(mmap_size, p->nr));\n+\n+\t\t\tp->mmap = mmap;\n+\t\t\tp->start_offset = src_offset;\n+\t\t\tif (previous_name) {\n+\t\t\t\tstrbuf_addbuf(&p->previous_name_buf, previous_name);\n+\t\t\t\tp->previous_name = &p->previous_name_buf;\n+\t\t\t}\n+\n+\t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n+\t\t\t\tdie(\"unable to create load_cache_entries_thread\");\n+\t\t\tif (++thread == threads || p->nr != thread_nr)\n+\t\t\t\tbreak;\n+\t\t}\n+\n+\t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\n+\t\t/* On-disk flags are just 16 bits */\n+\t\tflags = get_be16(&ondisk->flags);\n+\n+\t\tif (flags & CE_EXTENDED) {\n+\t\t\tstruct ondisk_cache_entry_extended *ondisk2;\n+\t\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n+\t\t\tname = ondisk2->name;\n+\t\t} else\n+\t\t\tname = ondisk->name;\n+\n+\t\tif (!previous_name) {\n+\t\t\tsize_t len;\n+\n+\t\t\t/* v3 and earlier */\n+\t\t\tlen = flags & CE_NAMEMASK;\n+\t\t\tif (len == CE_NAMEMASK)\n+\t\t\t\tlen = strlen(name);\n+\t\t\tsrc_offset += (flags & CE_EXTENDED) ?\n+\t\t\t\tondisk_cache_entry_extended_size(len) :\n+\t\t\t\tondisk_cache_entry_size(len);\n+\t\t} else\n+\t\t\tsrc_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n+\t}\n+\n+\tfor (i = 0; i < threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = data + i;\n+\t\tif (pthread_join(p->pthread, NULL))\n+\t\t\tdie(\"unable to join load_cache_entries_thread\");\n+\t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n+\t\tstrbuf_release(&p->previous_name_buf);\n+\t\tconsumed += p->consumed;\n+\t}\n+\n+\tfree(data);\n+\tstrbuf_release(&previous_name_buf);\n+\n+\treturn consumed;\n+}\n+\n+#endif\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n-\tint fd, i;\n+\tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1935,29 +2126,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n \tistate->initialized = 1;\n \n-\tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n-\t} else {\n-\t\tprevious_name = NULL;\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n-\t}\n-\n \tsrc_offset = sizeof(*hdr);\n-\tfor (i = 0; i < istate->cache_nr; i++) {\n-\t\tstruct ondisk_cache_entry *disk_ce;\n-\t\tstruct cache_entry *ce;\n-\t\tunsigned long consumed;\n-\n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_name);\n-\t\tset_index_entry(istate, i, ce);\n-\n-\t\tsrc_offset += consumed;\n-\t}\n-\tstrbuf_release(&previous_name_buf);\n+\tsrc_offset += load_cache_entries(istate, mmap, mmap_size, src_offset);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n\nbase-commit: 29d9e3e2c47dd4b5053b0a98c891878d398463e3\n-- \n2.18.0.windows.1\n\n"},{"id":"356387","messageId":"CAGZ79kbXfPPvcQ1rnUdiOqWs5wC2qccGCnf8DvCVnp8QV126MA@mail.gmail.com","threadId":"49204","inReplyTo":"20180823154053.20212-1-benpeart@microsoft.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-08-23T17:31:35Z","receivedAt":"2018-08-23T17:31:49Z","isPatch":true,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Thu, Aug 23, 2018 at 8:45 AM Ben Peart <Ben.Peart@microsoft.com> wrote:\n>\n> This patch helps address the CPU cost of loading the index by creating\n> multiple threads to divide the work of loading and converting the cache\n> entries across all available CPU cores.\n>\n> It accomplishes this by having the primary thread loop across the index file\n> tracking the offset and (for V4 indexes) expanding the name. It creates a\n> thread to process each block of entries as it comes to them. Once the\n> threads are complete and the cache entries are loaded, the rest of the\n> extensions can be loaded and processed normally on the primary thread.\n>\n> Performance impact:\n>\n> read cache .git/index times on a synthetic repo with:\n>\n> 100,000 entries\n> FALSE       TRUE        Savings     %Savings\n> 0.014798767 0.009580433 0.005218333 35.26%\n>\n> 1,000,000 entries\n> FALSE       TRUE        Savings     %Savings\n> 0.240896533 0.1751243   0.065772233 27.30%\n>\n> read cache .git/index times on an actual repo with:\n>\n> ~3M entries\n> FALSE       TRUE        Savings     %Savings\n> 0.59898098  0.4513169   0.14766408  24.65%\n>\n> Signed-off-by: Ben Peart <Ben.Peart@microsoft.com>\n> ---\n>\n> Notes:\n>     Base Ref: master\n>     Web-Diff: https://github.com/benpeart/git/commit/67a700419b\n>     Checkout: git fetch https://github.com/benpeart/git read-index-multithread-v1 && git checkout 67a700419b\n>\n>  Documentation/config.txt |   8 ++\n>  config.c                 |  13 +++\n>  config.h                 |   1 +\n>  read-cache.c             | 218 ++++++++++++++++++++++++++++++++++-----\n>  4 files changed, 216 insertions(+), 24 deletions(-)\n>\n> diff --git a/Documentation/config.txt b/Documentation/config.txt\n> index 1c42364988..3344685cc4 100644\n> --- a/Documentation/config.txt\n> +++ b/Documentation/config.txt\n> @@ -899,6 +899,14 @@ relatively high IO latencies.  When enabled, Git will do the\n>  index comparison to the filesystem data in parallel, allowing\n>  overlapping IO's.  Defaults to true.\n>\n> +core.fastIndex::\n> +       Enable parallel index loading\n> ++\n> +This can speed up operations like 'git diff' and 'git status' especially\n> +when the index is very large.  When enabled, Git will do the index\n> +loading from the on disk format to the in-memory format in parallel.\n> +Defaults to true.\n\n\"fast\" is a non-descriptive word as we try to be fast in any operation?\nMaybe core.parallelIndexReading as that just describes what it\nturns on/off, without second guessing its effects?\n(Are there still computers with just a single CPU, where this would not\nmake it faster? ;-))\n\n\n> +int git_config_get_fast_index(void)\n> +{\n> +       int val;\n> +\n> +       if (!git_config_get_maybe_bool(\"core.fastindex\", &val))\n> +               return val;\n> +\n> +       if (getenv(\"GIT_FASTINDEX_TEST\"))\n> +               return 1;\n\nWe look at this env value just before calling this function,\ncan be write it to only look at the evn variable once?\n\n> +++ b/config.h\n> @@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n>  extern int git_config_get_split_index(void);\n>  extern int git_config_get_max_percent_split_change(void);\n>  extern int git_config_get_fsmonitor(void);\n> +extern int git_config_get_fast_index(void);\n\nOh. nd/no-extern did not cover config.h\n\n\n>\n> +#ifndef min\n> +#define min(a,b) (((a) < (b)) ? (a) : (b))\n> +#endif\n\nWe do not have a minimum function in the tree,\nexcept for xdiff/xmacros.h:29: XDL_MIN.\nI wonder what the rationale is for not having a MIN()\ndefinition, I think we discussed that on the list a couple\ntimes but the rationale escaped me.\n\nIf we introduce a min/max macro, can we put it somewhere\nmore prominent? (I would find it useful elsewhere)\n\n> +/*\n> +* Mostly randomly chosen maximum thread counts: we\n> +* cap the parallelism to online_cpus() threads, and we want\n> +* to have at least 7500 cache entries per thread for it to\n> +* be worth starting a thread.\n> +*/\n> +#define THREAD_COST            (7500)\n\nThis reads very similar to preload-index.c THREAD_COST\n\n> +       /* loop through index entries starting a thread for every thread_nr entries */\n> +       consumed = thread = 0;\n> +       for (i = 0; ; i++) {\n> +               struct ondisk_cache_entry *ondisk;\n> +               const char *name;\n> +               unsigned int flags;\n> +\n> +               /* we've reached the begining of a block of cache entries, kick off a thread to process them */\n\nbeginning\n\nThanks,\nStefan\n"},{"id":"356389","messageId":"xmqqin41hs8x.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180823154053.20212-1-benpeart@microsoft.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-23T18:06:54Z","receivedAt":"2018-08-23T18:06:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <Ben.Peart@microsoft.com> writes:\n\n> This patch helps address the CPU cost of loading the index by creating\n> multiple threads to divide the work of loading and converting the cache\n> entries across all available CPU cores.\n\nNice.\n\n> +int git_config_get_fast_index(void)\n> +{\n> +\tint val;\n> +\n> +\tif (!git_config_get_maybe_bool(\"core.fastindex\", &val))\n> +\t\treturn val;\n> +\n> +\tif (getenv(\"GIT_FASTINDEX_TEST\"))\n> +\t\treturn 1;\n\nIt probably makes sense to use git_env_bool() to be consistent,\nwhich allows GIT_FASTINDEX_TEST=0 to turn it off after this becomes\nthe default.\n\n> diff --git a/read-cache.c b/read-cache.c\n> index 7b1354d759..0fa7e1a04c 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -24,6 +24,10 @@\n>  #include \"utf8.h\"\n>  #include \"fsmonitor.h\"\n>  \n> +#ifndef min\n> +#define min(a,b) (((a) < (b)) ? (a) : (b))\n> +#endif\n\nLet's lose this, which is used only once, even though it could be\nused elsewhere but not used (e.g. threads vs cpus near the beginning\nof load_cache_entries()).\n\n> +static unsigned long load_cache_entry_block(struct index_state *istate, struct mem_pool *ce_mem_pool, int offset, int nr, void *mmap, unsigned long start_offset, struct strbuf *previous_name)\n\nWrap and possibly add comment before the function to describe what\nit does and what its parameters mean?\n\n> +{\n> +\tint i;\n> +\tunsigned long src_offset = start_offset;\n> +\n> +\tfor (i = offset; i < offset + nr; i++) {\n> +\t\tstruct ondisk_cache_entry *disk_ce;\n> +\t\tstruct cache_entry *ce;\n> +\t\tunsigned long consumed;\n> +\n> +\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n> +\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n> +\t\tset_index_entry(istate, i, ce);\n> +\n> +\t\tsrc_offset += consumed;\n> +\t}\n> +\treturn src_offset - start_offset;\n> +}\n\nOK.\n\n> +static unsigned long load_all_cache_entries(struct index_state *istate, void *mmap, size_t mmap_size, unsigned long src_offset)\n> +{\n\n(following aloud) This \"all\" variant is \"one thread does all\", iow,\nunthreaded version.  Makes sense.\n\n> +\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n> +\tunsigned long consumed;\n> +\n> +\tif (istate->version == 4) {\n> +\t\tprevious_name = &previous_name_buf;\n> +\t\tmem_pool_init(&istate->ce_mem_pool,\n> +\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n> +\t} else {\n> +\t\tprevious_name = NULL;\n> +\t\tmem_pool_init(&istate->ce_mem_pool,\n> +\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n> +\t}\n\nI count there are three instances of \"if version 4 use the strbuf\nfor name-buf, otherwise...\" in this patch, which made me wonder if\nwe can make them shared more and/or if it makes sense to attempt to\ndo so.\n\n> +\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool, 0, istate->cache_nr, mmap, src_offset, previous_name);\n> +\tstrbuf_release(&previous_name_buf);\n> +\treturn consumed;\n> +}\n> +\n> +#ifdef NO_PTHREADS\n> +\n> +#define load_cache_entries load_all_cache_entries\n> +\n> +#else\n> +\n> +#include \"thread-utils.h\"\n> +\n> +/*\n> +* Mostly randomly chosen maximum thread counts: we\n> +* cap the parallelism to online_cpus() threads, and we want\n> +* to have at least 7500 cache entries per thread for it to\n> +* be worth starting a thread.\n> +*/\n> +#define THREAD_COST\t\t(7500)\n> +\n> +struct load_cache_entries_thread_data\n> +{\n> +\tpthread_t pthread;\n> +\tstruct index_state *istate;\n> +\tstruct mem_pool *ce_mem_pool;\n> +\tint offset, nr;\n> +\tvoid *mmap;\n> +\tunsigned long start_offset;\n> +\tstruct strbuf previous_name_buf;\n> +\tstruct strbuf *previous_name;\n> +\tunsigned long consumed;\t/* return # of bytes in index file processed */\n> +};\n> +\n> +/*\n> +* A thread proc to run the load_cache_entries() computation\n> +* across multiple background threads.\n> +*/\n> +static void *load_cache_entries_thread(void *_data)\n> +{\n> +\tstruct load_cache_entries_thread_data *p = _data;\n> +\n> +\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool, p->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n> +\treturn NULL;\n> +}\n\n(following aloud) And the threaded version chews the block of ce's\ngiven to each thread.  Makes sense.\n\n> +static unsigned long load_cache_entries(struct index_state *istate, void *mmap, size_t mmap_size, unsigned long src_offset)\n> +{\n> +\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n> +\tstruct load_cache_entries_thread_data *data;\n> +\tint threads, cpus, thread_nr;\n> +\tunsigned long consumed;\n> +\tint i, thread;\n> +\n> +\tcpus = online_cpus();\n> +\tthreads = istate->cache_nr / THREAD_COST;\n> +\tif (threads > cpus)\n> +\t\tthreads = cpus;\n\nNo other caller of online_cpus() is prepared to deal with faulty\nreturn from the function (e.g. 0 or negative), so it is perfectly\nfine for this caller to trust it would return at least 1.  OK.\n\nNot using min() and it still is very readable ;-).\n\n> +\t/* enable testing with fewer than default minimum of entries */\n> +\tif ((istate->cache_nr > 1) && (threads < 2) && getenv(\"GIT_FASTINDEX_TEST\"))\n> +\t\tthreads = 2;\n\nAnother good place to use git_env_bool().\n\n> +\tif (threads < 2 || !git_config_get_fast_index())\n> +\t\treturn load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n\nconfig_get_fast_index() can return -1 to signal \"no strong\npreference either way\".  A caller that negates the value without\npaying special attention to negative return makes the reader wonder\nif the code is buggy or actively interpreting \"do not care\" as \"I do\nnot mind if you use it\" (it is the latter in this case).\n\nI actually think git_config_get_fast_index() is a helper that does a\nbit too little.  Perhaps the above two if() statements can be\ncombined into a single call to\n\n\tthreads = use_fast_index(istate);\n\tif (threads < 2)\n\t\treturn load_all_cache_entries(...);\n\nand let it call online_cpus(), determination of thread-count taking\nTHREADS_COST into account, and also reading the configuration\nvariable?  The configuration variable might even want to say how\nmany threads it wants to cap us at maximum in the future.\n\n> +\tmem_pool_init(&istate->ce_mem_pool, 0);\n> +\tif (istate->version == 4)\n> +\t\tprevious_name = &previous_name_buf;\n> +\telse\n> +\t\tprevious_name = NULL;\n> +\n> +\tthread_nr = (istate->cache_nr + threads - 1) / threads;\n\n(following aloud) threads is the number of threads that we are going\nto spawn.  thread_nr is not any number about threads---it is number\nof cache entries each thread will work on.  The latter is\nconfusingly named.\n\nce_per_thread perhaps?\n\nAs the division is rounded up, among \"threads\" threads, we know we\nwill cover all \"cache_nr\" cache entries.  The last thread may handle\nfewer than \"thread_nr\" entries, or even just a single entry in the\nworst case.\n\nWhen cache_nr == 1 and FASTINDEX_TEST tells us to use threads == 2,\nthen thread_nr = (1 + 2 - 1) / 2 = 1.\n\nThe first one in the loop is given (offset, nr) = (0, 1) in the loop\nThe second one is given (offset, nr) = (1, 0) in the loop.  Two\nquestions come to mind:\n\n - Is load_cache_entries_thread() prepared to be given offset that\n   is beyond the end of istate->cache[] and become a no-op?\n\n - Does the next loop even terminate without running beyond the end\n   of istate->cache[]?\n\n> +\tdata = xcalloc(threads, sizeof(struct load_cache_entries_thread_data));\n> +\n> +\t/* loop through index entries starting a thread for every thread_nr entries */\n> +\tconsumed = thread = 0;\n> +\tfor (i = 0; ; i++) {\n\nUncapped for() loop makes readers a bit nervous.\nAn extra \"i < istate->cache_nr\" would not hurt, perhaps?\n\n> +\t\tstruct ondisk_cache_entry *ondisk;\n> +\t\tconst char *name;\n> +\t\tunsigned int flags;\n> +\n> +\t\t/* we've reached the begining of a block of cache entries, kick off a thread to process them */\n> +\t\tif (0 == i % thread_nr) {\n> +\t\t\tstruct load_cache_entries_thread_data *p = &data[thread];\n> +\n> +\t\t\tp->istate = istate;\n> +\t\t\tp->offset = i;\n> +\t\t\tp->nr = min(thread_nr, istate->cache_nr - i);\n\n(following aloud) p->nr is the number of entries this thread will\nwork on.\n\n> +\t\t\t/* create a mem_pool for each thread */\n> +\t\t\tif (istate->version == 4)\n> +\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n> +\t\t\t\t\t\t  estimate_cache_size_from_compressed(p->nr));\n> +\t\t\telse\n> +\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n> +\t\t\t\t\t\t  estimate_cache_size(mmap_size, p->nr));\n> +\n> +\t\t\tp->mmap = mmap;\n> +\t\t\tp->start_offset = src_offset;\n> +\t\t\tif (previous_name) {\n> +\t\t\t\tstrbuf_addbuf(&p->previous_name_buf, previous_name);\n> +\t\t\t\tp->previous_name = &p->previous_name_buf;\n> +\t\t\t}\n> +\n> +\t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n> +\t\t\t\tdie(\"unable to create load_cache_entries_thread\");\n> +\t\t\tif (++thread == threads || p->nr != thread_nr)\n> +\t\t\t\tbreak;\n> +\t\t}\n> +\n> +\t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n> +\n> +\t\t/* On-disk flags are just 16 bits */\n> +\t\tflags = get_be16(&ondisk->flags);\n> +\n> +\t\tif (flags & CE_EXTENDED) {\n> +\t\t\tstruct ondisk_cache_entry_extended *ondisk2;\n> +\t\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n> +\t\t\tname = ondisk2->name;\n> +\t\t} else\n> +\t\t\tname = ondisk->name;\n> +\n> +\t\tif (!previous_name) {\n> +\t\t\tsize_t len;\n> +\n> +\t\t\t/* v3 and earlier */\n> +\t\t\tlen = flags & CE_NAMEMASK;\n> +\t\t\tif (len == CE_NAMEMASK)\n> +\t\t\t\tlen = strlen(name);\n> +\t\t\tsrc_offset += (flags & CE_EXTENDED) ?\n> +\t\t\t\tondisk_cache_entry_extended_size(len) :\n> +\t\t\t\tondisk_cache_entry_size(len);\n> +\t\t} else\n> +\t\t\tsrc_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n\nNice to see this done without a new index extension that records\noffsets, so that we can load existing index files in parallel.\n\n> +\t}\n> +\n> +\tfor (i = 0; i < threads; i++) {\n> +\t\tstruct load_cache_entries_thread_data *p = data + i;\n> +\t\tif (pthread_join(p->pthread, NULL))\n> +\t\t\tdie(\"unable to join load_cache_entries_thread\");\n> +\t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n> +\t\tstrbuf_release(&p->previous_name_buf);\n> +\t\tconsumed += p->consumed;\n> +\t}\n> +\n> +\tfree(data);\n> +\tstrbuf_release(&previous_name_buf);\n> +\n> +\treturn consumed;\n> +}\n> +\n> +#endif\n"},{"id":"356394","messageId":"12b29bbf-4565-fe0e-97d4-19da2b4ddf5e@gmail.com","threadId":"49204","inReplyTo":"CAGZ79kbXfPPvcQ1rnUdiOqWs5wC2qccGCnf8DvCVnp8QV126MA@mail.gmail.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-23T19:44:58Z","receivedAt":"2018-08-23T19:45:04Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/23/2018 1:31 PM, Stefan Beller wrote:\n> On Thu, Aug 23, 2018 at 8:45 AM Ben Peart <Ben.Peart@microsoft.com> wrote:\n>>\n>> This patch helps address the CPU cost of loading the index by creating\n>> multiple threads to divide the work of loading and converting the cache\n>> entries across all available CPU cores.\n>>\n>> It accomplishes this by having the primary thread loop across the index file\n>> tracking the offset and (for V4 indexes) expanding the name. It creates a\n>> thread to process each block of entries as it comes to them. Once the\n>> threads are complete and the cache entries are loaded, the rest of the\n>> extensions can be loaded and processed normally on the primary thread.\n>>\n>> Performance impact:\n>>\n>> read cache .git/index times on a synthetic repo with:\n>>\n>> 100,000 entries\n>> FALSE       TRUE        Savings     %Savings\n>> 0.014798767 0.009580433 0.005218333 35.26%\n>>\n>> 1,000,000 entries\n>> FALSE       TRUE        Savings     %Savings\n>> 0.240896533 0.1751243   0.065772233 27.30%\n>>\n>> read cache .git/index times on an actual repo with:\n>>\n>> ~3M entries\n>> FALSE       TRUE        Savings     %Savings\n>> 0.59898098  0.4513169   0.14766408  24.65%\n>>\n>> Signed-off-by: Ben Peart <Ben.Peart@microsoft.com>\n>> ---\n>>\n>> Notes:\n>>      Base Ref: master\n>>      Web-Diff: https://github.com/benpeart/git/commit/67a700419b\n>>      Checkout: git fetch https://github.com/benpeart/git read-index-multithread-v1 && git checkout 67a700419b\n>>\n>>   Documentation/config.txt |   8 ++\n>>   config.c                 |  13 +++\n>>   config.h                 |   1 +\n>>   read-cache.c             | 218 ++++++++++++++++++++++++++++++++++-----\n>>   4 files changed, 216 insertions(+), 24 deletions(-)\n>>\n>> diff --git a/Documentation/config.txt b/Documentation/config.txt\n>> index 1c42364988..3344685cc4 100644\n>> --- a/Documentation/config.txt\n>> +++ b/Documentation/config.txt\n>> @@ -899,6 +899,14 @@ relatively high IO latencies.  When enabled, Git will do the\n>>   index comparison to the filesystem data in parallel, allowing\n>>   overlapping IO's.  Defaults to true.\n>>\n>> +core.fastIndex::\n>> +       Enable parallel index loading\n>> ++\n>> +This can speed up operations like 'git diff' and 'git status' especially\n>> +when the index is very large.  When enabled, Git will do the index\n>> +loading from the on disk format to the in-memory format in parallel.\n>> +Defaults to true.\n> \n> \"fast\" is a non-descriptive word as we try to be fast in any operation?\n> Maybe core.parallelIndexReading as that just describes what it\n> turns on/off, without second guessing its effects?\n> (Are there still computers with just a single CPU, where this would not\n> make it faster? ;-))\n> \n\nHow about core.parallelReadIndex?  Slightly shorter and matches the \nfunction names better.\n\n> \n>> +int git_config_get_fast_index(void)\n>> +{\n>> +       int val;\n>> +\n>> +       if (!git_config_get_maybe_bool(\"core.fastindex\", &val))\n>> +               return val;\n>> +\n>> +       if (getenv(\"GIT_FASTINDEX_TEST\"))\n>> +               return 1;\n> \n> We look at this env value just before calling this function,\n> can be write it to only look at the evn variable once?\n> \n\nSure, I didn't like the fact that it was called twice but didn't get \naround to cleaning it up.\n\n>> +++ b/config.h\n>> @@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n>>   extern int git_config_get_split_index(void);\n>>   extern int git_config_get_max_percent_split_change(void);\n>>   extern int git_config_get_fsmonitor(void);\n>> +extern int git_config_get_fast_index(void);\n> \n> Oh. nd/no-extern did not cover config.h\n> \n> \n>>\n>> +#ifndef min\n>> +#define min(a,b) (((a) < (b)) ? (a) : (b))\n>> +#endif\n> \n> We do not have a minimum function in the tree,\n> except for xdiff/xmacros.h:29: XDL_MIN.\n> I wonder what the rationale is for not having a MIN()\n> definition, I think we discussed that on the list a couple\n> times but the rationale escaped me.\n> \n> If we introduce a min/max macro, can we put it somewhere\n> more prominent? (I would find it useful elsewhere)\n>\n\nI'll avoid that particular rabbit hole and just remove the min macro \ndefinition.  ;-)\n\n>> +/*\n>> +* Mostly randomly chosen maximum thread counts: we\n>> +* cap the parallelism to online_cpus() threads, and we want\n>> +* to have at least 7500 cache entries per thread for it to\n>> +* be worth starting a thread.\n>> +*/\n>> +#define THREAD_COST            (7500)\n> \n> This reads very similar to preload-index.c THREAD_COST\n> \n>> +       /* loop through index entries starting a thread for every thread_nr entries */\n>> +       consumed = thread = 0;\n>> +       for (i = 0; ; i++) {\n>> +               struct ondisk_cache_entry *ondisk;\n>> +               const char *name;\n>> +               unsigned int flags;\n>> +\n>> +               /* we've reached the begining of a block of cache entries, kick off a thread to process them */\n> \n> beginning\n> \n\nThanks\n\n> Thanks,\n> Stefan\n> \n"},{"id":"356395","messageId":"4c70ea50-5b43-8696-3c46-cf3d658a0ef8@gmail.com","threadId":"49204","inReplyTo":"xmqqin41hs8x.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-23T20:33:48Z","receivedAt":"2018-08-23T20:33:53Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/23/2018 2:06 PM, Junio C Hamano wrote:\n> Ben Peart <Ben.Peart@microsoft.com> writes:\n> \n>> This patch helps address the CPU cost of loading the index by creating\n>> multiple threads to divide the work of loading and converting the cache\n>> entries across all available CPU cores.\n> \n> Nice.\n> \n>> +int git_config_get_fast_index(void)\n>> +{\n>> +\tint val;\n>> +\n>> +\tif (!git_config_get_maybe_bool(\"core.fastindex\", &val))\n>> +\t\treturn val;\n>> +\n>> +\tif (getenv(\"GIT_FASTINDEX_TEST\"))\n>> +\t\treturn 1;\n> \n> It probably makes sense to use git_env_bool() to be consistent,\n> which allows GIT_FASTINDEX_TEST=0 to turn it off after this becomes\n> the default.\n> \n>> diff --git a/read-cache.c b/read-cache.c\n>> index 7b1354d759..0fa7e1a04c 100644\n>> --- a/read-cache.c\n>> +++ b/read-cache.c\n>> @@ -24,6 +24,10 @@\n>>   #include \"utf8.h\"\n>>   #include \"fsmonitor.h\"\n>>   \n>> +#ifndef min\n>> +#define min(a,b) (((a) < (b)) ? (a) : (b))\n>> +#endif\n> \n> Let's lose this, which is used only once, even though it could be\n> used elsewhere but not used (e.g. threads vs cpus near the beginning\n> of load_cache_entries()).\n> \n\nI didn't have it, then added it to make it trivial to see what was \nactually happening.  I can switch back.\n\n>> +static unsigned long load_cache_entry_block(struct index_state *istate, struct mem_pool *ce_mem_pool, int offset, int nr, void *mmap, unsigned long start_offset, struct strbuf *previous_name)\n> \n> Wrap and possibly add comment before the function to describe what\n> it does and what its parameters mean?\n> \n>> +{\n>> +\tint i;\n>> +\tunsigned long src_offset = start_offset;\n>> +\n>> +\tfor (i = offset; i < offset + nr; i++) {\n>> +\t\tstruct ondisk_cache_entry *disk_ce;\n>> +\t\tstruct cache_entry *ce;\n>> +\t\tunsigned long consumed;\n>> +\n>> +\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n>> +\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n>> +\t\tset_index_entry(istate, i, ce);\n>> +\n>> +\t\tsrc_offset += consumed;\n>> +\t}\n>> +\treturn src_offset - start_offset;\n>> +}\n> \n> OK.\n> \n>> +static unsigned long load_all_cache_entries(struct index_state *istate, void *mmap, size_t mmap_size, unsigned long src_offset)\n>> +{\n> \n> (following aloud) This \"all\" variant is \"one thread does all\", iow,\n> unthreaded version.  Makes sense.\n> \n>> +\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n>> +\tunsigned long consumed;\n>> +\n>> +\tif (istate->version == 4) {\n>> +\t\tprevious_name = &previous_name_buf;\n>> +\t\tmem_pool_init(&istate->ce_mem_pool,\n>> +\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n>> +\t} else {\n>> +\t\tprevious_name = NULL;\n>> +\t\tmem_pool_init(&istate->ce_mem_pool,\n>> +\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n>> +\t}\n> \n> I count there are three instances of \"if version 4 use the strbuf\n> for name-buf, otherwise...\" in this patch, which made me wonder if\n> we can make them shared more and/or if it makes sense to attempt to\n> do so.\n> \n\nActually, they are all different and all required.  One sets it up for \nthe \"do it all on one thread\" path.  One sets it up for each thread. The \nlast one is used by the primary thread when scanning for blocks to hand \noff to the child threads.\n\n>> +\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool, 0, istate->cache_nr, mmap, src_offset, previous_name);\n>> +\tstrbuf_release(&previous_name_buf);\n>> +\treturn consumed;\n>> +}\n>> +\n>> +#ifdef NO_PTHREADS\n>> +\n>> +#define load_cache_entries load_all_cache_entries\n>> +\n>> +#else\n>> +\n>> +#include \"thread-utils.h\"\n>> +\n>> +/*\n>> +* Mostly randomly chosen maximum thread counts: we\n>> +* cap the parallelism to online_cpus() threads, and we want\n>> +* to have at least 7500 cache entries per thread for it to\n>> +* be worth starting a thread.\n>> +*/\n>> +#define THREAD_COST\t\t(7500)\n>> +\n>> +struct load_cache_entries_thread_data\n>> +{\n>> +\tpthread_t pthread;\n>> +\tstruct index_state *istate;\n>> +\tstruct mem_pool *ce_mem_pool;\n>> +\tint offset, nr;\n>> +\tvoid *mmap;\n>> +\tunsigned long start_offset;\n>> +\tstruct strbuf previous_name_buf;\n>> +\tstruct strbuf *previous_name;\n>> +\tunsigned long consumed;\t/* return # of bytes in index file processed */\n>> +};\n>> +\n>> +/*\n>> +* A thread proc to run the load_cache_entries() computation\n>> +* across multiple background threads.\n>> +*/\n>> +static void *load_cache_entries_thread(void *_data)\n>> +{\n>> +\tstruct load_cache_entries_thread_data *p = _data;\n>> +\n>> +\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool, p->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n>> +\treturn NULL;\n>> +}\n> \n> (following aloud) And the threaded version chews the block of ce's\n> given to each thread.  Makes sense.\n> \n>> +static unsigned long load_cache_entries(struct index_state *istate, void *mmap, size_t mmap_size, unsigned long src_offset)\n>> +{\n>> +\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n>> +\tstruct load_cache_entries_thread_data *data;\n>> +\tint threads, cpus, thread_nr;\n>> +\tunsigned long consumed;\n>> +\tint i, thread;\n>> +\n>> +\tcpus = online_cpus();\n>> +\tthreads = istate->cache_nr / THREAD_COST;\n>> +\tif (threads > cpus)\n>> +\t\tthreads = cpus;\n> \n> No other caller of online_cpus() is prepared to deal with faulty\n> return from the function (e.g. 0 or negative), so it is perfectly\n> fine for this caller to trust it would return at least 1.  OK.\n> \n> Not using min() and it still is very readable ;-).\n> \n>> +\t/* enable testing with fewer than default minimum of entries */\n>> +\tif ((istate->cache_nr > 1) && (threads < 2) && getenv(\"GIT_FASTINDEX_TEST\"))\n>> +\t\tthreads = 2;\n> \n> Another good place to use git_env_bool().\n> \n>> +\tif (threads < 2 || !git_config_get_fast_index())\n>> +\t\treturn load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n> \n> config_get_fast_index() can return -1 to signal \"no strong\n> preference either way\".  A caller that negates the value without\n> paying special attention to negative return makes the reader wonder\n> if the code is buggy or actively interpreting \"do not care\" as \"I do\n> not mind if you use it\" (it is the latter in this case).\n> \n> I actually think git_config_get_fast_index() is a helper that does a\n> bit too little.  Perhaps the above two if() statements can be\n> combined into a single call to\n> \n> \tthreads = use_fast_index(istate);\n> \tif (threads < 2)\n> \t\treturn load_all_cache_entries(...);\n> \n> and let it call online_cpus(), determination of thread-count taking\n> THREADS_COST into account, and also reading the configuration\n> variable?  The configuration variable might even want to say how\n> many threads it wants to cap us at maximum in the future.\n> \n\nI reworked this a bit.\n\ngit_config_get_parallel_read_index() still just deals with the config \nvalue (I had to read it this way as in some code paths, the global \nconfig settings in environment.c haven't been read yet).\n\nAll the logic about whether to use threads and how many to use is \ncentralized here along with the environment variable to override the \ndefault behavior.\n\n>> +\tmem_pool_init(&istate->ce_mem_pool, 0);\n>> +\tif (istate->version == 4)\n>> +\t\tprevious_name = &previous_name_buf;\n>> +\telse\n>> +\t\tprevious_name = NULL;\n>> +\n>> +\tthread_nr = (istate->cache_nr + threads - 1) / threads;\n> \n> (following aloud) threads is the number of threads that we are going\n> to spawn.  thread_nr is not any number about threads---it is number\n> of cache entries each thread will work on.  The latter is\n> confusingly named.\n> \n> ce_per_thread perhaps?\n> \n\nSure\n\n> As the division is rounded up, among \"threads\" threads, we know we\n> will cover all \"cache_nr\" cache entries.  The last thread may handle\n> fewer than \"thread_nr\" entries, or even just a single entry in the\n> worst case.\n> \n\nIt's divided by the number of threads so will only be up to 1 less than \nthe other threads.  Given the minimum # of entries per thread is 7500, \nyou'd never end up with just a single entry (unless using the \nGIT_PARALLELREADINDEX_TEST override).\n\n> When cache_nr == 1 and FASTINDEX_TEST tells us to use threads == 2,\n> then thread_nr = (1 + 2 - 1) / 2 = 1.\n> \n> The first one in the loop is given (offset, nr) = (0, 1) in the loop\n> The second one is given (offset, nr) = (1, 0) in the loop.  Two\n> questions come to mind:\n> \n>   - Is load_cache_entries_thread() prepared to be given offset that\n>     is beyond the end of istate->cache[] and become a no-op?\n> \n>   - Does the next loop even terminate without running beyond the end\n>     of istate->cache[]?\n> \n>> +\tdata = xcalloc(threads, sizeof(struct load_cache_entries_thread_data));\n>> +\n>> +\t/* loop through index entries starting a thread for every thread_nr entries */\n>> +\tconsumed = thread = 0;\n>> +\tfor (i = 0; ; i++) {\n> \n> Uncapped for() loop makes readers a bit nervous.\n> An extra \"i < istate->cache_nr\" would not hurt, perhaps?\n> \n\nWe don't need or want to run through _all_ the entries, only to the \nfirst entry of the last block.  I'd prefer to leave that extra test out \nas it implies that we are going to loop through them all. I'll add a \ncomment to make it more obvious what is happening.\n\n>> +\t\tstruct ondisk_cache_entry *ondisk;\n>> +\t\tconst char *name;\n>> +\t\tunsigned int flags;\n>> +\n>> +\t\t/* we've reached the begining of a block of cache entries, kick off a thread to process them */\n>> +\t\tif (0 == i % thread_nr) {\n>> +\t\t\tstruct load_cache_entries_thread_data *p = &data[thread];\n>> +\n>> +\t\t\tp->istate = istate;\n>> +\t\t\tp->offset = i;\n>> +\t\t\tp->nr = min(thread_nr, istate->cache_nr - i);\n> \n> (following aloud) p->nr is the number of entries this thread will\n> work on.\n> \n>> +\t\t\t/* create a mem_pool for each thread */\n>> +\t\t\tif (istate->version == 4)\n>> +\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n>> +\t\t\t\t\t\t  estimate_cache_size_from_compressed(p->nr));\n>> +\t\t\telse\n>> +\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n>> +\t\t\t\t\t\t  estimate_cache_size(mmap_size, p->nr));\n>> +\n>> +\t\t\tp->mmap = mmap;\n>> +\t\t\tp->start_offset = src_offset;\n>> +\t\t\tif (previous_name) {\n>> +\t\t\t\tstrbuf_addbuf(&p->previous_name_buf, previous_name);\n>> +\t\t\t\tp->previous_name = &p->previous_name_buf;\n>> +\t\t\t}\n>> +\n>> +\t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n>> +\t\t\t\tdie(\"unable to create load_cache_entries_thread\");\n>> +\t\t\tif (++thread == threads || p->nr != thread_nr)\n>> +\t\t\t\tbreak;\n>> +\t\t}\n>> +\n>> +\t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n>> +\n>> +\t\t/* On-disk flags are just 16 bits */\n>> +\t\tflags = get_be16(&ondisk->flags);\n>> +\n>> +\t\tif (flags & CE_EXTENDED) {\n>> +\t\t\tstruct ondisk_cache_entry_extended *ondisk2;\n>> +\t\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n>> +\t\t\tname = ondisk2->name;\n>> +\t\t} else\n>> +\t\t\tname = ondisk->name;\n>> +\n>> +\t\tif (!previous_name) {\n>> +\t\t\tsize_t len;\n>> +\n>> +\t\t\t/* v3 and earlier */\n>> +\t\t\tlen = flags & CE_NAMEMASK;\n>> +\t\t\tif (len == CE_NAMEMASK)\n>> +\t\t\t\tlen = strlen(name);\n>> +\t\t\tsrc_offset += (flags & CE_EXTENDED) ?\n>> +\t\t\t\tondisk_cache_entry_extended_size(len) :\n>> +\t\t\t\tondisk_cache_entry_size(len);\n>> +\t\t} else\n>> +\t\t\tsrc_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n> \n> Nice to see this done without a new index extension that records\n> offsets, so that we can load existing index files in parallel.\n> \n\nYes, I prefer this simpler model as well.  I wasn't sure it would \nproduce a significant improvement given the primary thread still has to \nrun through the variable length cache entries but was pleasantly surprised.\n\nThe recent mem_pool changes really helped as well as it removed all \nthread contention in the heap that was happening before.\n\n>> +\t}\n>> +\n>> +\tfor (i = 0; i < threads; i++) {\n>> +\t\tstruct load_cache_entries_thread_data *p = data + i;\n>> +\t\tif (pthread_join(p->pthread, NULL))\n>> +\t\t\tdie(\"unable to join load_cache_entries_thread\");\n>> +\t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n>> +\t\tstrbuf_release(&p->previous_name_buf);\n>> +\t\tconsumed += p->consumed;\n>> +\t}\n>> +\n>> +\tfree(data);\n>> +\tstrbuf_release(&previous_name_buf);\n>> +\n>> +\treturn consumed;\n>> +}\n>> +\n>> +#endif\n"},{"id":"356470","messageId":"CACsJy8CUPGUhR3girstdqD6YVxOQ6_xE+gacT98KXgqOSPz0dw@mail.gmail.com","threadId":"49204","inReplyTo":"4c70ea50-5b43-8696-3c46-cf3d658a0ef8@gmail.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-24T15:37:20Z","receivedAt":"2018-08-24T15:37:48Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Since we're cutting corners to speed things up, could you try\nsomething like this?\n\nI notice that reading v4 is significantly slower than v2 and\napparently strlen() (at least from glibc) is much cleverer and at\nleast gives me a few percentage time saving.\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 7b1354d759..d10cccaed0 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1755,8 +1755,7 @@ static unsigned long expand_name_field(struct\nstrbuf *name, const char *cp_)\n        if (name->len < len)\n                die(\"malformed name field in the index\");\n        strbuf_remove(name, name->len - len, len);\n-       for (ep = cp; *ep; ep++)\n-               ; /* find the end */\n+       ep = cp + strlen(cp);\n        strbuf_add(name, cp, ep - cp);\n        return (const char *)ep + 1 - cp_;\n }\n\nOn Thu, Aug 23, 2018 at 10:36 PM Ben Peart <peartben@gmail.com> wrote:\n> > Nice to see this done without a new index extension that records\n> > offsets, so that we can load existing index files in parallel.\n> >\n>\n> Yes, I prefer this simpler model as well.  I wasn't sure it would\n> produce a significant improvement given the primary thread still has to\n> run through the variable length cache entries but was pleasantly surprised.\n\nOut of curiosity, how much time saving could we gain by recording\noffsets as an extension (I assume we need, like 4 offsets if the\nsystem has 4 cores)? Much much more than this simpler model (which may\njustify the complexity) or just \"meh\" compared to this?\n-- \nDuy\n"},{"id":"356471","messageId":"20180824155734.GA6170@duynguyen.home","threadId":"49204","inReplyTo":"CACsJy8CUPGUhR3girstdqD6YVxOQ6_xE+gacT98KXgqOSPz0dw@mail.gmail.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-24T15:57:34Z","receivedAt":"2018-08-24T15:57:41Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Aug 24, 2018 at 05:37:20PM +0200, Duy Nguyen wrote:\n> Since we're cutting corners to speed things up, could you try\n> something like this?\n> \n> I notice that reading v4 is significantly slower than v2 and\n> apparently strlen() (at least from glibc) is much cleverer and at\n> least gives me a few percentage time saving.\n> \n> diff --git a/read-cache.c b/read-cache.c\n> index 7b1354d759..d10cccaed0 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1755,8 +1755,7 @@ static unsigned long expand_name_field(struct\n> strbuf *name, const char *cp_)\n>         if (name->len < len)\n>                 die(\"malformed name field in the index\");\n>         strbuf_remove(name, name->len - len, len);\n> -       for (ep = cp; *ep; ep++)\n> -               ; /* find the end */\n> +       ep = cp + strlen(cp);\n>         strbuf_add(name, cp, ep - cp);\n>         return (const char *)ep + 1 - cp_;\n>  }\n\nNo try this instead. It's half way back to v2 numbers for me (tested\nwith \"test-tool read-cache 100\" on webkit.git). For the record, v4 is\nabout 30% slower than v2 in my tests.\n\nWe could probably do better too. Instead of preparing the string in a\nseparate buffer (previous_name_buf), we could just assemble it directly\nto the newly allocated \"ce\".\n\n-- 8< --\ndiff --git a/read-cache.c b/read-cache.c\nindex 7b1354d759..237f60a76c 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1754,9 +1754,8 @@ static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n \n \tif (name->len < len)\n \t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n+\tstrbuf_setlen(name, name->len - len);\n+\tep = cp + strlen(cp);\n \tstrbuf_add(name, cp, ep - cp);\n \treturn (const char *)ep + 1 - cp_;\n }\n-- 8< --\n"},{"id":"356481","messageId":"3e82ee23-6a04-6714-4d51-68bb2ef4a123@gmail.com","threadId":"49204","inReplyTo":"20180824155734.GA6170@duynguyen.home","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-24T17:28:56Z","receivedAt":"2018-08-24T17:29:00Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/24/2018 11:57 AM, Duy Nguyen wrote:\n> On Fri, Aug 24, 2018 at 05:37:20PM +0200, Duy Nguyen wrote:\n>> Since we're cutting corners to speed things up, could you try\n>> something like this?\n>>\n>> I notice that reading v4 is significantly slower than v2 and\n>> apparently strlen() (at least from glibc) is much cleverer and at\n>> least gives me a few percentage time saving.\n>>\n>> diff --git a/read-cache.c b/read-cache.c\n>> index 7b1354d759..d10cccaed0 100644\n>> --- a/read-cache.c\n>> +++ b/read-cache.c\n>> @@ -1755,8 +1755,7 @@ static unsigned long expand_name_field(struct\n>> strbuf *name, const char *cp_)\n>>          if (name->len < len)\n>>                  die(\"malformed name field in the index\");\n>>          strbuf_remove(name, name->len - len, len);\n>> -       for (ep = cp; *ep; ep++)\n>> -               ; /* find the end */\n>> +       ep = cp + strlen(cp);\n>>          strbuf_add(name, cp, ep - cp);\n>>          return (const char *)ep + 1 - cp_;\n>>   }\n> \n> No try this instead. It's half way back to v2 numbers for me (tested\n> with \"test-tool read-cache 100\" on webkit.git). For the record, v4 is\n> about 30% slower than v2 in my tests.\n> \n\nThanks Duy, this helped on my system as well.\n\nInterestingly, simply reading the cache tree extension in read_one() now \ntakes about double the CPU on the primary thread as does \nload_cache_entries().\n\nHmm, that gives me an idea.  I could kick off another thread to load \nthat extension in parallel and cut off another ~160 ms.  I'll add that \nto my list of future patches to investigate...\n\n> We could probably do better too. Instead of preparing the string in a\n> separate buffer (previous_name_buf), we could just assemble it directly\n> to the newly allocated \"ce\".\n> \n> -- 8< --\n> diff --git a/read-cache.c b/read-cache.c\n> index 7b1354d759..237f60a76c 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1754,9 +1754,8 @@ static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n>   \n>   \tif (name->len < len)\n>   \t\tdie(\"malformed name field in the index\");\n> -\tstrbuf_remove(name, name->len - len, len);\n> -\tfor (ep = cp; *ep; ep++)\n> -\t\t; /* find the end */\n> +\tstrbuf_setlen(name, name->len - len);\n> +\tep = cp + strlen(cp);\n>   \tstrbuf_add(name, cp, ep - cp);\n>   \treturn (const char *)ep + 1 - cp_;\n>   }\n> -- 8< --\n> \n"},{"id":"356482","messageId":"CACsJy8BXy_7QbDtF8bY5YzwJf=JUwiODv0zKxoSXeu4rJ+xjwg@mail.gmail.com","threadId":"49204","inReplyTo":"CACsJy8CUPGUhR3girstdqD6YVxOQ6_xE+gacT98KXgqOSPz0dw@mail.gmail.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-24T18:20:37Z","receivedAt":"2018-08-24T18:21:06Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Aug 24, 2018 at 5:37 PM Duy Nguyen <pclouds@gmail.com> wrote:\n> On Thu, Aug 23, 2018 at 10:36 PM Ben Peart <peartben@gmail.com> wrote:\n> > > Nice to see this done without a new index extension that records\n> > > offsets, so that we can load existing index files in parallel.\n> > >\n> >\n> > Yes, I prefer this simpler model as well.  I wasn't sure it would\n> > produce a significant improvement given the primary thread still has to\n> > run through the variable length cache entries but was pleasantly surprised.\n>\n> Out of curiosity, how much time saving could we gain by recording\n> offsets as an extension (I assume we need, like 4 offsets if the\n> system has 4 cores)? Much much more than this simpler model (which may\n> justify the complexity) or just \"meh\" compared to this?\n\nTo answer my own question, I ran a patched git to precalculate\nindividual thread parameters, removed the scheduler code and hard\ncoded these parameters (I ran just 4 threads, one per core). I got\n0m2.949s (webkit.git, 275k files, 100 read-cache runs). Compared to\n0m4.996s from Ben's patch (same test settings of course) I think it's\ndefinitely worth adding some extra complexity.\n-- \nDuy\n"},{"id":"356483","messageId":"CACsJy8Cnxz0w0g53Gb=_iXEdbSUFgssTozfxea0H52mWJ-RmTg@mail.gmail.com","threadId":"49204","inReplyTo":"CAGZ79kbXfPPvcQ1rnUdiOqWs5wC2qccGCnf8DvCVnp8QV126MA@mail.gmail.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-24T18:40:09Z","receivedAt":"2018-08-24T18:40:38Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Aug 23, 2018 at 7:33 PM Stefan Beller <sbeller@google.com> wrote:\n> > +core.fastIndex::\n> > +       Enable parallel index loading\n> > ++\n> > +This can speed up operations like 'git diff' and 'git status' especially\n> > +when the index is very large.  When enabled, Git will do the index\n> > +loading from the on disk format to the in-memory format in parallel.\n> > +Defaults to true.\n>\n> \"fast\" is a non-descriptive word as we try to be fast in any operation?\n> Maybe core.parallelIndexReading as that just describes what it\n> turns on/off, without second guessing its effects?\n\nAnother option is index.threads (the \"index\" section currently only\nhas one item, index.version). The value could be the same as\ngrep.threads or pack.threads.\n\n(and if you're thinking about parallelizing write as well but it\nshould be tuned differently, then perhaps index.readThreads, but I\ndon't think we need to go that far)\n-- \nDuy\n"},{"id":"356484","messageId":"2ba0a9f7-8073-e606-d433-490ea605466b@gmail.com","threadId":"49204","inReplyTo":"CACsJy8BXy_7QbDtF8bY5YzwJf=JUwiODv0zKxoSXeu4rJ+xjwg@mail.gmail.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-24T18:40:50Z","receivedAt":"2018-08-24T18:40:55Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/24/2018 2:20 PM, Duy Nguyen wrote:\n> On Fri, Aug 24, 2018 at 5:37 PM Duy Nguyen <pclouds@gmail.com> wrote:\n>> On Thu, Aug 23, 2018 at 10:36 PM Ben Peart <peartben@gmail.com> wrote:\n>>>> Nice to see this done without a new index extension that records\n>>>> offsets, so that we can load existing index files in parallel.\n>>>>\n>>>\n>>> Yes, I prefer this simpler model as well.  I wasn't sure it would\n>>> produce a significant improvement given the primary thread still has to\n>>> run through the variable length cache entries but was pleasantly surprised.\n>>\n>> Out of curiosity, how much time saving could we gain by recording\n>> offsets as an extension (I assume we need, like 4 offsets if the\n>> system has 4 cores)? Much much more than this simpler model (which may\n>> justify the complexity) or just \"meh\" compared to this?\n> \n> To answer my own question, I ran a patched git to precalculate\n> individual thread parameters, removed the scheduler code and hard\n> coded these parameters (I ran just 4 threads, one per core). I got\n> 0m2.949s (webkit.git, 275k files, 100 read-cache runs). Compared to\n> 0m4.996s from Ben's patch (same test settings of course) I think it's\n> definitely worth adding some extra complexity.\n> \n\nI took a run at doing that last year [1] but that was before the \nmem_pool work that allowed us to avoid the thread contention on the heap \nso the numbers aren't an apples to apples comparison (they would be \nbetter today).\n\nThe trade-off is the additional complexity to be able to load the index \nextension without having to parse through all the variable length cache \nentries.  My patch worked but there was feedback requested to make it \nmore generic and robust that I haven't gotten around to yet.\n\nThis patch series went for simplicity over absolutely the best possible \nperformance.\n\n[1] \nhttps://public-inbox.org/git/20171109141737.47976-1-benpeart@microsoft.com/\n"},{"id":"356485","messageId":"CACsJy8BtSYm0+Ku_+_F3S3aH1vMv5LPb=U4XCn-P-bvn-6yhjw@mail.gmail.com","threadId":"49204","inReplyTo":"2ba0a9f7-8073-e606-d433-490ea605466b@gmail.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-24T19:00:36Z","receivedAt":"2018-08-24T19:01:06Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Aug 24, 2018 at 8:40 PM Ben Peart <peartben@gmail.com> wrote:\n>\n>\n>\n> On 8/24/2018 2:20 PM, Duy Nguyen wrote:\n> > On Fri, Aug 24, 2018 at 5:37 PM Duy Nguyen <pclouds@gmail.com> wrote:\n> >> On Thu, Aug 23, 2018 at 10:36 PM Ben Peart <peartben@gmail.com> wrote:\n> >>>> Nice to see this done without a new index extension that records\n> >>>> offsets, so that we can load existing index files in parallel.\n> >>>>\n> >>>\n> >>> Yes, I prefer this simpler model as well.  I wasn't sure it would\n> >>> produce a significant improvement given the primary thread still has to\n> >>> run through the variable length cache entries but was pleasantly surprised.\n> >>\n> >> Out of curiosity, how much time saving could we gain by recording\n> >> offsets as an extension (I assume we need, like 4 offsets if the\n> >> system has 4 cores)? Much much more than this simpler model (which may\n> >> justify the complexity) or just \"meh\" compared to this?\n> >\n> > To answer my own question, I ran a patched git to precalculate\n> > individual thread parameters, removed the scheduler code and hard\n> > coded these parameters (I ran just 4 threads, one per core). I got\n> > 0m2.949s (webkit.git, 275k files, 100 read-cache runs). Compared to\n> > 0m4.996s from Ben's patch (same test settings of course) I think it's\n> > definitely worth adding some extra complexity.\n> >\n>\n> I took a run at doing that last year [1] but that was before the\n> mem_pool work that allowed us to avoid the thread contention on the heap\n> so the numbers aren't an apples to apples comparison (they would be\n> better today).\n\nAh.. sorry I was not aware. A big chunk of 2017 is blank to me when it\ncomes to git.\n\n> The trade-off is the additional complexity to be able to load the index\n> extension without having to parse through all the variable length cache\n> entries.  My patch worked but there was feedback requested to make it\n> more generic and robust that I haven't gotten around to yet.\n\nOne more comment. Instead of forcing this special index at the bottom,\nadd a generic one that gives positions of all extensions and put that\none at the bottom. Then you can still quickly locate your offset table\nextension, and you could load UNTR and TREE extensions in parallel too\n(those scale up to worktree size)\n\n> This patch series went for simplicity over absolutely the best possible\n> performance.\n\nWell, you know my stance on this now :) Not that it really matters.\n\n> [1]\n> https://public-inbox.org/git/20171109141737.47976-1-benpeart@microsoft.com/\n\nPS. I still think it's worth bring v4's performance back to v2. It's\nlow hanging fruit because I'm pretty sure Junio did not add v4 code\nwith cpu performance in mind. It was about file size at that time and\ncpu consumption was still dwarfed by hashing.\n-- \nDuy\n"},{"id":"356486","messageId":"c2994cf0-2530-52ba-bbc0-ba0b61af54c2@gmail.com","threadId":"49204","inReplyTo":"CACsJy8BtSYm0+Ku_+_F3S3aH1vMv5LPb=U4XCn-P-bvn-6yhjw@mail.gmail.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-24T19:57:40Z","receivedAt":"2018-08-24T19:57:45Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/24/2018 3:00 PM, Duy Nguyen wrote:\n> On Fri, Aug 24, 2018 at 8:40 PM Ben Peart <peartben@gmail.com> wrote:\n>>\n>>\n>>\n>> On 8/24/2018 2:20 PM, Duy Nguyen wrote:\n>>> On Fri, Aug 24, 2018 at 5:37 PM Duy Nguyen <pclouds@gmail.com> wrote:\n>>>> On Thu, Aug 23, 2018 at 10:36 PM Ben Peart <peartben@gmail.com> wrote:\n>>>>>> Nice to see this done without a new index extension that records\n>>>>>> offsets, so that we can load existing index files in parallel.\n>>>>>>\n>>>>>\n>>>>> Yes, I prefer this simpler model as well.  I wasn't sure it would\n>>>>> produce a significant improvement given the primary thread still has to\n>>>>> run through the variable length cache entries but was pleasantly surprised.\n>>>>\n>>>> Out of curiosity, how much time saving could we gain by recording\n>>>> offsets as an extension (I assume we need, like 4 offsets if the\n>>>> system has 4 cores)? Much much more than this simpler model (which may\n>>>> justify the complexity) or just \"meh\" compared to this?\n>>>\n>>> To answer my own question, I ran a patched git to precalculate\n>>> individual thread parameters, removed the scheduler code and hard\n>>> coded these parameters (I ran just 4 threads, one per core). I got\n>>> 0m2.949s (webkit.git, 275k files, 100 read-cache runs). Compared to\n>>> 0m4.996s from Ben's patch (same test settings of course) I think it's\n>>> definitely worth adding some extra complexity.\n>>>\n>>\n>> I took a run at doing that last year [1] but that was before the\n>> mem_pool work that allowed us to avoid the thread contention on the heap\n>> so the numbers aren't an apples to apples comparison (they would be\n>> better today).\n> \n> Ah.. sorry I was not aware. A big chunk of 2017 is blank to me when it\n> comes to git.\n> \n>> The trade-off is the additional complexity to be able to load the index\n>> extension without having to parse through all the variable length cache\n>> entries.  My patch worked but there was feedback requested to make it\n>> more generic and robust that I haven't gotten around to yet.\n> \n> One more comment. Instead of forcing this special index at the bottom,\n> add a generic one that gives positions of all extensions and put that\n> one at the bottom. Then you can still quickly locate your offset table\n> extension, and you could load UNTR and TREE extensions in parallel too\n> (those scale up to worktree size)\n> \n\nThat is pretty much what Junio's feedback was and what I was referring \nto as making it \"more generic.\"  The \"more robust\" was the request to \nadd a SHA to the extension ensure it wasn't corrupt and was a valid \nextension.\n\n>> This patch series went for simplicity over absolutely the best possible\n>> performance.\n> \n> Well, you know my stance on this now :) Not that it really matters.\n> \n>> [1]\n>> https://public-inbox.org/git/20171109141737.47976-1-benpeart@microsoft.com/\n> \n> PS. I still think it's worth bring v4's performance back to v2. It's\n> low hanging fruit because I'm pretty sure Junio did not add v4 code\n> with cpu performance in mind. It was about file size at that time and\n> cpu consumption was still dwarfed by hashing.\n> \n\nI see that as a nice follow up patch.  If the extension exists, use it \nand jump directly to the blocks and spin up threads.  If it doesn't \nexist, fall back to the code in this patch that has to find/compute the \nblocks on the fly.\n\n"},{"id":"356501","messageId":"20180825064458.28484-1-pclouds@gmail.com","threadId":"49204","inReplyTo":"20180824155734.GA6170@duynguyen.home","subject":"[PATCH] read-cache.c: optimize reading index format v4","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-25T06:44:58Z","receivedAt":"2018-08-25T06:45:44Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Index format v4 requires some more computation to assemble a path\nbased on a previous one. The current code is not very efficient\nbecause\n\n - it doubles memory copy, we assemble the final path in a temporary\n   first before putting it back to a cache_entry\n\n - strbuf_remove() in expand_name_field() is not exactly a good fit\n   for stripping a part at the end, _setlen() would do the same job\n   and is much cheaper.\n\n - the open-coded loop to find the end of the string in \n   expand_name_field() can't beat an optimized strlen()\n\nThis patch avoids the temporary buffer and writes directly to the new\ncache_entry, which addresses the first two points. The last point\ncould also be avoided if the total string length fits in the first 12\nbits of ce_flags, if not we fall back to strlen().\n\nRunning \"test-tool read-cache 100\" on webkit.git (275k files), reading\nv2 only takes 4.226 seconds, while v4 takes 5.711 seconds, 35% more\ntime. The patch reduces read time on v4 to 4.319 seconds.\n\nPS. I notice that v4 does not pad to align entries at 4 byte boundary\nlike v2/v3. This could cause a slight slow down on x86 and segfault on\nsome other platforms. We need to fix this in v5 when we introduce\nSHA-256 support in the index.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c | 124 +++++++++++++++++++++++----------------------------\n 1 file changed, 56 insertions(+), 68 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 7b1354d759..5c04c8f200 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1713,63 +1713,16 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *cache_entry_from_ondisk(struct mem_pool *mem_pool,\n-\t\t\t\t\t\t   struct ondisk_cache_entry *ondisk,\n-\t\t\t\t\t\t   unsigned int flags,\n-\t\t\t\t\t\t   const char *name,\n-\t\t\t\t\t\t   size_t len)\n-{\n-\tstruct cache_entry *ce = mem_pool__ce_alloc(mem_pool, len);\n-\n-\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n-\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n-\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n-\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n-\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n-\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n-\tce->ce_mode  = get_be32(&ondisk->mode);\n-\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n-\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n-\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n-\tce->ce_flags = flags & ~CE_NAMEMASK;\n-\tce->ce_namelen = len;\n-\tce->index = 0;\n-\thashcpy(ce->oid.hash, ondisk->sha1);\n-\tmemcpy(ce->name, name, len);\n-\tce->name[len] = '\\0';\n-\treturn ce;\n-}\n-\n-/*\n- * Adjacent cache entries tend to share the leading paths, so it makes\n- * sense to only store the differences in later entries.  In the v4\n- * on-disk format of the index, each on-disk cache entry stores the\n- * number of bytes to be stripped from the end of the previous name,\n- * and the bytes to append to the result, to come up with its name.\n- */\n-static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n-{\n-\tconst unsigned char *ep, *cp = (const unsigned char *)cp_;\n-\tsize_t len = decode_varint(&cp);\n-\n-\tif (name->len < len)\n-\t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n-\tstrbuf_add(name, cp, ep - cp);\n-\treturn (const char *)ep + 1 - cp_;\n-}\n-\n static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n-\t\t\t\t\t    struct strbuf *previous_name)\n+\t\t\t\t\t    const struct cache_entry *previous_ce)\n {\n \tstruct cache_entry *ce;\n \tsize_t len;\n \tconst char *name;\n \tunsigned int flags;\n+\tsize_t strip_len;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1782,28 +1735,61 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \t\textended_flags = get_be16(&ondisk2->flags2) << 16;\n \t\t/* We do not yet understand any bit out of CE_EXTENDED_FLAGS */\n \t\tif (extended_flags & ~CE_EXTENDED_FLAGS)\n-\t\t\tdie(\"Unknown index entry format %08x\", extended_flags);\n+\t\t\tdie(_(\"unknown index entry format %08x\"), extended_flags);\n \t\tflags |= extended_flags;\n \t\tname = ondisk2->name;\n \t}\n \telse\n \t\tname = ondisk->name;\n \n-\tif (!previous_name) {\n-\t\t/* v3 and earlier */\n-\t\tif (len == CE_NAMEMASK)\n-\t\t\tlen = strlen(name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags, name, len);\n+\t/*\n+\t * Adjacent cache entries tend to share the leading paths, so it makes\n+\t * sense to only store the differences in later entries.  In the v4\n+\t * on-disk format of the index, each on-disk cache entry stores the\n+\t * number of bytes to be stripped from the end of the previous name,\n+\t * and the bytes to append to the result, to come up with its name.\n+\t */\n+\tif (previous_ce) {\n+\t\tconst unsigned char *cp = (const unsigned char *)name;\n \n-\t\t*ent_size = ondisk_ce_size(ce);\n-\t} else {\n-\t\tunsigned long consumed;\n-\t\tconsumed = expand_name_field(previous_name, name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags,\n-\t\t\t\t\t     previous_name->buf,\n-\t\t\t\t\t     previous_name->len);\n+\t\tstrip_len = decode_varint(&cp);\n+\t\tif (previous_ce->ce_namelen < strip_len)\n+\t\t\tdie(_(\"malformed name field in the index, path '%s'\"),\n+\t\t\t    previous_ce->name);\n+\t\tname = (const char *)cp;\n+\t}\n \n-\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n+\tif (len == CE_NAMEMASK) {\n+\t\tlen = strlen(name);\n+\t\tif (previous_ce)\n+\t\t\tlen += previous_ce->ce_namelen - strip_len;\n+\t}\n+\n+\tce = mem_pool__ce_alloc(mem_pool, len);\n+\n+\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n+\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n+\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n+\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n+\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n+\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n+\tce->ce_mode  = get_be32(&ondisk->mode);\n+\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n+\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n+\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n+\tce->ce_flags = flags & ~CE_NAMEMASK;\n+\tce->ce_namelen = len;\n+\tce->index = 0;\n+\thashcpy(ce->oid.hash, ondisk->sha1);\n+\n+\tif (previous_ce) {\n+\t\tsize_t copy_len = previous_ce->ce_namelen - strip_len;\n+\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n+\t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n+\t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n+\t} else {\n+\t\tmemcpy(ce->name, name, len + 1);\n+\t\t*ent_size = ondisk_ce_size(ce);\n \t}\n \treturn ce;\n }\n@@ -1898,7 +1884,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tconst struct cache_entry *previous_ce = NULL;\n+\tstruct cache_entry *dummy_entry = NULL;\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1936,11 +1923,10 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->initialized = 1;\n \n \tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n+\t\tprevious_ce = dummy_entry = make_empty_transient_cache_entry(0);\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n \t} else {\n-\t\tprevious_name = NULL;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n \t}\n@@ -1952,12 +1938,14 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tunsigned long consumed;\n \n \t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n \t\tsrc_offset += consumed;\n+\t\tif (previous_ce)\n+\t\t\tprevious_ce = ce;\n \t}\n-\tstrbuf_release(&previous_name_buf);\n+\tfree(dummy_entry);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-- \n2.19.0.rc0.337.ge906d732e7\n\n"},{"id":"356605","messageId":"xmqqwosbiouc.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180825064458.28484-1-pclouds@gmail.com","subject":"Re: [PATCH] read-cache.c: optimize reading index format v4","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-27T19:36:27Z","receivedAt":"2018-08-27T19:36:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> Running \"test-tool read-cache 100\" on webkit.git (275k files), reading\n> v2 only takes 4.226 seconds, while v4 takes 5.711 seconds, 35% more\n> time. The patch reduces read time on v4 to 4.319 seconds.\n\nNice.\n\n> PS. I notice that v4 does not pad to align entries at 4 byte boundary\n> like v2/v3. This could cause a slight slow down on x86 and segfault on\n> some other platforms.\n\nCare to elaborate?  \n\nLong time ago, we used to mmap and read directly from the index file\ncontents, requiring either an unaligned read or padded entries.  But\nthat was eons ago and we first read and convert from on-disk using\nget_be32() etc. to in-core structure, so I am not sure what you mean\nby \"segfault\" here.\n\n> @@ -1782,28 +1735,61 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n>  \t\textended_flags = get_be16(&ondisk2->flags2) << 16;\n>  \t\t/* We do not yet understand any bit out of CE_EXTENDED_FLAGS */\n>  \t\tif (extended_flags & ~CE_EXTENDED_FLAGS)\n> -\t\t\tdie(\"Unknown index entry format %08x\", extended_flags);\n> +\t\t\tdie(_(\"unknown index entry format %08x\"), extended_flags);\n\nDo this as a separate preparation patch that is not controversial\nand can sail through without waiting for the rest of this patch.\n\nIn other words, don't slip in unreleted changes.\n\n> -\tif (!previous_name) {\n> -\t\t/* v3 and earlier */\n> -\t\tif (len == CE_NAMEMASK)\n> -\t\t\tlen = strlen(name);\n> -\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags, name, len);\n> +\t/*\n> +\t * Adjacent cache entries tend to share the leading paths, so it makes\n> +\t * sense to only store the differences in later entries.  In the v4\n> +\t * on-disk format of the index, each on-disk cache entry stores the\n> +\t * number of bytes to be stripped from the end of the previous name,\n> +\t * and the bytes to append to the result, to come up with its name.\n> +\t */\n> +\tif (previous_ce) {\n> +\t\tconst unsigned char *cp = (const unsigned char *)name;\n>  \n> -\t\t*ent_size = ondisk_ce_size(ce);\n> -\t} else {\n> -\t\tunsigned long consumed;\n> -\t\tconsumed = expand_name_field(previous_name, name);\n> -\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags,\n> -\t\t\t\t\t     previous_name->buf,\n> -\t\t\t\t\t     previous_name->len);\n> +\t\tstrip_len = decode_varint(&cp);\n> +\t\tif (previous_ce->ce_namelen < strip_len)\n> +\t\t\tdie(_(\"malformed name field in the index, path '%s'\"),\n> +\t\t\t    previous_ce->name);\n\nThe message is misleading; the previous is not the problematic one,\nbut the one that comes after it is.  Perhaps s/, path/, near path/\nor something.\n\n> +\t\tname = (const char *)cp;\n> +\t}\n>  \n> -\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n> +\tif (len == CE_NAMEMASK) {\n> +\t\tlen = strlen(name);\n> +\t\tif (previous_ce)\n> +\t\t\tlen += previous_ce->ce_namelen - strip_len;\n\nNicely done.  If the result fits in that 12-bit truncated name, then\nit is full so we do not need to adjust for strip.  Otherwise, we\nknow the length of this name is the sum of the part that is shared\nwith the previous one and the part that is unique to this one.\n\n> +\t}\n> +\n> +\tce = mem_pool__ce_alloc(mem_pool, len);\n> +\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n> +\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n> +\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n> +\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n> +\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n> +\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n> +\tce->ce_mode  = get_be32(&ondisk->mode);\n> +\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n> +\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n> +\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n> +\tce->ce_flags = flags & ~CE_NAMEMASK;\n> +\tce->ce_namelen = len;\n> +\tce->index = 0;\n> +\thashcpy(ce->oid.hash, ondisk->sha1);\n\nAgain, nice.  Now two callsites (both in this function) that call\ncache_entry_from_ondisk() with slightly different parameters are\nunified, there is no strong reason to have it as a single caller\nhelper function.\n\n> @@ -1898,7 +1884,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>  \tstruct cache_header *hdr;\n>  \tvoid *mmap;\n>  \tsize_t mmap_size;\n> -\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n> +\tconst struct cache_entry *previous_ce = NULL;\n> +\tstruct cache_entry *dummy_entry = NULL;\n>  \n>  \tif (istate->initialized)\n>  \t\treturn istate->cache_nr;\n> @@ -1936,11 +1923,10 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>  \tistate->initialized = 1;\n>  \n>  \tif (istate->version == 4) {\n> -\t\tprevious_name = &previous_name_buf;\n> +\t\tprevious_ce = dummy_entry = make_empty_transient_cache_entry(0);\n\nI do like the idea of passing the previous ce around to tell the\nnext one what the previous name was, but I would have preferred to\nsee this done a bit more cleanly without requiring us to support \"a\ndummy entry with name whose length is 0\"; a real cache entry never\nhas zero-length name, and our code may want to enforce it as a\nsanity check.\n\nI think we can just call create_from_disk() with NULL set to\nprevious_ce in the first round; of course, the logic to assign the\none we just created to previous_ce must check istate->version,\ninstead of \"is previous_ce NULL?\" (which is an indirect way to check\nthe same thing used in this patch).\n\nOther than that, looks quite nice.\n\n"},{"id":"356698","messageId":"00c3b1de-4fcf-ba03-0f8d-9ea2540ba657@gmail.com","threadId":"49204","inReplyTo":"CACsJy8Cnxz0w0g53Gb=_iXEdbSUFgssTozfxea0H52mWJ-RmTg@mail.gmail.com","subject":"Re: [PATCH v1] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-28T14:53:34Z","receivedAt":"2018-08-28T14:52:52Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/24/2018 2:40 PM, Duy Nguyen wrote:\n> On Thu, Aug 23, 2018 at 7:33 PM Stefan Beller <sbeller@google.com> wrote:\n>>> +core.fastIndex::\n>>> +       Enable parallel index loading\n>>> ++\n>>> +This can speed up operations like 'git diff' and 'git status' especially\n>>> +when the index is very large.  When enabled, Git will do the index\n>>> +loading from the on disk format to the in-memory format in parallel.\n>>> +Defaults to true.\n>> \"fast\" is a non-descriptive word as we try to be fast in any operation?\n>> Maybe core.parallelIndexReading as that just describes what it\n>> turns on/off, without second guessing its effects?\n> Another option is index.threads (the \"index\" section currently only\n> has one item, index.version). The value could be the same as\n> grep.threads or pack.threads.\n>\n> (and if you're thinking about parallelizing write as well but it\n> should be tuned differently, then perhaps index.readThreads, but I\n> don't think we need to go that far)\n\nI like that.  I'll switch to index.threads and make 'true' or '0' mean \n\"automatically determine the number of threads to use\" similar to \npack.threads.\n"},{"id":"356720","messageId":"CACsJy8B38QAW8qq-CctLJyJNaC329o6Rr1gs0kd=EkV+ARAaVw@mail.gmail.com","threadId":"49204","inReplyTo":"xmqqwosbiouc.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH] read-cache.c: optimize reading index format v4","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-08-28T19:25:19Z","receivedAt":"2018-08-28T19:25:48Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 27, 2018 at 9:36 PM Junio C Hamano <gitster@pobox.com> wrote:\n> > PS. I notice that v4 does not pad to align entries at 4 byte boundary\n> > like v2/v3. This could cause a slight slow down on x86 and segfault on\n> > some other platforms.\n>\n> Care to elaborate?\n>\n> Long time ago, we used to mmap and read directly from the index file\n> contents, requiring either an unaligned read or padded entries.  But\n> that was eons ago and we first read and convert from on-disk using\n> get_be32() etc. to in-core structure, so I am not sure what you mean\n> by \"segfault\" here.\n>\n\nMy bad. I saw this line\n\n#define get_be16(p) ntohs(*(unsigned short *)(p))\n\nand jumped to conclusion without realizing that block is for safe\nunaligned access.\n\n> > @@ -1898,7 +1884,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n> >       struct cache_header *hdr;\n> >       void *mmap;\n> >       size_t mmap_size;\n> > -     struct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n> > +     const struct cache_entry *previous_ce = NULL;\n> > +     struct cache_entry *dummy_entry = NULL;\n> >\n> >       if (istate->initialized)\n> >               return istate->cache_nr;\n> > @@ -1936,11 +1923,10 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n> >       istate->initialized = 1;\n> >\n> >       if (istate->version == 4) {\n> > -             previous_name = &previous_name_buf;\n> > +             previous_ce = dummy_entry = make_empty_transient_cache_entry(0);\n>\n> I do like the idea of passing the previous ce around to tell the\n> next one what the previous name was, but I would have preferred to\n> see this done a bit more cleanly without requiring us to support \"a\n> dummy entry with name whose length is 0\"; a real cache entry never\n> has zero-length name, and our code may want to enforce it as a\n> sanity check.\n>\n> I think we can just call create_from_disk() with NULL set to\n> previous_ce in the first round; of course, the logic to assign the\n> one we just created to previous_ce must check istate->version,\n> instead of \"is previous_ce NULL?\" (which is an indirect way to check\n> the same thing used in this patch).\n\nYeah I kinda hated dummy_entry too but the feeling wasn't strong\nenough to move towards the index->version check. I guess I'm going to\ndo it now.\n-- \nDuy\n"},{"id":"356780","messageId":"b3898649-47ee-adcc-9e22-68b23ffde5d1@gmail.com","threadId":"49204","inReplyTo":"CACsJy8B38QAW8qq-CctLJyJNaC329o6Rr1gs0kd=EkV+ARAaVw@mail.gmail.com","subject":"Re: [PATCH] read-cache.c: optimize reading index format v4","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-28T23:54:35Z","receivedAt":"2018-08-28T23:54:40Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/28/2018 3:25 PM, Duy Nguyen wrote:\n> On Mon, Aug 27, 2018 at 9:36 PM Junio C Hamano <gitster@pobox.com> wrote:\n>>> PS. I notice that v4 does not pad to align entries at 4 byte boundary\n>>> like v2/v3. This could cause a slight slow down on x86 and segfault on\n>>> some other platforms.\n>>\n>> Care to elaborate?\n>>\n>> Long time ago, we used to mmap and read directly from the index file\n>> contents, requiring either an unaligned read or padded entries.  But\n>> that was eons ago and we first read and convert from on-disk using\n>> get_be32() etc. to in-core structure, so I am not sure what you mean\n>> by \"segfault\" here.\n>>\n> \n> My bad. I saw this line\n> \n> #define get_be16(p) ntohs(*(unsigned short *)(p))\n> \n> and jumped to conclusion without realizing that block is for safe\n> unaligned access.\n> \n>>> @@ -1898,7 +1884,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>>>        struct cache_header *hdr;\n>>>        void *mmap;\n>>>        size_t mmap_size;\n>>> -     struct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n>>> +     const struct cache_entry *previous_ce = NULL;\n>>> +     struct cache_entry *dummy_entry = NULL;\n>>>\n>>>        if (istate->initialized)\n>>>                return istate->cache_nr;\n>>> @@ -1936,11 +1923,10 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>>>        istate->initialized = 1;\n>>>\n>>>        if (istate->version == 4) {\n>>> -             previous_name = &previous_name_buf;\n>>> +             previous_ce = dummy_entry = make_empty_transient_cache_entry(0);\n>>\n>> I do like the idea of passing the previous ce around to tell the\n>> next one what the previous name was, but I would have preferred to\n>> see this done a bit more cleanly without requiring us to support \"a\n>> dummy entry with name whose length is 0\"; a real cache entry never\n>> has zero-length name, and our code may want to enforce it as a\n>> sanity check.\n>>\n>> I think we can just call create_from_disk() with NULL set to\n>> previous_ce in the first round; of course, the logic to assign the\n>> one we just created to previous_ce must check istate->version,\n>> instead of \"is previous_ce NULL?\" (which is an indirect way to check\n>> the same thing used in this patch).\n> \n> Yeah I kinda hated dummy_entry too but the feeling wasn't strong\n> enough to move towards the index->version check. I guess I'm going to\n> do it now.\n> \n\nI ran some perf tests using p0002-read-cache.sh to compare V4 \nperformance before and after this patch so I could get a feel for how \nmuch it helps.\n\n100,000 files\n\nTest                                  HEAD~1   HEAD\n------------------------------------------------------------\nread_cache/discard_cache 1000 times    14.12    10.75 -23.9%\n\n1,000,000 files\n\nTest                                  HEAD~1   HEAD\n------------------------------------------------------------\nread_cache/discard_cache 1000 times   202.81   170.33 -16.0%\n\n\nThis provides a nice speedup and IMO simplifies the code as well. \nNicely done.\n"},{"id":"356846","messageId":"20180829152500.46640-3-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180829152500.46640-1-benpeart@microsoft.com","subject":"[PATCH v2 2/3] read-cache: load cache extensions on worker thread","fromName":"Ben Peart","fromEmail":"ben.peart@microsoft.com","sentAt":"2018-08-29T15:25:20Z","receivedAt":"2018-08-29T15:25:40Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by loading\nthe cache extensions on a worker thread in parallel with loading the cache\nentries.\n\nThis is possible because the current extensions don't access the cache\nentries in the index_state structure so are OK that they don't all exist\nyet.\n\nThe CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\nextensions don't even get a pointer to the index so don't have access to the\ncache entries.\n\nCACHE_EXT_LINK only uses the index_state to initialize the split index.\nCACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\nupdate and dirty flags.\n\nI used p0002-read-cache.sh to generate some performance data on the\ncumulative impact:\n\n100,000 entries\n\nTest                                HEAD~3           HEAD~2\n---------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 14.08(0.01+0.10) 9.72(0.03+0.06) -31.0%\n\n1,000,000 entries\n\nTest                                HEAD~3            HEAD~2\n------------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 202.95(0.01+0.07) 154.14(0.03+0.06) -24.1%\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 60 +++++++++++++++++++++++++++++++++++++++++-----------\n 1 file changed, 48 insertions(+), 12 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex c30346388a..f768004617 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1959,16 +1959,13 @@ struct load_cache_entries_thread_data\n \tstruct mem_pool *ce_mem_pool;\n \tint offset, nr;\n \tvoid *mmap;\n+\tsize_t mmap_size;\n \tunsigned long start_offset;\n \tstruct strbuf previous_name_buf;\n \tstruct strbuf *previous_name;\n \tunsigned long consumed;\t/* return # of bytes in index file processed */\n };\n \n-/*\n-* A thread proc to run the load_cache_entries() computation\n-* across multiple background threads.\n-*/\n static void *load_cache_entries_thread(void *_data)\n {\n \tstruct load_cache_entries_thread_data *p = _data;\n@@ -1978,6 +1975,36 @@ static void *load_cache_entries_thread(void *_data)\n \treturn NULL;\n }\n \n+static void *load_index_extensions_thread(void *_data)\n+{\n+\tstruct load_cache_entries_thread_data *p = _data;\n+\tunsigned long src_offset = p->start_offset;\n+\n+\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\t\tif (read_index_extension(p->istate,\n+\t\t\t\t\t\t\t\t(const char *)p->mmap + src_offset,\n+\t\t\t\t\t\t\t\t(char *)p->mmap + src_offset + 8,\n+\t\t\t\t\t\t\t\textsize) < 0) {\n+\t\t\tmunmap(p->mmap, p->mmap_size);\n+\t\t\tdie(\"index file corrupt\");\n+\t\t}\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\tp->consumed += src_offset - p->start_offset;\n+\n+\treturn NULL;\n+}\n+\n static unsigned long load_cache_entries(struct index_state *istate,\n \t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n {\n@@ -2012,16 +2039,16 @@ static unsigned long load_cache_entries(struct index_state *istate,\n \telse\n \t\tprevious_name = NULL;\n \n+\t/* allocate an extra thread for loading the index extensions */\n \tce_per_thread = DIV_ROUND_UP(istate->cache_nr, nr_threads);\n-\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n+\tdata = xcalloc(nr_threads + 1, sizeof(struct load_cache_entries_thread_data));\n \n \t/*\n \t * Loop through index entries starting a thread for every ce_per_thread\n-\t * entries. Exit the loop when we've created the final thread (no need\n-\t * to parse the remaining entries.\n+\t * entries.\n \t */\n \tconsumed = thread = 0;\n-\tfor (i = 0; ; i++) {\n+\tfor (i = 0; i < istate->cache_nr; i++) {\n \t\tstruct ondisk_cache_entry *ondisk;\n \t\tconst char *name;\n \t\tunsigned int flags;\n@@ -2055,9 +2082,7 @@ static unsigned long load_cache_entries(struct index_state *istate,\n \t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n \t\t\t\tdie(\"unable to create load_cache_entries_thread\");\n \n-\t\t\t/* exit the loop when we've created the last thread */\n-\t\t\tif (++thread == nr_threads)\n-\t\t\t\tbreak;\n+\t\t\t++thread;\n \t\t}\n \n \t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n@@ -2086,7 +2111,18 @@ static unsigned long load_cache_entries(struct index_state *istate,\n \t\t\tsrc_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n \t}\n \n-\tfor (i = 0; i < nr_threads; i++) {\n+\t/* create a thread to load the index extensions */\n+\tstruct load_cache_entries_thread_data *p = &data[thread];\n+\tp->istate = istate;\n+\tmem_pool_init(&p->ce_mem_pool, 0);\n+\tp->mmap = mmap;\n+\tp->mmap_size = mmap_size;\n+\tp->start_offset = src_offset;\n+\n+\tif (pthread_create(&p->pthread, NULL, load_index_extensions_thread, p))\n+\t\tdie(\"unable to create load_index_extensions_thread\");\n+\n+\tfor (i = 0; i < nr_threads + 1; i++) {\n \t\tstruct load_cache_entries_thread_data *p = data + i;\n \t\tif (pthread_join(p->pthread, NULL))\n \t\t\tdie(\"unable to join load_cache_entries_thread\");\n-- \n2.18.0.windows.1\n\n"},{"id":"356847","messageId":"20180829152500.46640-4-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180829152500.46640-1-benpeart@microsoft.com","subject":"[PATCH v2 3/3] read-cache: micro-optimize expand_name_field() to speed up V4 index parsing.","fromName":"Ben Peart","fromEmail":"ben.peart@microsoft.com","sentAt":"2018-08-29T15:25:21Z","receivedAt":"2018-08-29T15:25:42Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":" - strbuf_remove() in expand_name_field() is not exactly a good fit\n   for stripping a part at the end, _setlen() would do the same job\n   and is much cheaper.\n\n - the open-coded loop to find the end of the string in\n   expand_name_field() can't beat an optimized strlen()\n\nI used p0002-read-cache.sh to generate some performance data on the\ncumulative impact:\n\n100,000 files\n\nTest                                HEAD~3           HEAD\n---------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 14.08(0.03+0.09) 8.71(0.01+0.09) -38.1%\n\n1,000,000 files\n\nTest                                HEAD~3            HEAD\n------------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 201.77(0.03+0.07) 149.68(0.04+0.07) -25.8%\n\nSuggested by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 5 ++---\n 1 file changed, 2 insertions(+), 3 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex f768004617..f5e7c86c42 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1754,9 +1754,8 @@ static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n \n \tif (name->len < len)\n \t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n+\tstrbuf_setlen(name, name->len - len);\n+\tep = cp + strlen((const char *)cp);\n \tstrbuf_add(name, cp, ep - cp);\n \treturn (const char *)ep + 1 - cp_;\n }\n-- \n2.18.0.windows.1\n\n"},{"id":"356848","messageId":"20180829152500.46640-2-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180829152500.46640-1-benpeart@microsoft.com","subject":"[PATCH v2 1/3] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"ben.peart@microsoft.com","sentAt":"2018-08-29T15:25:19Z","receivedAt":"2018-08-29T15:25:58Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by creating\nmultiple threads to divide the work of loading and converting the cache\nentries across all available CPU cores.\n\nIt accomplishes this by having the primary thread loop across the index file\ntracking the offset and (for V4 indexes) expanding the name. It creates a\nthread to process each block of entries as it comes to them. Once the\nthreads are complete and the cache entries are loaded, the rest of the\nextensions can be loaded and processed normally on the primary thread.\n\nI used p0002-read-cache.sh to generate some performance data:\n\n100,000 entries\n\nTest                                HEAD~3           HEAD~2\n---------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 14.02(0.01+0.12) 9.81(0.01+0.07) -30.0%\n\n1,000,000 entries\n\nTest                                HEAD~3            HEAD~2\n------------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 202.06(0.06+0.09) 155.72(0.03+0.06) -22.9%\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n Documentation/config.txt |   6 +\n config.c                 |  14 +++\n config.h                 |   1 +\n read-cache.c             | 240 +++++++++++++++++++++++++++++++++++----\n 4 files changed, 237 insertions(+), 24 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex 1c42364988..79f8296d9c 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -2391,6 +2391,12 @@ imap::\n \tThe configuration variables in the 'imap' section are described\n \tin linkgit:git-imap-send[1].\n \n+index.threads::\n+\tSpecifies the number of threads to spawn when loading the index.\n+\tThis is meant to reduce index load time on multiprocessor machines.\n+\tSpecifying 0 or 'true' will cause Git to auto-detect the number of\n+\tCPU's and set the number of threads accordingly. Defaults to 'true'.\n+\n index.version::\n \tSpecify the version with which new index files should be\n \tinitialized.  This does not affect existing repositories.\ndiff --git a/config.c b/config.c\nindex 9a0b10d4bc..3bda124550 100644\n--- a/config.c\n+++ b/config.c\n@@ -2289,6 +2289,20 @@ int git_config_get_fsmonitor(void)\n \treturn 0;\n }\n \n+int git_config_get_index_threads(void)\n+{\n+\tint is_bool, val;\n+\n+\tif (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n+\t\tif (is_bool)\n+\t\t\treturn val ? 0 : 1;\n+\t\telse\n+\t\t\treturn val;\n+\t}\n+\n+\treturn 0; /* auto-detect */\n+}\n+\n NORETURN\n void git_die_config_linenr(const char *key, const char *filename, int linenr)\n {\ndiff --git a/config.h b/config.h\nindex ab46e0165d..a06027e69b 100644\n--- a/config.h\n+++ b/config.h\n@@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n extern int git_config_get_split_index(void);\n extern int git_config_get_max_percent_split_change(void);\n extern int git_config_get_fsmonitor(void);\n+extern int git_config_get_index_threads(void);\n \n /* This dies if the configured or default date is in the future */\n extern int git_config_get_expiry(const char *key, const char **output);\ndiff --git a/read-cache.c b/read-cache.c\nindex 7b1354d759..c30346388a 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1889,16 +1889,229 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+/*\n+ * A helper function that will load the specified range of cache entries\n+ * from the memory mapped file and add them to the given index.\n+ */\n+static unsigned long load_cache_entry_block(struct index_state *istate,\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n+\t\t\tunsigned long start_offset, struct strbuf *previous_name)\n+{\n+\tint i;\n+\tunsigned long src_offset = start_offset;\n+\n+\tfor (i = offset; i < offset + nr; i++) {\n+\t\tstruct ondisk_cache_entry *disk_ce;\n+\t\tstruct cache_entry *ce;\n+\t\tunsigned long consumed;\n+\n+\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tset_index_entry(istate, i, ce);\n+\n+\t\tsrc_offset += consumed;\n+\t}\n+\treturn src_offset - start_offset;\n+}\n+\n+static unsigned long load_all_cache_entries(struct index_state *istate,\n+\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tunsigned long consumed;\n+\n+\tif (istate->version == 4) {\n+\t\tprevious_name = &previous_name_buf;\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n+\t} else {\n+\t\tprevious_name = NULL;\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n+\t}\n+\n+\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, previous_name);\n+\tstrbuf_release(&previous_name_buf);\n+\treturn consumed;\n+}\n+\n+#ifdef NO_PTHREADS\n+\n+#define load_cache_entries load_all_cache_entries\n+\n+#else\n+\n+#include \"thread-utils.h\"\n+\n+/*\n+* Mostly randomly chosen maximum thread counts: we\n+* cap the parallelism to online_cpus() threads, and we want\n+* to have at least 7500 cache entries per thread for it to\n+* be worth starting a thread.\n+*/\n+#define THREAD_COST\t\t(7500)\n+\n+struct load_cache_entries_thread_data\n+{\n+\tpthread_t pthread;\n+\tstruct index_state *istate;\n+\tstruct mem_pool *ce_mem_pool;\n+\tint offset, nr;\n+\tvoid *mmap;\n+\tunsigned long start_offset;\n+\tstruct strbuf previous_name_buf;\n+\tstruct strbuf *previous_name;\n+\tunsigned long consumed;\t/* return # of bytes in index file processed */\n+};\n+\n+/*\n+* A thread proc to run the load_cache_entries() computation\n+* across multiple background threads.\n+*/\n+static void *load_cache_entries_thread(void *_data)\n+{\n+\tstruct load_cache_entries_thread_data *p = _data;\n+\n+\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n+\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n+\treturn NULL;\n+}\n+\n+static unsigned long load_cache_entries(struct index_state *istate,\n+\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tstruct load_cache_entries_thread_data *data;\n+\tint nr_threads, cpus, ce_per_thread;\n+\tunsigned long consumed;\n+\tint i, thread;\n+\n+\tnr_threads = git_config_get_index_threads();\n+\tif (!nr_threads) {\n+\t\tcpus = online_cpus();\n+\t\tnr_threads = istate->cache_nr / THREAD_COST;\n+\t\tif (nr_threads > cpus)\n+\t\t\tnr_threads = cpus;\n+\t}\n+\n+\t/* enable testing with fewer than default minimum of entries */\n+\tif ((istate->cache_nr > 1) && (nr_threads < 2) && git_env_bool(\"GIT_INDEX_THREADS_TEST\", 0))\n+\t\tnr_threads = 2;\n+\n+\tif (nr_threads < 2)\n+\t\treturn load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+\n+\t/* a little sanity checking */\n+\tif (istate->name_hash_initialized)\n+\t\tdie(\"the name hash isn't thread safe\");\n+\n+\tmem_pool_init(&istate->ce_mem_pool, 0);\n+\tif (istate->version == 4)\n+\t\tprevious_name = &previous_name_buf;\n+\telse\n+\t\tprevious_name = NULL;\n+\n+\tce_per_thread = DIV_ROUND_UP(istate->cache_nr, nr_threads);\n+\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n+\n+\t/*\n+\t * Loop through index entries starting a thread for every ce_per_thread\n+\t * entries. Exit the loop when we've created the final thread (no need\n+\t * to parse the remaining entries.\n+\t */\n+\tconsumed = thread = 0;\n+\tfor (i = 0; ; i++) {\n+\t\tstruct ondisk_cache_entry *ondisk;\n+\t\tconst char *name;\n+\t\tunsigned int flags;\n+\n+\t\t/*\n+\t\t * we've reached the beginning of a block of cache entries,\n+\t\t * kick off a thread to process them\n+\t\t */\n+\t\tif (0 == i % ce_per_thread) {\n+\t\t\tstruct load_cache_entries_thread_data *p = &data[thread];\n+\n+\t\t\tp->istate = istate;\n+\t\t\tp->offset = i;\n+\t\t\tp->nr = ce_per_thread < istate->cache_nr - i ? ce_per_thread : istate->cache_nr - i;\n+\n+\t\t\t/* create a mem_pool for each thread */\n+\t\t\tif (istate->version == 4)\n+\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\t\t\t  estimate_cache_size_from_compressed(p->nr));\n+\t\t\telse\n+\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\t\t\t  estimate_cache_size(mmap_size, p->nr));\n+\n+\t\t\tp->mmap = mmap;\n+\t\t\tp->start_offset = src_offset;\n+\t\t\tif (previous_name) {\n+\t\t\t\tstrbuf_addbuf(&p->previous_name_buf, previous_name);\n+\t\t\t\tp->previous_name = &p->previous_name_buf;\n+\t\t\t}\n+\n+\t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n+\t\t\t\tdie(\"unable to create load_cache_entries_thread\");\n+\n+\t\t\t/* exit the loop when we've created the last thread */\n+\t\t\tif (++thread == nr_threads)\n+\t\t\t\tbreak;\n+\t\t}\n+\n+\t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\n+\t\t/* On-disk flags are just 16 bits */\n+\t\tflags = get_be16(&ondisk->flags);\n+\n+\t\tif (flags & CE_EXTENDED) {\n+\t\t\tstruct ondisk_cache_entry_extended *ondisk2;\n+\t\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n+\t\t\tname = ondisk2->name;\n+\t\t} else\n+\t\t\tname = ondisk->name;\n+\n+\t\tif (!previous_name) {\n+\t\t\tsize_t len;\n+\n+\t\t\t/* v3 and earlier */\n+\t\t\tlen = flags & CE_NAMEMASK;\n+\t\t\tif (len == CE_NAMEMASK)\n+\t\t\t\tlen = strlen(name);\n+\t\t\tsrc_offset += (flags & CE_EXTENDED) ?\n+\t\t\t\tondisk_cache_entry_extended_size(len) :\n+\t\t\t\tondisk_cache_entry_size(len);\n+\t\t} else\n+\t\t\tsrc_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n+\t}\n+\n+\tfor (i = 0; i < nr_threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = data + i;\n+\t\tif (pthread_join(p->pthread, NULL))\n+\t\t\tdie(\"unable to join load_cache_entries_thread\");\n+\t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n+\t\tstrbuf_release(&p->previous_name_buf);\n+\t\tconsumed += p->consumed;\n+\t}\n+\n+\tfree(data);\n+\tstrbuf_release(&previous_name_buf);\n+\n+\treturn consumed;\n+}\n+\n+#endif\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n-\tint fd, i;\n+\tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1935,29 +2148,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n \tistate->initialized = 1;\n \n-\tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n-\t} else {\n-\t\tprevious_name = NULL;\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n-\t}\n-\n \tsrc_offset = sizeof(*hdr);\n-\tfor (i = 0; i < istate->cache_nr; i++) {\n-\t\tstruct ondisk_cache_entry *disk_ce;\n-\t\tstruct cache_entry *ce;\n-\t\tunsigned long consumed;\n-\n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_name);\n-\t\tset_index_entry(istate, i, ce);\n-\n-\t\tsrc_offset += consumed;\n-\t}\n-\tstrbuf_release(&previous_name_buf);\n+\tsrc_offset += load_cache_entries(istate, mmap, mmap_size, src_offset);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-- \n2.18.0.windows.1\n\n"},{"id":"356849","messageId":"20180829152500.46640-1-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180823154053.20212-1-benpeart@microsoft.com","subject":"[PATCH v2 0/3] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"ben.peart@microsoft.com","sentAt":"2018-08-29T15:25:18Z","receivedAt":"2018-08-29T15:26:27Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"The big changes in this itteration are:\n\n- Switched to index.threads to provide control over the use of threading\n\n- Added another worker thread to load the index extensions in parallel\n\n- Applied optimization expand_name_field() suggested by Duy\n\nThe net result of these optimizations is a savings of 25.8% (1,000,000 files)\nto 38.1% (100,000 files) as measured by p0002-read-cache.sh.\n\nThis patch conflicts with Duy's patch to remove the double memory copy and\npass in the previous ce instead.  The two will need to be merged/reconciled\nonce they settle down a bit.\n\n\nBase Ref: master\nWeb-Diff: https://github.com/benpeart/git/commit/39f2b0f5fe\nCheckout: git fetch https://github.com/benpeart/git read-index-multithread-v2 && git checkout 39f2b0f5fe\n\n\n### Interdiff (v1..v2):\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex 3344685cc4..79f8296d9c 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -899,14 +899,6 @@ relatively high IO latencies.  When enabled, Git will do the\n index comparison to the filesystem data in parallel, allowing\n overlapping IO's.  Defaults to true.\n \n-core.fastIndex::\n-       Enable parallel index loading\n-+\n-This can speed up operations like 'git diff' and 'git status' especially\n-when the index is very large.  When enabled, Git will do the index\n-loading from the on disk format to the in-memory format in parallel.\n-Defaults to true.\n-\n core.createObject::\n \tYou can set this to 'link', in which case a hardlink followed by\n \ta delete of the source are used to make sure that object creation\n@@ -2399,6 +2391,12 @@ imap::\n \tThe configuration variables in the 'imap' section are described\n \tin linkgit:git-imap-send[1].\n \n+index.threads::\n+\tSpecifies the number of threads to spawn when loading the index.\n+\tThis is meant to reduce index load time on multiprocessor machines.\n+\tSpecifying 0 or 'true' will cause Git to auto-detect the number of\n+\tCPU's and set the number of threads accordingly. Defaults to 'true'.\n+\n index.version::\n \tSpecify the version with which new index files should be\n \tinitialized.  This does not affect existing repositories.\ndiff --git a/config.c b/config.c\nindex 883092fdd3..3bda124550 100644\n--- a/config.c\n+++ b/config.c\n@@ -2289,17 +2289,18 @@ int git_config_get_fsmonitor(void)\n \treturn 0;\n }\n \n-int git_config_get_fast_index(void)\n+int git_config_get_index_threads(void)\n {\n-\tint val;\n+\tint is_bool, val;\n \n-\tif (!git_config_get_maybe_bool(\"core.fastindex\", &val))\n+\tif (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n+\t\tif (is_bool)\n+\t\t\treturn val ? 0 : 1;\n+\t\telse\n \t\t\treturn val;\n+\t}\n \n-\tif (getenv(\"GIT_FASTINDEX_TEST\"))\n-\t\treturn 1;\n-\n-\treturn -1; /* default value */\n+\treturn 0; /* auto-detect */\n }\n \n NORETURN\ndiff --git a/config.h b/config.h\nindex 74ca4e7db5..a06027e69b 100644\n--- a/config.h\n+++ b/config.h\n@@ -250,7 +250,7 @@ extern int git_config_get_untracked_cache(void);\n extern int git_config_get_split_index(void);\n extern int git_config_get_max_percent_split_change(void);\n extern int git_config_get_fsmonitor(void);\n-extern int git_config_get_fast_index(void);\n+extern int git_config_get_index_threads(void);\n \n /* This dies if the configured or default date is in the future */\n extern int git_config_get_expiry(const char *key, const char **output);\ndiff --git a/read-cache.c b/read-cache.c\nindex 0fa7e1a04c..f5e7c86c42 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -24,10 +24,6 @@\n #include \"utf8.h\"\n #include \"fsmonitor.h\"\n \n-#ifndef min\n-#define min(a,b) (((a) < (b)) ? (a) : (b))\n-#endif\n-\n /* Mask for the name length in ce_flags in the on-disk index */\n \n #define CE_NAMEMASK  (0x0fff)\n@@ -1758,9 +1754,8 @@ static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n \n \tif (name->len < len)\n \t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n+\tstrbuf_setlen(name, name->len - len);\n+\tep = cp + strlen((const char *)cp);\n \tstrbuf_add(name, cp, ep - cp);\n \treturn (const char *)ep + 1 - cp_;\n }\n@@ -1893,7 +1888,13 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n-static unsigned long load_cache_entry_block(struct index_state *istate, struct mem_pool *ce_mem_pool, int offset, int nr, void *mmap, unsigned long start_offset, struct strbuf *previous_name)\n+/*\n+ * A helper function that will load the specified range of cache entries\n+ * from the memory mapped file and add them to the given index.\n+ */\n+static unsigned long load_cache_entry_block(struct index_state *istate,\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n+\t\t\tunsigned long start_offset, struct strbuf *previous_name)\n {\n \tint i;\n \tunsigned long src_offset = start_offset;\n@@ -1912,7 +1913,8 @@ static unsigned long load_cache_entry_block(struct index_state *istate, struct m\n \treturn src_offset - start_offset;\n }\n \n-static unsigned long load_all_cache_entries(struct index_state *istate, void *mmap, size_t mmap_size, unsigned long src_offset)\n+static unsigned long load_all_cache_entries(struct index_state *istate,\n+\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n {\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tunsigned long consumed;\n@@ -1927,7 +1929,8 @@ static unsigned long load_all_cache_entries(struct index_state *istate, void *mm\n \t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n \t}\n \n-\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool, 0, istate->cache_nr, mmap, src_offset, previous_name);\n+\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, previous_name);\n \tstrbuf_release(&previous_name_buf);\n \treturn consumed;\n }\n@@ -1955,67 +1958,110 @@ struct load_cache_entries_thread_data\n \tstruct mem_pool *ce_mem_pool;\n \tint offset, nr;\n \tvoid *mmap;\n+\tsize_t mmap_size;\n \tunsigned long start_offset;\n \tstruct strbuf previous_name_buf;\n \tstruct strbuf *previous_name;\n \tunsigned long consumed;\t/* return # of bytes in index file processed */\n };\n \n-/*\n-* A thread proc to run the load_cache_entries() computation\n-* across multiple background threads.\n-*/\n static void *load_cache_entries_thread(void *_data)\n {\n \tstruct load_cache_entries_thread_data *p = _data;\n \n-\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool, p->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n+\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n+\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n+\treturn NULL;\n+}\n+\n+static void *load_index_extensions_thread(void *_data)\n+{\n+\tstruct load_cache_entries_thread_data *p = _data;\n+\tunsigned long src_offset = p->start_offset;\n+\n+\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\t\tif (read_index_extension(p->istate,\n+\t\t\t\t\t\t\t\t(const char *)p->mmap + src_offset,\n+\t\t\t\t\t\t\t\t(char *)p->mmap + src_offset + 8,\n+\t\t\t\t\t\t\t\textsize) < 0) {\n+\t\t\tmunmap(p->mmap, p->mmap_size);\n+\t\t\tdie(\"index file corrupt\");\n+\t\t}\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\tp->consumed += src_offset - p->start_offset;\n+\n \treturn NULL;\n }\n \n-static unsigned long load_cache_entries(struct index_state *istate, void *mmap, size_t mmap_size, unsigned long src_offset)\n+static unsigned long load_cache_entries(struct index_state *istate,\n+\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n {\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tstruct load_cache_entries_thread_data *data;\n-\tint threads, cpus, thread_nr;\n+\tint nr_threads, cpus, ce_per_thread;\n \tunsigned long consumed;\n \tint i, thread;\n \n+\tnr_threads = git_config_get_index_threads();\n+\tif (!nr_threads) {\n \t\tcpus = online_cpus();\n-\tthreads = istate->cache_nr / THREAD_COST;\n-\tif (threads > cpus)\n-\t\tthreads = cpus;\n+\t\tnr_threads = istate->cache_nr / THREAD_COST;\n+\t\tif (nr_threads > cpus)\n+\t\t\tnr_threads = cpus;\n+\t}\n \n \t/* enable testing with fewer than default minimum of entries */\n-\tif ((istate->cache_nr > 1) && (threads < 2) && getenv(\"GIT_FASTINDEX_TEST\"))\n-\t\tthreads = 2;\n+\tif ((istate->cache_nr > 1) && (nr_threads < 2) && git_env_bool(\"GIT_INDEX_THREADS_TEST\", 0))\n+\t\tnr_threads = 2;\n \n-\tif (threads < 2 || !git_config_get_fast_index())\n+\tif (nr_threads < 2)\n \t\treturn load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n \n+\t/* a little sanity checking */\n+\tif (istate->name_hash_initialized)\n+\t\tdie(\"the name hash isn't thread safe\");\n+\n \tmem_pool_init(&istate->ce_mem_pool, 0);\n \tif (istate->version == 4)\n \t\tprevious_name = &previous_name_buf;\n \telse\n \t\tprevious_name = NULL;\n \n-\tthread_nr = (istate->cache_nr + threads - 1) / threads;\n-\tdata = xcalloc(threads, sizeof(struct load_cache_entries_thread_data));\n+\t/* allocate an extra thread for loading the index extensions */\n+\tce_per_thread = DIV_ROUND_UP(istate->cache_nr, nr_threads);\n+\tdata = xcalloc(nr_threads + 1, sizeof(struct load_cache_entries_thread_data));\n \n-\t/* loop through index entries starting a thread for every thread_nr entries */\n+\t/*\n+\t * Loop through index entries starting a thread for every ce_per_thread\n+\t * entries.\n+\t */\n \tconsumed = thread = 0;\n-\tfor (i = 0; ; i++) {\n+\tfor (i = 0; i < istate->cache_nr; i++) {\n \t\tstruct ondisk_cache_entry *ondisk;\n \t\tconst char *name;\n \t\tunsigned int flags;\n \n-\t\t/* we've reached the begining of a block of cache entries, kick off a thread to process them */\n-\t\tif (0 == i % thread_nr) {\n+\t\t/*\n+\t\t * we've reached the beginning of a block of cache entries,\n+\t\t * kick off a thread to process them\n+\t\t */\n+\t\tif (0 == i % ce_per_thread) {\n \t\t\tstruct load_cache_entries_thread_data *p = &data[thread];\n \n \t\t\tp->istate = istate;\n \t\t\tp->offset = i;\n-\t\t\tp->nr = min(thread_nr, istate->cache_nr - i);\n+\t\t\tp->nr = ce_per_thread < istate->cache_nr - i ? ce_per_thread : istate->cache_nr - i;\n \n \t\t\t/* create a mem_pool for each thread */\n \t\t\tif (istate->version == 4)\n@@ -2034,8 +2080,8 @@ static unsigned long load_cache_entries(struct index_state *istate, void *mmap,\n \n \t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n \t\t\t\tdie(\"unable to create load_cache_entries_thread\");\n-\t\t\tif (++thread == threads || p->nr != thread_nr)\n-\t\t\t\tbreak;\n+\n+\t\t\t++thread;\n \t\t}\n \n \t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n@@ -2064,7 +2110,18 @@ static unsigned long load_cache_entries(struct index_state *istate, void *mmap,\n \t\t\tsrc_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n \t}\n \n-\tfor (i = 0; i < threads; i++) {\n+\t/* create a thread to load the index extensions */\n+\tstruct load_cache_entries_thread_data *p = &data[thread];\n+\tp->istate = istate;\n+\tmem_pool_init(&p->ce_mem_pool, 0);\n+\tp->mmap = mmap;\n+\tp->mmap_size = mmap_size;\n+\tp->start_offset = src_offset;\n+\n+\tif (pthread_create(&p->pthread, NULL, load_index_extensions_thread, p))\n+\t\tdie(\"unable to create load_index_extensions_thread\");\n+\n+\tfor (i = 0; i < nr_threads + 1; i++) {\n \t\tstruct load_cache_entries_thread_data *p = data + i;\n \t\tif (pthread_join(p->pthread, NULL))\n \t\t\tdie(\"unable to join load_cache_entries_thread\");\n\n\n### Patches\n\nBen Peart (3):\n  read-cache: speed up index load through parallelization\n  read-cache: load cache extensions on worker thread\n  read-cache: micro-optimize expand_name_field() to speed up V4 index\n    parsing.\n\n Documentation/config.txt |   6 +\n config.c                 |  14 ++\n config.h                 |   1 +\n read-cache.c             | 281 +++++++++++++++++++++++++++++++++++----\n 4 files changed, 275 insertions(+), 27 deletions(-)\n\n\nbase-commit: 29d9e3e2c47dd4b5053b0a98c891878d398463e3\n-- \n2.18.0.windows.1\n\n\n"},{"id":"356857","messageId":"xmqqk1o9cd18.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180829152500.46640-3-benpeart@microsoft.com","subject":"Re: [PATCH v2 2/3] read-cache: load cache extensions on worker thread","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-29T17:12:35Z","receivedAt":"2018-08-29T17:12:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <Ben.Peart@microsoft.com> writes:\n\n> This is possible because the current extensions don't access the cache\n> entries in the index_state structure so are OK that they don't all exist\n> yet.\n>\n> The CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\n> extensions don't even get a pointer to the index so don't have access to the\n> cache entries.\n>\n> CACHE_EXT_LINK only uses the index_state to initialize the split index.\n> CACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\n> update and dirty flags.\n\nGood to see such an analysis here.  Once we define an extension\nsection, which requires us to have the cache entries before\npopulating it, this scheme would falls down, of course, but the\nextension mechanism is all about protecting ourselves from the\nfuture changes, so we'd at least need a good feel for how we read an\nunknown extension from the future with the current code.  Perhaps\njust like the main cache entries were pre-scanned to apportion them\nto worker threads, we can pre-scan the sections and compare them\nwith a white-list built into our binary before deciding that it is\nsafe to read them in parallel (and otherwise, we ask the last thread\nfor reading extensions to wait until the workers that read the main\nindex all return)?\n\n> -/*\n> -* A thread proc to run the load_cache_entries() computation\n> -* across multiple background threads.\n> -*/\n\nThis one was mis-indented (lacking SP before '*') but they are gone\nso ... ;-)\n\n> @@ -1978,6 +1975,36 @@ static void *load_cache_entries_thread(void *_data)\n>  \treturn NULL;\n>  }\n>  \n> +static void *load_index_extensions_thread(void *_data)\n> +{\n> +\tstruct load_cache_entries_thread_data *p = _data;\n> +\tunsigned long src_offset = p->start_offset;\n> +\n> +\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n> +\t\t/* After an array of active_nr index entries,\n> +\t\t * there can be arbitrary number of extended\n> +\t\t * sections, each of which is prefixed with\n> +\t\t * extension name (4-byte) and section length\n> +\t\t * in 4-byte network byte order.\n> +\t\t */\n> +\t\tuint32_t extsize;\n> +\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n> +\t\textsize = ntohl(extsize);\n> +\t\tif (read_index_extension(p->istate,\n> +\t\t\t\t\t\t\t\t(const char *)p->mmap + src_offset,\n> +\t\t\t\t\t\t\t\t(char *)p->mmap + src_offset + 8,\n> +\t\t\t\t\t\t\t\textsize) < 0) {\n\nOverly deep indentation.  Used a wrong tab-width?\n\n> +\t/* allocate an extra thread for loading the index extensions */\n>  \tce_per_thread = DIV_ROUND_UP(istate->cache_nr, nr_threads);\n> -\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n> +\tdata = xcalloc(nr_threads + 1, sizeof(struct load_cache_entries_thread_data));\n>  \n>  \t/*\n>  \t * Loop through index entries starting a thread for every ce_per_thread\n> -\t * entries. Exit the loop when we've created the final thread (no need\n> -\t * to parse the remaining entries.\n> +\t * entries.\n>  \t */\n\nI see.  Now the pre-parsing process needs to go through all the\ncache entries to find the beginning of the extensions section.\n\n>  \tconsumed = thread = 0;\n> -\tfor (i = 0; ; i++) {\n> +\tfor (i = 0; i < istate->cache_nr; i++) {\n>  \t\tstruct ondisk_cache_entry *ondisk;\n>  \t\tconst char *name;\n>  \t\tunsigned int flags;\n> @@ -2055,9 +2082,7 @@ static unsigned long load_cache_entries(struct index_state *istate,\n>  \t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n>  \t\t\t\tdie(\"unable to create load_cache_entries_thread\");\n>  \n> -\t\t\t/* exit the loop when we've created the last thread */\n> -\t\t\tif (++thread == nr_threads)\n> -\t\t\t\tbreak;\n> +\t\t\t++thread;\n\nThis is not C++, and in (void) context, the codebase always prefers\npost-increment.\n\n> @@ -2086,7 +2111,18 @@ static unsigned long load_cache_entries(struct index_state *istate,\n>  \t\t\tsrc_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n>  \t}\n>  \n> -\tfor (i = 0; i < nr_threads; i++) {\n> +\t/* create a thread to load the index extensions */\n> +\tstruct load_cache_entries_thread_data *p = &data[thread];\n\nThis probably triggers decl-after-statement.\n\n> +\tp->istate = istate;\n> +\tmem_pool_init(&p->ce_mem_pool, 0);\n> +\tp->mmap = mmap;\n> +\tp->mmap_size = mmap_size;\n> +\tp->start_offset = src_offset;\n> +\n> +\tif (pthread_create(&p->pthread, NULL, load_index_extensions_thread, p))\n> +\t\tdie(\"unable to create load_index_extensions_thread\");\n> +\n> +\tfor (i = 0; i < nr_threads + 1; i++) {\n>  \t\tstruct load_cache_entries_thread_data *p = data + i;\n>  \t\tif (pthread_join(p->pthread, NULL))\n>  \t\t\tdie(\"unable to join load_cache_entries_thread\");\n"},{"id":"356858","messageId":"xmqqd0u1ccyd.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"CACsJy8B38QAW8qq-CctLJyJNaC329o6Rr1gs0kd=EkV+ARAaVw@mail.gmail.com","subject":"Re: [PATCH] read-cache.c: optimize reading index format v4","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-29T17:14:18Z","receivedAt":"2018-08-29T17:14:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> Yeah I kinda hated dummy_entry too but the feeling wasn't strong\n> enough to move towards the index->version check. I guess I'm going to\n> do it now.\n\nSounds like a plan.  Thanks again for a pleasant read.\n"},{"id":"356859","messageId":"xmqq5zztccy7.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180829152500.46640-2-benpeart@microsoft.com","subject":"Re: [PATCH v2 1/3] read-cache: speed up index load through parallelization","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-29T17:14:24Z","receivedAt":"2018-08-29T17:14:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <Ben.Peart@microsoft.com> writes:\n\n> diff --git a/Documentation/config.txt b/Documentation/config.txt\n> index 1c42364988..79f8296d9c 100644\n> --- a/Documentation/config.txt\n> +++ b/Documentation/config.txt\n> @@ -2391,6 +2391,12 @@ imap::\n>  \tThe configuration variables in the 'imap' section are described\n>  \tin linkgit:git-imap-send[1].\n>  \n> +index.threads::\n> +\tSpecifies the number of threads to spawn when loading the index.\n> +\tThis is meant to reduce index load time on multiprocessor machines.\n> +\tSpecifying 0 or 'true' will cause Git to auto-detect the number of\n> +\tCPU's and set the number of threads accordingly. Defaults to 'true'.\n\n\"0 or 'true' means 'auto'\" made me go \"Huh?\"\n\nThe \"Huh?\"  I initially felt comes from the fact that usually 0 and\nfalse are interchangeable, but for this particular application,\n\"disabling\" the threading means setting the count to one (not zero),\nleaving us zero as a usable \"special value\" to signal 'auto'.\n\nSo the end result does make sense, especially with this bit ...\n\n> diff --git a/config.c b/config.c\n> index 9a0b10d4bc..3bda124550 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -2289,6 +2289,20 @@ int git_config_get_fsmonitor(void)\n> ...\n> +\tif (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n> +\t\tif (is_bool)\n> +\t\t\treturn val ? 0 : 1;\n> +\t\telse\n> +\t\t\treturn val;\n\n... which says \"'0' and 'true' are the same and yields 0, '1' and\n'false' yields 1, and '2' and above will give the int\".  \n\nAdding something like\n\n\tYou can disable multi-threaded code by setting this variable\n\tto 'false' (or 1).\n\nmay reduce the risk of a similar \"Huh?\" reaction by other readers.\n\n> +struct load_cache_entries_thread_data\n> +{\n> +\tpthread_t pthread;\n> +\tstruct index_state *istate;\n> +\tstruct mem_pool *ce_mem_pool;\n> +\tint offset, nr;\n> +\tvoid *mmap;\n> +\tunsigned long start_offset;\n> +\tstruct strbuf previous_name_buf;\n> +\tstruct strbuf *previous_name;\n> +\tunsigned long consumed;\t/* return # of bytes in index file processed */\n> +};\n\nWe saw that Duy's \"let's not use strbuf to remember the previous\nname but instead use the previous ce\" approach gave us a nice\nperformance boost; I wonder if we can build on that idea here?\n\nOne possible approach might be to create one ce per \"block\" in the\npre-scanning thread and use that ce as the \"previous one\" in the\nper-thread data before spawning a worker.\n\n> +static unsigned long load_cache_entries(struct index_state *istate,\n> +\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n> +{\n> +\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n> +\tstruct load_cache_entries_thread_data *data;\n> +\tint nr_threads, cpus, ce_per_thread;\n> +\tunsigned long consumed;\n> +\tint i, thread;\n> +\n> +\tnr_threads = git_config_get_index_threads();\n> +\tif (!nr_threads) {\n> +\t\tcpus = online_cpus();\n> +\t\tnr_threads = istate->cache_nr / THREAD_COST;\n\nHere, nr_threads could become 0 with a small index, but any value\nbelow 2 makes us call load_all_cache_entries() by the main thread\n(and the value of nr_thread is not used anyore), it is fine.  Of\ncourse, forced test will set it to 2 so there is no problem, either.\n\nOK.\n\n> +\t/* a little sanity checking */\n> +\tif (istate->name_hash_initialized)\n> +\t\tdie(\"the name hash isn't thread safe\");\n\nIf it is a programming error to call into this codepath without\ninitializing the name_hash, which I think is the case, this is\nbetter done with BUG(\"\").\n\nThe remainder of the patch looked good.  Thanks.\n"},{"id":"356895","messageId":"f1a703e8-27b9-3f9e-9bde-a7b74659b4b3@gmail.com","threadId":"49204","inReplyTo":"xmqq5zztccy7.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v2 1/3] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-29T21:35:21Z","receivedAt":"2018-08-29T21:35:26Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/29/2018 1:14 PM, Junio C Hamano wrote:\n> Ben Peart <Ben.Peart@microsoft.com> writes:\n> \n>> diff --git a/Documentation/config.txt b/Documentation/config.txt\n>> index 1c42364988..79f8296d9c 100644\n>> --- a/Documentation/config.txt\n>> +++ b/Documentation/config.txt\n>> @@ -2391,6 +2391,12 @@ imap::\n>>   \tThe configuration variables in the 'imap' section are described\n>>   \tin linkgit:git-imap-send[1].\n>>   \n\n> Adding something like\n> \n> \tYou can disable multi-threaded code by setting this variable\n> \tto 'false' (or 1).\n> \n> may reduce the risk of a similar \"Huh?\" reaction by other readers.\n> \n\nWill do\n\n>> +struct load_cache_entries_thread_data\n>> +{\n>> +\tpthread_t pthread;\n>> +\tstruct index_state *istate;\n>> +\tstruct mem_pool *ce_mem_pool;\n>> +\tint offset, nr;\n>> +\tvoid *mmap;\n>> +\tunsigned long start_offset;\n>> +\tstruct strbuf previous_name_buf;\n>> +\tstruct strbuf *previous_name;\n>> +\tunsigned long consumed;\t/* return # of bytes in index file processed */\n>> +};\n> \n> We saw that Duy's \"let's not use strbuf to remember the previous\n> name but instead use the previous ce\" approach gave us a nice\n> performance boost; I wonder if we can build on that idea here?\n> \n> One possible approach might be to create one ce per \"block\" in the\n> pre-scanning thread and use that ce as the \"previous one\" in the\n> per-thread data before spawning a worker.\n> \n\nYes, I believe this can be done.  I was planning to wait until both \npatches settled down a bit before adapting it to threads.  It's a little \ntrickier because the previous ce doesn't yet exist but I believe one can \nbe fabricated enough to make the optimization work.\n\n>> +static unsigned long load_cache_entries(struct index_state *istate,\n>> +\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n>> +{\n>> +\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n>> +\tstruct load_cache_entries_thread_data *data;\n>> +\tint nr_threads, cpus, ce_per_thread;\n>> +\tunsigned long consumed;\n>> +\tint i, thread;\n>> +\n>> +\tnr_threads = git_config_get_index_threads();\n>> +\tif (!nr_threads) {\n>> +\t\tcpus = online_cpus();\n>> +\t\tnr_threads = istate->cache_nr / THREAD_COST;\n> \n> Here, nr_threads could become 0 with a small index, but any value\n> below 2 makes us call load_all_cache_entries() by the main thread\n> (and the value of nr_thread is not used anyore), it is fine.  Of\n> course, forced test will set it to 2 so there is no problem, either.\n> \n> OK.\n> \n>> +\t/* a little sanity checking */\n>> +\tif (istate->name_hash_initialized)\n>> +\t\tdie(\"the name hash isn't thread safe\");\n> \n> If it is a programming error to call into this codepath without\n> initializing the name_hash, which I think is the case, this is\n> better done with BUG(\"\").\n> \n\nWill do\n\n> The remainder of the patch looked good.  Thanks.\n> \n"},{"id":"356897","messageId":"606bb7e9-0d58-ec86-6a3c-8a123679f9f4@gmail.com","threadId":"49204","inReplyTo":"xmqqk1o9cd18.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v2 2/3] read-cache: load cache extensions on worker thread","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-08-29T21:42:56Z","receivedAt":"2018-08-29T21:43:01Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 8/29/2018 1:12 PM, Junio C Hamano wrote:\n> Ben Peart <Ben.Peart@microsoft.com> writes:\n> \n>> This is possible because the current extensions don't access the cache\n>> entries in the index_state structure so are OK that they don't all exist\n>> yet.\n>>\n>> The CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\n>> extensions don't even get a pointer to the index so don't have access to the\n>> cache entries.\n>>\n>> CACHE_EXT_LINK only uses the index_state to initialize the split index.\n>> CACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\n>> update and dirty flags.\n> \n> Good to see such an analysis here.  Once we define an extension\n> section, which requires us to have the cache entries before\n> populating it, this scheme would falls down, of course, but the\n> extension mechanism is all about protecting ourselves from the\n> future changes, so we'd at least need a good feel for how we read an\n> unknown extension from the future with the current code.  Perhaps\n> just like the main cache entries were pre-scanned to apportion them\n> to worker threads, we can pre-scan the sections and compare them\n> with a white-list built into our binary before deciding that it is\n> safe to read them in parallel (and otherwise, we ask the last thread\n> for reading extensions to wait until the workers that read the main\n> index all return)?\n> \n\nYes, when we add a new extension that requires the cache entries to \nexist and be parsed, we will need to add a mechanism to ensure that \nhappens for that extension.  I agree a white list is probably the right \nway to deal with it.  Until we have that need, it would just add \nunnecessary complexity so I think we should wait till it is actually needed.\n\nThere isn't any change in behavior with unknown extensions and this \npatch.  If an unknown extension exists it will just get ignored and \nreported as an \"unknown extension\" or \"die\" if it is marked as \"required.\"\n\nI'll fix the rest of your suggestions - thanks for the close review.\n"},{"id":"356910","messageId":"xmqqd0u094ve.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"606bb7e9-0d58-ec86-6a3c-8a123679f9f4@gmail.com","subject":"Re: [PATCH v2 2/3] read-cache: load cache extensions on worker thread","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-29T22:19:18Z","receivedAt":"2018-08-29T22:37:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <peartben@gmail.com> writes:\n\n> There isn't any change in behavior with unknown extensions and this\n> patch.  If an unknown extension exists it will just get ignored and\n> reported as an \"unknown extension\" or \"die\" if it is marked as\n> \"required.\"\n\nOK.\n"},{"id":"357189","messageId":"20180902131933.27484-1-pclouds@gmail.com","threadId":"49204","inReplyTo":"20180825064458.28484-1-pclouds@gmail.com","subject":"[PATCH v2 0/1] optimize reading index format v4","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-02T13:19:32Z","receivedAt":"2018-09-02T13:19:40Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"v2 removes unrelated changes and the dummy_entry. strip_len is also\nreplaced with copy_len to reduce repeated subtraction calculation.\nDiff: \n\ndiff --git a/read-cache.c b/read-cache.c\nindex 5c04c8f200..8628d0f3a8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1713,7 +1713,7 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n+static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n \t\t\t\t\t    const struct cache_entry *previous_ce)\n@@ -1722,7 +1722,15 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \tsize_t len;\n \tconst char *name;\n \tunsigned int flags;\n-\tsize_t strip_len;\n+\tsize_t copy_len;\n+\t/*\n+\t * Adjacent cache entries tend to share the leading paths, so it makes\n+\t * sense to only store the differences in later entries.  In the v4\n+\t * on-disk format of the index, each on-disk cache entry stores the\n+\t * number of bytes to be stripped from the end of the previous name,\n+\t * and the bytes to append to the result, to come up with its name.\n+\t */\n+\tint expand_name_field = istate->version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1735,37 +1743,37 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \t\textended_flags = get_be16(&ondisk2->flags2) << 16;\n \t\t/* We do not yet understand any bit out of CE_EXTENDED_FLAGS */\n \t\tif (extended_flags & ~CE_EXTENDED_FLAGS)\n-\t\t\tdie(_(\"unknown index entry format %08x\"), extended_flags);\n+\t\t\tdie(\"Unknown index entry format %08x\", extended_flags);\n \t\tflags |= extended_flags;\n \t\tname = ondisk2->name;\n \t}\n \telse\n \t\tname = ondisk->name;\n \n-\t/*\n-\t * Adjacent cache entries tend to share the leading paths, so it makes\n-\t * sense to only store the differences in later entries.  In the v4\n-\t * on-disk format of the index, each on-disk cache entry stores the\n-\t * number of bytes to be stripped from the end of the previous name,\n-\t * and the bytes to append to the result, to come up with its name.\n-\t */\n-\tif (previous_ce) {\n+\tif (expand_name_field) {\n \t\tconst unsigned char *cp = (const unsigned char *)name;\n+\t\tsize_t strip_len, previous_len;\n \n+\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n \t\tstrip_len = decode_varint(&cp);\n-\t\tif (previous_ce->ce_namelen < strip_len)\n-\t\t\tdie(_(\"malformed name field in the index, path '%s'\"),\n-\t\t\t    previous_ce->name);\n+\t\tif (previous_len < strip_len) {\n+\t\t\tif (previous_ce)\n+\t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n+\t\t\t\t    previous_ce->name);\n+\t\t\telse\n+\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t}\n+\t\tcopy_len = previous_len - strip_len;\n \t\tname = (const char *)cp;\n \t}\n \n \tif (len == CE_NAMEMASK) {\n \t\tlen = strlen(name);\n-\t\tif (previous_ce)\n-\t\t\tlen += previous_ce->ce_namelen - strip_len;\n+\t\tif (expand_name_field)\n+\t\t\tlen += copy_len;\n \t}\n \n-\tce = mem_pool__ce_alloc(mem_pool, len);\n+\tce = mem_pool__ce_alloc(istate->ce_mem_pool, len);\n \n \tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n \tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n@@ -1782,9 +1790,9 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \tce->index = 0;\n \thashcpy(ce->oid.hash, ondisk->sha1);\n \n-\tif (previous_ce) {\n-\t\tsize_t copy_len = previous_ce->ce_namelen - strip_len;\n-\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n+\tif (expand_name_field) {\n+\t\tif (copy_len)\n+\t\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n \t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n \t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n \t} else {\n@@ -1885,7 +1893,6 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tvoid *mmap;\n \tsize_t mmap_size;\n \tconst struct cache_entry *previous_ce = NULL;\n-\tstruct cache_entry *dummy_entry = NULL;\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1923,7 +1930,6 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->initialized = 1;\n \n \tif (istate->version == 4) {\n-\t\tprevious_ce = dummy_entry = make_empty_transient_cache_entry(0);\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n \t} else {\n@@ -1938,14 +1944,12 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tunsigned long consumed;\n \n \t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_ce);\n+\t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n \t\tsrc_offset += consumed;\n-\t\tif (previous_ce)\n-\t\t\tprevious_ce = ce;\n+\t\tprevious_ce = ce;\n \t}\n-\tfree(dummy_entry);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n\n\nNguyễn Thái Ngọc Duy (1):\n  read-cache.c: optimize reading index format v4\n\n read-cache.c | 128 ++++++++++++++++++++++++---------------------------\n 1 file changed, 60 insertions(+), 68 deletions(-)\n\n-- \n2.19.0.rc0.337.ge906d732e7\n\n"},{"id":"357190","messageId":"20180902131933.27484-2-pclouds@gmail.com","threadId":"49204","inReplyTo":"20180902131933.27484-1-pclouds@gmail.com","subject":"[PATCH v2 1/1] read-cache.c: optimize reading index format v4","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-02T13:19:33Z","receivedAt":"2018-09-02T13:19:43Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Index format v4 requires some more computation to assemble a path\nbased on a previous one. The current code is not very efficient\nbecause\n\n - it doubles memory copy, we assemble the final path in a temporary\n   first before putting it back to a cache_entry\n\n - strbuf_remove() in expand_name_field() is not exactly a good fit\n   for stripping a part at the end, _setlen() would do the same job\n   and is much cheaper.\n\n - the open-coded loop to find the end of the string in\n   expand_name_field() can't beat an optimized strlen()\n\nThis patch avoids the temporary buffer and writes directly to the new\ncache_entry, which addresses the first two points. The last point\ncould also be avoided if the total string length fits in the first 12\nbits of ce_flags, if not we fall back to strlen().\n\nRunning \"test-tool read-cache 100\" on webkit.git (275k files), reading\nv2 only takes 4.226 seconds, while v4 takes 5.711 seconds, 35% more\ntime. The patch reduces read time on v4 to 4.319 seconds.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c | 128 ++++++++++++++++++++++++---------------------------\n 1 file changed, 60 insertions(+), 68 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 7b1354d759..8628d0f3a8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1713,63 +1713,24 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *cache_entry_from_ondisk(struct mem_pool *mem_pool,\n-\t\t\t\t\t\t   struct ondisk_cache_entry *ondisk,\n-\t\t\t\t\t\t   unsigned int flags,\n-\t\t\t\t\t\t   const char *name,\n-\t\t\t\t\t\t   size_t len)\n-{\n-\tstruct cache_entry *ce = mem_pool__ce_alloc(mem_pool, len);\n-\n-\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n-\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n-\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n-\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n-\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n-\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n-\tce->ce_mode  = get_be32(&ondisk->mode);\n-\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n-\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n-\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n-\tce->ce_flags = flags & ~CE_NAMEMASK;\n-\tce->ce_namelen = len;\n-\tce->index = 0;\n-\thashcpy(ce->oid.hash, ondisk->sha1);\n-\tmemcpy(ce->name, name, len);\n-\tce->name[len] = '\\0';\n-\treturn ce;\n-}\n-\n-/*\n- * Adjacent cache entries tend to share the leading paths, so it makes\n- * sense to only store the differences in later entries.  In the v4\n- * on-disk format of the index, each on-disk cache entry stores the\n- * number of bytes to be stripped from the end of the previous name,\n- * and the bytes to append to the result, to come up with its name.\n- */\n-static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n-{\n-\tconst unsigned char *ep, *cp = (const unsigned char *)cp_;\n-\tsize_t len = decode_varint(&cp);\n-\n-\tif (name->len < len)\n-\t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n-\tstrbuf_add(name, cp, ep - cp);\n-\treturn (const char *)ep + 1 - cp_;\n-}\n-\n-static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n+static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n-\t\t\t\t\t    struct strbuf *previous_name)\n+\t\t\t\t\t    const struct cache_entry *previous_ce)\n {\n \tstruct cache_entry *ce;\n \tsize_t len;\n \tconst char *name;\n \tunsigned int flags;\n+\tsize_t copy_len;\n+\t/*\n+\t * Adjacent cache entries tend to share the leading paths, so it makes\n+\t * sense to only store the differences in later entries.  In the v4\n+\t * on-disk format of the index, each on-disk cache entry stores the\n+\t * number of bytes to be stripped from the end of the previous name,\n+\t * and the bytes to append to the result, to come up with its name.\n+\t */\n+\tint expand_name_field = istate->version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1789,21 +1750,54 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \telse\n \t\tname = ondisk->name;\n \n-\tif (!previous_name) {\n-\t\t/* v3 and earlier */\n-\t\tif (len == CE_NAMEMASK)\n-\t\t\tlen = strlen(name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags, name, len);\n+\tif (expand_name_field) {\n+\t\tconst unsigned char *cp = (const unsigned char *)name;\n+\t\tsize_t strip_len, previous_len;\n \n-\t\t*ent_size = ondisk_ce_size(ce);\n-\t} else {\n-\t\tunsigned long consumed;\n-\t\tconsumed = expand_name_field(previous_name, name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags,\n-\t\t\t\t\t     previous_name->buf,\n-\t\t\t\t\t     previous_name->len);\n+\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n+\t\tstrip_len = decode_varint(&cp);\n+\t\tif (previous_len < strip_len) {\n+\t\t\tif (previous_ce)\n+\t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n+\t\t\t\t    previous_ce->name);\n+\t\t\telse\n+\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t}\n+\t\tcopy_len = previous_len - strip_len;\n+\t\tname = (const char *)cp;\n+\t}\n+\n+\tif (len == CE_NAMEMASK) {\n+\t\tlen = strlen(name);\n+\t\tif (expand_name_field)\n+\t\t\tlen += copy_len;\n+\t}\n+\n+\tce = mem_pool__ce_alloc(istate->ce_mem_pool, len);\n+\n+\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n+\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n+\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n+\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n+\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n+\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n+\tce->ce_mode  = get_be32(&ondisk->mode);\n+\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n+\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n+\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n+\tce->ce_flags = flags & ~CE_NAMEMASK;\n+\tce->ce_namelen = len;\n+\tce->index = 0;\n+\thashcpy(ce->oid.hash, ondisk->sha1);\n \n-\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n+\tif (expand_name_field) {\n+\t\tif (copy_len)\n+\t\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n+\t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n+\t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n+\t} else {\n+\t\tmemcpy(ce->name, name, len + 1);\n+\t\t*ent_size = ondisk_ce_size(ce);\n \t}\n \treturn ce;\n }\n@@ -1898,7 +1892,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tconst struct cache_entry *previous_ce = NULL;\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1936,11 +1930,9 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->initialized = 1;\n \n \tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n \t} else {\n-\t\tprevious_name = NULL;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n \t}\n@@ -1952,12 +1944,12 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tunsigned long consumed;\n \n \t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n \t\tsrc_offset += consumed;\n+\t\tprevious_ce = ce;\n \t}\n-\tstrbuf_release(&previous_name_buf);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-- \n2.19.0.rc0.337.ge906d732e7\n\n"},{"id":"357253","messageId":"CACsJy8BK5a1DgY4o5Xxrdjyg1QYxW-kCPNmAXkdudCD78O7dTg@mail.gmail.com","threadId":"49204","inReplyTo":"20180829152500.46640-2-benpeart@microsoft.com","subject":"Re: [PATCH v2 1/3] read-cache: speed up index load through parallelization","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-03T19:16:35Z","receivedAt":"2018-09-03T19:17:03Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Aug 29, 2018 at 5:25 PM Ben Peart <Ben.Peart@microsoft.com> wrote:\n> diff --git a/read-cache.c b/read-cache.c\n> index 7b1354d759..c30346388a 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1889,16 +1889,229 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n>         return ondisk_size + entries * per_entry;\n>  }\n>\n> +/*\n> + * A helper function that will load the specified range of cache entries\n> + * from the memory mapped file and add them to the given index.\n> + */\n> +static unsigned long load_cache_entry_block(struct index_state *istate,\n> +                       struct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n> +                       unsigned long start_offset, struct strbuf *previous_name)\n> +{\n> +       int i;\n> +       unsigned long src_offset = start_offset;\n> +\n> +       for (i = offset; i < offset + nr; i++) {\n\nIt may be micro optimization, but since we're looping a lot and can't\ntrust the compiler to optimize this, maybe just calculate this upper\nlimit and store in a local variable to make it clear the upper limit\nis known, no point of recalculating it at every iteration.\n\n> +               struct ondisk_cache_entry *disk_ce;\n> +               struct cache_entry *ce;\n> +               unsigned long consumed;\n> +\n> +               disk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n> +               ce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n> +               set_index_entry(istate, i, ce);\n> +\n> +               src_offset += consumed;\n> +       }\n> +       return src_offset - start_offset;\n> +}\n> +\n> +static unsigned long load_all_cache_entries(struct index_state *istate,\n> +                       void *mmap, size_t mmap_size, unsigned long src_offset)\n> +{\n> +       struct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n> +       unsigned long consumed;\n> +\n> +       if (istate->version == 4) {\n> +               previous_name = &previous_name_buf;\n> +               mem_pool_init(&istate->ce_mem_pool,\n> +                             estimate_cache_size_from_compressed(istate->cache_nr));\n> +       } else {\n> +               previous_name = NULL;\n> +               mem_pool_init(&istate->ce_mem_pool,\n> +                             estimate_cache_size(mmap_size, istate->cache_nr));\n> +       }\n> +\n> +       consumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n> +                                       0, istate->cache_nr, mmap, src_offset, previous_name);\n> +       strbuf_release(&previous_name_buf);\n> +       return consumed;\n> +}\n> +\n> +#ifdef NO_PTHREADS\n> +\n> +#define load_cache_entries load_all_cache_entries\n> +\n> +#else\n> +\n> +#include \"thread-utils.h\"\n\nDon't include files in a middle of a file.\n\n> +\n> +/*\n> +* Mostly randomly chosen maximum thread counts: we\n> +* cap the parallelism to online_cpus() threads, and we want\n> +* to have at least 7500 cache entries per thread for it to\n> +* be worth starting a thread.\n> +*/\n> +#define THREAD_COST            (7500)\n\nIsn't 7500 a bit too low? I'm still basing on webkit.git,  and 7500\nentries take about 1.2ms on average. 100k files would take about 16ms\nand may be more reasonable (still too low in my opinion).\n\n> +\n> +struct load_cache_entries_thread_data\n> +{\n> +       pthread_t pthread;\n> +       struct index_state *istate;\n> +       struct mem_pool *ce_mem_pool;\n> +       int offset, nr;\n> +       void *mmap;\n> +       unsigned long start_offset;\n> +       struct strbuf previous_name_buf;\n> +       struct strbuf *previous_name;\n> +       unsigned long consumed; /* return # of bytes in index file processed */\n> +};\n> +\n> +/*\n> +* A thread proc to run the load_cache_entries() computation\n> +* across multiple background threads.\n> +*/\n> +static void *load_cache_entries_thread(void *_data)\n> +{\n> +       struct load_cache_entries_thread_data *p = _data;\n> +\n> +       p->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n> +               p->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n> +       return NULL;\n> +}\n> +\n> +static unsigned long load_cache_entries(struct index_state *istate,\n> +                       void *mmap, size_t mmap_size, unsigned long src_offset)\n> +{\n> +       struct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n> +       struct load_cache_entries_thread_data *data;\n> +       int nr_threads, cpus, ce_per_thread;\n> +       unsigned long consumed;\n> +       int i, thread;\n> +\n> +       nr_threads = git_config_get_index_threads();\n> +       if (!nr_threads) {\n> +               cpus = online_cpus();\n> +               nr_threads = istate->cache_nr / THREAD_COST;\n> +               if (nr_threads > cpus)\n> +                       nr_threads = cpus;\n> +       }\n> +\n> +       /* enable testing with fewer than default minimum of entries */\n> +       if ((istate->cache_nr > 1) && (nr_threads < 2) && git_env_bool(\"GIT_INDEX_THREADS_TEST\", 0))\n> +               nr_threads = 2;\n\nPlease don't add more '()' than necessary. It's just harder to read.\nMaybe break that \"if\" into two lines since it's getting long.\n\n> +\n> +       if (nr_threads < 2)\n> +               return load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n> +\n> +       /* a little sanity checking */\n> +       if (istate->name_hash_initialized)\n> +               die(\"the name hash isn't thread safe\");\n> +\n> +       mem_pool_init(&istate->ce_mem_pool, 0);\n> +       if (istate->version == 4)\n> +               previous_name = &previous_name_buf;\n> +       else\n> +               previous_name = NULL;\n> +\n> +       ce_per_thread = DIV_ROUND_UP(istate->cache_nr, nr_threads);\n> +       data = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n> +\n> +       /*\n> +        * Loop through index entries starting a thread for every ce_per_thread\n> +        * entries. Exit the loop when we've created the final thread (no need\n> +        * to parse the remaining entries.\n> +        */\n> +       consumed = thread = 0;\n> +       for (i = 0; ; i++) {\n> +               struct ondisk_cache_entry *ondisk;\n> +               const char *name;\n> +               unsigned int flags;\n> +\n> +               /*\n> +                * we've reached the beginning of a block of cache entries,\n> +                * kick off a thread to process them\n> +                */\n> +               if (0 == i % ce_per_thread) {\n\nI don't get why people keep putting constants in reversed order like\nthis. Perhaps in the old days, it helps catch \"a = 0\" mistakes, but\ncompilers nowadays are smart enough to complain about that and this is\njust hard to read.\n\n> +                       struct load_cache_entries_thread_data *p = &data[thread];\n> +\n> +                       p->istate = istate;\n> +                       p->offset = i;\n> +                       p->nr = ce_per_thread < istate->cache_nr - i ? ce_per_thread : istate->cache_nr - i;\n> +\n> +                       /* create a mem_pool for each thread */\n> +                       if (istate->version == 4)\n> +                               mem_pool_init(&p->ce_mem_pool,\n> +                                                 estimate_cache_size_from_compressed(p->nr));\n> +                       else\n> +                               mem_pool_init(&p->ce_mem_pool,\n> +                                                 estimate_cache_size(mmap_size, p->nr));\n> +\n> +                       p->mmap = mmap;\n> +                       p->start_offset = src_offset;\n> +                       if (previous_name) {\n> +                               strbuf_addbuf(&p->previous_name_buf, previous_name);\n> +                               p->previous_name = &p->previous_name_buf;\n> +                       }\n> +\n> +                       if (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n> +                               die(\"unable to create load_cache_entries_thread\");\n> +\n> +                       /* exit the loop when we've created the last thread */\n> +                       if (++thread == nr_threads)\n> +                               break;\n\nI still think it's better to have an extension to avoid looping\nthrough like this. How much time does this \"for (i = 0; ; i++)\" loop\ncost? The first thread can't start until you've scanned to the second\nblock, when you have zillion of entries and about 4 cores, that could\nbe significant delay. Unless you break smaller blocks and have one\nthread handles multiple blocks, but then you pay the cost for\nsynchronization. Other threads may overlap a bit, but starting all\nthreads at the same time would benefit more. You also can't start\nloading the extensions until you've scanned through all this.\n\n> +               }\n> +\n> +               ondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n> +\n> +               /* On-disk flags are just 16 bits */\n> +               flags = get_be16(&ondisk->flags);\n> +\n> +               if (flags & CE_EXTENDED) {\n> +                       struct ondisk_cache_entry_extended *ondisk2;\n> +                       ondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n> +                       name = ondisk2->name;\n> +               } else\n> +                       name = ondisk->name;\n> +\n> +               if (!previous_name) {\n> +                       size_t len;\n> +\n> +                       /* v3 and earlier */\n> +                       len = flags & CE_NAMEMASK;\n> +                       if (len == CE_NAMEMASK)\n> +                               len = strlen(name);\n> +                       src_offset += (flags & CE_EXTENDED) ?\n> +                               ondisk_cache_entry_extended_size(len) :\n> +                               ondisk_cache_entry_size(len);\n> +               } else\n> +                       src_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n> +       }\n> +\n> +       for (i = 0; i < nr_threads; i++) {\n> +               struct load_cache_entries_thread_data *p = data + i;\n> +               if (pthread_join(p->pthread, NULL))\n> +                       die(\"unable to join load_cache_entries_thread\");\n\n_()\n\n> +               mem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n> +               strbuf_release(&p->previous_name_buf);\n> +               consumed += p->consumed;\n> +       }\n> +\n> +       free(data);\n> +       strbuf_release(&previous_name_buf);\n> +\n> +       return consumed;\n> +}\n> +\n> +#endif\n> +\n-- \nDuy\n"},{"id":"357254","messageId":"CACsJy8AC8VT=DEcmqAtW26pYKRVT1Kz=pVyj-Wnu3uOsKwWGTw@mail.gmail.com","threadId":"49204","inReplyTo":"20180829152500.46640-3-benpeart@microsoft.com","subject":"Re: [PATCH v2 2/3] read-cache: load cache extensions on worker thread","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-03T19:21:55Z","receivedAt":"2018-09-03T19:22:23Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Aug 29, 2018 at 5:25 PM Ben Peart <Ben.Peart@microsoft.com> wrote:\n>\n> This patch helps address the CPU cost of loading the index by loading\n> the cache extensions on a worker thread in parallel with loading the cache\n> entries.\n>\n> This is possible because the current extensions don't access the cache\n> entries in the index_state structure so are OK that they don't all exist\n> yet.\n>\n> The CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\n> extensions don't even get a pointer to the index so don't have access to the\n> cache entries.\n>\n> CACHE_EXT_LINK only uses the index_state to initialize the split index.\n> CACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\n> update and dirty flags.\n>\n> I used p0002-read-cache.sh to generate some performance data on the\n> cumulative impact:\n>\n> 100,000 entries\n>\n> Test                                HEAD~3           HEAD~2\n> ---------------------------------------------------------------------------\n> read_cache/discard_cache 1000 times 14.08(0.01+0.10) 9.72(0.03+0.06) -31.0%\n\nThis is misleading (if I read it correctly). 1/3 already drops\nexecution time down to 9.81, so this patch alone only has about 6%\nsaving. Have you measured how much time is spent on loading extensions\nin single threaded mode? I'm just curious if we could hide that\ncompletely (provided that we have enough cores) while we load the\nindex.\n-- \nDuy\n"},{"id":"357255","messageId":"CACsJy8Dm-jRYfp1UVu8O+NiHf61RVY5t1vvsbXan0YrhsDabbA@mail.gmail.com","threadId":"49204","inReplyTo":"CACsJy8AC8VT=DEcmqAtW26pYKRVT1Kz=pVyj-Wnu3uOsKwWGTw@mail.gmail.com","subject":"Re: [PATCH v2 2/3] read-cache: load cache extensions on worker thread","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-03T19:27:57Z","receivedAt":"2018-09-03T19:28:25Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Sep 3, 2018 at 9:21 PM Duy Nguyen <pclouds@gmail.com> wrote:\n> > I used p0002-read-cache.sh to generate some performance data on the\n> > cumulative impact:\n> >\n> > 100,000 entries\n> >\n> > Test                                HEAD~3           HEAD~2\n> > ---------------------------------------------------------------------------\n> > read_cache/discard_cache 1000 times 14.08(0.01+0.10) 9.72(0.03+0.06) -31.0%\n>\n> This is misleading (if I read it correctly). 1/3 already drops\n> execution time down to 9.81, so this patch alone only has about 6%\n> saving.\n\nI may have miscalculated that. 1/3 says -30% saving, here it's -31%,\nso I guess it's 1% extra saving (or ~3% on 1m entries)? That's\ndefinitely not worth doing.\n-- \nDuy\n"},{"id":"357300","messageId":"20180904160840.GA24679@duynguyen.home","threadId":"49204","inReplyTo":"xmqqwosbiouc.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH] read-cache.c: optimize reading index format v4","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-04T16:08:40Z","receivedAt":"2018-09-04T16:08:46Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Aug 27, 2018 at 12:36:27PM -0700, Junio C Hamano wrote:\n> > PS. I notice that v4 does not pad to align entries at 4 byte boundary\n> > like v2/v3. This could cause a slight slow down on x86 and segfault on\n> > some other platforms.\n> \n> Care to elaborate?  \n> \n> Long time ago, we used to mmap and read directly from the index file\n> contents, requiring either an unaligned read or padded entries.  But\n> that was eons ago and we first read and convert from on-disk using\n> get_be32() etc. to in-core structure, so I am not sure what you mean\n> by \"segfault\" here.\n\nTo conclude this unalignment thing (since I plan more changes in the\nindex to keep its size down, which may increase unaligned access), I\nran with the following patch on amd64 (still webkit.git, 275k files,\n100 runs), the index version that does not make unaligned access does\nnot give noticeable differences. Still roughly around 4.2s.\n\nRunning with NO_UNALIGNED_LOADS defined is clearly slower, in 4.3s\nrange. So in theory if we avoid unaligned access in the index and\navoid slow get_beXX versions, we could bring performance back to 4.2s\nrange for those platforms.\n\nBut on the other hand, padding the index increases the index size by\n~1MB (v4 version before padding is 21MB) and this may add more cost at\nupdate time because of the trailer hash.\n\nSo, yeah it's probably ok to keep living with unaligned access and not\npad more. At least until those on \"no unaligned access\" platforms yell\nup.\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 8628d0f3a8..33ee35fb81 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1794,7 +1794,7 @@ static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\tif (copy_len)\n \t\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n \t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n-\t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n+\t\t*ent_size = ((name - ((char *)ondisk)) + len - copy_len + 8) & ~7;\n \t} else {\n \t\tmemcpy(ce->name, name, len + 1);\n \t\t*ent_size = ondisk_ce_size(ce);\n@@ -2345,8 +2345,10 @@ static int ce_write_entry(git_hash_ctx *c, int fd, struct cache_entry *ce,\n \t\t\tresult = ce_write(c, fd, to_remove_vi, prefix_size);\n \t\tif (!result)\n \t\t\tresult = ce_write(c, fd, ce->name + common, ce_namelen(ce) - common);\n-\t\tif (!result)\n-\t\t\tresult = ce_write(c, fd, padding, 1);\n+\t\tif (!result) {\n+\t\t\tint len = prefix_size + ce_namelen(ce) - common;\n+\t\t\tresult = ce_write(c, fd, padding, align_padding_size(size, len));\n+\t\t}\n \n \t\tstrbuf_splice(previous_name, common, to_remove,\n \t\t\t      ce->name + common, ce_namelen(ce) - common);\n\n--\nDuy\n"},{"id":"357325","messageId":"xmqqk1o19jj8.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180902131933.27484-2-pclouds@gmail.com","subject":"Re: [PATCH v2 1/1] read-cache.c: optimize reading index format v4","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-04T18:58:35Z","receivedAt":"2018-09-04T18:58:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> +static struct cache_entry *create_from_disk(struct index_state *istate,\n>  \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n>  \t\t\t\t\t    unsigned long *ent_size,\n> -\t\t\t\t\t    struct strbuf *previous_name)\n> +\t\t\t\t\t    const struct cache_entry *previous_ce)\n>  {\n>  \tstruct cache_entry *ce;\n>  \tsize_t len;\n>  \tconst char *name;\n>  \tunsigned int flags;\n> +\tsize_t copy_len;\n\nWe should not have to, but let's initialize it to 0 here, because\n...\n\n> +\tif (expand_name_field) {\n> +...\n> +\t\tcopy_len = previous_len - strip_len;\n> +\t\tname = (const char *)cp;\n> +\t}\n> +\n> +\tif (len == CE_NAMEMASK) {\n> +\t\tlen = strlen(name);\n> +\t\tif (expand_name_field)\n> +\t\t\tlen += copy_len;\n> ...\n> +\t}\n> +\tif (expand_name_field) {\n> +\t\tif (copy_len)\n> +\t\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n> +\t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n> +\t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n\nI am seeing a compiler getting confused, thinking that copy_len\ncould be used before getting assigned.\n\nHumans can see that reference to copy_len are made only inside \"if\n(expand_name_field)\", so we shouldn't have to.\n\n"},{"id":"357329","messageId":"xmqq5zzl9hzp.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180902131933.27484-2-pclouds@gmail.com","subject":"Re: [PATCH v2 1/1] read-cache.c: optimize reading index format v4","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-04T19:31:54Z","receivedAt":"2018-09-04T19:31:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> Index format v4 requires some more computation to assemble a path\n> based on a previous one. The current code is not very efficient\n> because\n>\n>  - it doubles memory copy, we assemble the final path in a temporary\n>    first before putting it back to a cache_entry\n>\n>  - strbuf_remove() in expand_name_field() is not exactly a good fit\n>    for stripping a part at the end, _setlen() would do the same job\n>    and is much cheaper.\n>\n>  - the open-coded loop to find the end of the string in\n>    expand_name_field() can't beat an optimized strlen()\n>\n> This patch avoids the temporary buffer and writes directly to the new\n> cache_entry, which addresses the first two points. The last point\n> could also be avoided if the total string length fits in the first 12\n> bits of ce_flags, if not we fall back to strlen().\n>\n> Running \"test-tool read-cache 100\" on webkit.git (275k files), reading\n> v2 only takes 4.226 seconds, while v4 takes 5.711 seconds, 35% more\n> time. The patch reduces read time on v4 to 4.319 seconds.\n>\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  read-cache.c | 128 ++++++++++++++++++++++++---------------------------\n>  1 file changed, 60 insertions(+), 68 deletions(-)\n\nThanks; this round is much easier to read with a clearly named\n\"expand_name_field\" boolean variable, etc.\n\n"},{"id":"357578","messageId":"20180906210227.54368-1-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180823154053.20212-1-benpeart@microsoft.com","subject":"[PATCH v3 0/4] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-06T21:03:53Z","receivedAt":"2018-09-06T21:03:59Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"On further investigation with the previous patch, I noticed that my test\nrepos didn't contain the cache tree extension in their index. After doing a\ncommit to ensure they existed, I realized that in some instances, the time\nto load the cache tree exceeded the time to load all the cache entries in\nparallel.  Because the thread to read the cache tree was started last (due\nto having to parse through all the cache entries first) we weren't always\ngetting optimal performance.\n\nTo better optimize for this case, I decided to write the EOIE extension\nas suggested by Junio [1] in response to my earlier multithreading patch\nseries [2].  This enables me to spin up the thread to load the extensions\nearlier as it no longer has to parse through all the cache entries first.\n\nThe big changes in this iteration are:\n\n- add the EOIE extension\n- update the index extension worker thread to start first\n\nThe absolute perf numbers don't look as good as the previous iteration\nbecause not loading the cache tree at all is a lot faster than loading it in\nparallel. These were measured with a V4 index that included a cache tree\nextension.\n\nI used p0002-read-cache.sh to generate some performance data on how the three\nperformance patches help:\n\np0002-read-cache.sh w/100,000 files                        \nBaseline         expand_name_field()    Thread extensions       Thread entries\n---------------------------------------------------------------------------------------\n22.34(0.01+0.12) 21.14(0.03+0.01) -5.4% 20.71(0.03+0.03) -7.3%\t13.93(0.04+0.04) -37.6%\n\np0002-read-cache.sh w/1,000,000 files                        \nBaseline          expand_name_field()     Thread extensions        Thread entries\n-------------------------------------------------------------------------------------------\n306.44(0.04+0.07) 295.42(0.01+0.07) -3.6% 217.60(0.03+0.04) -29.0% 199.00(0.00+0.10) -35.1%\n\nThis patch conflicts with Duy's patch to remove the double memory copy and\npass in the previous ce instead.  The two will need to be merged/reconciled\nonce they settle down a bit.\n\n[1] https://public-inbox.org/git/xmqq1sl017dw.fsf@gitster.mtv.corp.google.com/\n[2] https://public-inbox.org/git/20171109141737.47976-1-benpeart@microsoft.com/\n\n\nBase Ref: master\nWeb-Diff: https://github.com/benpeart/git/commit/325ec69299\nCheckout: git fetch https://github.com/benpeart/git read-index-multithread-v3 && git checkout 325ec69299\n\n\n### Patches\n\nBen Peart (4):\n  read-cache: optimize expand_name_field() to speed up V4 index parsing.\n  eoie: add End of Index Entry (EOIE) extension\n  read-cache: load cache extensions on a worker thread\n  read-cache: speed up index load through parallelization\n\n Documentation/config.txt                 |   6 +\n Documentation/technical/index-format.txt |  23 ++\n config.c                                 |  18 +\n config.h                                 |   1 +\n read-cache.c                             | 476 ++++++++++++++++++++---\n t/README                                 |  11 +\n t/t1700-split-index.sh                   |   1 +\n 7 files changed, 487 insertions(+), 49 deletions(-)\n\n\nbase-commit: 29d9e3e2c47dd4b5053b0a98c891878d398463e3\n-- \n2.18.0.windows.1\n\n\n"},{"id":"357579","messageId":"20180906210227.54368-2-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180906210227.54368-1-benpeart@microsoft.com","subject":"[PATCH v3 1/4] read-cache: optimize expand_name_field() to speed up V4 index parsing.","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-06T21:03:54Z","receivedAt":"2018-09-06T21:04:01Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"Optimize expand_name_field() to speed up V4 index parsing.\n\n - strbuf_remove() in expand_name_field() is not exactly a good fit\n   for stripping a part at the end, _setlen() would do the same job\n   and is much cheaper.\n\n - the open-coded loop to find the end of the string in\n   expand_name_field() can't beat an optimized strlen()\n\nI used p0002-read-cache.sh to generate some performance data:\n\np0002-read-cache.sh w/100,000 files\nBaseline         expand_name_field()\n---------------------------------------\n22.34(0.01+0.12) 21.14(0.03+0.01) -5.4%\n\np0002-read-cache.sh w/1,000,000 files\nBaseline          expand_name_field()\n-----------------------------------------\n306.44(0.04+0.07) 295.42(0.01+0.07) -3.6%\n\nSuggested by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 5 ++---\n 1 file changed, 2 insertions(+), 3 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 7b1354d759..382cc16bdc 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1754,9 +1754,8 @@ static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n \n \tif (name->len < len)\n \t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n+\tstrbuf_setlen(name, name->len - len);\n+\tep = cp + strlen((const char *)cp);\n \tstrbuf_add(name, cp, ep - cp);\n \treturn (const char *)ep + 1 - cp_;\n }\n-- \n2.18.0.windows.1\n\n"},{"id":"357580","messageId":"20180906210227.54368-3-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180906210227.54368-1-benpeart@microsoft.com","subject":"[PATCH v3 2/4] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-06T21:03:56Z","receivedAt":"2018-09-06T21:04:05Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"The End of Index Entry (EOIE) is used to locate the end of the variable\nlength index entries and the beginning of the extensions. Code can take\nadvantage of this to quickly locate the index extensions without having\nto parse through all of the index entries.\n\nBecause it must be able to be loaded before the variable length cache\nentries and other index extensions, this extension must be written last.\nThe signature for this extension is { 'E', 'O', 'I', 'E' }.\n\nThe extension consists of:\n\n- 32-bit offset to the end of the index entries\n\n- 160-bit SHA-1 over the extension types and their sizes (but not\ntheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\nlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\nthen the hash would be:\n\nSHA-1(\"TREE\" + <binary representation of N> +\n\t\"REUC\" + <binary representation of M>)\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n Documentation/technical/index-format.txt |  23 ++++\n read-cache.c                             | 149 +++++++++++++++++++++--\n t/README                                 |   5 +\n t/t1700-split-index.sh                   |   1 +\n 4 files changed, 170 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\nindex db3572626b..6bc2d90f7f 100644\n--- a/Documentation/technical/index-format.txt\n+++ b/Documentation/technical/index-format.txt\n@@ -314,3 +314,26 @@ The remaining data of each directory block is grouped by type:\n \n   - An ewah bitmap, the n-th bit indicates whether the n-th index entry\n     is not CE_FSMONITOR_VALID.\n+\n+== End of Index Entry\n+\n+  The End of Index Entry (EOIE) is used to locate the end of the variable\n+  length index entries and the begining of the extensions. Code can take\n+  advantage of this to quickly locate the index extensions without having\n+  to parse through all of the index entries.\n+\n+  Because it must be able to be loaded before the variable length cache\n+  entries and other index extensions, this extension must be written last.\n+  The signature for this extension is { 'E', 'O', 'I', 'E' }.\n+\n+  The extension consists of:\n+\n+  - 32-bit offset to the end of the index entries\n+\n+  - 160-bit SHA-1 over the extension types and their sizes (but not\n+\ttheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\tlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\tthen the hash would be:\n+\n+\tSHA-1(\"TREE\" + <binary representation of N> +\n+\t\t\"REUC\" + <binary representation of M>)\ndiff --git a/read-cache.c b/read-cache.c\nindex 382cc16bdc..d0d2793780 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -43,6 +43,7 @@\n #define CACHE_EXT_LINK 0x6c696e6b\t  /* \"link\" */\n #define CACHE_EXT_UNTRACKED 0x554E5452\t  /* \"UNTR\" */\n #define CACHE_EXT_FSMONITOR 0x46534D4E\t  /* \"FSMN\" */\n+#define CACHE_EXT_ENDOFINDEXENTRIES 0x454F4945\t/* \"EOIE\" */\n \n /* changes that can be kept in $GIT_DIR/index (basically all extensions) */\n #define EXTMASK (RESOLVE_UNDO_CHANGED | CACHE_TREE_CHANGED | \\\n@@ -1693,6 +1694,9 @@ static int read_index_extension(struct index_state *istate,\n \tcase CACHE_EXT_FSMONITOR:\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n+\tcase CACHE_EXT_ENDOFINDEXENTRIES:\n+\t\t/* already handled in do_read_index() */\n+\t\tbreak;\n \tdefault:\n \t\tif (*ext < 'A' || 'Z' < *ext)\n \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n@@ -1888,6 +1892,11 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+#ifndef NO_PTHREADS\n+static unsigned long read_eoie_extension(void *mmap, size_t mmap_size);\n+#endif\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -2197,11 +2206,15 @@ static int ce_write(git_hash_ctx *context, int fd, void *data, unsigned int len)\n \treturn 0;\n }\n \n-static int write_index_ext_header(git_hash_ctx *context, int fd,\n-\t\t\t\t  unsigned int ext, unsigned int sz)\n+static int write_index_ext_header(git_hash_ctx *context, git_hash_ctx *eoie_context,\n+\t\t\t\t  int fd, unsigned int ext, unsigned int sz)\n {\n \text = htonl(ext);\n \tsz = htonl(sz);\n+\tif (eoie_context) {\n+\t\tthe_hash_algo->update_fn(eoie_context, &ext, 4);\n+\t\tthe_hash_algo->update_fn(eoie_context, &sz, 4);\n+\t}\n \treturn ((ce_write(context, fd, &ext, 4) < 0) ||\n \t\t(ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n }\n@@ -2444,7 +2457,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n {\n \tuint64_t start = getnanotime();\n \tint newfd = tempfile->fd;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, eoie_c;\n \tstruct cache_header hdr;\n \tint i, err = 0, removed, extended, hdr_version;\n \tstruct cache_entry **cache = istate->cache;\n@@ -2453,6 +2466,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct ondisk_cache_entry_extended ondisk;\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n+\tunsigned long offset;\n \n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n@@ -2519,11 +2533,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\treturn err;\n \n \t/* Write extension data here */\n+\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\tthe_hash_algo->init_fn(&eoie_c);\n \tif (!strip_extensions && istate->split_index) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\terr = write_link_extension(&sb, istate) < 0 ||\n-\t\t\twrite_index_ext_header(&c, newfd, CACHE_EXT_LINK,\n+\t\t\twrite_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_LINK,\n \t\t\t\t\t       sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2534,7 +2550,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tcache_tree_write(&sb, istate->cache_tree);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_TREE, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_TREE, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2544,7 +2560,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tresolve_undo_write(&sb, istate->resolve_undo);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_RESOLVE_UNDO,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_RESOLVE_UNDO,\n \t\t\t\t\t     sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2555,7 +2571,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_untracked_extension(&sb, istate->untracked);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_UNTRACKED,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_UNTRACKED,\n \t\t\t\t\t     sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2566,7 +2582,23 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_fsmonitor_extension(&sb, istate);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+\n+\t/*\n+\t * CACHE_EXT_ENDOFINDEXENTRIES must be written as the last entry before the SHA1\n+\t * so that it can be found and processed before all the index entries are\n+\t * read.\n+\t */\n+\tif (!strip_extensions && offset && !git_env_bool(\"GIT_TEST_DISABLE_EOIE\", 0)) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\twrite_eoie_extension(&sb, &eoie_c, offset);\n+\t\terr = write_index_ext_header(&c, NULL, newfd, CACHE_EXT_ENDOFINDEXENTRIES, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2977,3 +3009,104 @@ int should_validate_cache_entries(void)\n \n \treturn validate_index_cache_entries;\n }\n+\n+#define EOIE_SIZE 24 /* <4-byte offset> + <20-byte hash> */\n+#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n+\n+#ifndef NO_PTHREADS\n+static unsigned long read_eoie_extension(void *mmap, size_t mmap_size)\n+{\n+\t/*\n+\t * The end of index entries (EOIE) extension is guaranteed to be last\n+\t * so that it can be found by scanning backwards from the EOF.\n+\t *\n+\t * \"EOIE\"\n+\t * <4-byte length>\n+\t * <4-byte offset>\n+\t * <20-byte hash>\n+\t */\n+\tconst char *index, *eoie = (const char *)mmap + mmap_size - GIT_SHA1_RAWSZ - EOIE_SIZE_WITH_HEADER;\n+\tuint32_t extsize;\n+\tunsigned long offset, src_offset;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\tgit_hash_ctx c;\n+\n+\t/* validate the extension signature */\n+\tindex = eoie;\n+\tif (CACHE_EXT(index) != CACHE_EXT_ENDOFINDEXENTRIES)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/* validate the extension size */\n+\textsize = get_be32(index);\n+\tif (extsize != EOIE_SIZE)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * Validate the offset we're going to look for the first extension\n+\t * signature is after the index header and before the eoie extension.\n+\t */\n+\toffset = get_be32(index);\n+\tif ((const char *)mmap + offset < (const char *)mmap + sizeof(struct cache_header))\n+\t\treturn 0;\n+\tif ((const char *)mmap + offset >= eoie)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * The hash is computed over extension types and their sizes (but not\n+\t * their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\t * long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\t * then the hash would be:\n+\t *\n+\t * SHA-1(\"TREE\" + <binary representation of N> +\n+\t *               \"REUC\" + <binary representation of M>)\n+\t */\n+\tsrc_offset = offset;\n+\tthe_hash_algo->init_fn(&c);\n+\twhile (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\n+\t\t/* verify the extension size isn't so large it will wrap around */\n+\t\tif (src_offset + 8 + extsize < src_offset)\n+\t\t\treturn 0;\n+\n+\t\tthe_hash_algo->update_fn(&c, (const char *)mmap + src_offset, 8);\n+\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\tthe_hash_algo->final_fn(hash, &c);\n+\tif (hashcmp(hash, (unsigned char *)index))\n+\t\treturn 0;\n+\n+\t/* Validate that the extension offsets returned us back to the eoie extension. */\n+\tif (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n+\t\treturn 0;\n+\n+\treturn offset;\n+}\n+#endif\n+\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset)\n+{\n+\tuint32_t buffer;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\n+\t/* offset */\n+\tput_be32(&buffer, offset);\n+\tstrbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+\t/* hash */\n+\tthe_hash_algo->final_fn(hash, eoie_context);\n+\tstrbuf_add(sb, hash, the_hash_algo->rawsz);\n+}\ndiff --git a/t/README b/t/README\nindex 9028b47d92..d8754dd23a 100644\n--- a/t/README\n+++ b/t/README\n@@ -319,6 +319,11 @@ GIT_TEST_OE_DELTA_SIZE=<n> exercises the uncomon pack-objects code\n path where deltas larger than this limit require extra memory\n allocation for bookkeeping.\n \n+GIT_TEST_DISABLE_EOIE=<boolean> disables writing the EOIE extension.\n+This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n+as they currently hard code SHA values for the index which are no longer\n+valid due to the addition of the EOIE extension.\n+\n Naming Tests\n ------------\n \ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex 39133bcbc8..f613dd72e3 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -7,6 +7,7 @@ test_description='split index mode tests'\n # We need total control of index splitting here\n sane_unset GIT_TEST_SPLIT_INDEX\n sane_unset GIT_FSMONITOR_TEST\n+export GIT_TEST_DISABLE_EOIE=true\n \n test_expect_success 'enable split index' '\n \tgit config splitIndex.maxPercentChange 100 &&\n-- \n2.18.0.windows.1\n\n"},{"id":"357581","messageId":"20180906210227.54368-4-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180906210227.54368-1-benpeart@microsoft.com","subject":"[PATCH v3 3/4] read-cache: load cache extensions on a worker thread","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-06T21:03:58Z","receivedAt":"2018-09-06T21:04:06Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by loading\nthe cache extensions on a worker thread in parallel with loading the cache\nentries.\n\nIn some cases, loading the extensions takes longer than loading the\ncache entries so this patch utilizes the new EOIE to start the thread to\nload the extensions before loading all the cache entries in parallel.\n\nThis is possible because the current extensions don't access the cache\nentries in the index_state structure so are OK that they don't all exist\nyet.\n\nThe CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\nextensions don't even get a pointer to the index so don't have access to the\ncache entries.\n\nCACHE_EXT_LINK only uses the index_state to initialize the split index.\nCACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\nupdate and dirty flags.\n\nI used p0002-read-cache.sh to generate some performance data:\n\np0002-read-cache.sh w/100,000 files\nBaseline         Thread extensions\n---------------------------------------\n21.14(0.03+0.01) 20.71(0.03+0.03) -2.0%\n\np0002-read-cache.sh w/1,000,000 files\nBaseline          Thread extensions\n------------------------------------------\n295.42(0.01+0.07) 217.60(0.03+0.04) -26.3%\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n Documentation/config.txt |  6 +++\n config.c                 | 18 ++++++++\n config.h                 |  1 +\n read-cache.c             | 94 ++++++++++++++++++++++++++++++++--------\n 4 files changed, 102 insertions(+), 17 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex 1c42364988..79f8296d9c 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -2391,6 +2391,12 @@ imap::\n \tThe configuration variables in the 'imap' section are described\n \tin linkgit:git-imap-send[1].\n \n+index.threads::\n+\tSpecifies the number of threads to spawn when loading the index.\n+\tThis is meant to reduce index load time on multiprocessor machines.\n+\tSpecifying 0 or 'true' will cause Git to auto-detect the number of\n+\tCPU's and set the number of threads accordingly. Defaults to 'true'.\n+\n index.version::\n \tSpecify the version with which new index files should be\n \tinitialized.  This does not affect existing repositories.\ndiff --git a/config.c b/config.c\nindex 9a0b10d4bc..9bd79fb165 100644\n--- a/config.c\n+++ b/config.c\n@@ -2289,6 +2289,24 @@ int git_config_get_fsmonitor(void)\n \treturn 0;\n }\n \n+/*\n+ * You can disable multi-threaded code by setting index.threads\n+ * to 'false' (or 1)\n+ */\n+int git_config_get_index_threads(void)\n+{\n+\tint is_bool, val;\n+\n+\tif (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n+\t\tif (is_bool)\n+\t\t\treturn val ? 0 : 1;\n+\t\telse\n+\t\t\treturn val;\n+\t}\n+\n+\treturn 0; /* auto-detect */\n+}\n+\n NORETURN\n void git_die_config_linenr(const char *key, const char *filename, int linenr)\n {\ndiff --git a/config.h b/config.h\nindex ab46e0165d..a06027e69b 100644\n--- a/config.h\n+++ b/config.h\n@@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n extern int git_config_get_split_index(void);\n extern int git_config_get_max_percent_split_change(void);\n extern int git_config_get_fsmonitor(void);\n+extern int git_config_get_index_threads(void);\n \n /* This dies if the configured or default date is in the future */\n extern int git_config_get_expiry(const char *key, const char **output);\ndiff --git a/read-cache.c b/read-cache.c\nindex d0d2793780..fcc776aaf0 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -23,6 +23,10 @@\n #include \"split-index.h\"\n #include \"utf8.h\"\n #include \"fsmonitor.h\"\n+#ifndef NO_PTHREADS\n+#include <pthread.h>\n+#include <thread-utils.h>\n+#endif\n \n /* Mask for the name length in ce_flags in the on-disk index */\n \n@@ -1897,6 +1901,46 @@ static unsigned long read_eoie_extension(void *mmap, size_t mmap_size);\n #endif\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n \n+struct load_index_extensions\n+{\n+#ifndef NO_PTHREADS\n+\tpthread_t pthread;\n+#endif\n+\tstruct index_state *istate;\n+\tvoid *mmap;\n+\tsize_t mmap_size;\n+\tunsigned long src_offset;\n+ };\n+\n+static void *load_index_extensions(void *_data)\n+{\n+\tstruct load_index_extensions *p = _data;\n+\tunsigned long src_offset = p->src_offset;\n+\n+\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\t\tif (read_index_extension(p->istate,\n+\t\t\t(const char *)p->mmap + src_offset,\n+\t\t\t(char *)p->mmap + src_offset + 8,\n+\t\t\textsize) < 0) {\n+\t\t\tmunmap(p->mmap, p->mmap_size);\n+\t\t\tdie(\"index file corrupt\");\n+\t\t}\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\n+\treturn NULL;\n+}\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -1907,6 +1951,11 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tvoid *mmap;\n \tsize_t mmap_size;\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tstruct load_index_extensions p = { 0 };\n+\tunsigned long extension_offset = 0;\n+#ifndef NO_PTHREADS\n+\tint nr_threads;\n+#endif\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1943,6 +1992,26 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n \tistate->initialized = 1;\n \n+\tp.istate = istate;\n+\tp.mmap = mmap;\n+\tp.mmap_size = mmap_size;\n+\n+#ifndef NO_PTHREADS\n+\tnr_threads = git_config_get_index_threads();\n+\tif (!nr_threads)\n+\t\tnr_threads = online_cpus();\n+\n+\tif (nr_threads >= 2) {\n+\t\textension_offset = read_eoie_extension(mmap, mmap_size);\n+\t\tif (extension_offset) {\n+\t\t\t/* create a thread to load the index extensions */\n+\t\t\tp.src_offset = extension_offset;\n+\t\t\tif (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n+\t\t\t\tdie(_(\"unable to create load_index_extensions_thread\"));\n+\t\t}\n+\t}\n+#endif\n+\n \tif (istate->version == 4) {\n \t\tprevious_name = &previous_name_buf;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n@@ -1969,23 +2038,14 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-\twhile (src_offset <= mmap_size - the_hash_algo->rawsz - 8) {\n-\t\t/* After an array of active_nr index entries,\n-\t\t * there can be arbitrary number of extended\n-\t\t * sections, each of which is prefixed with\n-\t\t * extension name (4-byte) and section length\n-\t\t * in 4-byte network byte order.\n-\t\t */\n-\t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n-\t\tif (read_index_extension(istate,\n-\t\t\t\t\t (const char *) mmap + src_offset,\n-\t\t\t\t\t (char *) mmap + src_offset + 8,\n-\t\t\t\t\t extsize) < 0)\n-\t\t\tgoto unmap;\n-\t\tsrc_offset += 8;\n-\t\tsrc_offset += extsize;\n+\t/* if we created a thread, join it otherwise load the extensions on the primary thread */\n+#ifndef NO_PTHREADS\n+\tif (extension_offset && pthread_join(p.pthread, NULL))\n+\t\tdie(_(\"unable to join load_index_extensions_thread\"));\n+#endif\n+\tif (!extension_offset) {\n+\t\tp.src_offset = src_offset;\n+\t\tload_index_extensions(&p);\n \t}\n \tmunmap(mmap, mmap_size);\n \treturn istate->cache_nr;\n-- \n2.18.0.windows.1\n\n"},{"id":"357582","messageId":"20180906210227.54368-5-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180906210227.54368-1-benpeart@microsoft.com","subject":"[PATCH v3 4/4] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-06T21:03:59Z","receivedAt":"2018-09-06T21:04:09Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by creating\nmultiple threads to divide the work of loading and converting the cache\nentries across all available CPU cores.\n\nIt accomplishes this by having the primary thread loop across the index file\ntracking the offset and (for V4 indexes) expanding the name. It creates a\nthread to process each block of entries as it comes to them.\n\nI used p0002-read-cache.sh to generate some performance data:\n\np0002-read-cache.sh w/100,000 files\nBaseline           Thread entries\n------------------------------------------\n20.71(0.03+0.03)   13.93(0.04+0.04) -32.7%\n\np0002-read-cache.sh w/1,000,000 files\nBaseline            Thread entries\n-------------------------------------------\n217.60(0.03+0.04)   199.00(0.00+0.10) -8.6%\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 242 +++++++++++++++++++++++++++++++++++++++++++++------\n t/README     |   6 ++\n 2 files changed, 220 insertions(+), 28 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex fcc776aaf0..8537a55750 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1941,20 +1941,212 @@ static void *load_index_extensions(void *_data)\n \treturn NULL;\n }\n \n+/*\n+ * A helper function that will load the specified range of cache entries\n+ * from the memory mapped file and add them to the given index.\n+ */\n+static unsigned long load_cache_entry_block(struct index_state *istate,\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n+\t\t\tunsigned long start_offset, struct strbuf *previous_name)\n+{\n+\tint i;\n+\tunsigned long src_offset = start_offset;\n+\n+\tfor (i = offset; i < offset + nr; i++) {\n+\t\tstruct ondisk_cache_entry *disk_ce;\n+\t\tstruct cache_entry *ce;\n+\t\tunsigned long consumed;\n+\n+\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tset_index_entry(istate, i, ce);\n+\n+\t\tsrc_offset += consumed;\n+\t}\n+\treturn src_offset - start_offset;\n+}\n+\n+static unsigned long load_all_cache_entries(struct index_state *istate,\n+\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tunsigned long consumed;\n+\n+\tif (istate->version == 4) {\n+\t\tprevious_name = &previous_name_buf;\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n+\t} else {\n+\t\tprevious_name = NULL;\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t\testimate_cache_size(mmap_size, istate->cache_nr));\n+\t}\n+\n+\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, previous_name);\n+\tstrbuf_release(&previous_name_buf);\n+\treturn consumed;\n+}\n+\n+#ifndef NO_PTHREADS\n+\n+/*\n+ * Mostly randomly chosen maximum thread counts: we\n+ * cap the parallelism to online_cpus() threads, and we want\n+ * to have at least 100000 cache entries per thread for it to\n+ * be worth starting a thread.\n+ */\n+#define THREAD_COST\t\t(10000)\n+\n+struct load_cache_entries_thread_data\n+{\n+\tpthread_t pthread;\n+\tstruct index_state *istate;\n+\tstruct mem_pool *ce_mem_pool;\n+\tint offset, nr;\n+\tvoid *mmap;\n+\tunsigned long start_offset;\n+\tstruct strbuf previous_name_buf;\n+\tstruct strbuf *previous_name;\n+\tunsigned long consumed;\t/* return # of bytes in index file processed */\n+};\n+\n+/*\n+ * A thread proc to run the load_cache_entries() computation\n+ * across multiple background threads.\n+ */\n+static void *load_cache_entries_thread(void *_data)\n+{\n+\tstruct load_cache_entries_thread_data *p = _data;\n+\n+\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n+\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n+\treturn NULL;\n+}\n+\n+static unsigned long load_cache_entries_threaded(int nr_threads, struct index_state *istate,\n+\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tstruct load_cache_entries_thread_data *data;\n+\tint ce_per_thread;\n+\tunsigned long consumed;\n+\tint i, thread;\n+\n+\t/* a little sanity checking */\n+\tif (istate->name_hash_initialized)\n+\t\tBUG(\"the name hash isn't thread safe\");\n+\n+\tmem_pool_init(&istate->ce_mem_pool, 0);\n+\tif (istate->version == 4)\n+\t\tprevious_name = &previous_name_buf;\n+\telse\n+\t\tprevious_name = NULL;\n+\n+\tce_per_thread = DIV_ROUND_UP(istate->cache_nr, nr_threads);\n+\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n+\n+\t/*\n+\t * Loop through index entries starting a thread for every ce_per_thread\n+\t * entries. Exit the loop when we've created the final thread (no need\n+\t * to parse the remaining entries.\n+\t */\n+\tconsumed = thread = 0;\n+\tfor (i = 0; ; i++) {\n+\t\tstruct ondisk_cache_entry *ondisk;\n+\t\tconst char *name;\n+\t\tunsigned int flags;\n+\n+\t\t/*\n+\t\t * we've reached the beginning of a block of cache entries,\n+\t\t * kick off a thread to process them\n+\t\t */\n+\t\tif (i % ce_per_thread == 0) {\n+\t\t\tstruct load_cache_entries_thread_data *p = &data[thread];\n+\n+\t\t\tp->istate = istate;\n+\t\t\tp->offset = i;\n+\t\t\tp->nr = ce_per_thread < istate->cache_nr - i ? ce_per_thread : istate->cache_nr - i;\n+\n+\t\t\t/* create a mem_pool for each thread */\n+\t\t\tif (istate->version == 4)\n+\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\t\t\testimate_cache_size_from_compressed(p->nr));\n+\t\t\telse\n+\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\t\t\testimate_cache_size(mmap_size, p->nr));\n+\n+\t\t\tp->mmap = mmap;\n+\t\t\tp->start_offset = src_offset;\n+\t\t\tif (previous_name) {\n+\t\t\t\tstrbuf_addbuf(&p->previous_name_buf, previous_name);\n+\t\t\t\tp->previous_name = &p->previous_name_buf;\n+\t\t\t}\n+\n+\t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n+\t\t\t\tdie(\"unable to create load_cache_entries_thread\");\n+\n+\t\t\t/* exit the loop when we've created the last thread */\n+\t\t\tif (++thread == nr_threads)\n+\t\t\t\tbreak;\n+\t\t}\n+\n+\t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\n+\t\t/* On-disk flags are just 16 bits */\n+\t\tflags = get_be16(&ondisk->flags);\n+\n+\t\tif (flags & CE_EXTENDED) {\n+\t\t\tstruct ondisk_cache_entry_extended *ondisk2;\n+\t\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n+\t\t\tname = ondisk2->name;\n+\t\t} else\n+\t\t\tname = ondisk->name;\n+\n+\t\tif (!previous_name) {\n+\t\t\tsize_t len;\n+\n+\t\t\t/* v3 and earlier */\n+\t\t\tlen = flags & CE_NAMEMASK;\n+\t\t\tif (len == CE_NAMEMASK)\n+\t\t\t\tlen = strlen(name);\n+\t\t\tsrc_offset += (flags & CE_EXTENDED) ?\n+\t\t\t\tondisk_cache_entry_extended_size(len) :\n+\t\t\t\tondisk_cache_entry_size(len);\n+\t\t} else\n+\t\t\tsrc_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n+\t}\n+\n+\tfor (i = 0; i < nr_threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = data + i;\n+\t\tif (pthread_join(p->pthread, NULL))\n+\t\t\tdie(\"unable to join load_cache_entries_thread\");\n+\t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n+\t\tstrbuf_release(&p->previous_name_buf);\n+\t\tconsumed += p->consumed;\n+\t}\n+\n+\tfree(data);\n+\tstrbuf_release(&previous_name_buf);\n+\n+\treturn consumed;\n+}\n+\n+#endif\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n-\tint fd, i;\n+\tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tstruct load_index_extensions p = { 0 };\n \tunsigned long extension_offset = 0;\n #ifndef NO_PTHREADS\n-\tint nr_threads;\n+\tint cpus, nr_threads;\n #endif\n \n \tif (istate->initialized)\n@@ -1996,10 +2188,20 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tp.mmap = mmap;\n \tp.mmap_size = mmap_size;\n \n+\tsrc_offset = sizeof(*hdr);\n+\n #ifndef NO_PTHREADS\n \tnr_threads = git_config_get_index_threads();\n-\tif (!nr_threads)\n-\t\tnr_threads = online_cpus();\n+\tif (!nr_threads) {\n+\t\tcpus = online_cpus();\n+\t\tnr_threads = istate->cache_nr / THREAD_COST;\n+\t\tif (nr_threads > cpus)\n+\t\t\tnr_threads = cpus;\n+\t}\n+\n+\t/* enable testing with fewer than default minimum of entries */\n+\tif (istate->cache_nr > 1 && nr_threads < 3 && git_env_bool(\"GIT_TEST_INDEX_THREADS\", 0))\n+\t\tnr_threads = 3;\n \n \tif (nr_threads >= 2) {\n \t\textension_offset = read_eoie_extension(mmap, mmap_size);\n@@ -2008,33 +2210,17 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\t\tp.src_offset = extension_offset;\n \t\t\tif (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n \t\t\t\tdie(_(\"unable to create load_index_extensions_thread\"));\n+\t\t\tnr_threads--;\n \t\t}\n \t}\n+\tif (nr_threads >= 2)\n+\t\tsrc_offset += load_cache_entries_threaded(nr_threads, istate, mmap, mmap_size, src_offset);\n+\telse\n+\t\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+#else\n+\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n #endif\n \n-\tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n-\t} else {\n-\t\tprevious_name = NULL;\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n-\t}\n-\n-\tsrc_offset = sizeof(*hdr);\n-\tfor (i = 0; i < istate->cache_nr; i++) {\n-\t\tstruct ondisk_cache_entry *disk_ce;\n-\t\tstruct cache_entry *ce;\n-\t\tunsigned long consumed;\n-\n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_name);\n-\t\tset_index_entry(istate, i, ce);\n-\n-\t\tsrc_offset += consumed;\n-\t}\n-\tstrbuf_release(&previous_name_buf);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \ndiff --git a/t/README b/t/README\nindex d8754dd23a..59015f7150 100644\n--- a/t/README\n+++ b/t/README\n@@ -324,6 +324,12 @@ This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n as they currently hard code SHA values for the index which are no longer\n valid due to the addition of the EOIE extension.\n \n+GIT_TEST_INDEX_THREADS=<boolean> forces multi-threaded loading of\n+the index cache entries and extensions for the whole test suite.\n+\n Naming Tests\n ------------\n \n-- \n2.18.0.windows.1\n\n"},{"id":"357605","messageId":"20180907041657.GA12835@tor.lan","threadId":"49204","inReplyTo":"20180906210227.54368-5-benpeart@microsoft.com","subject":"Re: [PATCH v3 4/4] read-cache: speed up index load through parallelization","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2018-09-07T04:16:57Z","receivedAt":"2018-09-07T04:17:07Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"\n> diff --git a/read-cache.c b/read-cache.c\n> index fcc776aaf0..8537a55750 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1941,20 +1941,212 @@ static void *load_index_extensions(void *_data)\n>  \treturn NULL;\n>  }\n>  \n> +/*\n> + * A helper function that will load the specified range of cache entries\n> + * from the memory mapped file and add them to the given index.\n> + */\n> +static unsigned long load_cache_entry_block(struct index_state *istate,\n> +\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n> +\t\t\tunsigned long start_offset, struct strbuf *previous_name)\n> +{\n> +\tint i;\n> +\tunsigned long src_offset = start_offset;\n\nI read an unsigned long here:\nshould that be a size_t instead ?\n\n(And probably even everywhere else in this patch)\n\n> +\n> +\tfor (i = offset; i < offset + nr; i++) {\n> +\t\tstruct ondisk_cache_entry *disk_ce;\n> +\t\tstruct cache_entry *ce;\n> +\t\tunsigned long consumed;\n> +\n> +\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n> +\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n> +\t\tset_index_entry(istate, i, ce);\n> +\n> +\t\tsrc_offset += consumed;\n> +\t}\n> +\treturn src_offset - start_offset;\n> +}\n> +\n> +static unsigned long load_all_cache_entries(struct index_state *istate,\n> +\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n> +{\n> +\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n> +\tunsigned long consumed;\n> +\n> +\tif (istate->version == 4) {\n> +\t\tprevious_name = &previous_name_buf;\n> +\t\tmem_pool_init(&istate->ce_mem_pool,\n> +\t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n> +\t} else {\n> +\t\tprevious_name = NULL;\n> +\t\tmem_pool_init(&istate->ce_mem_pool,\n> +\t\t\t\testimate_cache_size(mmap_size, istate->cache_nr));\n> +\t}\n> +\n> +\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n> +\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, previous_name);\n> +\tstrbuf_release(&previous_name_buf);\n> +\treturn consumed;\n> +}\n> +\n> +#ifndef NO_PTHREADS\n> +\n"},{"id":"357623","messageId":"18022c85-cf2d-026e-d643-2b872f3cfbe8@gmail.com","threadId":"49204","inReplyTo":"20180907041657.GA12835@tor.lan","subject":"Re: [PATCH v3 4/4] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-07T13:43:39Z","receivedAt":"2018-09-07T13:43:45Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/7/2018 12:16 AM, Torsten Bögershausen wrote:\n> \n>> diff --git a/read-cache.c b/read-cache.c\n>> index fcc776aaf0..8537a55750 100644\n>> --- a/read-cache.c\n>> +++ b/read-cache.c\n>> @@ -1941,20 +1941,212 @@ static void *load_index_extensions(void *_data)\n>>   \treturn NULL;\n>>   }\n>>   \n>> +/*\n>> + * A helper function that will load the specified range of cache entries\n>> + * from the memory mapped file and add them to the given index.\n>> + */\n>> +static unsigned long load_cache_entry_block(struct index_state *istate,\n>> +\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n>> +\t\t\tunsigned long start_offset, struct strbuf *previous_name)\n>> +{\n>> +\tint i;\n>> +\tunsigned long src_offset = start_offset;\n> \n> I read an unsigned long here:\n> should that be a size_t instead ?\n> \n> (And probably even everywhere else in this patch)\n> \n\nIt's a fair question.  The pre-patch code had a mix of unsigned long and \nsize_t.  Both src_offset and consumed were unsigned long but mmap_size \nwas a size_t.  I stuck with that pattern for consistency.\n\nWhile it would be possible to convert everything to size_t as a step to \nenable index files >4 GB, I have a hard time believing that will be \nnecessary for a very long time and would likely require more substantial \nchanges to enable that to work.\n\n>> +\n>> +\tfor (i = offset; i < offset + nr; i++) {\n>> +\t\tstruct ondisk_cache_entry *disk_ce;\n>> +\t\tstruct cache_entry *ce;\n>> +\t\tunsigned long consumed;\n>> +\n>> +\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n>> +\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n>> +\t\tset_index_entry(istate, i, ce);\n>> +\n>> +\t\tsrc_offset += consumed;\n>> +\t}\n>> +\treturn src_offset - start_offset;\n>> +}\n>> +\n>> +static unsigned long load_all_cache_entries(struct index_state *istate,\n>> +\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n>> +{\n>> +\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n>> +\tunsigned long consumed;\n>> +\n>> +\tif (istate->version == 4) {\n>> +\t\tprevious_name = &previous_name_buf;\n>> +\t\tmem_pool_init(&istate->ce_mem_pool,\n>> +\t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n>> +\t} else {\n>> +\t\tprevious_name = NULL;\n>> +\t\tmem_pool_init(&istate->ce_mem_pool,\n>> +\t\t\t\testimate_cache_size(mmap_size, istate->cache_nr));\n>> +\t}\n>> +\n>> +\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n>> +\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, previous_name);\n>> +\tstrbuf_release(&previous_name_buf);\n>> +\treturn consumed;\n>> +}\n>> +\n>> +#ifndef NO_PTHREADS\n>> +\n"},{"id":"357635","messageId":"xmqq5zzhxlxm.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180906210227.54368-1-benpeart@microsoft.com","subject":"Re: [PATCH v3 0/4] read-cache: speed up index load through parallelization","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-07T17:21:57Z","receivedAt":"2018-09-07T17:22:03Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <benpeart@microsoft.com> writes:\n\n> On further investigation with the previous patch, I noticed that my test\n> repos didn't contain the cache tree extension in their index. After doing a\n> commit to ensure they existed, I realized that in some instances, the time\n> to load the cache tree exceeded the time to load all the cache entries in\n> parallel.  Because the thread to read the cache tree was started last (due\n> to having to parse through all the cache entries first) we weren't always\n> getting optimal performance.\n>\n> To better optimize for this case, I decided to write the EOIE extension\n> as suggested by Junio [1] in response to my earlier multithreading patch\n> series [2].  This enables me to spin up the thread to load the extensions\n> earlier as it no longer has to parse through all the cache entries first.\n\nHmph. I kinda liked the simplicity of the previous one, but if we\nneed to start reading the extension sections sooner by eliminating\nthe overhead to scan the cache entries, perhaps we should bite the\nbullet now.\n\n> The big changes in this iteration are:\n>\n> - add the EOIE extension\n> - update the index extension worker thread to start first\n\nI guess I'd need to see the actual patch to find this out, but once\nwe rely on a new extension, then we could omit scanning the main\nindex even to partition the work among workers (i.e. like the topic\nlong ago, you can have list of pointers into the main index to help\npartitioning, plus reset the prefix compression used in v4).  I\nthink you didn't get that far in this round, which is good.  If the\ngain with EOIE alone (and starting the worker for the extension\nsection early) is large enough without such a pre-computed work\npartition table, the simplicity of this round may give us a good\nstopping point.\n\n> This patch conflicts with Duy's patch to remove the double memory copy and\n> pass in the previous ce instead.  The two will need to be merged/reconciled\n> once they settle down a bit.\n\nThanks.  I have a feeling that 67922abb (\"read-cache.c: optimize\nreading index format v4\", 2018-09-02) is already 'next'-worthy\nand ready to be built on, but I'd prefer to hear from Duy to double\ncheck.\n\n"},{"id":"357639","messageId":"xmqqpnxpw5sn.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180906210227.54368-3-benpeart@microsoft.com","subject":"Re: [PATCH v3 2/4] eoie: add End of Index Entry (EOIE) extension","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-07T17:55:52Z","receivedAt":"2018-09-07T17:55:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <benpeart@microsoft.com> writes:\n\n> The extension consists of:\n>\n> - 32-bit offset to the end of the index entries\n>\n> - 160-bit SHA-1 over the extension types and their sizes (but not\n> their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n> long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n> then the hash would be:\n>\n> SHA-1(\"TREE\" + <binary representation of N> +\n> \t\"REUC\" + <binary representation of M>)\n\nI didn't look at the documentation patch in the larger context, but\nplease make sure that it is clear to the readers that these fixed\nwidth integers \"binary representations\" use network byte order.\n\nI briefly wondered if the above should include\n\n    + \"EOIE\" + <binary representation of (32+160)/8 = 24>\n\nas it is pretty much common file format design to include the header\npart of the checksum record (with checksum values padded out with NUL\nbytes) when you define a record to hold the checksum of the entire\nfile.  Since this does not protect the contents of each section from\nbit-flipping, adding the data for EOIE itself in the sum does not\ngive us much (iow, what I am adding above is a constant that does\nnot contribute any entropy).\n\nHow big is a typical TREE extension in _your_ work repository\nhousing Windows sources?  I am guessing that replacing SHA-1 with\nsomething faster (as this is not about security but is about disk\ncorruption) and instead hash also the contents of these sections\nwould NOT help all that much in the performance department, as\nhaving to page them in to read through would already consume\nnon-trivial amount of time, and that is why you are not hashing the\ncontents.\n\n> +\t/*\n> +\t * CACHE_EXT_ENDOFINDEXENTRIES must be written as the last entry before the SHA1\n\ns/SHA1/trailing checksum/ or something so that we can withstand\nNewHash world order?\n\n> +\t * so that it can be found and processed before all the index entries are\n> +\t * read.\n> +\t */\n> +\tif (!strip_extensions && offset && !git_env_bool(\"GIT_TEST_DISABLE_EOIE\", 0)) {\n> +\t\tstruct strbuf sb = STRBUF_INIT;\n> +\n> +\t\twrite_eoie_extension(&sb, &eoie_c, offset);\n> +\t\terr = write_index_ext_header(&c, NULL, newfd, CACHE_EXT_ENDOFINDEXENTRIES, sb.len) < 0\n>  \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n>  \t\tstrbuf_release(&sb);\n>  \t\tif (err)\n\nOK.\n\n> +#define EOIE_SIZE 24 /* <4-byte offset> + <20-byte hash> */\n> +#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n> +\n> +#ifndef NO_PTHREADS\n> +static unsigned long read_eoie_extension(void *mmap, size_t mmap_size)\n> +{\n> +\t/*\n> +\t * The end of index entries (EOIE) extension is guaranteed to be last\n> +\t * so that it can be found by scanning backwards from the EOF.\n> +\t *\n> +\t * \"EOIE\"\n> +\t * <4-byte length>\n> +\t * <4-byte offset>\n> +\t * <20-byte hash>\n> +\t */\n> +\tconst char *index, *eoie = (const char *)mmap + mmap_size - GIT_SHA1_RAWSZ - EOIE_SIZE_WITH_HEADER;\n> +\tuint32_t extsize;\n> +\tunsigned long offset, src_offset;\n> +\tunsigned char hash[GIT_MAX_RAWSZ];\n> +\tgit_hash_ctx c;\n> +\n> +\t/* validate the extension signature */\n> +\tindex = eoie;\n> +\tif (CACHE_EXT(index) != CACHE_EXT_ENDOFINDEXENTRIES)\n> +\t\treturn 0;\n> +\tindex += sizeof(uint32_t);\n> +\n> +\t/* validate the extension size */\n> +\textsize = get_be32(index);\n> +\tif (extsize != EOIE_SIZE)\n> +\t\treturn 0;\n> +\tindex += sizeof(uint32_t);\n\nDo we know we have at least 8-byte to consume to perform the above\ntwo checks, or is that something we need to verify at the beginning\nof this function?  Better yet, as we know that a correct EOIE with\nits own header is 28-byte long, we probably should abort if\nmmap_size is smaller than that.\n\n> +\t/*\n> +\t * Validate the offset we're going to look for the first extension\n> +\t * signature is after the index header and before the eoie extension.\n> +\t */\n> +\toffset = get_be32(index);\n> +\tif ((const char *)mmap + offset < (const char *)mmap + sizeof(struct cache_header))\n> +\t\treturn 0;\n\nClaims that the end is before the beginning, which we reject as bogus.  Good.\n\n> +\tif ((const char *)mmap + offset >= eoie)\n> +\t\treturn 0;\n\nClaims that the end is beyond the EOIE marker we should have placed\nafter it, which we reject as bogus.  Good.\n\n> +\tindex += sizeof(uint32_t);\n> +\n> +\t/*\n> +\t * The hash is computed over extension types and their sizes (but not\n> +\t * their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n> +\t * long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n> +\t * then the hash would be:\n> +\t *\n> +\t * SHA-1(\"TREE\" + <binary representation of N> +\n> +\t *               \"REUC\" + <binary representation of M>)\n> +\t */\n> +\tsrc_offset = offset;\n> +\tthe_hash_algo->init_fn(&c);\n> +\twhile (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n> +\t\t/* After an array of active_nr index entries,\n(Style nit).\n> +\t\t * there can be arbitrary number of extended\n> +\t\t * sections, each of which is prefixed with\n> +\t\t * extension name (4-byte) and section length\n> +\t\t * in 4-byte network byte order.\n> +\t\t */\n> +\t\tuint32_t extsize;\n> +\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n> +\t\textsize = ntohl(extsize);\n\nEarlier we were using get_be32() but now we use memcpy with ntohl()?\nHow are we choosing which one to use?\n\nI think you meant to cast mmap to (const char *) here.  It may make it\neasier to write and read if we started this function like so:\n\n\tstatic unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n\t{\n\t\tconst char *mmap = mmap_;\n\nthen we do not have to keep casting mmap and cast to a wrong type by\nmistake.\n\n> +\n> +\t\t/* verify the extension size isn't so large it will wrap around */\n> +\t\tif (src_offset + 8 + extsize < src_offset)\n> +\t\t\treturn 0;\n\nGood.\n\n> +\t\tthe_hash_algo->update_fn(&c, (const char *)mmap + src_offset, 8);\n> +\n> +\t\tsrc_offset += 8;\n> +\t\tsrc_offset += extsize;\n> +\t}\n> +\tthe_hash_algo->final_fn(hash, &c);\n> +\tif (hashcmp(hash, (unsigned char *)index))\n> +\t\treturn 0;\n> +\n> +\t/* Validate that the extension offsets returned us back to the eoie extension. */\n> +\tif (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n> +\t\treturn 0;\n\nVery good.\n\n> +\treturn offset;\n> +}\n> +#endif\n\nOverall it looks like it is carefully done.\nThanks.\n"},{"id":"357647","messageId":"6ef05d29-b303-0282-7ca7-6e51efded005@gmail.com","threadId":"49204","inReplyTo":"xmqq5zzhxlxm.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 0/4] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-07T18:31:54Z","receivedAt":"2018-09-07T18:31:59Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/7/2018 1:21 PM, Junio C Hamano wrote:\n> Ben Peart <benpeart@microsoft.com> writes:\n> \n>> On further investigation with the previous patch, I noticed that my test\n>> repos didn't contain the cache tree extension in their index. After doing a\n>> commit to ensure they existed, I realized that in some instances, the time\n>> to load the cache tree exceeded the time to load all the cache entries in\n>> parallel.  Because the thread to read the cache tree was started last (due\n>> to having to parse through all the cache entries first) we weren't always\n>> getting optimal performance.\n>>\n>> To better optimize for this case, I decided to write the EOIE extension\n>> as suggested by Junio [1] in response to my earlier multithreading patch\n>> series [2].  This enables me to spin up the thread to load the extensions\n>> earlier as it no longer has to parse through all the cache entries first.\n> \n> Hmph. I kinda liked the simplicity of the previous one, but if we\n> need to start reading the extension sections sooner by eliminating\n> the overhead to scan the cache entries, perhaps we should bite the\n> bullet now.\n> \n\nI preferred the simplicity as well but when I was profiling the code and \nfound out that loading the extensions was most often the last thread to \ncomplete, I took this intermediate step to speed things up.\n\n>> The big changes in this iteration are:\n>>\n>> - add the EOIE extension\n>> - update the index extension worker thread to start first\n> \n> I guess I'd need to see the actual patch to find this out, but once\n> we rely on a new extension, then we could omit scanning the main\n> index even to partition the work among workers (i.e. like the topic\n> long ago, you can have list of pointers into the main index to help\n> partitioning, plus reset the prefix compression used in v4).  I\n> think you didn't get that far in this round, which is good.  If the\n> gain with EOIE alone (and starting the worker for the extension\n> section early) is large enough without such a pre-computed work\n> partition table, the simplicity of this round may give us a good\n> stopping point.\n> \n\nAgreed.  I didn't go that far in this series as it doesn't appear to be \nnecessary.  We could always add that later if it turned out to be worth \nthe additional complexity.\n\n>> This patch conflicts with Duy's patch to remove the double memory copy and\n>> pass in the previous ce instead.  The two will need to be merged/reconciled\n>> once they settle down a bit.\n> \n> Thanks.  I have a feeling that 67922abb (\"read-cache.c: optimize\n> reading index format v4\", 2018-09-02) is already 'next'-worthy\n> and ready to be built on, but I'd prefer to hear from Duy to double\n> check.\n> \n\nI'll take a closer look at what this will entail. It gets more \ncomplicated as we don't actually have a previous expanded cache entry \nwhen starting each thread.  I'll see how complex it makes the code and \nhow much additional performance it gives.\n"},{"id":"357653","messageId":"fc531863-c46c-6d27-4749-c6b092a14a6f@gmail.com","threadId":"49204","inReplyTo":"xmqqpnxpw5sn.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 2/4] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-07T20:23:51Z","receivedAt":"2018-09-07T20:23:56Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/7/2018 1:55 PM, Junio C Hamano wrote:\n> Ben Peart <benpeart@microsoft.com> writes:\n> \n>> The extension consists of:\n>>\n>> - 32-bit offset to the end of the index entries\n>>\n>> - 160-bit SHA-1 over the extension types and their sizes (but not\n>> their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n>> long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n>> then the hash would be:\n>>\n>> SHA-1(\"TREE\" + <binary representation of N> +\n>> \t\"REUC\" + <binary representation of M>)\n> \n> I didn't look at the documentation patch in the larger context, but\n> please make sure that it is clear to the readers that these fixed\n> width integers \"binary representations\" use network byte order.\n> \n\nAt the top of the documentation it says \"All binary numbers are in \nnetwork byte order\" and that is not repeated for any of the other \nsections that are documenting the file format.\n\n> I briefly wondered if the above should include\n> \n>      + \"EOIE\" + <binary representation of (32+160)/8 = 24>\n> \n> as it is pretty much common file format design to include the header\n> part of the checksum record (with checksum values padded out with NUL\n> bytes) when you define a record to hold the checksum of the entire\n> file.  Since this does not protect the contents of each section from\n> bit-flipping, adding the data for EOIE itself in the sum does not\n> give us much (iow, what I am adding above is a constant that does\n> not contribute any entropy).\n> \n> How big is a typical TREE extension in _your_ work repository\n> housing Windows sources?  I am guessing that replacing SHA-1 with\n> something faster (as this is not about security but is about disk\n> corruption) and instead hash also the contents of these sections\n> would NOT help all that much in the performance department, as\n> having to page them in to read through would already consume\n> non-trivial amount of time, and that is why you are not hashing the\n> contents.\n> \n\nThe purpose of the SHA isn't to detect disk corruption (we already have \na SHA for the entire index that can serve that purpose) but to help \nensure that this was actually a valid EOIE extension and not a lucky \nrandom set of bytes.  I had used leading and trailing signature bytes \nalong with the length and version bytes to validate it was an actual \nEOIE extension but you suggested [1] that I use a SHA of the 4 byte \nextension type + 4 byte extension length instead so I rewrote it that way.\n\n[1] \nhttps://public-inbox.org/git/xmqq1sl017dw.fsf@gitster.mtv.corp.google.com/\n\n>> +\t/*\n>> +\t * CACHE_EXT_ENDOFINDEXENTRIES must be written as the last entry before the SHA1\n> \n> s/SHA1/trailing checksum/ or something so that we can withstand\n> NewHash world order?\n> \n\nI thought about this but in the document elsewhere it refers to it as \n\"160-bit SHA-1 over the content of the index file before this checksum.\" \nand there are at least a dozen other references to \"SHA-1\" so I figured \nwe can fix them all at the same time when we have a new/better name. :-)\n\n>> +\t * so that it can be found and processed before all the index entries are\n>> +\t * read.\n>> +\t */\n>> +\tif (!strip_extensions && offset && !git_env_bool(\"GIT_TEST_DISABLE_EOIE\", 0)) {\n>> +\t\tstruct strbuf sb = STRBUF_INIT;\n>> +\n>> +\t\twrite_eoie_extension(&sb, &eoie_c, offset);\n>> +\t\terr = write_index_ext_header(&c, NULL, newfd, CACHE_EXT_ENDOFINDEXENTRIES, sb.len) < 0\n>>   \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n>>   \t\tstrbuf_release(&sb);\n>>   \t\tif (err)\n> \n> OK.\n> \n>> +#define EOIE_SIZE 24 /* <4-byte offset> + <20-byte hash> */\n>> +#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n>> +\n>> +#ifndef NO_PTHREADS\n>> +static unsigned long read_eoie_extension(void *mmap, size_t mmap_size)\n>> +{\n>> +\t/*\n>> +\t * The end of index entries (EOIE) extension is guaranteed to be last\n>> +\t * so that it can be found by scanning backwards from the EOF.\n>> +\t *\n>> +\t * \"EOIE\"\n>> +\t * <4-byte length>\n>> +\t * <4-byte offset>\n>> +\t * <20-byte hash>\n>> +\t */\n>> +\tconst char *index, *eoie = (const char *)mmap + mmap_size - GIT_SHA1_RAWSZ - EOIE_SIZE_WITH_HEADER;\n>> +\tuint32_t extsize;\n>> +\tunsigned long offset, src_offset;\n>> +\tunsigned char hash[GIT_MAX_RAWSZ];\n>> +\tgit_hash_ctx c;\n>> +\n>> +\t/* validate the extension signature */\n>> +\tindex = eoie;\n>> +\tif (CACHE_EXT(index) != CACHE_EXT_ENDOFINDEXENTRIES)\n>> +\t\treturn 0;\n>> +\tindex += sizeof(uint32_t);\n>> +\n>> +\t/* validate the extension size */\n>> +\textsize = get_be32(index);\n>> +\tif (extsize != EOIE_SIZE)\n>> +\t\treturn 0;\n>> +\tindex += sizeof(uint32_t);\n> \n> Do we know we have at least 8-byte to consume to perform the above\n> two checks, or is that something we need to verify at the beginning\n> of this function?  Better yet, as we know that a correct EOIE with\n> its own header is 28-byte long, we probably should abort if\n> mmap_size is smaller than that.\n> \n\nI'll add that additional test.\n\n>> +\t/*\n>> +\t * Validate the offset we're going to look for the first extension\n>> +\t * signature is after the index header and before the eoie extension.\n>> +\t */\n>> +\toffset = get_be32(index);\n>> +\tif ((const char *)mmap + offset < (const char *)mmap + sizeof(struct cache_header))\n>> +\t\treturn 0;\n> \n> Claims that the end is before the beginning, which we reject as bogus.  Good.\n> \n>> +\tif ((const char *)mmap + offset >= eoie)\n>> +\t\treturn 0;\n> \n> Claims that the end is beyond the EOIE marker we should have placed\n> after it, which we reject as bogus.  Good.\n> \n>> +\tindex += sizeof(uint32_t);\n>> +\n>> +\t/*\n>> +\t * The hash is computed over extension types and their sizes (but not\n>> +\t * their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n>> +\t * long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n>> +\t * then the hash would be:\n>> +\t *\n>> +\t * SHA-1(\"TREE\" + <binary representation of N> +\n>> +\t *               \"REUC\" + <binary representation of M>)\n>> +\t */\n>> +\tsrc_offset = offset;\n>> +\tthe_hash_algo->init_fn(&c);\n>> +\twhile (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n>> +\t\t/* After an array of active_nr index entries,\n> (Style nit).\n>> +\t\t * there can be arbitrary number of extended\n>> +\t\t * sections, each of which is prefixed with\n>> +\t\t * extension name (4-byte) and section length\n>> +\t\t * in 4-byte network byte order.\n>> +\t\t */\n>> +\t\tuint32_t extsize;\n>> +\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n>> +\t\textsize = ntohl(extsize);\n> \n> Earlier we were using get_be32() but now we use memcpy with ntohl()?\n> How are we choosing which one to use?\n> \n\nI literally copy/pasted this logic from the code that actually loads the \nextensions then removed the call to load the extension and replaced it \nwith the call to update the hash.  I kept it the same to facilitate \nconsistency for any future fixes or changes.\n\n> I think you meant to cast mmap to (const char *) here.  It may make it\n> easier to write and read if we started this function like so:\n> \n> \tstatic unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n> \t{\n> \t\tconst char *mmap = mmap_;\n> \n> then we do not have to keep casting mmap and cast to a wrong type by\n> mistake.\n> \n\nGood suggestion.\n\n>> +\n>> +\t\t/* verify the extension size isn't so large it will wrap around */\n>> +\t\tif (src_offset + 8 + extsize < src_offset)\n>> +\t\t\treturn 0;\n> \n> Good.\n> \n>> +\t\tthe_hash_algo->update_fn(&c, (const char *)mmap + src_offset, 8);\n>> +\n>> +\t\tsrc_offset += 8;\n>> +\t\tsrc_offset += extsize;\n>> +\t}\n>> +\tthe_hash_algo->final_fn(hash, &c);\n>> +\tif (hashcmp(hash, (unsigned char *)index))\n>> +\t\treturn 0;\n>> +\n>> +\t/* Validate that the extension offsets returned us back to the eoie extension. */\n>> +\tif (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n>> +\t\treturn 0;\n> \n> Very good.\n> \n>> +\treturn offset;\n>> +}\n>> +#endif\n> \n> Overall it looks like it is carefully done.\n\nThanks for the careful review!\n\n> Thanks.\n> \n"},{"id":"357655","messageId":"xmqqbm99vwsy.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180906210227.54368-4-benpeart@microsoft.com","subject":"Re: [PATCH v3 3/4] read-cache: load cache extensions on a worker thread","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-07T21:10:05Z","receivedAt":"2018-09-07T21:10:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <benpeart@microsoft.com> writes:\n\n> +struct load_index_extensions\n> +{\n> +#ifndef NO_PTHREADS\n> +\tpthread_t pthread;\n> +#endif\n> +\tstruct index_state *istate;\n> +\tvoid *mmap;\n> +\tsize_t mmap_size;\n> +\tunsigned long src_offset;\n\nIf the file format only allows uint32_t on any platform, perhaps\nthis is better specified as uint32_t?  Or if this is offset into\na mmap'ed region of memory, size_t may be more appropriate.\n\nSame comment applies to \"extension_offset\" we see below (which in\nturn means the returned type of read_eoie_extension() function may\nwant to match).\n\n> + };\n\nSpace before '}'??\n\n> +\n> +static void *load_index_extensions(void *_data)\n> +{\n> +\tstruct load_index_extensions *p = _data;\n\nPerhaps we are being superstitious, but I think our code try to\navoid leading underscore when able, i.e.\n\n\tload_index_extensions(void *data_)\n\t{\n\t\tstruct load_index_extensions *p = data;\n\n> +\tunsigned long src_offset = p->src_offset;\n> +\n> +\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n> +\t\t/* After an array of active_nr index entries,\n> +\t\t * there can be arbitrary number of extended\n> +\t\t * sections, each of which is prefixed with\n> +\t\t * extension name (4-byte) and section length\n> +\t\t * in 4-byte network byte order.\n> +\t\t */\n> +\t\tuint32_t extsize;\n> +\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n> +\t\textsize = ntohl(extsize);\n\nThe same \"ntohl(), not get_be32()?\" question as the one for the\nprevious step applies here, too.  I think the answer is \"the\noriginal was written that way\" and that is acceptable, but once this\nseries lands, we may want to review the whole file and see if it is\nworth making them consistent with a separate clean-up patch.\n\nI think mmap() and munmap() are the only places that wants p->mmap\nand mmap parameters passed around in various callchains to be of\ntype \"void *\"---I wonder if it is simpler to use \"const char *\"\nthroughout and only cast it to \"void *\" when necessary (I suspect\nthat there is nowhere we need to cast to or from \"void *\" explicitly\nif we did so---assignment and argument passing would give us an\nappropriate cast for free)?\n\n> +\t\tif (read_index_extension(p->istate,\n> +\t\t\t(const char *)p->mmap + src_offset,\n> +\t\t\t(char *)p->mmap + src_offset + 8,\n> +\t\t\textsize) < 0) {\n> +\t\t\tmunmap(p->mmap, p->mmap_size);\n> +\t\t\tdie(\"index file corrupt\");\n> +\t\t}\n> +\t...\n> @@ -1907,6 +1951,11 @@ ...\n> ...\n> +\tp.mmap = mmap;\n> +\tp.mmap_size = mmap_size;\n> +\n> +#ifndef NO_PTHREADS\n> +\tnr_threads = git_config_get_index_threads();\n> +\tif (!nr_threads)\n> +\t\tnr_threads = online_cpus();\n> +\n> +\tif (nr_threads >= 2) {\n> +\t\textension_offset = read_eoie_extension(mmap, mmap_size);\n> +\t\tif (extension_offset) {\n> +\t\t\t/* create a thread to load the index extensions */\n> +\t\t\tp.src_offset = extension_offset;\n> +\t\t\tif (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n> +\t\t\t\tdie(_(\"unable to create load_index_extensions_thread\"));\n> +\t\t}\n> +\t}\n> +#endif\n\nMakes sense.\n"},{"id":"357680","messageId":"CAN0heSreAfMsseZcxR75CFDph-n1b8EUNsRhpFsVqxMLc0hvpA@mail.gmail.com","threadId":"49204","inReplyTo":"fc531863-c46c-6d27-4749-c6b092a14a6f@gmail.com","subject":"Re: [PATCH v3 2/4] eoie: add End of Index Entry (EOIE) extension","fromName":"Martin Ågren","fromEmail":"martin.agren@gmail.com","sentAt":"2018-09-08T06:29:03Z","receivedAt":"2018-09-08T06:31:35Z","isPatch":true,"sender":{"key":"martin.agren@gmail.com","avatar":null},"body":"On Fri, 7 Sep 2018 at 22:24, Ben Peart <peartben@gmail.com> wrote:\n> > Ben Peart <benpeart@microsoft.com> writes:\n\n> >> - 160-bit SHA-1 over the extension types and their sizes (but not\n> >> their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n> >> long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n> >> then the hash would be:\n\n> The purpose of the SHA isn't to detect disk corruption (we already have\n> a SHA for the entire index that can serve that purpose) but to help\n> ensure that this was actually a valid EOIE extension and not a lucky\n> random set of bytes. [...]\n\n> >> +#define EOIE_SIZE 24 /* <4-byte offset> + <20-byte hash> */\n\n> >> +    the_hash_algo->init_fn(&c);\n> >> +    while (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n[...]\n> >> +    the_hash_algo->final_fn(hash, &c);\n> >> +    if (hashcmp(hash, (unsigned char *)index))\n> >> +            return 0;\n> >> +\n> >> +    /* Validate that the extension offsets returned us back to the eoie extension. */\n> >> +    if (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n> >> +            return 0;\n\nBesides the issue you and Junio discussed with \"should we document this\nas being SHA-1 or NewHash\" (or \"the hash algo\"), it seems to me that\nthis implementation is living somewhere between using SHA-1 and \"the\nhash algo\". The hashing uses `the_hash_algo`, but the hash size is\nhardcoded at 20 bytes.\n\nMaybe it all works out, e.g., so that when someone (brian) merges a\nNewHash and runs the testsuite, this will fail consistently and in a\nsafe way. But I wonder if it would be too hard to avoid the hardcoded 24\nalready now.\n\nMartin\n"},{"id":"357688","messageId":"CACsJy8DF1YgXNixLxhQXkTyGdmzhF=GWyJR2_1-BxmmO0wc+2w@mail.gmail.com","threadId":"49204","inReplyTo":"xmqq5zzhxlxm.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 0/4] read-cache: speed up index load through parallelization","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-08T13:18:47Z","receivedAt":"2018-09-08T13:19:16Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Sep 7, 2018 at 7:21 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Ben Peart <benpeart@microsoft.com> writes:\n>\n> > On further investigation with the previous patch, I noticed that my test\n> > repos didn't contain the cache tree extension in their index. After doing a\n> > commit to ensure they existed, I realized that in some instances, the time\n> > to load the cache tree exceeded the time to load all the cache entries in\n> > parallel.  Because the thread to read the cache tree was started last (due\n> > to having to parse through all the cache entries first) we weren't always\n> > getting optimal performance.\n> >\n> > To better optimize for this case, I decided to write the EOIE extension\n> > as suggested by Junio [1] in response to my earlier multithreading patch\n> > series [2].  This enables me to spin up the thread to load the extensions\n> > earlier as it no longer has to parse through all the cache entries first.\n>\n> Hmph. I kinda liked the simplicity of the previous one, but if we\n> need to start reading the extension sections sooner by eliminating\n> the overhead to scan the cache entries, perhaps we should bite the\n> bullet now.\n\nMy view is slightly different. If we have to optimize might as well\ntry to squeeze the best out of it. Simplicity is already out of the\nwindow at this point (but maintainability remains).\n\n> > The big changes in this iteration are:\n> >\n> > - add the EOIE extension\n> > - update the index extension worker thread to start first\n>\n> I guess I'd need to see the actual patch to find this out, but once\n> we rely on a new extension, then we could omit scanning the main\n> index even to partition the work among workers (i.e. like the topic\n> long ago, you can have list of pointers into the main index to help\n> partitioning, plus reset the prefix compression used in v4).  I\n> think you didn't get that far in this round, which is good.  If the\n> gain with EOIE alone (and starting the worker for the extension\n> section early) is large enough without such a pre-computed work\n> partition table, the simplicity of this round may give us a good\n> stopping point.\n\nI suspect the reduced gain in 1M files case compared to 100k files in\n4/4 is because of scanning the index to split work to worker threads.\nSince the index is now larger, the scanning takes more time before we\ncan start worker threads and we gain less from parallelization. I have\nnot experimented to see if this is true or there is something else.\n\n> > This patch conflicts with Duy's patch to remove the double memory copy and\n> > pass in the previous ce instead.  The two will need to be merged/reconciled\n> > once they settle down a bit.\n>\n> Thanks.  I have a feeling that 67922abb (\"read-cache.c: optimize\n> reading index format v4\", 2018-09-02) is already 'next'-worthy\n> and ready to be built on, but I'd prefer to hear from Duy to double\n> check.\n\nYes I think it's good. I ran the entire test suite with v4 just to\ndouble check (and thinking of testing v4 version in travis too).\n-- \nDuy\n"},{"id":"357694","messageId":"ba1c8611-5480-deae-2b45-75fc9943086c@gmail.com","threadId":"49204","inReplyTo":"CAN0heSreAfMsseZcxR75CFDph-n1b8EUNsRhpFsVqxMLc0hvpA@mail.gmail.com","subject":"Re: [PATCH v3 2/4] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-08T14:03:59Z","receivedAt":"2018-09-08T14:04:05Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/8/2018 2:29 AM, Martin Ågren wrote:\n> On Fri, 7 Sep 2018 at 22:24, Ben Peart <peartben@gmail.com> wrote:\n>>> Ben Peart <benpeart@microsoft.com> writes:\n> \n>>>> - 160-bit SHA-1 over the extension types and their sizes (but not\n>>>> their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n>>>> long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n>>>> then the hash would be:\n> \n>> The purpose of the SHA isn't to detect disk corruption (we already have\n>> a SHA for the entire index that can serve that purpose) but to help\n>> ensure that this was actually a valid EOIE extension and not a lucky\n>> random set of bytes. [...]\n> \n>>>> +#define EOIE_SIZE 24 /* <4-byte offset> + <20-byte hash> */\n> \n>>>> +    the_hash_algo->init_fn(&c);\n>>>> +    while (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n> [...]\n>>>> +    the_hash_algo->final_fn(hash, &c);\n>>>> +    if (hashcmp(hash, (unsigned char *)index))\n>>>> +            return 0;\n>>>> +\n>>>> +    /* Validate that the extension offsets returned us back to the eoie extension. */\n>>>> +    if (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n>>>> +            return 0;\n> \n> Besides the issue you and Junio discussed with \"should we document this\n> as being SHA-1 or NewHash\" (or \"the hash algo\"), it seems to me that\n> this implementation is living somewhere between using SHA-1 and \"the\n> hash algo\". The hashing uses `the_hash_algo`, but the hash size is\n> hardcoded at 20 bytes.\n> \n> Maybe it all works out, e.g., so that when someone (brian) merges a\n> NewHash and runs the testsuite, this will fail consistently and in a\n> safe way. But I wonder if it would be too hard to avoid the hardcoded 24\n> already now.\n> \n> Martin\n> \n\nI can certainly change this to be:\n\n#define EOIE_SIZE (4 + GIT_SHA1_RAWSZ)\n\nwhich should (hopefully) make it easier to find this hard coded hash \nlength in the sea of hard coded \"20\" and \"160\" (bits) littered through \nthe codebase.\n"},{"id":"357695","messageId":"447feebe-c99a-dc53-21ae-d33541baf30d@gmail.com","threadId":"49204","inReplyTo":"xmqqbm99vwsy.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 3/4] read-cache: load cache extensions on a worker thread","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-08T14:56:22Z","receivedAt":"2018-09-08T14:56:28Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/7/2018 5:10 PM, Junio C Hamano wrote:\n> Ben Peart <benpeart@microsoft.com> writes:\n> \n>> +struct load_index_extensions\n>> +{\n>> +#ifndef NO_PTHREADS\n>> +\tpthread_t pthread;\n>> +#endif\n>> +\tstruct index_state *istate;\n>> +\tvoid *mmap;\n>> +\tsize_t mmap_size;\n>> +\tunsigned long src_offset;\n> \n> If the file format only allows uint32_t on any platform, perhaps\n> this is better specified as uint32_t?  Or if this is offset into\n> a mmap'ed region of memory, size_t may be more appropriate.\n> \n> Same comment applies to \"extension_offset\" we see below (which in\n> turn means the returned type of read_eoie_extension() function may\n> want to match).\n> \n>> + };\n> \n> Space before '}'??\n> \n>> +\n>> +static void *load_index_extensions(void *_data)\n>> +{\n>> +\tstruct load_index_extensions *p = _data;\n> \n> Perhaps we are being superstitious, but I think our code try to\n> avoid leading underscore when able, i.e.\n> \n> \tload_index_extensions(void *data_)\n> \t{\n> \t\tstruct load_index_extensions *p = data;\n\nThat's what I get for copying code from elsewhere in the source. :-)\n\nstatic void *preload_thread(void *_data)\n{\n\tint nr;\n\tstruct thread_data *p = _data;\n\nsince there isn't any need for the underscore at all, I'll just make it:\n\nstatic void *load_index_extensions(void *data)\n{\n\tstruct load_index_extensions *p = data;\n\n> \n>> +\tunsigned long src_offset = p->src_offset;\n>> +\n>> +\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n>> +\t\t/* After an array of active_nr index entries,\n>> +\t\t * there can be arbitrary number of extended\n>> +\t\t * sections, each of which is prefixed with\n>> +\t\t * extension name (4-byte) and section length\n>> +\t\t * in 4-byte network byte order.\n>> +\t\t */\n>> +\t\tuint32_t extsize;\n>> +\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n>> +\t\textsize = ntohl(extsize);\n> \n> The same \"ntohl(), not get_be32()?\" question as the one for the\n> previous step applies here, too.  I think the answer is \"the\n> original was written that way\" and that is acceptable, but once this\n> series lands, we may want to review the whole file and see if it is\n> worth making them consistent with a separate clean-up patch.\n> \n\nMakes sense, I'll add a cleanup patch to fix the inconsistency and have \nthem use get_be32().\n\n> I think mmap() and munmap() are the only places that wants p->mmap\n> and mmap parameters passed around in various callchains to be of\n> type \"void *\"---I wonder if it is simpler to use \"const char *\"\n> throughout and only cast it to \"void *\" when necessary (I suspect\n> that there is nowhere we need to cast to or from \"void *\" explicitly\n> if we did so---assignment and argument passing would give us an\n> appropriate cast for free)?\n\nSure, I'll add minimizing the casting to the clean up patch.\n\n> \n>> +\t\tif (read_index_extension(p->istate,\n>> +\t\t\t(const char *)p->mmap + src_offset,\n>> +\t\t\t(char *)p->mmap + src_offset + 8,\n>> +\t\t\textsize) < 0) {\n>> +\t\t\tmunmap(p->mmap, p->mmap_size);\n>> +\t\t\tdie(\"index file corrupt\");\n>> +\t\t}\n>> +\t...\n>> @@ -1907,6 +1951,11 @@ ...\n>> ...\n>> +\tp.mmap = mmap;\n>> +\tp.mmap_size = mmap_size;\n>> +\n>> +#ifndef NO_PTHREADS\n>> +\tnr_threads = git_config_get_index_threads();\n>> +\tif (!nr_threads)\n>> +\t\tnr_threads = online_cpus();\n>> +\n>> +\tif (nr_threads >= 2) {\n>> +\t\textension_offset = read_eoie_extension(mmap, mmap_size);\n>> +\t\tif (extension_offset) {\n>> +\t\t\t/* create a thread to load the index extensions */\n>> +\t\t\tp.src_offset = extension_offset;\n>> +\t\t\tif (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n>> +\t\t\t\tdie(_(\"unable to create load_index_extensions_thread\"));\n>> +\t\t}\n>> +\t}\n>> +#endif\n> \n> Makes sense.\n> \n"},{"id":"357703","messageId":"CAN0heSqiMQ0iW-pvsfDrf+NDjhnY+dznWHHL5H2k+c2MBpJp8g@mail.gmail.com","threadId":"49204","inReplyTo":"ba1c8611-5480-deae-2b45-75fc9943086c@gmail.com","subject":"Re: [PATCH v3 2/4] eoie: add End of Index Entry (EOIE) extension","fromName":"Martin Ågren","fromEmail":"martin.agren@gmail.com","sentAt":"2018-09-08T17:08:11Z","receivedAt":"2018-09-08T17:08:25Z","isPatch":true,"sender":{"key":"martin.agren@gmail.com","avatar":null},"body":"On Sat, 8 Sep 2018 at 16:04, Ben Peart <peartben@gmail.com> wrote:\n> On 9/8/2018 2:29 AM, Martin Ågren wrote:\n> > Maybe it all works out, e.g., so that when someone (brian) merges a\n> > NewHash and runs the testsuite, this will fail consistently and in a\n> > safe way. But I wonder if it would be too hard to avoid the hardcoded 24\n> > already now.\n>\n> I can certainly change this to be:\n>\n> #define EOIE_SIZE (4 + GIT_SHA1_RAWSZ)\n>\n> which should (hopefully) make it easier to find this hard coded hash\n> length in the sea of hard coded \"20\" and \"160\" (bits) littered through\n> the codebase.\n\nYeah, that seems more grep-friendly.\n\nMartin\n"},{"id":"357923","messageId":"20180911232615.35904-1-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180823154053.20212-1-benpeart@microsoft.com","subject":"[PATCH v4 0/5] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-11T23:26:34Z","receivedAt":"2018-09-11T23:26:40Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This version of the patch merges in Duy's work to speed up index v4 decoding.\nI had to massage it a bit to get it to work with the multi-threading but its\nstill largely his code. It helps a little (3%-4%) when the cache entry thread(s)\ntake the longest and not when the index extensions loading is the long thread.\n\nI also added a minor cleanup patch to minimize the casting required when\nworking with the memory mapped index and other minor changes based on the\nfeedback received.\n\nBase Ref: v2.19.0\nWeb-Diff: https://github.com/benpeart/git/commit/9d31d5fb20\nCheckout: git fetch https://github.com/benpeart/git read-index-multithread-v4 && git checkout 9d31d5fb20\n\n\n### Patches\n\nBen Peart (4):\n  eoie: add End of Index Entry (EOIE) extension\n  read-cache: load cache extensions on a worker thread\n  read-cache: speed up index load through parallelization\n  read-cache: clean up casting and byte decoding\n\nNguyễn Thái Ngọc Duy (1):\n  read-cache.c: optimize reading index format v4\n\n Documentation/config.txt                 |   6 +\n Documentation/technical/index-format.txt |  23 +\n config.c                                 |  18 +\n config.h                                 |   1 +\n read-cache.c                             | 581 +++++++++++++++++++----\n 5 files changed, 531 insertions(+), 98 deletions(-)\n\n\nbase-commit: 1d4361b0f344188ab5eec6dcea01f61a3a3a1670\n-- \n2.18.0.windows.1\n\n\n"},{"id":"357924","messageId":"20180911232615.35904-2-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180911232615.35904-1-benpeart@microsoft.com","subject":"[PATCH v4 1/5] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-11T23:26:36Z","receivedAt":"2018-09-11T23:26:43Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"The End of Index Entry (EOIE) is used to locate the end of the variable\nlength index entries and the beginning of the extensions. Code can take\nadvantage of this to quickly locate the index extensions without having\nto parse through all of the index entries.\n\nBecause it must be able to be loaded before the variable length cache\nentries and other index extensions, this extension must be written last.\nThe signature for this extension is { 'E', 'O', 'I', 'E' }.\n\nThe extension consists of:\n\n- 32-bit offset to the end of the index entries\n\n- 160-bit SHA-1 over the extension types and their sizes (but not\ntheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\nlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\nthen the hash would be:\n\nSHA-1(\"TREE\" + <binary representation of N> +\n\t\"REUC\" + <binary representation of M>)\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n Documentation/technical/index-format.txt |  23 ++++\n read-cache.c                             | 154 +++++++++++++++++++++--\n 2 files changed, 169 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\nindex db3572626b..6bc2d90f7f 100644\n--- a/Documentation/technical/index-format.txt\n+++ b/Documentation/technical/index-format.txt\n@@ -314,3 +314,26 @@ The remaining data of each directory block is grouped by type:\n \n   - An ewah bitmap, the n-th bit indicates whether the n-th index entry\n     is not CE_FSMONITOR_VALID.\n+\n+== End of Index Entry\n+\n+  The End of Index Entry (EOIE) is used to locate the end of the variable\n+  length index entries and the begining of the extensions. Code can take\n+  advantage of this to quickly locate the index extensions without having\n+  to parse through all of the index entries.\n+\n+  Because it must be able to be loaded before the variable length cache\n+  entries and other index extensions, this extension must be written last.\n+  The signature for this extension is { 'E', 'O', 'I', 'E' }.\n+\n+  The extension consists of:\n+\n+  - 32-bit offset to the end of the index entries\n+\n+  - 160-bit SHA-1 over the extension types and their sizes (but not\n+\ttheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\tlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\tthen the hash would be:\n+\n+\tSHA-1(\"TREE\" + <binary representation of N> +\n+\t\t\"REUC\" + <binary representation of M>)\ndiff --git a/read-cache.c b/read-cache.c\nindex 7b1354d759..2abac0a7a2 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -43,6 +43,7 @@\n #define CACHE_EXT_LINK 0x6c696e6b\t  /* \"link\" */\n #define CACHE_EXT_UNTRACKED 0x554E5452\t  /* \"UNTR\" */\n #define CACHE_EXT_FSMONITOR 0x46534D4E\t  /* \"FSMN\" */\n+#define CACHE_EXT_ENDOFINDEXENTRIES 0x454F4945\t/* \"EOIE\" */\n \n /* changes that can be kept in $GIT_DIR/index (basically all extensions) */\n #define EXTMASK (RESOLVE_UNDO_CHANGED | CACHE_TREE_CHANGED | \\\n@@ -1693,6 +1694,9 @@ static int read_index_extension(struct index_state *istate,\n \tcase CACHE_EXT_FSMONITOR:\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n+\tcase CACHE_EXT_ENDOFINDEXENTRIES:\n+\t\t/* already handled in do_read_index() */\n+\t\tbreak;\n \tdefault:\n \t\tif (*ext < 'A' || 'Z' < *ext)\n \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n@@ -1889,6 +1893,11 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+#ifndef NO_PTHREADS\n+static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n+#endif\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -2198,11 +2207,15 @@ static int ce_write(git_hash_ctx *context, int fd, void *data, unsigned int len)\n \treturn 0;\n }\n \n-static int write_index_ext_header(git_hash_ctx *context, int fd,\n-\t\t\t\t  unsigned int ext, unsigned int sz)\n+static int write_index_ext_header(git_hash_ctx *context, git_hash_ctx *eoie_context,\n+\t\t\t\t  int fd, unsigned int ext, unsigned int sz)\n {\n \text = htonl(ext);\n \tsz = htonl(sz);\n+\tif (eoie_context) {\n+\t\tthe_hash_algo->update_fn(eoie_context, &ext, 4);\n+\t\tthe_hash_algo->update_fn(eoie_context, &sz, 4);\n+\t}\n \treturn ((ce_write(context, fd, &ext, 4) < 0) ||\n \t\t(ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n }\n@@ -2445,7 +2458,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n {\n \tuint64_t start = getnanotime();\n \tint newfd = tempfile->fd;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, eoie_c;\n \tstruct cache_header hdr;\n \tint i, err = 0, removed, extended, hdr_version;\n \tstruct cache_entry **cache = istate->cache;\n@@ -2454,6 +2467,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct ondisk_cache_entry_extended ondisk;\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n+\tunsigned long offset;\n \n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n@@ -2520,11 +2534,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\treturn err;\n \n \t/* Write extension data here */\n+\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\tthe_hash_algo->init_fn(&eoie_c);\n \tif (!strip_extensions && istate->split_index) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\terr = write_link_extension(&sb, istate) < 0 ||\n-\t\t\twrite_index_ext_header(&c, newfd, CACHE_EXT_LINK,\n+\t\t\twrite_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_LINK,\n \t\t\t\t\t       sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2535,7 +2551,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tcache_tree_write(&sb, istate->cache_tree);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_TREE, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_TREE, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2545,7 +2561,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tresolve_undo_write(&sb, istate->resolve_undo);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_RESOLVE_UNDO,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_RESOLVE_UNDO,\n \t\t\t\t\t     sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2556,7 +2572,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_untracked_extension(&sb, istate->untracked);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_UNTRACKED,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_UNTRACKED,\n \t\t\t\t\t     sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2567,7 +2583,23 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_fsmonitor_extension(&sb, istate);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+\n+\t/*\n+\t * CACHE_EXT_ENDOFINDEXENTRIES must be written as the last entry before the SHA1\n+\t * so that it can be found and processed before all the index entries are\n+\t * read.\n+\t */\n+\tif (!strip_extensions && offset) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\twrite_eoie_extension(&sb, &eoie_c, offset);\n+\t\terr = write_index_ext_header(&c, NULL, newfd, CACHE_EXT_ENDOFINDEXENTRIES, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2978,3 +3010,109 @@ int should_validate_cache_entries(void)\n \n \treturn validate_index_cache_entries;\n }\n+\n+#define EOIE_SIZE (4 + GIT_SHA1_RAWSZ) /* <4-byte offset> + <20-byte hash> */\n+#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n+\n+#ifndef NO_PTHREADS\n+static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n+{\n+\t/*\n+\t * The end of index entries (EOIE) extension is guaranteed to be last\n+\t * so that it can be found by scanning backwards from the EOF.\n+\t *\n+\t * \"EOIE\"\n+\t * <4-byte length>\n+\t * <4-byte offset>\n+\t * <20-byte hash>\n+\t */\n+\tconst char *mmap = mmap_;\n+\tconst char *index, *eoie;\n+\tuint32_t extsize;\n+\tunsigned long offset, src_offset;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\tgit_hash_ctx c;\n+\n+\t/* ensure we have an index big enough to contain an EOIE extension */\n+\tif (mmap_size < sizeof(struct cache_header) + EOIE_SIZE_WITH_HEADER + the_hash_algo->rawsz)\n+\t\treturn 0;\n+\n+\t/* validate the extension signature */\n+\tindex = eoie = mmap + mmap_size - EOIE_SIZE_WITH_HEADER - the_hash_algo->rawsz;\n+\tif (CACHE_EXT(index) != CACHE_EXT_ENDOFINDEXENTRIES)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/* validate the extension size */\n+\textsize = get_be32(index);\n+\tif (extsize != EOIE_SIZE)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * Validate the offset we're going to look for the first extension\n+\t * signature is after the index header and before the eoie extension.\n+\t */\n+\toffset = get_be32(index);\n+\tif (mmap + offset < mmap + sizeof(struct cache_header))\n+\t\treturn 0;\n+\tif (mmap + offset >= eoie)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * The hash is computed over extension types and their sizes (but not\n+\t * their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\t * long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\t * then the hash would be:\n+\t *\n+\t * SHA-1(\"TREE\" + <binary representation of N> +\n+\t *               \"REUC\" + <binary representation of M>)\n+\t */\n+\tsrc_offset = offset;\n+\tthe_hash_algo->init_fn(&c);\n+\twhile (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\n+\t\t/* verify the extension size isn't so large it will wrap around */\n+\t\tif (src_offset + 8 + extsize < src_offset)\n+\t\t\treturn 0;\n+\n+\t\tthe_hash_algo->update_fn(&c, mmap + src_offset, 8);\n+\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\tthe_hash_algo->final_fn(hash, &c);\n+\tif (hashcmp(hash, (const unsigned char *)index))\n+\t\treturn 0;\n+\n+\t/* Validate that the extension offsets returned us back to the eoie extension. */\n+\tif (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n+\t\treturn 0;\n+\n+\treturn offset;\n+}\n+#endif\n+\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset)\n+{\n+\tuint32_t buffer;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\n+\t/* offset */\n+\tput_be32(&buffer, offset);\n+\tstrbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+\t/* hash */\n+\tthe_hash_algo->final_fn(hash, eoie_context);\n+\tstrbuf_add(sb, hash, the_hash_algo->rawsz);\n+}\n-- \n2.18.0.windows.1\n\n"},{"id":"357925","messageId":"20180911232615.35904-3-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180911232615.35904-1-benpeart@microsoft.com","subject":"[PATCH v4 2/5] read-cache: load cache extensions on a worker thread","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-11T23:26:37Z","receivedAt":"2018-09-11T23:26:44Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by loading\nthe cache extensions on a worker thread in parallel with loading the cache\nentries.\n\nIn some cases, loading the extensions takes longer than loading the\ncache entries so this patch utilizes the new EOIE to start the thread to\nload the extensions before loading all the cache entries in parallel.\n\nThis is possible because the current extensions don't access the cache\nentries in the index_state structure so are OK that they don't all exist\nyet.\n\nThe CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\nextensions don't even get a pointer to the index so don't have access to the\ncache entries.\n\nCACHE_EXT_LINK only uses the index_state to initialize the split index.\nCACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\nupdate and dirty flags.\n\nI used p0002-read-cache.sh to generate some performance data:\n\nTest w/100,000 files                Baseline         Parallel Extensions\n---------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 14.08(0.01+0.10) 9.72(0.03+0.06) -31.0%\n\nTest w/1,000,000 files              Baseline         Parallel Extensions\n------------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 202.95(0.01+0.07) 154.14(0.03+0.06) -24.1%\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n Documentation/config.txt |  6 +++\n config.c                 | 18 ++++++++\n config.h                 |  1 +\n read-cache.c             | 94 ++++++++++++++++++++++++++++++++--------\n 4 files changed, 102 insertions(+), 17 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex eb66a11975..d0d8075978 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -2400,6 +2400,12 @@ imap::\n \tThe configuration variables in the 'imap' section are described\n \tin linkgit:git-imap-send[1].\n \n+index.threads::\n+\tSpecifies the number of threads to spawn when loading the index.\n+\tThis is meant to reduce index load time on multiprocessor machines.\n+\tSpecifying 0 or 'true' will cause Git to auto-detect the number of\n+\tCPU's and set the number of threads accordingly. Defaults to 'true'.\n+\n index.version::\n \tSpecify the version with which new index files should be\n \tinitialized.  This does not affect existing repositories.\ndiff --git a/config.c b/config.c\nindex 3461993f0a..f7ebf149fc 100644\n--- a/config.c\n+++ b/config.c\n@@ -2289,6 +2289,24 @@ int git_config_get_fsmonitor(void)\n \treturn 0;\n }\n \n+/*\n+ * You can disable multi-threaded code by setting index.threads\n+ * to 'false' (or 1)\n+ */\n+int git_config_get_index_threads(void)\n+{\n+\tint is_bool, val;\n+\n+\tif (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n+\t\tif (is_bool)\n+\t\t\treturn val ? 0 : 1;\n+\t\telse\n+\t\t\treturn val;\n+\t}\n+\n+\treturn 0; /* auto-detect */\n+}\n+\n NORETURN\n void git_die_config_linenr(const char *key, const char *filename, int linenr)\n {\ndiff --git a/config.h b/config.h\nindex ab46e0165d..a06027e69b 100644\n--- a/config.h\n+++ b/config.h\n@@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n extern int git_config_get_split_index(void);\n extern int git_config_get_max_percent_split_change(void);\n extern int git_config_get_fsmonitor(void);\n+extern int git_config_get_index_threads(void);\n \n /* This dies if the configured or default date is in the future */\n extern int git_config_get_expiry(const char *key, const char **output);\ndiff --git a/read-cache.c b/read-cache.c\nindex 2abac0a7a2..9b97c29f5b 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -23,6 +23,10 @@\n #include \"split-index.h\"\n #include \"utf8.h\"\n #include \"fsmonitor.h\"\n+#ifndef NO_PTHREADS\n+#include <pthread.h>\n+#include <thread-utils.h>\n+#endif\n \n /* Mask for the name length in ce_flags in the on-disk index */\n \n@@ -1898,6 +1902,46 @@ static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n #endif\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n \n+struct load_index_extensions\n+{\n+#ifndef NO_PTHREADS\n+\tpthread_t pthread;\n+#endif\n+\tstruct index_state *istate;\n+\tvoid *mmap;\n+\tsize_t mmap_size;\n+\tunsigned long src_offset;\n+};\n+\n+static void *load_index_extensions(void *_data)\n+{\n+\tstruct load_index_extensions *p = _data;\n+\tunsigned long src_offset = p->src_offset;\n+\n+\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\t\tif (read_index_extension(p->istate,\n+\t\t\t(const char *)p->mmap + src_offset,\n+\t\t\t(char *)p->mmap + src_offset + 8,\n+\t\t\textsize) < 0) {\n+\t\t\tmunmap(p->mmap, p->mmap_size);\n+\t\t\tdie(\"index file corrupt\");\n+\t\t}\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\n+\treturn NULL;\n+}\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -1908,6 +1952,11 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tvoid *mmap;\n \tsize_t mmap_size;\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tstruct load_index_extensions p = { 0 };\n+\tunsigned long extension_offset = 0;\n+#ifndef NO_PTHREADS\n+\tint nr_threads;\n+#endif\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1944,6 +1993,26 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n \tistate->initialized = 1;\n \n+\tp.istate = istate;\n+\tp.mmap = mmap;\n+\tp.mmap_size = mmap_size;\n+\n+#ifndef NO_PTHREADS\n+\tnr_threads = git_config_get_index_threads();\n+\tif (!nr_threads)\n+\t\tnr_threads = online_cpus();\n+\n+\tif (nr_threads >= 2) {\n+\t\textension_offset = read_eoie_extension(mmap, mmap_size);\n+\t\tif (extension_offset) {\n+\t\t\t/* create a thread to load the index extensions */\n+\t\t\tp.src_offset = extension_offset;\n+\t\t\tif (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n+\t\t\t\tdie(_(\"unable to create load_index_extensions_thread\"));\n+\t\t}\n+\t}\n+#endif\n+\n \tif (istate->version == 4) {\n \t\tprevious_name = &previous_name_buf;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n@@ -1970,23 +2039,14 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-\twhile (src_offset <= mmap_size - the_hash_algo->rawsz - 8) {\n-\t\t/* After an array of active_nr index entries,\n-\t\t * there can be arbitrary number of extended\n-\t\t * sections, each of which is prefixed with\n-\t\t * extension name (4-byte) and section length\n-\t\t * in 4-byte network byte order.\n-\t\t */\n-\t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n-\t\tif (read_index_extension(istate,\n-\t\t\t\t\t (const char *) mmap + src_offset,\n-\t\t\t\t\t (char *) mmap + src_offset + 8,\n-\t\t\t\t\t extsize) < 0)\n-\t\t\tgoto unmap;\n-\t\tsrc_offset += 8;\n-\t\tsrc_offset += extsize;\n+\t/* if we created a thread, join it otherwise load the extensions on the primary thread */\n+#ifndef NO_PTHREADS\n+\tif (extension_offset && pthread_join(p.pthread, NULL))\n+\t\tdie(_(\"unable to join load_index_extensions_thread\"));\n+#endif\n+\tif (!extension_offset) {\n+\t\tp.src_offset = src_offset;\n+\t\tload_index_extensions(&p);\n \t}\n \tmunmap(mmap, mmap_size);\n \treturn istate->cache_nr;\n-- \n2.18.0.windows.1\n\n"},{"id":"357926","messageId":"20180911232615.35904-4-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180911232615.35904-1-benpeart@microsoft.com","subject":"[PATCH v4 3/5] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-11T23:26:38Z","receivedAt":"2018-09-11T23:26:48Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by creating\nmultiple threads to divide the work of loading and converting the cache\nentries across all available CPU cores.\n\nIt accomplishes this by having the primary thread loop across the index file\ntracking the offset and (for V4 indexes) expanding the name. It creates a\nthread to process each block of entries as it comes to them.\n\nI used p0002-read-cache.sh to generate some performance data:\n\nTest w/100,000 files                Baseline         Parallel entries\n---------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 14.08(0.01+0.10) 9.72(0.03+0.06) -31.0%\n\nTest w/1,000,000 files              Baseline         Parallel entries\n------------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 202.95(0.01+0.07) 154.14(0.03+0.06) -24.1%\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 240 +++++++++++++++++++++++++++++++++++++++++++++------\n 1 file changed, 212 insertions(+), 28 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 9b97c29f5b..c01d34a71d 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1942,20 +1942,210 @@ static void *load_index_extensions(void *_data)\n \treturn NULL;\n }\n \n+/*\n+ * A helper function that will load the specified range of cache entries\n+ * from the memory mapped file and add them to the given index.\n+ */\n+static unsigned long load_cache_entry_block(struct index_state *istate,\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n+\t\t\tunsigned long start_offset, struct strbuf *previous_name)\n+{\n+\tint i;\n+\tunsigned long src_offset = start_offset;\n+\n+\tfor (i = offset; i < offset + nr; i++) {\n+\t\tstruct ondisk_cache_entry *disk_ce;\n+\t\tstruct cache_entry *ce;\n+\t\tunsigned long consumed;\n+\n+\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tset_index_entry(istate, i, ce);\n+\n+\t\tsrc_offset += consumed;\n+\t}\n+\treturn src_offset - start_offset;\n+}\n+\n+static unsigned long load_all_cache_entries(struct index_state *istate,\n+\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tunsigned long consumed;\n+\n+\tif (istate->version == 4) {\n+\t\tprevious_name = &previous_name_buf;\n+\t\tmem_pool_init(&istate->ce_mem_pool, istate->cache_nr * (sizeof(struct cache_entry) + CACHE_ENTRY_PATH_LENGTH));\n+\t} else {\n+\t\tprevious_name = NULL;\n+\t\tmem_pool_init(&istate->ce_mem_pool, estimate_cache_size(mmap_size, istate->cache_nr));\n+\t}\n+\n+\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, previous_name);\n+\tstrbuf_release(&previous_name_buf);\n+\treturn consumed;\n+}\n+\n+#ifndef NO_PTHREADS\n+\n+/*\n+ * Mostly randomly chosen maximum thread counts: we\n+ * cap the parallelism to online_cpus() threads, and we want\n+ * to have at least 10000 cache entries per thread for it to\n+ * be worth starting a thread.\n+ */\n+#define THREAD_COST\t\t(10000)\n+\n+struct load_cache_entries_thread_data\n+{\n+\tpthread_t pthread;\n+\tstruct index_state *istate;\n+\tstruct mem_pool *ce_mem_pool;\n+\tint offset, nr;\n+\tvoid *mmap;\n+\tunsigned long start_offset;\n+\tstruct strbuf previous_name_buf;\n+\tstruct strbuf *previous_name;\n+\tunsigned long consumed;\t/* return # of bytes in index file processed */\n+};\n+\n+/*\n+ * A thread proc to run the load_cache_entries() computation\n+ * across multiple background threads.\n+ */\n+static void *load_cache_entries_thread(void *_data)\n+{\n+\tstruct load_cache_entries_thread_data *p = _data;\n+\n+\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n+\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n+\treturn NULL;\n+}\n+\n+static unsigned long load_cache_entries_threaded(int nr_threads, struct index_state *istate,\n+\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tstruct load_cache_entries_thread_data *data;\n+\tint ce_per_thread;\n+\tunsigned long consumed;\n+\tint i, thread;\n+\n+\t/* a little sanity checking */\n+\tif (istate->name_hash_initialized)\n+\t\tBUG(\"the name hash isn't thread safe\");\n+\n+\tmem_pool_init(&istate->ce_mem_pool, 0);\n+\tif (istate->version == 4)\n+\t\tprevious_name = &previous_name_buf;\n+\telse\n+\t\tprevious_name = NULL;\n+\n+\tce_per_thread = DIV_ROUND_UP(istate->cache_nr, nr_threads);\n+\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n+\n+\t/*\n+\t * Loop through index entries starting a thread for every ce_per_thread\n+\t * entries. Exit the loop when we've created the final thread (no need\n+\t * to parse the remaining entries.\n+\t */\n+\tconsumed = thread = 0;\n+\tfor (i = 0; ; i++) {\n+\t\tstruct ondisk_cache_entry *ondisk;\n+\t\tconst char *name;\n+\t\tunsigned int flags;\n+\n+\t\t/*\n+\t\t * we've reached the beginning of a block of cache entries,\n+\t\t * kick off a thread to process them\n+\t\t */\n+\t\tif (i % ce_per_thread == 0) {\n+\t\t\tstruct load_cache_entries_thread_data *p = &data[thread];\n+\n+\t\t\tp->istate = istate;\n+\t\t\tp->offset = i;\n+\t\t\tp->nr = ce_per_thread < istate->cache_nr - i ? ce_per_thread : istate->cache_nr - i;\n+\n+\t\t\t/* create a mem_pool for each thread */\n+\t\t\tif (istate->version == 4)\n+\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\t\testimate_cache_size_from_compressed(p->nr));\n+\t\t\telse\n+\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\t\testimate_cache_size(mmap_size, p->nr));\n+\n+\t\t\tp->mmap = mmap;\n+\t\t\tp->start_offset = src_offset;\n+\t\t\tif (previous_name) {\n+\t\t\t\tstrbuf_addbuf(&p->previous_name_buf, previous_name);\n+\t\t\t\tp->previous_name = &p->previous_name_buf;\n+\t\t\t}\n+\n+\t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n+\t\t\t\tdie(\"unable to create load_cache_entries_thread\");\n+\n+\t\t\t/* exit the loop when we've created the last thread */\n+\t\t\tif (++thread == nr_threads)\n+\t\t\t\tbreak;\n+\t\t}\n+\n+\t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\n+\t\t/* On-disk flags are just 16 bits */\n+\t\tflags = get_be16(&ondisk->flags);\n+\n+\t\tif (flags & CE_EXTENDED) {\n+\t\t\tstruct ondisk_cache_entry_extended *ondisk2;\n+\t\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n+\t\t\tname = ondisk2->name;\n+\t\t} else\n+\t\t\tname = ondisk->name;\n+\n+\t\tif (!previous_name) {\n+\t\t\tsize_t len;\n+\n+\t\t\t/* v3 and earlier */\n+\t\t\tlen = flags & CE_NAMEMASK;\n+\t\t\tif (len == CE_NAMEMASK)\n+\t\t\t\tlen = strlen(name);\n+\t\t\tsrc_offset += (flags & CE_EXTENDED) ?\n+\t\t\t\tondisk_cache_entry_extended_size(len) :\n+\t\t\t\tondisk_cache_entry_size(len);\n+\t\t} else\n+\t\t\tsrc_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n+\t}\n+\n+\tfor (i = 0; i < nr_threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = data + i;\n+\t\tif (pthread_join(p->pthread, NULL))\n+\t\t\tdie(\"unable to join load_cache_entries_thread\");\n+\t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n+\t\tstrbuf_release(&p->previous_name_buf);\n+\t\tconsumed += p->consumed;\n+\t}\n+\n+\tfree(data);\n+\tstrbuf_release(&previous_name_buf);\n+\n+\treturn consumed;\n+}\n+\n+#endif\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n-\tint fd, i;\n+\tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tstruct load_index_extensions p = { 0 };\n \tunsigned long extension_offset = 0;\n #ifndef NO_PTHREADS\n-\tint nr_threads;\n+\tint cpus, nr_threads;\n #endif\n \n \tif (istate->initialized)\n@@ -1997,10 +2187,20 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tp.mmap = mmap;\n \tp.mmap_size = mmap_size;\n \n+\tsrc_offset = sizeof(*hdr);\n+\n #ifndef NO_PTHREADS\n \tnr_threads = git_config_get_index_threads();\n-\tif (!nr_threads)\n-\t\tnr_threads = online_cpus();\n+\tif (!nr_threads) {\n+\t\tcpus = online_cpus();\n+\t\tnr_threads = istate->cache_nr / THREAD_COST;\n+\t\tif (nr_threads > cpus)\n+\t\t\tnr_threads = cpus;\n+\t}\n+\n+\t/* enable testing with fewer than default minimum of entries */\n+\tif (istate->cache_nr > 1 && nr_threads < 3 && git_env_bool(\"GIT_INDEX_THREADS_TEST\", 0))\n+\t\tnr_threads = 3;\n \n \tif (nr_threads >= 2) {\n \t\textension_offset = read_eoie_extension(mmap, mmap_size);\n@@ -2009,33 +2209,17 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\t\tp.src_offset = extension_offset;\n \t\t\tif (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n \t\t\t\tdie(_(\"unable to create load_index_extensions_thread\"));\n+\t\t\tnr_threads--;\n \t\t}\n \t}\n+\tif (nr_threads >= 2)\n+\t\tsrc_offset += load_cache_entries_threaded(nr_threads, istate, mmap, mmap_size, src_offset);\n+\telse\n+\t\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+#else\n+\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n #endif\n \n-\tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n-\t} else {\n-\t\tprevious_name = NULL;\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n-\t}\n-\n-\tsrc_offset = sizeof(*hdr);\n-\tfor (i = 0; i < istate->cache_nr; i++) {\n-\t\tstruct ondisk_cache_entry *disk_ce;\n-\t\tstruct cache_entry *ce;\n-\t\tunsigned long consumed;\n-\n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_name);\n-\t\tset_index_entry(istate, i, ce);\n-\n-\t\tsrc_offset += consumed;\n-\t}\n-\tstrbuf_release(&previous_name_buf);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-- \n2.18.0.windows.1\n\n"},{"id":"357927","messageId":"20180911232615.35904-5-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180911232615.35904-1-benpeart@microsoft.com","subject":"[PATCH v4 4/5] read-cache.c: optimize reading index format v4","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-11T23:26:40Z","receivedAt":"2018-09-11T23:26:50Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\nIndex format v4 requires some more computation to assemble a path\nbased on a previous one. The current code is not very efficient\nbecause\n\n - it doubles memory copy, we assemble the final path in a temporary\n   first before putting it back to a cache_entry\n\n - strbuf_remove() in expand_name_field() is not exactly a good fit\n   for stripping a part at the end, _setlen() would do the same job\n   and is much cheaper.\n\n - the open-coded loop to find the end of the string in\n   expand_name_field() can't beat an optimized strlen()\n\nThis patch avoids the temporary buffer and writes directly to the new\ncache_entry, which addresses the first two points. The last point\ncould also be avoided if the total string length fits in the first 12\nbits of ce_flags, if not we fall back to strlen().\n\nRunning \"test-tool read-cache 100\" on webkit.git (275k files), reading\nv2 only takes 4.226 seconds, while v4 takes 5.711 seconds, 35% more\ntime. The patch reduces read time on v4 to 4.319 seconds.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Ben Peart <benpeart@microsoft.com>\n---\n read-cache.c | 136 +++++++++++++++++++++++++++------------------------\n 1 file changed, 71 insertions(+), 65 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex c01d34a71d..d21ccb5e67 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1721,33 +1721,6 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *cache_entry_from_ondisk(struct mem_pool *mem_pool,\n-\t\t\t\t\t\t   struct ondisk_cache_entry *ondisk,\n-\t\t\t\t\t\t   unsigned int flags,\n-\t\t\t\t\t\t   const char *name,\n-\t\t\t\t\t\t   size_t len)\n-{\n-\tstruct cache_entry *ce = mem_pool__ce_alloc(mem_pool, len);\n-\n-\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n-\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n-\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n-\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n-\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n-\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n-\tce->ce_mode  = get_be32(&ondisk->mode);\n-\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n-\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n-\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n-\tce->ce_flags = flags & ~CE_NAMEMASK;\n-\tce->ce_namelen = len;\n-\tce->index = 0;\n-\thashcpy(ce->oid.hash, ondisk->sha1);\n-\tmemcpy(ce->name, name, len);\n-\tce->name[len] = '\\0';\n-\treturn ce;\n-}\n-\n /*\n  * Adjacent cache entries tend to share the leading paths, so it makes\n  * sense to only store the differences in later entries.  In the v4\n@@ -1762,22 +1735,24 @@ static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n \n \tif (name->len < len)\n \t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n+\tstrbuf_setlen(name, name->len - len);\n+\tep = cp + strlen((const char *)cp);\n \tstrbuf_add(name, cp, ep - cp);\n \treturn (const char *)ep + 1 - cp_;\n }\n \n-static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n+static struct cache_entry *create_from_disk(struct mem_pool *ce_mem_pool,\n+\t\t\t\t\t    unsigned int version,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n-\t\t\t\t\t    struct strbuf *previous_name)\n+\t\t\t\t\t    const struct cache_entry *previous_ce)\n {\n \tstruct cache_entry *ce;\n \tsize_t len;\n \tconst char *name;\n \tunsigned int flags;\n+\tsize_t copy_len;\n+\tint expand_name_field = version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1797,21 +1772,54 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \telse\n \t\tname = ondisk->name;\n \n-\tif (!previous_name) {\n-\t\t/* v3 and earlier */\n-\t\tif (len == CE_NAMEMASK)\n-\t\t\tlen = strlen(name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags, name, len);\n+\tif (expand_name_field) {\n+\t\tconst unsigned char *cp = (const unsigned char *)name;\n+\t\tsize_t strip_len, previous_len;\n \n-\t\t*ent_size = ondisk_ce_size(ce);\n-\t} else {\n-\t\tunsigned long consumed;\n-\t\tconsumed = expand_name_field(previous_name, name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags,\n-\t\t\t\t\t     previous_name->buf,\n-\t\t\t\t\t     previous_name->len);\n+\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n+\t\tstrip_len = decode_varint(&cp);\n+\t\tif (previous_len < strip_len) {\n+\t\t\tif (previous_ce)\n+\t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n+\t\t\t\t    previous_ce->name);\n+\t\t\telse\n+\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t}\n+\t\tcopy_len = previous_len - strip_len;\n+\t\tname = (const char *)cp;\n+\t}\n \n-\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n+\tif (len == CE_NAMEMASK) {\n+\t\tlen = strlen(name);\n+\t\tif (expand_name_field)\n+\t\t\tlen += copy_len;\n+\t}\n+\n+\tce = mem_pool__ce_alloc(ce_mem_pool, len);\n+\n+\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n+\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n+\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n+\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n+\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n+\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n+\tce->ce_mode  = get_be32(&ondisk->mode);\n+\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n+\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n+\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n+\tce->ce_flags = flags & ~CE_NAMEMASK;\n+\tce->ce_namelen = len;\n+\tce->index = 0;\n+\thashcpy(ce->oid.hash, ondisk->sha1);\n+\n+\tif (expand_name_field) {\n+\t\tif (copy_len)\n+\t\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n+\t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n+\t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n+\t} else {\n+\t\tmemcpy(ce->name, name, len + 1);\n+\t\t*ent_size = ondisk_ce_size(ce);\n \t}\n \treturn ce;\n }\n@@ -1948,7 +1956,7 @@ static void *load_index_extensions(void *_data)\n  */\n static unsigned long load_cache_entry_block(struct index_state *istate,\n \t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n-\t\t\tunsigned long start_offset, struct strbuf *previous_name)\n+\t\t\tunsigned long start_offset, const struct cache_entry *previous_ce)\n {\n \tint i;\n \tunsigned long src_offset = start_offset;\n@@ -1959,10 +1967,11 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n \t\tunsigned long consumed;\n \n \t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tce = create_from_disk(ce_mem_pool, istate->version, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n \t\tsrc_offset += consumed;\n+\t\tprevious_ce = ce;\n \t}\n \treturn src_offset - start_offset;\n }\n@@ -1970,20 +1979,16 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n static unsigned long load_all_cache_entries(struct index_state *istate,\n \t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n {\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tunsigned long consumed;\n \n \tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n \t\tmem_pool_init(&istate->ce_mem_pool, istate->cache_nr * (sizeof(struct cache_entry) + CACHE_ENTRY_PATH_LENGTH));\n \t} else {\n-\t\tprevious_name = NULL;\n \t\tmem_pool_init(&istate->ce_mem_pool, estimate_cache_size(mmap_size, istate->cache_nr));\n \t}\n \n \tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n-\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, previous_name);\n-\tstrbuf_release(&previous_name_buf);\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, NULL);\n \treturn consumed;\n }\n \n@@ -2005,8 +2010,7 @@ struct load_cache_entries_thread_data\n \tint offset, nr;\n \tvoid *mmap;\n \tunsigned long start_offset;\n-\tstruct strbuf previous_name_buf;\n-\tstruct strbuf *previous_name;\n+\tstruct cache_entry *previous_ce;\n \tunsigned long consumed;\t/* return # of bytes in index file processed */\n };\n \n@@ -2019,7 +2023,7 @@ static void *load_cache_entries_thread(void *_data)\n \tstruct load_cache_entries_thread_data *p = _data;\n \n \tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n-\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n+\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_ce);\n \treturn NULL;\n }\n \n@@ -2066,20 +2070,23 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\t\tp->istate = istate;\n \t\t\tp->offset = i;\n \t\t\tp->nr = ce_per_thread < istate->cache_nr - i ? ce_per_thread : istate->cache_nr - i;\n+\t\t\tp->mmap = mmap;\n+\t\t\tp->start_offset = src_offset;\n \n \t\t\t/* create a mem_pool for each thread */\n-\t\t\tif (istate->version == 4)\n+\t\t\tif (istate->version == 4) {\n \t\t\t\tmem_pool_init(&p->ce_mem_pool,\n \t\t\t\t\testimate_cache_size_from_compressed(p->nr));\n-\t\t\telse\n+\n+\t\t\t\t/* create a previous ce entry for this block of cache entries */\n+\t\t\t\tif (previous_name->len) {\n+\t\t\t\t\tp->previous_ce = mem_pool__ce_alloc(p->ce_mem_pool, previous_name->len);\n+\t\t\t\t\tp->previous_ce->ce_namelen = previous_name->len;\n+\t\t\t\t\tmemcpy(p->previous_ce->name, previous_name->buf, previous_name->len);\n+\t\t\t\t}\n+\t\t\t} else {\n \t\t\t\tmem_pool_init(&p->ce_mem_pool,\n \t\t\t\t\testimate_cache_size(mmap_size, p->nr));\n-\n-\t\t\tp->mmap = mmap;\n-\t\t\tp->start_offset = src_offset;\n-\t\t\tif (previous_name) {\n-\t\t\t\tstrbuf_addbuf(&p->previous_name_buf, previous_name);\n-\t\t\t\tp->previous_name = &p->previous_name_buf;\n \t\t\t}\n \n \t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n@@ -2102,7 +2109,7 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\t} else\n \t\t\tname = ondisk->name;\n \n-\t\tif (!previous_name) {\n+\t\tif (istate->version != 4) {\n \t\t\tsize_t len;\n \n \t\t\t/* v3 and earlier */\n@@ -2121,7 +2128,6 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\tif (pthread_join(p->pthread, NULL))\n \t\t\tdie(\"unable to join load_cache_entries_thread\");\n \t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n-\t\tstrbuf_release(&p->previous_name_buf);\n \t\tconsumed += p->consumed;\n \t}\n \n-- \n2.18.0.windows.1\n\n"},{"id":"357928","messageId":"20180911232615.35904-6-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180911232615.35904-1-benpeart@microsoft.com","subject":"[PATCH v4 5/5] read-cache: clean up casting and byte decoding","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-11T23:26:41Z","receivedAt":"2018-09-11T23:26:54Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch does a clean up pass to minimize the casting required to work\nwith the memory mapped index (mmap).\n\nIt also makes the decoding of network byte order more consistent by using\nget_be32() where possible.\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 49 +++++++++++++++++++++++--------------------------\n 1 file changed, 23 insertions(+), 26 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex d21ccb5e67..6220abc491 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1655,7 +1655,7 @@ int verify_index_checksum;\n /* Allow fsck to force verification of the cache entry order. */\n int verify_ce_order;\n \n-static int verify_hdr(struct cache_header *hdr, unsigned long size)\n+static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n {\n \tgit_hash_ctx c;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n@@ -1679,7 +1679,7 @@ static int verify_hdr(struct cache_header *hdr, unsigned long size)\n }\n \n static int read_index_extension(struct index_state *istate,\n-\t\t\t\tconst char *ext, void *data, unsigned long sz)\n+\t\t\t\tconst char *ext, const char *data, unsigned long sz)\n {\n \tswitch (CACHE_EXT(ext)) {\n \tcase CACHE_EXT_TREE:\n@@ -1906,7 +1906,7 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n }\n \n #ifndef NO_PTHREADS\n-static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n+static unsigned long read_eoie_extension(const char *mmap, size_t mmap_size);\n #endif\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n \n@@ -1916,14 +1916,14 @@ struct load_index_extensions\n \tpthread_t pthread;\n #endif\n \tstruct index_state *istate;\n-\tvoid *mmap;\n+\tconst char *mmap;\n \tsize_t mmap_size;\n \tunsigned long src_offset;\n };\n \n-static void *load_index_extensions(void *_data)\n+static void *load_index_extensions(void *data)\n {\n-\tstruct load_index_extensions *p = _data;\n+\tstruct load_index_extensions *p = data;\n \tunsigned long src_offset = p->src_offset;\n \n \twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n@@ -1934,13 +1934,12 @@ static void *load_index_extensions(void *_data)\n \t\t * in 4-byte network byte order.\n \t\t */\n \t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n+\t\textsize = get_be32(p->mmap + src_offset + 4);\n \t\tif (read_index_extension(p->istate,\n-\t\t\t(const char *)p->mmap + src_offset,\n-\t\t\t(char *)p->mmap + src_offset + 8,\n+\t\t\tp->mmap + src_offset,\n+\t\t\tp->mmap + src_offset + 8,\n \t\t\textsize) < 0) {\n-\t\t\tmunmap(p->mmap, p->mmap_size);\n+\t\t\tmunmap((void *)p->mmap, p->mmap_size);\n \t\t\tdie(\"index file corrupt\");\n \t\t}\n \t\tsrc_offset += 8;\n@@ -1955,7 +1954,7 @@ static void *load_index_extensions(void *_data)\n  * from the memory mapped file and add them to the given index.\n  */\n static unsigned long load_cache_entry_block(struct index_state *istate,\n-\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, const char *mmap,\n \t\t\tunsigned long start_offset, const struct cache_entry *previous_ce)\n {\n \tint i;\n@@ -1966,7 +1965,7 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n \t\tstruct cache_entry *ce;\n \t\tunsigned long consumed;\n \n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n \t\tce = create_from_disk(ce_mem_pool, istate->version, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n@@ -1977,7 +1976,7 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n }\n \n static unsigned long load_all_cache_entries(struct index_state *istate,\n-\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+\t\t\tconst char *mmap, size_t mmap_size, unsigned long src_offset)\n {\n \tunsigned long consumed;\n \n@@ -2008,7 +2007,7 @@ struct load_cache_entries_thread_data\n \tstruct index_state *istate;\n \tstruct mem_pool *ce_mem_pool;\n \tint offset, nr;\n-\tvoid *mmap;\n+\tconst char *mmap;\n \tunsigned long start_offset;\n \tstruct cache_entry *previous_ce;\n \tunsigned long consumed;\t/* return # of bytes in index file processed */\n@@ -2028,7 +2027,7 @@ static void *load_cache_entries_thread(void *_data)\n }\n \n static unsigned long load_cache_entries_threaded(int nr_threads, struct index_state *istate,\n-\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+\t\t\tconst char *mmap, size_t mmap_size, unsigned long src_offset)\n {\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tstruct load_cache_entries_thread_data *data;\n@@ -2097,7 +2096,7 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\t\t\tbreak;\n \t\t}\n \n-\t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tondisk = (struct ondisk_cache_entry *)(mmap + src_offset);\n \n \t\t/* On-disk flags are just 16 bits */\n \t\tflags = get_be16(&ondisk->flags);\n@@ -2145,8 +2144,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n-\tstruct cache_header *hdr;\n-\tvoid *mmap;\n+\tconst struct cache_header *hdr;\n+\tconst char *mmap;\n \tsize_t mmap_size;\n \tstruct load_index_extensions p = { 0 };\n \tunsigned long extension_offset = 0;\n@@ -2178,7 +2177,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tdie_errno(\"unable to map index file\");\n \tclose(fd);\n \n-\thdr = mmap;\n+\thdr = (const struct cache_header *)mmap;\n \tif (verify_hdr(hdr, mmap_size) < 0)\n \t\tgoto unmap;\n \n@@ -2238,11 +2237,11 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tp.src_offset = src_offset;\n \t\tload_index_extensions(&p);\n \t}\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \treturn istate->cache_nr;\n \n unmap:\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \tdie(\"index file corrupt\");\n }\n \n@@ -3265,7 +3264,7 @@ int should_validate_cache_entries(void)\n #define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n \n #ifndef NO_PTHREADS\n-static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n+static unsigned long read_eoie_extension(const char *mmap, size_t mmap_size)\n {\n \t/*\n \t * The end of index entries (EOIE) extension is guaranteed to be last\n@@ -3276,7 +3275,6 @@ static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n \t * <4-byte offset>\n \t * <20-byte hash>\n \t */\n-\tconst char *mmap = mmap_;\n \tconst char *index, *eoie;\n \tuint32_t extsize;\n \tunsigned long offset, src_offset;\n@@ -3329,8 +3327,7 @@ static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n \t\t * in 4-byte network byte order.\n \t\t */\n \t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n+\t\textsize = get_be32(mmap + src_offset + 4);\n \n \t\t/* verify the extension size isn't so large it will wrap around */\n \t\tif (src_offset + 8 + extsize < src_offset)\n-- \n2.18.0.windows.1\n\n"},{"id":"357969","messageId":"ad06ba0f-b436-ff6a-8ae7-21054b9c5b6c@gmail.com","threadId":"49204","inReplyTo":"20180911232615.35904-1-benpeart@microsoft.com","subject":"Re: [PATCH v4 0/5] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-12T14:34:06Z","receivedAt":"2018-09-12T14:34:13Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/11/2018 7:26 PM, Ben Peart wrote:\n> This version of the patch merges in Duy's work to speed up index v4 decoding.\n> I had to massage it a bit to get it to work with the multi-threading but its\n> still largely his code. It helps a little (3%-4%) when the cache entry thread(s)\n> take the longest and not when the index extensions loading is the long thread.\n> \n> I also added a minor cleanup patch to minimize the casting required when\n> working with the memory mapped index and other minor changes based on the\n> feedback received.\n> \n> Base Ref: v2.19.0\n> Web-Diff: https://github.com/benpeart/git/commit/9d31d5fb20\n> Checkout: git fetch https://github.com/benpeart/git read-index-multithread-v4 && git checkout 9d31d5fb20\n> \n> \n\nA bad merge (mistake on my part, not a bug) means this is missing some \nof the changes from V3.  Please ignore, I'll send an updated series to \naddress it.\n\n> ### Patches\n> \n> Ben Peart (4):\n>    eoie: add End of Index Entry (EOIE) extension\n>    read-cache: load cache extensions on a worker thread\n>    read-cache: speed up index load through parallelization\n>    read-cache: clean up casting and byte decoding\n> \n> Nguyễn Thái Ngọc Duy (1):\n>    read-cache.c: optimize reading index format v4\n> \n>   Documentation/config.txt                 |   6 +\n>   Documentation/technical/index-format.txt |  23 +\n>   config.c                                 |  18 +\n>   config.h                                 |   1 +\n>   read-cache.c                             | 581 +++++++++++++++++++----\n>   5 files changed, 531 insertions(+), 98 deletions(-)\n> \n> \n> base-commit: 1d4361b0f344188ab5eec6dcea01f61a3a3a1670\n> \n"},{"id":"357976","messageId":"20180912161832.55324-1-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180823154053.20212-1-benpeart@microsoft.com","subject":"[PATCH v5 0/5] read-cache: speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-12T16:18:50Z","receivedAt":"2018-09-12T16:19:15Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This version of the patch merges in Duy's work to speed up index v4 decoding.\nI had to massage it a bit to get it to work with the multi-threading but it is\nstill largely his code. I also responded to Junio's feedback on initializing\ncopy_len to avoid compiler warnings.\n\nI also added a minor cleanup patch to minimize the casting required when\nworking with the memory mapped index and other minor changes based on the\nfeedback received.\n\nBase Ref: master\nWeb-Diff: https://github.com/benpeart/git/commit/dcf62005f8\nCheckout: git fetch https://github.com/benpeart/git read-index-multithread-v5 && git checkout dcf62005f8\n\n\n### Interdiff (v3..v5):\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 8537a55750..c05e887fc9 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1655,7 +1655,7 @@ int verify_index_checksum;\n /* Allow fsck to force verification of the cache entry order. */\n int verify_ce_order;\n \n-static int verify_hdr(struct cache_header *hdr, unsigned long size)\n+static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n {\n \tgit_hash_ctx c;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n@@ -1679,7 +1679,7 @@ static int verify_hdr(struct cache_header *hdr, unsigned long size)\n }\n \n static int read_index_extension(struct index_state *istate,\n-\t\t\t\tconst char *ext, void *data, unsigned long sz)\n+\t\t\t\tconst char *ext, const char *data, unsigned long sz)\n {\n \tswitch (CACHE_EXT(ext)) {\n \tcase CACHE_EXT_TREE:\n@@ -1721,33 +1721,6 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *cache_entry_from_ondisk(struct mem_pool *mem_pool,\n-\t\t\t\t\t\t   struct ondisk_cache_entry *ondisk,\n-\t\t\t\t\t\t   unsigned int flags,\n-\t\t\t\t\t\t   const char *name,\n-\t\t\t\t\t\t   size_t len)\n-{\n-\tstruct cache_entry *ce = mem_pool__ce_alloc(mem_pool, len);\n-\n-\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n-\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n-\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n-\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n-\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n-\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n-\tce->ce_mode  = get_be32(&ondisk->mode);\n-\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n-\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n-\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n-\tce->ce_flags = flags & ~CE_NAMEMASK;\n-\tce->ce_namelen = len;\n-\tce->index = 0;\n-\thashcpy(ce->oid.hash, ondisk->sha1);\n-\tmemcpy(ce->name, name, len);\n-\tce->name[len] = '\\0';\n-\treturn ce;\n-}\n-\n /*\n  * Adjacent cache entries tend to share the leading paths, so it makes\n  * sense to only store the differences in later entries.  In the v4\n@@ -1768,15 +1741,18 @@ static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n \treturn (const char *)ep + 1 - cp_;\n }\n \n-static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n+static struct cache_entry *create_from_disk(struct mem_pool *ce_mem_pool,\n+\t\t\t\t\t    unsigned int version,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n-\t\t\t\t\t    struct strbuf *previous_name)\n+\t\t\t\t\t    const struct cache_entry *previous_ce)\n {\n \tstruct cache_entry *ce;\n \tsize_t len;\n \tconst char *name;\n \tunsigned int flags;\n+\tsize_t copy_len = 0;\n+\tint expand_name_field = version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1796,21 +1772,50 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \telse\n \t\tname = ondisk->name;\n \n-\tif (!previous_name) {\n-\t\t/* v3 and earlier */\n+\tif (expand_name_field) {\n+\t\tconst unsigned char *cp = (const unsigned char *)name;\n+\t\tsize_t strip_len, previous_len;\n+\n+\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n+\t\tstrip_len = decode_varint(&cp);\n+\t\tif (previous_len < strip_len) {\n+\t\t\tif (previous_ce)\n+\t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n+\t\t\t\t    previous_ce->name);\n+\t\t\telse\n+\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t}\n+\t\tcopy_len = previous_len - strip_len;\n+\t\tname = (const char *)cp;\n+\t}\n+\n \tif (len == CE_NAMEMASK)\n-\t\t\tlen = strlen(name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags, name, len);\n+\t\tlen = strlen(name) + copy_len;\n \n-\t\t*ent_size = ondisk_ce_size(ce);\n-\t} else {\n-\t\tunsigned long consumed;\n-\t\tconsumed = expand_name_field(previous_name, name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags,\n-\t\t\t\t\t     previous_name->buf,\n-\t\t\t\t\t     previous_name->len);\n+\tce = mem_pool__ce_alloc(ce_mem_pool, len);\n+\n+\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n+\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n+\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n+\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n+\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n+\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n+\tce->ce_mode  = get_be32(&ondisk->mode);\n+\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n+\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n+\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n+\tce->ce_flags = flags & ~CE_NAMEMASK;\n+\tce->ce_namelen = len;\n+\tce->index = 0;\n+\thashcpy(ce->oid.hash, ondisk->sha1);\n \n-\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n+\tif (expand_name_field) {\n+\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n+\t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n+\t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n+\t} else {\n+\t\tmemcpy(ce->name, name, len + 1);\n+\t\t*ent_size = ondisk_ce_size(ce);\n \t}\n \treturn ce;\n }\n@@ -1897,7 +1902,7 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n }\n \n #ifndef NO_PTHREADS\n-static unsigned long read_eoie_extension(void *mmap, size_t mmap_size);\n+static unsigned long read_eoie_extension(const char *mmap, size_t mmap_size);\n #endif\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n \n@@ -1907,14 +1912,14 @@ struct load_index_extensions\n \tpthread_t pthread;\n #endif\n \tstruct index_state *istate;\n-\tvoid *mmap;\n+\tconst char *mmap;\n \tsize_t mmap_size;\n \tunsigned long src_offset;\n };\n \n-static void *load_index_extensions(void *_data)\n+static void *load_index_extensions(void *data)\n {\n-\tstruct load_index_extensions *p = _data;\n+\tstruct load_index_extensions *p = data;\n \tunsigned long src_offset = p->src_offset;\n \n \twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n@@ -1925,13 +1930,12 @@ static void *load_index_extensions(void *_data)\n \t\t * in 4-byte network byte order.\n \t\t */\n \t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n+\t\textsize = get_be32(p->mmap + src_offset + 4);\n \t\tif (read_index_extension(p->istate,\n-\t\t\t(const char *)p->mmap + src_offset,\n-\t\t\t(char *)p->mmap + src_offset + 8,\n+\t\t\tp->mmap + src_offset,\n+\t\t\tp->mmap + src_offset + 8,\n \t\t\textsize) < 0) {\n-\t\t\tmunmap(p->mmap, p->mmap_size);\n+\t\t\tmunmap((void *)p->mmap, p->mmap_size);\n \t\t\tdie(\"index file corrupt\");\n \t\t}\n \t\tsrc_offset += 8;\n@@ -1946,8 +1950,8 @@ static void *load_index_extensions(void *_data)\n  * from the memory mapped file and add them to the given index.\n  */\n static unsigned long load_cache_entry_block(struct index_state *istate,\n-\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n-\t\t\tunsigned long start_offset, struct strbuf *previous_name)\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, const char *mmap,\n+\t\t\tunsigned long start_offset, const struct cache_entry *previous_ce)\n {\n \tint i;\n \tunsigned long src_offset = start_offset;\n@@ -1957,34 +1961,31 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n \t\tstruct cache_entry *ce;\n \t\tunsigned long consumed;\n \n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n+\t\tce = create_from_disk(ce_mem_pool, istate->version, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n \t\tsrc_offset += consumed;\n+\t\tprevious_ce = ce;\n \t}\n \treturn src_offset - start_offset;\n }\n \n static unsigned long load_all_cache_entries(struct index_state *istate,\n-\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+\t\t\tconst char *mmap, size_t mmap_size, unsigned long src_offset)\n {\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tunsigned long consumed;\n \n \tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n \t} else {\n-\t\tprevious_name = NULL;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t\testimate_cache_size(mmap_size, istate->cache_nr));\n \t}\n \n \tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n-\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, previous_name);\n-\tstrbuf_release(&previous_name_buf);\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, NULL);\n \treturn consumed;\n }\n \n@@ -1993,7 +1994,7 @@ static unsigned long load_all_cache_entries(struct index_state *istate,\n /*\n  * Mostly randomly chosen maximum thread counts: we\n  * cap the parallelism to online_cpus() threads, and we want\n- * to have at least 100000 cache entries per thread for it to\n+ * to have at least 10000 cache entries per thread for it to\n  * be worth starting a thread.\n  */\n #define THREAD_COST\t\t(10000)\n@@ -2004,10 +2005,9 @@ struct load_cache_entries_thread_data\n \tstruct index_state *istate;\n \tstruct mem_pool *ce_mem_pool;\n \tint offset, nr;\n-\tvoid *mmap;\n+\tconst char *mmap;\n \tunsigned long start_offset;\n-\tstruct strbuf previous_name_buf;\n-\tstruct strbuf *previous_name;\n+\tstruct cache_entry *previous_ce;\n \tunsigned long consumed;\t/* return # of bytes in index file processed */\n };\n \n@@ -2020,12 +2020,12 @@ static void *load_cache_entries_thread(void *_data)\n \tstruct load_cache_entries_thread_data *p = _data;\n \n \tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n-\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n+\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_ce);\n \treturn NULL;\n }\n \n static unsigned long load_cache_entries_threaded(int nr_threads, struct index_state *istate,\n-\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+\t\t\tconst char *mmap, size_t mmap_size, unsigned long src_offset)\n {\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tstruct load_cache_entries_thread_data *data;\n@@ -2067,20 +2067,23 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\t\tp->istate = istate;\n \t\t\tp->offset = i;\n \t\t\tp->nr = ce_per_thread < istate->cache_nr - i ? ce_per_thread : istate->cache_nr - i;\n+\t\t\tp->mmap = mmap;\n+\t\t\tp->start_offset = src_offset;\n \n \t\t\t/* create a mem_pool for each thread */\n-\t\t\tif (istate->version == 4)\n+\t\t\tif (istate->version == 4) {\n \t\t\t\tmem_pool_init(&p->ce_mem_pool,\n \t\t\t\t\testimate_cache_size_from_compressed(p->nr));\n-\t\t\telse\n+\n+\t\t\t\t/* create a previous ce entry for this block of cache entries */\n+\t\t\t\tif (previous_name->len) {\n+\t\t\t\t\tp->previous_ce = mem_pool__ce_alloc(p->ce_mem_pool, previous_name->len);\n+\t\t\t\t\tp->previous_ce->ce_namelen = previous_name->len;\n+\t\t\t\t\tmemcpy(p->previous_ce->name, previous_name->buf, previous_name->len);\n+\t\t\t\t}\n+\t\t\t} else {\n \t\t\t\tmem_pool_init(&p->ce_mem_pool,\n \t\t\t\t\testimate_cache_size(mmap_size, p->nr));\n-\n-\t\t\tp->mmap = mmap;\n-\t\t\tp->start_offset = src_offset;\n-\t\t\tif (previous_name) {\n-\t\t\t\tstrbuf_addbuf(&p->previous_name_buf, previous_name);\n-\t\t\t\tp->previous_name = &p->previous_name_buf;\n \t\t\t}\n \n \t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n@@ -2091,7 +2094,7 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\t\t\tbreak;\n \t\t}\n \n-\t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tondisk = (struct ondisk_cache_entry *)(mmap + src_offset);\n \n \t\t/* On-disk flags are just 16 bits */\n \t\tflags = get_be16(&ondisk->flags);\n@@ -2103,7 +2106,7 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\t} else\n \t\t\tname = ondisk->name;\n \n-\t\tif (!previous_name) {\n+\t\tif (istate->version != 4) {\n \t\t\tsize_t len;\n \n \t\t\t/* v3 and earlier */\n@@ -2122,7 +2125,6 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\tif (pthread_join(p->pthread, NULL))\n \t\t\tdie(\"unable to join load_cache_entries_thread\");\n \t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n-\t\tstrbuf_release(&p->previous_name_buf);\n \t\tconsumed += p->consumed;\n \t}\n \n@@ -2140,8 +2142,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n-\tstruct cache_header *hdr;\n-\tvoid *mmap;\n+\tconst struct cache_header *hdr;\n+\tconst char *mmap;\n \tsize_t mmap_size;\n \tstruct load_index_extensions p = { 0 };\n \tunsigned long extension_offset = 0;\n@@ -2173,7 +2175,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tdie_errno(\"unable to map index file\");\n \tclose(fd);\n \n-\thdr = mmap;\n+\thdr = (const struct cache_header *)mmap;\n \tif (verify_hdr(hdr, mmap_size) < 0)\n \t\tgoto unmap;\n \n@@ -2233,11 +2235,11 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tp.src_offset = src_offset;\n \t\tload_index_extensions(&p);\n \t}\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \treturn istate->cache_nr;\n \n unmap:\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \tdie(\"index file corrupt\");\n }\n \n@@ -3256,11 +3258,11 @@ int should_validate_cache_entries(void)\n \treturn validate_index_cache_entries;\n }\n \n-#define EOIE_SIZE 24 /* <4-byte offset> + <20-byte hash> */\n+#define EOIE_SIZE (4 + GIT_SHA1_RAWSZ) /* <4-byte offset> + <20-byte hash> */\n #define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n \n #ifndef NO_PTHREADS\n-static unsigned long read_eoie_extension(void *mmap, size_t mmap_size)\n+static unsigned long read_eoie_extension(const char *mmap, size_t mmap_size)\n {\n \t/*\n \t * The end of index entries (EOIE) extension is guaranteed to be last\n@@ -3271,14 +3273,18 @@ static unsigned long read_eoie_extension(void *mmap, size_t mmap_size)\n \t * <4-byte offset>\n \t * <20-byte hash>\n \t */\n-\tconst char *index, *eoie = (const char *)mmap + mmap_size - GIT_SHA1_RAWSZ - EOIE_SIZE_WITH_HEADER;\n+\tconst char *index, *eoie;\n \tuint32_t extsize;\n \tunsigned long offset, src_offset;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n \tgit_hash_ctx c;\n \n+\t/* ensure we have an index big enough to contain an EOIE extension */\n+\tif (mmap_size < sizeof(struct cache_header) + EOIE_SIZE_WITH_HEADER + the_hash_algo->rawsz)\n+\t\treturn 0;\n+\n \t/* validate the extension signature */\n-\tindex = eoie;\n+\tindex = eoie = mmap + mmap_size - EOIE_SIZE_WITH_HEADER - the_hash_algo->rawsz;\n \tif (CACHE_EXT(index) != CACHE_EXT_ENDOFINDEXENTRIES)\n \t\treturn 0;\n \tindex += sizeof(uint32_t);\n@@ -3294,9 +3300,9 @@ static unsigned long read_eoie_extension(void *mmap, size_t mmap_size)\n \t * signature is after the index header and before the eoie extension.\n \t */\n \toffset = get_be32(index);\n-\tif ((const char *)mmap + offset < (const char *)mmap + sizeof(struct cache_header))\n+\tif (mmap + offset < mmap + sizeof(struct cache_header))\n \t\treturn 0;\n-\tif ((const char *)mmap + offset >= eoie)\n+\tif (mmap + offset >= eoie)\n \t\treturn 0;\n \tindex += sizeof(uint32_t);\n \n@@ -3319,20 +3325,19 @@ static unsigned long read_eoie_extension(void *mmap, size_t mmap_size)\n \t\t * in 4-byte network byte order.\n \t\t */\n \t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n+\t\textsize = get_be32(mmap + src_offset + 4);\n \n \t\t/* verify the extension size isn't so large it will wrap around */\n \t\tif (src_offset + 8 + extsize < src_offset)\n \t\t\treturn 0;\n \n-\t\tthe_hash_algo->update_fn(&c, (const char *)mmap + src_offset, 8);\n+\t\tthe_hash_algo->update_fn(&c, mmap + src_offset, 8);\n \n \t\tsrc_offset += 8;\n \t\tsrc_offset += extsize;\n \t}\n \tthe_hash_algo->final_fn(hash, &c);\n-\tif (hashcmp(hash, (unsigned char *)index))\n+\tif (hashcmp(hash, (const unsigned char *)index))\n \t\treturn 0;\n \n \t/* Validate that the extension offsets returned us back to the eoie extension. */\ndiff --git a/t/README b/t/README\nindex 59015f7150..69c695ad8e 100644\n--- a/t/README\n+++ b/t/README\n@@ -326,9 +326,6 @@ valid due to the addition of the EOIE extension.\n \n GIT_TEST_INDEX_THREADS=<boolean> forces multi-threaded loading of\n the index cache entries and extensions for the whole test suite.\n-Currently tests 1, 4-9 in t1700-split-index.sh fail as they hard\n-code SHA values for the index which are no longer valid due to the\n-addition of the EOIE extension.\n \n Naming Tests\n ------------\n\n\n### Patches\n\nBen Peart (4):\n  eoie: add End of Index Entry (EOIE) extension\n  read-cache: load cache extensions on a worker thread\n  read-cache: load cache entries on worker threads\n  read-cache: clean up casting and byte decoding\n\nNguyễn Thái Ngọc Duy (1):\n  read-cache.c: optimize reading index format v4\n\n Documentation/config.txt                 |   6 +\n Documentation/technical/index-format.txt |  23 +\n config.c                                 |  18 +\n config.h                                 |   1 +\n read-cache.c                             | 579 +++++++++++++++++++----\n t/README                                 |   8 +\n t/t1700-split-index.sh                   |   1 +\n 7 files changed, 538 insertions(+), 98 deletions(-)\n\n\nbase-commit: 29d9e3e2c47dd4b5053b0a98c891878d398463e3\n-- \n2.18.0.windows.1\n\n\n"},{"id":"357977","messageId":"20180912161832.55324-3-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180912161832.55324-1-benpeart@microsoft.com","subject":"[PATCH v5 2/5] read-cache: load cache extensions on a worker thread","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-12T16:18:55Z","receivedAt":"2018-09-12T16:19:35Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by loading\nthe cache extensions on a worker thread in parallel with loading the cache\nentries.\n\nIn some cases, loading the extensions takes longer than loading the\ncache entries so this patch utilizes the new EOIE to start the thread to\nload the extensions before loading all the cache entries in parallel.\n\nThis is possible because the current extensions don't access the cache\nentries in the index_state structure so are OK that they don't all exist\nyet.\n\nThe CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\nextensions don't even get a pointer to the index so don't have access to the\ncache entries.\n\nCACHE_EXT_LINK only uses the index_state to initialize the split index.\nCACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\nupdate and dirty flags.\n\nI used p0002-read-cache.sh to generate some performance data:\n\nTest w/100,000 files                Baseline         Parallel Extensions\n---------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 14.08(0.01+0.10) 9.72(0.03+0.06) -31.0%\n\nTest w/1,000,000 files              Baseline         Parallel Extensions\n------------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 202.95(0.01+0.07) 154.14(0.03+0.06) -24.1%\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n Documentation/config.txt |  6 +++\n config.c                 | 18 ++++++++\n config.h                 |  1 +\n read-cache.c             | 94 ++++++++++++++++++++++++++++++++--------\n 4 files changed, 102 insertions(+), 17 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex 1c42364988..79f8296d9c 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -2391,6 +2391,12 @@ imap::\n \tThe configuration variables in the 'imap' section are described\n \tin linkgit:git-imap-send[1].\n \n+index.threads::\n+\tSpecifies the number of threads to spawn when loading the index.\n+\tThis is meant to reduce index load time on multiprocessor machines.\n+\tSpecifying 0 or 'true' will cause Git to auto-detect the number of\n+\tCPU's and set the number of threads accordingly. Defaults to 'true'.\n+\n index.version::\n \tSpecify the version with which new index files should be\n \tinitialized.  This does not affect existing repositories.\ndiff --git a/config.c b/config.c\nindex 9a0b10d4bc..9bd79fb165 100644\n--- a/config.c\n+++ b/config.c\n@@ -2289,6 +2289,24 @@ int git_config_get_fsmonitor(void)\n \treturn 0;\n }\n \n+/*\n+ * You can disable multi-threaded code by setting index.threads\n+ * to 'false' (or 1)\n+ */\n+int git_config_get_index_threads(void)\n+{\n+\tint is_bool, val;\n+\n+\tif (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n+\t\tif (is_bool)\n+\t\t\treturn val ? 0 : 1;\n+\t\telse\n+\t\t\treturn val;\n+\t}\n+\n+\treturn 0; /* auto-detect */\n+}\n+\n NORETURN\n void git_die_config_linenr(const char *key, const char *filename, int linenr)\n {\ndiff --git a/config.h b/config.h\nindex ab46e0165d..a06027e69b 100644\n--- a/config.h\n+++ b/config.h\n@@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n extern int git_config_get_split_index(void);\n extern int git_config_get_max_percent_split_change(void);\n extern int git_config_get_fsmonitor(void);\n+extern int git_config_get_index_threads(void);\n \n /* This dies if the configured or default date is in the future */\n extern int git_config_get_expiry(const char *key, const char **output);\ndiff --git a/read-cache.c b/read-cache.c\nindex 858935f123..b203eebb44 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -23,6 +23,10 @@\n #include \"split-index.h\"\n #include \"utf8.h\"\n #include \"fsmonitor.h\"\n+#ifndef NO_PTHREADS\n+#include <pthread.h>\n+#include <thread-utils.h>\n+#endif\n \n /* Mask for the name length in ce_flags in the on-disk index */\n \n@@ -1898,6 +1902,46 @@ static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n #endif\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n \n+struct load_index_extensions\n+{\n+#ifndef NO_PTHREADS\n+\tpthread_t pthread;\n+#endif\n+\tstruct index_state *istate;\n+\tvoid *mmap;\n+\tsize_t mmap_size;\n+\tunsigned long src_offset;\n+};\n+\n+static void *load_index_extensions(void *_data)\n+{\n+\tstruct load_index_extensions *p = _data;\n+\tunsigned long src_offset = p->src_offset;\n+\n+\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\t\tif (read_index_extension(p->istate,\n+\t\t\t(const char *)p->mmap + src_offset,\n+\t\t\t(char *)p->mmap + src_offset + 8,\n+\t\t\textsize) < 0) {\n+\t\t\tmunmap(p->mmap, p->mmap_size);\n+\t\t\tdie(\"index file corrupt\");\n+\t\t}\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\n+\treturn NULL;\n+}\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -1908,6 +1952,11 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tvoid *mmap;\n \tsize_t mmap_size;\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tstruct load_index_extensions p = { 0 };\n+\tunsigned long extension_offset = 0;\n+#ifndef NO_PTHREADS\n+\tint nr_threads;\n+#endif\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1944,6 +1993,26 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n \tistate->initialized = 1;\n \n+\tp.istate = istate;\n+\tp.mmap = mmap;\n+\tp.mmap_size = mmap_size;\n+\n+#ifndef NO_PTHREADS\n+\tnr_threads = git_config_get_index_threads();\n+\tif (!nr_threads)\n+\t\tnr_threads = online_cpus();\n+\n+\tif (nr_threads >= 2) {\n+\t\textension_offset = read_eoie_extension(mmap, mmap_size);\n+\t\tif (extension_offset) {\n+\t\t\t/* create a thread to load the index extensions */\n+\t\t\tp.src_offset = extension_offset;\n+\t\t\tif (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n+\t\t\t\tdie(_(\"unable to create load_index_extensions_thread\"));\n+\t\t}\n+\t}\n+#endif\n+\n \tif (istate->version == 4) {\n \t\tprevious_name = &previous_name_buf;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n@@ -1970,23 +2039,14 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-\twhile (src_offset <= mmap_size - the_hash_algo->rawsz - 8) {\n-\t\t/* After an array of active_nr index entries,\n-\t\t * there can be arbitrary number of extended\n-\t\t * sections, each of which is prefixed with\n-\t\t * extension name (4-byte) and section length\n-\t\t * in 4-byte network byte order.\n-\t\t */\n-\t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n-\t\tif (read_index_extension(istate,\n-\t\t\t\t\t (const char *) mmap + src_offset,\n-\t\t\t\t\t (char *) mmap + src_offset + 8,\n-\t\t\t\t\t extsize) < 0)\n-\t\t\tgoto unmap;\n-\t\tsrc_offset += 8;\n-\t\tsrc_offset += extsize;\n+\t/* if we created a thread, join it otherwise load the extensions on the primary thread */\n+#ifndef NO_PTHREADS\n+\tif (extension_offset && pthread_join(p.pthread, NULL))\n+\t\tdie(_(\"unable to join load_index_extensions_thread\"));\n+#endif\n+\tif (!extension_offset) {\n+\t\tp.src_offset = src_offset;\n+\t\tload_index_extensions(&p);\n \t}\n \tmunmap(mmap, mmap_size);\n \treturn istate->cache_nr;\n-- \n2.18.0.windows.1\n\n"},{"id":"357978","messageId":"20180912161832.55324-4-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180912161832.55324-1-benpeart@microsoft.com","subject":"[PATCH v5 3/5] read-cache: load cache entries on worker threads","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-12T16:18:56Z","receivedAt":"2018-09-12T16:19:36Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by creating\nmultiple threads to divide the work of loading and converting the cache\nentries across all available CPU cores.\n\nIt accomplishes this by having the primary thread loop across the index file\ntracking the offset and (for V4 indexes) expanding the name. It creates a\nthread to process each block of entries as it comes to them.\n\nI used p0002-read-cache.sh to generate some performance data:\n\nTest w/100,000 files                Baseline         Parallel entries\n---------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 14.08(0.01+0.10) 9.72(0.03+0.06) -31.0%\n\nTest w/1,000,000 files              Baseline         Parallel entries\n------------------------------------------------------------------------------\nread_cache/discard_cache 1000 times 202.95(0.01+0.07) 154.14(0.03+0.06) -24.1%\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 242 +++++++++++++++++++++++++++++++++++++++++++++------\n t/README     |   3 +\n 2 files changed, 217 insertions(+), 28 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex b203eebb44..880f627b4c 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1942,20 +1942,212 @@ static void *load_index_extensions(void *_data)\n \treturn NULL;\n }\n \n+/*\n+ * A helper function that will load the specified range of cache entries\n+ * from the memory mapped file and add them to the given index.\n+ */\n+static unsigned long load_cache_entry_block(struct index_state *istate,\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n+\t\t\tunsigned long start_offset, struct strbuf *previous_name)\n+{\n+\tint i;\n+\tunsigned long src_offset = start_offset;\n+\n+\tfor (i = offset; i < offset + nr; i++) {\n+\t\tstruct ondisk_cache_entry *disk_ce;\n+\t\tstruct cache_entry *ce;\n+\t\tunsigned long consumed;\n+\n+\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tset_index_entry(istate, i, ce);\n+\n+\t\tsrc_offset += consumed;\n+\t}\n+\treturn src_offset - start_offset;\n+}\n+\n+static unsigned long load_all_cache_entries(struct index_state *istate,\n+\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tunsigned long consumed;\n+\n+\tif (istate->version == 4) {\n+\t\tprevious_name = &previous_name_buf;\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n+\t} else {\n+\t\tprevious_name = NULL;\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t\testimate_cache_size(mmap_size, istate->cache_nr));\n+\t}\n+\n+\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, previous_name);\n+\tstrbuf_release(&previous_name_buf);\n+\treturn consumed;\n+}\n+\n+#ifndef NO_PTHREADS\n+\n+/*\n+ * Mostly randomly chosen maximum thread counts: we\n+ * cap the parallelism to online_cpus() threads, and we want\n+ * to have at least 10000 cache entries per thread for it to\n+ * be worth starting a thread.\n+ */\n+#define THREAD_COST\t\t(10000)\n+\n+struct load_cache_entries_thread_data\n+{\n+\tpthread_t pthread;\n+\tstruct index_state *istate;\n+\tstruct mem_pool *ce_mem_pool;\n+\tint offset, nr;\n+\tvoid *mmap;\n+\tunsigned long start_offset;\n+\tstruct strbuf previous_name_buf;\n+\tstruct strbuf *previous_name;\n+\tunsigned long consumed;\t/* return # of bytes in index file processed */\n+};\n+\n+/*\n+ * A thread proc to run the load_cache_entries() computation\n+ * across multiple background threads.\n+ */\n+static void *load_cache_entries_thread(void *_data)\n+{\n+\tstruct load_cache_entries_thread_data *p = _data;\n+\n+\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n+\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n+\treturn NULL;\n+}\n+\n+static unsigned long load_cache_entries_threaded(int nr_threads, struct index_state *istate,\n+\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tstruct load_cache_entries_thread_data *data;\n+\tint ce_per_thread;\n+\tunsigned long consumed;\n+\tint i, thread;\n+\n+\t/* a little sanity checking */\n+\tif (istate->name_hash_initialized)\n+\t\tBUG(\"the name hash isn't thread safe\");\n+\n+\tmem_pool_init(&istate->ce_mem_pool, 0);\n+\tif (istate->version == 4)\n+\t\tprevious_name = &previous_name_buf;\n+\telse\n+\t\tprevious_name = NULL;\n+\n+\tce_per_thread = DIV_ROUND_UP(istate->cache_nr, nr_threads);\n+\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n+\n+\t/*\n+\t * Loop through index entries starting a thread for every ce_per_thread\n+\t * entries. Exit the loop when we've created the final thread (no need\n+\t * to parse the remaining entries.\n+\t */\n+\tconsumed = thread = 0;\n+\tfor (i = 0; ; i++) {\n+\t\tstruct ondisk_cache_entry *ondisk;\n+\t\tconst char *name;\n+\t\tunsigned int flags;\n+\n+\t\t/*\n+\t\t * we've reached the beginning of a block of cache entries,\n+\t\t * kick off a thread to process them\n+\t\t */\n+\t\tif (i % ce_per_thread == 0) {\n+\t\t\tstruct load_cache_entries_thread_data *p = &data[thread];\n+\n+\t\t\tp->istate = istate;\n+\t\t\tp->offset = i;\n+\t\t\tp->nr = ce_per_thread < istate->cache_nr - i ? ce_per_thread : istate->cache_nr - i;\n+\n+\t\t\t/* create a mem_pool for each thread */\n+\t\t\tif (istate->version == 4)\n+\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\t\testimate_cache_size_from_compressed(p->nr));\n+\t\t\telse\n+\t\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\t\testimate_cache_size(mmap_size, p->nr));\n+\n+\t\t\tp->mmap = mmap;\n+\t\t\tp->start_offset = src_offset;\n+\t\t\tif (previous_name) {\n+\t\t\t\tstrbuf_addbuf(&p->previous_name_buf, previous_name);\n+\t\t\t\tp->previous_name = &p->previous_name_buf;\n+\t\t\t}\n+\n+\t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n+\t\t\t\tdie(\"unable to create load_cache_entries_thread\");\n+\n+\t\t\t/* exit the loop when we've created the last thread */\n+\t\t\tif (++thread == nr_threads)\n+\t\t\t\tbreak;\n+\t\t}\n+\n+\t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\n+\t\t/* On-disk flags are just 16 bits */\n+\t\tflags = get_be16(&ondisk->flags);\n+\n+\t\tif (flags & CE_EXTENDED) {\n+\t\t\tstruct ondisk_cache_entry_extended *ondisk2;\n+\t\t\tondisk2 = (struct ondisk_cache_entry_extended *)ondisk;\n+\t\t\tname = ondisk2->name;\n+\t\t} else\n+\t\t\tname = ondisk->name;\n+\n+\t\tif (!previous_name) {\n+\t\t\tsize_t len;\n+\n+\t\t\t/* v3 and earlier */\n+\t\t\tlen = flags & CE_NAMEMASK;\n+\t\t\tif (len == CE_NAMEMASK)\n+\t\t\t\tlen = strlen(name);\n+\t\t\tsrc_offset += (flags & CE_EXTENDED) ?\n+\t\t\t\tondisk_cache_entry_extended_size(len) :\n+\t\t\t\tondisk_cache_entry_size(len);\n+\t\t} else\n+\t\t\tsrc_offset += (name - ((char *)ondisk)) + expand_name_field(previous_name, name);\n+\t}\n+\n+\tfor (i = 0; i < nr_threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = data + i;\n+\t\tif (pthread_join(p->pthread, NULL))\n+\t\t\tdie(\"unable to join load_cache_entries_thread\");\n+\t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n+\t\tstrbuf_release(&p->previous_name_buf);\n+\t\tconsumed += p->consumed;\n+\t}\n+\n+\tfree(data);\n+\tstrbuf_release(&previous_name_buf);\n+\n+\treturn consumed;\n+}\n+\n+#endif\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n-\tint fd, i;\n+\tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tstruct load_index_extensions p = { 0 };\n \tunsigned long extension_offset = 0;\n #ifndef NO_PTHREADS\n-\tint nr_threads;\n+\tint cpus, nr_threads;\n #endif\n \n \tif (istate->initialized)\n@@ -1997,10 +2189,20 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tp.mmap = mmap;\n \tp.mmap_size = mmap_size;\n \n+\tsrc_offset = sizeof(*hdr);\n+\n #ifndef NO_PTHREADS\n \tnr_threads = git_config_get_index_threads();\n-\tif (!nr_threads)\n-\t\tnr_threads = online_cpus();\n+\tif (!nr_threads) {\n+\t\tcpus = online_cpus();\n+\t\tnr_threads = istate->cache_nr / THREAD_COST;\n+\t\tif (nr_threads > cpus)\n+\t\t\tnr_threads = cpus;\n+\t}\n+\n+\t/* enable testing with fewer than default minimum of entries */\n+\tif (istate->cache_nr > 1 && nr_threads < 3 && git_env_bool(\"GIT_TEST_INDEX_THREADS\", 0))\n+\t\tnr_threads = 3;\n \n \tif (nr_threads >= 2) {\n \t\textension_offset = read_eoie_extension(mmap, mmap_size);\n@@ -2009,33 +2211,17 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\t\tp.src_offset = extension_offset;\n \t\t\tif (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n \t\t\t\tdie(_(\"unable to create load_index_extensions_thread\"));\n+\t\t\tnr_threads--;\n \t\t}\n \t}\n+\tif (nr_threads >= 2)\n+\t\tsrc_offset += load_cache_entries_threaded(nr_threads, istate, mmap, mmap_size, src_offset);\n+\telse\n+\t\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+#else\n+\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n #endif\n \n-\tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n-\t} else {\n-\t\tprevious_name = NULL;\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n-\t}\n-\n-\tsrc_offset = sizeof(*hdr);\n-\tfor (i = 0; i < istate->cache_nr; i++) {\n-\t\tstruct ondisk_cache_entry *disk_ce;\n-\t\tstruct cache_entry *ce;\n-\t\tunsigned long consumed;\n-\n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_name);\n-\t\tset_index_entry(istate, i, ce);\n-\n-\t\tsrc_offset += consumed;\n-\t}\n-\tstrbuf_release(&previous_name_buf);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \ndiff --git a/t/README b/t/README\nindex d8754dd23a..69c695ad8e 100644\n--- a/t/README\n+++ b/t/README\n@@ -324,6 +324,9 @@ This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n as they currently hard code SHA values for the index which are no longer\n valid due to the addition of the EOIE extension.\n \n+GIT_TEST_INDEX_THREADS=<boolean> forces multi-threaded loading of\n+the index cache entries and extensions for the whole test suite.\n+\n Naming Tests\n ------------\n \n-- \n2.18.0.windows.1\n\n"},{"id":"357979","messageId":"20180912161832.55324-5-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180912161832.55324-1-benpeart@microsoft.com","subject":"[PATCH v5 4/5] read-cache.c: optimize reading index format v4","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-12T16:18:57Z","receivedAt":"2018-09-12T16:19:41Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\nIndex format v4 requires some more computation to assemble a path\nbased on a previous one. The current code is not very efficient\nbecause\n\n - it doubles memory copy, we assemble the final path in a temporary\n   first before putting it back to a cache_entry\n\n - strbuf_remove() in expand_name_field() is not exactly a good fit\n   for stripping a part at the end, _setlen() would do the same job\n   and is much cheaper.\n\n - the open-coded loop to find the end of the string in\n   expand_name_field() can't beat an optimized strlen()\n\nThis patch avoids the temporary buffer and writes directly to the new\ncache_entry, which addresses the first two points. The last point\ncould also be avoided if the total string length fits in the first 12\nbits of ce_flags, if not we fall back to strlen().\n\nRunning \"test-tool read-cache 100\" on webkit.git (275k files), reading\nv2 only takes 4.226 seconds, while v4 takes 5.711 seconds, 35% more\ntime. The patch reduces read time on v4 to 4.319 seconds.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\nSigned-off-by: Ben Peart <benpeart@microsoft.com>\n---\n read-cache.c | 132 ++++++++++++++++++++++++++-------------------------\n 1 file changed, 67 insertions(+), 65 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 880f627b4c..40dc4723b2 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1721,33 +1721,6 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *cache_entry_from_ondisk(struct mem_pool *mem_pool,\n-\t\t\t\t\t\t   struct ondisk_cache_entry *ondisk,\n-\t\t\t\t\t\t   unsigned int flags,\n-\t\t\t\t\t\t   const char *name,\n-\t\t\t\t\t\t   size_t len)\n-{\n-\tstruct cache_entry *ce = mem_pool__ce_alloc(mem_pool, len);\n-\n-\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n-\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n-\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n-\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n-\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n-\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n-\tce->ce_mode  = get_be32(&ondisk->mode);\n-\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n-\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n-\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n-\tce->ce_flags = flags & ~CE_NAMEMASK;\n-\tce->ce_namelen = len;\n-\tce->index = 0;\n-\thashcpy(ce->oid.hash, ondisk->sha1);\n-\tmemcpy(ce->name, name, len);\n-\tce->name[len] = '\\0';\n-\treturn ce;\n-}\n-\n /*\n  * Adjacent cache entries tend to share the leading paths, so it makes\n  * sense to only store the differences in later entries.  In the v4\n@@ -1762,22 +1735,24 @@ static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n \n \tif (name->len < len)\n \t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n+\tstrbuf_setlen(name, name->len - len);\n+\tep = cp + strlen((const char *)cp);\n \tstrbuf_add(name, cp, ep - cp);\n \treturn (const char *)ep + 1 - cp_;\n }\n \n-static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n+static struct cache_entry *create_from_disk(struct mem_pool *ce_mem_pool,\n+\t\t\t\t\t    unsigned int version,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n-\t\t\t\t\t    struct strbuf *previous_name)\n+\t\t\t\t\t    const struct cache_entry *previous_ce)\n {\n \tstruct cache_entry *ce;\n \tsize_t len;\n \tconst char *name;\n \tunsigned int flags;\n+\tsize_t copy_len = 0;\n+\tint expand_name_field = version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1797,21 +1772,50 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \telse\n \t\tname = ondisk->name;\n \n-\tif (!previous_name) {\n-\t\t/* v3 and earlier */\n-\t\tif (len == CE_NAMEMASK)\n-\t\t\tlen = strlen(name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags, name, len);\n+\tif (expand_name_field) {\n+\t\tconst unsigned char *cp = (const unsigned char *)name;\n+\t\tsize_t strip_len, previous_len;\n \n-\t\t*ent_size = ondisk_ce_size(ce);\n-\t} else {\n-\t\tunsigned long consumed;\n-\t\tconsumed = expand_name_field(previous_name, name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags,\n-\t\t\t\t\t     previous_name->buf,\n-\t\t\t\t\t     previous_name->len);\n+\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n+\t\tstrip_len = decode_varint(&cp);\n+\t\tif (previous_len < strip_len) {\n+\t\t\tif (previous_ce)\n+\t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n+\t\t\t\t    previous_ce->name);\n+\t\t\telse\n+\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t}\n+\t\tcopy_len = previous_len - strip_len;\n+\t\tname = (const char *)cp;\n+\t}\n+\n+\tif (len == CE_NAMEMASK)\n+\t\tlen = strlen(name) + copy_len;\n+\n+\tce = mem_pool__ce_alloc(ce_mem_pool, len);\n+\n+\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n+\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n+\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n+\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n+\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n+\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n+\tce->ce_mode  = get_be32(&ondisk->mode);\n+\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n+\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n+\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n+\tce->ce_flags = flags & ~CE_NAMEMASK;\n+\tce->ce_namelen = len;\n+\tce->index = 0;\n+\thashcpy(ce->oid.hash, ondisk->sha1);\n \n-\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n+\tif (expand_name_field) {\n+\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n+\t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n+\t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n+\t} else {\n+\t\tmemcpy(ce->name, name, len + 1);\n+\t\t*ent_size = ondisk_ce_size(ce);\n \t}\n \treturn ce;\n }\n@@ -1948,7 +1952,7 @@ static void *load_index_extensions(void *_data)\n  */\n static unsigned long load_cache_entry_block(struct index_state *istate,\n \t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n-\t\t\tunsigned long start_offset, struct strbuf *previous_name)\n+\t\t\tunsigned long start_offset, const struct cache_entry *previous_ce)\n {\n \tint i;\n \tunsigned long src_offset = start_offset;\n@@ -1959,10 +1963,11 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n \t\tunsigned long consumed;\n \n \t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tce = create_from_disk(ce_mem_pool, istate->version, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n \t\tsrc_offset += consumed;\n+\t\tprevious_ce = ce;\n \t}\n \treturn src_offset - start_offset;\n }\n@@ -1970,22 +1975,18 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n static unsigned long load_all_cache_entries(struct index_state *istate,\n \t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n {\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tunsigned long consumed;\n \n \tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n \t} else {\n-\t\tprevious_name = NULL;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t\testimate_cache_size(mmap_size, istate->cache_nr));\n \t}\n \n \tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n-\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, previous_name);\n-\tstrbuf_release(&previous_name_buf);\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, NULL);\n \treturn consumed;\n }\n \n@@ -2007,8 +2008,7 @@ struct load_cache_entries_thread_data\n \tint offset, nr;\n \tvoid *mmap;\n \tunsigned long start_offset;\n-\tstruct strbuf previous_name_buf;\n-\tstruct strbuf *previous_name;\n+\tstruct cache_entry *previous_ce;\n \tunsigned long consumed;\t/* return # of bytes in index file processed */\n };\n \n@@ -2021,7 +2021,7 @@ static void *load_cache_entries_thread(void *_data)\n \tstruct load_cache_entries_thread_data *p = _data;\n \n \tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n-\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_name);\n+\t\tp->offset, p->nr, p->mmap, p->start_offset, p->previous_ce);\n \treturn NULL;\n }\n \n@@ -2068,20 +2068,23 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\t\tp->istate = istate;\n \t\t\tp->offset = i;\n \t\t\tp->nr = ce_per_thread < istate->cache_nr - i ? ce_per_thread : istate->cache_nr - i;\n+\t\t\tp->mmap = mmap;\n+\t\t\tp->start_offset = src_offset;\n \n \t\t\t/* create a mem_pool for each thread */\n-\t\t\tif (istate->version == 4)\n+\t\t\tif (istate->version == 4) {\n \t\t\t\tmem_pool_init(&p->ce_mem_pool,\n \t\t\t\t\testimate_cache_size_from_compressed(p->nr));\n-\t\t\telse\n+\n+\t\t\t\t/* create a previous ce entry for this block of cache entries */\n+\t\t\t\tif (previous_name->len) {\n+\t\t\t\t\tp->previous_ce = mem_pool__ce_alloc(p->ce_mem_pool, previous_name->len);\n+\t\t\t\t\tp->previous_ce->ce_namelen = previous_name->len;\n+\t\t\t\t\tmemcpy(p->previous_ce->name, previous_name->buf, previous_name->len);\n+\t\t\t\t}\n+\t\t\t} else {\n \t\t\t\tmem_pool_init(&p->ce_mem_pool,\n \t\t\t\t\testimate_cache_size(mmap_size, p->nr));\n-\n-\t\t\tp->mmap = mmap;\n-\t\t\tp->start_offset = src_offset;\n-\t\t\tif (previous_name) {\n-\t\t\t\tstrbuf_addbuf(&p->previous_name_buf, previous_name);\n-\t\t\t\tp->previous_name = &p->previous_name_buf;\n \t\t\t}\n \n \t\t\tif (pthread_create(&p->pthread, NULL, load_cache_entries_thread, p))\n@@ -2104,7 +2107,7 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\t} else\n \t\t\tname = ondisk->name;\n \n-\t\tif (!previous_name) {\n+\t\tif (istate->version != 4) {\n \t\t\tsize_t len;\n \n \t\t\t/* v3 and earlier */\n@@ -2123,7 +2126,6 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\tif (pthread_join(p->pthread, NULL))\n \t\t\tdie(\"unable to join load_cache_entries_thread\");\n \t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n-\t\tstrbuf_release(&p->previous_name_buf);\n \t\tconsumed += p->consumed;\n \t}\n \n-- \n2.18.0.windows.1\n\n"},{"id":"357980","messageId":"20180912161832.55324-6-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180912161832.55324-1-benpeart@microsoft.com","subject":"[PATCH v5 5/5] read-cache: clean up casting and byte decoding","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-12T16:18:59Z","receivedAt":"2018-09-12T16:19:44Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch does a clean up pass to minimize the casting required to work\nwith the memory mapped index (mmap).\n\nIt also makes the decoding of network byte order more consistent by using\nget_be32() where possible.\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 49 +++++++++++++++++++++++--------------------------\n 1 file changed, 23 insertions(+), 26 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 40dc4723b2..c05e887fc9 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1655,7 +1655,7 @@ int verify_index_checksum;\n /* Allow fsck to force verification of the cache entry order. */\n int verify_ce_order;\n \n-static int verify_hdr(struct cache_header *hdr, unsigned long size)\n+static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n {\n \tgit_hash_ctx c;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n@@ -1679,7 +1679,7 @@ static int verify_hdr(struct cache_header *hdr, unsigned long size)\n }\n \n static int read_index_extension(struct index_state *istate,\n-\t\t\t\tconst char *ext, void *data, unsigned long sz)\n+\t\t\t\tconst char *ext, const char *data, unsigned long sz)\n {\n \tswitch (CACHE_EXT(ext)) {\n \tcase CACHE_EXT_TREE:\n@@ -1902,7 +1902,7 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n }\n \n #ifndef NO_PTHREADS\n-static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n+static unsigned long read_eoie_extension(const char *mmap, size_t mmap_size);\n #endif\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n \n@@ -1912,14 +1912,14 @@ struct load_index_extensions\n \tpthread_t pthread;\n #endif\n \tstruct index_state *istate;\n-\tvoid *mmap;\n+\tconst char *mmap;\n \tsize_t mmap_size;\n \tunsigned long src_offset;\n };\n \n-static void *load_index_extensions(void *_data)\n+static void *load_index_extensions(void *data)\n {\n-\tstruct load_index_extensions *p = _data;\n+\tstruct load_index_extensions *p = data;\n \tunsigned long src_offset = p->src_offset;\n \n \twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n@@ -1930,13 +1930,12 @@ static void *load_index_extensions(void *_data)\n \t\t * in 4-byte network byte order.\n \t\t */\n \t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n+\t\textsize = get_be32(p->mmap + src_offset + 4);\n \t\tif (read_index_extension(p->istate,\n-\t\t\t(const char *)p->mmap + src_offset,\n-\t\t\t(char *)p->mmap + src_offset + 8,\n+\t\t\tp->mmap + src_offset,\n+\t\t\tp->mmap + src_offset + 8,\n \t\t\textsize) < 0) {\n-\t\t\tmunmap(p->mmap, p->mmap_size);\n+\t\t\tmunmap((void *)p->mmap, p->mmap_size);\n \t\t\tdie(\"index file corrupt\");\n \t\t}\n \t\tsrc_offset += 8;\n@@ -1951,7 +1950,7 @@ static void *load_index_extensions(void *_data)\n  * from the memory mapped file and add them to the given index.\n  */\n static unsigned long load_cache_entry_block(struct index_state *istate,\n-\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, void *mmap,\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, const char *mmap,\n \t\t\tunsigned long start_offset, const struct cache_entry *previous_ce)\n {\n \tint i;\n@@ -1962,7 +1961,7 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n \t\tstruct cache_entry *ce;\n \t\tunsigned long consumed;\n \n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n \t\tce = create_from_disk(ce_mem_pool, istate->version, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n@@ -1973,7 +1972,7 @@ static unsigned long load_cache_entry_block(struct index_state *istate,\n }\n \n static unsigned long load_all_cache_entries(struct index_state *istate,\n-\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+\t\t\tconst char *mmap, size_t mmap_size, unsigned long src_offset)\n {\n \tunsigned long consumed;\n \n@@ -2006,7 +2005,7 @@ struct load_cache_entries_thread_data\n \tstruct index_state *istate;\n \tstruct mem_pool *ce_mem_pool;\n \tint offset, nr;\n-\tvoid *mmap;\n+\tconst char *mmap;\n \tunsigned long start_offset;\n \tstruct cache_entry *previous_ce;\n \tunsigned long consumed;\t/* return # of bytes in index file processed */\n@@ -2026,7 +2025,7 @@ static void *load_cache_entries_thread(void *_data)\n }\n \n static unsigned long load_cache_entries_threaded(int nr_threads, struct index_state *istate,\n-\t\t\tvoid *mmap, size_t mmap_size, unsigned long src_offset)\n+\t\t\tconst char *mmap, size_t mmap_size, unsigned long src_offset)\n {\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tstruct load_cache_entries_thread_data *data;\n@@ -2095,7 +2094,7 @@ static unsigned long load_cache_entries_threaded(int nr_threads, struct index_st\n \t\t\t\tbreak;\n \t\t}\n \n-\t\tondisk = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tondisk = (struct ondisk_cache_entry *)(mmap + src_offset);\n \n \t\t/* On-disk flags are just 16 bits */\n \t\tflags = get_be16(&ondisk->flags);\n@@ -2143,8 +2142,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n-\tstruct cache_header *hdr;\n-\tvoid *mmap;\n+\tconst struct cache_header *hdr;\n+\tconst char *mmap;\n \tsize_t mmap_size;\n \tstruct load_index_extensions p = { 0 };\n \tunsigned long extension_offset = 0;\n@@ -2176,7 +2175,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tdie_errno(\"unable to map index file\");\n \tclose(fd);\n \n-\thdr = mmap;\n+\thdr = (const struct cache_header *)mmap;\n \tif (verify_hdr(hdr, mmap_size) < 0)\n \t\tgoto unmap;\n \n@@ -2236,11 +2235,11 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tp.src_offset = src_offset;\n \t\tload_index_extensions(&p);\n \t}\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \treturn istate->cache_nr;\n \n unmap:\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \tdie(\"index file corrupt\");\n }\n \n@@ -3263,7 +3262,7 @@ int should_validate_cache_entries(void)\n #define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n \n #ifndef NO_PTHREADS\n-static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n+static unsigned long read_eoie_extension(const char *mmap, size_t mmap_size)\n {\n \t/*\n \t * The end of index entries (EOIE) extension is guaranteed to be last\n@@ -3274,7 +3273,6 @@ static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n \t * <4-byte offset>\n \t * <20-byte hash>\n \t */\n-\tconst char *mmap = mmap_;\n \tconst char *index, *eoie;\n \tuint32_t extsize;\n \tunsigned long offset, src_offset;\n@@ -3327,8 +3325,7 @@ static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n \t\t * in 4-byte network byte order.\n \t\t */\n \t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n+\t\textsize = get_be32(mmap + src_offset + 4);\n \n \t\t/* verify the extension size isn't so large it will wrap around */\n \t\tif (src_offset + 8 + extsize < src_offset)\n-- \n2.18.0.windows.1\n\n"},{"id":"357982","messageId":"20180912161832.55324-2-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180912161832.55324-1-benpeart@microsoft.com","subject":"[PATCH v5 1/5] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"benpeart@microsoft.com","sentAt":"2018-09-12T16:18:53Z","receivedAt":"2018-09-12T16:21:04Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"The End of Index Entry (EOIE) is used to locate the end of the variable\nlength index entries and the beginning of the extensions. Code can take\nadvantage of this to quickly locate the index extensions without having\nto parse through all of the index entries.\n\nBecause it must be able to be loaded before the variable length cache\nentries and other index extensions, this extension must be written last.\nThe signature for this extension is { 'E', 'O', 'I', 'E' }.\n\nThe extension consists of:\n\n- 32-bit offset to the end of the index entries\n\n- 160-bit SHA-1 over the extension types and their sizes (but not\ntheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\nlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\nthen the hash would be:\n\nSHA-1(\"TREE\" + <binary representation of N> +\n\t\"REUC\" + <binary representation of M>)\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n Documentation/technical/index-format.txt |  23 ++++\n read-cache.c                             | 154 +++++++++++++++++++++--\n t/README                                 |   5 +\n t/t1700-split-index.sh                   |   1 +\n 4 files changed, 175 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\nindex db3572626b..6bc2d90f7f 100644\n--- a/Documentation/technical/index-format.txt\n+++ b/Documentation/technical/index-format.txt\n@@ -314,3 +314,26 @@ The remaining data of each directory block is grouped by type:\n \n   - An ewah bitmap, the n-th bit indicates whether the n-th index entry\n     is not CE_FSMONITOR_VALID.\n+\n+== End of Index Entry\n+\n+  The End of Index Entry (EOIE) is used to locate the end of the variable\n+  length index entries and the begining of the extensions. Code can take\n+  advantage of this to quickly locate the index extensions without having\n+  to parse through all of the index entries.\n+\n+  Because it must be able to be loaded before the variable length cache\n+  entries and other index extensions, this extension must be written last.\n+  The signature for this extension is { 'E', 'O', 'I', 'E' }.\n+\n+  The extension consists of:\n+\n+  - 32-bit offset to the end of the index entries\n+\n+  - 160-bit SHA-1 over the extension types and their sizes (but not\n+\ttheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\tlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\tthen the hash would be:\n+\n+\tSHA-1(\"TREE\" + <binary representation of N> +\n+\t\t\"REUC\" + <binary representation of M>)\ndiff --git a/read-cache.c b/read-cache.c\nindex 7b1354d759..858935f123 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -43,6 +43,7 @@\n #define CACHE_EXT_LINK 0x6c696e6b\t  /* \"link\" */\n #define CACHE_EXT_UNTRACKED 0x554E5452\t  /* \"UNTR\" */\n #define CACHE_EXT_FSMONITOR 0x46534D4E\t  /* \"FSMN\" */\n+#define CACHE_EXT_ENDOFINDEXENTRIES 0x454F4945\t/* \"EOIE\" */\n \n /* changes that can be kept in $GIT_DIR/index (basically all extensions) */\n #define EXTMASK (RESOLVE_UNDO_CHANGED | CACHE_TREE_CHANGED | \\\n@@ -1693,6 +1694,9 @@ static int read_index_extension(struct index_state *istate,\n \tcase CACHE_EXT_FSMONITOR:\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n+\tcase CACHE_EXT_ENDOFINDEXENTRIES:\n+\t\t/* already handled in do_read_index() */\n+\t\tbreak;\n \tdefault:\n \t\tif (*ext < 'A' || 'Z' < *ext)\n \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n@@ -1889,6 +1893,11 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+#ifndef NO_PTHREADS\n+static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n+#endif\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -2198,11 +2207,15 @@ static int ce_write(git_hash_ctx *context, int fd, void *data, unsigned int len)\n \treturn 0;\n }\n \n-static int write_index_ext_header(git_hash_ctx *context, int fd,\n-\t\t\t\t  unsigned int ext, unsigned int sz)\n+static int write_index_ext_header(git_hash_ctx *context, git_hash_ctx *eoie_context,\n+\t\t\t\t  int fd, unsigned int ext, unsigned int sz)\n {\n \text = htonl(ext);\n \tsz = htonl(sz);\n+\tif (eoie_context) {\n+\t\tthe_hash_algo->update_fn(eoie_context, &ext, 4);\n+\t\tthe_hash_algo->update_fn(eoie_context, &sz, 4);\n+\t}\n \treturn ((ce_write(context, fd, &ext, 4) < 0) ||\n \t\t(ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n }\n@@ -2445,7 +2458,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n {\n \tuint64_t start = getnanotime();\n \tint newfd = tempfile->fd;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, eoie_c;\n \tstruct cache_header hdr;\n \tint i, err = 0, removed, extended, hdr_version;\n \tstruct cache_entry **cache = istate->cache;\n@@ -2454,6 +2467,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct ondisk_cache_entry_extended ondisk;\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n+\tunsigned long offset;\n \n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n@@ -2520,11 +2534,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\treturn err;\n \n \t/* Write extension data here */\n+\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\tthe_hash_algo->init_fn(&eoie_c);\n \tif (!strip_extensions && istate->split_index) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\terr = write_link_extension(&sb, istate) < 0 ||\n-\t\t\twrite_index_ext_header(&c, newfd, CACHE_EXT_LINK,\n+\t\t\twrite_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_LINK,\n \t\t\t\t\t       sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2535,7 +2551,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tcache_tree_write(&sb, istate->cache_tree);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_TREE, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_TREE, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2545,7 +2561,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tresolve_undo_write(&sb, istate->resolve_undo);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_RESOLVE_UNDO,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_RESOLVE_UNDO,\n \t\t\t\t\t     sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2556,7 +2572,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_untracked_extension(&sb, istate->untracked);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_UNTRACKED,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_UNTRACKED,\n \t\t\t\t\t     sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2567,7 +2583,23 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_fsmonitor_extension(&sb, istate);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+\n+\t/*\n+\t * CACHE_EXT_ENDOFINDEXENTRIES must be written as the last entry before the SHA1\n+\t * so that it can be found and processed before all the index entries are\n+\t * read.\n+\t */\n+\tif (!strip_extensions && offset && !git_env_bool(\"GIT_TEST_DISABLE_EOIE\", 0)) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\twrite_eoie_extension(&sb, &eoie_c, offset);\n+\t\terr = write_index_ext_header(&c, NULL, newfd, CACHE_EXT_ENDOFINDEXENTRIES, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2978,3 +3010,109 @@ int should_validate_cache_entries(void)\n \n \treturn validate_index_cache_entries;\n }\n+\n+#define EOIE_SIZE (4 + GIT_SHA1_RAWSZ) /* <4-byte offset> + <20-byte hash> */\n+#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n+\n+#ifndef NO_PTHREADS\n+static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n+{\n+\t/*\n+\t * The end of index entries (EOIE) extension is guaranteed to be last\n+\t * so that it can be found by scanning backwards from the EOF.\n+\t *\n+\t * \"EOIE\"\n+\t * <4-byte length>\n+\t * <4-byte offset>\n+\t * <20-byte hash>\n+\t */\n+\tconst char *mmap = mmap_;\n+\tconst char *index, *eoie;\n+\tuint32_t extsize;\n+\tunsigned long offset, src_offset;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\tgit_hash_ctx c;\n+\n+\t/* ensure we have an index big enough to contain an EOIE extension */\n+\tif (mmap_size < sizeof(struct cache_header) + EOIE_SIZE_WITH_HEADER + the_hash_algo->rawsz)\n+\t\treturn 0;\n+\n+\t/* validate the extension signature */\n+\tindex = eoie = mmap + mmap_size - EOIE_SIZE_WITH_HEADER - the_hash_algo->rawsz;\n+\tif (CACHE_EXT(index) != CACHE_EXT_ENDOFINDEXENTRIES)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/* validate the extension size */\n+\textsize = get_be32(index);\n+\tif (extsize != EOIE_SIZE)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * Validate the offset we're going to look for the first extension\n+\t * signature is after the index header and before the eoie extension.\n+\t */\n+\toffset = get_be32(index);\n+\tif (mmap + offset < mmap + sizeof(struct cache_header))\n+\t\treturn 0;\n+\tif (mmap + offset >= eoie)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * The hash is computed over extension types and their sizes (but not\n+\t * their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\t * long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\t * then the hash would be:\n+\t *\n+\t * SHA-1(\"TREE\" + <binary representation of N> +\n+\t *               \"REUC\" + <binary representation of M>)\n+\t */\n+\tsrc_offset = offset;\n+\tthe_hash_algo->init_fn(&c);\n+\twhile (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\n+\t\t/* verify the extension size isn't so large it will wrap around */\n+\t\tif (src_offset + 8 + extsize < src_offset)\n+\t\t\treturn 0;\n+\n+\t\tthe_hash_algo->update_fn(&c, mmap + src_offset, 8);\n+\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\tthe_hash_algo->final_fn(hash, &c);\n+\tif (hashcmp(hash, (const unsigned char *)index))\n+\t\treturn 0;\n+\n+\t/* Validate that the extension offsets returned us back to the eoie extension. */\n+\tif (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n+\t\treturn 0;\n+\n+\treturn offset;\n+}\n+#endif\n+\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset)\n+{\n+\tuint32_t buffer;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\n+\t/* offset */\n+\tput_be32(&buffer, offset);\n+\tstrbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+\t/* hash */\n+\tthe_hash_algo->final_fn(hash, eoie_context);\n+\tstrbuf_add(sb, hash, the_hash_algo->rawsz);\n+}\ndiff --git a/t/README b/t/README\nindex 9028b47d92..d8754dd23a 100644\n--- a/t/README\n+++ b/t/README\n@@ -319,6 +319,11 @@ GIT_TEST_OE_DELTA_SIZE=<n> exercises the uncomon pack-objects code\n path where deltas larger than this limit require extra memory\n allocation for bookkeeping.\n \n+GIT_TEST_DISABLE_EOIE=<boolean> disables writing the EOIE extension.\n+This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n+as they currently hard code SHA values for the index which are no longer\n+valid due to the addition of the EOIE extension.\n+\n Naming Tests\n ------------\n \ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex 39133bcbc8..f613dd72e3 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -7,6 +7,7 @@ test_description='split index mode tests'\n # We need total control of index splitting here\n sane_unset GIT_TEST_SPLIT_INDEX\n sane_unset GIT_FSMONITOR_TEST\n+export GIT_TEST_DISABLE_EOIE=true\n \n test_expect_success 'enable split index' '\n \tgit config splitIndex.maxPercentChange 100 &&\n-- \n2.18.0.windows.1\n\n"},{"id":"358100","messageId":"xmqq8t45dnlh.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180912161832.55324-2-benpeart@microsoft.com","subject":"Re: [PATCH v5 1/5] eoie: add End of Index Entry (EOIE) extension","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-13T22:44:26Z","receivedAt":"2018-09-13T22:44:31Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <benpeart@microsoft.com> writes:\n\n> diff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\n> index 39133bcbc8..f613dd72e3 100755\n> --- a/t/t1700-split-index.sh\n> +++ b/t/t1700-split-index.sh\n> @@ -7,6 +7,7 @@ test_description='split index mode tests'\n>  # We need total control of index splitting here\n>  sane_unset GIT_TEST_SPLIT_INDEX\n>  sane_unset GIT_FSMONITOR_TEST\n> +export GIT_TEST_DISABLE_EOIE=true\n>  \n>  test_expect_success 'enable split index' '\n>  \tgit config splitIndex.maxPercentChange 100 &&\n\nIt is safer to squash the following in; we may want to revisit the\ndecision test-lint makes on this issue later, though.\n\n-- >8 --\nSubject: [PATCH] SQUASH???\n\nhttp://pubs.opengroup.org/onlinepubs/9699919799/utilities/V3_chap02.html#export\n\nspecifies how \"export name[=word]\" ought to work, but because\nwriting \"name=word; export name\" is not so much more cumbersome\nand some older shells that do not understand the former do grok\nthe latter.  test-lint also recommends spelling it this way.\n---\n t/t1700-split-index.sh | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex f613dd72e3..dab97c2187 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -7,7 +7,7 @@ test_description='split index mode tests'\n # We need total control of index splitting here\n sane_unset GIT_TEST_SPLIT_INDEX\n sane_unset GIT_FSMONITOR_TEST\n-export GIT_TEST_DISABLE_EOIE=true\n+GIT_TEST_DISABLE_EOIE=true; export GIT_TEST_DISABLE_EOIE\n \n test_expect_success 'enable split index' '\n \tgit config splitIndex.maxPercentChange 100 &&\n-- \n2.19.0\n\n"},{"id":"358180","messageId":"CACsJy8B51s2j0aR69UdwtpSbRN6qdLy--am_tyP5Xqo=5Zm_7g@mail.gmail.com","threadId":"49204","inReplyTo":"20180912161832.55324-2-benpeart@microsoft.com","subject":"Re: [PATCH v5 1/5] eoie: add End of Index Entry (EOIE) extension","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-15T10:02:35Z","receivedAt":"2018-09-15T10:03:06Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Sep 12, 2018 at 6:18 PM Ben Peart <benpeart@microsoft.com> wrote:\n>\n> The End of Index Entry (EOIE) is used to locate the end of the variable\n> length index entries and the beginning of the extensions. Code can take\n> advantage of this to quickly locate the index extensions without having\n> to parse through all of the index entries.\n>\n> Because it must be able to be loaded before the variable length cache\n> entries and other index extensions, this extension must be written last.\n> The signature for this extension is { 'E', 'O', 'I', 'E' }.\n>\n> The extension consists of:\n>\n> - 32-bit offset to the end of the index entries\n>\n> - 160-bit SHA-1 over the extension types and their sizes (but not\n> their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n> long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n> then the hash would be:\n>\n> SHA-1(\"TREE\" + <binary representation of N> +\n>         \"REUC\" + <binary representation of M>)\n>\n> Signed-off-by: Ben Peart <Ben.Peart@microsoft.com>\n> ---\n>  Documentation/technical/index-format.txt |  23 ++++\n>  read-cache.c                             | 154 +++++++++++++++++++++--\n>  t/README                                 |   5 +\n>  t/t1700-split-index.sh                   |   1 +\n>  4 files changed, 175 insertions(+), 8 deletions(-)\n>\n> diff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\n> index db3572626b..6bc2d90f7f 100644\n> --- a/Documentation/technical/index-format.txt\n> +++ b/Documentation/technical/index-format.txt\n> @@ -314,3 +314,26 @@ The remaining data of each directory block is grouped by type:\n>\n>    - An ewah bitmap, the n-th bit indicates whether the n-th index entry\n>      is not CE_FSMONITOR_VALID.\n> +\n> +== End of Index Entry\n> +\n> +  The End of Index Entry (EOIE) is used to locate the end of the variable\n> +  length index entries and the begining of the extensions. Code can take\n> +  advantage of this to quickly locate the index extensions without having\n> +  to parse through all of the index entries.\n> +\n> +  Because it must be able to be loaded before the variable length cache\n> +  entries and other index extensions, this extension must be written last.\n> +  The signature for this extension is { 'E', 'O', 'I', 'E' }.\n> +\n> +  The extension consists of:\n> +\n> +  - 32-bit offset to the end of the index entries\n> +\n> +  - 160-bit SHA-1 over the extension types and their sizes (but not\n> +       their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n> +       long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n> +       then the hash would be:\n> +\n> +       SHA-1(\"TREE\" + <binary representation of N> +\n> +               \"REUC\" + <binary representation of M>)\n> diff --git a/read-cache.c b/read-cache.c\n> index 7b1354d759..858935f123 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -43,6 +43,7 @@\n>  #define CACHE_EXT_LINK 0x6c696e6b        /* \"link\" */\n>  #define CACHE_EXT_UNTRACKED 0x554E5452   /* \"UNTR\" */\n>  #define CACHE_EXT_FSMONITOR 0x46534D4E   /* \"FSMN\" */\n> +#define CACHE_EXT_ENDOFINDEXENTRIES 0x454F4945 /* \"EOIE\" */\n>\n>  /* changes that can be kept in $GIT_DIR/index (basically all extensions) */\n>  #define EXTMASK (RESOLVE_UNDO_CHANGED | CACHE_TREE_CHANGED | \\\n> @@ -1693,6 +1694,9 @@ static int read_index_extension(struct index_state *istate,\n>         case CACHE_EXT_FSMONITOR:\n>                 read_fsmonitor_extension(istate, data, sz);\n>                 break;\n> +       case CACHE_EXT_ENDOFINDEXENTRIES:\n> +               /* already handled in do_read_index() */\n> +               break;\n\nPerhaps catch this extension when it's not written at the end (e.g. by\nsome other git implementation) and warn.\n\n>         default:\n>                 if (*ext < 'A' || 'Z' < *ext)\n>                         return error(\"index uses %.4s extension, which we do not understand\",\n> @@ -1889,6 +1893,11 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n>         return ondisk_size + entries * per_entry;\n>  }\n>\n> +#ifndef NO_PTHREADS\n> +static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n> +#endif\n\nKeep functions unconditionally built as much as possible. I don't see\nwhy this read_eoie_extension() must be built only on multithread\nplatforms.\n\n> +static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n> +\n>  /* remember to discard_cache() before reading a different cache! */\n>  int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>  {\n> @@ -2198,11 +2207,15 @@ static int ce_write(git_hash_ctx *context, int fd, void *data, unsigned int len)\n>         return 0;\n>  }\n>\n> -static int write_index_ext_header(git_hash_ctx *context, int fd,\n> -                                 unsigned int ext, unsigned int sz)\n> +static int write_index_ext_header(git_hash_ctx *context, git_hash_ctx *eoie_context,\n> +                                 int fd, unsigned int ext, unsigned int sz)\n>  {\n>         ext = htonl(ext);\n>         sz = htonl(sz);\n> +       if (eoie_context) {\n> +               the_hash_algo->update_fn(eoie_context, &ext, 4);\n> +               the_hash_algo->update_fn(eoie_context, &sz, 4);\n> +       }\n>         return ((ce_write(context, fd, &ext, 4) < 0) ||\n>                 (ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n>  }\n> @@ -2445,7 +2458,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>  {\n>         uint64_t start = getnanotime();\n>         int newfd = tempfile->fd;\n> -       git_hash_ctx c;\n> +       git_hash_ctx c, eoie_c;\n>         struct cache_header hdr;\n>         int i, err = 0, removed, extended, hdr_version;\n>         struct cache_entry **cache = istate->cache;\n> @@ -2454,6 +2467,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>         struct ondisk_cache_entry_extended ondisk;\n>         struct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n>         int drop_cache_tree = istate->drop_cache_tree;\n> +       unsigned long offset;\n>\n>         for (i = removed = extended = 0; i < entries; i++) {\n>                 if (cache[i]->ce_flags & CE_REMOVE)\n> @@ -2520,11 +2534,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>                 return err;\n>\n>         /* Write extension data here */\n> +       offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n> +       the_hash_algo->init_fn(&eoie_c);\n\nDon't write (or even calculate to write it) unless it's needed. Which\nmeans only do this when parallel reading is enabled and the index size\nlarge enough, or when a test variable is set so you can force writing\nthis extension.\n\nI briefly wondered if we should continue writing the extension if it's\nalready written. This way you can manually enable it with \"git\nupdate-index\". But I don't think it's worth the complexity.\n\n>         if (!strip_extensions && istate->split_index) {\n>                 struct strbuf sb = STRBUF_INIT;\n>\n>                 err = write_link_extension(&sb, istate) < 0 ||\n> -                       write_index_ext_header(&c, newfd, CACHE_EXT_LINK,\n> +                       write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_LINK,\n>                                                sb.len) < 0 ||\n>                         ce_write(&c, newfd, sb.buf, sb.len) < 0;\n>                 strbuf_release(&sb);\n> @@ -2535,7 +2551,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>                 struct strbuf sb = STRBUF_INIT;\n>\n>                 cache_tree_write(&sb, istate->cache_tree);\n> -               err = write_index_ext_header(&c, newfd, CACHE_EXT_TREE, sb.len) < 0\n> +               err = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_TREE, sb.len) < 0\n>                         || ce_write(&c, newfd, sb.buf, sb.len) < 0;\n>                 strbuf_release(&sb);\n>                 if (err)\n> @@ -2545,7 +2561,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>                 struct strbuf sb = STRBUF_INIT;\n>\n>                 resolve_undo_write(&sb, istate->resolve_undo);\n> -               err = write_index_ext_header(&c, newfd, CACHE_EXT_RESOLVE_UNDO,\n> +               err = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_RESOLVE_UNDO,\n>                                              sb.len) < 0\n>                         || ce_write(&c, newfd, sb.buf, sb.len) < 0;\n>                 strbuf_release(&sb);\n> @@ -2556,7 +2572,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>                 struct strbuf sb = STRBUF_INIT;\n>\n>                 write_untracked_extension(&sb, istate->untracked);\n> -               err = write_index_ext_header(&c, newfd, CACHE_EXT_UNTRACKED,\n> +               err = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_UNTRACKED,\n>                                              sb.len) < 0 ||\n>                         ce_write(&c, newfd, sb.buf, sb.len) < 0;\n>                 strbuf_release(&sb);\n> @@ -2567,7 +2583,23 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>                 struct strbuf sb = STRBUF_INIT;\n>\n>                 write_fsmonitor_extension(&sb, istate);\n> -               err = write_index_ext_header(&c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n> +               err = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n> +                       || ce_write(&c, newfd, sb.buf, sb.len) < 0;\n> +               strbuf_release(&sb);\n> +               if (err)\n> +                       return -1;\n> +       }\n> +\n> +       /*\n> +        * CACHE_EXT_ENDOFINDEXENTRIES must be written as the last entry before the SHA1\n> +        * so that it can be found and processed before all the index entries are\n> +        * read.\n> +        */\n> +       if (!strip_extensions && offset && !git_env_bool(\"GIT_TEST_DISABLE_EOIE\", 0)) {\n> +               struct strbuf sb = STRBUF_INIT;\n> +\n> +               write_eoie_extension(&sb, &eoie_c, offset);\n> +               err = write_index_ext_header(&c, NULL, newfd, CACHE_EXT_ENDOFINDEXENTRIES, sb.len) < 0\n>                         || ce_write(&c, newfd, sb.buf, sb.len) < 0;\n>                 strbuf_release(&sb);\n>                 if (err)\n> @@ -2978,3 +3010,109 @@ int should_validate_cache_entries(void)\n>\n>         return validate_index_cache_entries;\n>  }\n> +\n> +#define EOIE_SIZE (4 + GIT_SHA1_RAWSZ) /* <4-byte offset> + <20-byte hash> */\n> +#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n> +\n> +#ifndef NO_PTHREADS\n> +static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size)\n> +{\n> +       /*\n> +        * The end of index entries (EOIE) extension is guaranteed to be last\n> +        * so that it can be found by scanning backwards from the EOF.\n> +        *\n> +        * \"EOIE\"\n> +        * <4-byte length>\n> +        * <4-byte offset>\n> +        * <20-byte hash>\n> +        */\n> +       const char *mmap = mmap_;\n> +       const char *index, *eoie;\n> +       uint32_t extsize;\n> +       unsigned long offset, src_offset;\n> +       unsigned char hash[GIT_MAX_RAWSZ];\n> +       git_hash_ctx c;\n> +\n> +       /* ensure we have an index big enough to contain an EOIE extension */\n> +       if (mmap_size < sizeof(struct cache_header) + EOIE_SIZE_WITH_HEADER + the_hash_algo->rawsz)\n> +               return 0;\n\nAll these \"return 0\" indicates an error in EOIE extension. You\nprobably want to print some warning (much easier to track down why\nparallel reading does not happen).\n\n> +\n> +       /* validate the extension signature */\n> +       index = eoie = mmap + mmap_size - EOIE_SIZE_WITH_HEADER - the_hash_algo->rawsz;\n> +       if (CACHE_EXT(index) != CACHE_EXT_ENDOFINDEXENTRIES)\n> +               return 0;\n> +       index += sizeof(uint32_t);\n> +\n> +       /* validate the extension size */\n> +       extsize = get_be32(index);\n> +       if (extsize != EOIE_SIZE)\n> +               return 0;\n> +       index += sizeof(uint32_t);\n> +\n> +       /*\n> +        * Validate the offset we're going to look for the first extension\n> +        * signature is after the index header and before the eoie extension.\n> +        */\n> +       offset = get_be32(index);\n> +       if (mmap + offset < mmap + sizeof(struct cache_header))\n> +               return 0;\n> +       if (mmap + offset >= eoie)\n> +               return 0;\n> +       index += sizeof(uint32_t);\n> +\n> +       /*\n> +        * The hash is computed over extension types and their sizes (but not\n> +        * their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n> +        * long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n> +        * then the hash would be:\n> +        *\n> +        * SHA-1(\"TREE\" + <binary representation of N> +\n> +        *               \"REUC\" + <binary representation of M>)\n> +        */\n> +       src_offset = offset;\n> +       the_hash_algo->init_fn(&c);\n> +       while (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n> +               /* After an array of active_nr index entries,\n> +                * there can be arbitrary number of extended\n> +                * sections, each of which is prefixed with\n> +                * extension name (4-byte) and section length\n> +                * in 4-byte network byte order.\n> +                */\n> +               uint32_t extsize;\n> +               memcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n> +               extsize = ntohl(extsize);\n> +\n> +               /* verify the extension size isn't so large it will wrap around */\n> +               if (src_offset + 8 + extsize < src_offset)\n> +                       return 0;\n> +\n> +               the_hash_algo->update_fn(&c, mmap + src_offset, 8);\n> +\n> +               src_offset += 8;\n> +               src_offset += extsize;\n> +       }\n> +       the_hash_algo->final_fn(hash, &c);\n> +       if (hashcmp(hash, (const unsigned char *)index))\n> +               return 0;\n> +\n> +       /* Validate that the extension offsets returned us back to the eoie extension. */\n> +       if (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n> +               return 0;\n> +\n> +       return offset;\n> +}\n> +#endif\n> +\n> +static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset)\n\nWe normally just put function implementations before it's used to\navoid static forward declaration. Any special reason why it's not done\nhere?\n\n> +{\n> +       uint32_t buffer;\n> +       unsigned char hash[GIT_MAX_RAWSZ];\n> +\n> +       /* offset */\n> +       put_be32(&buffer, offset);\n> +       strbuf_add(sb, &buffer, sizeof(uint32_t));\n> +\n> +       /* hash */\n> +       the_hash_algo->final_fn(hash, eoie_context);\n> +       strbuf_add(sb, hash, the_hash_algo->rawsz);\n> +}\n> diff --git a/t/README b/t/README\n> index 9028b47d92..d8754dd23a 100644\n> --- a/t/README\n> +++ b/t/README\n> @@ -319,6 +319,11 @@ GIT_TEST_OE_DELTA_SIZE=<n> exercises the uncomon pack-objects code\n>  path where deltas larger than this limit require extra memory\n>  allocation for bookkeeping.\n>\n> +GIT_TEST_DISABLE_EOIE=<boolean> disables writing the EOIE extension.\n> +This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n\nI have a feeling that you won't have problems if you don't write eoie\nextension by default in the first place. Then this could be switched\nto GIT_TEST_ENABLE_EOIE instead. We may still have problem if both\neoie and split index are forced on when running through the test\nsuite, but that should be an easy fix.\n\n> +as they currently hard code SHA values for the index which are no longer\n> +valid due to the addition of the EOIE extension.\n> +\n>  Naming Tests\n>  ------------\n>\n> diff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\n> index 39133bcbc8..f613dd72e3 100755\n> --- a/t/t1700-split-index.sh\n> +++ b/t/t1700-split-index.sh\n> @@ -7,6 +7,7 @@ test_description='split index mode tests'\n>  # We need total control of index splitting here\n>  sane_unset GIT_TEST_SPLIT_INDEX\n>  sane_unset GIT_FSMONITOR_TEST\n> +export GIT_TEST_DISABLE_EOIE=true\n>\n>  test_expect_success 'enable split index' '\n>         git config splitIndex.maxPercentChange 100 &&\n> --\n> 2.18.0.windows.1\n>\n\n\n-- \nDuy\n"},{"id":"358181","messageId":"CACsJy8ATsS6S5zib2FqJf1stPcGwSTO1qYBSz514Xu2GfJ4Apw@mail.gmail.com","threadId":"49204","inReplyTo":"20180912161832.55324-3-benpeart@microsoft.com","subject":"Re: [PATCH v5 2/5] read-cache: load cache extensions on a worker thread","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-15T10:22:02Z","receivedAt":"2018-09-15T10:22:34Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Sep 12, 2018 at 6:18 PM Ben Peart <benpeart@microsoft.com> wrote:\n>\n> This patch helps address the CPU cost of loading the index by loading\n> the cache extensions on a worker thread in parallel with loading the cache\n> entries.\n>\n> In some cases, loading the extensions takes longer than loading the\n> cache entries so this patch utilizes the new EOIE to start the thread to\n> load the extensions before loading all the cache entries in parallel.\n>\n> This is possible because the current extensions don't access the cache\n> entries in the index_state structure so are OK that they don't all exist\n> yet.\n>\n> The CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\n> extensions don't even get a pointer to the index so don't have access to the\n> cache entries.\n>\n> CACHE_EXT_LINK only uses the index_state to initialize the split index.\n> CACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\n> update and dirty flags.\n>\n> I used p0002-read-cache.sh to generate some performance data:\n>\n> Test w/100,000 files                Baseline         Parallel Extensions\n> ---------------------------------------------------------------------------\n> read_cache/discard_cache 1000 times 14.08(0.01+0.10) 9.72(0.03+0.06) -31.0%\n>\n> Test w/1,000,000 files              Baseline         Parallel Extensions\n> ------------------------------------------------------------------------------\n> read_cache/discard_cache 1000 times 202.95(0.01+0.07) 154.14(0.03+0.06) -24.1%\n>\n> Signed-off-by: Ben Peart <Ben.Peart@microsoft.com>\n> ---\n>  Documentation/config.txt |  6 +++\n>  config.c                 | 18 ++++++++\n>  config.h                 |  1 +\n>  read-cache.c             | 94 ++++++++++++++++++++++++++++++++--------\n>  4 files changed, 102 insertions(+), 17 deletions(-)\n>\n> diff --git a/Documentation/config.txt b/Documentation/config.txt\n> index 1c42364988..79f8296d9c 100644\n> --- a/Documentation/config.txt\n> +++ b/Documentation/config.txt\n> @@ -2391,6 +2391,12 @@ imap::\n>         The configuration variables in the 'imap' section are described\n>         in linkgit:git-imap-send[1].\n>\n> +index.threads::\n> +       Specifies the number of threads to spawn when loading the index.\n> +       This is meant to reduce index load time on multiprocessor machines.\n> +       Specifying 0 or 'true' will cause Git to auto-detect the number of\n> +       CPU's and set the number of threads accordingly. Defaults to 'true'.\n\nI'd rather this variable defaults to 0. Spawning threads have\nassociated cost and most projects out there are small enough that this\nmulti threading could just add more cost than gain. It only makes\nsense to enable this on huge repos.\n\nWait there's no way to disable this parallel reading? Does not sound\nright. And  if ordinary numbers mean the number of threads then 0\nshould mean no threading. Auto detection could have a new keyword,\nlike 'auto'.\n\n> +\n>  index.version::\n>         Specify the version with which new index files should be\n>         initialized.  This does not affect existing repositories.\n> diff --git a/config.c b/config.c\n> index 9a0b10d4bc..9bd79fb165 100644\n> --- a/config.c\n> +++ b/config.c\n> @@ -2289,6 +2289,24 @@ int git_config_get_fsmonitor(void)\n>         return 0;\n>  }\n>\n> +/*\n> + * You can disable multi-threaded code by setting index.threads\n> + * to 'false' (or 1)\n> + */\n> +int git_config_get_index_threads(void)\n> +{\n> +       int is_bool, val;\n> +\n> +       if (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n> +               if (is_bool)\n> +                       return val ? 0 : 1;\n> +               else\n> +                       return val;\n> +       }\n> +\n> +       return 0; /* auto-detect */\n> +}\n> +\n>  NORETURN\n>  void git_die_config_linenr(const char *key, const char *filename, int linenr)\n>  {\n> diff --git a/config.h b/config.h\n> index ab46e0165d..a06027e69b 100644\n> --- a/config.h\n> +++ b/config.h\n> @@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n>  extern int git_config_get_split_index(void);\n>  extern int git_config_get_max_percent_split_change(void);\n>  extern int git_config_get_fsmonitor(void);\n> +extern int git_config_get_index_threads(void);\n>\n>  /* This dies if the configured or default date is in the future */\n>  extern int git_config_get_expiry(const char *key, const char **output);\n> diff --git a/read-cache.c b/read-cache.c\n> index 858935f123..b203eebb44 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -23,6 +23,10 @@\n>  #include \"split-index.h\"\n>  #include \"utf8.h\"\n>  #include \"fsmonitor.h\"\n> +#ifndef NO_PTHREADS\n> +#include <pthread.h>\n> +#include <thread-utils.h>\n> +#endif\n\nI don't think you're supposed to include system header files after\n\"cache.h\". Including thread-utils.h should be enough (and it keeps the\nexception of inclduing pthread.h in just one place). Please use\n\"pthread-utils.h\" instead of <pthread-utils.h> which is usually for\nsystem header files. And include ptherad-utils.h unconditionally.\n\n>\n>  /* Mask for the name length in ce_flags in the on-disk index */\n>\n> @@ -1898,6 +1902,46 @@ static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n>  #endif\n>  static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n>\n> +struct load_index_extensions\n> +{\n> +#ifndef NO_PTHREADS\n> +       pthread_t pthread;\n> +#endif\n> +       struct index_state *istate;\n> +       void *mmap;\n> +       size_t mmap_size;\n> +       unsigned long src_offset;\n> +};\n> +\n> +static void *load_index_extensions(void *_data)\n> +{\n> +       struct load_index_extensions *p = _data;\n> +       unsigned long src_offset = p->src_offset;\n> +\n> +       while (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n> +               /* After an array of active_nr index entries,\n> +                * there can be arbitrary number of extended\n> +                * sections, each of which is prefixed with\n> +                * extension name (4-byte) and section length\n> +                * in 4-byte network byte order.\n> +                */\n> +               uint32_t extsize;\n> +               memcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n> +               extsize = ntohl(extsize);\n> +               if (read_index_extension(p->istate,\n> +                       (const char *)p->mmap + src_offset,\n> +                       (char *)p->mmap + src_offset + 8,\n> +                       extsize) < 0) {\n> +                       munmap(p->mmap, p->mmap_size);\n> +                       die(\"index file corrupt\");\n\n_()\n\n> +               }\n> +               src_offset += 8;\n> +               src_offset += extsize;\n> +       }\n> +\n> +       return NULL;\n> +}\n> +\n>  /* remember to discard_cache() before reading a different cache! */\n>  int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>  {\n> @@ -1908,6 +1952,11 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>         void *mmap;\n>         size_t mmap_size;\n>         struct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n> +       struct load_index_extensions p = { 0 };\n> +       unsigned long extension_offset = 0;\n> +#ifndef NO_PTHREADS\n> +       int nr_threads;\n> +#endif\n>\n>         if (istate->initialized)\n>                 return istate->cache_nr;\n> @@ -1944,6 +1993,26 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>         istate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n>         istate->initialized = 1;\n>\n> +       p.istate = istate;\n> +       p.mmap = mmap;\n> +       p.mmap_size = mmap_size;\n> +\n> +#ifndef NO_PTHREADS\n> +       nr_threads = git_config_get_index_threads();\n> +       if (!nr_threads)\n> +               nr_threads = online_cpus();\n> +\n> +       if (nr_threads >= 2) {\n> +               extension_offset = read_eoie_extension(mmap, mmap_size);\n> +               if (extension_offset) {\n> +                       /* create a thread to load the index extensions */\n\nPointless comment. It's pretty clear from the pthread_create() below\nthanks to good function naming. Please remove.\n\n> +                       p.src_offset = extension_offset;\n> +                       if (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n> +                               die(_(\"unable to create load_index_extensions_thread\"));\n> +               }\n> +       }\n> +#endif\n> +\n>         if (istate->version == 4) {\n>                 previous_name = &previous_name_buf;\n>                 mem_pool_init(&istate->ce_mem_pool,\n> @@ -1970,23 +2039,14 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>         istate->timestamp.sec = st.st_mtime;\n>         istate->timestamp.nsec = ST_MTIME_NSEC(st);\n>\n> -       while (src_offset <= mmap_size - the_hash_algo->rawsz - 8) {\n> -               /* After an array of active_nr index entries,\n> -                * there can be arbitrary number of extended\n> -                * sections, each of which is prefixed with\n> -                * extension name (4-byte) and section length\n> -                * in 4-byte network byte order.\n> -                */\n> -               uint32_t extsize;\n> -               memcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n> -               extsize = ntohl(extsize);\n> -               if (read_index_extension(istate,\n> -                                        (const char *) mmap + src_offset,\n> -                                        (char *) mmap + src_offset + 8,\n> -                                        extsize) < 0)\n> -                       goto unmap;\n> -               src_offset += 8;\n> -               src_offset += extsize;\n> +       /* if we created a thread, join it otherwise load the extensions on the primary thread */\n> +#ifndef NO_PTHREADS\n> +       if (extension_offset && pthread_join(p.pthread, NULL))\n> +               die(_(\"unable to join load_index_extensions_thread\"));\n\nI guess the last _ is a typo and you wanted \"unable to join\nload_index_extensions thread\". Please use die_errno() instead.\n\n> +#endif\n> +       if (!extension_offset) {\n> +               p.src_offset = src_offset;\n> +               load_index_extensions(&p);\n>         }\n>         munmap(mmap, mmap_size);\n>         return istate->cache_nr;\n> --\n> 2.18.0.windows.1\n>\n\n\n-- \nDuy\n"},{"id":"358182","messageId":"CACsJy8D-sM2SSfTMmsR0uKnP0FM9fGrN0Z-pH5irMsH1=-jrmQ@mail.gmail.com","threadId":"49204","inReplyTo":"CACsJy8ATsS6S5zib2FqJf1stPcGwSTO1qYBSz514Xu2GfJ4Apw@mail.gmail.com","subject":"Re: [PATCH v5 2/5] read-cache: load cache extensions on a worker thread","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-15T10:24:48Z","receivedAt":"2018-09-15T10:25:17Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sat, Sep 15, 2018 at 12:22 PM Duy Nguyen <pclouds@gmail.com> wrote:\n> > @@ -1944,6 +1993,26 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n> >         istate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n> >         istate->initialized = 1;\n> >\n> > +       p.istate = istate;\n> > +       p.mmap = mmap;\n> > +       p.mmap_size = mmap_size;\n> > +\n> > +#ifndef NO_PTHREADS\n> > +       nr_threads = git_config_get_index_threads();\n> > +       if (!nr_threads)\n> > +               nr_threads = online_cpus();\n> > +\n> > +       if (nr_threads >= 2) {\n> > +               extension_offset = read_eoie_extension(mmap, mmap_size);\n> > +               if (extension_offset) {\n\nOne more thing I forgot. If the extension area is small enough, then\nwe should not need to create a thread to parse extensions in parallel.\nWe should know roughly how much work we need because we know the total\nsize of all extensions.\n\n> > +                       /* create a thread to load the index extensions */\n>\n> Pointless comment. It's pretty clear from the pthread_create() below\n> thanks to good function naming. Please remove.\n>\n> > +                       p.src_offset = extension_offset;\n> > +                       if (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n> > +                               die(_(\"unable to create load_index_extensions_thread\"));\n> > +               }\n> > +       }\n> > +#endif\n> > +\n> >         if (istate->version == 4) {\n> >                 previous_name = &previous_name_buf;\n> >                 mem_pool_init(&istate->ce_mem_pool,\n-- \nDuy\n"},{"id":"358183","messageId":"CACsJy8CjWD_CHwv5NoURLt9is7-UUDpKDo-3EcM28imznZAOpA@mail.gmail.com","threadId":"49204","inReplyTo":"20180912161832.55324-4-benpeart@microsoft.com","subject":"Re: [PATCH v5 3/5] read-cache: load cache entries on worker threads","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-15T10:31:29Z","receivedAt":"2018-09-15T10:31:58Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Sep 12, 2018 at 6:18 PM Ben Peart <benpeart@microsoft.com> wrote:\n>\n> This patch helps address the CPU cost of loading the index by creating\n> multiple threads to divide the work of loading and converting the cache\n> entries across all available CPU cores.\n>\n> It accomplishes this by having the primary thread loop across the index file\n> tracking the offset and (for V4 indexes) expanding the name. It creates a\n> thread to process each block of entries as it comes to them.\n>\n> I used p0002-read-cache.sh to generate some performance data:\n>\n> Test w/100,000 files                Baseline         Parallel entries\n> ---------------------------------------------------------------------------\n> read_cache/discard_cache 1000 times 14.08(0.01+0.10) 9.72(0.03+0.06) -31.0%\n>\n> Test w/1,000,000 files              Baseline         Parallel entries\n> ------------------------------------------------------------------------------\n> read_cache/discard_cache 1000 times 202.95(0.01+0.07) 154.14(0.03+0.06) -24.1%\n\nThe numbers here and the previous patch to load extensions in parallel\nare exactly the same. What do these numbers mean? With both changes?\n-- \nDuy\n"},{"id":"358184","messageId":"CACsJy8CUsOLy_OWdJbx5TqyzPJ5eZ0QcrJniQ82nAiwwtk9iyA@mail.gmail.com","threadId":"49204","inReplyTo":"20180912161832.55324-4-benpeart@microsoft.com","subject":"Re: [PATCH v5 3/5] read-cache: load cache entries on worker threads","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-15T11:07:46Z","receivedAt":"2018-09-15T11:08:15Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Sep 12, 2018 at 6:18 PM Ben Peart <benpeart@microsoft.com> wrote:\n>\n> This patch helps address the CPU cost of loading the index by creating\n> multiple threads to divide the work of loading and converting the cache\n> entries across all available CPU cores.\n>\n> It accomplishes this by having the primary thread loop across the index file\n> tracking the offset and (for V4 indexes) expanding the name. It creates a\n> thread to process each block of entries as it comes to them.\n\nI added a couple trace_printf() to see how time is spent. This is with\na 1m entry index (basically my webkit.git index repeated 4 times)\n\n12:50:00.084237 read-cache.c:1721       start loading index\n12:50:00.119941 read-cache.c:1943       performance: 0.034778758 s:\nloaded all extensions (1667075 bytes)\n12:50:00.185352 read-cache.c:2029       performance: 0.100152079 s:\nloaded 367110 entries\n12:50:00.189683 read-cache.c:2126       performance: 0.104566615 s:\nfinished scanning all entries\n12:50:00.217900 read-cache.c:2029       performance: 0.082309193 s:\nloaded 367110 entries\n12:50:00.259969 read-cache.c:2029       performance: 0.070257130 s:\nloaded 367108 entries\n12:50:00.263662 read-cache.c:2278       performance: 0.179344458 s:\nread cache .git/index\n\nTwo observations:\n\n- the extension thread finishes up quickly (this is with TREE\nextension alone). We could use that spare core to parse some more\nentries.\n\n- the main \"scanning and allocating\" thread does hold up the two\nremaining threads. You can see the first index entry thread is\nfinished even before the scanning thread. And this scanning thread\ntakes a lot of cpu.\n\nIf all index entry threads start at the same time, based on these\nnumbers we would be finished around 12:50:00.185352 mark, cutting\nloading time by half.\n\nCould you go back to your original solution? If you don't want to\nspend more time on this, I offer to rewrite this patch.\n-- \nDuy\n"},{"id":"358185","messageId":"20180915110951.GA17817@duynguyen.home","threadId":"49204","inReplyTo":"CACsJy8CUsOLy_OWdJbx5TqyzPJ5eZ0QcrJniQ82nAiwwtk9iyA@mail.gmail.com","subject":"Re: [PATCH v5 3/5] read-cache: load cache entries on worker threads","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-15T11:09:52Z","receivedAt":"2018-09-15T11:09:59Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sat, Sep 15, 2018 at 01:07:46PM +0200, Duy Nguyen wrote:\n> 12:50:00.084237 read-cache.c:1721       start loading index\n> 12:50:00.119941 read-cache.c:1943       performance: 0.034778758 s: loaded all extensions (1667075 bytes)\n> 12:50:00.185352 read-cache.c:2029       performance: 0.100152079 s: loaded 367110 entries\n> 12:50:00.189683 read-cache.c:2126       performance: 0.104566615 s: finished scanning all entries\n> 12:50:00.217900 read-cache.c:2029       performance: 0.082309193 s: loaded 367110 entries\n> 12:50:00.259969 read-cache.c:2029       performance: 0.070257130 s: loaded 367108 entries\n> 12:50:00.263662 read-cache.c:2278       performance: 0.179344458 s: read cache .git/index\n\nThe previous mail wraps these lines and make it a bit hard to read. Corrected now.\n\n--\nDuy\n"},{"id":"358186","messageId":"CACsJy8AhNQhFa1ONsmnLOjznbZss2=L=xCPD5fa2vdHw79x0ag@mail.gmail.com","threadId":"49204","inReplyTo":"20180912161832.55324-4-benpeart@microsoft.com","subject":"Re: [PATCH v5 3/5] read-cache: load cache entries on worker threads","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-15T11:29:41Z","receivedAt":"2018-09-15T11:35:18Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Sep 12, 2018 at 6:18 PM Ben Peart <benpeart@microsoft.com> wrote:\n>  #ifndef NO_PTHREADS\n>         nr_threads = git_config_get_index_threads();\n> -       if (!nr_threads)\n> -               nr_threads = online_cpus();\n> +       if (!nr_threads) {\n> +               cpus = online_cpus();\n> +               nr_threads = istate->cache_nr / THREAD_COST;\n> +               if (nr_threads > cpus)\n> +                       nr_threads = cpus;\n\nIt seems like overcommitting cpu does reduce time. With this patch\n(and a 4 core system), I got\n\n$ test-tool read-cache 100\nreal    0m36.270s\nuser    0m54.193s\nsys     0m17.346s\n\nif I force nr_threads to 9 (even though cpus is 4)\n\n$ test-tool read-cache 100\nreal    0m33.592s\nuser    1m4.230s\nsys     0m18.380s\n\nEven though we use more cpus, real time is shorter. I guess these\nthreads still sleep a bit due to I/O and having more threads than\ncores will utilize those idle cycles.\n--\nDuy\n"},{"id":"358214","messageId":"CACsJy8DtYrB99_7GyZaHsLTG8Ff7Mt_hENwhLE6_x6b5zzQ9Hg@mail.gmail.com","threadId":"49204","inReplyTo":"CACsJy8ATsS6S5zib2FqJf1stPcGwSTO1qYBSz514Xu2GfJ4Apw@mail.gmail.com","subject":"Re: [PATCH v5 2/5] read-cache: load cache extensions on a worker thread","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-15T16:23:59Z","receivedAt":"2018-09-15T16:24:27Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sat, Sep 15, 2018 at 12:22 PM Duy Nguyen <pclouds@gmail.com> wrote:\n> Wait there's no way to disable this parallel reading? Does not sound\n> right. And  if ordinary numbers mean the number of threads then 0\n> should mean no threading. Auto detection could have a new keyword,\n> like 'auto'.\n\nMy bad. Disabling threading means _1_ thread. What was I thinking...\n-- \nDuy\n"},{"id":"358271","messageId":"f7250999-71a3-0113-2858-e66bad283db3@gmail.com","threadId":"49204","inReplyTo":"CACsJy8B51s2j0aR69UdwtpSbRN6qdLy--am_tyP5Xqo=5Zm_7g@mail.gmail.com","subject":"Re: [PATCH v5 1/5] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-17T14:54:59Z","receivedAt":"2018-09-17T14:55:06Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/15/2018 6:02 AM, Duy Nguyen wrote:\n\n>>          default:\n>>                  if (*ext < 'A' || 'Z' < *ext)\n>>                          return error(\"index uses %.4s extension, which we do not understand\",\n>> @@ -1889,6 +1893,11 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n>>          return ondisk_size + entries * per_entry;\n>>   }\n>>\n>> +#ifndef NO_PTHREADS\n>> +static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n>> +#endif\n> \n> Keep functions unconditionally built as much as possible. I don't see\n> why this read_eoie_extension() must be built only on multithread\n> platforms.\n> \n\nThis is conditional to avoid generating a warning on single threaded \nplatforms where the function is currently unused.  That seemed like a \nbetter choice than calling it and ignoring it on single threaded \nplatforms just to avoid a compiler warning.\n\n>> +static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n>> +\n>>   /* remember to discard_cache() before reading a different cache! */\n>>   int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>>   {\n>> @@ -2198,11 +2207,15 @@ static int ce_write(git_hash_ctx *context, int fd, void *data, unsigned int len)\n>>          return 0;\n>>   }\n>>\n>> -static int write_index_ext_header(git_hash_ctx *context, int fd,\n>> -                                 unsigned int ext, unsigned int sz)\n>> +static int write_index_ext_header(git_hash_ctx *context, git_hash_ctx *eoie_context,\n>> +                                 int fd, unsigned int ext, unsigned int sz)\n>>   {\n>>          ext = htonl(ext);\n>>          sz = htonl(sz);\n>> +       if (eoie_context) {\n>> +               the_hash_algo->update_fn(eoie_context, &ext, 4);\n>> +               the_hash_algo->update_fn(eoie_context, &sz, 4);\n>> +       }\n>>          return ((ce_write(context, fd, &ext, 4) < 0) ||\n>>                  (ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n>>   }\n>> @@ -2445,7 +2458,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>>   {\n>>          uint64_t start = getnanotime();\n>>          int newfd = tempfile->fd;\n>> -       git_hash_ctx c;\n>> +       git_hash_ctx c, eoie_c;\n>>          struct cache_header hdr;\n>>          int i, err = 0, removed, extended, hdr_version;\n>>          struct cache_entry **cache = istate->cache;\n>> @@ -2454,6 +2467,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>>          struct ondisk_cache_entry_extended ondisk;\n>>          struct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n>>          int drop_cache_tree = istate->drop_cache_tree;\n>> +       unsigned long offset;\n>>\n>>          for (i = removed = extended = 0; i < entries; i++) {\n>>                  if (cache[i]->ce_flags & CE_REMOVE)\n>> @@ -2520,11 +2534,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>>                  return err;\n>>\n>>          /* Write extension data here */\n>> +       offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n>> +       the_hash_algo->init_fn(&eoie_c);\n> \n> Don't write (or even calculate to write it) unless it's needed. Which\n> means only do this when parallel reading is enabled and the index size\n> large enough, or when a test variable is set so you can force writing\n> this extension.\n\nI made the logic always write the extension based on the earlier \ndiscussion [1] where it was suggested this should have been part of the \noriginal index format for extensions from the beginning.  This helps \nensure it is available for current and future uses we haven't even \ndiscovered yet.\n\n[1] \nhttps://public-inbox.org/git/xmqqwp2s1h1x.fsf@gitster.mtv.corp.google.com/\n\n\n>> +\n>> +static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset)\n> \n> We normally just put function implementations before it's used to\n> avoid static forward declaration. Any special reason why it's not done\n> here?\n> \n\nThis was done to promote readability of the (already large) read-cache.c \nfile.  I first considered moving the EOIE read/write functions into a \nseparate file entirely but they need access to information only \navailable within read-cache.c so I compromised and moved them to the end \nof the file instead.\n\n>> +{\n>> +       uint32_t buffer;\n>> +       unsigned char hash[GIT_MAX_RAWSZ];\n>> +\n>> +       /* offset */\n>> +       put_be32(&buffer, offset);\n>> +       strbuf_add(sb, &buffer, sizeof(uint32_t));\n>> +\n>> +       /* hash */\n>> +       the_hash_algo->final_fn(hash, eoie_context);\n>> +       strbuf_add(sb, hash, the_hash_algo->rawsz);\n>> +}\n>> diff --git a/t/README b/t/README\n>> index 9028b47d92..d8754dd23a 100644\n>> --- a/t/README\n>> +++ b/t/README\n>> @@ -319,6 +319,11 @@ GIT_TEST_OE_DELTA_SIZE=<n> exercises the uncomon pack-objects code\n>>   path where deltas larger than this limit require extra memory\n>>   allocation for bookkeeping.\n>>\n>> +GIT_TEST_DISABLE_EOIE=<boolean> disables writing the EOIE extension.\n>> +This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n> \n> I have a feeling that you won't have problems if you don't write eoie\n> extension by default in the first place. Then this could be switched\n> to GIT_TEST_ENABLE_EOIE instead. We may still have problem if both\n> eoie and split index are forced on when running through the test\n> suite, but that should be an easy fix.\n> \n>> +as they currently hard code SHA values for the index which are no longer\n>> +valid due to the addition of the EOIE extension.\n>> +\n>>   Naming Tests\n>>   ------------\n>>\n>> diff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\n>> index 39133bcbc8..f613dd72e3 100755\n>> --- a/t/t1700-split-index.sh\n>> +++ b/t/t1700-split-index.sh\n>> @@ -7,6 +7,7 @@ test_description='split index mode tests'\n>>   # We need total control of index splitting here\n>>   sane_unset GIT_TEST_SPLIT_INDEX\n>>   sane_unset GIT_FSMONITOR_TEST\n>> +export GIT_TEST_DISABLE_EOIE=true\n>>\n>>   test_expect_success 'enable split index' '\n>>          git config splitIndex.maxPercentChange 100 &&\n>> --\n>> 2.18.0.windows.1\n>>\n> \n> \n"},{"id":"358282","messageId":"CACsJy8DEvLJYBm0P1VtvKFD-CAo6_4Z13dBHWkuuAavghbGBag@mail.gmail.com","threadId":"49204","inReplyTo":"f7250999-71a3-0113-2858-e66bad283db3@gmail.com","subject":"Re: [PATCH v5 1/5] eoie: add End of Index Entry (EOIE) extension","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-17T16:05:10Z","receivedAt":"2018-09-17T16:05:40Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Sep 17, 2018 at 4:55 PM Ben Peart <peartben@gmail.com> wrote:\n> On 9/15/2018 6:02 AM, Duy Nguyen wrote:\n>\n> >>          default:\n> >>                  if (*ext < 'A' || 'Z' < *ext)\n> >>                          return error(\"index uses %.4s extension, which we do not understand\",\n> >> @@ -1889,6 +1893,11 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n> >>          return ondisk_size + entries * per_entry;\n> >>   }\n> >>\n> >> +#ifndef NO_PTHREADS\n> >> +static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n> >> +#endif\n> >\n> > Keep functions unconditionally built as much as possible. I don't see\n> > why this read_eoie_extension() must be built only on multithread\n> > platforms.\n> >\n>\n> This is conditional to avoid generating a warning on single threaded\n> platforms where the function is currently unused.  That seemed like a\n> better choice than calling it and ignoring it on single threaded\n> platforms just to avoid a compiler warning.\n\nThe third option is ignore the compiler. I consider that warning a\nhelpful suggestion, not a strict rule.\n\nMost devs don't run single thread builds (I think) so is this function\nis updated in a way that breaks single thread mode, it can only be\nfound out when this function is used in single thread mode. At that\npoint the function may have changed a lot. If it's built\nunconditionally, at least single thread users will yell up much sooner\nand we could fix it much earlier.\n\n> >> @@ -2520,11 +2534,13 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n> >>                  return err;\n> >>\n> >>          /* Write extension data here */\n> >> +       offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n> >> +       the_hash_algo->init_fn(&eoie_c);\n> >\n> > Don't write (or even calculate to write it) unless it's needed. Which\n> > means only do this when parallel reading is enabled and the index size\n> > large enough, or when a test variable is set so you can force writing\n> > this extension.\n>\n> I made the logic always write the extension based on the earlier\n> discussion [1] where it was suggested this should have been part of the\n> original index format for extensions from the beginning.  This helps\n> ensure it is available for current and future uses we haven't even\n> discovered yet.\n\nBut it _is_ available now. If you need it, you write the extension\nout. If we make this part of index version 5 (and make it not an\nextension anymore) then I buy that argument. As it is, it's an\noptional extension.\n\n> [1] https://public-inbox.org/git/xmqqwp2s1h1x.fsf@gitster.mtv.corp.google.com/\n>\n>\n> >> +\n> >> +static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset)\n> >\n> > We normally just put function implementations before it's used to\n> > avoid static forward declaration. Any special reason why it's not done\n> > here?\n> >\n>\n> This was done to promote readability of the (already large) read-cache.c\n> file.  I first considered moving the EOIE read/write functions into a\n> separate file entirely but they need access to information only\n> available within read-cache.c so I compromised and moved them to the end\n> of the file instead.\n\nI consider grouping extension related functions closer to\nread_index_extension gives better readability, or at least better than\njust putting new functions at the end in no particular order. But I\nguess this is personal view.\n-- \nDuy\n"},{"id":"358285","messageId":"78f62979-18a7-2fc1-6f26-c4f84e19424f@gmail.com","threadId":"49204","inReplyTo":"CACsJy8ATsS6S5zib2FqJf1stPcGwSTO1qYBSz514Xu2GfJ4Apw@mail.gmail.com","subject":"Re: [PATCH v5 2/5] read-cache: load cache extensions on a worker thread","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-17T16:26:20Z","receivedAt":"2018-09-17T16:26:25Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/15/2018 6:22 AM, Duy Nguyen wrote:\n>> +index.threads::\n>> +       Specifies the number of threads to spawn when loading the index.\n>> +       This is meant to reduce index load time on multiprocessor machines.\n>> +       Specifying 0 or 'true' will cause Git to auto-detect the number of\n>> +       CPU's and set the number of threads accordingly. Defaults to 'true'.\n> \n> I'd rather this variable defaults to 0. Spawning threads have\n> associated cost and most projects out there are small enough that this\n> multi threading could just add more cost than gain. It only makes\n> sense to enable this on huge repos.\n> \n> Wait there's no way to disable this parallel reading? Does not sound\n> right. And  if ordinary numbers mean the number of threads then 0\n> should mean no threading. Auto detection could have a new keyword,\n> like 'auto'.\n> \n\nThe index.threads setting is patterned after the pack.threads setting \nfor consistency.  Specifying 1 (or 'false') will disable multithreading \nbut I will call that out explicitly in the documentation to make it more \nobvious.\n\nThe THREAD_COST logic is designed to ensure small repos don't incur more \ncost than gain.  If you have data on that logic that shows it isn't \nworking properly, I'm happy to change the logic as necessary.\n\n>> --- a/read-cache.c\n>> +++ b/read-cache.c\n>> @@ -23,6 +23,10 @@\n>>   #include \"split-index.h\"\n>>   #include \"utf8.h\"\n>>   #include \"fsmonitor.h\"\n>> +#ifndef NO_PTHREADS\n>> +#include <pthread.h>\n>> +#include <thread-utils.h>\n>> +#endif\n> \n> I don't think you're supposed to include system header files after\n> \"cache.h\". Including thread-utils.h should be enough (and it keeps the\n> exception of inclduing pthread.h in just one place). Please use\n> \"pthread-utils.h\" instead of <pthread-utils.h> which is usually for\n> system header files. And include ptherad-utils.h unconditionally.\n> \n\nThanks, I'll fix that.\n\n>>\n>>   /* Mask for the name length in ce_flags in the on-disk index */\n>>\n>> @@ -1898,6 +1902,46 @@ static unsigned long read_eoie_extension(void *mmap_, size_t mmap_size);\n>>   #endif\n>>   static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, unsigned long offset);\n>>\n>> +struct load_index_extensions\n>> +{\n>> +#ifndef NO_PTHREADS\n>> +       pthread_t pthread;\n>> +#endif\n>> +       struct index_state *istate;\n>> +       void *mmap;\n>> +       size_t mmap_size;\n>> +       unsigned long src_offset;\n>> +};\n>> +\n>> +static void *load_index_extensions(void *_data)\n>> +{\n>> +       struct load_index_extensions *p = _data;\n>> +       unsigned long src_offset = p->src_offset;\n>> +\n>> +       while (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n>> +               /* After an array of active_nr index entries,\n>> +                * there can be arbitrary number of extended\n>> +                * sections, each of which is prefixed with\n>> +                * extension name (4-byte) and section length\n>> +                * in 4-byte network byte order.\n>> +                */\n>> +               uint32_t extsize;\n>> +               memcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n>> +               extsize = ntohl(extsize);\n>> +               if (read_index_extension(p->istate,\n>> +                       (const char *)p->mmap + src_offset,\n>> +                       (char *)p->mmap + src_offset + 8,\n>> +                       extsize) < 0) {\n>> +                       munmap(p->mmap, p->mmap_size);\n>> +                       die(\"index file corrupt\");\n> \n> _()\n> \n\nYou're feedback style can be a bit abrupt and terse.  I _think_ what you \nare trying to say here is that the \"die\" call should use the _() macro \naround the string.\n\nThis is an edit of the previous code that loaded index extensions and \ndoesn't change the use of _(). I don't know the rules for when _() \nshould be used and didn't have any luck finding where it was documented \nso left it unchanged.\n\nFWIW, in this file alone there are 20 existing instances of die() or \ndie_errorno() and only two that use the _() macro.  A quick grep through \nthe source code shows thousands of die() calls the vast majority of \nwhich do not use the _() macro.  This appears to be an area that is \nunclear and inconsistent and could use some attention in a separate patch.\n\n\n>> +       /* if we created a thread, join it otherwise load the extensions on the primary thread */\n>> +#ifndef NO_PTHREADS\n>> +       if (extension_offset && pthread_join(p.pthread, NULL))\n>> +               die(_(\"unable to join load_index_extensions_thread\"));\n> \n> I guess the last _ is a typo and you wanted \"unable to join\n> load_index_extensions thread\". Please use die_errno() instead.\n> \n\nWhy should this be die_errorno() here?  All other instances of \npthread_join() failing in a fatal way use die(), not die_errorno().\n\n>> +#endif\n>> +       if (!extension_offset) {\n>> +               p.src_offset = src_offset;\n>> +               load_index_extensions(&p);\n>>          }\n>>          munmap(mmap, mmap_size);\n>>          return istate->cache_nr;\n>> --\n>> 2.18.0.windows.1\n>>\n> \n> \n"},{"id":"358290","messageId":"bfb8d836-86a4-a4f0-cdaf-aba25424d4cf@gmail.com","threadId":"49204","inReplyTo":"CACsJy8D-sM2SSfTMmsR0uKnP0FM9fGrN0Z-pH5irMsH1=-jrmQ@mail.gmail.com","subject":"Re: [PATCH v5 2/5] read-cache: load cache extensions on a worker thread","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-17T16:38:23Z","receivedAt":"2018-09-17T16:38:28Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/15/2018 6:24 AM, Duy Nguyen wrote:\n> On Sat, Sep 15, 2018 at 12:22 PM Duy Nguyen <pclouds@gmail.com> wrote:\n>>> @@ -1944,6 +1993,26 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>>>          istate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n>>>          istate->initialized = 1;\n>>>\n>>> +       p.istate = istate;\n>>> +       p.mmap = mmap;\n>>> +       p.mmap_size = mmap_size;\n>>> +\n>>> +#ifndef NO_PTHREADS\n>>> +       nr_threads = git_config_get_index_threads();\n>>> +       if (!nr_threads)\n>>> +               nr_threads = online_cpus();\n>>> +\n>>> +       if (nr_threads >= 2) {\n>>> +               extension_offset = read_eoie_extension(mmap, mmap_size);\n>>> +               if (extension_offset) {\n> \n> One more thing I forgot. If the extension area is small enough, then\n> we should not need to create a thread to parse extensions in parallel.\n> We should know roughly how much work we need because we know the total\n> size of all extensions.\n> \n\nThe only extensions I found to be significant enough to be helped by a \nseparate thread was the cache tree.  Since the size of the cache tree is \ndriven by the number of files in the repo, I think the existing \nTHREAD_COST logic (that comes in the next patch of the series) is a \nsufficient proxy.  Basically, if you have enough cache entries to be \nbenefited by threading, your extensions (driven by the cache tree) are \nprobably also big enough to warrant a thread.\n\n>>> +                       /* create a thread to load the index extensions */\n>>\n>> Pointless comment. It's pretty clear from the pthread_create() below\n>> thanks to good function naming. Please remove.\n>>\n>>> +                       p.src_offset = extension_offset;\n>>> +                       if (pthread_create(&p.pthread, NULL, load_index_extensions, &p))\n>>> +                               die(_(\"unable to create load_index_extensions_thread\"));\n>>> +               }\n>>> +       }\n>>> +#endif\n>>> +\n>>>          if (istate->version == 4) {\n>>>                  previous_name = &previous_name_buf;\n>>>                  mem_pool_init(&istate->ce_mem_pool,\n"},{"id":"358292","messageId":"CACsJy8AYq=FivKZ869tvjwuSc70tuaPV0HJ0aRp=VFbJBSpm=A@mail.gmail.com","threadId":"49204","inReplyTo":"78f62979-18a7-2fc1-6f26-c4f84e19424f@gmail.com","subject":"Re: [PATCH v5 2/5] read-cache: load cache extensions on a worker thread","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-17T16:45:32Z","receivedAt":"2018-09-17T16:46:01Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Sep 17, 2018 at 6:26 PM Ben Peart <peartben@gmail.com> wrote:\n>\n>\n>\n> On 9/15/2018 6:22 AM, Duy Nguyen wrote:\n> >> +index.threads::\n> >> +       Specifies the number of threads to spawn when loading the index.\n> >> +       This is meant to reduce index load time on multiprocessor machines.\n> >> +       Specifying 0 or 'true' will cause Git to auto-detect the number of\n> >> +       CPU's and set the number of threads accordingly. Defaults to 'true'.\n> >\n> > I'd rather this variable defaults to 0. Spawning threads have\n> > associated cost and most projects out there are small enough that this\n> > multi threading could just add more cost than gain. It only makes\n> > sense to enable this on huge repos.\n> >\n> > Wait there's no way to disable this parallel reading? Does not sound\n> > right. And  if ordinary numbers mean the number of threads then 0\n> > should mean no threading. Auto detection could have a new keyword,\n> > like 'auto'.\n> >\n>\n> The index.threads setting is patterned after the pack.threads setting\n> for consistency.  Specifying 1 (or 'false') will disable multithreading\n> but I will call that out explicitly in the documentation to make it more\n> obvious.\n>\n> The THREAD_COST logic is designed to ensure small repos don't incur more\n> cost than gain.  If you have data on that logic that shows it isn't\n> working properly, I'm happy to change the logic as necessary.\n\nTHREAD_COST does not apply to this extension thread if I remember correctly.\n\n> >> +static void *load_index_extensions(void *_data)\n> >> +{\n> >> +       struct load_index_extensions *p = _data;\n> >> +       unsigned long src_offset = p->src_offset;\n> >> +\n> >> +       while (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n> >> +               /* After an array of active_nr index entries,\n> >> +                * there can be arbitrary number of extended\n> >> +                * sections, each of which is prefixed with\n> >> +                * extension name (4-byte) and section length\n> >> +                * in 4-byte network byte order.\n> >> +                */\n> >> +               uint32_t extsize;\n> >> +               memcpy(&extsize, (char *)p->mmap + src_offset + 4, 4);\n> >> +               extsize = ntohl(extsize);\n> >> +               if (read_index_extension(p->istate,\n> >> +                       (const char *)p->mmap + src_offset,\n> >> +                       (char *)p->mmap + src_offset + 8,\n> >> +                       extsize) < 0) {\n> >> +                       munmap(p->mmap, p->mmap_size);\n> >> +                       die(\"index file corrupt\");\n> >\n> > _()\n> >\n>\n> You're feedback style can be a bit abrupt and terse.  I _think_ what you\n> are trying to say here is that the \"die\" call should use the _() macro\n> around the string.\n\nYes. Sorry I should have explained a bit better.\n\n> This is an edit of the previous code that loaded index extensions and\n> doesn't change the use of _(). I don't know the rules for when _()\n> should be used and didn't have any luck finding where it was documented\n> so left it unchanged.\n>\n> FWIW, in this file alone there are 20 existing instances of die() or\n> die_errorno() and only two that use the _() macro.  A quick grep through\n> the source code shows thousands of die() calls the vast majority of\n> which do not use the _() macro.  This appears to be an area that is\n> unclear and inconsistent and could use some attention in a separate patch.\n\nThis is one of the gray areas where we have to determine if the\nmessage should be translated or not. And it should be translated\nunless it's part of the plumbing output, to be consumed by scripts.\n\nI know there's lots of messages still untranslated. I'm trying to do\nsomething about that. But I cannot just go fix up all strings when you\nall keep adding more strings for me to go fix. When you add a new\nstring, please consider if it should be translated or not. In this\ncase since it already receives reviewer attention we should be able to\ndetermine it now, instead of delaying it for later.\n\n> >> +       /* if we created a thread, join it otherwise load the extensions on the primary thread */\n> >> +#ifndef NO_PTHREADS\n> >> +       if (extension_offset && pthread_join(p.pthread, NULL))\n> >> +               die(_(\"unable to join load_index_extensions_thread\"));\n> >\n> > I guess the last _ is a typo and you wanted \"unable to join\n> > load_index_extensions thread\". Please use die_errno() instead.\n> >\n>\n> Why should this be die_errorno() here?  All other instances of\n> pthread_join() failing in a fatal way use die(), not die_errorno().\n\nThat argument does not fly well in my opinion. I read the man page and\nit listed the error codes, which made me think that we need to use\ndie_errno() to show the error. My mistake though is the error is\nreturned as the return value, not in errno, so die_errno() would not\ncatch it. But we could still do something like\n\n    int ret = pthread_join();\n    die(_(\"blah blah: %s\"), strerror(ret));\n\nOther code can also be improved, but that's a separate issue.\n-- \nDuy\n"},{"id":"358297","messageId":"xmqqwork6nyn.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"CACsJy8DtYrB99_7GyZaHsLTG8Ff7Mt_hENwhLE6_x6b5zzQ9Hg@mail.gmail.com","subject":"Re: [PATCH v5 2/5] read-cache: load cache extensions on a worker thread","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-17T17:19:44Z","receivedAt":"2018-09-17T17:19:50Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> On Sat, Sep 15, 2018 at 12:22 PM Duy Nguyen <pclouds@gmail.com> wrote:\n>> Wait there's no way to disable this parallel reading? Does not sound\n>> right. And  if ordinary numbers mean the number of threads then 0\n>> should mean no threading. Auto detection could have a new keyword,\n>> like 'auto'.\n>\n> My bad. Disabling threading means _1_ thread. What was I thinking...\n\nI did the same during my earlier review.  It seems that it somehow\nis unintuitive to us that we do not specify how many _extra_ threads\nof control we dedicate to ;-).\n"},{"id":"358299","messageId":"e021377f-fbde-e3e9-7b62-c2a2d33cf1eb@gmail.com","threadId":"49204","inReplyTo":"CACsJy8CjWD_CHwv5NoURLt9is7-UUDpKDo-3EcM28imznZAOpA@mail.gmail.com","subject":"Re: [PATCH v5 3/5] read-cache: load cache entries on worker threads","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-17T17:25:42Z","receivedAt":"2018-09-17T17:25:46Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/15/2018 6:31 AM, Duy Nguyen wrote:\n> On Wed, Sep 12, 2018 at 6:18 PM Ben Peart <benpeart@microsoft.com> wrote:\n>>\n>> This patch helps address the CPU cost of loading the index by creating\n>> multiple threads to divide the work of loading and converting the cache\n>> entries across all available CPU cores.\n>>\n>> It accomplishes this by having the primary thread loop across the index file\n>> tracking the offset and (for V4 indexes) expanding the name. It creates a\n>> thread to process each block of entries as it comes to them.\n>>\n>> I used p0002-read-cache.sh to generate some performance data:\n>>\n>> Test w/100,000 files                Baseline         Parallel entries\n>> ---------------------------------------------------------------------------\n>> read_cache/discard_cache 1000 times 14.08(0.01+0.10) 9.72(0.03+0.06) -31.0%\n>>\n>> Test w/1,000,000 files              Baseline         Parallel entries\n>> ------------------------------------------------------------------------------\n>> read_cache/discard_cache 1000 times 202.95(0.01+0.07) 154.14(0.03+0.06) -24.1%\n> \n> The numbers here and the previous patch to load extensions in parallel\n> are exactly the same. What do these numbers mean? With both changes?\n> \n\nIt means I messed up when creating my commit message for the extension \npatch and copy/pasted the wrong numbers.  Yes, these numbers are with \nboth changes (the correct numbers for the extension only are not as good).\n"},{"id":"358301","messageId":"xmqqsh286nfs.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"CACsJy8DEvLJYBm0P1VtvKFD-CAo6_4Z13dBHWkuuAavghbGBag@mail.gmail.com","subject":"Re: [PATCH v5 1/5] eoie: add End of Index Entry (EOIE) extension","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-17T17:31:03Z","receivedAt":"2018-09-17T17:31:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> But it _is_ available now. If you need it, you write the extension\n> out.\n\nAre you arguing for making it omitted when it is not needed (e.g.\nsmall enough index file)?  IOW, did you mean \"If you do not need it,\nyou do not write it out\" by the above?\n\nI do not think overhead of writing (or preparing to write) the\nextension for a small index file is by definition small enough ;-).\n\nI do not think the configuration that decides if the reader side\nuses parallel reading should have any say in the decision to write\n(or omit) the extension, by the way.\n\n\n"},{"id":"358302","messageId":"CACsJy8CqaEGDaEAgp1EspR+BwyHB6YSPoppZ2t5M+qxg77hULg@mail.gmail.com","threadId":"49204","inReplyTo":"xmqqsh286nfs.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v5 1/5] eoie: add End of Index Entry (EOIE) extension","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-17T17:38:43Z","receivedAt":"2018-09-17T17:39:12Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Sep 17, 2018 at 7:31 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Duy Nguyen <pclouds@gmail.com> writes:\n>\n> > But it _is_ available now. If you need it, you write the extension\n> > out.\n>\n> Are you arguing for making it omitted when it is not needed (e.g.\n> small enough index file)?  IOW, did you mean \"If you do not need it,\n> you do not write it out\" by the above?\n\nYes I did.\n\n> I do not think overhead of writing (or preparing to write) the\n> extension for a small index file is by definition small enough ;-).\n\nGood point.\n\nI get annoyed by the \"ignoring unknown extension xxx\" messages while\ntesting though (not just this extension) and I think it will be the\nsame for other git implementations. But perhaps other implementations\njust silently drop the extension. Most of the extensions we have added\nso far (except the ancient 'TREE') are optional and are probably not\npresent 99% of time when a different git impl reads an index created\nby C Git. This 'EIOE' may be a good test then to see if they follow\nthe \"ignore optional extensions\" rule since it will always appear in\nnew C Git releases.\n-- \nDuy\n"},{"id":"358316","messageId":"f8f021ce-5a3d-a5f3-ef47-e9cc94460b24@gmail.com","threadId":"49204","inReplyTo":"20180915110951.GA17817@duynguyen.home","subject":"Re: [PATCH v5 3/5] read-cache: load cache entries on worker threads","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-17T18:52:05Z","receivedAt":"2018-09-17T18:52:10Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/15/2018 7:09 AM, Duy Nguyen wrote:\n> On Sat, Sep 15, 2018 at 01:07:46PM +0200, Duy Nguyen wrote:\n>> 12:50:00.084237 read-cache.c:1721       start loading index\n>> 12:50:00.119941 read-cache.c:1943       performance: 0.034778758 s: loaded all extensions (1667075 bytes)\n>> 12:50:00.185352 read-cache.c:2029       performance: 0.100152079 s: loaded 367110 entries\n>> 12:50:00.189683 read-cache.c:2126       performance: 0.104566615 s: finished scanning all entries\n>> 12:50:00.217900 read-cache.c:2029       performance: 0.082309193 s: loaded 367110 entries\n>> 12:50:00.259969 read-cache.c:2029       performance: 0.070257130 s: loaded 367108 entries\n>> 12:50:00.263662 read-cache.c:2278       performance: 0.179344458 s: read cache .git/index\n> \n> The previous mail wraps these lines and make it a bit hard to read. Corrected now.\n> \n> --\n> Duy\n> \n\nInteresting!  Clearly the data shape makes a big difference here as I \nhad run a similar test but in my case, the extensions thread actually \nfinished last (and it's cost is what drove me to move that onto a \nseparate thread that starts first).\n\nPurpose\t    \t\t\tFirst\tLast\tDuration\nload_index_extensions_thread\t719.40\t968.50\t249.10\nload_cache_entries_thread\t718.89\t738.65\t19.76\nload_cache_entries_thread\t730.39\t753.83\t23.43\nload_cache_entries_thread\t741.23\t751.23\t10.00\nload_cache_entries_thread\t751.93\t780.88\t28.95\nload_cache_entries_thread\t763.60\t791.31\t27.72\nload_cache_entries_thread\t773.46\t783.46\t10.00\nload_cache_entries_thread\t783.96\t794.28\t10.32\nload_cache_entries_thread\t795.61\t805.52\t9.91\nload_cache_entries_thread\t805.99\t827.21\t21.22\nload_cache_entries_thread\t816.85\t826.85\t10.00\nload_cache_entries_thread\t827.03\t837.96\t10.93\n\nIn my tests, the scanning thread clearly delayed the later ce threads \nbut given the extension was so slow, it didn't impact the overall time \nnearly as much as your case.\n\nI completely agree that the optimal solution would be to go back to my \noriginal patch/design.  It eliminates the overhead of the scanning \nthread entirely and allows all threads to start at the same time. This \nwould ensure the best performance whether the extensions were the \nlongest thread or the cache entry threads took the longest.\n\nI ran out of time and energy last year so dropped it to work on other \ntasks.  I appreciate your offer of help. Perhaps between the two of us \nwe could successfully get it through the mailing list this time. :-) \nLet me go back and see what it would take to combine the current EOIE \npatch with the older IEOT patch.\n\nI'm also intrigued with your observation that over committing the cpu \nactually results in time savings.  I hadn't tested that.  It looks like \nthat could have a positive impact on the overall time and warrant a \nchange to the default nr_threads logic.\n"},{"id":"358318","messageId":"xmqqfty86iwv.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"CACsJy8CqaEGDaEAgp1EspR+BwyHB6YSPoppZ2t5M+qxg77hULg@mail.gmail.com","subject":"Re: [PATCH v5 1/5] eoie: add End of Index Entry (EOIE) extension","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-17T19:08:48Z","receivedAt":"2018-09-17T19:08:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> I get annoyed by the \"ignoring unknown extension xxx\" messages while\n> testing though (not just this extension) and I think it will be the\n> same for other git implementations. But perhaps other implementations\n> just silently drop the extension. Most of the extensions we have added\n> so far (except the ancient 'TREE') are optional and are probably not\n\nMost of the index extensions are optional, including TREE.  I think\n\"link\" is the only one that the readers that do not understand it\nare told to abort without causing damage.\n\n> present 99% of time when a different git impl reads an index created\n> by C Git. This 'EIOE' may be a good test then to see if they follow\n> the \"ignore optional extensions\" rule since it will always appear in\n> new C Git releases.\n\nI think we probably should squelch \"ignoring unknown\" unless some\nsort of GIT_TRACE/DEBUG switch is set.\n\nPatches welcome ;-)\n\nThanks.\n\n"},{"id":"358337","messageId":"xmqqh8in6c98.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"CACsJy8ATsS6S5zib2FqJf1stPcGwSTO1qYBSz514Xu2GfJ4Apw@mail.gmail.com","subject":"Re: [PATCH v5 2/5] read-cache: load cache extensions on a worker thread","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-17T21:32:35Z","receivedAt":"2018-09-17T21:32:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n>> diff --git a/read-cache.c b/read-cache.c\n>> index 858935f123..b203eebb44 100644\n>> --- a/read-cache.c\n>> +++ b/read-cache.c\n>> @@ -23,6 +23,10 @@\n>>  #include \"split-index.h\"\n>>  #include \"utf8.h\"\n>>  #include \"fsmonitor.h\"\n>> +#ifndef NO_PTHREADS\n>> +#include <pthread.h>\n>> +#include <thread-utils.h>\n>> +#endif\n>\n> I don't think you're supposed to include system header files after\n> \"cache.h\". Including thread-utils.h should be enough (and it keeps the\n> exception of inclduing pthread.h in just one place). Please use\n> \"pthread-utils.h\" instead of <pthread-utils.h> which is usually for\n> system header files. And include ptherad-utils.h unconditionally.\n\nAll correct except for s/p\\(thread-utils\\)/\\1/g;\nSorry for missing this during my earlier review.\n\nThanks.\n\n\n"},{"id":"358989","messageId":"20180926195442.1380-1-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180823154053.20212-1-benpeart@microsoft.com","subject":"[PATCH v6 0/7] speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-26T19:54:35Z","receivedAt":"2018-09-26T19:54:56Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\nBase Ref: master\nWeb-Diff: https://github.com/benpeart/git/commit/a0300882d4\nCheckout: git fetch https://github.com/benpeart/git read-index-multithread-v6 && git checkout a0300882d4\n\n\nThis iteration brings back the Index Entry Offset Table (IEOT) extension\nwhich enables us to multi-thread the cache entry parsing without having\nthe primary thread have to scan all the entries first.  In cases where the\ncache entry parsing is the most expensive part, this yields some additional\nsavings.\n\nUsing p0002-read-cache.sh to generate some performance numbers shows how\neach of the various patches contribute to the overall performance win.\n\n\nTest w/100,000 files    Baseline  Optimize V4    Extensions     Entries\n----------------------------------------------------------------------------\n0002.1: read_cache      22.36     18.74 -16.2%   18.64 -16.6%   12.63 -43.5%\n\nTest w/1,000,000 files  Baseline  Optimize V4    Extensions     Entries\n-----------------------------------------------------------------------------\n0002.1: read_cache      304.40    270.70 -11.1%  195.50 -35.8%  204.82 -32.7%\n\nNote that on the 1,000,000 files case, multi-threading the cache entry parsing\ndoes not yield a performance win.  This is because the cost to parse the\nindex extensions in this repo, far outweigh the cost of loading the cache\nentries.\n\nName                            First    Last\t  Elapsed\t\nload_index_extensions()\t\t629.001  870.244  241.243\t\nload_cache_entries_thread()\t683.911  723.199  39.288\t\nload_cache_entries_thread()\t686.206  723.512  37.306\t\nload_cache_entries_thread()\t686.43   722.596  36.166\t\nload_cache_entries_thread()\t684.998  718.74   33.742\t\nload_cache_entries_thread()\t685.035  718.698  33.663\t\nload_cache_entries_thread()\t686.557  709.545  22.988\t\nload_cache_entries_thread()\t684.533  703.536  19.003\t\nload_cache_entries_thread()\t684.537  703.521  18.984\t\nload_cache_entries_thread()\t685.062  703.774  18.712\t\nload_cache_entries_thread()\t685.42   703.416  17.996\t\nload_cache_entries_thread()\t648.604  664.496  15.892\t\n\t\t\t\t\n293.74 Total load_cache_entries_thread()\n\nThe high cost of parsing the index extensions is driven by the cache tree\nand the untracked cache extensions. As this is currently the longest pole,\nany reduction in this time will reduce the overall index load times so is\nworth further investigation in another patch series.\n\nName                                    First    Last     Elapsed\n|   + git!read_index_extension     \t684.052  870.244  186.192\n|    + git!cache_tree_read         \t684.052  797.801  113.749\n|    + git!read_untracked_extension\t797.801  870.244  72.443\n\nOne option would be to load each extension on a separate thread but I\nbelieve that is overkill for the vast majority of repos.  Instead, some\noptimization of the loading code for these two extensions is probably worth\nlooking into as a quick examination shows that the bulk of the time for both\nof them is spent in xcalloc().\n\n\n### Patches\n\nBen Peart (6):\n  read-cache: clean up casting and byte decoding\n  eoie: add End of Index Entry (EOIE) extension\n  config: add new index.threads config setting\n  read-cache: load cache extensions on a worker thread\n  ieot: add Index Entry Offset Table (IEOT) extension\n  read-cache: load cache entries on worker threads\n\nNguyễn Thái Ngọc Duy (1):\n  read-cache.c: optimize reading index format v4\n\n Documentation/config.txt                 |   7 +\n Documentation/technical/index-format.txt |  41 ++\n config.c                                 |  18 +\n config.h                                 |   1 +\n read-cache.c                             | 741 +++++++++++++++++++----\n t/README                                 |  10 +\n t/t1700-split-index.sh                   |   2 +\n 7 files changed, 705 insertions(+), 115 deletions(-)\n\n\nbase-commit: fe8321ec057f9231c26c29b364721568e58040f7\n-- \n2.18.0.windows.1\n\n\n"},{"id":"358990","messageId":"20180926195442.1380-2-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180926195442.1380-1-benpeart@microsoft.com","subject":"[PATCH v6 1/7] read-cache.c: optimize reading index format v4","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-26T19:54:36Z","receivedAt":"2018-09-26T19:54:58Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\nIndex format v4 requires some more computation to assemble a path\nbased on a previous one. The current code is not very efficient\nbecause\n\n - it doubles memory copy, we assemble the final path in a temporary\n   first before putting it back to a cache_entry\n\n - strbuf_remove() in expand_name_field() is not exactly a good fit\n   for stripping a part at the end, _setlen() would do the same job\n   and is much cheaper.\n\n - the open-coded loop to find the end of the string in\n   expand_name_field() can't beat an optimized strlen()\n\nThis patch avoids the temporary buffer and writes directly to the new\ncache_entry, which addresses the first two points. The last point\ncould also be avoided if the total string length fits in the first 12\nbits of ce_flags, if not we fall back to strlen().\n\nRunning \"test-tool read-cache 100\" on webkit.git (275k files), reading\nv2 only takes 4.226 seconds, while v4 takes 5.711 seconds, 35% more\ntime. The patch reduces read time on v4 to 4.319 seconds.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c | 128 ++++++++++++++++++++++++---------------------------\n 1 file changed, 60 insertions(+), 68 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 8d04d78a58..583a4fb1f8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1713,63 +1713,24 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *cache_entry_from_ondisk(struct mem_pool *mem_pool,\n-\t\t\t\t\t\t   struct ondisk_cache_entry *ondisk,\n-\t\t\t\t\t\t   unsigned int flags,\n-\t\t\t\t\t\t   const char *name,\n-\t\t\t\t\t\t   size_t len)\n-{\n-\tstruct cache_entry *ce = mem_pool__ce_alloc(mem_pool, len);\n-\n-\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n-\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n-\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n-\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n-\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n-\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n-\tce->ce_mode  = get_be32(&ondisk->mode);\n-\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n-\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n-\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n-\tce->ce_flags = flags & ~CE_NAMEMASK;\n-\tce->ce_namelen = len;\n-\tce->index = 0;\n-\thashcpy(ce->oid.hash, ondisk->sha1);\n-\tmemcpy(ce->name, name, len);\n-\tce->name[len] = '\\0';\n-\treturn ce;\n-}\n-\n-/*\n- * Adjacent cache entries tend to share the leading paths, so it makes\n- * sense to only store the differences in later entries.  In the v4\n- * on-disk format of the index, each on-disk cache entry stores the\n- * number of bytes to be stripped from the end of the previous name,\n- * and the bytes to append to the result, to come up with its name.\n- */\n-static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n-{\n-\tconst unsigned char *ep, *cp = (const unsigned char *)cp_;\n-\tsize_t len = decode_varint(&cp);\n-\n-\tif (name->len < len)\n-\t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n-\tstrbuf_add(name, cp, ep - cp);\n-\treturn (const char *)ep + 1 - cp_;\n-}\n-\n-static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n+static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n-\t\t\t\t\t    struct strbuf *previous_name)\n+\t\t\t\t\t    const struct cache_entry *previous_ce)\n {\n \tstruct cache_entry *ce;\n \tsize_t len;\n \tconst char *name;\n \tunsigned int flags;\n+\tsize_t copy_len;\n+\t/*\n+\t * Adjacent cache entries tend to share the leading paths, so it makes\n+\t * sense to only store the differences in later entries.  In the v4\n+\t * on-disk format of the index, each on-disk cache entry stores the\n+\t * number of bytes to be stripped from the end of the previous name,\n+\t * and the bytes to append to the result, to come up with its name.\n+\t */\n+\tint expand_name_field = istate->version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1789,21 +1750,54 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \telse\n \t\tname = ondisk->name;\n \n-\tif (!previous_name) {\n-\t\t/* v3 and earlier */\n-\t\tif (len == CE_NAMEMASK)\n-\t\t\tlen = strlen(name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags, name, len);\n+\tif (expand_name_field) {\n+\t\tconst unsigned char *cp = (const unsigned char *)name;\n+\t\tsize_t strip_len, previous_len;\n \n-\t\t*ent_size = ondisk_ce_size(ce);\n-\t} else {\n-\t\tunsigned long consumed;\n-\t\tconsumed = expand_name_field(previous_name, name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags,\n-\t\t\t\t\t     previous_name->buf,\n-\t\t\t\t\t     previous_name->len);\n+\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n+\t\tstrip_len = decode_varint(&cp);\n+\t\tif (previous_len < strip_len) {\n+\t\t\tif (previous_ce)\n+\t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n+\t\t\t\t    previous_ce->name);\n+\t\t\telse\n+\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t}\n+\t\tcopy_len = previous_len - strip_len;\n+\t\tname = (const char *)cp;\n+\t}\n+\n+\tif (len == CE_NAMEMASK) {\n+\t\tlen = strlen(name);\n+\t\tif (expand_name_field)\n+\t\t\tlen += copy_len;\n+\t}\n+\n+\tce = mem_pool__ce_alloc(istate->ce_mem_pool, len);\n+\n+\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n+\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n+\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n+\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n+\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n+\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n+\tce->ce_mode  = get_be32(&ondisk->mode);\n+\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n+\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n+\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n+\tce->ce_flags = flags & ~CE_NAMEMASK;\n+\tce->ce_namelen = len;\n+\tce->index = 0;\n+\thashcpy(ce->oid.hash, ondisk->sha1);\n \n-\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n+\tif (expand_name_field) {\n+\t\tif (copy_len)\n+\t\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n+\t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n+\t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n+\t} else {\n+\t\tmemcpy(ce->name, name, len + 1);\n+\t\t*ent_size = ondisk_ce_size(ce);\n \t}\n \treturn ce;\n }\n@@ -1898,7 +1892,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tconst struct cache_entry *previous_ce = NULL;\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1936,11 +1930,9 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->initialized = 1;\n \n \tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n \t} else {\n-\t\tprevious_name = NULL;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n \t}\n@@ -1952,12 +1944,12 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tunsigned long consumed;\n \n \t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n \t\tsrc_offset += consumed;\n+\t\tprevious_ce = ce;\n \t}\n-\tstrbuf_release(&previous_name_buf);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-- \n2.18.0.windows.1\n\n"},{"id":"358991","messageId":"20180926195442.1380-3-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180926195442.1380-1-benpeart@microsoft.com","subject":"[PATCH v6 2/7] read-cache: clean up casting and byte decoding","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-26T19:54:37Z","receivedAt":"2018-09-26T19:54:58Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch does a clean up pass to minimize the casting required to work\nwith the memory mapped index (mmap).\n\nIt also makes the decoding of network byte order more consistent by using\nget_be32() where possible.\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 23 +++++++++++------------\n 1 file changed, 11 insertions(+), 12 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 583a4fb1f8..6ba99e2c96 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1650,7 +1650,7 @@ int verify_index_checksum;\n /* Allow fsck to force verification of the cache entry order. */\n int verify_ce_order;\n \n-static int verify_hdr(struct cache_header *hdr, unsigned long size)\n+static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n {\n \tgit_hash_ctx c;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n@@ -1674,7 +1674,7 @@ static int verify_hdr(struct cache_header *hdr, unsigned long size)\n }\n \n static int read_index_extension(struct index_state *istate,\n-\t\t\t\tconst char *ext, void *data, unsigned long sz)\n+\t\t\t\tconst char *ext, const char *data, unsigned long sz)\n {\n \tswitch (CACHE_EXT(ext)) {\n \tcase CACHE_EXT_TREE:\n@@ -1889,8 +1889,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tint fd, i;\n \tstruct stat st;\n \tunsigned long src_offset;\n-\tstruct cache_header *hdr;\n-\tvoid *mmap;\n+\tconst struct cache_header *hdr;\n+\tconst char *mmap;\n \tsize_t mmap_size;\n \tconst struct cache_entry *previous_ce = NULL;\n \n@@ -1918,7 +1918,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tdie_errno(\"unable to map index file\");\n \tclose(fd);\n \n-\thdr = mmap;\n+\thdr = (const struct cache_header *)mmap;\n \tif (verify_hdr(hdr, mmap_size) < 0)\n \t\tgoto unmap;\n \n@@ -1943,7 +1943,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tstruct cache_entry *ce;\n \t\tunsigned long consumed;\n \n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n \t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n@@ -1961,21 +1961,20 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\t * in 4-byte network byte order.\n \t\t */\n \t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n+\t\textsize = get_be32(mmap + src_offset + 4);\n \t\tif (read_index_extension(istate,\n-\t\t\t\t\t (const char *) mmap + src_offset,\n-\t\t\t\t\t (char *) mmap + src_offset + 8,\n+\t\t\t\t\t mmap + src_offset,\n+\t\t\t\t\t mmap + src_offset + 8,\n \t\t\t\t\t extsize) < 0)\n \t\t\tgoto unmap;\n \t\tsrc_offset += 8;\n \t\tsrc_offset += extsize;\n \t}\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \treturn istate->cache_nr;\n \n unmap:\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \tdie(\"index file corrupt\");\n }\n \n-- \n2.18.0.windows.1\n\n"},{"id":"358992","messageId":"20180926195442.1380-4-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180926195442.1380-1-benpeart@microsoft.com","subject":"[PATCH v6 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-26T19:54:38Z","receivedAt":"2018-09-26T19:55:00Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"The End of Index Entry (EOIE) is used to locate the end of the variable\nlength index entries and the beginning of the extensions. Code can take\nadvantage of this to quickly locate the index extensions without having\nto parse through all of the index entries.\n\nBecause it must be able to be loaded before the variable length cache\nentries and other index extensions, this extension must be written last.\nThe signature for this extension is { 'E', 'O', 'I', 'E' }.\n\nThe extension consists of:\n\n- 32-bit offset to the end of the index entries\n\n- 160-bit SHA-1 over the extension types and their sizes (but not\ntheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\nlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\nthen the hash would be:\n\nSHA-1(\"TREE\" + <binary representation of N> +\n\t\"REUC\" + <binary representation of M>)\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n Documentation/technical/index-format.txt |  23 ++++\n read-cache.c                             | 151 +++++++++++++++++++++--\n t/README                                 |   5 +\n t/t1700-split-index.sh                   |   1 +\n 4 files changed, 172 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\nindex db3572626b..6bc2d90f7f 100644\n--- a/Documentation/technical/index-format.txt\n+++ b/Documentation/technical/index-format.txt\n@@ -314,3 +314,26 @@ The remaining data of each directory block is grouped by type:\n \n   - An ewah bitmap, the n-th bit indicates whether the n-th index entry\n     is not CE_FSMONITOR_VALID.\n+\n+== End of Index Entry\n+\n+  The End of Index Entry (EOIE) is used to locate the end of the variable\n+  length index entries and the begining of the extensions. Code can take\n+  advantage of this to quickly locate the index extensions without having\n+  to parse through all of the index entries.\n+\n+  Because it must be able to be loaded before the variable length cache\n+  entries and other index extensions, this extension must be written last.\n+  The signature for this extension is { 'E', 'O', 'I', 'E' }.\n+\n+  The extension consists of:\n+\n+  - 32-bit offset to the end of the index entries\n+\n+  - 160-bit SHA-1 over the extension types and their sizes (but not\n+\ttheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\tlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\tthen the hash would be:\n+\n+\tSHA-1(\"TREE\" + <binary representation of N> +\n+\t\t\"REUC\" + <binary representation of M>)\ndiff --git a/read-cache.c b/read-cache.c\nindex 6ba99e2c96..80255d3088 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -43,6 +43,7 @@\n #define CACHE_EXT_LINK 0x6c696e6b\t  /* \"link\" */\n #define CACHE_EXT_UNTRACKED 0x554E5452\t  /* \"UNTR\" */\n #define CACHE_EXT_FSMONITOR 0x46534D4E\t  /* \"FSMN\" */\n+#define CACHE_EXT_ENDOFINDEXENTRIES 0x454F4945\t/* \"EOIE\" */\n \n /* changes that can be kept in $GIT_DIR/index (basically all extensions) */\n #define EXTMASK (RESOLVE_UNDO_CHANGED | CACHE_TREE_CHANGED | \\\n@@ -1693,6 +1694,9 @@ static int read_index_extension(struct index_state *istate,\n \tcase CACHE_EXT_FSMONITOR:\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n+\tcase CACHE_EXT_ENDOFINDEXENTRIES:\n+\t\t/* already handled in do_read_index() */\n+\t\tbreak;\n \tdefault:\n \t\tif (*ext < 'A' || 'Z' < *ext)\n \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n@@ -1883,6 +1887,9 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -2190,11 +2197,15 @@ static int ce_write(git_hash_ctx *context, int fd, void *data, unsigned int len)\n \treturn 0;\n }\n \n-static int write_index_ext_header(git_hash_ctx *context, int fd,\n-\t\t\t\t  unsigned int ext, unsigned int sz)\n+static int write_index_ext_header(git_hash_ctx *context, git_hash_ctx *eoie_context,\n+\t\t\t\t  int fd, unsigned int ext, unsigned int sz)\n {\n \text = htonl(ext);\n \tsz = htonl(sz);\n+\tif (eoie_context) {\n+\t\tthe_hash_algo->update_fn(eoie_context, &ext, 4);\n+\t\tthe_hash_algo->update_fn(eoie_context, &sz, 4);\n+\t}\n \treturn ((ce_write(context, fd, &ext, 4) < 0) ||\n \t\t(ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n }\n@@ -2437,7 +2448,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n {\n \tuint64_t start = getnanotime();\n \tint newfd = tempfile->fd;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, eoie_c;\n \tstruct cache_header hdr;\n \tint i, err = 0, removed, extended, hdr_version;\n \tstruct cache_entry **cache = istate->cache;\n@@ -2446,6 +2457,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct ondisk_cache_entry_extended ondisk;\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n+\toff_t offset;\n \n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n@@ -2479,6 +2491,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tif (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n \t\treturn -1;\n \n+\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n \tprevious_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n \n \tfor (i = 0; i < entries; i++) {\n@@ -2512,11 +2525,14 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\treturn err;\n \n \t/* Write extension data here */\n+\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\tthe_hash_algo->init_fn(&eoie_c);\n+\n \tif (!strip_extensions && istate->split_index) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\terr = write_link_extension(&sb, istate) < 0 ||\n-\t\t\twrite_index_ext_header(&c, newfd, CACHE_EXT_LINK,\n+\t\t\twrite_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_LINK,\n \t\t\t\t\t       sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2527,7 +2543,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tcache_tree_write(&sb, istate->cache_tree);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_TREE, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_TREE, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2537,7 +2553,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tresolve_undo_write(&sb, istate->resolve_undo);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_RESOLVE_UNDO,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_RESOLVE_UNDO,\n \t\t\t\t\t     sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2548,7 +2564,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_untracked_extension(&sb, istate->untracked);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_UNTRACKED,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_UNTRACKED,\n \t\t\t\t\t     sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2559,7 +2575,23 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_fsmonitor_extension(&sb, istate);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+\n+\t/*\n+\t * CACHE_EXT_ENDOFINDEXENTRIES must be written as the last entry before the SHA1\n+\t * so that it can be found and processed before all the index entries are\n+\t * read.\n+\t */\n+\tif (!strip_extensions && offset && !git_env_bool(\"GIT_TEST_DISABLE_EOIE\", 0)) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\twrite_eoie_extension(&sb, &eoie_c, offset);\n+\t\terr = write_index_ext_header(&c, NULL, newfd, CACHE_EXT_ENDOFINDEXENTRIES, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2975,3 +3007,106 @@ int should_validate_cache_entries(void)\n \n \treturn validate_index_cache_entries;\n }\n+\n+#define EOIE_SIZE (4 + GIT_SHA1_RAWSZ) /* <4-byte offset> + <20-byte hash> */\n+#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n+\n+static size_t read_eoie_extension(const char *mmap, size_t mmap_size)\n+{\n+\t/*\n+\t * The end of index entries (EOIE) extension is guaranteed to be last\n+\t * so that it can be found by scanning backwards from the EOF.\n+\t *\n+\t * \"EOIE\"\n+\t * <4-byte length>\n+\t * <4-byte offset>\n+\t * <20-byte hash>\n+\t */\n+\tconst char *index, *eoie;\n+\tuint32_t extsize;\n+\tsize_t offset, src_offset;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\tgit_hash_ctx c;\n+\n+\t/* ensure we have an index big enough to contain an EOIE extension */\n+\tif (mmap_size < sizeof(struct cache_header) + EOIE_SIZE_WITH_HEADER + the_hash_algo->rawsz)\n+\t\treturn 0;\n+\n+\t/* validate the extension signature */\n+\tindex = eoie = mmap + mmap_size - EOIE_SIZE_WITH_HEADER - the_hash_algo->rawsz;\n+\tif (CACHE_EXT(index) != CACHE_EXT_ENDOFINDEXENTRIES)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/* validate the extension size */\n+\textsize = get_be32(index);\n+\tif (extsize != EOIE_SIZE)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * Validate the offset we're going to look for the first extension\n+\t * signature is after the index header and before the eoie extension.\n+\t */\n+\toffset = get_be32(index);\n+\tif (mmap + offset < mmap + sizeof(struct cache_header))\n+\t\treturn 0;\n+\tif (mmap + offset >= eoie)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * The hash is computed over extension types and their sizes (but not\n+\t * their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\t * long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\t * then the hash would be:\n+\t *\n+\t * SHA-1(\"TREE\" + <binary representation of N> +\n+\t *\t \"REUC\" + <binary representation of M>)\n+\t */\n+\tsrc_offset = offset;\n+\tthe_hash_algo->init_fn(&c);\n+\twhile (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\n+\t\t/* verify the extension size isn't so large it will wrap around */\n+\t\tif (src_offset + 8 + extsize < src_offset)\n+\t\t\treturn 0;\n+\n+\t\tthe_hash_algo->update_fn(&c, mmap + src_offset, 8);\n+\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\tthe_hash_algo->final_fn(hash, &c);\n+\tif (hashcmp(hash, (const unsigned char *)index))\n+\t\treturn 0;\n+\n+\t/* Validate that the extension offsets returned us back to the eoie extension. */\n+\tif (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n+\t\treturn 0;\n+\n+\treturn offset;\n+}\n+\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset)\n+{\n+\tuint32_t buffer;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\n+\t/* offset */\n+\tput_be32(&buffer, offset);\n+\tstrbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+\t/* hash */\n+\tthe_hash_algo->final_fn(hash, eoie_context);\n+\tstrbuf_add(sb, hash, the_hash_algo->rawsz);\n+}\ndiff --git a/t/README b/t/README\nindex 3ea6c85460..aa33ac4f26 100644\n--- a/t/README\n+++ b/t/README\n@@ -327,6 +327,11 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_DISABLE_EOIE=<boolean> disables writing the EOIE extension.\n+This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n+as they currently hard code SHA values for the index which are no longer\n+valid due to the addition of the EOIE extension.\n+\n Naming Tests\n ------------\n \ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex be22398a85..1f168378c8 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -7,6 +7,7 @@ test_description='split index mode tests'\n # We need total control of index splitting here\n sane_unset GIT_TEST_SPLIT_INDEX\n sane_unset GIT_FSMONITOR_TEST\n+GIT_TEST_DISABLE_EOIE=true; export GIT_TEST_DISABLE_EOIE\n \n test_expect_success 'enable split index' '\n \tgit config splitIndex.maxPercentChange 100 &&\n-- \n2.18.0.windows.1\n\n"},{"id":"358993","messageId":"20180926195442.1380-5-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180926195442.1380-1-benpeart@microsoft.com","subject":"[PATCH v6 4/7] config: add new index.threads config setting","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-26T19:54:39Z","receivedAt":"2018-09-26T19:55:01Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"Add support for a new index.threads config setting which will be used to\ncontrol the threading code in do_read_index().  A value of 0 will tell the\nindex code to automatically determine the correct number of threads to use.\nA value of 1 will make the code single threaded.  A value greater than 1\nwill set the maximum number of threads to use.\n\nFor testing purposes, this setting can be overwritten by setting the\nGIT_TEST_INDEX_THREADS=<n> environment variable to a value greater than 0.\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n Documentation/config.txt |  7 +++++++\n config.c                 | 18 ++++++++++++++++++\n config.h                 |  1 +\n t/README                 |  5 +++++\n t/t1700-split-index.sh   |  1 +\n 5 files changed, 32 insertions(+)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex ad0f4510c3..8fd973b76b 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -2413,6 +2413,13 @@ imap::\n \tThe configuration variables in the 'imap' section are described\n \tin linkgit:git-imap-send[1].\n \n+index.threads::\n+\tSpecifies the number of threads to spawn when loading the index.\n+\tThis is meant to reduce index load time on multiprocessor machines.\n+\tSpecifying 0 or 'true' will cause Git to auto-detect the number of\n+\tCPU's and set the number of threads accordingly. Specifying 1 or\n+\t'false' will disable multithreading. Defaults to 'true'.\n+\n index.version::\n \tSpecify the version with which new index files should be\n \tinitialized.  This does not affect existing repositories.\ndiff --git a/config.c b/config.c\nindex 3461993f0a..2ee29f6f86 100644\n--- a/config.c\n+++ b/config.c\n@@ -2289,6 +2289,24 @@ int git_config_get_fsmonitor(void)\n \treturn 0;\n }\n \n+int git_config_get_index_threads(void)\n+{\n+\tint is_bool, val = 0;\n+\n+\tval = git_env_ulong(\"GIT_TEST_INDEX_THREADS\", 0);\n+\tif (val)\n+\t\treturn val;\n+\n+\tif (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n+\t\tif (is_bool)\n+\t\t\treturn val ? 0 : 1;\n+\t\telse\n+\t\t\treturn val;\n+\t}\n+\n+\treturn 0; /* auto */\n+}\n+\n NORETURN\n void git_die_config_linenr(const char *key, const char *filename, int linenr)\n {\ndiff --git a/config.h b/config.h\nindex ab46e0165d..a06027e69b 100644\n--- a/config.h\n+++ b/config.h\n@@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n extern int git_config_get_split_index(void);\n extern int git_config_get_max_percent_split_change(void);\n extern int git_config_get_fsmonitor(void);\n+extern int git_config_get_index_threads(void);\n \n /* This dies if the configured or default date is in the future */\n extern int git_config_get_expiry(const char *key, const char **output);\ndiff --git a/t/README b/t/README\nindex aa33ac4f26..0fcecf4500 100644\n--- a/t/README\n+++ b/t/README\n@@ -332,6 +332,11 @@ This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n as they currently hard code SHA values for the index which are no longer\n valid due to the addition of the EOIE extension.\n \n+GIT_TEST_INDEX_THREADS=<n> enables exercising the multi-threaded loading\n+of the index for the whole test suite by bypassing the default number of\n+cache entries and thread minimums. Settting this to 1 will make the\n+index loading single threaded.\n+\n Naming Tests\n ------------\n \ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex 1f168378c8..ab205954cf 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -8,6 +8,7 @@ test_description='split index mode tests'\n sane_unset GIT_TEST_SPLIT_INDEX\n sane_unset GIT_FSMONITOR_TEST\n GIT_TEST_DISABLE_EOIE=true; export GIT_TEST_DISABLE_EOIE\n+GIT_TEST_INDEX_THREADS=1; export GIT_TEST_INDEX_THREADS\n \n test_expect_success 'enable split index' '\n \tgit config splitIndex.maxPercentChange 100 &&\n-- \n2.18.0.windows.1\n\n"},{"id":"358994","messageId":"20180926195442.1380-6-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180926195442.1380-1-benpeart@microsoft.com","subject":"[PATCH v6 5/7] read-cache: load cache extensions on a worker thread","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-26T19:54:40Z","receivedAt":"2018-09-26T19:55:02Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by loading\nthe cache extensions on a worker thread in parallel with loading the cache\nentries.\n\nIn some cases, loading the extensions takes longer than loading the\ncache entries so this patch utilizes the new EOIE to start the thread to\nload the extensions before loading all the cache entries in parallel.\n\nThis is possible because the current extensions don't access the cache\nentries in the index_state structure so are OK that they don't all exist\nyet.\n\nThe CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\nextensions don't even get a pointer to the index so don't have access to the\ncache entries.\n\nCACHE_EXT_LINK only uses the index_state to initialize the split index.\nCACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\nupdate and dirty flags.\n\nI used p0002-read-cache.sh to generate some performance data:\n\n\tTest w/100,000 files reduced the time by 0.53%\n\tTest w/1,000,000 files reduced the time by 27.78%\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 97 +++++++++++++++++++++++++++++++++++++++++++---------\n 1 file changed, 81 insertions(+), 16 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 80255d3088..8da21c9273 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -23,6 +23,7 @@\n #include \"split-index.h\"\n #include \"utf8.h\"\n #include \"fsmonitor.h\"\n+#include \"thread-utils.h\"\n \n /* Mask for the name length in ce_flags in the on-disk index */\n \n@@ -1890,6 +1891,46 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n \n+struct load_index_extensions\n+{\n+#ifndef NO_PTHREADS\n+\tpthread_t pthread;\n+#endif\n+\tstruct index_state *istate;\n+\tconst char *mmap;\n+\tsize_t mmap_size;\n+\tunsigned long src_offset;\n+};\n+\n+static void *load_index_extensions(void *_data)\n+{\n+\tstruct load_index_extensions *p = _data;\n+\tunsigned long src_offset = p->src_offset;\n+\n+\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, p->mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\t\tif (read_index_extension(p->istate,\n+\t\t\tp->mmap + src_offset,\n+\t\t\tp->mmap + src_offset + 8,\n+\t\t\textsize) < 0) {\n+\t\t\tmunmap((void *)p->mmap, p->mmap_size);\n+\t\t\tdie(_(\"index file corrupt\"));\n+\t\t}\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\n+\treturn NULL;\n+}\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -1900,6 +1941,11 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tconst char *mmap;\n \tsize_t mmap_size;\n \tconst struct cache_entry *previous_ce = NULL;\n+\tstruct load_index_extensions p;\n+\tsize_t extension_offset = 0;\n+#ifndef NO_PTHREADS\n+\tint nr_threads;\n+#endif\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1936,6 +1982,30 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n \tistate->initialized = 1;\n \n+\tp.istate = istate;\n+\tp.mmap = mmap;\n+\tp.mmap_size = mmap_size;\n+\n+#ifndef NO_PTHREADS\n+\tnr_threads = git_config_get_index_threads();\n+\tif (!nr_threads)\n+\t\tnr_threads = online_cpus();\n+\n+\tif (nr_threads > 1) {\n+\t\textension_offset = read_eoie_extension(mmap, mmap_size);\n+\t\tif (extension_offset) {\n+\t\t\tint err;\n+\n+\t\t\tp.src_offset = extension_offset;\n+\t\t\terr = pthread_create(&p.pthread, NULL, load_index_extensions, &p);\n+\t\t\tif (err)\n+\t\t\t\tdie(_(\"unable to create load_index_extensions thread: %s\"), strerror(err));\n+\n+\t\t\tnr_threads--;\n+\t\t}\n+\t}\n+#endif\n+\n \tif (istate->version == 4) {\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n@@ -1960,22 +2030,17 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-\twhile (src_offset <= mmap_size - the_hash_algo->rawsz - 8) {\n-\t\t/* After an array of active_nr index entries,\n-\t\t * there can be arbitrary number of extended\n-\t\t * sections, each of which is prefixed with\n-\t\t * extension name (4-byte) and section length\n-\t\t * in 4-byte network byte order.\n-\t\t */\n-\t\tuint32_t extsize;\n-\t\textsize = get_be32(mmap + src_offset + 4);\n-\t\tif (read_index_extension(istate,\n-\t\t\t\t\t mmap + src_offset,\n-\t\t\t\t\t mmap + src_offset + 8,\n-\t\t\t\t\t extsize) < 0)\n-\t\t\tgoto unmap;\n-\t\tsrc_offset += 8;\n-\t\tsrc_offset += extsize;\n+\t/* if we created a thread, join it otherwise load the extensions on the primary thread */\n+#ifndef NO_PTHREADS\n+\tif (extension_offset) {\n+\t\tint ret = pthread_join(p.pthread, NULL);\n+\t\tif (ret)\n+\t\t\tdie(_(\"unable to join load_index_extensions thread: %s\"), strerror(ret));\n+\t}\n+#endif\n+\tif (!extension_offset) {\n+\t\tp.src_offset = src_offset;\n+\t\tload_index_extensions(&p);\n \t}\n \tmunmap((void *)mmap, mmap_size);\n \treturn istate->cache_nr;\n-- \n2.18.0.windows.1\n\n"},{"id":"358995","messageId":"20180926195442.1380-7-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180926195442.1380-1-benpeart@microsoft.com","subject":"[PATCH v6 6/7] ieot: add Index Entry Offset Table (IEOT) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-26T19:54:41Z","receivedAt":"2018-09-26T19:55:04Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch enables addressing the CPU cost of loading the index by adding\nadditional data to the index that will allow us to efficiently multi-\nthread the loading and conversion of cache entries.\n\nIt accomplishes this by adding an (optional) index extension that is a\ntable of offsets to blocks of cache entries in the index file.  To make\nthis work for V4 indexes, when writing the cache entries, it periodically\n\"resets\" the prefix-compression by encoding the current entry as if the\npath name for the previous entry is completely different and saves the\noffset of that entry in the IEOT.  Basically, with V4 indexes, it\ngenerates offsets into blocks of prefix-compressed entries.\n\nSigned-off-by: Ben Peart <benpeart@microsoft.com>\n---\n Documentation/technical/index-format.txt |  18 +++\n read-cache.c                             | 166 +++++++++++++++++++++++\n 2 files changed, 184 insertions(+)\n\ndiff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\nindex 6bc2d90f7f..7c4d67aa6a 100644\n--- a/Documentation/technical/index-format.txt\n+++ b/Documentation/technical/index-format.txt\n@@ -337,3 +337,21 @@ The remaining data of each directory block is grouped by type:\n \n \tSHA-1(\"TREE\" + <binary representation of N> +\n \t\t\"REUC\" + <binary representation of M>)\n+\n+== Index Entry Offset Table\n+\n+  The Index Entry Offset Table (IEOT) is used to help address the CPU\n+  cost of loading the index by enabling multi-threading the process of\n+  converting cache entries from the on-disk format to the in-memory format.\n+  The signature for this extension is { 'I', 'E', 'O', 'T' }.\n+\n+  The extension consists of:\n+\n+  - 32-bit version (currently 1)\n+\n+  - A number of index offset entries each consisting of:\n+\n+    - 32-bit offset from the begining of the file to the first cache entry\n+\tin this block of entries.\n+\n+    - 32-bit count of cache entries in this block\ndiff --git a/read-cache.c b/read-cache.c\nindex 8da21c9273..9b0554d4e6 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -45,6 +45,7 @@\n #define CACHE_EXT_UNTRACKED 0x554E5452\t  /* \"UNTR\" */\n #define CACHE_EXT_FSMONITOR 0x46534D4E\t  /* \"FSMN\" */\n #define CACHE_EXT_ENDOFINDEXENTRIES 0x454F4945\t/* \"EOIE\" */\n+#define CACHE_EXT_INDEXENTRYOFFSETTABLE 0x49454F54 /* \"IEOT\" */\n \n /* changes that can be kept in $GIT_DIR/index (basically all extensions) */\n #define EXTMASK (RESOLVE_UNDO_CHANGED | CACHE_TREE_CHANGED | \\\n@@ -1696,6 +1697,7 @@ static int read_index_extension(struct index_state *istate,\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n \tcase CACHE_EXT_ENDOFINDEXENTRIES:\n+\tcase CACHE_EXT_INDEXENTRYOFFSETTABLE:\n \t\t/* already handled in do_read_index() */\n \t\tbreak;\n \tdefault:\n@@ -1888,6 +1890,23 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+struct index_entry_offset\n+{\n+\t/* starting byte offset into index file, count of index entries in this block */\n+\tint offset, nr;\n+};\n+\n+struct index_entry_offset_table\n+{\n+\tint nr;\n+\tstruct index_entry_offset entries[0];\n+};\n+\n+#ifndef NO_PTHREADS\n+static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset);\n+static void write_ieot_extension(struct strbuf *sb, struct index_entry_offset_table *ieot);\n+#endif\n+\n static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n \n@@ -1931,6 +1950,15 @@ static void *load_index_extensions(void *_data)\n \treturn NULL;\n }\n \n+/*\n+ * Mostly randomly chosen maximum thread counts: we\n+ * cap the parallelism to online_cpus() threads, and we want\n+ * to have at least 10000 cache entries per thread for it to\n+ * be worth starting a thread.\n+ */\n+\n+#define THREAD_COST\t\t(10000)\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -2523,6 +2551,9 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n+\tint ieot_work = 1;\n+\tstruct index_entry_offset_table *ieot = NULL;\n+\tint nr;\n \n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n@@ -2556,7 +2587,33 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tif (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n \t\treturn -1;\n \n+#ifndef NO_PTHREADS\n+\tif (!strip_extensions && (nr = git_config_get_index_threads()) != 1) {\n+\t\tint ieot_blocks, cpus;\n+\n+\t\t/*\n+\t\t * ensure default number of ieot blocks maps evenly to the\n+\t\t * default number of threads that will process them\n+\t\t */\n+\t\tif (!nr) {\n+\t\t\tieot_blocks = istate->cache_nr / THREAD_COST;\n+\t\t\tif (ieot_blocks < 1)\n+\t\t\t\tieot_blocks = 1;\n+\t\t\tcpus = online_cpus();\n+\t\t\tif (ieot_blocks > cpus - 1)\n+\t\t\t\tieot_blocks = cpus - 1;\n+\t\t} else {\n+\t\t\tieot_blocks = nr;\n+\t\t}\n+\t\tieot = xcalloc(1, sizeof(struct index_entry_offset_table)\n+\t\t\t+ (ieot_blocks * sizeof(struct index_entry_offset)));\n+\t\tieot->nr = 0;\n+\t\tieot_work = DIV_ROUND_UP(entries, ieot_blocks);\n+\t}\n+#endif\n+\n \toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\tnr = 0;\n \tprevious_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n \n \tfor (i = 0; i < entries; i++) {\n@@ -2578,11 +2635,31 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \n \t\t\tdrop_cache_tree = 1;\n \t\t}\n+\t\tif (ieot && i && (i % ieot_work == 0)) {\n+\t\t\tieot->entries[ieot->nr].nr = nr;\n+\t\t\tieot->entries[ieot->nr].offset = offset;\n+\t\t\tieot->nr++;\n+\t\t\t/*\n+\t\t\t * If we have a V4 index, set the first byte to an invalid\n+\t\t\t * character to ensure there is nothing common with the previous\n+\t\t\t * entry\n+\t\t\t */\n+\t\t\tif (previous_name)\n+\t\t\t\tprevious_name->buf[0] = 0;\n+\t\t\tnr = 0;\n+\t\t\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\t\t}\n \t\tif (ce_write_entry(&c, newfd, ce, previous_name, (struct ondisk_cache_entry *)&ondisk) < 0)\n \t\t\terr = -1;\n \n \t\tif (err)\n \t\t\tbreak;\n+\t\tnr++;\n+\t}\n+\tif (ieot && nr) {\n+\t\tieot->entries[ieot->nr].nr = nr;\n+\t\tieot->entries[ieot->nr].offset = offset;\n+\t\tieot->nr++;\n \t}\n \tstrbuf_release(&previous_name_buf);\n \n@@ -2593,6 +2670,24 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n \tthe_hash_algo->init_fn(&eoie_c);\n \n+\t/*\n+\t * Lets write out CACHE_EXT_INDEXENTRYOFFSETTABLE first so that we\n+\t * can minimze the number of extensions we have to scan through to\n+\t * find it during load.\n+\t */\n+#ifndef NO_PTHREADS\n+\tif (!strip_extensions && ieot) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\twrite_ieot_extension(&sb, ieot);\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_INDEXENTRYOFFSETTABLE, sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+#endif\n+\n \tif (!strip_extensions && istate->split_index) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n@@ -3175,3 +3270,74 @@ static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context,\n \tthe_hash_algo->final_fn(hash, eoie_context);\n \tstrbuf_add(sb, hash, the_hash_algo->rawsz);\n }\n+\n+#ifndef NO_PTHREADS\n+#define IEOT_VERSION\t(1)\n+\n+static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset)\n+{\n+       const char *index = NULL;\n+       uint32_t extsize, ext_version;\n+       struct index_entry_offset_table *ieot;\n+       int i, nr;\n+\n+       /* find the IEOT extension */\n+       if (!offset)\n+\t       return NULL;\n+       while (offset <= mmap_size - the_hash_algo->rawsz - 8) {\n+\t       extsize = get_be32(mmap + offset + 4);\n+\t       if (CACHE_EXT((mmap + offset)) == CACHE_EXT_INDEXENTRYOFFSETTABLE) {\n+\t\t       index = mmap + offset + 4 + 4;\n+\t\t       break;\n+\t       }\n+\t       offset += 8;\n+\t       offset += extsize;\n+       }\n+       if (!index)\n+\t       return NULL;\n+\n+       /* validate the version is IEOT_VERSION */\n+       ext_version = get_be32(index);\n+       if (ext_version != IEOT_VERSION)\n+\t       return NULL;\n+       index += sizeof(uint32_t);\n+\n+       /* extension size - version bytes / bytes per entry */\n+       nr = (extsize - sizeof(uint32_t)) / (sizeof(uint32_t) + sizeof(uint32_t));\n+       if (!nr)\n+\t       return NULL;\n+       ieot = xmalloc(sizeof(struct index_entry_offset_table)\n+\t       + (nr * sizeof(struct index_entry_offset)));\n+       ieot->nr = nr;\n+       for (i = 0; i < nr; i++) {\n+\t       ieot->entries[i].offset = get_be32(index);\n+\t       index += sizeof(uint32_t);\n+\t       ieot->entries[i].nr = get_be32(index);\n+\t       index += sizeof(uint32_t);\n+       }\n+\n+       return ieot;\n+}\n+\n+static void write_ieot_extension(struct strbuf *sb, struct index_entry_offset_table *ieot)\n+{\n+       uint32_t buffer;\n+       int i;\n+\n+       /* version */\n+       put_be32(&buffer, IEOT_VERSION);\n+       strbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+       /* ieot */\n+       for (i = 0; i < ieot->nr; i++) {\n+\n+\t       /* offset */\n+\t       put_be32(&buffer, ieot->entries[i].offset);\n+\t       strbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+\t       /* count */\n+\t       put_be32(&buffer, ieot->entries[i].nr);\n+\t       strbuf_add(sb, &buffer, sizeof(uint32_t));\n+       }\n+}\n+#endif\n-- \n2.18.0.windows.1\n\n"},{"id":"358996","messageId":"20180926195442.1380-8-benpeart@microsoft.com","threadId":"49204","inReplyTo":"20180926195442.1380-1-benpeart@microsoft.com","subject":"[PATCH v6 7/7] read-cache: load cache entries on worker threads","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-26T19:54:42Z","receivedAt":"2018-09-26T19:55:05Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"This patch helps address the CPU cost of loading the index by utilizing\nthe Index Entry Offset Table (IEOT) to divide loading and conversion of\nthe cache entries across multiple threads in parallel.\n\nI used p0002-read-cache.sh to generate some performance data:\n\nTest w/100,000 files reduced the time by 32.24%\nTest w/1,000,000 files reduced the time by -4.77%\n\nNote that on the 1,000,000 files case, multi-threading the cache entry parsing\ndoes not yield a performance win.  This is because the cost to parse the\nindex extensions in this repo, far outweigh the cost of loading the cache\nentries.\n\nThe high cost of parsing the index extensions is driven by the cache tree\nand the untracked cache extensions. As this is currently the longest pole,\nany reduction in this time will reduce the overall index load times so is\nworth further investigation in another patch series.\n\nSigned-off-by: Ben Peart <Ben.Peart@microsoft.com>\n---\n read-cache.c | 224 +++++++++++++++++++++++++++++++++++++++++++--------\n 1 file changed, 189 insertions(+), 35 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 9b0554d4e6..f5d766088d 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1720,7 +1720,8 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *create_from_disk(struct index_state *istate,\n+static struct cache_entry *create_from_disk(struct mem_pool *ce_mem_pool,\n+\t\t\t\t\t    unsigned int version,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n \t\t\t\t\t    const struct cache_entry *previous_ce)\n@@ -1737,7 +1738,7 @@ static struct cache_entry *create_from_disk(struct index_state *istate,\n \t * number of bytes to be stripped from the end of the previous name,\n \t * and the bytes to append to the result, to come up with its name.\n \t */\n-\tint expand_name_field = istate->version == 4;\n+\tint expand_name_field = version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1761,16 +1762,17 @@ static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\tconst unsigned char *cp = (const unsigned char *)name;\n \t\tsize_t strip_len, previous_len;\n \n-\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n+\t\t/* If we're at the begining of a block, ignore the previous name */\n \t\tstrip_len = decode_varint(&cp);\n-\t\tif (previous_len < strip_len) {\n-\t\t\tif (previous_ce)\n+\t\tif (previous_ce) {\n+\t\t\tprevious_len = previous_ce->ce_namelen;\n+\t\t\tif (previous_len < strip_len)\n \t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n-\t\t\t\t    previous_ce->name);\n-\t\t\telse\n-\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t\t\t\tprevious_ce->name);\n+\t\t\tcopy_len = previous_len - strip_len;\n+\t\t} else {\n+\t\t\tcopy_len = 0;\n \t\t}\n-\t\tcopy_len = previous_len - strip_len;\n \t\tname = (const char *)cp;\n \t}\n \n@@ -1780,7 +1782,7 @@ static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\t\tlen += copy_len;\n \t}\n \n-\tce = mem_pool__ce_alloc(istate->ce_mem_pool, len);\n+\tce = mem_pool__ce_alloc(ce_mem_pool, len);\n \n \tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n \tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n@@ -1950,6 +1952,52 @@ static void *load_index_extensions(void *_data)\n \treturn NULL;\n }\n \n+/*\n+ * A helper function that will load the specified range of cache entries\n+ * from the memory mapped file and add them to the given index.\n+ */\n+static unsigned long load_cache_entry_block(struct index_state *istate,\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, const char *mmap,\n+\t\t\tunsigned long start_offset, const struct cache_entry *previous_ce)\n+{\n+\tint i;\n+\tunsigned long src_offset = start_offset;\n+\n+\tfor (i = offset; i < offset + nr; i++) {\n+\t\tstruct ondisk_cache_entry *disk_ce;\n+\t\tstruct cache_entry *ce;\n+\t\tunsigned long consumed;\n+\n+\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n+\t\tce = create_from_disk(ce_mem_pool, istate->version, disk_ce, &consumed, previous_ce);\n+\t\tset_index_entry(istate, i, ce);\n+\n+\t\tsrc_offset += consumed;\n+\t\tprevious_ce = ce;\n+\t}\n+\treturn src_offset - start_offset;\n+}\n+\n+static unsigned long load_all_cache_entries(struct index_state *istate,\n+\t\t\tconst char *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tunsigned long consumed;\n+\n+\tif (istate->version == 4) {\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n+\t} else {\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t\testimate_cache_size(mmap_size, istate->cache_nr));\n+\t}\n+\n+\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, NULL);\n+\treturn consumed;\n+}\n+\n+#ifndef NO_PTHREADS\n+\n /*\n  * Mostly randomly chosen maximum thread counts: we\n  * cap the parallelism to online_cpus() threads, and we want\n@@ -1959,20 +2007,125 @@ static void *load_index_extensions(void *_data)\n \n #define THREAD_COST\t\t(10000)\n \n+struct load_cache_entries_thread_data\n+{\n+\tpthread_t pthread;\n+\tstruct index_state *istate;\n+\tstruct mem_pool *ce_mem_pool;\n+\tint offset;\n+\tconst char *mmap;\n+\tstruct index_entry_offset_table *ieot;\n+\tint ieot_offset;        /* starting index into the ieot array */\n+\tint ieot_work;          /* count of ieot entries to process */\n+\tunsigned long consumed;\t/* return # of bytes in index file processed */\n+};\n+\n+/*\n+ * A thread proc to run the load_cache_entries() computation\n+ * across multiple background threads.\n+ */\n+static void *load_cache_entries_thread(void *_data)\n+{\n+\tstruct load_cache_entries_thread_data *p = _data;\n+\tint i;\n+\n+\t/* iterate across all ieot blocks assigned to this thread */\n+\tfor (i = p->ieot_offset; i < p->ieot_offset + p->ieot_work; i++) {\n+\t\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool, p->offset, p->ieot->entries[i].nr, p->mmap, p->ieot->entries[i].offset, NULL);\n+\t\tp->offset += p->ieot->entries[i].nr;\n+\t}\n+\treturn NULL;\n+}\n+\n+static unsigned long load_cache_entries_threaded(struct index_state *istate, const char *mmap, size_t mmap_size,\n+\t\t\tunsigned long src_offset, int nr_threads, struct index_entry_offset_table *ieot)\n+{\n+\tint i, offset, ieot_work, ieot_offset, err;\n+\tstruct load_cache_entries_thread_data *data;\n+\tunsigned long consumed = 0;\n+\tint nr;\n+\n+\t/* a little sanity checking */\n+\tif (istate->name_hash_initialized)\n+\t\tBUG(\"the name hash isn't thread safe\");\n+\n+\tmem_pool_init(&istate->ce_mem_pool, 0);\n+\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n+\n+\t/* ensure we have no more threads than we have blocks to process */\n+\tif (nr_threads > ieot->nr)\n+\t\tnr_threads = ieot->nr;\n+\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n+\n+\toffset = ieot_offset = 0;\n+\tieot_work = DIV_ROUND_UP(ieot->nr, nr_threads);\n+\tfor (i = 0; i < nr_threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = &data[i];\n+\t\tint j;\n+\n+\t\tif (ieot_offset + ieot_work > ieot->nr)\n+\t\t\tieot_work = ieot->nr - ieot_offset;\n+\n+\t\tp->istate = istate;\n+\t\tp->offset = offset;\n+\t\tp->mmap = mmap;\n+\t\tp->ieot = ieot;\n+\t\tp->ieot_offset = ieot_offset;\n+\t\tp->ieot_work = ieot_work;\n+\n+\t\t/* create a mem_pool for each thread */\n+\t\tnr = 0;\n+\t\tfor (j = p->ieot_offset; j < p->ieot_offset + p->ieot_work; j++)\n+\t\t\tnr += p->ieot->entries[j].nr;\n+\t\tif (istate->version == 4) {\n+\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\testimate_cache_size_from_compressed(nr));\n+\t\t}\n+\t\telse {\n+\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\testimate_cache_size(mmap_size, nr));\n+\t\t}\n+\n+\t\terr = pthread_create(&p->pthread, NULL, load_cache_entries_thread, p);\n+\t\tif (err)\n+\t\t\tdie(_(\"unable to create load_cache_entries thread: %s\"), strerror(err));\n+\n+\t\t/* increment by the number of cache entries in the ieot block being processed */\n+\t\tfor (j = 0; j < ieot_work; j++)\n+\t\t\toffset += ieot->entries[ieot_offset + j].nr;\n+\t\tieot_offset += ieot_work;\n+\t}\n+\n+\tfor (i = 0; i < nr_threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = &data[i];\n+\n+\t\terr = pthread_join(p->pthread, NULL);\n+\t\tif (err)\n+\t\t\tdie(_(\"unable to join load_cache_entries thread: %s\"), strerror(err));\n+\t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n+\t\tconsumed += p->consumed;\n+\t}\n+\n+\tfree(data);\n+\n+\treturn consumed;\n+}\n+#endif\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n-\tint fd, i;\n+\tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n \tconst struct cache_header *hdr;\n \tconst char *mmap;\n \tsize_t mmap_size;\n-\tconst struct cache_entry *previous_ce = NULL;\n \tstruct load_index_extensions p;\n \tsize_t extension_offset = 0;\n #ifndef NO_PTHREADS\n-\tint nr_threads;\n+\tint nr_threads, cpus;\n+\tstruct index_entry_offset_table *ieot = 0;\n #endif\n \n \tif (istate->initialized)\n@@ -2014,10 +2167,18 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tp.mmap = mmap;\n \tp.mmap_size = mmap_size;\n \n+\tsrc_offset = sizeof(*hdr);\n+\n #ifndef NO_PTHREADS\n \tnr_threads = git_config_get_index_threads();\n-\tif (!nr_threads)\n-\t\tnr_threads = online_cpus();\n+\n+\t/* TODO: does creating more threads than cores help? */\n+\tif (!nr_threads) {\n+\t\tnr_threads = istate->cache_nr / THREAD_COST;\n+\t\tcpus = online_cpus();\n+\t\tif (nr_threads > cpus)\n+\t\t\tnr_threads = cpus;\n+\t}\n \n \tif (nr_threads > 1) {\n \t\textension_offset = read_eoie_extension(mmap, mmap_size);\n@@ -2032,29 +2193,22 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\t\tnr_threads--;\n \t\t}\n \t}\n-#endif\n-\n-\tif (istate->version == 4) {\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n-\t} else {\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n-\t}\n \n-\tsrc_offset = sizeof(*hdr);\n-\tfor (i = 0; i < istate->cache_nr; i++) {\n-\t\tstruct ondisk_cache_entry *disk_ce;\n-\t\tstruct cache_entry *ce;\n-\t\tunsigned long consumed;\n+\t/*\n+\t * Locate and read the index entry offset table so that we can use it\n+\t * to multi-thread the reading of the cache entries.\n+\t */\n+\tif (extension_offset && nr_threads > 1)\n+\t\tieot = read_ieot_extension(mmap, mmap_size, extension_offset);\n \n-\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n-\t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n-\t\tset_index_entry(istate, i, ce);\n+\tif (ieot)\n+\t\tsrc_offset += load_cache_entries_threaded(istate, mmap, mmap_size, src_offset, nr_threads, ieot);\n+\telse\n+\t\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+#else\n+\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+#endif\n \n-\t\tsrc_offset += consumed;\n-\t\tprevious_ce = ce;\n-\t}\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-- \n2.18.0.windows.1\n\n"},{"id":"359013","messageId":"xmqq5zyrkj5q.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180926195442.1380-1-benpeart@microsoft.com","subject":"Re: [PATCH v6 0/7] speed up index load through parallelization","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-26T22:06:57Z","receivedAt":"2018-09-26T22:07:03Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <peartben@gmail.com> writes:\n\n> Base Ref: master\n> Web-Diff: https://github.com/benpeart/git/commit/a0300882d4\n> Checkout: git fetch https://github.com/benpeart/git read-index-multithread-v6 && git checkout a0300882d4\n>\n>\n> This iteration brings back the Index Entry Offset Table (IEOT) extension\n> which enables us to multi-thread the cache entry parsing without having\n> the primary thread have to scan all the entries first.  In cases where the\n> cache entry parsing is the most expensive part, this yields some additional\n> savings.\n\nNice.\n\n> Test w/100,000 files    Baseline  Optimize V4    Extensions     Entries\n> ----------------------------------------------------------------------------\n> 0002.1: read_cache      22.36     18.74 -16.2%   18.64 -16.6%   12.63 -43.5%\n>\n> Test w/1,000,000 files  Baseline  Optimize V4    Extensions     Entries\n> -----------------------------------------------------------------------------\n> 0002.1: read_cache      304.40    270.70 -11.1%  195.50 -35.8%  204.82 -32.7%\n>\n> Note that on the 1,000,000 files case, multi-threading the cache entry parsing\n> does not yield a performance win.  This is because the cost to parse the\n> index extensions in this repo, far outweigh the cost of loading the cache\n> entries.\n> ...\n> The high cost of parsing the index extensions is driven by the cache tree\n> and the untracked cache extensions. As this is currently the longest pole,\n> any reduction in this time will reduce the overall index load times so is\n> worth further investigation in another patch series.\n\nInteresting.\n\n> One option would be to load each extension on a separate thread but I\n> believe that is overkill for the vast majority of repos.  Instead, some\n> optimization of the loading code for these two extensions is probably worth\n> looking into as a quick examination shows that the bulk of the time for both\n> of them is spent in xcalloc().\n\nThanks.  Looking forward to block some quality time off to read this\nthrough, but from the cursory look (read: diff between the previous\nround), this looks quite promising.\n"},{"id":"359071","messageId":"CACsJy8BdKo-TQi2YuPjyoTb3uyqKLvJKNOSD5cdopygm3y4CcA@mail.gmail.com","threadId":"49204","inReplyTo":"20180926195442.1380-1-benpeart@microsoft.com","subject":"Re: [PATCH v6 0/7] speed up index load through parallelization","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-27T17:13:08Z","receivedAt":"2018-09-27T17:13:36Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Sep 26, 2018 at 9:54 PM Ben Peart <peartben@gmail.com> wrote:\n> The high cost of parsing the index extensions is driven by the cache tree\n> and the untracked cache extensions. As this is currently the longest pole,\n> any reduction in this time will reduce the overall index load times so is\n> worth further investigation in another patch series.\n>\n> Name                                    First    Last     Elapsed\n> |   + git!read_index_extension          684.052  870.244  186.192\n> |    + git!cache_tree_read              684.052  797.801  113.749\n> |    + git!read_untracked_extension     797.801  870.244  72.443\n>\n> One option would be to load each extension on a separate thread but I\n> believe that is overkill for the vast majority of repos.\n\nThey both grow proportional to the number of trees in worktree, which\nprobably also scales to the worktree size. Frankly I think the\nparallel index loading is already overkill for the majority of repos,\nso speeding up further of the 1% giant repos does not sound that bad.\nAnd I think you already lay the foundation for loading index stuff in\nparallel with this series.\n\n> Instead, some\n> optimization of the loading code for these two extensions is probably worth\n> looking into as a quick examination shows that the bulk of the time for both\n> of them is spent in xcalloc().\n\nAnother easy \"optimization\" is delaying loading these until we need\nthem (or load them in background, read_index() returns even before\nthese extensions are finished, but this is of course trickier).\n\nUNTR extension for example is only useful for \"git status\" (and maybe\none or two other use cases). Not having to load them all the time is\nlikely a win. The role of TREE extension has grown bigger these days\nso it's still maybe worth putting more effort into making it load it\nfaster rather than just hiding the cost.\n-- \nDuy\n"},{"id":"359164","messageId":"20180928001929.GN27036@localhost","threadId":"49204","inReplyTo":"20180926195442.1380-4-benpeart@microsoft.com","subject":"Re: [PATCH v6 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2018-09-28T00:19:29Z","receivedAt":"2018-09-28T00:19:37Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Wed, Sep 26, 2018 at 03:54:38PM -0400, Ben Peart wrote:\n> The End of Index Entry (EOIE) is used to locate the end of the variable\n\nNit: perhaps start with: \n\n  The End of Index Entry (EOIE) optional extension can be used to ...\n\nto make it clearer for those who don't immediately realize the\nsignificance of the upper case 'E' in the extension's signature.\n\n> length index entries and the beginning of the extensions. Code can take\n> advantage of this to quickly locate the index extensions without having\n> to parse through all of the index entries.\n> \n> Because it must be able to be loaded before the variable length cache\n> entries and other index extensions, this extension must be written last.\n> The signature for this extension is { 'E', 'O', 'I', 'E' }.\n> \n> The extension consists of:\n> \n> - 32-bit offset to the end of the index entries\n> \n> - 160-bit SHA-1 over the extension types and their sizes (but not\n> their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n> long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n> then the hash would be:\n> \n> SHA-1(\"TREE\" + <binary representation of N> +\n> \t\"REUC\" + <binary representation of M>)\n> \n> Signed-off-by: Ben Peart <Ben.Peart@microsoft.com>\n> ---\n>  Documentation/technical/index-format.txt |  23 ++++\n>  read-cache.c                             | 151 +++++++++++++++++++++--\n>  t/README                                 |   5 +\n>  t/t1700-split-index.sh                   |   1 +\n>  4 files changed, 172 insertions(+), 8 deletions(-)\n> \n\n> diff --git a/t/README b/t/README\n> index 3ea6c85460..aa33ac4f26 100644\n> --- a/t/README\n> +++ b/t/README\n> @@ -327,6 +327,11 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n>  be written after every 'git commit' command, and overrides the\n>  'core.commitGraph' setting to true.\n>  \n> +GIT_TEST_DISABLE_EOIE=<boolean> disables writing the EOIE extension.\n> +This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n> +as they currently hard code SHA values for the index which are no longer\n> +valid due to the addition of the EOIE extension.\n\nIs this extension enabled by default?  The commit message doesn't\nexplicitly say so, but I don't see any way to turn it on or off, while\nthere is this new GIT_TEST environment variable to disable it for one\nparticular test, so it seems so.  If that's indeed the case, then\nwouldn't it be better to update those hard-coded SHA1 values in t1700\ninstead?\n\n>  Naming Tests\n>  ------------\n>  \n> diff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\n> index be22398a85..1f168378c8 100755\n> --- a/t/t1700-split-index.sh\n> +++ b/t/t1700-split-index.sh\n> @@ -7,6 +7,7 @@ test_description='split index mode tests'\n>  # We need total control of index splitting here\n>  sane_unset GIT_TEST_SPLIT_INDEX\n>  sane_unset GIT_FSMONITOR_TEST\n> +GIT_TEST_DISABLE_EOIE=true; export GIT_TEST_DISABLE_EOIE\n>  \n>  test_expect_success 'enable split index' '\n>  \tgit config splitIndex.maxPercentChange 100 &&\n> -- \n> 2.18.0.windows.1\n> \n"},{"id":"359165","messageId":"20180928002627.GO27036@localhost","threadId":"49204","inReplyTo":"20180926195442.1380-5-benpeart@microsoft.com","subject":"Re: [PATCH v6 4/7] config: add new index.threads config setting","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2018-09-28T00:26:27Z","receivedAt":"2018-09-28T00:26:34Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Wed, Sep 26, 2018 at 03:54:39PM -0400, Ben Peart wrote:\n> Add support for a new index.threads config setting which will be used to\n> control the threading code in do_read_index().  A value of 0 will tell the\n> index code to automatically determine the correct number of threads to use.\n> A value of 1 will make the code single threaded.  A value greater than 1\n> will set the maximum number of threads to use.\n> \n> For testing purposes, this setting can be overwritten by setting the\n> GIT_TEST_INDEX_THREADS=<n> environment variable to a value greater than 0.\n> \n> Signed-off-by: Ben Peart <Ben.Peart@microsoft.com>\n> ---\n\n> diff --git a/t/README b/t/README\n> index aa33ac4f26..0fcecf4500 100644\n> --- a/t/README\n> +++ b/t/README\n> @@ -332,6 +332,11 @@ This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n>  as they currently hard code SHA values for the index which are no longer\n>  valid due to the addition of the EOIE extension.\n>  \n> +GIT_TEST_INDEX_THREADS=<n> enables exercising the multi-threaded loading\n> +of the index for the whole test suite by bypassing the default number of\n> +cache entries and thread minimums. Settting this to 1 will make the\n\ns/ttt/tt/\n\n> +index loading single threaded.\n> +\n>  Naming Tests\n>  ------------\n>  \n> diff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\n> index 1f168378c8..ab205954cf 100755\n> --- a/t/t1700-split-index.sh\n> +++ b/t/t1700-split-index.sh\n> @@ -8,6 +8,7 @@ test_description='split index mode tests'\n>  sane_unset GIT_TEST_SPLIT_INDEX\n>  sane_unset GIT_FSMONITOR_TEST\n>  GIT_TEST_DISABLE_EOIE=true; export GIT_TEST_DISABLE_EOIE\n> +GIT_TEST_INDEX_THREADS=1; export GIT_TEST_INDEX_THREADS\n\nWhy does multithreading have to be disabled in this test?\n\n>  test_expect_success 'enable split index' '\n>  \tgit config splitIndex.maxPercentChange 100 &&\n> -- \n> 2.18.0.windows.1\n> \n"},{"id":"359200","messageId":"cbc48a95-62f5-a098-fb70-97b6cf241920@gmail.com","threadId":"49204","inReplyTo":"20180928002627.GO27036@localhost","subject":"Re: [PATCH v6 4/7] config: add new index.threads config setting","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-28T13:39:12Z","receivedAt":"2018-09-28T13:39:19Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/27/2018 8:26 PM, SZEDER Gábor wrote:\n> On Wed, Sep 26, 2018 at 03:54:39PM -0400, Ben Peart wrote:\n>> Add support for a new index.threads config setting which will be used to\n>> control the threading code in do_read_index().  A value of 0 will tell the\n>> index code to automatically determine the correct number of threads to use.\n>> A value of 1 will make the code single threaded.  A value greater than 1\n>> will set the maximum number of threads to use.\n>>\n>> For testing purposes, this setting can be overwritten by setting the\n>> GIT_TEST_INDEX_THREADS=<n> environment variable to a value greater than 0.\n>>\n>> Signed-off-by: Ben Peart <Ben.Peart@microsoft.com>\n>> ---\n> \n>> diff --git a/t/README b/t/README\n>> index aa33ac4f26..0fcecf4500 100644\n>> --- a/t/README\n>> +++ b/t/README\n>> @@ -332,6 +332,11 @@ This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n>>   as they currently hard code SHA values for the index which are no longer\n>>   valid due to the addition of the EOIE extension.\n>>   \n>> +GIT_TEST_INDEX_THREADS=<n> enables exercising the multi-threaded loading\n>> +of the index for the whole test suite by bypassing the default number of\n>> +cache entries and thread minimums. Settting this to 1 will make the\n> \n> s/ttt/tt/\n> \n>> +index loading single threaded.\n>> +\n>>   Naming Tests\n>>   ------------\n>>   \n>> diff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\n>> index 1f168378c8..ab205954cf 100755\n>> --- a/t/t1700-split-index.sh\n>> +++ b/t/t1700-split-index.sh\n>> @@ -8,6 +8,7 @@ test_description='split index mode tests'\n>>   sane_unset GIT_TEST_SPLIT_INDEX\n>>   sane_unset GIT_FSMONITOR_TEST\n>>   GIT_TEST_DISABLE_EOIE=true; export GIT_TEST_DISABLE_EOIE\n>> +GIT_TEST_INDEX_THREADS=1; export GIT_TEST_INDEX_THREADS\n> \n> Why does multithreading have to be disabled in this test?\n> \n\nIf multi-threading is enabled, it will write out the IEOT extension \nwhich changes the SHA and causes the test to fail.  I will update the \nlogic in this case to not write out the IEOT extension as it isn't needed.\n\n>>   test_expect_success 'enable split index' '\n>>   \tgit config splitIndex.maxPercentChange 100 &&\n>> -- \n>> 2.18.0.windows.1\n>>\n"},{"id":"359221","messageId":"xmqqsh1tczyz.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"cbc48a95-62f5-a098-fb70-97b6cf241920@gmail.com","subject":"Re: [PATCH v6 4/7] config: add new index.threads config setting","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-28T17:07:48Z","receivedAt":"2018-09-28T17:07:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <peartben@gmail.com> writes:\n\n>> Why does multithreading have to be disabled in this test?\n>\n> If multi-threading is enabled, it will write out the IEOT extension\n> which changes the SHA and causes the test to fail.\n\nI think it is a design mistake to let the writing processes's\ncapability decide what is written in the file to be read later by a\ndifferent process, which possibly may have different capability.  If\nyou are not writing with multiple threads, it should not matter if\nthat writer process is capable of and configured to spawn 8 threads\nif the process were reading the file---as it is not reading the file\nit is writing right now.\n\nI can understand if the design is to write IEOT only if the\nresulting index is expected to become large enough (above an\narbitrary threshold like 100k entries) to matter.  I also can\nunderstand if IEOT is omitted when the repository configuration says\nthat no process is allowed to read the index with multi-threaded\ncodepath in that repository.\n"},{"id":"359235","messageId":"8a12adb1-74f2-b824-b2c7-c82e0f36ad94@gmail.com","threadId":"49204","inReplyTo":"20180928001929.GN27036@localhost","subject":"Re: [PATCH v6 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-28T18:38:54Z","receivedAt":"2018-09-28T18:39:00Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/27/2018 8:19 PM, SZEDER Gábor wrote:\n> On Wed, Sep 26, 2018 at 03:54:38PM -0400, Ben Peart wrote:\n>> The End of Index Entry (EOIE) is used to locate the end of the variable\n> \n> Nit: perhaps start with:\n> \n>    The End of Index Entry (EOIE) optional extension can be used to ...\n> \n> to make it clearer for those who don't immediately realize the\n> significance of the upper case 'E' in the extension's signature.\n> \n>> length index entries and the beginning of the extensions. Code can take\n>> advantage of this to quickly locate the index extensions without having\n>> to parse through all of the index entries.\n>>\n>> Because it must be able to be loaded before the variable length cache\n>> entries and other index extensions, this extension must be written last.\n>> The signature for this extension is { 'E', 'O', 'I', 'E' }.\n>>\n>> The extension consists of:\n>>\n>> - 32-bit offset to the end of the index entries\n>>\n>> - 160-bit SHA-1 over the extension types and their sizes (but not\n>> their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n>> long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n>> then the hash would be:\n>>\n>> SHA-1(\"TREE\" + <binary representation of N> +\n>> \t\"REUC\" + <binary representation of M>)\n>>\n>> Signed-off-by: Ben Peart <Ben.Peart@microsoft.com>\n>> ---\n>>   Documentation/technical/index-format.txt |  23 ++++\n>>   read-cache.c                             | 151 +++++++++++++++++++++--\n>>   t/README                                 |   5 +\n>>   t/t1700-split-index.sh                   |   1 +\n>>   4 files changed, 172 insertions(+), 8 deletions(-)\n>>\n> \n>> diff --git a/t/README b/t/README\n>> index 3ea6c85460..aa33ac4f26 100644\n>> --- a/t/README\n>> +++ b/t/README\n>> @@ -327,6 +327,11 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n>>   be written after every 'git commit' command, and overrides the\n>>   'core.commitGraph' setting to true.\n>>   \n>> +GIT_TEST_DISABLE_EOIE=<boolean> disables writing the EOIE extension.\n>> +This is used to allow tests 1, 4-9 in t1700-split-index.sh to succeed\n>> +as they currently hard code SHA values for the index which are no longer\n>> +valid due to the addition of the EOIE extension.\n> \n> Is this extension enabled by default?  The commit message doesn't\n> explicitly say so, but I don't see any way to turn it on or off, while\n> there is this new GIT_TEST environment variable to disable it for one\n> particular test, so it seems so.  If that's indeed the case, then\n> wouldn't it be better to update those hard-coded SHA1 values in t1700\n> instead?\n> \n\nYes, it is enabled by default and the only way to disable it is the \nGIT_TEST_DISABLE_EOIE environment variable.\n\nThe tests in t1700-split-index.sh assume that there are no extensions in \nthe index file so anything that adds an extension, will break one or \nmore of the tests.\n\nFirst in 'enable split index', they hard code SHA values assuming there \nare no extensions. If some option adds an extension, these hard coded \nvalues no longer match and the test fails.\n\nLater in 'disable split index' they save off the SHA of the index with \nsplit-index turned off and then in later tests, compare it to the SHA of \nthe shared index.  Because extensions are stripped when the shared index \nis written out this only works if there were not extensions in the \noriginal index.\n\nI'll document this behavior and reasoning in the test directly.\n\nThis did cause me to reexamine how EOIE and IEOT behave when split index \nis turned on.  These two extensions help most with a large index.  When \nsplit index is turned on, the large index is actually the shared index \nas the index is now the smaller set of deltas.\n\nCurrently, the extensions are stripped out of the shared index which \nmeans they are not available when they are needed to quickly load the \nshared index.  I'll see if I can update the patch so that these \nextensions are still written out and available in the shared index to \nspeed up when it is loaded.\n\nThanks!\n\n>>   Naming Tests\n>>   ------------\n>>   \n>> diff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\n>> index be22398a85..1f168378c8 100755\n>> --- a/t/t1700-split-index.sh\n>> +++ b/t/t1700-split-index.sh\n>> @@ -7,6 +7,7 @@ test_description='split index mode tests'\n>>   # We need total control of index splitting here\n>>   sane_unset GIT_TEST_SPLIT_INDEX\n>>   sane_unset GIT_FSMONITOR_TEST\n>> +GIT_TEST_DISABLE_EOIE=true; export GIT_TEST_DISABLE_EOIE\n>>   \n>>   test_expect_success 'enable split index' '\n>>   \tgit config splitIndex.maxPercentChange 100 &&\n>> -- \n>> 2.18.0.windows.1\n>>\n"},{"id":"359241","messageId":"a58a5cce-b3c2-62a2-598b-6b7dbe1a86fc@gmail.com","threadId":"49204","inReplyTo":"xmqqsh1tczyz.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v6 4/7] config: add new index.threads config setting","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-09-28T19:41:28Z","receivedAt":"2018-09-28T19:41:34Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/28/2018 1:07 PM, Junio C Hamano wrote:\n> Ben Peart <peartben@gmail.com> writes:\n> \n>>> Why does multithreading have to be disabled in this test?\n>>\n>> If multi-threading is enabled, it will write out the IEOT extension\n>> which changes the SHA and causes the test to fail.\n> \n> I think it is a design mistake to let the writing processes's\n> capability decide what is written in the file to be read later by a\n> different process, which possibly may have different capability.  If\n> you are not writing with multiple threads, it should not matter if\n> that writer process is capable of and configured to spawn 8 threads\n> if the process were reading the file---as it is not reading the file\n> it is writing right now.\n> \n> I can understand if the design is to write IEOT only if the\n> resulting index is expected to become large enough (above an\n> arbitrary threshold like 100k entries) to matter.  I also can\n> understand if IEOT is omitted when the repository configuration says\n> that no process is allowed to read the index with multi-threaded\n> codepath in that repository.\n> \n\nThere are two different paths which determine how many blocks are \nwritten to the IEOT.  The first is the default path.  On this path, the \nnumber of blocks is determined by the number of cache entries divided by \nthe THREAD_COST.  If there are sufficient entries to make it faster to \nuse threading, then it will automatically use enough blocks to optimize \nthe performance of reading the entries across multiple threads.\n\nI currently cap the maximum number of blocks to be the number of cores \nthat would be available to process them on that same machine purely as \nan optimization.  The majority of the time, the index will be read from \nthe same machine that it was written on so this works well.  Before I \nadded that logic, you would usually end up with more blocks than \navailable threads which meant some threads had more to do than the other \nthreads and resulted in worse performance.  For example, 4 blocks across \n3 threads results in the 1st thread having twice as much work to do as \nthe other threads.\n\nIf the index is copied to a machine with a different number of cores, it \nwill still all work - it just may not be optimal for that machine.  This \nis self correcting because as soon as the index is written out, it will \nbe optimized for that machine.\n\nIf the \"automatically try to make it perform optimally\" logic doesn't \nwork for some reason, we have path #2.\n\nThe second path is when the user specifies a specific number of blocks \nvia the GIT_TEST_INDEX_THREADS=<n> environment variable or the \nindex.threads=<n> config setting.  If they ask for n blocks, they will \nget n blocks.  This is the \"I know what I'm doing and want to control \nthe behavior\" path.\n\nI just added one additional test (see patch below) to avoid a divide by \nzero bug and simplify things a bit.  With this change, if there are \nfewer than two blocks, the IEOT extension is not written out as it isn't \nneeded.  The load would be single threaded anyway so there is no reason \nto write out a IEOT extensions that won't be used.\n\n\n\ndiff --git a/read-cache.c b/read-cache.c\nindex f5d766088d..a1006fa824 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2751,18 +2751,23 @@ static int do_write_index(struct index_state \n*istate, struct tempfile *tempfil\ne,\n                  */\n                 if (!nr) {\n                         ieot_blocks = istate->cache_nr / THREAD_COST;\n-                       if (ieot_blocks < 1)\n-                               ieot_blocks = 1;\n                         cpus = online_cpus();\n                         if (ieot_blocks > cpus - 1)\n                                 ieot_blocks = cpus - 1;\n                 } else {\n                         ieot_blocks = nr;\n                 }\n-               ieot = xcalloc(1, sizeof(struct index_entry_offset_table)\n-                       + (ieot_blocks * sizeof(struct \nindex_entry_offset)));\n-               ieot->nr = 0;\n-               ieot_work = DIV_ROUND_UP(entries, ieot_blocks);\n+\n+               /*\n+                * no reason to write out the IEOT extension if we don't\n+                * have enough blocks to utilize multi-threading\n+                */\n+               if (ieot_blocks > 1) {\n+                       ieot = xcalloc(1, sizeof(struct \nindex_entry_offset_table)\n+                               + (ieot_blocks * sizeof(struct \nindex_entry_offset)));\n+                       ieot->nr = 0;\n+                       ieot_work = DIV_ROUND_UP(entries, ieot_blocks);\n+               }\n         }\n  #endif\n\n"},{"id":"359247","messageId":"bf0c24ac-6e2a-9a3e-835f-f21e763ab2c7@ramsayjones.plus.com","threadId":"49204","inReplyTo":"a58a5cce-b3c2-62a2-598b-6b7dbe1a86fc@gmail.com","subject":"Re: [PATCH v6 4/7] config: add new index.threads config setting","fromName":"Ramsay Jones","fromEmail":"ramsay@ramsayjones.plus.com","sentAt":"2018-09-28T20:30:02Z","receivedAt":"2018-09-28T20:30:09Z","isPatch":true,"sender":{"key":"ramsay@ramsayjones.plus.com","avatar":"https://avatars.githubusercontent.com/u/33702710?v=4"},"body":"\n\nOn 28/09/18 20:41, Ben Peart wrote:\n> \n> \n> On 9/28/2018 1:07 PM, Junio C Hamano wrote:\n>> Ben Peart <peartben@gmail.com> writes:\n>>\n>>>> Why does multithreading have to be disabled in this test?\n>>>\n>>> If multi-threading is enabled, it will write out the IEOT extension\n>>> which changes the SHA and causes the test to fail.\n>>\n>> I think it is a design mistake to let the writing processes's\n>> capability decide what is written in the file to be read later by a\n>> different process, which possibly may have different capability.  If\n>> you are not writing with multiple threads, it should not matter if\n>> that writer process is capable of and configured to spawn 8 threads\n>> if the process were reading the file---as it is not reading the file\n>> it is writing right now.\n>>\n>> I can understand if the design is to write IEOT only if the\n>> resulting index is expected to become large enough (above an\n>> arbitrary threshold like 100k entries) to matter.  I also can\n>> understand if IEOT is omitted when the repository configuration says\n>> that no process is allowed to read the index with multi-threaded\n>> codepath in that repository.\n>>\n> \n> There are two different paths which determine how many blocks are written to the IEOT.  The first is the default path.  On this path, the number of blocks is determined by the number of cache entries divided by the THREAD_COST.  If there are sufficient entries to make it faster to use threading, then it will automatically use enough blocks to optimize the performance of reading the entries across multiple threads.\n> \n> I currently cap the maximum number of blocks to be the number of cores that would be available to process them on that same machine purely as an optimization.  The majority of the time, the index will be read from the same machine that it was written on so this works well.  Before I added that logic, you would usually end up with more blocks than available threads which meant some threads had more to do than the other threads and resulted in worse performance.  For example, 4 blocks across 3 threads results in the 1st thread having twice as much work to do as the other threads.\n> \n> If the index is copied to a machine with a different number of cores, it will still all work - it just may not be optimal for that machine.  This is self correcting because as soon as the index is written out, it will be optimized for that machine.\n> \n> If the \"automatically try to make it perform optimally\" logic doesn't work for some reason, we have path #2.\n> \n> The second path is when the user specifies a specific number of blocks via the GIT_TEST_INDEX_THREADS=<n> environment variable or the index.threads=<n> config setting.  If they ask for n blocks, they will get n blocks.  This is the \"I know what I'm doing and want to control the behavior\" path.\n> \n> I just added one additional test (see patch below) to avoid a divide by zero bug and simplify things a bit.  With this change, if there are fewer than two blocks, the IEOT extension is not written out as it isn't needed.  The load would be single threaded anyway so there is no reason to write out a IEOT extensions that won't be used.\n> \n> \n> \n> diff --git a/read-cache.c b/read-cache.c\n> index f5d766088d..a1006fa824 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -2751,18 +2751,23 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfil\n> e,\n>                  */\n>                 if (!nr) {\n>                         ieot_blocks = istate->cache_nr / THREAD_COST;\n> -                       if (ieot_blocks < 1)\n> -                               ieot_blocks = 1;\n>                         cpus = online_cpus();\n>                         if (ieot_blocks > cpus - 1)\n>                                 ieot_blocks = cpus - 1;\n\nSo, am I reading this correctly - you need cpus > 2 before an\nIEOT extension block is written out?\n\nOK.\n\nATB,\nRamsay Jones\n\n>                 } else {\n>                         ieot_blocks = nr;\n>                 }\n> -               ieot = xcalloc(1, sizeof(struct index_entry_offset_table)\n> -                       + (ieot_blocks * sizeof(struct index_entry_offset)));\n> -               ieot->nr = 0;\n> -               ieot_work = DIV_ROUND_UP(entries, ieot_blocks);\n> +\n> +               /*\n> +                * no reason to write out the IEOT extension if we don't\n> +                * have enough blocks to utilize multi-threading\n> +                */\n> +               if (ieot_blocks > 1) {\n> +                       ieot = xcalloc(1, sizeof(struct index_entry_offset_table)\n> +                               + (ieot_blocks * sizeof(struct index_entry_offset)));\n> +                       ieot->nr = 0;\n> +                       ieot_work = DIV_ROUND_UP(entries, ieot_blocks);\n> +               }\n>         }\n>  #endif\n> \n> \n"},{"id":"359259","messageId":"xmqqo9ch9slw.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"bf0c24ac-6e2a-9a3e-835f-f21e763ab2c7@ramsayjones.plus.com","subject":"Re: [PATCH v6 4/7] config: add new index.threads config setting","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-28T22:15:07Z","receivedAt":"2018-09-28T22:15:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ramsay Jones <ramsay@ramsayjones.plus.com> writes:\n\n>>                 if (!nr) {\n>>                         ieot_blocks = istate->cache_nr / THREAD_COST;\n>> -                       if (ieot_blocks < 1)\n>> -                               ieot_blocks = 1;\n>>                         cpus = online_cpus();\n>>                         if (ieot_blocks > cpus - 1)\n>>                                 ieot_blocks = cpus - 1;\n>\n> So, am I reading this correctly - you need cpus > 2 before an\n> IEOT extension block is written out?\n>\n> OK.\n\nWhy should we be even calling online_cpus() in this codepath to\nwrite the index in a single thread to begin with?\n\nThe number of cpus that readers would use to read this index file\nhas nothing to do with the number of cpus available to this\nparticular writer process.  \n\n"},{"id":"359264","messageId":"20180929005128.GD23446@localhost","threadId":"49204","inReplyTo":"20180926195442.1380-4-benpeart@microsoft.com","subject":"Re: [PATCH v6 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2018-09-29T00:51:28Z","receivedAt":"2018-09-29T00:51:35Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Wed, Sep 26, 2018 at 03:54:38PM -0400, Ben Peart wrote:\n> diff --git a/read-cache.c b/read-cache.c\n> index 6ba99e2c96..80255d3088 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n\n> +static size_t read_eoie_extension(const char *mmap, size_t mmap_size)\n> +{\n\n<....>\n\n> +\tthe_hash_algo->final_fn(hash, &c);\n> +\tif (hashcmp(hash, (const unsigned char *)index))\n> +\t\treturn 0;\n\nPlease use !hasheq() instead of hashcmp().\n\n"},{"id":"359268","messageId":"20180929054520.GA21901@duynguyen.home","threadId":"49204","inReplyTo":"20180926195442.1380-4-benpeart@microsoft.com","subject":"Re: [PATCH v6 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-09-29T05:45:20Z","receivedAt":"2018-09-29T05:45:26Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Sep 26, 2018 at 03:54:38PM -0400, Ben Peart wrote:\n> +\n> +#define EOIE_SIZE (4 + GIT_SHA1_RAWSZ) /* <4-byte offset> + <20-byte hash> */\n> +#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n\nIf you make these variables instead of macros, you can use\nthe_hash_algo, which makes this code sha256-friendlier and probably\ncan explain less, e.g. ...\n\n> +\n> +static size_t read_eoie_extension(const char *mmap, size_t mmap_size)\n> +{\n> +\t/*\n> +\t * The end of index entries (EOIE) extension is guaranteed to be last\n> +\t * so that it can be found by scanning backwards from the EOF.\n> +\t *\n> +\t * \"EOIE\"\n> +\t * <4-byte length>\n> +\t * <4-byte offset>\n> +\t * <20-byte hash>\n> +\t */\n\n\tuint32_t EOIE_SIZE = 4 + the_hash_algo->rawsz;\n\tuint32_t EOIE_SIZE_WITH_HEADER = 4 + 4 + EOIE_SIZE;\n\n> +\tconst char *index, *eoie;\n> +\tuint32_t extsize;\n> +\tsize_t offset, src_offset;\n> +\tunsigned char hash[GIT_MAX_RAWSZ];\n> +\tgit_hash_ctx c;\n--\nDuy\n"},{"id":"359302","messageId":"xmqqzhw088mj.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20180929054520.GA21901@duynguyen.home","subject":"Re: [PATCH v6 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-09-29T18:24:20Z","receivedAt":"2018-09-29T18:24:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Duy Nguyen <pclouds@gmail.com> writes:\n\n> On Wed, Sep 26, 2018 at 03:54:38PM -0400, Ben Peart wrote:\n>> +\n>> +#define EOIE_SIZE (4 + GIT_SHA1_RAWSZ) /* <4-byte offset> + <20-byte hash> */\n>> +#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n>\n> If you make these variables instead of macros, you can use\n> the_hash_algo, which makes this code sha256-friendlier and probably\n> can explain less, e.g. ...\n>\n>> +\n>> +static size_t read_eoie_extension(const char *mmap, size_t mmap_size)\n>> +{\n>> +\t/*\n>> +\t * The end of index entries (EOIE) extension is guaranteed to be last\n>> +\t * so that it can be found by scanning backwards from the EOF.\n>> +\t *\n>> +\t * \"EOIE\"\n>> +\t * <4-byte length>\n>> +\t * <4-byte offset>\n>> +\t * <20-byte hash>\n\n20? ;-)\n\n>> +\t */\n>\n> \tuint32_t EOIE_SIZE = 4 + the_hash_algo->rawsz;\n> \tuint32_t EOIE_SIZE_WITH_HEADER = 4 + 4 + EOIE_SIZE;\n>\n>> +\tconst char *index, *eoie;\n>> +\tuint32_t extsize;\n>> +\tsize_t offset, src_offset;\n>> +\tunsigned char hash[GIT_MAX_RAWSZ];\n>> +\tgit_hash_ctx c;\n> --\n> Duy\n"},{"id":"359358","messageId":"48d6e62f-bf44-ffa1-befb-fef33ad00411@gmail.com","threadId":"49204","inReplyTo":"xmqqo9ch9slw.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v6 4/7] config: add new index.threads config setting","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-01T13:17:53Z","receivedAt":"2018-10-01T13:18:14Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 9/28/2018 6:15 PM, Junio C Hamano wrote:\n> Ramsay Jones <ramsay@ramsayjones.plus.com> writes:\n> \n>>>                  if (!nr) {\n>>>                          ieot_blocks = istate->cache_nr / THREAD_COST;\n>>> -                       if (ieot_blocks < 1)\n>>> -                               ieot_blocks = 1;\n>>>                          cpus = online_cpus();\n>>>                          if (ieot_blocks > cpus - 1)\n>>>                                  ieot_blocks = cpus - 1;\n>>\n>> So, am I reading this correctly - you need cpus > 2 before an\n>> IEOT extension block is written out?\n>>\n>> OK.\n> \n> Why should we be even calling online_cpus() in this codepath to\n> write the index in a single thread to begin with?\n> \n> The number of cpus that readers would use to read this index file\n> has nothing to do with the number of cpus available to this\n> particular writer process.\n> \n\nAs I mentioned in my other reply, this is optimizing for the most common \ncase where the index is read from the same machine that wrote it and the \nuser is taking the default settings (ie index.threads=true).\n\nAligning the number of blocks to the number of threads that will be \nprocessing them avoids situations where one thread may have up to double \nthe work to do as the other threads (for example, if there were 3 blocks \nto be processed by 2 threads).\n"},{"id":"359359","messageId":"20181001134556.33232-1-peartben@gmail.com","threadId":"49204","inReplyTo":"20180823154053.20212-1-benpeart@microsoft.com","subject":"[PATCH v7 0/7] speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-01T13:45:49Z","receivedAt":"2018-10-01T13:46:12Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"Thanks for all the feedback.\n\nThe biggest change since the last version is how this patch series interacts\nwith the split-index feature.  With a split index, most of the cache entries\nare stored in the shared index so would benefit from multi-threaded parsing.\nTo enable that, the EOIE and IEOT extensions are now written into the shared\nindex (rather than being stripped out like the other extensions).\n\nBecause of this, I can now update the tests in t1700-split-index.sh to have\nupdated SHA values that include the EOIE extension instead of disabling the\nextension.\n\nUsing p0002-read-cache.sh to generate some performance numbers shows how\neach of the various patches contribute to the overall performance win on a\nparticularly large repo.\n\nRepo w/3M files      Baseline  Optimize V4   Extensions      Entries\n--------------------------------------------------------------------------\n0002.1: read_cache   693.29    655.65 -5.4%  470.71 -32.1%   399.62 -42.4%\n\nNote how this cuts nearly 300ms off the index load time!\n\nBase Ref: master\nWeb-Diff: https://github.com/benpeart/git/commit/c1125a5d9a\nCheckout: git fetch https://github.com/benpeart/git read-index-multithread-v7 && git checkout c1125a5d9a\n\n\n### Patches\n\nBen Peart (6):\n  read-cache: clean up casting and byte decoding\n  eoie: add End of Index Entry (EOIE) extension\n  config: add new index.threads config setting\n  read-cache: load cache extensions on a worker thread\n  ieot: add Index Entry Offset Table (IEOT) extension\n  read-cache: load cache entries on worker threads\n\nNguyễn Thái Ngọc Duy (1):\n  read-cache.c: optimize reading index format v4\n\n Documentation/config.txt                 |   7 +\n Documentation/technical/index-format.txt |  41 ++\n config.c                                 |  18 +\n config.h                                 |   1 +\n read-cache.c                             | 749 +++++++++++++++++++----\n t/README                                 |   5 +\n t/t1700-split-index.sh                   |  13 +-\n 7 files changed, 715 insertions(+), 119 deletions(-)\n\n\nbase-commit: fe8321ec057f9231c26c29b364721568e58040f7\n-- \n2.18.0.windows.1\n\n\n"},{"id":"359360","messageId":"20181001134556.33232-2-peartben@gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-1-peartben@gmail.com","subject":"[PATCH v7 1/7] read-cache.c: optimize reading index format v4","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-01T13:45:50Z","receivedAt":"2018-10-01T13:46:14Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\nIndex format v4 requires some more computation to assemble a path\nbased on a previous one. The current code is not very efficient\nbecause\n\n - it doubles memory copy, we assemble the final path in a temporary\n   first before putting it back to a cache_entry\n\n - strbuf_remove() in expand_name_field() is not exactly a good fit\n   for stripping a part at the end, _setlen() would do the same job\n   and is much cheaper.\n\n - the open-coded loop to find the end of the string in\n   expand_name_field() can't beat an optimized strlen()\n\nThis patch avoids the temporary buffer and writes directly to the new\ncache_entry, which addresses the first two points. The last point\ncould also be avoided if the total string length fits in the first 12\nbits of ce_flags, if not we fall back to strlen().\n\nRunning \"test-tool read-cache 100\" on webkit.git (275k files), reading\nv2 only takes 4.226 seconds, while v4 takes 5.711 seconds, 35% more\ntime. The patch reduces read time on v4 to 4.319 seconds.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c | 128 ++++++++++++++++++++++++---------------------------\n 1 file changed, 60 insertions(+), 68 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 8d04d78a58..583a4fb1f8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1713,63 +1713,24 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *cache_entry_from_ondisk(struct mem_pool *mem_pool,\n-\t\t\t\t\t\t   struct ondisk_cache_entry *ondisk,\n-\t\t\t\t\t\t   unsigned int flags,\n-\t\t\t\t\t\t   const char *name,\n-\t\t\t\t\t\t   size_t len)\n-{\n-\tstruct cache_entry *ce = mem_pool__ce_alloc(mem_pool, len);\n-\n-\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n-\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n-\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n-\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n-\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n-\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n-\tce->ce_mode  = get_be32(&ondisk->mode);\n-\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n-\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n-\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n-\tce->ce_flags = flags & ~CE_NAMEMASK;\n-\tce->ce_namelen = len;\n-\tce->index = 0;\n-\thashcpy(ce->oid.hash, ondisk->sha1);\n-\tmemcpy(ce->name, name, len);\n-\tce->name[len] = '\\0';\n-\treturn ce;\n-}\n-\n-/*\n- * Adjacent cache entries tend to share the leading paths, so it makes\n- * sense to only store the differences in later entries.  In the v4\n- * on-disk format of the index, each on-disk cache entry stores the\n- * number of bytes to be stripped from the end of the previous name,\n- * and the bytes to append to the result, to come up with its name.\n- */\n-static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n-{\n-\tconst unsigned char *ep, *cp = (const unsigned char *)cp_;\n-\tsize_t len = decode_varint(&cp);\n-\n-\tif (name->len < len)\n-\t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n-\tstrbuf_add(name, cp, ep - cp);\n-\treturn (const char *)ep + 1 - cp_;\n-}\n-\n-static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n+static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n-\t\t\t\t\t    struct strbuf *previous_name)\n+\t\t\t\t\t    const struct cache_entry *previous_ce)\n {\n \tstruct cache_entry *ce;\n \tsize_t len;\n \tconst char *name;\n \tunsigned int flags;\n+\tsize_t copy_len;\n+\t/*\n+\t * Adjacent cache entries tend to share the leading paths, so it makes\n+\t * sense to only store the differences in later entries.  In the v4\n+\t * on-disk format of the index, each on-disk cache entry stores the\n+\t * number of bytes to be stripped from the end of the previous name,\n+\t * and the bytes to append to the result, to come up with its name.\n+\t */\n+\tint expand_name_field = istate->version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1789,21 +1750,54 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \telse\n \t\tname = ondisk->name;\n \n-\tif (!previous_name) {\n-\t\t/* v3 and earlier */\n-\t\tif (len == CE_NAMEMASK)\n-\t\t\tlen = strlen(name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags, name, len);\n+\tif (expand_name_field) {\n+\t\tconst unsigned char *cp = (const unsigned char *)name;\n+\t\tsize_t strip_len, previous_len;\n \n-\t\t*ent_size = ondisk_ce_size(ce);\n-\t} else {\n-\t\tunsigned long consumed;\n-\t\tconsumed = expand_name_field(previous_name, name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags,\n-\t\t\t\t\t     previous_name->buf,\n-\t\t\t\t\t     previous_name->len);\n+\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n+\t\tstrip_len = decode_varint(&cp);\n+\t\tif (previous_len < strip_len) {\n+\t\t\tif (previous_ce)\n+\t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n+\t\t\t\t    previous_ce->name);\n+\t\t\telse\n+\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t}\n+\t\tcopy_len = previous_len - strip_len;\n+\t\tname = (const char *)cp;\n+\t}\n+\n+\tif (len == CE_NAMEMASK) {\n+\t\tlen = strlen(name);\n+\t\tif (expand_name_field)\n+\t\t\tlen += copy_len;\n+\t}\n+\n+\tce = mem_pool__ce_alloc(istate->ce_mem_pool, len);\n+\n+\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n+\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n+\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n+\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n+\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n+\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n+\tce->ce_mode  = get_be32(&ondisk->mode);\n+\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n+\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n+\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n+\tce->ce_flags = flags & ~CE_NAMEMASK;\n+\tce->ce_namelen = len;\n+\tce->index = 0;\n+\thashcpy(ce->oid.hash, ondisk->sha1);\n \n-\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n+\tif (expand_name_field) {\n+\t\tif (copy_len)\n+\t\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n+\t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n+\t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n+\t} else {\n+\t\tmemcpy(ce->name, name, len + 1);\n+\t\t*ent_size = ondisk_ce_size(ce);\n \t}\n \treturn ce;\n }\n@@ -1898,7 +1892,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tconst struct cache_entry *previous_ce = NULL;\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1936,11 +1930,9 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->initialized = 1;\n \n \tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n \t} else {\n-\t\tprevious_name = NULL;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n \t}\n@@ -1952,12 +1944,12 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tunsigned long consumed;\n \n \t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n \t\tsrc_offset += consumed;\n+\t\tprevious_ce = ce;\n \t}\n-\tstrbuf_release(&previous_name_buf);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-- \n2.18.0.windows.1\n\n"},{"id":"359361","messageId":"20181001134556.33232-3-peartben@gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-1-peartben@gmail.com","subject":"[PATCH v7 2/7] read-cache: clean up casting and byte decoding","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-01T13:45:51Z","receivedAt":"2018-10-01T13:46:15Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nThis patch does a clean up pass to minimize the casting required to work\nwith the memory mapped index (mmap).\n\nIt also makes the decoding of network byte order more consistent by using\nget_be32() where possible.\n\nSigned-off-by: Ben Peart <peartben@gmail.com>\n---\n read-cache.c | 23 +++++++++++------------\n 1 file changed, 11 insertions(+), 12 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 583a4fb1f8..6ba99e2c96 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1650,7 +1650,7 @@ int verify_index_checksum;\n /* Allow fsck to force verification of the cache entry order. */\n int verify_ce_order;\n \n-static int verify_hdr(struct cache_header *hdr, unsigned long size)\n+static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n {\n \tgit_hash_ctx c;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n@@ -1674,7 +1674,7 @@ static int verify_hdr(struct cache_header *hdr, unsigned long size)\n }\n \n static int read_index_extension(struct index_state *istate,\n-\t\t\t\tconst char *ext, void *data, unsigned long sz)\n+\t\t\t\tconst char *ext, const char *data, unsigned long sz)\n {\n \tswitch (CACHE_EXT(ext)) {\n \tcase CACHE_EXT_TREE:\n@@ -1889,8 +1889,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tint fd, i;\n \tstruct stat st;\n \tunsigned long src_offset;\n-\tstruct cache_header *hdr;\n-\tvoid *mmap;\n+\tconst struct cache_header *hdr;\n+\tconst char *mmap;\n \tsize_t mmap_size;\n \tconst struct cache_entry *previous_ce = NULL;\n \n@@ -1918,7 +1918,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tdie_errno(\"unable to map index file\");\n \tclose(fd);\n \n-\thdr = mmap;\n+\thdr = (const struct cache_header *)mmap;\n \tif (verify_hdr(hdr, mmap_size) < 0)\n \t\tgoto unmap;\n \n@@ -1943,7 +1943,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tstruct cache_entry *ce;\n \t\tunsigned long consumed;\n \n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n \t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n@@ -1961,21 +1961,20 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\t * in 4-byte network byte order.\n \t\t */\n \t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n+\t\textsize = get_be32(mmap + src_offset + 4);\n \t\tif (read_index_extension(istate,\n-\t\t\t\t\t (const char *) mmap + src_offset,\n-\t\t\t\t\t (char *) mmap + src_offset + 8,\n+\t\t\t\t\t mmap + src_offset,\n+\t\t\t\t\t mmap + src_offset + 8,\n \t\t\t\t\t extsize) < 0)\n \t\t\tgoto unmap;\n \t\tsrc_offset += 8;\n \t\tsrc_offset += extsize;\n \t}\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \treturn istate->cache_nr;\n \n unmap:\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \tdie(\"index file corrupt\");\n }\n \n-- \n2.18.0.windows.1\n\n"},{"id":"359362","messageId":"20181001134556.33232-4-peartben@gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-1-peartben@gmail.com","subject":"[PATCH v7 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-01T13:45:52Z","receivedAt":"2018-10-01T13:46:17Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nThe End of Index Entry (EOIE) is used to locate the end of the variable\nlength index entries and the beginning of the extensions. Code can take\nadvantage of this to quickly locate the index extensions without having\nto parse through all of the index entries.\n\nBecause it must be able to be loaded before the variable length cache\nentries and other index extensions, this extension must be written last.\nThe signature for this extension is { 'E', 'O', 'I', 'E' }.\n\nThe extension consists of:\n\n- 32-bit offset to the end of the index entries\n\n- 160-bit SHA-1 over the extension types and their sizes (but not\ntheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\nlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\nthen the hash would be:\n\nSHA-1(\"TREE\" + <binary representation of N> +\n\t\"REUC\" + <binary representation of M>)\n\nSigned-off-by: Ben Peart <peartben@gmail.com>\n---\n Documentation/technical/index-format.txt |  23 ++++\n read-cache.c                             | 152 +++++++++++++++++++++--\n t/t1700-split-index.sh                   |   8 +-\n 3 files changed, 171 insertions(+), 12 deletions(-)\n\ndiff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\nindex db3572626b..6bc2d90f7f 100644\n--- a/Documentation/technical/index-format.txt\n+++ b/Documentation/technical/index-format.txt\n@@ -314,3 +314,26 @@ The remaining data of each directory block is grouped by type:\n \n   - An ewah bitmap, the n-th bit indicates whether the n-th index entry\n     is not CE_FSMONITOR_VALID.\n+\n+== End of Index Entry\n+\n+  The End of Index Entry (EOIE) is used to locate the end of the variable\n+  length index entries and the begining of the extensions. Code can take\n+  advantage of this to quickly locate the index extensions without having\n+  to parse through all of the index entries.\n+\n+  Because it must be able to be loaded before the variable length cache\n+  entries and other index extensions, this extension must be written last.\n+  The signature for this extension is { 'E', 'O', 'I', 'E' }.\n+\n+  The extension consists of:\n+\n+  - 32-bit offset to the end of the index entries\n+\n+  - 160-bit SHA-1 over the extension types and their sizes (but not\n+\ttheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\tlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\tthen the hash would be:\n+\n+\tSHA-1(\"TREE\" + <binary representation of N> +\n+\t\t\"REUC\" + <binary representation of M>)\ndiff --git a/read-cache.c b/read-cache.c\nindex 6ba99e2c96..af2605a168 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -43,6 +43,7 @@\n #define CACHE_EXT_LINK 0x6c696e6b\t  /* \"link\" */\n #define CACHE_EXT_UNTRACKED 0x554E5452\t  /* \"UNTR\" */\n #define CACHE_EXT_FSMONITOR 0x46534D4E\t  /* \"FSMN\" */\n+#define CACHE_EXT_ENDOFINDEXENTRIES 0x454F4945\t/* \"EOIE\" */\n \n /* changes that can be kept in $GIT_DIR/index (basically all extensions) */\n #define EXTMASK (RESOLVE_UNDO_CHANGED | CACHE_TREE_CHANGED | \\\n@@ -1693,6 +1694,9 @@ static int read_index_extension(struct index_state *istate,\n \tcase CACHE_EXT_FSMONITOR:\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n+\tcase CACHE_EXT_ENDOFINDEXENTRIES:\n+\t\t/* already handled in do_read_index() */\n+\t\tbreak;\n \tdefault:\n \t\tif (*ext < 'A' || 'Z' < *ext)\n \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n@@ -1883,6 +1887,9 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -2190,11 +2197,15 @@ static int ce_write(git_hash_ctx *context, int fd, void *data, unsigned int len)\n \treturn 0;\n }\n \n-static int write_index_ext_header(git_hash_ctx *context, int fd,\n-\t\t\t\t  unsigned int ext, unsigned int sz)\n+static int write_index_ext_header(git_hash_ctx *context, git_hash_ctx *eoie_context,\n+\t\t\t\t  int fd, unsigned int ext, unsigned int sz)\n {\n \text = htonl(ext);\n \tsz = htonl(sz);\n+\tif (eoie_context) {\n+\t\tthe_hash_algo->update_fn(eoie_context, &ext, 4);\n+\t\tthe_hash_algo->update_fn(eoie_context, &sz, 4);\n+\t}\n \treturn ((ce_write(context, fd, &ext, 4) < 0) ||\n \t\t(ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n }\n@@ -2437,7 +2448,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n {\n \tuint64_t start = getnanotime();\n \tint newfd = tempfile->fd;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, eoie_c;\n \tstruct cache_header hdr;\n \tint i, err = 0, removed, extended, hdr_version;\n \tstruct cache_entry **cache = istate->cache;\n@@ -2446,6 +2457,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct ondisk_cache_entry_extended ondisk;\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n+\toff_t offset;\n \n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n@@ -2479,6 +2491,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tif (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n \t\treturn -1;\n \n+\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n \tprevious_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n \n \tfor (i = 0; i < entries; i++) {\n@@ -2512,11 +2525,14 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\treturn err;\n \n \t/* Write extension data here */\n+\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\tthe_hash_algo->init_fn(&eoie_c);\n+\n \tif (!strip_extensions && istate->split_index) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\terr = write_link_extension(&sb, istate) < 0 ||\n-\t\t\twrite_index_ext_header(&c, newfd, CACHE_EXT_LINK,\n+\t\t\twrite_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_LINK,\n \t\t\t\t\t       sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2527,7 +2543,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tcache_tree_write(&sb, istate->cache_tree);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_TREE, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_TREE, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2537,7 +2553,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tresolve_undo_write(&sb, istate->resolve_undo);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_RESOLVE_UNDO,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_RESOLVE_UNDO,\n \t\t\t\t\t     sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2548,7 +2564,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_untracked_extension(&sb, istate->untracked);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_UNTRACKED,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_UNTRACKED,\n \t\t\t\t\t     sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2559,7 +2575,24 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_fsmonitor_extension(&sb, istate);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+\n+\t/*\n+\t * CACHE_EXT_ENDOFINDEXENTRIES must be written as the last entry before the SHA1\n+\t * so that it can be found and processed before all the index entries are\n+\t * read.  Write it out regardless of the strip_extensions parameter as we need it\n+\t * when loading the shared index.\n+\t */\n+\tif (offset) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\twrite_eoie_extension(&sb, &eoie_c, offset);\n+\t\terr = write_index_ext_header(&c, NULL, newfd, CACHE_EXT_ENDOFINDEXENTRIES, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2975,3 +3008,106 @@ int should_validate_cache_entries(void)\n \n \treturn validate_index_cache_entries;\n }\n+\n+#define EOIE_SIZE (4 + GIT_SHA1_RAWSZ) /* <4-byte offset> + <20-byte hash> */\n+#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n+\n+static size_t read_eoie_extension(const char *mmap, size_t mmap_size)\n+{\n+\t/*\n+\t * The end of index entries (EOIE) extension is guaranteed to be last\n+\t * so that it can be found by scanning backwards from the EOF.\n+\t *\n+\t * \"EOIE\"\n+\t * <4-byte length>\n+\t * <4-byte offset>\n+\t * <20-byte hash>\n+\t */\n+\tconst char *index, *eoie;\n+\tuint32_t extsize;\n+\tsize_t offset, src_offset;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\tgit_hash_ctx c;\n+\n+\t/* ensure we have an index big enough to contain an EOIE extension */\n+\tif (mmap_size < sizeof(struct cache_header) + EOIE_SIZE_WITH_HEADER + the_hash_algo->rawsz)\n+\t\treturn 0;\n+\n+\t/* validate the extension signature */\n+\tindex = eoie = mmap + mmap_size - EOIE_SIZE_WITH_HEADER - the_hash_algo->rawsz;\n+\tif (CACHE_EXT(index) != CACHE_EXT_ENDOFINDEXENTRIES)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/* validate the extension size */\n+\textsize = get_be32(index);\n+\tif (extsize != EOIE_SIZE)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * Validate the offset we're going to look for the first extension\n+\t * signature is after the index header and before the eoie extension.\n+\t */\n+\toffset = get_be32(index);\n+\tif (mmap + offset < mmap + sizeof(struct cache_header))\n+\t\treturn 0;\n+\tif (mmap + offset >= eoie)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * The hash is computed over extension types and their sizes (but not\n+\t * their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\t * long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\t * then the hash would be:\n+\t *\n+\t * SHA-1(\"TREE\" + <binary representation of N> +\n+\t *\t \"REUC\" + <binary representation of M>)\n+\t */\n+\tsrc_offset = offset;\n+\tthe_hash_algo->init_fn(&c);\n+\twhile (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\n+\t\t/* verify the extension size isn't so large it will wrap around */\n+\t\tif (src_offset + 8 + extsize < src_offset)\n+\t\t\treturn 0;\n+\n+\t\tthe_hash_algo->update_fn(&c, mmap + src_offset, 8);\n+\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\tthe_hash_algo->final_fn(hash, &c);\n+\tif (!hasheq(hash, (const unsigned char *)index))\n+\t\treturn 0;\n+\n+\t/* Validate that the extension offsets returned us back to the eoie extension. */\n+\tif (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n+\t\treturn 0;\n+\n+\treturn offset;\n+}\n+\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset)\n+{\n+\tuint32_t buffer;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\n+\t/* offset */\n+\tput_be32(&buffer, offset);\n+\tstrbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+\t/* hash */\n+\tthe_hash_algo->final_fn(hash, eoie_context);\n+\tstrbuf_add(sb, hash, the_hash_algo->rawsz);\n+}\ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex be22398a85..8e17f8e7a0 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -15,11 +15,11 @@ test_expect_success 'enable split index' '\n \tindexversion=$(test-tool index-version <.git/index) &&\n \tif test \"$indexversion\" = \"4\"\n \tthen\n-\t\town=432ef4b63f32193984f339431fd50ca796493569\n-\t\tbase=508851a7f0dfa8691e9f69c7f055865389012491\n+\t\town=3527df833c6c100d3d1d921a9a782d62a8be4b58\n+\t\tbase=746f7ab2ed44fb839efdfbffcf399d0b113fb4cb\n \telse\n-\t\town=8299b0bcd1ac364e5f1d7768efb62fa2da79a339\n-\t\tbase=39d890139ee5356c7ef572216cebcd27aa41f9df\n+\t\town=5e9b60117ece18da410ddecc8b8d43766a0e4204\n+\t\tbase=4370042739b31cd17a5c5cd6043a77c9a00df113\n \tfi &&\n \tcat >expect <<-EOF &&\n \town $own\n-- \n2.18.0.windows.1\n\n"},{"id":"359363","messageId":"20181001134556.33232-5-peartben@gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-1-peartben@gmail.com","subject":"[PATCH v7 4/7] config: add new index.threads config setting","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-01T13:45:53Z","receivedAt":"2018-10-01T13:46:19Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nAdd support for a new index.threads config setting which will be used to\ncontrol the threading code in do_read_index().  A value of 0 will tell the\nindex code to automatically determine the correct number of threads to use.\nA value of 1 will make the code single threaded.  A value greater than 1\nwill set the maximum number of threads to use.\n\nFor testing purposes, this setting can be overwritten by setting the\nGIT_TEST_INDEX_THREADS=<n> environment variable to a value greater than 0.\n\nSigned-off-by: Ben Peart <peartben@gmail.com>\n---\n Documentation/config.txt |  7 +++++++\n config.c                 | 18 ++++++++++++++++++\n config.h                 |  1 +\n t/README                 |  5 +++++\n t/t1700-split-index.sh   |  5 +++++\n 5 files changed, 36 insertions(+)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex ad0f4510c3..8fd973b76b 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -2413,6 +2413,13 @@ imap::\n \tThe configuration variables in the 'imap' section are described\n \tin linkgit:git-imap-send[1].\n \n+index.threads::\n+\tSpecifies the number of threads to spawn when loading the index.\n+\tThis is meant to reduce index load time on multiprocessor machines.\n+\tSpecifying 0 or 'true' will cause Git to auto-detect the number of\n+\tCPU's and set the number of threads accordingly. Specifying 1 or\n+\t'false' will disable multithreading. Defaults to 'true'.\n+\n index.version::\n \tSpecify the version with which new index files should be\n \tinitialized.  This does not affect existing repositories.\ndiff --git a/config.c b/config.c\nindex 3461993f0a..2ee29f6f86 100644\n--- a/config.c\n+++ b/config.c\n@@ -2289,6 +2289,24 @@ int git_config_get_fsmonitor(void)\n \treturn 0;\n }\n \n+int git_config_get_index_threads(void)\n+{\n+\tint is_bool, val = 0;\n+\n+\tval = git_env_ulong(\"GIT_TEST_INDEX_THREADS\", 0);\n+\tif (val)\n+\t\treturn val;\n+\n+\tif (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n+\t\tif (is_bool)\n+\t\t\treturn val ? 0 : 1;\n+\t\telse\n+\t\t\treturn val;\n+\t}\n+\n+\treturn 0; /* auto */\n+}\n+\n NORETURN\n void git_die_config_linenr(const char *key, const char *filename, int linenr)\n {\ndiff --git a/config.h b/config.h\nindex ab46e0165d..a06027e69b 100644\n--- a/config.h\n+++ b/config.h\n@@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n extern int git_config_get_split_index(void);\n extern int git_config_get_max_percent_split_change(void);\n extern int git_config_get_fsmonitor(void);\n+extern int git_config_get_index_threads(void);\n \n /* This dies if the configured or default date is in the future */\n extern int git_config_get_expiry(const char *key, const char **output);\ndiff --git a/t/README b/t/README\nindex 3ea6c85460..8f5c0620ea 100644\n--- a/t/README\n+++ b/t/README\n@@ -327,6 +327,11 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_INDEX_THREADS=<n> enables exercising the multi-threaded loading\n+of the index for the whole test suite by bypassing the default number of\n+cache entries and thread minimums. Setting this to 1 will make the\n+index loading single threaded.\n+\n Naming Tests\n ------------\n \ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex 8e17f8e7a0..ef9349bd70 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -6,7 +6,12 @@ test_description='split index mode tests'\n \n # We need total control of index splitting here\n sane_unset GIT_TEST_SPLIT_INDEX\n+\n+# Testing a hard coded SHA against an index with an extension\n+# that can vary from run to run is problematic so we disable\n+# those extensions.\n sane_unset GIT_FSMONITOR_TEST\n+sane_unset GIT_TEST_INDEX_THREADS\n \n test_expect_success 'enable split index' '\n \tgit config splitIndex.maxPercentChange 100 &&\n-- \n2.18.0.windows.1\n\n"},{"id":"359364","messageId":"20181001134556.33232-6-peartben@gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-1-peartben@gmail.com","subject":"[PATCH v7 5/7] read-cache: load cache extensions on a worker thread","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-01T13:45:54Z","receivedAt":"2018-10-01T13:46:21Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nThis patch helps address the CPU cost of loading the index by loading\nthe cache extensions on a worker thread in parallel with loading the cache\nentries.\n\nIn some cases, loading the extensions takes longer than loading the\ncache entries so this patch utilizes the new EOIE to start the thread to\nload the extensions before loading all the cache entries in parallel.\n\nThis is possible because the current extensions don't access the cache\nentries in the index_state structure so are OK that they don't all exist\nyet.\n\nThe CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\nextensions don't even get a pointer to the index so don't have access to the\ncache entries.\n\nCACHE_EXT_LINK only uses the index_state to initialize the split index.\nCACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\nupdate and dirty flags.\n\nI used p0002-read-cache.sh to generate some performance data:\n\n\tTest w/100,000 files reduced the time by 0.53%\n\tTest w/1,000,000 files reduced the time by 27.78%\n\nSigned-off-by: Ben Peart <peartben@gmail.com>\n---\n read-cache.c | 97 +++++++++++++++++++++++++++++++++++++++++++---------\n 1 file changed, 81 insertions(+), 16 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex af2605a168..77083ab8bb 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -23,6 +23,7 @@\n #include \"split-index.h\"\n #include \"utf8.h\"\n #include \"fsmonitor.h\"\n+#include \"thread-utils.h\"\n \n /* Mask for the name length in ce_flags in the on-disk index */\n \n@@ -1890,6 +1891,46 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n \n+struct load_index_extensions\n+{\n+#ifndef NO_PTHREADS\n+\tpthread_t pthread;\n+#endif\n+\tstruct index_state *istate;\n+\tconst char *mmap;\n+\tsize_t mmap_size;\n+\tunsigned long src_offset;\n+};\n+\n+static void *load_index_extensions(void *_data)\n+{\n+\tstruct load_index_extensions *p = _data;\n+\tunsigned long src_offset = p->src_offset;\n+\n+\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, p->mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\t\tif (read_index_extension(p->istate,\n+\t\t\tp->mmap + src_offset,\n+\t\t\tp->mmap + src_offset + 8,\n+\t\t\textsize) < 0) {\n+\t\t\tmunmap((void *)p->mmap, p->mmap_size);\n+\t\t\tdie(_(\"index file corrupt\"));\n+\t\t}\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\n+\treturn NULL;\n+}\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -1900,6 +1941,11 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tconst char *mmap;\n \tsize_t mmap_size;\n \tconst struct cache_entry *previous_ce = NULL;\n+\tstruct load_index_extensions p;\n+\tsize_t extension_offset = 0;\n+#ifndef NO_PTHREADS\n+\tint nr_threads;\n+#endif\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1936,6 +1982,30 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n \tistate->initialized = 1;\n \n+\tp.istate = istate;\n+\tp.mmap = mmap;\n+\tp.mmap_size = mmap_size;\n+\n+#ifndef NO_PTHREADS\n+\tnr_threads = git_config_get_index_threads();\n+\tif (!nr_threads)\n+\t\tnr_threads = online_cpus();\n+\n+\tif (nr_threads > 1) {\n+\t\textension_offset = read_eoie_extension(mmap, mmap_size);\n+\t\tif (extension_offset) {\n+\t\t\tint err;\n+\n+\t\t\tp.src_offset = extension_offset;\n+\t\t\terr = pthread_create(&p.pthread, NULL, load_index_extensions, &p);\n+\t\t\tif (err)\n+\t\t\t\tdie(_(\"unable to create load_index_extensions thread: %s\"), strerror(err));\n+\n+\t\t\tnr_threads--;\n+\t\t}\n+\t}\n+#endif\n+\n \tif (istate->version == 4) {\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n@@ -1960,22 +2030,17 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-\twhile (src_offset <= mmap_size - the_hash_algo->rawsz - 8) {\n-\t\t/* After an array of active_nr index entries,\n-\t\t * there can be arbitrary number of extended\n-\t\t * sections, each of which is prefixed with\n-\t\t * extension name (4-byte) and section length\n-\t\t * in 4-byte network byte order.\n-\t\t */\n-\t\tuint32_t extsize;\n-\t\textsize = get_be32(mmap + src_offset + 4);\n-\t\tif (read_index_extension(istate,\n-\t\t\t\t\t mmap + src_offset,\n-\t\t\t\t\t mmap + src_offset + 8,\n-\t\t\t\t\t extsize) < 0)\n-\t\t\tgoto unmap;\n-\t\tsrc_offset += 8;\n-\t\tsrc_offset += extsize;\n+\t/* if we created a thread, join it otherwise load the extensions on the primary thread */\n+#ifndef NO_PTHREADS\n+\tif (extension_offset) {\n+\t\tint ret = pthread_join(p.pthread, NULL);\n+\t\tif (ret)\n+\t\t\tdie(_(\"unable to join load_index_extensions thread: %s\"), strerror(ret));\n+\t}\n+#endif\n+\tif (!extension_offset) {\n+\t\tp.src_offset = src_offset;\n+\t\tload_index_extensions(&p);\n \t}\n \tmunmap((void *)mmap, mmap_size);\n \treturn istate->cache_nr;\n-- \n2.18.0.windows.1\n\n"},{"id":"359365","messageId":"20181001134556.33232-7-peartben@gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-1-peartben@gmail.com","subject":"[PATCH v7 6/7] ieot: add Index Entry Offset Table (IEOT) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-01T13:45:55Z","receivedAt":"2018-10-01T13:46:24Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nThis patch enables addressing the CPU cost of loading the index by adding\nadditional data to the index that will allow us to efficiently multi-\nthread the loading and conversion of cache entries.\n\nIt accomplishes this by adding an (optional) index extension that is a\ntable of offsets to blocks of cache entries in the index file.  To make\nthis work for V4 indexes, when writing the cache entries, it periodically\n\"resets\" the prefix-compression by encoding the current entry as if the\npath name for the previous entry is completely different and saves the\noffset of that entry in the IEOT.  Basically, with V4 indexes, it\ngenerates offsets into blocks of prefix-compressed entries.\n\nSigned-off-by: Ben Peart <peartben@gmail.com>\n---\n Documentation/technical/index-format.txt |  18 +++\n read-cache.c                             | 173 +++++++++++++++++++++++\n 2 files changed, 191 insertions(+)\n\ndiff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\nindex 6bc2d90f7f..7c4d67aa6a 100644\n--- a/Documentation/technical/index-format.txt\n+++ b/Documentation/technical/index-format.txt\n@@ -337,3 +337,21 @@ The remaining data of each directory block is grouped by type:\n \n \tSHA-1(\"TREE\" + <binary representation of N> +\n \t\t\"REUC\" + <binary representation of M>)\n+\n+== Index Entry Offset Table\n+\n+  The Index Entry Offset Table (IEOT) is used to help address the CPU\n+  cost of loading the index by enabling multi-threading the process of\n+  converting cache entries from the on-disk format to the in-memory format.\n+  The signature for this extension is { 'I', 'E', 'O', 'T' }.\n+\n+  The extension consists of:\n+\n+  - 32-bit version (currently 1)\n+\n+  - A number of index offset entries each consisting of:\n+\n+    - 32-bit offset from the begining of the file to the first cache entry\n+\tin this block of entries.\n+\n+    - 32-bit count of cache entries in this block\ndiff --git a/read-cache.c b/read-cache.c\nindex 77083ab8bb..9557376e78 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -45,6 +45,7 @@\n #define CACHE_EXT_UNTRACKED 0x554E5452\t  /* \"UNTR\" */\n #define CACHE_EXT_FSMONITOR 0x46534D4E\t  /* \"FSMN\" */\n #define CACHE_EXT_ENDOFINDEXENTRIES 0x454F4945\t/* \"EOIE\" */\n+#define CACHE_EXT_INDEXENTRYOFFSETTABLE 0x49454F54 /* \"IEOT\" */\n \n /* changes that can be kept in $GIT_DIR/index (basically all extensions) */\n #define EXTMASK (RESOLVE_UNDO_CHANGED | CACHE_TREE_CHANGED | \\\n@@ -1696,6 +1697,7 @@ static int read_index_extension(struct index_state *istate,\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n \tcase CACHE_EXT_ENDOFINDEXENTRIES:\n+\tcase CACHE_EXT_INDEXENTRYOFFSETTABLE:\n \t\t/* already handled in do_read_index() */\n \t\tbreak;\n \tdefault:\n@@ -1888,6 +1890,23 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+struct index_entry_offset\n+{\n+\t/* starting byte offset into index file, count of index entries in this block */\n+\tint offset, nr;\n+};\n+\n+struct index_entry_offset_table\n+{\n+\tint nr;\n+\tstruct index_entry_offset entries[0];\n+};\n+\n+#ifndef NO_PTHREADS\n+static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset);\n+static void write_ieot_extension(struct strbuf *sb, struct index_entry_offset_table *ieot);\n+#endif\n+\n static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n \n@@ -1931,6 +1950,15 @@ static void *load_index_extensions(void *_data)\n \treturn NULL;\n }\n \n+/*\n+ * Mostly randomly chosen maximum thread counts: we\n+ * cap the parallelism to online_cpus() threads, and we want\n+ * to have at least 10000 cache entries per thread for it to\n+ * be worth starting a thread.\n+ */\n+\n+#define THREAD_COST\t\t(10000)\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -2523,6 +2551,9 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n+\tint ieot_work = 1;\n+\tstruct index_entry_offset_table *ieot = NULL;\n+\tint nr;\n \n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n@@ -2556,7 +2587,38 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tif (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n \t\treturn -1;\n \n+#ifndef NO_PTHREADS\n+\tif ((nr = git_config_get_index_threads()) != 1) {\n+\t\tint ieot_blocks, cpus;\n+\n+\t\t/*\n+\t\t * ensure default number of ieot blocks maps evenly to the\n+\t\t * default number of threads that will process them\n+\t\t */\n+\t\tif (!nr) {\n+\t\t\tieot_blocks = istate->cache_nr / THREAD_COST;\n+\t\t\tcpus = online_cpus();\n+\t\t\tif (ieot_blocks > cpus - 1)\n+\t\t\t\tieot_blocks = cpus - 1;\n+\t\t} else {\n+\t\t\tieot_blocks = nr;\n+\t\t}\n+\n+\t\t/*\n+\t\t * no reason to write out the IEOT extension if we don't\n+\t\t * have enough blocks to utilize multi-threading\n+\t\t */\n+\t\tif (ieot_blocks > 1) {\n+\t\t\tieot = xcalloc(1, sizeof(struct index_entry_offset_table)\n+\t\t\t\t+ (ieot_blocks * sizeof(struct index_entry_offset)));\n+\t\t\tieot->nr = 0;\n+\t\t\tieot_work = DIV_ROUND_UP(entries, ieot_blocks);\n+\t\t}\n+\t}\n+#endif\n+\n \toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\tnr = 0;\n \tprevious_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n \n \tfor (i = 0; i < entries; i++) {\n@@ -2578,11 +2640,31 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \n \t\t\tdrop_cache_tree = 1;\n \t\t}\n+\t\tif (ieot && i && (i % ieot_work == 0)) {\n+\t\t\tieot->entries[ieot->nr].nr = nr;\n+\t\t\tieot->entries[ieot->nr].offset = offset;\n+\t\t\tieot->nr++;\n+\t\t\t/*\n+\t\t\t * If we have a V4 index, set the first byte to an invalid\n+\t\t\t * character to ensure there is nothing common with the previous\n+\t\t\t * entry\n+\t\t\t */\n+\t\t\tif (previous_name)\n+\t\t\t\tprevious_name->buf[0] = 0;\n+\t\t\tnr = 0;\n+\t\t\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\t\t}\n \t\tif (ce_write_entry(&c, newfd, ce, previous_name, (struct ondisk_cache_entry *)&ondisk) < 0)\n \t\t\terr = -1;\n \n \t\tif (err)\n \t\t\tbreak;\n+\t\tnr++;\n+\t}\n+\tif (ieot && nr) {\n+\t\tieot->entries[ieot->nr].nr = nr;\n+\t\tieot->entries[ieot->nr].offset = offset;\n+\t\tieot->nr++;\n \t}\n \tstrbuf_release(&previous_name_buf);\n \n@@ -2593,6 +2675,26 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n \tthe_hash_algo->init_fn(&eoie_c);\n \n+\t/*\n+\t * Lets write out CACHE_EXT_INDEXENTRYOFFSETTABLE first so that we\n+\t * can minimze the number of extensions we have to scan through to\n+\t * find it during load.  Write it out regardless of the\n+\t * strip_extensions parameter as we need it when loading the shared\n+\t * index.\n+\t */\n+#ifndef NO_PTHREADS\n+\tif (ieot) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\twrite_ieot_extension(&sb, ieot);\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_INDEXENTRYOFFSETTABLE, sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+#endif\n+\n \tif (!strip_extensions && istate->split_index) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n@@ -3176,3 +3278,74 @@ static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context,\n \tthe_hash_algo->final_fn(hash, eoie_context);\n \tstrbuf_add(sb, hash, the_hash_algo->rawsz);\n }\n+\n+#ifndef NO_PTHREADS\n+#define IEOT_VERSION\t(1)\n+\n+static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset)\n+{\n+       const char *index = NULL;\n+       uint32_t extsize, ext_version;\n+       struct index_entry_offset_table *ieot;\n+       int i, nr;\n+\n+       /* find the IEOT extension */\n+       if (!offset)\n+\t       return NULL;\n+       while (offset <= mmap_size - the_hash_algo->rawsz - 8) {\n+\t       extsize = get_be32(mmap + offset + 4);\n+\t       if (CACHE_EXT((mmap + offset)) == CACHE_EXT_INDEXENTRYOFFSETTABLE) {\n+\t\t       index = mmap + offset + 4 + 4;\n+\t\t       break;\n+\t       }\n+\t       offset += 8;\n+\t       offset += extsize;\n+       }\n+       if (!index)\n+\t       return NULL;\n+\n+       /* validate the version is IEOT_VERSION */\n+       ext_version = get_be32(index);\n+       if (ext_version != IEOT_VERSION)\n+\t       return NULL;\n+       index += sizeof(uint32_t);\n+\n+       /* extension size - version bytes / bytes per entry */\n+       nr = (extsize - sizeof(uint32_t)) / (sizeof(uint32_t) + sizeof(uint32_t));\n+       if (!nr)\n+\t       return NULL;\n+       ieot = xmalloc(sizeof(struct index_entry_offset_table)\n+\t       + (nr * sizeof(struct index_entry_offset)));\n+       ieot->nr = nr;\n+       for (i = 0; i < nr; i++) {\n+\t       ieot->entries[i].offset = get_be32(index);\n+\t       index += sizeof(uint32_t);\n+\t       ieot->entries[i].nr = get_be32(index);\n+\t       index += sizeof(uint32_t);\n+       }\n+\n+       return ieot;\n+}\n+\n+static void write_ieot_extension(struct strbuf *sb, struct index_entry_offset_table *ieot)\n+{\n+       uint32_t buffer;\n+       int i;\n+\n+       /* version */\n+       put_be32(&buffer, IEOT_VERSION);\n+       strbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+       /* ieot */\n+       for (i = 0; i < ieot->nr; i++) {\n+\n+\t       /* offset */\n+\t       put_be32(&buffer, ieot->entries[i].offset);\n+\t       strbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+\t       /* count */\n+\t       put_be32(&buffer, ieot->entries[i].nr);\n+\t       strbuf_add(sb, &buffer, sizeof(uint32_t));\n+       }\n+}\n+#endif\n-- \n2.18.0.windows.1\n\n"},{"id":"359366","messageId":"20181001134556.33232-8-peartben@gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-1-peartben@gmail.com","subject":"[PATCH v7 7/7] read-cache: load cache entries on worker threads","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-01T13:45:56Z","receivedAt":"2018-10-01T13:46:25Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nThis patch helps address the CPU cost of loading the index by utilizing\nthe Index Entry Offset Table (IEOT) to divide loading and conversion of\nthe cache entries across multiple threads in parallel.\n\nI used p0002-read-cache.sh to generate some performance data:\n\nTest w/100,000 files reduced the time by 32.24%\nTest w/1,000,000 files reduced the time by -4.77%\n\nNote that on the 1,000,000 files case, multi-threading the cache entry parsing\ndoes not yield a performance win.  This is because the cost to parse the\nindex extensions in this repo, far outweigh the cost of loading the cache\nentries.\n\nThe high cost of parsing the index extensions is driven by the cache tree\nand the untracked cache extensions. As this is currently the longest pole,\nany reduction in this time will reduce the overall index load times so is\nworth further investigation in another patch series.\n\nSigned-off-by: Ben Peart <peartben@gmail.com>\n---\n read-cache.c | 224 +++++++++++++++++++++++++++++++++++++++++++--------\n 1 file changed, 189 insertions(+), 35 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 9557376e78..14402a0738 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1720,7 +1720,8 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *create_from_disk(struct index_state *istate,\n+static struct cache_entry *create_from_disk(struct mem_pool *ce_mem_pool,\n+\t\t\t\t\t    unsigned int version,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n \t\t\t\t\t    const struct cache_entry *previous_ce)\n@@ -1737,7 +1738,7 @@ static struct cache_entry *create_from_disk(struct index_state *istate,\n \t * number of bytes to be stripped from the end of the previous name,\n \t * and the bytes to append to the result, to come up with its name.\n \t */\n-\tint expand_name_field = istate->version == 4;\n+\tint expand_name_field = version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1761,16 +1762,17 @@ static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\tconst unsigned char *cp = (const unsigned char *)name;\n \t\tsize_t strip_len, previous_len;\n \n-\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n+\t\t/* If we're at the begining of a block, ignore the previous name */\n \t\tstrip_len = decode_varint(&cp);\n-\t\tif (previous_len < strip_len) {\n-\t\t\tif (previous_ce)\n+\t\tif (previous_ce) {\n+\t\t\tprevious_len = previous_ce->ce_namelen;\n+\t\t\tif (previous_len < strip_len)\n \t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n-\t\t\t\t    previous_ce->name);\n-\t\t\telse\n-\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t\t\t\tprevious_ce->name);\n+\t\t\tcopy_len = previous_len - strip_len;\n+\t\t} else {\n+\t\t\tcopy_len = 0;\n \t\t}\n-\t\tcopy_len = previous_len - strip_len;\n \t\tname = (const char *)cp;\n \t}\n \n@@ -1780,7 +1782,7 @@ static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\t\tlen += copy_len;\n \t}\n \n-\tce = mem_pool__ce_alloc(istate->ce_mem_pool, len);\n+\tce = mem_pool__ce_alloc(ce_mem_pool, len);\n \n \tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n \tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n@@ -1950,6 +1952,52 @@ static void *load_index_extensions(void *_data)\n \treturn NULL;\n }\n \n+/*\n+ * A helper function that will load the specified range of cache entries\n+ * from the memory mapped file and add them to the given index.\n+ */\n+static unsigned long load_cache_entry_block(struct index_state *istate,\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, const char *mmap,\n+\t\t\tunsigned long start_offset, const struct cache_entry *previous_ce)\n+{\n+\tint i;\n+\tunsigned long src_offset = start_offset;\n+\n+\tfor (i = offset; i < offset + nr; i++) {\n+\t\tstruct ondisk_cache_entry *disk_ce;\n+\t\tstruct cache_entry *ce;\n+\t\tunsigned long consumed;\n+\n+\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n+\t\tce = create_from_disk(ce_mem_pool, istate->version, disk_ce, &consumed, previous_ce);\n+\t\tset_index_entry(istate, i, ce);\n+\n+\t\tsrc_offset += consumed;\n+\t\tprevious_ce = ce;\n+\t}\n+\treturn src_offset - start_offset;\n+}\n+\n+static unsigned long load_all_cache_entries(struct index_state *istate,\n+\t\t\tconst char *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tunsigned long consumed;\n+\n+\tif (istate->version == 4) {\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n+\t} else {\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t\testimate_cache_size(mmap_size, istate->cache_nr));\n+\t}\n+\n+\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, NULL);\n+\treturn consumed;\n+}\n+\n+#ifndef NO_PTHREADS\n+\n /*\n  * Mostly randomly chosen maximum thread counts: we\n  * cap the parallelism to online_cpus() threads, and we want\n@@ -1959,20 +2007,125 @@ static void *load_index_extensions(void *_data)\n \n #define THREAD_COST\t\t(10000)\n \n+struct load_cache_entries_thread_data\n+{\n+\tpthread_t pthread;\n+\tstruct index_state *istate;\n+\tstruct mem_pool *ce_mem_pool;\n+\tint offset;\n+\tconst char *mmap;\n+\tstruct index_entry_offset_table *ieot;\n+\tint ieot_offset;        /* starting index into the ieot array */\n+\tint ieot_work;          /* count of ieot entries to process */\n+\tunsigned long consumed;\t/* return # of bytes in index file processed */\n+};\n+\n+/*\n+ * A thread proc to run the load_cache_entries() computation\n+ * across multiple background threads.\n+ */\n+static void *load_cache_entries_thread(void *_data)\n+{\n+\tstruct load_cache_entries_thread_data *p = _data;\n+\tint i;\n+\n+\t/* iterate across all ieot blocks assigned to this thread */\n+\tfor (i = p->ieot_offset; i < p->ieot_offset + p->ieot_work; i++) {\n+\t\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool, p->offset, p->ieot->entries[i].nr, p->mmap, p->ieot->entries[i].offset, NULL);\n+\t\tp->offset += p->ieot->entries[i].nr;\n+\t}\n+\treturn NULL;\n+}\n+\n+static unsigned long load_cache_entries_threaded(struct index_state *istate, const char *mmap, size_t mmap_size,\n+\t\t\tunsigned long src_offset, int nr_threads, struct index_entry_offset_table *ieot)\n+{\n+\tint i, offset, ieot_work, ieot_offset, err;\n+\tstruct load_cache_entries_thread_data *data;\n+\tunsigned long consumed = 0;\n+\tint nr;\n+\n+\t/* a little sanity checking */\n+\tif (istate->name_hash_initialized)\n+\t\tBUG(\"the name hash isn't thread safe\");\n+\n+\tmem_pool_init(&istate->ce_mem_pool, 0);\n+\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n+\n+\t/* ensure we have no more threads than we have blocks to process */\n+\tif (nr_threads > ieot->nr)\n+\t\tnr_threads = ieot->nr;\n+\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n+\n+\toffset = ieot_offset = 0;\n+\tieot_work = DIV_ROUND_UP(ieot->nr, nr_threads);\n+\tfor (i = 0; i < nr_threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = &data[i];\n+\t\tint j;\n+\n+\t\tif (ieot_offset + ieot_work > ieot->nr)\n+\t\t\tieot_work = ieot->nr - ieot_offset;\n+\n+\t\tp->istate = istate;\n+\t\tp->offset = offset;\n+\t\tp->mmap = mmap;\n+\t\tp->ieot = ieot;\n+\t\tp->ieot_offset = ieot_offset;\n+\t\tp->ieot_work = ieot_work;\n+\n+\t\t/* create a mem_pool for each thread */\n+\t\tnr = 0;\n+\t\tfor (j = p->ieot_offset; j < p->ieot_offset + p->ieot_work; j++)\n+\t\t\tnr += p->ieot->entries[j].nr;\n+\t\tif (istate->version == 4) {\n+\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\testimate_cache_size_from_compressed(nr));\n+\t\t}\n+\t\telse {\n+\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\testimate_cache_size(mmap_size, nr));\n+\t\t}\n+\n+\t\terr = pthread_create(&p->pthread, NULL, load_cache_entries_thread, p);\n+\t\tif (err)\n+\t\t\tdie(_(\"unable to create load_cache_entries thread: %s\"), strerror(err));\n+\n+\t\t/* increment by the number of cache entries in the ieot block being processed */\n+\t\tfor (j = 0; j < ieot_work; j++)\n+\t\t\toffset += ieot->entries[ieot_offset + j].nr;\n+\t\tieot_offset += ieot_work;\n+\t}\n+\n+\tfor (i = 0; i < nr_threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = &data[i];\n+\n+\t\terr = pthread_join(p->pthread, NULL);\n+\t\tif (err)\n+\t\t\tdie(_(\"unable to join load_cache_entries thread: %s\"), strerror(err));\n+\t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n+\t\tconsumed += p->consumed;\n+\t}\n+\n+\tfree(data);\n+\n+\treturn consumed;\n+}\n+#endif\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n-\tint fd, i;\n+\tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n \tconst struct cache_header *hdr;\n \tconst char *mmap;\n \tsize_t mmap_size;\n-\tconst struct cache_entry *previous_ce = NULL;\n \tstruct load_index_extensions p;\n \tsize_t extension_offset = 0;\n #ifndef NO_PTHREADS\n-\tint nr_threads;\n+\tint nr_threads, cpus;\n+\tstruct index_entry_offset_table *ieot = NULL;\n #endif\n \n \tif (istate->initialized)\n@@ -2014,10 +2167,18 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tp.mmap = mmap;\n \tp.mmap_size = mmap_size;\n \n+\tsrc_offset = sizeof(*hdr);\n+\n #ifndef NO_PTHREADS\n \tnr_threads = git_config_get_index_threads();\n-\tif (!nr_threads)\n-\t\tnr_threads = online_cpus();\n+\n+\t/* TODO: does creating more threads than cores help? */\n+\tif (!nr_threads) {\n+\t\tnr_threads = istate->cache_nr / THREAD_COST;\n+\t\tcpus = online_cpus();\n+\t\tif (nr_threads > cpus)\n+\t\t\tnr_threads = cpus;\n+\t}\n \n \tif (nr_threads > 1) {\n \t\textension_offset = read_eoie_extension(mmap, mmap_size);\n@@ -2032,29 +2193,22 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\t\tnr_threads--;\n \t\t}\n \t}\n-#endif\n-\n-\tif (istate->version == 4) {\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n-\t} else {\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n-\t}\n \n-\tsrc_offset = sizeof(*hdr);\n-\tfor (i = 0; i < istate->cache_nr; i++) {\n-\t\tstruct ondisk_cache_entry *disk_ce;\n-\t\tstruct cache_entry *ce;\n-\t\tunsigned long consumed;\n+\t/*\n+\t * Locate and read the index entry offset table so that we can use it\n+\t * to multi-thread the reading of the cache entries.\n+\t */\n+\tif (extension_offset && nr_threads > 1)\n+\t\tieot = read_ieot_extension(mmap, mmap_size, extension_offset);\n \n-\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n-\t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n-\t\tset_index_entry(istate, i, ce);\n+\tif (ieot)\n+\t\tsrc_offset += load_cache_entries_threaded(istate, mmap, mmap_size, src_offset, nr_threads, ieot);\n+\telse\n+\t\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+#else\n+\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+#endif\n \n-\t\tsrc_offset += consumed;\n-\t\tprevious_ce = ce;\n-\t}\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-- \n2.18.0.windows.1\n\n"},{"id":"359370","messageId":"20181001150613.GK23446@localhost","threadId":"49204","inReplyTo":"48d6e62f-bf44-ffa1-befb-fef33ad00411@gmail.com","subject":"Re: [PATCH v6 4/7] config: add new index.threads config setting","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2018-10-01T15:06:13Z","receivedAt":"2018-10-01T15:06:20Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Oct 01, 2018 at 09:17:53AM -0400, Ben Peart wrote:\n> \n> \n> On 9/28/2018 6:15 PM, Junio C Hamano wrote:\n> >Ramsay Jones <ramsay@ramsayjones.plus.com> writes:\n> >\n> >>>                 if (!nr) {\n> >>>                         ieot_blocks = istate->cache_nr / THREAD_COST;\n> >>>-                       if (ieot_blocks < 1)\n> >>>-                               ieot_blocks = 1;\n> >>>                         cpus = online_cpus();\n> >>>                         if (ieot_blocks > cpus - 1)\n> >>>                                 ieot_blocks = cpus - 1;\n> >>\n> >>So, am I reading this correctly - you need cpus > 2 before an\n> >>IEOT extension block is written out?\n> >>\n> >>OK.\n> >\n> >Why should we be even calling online_cpus() in this codepath to\n> >write the index in a single thread to begin with?\n> >\n> >The number of cpus that readers would use to read this index file\n> >has nothing to do with the number of cpus available to this\n> >particular writer process.\n> >\n> \n> As I mentioned in my other reply, this is optimizing for the most common\n> case where the index is read from the same machine that wrote it and the\n> user is taking the default settings (ie index.threads=true).\n\nI think this is a reasonable assumption to make, but it should be\nmentioned in the relevant commit message.  Alas, as far as I can tell,\nnot a single commit message has been updated in v7.\n\n> Aligning the number of blocks to the number of threads that will be\n> processing them avoids situations where one thread may have up to double the\n> work to do as the other threads (for example, if there were 3 blocks to be\n> processed by 2 threads).\n"},{"id":"359371","messageId":"CACsJy8DRLPmrGD1podPJ12G3VitsK3dQnq+2sOjCiQj6N4ayTQ@mail.gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-3-peartben@gmail.com","subject":"Re: [PATCH v7 2/7] read-cache: clean up casting and byte decoding","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-10-01T15:10:18Z","receivedAt":"2018-10-01T15:10:47Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Oct 1, 2018 at 3:46 PM Ben Peart <peartben@gmail.com> wrote:\n>\n> From: Ben Peart <benpeart@microsoft.com>\n>\n> This patch does a clean up pass to minimize the casting required to work\n> with the memory mapped index (mmap).\n>\n> It also makes the decoding of network byte order more consistent by using\n> get_be32() where possible.\n>\n> Signed-off-by: Ben Peart <peartben@gmail.com>\n> ---\n>  read-cache.c | 23 +++++++++++------------\n>  1 file changed, 11 insertions(+), 12 deletions(-)\n>\n> diff --git a/read-cache.c b/read-cache.c\n> index 583a4fb1f8..6ba99e2c96 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1650,7 +1650,7 @@ int verify_index_checksum;\n>  /* Allow fsck to force verification of the cache entry order. */\n>  int verify_ce_order;\n>\n> -static int verify_hdr(struct cache_header *hdr, unsigned long size)\n> +static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n\nOK more constness. Good.\n\n>  {\n>         git_hash_ctx c;\n>         unsigned char hash[GIT_MAX_RAWSZ];\n> @@ -1674,7 +1674,7 @@ static int verify_hdr(struct cache_header *hdr, unsigned long size)\n>  }\n>\n>  static int read_index_extension(struct index_state *istate,\n> -                               const char *ext, void *data, unsigned long sz)\n> +                               const char *ext, const char *data, unsigned long sz)\n\nBut it's not clear why you need to change the data type from void * to\nchar * here. I guess all the consumer functions take 'const char *'\nanyway, so it's best to use 'const char *'?\n\nNot worth a reroll (to give a reason why you do this in the commit\nmessage), unless there are other changes.\n-- \nDuy\n"},{"id":"359372","messageId":"20181001151716.GL23446@localhost","threadId":"49204","inReplyTo":"20181001134556.33232-4-peartben@gmail.com","subject":"Re: [PATCH v7 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2018-10-01T15:17:16Z","receivedAt":"2018-10-01T15:17:22Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Oct 01, 2018 at 09:45:52AM -0400, Ben Peart wrote:\n> From: Ben Peart <benpeart@microsoft.com>\n> \n> The End of Index Entry (EOIE) is used to locate the end of the variable\n> length index entries and the beginning of the extensions. Code can take\n> advantage of this to quickly locate the index extensions without having\n> to parse through all of the index entries.\n> \n> Because it must be able to be loaded before the variable length cache\n> entries and other index extensions, this extension must be written last.\n> The signature for this extension is { 'E', 'O', 'I', 'E' }.\n> \n> The extension consists of:\n> \n> - 32-bit offset to the end of the index entries\n> \n> - 160-bit SHA-1 over the extension types and their sizes (but not\n> their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n> long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n> then the hash would be:\n> \n> SHA-1(\"TREE\" + <binary representation of N> +\n> \t\"REUC\" + <binary representation of M>)\n> \n> Signed-off-by: Ben Peart <peartben@gmail.com>\n\nI think the commit message should explicitly mention that this this\nextension\n\n  - will always be written and why,\n  - but is optional, so other Git implementations not supporting it will\n    have no troubles reading the index,\n  - and that it is written even to the shared index and why, and that\n    because of this the index checksums in t1700 had to be updated.\n\n"},{"id":"359373","messageId":"CACsJy8CX0TwVydzmqsjHK+W7tcaDgRCgeU3Gmc-bA1Ecf=Yz6A@mail.gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-4-peartben@gmail.com","subject":"Re: [PATCH v7 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-10-01T15:30:25Z","receivedAt":"2018-10-01T15:30:54Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Oct 1, 2018 at 3:46 PM Ben Peart <peartben@gmail.com> wrote:\n> @@ -2479,6 +2491,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>         if (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n>                 return -1;\n>\n> +       offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n\nNote, lseek() could in theory return -1 on error. Looking at the error\ncode list in the man page it's pretty unlikely though, unless\n\n> +static size_t read_eoie_extension(const char *mmap, size_t mmap_size)\n> +{\n> +       /*\n> +        * The end of index entries (EOIE) extension is guaranteed to be last\n> +        * so that it can be found by scanning backwards from the EOF.\n> +        *\n> +        * \"EOIE\"\n> +        * <4-byte length>\n> +        * <4-byte offset>\n> +        * <20-byte hash>\n> +        */\n> +       const char *index, *eoie;\n> +       uint32_t extsize;\n> +       size_t offset, src_offset;\n> +       unsigned char hash[GIT_MAX_RAWSZ];\n> +       git_hash_ctx c;\n> +\n> +       /* ensure we have an index big enough to contain an EOIE extension */\n> +       if (mmap_size < sizeof(struct cache_header) + EOIE_SIZE_WITH_HEADER + the_hash_algo->rawsz)\n\nUsing sizeof() for on-disk structures could be dangerous because you\ndon't know how much padding there could be (I'm not sure if it's\nactually specified in the C language spec). I've checked, on at least\nx86 and amd64, sizeof(struct cache_header) is 12 bytes, but I don't\nknow if there are any crazy architectures out there that set higher\npadding.\n-- \nDuy\n"},{"id":"359375","messageId":"CACsJy8A2+P6RM5OOhke=Ptc2iPB81fGu0BF-Ven9am_UEThB8A@mail.gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-6-peartben@gmail.com","subject":"Re: [PATCH v7 5/7] read-cache: load cache extensions on a worker thread","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-10-01T15:50:39Z","receivedAt":"2018-10-01T15:51:08Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Oct 1, 2018 at 3:46 PM Ben Peart <peartben@gmail.com> wrote:\n> @@ -1890,6 +1891,46 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n>  static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n>  static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n>\n> +struct load_index_extensions\n> +{\n> +#ifndef NO_PTHREADS\n> +       pthread_t pthread;\n> +#endif\n> +       struct index_state *istate;\n> +       const char *mmap;\n> +       size_t mmap_size;\n> +       unsigned long src_offset;\n> +};\n> +\n> +static void *load_index_extensions(void *_data)\n> +{\n> +       struct load_index_extensions *p = _data;\n> +       unsigned long src_offset = p->src_offset;\n> +\n> +       while (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n> +               /* After an array of active_nr index entries,\n> +                * there can be arbitrary number of extended\n> +                * sections, each of which is prefixed with\n> +                * extension name (4-byte) and section length\n> +                * in 4-byte network byte order.\n> +                */\n> +               uint32_t extsize;\n> +               memcpy(&extsize, p->mmap + src_offset + 4, 4);\n> +               extsize = ntohl(extsize);\n\nThis could be get_be32() so that the next person will not need to do\nanother cleanup patch.\n\n> +               if (read_index_extension(p->istate,\n> +                       p->mmap + src_offset,\n> +                       p->mmap + src_offset + 8,\n> +                       extsize) < 0) {\n\nThis alignment is misleading because the conditions are aligned with\nthe code block below. If you can't align it with the '(', then just\nadd another tab.\n\n> +                       munmap((void *)p->mmap, p->mmap_size);\n\nThis made me pause for a bit since we should not need to cast back to\nvoid *. It turns out you need this because mmap pointer is const. But\nyou don't even need to munmap here. We're dying, the OS will clean\neverything up.\n\n> +                       die(_(\"index file corrupt\"));\n> +               }\n> +               src_offset += 8;\n> +               src_offset += extsize;\n> +       }\n> +\n> +       return NULL;\n> +}\n-- \nDuy\n"},{"id":"359377","messageId":"CACsJy8B9dd-N=w13XP2FuHRfqK2tmzuNx0WN-ZhuchssG6RUdg@mail.gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-7-peartben@gmail.com","subject":"Re: [PATCH v7 6/7] ieot: add Index Entry Offset Table (IEOT) extension","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-10-01T16:27:17Z","receivedAt":"2018-10-01T16:27:46Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Oct 1, 2018 at 3:46 PM Ben Peart <peartben@gmail.com> wrote:\n> @@ -1888,6 +1890,23 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n>         return ondisk_size + entries * per_entry;\n>  }\n>\n> +struct index_entry_offset\n> +{\n> +       /* starting byte offset into index file, count of index entries in this block */\n> +       int offset, nr;\n\nuint32_t?\n\n> +};\n> +\n> +struct index_entry_offset_table\n> +{\n> +       int nr;\n> +       struct index_entry_offset entries[0];\n\nUse FLEX_ARRAY. Some compilers are not happy with an array of zero\nitems if I remember correctly.\n\n> @@ -2523,6 +2551,9 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>         struct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n>         int drop_cache_tree = istate->drop_cache_tree;\n>         off_t offset;\n> +       int ieot_work = 1;\n> +       struct index_entry_offset_table *ieot = NULL;\n> +       int nr;\n\nThere are a bunch of stuff going on in this function, maybe rename\nthis to nr_threads or nr_blocks to be less generic.\n\n>\n>         for (i = removed = extended = 0; i < entries; i++) {\n>                 if (cache[i]->ce_flags & CE_REMOVE)\n> @@ -2556,7 +2587,38 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>         if (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n>                 return -1;\n>\n> +#ifndef NO_PTHREADS\n> +       if ((nr = git_config_get_index_threads()) != 1) {\n\nMaybe keep this assignment out of \"if\".\n\n> +               int ieot_blocks, cpus;\n> +\n> +               /*\n> +                * ensure default number of ieot blocks maps evenly to the\n> +                * default number of threads that will process them\n> +                */\n> +               if (!nr) {\n> +                       ieot_blocks = istate->cache_nr / THREAD_COST;\n> +                       cpus = online_cpus();\n> +                       if (ieot_blocks > cpus - 1)\n> +                               ieot_blocks = cpus - 1;\n\nThe \" - 1\" here is for extension thread, yes? Probably worth a comment.\n\n> +               } else {\n> +                       ieot_blocks = nr;\n> +               }\n> +\n> +               /*\n> +                * no reason to write out the IEOT extension if we don't\n> +                * have enough blocks to utilize multi-threading\n> +                */\n> +               if (ieot_blocks > 1) {\n> +                       ieot = xcalloc(1, sizeof(struct index_entry_offset_table)\n> +                               + (ieot_blocks * sizeof(struct index_entry_offset)));\n\nUse FLEX_ALLOC_MEM() after you declare ..._table with FLEX_ARRAY.\n\nThis ieot needs to be freed also and should be before any \"return -1\"\nin this function.\n\n> +                       ieot->nr = 0;\n> +                       ieot_work = DIV_ROUND_UP(entries, ieot_blocks);\n\nPerhaps a better name for ioet_work? This looks like the number of\ncache entries per block.\n\n> +               }\n> +       }\n> +#endif\n> +\n>         offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n> +       nr = 0;\n\nEh.. repurpose nr to count cache entries now? It's kinda hard to follow.\n\n>         previous_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n>\n>         for (i = 0; i < entries; i++) {\n> @@ -2578,11 +2640,31 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>\n>                         drop_cache_tree = 1;\n>                 }\n> +               if (ieot && i && (i % ieot_work == 0)) {\n> +                       ieot->entries[ieot->nr].nr = nr;\n> +                       ieot->entries[ieot->nr].offset = offset;\n> +                       ieot->nr++;\n> +                       /*\n> +                        * If we have a V4 index, set the first byte to an invalid\n> +                        * character to ensure there is nothing common with the previous\n> +                        * entry\n> +                        */\n> +                       if (previous_name)\n> +                               previous_name->buf[0] = 0;\n> +                       nr = 0;\n> +                       offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n\nThis only works correctly if the ce_write_entry() from the last\niteration has flushed everything to out to newfd. Maybe it does, but\nit's error prone to rely on that in my opinion. Maybe we need an\nexplicit ce_write_flush() here to make sure.\n\n> +               }\n>                 if (ce_write_entry(&c, newfd, ce, previous_name, (struct ondisk_cache_entry *)&ondisk) < 0)\n>                         err = -1;\n>\n>                 if (err)\n>                         break;\n> +               nr++;\n> +       }\n> +       if (ieot && nr) {\n> +               ieot->entries[ieot->nr].nr = nr;\n> +               ieot->entries[ieot->nr].offset = offset;\n> +               ieot->nr++;\n>         }\n>         strbuf_release(&previous_name_buf);\n>\n> @@ -2593,6 +2675,26 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>         offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n>         the_hash_algo->init_fn(&eoie_c);\n>\n> +       /*\n> +        * Lets write out CACHE_EXT_INDEXENTRYOFFSETTABLE first so that we\n> +        * can minimze the number of extensions we have to scan through to\n\ns/minimze/minimize/\n\n> +        * find it during load.  Write it out regardless of the\n> +        * strip_extensions parameter as we need it when loading the shared\n> +        * index.\n> +        */\n> +#ifndef NO_PTHREADS\n> +       if (ieot) {\n> +               struct strbuf sb = STRBUF_INIT;\n> +\n> +               write_ieot_extension(&sb, ieot);\n> +               err = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_INDEXENTRYOFFSETTABLE, sb.len) < 0\n> +                       || ce_write(&c, newfd, sb.buf, sb.len) < 0;\n> +               strbuf_release(&sb);\n> +               if (err)\n> +                       return -1;\n> +       }\n> +#endif\n> +\n>         if (!strip_extensions && istate->split_index) {\n>                 struct strbuf sb = STRBUF_INIT;\n>\n> @@ -3176,3 +3278,74 @@ static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context,\n>         the_hash_algo->final_fn(hash, eoie_context);\n>         strbuf_add(sb, hash, the_hash_algo->rawsz);\n>  }\n> +\n> +#ifndef NO_PTHREADS\n> +#define IEOT_VERSION   (1)\n> +\n> +static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset)\n> +{\n> +       const char *index = NULL;\n> +       uint32_t extsize, ext_version;\n> +       struct index_entry_offset_table *ieot;\n> +       int i, nr;\n> +\n> +       /* find the IEOT extension */\n> +       if (!offset)\n> +              return NULL;\n> +       while (offset <= mmap_size - the_hash_algo->rawsz - 8) {\n> +              extsize = get_be32(mmap + offset + 4);\n> +              if (CACHE_EXT((mmap + offset)) == CACHE_EXT_INDEXENTRYOFFSETTABLE) {\n> +                      index = mmap + offset + 4 + 4;\n> +                      break;\n> +              }\n> +              offset += 8;\n> +              offset += extsize;\n> +       }\n\nMaybe refactor this loop. I think I've seen this in at least two\nplaces now. Probably three?\n\n> +       if (!index)\n> +              return NULL;\n> +\n> +       /* validate the version is IEOT_VERSION */\n> +       ext_version = get_be32(index);\n> +       if (ext_version != IEOT_VERSION)\n> +              return NULL;\n\nReport the error (e.g. \"unsupported version\" or something)\n\n> +       index += sizeof(uint32_t);\n> +\n> +       /* extension size - version bytes / bytes per entry */\n> +       nr = (extsize - sizeof(uint32_t)) / (sizeof(uint32_t) + sizeof(uint32_t));\n\nDo we need to check if \"(extsize - version) % sizeof(entry) == 0\"?\n\n> +       if (!nr)\n> +              return NULL;\n> +       ieot = xmalloc(sizeof(struct index_entry_offset_table)\n> +              + (nr * sizeof(struct index_entry_offset)));\n> +       ieot->nr = nr;\n> +       for (i = 0; i < nr; i++) {\n> +              ieot->entries[i].offset = get_be32(index);\n> +              index += sizeof(uint32_t);\n> +              ieot->entries[i].nr = get_be32(index);\n> +              index += sizeof(uint32_t);\n> +       }\n> +\n> +       return ieot;\n> +}\n> +\n> +static void write_ieot_extension(struct strbuf *sb, struct index_entry_offset_table *ieot)\n> +{\n> +       uint32_t buffer;\n> +       int i;\n> +\n> +       /* version */\n> +       put_be32(&buffer, IEOT_VERSION);\n> +       strbuf_add(sb, &buffer, sizeof(uint32_t));\n> +\n> +       /* ieot */\n> +       for (i = 0; i < ieot->nr; i++) {\n> +\n> +              /* offset */\n> +              put_be32(&buffer, ieot->entries[i].offset);\n> +              strbuf_add(sb, &buffer, sizeof(uint32_t));\n> +\n> +              /* count */\n> +              put_be32(&buffer, ieot->entries[i].nr);\n> +              strbuf_add(sb, &buffer, sizeof(uint32_t));\n> +       }\n> +}\n> +#endif\n> --\n> 2.18.0.windows.1\n>\n\n\n-- \nDuy\n"},{"id":"359378","messageId":"CACsJy8D7Pbg6xMZBfCiz_7_=reY3Os4R_70wc65VMxbu2=Kqjw@mail.gmail.com","threadId":"49204","inReplyTo":"20181001134556.33232-8-peartben@gmail.com","subject":"Re: [PATCH v7 7/7] read-cache: load cache entries on worker threads","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-10-01T17:09:46Z","receivedAt":"2018-10-01T17:10:16Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Oct 1, 2018 at 3:46 PM Ben Peart <peartben@gmail.com> wrote:\n> +/*\n> + * A helper function that will load the specified range of cache entries\n> + * from the memory mapped file and add them to the given index.\n> + */\n> +static unsigned long load_cache_entry_block(struct index_state *istate,\n> +                       struct mem_pool *ce_mem_pool, int offset, int nr, const char *mmap,\n\nPlease use unsigned long for offset (here and in the thread_data\nstruct). We should use off_t instead, but that's out of scope. At\nleast keep offset type consistent in here.\n\n> +                       unsigned long start_offset, const struct cache_entry *previous_ce)\n\nI don't think you want to pass previous_ce in. You always pass NULL\nanyway. And if this function is about loading a block (i.e. at block\nboundary) then initial previous_ce _must_ be NULL or things break\nhorribly.\n\n> @@ -1959,20 +2007,125 @@ static void *load_index_extensions(void *_data)\n>\n>  #define THREAD_COST            (10000)\n>\n> +struct load_cache_entries_thread_data\n> +{\n> +       pthread_t pthread;\n> +       struct index_state *istate;\n> +       struct mem_pool *ce_mem_pool;\n> +       int offset;\n> +       const char *mmap;\n> +       struct index_entry_offset_table *ieot;\n> +       int ieot_offset;        /* starting index into the ieot array */\n\nIf it's an index, maybe just name it ieot_index and we can get rid of\nthe comment.\n\n> +       int ieot_work;          /* count of ieot entries to process */\n\nMaybe instead of saving the whole \"ieot\" table here. Add\n\n     struct index_entry_offset *blocks;\n\nwhich points to the starting block for this thread and rename that\nmysterious (to me) ieot_work to nr_blocks. The thread will have access\nfrom blocks[0] to blocks[nr_blocks - 1]\n\n> +       unsigned long consumed; /* return # of bytes in index file processed */\n> +};\n> +\n> +/*\n> + * A thread proc to run the load_cache_entries() computation\n> + * across multiple background threads.\n> + */\n> +static void *load_cache_entries_thread(void *_data)\n> +{\n> +       struct load_cache_entries_thread_data *p = _data;\n> +       int i;\n> +\n> +       /* iterate across all ieot blocks assigned to this thread */\n> +       for (i = p->ieot_offset; i < p->ieot_offset + p->ieot_work; i++) {\n> +               p->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool, p->offset, p->ieot->entries[i].nr, p->mmap, p->ieot->entries[i].offset, NULL);\n\nPlease wrap this long line.\n\n> +               p->offset += p->ieot->entries[i].nr;\n> +       }\n> +       return NULL;\n> +}\n> +\n> +static unsigned long load_cache_entries_threaded(struct index_state *istate, const char *mmap, size_t mmap_size,\n> +                       unsigned long src_offset, int nr_threads, struct index_entry_offset_table *ieot)\n> +{\n> +       int i, offset, ieot_work, ieot_offset, err;\n> +       struct load_cache_entries_thread_data *data;\n> +       unsigned long consumed = 0;\n> +       int nr;\n> +\n> +       /* a little sanity checking */\n> +       if (istate->name_hash_initialized)\n> +               BUG(\"the name hash isn't thread safe\");\n> +\n> +       mem_pool_init(&istate->ce_mem_pool, 0);\n> +       data = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n\nwe normally use sizeof(*data) instead of sizeof(struct ...)\n\n> +\n> +       /* ensure we have no more threads than we have blocks to process */\n> +       if (nr_threads > ieot->nr)\n> +               nr_threads = ieot->nr;\n> +       data = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n\neh.. reallocate the same \"data\"?\n\n> +\n> +       offset = ieot_offset = 0;\n> +       ieot_work = DIV_ROUND_UP(ieot->nr, nr_threads);\n> +       for (i = 0; i < nr_threads; i++) {\n> +               struct load_cache_entries_thread_data *p = &data[i];\n> +               int j;\n> +\n> +               if (ieot_offset + ieot_work > ieot->nr)\n> +                       ieot_work = ieot->nr - ieot_offset;\n> +\n> +               p->istate = istate;\n> +               p->offset = offset;\n> +               p->mmap = mmap;\n> +               p->ieot = ieot;\n> +               p->ieot_offset = ieot_offset;\n> +               p->ieot_work = ieot_work;\n> +\n> +               /* create a mem_pool for each thread */\n> +               nr = 0;\n\nSince nr is only used in this for loop. Declare it in this scope\ninstead of declaring it for the whole function.\n\n> +               for (j = p->ieot_offset; j < p->ieot_offset + p->ieot_work; j++)\n> +                       nr += p->ieot->entries[j].nr;\n> +               if (istate->version == 4) {\n> +                       mem_pool_init(&p->ce_mem_pool,\n> +                               estimate_cache_size_from_compressed(nr));\n> +               }\n> +               else {\n> +                       mem_pool_init(&p->ce_mem_pool,\n> +                               estimate_cache_size(mmap_size, nr));\n> +               }\n\nMaybe keep this mem_pool_init code inside load_cache_entries_thread(),\nsimilar to how you do it for load_cache_entries_thread(). It's mostly\nto keep this loop shorter to see (and understand), of course\nparallelizing this mem_pool_init() is just noise.\n\n> +\n> +               err = pthread_create(&p->pthread, NULL, load_cache_entries_thread, p);\n> +               if (err)\n> +                       die(_(\"unable to create load_cache_entries thread: %s\"), strerror(err));\n> +\n> +               /* increment by the number of cache entries in the ieot block being processed */\n> +               for (j = 0; j < ieot_work; j++)\n> +                       offset += ieot->entries[ieot_offset + j].nr;\n\nI wonder if it makes things simpler if you store cache_entry _index_\nin entrie[] array instead of storing the number of entries. You can\neasily calculate nr then by doing entries[i].index -\nentries[i-1].index. And you can count multiple blocks the same way,\nwithout looping like this.\n\n> +               ieot_offset += ieot_work;\n> +       }\n> +\n> +       for (i = 0; i < nr_threads; i++) {\n> +               struct load_cache_entries_thread_data *p = &data[i];\n> +\n> +               err = pthread_join(p->pthread, NULL);\n> +               if (err)\n> +                       die(_(\"unable to join load_cache_entries thread: %s\"), strerror(err));\n> +               mem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n> +               consumed += p->consumed;\n> +       }\n> +\n> +       free(data);\n> +\n> +       return consumed;\n> +}\n> +#endif\n> +\n>  /* remember to discard_cache() before reading a different cache! */\n>  int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>  {\n> -       int fd, i;\n> +       int fd;\n>         struct stat st;\n>         unsigned long src_offset;\n>         const struct cache_header *hdr;\n>         const char *mmap;\n>         size_t mmap_size;\n> -       const struct cache_entry *previous_ce = NULL;\n>         struct load_index_extensions p;\n>         size_t extension_offset = 0;\n>  #ifndef NO_PTHREADS\n> -       int nr_threads;\n> +       int nr_threads, cpus;\n> +       struct index_entry_offset_table *ieot = NULL;\n>  #endif\n>\n>         if (istate->initialized)\n> @@ -2014,10 +2167,18 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>         p.mmap = mmap;\n>         p.mmap_size = mmap_size;\n>\n> +       src_offset = sizeof(*hdr);\n\nOK we've been doing this since forever, sizeof(struct cache_header)\nprobably does not have extra padding on any supported platform.\n\n> @@ -2032,29 +2193,22 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>                         nr_threads--;\n>                 }\n>         }\n> -#endif\n> -\n> -       if (istate->version == 4) {\n> -               mem_pool_init(&istate->ce_mem_pool,\n> -                             estimate_cache_size_from_compressed(istate->cache_nr));\n> -       } else {\n> -               mem_pool_init(&istate->ce_mem_pool,\n> -                             estimate_cache_size(mmap_size, istate->cache_nr));\n> -       }\n>\n> -       src_offset = sizeof(*hdr);\n> -       for (i = 0; i < istate->cache_nr; i++) {\n> -               struct ondisk_cache_entry *disk_ce;\n> -               struct cache_entry *ce;\n> -               unsigned long consumed;\n> +       /*\n> +        * Locate and read the index entry offset table so that we can use it\n> +        * to multi-thread the reading of the cache entries.\n> +        */\n> +       if (extension_offset && nr_threads > 1)\n> +               ieot = read_ieot_extension(mmap, mmap_size, extension_offset);\n\nYou need to free ieot at some point.\n\n>\n> -               disk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n> -               ce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n> -               set_index_entry(istate, i, ce);\n> +       if (ieot)\n> +               src_offset += load_cache_entries_threaded(istate, mmap, mmap_size, src_offset, nr_threads, ieot);\n> +       else\n> +               src_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n> +#else\n> +       src_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n> +#endif\n>\n> -               src_offset += consumed;\n> -               previous_ce = ce;\n> -       }\n>         istate->timestamp.sec = st.st_mtime;\n>         istate->timestamp.nsec = ST_MTIME_NSEC(st);\n>\n> --\n> 2.18.0.windows.1\n>\n-- \nDuy\n"},{"id":"359417","messageId":"dd8622f1-9005-a989-6934-2b9cd24c87a9@gmail.com","threadId":"49204","inReplyTo":"20181001151716.GL23446@localhost","subject":"Re: [PATCH v7 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-02T14:34:07Z","receivedAt":"2018-10-02T14:34:14Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 10/1/2018 11:17 AM, SZEDER Gábor wrote:\n> On Mon, Oct 01, 2018 at 09:45:52AM -0400, Ben Peart wrote:\n>> From: Ben Peart <benpeart@microsoft.com>\n>>\n>> The End of Index Entry (EOIE) is used to locate the end of the variable\n>> length index entries and the beginning of the extensions. Code can take\n>> advantage of this to quickly locate the index extensions without having\n>> to parse through all of the index entries.\n>>\n>> Because it must be able to be loaded before the variable length cache\n>> entries and other index extensions, this extension must be written last.\n>> The signature for this extension is { 'E', 'O', 'I', 'E' }.\n>>\n>> The extension consists of:\n>>\n>> - 32-bit offset to the end of the index entries\n>>\n>> - 160-bit SHA-1 over the extension types and their sizes (but not\n>> their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n>> long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n>> then the hash would be:\n>>\n>> SHA-1(\"TREE\" + <binary representation of N> +\n>> \t\"REUC\" + <binary representation of M>)\n>>\n>> Signed-off-by: Ben Peart <peartben@gmail.com>\n> \n> I think the commit message should explicitly mention that this this\n> extension\n> \n>    - will always be written and why,\n>    - but is optional, so other Git implementations not supporting it will\n>      have no troubles reading the index,\n>    - and that it is written even to the shared index and why, and that\n>      because of this the index checksums in t1700 had to be updated.\n> \n\nSure, I'll add that additional information to the commit message on the \nnext spin.\n"},{"id":"359425","messageId":"8fed4b13-71f3-4657-8058-39400009b32a@gmail.com","threadId":"49204","inReplyTo":"CACsJy8A2+P6RM5OOhke=Ptc2iPB81fGu0BF-Ven9am_UEThB8A@mail.gmail.com","subject":"Re: [PATCH v7 5/7] read-cache: load cache extensions on a worker thread","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-02T15:00:58Z","receivedAt":"2018-10-02T15:01:05Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 10/1/2018 11:50 AM, Duy Nguyen wrote:\n> On Mon, Oct 1, 2018 at 3:46 PM Ben Peart <peartben@gmail.com> wrote:\n>> @@ -1890,6 +1891,46 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n>>   static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n>>   static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n>>\n>> +struct load_index_extensions\n>> +{\n>> +#ifndef NO_PTHREADS\n>> +       pthread_t pthread;\n>> +#endif\n>> +       struct index_state *istate;\n>> +       const char *mmap;\n>> +       size_t mmap_size;\n>> +       unsigned long src_offset;\n>> +};\n>> +\n>> +static void *load_index_extensions(void *_data)\n>> +{\n>> +       struct load_index_extensions *p = _data;\n>> +       unsigned long src_offset = p->src_offset;\n>> +\n>> +       while (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n>> +               /* After an array of active_nr index entries,\n>> +                * there can be arbitrary number of extended\n>> +                * sections, each of which is prefixed with\n>> +                * extension name (4-byte) and section length\n>> +                * in 4-byte network byte order.\n>> +                */\n>> +               uint32_t extsize;\n>> +               memcpy(&extsize, p->mmap + src_offset + 4, 4);\n>> +               extsize = ntohl(extsize);\n> \n> This could be get_be32() so that the next person will not need to do\n> another cleanup patch.\n> \n\nGood point, it was existing code so I focused on doing the minimal \nchange possible but I can clean it up since I'm touching it already.\n\n>> +               if (read_index_extension(p->istate,\n>> +                       p->mmap + src_offset,\n>> +                       p->mmap + src_offset + 8,\n>> +                       extsize) < 0) {\n> \n> This alignment is misleading because the conditions are aligned with\n> the code block below. If you can't align it with the '(', then just\n> add another tab.\n> \n\nDitto. I'll make it:\n\n\t\tuint32_t extsize = get_be32(p->mmap + src_offset + 4);\n\t\tif (read_index_extension(p->istate,\n\t\t\t\t\t p->mmap + src_offset,\n\t\t\t\t\t p->mmap + src_offset + 8,\n\t\t\t\t\t extsize) < 0) {\n\t\t\tmunmap((void *)p->mmap, p->mmap_size);\n\t\t\tdie(_(\"index file corrupt\"));\n\t\t}\n\n\n>> +                       munmap((void *)p->mmap, p->mmap_size);\n> \n> This made me pause for a bit since we should not need to cast back to\n> void *. It turns out you need this because mmap pointer is const. But\n> you don't even need to munmap here. We're dying, the OS will clean\n> everything up.\n> \n\nI had the same thought about \"we're about to die so why bother calling \nmunmap() here\" but I decided rather than change it, I'd follow the \nexisting pattern just in case there was some platform/bug that required \nit.  I apparently doesn't cause harm as it's been that way a long time.\n\n>> +                       die(_(\"index file corrupt\"));\n>> +               }\n>> +               src_offset += 8;\n>> +               src_offset += extsize;\n>> +       }\n>> +\n>> +       return NULL;\n>> +}\n"},{"id":"359427","messageId":"424326cc-0cfe-3af8-d2f7-6c124cb077cd@gmail.com","threadId":"49204","inReplyTo":"CACsJy8CX0TwVydzmqsjHK+W7tcaDgRCgeU3Gmc-bA1Ecf=Yz6A@mail.gmail.com","subject":"Re: [PATCH v7 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-02T15:13:03Z","receivedAt":"2018-10-02T15:13:11Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 10/1/2018 11:30 AM, Duy Nguyen wrote:\n> On Mon, Oct 1, 2018 at 3:46 PM Ben Peart <peartben@gmail.com> wrote:\n>> @@ -2479,6 +2491,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>>          if (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n>>                  return -1;\n>>\n>> +       offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n> \n> Note, lseek() could in theory return -1 on error. Looking at the error\n> code list in the man page it's pretty unlikely though, unless\n> \n\nGood catch. I'll add the logic to check for an error.\n\n>> +static size_t read_eoie_extension(const char *mmap, size_t mmap_size)\n>> +{\n>> +       /*\n>> +        * The end of index entries (EOIE) extension is guaranteed to be last\n>> +        * so that it can be found by scanning backwards from the EOF.\n>> +        *\n>> +        * \"EOIE\"\n>> +        * <4-byte length>\n>> +        * <4-byte offset>\n>> +        * <20-byte hash>\n>> +        */\n>> +       const char *index, *eoie;\n>> +       uint32_t extsize;\n>> +       size_t offset, src_offset;\n>> +       unsigned char hash[GIT_MAX_RAWSZ];\n>> +       git_hash_ctx c;\n>> +\n>> +       /* ensure we have an index big enough to contain an EOIE extension */\n>> +       if (mmap_size < sizeof(struct cache_header) + EOIE_SIZE_WITH_HEADER + the_hash_algo->rawsz)\n> \n> Using sizeof() for on-disk structures could be dangerous because you\n> don't know how much padding there could be (I'm not sure if it's\n> actually specified in the C language spec). I've checked, on at least\n> x86 and amd64, sizeof(struct cache_header) is 12 bytes, but I don't\n> know if there are any crazy architectures out there that set higher\n> padding.\n> \n\nThis must be safe as the same code has been in do_read_index() and \nverify_index_from() for a long time.\n"},{"id":"359434","messageId":"351b9746-6c2e-a658-3f51-71c1f4cbc3ac@gmail.com","threadId":"49204","inReplyTo":"CACsJy8B9dd-N=w13XP2FuHRfqK2tmzuNx0WN-ZhuchssG6RUdg@mail.gmail.com","subject":"Re: [PATCH v7 6/7] ieot: add Index Entry Offset Table (IEOT) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-02T16:34:21Z","receivedAt":"2018-10-02T16:34:28Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 10/1/2018 12:27 PM, Duy Nguyen wrote:\n> On Mon, Oct 1, 2018 at 3:46 PM Ben Peart <peartben@gmail.com> wrote:\n>> @@ -1888,6 +1890,23 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n>>          return ondisk_size + entries * per_entry;\n>>   }\n>>\n>> +struct index_entry_offset\n>> +{\n>> +       /* starting byte offset into index file, count of index entries in this block */\n>> +       int offset, nr;\n> \n> uint32_t?\n> \n>> +};\n>> +\n>> +struct index_entry_offset_table\n>> +{\n>> +       int nr;\n>> +       struct index_entry_offset entries[0];\n> \n> Use FLEX_ARRAY. Some compilers are not happy with an array of zero\n> items if I remember correctly.\n> \n\nThanks for the warning, I'll update that.\n\n>> @@ -2523,6 +2551,9 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>>          struct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n>>          int drop_cache_tree = istate->drop_cache_tree;\n>>          off_t offset;\n>> +       int ieot_work = 1;\n>> +       struct index_entry_offset_table *ieot = NULL;\n>> +       int nr;\n> \n> There are a bunch of stuff going on in this function, maybe rename\n> this to nr_threads or nr_blocks to be less generic.\n> \n\nI can add a nr_threads variable to make this more obvious.\n\n>>\n>>          for (i = removed = extended = 0; i < entries; i++) {\n>>                  if (cache[i]->ce_flags & CE_REMOVE)\n>> @@ -2556,7 +2587,38 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>>          if (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n>>                  return -1;\n>>\n>> +#ifndef NO_PTHREADS\n>> +       if ((nr = git_config_get_index_threads()) != 1) {\n> \n> Maybe keep this assignment out of \"if\".\n> \n>> +               int ieot_blocks, cpus;\n>> +\n>> +               /*\n>> +                * ensure default number of ieot blocks maps evenly to the\n>> +                * default number of threads that will process them\n>> +                */\n>> +               if (!nr) {\n>> +                       ieot_blocks = istate->cache_nr / THREAD_COST;\n>> +                       cpus = online_cpus();\n>> +                       if (ieot_blocks > cpus - 1)\n>> +                               ieot_blocks = cpus - 1;\n> \n> The \" - 1\" here is for extension thread, yes? Probably worth a comment.\n> \n>> +               } else {\n>> +                       ieot_blocks = nr;\n>> +               }\n>> +\n>> +               /*\n>> +                * no reason to write out the IEOT extension if we don't\n>> +                * have enough blocks to utilize multi-threading\n>> +                */\n>> +               if (ieot_blocks > 1) {\n>> +                       ieot = xcalloc(1, sizeof(struct index_entry_offset_table)\n>> +                               + (ieot_blocks * sizeof(struct index_entry_offset)));\n> \n> Use FLEX_ALLOC_MEM() after you declare ..._table with FLEX_ARRAY.\n> \n\nFLEX_ALLOC_MEM() is focused on variable length \"char\" data.  All uses of \nFLEX_ARRAY with non char data did the allocation themselves to avoid the \nunnecessary memcpy() that comes with FLEX_ALLOC_MEM.\n\n> This ieot needs to be freed also and should be before any \"return -1\"\n> in this function.\n> \n\nGood catch. Will do.\n\n>> +                       ieot->nr = 0;\n>> +                       ieot_work = DIV_ROUND_UP(entries, ieot_blocks);\n> \n> Perhaps a better name for ioet_work? This looks like the number of\n> cache entries per block.\n> \n>> +               }\n>> +       }\n>> +#endif\n>> +\n>>          offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n>> +       nr = 0;\n> \n> Eh.. repurpose nr to count cache entries now? It's kinda hard to follow.\n> \n>>          previous_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n>>\n>>          for (i = 0; i < entries; i++) {\n>> @@ -2578,11 +2640,31 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>>\n>>                          drop_cache_tree = 1;\n>>                  }\n>> +               if (ieot && i && (i % ieot_work == 0)) {\n>> +                       ieot->entries[ieot->nr].nr = nr;\n>> +                       ieot->entries[ieot->nr].offset = offset;\n>> +                       ieot->nr++;\n>> +                       /*\n>> +                        * If we have a V4 index, set the first byte to an invalid\n>> +                        * character to ensure there is nothing common with the previous\n>> +                        * entry\n>> +                        */\n>> +                       if (previous_name)\n>> +                               previous_name->buf[0] = 0;\n>> +                       nr = 0;\n>> +                       offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n> \n> This only works correctly if the ce_write_entry() from the last\n> iteration has flushed everything to out to newfd. Maybe it does, but\n> it's error prone to rely on that in my opinion. Maybe we need an\n> explicit ce_write_flush() here to make sure.\n> \n\nThis logic already takes any unflushed data into account - the offset is \nwhat has been flushed to disk (lseek) plus the amount still in the \nbuffer (write_buffer_len) waiting to be flushed.  I don't see any need \nto force an additional flush and adding one could have a negative impact \non performance.\n\n>> +               }\n>>                  if (ce_write_entry(&c, newfd, ce, previous_name, (struct ondisk_cache_entry *)&ondisk) < 0)\n>>                          err = -1;\n>>\n>>                  if (err)\n>>                          break;\n>> +               nr++;\n>> +       }\n>> +       if (ieot && nr) {\n>> +               ieot->entries[ieot->nr].nr = nr;\n>> +               ieot->entries[ieot->nr].offset = offset;\n>> +               ieot->nr++;\n>>          }\n>>          strbuf_release(&previous_name_buf);\n>>\n>> @@ -2593,6 +2675,26 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>>          offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n>>          the_hash_algo->init_fn(&eoie_c);\n>>\n>> +       /*\n>> +        * Lets write out CACHE_EXT_INDEXENTRYOFFSETTABLE first so that we\n>> +        * can minimze the number of extensions we have to scan through to\n> \n> s/minimze/minimize/\n> \n>> +        * find it during load.  Write it out regardless of the\n>> +        * strip_extensions parameter as we need it when loading the shared\n>> +        * index.\n>> +        */\n>> +#ifndef NO_PTHREADS\n>> +       if (ieot) {\n>> +               struct strbuf sb = STRBUF_INIT;\n>> +\n>> +               write_ieot_extension(&sb, ieot);\n>> +               err = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_INDEXENTRYOFFSETTABLE, sb.len) < 0\n>> +                       || ce_write(&c, newfd, sb.buf, sb.len) < 0;\n>> +               strbuf_release(&sb);\n>> +               if (err)\n>> +                       return -1;\n>> +       }\n>> +#endif\n>> +\n>>          if (!strip_extensions && istate->split_index) {\n>>                  struct strbuf sb = STRBUF_INIT;\n>>\n>> @@ -3176,3 +3278,74 @@ static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context,\n>>          the_hash_algo->final_fn(hash, eoie_context);\n>>          strbuf_add(sb, hash, the_hash_algo->rawsz);\n>>   }\n>> +\n>> +#ifndef NO_PTHREADS\n>> +#define IEOT_VERSION   (1)\n>> +\n>> +static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset)\n>> +{\n>> +       const char *index = NULL;\n>> +       uint32_t extsize, ext_version;\n>> +       struct index_entry_offset_table *ieot;\n>> +       int i, nr;\n>> +\n>> +       /* find the IEOT extension */\n>> +       if (!offset)\n>> +              return NULL;\n>> +       while (offset <= mmap_size - the_hash_algo->rawsz - 8) {\n>> +              extsize = get_be32(mmap + offset + 4);\n>> +              if (CACHE_EXT((mmap + offset)) == CACHE_EXT_INDEXENTRYOFFSETTABLE) {\n>> +                      index = mmap + offset + 4 + 4;\n>> +                      break;\n>> +              }\n>> +              offset += 8;\n>> +              offset += extsize;\n>> +       }\n> \n> Maybe refactor this loop. I think I've seen this in at least two\n> places now. Probably three?\n> \n>> +       if (!index)\n>> +              return NULL;\n>> +\n>> +       /* validate the version is IEOT_VERSION */\n>> +       ext_version = get_be32(index);\n>> +       if (ext_version != IEOT_VERSION)\n>> +              return NULL;\n> \n> Report the error (e.g. \"unsupported version\" or something)\n> \n\nSure.  I'll add reporting here and in the error check below.\n\n>> +       index += sizeof(uint32_t);\n>> +\n>> +       /* extension size - version bytes / bytes per entry */\n>> +       nr = (extsize - sizeof(uint32_t)) / (sizeof(uint32_t) + sizeof(uint32_t));\n> \n> Do we need to check if \"(extsize - version) % sizeof(entry) == 0\"?\n> \n>> +       if (!nr)\n>> +              return NULL;\n>> +       ieot = xmalloc(sizeof(struct index_entry_offset_table)\n>> +              + (nr * sizeof(struct index_entry_offset)));\n>> +       ieot->nr = nr;\n>> +       for (i = 0; i < nr; i++) {\n>> +              ieot->entries[i].offset = get_be32(index);\n>> +              index += sizeof(uint32_t);\n>> +              ieot->entries[i].nr = get_be32(index);\n>> +              index += sizeof(uint32_t);\n>> +       }\n>> +\n>> +       return ieot;\n>> +}\n>> +\n>> +static void write_ieot_extension(struct strbuf *sb, struct index_entry_offset_table *ieot)\n>> +{\n>> +       uint32_t buffer;\n>> +       int i;\n>> +\n>> +       /* version */\n>> +       put_be32(&buffer, IEOT_VERSION);\n>> +       strbuf_add(sb, &buffer, sizeof(uint32_t));\n>> +\n>> +       /* ieot */\n>> +       for (i = 0; i < ieot->nr; i++) {\n>> +\n>> +              /* offset */\n>> +              put_be32(&buffer, ieot->entries[i].offset);\n>> +              strbuf_add(sb, &buffer, sizeof(uint32_t));\n>> +\n>> +              /* count */\n>> +              put_be32(&buffer, ieot->entries[i].nr);\n>> +              strbuf_add(sb, &buffer, sizeof(uint32_t));\n>> +       }\n>> +}\n>> +#endif\n>> --\n>> 2.18.0.windows.1\n>>\n> \n> \n"},{"id":"359435","messageId":"CACsJy8BEk9q0DNzXLUo0BgN3DUYYk=vSQcLqtQY=-JLAn74RXg@mail.gmail.com","threadId":"49204","inReplyTo":"351b9746-6c2e-a658-3f51-71c1f4cbc3ac@gmail.com","subject":"Re: [PATCH v7 6/7] ieot: add Index Entry Offset Table (IEOT) extension","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-10-02T17:02:09Z","receivedAt":"2018-10-02T17:02:37Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Oct 2, 2018 at 6:34 PM Ben Peart <peartben@gmail.com> wrote:\n> >> +                       offset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n> >\n> > This only works correctly if the ce_write_entry() from the last\n> > iteration has flushed everything to out to newfd. Maybe it does, but\n> > it's error prone to rely on that in my opinion. Maybe we need an\n> > explicit ce_write_flush() here to make sure.\n> >\n>\n> This logic already takes any unflushed data into account - the offset is\n> what has been flushed to disk (lseek) plus the amount still in the\n> buffer (write_buffer_len) waiting to be flushed.  I don't see any need\n> to force an additional flush and adding one could have a negative impact\n> on performance.\n\nEck! How did I miss that write_buffer_len :P\n-- \nDuy\n"},{"id":"359449","messageId":"8f00ea21-07dc-4d19-7bd8-16dee53eba67@gmail.com","threadId":"49204","inReplyTo":"CACsJy8D7Pbg6xMZBfCiz_7_=reY3Os4R_70wc65VMxbu2=Kqjw@mail.gmail.com","subject":"Re: [PATCH v7 7/7] read-cache: load cache entries on worker threads","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-02T19:09:24Z","receivedAt":"2018-10-02T19:09:30Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 10/1/2018 1:09 PM, Duy Nguyen wrote:\n> On Mon, Oct 1, 2018 at 3:46 PM Ben Peart <peartben@gmail.com> wrote:\n>> +/*\n>> + * A helper function that will load the specified range of cache entries\n>> + * from the memory mapped file and add them to the given index.\n>> + */\n>> +static unsigned long load_cache_entry_block(struct index_state *istate,\n>> +                       struct mem_pool *ce_mem_pool, int offset, int nr, const char *mmap,\n> \n> Please use unsigned long for offset (here and in the thread_data\n> struct). We should use off_t instead, but that's out of scope. At\n> least keep offset type consistent in here.\n> \n\nUnfortunately, this code is littered with different types for size and \noffset.  \"int\" is the most common but there are also off_t, size_t and \nsome unsigned long as well.  Currently all of them are at least 32 bits \nso until we need to have an index larger than 32 bits, we should be OK. \nI agree, fixing them all is outside the scope of this patch.\n\n>> +                       unsigned long start_offset, const struct cache_entry *previous_ce)\n> \n> I don't think you want to pass previous_ce in. You always pass NULL\n> anyway. And if this function is about loading a block (i.e. at block\n> boundary) then initial previous_ce _must_ be NULL or things break\n> horribly.\n> \n\nThe function as written can load any arbitrary subset of cache entries \nas long as previous_ce is set correctly.  I currently only use it on \nblock boundaries but I don't see any good reason to limit its \ncapabilities by moving what code passes the NULL in one function deeper.\n\n>> @@ -1959,20 +2007,125 @@ static void *load_index_extensions(void *_data)\n>>\n>>   #define THREAD_COST            (10000)\n>>\n>> +struct load_cache_entries_thread_data\n>> +{\n>> +       pthread_t pthread;\n>> +       struct index_state *istate;\n>> +       struct mem_pool *ce_mem_pool;\n>> +       int offset;\n>> +       const char *mmap;\n>> +       struct index_entry_offset_table *ieot;\n>> +       int ieot_offset;        /* starting index into the ieot array */\n> \n> If it's an index, maybe just name it ieot_index and we can get rid of\n> the comment.\n> \n>> +       int ieot_work;          /* count of ieot entries to process */\n> \n> Maybe instead of saving the whole \"ieot\" table here. Add\n> \n>       struct index_entry_offset *blocks;\n> \n> which points to the starting block for this thread and rename that\n> mysterious (to me) ieot_work to nr_blocks. The thread will have access\n> from blocks[0] to blocks[nr_blocks - 1]\n> \n\nMeh. Either way you have to figure out there are a block of entries and \neach thread is going to process some subset of those entries.  You can \ndo the base + offset math here or down in the calling function but it \nhas to happen (and be understood) either way.\n\nI'll rename ieot_offset to ieot_start and ieot_work to ieot_blocks which \nshould hopefully help make it more obvious what they do.\n\n>> +       unsigned long consumed; /* return # of bytes in index file processed */\n>> +};\n>> +\n>> +/*\n>> + * A thread proc to run the load_cache_entries() computation\n>> + * across multiple background threads.\n>> + */\n>> +static void *load_cache_entries_thread(void *_data)\n>> +{\n>> +       struct load_cache_entries_thread_data *p = _data;\n>> +       int i;\n>> +\n>> +       /* iterate across all ieot blocks assigned to this thread */\n>> +       for (i = p->ieot_offset; i < p->ieot_offset + p->ieot_work; i++) {\n>> +               p->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool, p->offset, p->ieot->entries[i].nr, p->mmap, p->ieot->entries[i].offset, NULL);\n> \n> Please wrap this long line.\n> \n>> +               p->offset += p->ieot->entries[i].nr;\n>> +       }\n>> +       return NULL;\n>> +}\n>> +\n>> +static unsigned long load_cache_entries_threaded(struct index_state *istate, const char *mmap, size_t mmap_size,\n>> +                       unsigned long src_offset, int nr_threads, struct index_entry_offset_table *ieot)\n>> +{\n>> +       int i, offset, ieot_work, ieot_offset, err;\n>> +       struct load_cache_entries_thread_data *data;\n>> +       unsigned long consumed = 0;\n>> +       int nr;\n>> +\n>> +       /* a little sanity checking */\n>> +       if (istate->name_hash_initialized)\n>> +               BUG(\"the name hash isn't thread safe\");\n>> +\n>> +       mem_pool_init(&istate->ce_mem_pool, 0);\n>> +       data = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n> \n> we normally use sizeof(*data) instead of sizeof(struct ...)\n> \n>> +\n>> +       /* ensure we have no more threads than we have blocks to process */\n>> +       if (nr_threads > ieot->nr)\n>> +               nr_threads = ieot->nr;\n>> +       data = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n> \n> eh.. reallocate the same \"data\"?\n> \n\nThanks, good catch - I hate leaky code.\n\n>> +\n>> +       offset = ieot_offset = 0;\n>> +       ieot_work = DIV_ROUND_UP(ieot->nr, nr_threads);\n>> +       for (i = 0; i < nr_threads; i++) {\n>> +               struct load_cache_entries_thread_data *p = &data[i];\n>> +               int j;\n>> +\n>> +               if (ieot_offset + ieot_work > ieot->nr)\n>> +                       ieot_work = ieot->nr - ieot_offset;\n>> +\n>> +               p->istate = istate;\n>> +               p->offset = offset;\n>> +               p->mmap = mmap;\n>> +               p->ieot = ieot;\n>> +               p->ieot_offset = ieot_offset;\n>> +               p->ieot_work = ieot_work;\n>> +\n>> +               /* create a mem_pool for each thread */\n>> +               nr = 0;\n> \n> Since nr is only used in this for loop. Declare it in this scope\n> instead of declaring it for the whole function.\n> \n>> +               for (j = p->ieot_offset; j < p->ieot_offset + p->ieot_work; j++)\n>> +                       nr += p->ieot->entries[j].nr;\n>> +               if (istate->version == 4) {\n>> +                       mem_pool_init(&p->ce_mem_pool,\n>> +                               estimate_cache_size_from_compressed(nr));\n>> +               }\n>> +               else {\n>> +                       mem_pool_init(&p->ce_mem_pool,\n>> +                               estimate_cache_size(mmap_size, nr));\n>> +               }\n> \n> Maybe keep this mem_pool_init code inside load_cache_entries_thread(),\n> similar to how you do it for load_cache_entries_thread(). It's mostly\n> to keep this loop shorter to see (and understand), of course\n> parallelizing this mem_pool_init() is just noise.\n> \n\nI understand the desire to get that part of the thread initialization \nout of the main line of this function (it's a bit messy between the \nentry counting and version differences) but I prefer to have all the \nthread initialization completed before creating the thread.  That allows \nfor simpler error handling and helps minimize the state you have to pass \ninto the thread (mmap_size in this case).\n\n>> +\n>> +               err = pthread_create(&p->pthread, NULL, load_cache_entries_thread, p);\n>> +               if (err)\n>> +                       die(_(\"unable to create load_cache_entries thread: %s\"), strerror(err));\n>> +\n>> +               /* increment by the number of cache entries in the ieot block being processed */\n>> +               for (j = 0; j < ieot_work; j++)\n>> +                       offset += ieot->entries[ieot_offset + j].nr;\n> \n> I wonder if it makes things simpler if you store cache_entry _index_\n> in entrie[] array instead of storing the number of entries. You can\n> easily calculate nr then by doing entries[i].index -\n> entries[i-1].index. And you can count multiple blocks the same way,\n> without looping like this.\n> \n>> +               ieot_offset += ieot_work;\n>> +       }\n>> +\n>> +       for (i = 0; i < nr_threads; i++) {\n>> +               struct load_cache_entries_thread_data *p = &data[i];\n>> +\n>> +               err = pthread_join(p->pthread, NULL);\n>> +               if (err)\n>> +                       die(_(\"unable to join load_cache_entries thread: %s\"), strerror(err));\n>> +               mem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n>> +               consumed += p->consumed;\n>> +       }\n>> +\n>> +       free(data);\n>> +\n>> +       return consumed;\n>> +}\n>> +#endif\n>> +\n>>   /* remember to discard_cache() before reading a different cache! */\n>>   int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>>   {\n>> -       int fd, i;\n>> +       int fd;\n>>          struct stat st;\n>>          unsigned long src_offset;\n>>          const struct cache_header *hdr;\n>>          const char *mmap;\n>>          size_t mmap_size;\n>> -       const struct cache_entry *previous_ce = NULL;\n>>          struct load_index_extensions p;\n>>          size_t extension_offset = 0;\n>>   #ifndef NO_PTHREADS\n>> -       int nr_threads;\n>> +       int nr_threads, cpus;\n>> +       struct index_entry_offset_table *ieot = NULL;\n>>   #endif\n>>\n>>          if (istate->initialized)\n>> @@ -2014,10 +2167,18 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>>          p.mmap = mmap;\n>>          p.mmap_size = mmap_size;\n>>\n>> +       src_offset = sizeof(*hdr);\n> \n> OK we've been doing this since forever, sizeof(struct cache_header)\n> probably does not have extra padding on any supported platform.\n> \n>> @@ -2032,29 +2193,22 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n>>                          nr_threads--;\n>>                  }\n>>          }\n>> -#endif\n>> -\n>> -       if (istate->version == 4) {\n>> -               mem_pool_init(&istate->ce_mem_pool,\n>> -                             estimate_cache_size_from_compressed(istate->cache_nr));\n>> -       } else {\n>> -               mem_pool_init(&istate->ce_mem_pool,\n>> -                             estimate_cache_size(mmap_size, istate->cache_nr));\n>> -       }\n>>\n>> -       src_offset = sizeof(*hdr);\n>> -       for (i = 0; i < istate->cache_nr; i++) {\n>> -               struct ondisk_cache_entry *disk_ce;\n>> -               struct cache_entry *ce;\n>> -               unsigned long consumed;\n>> +       /*\n>> +        * Locate and read the index entry offset table so that we can use it\n>> +        * to multi-thread the reading of the cache entries.\n>> +        */\n>> +       if (extension_offset && nr_threads > 1)\n>> +               ieot = read_ieot_extension(mmap, mmap_size, extension_offset);\n> \n> You need to free ieot at some point.\n> \n\nGood catch - I hate leaky code.\n\n>>\n>> -               disk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n>> -               ce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n>> -               set_index_entry(istate, i, ce);\n>> +       if (ieot)\n>> +               src_offset += load_cache_entries_threaded(istate, mmap, mmap_size, src_offset, nr_threads, ieot);\n>> +       else\n>> +               src_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n>> +#else\n>> +       src_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n>> +#endif\n>>\n>> -               src_offset += consumed;\n>> -               previous_ce = ce;\n>> -       }\n>>          istate->timestamp.sec = st.st_mtime;\n>>          istate->timestamp.nsec = ST_MTIME_NSEC(st);\n>>\n>> --\n>> 2.18.0.windows.1\n>>\n"},{"id":"360038","messageId":"20181010155938.20996-1-peartben@gmail.com","threadId":"49204","inReplyTo":"20180823154053.20212-1-benpeart@microsoft.com","subject":"[PATCH v8 0/7] speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-10T15:59:31Z","receivedAt":"2018-10-10T15:59:53Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nFixed issues identified in review the most impactful probably being plugging\nsome leaks and improved error handling.  Also added better error messages\nand some code cleanup to code I'd touched.\n\nThe biggest change in the interdiff is the impact of renaming ieot_offset to\nieot_start and ieot_work to ieot_blocks in hopes of making it easier to read\nand understand the code.\n\nBase Ref: master\nWeb-Diff: https://github.com/benpeart/git/commit/6caa0bac46\nCheckout: git fetch https://github.com/benpeart/git read-index-multithread-v8 && git checkout 6caa0bac46\n\n\n### Interdiff (v7..v8):\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 14402a0738..7acc2c86f4 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1901,7 +1901,7 @@ struct index_entry_offset\n struct index_entry_offset_table\n {\n \tint nr;\n-\tstruct index_entry_offset entries[0];\n+\tstruct index_entry_offset entries[FLEX_ARRAY];\n };\n \n #ifndef NO_PTHREADS\n@@ -1935,9 +1935,7 @@ static void *load_index_extensions(void *_data)\n \t\t * extension name (4-byte) and section length\n \t\t * in 4-byte network byte order.\n \t\t */\n-\t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, p->mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n+\t\tuint32_t extsize = get_be32(p->mmap + src_offset + 4);\n \t\tif (read_index_extension(p->istate,\n \t\t\t\t\t p->mmap + src_offset,\n \t\t\t\t\t p->mmap + src_offset + 8,\n@@ -2015,8 +2013,8 @@ struct load_cache_entries_thread_data\n \tint offset;\n \tconst char *mmap;\n \tstruct index_entry_offset_table *ieot;\n-\tint ieot_offset;        /* starting index into the ieot array */\n-\tint ieot_work;          /* count of ieot entries to process */\n+\tint ieot_start;\t\t/* starting index into the ieot array */\n+\tint ieot_blocks;\t/* count of ieot entries to process */\n \tunsigned long consumed;\t/* return # of bytes in index file processed */\n };\n \n@@ -2030,8 +2028,9 @@ static void *load_cache_entries_thread(void *_data)\n \tint i;\n \n \t/* iterate across all ieot blocks assigned to this thread */\n-\tfor (i = p->ieot_offset; i < p->ieot_offset + p->ieot_work; i++) {\n-\t\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool, p->offset, p->ieot->entries[i].nr, p->mmap, p->ieot->entries[i].offset, NULL);\n+\tfor (i = p->ieot_start; i < p->ieot_start + p->ieot_blocks; i++) {\n+\t\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n+\t\t\tp->offset, p->ieot->entries[i].nr, p->mmap, p->ieot->entries[i].offset, NULL);\n \t\tp->offset += p->ieot->entries[i].nr;\n \t}\n \treturn NULL;\n@@ -2040,48 +2039,45 @@ static void *load_cache_entries_thread(void *_data)\n static unsigned long load_cache_entries_threaded(struct index_state *istate, const char *mmap, size_t mmap_size,\n \t\t\tunsigned long src_offset, int nr_threads, struct index_entry_offset_table *ieot)\n {\n-\tint i, offset, ieot_work, ieot_offset, err;\n+\tint i, offset, ieot_blocks, ieot_start, err;\n \tstruct load_cache_entries_thread_data *data;\n \tunsigned long consumed = 0;\n-\tint nr;\n \n \t/* a little sanity checking */\n \tif (istate->name_hash_initialized)\n \t\tBUG(\"the name hash isn't thread safe\");\n \n \tmem_pool_init(&istate->ce_mem_pool, 0);\n-\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n \n \t/* ensure we have no more threads than we have blocks to process */\n \tif (nr_threads > ieot->nr)\n \t\tnr_threads = ieot->nr;\n-\tdata = xcalloc(nr_threads, sizeof(struct load_cache_entries_thread_data));\n+\tdata = xcalloc(nr_threads, sizeof(*data));\n \n-\toffset = ieot_offset = 0;\n-\tieot_work = DIV_ROUND_UP(ieot->nr, nr_threads);\n+\toffset = ieot_start = 0;\n+\tieot_blocks = DIV_ROUND_UP(ieot->nr, nr_threads);\n \tfor (i = 0; i < nr_threads; i++) {\n \t\tstruct load_cache_entries_thread_data *p = &data[i];\n-\t\tint j;\n+\t\tint nr, j;\n \n-\t\tif (ieot_offset + ieot_work > ieot->nr)\n-\t\t\tieot_work = ieot->nr - ieot_offset;\n+\t\tif (ieot_start + ieot_blocks > ieot->nr)\n+\t\t\tieot_blocks = ieot->nr - ieot_start;\n \n \t\tp->istate = istate;\n \t\tp->offset = offset;\n \t\tp->mmap = mmap;\n \t\tp->ieot = ieot;\n-\t\tp->ieot_offset = ieot_offset;\n-\t\tp->ieot_work = ieot_work;\n+\t\tp->ieot_start = ieot_start;\n+\t\tp->ieot_blocks = ieot_blocks;\n \n \t\t/* create a mem_pool for each thread */\n \t\tnr = 0;\n-\t\tfor (j = p->ieot_offset; j < p->ieot_offset + p->ieot_work; j++)\n+\t\tfor (j = p->ieot_start; j < p->ieot_start + p->ieot_blocks; j++)\n \t\t\tnr += p->ieot->entries[j].nr;\n \t\tif (istate->version == 4) {\n \t\t\tmem_pool_init(&p->ce_mem_pool,\n \t\t\t\testimate_cache_size_from_compressed(nr));\n-\t\t}\n-\t\telse {\n+\t\t} else {\n \t\t\tmem_pool_init(&p->ce_mem_pool,\n \t\t\t\testimate_cache_size(mmap_size, nr));\n \t\t}\n@@ -2091,9 +2087,9 @@ static unsigned long load_cache_entries_threaded(struct index_state *istate, con\n \t\t\tdie(_(\"unable to create load_cache_entries thread: %s\"), strerror(err));\n \n \t\t/* increment by the number of cache entries in the ieot block being processed */\n-\t\tfor (j = 0; j < ieot_work; j++)\n-\t\t\toffset += ieot->entries[ieot_offset + j].nr;\n-\t\tieot_offset += ieot_work;\n+\t\tfor (j = 0; j < ieot_blocks; j++)\n+\t\t\toffset += ieot->entries[ieot_start + j].nr;\n+\t\tieot_start += ieot_blocks;\n \t}\n \n \tfor (i = 0; i < nr_threads; i++) {\n@@ -2201,10 +2197,12 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tif (extension_offset && nr_threads > 1)\n \t\tieot = read_ieot_extension(mmap, mmap_size, extension_offset);\n \n-\tif (ieot)\n+\tif (ieot) {\n \t\tsrc_offset += load_cache_entries_threaded(istate, mmap, mmap_size, src_offset, nr_threads, ieot);\n-\telse\n+\t\tfree(ieot);\n+\t} else {\n \t\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+\t}\n #else\n \tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n #endif\n@@ -2705,9 +2703,9 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n-\tint ieot_work = 1;\n+\tint ieot_entries = 1;\n \tstruct index_entry_offset_table *ieot = NULL;\n-\tint nr;\n+\tint nr, nr_threads;\n \n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n@@ -2742,20 +2740,24 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\treturn -1;\n \n #ifndef NO_PTHREADS\n-\tif ((nr = git_config_get_index_threads()) != 1) {\n+\tnr_threads = git_config_get_index_threads();\n+\tif (nr_threads != 1) {\n \t\tint ieot_blocks, cpus;\n \n \t\t/*\n \t\t * ensure default number of ieot blocks maps evenly to the\n-\t\t * default number of threads that will process them\n+\t\t * default number of threads that will process them leaving\n+\t\t * room for the thread to load the index extensions.\n \t\t */\n-\t\tif (!nr) {\n+\t\tif (!nr_threads) {\n \t\t\tieot_blocks = istate->cache_nr / THREAD_COST;\n \t\t\tcpus = online_cpus();\n \t\t\tif (ieot_blocks > cpus - 1)\n \t\t\t\tieot_blocks = cpus - 1;\n \t\t} else {\n-\t\t\tieot_blocks = nr;\n+\t\t\tieot_blocks = nr_threads;\n+\t\t\tif (ieot_blocks > istate->cache_nr)\n+\t\t\t\tieot_blocks = istate->cache_nr;\n \t\t}\n \n \t\t/*\n@@ -2765,13 +2767,17 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tif (ieot_blocks > 1) {\n \t\t\tieot = xcalloc(1, sizeof(struct index_entry_offset_table)\n \t\t\t\t+ (ieot_blocks * sizeof(struct index_entry_offset)));\n-\t\t\tieot->nr = 0;\n-\t\t\tieot_work = DIV_ROUND_UP(entries, ieot_blocks);\n+\t\t\tieot_entries = DIV_ROUND_UP(entries, ieot_blocks);\n \t\t}\n \t}\n #endif\n \n-\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\toffset = lseek(newfd, 0, SEEK_CUR);\n+\tif (offset < 0) {\n+\t\tfree(ieot);\n+\t\treturn -1;\n+\t}\n+\toffset += write_buffer_len;\n \tnr = 0;\n \tprevious_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n \n@@ -2794,7 +2800,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \n \t\t\tdrop_cache_tree = 1;\n \t\t}\n-\t\tif (ieot && i && (i % ieot_work == 0)) {\n+\t\tif (ieot && i && (i % ieot_entries == 0)) {\n \t\t\tieot->entries[ieot->nr].nr = nr;\n \t\t\tieot->entries[ieot->nr].offset = offset;\n \t\t\tieot->nr++;\n@@ -2806,7 +2812,12 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\tif (previous_name)\n \t\t\t\tprevious_name->buf[0] = 0;\n \t\t\tnr = 0;\n-\t\t\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\t\t\toffset = lseek(newfd, 0, SEEK_CUR);\n+\t\t\tif (offset < 0) {\n+\t\t\t\tfree(ieot);\n+\t\t\t\treturn -1;\n+\t\t\t}\n+\t\t\toffset += write_buffer_len;\n \t\t}\n \t\tif (ce_write_entry(&c, newfd, ce, previous_name, (struct ondisk_cache_entry *)&ondisk) < 0)\n \t\t\terr = -1;\n@@ -2822,16 +2833,23 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t}\n \tstrbuf_release(&previous_name_buf);\n \n-\tif (err)\n+\tif (err) {\n+\t\tfree(ieot);\n \t\treturn err;\n+\t}\n \n \t/* Write extension data here */\n-\toffset = lseek(newfd, 0, SEEK_CUR) + write_buffer_len;\n+\toffset = lseek(newfd, 0, SEEK_CUR);\n+\tif (offset < 0) {\n+\t\tfree(ieot);\n+\t\treturn -1;\n+\t}\n+\toffset += write_buffer_len;\n \tthe_hash_algo->init_fn(&eoie_c);\n \n \t/*\n \t * Lets write out CACHE_EXT_INDEXENTRYOFFSETTABLE first so that we\n-\t * can minimze the number of extensions we have to scan through to\n+\t * can minimize the number of extensions we have to scan through to\n \t * find it during load.  Write it out regardless of the\n \t * strip_extensions parameter as we need it when loading the shared\n \t * index.\n@@ -2844,6 +2862,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_INDEXENTRYOFFSETTABLE, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n+\t\tfree(ieot);\n \t\tif (err)\n \t\t\treturn -1;\n \t}\n@@ -3460,14 +3479,18 @@ static struct index_entry_offset_table *read_ieot_extension(const char *mmap, si\n \n        /* validate the version is IEOT_VERSION */\n        ext_version = get_be32(index);\n-       if (ext_version != IEOT_VERSION)\n+       if (ext_version != IEOT_VERSION) {\n+\t       error(\"invalid IEOT version %d\", ext_version);\n \t       return NULL;\n+       }\n        index += sizeof(uint32_t);\n \n        /* extension size - version bytes / bytes per entry */\n        nr = (extsize - sizeof(uint32_t)) / (sizeof(uint32_t) + sizeof(uint32_t));\n-       if (!nr)\n+       if (!nr) {\n+\t       error(\"invalid number of IEOT entries %d\", nr);\n \t       return NULL;\n+       }\n        ieot = xmalloc(sizeof(struct index_entry_offset_table)\n \t       + (nr * sizeof(struct index_entry_offset)));\n        ieot->nr = nr;\n\n\n### Patches\n\nBen Peart (6):\n  read-cache: clean up casting and byte decoding\n  eoie: add End of Index Entry (EOIE) extension\n  config: add new index.threads config setting\n  read-cache: load cache extensions on a worker thread\n  ieot: add Index Entry Offset Table (IEOT) extension\n  read-cache: load cache entries on worker threads\n\nNguyễn Thái Ngọc Duy (1):\n  read-cache.c: optimize reading index format v4\n\n Documentation/config.txt                 |   7 +\n Documentation/technical/index-format.txt |  41 ++\n config.c                                 |  18 +\n config.h                                 |   1 +\n read-cache.c                             | 774 +++++++++++++++++++----\n t/README                                 |   5 +\n t/t1700-split-index.sh                   |  13 +-\n 7 files changed, 739 insertions(+), 120 deletions(-)\n\n\nbase-commit: fe8321ec057f9231c26c29b364721568e58040f7\n-- \n2.18.0.windows.1\n\n\n"},{"id":"360039","messageId":"20181010155938.20996-2-peartben@gmail.com","threadId":"49204","inReplyTo":"20181010155938.20996-1-peartben@gmail.com","subject":"[PATCH v8 1/7] read-cache.c: optimize reading index format v4","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-10T15:59:32Z","receivedAt":"2018-10-10T15:59:53Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n\nIndex format v4 requires some more computation to assemble a path\nbased on a previous one. The current code is not very efficient\nbecause\n\n - it doubles memory copy, we assemble the final path in a temporary\n   first before putting it back to a cache_entry\n\n - strbuf_remove() in expand_name_field() is not exactly a good fit\n   for stripping a part at the end, _setlen() would do the same job\n   and is much cheaper.\n\n - the open-coded loop to find the end of the string in\n   expand_name_field() can't beat an optimized strlen()\n\nThis patch avoids the temporary buffer and writes directly to the new\ncache_entry, which addresses the first two points. The last point\ncould also be avoided if the total string length fits in the first 12\nbits of ce_flags, if not we fall back to strlen().\n\nRunning \"test-tool read-cache 100\" on webkit.git (275k files), reading\nv2 only takes 4.226 seconds, while v4 takes 5.711 seconds, 35% more\ntime. The patch reduces read time on v4 to 4.319 seconds.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n read-cache.c | 128 ++++++++++++++++++++++++---------------------------\n 1 file changed, 60 insertions(+), 68 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 8d04d78a58..583a4fb1f8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1713,63 +1713,24 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *cache_entry_from_ondisk(struct mem_pool *mem_pool,\n-\t\t\t\t\t\t   struct ondisk_cache_entry *ondisk,\n-\t\t\t\t\t\t   unsigned int flags,\n-\t\t\t\t\t\t   const char *name,\n-\t\t\t\t\t\t   size_t len)\n-{\n-\tstruct cache_entry *ce = mem_pool__ce_alloc(mem_pool, len);\n-\n-\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n-\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n-\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n-\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n-\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n-\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n-\tce->ce_mode  = get_be32(&ondisk->mode);\n-\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n-\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n-\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n-\tce->ce_flags = flags & ~CE_NAMEMASK;\n-\tce->ce_namelen = len;\n-\tce->index = 0;\n-\thashcpy(ce->oid.hash, ondisk->sha1);\n-\tmemcpy(ce->name, name, len);\n-\tce->name[len] = '\\0';\n-\treturn ce;\n-}\n-\n-/*\n- * Adjacent cache entries tend to share the leading paths, so it makes\n- * sense to only store the differences in later entries.  In the v4\n- * on-disk format of the index, each on-disk cache entry stores the\n- * number of bytes to be stripped from the end of the previous name,\n- * and the bytes to append to the result, to come up with its name.\n- */\n-static unsigned long expand_name_field(struct strbuf *name, const char *cp_)\n-{\n-\tconst unsigned char *ep, *cp = (const unsigned char *)cp_;\n-\tsize_t len = decode_varint(&cp);\n-\n-\tif (name->len < len)\n-\t\tdie(\"malformed name field in the index\");\n-\tstrbuf_remove(name, name->len - len, len);\n-\tfor (ep = cp; *ep; ep++)\n-\t\t; /* find the end */\n-\tstrbuf_add(name, cp, ep - cp);\n-\treturn (const char *)ep + 1 - cp_;\n-}\n-\n-static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n+static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n-\t\t\t\t\t    struct strbuf *previous_name)\n+\t\t\t\t\t    const struct cache_entry *previous_ce)\n {\n \tstruct cache_entry *ce;\n \tsize_t len;\n \tconst char *name;\n \tunsigned int flags;\n+\tsize_t copy_len;\n+\t/*\n+\t * Adjacent cache entries tend to share the leading paths, so it makes\n+\t * sense to only store the differences in later entries.  In the v4\n+\t * on-disk format of the index, each on-disk cache entry stores the\n+\t * number of bytes to be stripped from the end of the previous name,\n+\t * and the bytes to append to the result, to come up with its name.\n+\t */\n+\tint expand_name_field = istate->version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1789,21 +1750,54 @@ static struct cache_entry *create_from_disk(struct mem_pool *mem_pool,\n \telse\n \t\tname = ondisk->name;\n \n-\tif (!previous_name) {\n-\t\t/* v3 and earlier */\n-\t\tif (len == CE_NAMEMASK)\n-\t\t\tlen = strlen(name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags, name, len);\n+\tif (expand_name_field) {\n+\t\tconst unsigned char *cp = (const unsigned char *)name;\n+\t\tsize_t strip_len, previous_len;\n \n-\t\t*ent_size = ondisk_ce_size(ce);\n-\t} else {\n-\t\tunsigned long consumed;\n-\t\tconsumed = expand_name_field(previous_name, name);\n-\t\tce = cache_entry_from_ondisk(mem_pool, ondisk, flags,\n-\t\t\t\t\t     previous_name->buf,\n-\t\t\t\t\t     previous_name->len);\n+\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n+\t\tstrip_len = decode_varint(&cp);\n+\t\tif (previous_len < strip_len) {\n+\t\t\tif (previous_ce)\n+\t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n+\t\t\t\t    previous_ce->name);\n+\t\t\telse\n+\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t}\n+\t\tcopy_len = previous_len - strip_len;\n+\t\tname = (const char *)cp;\n+\t}\n+\n+\tif (len == CE_NAMEMASK) {\n+\t\tlen = strlen(name);\n+\t\tif (expand_name_field)\n+\t\t\tlen += copy_len;\n+\t}\n+\n+\tce = mem_pool__ce_alloc(istate->ce_mem_pool, len);\n+\n+\tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n+\tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n+\tce->ce_stat_data.sd_ctime.nsec = get_be32(&ondisk->ctime.nsec);\n+\tce->ce_stat_data.sd_mtime.nsec = get_be32(&ondisk->mtime.nsec);\n+\tce->ce_stat_data.sd_dev   = get_be32(&ondisk->dev);\n+\tce->ce_stat_data.sd_ino   = get_be32(&ondisk->ino);\n+\tce->ce_mode  = get_be32(&ondisk->mode);\n+\tce->ce_stat_data.sd_uid   = get_be32(&ondisk->uid);\n+\tce->ce_stat_data.sd_gid   = get_be32(&ondisk->gid);\n+\tce->ce_stat_data.sd_size  = get_be32(&ondisk->size);\n+\tce->ce_flags = flags & ~CE_NAMEMASK;\n+\tce->ce_namelen = len;\n+\tce->index = 0;\n+\thashcpy(ce->oid.hash, ondisk->sha1);\n \n-\t\t*ent_size = (name - ((char *)ondisk)) + consumed;\n+\tif (expand_name_field) {\n+\t\tif (copy_len)\n+\t\t\tmemcpy(ce->name, previous_ce->name, copy_len);\n+\t\tmemcpy(ce->name + copy_len, name, len + 1 - copy_len);\n+\t\t*ent_size = (name - ((char *)ondisk)) + len + 1 - copy_len;\n+\t} else {\n+\t\tmemcpy(ce->name, name, len + 1);\n+\t\t*ent_size = ondisk_ce_size(ce);\n \t}\n \treturn ce;\n }\n@@ -1898,7 +1892,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tstruct cache_header *hdr;\n \tvoid *mmap;\n \tsize_t mmap_size;\n-\tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n+\tconst struct cache_entry *previous_ce = NULL;\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1936,11 +1930,9 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->initialized = 1;\n \n \tif (istate->version == 4) {\n-\t\tprevious_name = &previous_name_buf;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n \t} else {\n-\t\tprevious_name = NULL;\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n \t}\n@@ -1952,12 +1944,12 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tunsigned long consumed;\n \n \t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n-\t\tce = create_from_disk(istate->ce_mem_pool, disk_ce, &consumed, previous_name);\n+\t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n \t\tsrc_offset += consumed;\n+\t\tprevious_ce = ce;\n \t}\n-\tstrbuf_release(&previous_name_buf);\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-- \n2.18.0.windows.1\n\n"},{"id":"360040","messageId":"20181010155938.20996-3-peartben@gmail.com","threadId":"49204","inReplyTo":"20181010155938.20996-1-peartben@gmail.com","subject":"[PATCH v8 2/7] read-cache: clean up casting and byte decoding","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-10T15:59:33Z","receivedAt":"2018-10-10T15:59:55Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nThis patch does a clean up pass to minimize the casting required to work\nwith the memory mapped index (mmap).\n\nIt also makes the decoding of network byte order more consistent by using\nget_be32() where possible.\n\nSigned-off-by: Ben Peart <benpeart@microsoft.com>\n---\n read-cache.c | 23 +++++++++++------------\n 1 file changed, 11 insertions(+), 12 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 583a4fb1f8..6ba99e2c96 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1650,7 +1650,7 @@ int verify_index_checksum;\n /* Allow fsck to force verification of the cache entry order. */\n int verify_ce_order;\n \n-static int verify_hdr(struct cache_header *hdr, unsigned long size)\n+static int verify_hdr(const struct cache_header *hdr, unsigned long size)\n {\n \tgit_hash_ctx c;\n \tunsigned char hash[GIT_MAX_RAWSZ];\n@@ -1674,7 +1674,7 @@ static int verify_hdr(struct cache_header *hdr, unsigned long size)\n }\n \n static int read_index_extension(struct index_state *istate,\n-\t\t\t\tconst char *ext, void *data, unsigned long sz)\n+\t\t\t\tconst char *ext, const char *data, unsigned long sz)\n {\n \tswitch (CACHE_EXT(ext)) {\n \tcase CACHE_EXT_TREE:\n@@ -1889,8 +1889,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tint fd, i;\n \tstruct stat st;\n \tunsigned long src_offset;\n-\tstruct cache_header *hdr;\n-\tvoid *mmap;\n+\tconst struct cache_header *hdr;\n+\tconst char *mmap;\n \tsize_t mmap_size;\n \tconst struct cache_entry *previous_ce = NULL;\n \n@@ -1918,7 +1918,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tdie_errno(\"unable to map index file\");\n \tclose(fd);\n \n-\thdr = mmap;\n+\thdr = (const struct cache_header *)mmap;\n \tif (verify_hdr(hdr, mmap_size) < 0)\n \t\tgoto unmap;\n \n@@ -1943,7 +1943,7 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\tstruct cache_entry *ce;\n \t\tunsigned long consumed;\n \n-\t\tdisk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);\n+\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n \t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n \t\tset_index_entry(istate, i, ce);\n \n@@ -1961,21 +1961,20 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\t * in 4-byte network byte order.\n \t\t */\n \t\tuint32_t extsize;\n-\t\tmemcpy(&extsize, (char *)mmap + src_offset + 4, 4);\n-\t\textsize = ntohl(extsize);\n+\t\textsize = get_be32(mmap + src_offset + 4);\n \t\tif (read_index_extension(istate,\n-\t\t\t\t\t (const char *) mmap + src_offset,\n-\t\t\t\t\t (char *) mmap + src_offset + 8,\n+\t\t\t\t\t mmap + src_offset,\n+\t\t\t\t\t mmap + src_offset + 8,\n \t\t\t\t\t extsize) < 0)\n \t\t\tgoto unmap;\n \t\tsrc_offset += 8;\n \t\tsrc_offset += extsize;\n \t}\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \treturn istate->cache_nr;\n \n unmap:\n-\tmunmap(mmap, mmap_size);\n+\tmunmap((void *)mmap, mmap_size);\n \tdie(\"index file corrupt\");\n }\n \n-- \n2.18.0.windows.1\n\n"},{"id":"360042","messageId":"20181010155938.20996-4-peartben@gmail.com","threadId":"49204","inReplyTo":"20181010155938.20996-1-peartben@gmail.com","subject":"[PATCH v8 3/7] eoie: add End of Index Entry (EOIE) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-10T15:59:34Z","receivedAt":"2018-10-10T15:59:56Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nThe End of Index Entry (EOIE) is used to locate the end of the variable\nlength index entries and the beginning of the extensions. Code can take\nadvantage of this to quickly locate the index extensions without having\nto parse through all of the index entries.\n\nThe EOIE extension is always written out to the index file including to\nthe shared index when using the split index feature. Because it is always\nwritten out, the SHA checksums in t/t1700-split-index.sh were updated\nto reflect its inclusion.\n\nIt is written as an optional extension to ensure compatibility with other\ngit implementations that do not yet support it.  It is always written out\nto ensure it is available as often as possible to speed up index operations.\n\nBecause it must be able to be loaded before the variable length cache\nentries and other index extensions, this extension must be written last.\nThe signature for this extension is { 'E', 'O', 'I', 'E' }.\n\nThe extension consists of:\n\n- 32-bit offset to the end of the index entries\n\n- 160-bit SHA-1 over the extension types and their sizes (but not\ntheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\nlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\nthen the hash would be:\n\nSHA-1(\"TREE\" + <binary representation of N> +\n    \"REUC\" + <binary representation of M>)\n\nSigned-off-by: Ben Peart <benpeart@microsoft.com>\n---\n Documentation/technical/index-format.txt |  23 ++++\n read-cache.c                             | 158 +++++++++++++++++++++--\n t/t1700-split-index.sh                   |   8 +-\n 3 files changed, 177 insertions(+), 12 deletions(-)\n\ndiff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\nindex db3572626b..6bc2d90f7f 100644\n--- a/Documentation/technical/index-format.txt\n+++ b/Documentation/technical/index-format.txt\n@@ -314,3 +314,26 @@ The remaining data of each directory block is grouped by type:\n \n   - An ewah bitmap, the n-th bit indicates whether the n-th index entry\n     is not CE_FSMONITOR_VALID.\n+\n+== End of Index Entry\n+\n+  The End of Index Entry (EOIE) is used to locate the end of the variable\n+  length index entries and the begining of the extensions. Code can take\n+  advantage of this to quickly locate the index extensions without having\n+  to parse through all of the index entries.\n+\n+  Because it must be able to be loaded before the variable length cache\n+  entries and other index extensions, this extension must be written last.\n+  The signature for this extension is { 'E', 'O', 'I', 'E' }.\n+\n+  The extension consists of:\n+\n+  - 32-bit offset to the end of the index entries\n+\n+  - 160-bit SHA-1 over the extension types and their sizes (but not\n+\ttheir contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\tlong, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\tthen the hash would be:\n+\n+\tSHA-1(\"TREE\" + <binary representation of N> +\n+\t\t\"REUC\" + <binary representation of M>)\ndiff --git a/read-cache.c b/read-cache.c\nindex 6ba99e2c96..4781515252 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -43,6 +43,7 @@\n #define CACHE_EXT_LINK 0x6c696e6b\t  /* \"link\" */\n #define CACHE_EXT_UNTRACKED 0x554E5452\t  /* \"UNTR\" */\n #define CACHE_EXT_FSMONITOR 0x46534D4E\t  /* \"FSMN\" */\n+#define CACHE_EXT_ENDOFINDEXENTRIES 0x454F4945\t/* \"EOIE\" */\n \n /* changes that can be kept in $GIT_DIR/index (basically all extensions) */\n #define EXTMASK (RESOLVE_UNDO_CHANGED | CACHE_TREE_CHANGED | \\\n@@ -1693,6 +1694,9 @@ static int read_index_extension(struct index_state *istate,\n \tcase CACHE_EXT_FSMONITOR:\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n+\tcase CACHE_EXT_ENDOFINDEXENTRIES:\n+\t\t/* already handled in do_read_index() */\n+\t\tbreak;\n \tdefault:\n \t\tif (*ext < 'A' || 'Z' < *ext)\n \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n@@ -1883,6 +1887,9 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -2190,11 +2197,15 @@ static int ce_write(git_hash_ctx *context, int fd, void *data, unsigned int len)\n \treturn 0;\n }\n \n-static int write_index_ext_header(git_hash_ctx *context, int fd,\n-\t\t\t\t  unsigned int ext, unsigned int sz)\n+static int write_index_ext_header(git_hash_ctx *context, git_hash_ctx *eoie_context,\n+\t\t\t\t  int fd, unsigned int ext, unsigned int sz)\n {\n \text = htonl(ext);\n \tsz = htonl(sz);\n+\tif (eoie_context) {\n+\t\tthe_hash_algo->update_fn(eoie_context, &ext, 4);\n+\t\tthe_hash_algo->update_fn(eoie_context, &sz, 4);\n+\t}\n \treturn ((ce_write(context, fd, &ext, 4) < 0) ||\n \t\t(ce_write(context, fd, &sz, 4) < 0)) ? -1 : 0;\n }\n@@ -2437,7 +2448,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n {\n \tuint64_t start = getnanotime();\n \tint newfd = tempfile->fd;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, eoie_c;\n \tstruct cache_header hdr;\n \tint i, err = 0, removed, extended, hdr_version;\n \tstruct cache_entry **cache = istate->cache;\n@@ -2446,6 +2457,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct ondisk_cache_entry_extended ondisk;\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n+\toff_t offset;\n \n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n@@ -2479,6 +2491,10 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tif (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n \t\treturn -1;\n \n+\toffset = lseek(newfd, 0, SEEK_CUR);\n+\tif (offset < 0)\n+\t\treturn -1;\n+\toffset += write_buffer_len;\n \tprevious_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n \n \tfor (i = 0; i < entries; i++) {\n@@ -2512,11 +2528,17 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\treturn err;\n \n \t/* Write extension data here */\n+\toffset = lseek(newfd, 0, SEEK_CUR);\n+\tif (offset < 0)\n+\t\treturn -1;\n+\toffset += write_buffer_len;\n+\tthe_hash_algo->init_fn(&eoie_c);\n+\n \tif (!strip_extensions && istate->split_index) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\terr = write_link_extension(&sb, istate) < 0 ||\n-\t\t\twrite_index_ext_header(&c, newfd, CACHE_EXT_LINK,\n+\t\t\twrite_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_LINK,\n \t\t\t\t\t       sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2527,7 +2549,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tcache_tree_write(&sb, istate->cache_tree);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_TREE, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_TREE, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2537,7 +2559,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\tresolve_undo_write(&sb, istate->resolve_undo);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_RESOLVE_UNDO,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_RESOLVE_UNDO,\n \t\t\t\t\t     sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2548,7 +2570,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_untracked_extension(&sb, istate->untracked);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_UNTRACKED,\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_UNTRACKED,\n \t\t\t\t\t     sb.len) < 0 ||\n \t\t\tce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n@@ -2559,7 +2581,24 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_fsmonitor_extension(&sb, istate);\n-\t\terr = write_index_ext_header(&c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_FSMONITOR, sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+\n+\t/*\n+\t * CACHE_EXT_ENDOFINDEXENTRIES must be written as the last entry before the SHA1\n+\t * so that it can be found and processed before all the index entries are\n+\t * read.  Write it out regardless of the strip_extensions parameter as we need it\n+\t * when loading the shared index.\n+\t */\n+\tif (offset) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\twrite_eoie_extension(&sb, &eoie_c, offset);\n+\t\terr = write_index_ext_header(&c, NULL, newfd, CACHE_EXT_ENDOFINDEXENTRIES, sb.len) < 0\n \t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n \t\tstrbuf_release(&sb);\n \t\tif (err)\n@@ -2975,3 +3014,106 @@ int should_validate_cache_entries(void)\n \n \treturn validate_index_cache_entries;\n }\n+\n+#define EOIE_SIZE (4 + GIT_SHA1_RAWSZ) /* <4-byte offset> + <20-byte hash> */\n+#define EOIE_SIZE_WITH_HEADER (4 + 4 + EOIE_SIZE) /* <4-byte signature> + <4-byte length> + EOIE_SIZE */\n+\n+static size_t read_eoie_extension(const char *mmap, size_t mmap_size)\n+{\n+\t/*\n+\t * The end of index entries (EOIE) extension is guaranteed to be last\n+\t * so that it can be found by scanning backwards from the EOF.\n+\t *\n+\t * \"EOIE\"\n+\t * <4-byte length>\n+\t * <4-byte offset>\n+\t * <20-byte hash>\n+\t */\n+\tconst char *index, *eoie;\n+\tuint32_t extsize;\n+\tsize_t offset, src_offset;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\tgit_hash_ctx c;\n+\n+\t/* ensure we have an index big enough to contain an EOIE extension */\n+\tif (mmap_size < sizeof(struct cache_header) + EOIE_SIZE_WITH_HEADER + the_hash_algo->rawsz)\n+\t\treturn 0;\n+\n+\t/* validate the extension signature */\n+\tindex = eoie = mmap + mmap_size - EOIE_SIZE_WITH_HEADER - the_hash_algo->rawsz;\n+\tif (CACHE_EXT(index) != CACHE_EXT_ENDOFINDEXENTRIES)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/* validate the extension size */\n+\textsize = get_be32(index);\n+\tif (extsize != EOIE_SIZE)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * Validate the offset we're going to look for the first extension\n+\t * signature is after the index header and before the eoie extension.\n+\t */\n+\toffset = get_be32(index);\n+\tif (mmap + offset < mmap + sizeof(struct cache_header))\n+\t\treturn 0;\n+\tif (mmap + offset >= eoie)\n+\t\treturn 0;\n+\tindex += sizeof(uint32_t);\n+\n+\t/*\n+\t * The hash is computed over extension types and their sizes (but not\n+\t * their contents).  E.g. if we have \"TREE\" extension that is N-bytes\n+\t * long, \"REUC\" extension that is M-bytes long, followed by \"EOIE\",\n+\t * then the hash would be:\n+\t *\n+\t * SHA-1(\"TREE\" + <binary representation of N> +\n+\t *\t \"REUC\" + <binary representation of M>)\n+\t */\n+\tsrc_offset = offset;\n+\tthe_hash_algo->init_fn(&c);\n+\twhile (src_offset < mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize;\n+\t\tmemcpy(&extsize, mmap + src_offset + 4, 4);\n+\t\textsize = ntohl(extsize);\n+\n+\t\t/* verify the extension size isn't so large it will wrap around */\n+\t\tif (src_offset + 8 + extsize < src_offset)\n+\t\t\treturn 0;\n+\n+\t\tthe_hash_algo->update_fn(&c, mmap + src_offset, 8);\n+\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\tthe_hash_algo->final_fn(hash, &c);\n+\tif (!hasheq(hash, (const unsigned char *)index))\n+\t\treturn 0;\n+\n+\t/* Validate that the extension offsets returned us back to the eoie extension. */\n+\tif (src_offset != mmap_size - the_hash_algo->rawsz - EOIE_SIZE_WITH_HEADER)\n+\t\treturn 0;\n+\n+\treturn offset;\n+}\n+\n+static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset)\n+{\n+\tuint32_t buffer;\n+\tunsigned char hash[GIT_MAX_RAWSZ];\n+\n+\t/* offset */\n+\tput_be32(&buffer, offset);\n+\tstrbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+\t/* hash */\n+\tthe_hash_algo->final_fn(hash, eoie_context);\n+\tstrbuf_add(sb, hash, the_hash_algo->rawsz);\n+}\ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex be22398a85..8e17f8e7a0 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -15,11 +15,11 @@ test_expect_success 'enable split index' '\n \tindexversion=$(test-tool index-version <.git/index) &&\n \tif test \"$indexversion\" = \"4\"\n \tthen\n-\t\town=432ef4b63f32193984f339431fd50ca796493569\n-\t\tbase=508851a7f0dfa8691e9f69c7f055865389012491\n+\t\town=3527df833c6c100d3d1d921a9a782d62a8be4b58\n+\t\tbase=746f7ab2ed44fb839efdfbffcf399d0b113fb4cb\n \telse\n-\t\town=8299b0bcd1ac364e5f1d7768efb62fa2da79a339\n-\t\tbase=39d890139ee5356c7ef572216cebcd27aa41f9df\n+\t\town=5e9b60117ece18da410ddecc8b8d43766a0e4204\n+\t\tbase=4370042739b31cd17a5c5cd6043a77c9a00df113\n \tfi &&\n \tcat >expect <<-EOF &&\n \town $own\n-- \n2.18.0.windows.1\n\n"},{"id":"360041","messageId":"20181010155938.20996-5-peartben@gmail.com","threadId":"49204","inReplyTo":"20181010155938.20996-1-peartben@gmail.com","subject":"[PATCH v8 4/7] config: add new index.threads config setting","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-10T15:59:35Z","receivedAt":"2018-10-10T15:59:57Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nAdd support for a new index.threads config setting which will be used to\ncontrol the threading code in do_read_index().  A value of 0 will tell the\nindex code to automatically determine the correct number of threads to use.\nA value of 1 will make the code single threaded.  A value greater than 1\nwill set the maximum number of threads to use.\n\nFor testing purposes, this setting can be overwritten by setting the\nGIT_TEST_INDEX_THREADS=<n> environment variable to a value greater than 0.\n\nSigned-off-by: Ben Peart <benpeart@microsoft.com>\n---\n Documentation/config.txt |  7 +++++++\n config.c                 | 18 ++++++++++++++++++\n config.h                 |  1 +\n t/README                 |  5 +++++\n t/t1700-split-index.sh   |  5 +++++\n 5 files changed, 36 insertions(+)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex ad0f4510c3..8fd973b76b 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -2413,6 +2413,13 @@ imap::\n \tThe configuration variables in the 'imap' section are described\n \tin linkgit:git-imap-send[1].\n \n+index.threads::\n+\tSpecifies the number of threads to spawn when loading the index.\n+\tThis is meant to reduce index load time on multiprocessor machines.\n+\tSpecifying 0 or 'true' will cause Git to auto-detect the number of\n+\tCPU's and set the number of threads accordingly. Specifying 1 or\n+\t'false' will disable multithreading. Defaults to 'true'.\n+\n index.version::\n \tSpecify the version with which new index files should be\n \tinitialized.  This does not affect existing repositories.\ndiff --git a/config.c b/config.c\nindex 3461993f0a..2ee29f6f86 100644\n--- a/config.c\n+++ b/config.c\n@@ -2289,6 +2289,24 @@ int git_config_get_fsmonitor(void)\n \treturn 0;\n }\n \n+int git_config_get_index_threads(void)\n+{\n+\tint is_bool, val = 0;\n+\n+\tval = git_env_ulong(\"GIT_TEST_INDEX_THREADS\", 0);\n+\tif (val)\n+\t\treturn val;\n+\n+\tif (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n+\t\tif (is_bool)\n+\t\t\treturn val ? 0 : 1;\n+\t\telse\n+\t\t\treturn val;\n+\t}\n+\n+\treturn 0; /* auto */\n+}\n+\n NORETURN\n void git_die_config_linenr(const char *key, const char *filename, int linenr)\n {\ndiff --git a/config.h b/config.h\nindex ab46e0165d..a06027e69b 100644\n--- a/config.h\n+++ b/config.h\n@@ -250,6 +250,7 @@ extern int git_config_get_untracked_cache(void);\n extern int git_config_get_split_index(void);\n extern int git_config_get_max_percent_split_change(void);\n extern int git_config_get_fsmonitor(void);\n+extern int git_config_get_index_threads(void);\n \n /* This dies if the configured or default date is in the future */\n extern int git_config_get_expiry(const char *key, const char **output);\ndiff --git a/t/README b/t/README\nindex 3ea6c85460..8f5c0620ea 100644\n--- a/t/README\n+++ b/t/README\n@@ -327,6 +327,11 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_INDEX_THREADS=<n> enables exercising the multi-threaded loading\n+of the index for the whole test suite by bypassing the default number of\n+cache entries and thread minimums. Setting this to 1 will make the\n+index loading single threaded.\n+\n Naming Tests\n ------------\n \ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex 8e17f8e7a0..ef9349bd70 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -6,7 +6,12 @@ test_description='split index mode tests'\n \n # We need total control of index splitting here\n sane_unset GIT_TEST_SPLIT_INDEX\n+\n+# Testing a hard coded SHA against an index with an extension\n+# that can vary from run to run is problematic so we disable\n+# those extensions.\n sane_unset GIT_FSMONITOR_TEST\n+sane_unset GIT_TEST_INDEX_THREADS\n \n test_expect_success 'enable split index' '\n \tgit config splitIndex.maxPercentChange 100 &&\n-- \n2.18.0.windows.1\n\n"},{"id":"360043","messageId":"20181010155938.20996-6-peartben@gmail.com","threadId":"49204","inReplyTo":"20181010155938.20996-1-peartben@gmail.com","subject":"[PATCH v8 5/7] read-cache: load cache extensions on a worker thread","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-10T15:59:36Z","receivedAt":"2018-10-10T15:59:59Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nThis patch helps address the CPU cost of loading the index by loading\nthe cache extensions on a worker thread in parallel with loading the cache\nentries.\n\nIn some cases, loading the extensions takes longer than loading the\ncache entries so this patch utilizes the new EOIE to start the thread to\nload the extensions before loading all the cache entries in parallel.\n\nThis is possible because the current extensions don't access the cache\nentries in the index_state structure so are OK that they don't all exist\nyet.\n\nThe CACHE_EXT_TREE, CACHE_EXT_RESOLVE_UNDO, and CACHE_EXT_UNTRACKED\nextensions don't even get a pointer to the index so don't have access to the\ncache entries.\n\nCACHE_EXT_LINK only uses the index_state to initialize the split index.\nCACHE_EXT_FSMONITOR only uses the index_state to save the fsmonitor last\nupdate and dirty flags.\n\nI used p0002-read-cache.sh to generate some performance data:\n\n\tTest w/100,000 files reduced the time by 0.53%\n\tTest w/1,000,000 files reduced the time by 27.78%\n\nSigned-off-by: Ben Peart <benpeart@microsoft.com>\n---\n read-cache.c | 95 +++++++++++++++++++++++++++++++++++++++++++---------\n 1 file changed, 79 insertions(+), 16 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 4781515252..2214b3153d 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -23,6 +23,7 @@\n #include \"split-index.h\"\n #include \"utf8.h\"\n #include \"fsmonitor.h\"\n+#include \"thread-utils.h\"\n \n /* Mask for the name length in ce_flags in the on-disk index */\n \n@@ -1890,6 +1891,44 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n \n+struct load_index_extensions\n+{\n+#ifndef NO_PTHREADS\n+\tpthread_t pthread;\n+#endif\n+\tstruct index_state *istate;\n+\tconst char *mmap;\n+\tsize_t mmap_size;\n+\tunsigned long src_offset;\n+};\n+\n+static void *load_index_extensions(void *_data)\n+{\n+\tstruct load_index_extensions *p = _data;\n+\tunsigned long src_offset = p->src_offset;\n+\n+\twhile (src_offset <= p->mmap_size - the_hash_algo->rawsz - 8) {\n+\t\t/* After an array of active_nr index entries,\n+\t\t * there can be arbitrary number of extended\n+\t\t * sections, each of which is prefixed with\n+\t\t * extension name (4-byte) and section length\n+\t\t * in 4-byte network byte order.\n+\t\t */\n+\t\tuint32_t extsize = get_be32(p->mmap + src_offset + 4);\n+\t\tif (read_index_extension(p->istate,\n+\t\t\t\t\t p->mmap + src_offset,\n+\t\t\t\t\t p->mmap + src_offset + 8,\n+\t\t\t\t\t extsize) < 0) {\n+\t\t\tmunmap((void *)p->mmap, p->mmap_size);\n+\t\t\tdie(_(\"index file corrupt\"));\n+\t\t}\n+\t\tsrc_offset += 8;\n+\t\tsrc_offset += extsize;\n+\t}\n+\n+\treturn NULL;\n+}\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -1900,6 +1939,11 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tconst char *mmap;\n \tsize_t mmap_size;\n \tconst struct cache_entry *previous_ce = NULL;\n+\tstruct load_index_extensions p;\n+\tsize_t extension_offset = 0;\n+#ifndef NO_PTHREADS\n+\tint nr_threads;\n+#endif\n \n \tif (istate->initialized)\n \t\treturn istate->cache_nr;\n@@ -1936,6 +1980,30 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->cache = xcalloc(istate->cache_alloc, sizeof(*istate->cache));\n \tistate->initialized = 1;\n \n+\tp.istate = istate;\n+\tp.mmap = mmap;\n+\tp.mmap_size = mmap_size;\n+\n+#ifndef NO_PTHREADS\n+\tnr_threads = git_config_get_index_threads();\n+\tif (!nr_threads)\n+\t\tnr_threads = online_cpus();\n+\n+\tif (nr_threads > 1) {\n+\t\textension_offset = read_eoie_extension(mmap, mmap_size);\n+\t\tif (extension_offset) {\n+\t\t\tint err;\n+\n+\t\t\tp.src_offset = extension_offset;\n+\t\t\terr = pthread_create(&p.pthread, NULL, load_index_extensions, &p);\n+\t\t\tif (err)\n+\t\t\t\tdie(_(\"unable to create load_index_extensions thread: %s\"), strerror(err));\n+\n+\t\t\tnr_threads--;\n+\t\t}\n+\t}\n+#endif\n+\n \tif (istate->version == 4) {\n \t\tmem_pool_init(&istate->ce_mem_pool,\n \t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n@@ -1960,22 +2028,17 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n-\twhile (src_offset <= mmap_size - the_hash_algo->rawsz - 8) {\n-\t\t/* After an array of active_nr index entries,\n-\t\t * there can be arbitrary number of extended\n-\t\t * sections, each of which is prefixed with\n-\t\t * extension name (4-byte) and section length\n-\t\t * in 4-byte network byte order.\n-\t\t */\n-\t\tuint32_t extsize;\n-\t\textsize = get_be32(mmap + src_offset + 4);\n-\t\tif (read_index_extension(istate,\n-\t\t\t\t\t mmap + src_offset,\n-\t\t\t\t\t mmap + src_offset + 8,\n-\t\t\t\t\t extsize) < 0)\n-\t\t\tgoto unmap;\n-\t\tsrc_offset += 8;\n-\t\tsrc_offset += extsize;\n+\t/* if we created a thread, join it otherwise load the extensions on the primary thread */\n+#ifndef NO_PTHREADS\n+\tif (extension_offset) {\n+\t\tint ret = pthread_join(p.pthread, NULL);\n+\t\tif (ret)\n+\t\t\tdie(_(\"unable to join load_index_extensions thread: %s\"), strerror(ret));\n+\t}\n+#endif\n+\tif (!extension_offset) {\n+\t\tp.src_offset = src_offset;\n+\t\tload_index_extensions(&p);\n \t}\n \tmunmap((void *)mmap, mmap_size);\n \treturn istate->cache_nr;\n-- \n2.18.0.windows.1\n\n"},{"id":"360045","messageId":"20181010155938.20996-7-peartben@gmail.com","threadId":"49204","inReplyTo":"20181010155938.20996-1-peartben@gmail.com","subject":"[PATCH v8 6/7] ieot: add Index Entry Offset Table (IEOT) extension","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-10T15:59:37Z","receivedAt":"2018-10-10T16:00:00Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nThis patch enables addressing the CPU cost of loading the index by adding\nadditional data to the index that will allow us to efficiently multi-\nthread the loading and conversion of cache entries.\n\nIt accomplishes this by adding an (optional) index extension that is a\ntable of offsets to blocks of cache entries in the index file.  To make\nthis work for V4 indexes, when writing the cache entries, it periodically\n\"resets\" the prefix-compression by encoding the current entry as if the\npath name for the previous entry is completely different and saves the\noffset of that entry in the IEOT.  Basically, with V4 indexes, it\ngenerates offsets into blocks of prefix-compressed entries.\n\nSigned-off-by: Ben Peart <benpeart@microsoft.com>\n---\n Documentation/technical/index-format.txt |  18 +++\n read-cache.c                             | 196 ++++++++++++++++++++++-\n 2 files changed, 211 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/technical/index-format.txt b/Documentation/technical/index-format.txt\nindex 6bc2d90f7f..7c4d67aa6a 100644\n--- a/Documentation/technical/index-format.txt\n+++ b/Documentation/technical/index-format.txt\n@@ -337,3 +337,21 @@ The remaining data of each directory block is grouped by type:\n \n \tSHA-1(\"TREE\" + <binary representation of N> +\n \t\t\"REUC\" + <binary representation of M>)\n+\n+== Index Entry Offset Table\n+\n+  The Index Entry Offset Table (IEOT) is used to help address the CPU\n+  cost of loading the index by enabling multi-threading the process of\n+  converting cache entries from the on-disk format to the in-memory format.\n+  The signature for this extension is { 'I', 'E', 'O', 'T' }.\n+\n+  The extension consists of:\n+\n+  - 32-bit version (currently 1)\n+\n+  - A number of index offset entries each consisting of:\n+\n+    - 32-bit offset from the begining of the file to the first cache entry\n+\tin this block of entries.\n+\n+    - 32-bit count of cache entries in this block\ndiff --git a/read-cache.c b/read-cache.c\nindex 2214b3153d..3ace29d58f 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -45,6 +45,7 @@\n #define CACHE_EXT_UNTRACKED 0x554E5452\t  /* \"UNTR\" */\n #define CACHE_EXT_FSMONITOR 0x46534D4E\t  /* \"FSMN\" */\n #define CACHE_EXT_ENDOFINDEXENTRIES 0x454F4945\t/* \"EOIE\" */\n+#define CACHE_EXT_INDEXENTRYOFFSETTABLE 0x49454F54 /* \"IEOT\" */\n \n /* changes that can be kept in $GIT_DIR/index (basically all extensions) */\n #define EXTMASK (RESOLVE_UNDO_CHANGED | CACHE_TREE_CHANGED | \\\n@@ -1696,6 +1697,7 @@ static int read_index_extension(struct index_state *istate,\n \t\tread_fsmonitor_extension(istate, data, sz);\n \t\tbreak;\n \tcase CACHE_EXT_ENDOFINDEXENTRIES:\n+\tcase CACHE_EXT_INDEXENTRYOFFSETTABLE:\n \t\t/* already handled in do_read_index() */\n \t\tbreak;\n \tdefault:\n@@ -1888,6 +1890,23 @@ static size_t estimate_cache_size(size_t ondisk_size, unsigned int entries)\n \treturn ondisk_size + entries * per_entry;\n }\n \n+struct index_entry_offset\n+{\n+\t/* starting byte offset into index file, count of index entries in this block */\n+\tint offset, nr;\n+};\n+\n+struct index_entry_offset_table\n+{\n+\tint nr;\n+\tstruct index_entry_offset entries[FLEX_ARRAY];\n+};\n+\n+#ifndef NO_PTHREADS\n+static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset);\n+static void write_ieot_extension(struct strbuf *sb, struct index_entry_offset_table *ieot);\n+#endif\n+\n static size_t read_eoie_extension(const char *mmap, size_t mmap_size);\n static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context, size_t offset);\n \n@@ -1929,6 +1948,15 @@ static void *load_index_extensions(void *_data)\n \treturn NULL;\n }\n \n+/*\n+ * Mostly randomly chosen maximum thread counts: we\n+ * cap the parallelism to online_cpus() threads, and we want\n+ * to have at least 10000 cache entries per thread for it to\n+ * be worth starting a thread.\n+ */\n+\n+#define THREAD_COST\t\t(10000)\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n@@ -2521,6 +2549,9 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n+\tint ieot_blocks = 1;\n+\tstruct index_entry_offset_table *ieot = NULL;\n+\tint nr, nr_threads;\n \n \tfor (i = removed = extended = 0; i < entries; i++) {\n \t\tif (cache[i]->ce_flags & CE_REMOVE)\n@@ -2554,10 +2585,44 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tif (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n \t\treturn -1;\n \n+#ifndef NO_PTHREADS\n+\tnr_threads = git_config_get_index_threads();\n+\tif (nr_threads != 1) {\n+\t\tint ieot_blocks, cpus;\n+\n+\t\t/*\n+\t\t * ensure default number of ieot blocks maps evenly to the\n+\t\t * default number of threads that will process them leaving\n+\t\t * room for the thread to load the index extensions.\n+\t\t */\n+\t\tif (!nr_threads) {\n+\t\t\tieot_blocks = istate->cache_nr / THREAD_COST;\n+\t\t\tcpus = online_cpus();\n+\t\t\tif (ieot_blocks > cpus - 1)\n+\t\t\t\tieot_blocks = cpus - 1;\n+\t\t} else {\n+\t\t\tieot_blocks = nr_threads;\n+\t\t}\n+\n+\t\t/*\n+\t\t * no reason to write out the IEOT extension if we don't\n+\t\t * have enough blocks to utilize multi-threading\n+\t\t */\n+\t\tif (ieot_blocks > 1) {\n+\t\t\tieot = xcalloc(1, sizeof(struct index_entry_offset_table)\n+\t\t\t\t+ (ieot_blocks * sizeof(struct index_entry_offset)));\n+\t\t\tieot_blocks = DIV_ROUND_UP(entries, ieot_blocks);\n+\t\t}\n+\t}\n+#endif\n+\n \toffset = lseek(newfd, 0, SEEK_CUR);\n-\tif (offset < 0)\n+\tif (offset < 0) {\n+\t\tfree(ieot);\n \t\treturn -1;\n+\t}\n \toffset += write_buffer_len;\n+\tnr = 0;\n \tprevious_name = (hdr_version == 4) ? &previous_name_buf : NULL;\n \n \tfor (i = 0; i < entries; i++) {\n@@ -2579,24 +2644,74 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \n \t\t\tdrop_cache_tree = 1;\n \t\t}\n+\t\tif (ieot && i && (i % ieot_blocks == 0)) {\n+\t\t\tieot->entries[ieot->nr].nr = nr;\n+\t\t\tieot->entries[ieot->nr].offset = offset;\n+\t\t\tieot->nr++;\n+\t\t\t/*\n+\t\t\t * If we have a V4 index, set the first byte to an invalid\n+\t\t\t * character to ensure there is nothing common with the previous\n+\t\t\t * entry\n+\t\t\t */\n+\t\t\tif (previous_name)\n+\t\t\t\tprevious_name->buf[0] = 0;\n+\t\t\tnr = 0;\n+\t\t\toffset = lseek(newfd, 0, SEEK_CUR);\n+\t\t\tif (offset < 0) {\n+\t\t\t\tfree(ieot);\n+\t\t\t\treturn -1;\n+\t\t\t}\n+\t\t\toffset += write_buffer_len;\n+\t\t}\n \t\tif (ce_write_entry(&c, newfd, ce, previous_name, (struct ondisk_cache_entry *)&ondisk) < 0)\n \t\t\terr = -1;\n \n \t\tif (err)\n \t\t\tbreak;\n+\t\tnr++;\n+\t}\n+\tif (ieot && nr) {\n+\t\tieot->entries[ieot->nr].nr = nr;\n+\t\tieot->entries[ieot->nr].offset = offset;\n+\t\tieot->nr++;\n \t}\n \tstrbuf_release(&previous_name_buf);\n \n-\tif (err)\n+\tif (err) {\n+\t\tfree(ieot);\n \t\treturn err;\n+\t}\n \n \t/* Write extension data here */\n \toffset = lseek(newfd, 0, SEEK_CUR);\n-\tif (offset < 0)\n+\tif (offset < 0) {\n+\t\tfree(ieot);\n \t\treturn -1;\n+\t}\n \toffset += write_buffer_len;\n \tthe_hash_algo->init_fn(&eoie_c);\n \n+\t/*\n+\t * Lets write out CACHE_EXT_INDEXENTRYOFFSETTABLE first so that we\n+\t * can minimize the number of extensions we have to scan through to\n+\t * find it during load.  Write it out regardless of the\n+\t * strip_extensions parameter as we need it when loading the shared\n+\t * index.\n+\t */\n+#ifndef NO_PTHREADS\n+\tif (ieot) {\n+\t\tstruct strbuf sb = STRBUF_INIT;\n+\n+\t\twrite_ieot_extension(&sb, ieot);\n+\t\terr = write_index_ext_header(&c, &eoie_c, newfd, CACHE_EXT_INDEXENTRYOFFSETTABLE, sb.len) < 0\n+\t\t\t|| ce_write(&c, newfd, sb.buf, sb.len) < 0;\n+\t\tstrbuf_release(&sb);\n+\t\tfree(ieot);\n+\t\tif (err)\n+\t\t\treturn -1;\n+\t}\n+#endif\n+\n \tif (!strip_extensions && istate->split_index) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n@@ -3180,3 +3295,78 @@ static void write_eoie_extension(struct strbuf *sb, git_hash_ctx *eoie_context,\n \tthe_hash_algo->final_fn(hash, eoie_context);\n \tstrbuf_add(sb, hash, the_hash_algo->rawsz);\n }\n+\n+#ifndef NO_PTHREADS\n+#define IEOT_VERSION\t(1)\n+\n+static struct index_entry_offset_table *read_ieot_extension(const char *mmap, size_t mmap_size, size_t offset)\n+{\n+       const char *index = NULL;\n+       uint32_t extsize, ext_version;\n+       struct index_entry_offset_table *ieot;\n+       int i, nr;\n+\n+       /* find the IEOT extension */\n+       if (!offset)\n+\t       return NULL;\n+       while (offset <= mmap_size - the_hash_algo->rawsz - 8) {\n+\t       extsize = get_be32(mmap + offset + 4);\n+\t       if (CACHE_EXT((mmap + offset)) == CACHE_EXT_INDEXENTRYOFFSETTABLE) {\n+\t\t       index = mmap + offset + 4 + 4;\n+\t\t       break;\n+\t       }\n+\t       offset += 8;\n+\t       offset += extsize;\n+       }\n+       if (!index)\n+\t       return NULL;\n+\n+       /* validate the version is IEOT_VERSION */\n+       ext_version = get_be32(index);\n+       if (ext_version != IEOT_VERSION) {\n+\t       error(\"invalid IEOT version %d\", ext_version);\n+\t       return NULL;\n+       }\n+       index += sizeof(uint32_t);\n+\n+       /* extension size - version bytes / bytes per entry */\n+       nr = (extsize - sizeof(uint32_t)) / (sizeof(uint32_t) + sizeof(uint32_t));\n+       if (!nr) {\n+\t       error(\"invalid number of IEOT entries %d\", nr);\n+\t       return NULL;\n+       }\n+       ieot = xmalloc(sizeof(struct index_entry_offset_table)\n+\t       + (nr * sizeof(struct index_entry_offset)));\n+       ieot->nr = nr;\n+       for (i = 0; i < nr; i++) {\n+\t       ieot->entries[i].offset = get_be32(index);\n+\t       index += sizeof(uint32_t);\n+\t       ieot->entries[i].nr = get_be32(index);\n+\t       index += sizeof(uint32_t);\n+       }\n+\n+       return ieot;\n+}\n+\n+static void write_ieot_extension(struct strbuf *sb, struct index_entry_offset_table *ieot)\n+{\n+       uint32_t buffer;\n+       int i;\n+\n+       /* version */\n+       put_be32(&buffer, IEOT_VERSION);\n+       strbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+       /* ieot */\n+       for (i = 0; i < ieot->nr; i++) {\n+\n+\t       /* offset */\n+\t       put_be32(&buffer, ieot->entries[i].offset);\n+\t       strbuf_add(sb, &buffer, sizeof(uint32_t));\n+\n+\t       /* count */\n+\t       put_be32(&buffer, ieot->entries[i].nr);\n+\t       strbuf_add(sb, &buffer, sizeof(uint32_t));\n+       }\n+}\n+#endif\n-- \n2.18.0.windows.1\n\n"},{"id":"360044","messageId":"20181010155938.20996-8-peartben@gmail.com","threadId":"49204","inReplyTo":"20181010155938.20996-1-peartben@gmail.com","subject":"[PATCH v8 7/7] read-cache: load cache entries on worker threads","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-10T15:59:38Z","receivedAt":"2018-10-10T16:00:02Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"From: Ben Peart <benpeart@microsoft.com>\n\nThis patch helps address the CPU cost of loading the index by utilizing\nthe Index Entry Offset Table (IEOT) to divide loading and conversion of\nthe cache entries across multiple threads in parallel.\n\nI used p0002-read-cache.sh to generate some performance data:\n\nTest w/100,000 files reduced the time by 32.24%\nTest w/1,000,000 files reduced the time by -4.77%\n\nNote that on the 1,000,000 files case, multi-threading the cache entry parsing\ndoes not yield a performance win.  This is because the cost to parse the\nindex extensions in this repo, far outweigh the cost of loading the cache\nentries.\n\nThe high cost of parsing the index extensions is driven by the cache tree\nand the untracked cache extensions. As this is currently the longest pole,\nany reduction in this time will reduce the overall index load times so is\nworth further investigation in another patch series.\n\nSigned-off-by: Ben Peart <benpeart@microsoft.com>\n---\n read-cache.c | 230 ++++++++++++++++++++++++++++++++++++++++++---------\n 1 file changed, 193 insertions(+), 37 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 3ace29d58f..7acc2c86f4 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1720,7 +1720,8 @@ int read_index(struct index_state *istate)\n \treturn read_index_from(istate, get_index_file(), get_git_dir());\n }\n \n-static struct cache_entry *create_from_disk(struct index_state *istate,\n+static struct cache_entry *create_from_disk(struct mem_pool *ce_mem_pool,\n+\t\t\t\t\t    unsigned int version,\n \t\t\t\t\t    struct ondisk_cache_entry *ondisk,\n \t\t\t\t\t    unsigned long *ent_size,\n \t\t\t\t\t    const struct cache_entry *previous_ce)\n@@ -1737,7 +1738,7 @@ static struct cache_entry *create_from_disk(struct index_state *istate,\n \t * number of bytes to be stripped from the end of the previous name,\n \t * and the bytes to append to the result, to come up with its name.\n \t */\n-\tint expand_name_field = istate->version == 4;\n+\tint expand_name_field = version == 4;\n \n \t/* On-disk flags are just 16 bits */\n \tflags = get_be16(&ondisk->flags);\n@@ -1761,16 +1762,17 @@ static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\tconst unsigned char *cp = (const unsigned char *)name;\n \t\tsize_t strip_len, previous_len;\n \n-\t\tprevious_len = previous_ce ? previous_ce->ce_namelen : 0;\n+\t\t/* If we're at the begining of a block, ignore the previous name */\n \t\tstrip_len = decode_varint(&cp);\n-\t\tif (previous_len < strip_len) {\n-\t\t\tif (previous_ce)\n+\t\tif (previous_ce) {\n+\t\t\tprevious_len = previous_ce->ce_namelen;\n+\t\t\tif (previous_len < strip_len)\n \t\t\t\tdie(_(\"malformed name field in the index, near path '%s'\"),\n-\t\t\t\t    previous_ce->name);\n-\t\t\telse\n-\t\t\t\tdie(_(\"malformed name field in the index in the first path\"));\n+\t\t\t\t\tprevious_ce->name);\n+\t\t\tcopy_len = previous_len - strip_len;\n+\t\t} else {\n+\t\t\tcopy_len = 0;\n \t\t}\n-\t\tcopy_len = previous_len - strip_len;\n \t\tname = (const char *)cp;\n \t}\n \n@@ -1780,7 +1782,7 @@ static struct cache_entry *create_from_disk(struct index_state *istate,\n \t\t\tlen += copy_len;\n \t}\n \n-\tce = mem_pool__ce_alloc(istate->ce_mem_pool, len);\n+\tce = mem_pool__ce_alloc(ce_mem_pool, len);\n \n \tce->ce_stat_data.sd_ctime.sec = get_be32(&ondisk->ctime.sec);\n \tce->ce_stat_data.sd_mtime.sec = get_be32(&ondisk->mtime.sec);\n@@ -1948,6 +1950,52 @@ static void *load_index_extensions(void *_data)\n \treturn NULL;\n }\n \n+/*\n+ * A helper function that will load the specified range of cache entries\n+ * from the memory mapped file and add them to the given index.\n+ */\n+static unsigned long load_cache_entry_block(struct index_state *istate,\n+\t\t\tstruct mem_pool *ce_mem_pool, int offset, int nr, const char *mmap,\n+\t\t\tunsigned long start_offset, const struct cache_entry *previous_ce)\n+{\n+\tint i;\n+\tunsigned long src_offset = start_offset;\n+\n+\tfor (i = offset; i < offset + nr; i++) {\n+\t\tstruct ondisk_cache_entry *disk_ce;\n+\t\tstruct cache_entry *ce;\n+\t\tunsigned long consumed;\n+\n+\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n+\t\tce = create_from_disk(ce_mem_pool, istate->version, disk_ce, &consumed, previous_ce);\n+\t\tset_index_entry(istate, i, ce);\n+\n+\t\tsrc_offset += consumed;\n+\t\tprevious_ce = ce;\n+\t}\n+\treturn src_offset - start_offset;\n+}\n+\n+static unsigned long load_all_cache_entries(struct index_state *istate,\n+\t\t\tconst char *mmap, size_t mmap_size, unsigned long src_offset)\n+{\n+\tunsigned long consumed;\n+\n+\tif (istate->version == 4) {\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t\testimate_cache_size_from_compressed(istate->cache_nr));\n+\t} else {\n+\t\tmem_pool_init(&istate->ce_mem_pool,\n+\t\t\t\testimate_cache_size(mmap_size, istate->cache_nr));\n+\t}\n+\n+\tconsumed = load_cache_entry_block(istate, istate->ce_mem_pool,\n+\t\t\t\t\t0, istate->cache_nr, mmap, src_offset, NULL);\n+\treturn consumed;\n+}\n+\n+#ifndef NO_PTHREADS\n+\n /*\n  * Mostly randomly chosen maximum thread counts: we\n  * cap the parallelism to online_cpus() threads, and we want\n@@ -1957,20 +2005,123 @@ static void *load_index_extensions(void *_data)\n \n #define THREAD_COST\t\t(10000)\n \n+struct load_cache_entries_thread_data\n+{\n+\tpthread_t pthread;\n+\tstruct index_state *istate;\n+\tstruct mem_pool *ce_mem_pool;\n+\tint offset;\n+\tconst char *mmap;\n+\tstruct index_entry_offset_table *ieot;\n+\tint ieot_start;\t\t/* starting index into the ieot array */\n+\tint ieot_blocks;\t/* count of ieot entries to process */\n+\tunsigned long consumed;\t/* return # of bytes in index file processed */\n+};\n+\n+/*\n+ * A thread proc to run the load_cache_entries() computation\n+ * across multiple background threads.\n+ */\n+static void *load_cache_entries_thread(void *_data)\n+{\n+\tstruct load_cache_entries_thread_data *p = _data;\n+\tint i;\n+\n+\t/* iterate across all ieot blocks assigned to this thread */\n+\tfor (i = p->ieot_start; i < p->ieot_start + p->ieot_blocks; i++) {\n+\t\tp->consumed += load_cache_entry_block(p->istate, p->ce_mem_pool,\n+\t\t\tp->offset, p->ieot->entries[i].nr, p->mmap, p->ieot->entries[i].offset, NULL);\n+\t\tp->offset += p->ieot->entries[i].nr;\n+\t}\n+\treturn NULL;\n+}\n+\n+static unsigned long load_cache_entries_threaded(struct index_state *istate, const char *mmap, size_t mmap_size,\n+\t\t\tunsigned long src_offset, int nr_threads, struct index_entry_offset_table *ieot)\n+{\n+\tint i, offset, ieot_blocks, ieot_start, err;\n+\tstruct load_cache_entries_thread_data *data;\n+\tunsigned long consumed = 0;\n+\n+\t/* a little sanity checking */\n+\tif (istate->name_hash_initialized)\n+\t\tBUG(\"the name hash isn't thread safe\");\n+\n+\tmem_pool_init(&istate->ce_mem_pool, 0);\n+\n+\t/* ensure we have no more threads than we have blocks to process */\n+\tif (nr_threads > ieot->nr)\n+\t\tnr_threads = ieot->nr;\n+\tdata = xcalloc(nr_threads, sizeof(*data));\n+\n+\toffset = ieot_start = 0;\n+\tieot_blocks = DIV_ROUND_UP(ieot->nr, nr_threads);\n+\tfor (i = 0; i < nr_threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = &data[i];\n+\t\tint nr, j;\n+\n+\t\tif (ieot_start + ieot_blocks > ieot->nr)\n+\t\t\tieot_blocks = ieot->nr - ieot_start;\n+\n+\t\tp->istate = istate;\n+\t\tp->offset = offset;\n+\t\tp->mmap = mmap;\n+\t\tp->ieot = ieot;\n+\t\tp->ieot_start = ieot_start;\n+\t\tp->ieot_blocks = ieot_blocks;\n+\n+\t\t/* create a mem_pool for each thread */\n+\t\tnr = 0;\n+\t\tfor (j = p->ieot_start; j < p->ieot_start + p->ieot_blocks; j++)\n+\t\t\tnr += p->ieot->entries[j].nr;\n+\t\tif (istate->version == 4) {\n+\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\testimate_cache_size_from_compressed(nr));\n+\t\t} else {\n+\t\t\tmem_pool_init(&p->ce_mem_pool,\n+\t\t\t\testimate_cache_size(mmap_size, nr));\n+\t\t}\n+\n+\t\terr = pthread_create(&p->pthread, NULL, load_cache_entries_thread, p);\n+\t\tif (err)\n+\t\t\tdie(_(\"unable to create load_cache_entries thread: %s\"), strerror(err));\n+\n+\t\t/* increment by the number of cache entries in the ieot block being processed */\n+\t\tfor (j = 0; j < ieot_blocks; j++)\n+\t\t\toffset += ieot->entries[ieot_start + j].nr;\n+\t\tieot_start += ieot_blocks;\n+\t}\n+\n+\tfor (i = 0; i < nr_threads; i++) {\n+\t\tstruct load_cache_entries_thread_data *p = &data[i];\n+\n+\t\terr = pthread_join(p->pthread, NULL);\n+\t\tif (err)\n+\t\t\tdie(_(\"unable to join load_cache_entries thread: %s\"), strerror(err));\n+\t\tmem_pool_combine(istate->ce_mem_pool, p->ce_mem_pool);\n+\t\tconsumed += p->consumed;\n+\t}\n+\n+\tfree(data);\n+\n+\treturn consumed;\n+}\n+#endif\n+\n /* remember to discard_cache() before reading a different cache! */\n int do_read_index(struct index_state *istate, const char *path, int must_exist)\n {\n-\tint fd, i;\n+\tint fd;\n \tstruct stat st;\n \tunsigned long src_offset;\n \tconst struct cache_header *hdr;\n \tconst char *mmap;\n \tsize_t mmap_size;\n-\tconst struct cache_entry *previous_ce = NULL;\n \tstruct load_index_extensions p;\n \tsize_t extension_offset = 0;\n #ifndef NO_PTHREADS\n-\tint nr_threads;\n+\tint nr_threads, cpus;\n+\tstruct index_entry_offset_table *ieot = NULL;\n #endif\n \n \tif (istate->initialized)\n@@ -2012,10 +2163,18 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \tp.mmap = mmap;\n \tp.mmap_size = mmap_size;\n \n+\tsrc_offset = sizeof(*hdr);\n+\n #ifndef NO_PTHREADS\n \tnr_threads = git_config_get_index_threads();\n-\tif (!nr_threads)\n-\t\tnr_threads = online_cpus();\n+\n+\t/* TODO: does creating more threads than cores help? */\n+\tif (!nr_threads) {\n+\t\tnr_threads = istate->cache_nr / THREAD_COST;\n+\t\tcpus = online_cpus();\n+\t\tif (nr_threads > cpus)\n+\t\t\tnr_threads = cpus;\n+\t}\n \n \tif (nr_threads > 1) {\n \t\textension_offset = read_eoie_extension(mmap, mmap_size);\n@@ -2030,29 +2189,24 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \t\t\tnr_threads--;\n \t\t}\n \t}\n-#endif\n \n-\tif (istate->version == 4) {\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size_from_compressed(istate->cache_nr));\n+\t/*\n+\t * Locate and read the index entry offset table so that we can use it\n+\t * to multi-thread the reading of the cache entries.\n+\t */\n+\tif (extension_offset && nr_threads > 1)\n+\t\tieot = read_ieot_extension(mmap, mmap_size, extension_offset);\n+\n+\tif (ieot) {\n+\t\tsrc_offset += load_cache_entries_threaded(istate, mmap, mmap_size, src_offset, nr_threads, ieot);\n+\t\tfree(ieot);\n \t} else {\n-\t\tmem_pool_init(&istate->ce_mem_pool,\n-\t\t\t      estimate_cache_size(mmap_size, istate->cache_nr));\n+\t\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n \t}\n+#else\n+\tsrc_offset += load_all_cache_entries(istate, mmap, mmap_size, src_offset);\n+#endif\n \n-\tsrc_offset = sizeof(*hdr);\n-\tfor (i = 0; i < istate->cache_nr; i++) {\n-\t\tstruct ondisk_cache_entry *disk_ce;\n-\t\tstruct cache_entry *ce;\n-\t\tunsigned long consumed;\n-\n-\t\tdisk_ce = (struct ondisk_cache_entry *)(mmap + src_offset);\n-\t\tce = create_from_disk(istate, disk_ce, &consumed, previous_ce);\n-\t\tset_index_entry(istate, i, ce);\n-\n-\t\tsrc_offset += consumed;\n-\t\tprevious_ce = ce;\n-\t}\n \tistate->timestamp.sec = st.st_mtime;\n \tistate->timestamp.nsec = ST_MTIME_NSEC(st);\n \n@@ -2549,7 +2703,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tstruct strbuf previous_name_buf = STRBUF_INIT, *previous_name;\n \tint drop_cache_tree = istate->drop_cache_tree;\n \toff_t offset;\n-\tint ieot_blocks = 1;\n+\tint ieot_entries = 1;\n \tstruct index_entry_offset_table *ieot = NULL;\n \tint nr, nr_threads;\n \n@@ -2602,6 +2756,8 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\t\t\tieot_blocks = cpus - 1;\n \t\t} else {\n \t\t\tieot_blocks = nr_threads;\n+\t\t\tif (ieot_blocks > istate->cache_nr)\n+\t\t\t\tieot_blocks = istate->cache_nr;\n \t\t}\n \n \t\t/*\n@@ -2611,7 +2767,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t\tif (ieot_blocks > 1) {\n \t\t\tieot = xcalloc(1, sizeof(struct index_entry_offset_table)\n \t\t\t\t+ (ieot_blocks * sizeof(struct index_entry_offset)));\n-\t\t\tieot_blocks = DIV_ROUND_UP(entries, ieot_blocks);\n+\t\t\tieot_entries = DIV_ROUND_UP(entries, ieot_blocks);\n \t\t}\n \t}\n #endif\n@@ -2644,7 +2800,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \n \t\t\tdrop_cache_tree = 1;\n \t\t}\n-\t\tif (ieot && i && (i % ieot_blocks == 0)) {\n+\t\tif (ieot && i && (i % ieot_entries == 0)) {\n \t\t\tieot->entries[ieot->nr].nr = nr;\n \t\t\tieot->entries[ieot->nr].offset = offset;\n \t\t\tieot->nr++;\n-- \n2.18.0.windows.1\n\n"},{"id":"360267","messageId":"xmqqwoqndf8c.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20181010155938.20996-1-peartben@gmail.com","subject":"Re: [PATCH v8 0/7] speed up index load through parallelization","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-10-12T03:18:59Z","receivedAt":"2018-10-12T03:19:03Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <peartben@gmail.com> writes:\n\n> From: Ben Peart <benpeart@microsoft.com>\n>\n> Fixed issues identified in review the most impactful probably being plugging\n> some leaks and improved error handling.  Also added better error messages\n> and some code cleanup to code I'd touched.\n>\n> The biggest change in the interdiff is the impact of renaming ieot_offset to\n> ieot_start and ieot_work to ieot_blocks in hopes of making it easier to read\n> and understand the code.\n\nThanks, I think this one is ready to be in 'next' and any further\ntweaks can be done incrementally.\n\n"},{"id":"360406","messageId":"CACsJy8CyG0DWPyq5cSUteFUiz1ZCpmmVFjYjt8Gxm3Hnvd5q9g@mail.gmail.com","threadId":"49204","inReplyTo":"20181010155938.20996-1-peartben@gmail.com","subject":"Re: [PATCH v8 0/7] speed up index load through parallelization","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-10-14T12:28:05Z","receivedAt":"2018-10-14T12:30:14Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Oct 10, 2018 at 5:59 PM Ben Peart <peartben@gmail.com> wrote:\n> @@ -3460,14 +3479,18 @@ static struct index_entry_offset_table *read_ieot_extension(const char *mmap, si\n>\n>         /* validate the version is IEOT_VERSION */\n>         ext_version = get_be32(index);\n> -       if (ext_version != IEOT_VERSION)\n> +       if (ext_version != IEOT_VERSION) {\n> +              error(\"invalid IEOT version %d\", ext_version);\n\nPlease wrap this string in _() so that it can be translated.\n\n>                return NULL;\n> +       }\n>         index += sizeof(uint32_t);\n>\n>         /* extension size - version bytes / bytes per entry */\n>         nr = (extsize - sizeof(uint32_t)) / (sizeof(uint32_t) + sizeof(uint32_t));\n> -       if (!nr)\n> +       if (!nr) {\n> +              error(\"invalid number of IEOT entries %d\", nr);\n\nDitto. And reporting extsize may be more useful than nr, which we know\nis zero, but we don't know why it's calculated zero unless we know\nextsize.\n-- \nDuy\n"},{"id":"360548","messageId":"3a7fb135-95e3-9dab-ca31-8daadc5cf80c@gmail.com","threadId":"49204","inReplyTo":"CACsJy8CyG0DWPyq5cSUteFUiz1ZCpmmVFjYjt8Gxm3Hnvd5q9g@mail.gmail.com","subject":"Re: [PATCH v8 0/7] speed up index load through parallelization","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-15T17:33:32Z","receivedAt":"2018-10-15T17:33:38Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"fixup! IEOT error messages\n\nEnable localizing new error messages and improve the error message for\ninvalid IEOT extension sizes.\n\nSigned-off-by: Ben Peart <benpeart@microsoft.com>\n---\n  read-cache.c | 4 ++--\n  1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 7acc2c86f4..f9fa6a7979 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -3480,7 +3480,7 @@ static struct index_entry_offset_table \n*read_ieot_extension(const char *mmap, si\n         /* validate the version is IEOT_VERSION */\n         ext_version = get_be32(index);\n         if (ext_version != IEOT_VERSION) {\n-\t       error(\"invalid IEOT version %d\", ext_version);\n+\t       error(_(\"invalid IEOT version %d\"), ext_version);\n  \t       return NULL;\n         }\n         index += sizeof(uint32_t);\n@@ -3488,7 +3488,7 @@ static struct index_entry_offset_table \n*read_ieot_extension(const char *mmap, si\n         /* extension size - version bytes / bytes per entry */\n         nr = (extsize - sizeof(uint32_t)) / (sizeof(uint32_t) + \nsizeof(uint32_t));\n         if (!nr) {\n-\t       error(\"invalid number of IEOT entries %d\", nr);\n+\t       error(_(\"invalid IEOT extension size %d\"), extsize);\n  \t       return NULL;\n         }\n         ieot = xmalloc(sizeof(struct index_entry_offset_table)\n-- \n2.18.0.windows.1\n\n\n\nOn 10/14/2018 8:28 AM, Duy Nguyen wrote:\n> On Wed, Oct 10, 2018 at 5:59 PM Ben Peart <peartben@gmail.com> wrote:\n>> @@ -3460,14 +3479,18 @@ static struct index_entry_offset_table *read_ieot_extension(const char *mmap, si\n>>\n>>          /* validate the version is IEOT_VERSION */\n>>          ext_version = get_be32(index);\n>> -       if (ext_version != IEOT_VERSION)\n>> +       if (ext_version != IEOT_VERSION) {\n>> +              error(\"invalid IEOT version %d\", ext_version);\n> \n> Please wrap this string in _() so that it can be translated.\n> \n>>                 return NULL;\n>> +       }\n>>          index += sizeof(uint32_t);\n>>\n>>          /* extension size - version bytes / bytes per entry */\n>>          nr = (extsize - sizeof(uint32_t)) / (sizeof(uint32_t) + sizeof(uint32_t));\n>> -       if (!nr)\n>> +       if (!nr) {\n>> +              error(\"invalid number of IEOT entries %d\", nr);\n> \n> Ditto. And reporting extsize may be more useful than nr, which we know\n> is zero, but we don't know why it's calculated zero unless we know\n> extsize.\n> \n"},{"id":"360938","messageId":"20181019161118.GA8100@sigill.intra.peff.net","threadId":"49204","inReplyTo":"20181010155938.20996-8-peartben@gmail.com","subject":"Re: [PATCH v8 7/7] read-cache: load cache entries on worker threads","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-10-19T16:11:19Z","receivedAt":"2018-10-19T16:11:22Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Oct 10, 2018 at 11:59:38AM -0400, Ben Peart wrote:\n\n> +static unsigned long load_cache_entries_threaded(struct index_state *istate, const char *mmap, size_t mmap_size,\n> +\t\t\tunsigned long src_offset, int nr_threads, struct index_entry_offset_table *ieot)\n\nThe src_offset parameter isn't used in this function.\n\nIn early versions of the series, it was used to feed the p->start_offset\nfield of each load_cache_entries_thread_data. But after the switch to\nieot, we don't, and instead feed p->ieot_start. But we always begin that\nat 0.\n\nIs that right (and we can drop the parameter), or should this logic:\n\n> +\toffset = ieot_start = 0;\n> +\tieot_blocks = DIV_ROUND_UP(ieot->nr, nr_threads);\n> +\tfor (i = 0; i < nr_threads; i++) {\n> [...]\n\nbe starting at src_offset instead of 0?\n\n-Peff\n"},{"id":"361109","messageId":"xmqqsh0yybg7.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20181019161118.GA8100@sigill.intra.peff.net","subject":"Re: [PATCH v8 7/7] read-cache: load cache entries on worker threads","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-10-22T02:14:32Z","receivedAt":"2018-10-22T02:14:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Wed, Oct 10, 2018 at 11:59:38AM -0400, Ben Peart wrote:\n>\n>> +static unsigned long load_cache_entries_threaded(struct index_state *istate, const char *mmap, size_t mmap_size,\n>> +\t\t\tunsigned long src_offset, int nr_threads, struct index_entry_offset_table *ieot)\n>\n> The src_offset parameter isn't used in this function.\n>\n> In early versions of the series, it was used to feed the p->start_offset\n> field of each load_cache_entries_thread_data. But after the switch to\n> ieot, we don't, and instead feed p->ieot_start. But we always begin that\n> at 0.\n>\n> Is that right (and we can drop the parameter), or should this logic:\n>\n>> +\toffset = ieot_start = 0;\n>> +\tieot_blocks = DIV_ROUND_UP(ieot->nr, nr_threads);\n>> +\tfor (i = 0; i < nr_threads; i++) {\n>> [...]\n>\n> be starting at src_offset instead of 0?\n\nI think \"offset\" has nothing to do with the offset into the mmapped\nregion of memory.  It is an integer index into a (virtual) array\nthat is a concatenation of ieot->entries[].entries[], and it is\ncorrect to count from zero.  The value taken from that array using\nthe index is used to compute the offset into the mmapped region.\n\nUnlike load_all_cache_entries() called from the other side of the\nsame if() statement in the same caller, this does not depend on the\nfact that the first index entry in the mmapped region appears\nimmediately after the index-file header.  It goes from the offsets\ninto the file that are recorded in the entry offset table that is an\nindex extension, so the sizeof(*hdr) that initializes src_offset is\nnot used by the codepath.\n\nThe number of bytes consumed, i.e. its return value from the\nfunction, is not really used, either, as the caller does not use\nsrc_offset for anything other than updating it with \"+=\" and passing\nit to this function (which does not use it) when it calls this\nfunction (i.e. when ieot extension exists--and by definition when\nthat extension exists extension_offset is not 0, so we do not make\nthe final load_index_extensions() call in the caller that uses\nsrc_offset).\n"},{"id":"361163","messageId":"72c8938c-7ffa-a384-55dc-09d90d9cf7f8@gmail.com","threadId":"49204","inReplyTo":"xmqqsh0yybg7.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v8 7/7] read-cache: load cache entries on worker threads","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-10-22T14:40:00Z","receivedAt":"2018-10-22T14:40:06Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 10/21/2018 10:14 PM, Junio C Hamano wrote:\n> Jeff King <peff@peff.net> writes:\n> \n>> On Wed, Oct 10, 2018 at 11:59:38AM -0400, Ben Peart wrote:\n>>\n>>> +static unsigned long load_cache_entries_threaded(struct index_state *istate, const char *mmap, size_t mmap_size,\n>>> +\t\t\tunsigned long src_offset, int nr_threads, struct index_entry_offset_table *ieot)\n>>\n>> The src_offset parameter isn't used in this function.\n>>\n>> In early versions of the series, it was used to feed the p->start_offset\n>> field of each load_cache_entries_thread_data. But after the switch to\n>> ieot, we don't, and instead feed p->ieot_start. But we always begin that\n>> at 0.\n>>\n>> Is that right (and we can drop the parameter), or should this logic:\n>>\n>>> +\toffset = ieot_start = 0;\n>>> +\tieot_blocks = DIV_ROUND_UP(ieot->nr, nr_threads);\n>>> +\tfor (i = 0; i < nr_threads; i++) {\n>>> [...]\n>>\n>> be starting at src_offset instead of 0?\n> \n> I think \"offset\" has nothing to do with the offset into the mmapped\n> region of memory.  It is an integer index into a (virtual) array\n> that is a concatenation of ieot->entries[].entries[], and it is\n> correct to count from zero.  The value taken from that array using\n> the index is used to compute the offset into the mmapped region.\n> \n> Unlike load_all_cache_entries() called from the other side of the\n> same if() statement in the same caller, this does not depend on the\n> fact that the first index entry in the mmapped region appears\n> immediately after the index-file header.  It goes from the offsets\n> into the file that are recorded in the entry offset table that is an\n> index extension, so the sizeof(*hdr) that initializes src_offset is\n> not used by the codepath.\n> \n> The number of bytes consumed, i.e. its return value from the\n> function, is not really used, either, as the caller does not use\n> src_offset for anything other than updating it with \"+=\" and passing\n> it to this function (which does not use it) when it calls this\n> function (i.e. when ieot extension exists--and by definition when\n> that extension exists extension_offset is not 0, so we do not make\n> the final load_index_extensions() call in the caller that uses\n> src_offset).\n> \n\nThanks for discovering/analyzing this.  You're right, I missed removing \nthis when we switched from a single offset to an array of offsets via \nthe IEOT.  I'll send a patch to fix both issues shortly.\n"},{"id":"363106","messageId":"20181113003817.GA170017@google.com","threadId":"49204","inReplyTo":"20181010155938.20996-1-peartben@gmail.com","subject":"[PATCH 0/3] Avoid confusing messages from new index extensions (Re: [PATCH v8 0/7] speed up index load through parallelization)","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-13T00:38:17Z","receivedAt":"2018-11-13T00:38:23Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nBen Peart wrote:\n\n> Ben Peart (6):\n>   read-cache: clean up casting and byte decoding\n>   eoie: add End of Index Entry (EOIE) extension\n>   config: add new index.threads config setting\n>   read-cache: load cache extensions on a worker thread\n>   ieot: add Index Entry Offset Table (IEOT) extension\n>   read-cache: load cache entries on worker threads\n\nI love this, but when deploying it I ran into a problem.\n\nHow about these patches?\n\nThanks,\nJonathan Nieder (3):\n  eoie: default to not writing EOIE section\n  ieot: default to not writing IEOT section\n  index: do not warn about unrecognized extensions\n\n Documentation/config.txt | 14 ++++++++++++++\n read-cache.c             | 24 +++++++++++++++++++++---\n t/t1700-split-index.sh   | 11 +++++++----\n 3 files changed, 42 insertions(+), 7 deletions(-)\n"},{"id":"363107","messageId":"20181113003911.GB170017@google.com","threadId":"49204","inReplyTo":"20181113003817.GA170017@google.com","subject":"[PATCH 1/3] eoie: default to not writing EOIE section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-13T00:39:11Z","receivedAt":"2018-11-13T00:39:16Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Since 3b1d9e04 (eoie: add End of Index Entry (EOIE) extension,\n2018-10-10) Git defaults to writing the new EOIE section when writing\nout an index file.  Usually that is a good thing because it improves\nthreaded performance, but when a Git repository is shared with older\nversions of Git, it produces a confusing warning:\n\n  $ git status\n  ignoring EOIE extension\n  HEAD detached at 371ed0defa\n  nothing to commit, working tree clean\n\nLet's introduce the new index extension more gently.  First we'll roll\nout the new version of Git that understands it, and then once\nsufficiently many users are using such a version, we can flip the\ndefault to writing it by default.\n\nIntroduce a '[index] recordEndOfIndexEntries' configuration variable\nto allow interested users to benefit from this index extension early.\n\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n---\n Documentation/config.txt |  7 +++++++\n read-cache.c             | 11 ++++++++++-\n t/t1700-split-index.sh   | 11 +++++++----\n 3 files changed, 24 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex 41a9ff2b6a..d702379db4 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -2188,6 +2188,13 @@ imap::\n \tThe configuration variables in the 'imap' section are described\n \tin linkgit:git-imap-send[1].\n \n+index.recordEndOfIndexEntries::\n+\tSpecifies whether the index file should include an \"End Of Index\n+\tEntry\" section. This reduces index load time on multiprocessor\n+\tmachines but produces a message \"ignoring EOIE extension\" when\n+\treading the index using Git versions before 2.20. Defaults to\n+\t'false'.\n+\n index.threads::\n \tSpecifies the number of threads to spawn when loading the index.\n \tThis is meant to reduce index load time on multiprocessor machines.\ndiff --git a/read-cache.c b/read-cache.c\nindex f3a848d61c..4bfe93c4c2 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2698,6 +2698,15 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n \t\trollback_lock_file(lockfile);\n }\n \n+static int record_eoie(void)\n+{\n+\tint val;\n+\n+\tif (!git_config_get_bool(\"index.recordendofindexentries\", &val))\n+\t\treturn val;\n+\treturn 0;\n+}\n+\n /*\n  * On success, `tempfile` is closed. If it is the temporary file\n  * of a `struct lock_file`, we will therefore effectively perform\n@@ -2945,7 +2954,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t * read.  Write it out regardless of the strip_extensions parameter as we need it\n \t * when loading the shared index.\n \t */\n-\tif (offset) {\n+\tif (offset && record_eoie()) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_eoie_extension(&sb, &eoie_c, offset);\ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex 2ac47aa0e4..0cbac64e28 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -25,14 +25,17 @@ test_expect_success 'enable split index' '\n \tgit update-index --split-index &&\n \ttest-tool dump-split-index .git/index >actual &&\n \tindexversion=$(test-tool index-version <.git/index) &&\n+\n+\t# NEEDSWORK: Stop hard-coding checksums.\n \tif test \"$indexversion\" = \"4\"\n \tthen\n-\t\town=3527df833c6c100d3d1d921a9a782d62a8be4b58\n-\t\tbase=746f7ab2ed44fb839efdfbffcf399d0b113fb4cb\n+\t\town=432ef4b63f32193984f339431fd50ca796493569\n+\t\tbase=508851a7f0dfa8691e9f69c7f055865389012491\n \telse\n-\t\town=5e9b60117ece18da410ddecc8b8d43766a0e4204\n-\t\tbase=4370042739b31cd17a5c5cd6043a77c9a00df113\n+\t\town=8299b0bcd1ac364e5f1d7768efb62fa2da79a339\n+\t\tbase=39d890139ee5356c7ef572216cebcd27aa41f9df\n \tfi &&\n+\n \tcat >expect <<-EOF &&\n \town $own\n \tbase $base\n-- \n2.19.1.930.g4563a0d9d0\n\n"},{"id":"363108","messageId":"20181113003938.GC170017@google.com","threadId":"49204","inReplyTo":"20181113003817.GA170017@google.com","subject":"[PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-13T00:39:38Z","receivedAt":"2018-11-13T00:39:43Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"As with EOIE, popular versions of Git do not support the new IEOT\nextension yet.  When accessing a Git repository written by a more\nmodern version of Git, they correctly ignore the unrecognized section,\nbut in the process they loudly warn\n\n\tignoring IEOT extension\n\nresulting in confusion for users.  Introduce the index extension more\ngently by not writing it yet in this first version with support for\nit.  Soon, once sufficiently many users are running a modern version\nof Git, we can flip the default so users benefit from this index\nextension by default.\n\nIntroduce a '[index] recordOffsetTable' configuration variable to\ncontrol whether the new index extension is written.\n\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n---\n Documentation/config.txt |  7 +++++++\n read-cache.c             | 11 ++++++++++-\n 2 files changed, 17 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex d702379db4..cc66fb7de3 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -2195,6 +2195,13 @@ index.recordEndOfIndexEntries::\n \treading the index using Git versions before 2.20. Defaults to\n \t'false'.\n \n+index.recordOffsetTable::\n+\tSpecifies whether the index file should include an \"Index Entry\n+\tOffset Table\" section. This reduces index load time on\n+\tmultiprocessor machines but produces a message \"ignoring IEOT\n+\textension\" when reading the index using Git versions before 2.20.\n+\tDefaults to 'false'.\n+\n index.threads::\n \tSpecifies the number of threads to spawn when loading the index.\n \tThis is meant to reduce index load time on multiprocessor machines.\ndiff --git a/read-cache.c b/read-cache.c\nindex 4bfe93c4c2..290bd54708 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2707,6 +2707,15 @@ static int record_eoie(void)\n \treturn 0;\n }\n \n+static int record_ieot(void)\n+{\n+\tint val;\n+\n+\tif (!git_config_get_bool(\"index.recordoffsettable\", &val))\n+\t\treturn val;\n+\treturn 0;\n+}\n+\n /*\n  * On success, `tempfile` is closed. If it is the temporary file\n  * of a `struct lock_file`, we will therefore effectively perform\n@@ -2767,7 +2776,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \n #ifndef NO_PTHREADS\n \tnr_threads = git_config_get_index_threads();\n-\tif (nr_threads != 1) {\n+\tif (nr_threads != 1 && record_ieot()) {\n \t\tint ieot_blocks, cpus;\n \n \t\t/*\n-- \n2.19.1.930.g4563a0d9d0\n\n"},{"id":"363110","messageId":"20181113004019.GD170017@google.com","threadId":"49204","inReplyTo":"20181113003817.GA170017@google.com","subject":"[PATCH 3/3] index: do not warn about unrecognized extensions","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-13T00:40:19Z","receivedAt":"2018-11-13T00:40:24Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Documentation/technical/index-format explains:\n\n     4-byte extension signature. If the first byte is 'A'..'Z' the\n     extension is optional and can be ignored.\n\nThis allows gracefully introducing a new index extension without\nhaving to rely on all readers having support for it.  Mandatory\nextensions start with a lowercase letter and optional ones start with\na capital.  Thus the versions of Git acting on a shared local\nrepository do not have to upgrade in lockstep.\n\nWe almost obey that convention, but there is a problem: when\nencountering an unrecognized optional extension, we write\n\n\tignoring FNCY extension\n\nto stderr, which alarms users.  This means that in practice we have\nhad to introduce index extensions in two steps: first add read\nsupport, and then a while later, start writing by default.  This\ndelays when users can benefit from improvements to the index format.\n\nWe cannot change the past, but for index extensions of the future,\nthere is a straightforward improvement: silence that message except\nwhen tracing.  This way, the message is still available when\ndebugging, but in everyday use it does not show up so (once most Git\nusers have this patch) we can turn on new optional extensions right\naway without alarming people.\n\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n---\nThanks for reading.  Thoughts?\n\nSincerely,\nJonathan\n\n read-cache.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex 290bd54708..65530a68c2 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1720,7 +1720,7 @@ static int read_index_extension(struct index_state *istate,\n \t\tif (*ext < 'A' || 'Z' < *ext)\n \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n \t\t\t\t     ext);\n-\t\tfprintf(stderr, \"ignoring %.4s extension\\n\", ext);\n+\t\ttrace_printf(\"ignoring %.4s extension\\n\", ext);\n \t\tbreak;\n \t}\n \treturn 0;\n-- \n2.19.1.930.g4563a0d9d0\n\n"},{"id":"363111","messageId":"20181113005806.128469-1-jonathantanmy@google.com","threadId":"49204","inReplyTo":"20181113003938.GC170017@google.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2018-11-13T00:58:06Z","receivedAt":"2018-11-13T00:58:21Z","isPatch":true,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> +index.recordOffsetTable::\n> +\tSpecifies whether the index file should include an \"Index Entry\n> +\tOffset Table\" section. This reduces index load time on\n> +\tmultiprocessor machines but produces a message \"ignoring IEOT\n> +\textension\" when reading the index using Git versions before 2.20.\n> +\tDefaults to 'false'.\n\nProbably worth adding a test that exercises this new config option -\nsomehow create an index with index.recordOffsetTable=1, check that the\nindex contains the appropriate string (a few ways to do this, but I'm\nnot sure which are portable), and then run a Git command that reads the\nindex to make sure it is valid; then do the same except\nindex.recordOffsetTable=0.\n\nThe code itself looks good to me.\n\nSame comment for patch 1.\n"},{"id":"363112","messageId":"xmqqtvklzszv.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20181113003911.GB170017@google.com","subject":"Re: [PATCH 1/3] eoie: default to not writing EOIE section","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-13T01:05:56Z","receivedAt":"2018-11-13T01:06:06Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> Since 3b1d9e04 (eoie: add End of Index Entry (EOIE) extension,\n> 2018-10-10) Git defaults to writing the new EOIE section when writing\n> out an index file.  Usually that is a good thing because it improves\n> threaded performance, but when a Git repository is shared with older\n> versions of Git, it produces a confusing warning:\n>\n>   $ git status\n>   ignoring EOIE extension\n>   HEAD detached at 371ed0defa\n>   nothing to commit, working tree clean\n>\n> Let's introduce the new index extension more gently.  First we'll roll\n> out the new version of Git that understands it, and then once\n> sufficiently many users are using such a version, we can flip the\n> default to writing it by default.\n>\n> Introduce a '[index] recordEndOfIndexEntries' configuration variable\n> to allow interested users to benefit from this index extension early.\n\nThanks.  I am in principle OK with this approach.  In fact, I\nsuspect that the default may want to be dynamically determined, and\nwe give this knob to let the users further force their preference.\nWhen no extension that benefits from multi-threading is written, the\ndefault can stay \"no\" in future versions of Git, for example.\n\n> diff --git a/Documentation/config.txt b/Documentation/config.txt\n> index 41a9ff2b6a..d702379db4 100644\n\nThe timing is a bit unfortunate for any topic to touch this file,\nand contrib/cocci would not help us in this case X-<.\n\n> diff --git a/read-cache.c b/read-cache.c\n> index f3a848d61c..4bfe93c4c2 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -2698,6 +2698,15 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n>  \t\trollback_lock_file(lockfile);\n>  }\n>  \n> +static int record_eoie(void)\n> +{\n> +\tint val;\n> +\n> +\tif (!git_config_get_bool(\"index.recordendofindexentries\", &val))\n> +\t\treturn val;\n> +\treturn 0;\n> +}\n\nUnconditionally defaulting to no in this round is perfectly fine.\nLet's make a mental note that this is the place to decide dynamic\ndefault in the future when we want to.  It would probably have to\nask around various \"extension writing\" helpers if they want to have\na say in the outcome (e.g. if there are very many cache entries in\nthe istate, the entry offset table may want to be written and\notherwise not).\n\n> @@ -2945,7 +2954,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>  \t * read.  Write it out regardless of the strip_extensions parameter as we need it\n>  \t * when loading the shared index.\n>  \t */\n> -\tif (offset) {\n> +\tif (offset && record_eoie()) {\n>  \t\tstruct strbuf sb = STRBUF_INIT;\n>  \n>  \t\twrite_eoie_extension(&sb, &eoie_c, offset);\n> diff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\n> index 2ac47aa0e4..0cbac64e28 100755\n> --- a/t/t1700-split-index.sh\n> +++ b/t/t1700-split-index.sh\n> @@ -25,14 +25,17 @@ test_expect_success 'enable split index' '\n>  \tgit update-index --split-index &&\n>  \ttest-tool dump-split-index .git/index >actual &&\n>  \tindexversion=$(test-tool index-version <.git/index) &&\n> +\n> +\t# NEEDSWORK: Stop hard-coding checksums.\n\nAlso let's stop hard-coding the assumption that the new knob is off\nby default.  Ideally, you'd want to test both cases, right?\n\nPerhaps you'd call \"git update-index --split-index\" we see in the\nprecontext twice, with \"-c VAR=false\" and \"-c VAR=true\", to prepare\n\"actual.without-eoie\" and \"actual.with-eoie\", or something like\nthat?\n\nThanks.\n\n>  \tif test \"$indexversion\" = \"4\"\n>  \tthen\n> -\t\town=3527df833c6c100d3d1d921a9a782d62a8be4b58\n> -\t\tbase=746f7ab2ed44fb839efdfbffcf399d0b113fb4cb\n> +\t\town=432ef4b63f32193984f339431fd50ca796493569\n> +\t\tbase=508851a7f0dfa8691e9f69c7f055865389012491\n>  \telse\n> -\t\town=5e9b60117ece18da410ddecc8b8d43766a0e4204\n> -\t\tbase=4370042739b31cd17a5c5cd6043a77c9a00df113\n> +\t\town=8299b0bcd1ac364e5f1d7768efb62fa2da79a339\n> +\t\tbase=39d890139ee5356c7ef572216cebcd27aa41f9df\n>  \tfi &&\n> +\n>  \tcat >expect <<-EOF &&\n>  \town $own\n>  \tbase $base\n"},{"id":"363113","messageId":"xmqqpnv9zsu6.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20181113003938.GC170017@google.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-13T01:09:21Z","receivedAt":"2018-11-13T01:09:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> As with EOIE, popular versions of Git do not support the new IEOT\n> extension yet.  When accessing a Git repository written by a more\n> modern version of Git, they correctly ignore the unrecognized section,\n> but in the process they loudly warn\n>\n> \tignoring IEOT extension\n>\n> resulting in confusion for users.\n\nThen removing the message is throwing it with bathwater.  First\nthink about which part of the message is confusiong and then make it\nless confusing.\n\nHow about\n\n\thint: ignoring an optional IEOT extension\n\nto make it clear that it is totally harmless?\n\nWith that, we can add advise.unknownIndexExtension=false to turn all\nof them off with a single switch.\n\n\n"},{"id":"363114","messageId":"xmqqlg5xzssq.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20181113004019.GD170017@google.com","subject":"Re: [PATCH 3/3] index: do not warn about unrecognized extensions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-13T01:10:13Z","receivedAt":"2018-11-13T01:10:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> We almost obey that convention, but there is a problem: when\n> encountering an unrecognized optional extension, we write\n>\n> \tignoring FNCY extension\n>\n> to stderr, which alarms users.\n\nThen the same comment as 2/3 applies to this step.\n"},{"id":"363115","messageId":"20181113011207.GE170017@google.com","threadId":"49204","inReplyTo":"xmqqpnv9zsu6.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-13T01:12:07Z","receivedAt":"2018-11-13T01:12:12Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n\n> How about\n>\n> \thint: ignoring an optional IEOT extension\n>\n> to make it clear that it is totally harmless?\n>\n> With that, we can add advise.unknownIndexExtension=false to turn all\n> of them off with a single switch.\n\nI like it.  Expect a patch soon (tonight or tomorrow) that does that.\n\nWe'll have to find some appropriate place in the documentation to\nexplain what the message is about, still.\n\nThanks,\nJonathan\n"},{"id":"363165","messageId":"5fae19dc-2e77-1211-0086-e7aa9d30562f@gmail.com","threadId":"49204","inReplyTo":"xmqqtvklzszv.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 1/3] eoie: default to not writing EOIE section","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-13T15:14:26Z","receivedAt":"2018-11-13T15:14:31Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/12/2018 8:05 PM, Junio C Hamano wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n> \n>> Since 3b1d9e04 (eoie: add End of Index Entry (EOIE) extension,\n>> 2018-10-10) Git defaults to writing the new EOIE section when writing\n>> out an index file.  Usually that is a good thing because it improves\n>> threaded performance, but when a Git repository is shared with older\n>> versions of Git, it produces a confusing warning:\n>>\n>>    $ git status\n>>    ignoring EOIE extension\n>>    HEAD detached at 371ed0defa\n>>    nothing to commit, working tree clean\n>>\n>> Let's introduce the new index extension more gently.  First we'll roll\n>> out the new version of Git that understands it, and then once\n>> sufficiently many users are using such a version, we can flip the\n>> default to writing it by default.\n>>\n>> Introduce a '[index] recordEndOfIndexEntries' configuration variable\n>> to allow interested users to benefit from this index extension early.\n> \n> Thanks.  I am in principle OK with this approach.  In fact, I\n> suspect that the default may want to be dynamically determined, and\n> we give this knob to let the users further force their preference.\n> When no extension that benefits from multi-threading is written, the\n> default can stay \"no\" in future versions of Git, for example.\n> \n\nWhile I can understand the user confusion the warning about ignoring an \nextension could cause I guess I'm a little surprised that people would \nsee it that often.  To see the warning means they are running a new \nversion of git in the same repo as they are running an old version of \ngit.  I just haven't ever experienced that (I only ever have one copy of \ngit installed) so am surprised it comes up often enough to warrant this \nchange.\n\nThat said, if it _is_ that much of an issue, this patch makes sense and \nprovides a way to more gracefully transition into this feature.  Even if \nwe had some logic to dynamically determine whether to write it or not, \nwe'd still want to avoid confusing users when it did get written out.\n\n>> diff --git a/Documentation/config.txt b/Documentation/config.txt\n>> index 41a9ff2b6a..d702379db4 100644\n> \n> The timing is a bit unfortunate for any topic to touch this file,\n> and contrib/cocci would not help us in this case X-<.\n> \n>> diff --git a/read-cache.c b/read-cache.c\n>> index f3a848d61c..4bfe93c4c2 100644\n>> --- a/read-cache.c\n>> +++ b/read-cache.c\n>> @@ -2698,6 +2698,15 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n>>   \t\trollback_lock_file(lockfile);\n>>   }\n>>   \n>> +static int record_eoie(void)\n>> +{\n>> +\tint val;\n>> +\n>> +\tif (!git_config_get_bool(\"index.recordendofindexentries\", &val))\n>> +\t\treturn val;\n>> +\treturn 0;\n>> +}\n> \n> Unconditionally defaulting to no in this round is perfectly fine.\n> Let's make a mental note that this is the place to decide dynamic\n> default in the future when we want to.  It would probably have to\n> ask around various \"extension writing\" helpers if they want to have\n> a say in the outcome (e.g. if there are very many cache entries in\n> the istate, the entry offset table may want to be written and\n> otherwise not).\n> \n>> @@ -2945,7 +2954,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>>   \t * read.  Write it out regardless of the strip_extensions parameter as we need it\n>>   \t * when loading the shared index.\n>>   \t */\n>> -\tif (offset) {\n>> +\tif (offset && record_eoie()) {\n>>   \t\tstruct strbuf sb = STRBUF_INIT;\n>>   \n>>   \t\twrite_eoie_extension(&sb, &eoie_c, offset);\n>> diff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\n>> index 2ac47aa0e4..0cbac64e28 100755\n>> --- a/t/t1700-split-index.sh\n>> +++ b/t/t1700-split-index.sh\n>> @@ -25,14 +25,17 @@ test_expect_success 'enable split index' '\n>>   \tgit update-index --split-index &&\n>>   \ttest-tool dump-split-index .git/index >actual &&\n>>   \tindexversion=$(test-tool index-version <.git/index) &&\n>> +\n>> +\t# NEEDSWORK: Stop hard-coding checksums.\n> \n> Also let's stop hard-coding the assumption that the new knob is off\n> by default.  Ideally, you'd want to test both cases, right?\n> \n> Perhaps you'd call \"git update-index --split-index\" we see in the\n> precontext twice, with \"-c VAR=false\" and \"-c VAR=true\", to prepare\n> \"actual.without-eoie\" and \"actual.with-eoie\", or something like\n> that?\n> \n> Thanks.\n> \n>>   \tif test \"$indexversion\" = \"4\"\n>>   \tthen\n>> -\t\town=3527df833c6c100d3d1d921a9a782d62a8be4b58\n>> -\t\tbase=746f7ab2ed44fb839efdfbffcf399d0b113fb4cb\n>> +\t\town=432ef4b63f32193984f339431fd50ca796493569\n>> +\t\tbase=508851a7f0dfa8691e9f69c7f055865389012491\n>>   \telse\n>> -\t\town=5e9b60117ece18da410ddecc8b8d43766a0e4204\n>> -\t\tbase=4370042739b31cd17a5c5cd6043a77c9a00df113\n>> +\t\town=8299b0bcd1ac364e5f1d7768efb62fa2da79a339\n>> +\t\tbase=39d890139ee5356c7ef572216cebcd27aa41f9df\n>>   \tfi &&\n>> +\n>>   \tcat >expect <<-EOF &&\n>>   \town $own\n>>   \tbase $base\n"},{"id":"363168","messageId":"f2f8cec8-d770-a1e9-b5a1-83653575122e@gmail.com","threadId":"49204","inReplyTo":"20181113003938.GC170017@google.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-13T15:22:44Z","receivedAt":"2018-11-13T15:22:50Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/12/2018 7:39 PM, Jonathan Nieder wrote:\n> As with EOIE, popular versions of Git do not support the new IEOT\n> extension yet.  When accessing a Git repository written by a more\n> modern version of Git, they correctly ignore the unrecognized section,\n> but in the process they loudly warn\n> \n> \tignoring IEOT extension\n> \n> resulting in confusion for users.  Introduce the index extension more\n> gently by not writing it yet in this first version with support for\n> it.  Soon, once sufficiently many users are running a modern version\n> of Git, we can flip the default so users benefit from this index\n> extension by default.\n> \n> Introduce a '[index] recordOffsetTable' configuration variable to\n> control whether the new index extension is written.\n> \n\nWhy introduce a new setting to disable writing the IEOT extension \ninstead of just using the existing index.threads setting?  If \nindex.threads=1 then the IEOT extension isn't written which (I believe) \nwill accomplish the same goal.\n\n> Signed-off-by: Jonathan Nieder <jrnieder@gmail.com>\n> ---\n>   Documentation/config.txt |  7 +++++++\n>   read-cache.c             | 11 ++++++++++-\n>   2 files changed, 17 insertions(+), 1 deletion(-)\n> \n> diff --git a/Documentation/config.txt b/Documentation/config.txt\n> index d702379db4..cc66fb7de3 100644\n> --- a/Documentation/config.txt\n> +++ b/Documentation/config.txt\n> @@ -2195,6 +2195,13 @@ index.recordEndOfIndexEntries::\n>   \treading the index using Git versions before 2.20. Defaults to\n>   \t'false'.\n>   \n> +index.recordOffsetTable::\n> +\tSpecifies whether the index file should include an \"Index Entry\n> +\tOffset Table\" section. This reduces index load time on\n> +\tmultiprocessor machines but produces a message \"ignoring IEOT\n> +\textension\" when reading the index using Git versions before 2.20.\n> +\tDefaults to 'false'.\n> +\n>   index.threads::\n>   \tSpecifies the number of threads to spawn when loading the index.\n>   \tThis is meant to reduce index load time on multiprocessor machines.\n> diff --git a/read-cache.c b/read-cache.c\n> index 4bfe93c4c2..290bd54708 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -2707,6 +2707,15 @@ static int record_eoie(void)\n>   \treturn 0;\n>   }\n>   \n> +static int record_ieot(void)\n> +{\n> +\tint val;\n> +\n> +\tif (!git_config_get_bool(\"index.recordoffsettable\", &val))\n> +\t\treturn val;\n> +\treturn 0;\n> +}\n> +\n>   /*\n>    * On success, `tempfile` is closed. If it is the temporary file\n>    * of a `struct lock_file`, we will therefore effectively perform\n> @@ -2767,7 +2776,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>   \n>   #ifndef NO_PTHREADS\n>   \tnr_threads = git_config_get_index_threads();\n> -\tif (nr_threads != 1) {\n> +\tif (nr_threads != 1 && record_ieot()) {\n>   \t\tint ieot_blocks, cpus;\n>   \n>   \t\t/*\n> \n"},{"id":"363169","messageId":"8e7c3e05-ae60-0801-ab2d-5ead02192695@gmail.com","threadId":"49204","inReplyTo":"20181113004019.GD170017@google.com","subject":"Re: [PATCH 3/3] index: do not warn about unrecognized extensions","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-13T15:25:52Z","receivedAt":"2018-11-13T15:25:57Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/12/2018 7:40 PM, Jonathan Nieder wrote:\n> Documentation/technical/index-format explains:\n> \n>       4-byte extension signature. If the first byte is 'A'..'Z' the\n>       extension is optional and can be ignored.\n> \n> This allows gracefully introducing a new index extension without\n> having to rely on all readers having support for it.  Mandatory\n> extensions start with a lowercase letter and optional ones start with\n> a capital.  Thus the versions of Git acting on a shared local\n> repository do not have to upgrade in lockstep.\n> \n> We almost obey that convention, but there is a problem: when\n> encountering an unrecognized optional extension, we write\n> \n> \tignoring FNCY extension\n> \n> to stderr, which alarms users.  This means that in practice we have\n> had to introduce index extensions in two steps: first add read\n> support, and then a while later, start writing by default.  This\n> delays when users can benefit from improvements to the index format.\n> \n> We cannot change the past, but for index extensions of the future,\n> there is a straightforward improvement: silence that message except\n> when tracing.  This way, the message is still available when\n> debugging, but in everyday use it does not show up so (once most Git\n> users have this patch) we can turn on new optional extensions right\n> away without alarming people.\n> \n\nThe best patch of the bunch. Glad to see it.\n\nI'm fine with doing this via advise.unknownIndexExtension as well.  Who \nknows, someone may actually want to see this and not have tracing turned \non.  I don't know who but it is possible :-)\n\n> Signed-off-by: Jonathan Nieder <jrnieder@gmail.com>\n> ---\n> Thanks for reading.  Thoughts?\n> \n> Sincerely,\n> Jonathan\n> \n>   read-cache.c | 2 +-\n>   1 file changed, 1 insertion(+), 1 deletion(-)\n> \n> diff --git a/read-cache.c b/read-cache.c\n> index 290bd54708..65530a68c2 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -1720,7 +1720,7 @@ static int read_index_extension(struct index_state *istate,\n>   \t\tif (*ext < 'A' || 'Z' < *ext)\n>   \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n>   \t\t\t\t     ext);\n> -\t\tfprintf(stderr, \"ignoring %.4s extension\\n\", ext);\n> +\t\ttrace_printf(\"ignoring %.4s extension\\n\", ext);\n>   \t\tbreak;\n>   \t}\n>   \treturn 0;\n> \n"},{"id":"363171","messageId":"CACsJy8DNo1Q96jb5nwJ8vREFu=ZmbN29+uoH17N2VVy80k387Q@mail.gmail.com","threadId":"49204","inReplyTo":"20181113011207.GE170017@google.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2018-11-13T15:37:37Z","receivedAt":"2018-11-13T15:38:07Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, Nov 13, 2018 at 2:12 AM Jonathan Nieder <jrnieder@gmail.com> wrote:\n>\n> Junio C Hamano wrote:\n>\n> > How about\n> >\n> >       hint: ignoring an optional IEOT extension\n> >\n> > to make it clear that it is totally harmless?\n> >\n> > With that, we can add advise.unknownIndexExtension=false to turn all\n> > of them off with a single switch.\n>\n> I like it.  Expect a patch soon (tonight or tomorrow) that does that.\n>\n> We'll have to find some appropriate place in the documentation to\n> explain what the message is about, still.\n\nAlso from the last discussion, if I remember correctly, this\n\"ignoring\" is considered harmless and could be suppressed most of the\ntime. But commands like 'fsck' should always report it.\n-- \nDuy\n"},{"id":"363184","messageId":"20181113180956.GA68106@google.com","threadId":"49204","inReplyTo":"xmqqpnv9zsu6.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-13T18:09:56Z","receivedAt":"2018-11-13T18:10:01Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi again,\n\nJunio C Hamano wrote:\n\n> Then removing the message is throwing it with bathwater.  First\n> think about which part of the message is confusiong and then make it\n> less confusing.\n>\n> How about\n>\n> \thint: ignoring an optional IEOT extension\n>\n> to make it clear that it is totally harmless?\n>\n> With that, we can add advise.unknownIndexExtension=false to turn all\n> of them off with a single switch.\n\nAfter having slept on it, this doesn't seem like a good fit for the\nadvice subsystem.  The advice subsystem provides hints about suggested\nactions for new users to understand what to do about a condition.  In\nthis example, the message is not suggesting a particular user action\n--- instead, it's describing state, which would seem to be a better\nfit for tracing, as in the patch 3/3 I sent.\n\nAm I understanding correclty?  Can you give an example of when a user\nwould *want* to see this message and what they would do in response?\n\nThanks,\nJonathan\n"},{"id":"363185","messageId":"20181113181855.GB68106@google.com","threadId":"49204","inReplyTo":"f2f8cec8-d770-a1e9-b5a1-83653575122e@gmail.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-13T18:18:55Z","receivedAt":"2018-11-13T18:19:00Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nBen Peart wrote:\n> On 11/12/2018 7:39 PM, Jonathan Nieder wrote:\n\n>> As with EOIE, popular versions of Git do not support the new IEOT\n>> extension yet.  When accessing a Git repository written by a more\n>> modern version of Git, they correctly ignore the unrecognized section,\n>> but in the process they loudly warn\n>>\n>> \tignoring IEOT extension\n>>\n>> resulting in confusion for users.  Introduce the index extension more\n>> gently by not writing it yet in this first version with support for\n>> it.\n[...]\n>> Introduce a '[index] recordOffsetTable' configuration variable to\n>> control whether the new index extension is written.\n>\n> Why introduce a new setting to disable writing the IEOT extension instead of\n> just using the existing index.threads setting?  If index.threads=1 then the\n> IEOT extension isn't written which (I believe) will accomplish the same\n> goal.\n\nDo you mean defaulting to index.threads=1?  I don't think that would\nbe a good default, but if you have a different change in mind then I'd\nbe happy to hear it.\n\nOr do you mean that if the user has explicitly specified index.threads=true,\nthen that should imply index.recordOffsetTable=true so users only have\nto set one setting to turn it on?  I can imagine that working well.\n\nThanks,\nJonathan\n"},{"id":"363186","messageId":"20181113182502.GC68106@google.com","threadId":"49204","inReplyTo":"5fae19dc-2e77-1211-0086-e7aa9d30562f@gmail.com","subject":"Re: [PATCH 1/3] eoie: default to not writing EOIE section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-13T18:25:02Z","receivedAt":"2018-11-13T18:25:07Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nBen Peart wrote:\n\n> While I can understand the user confusion the warning about ignoring an\n> extension could cause I guess I'm a little surprised that people would see\n> it that often.  To see the warning means they are running a new version of\n> git in the same repo as they are running an old version of git.  I just\n> haven't ever experienced that (I only ever have one copy of git installed)\n> so am surprised it comes up often enough to warrant this change.\n\nGreat question.  There are a few contexts where it comes up:\n\n 1. Using multiple versions of Git on a single machine.  For example,\n    some IDEs bundle a particular version of Git, which can be a\n    different version from the system copy, or on a Mac, /usr/bin/git\n    quickly goes out of sync with the Homebrew git in\n    /usr/local/bin/git.\n\n 2. Sharing a single Git repository between multiple machines.  This is\n    not unusual, using NFS or sshfs, for example.\n\n 3. Downgrading after trying a new version of Git.\n\nTo support these, Git is generally careful to avoid writing\nrepositories that older versions of Git do not understand.  The EOIE\nextension was almost perfect in this respect: it works fine with older\nversions of Git, except for the alarming error message.\n\n> That said, if it _is_ that much of an issue, this patch makes sense and\n> provides a way to more gracefully transition into this feature.  Even if we\n> had some logic to dynamically determine whether to write it or not, we'd\n> still want to avoid confusing users when it did get written out.\n\nYes.  An earlier version of this patch defaulted to writing EOIE if\nand only if the .git/index file already has an EOIE extension.  There\nwere enough holes in that (commands like \"git reset\" that do not read\nthe existing index file) and enough complexity that it didn't seem\nworth it.\n\nReally in this series, patch 3/3 is the one I care most about.  I wish\nwe had had it years ago. :)  It would make patches 1 and 2\nunnecessary.\n\nThanks,\nJonathan\n"},{"id":"363199","messageId":"1b890149-ee7f-c391-9abc-46d120e4324c@gmail.com","threadId":"49204","inReplyTo":"20181113181855.GB68106@google.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-13T19:15:08Z","receivedAt":"2018-11-13T19:15:13Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/13/2018 1:18 PM, Jonathan Nieder wrote:\n> Hi,\n> \n> Ben Peart wrote:\n>> On 11/12/2018 7:39 PM, Jonathan Nieder wrote:\n> \n>>> As with EOIE, popular versions of Git do not support the new IEOT\n>>> extension yet.  When accessing a Git repository written by a more\n>>> modern version of Git, they correctly ignore the unrecognized section,\n>>> but in the process they loudly warn\n>>>\n>>> \tignoring IEOT extension\n>>>\n>>> resulting in confusion for users.  Introduce the index extension more\n>>> gently by not writing it yet in this first version with support for\n>>> it.\n> [...]\n>>> Introduce a '[index] recordOffsetTable' configuration variable to\n>>> control whether the new index extension is written.\n>>\n>> Why introduce a new setting to disable writing the IEOT extension instead of\n>> just using the existing index.threads setting?  If index.threads=1 then the\n>> IEOT extension isn't written which (I believe) will accomplish the same\n>> goal.\n> \n> Do you mean defaulting to index.threads=1?  I don't think that would\n> be a good default, but if you have a different change in mind then I'd\n> be happy to hear it.\n> \n> Or do you mean that if the user has explicitly specified index.threads=true,\n> then that should imply index.recordOffsetTable=true so users only have\n> to set one setting to turn it on?  I can imagine that working well.\n> \n\nReading the index with multiple threads requires the EOIE and IEOT \nextensions to exist in the index.  If either extension doesn't exist, \nthen the code falls back to the single threaded path.  That means you \ncan't have both 1) no warning for old versions of git and 2) \nmulti-threaded reading for new versions of git.\n\nIf you set index.threads=1, that will prevent the IEOT extension from \nbeing written and there will be no \"ignoring IEOT extension\" warning in \nolder versions of git.\n\nWith this patch 'as is' you would have to set both index.threads=true \nand index.recordOffsetTable=true to get multi-threaded index reads.  If \neither is set to false, it will silently drop back to single threaded reads.\n\n> Thanks,\n> Jonathan\n> \n"},{"id":"363217","messageId":"20181113210815.GD68106@google.com","threadId":"49204","inReplyTo":"1b890149-ee7f-c391-9abc-46d120e4324c@gmail.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-13T21:08:15Z","receivedAt":"2018-11-13T21:08:21Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi again,\n\nBen Peart wrote:\n> On 11/13/2018 1:18 PM, Jonathan Nieder wrote:\n>> Ben Peart wrote:\n\n>>> Why introduce a new setting to disable writing the IEOT extension instead of\n>>> just using the existing index.threads setting?  If index.threads=1 then the\n>>> IEOT extension isn't written which (I believe) will accomplish the same\n>>> goal.\n>>\n>> Do you mean defaulting to index.threads=1?  I don't think that would\n>> be a good default, but if you have a different change in mind then I'd\n>> be happy to hear it.\n>>\n>> Or do you mean that if the user has explicitly specified index.threads=true,\n>> then that should imply index.recordOffsetTable=true so users only have\n>> to set one setting to turn it on?  I can imagine that working well.\n>\n> Reading the index with multiple threads requires the EOIE and IEOT\n> extensions to exist in the index.  If either extension doesn't exist, then\n> the code falls back to the single threaded path.  That means you can't have\n> both 1) no warning for old versions of git and 2) multi-threaded reading for\n> new versions of git.\n>\n> If you set index.threads=1, that will prevent the IEOT extension from being\n> written and there will be no \"ignoring IEOT extension\" warning in older\n> versions of git.\n>\n> With this patch 'as is' you would have to set both index.threads=true and\n> index.recordOffsetTable=true to get multi-threaded index reads.  If either\n> is set to false, it will silently drop back to single threaded reads.\n\nSorry, I'm still not understanding what you're proposing.  What would be\n\n- the default behavior\n- the mechanism for changing that behavior\n\nunder your proposal?\n\nI consider index.threads=1 to be a bad default.  I would understand if\nyou are saying that that should be the default, and I tried to propose\na different way to achieve what you're looking for in the quoted reply\nabove (but I'm not understanding whether you like that proposal or\nnot).\n\nJonathan\n"},{"id":"363280","messageId":"xmqqzhuctp72.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20181113182502.GC68106@google.com","subject":"Re: [PATCH 1/3] eoie: default to not writing EOIE section","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-14T01:36:49Z","receivedAt":"2018-11-14T01:36:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n>  1. Using multiple versions of Git on a single machine.  For example,\n>     some IDEs bundle a particular version of Git, which can be a\n>     different version from the system copy, or on a Mac, /usr/bin/git\n>     quickly goes out of sync with the Homebrew git in\n>     /usr/local/bin/git.\n\nExactly this, especially the latter, is the answer to your \nquestion in an earlier message:\n\n>> Am I understanding correclty?  Can you give an example of when a user\n>> would *want* to see this message and what they would do in response?\n\nThe user may not be even aware of using another version of Git that\ndoes not know how to take advantage of the version of Git you have\nused in the repository, and it can be a mistake the user may want to\nfix (e.g. by futzing with PATH).  The message would help the user\nnotice the situation and take corrective action.  Users of IDEs that\nbundle stale version of Git cannot even bug the supplier of the IDE\nto make them more up-to-date if they aren't aware of it.\n"},{"id":"363289","messageId":"xmqqo9asqrxu.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"f2f8cec8-d770-a1e9-b5a1-83653575122e@gmail.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-14T03:05:49Z","receivedAt":"2018-11-14T03:05:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <peartben@gmail.com> writes:\n\n> Why introduce a new setting to disable writing the IEOT extension\n> instead of just using the existing index.threads setting?\n\nBut index.threads is about what the reader does, not about the\nwriter who does not even know who will be reading the resulting\nindex, no?\n"},{"id":"363290","messageId":"xmqqh8gkqr2l.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20181113004019.GD170017@google.com","subject":"Re: [PATCH 3/3] index: do not warn about unrecognized extensions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-14T03:24:34Z","receivedAt":"2018-11-14T03:24:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> We cannot change the past, but for index extensions of the future,\n> there is a straightforward improvement: silence that message except\n> when tracing.  This way, the message is still available when\n> debugging, but in everyday use it does not show up so (once most Git\n> users have this patch) we can turn on new optional extensions right\n> away without alarming people.\n\nThat argument ignores the \"let the users know they are using a stale\nversion when they did use (either by accident or deliberately) a\nmore recent one\" value, though.\n\nEven if we consider that this is only for debugging, I am not sure\nif tracing is the right place to add.  As long as the \"optional\nextensions can be ignored without affecting the correctness\" rule\nholds, there is nothing gained by letting these messages shown for\ndebugging purposes, and if there is such a bug (e.g. we introduced\nan optional extension but the code that wrote an index with an\noptional extension wrote the non-optional part in such a way that it\ncannot be correctly handled without the extension that is supposed\nto be optional) we'd probably want to let users notice without\nhaving to explicitly go into a debugging session.  If Googling for\n\"ignoring FNCY ext\" gives \"discard your index with 'reset HEAD',\nbecause an index file with FNCY ext cannot be read without\nunderstanding it\", that may prevent damages from happening in the\nfirst place.  On the other hand, hiding it behind tracing would mean\nthe user first need to exprience an unknown breakage first and then\nhas to enable tracing among other 47 different things to diagnose\nand drill down to the root cause.\n\n\n"},{"id":"363387","messageId":"75c91c81-f66f-ab2d-2b29-339deb3a6557@gmail.com","threadId":"49204","inReplyTo":"20181113210815.GD68106@google.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-14T18:09:35Z","receivedAt":"2018-11-14T18:09:42Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/13/2018 4:08 PM, Jonathan Nieder wrote:\n> Hi again,\n> \n> Ben Peart wrote:\n>> On 11/13/2018 1:18 PM, Jonathan Nieder wrote:\n>>> Ben Peart wrote:\n> \n>>>> Why introduce a new setting to disable writing the IEOT extension instead of\n>>>> just using the existing index.threads setting?  If index.threads=1 then the\n>>>> IEOT extension isn't written which (I believe) will accomplish the same\n>>>> goal.\n>>>\n>>> Do you mean defaulting to index.threads=1?  I don't think that would\n>>> be a good default, but if you have a different change in mind then I'd\n>>> be happy to hear it.\n>>>\n>>> Or do you mean that if the user has explicitly specified index.threads=true,\n>>> then that should imply index.recordOffsetTable=true so users only have\n>>> to set one setting to turn it on?  I can imagine that working well.\n>>\n>> Reading the index with multiple threads requires the EOIE and IEOT\n>> extensions to exist in the index.  If either extension doesn't exist, then\n>> the code falls back to the single threaded path.  That means you can't have\n>> both 1) no warning for old versions of git and 2) multi-threaded reading for\n>> new versions of git.\n>>\n>> If you set index.threads=1, that will prevent the IEOT extension from being\n>> written and there will be no \"ignoring IEOT extension\" warning in older\n>> versions of git.\n>>\n>> With this patch 'as is' you would have to set both index.threads=true and\n>> index.recordOffsetTable=true to get multi-threaded index reads.  If either\n>> is set to false, it will silently drop back to single threaded reads.\n> \n> Sorry, I'm still not understanding what you're proposing.  What would be\n> \n> - the default behavior\n> - the mechanism for changing that behavior\n> \n> under your proposal?\n> \n> I consider index.threads=1 to be a bad default.  I would understand if\n> you are saying that that should be the default, and I tried to propose\n> a different way to achieve what you're looking for in the quoted reply\n> above (but I'm not understanding whether you like that proposal or\n> not).\n> \n\nToday, both the write logic (ie should we write out the IEOT extension) \nand the read logic (should I use the IEOT, if available, and do \nmulti-threaded reading) are controlled by the single \"index.threads\" \nsetting.  I would like to keep the settings as simple as possible to \nprevent user confusion.\n\nIf we have two different settings (index.threads and \nindex.recordoffsettable) then the only combination that will result in \nthe user actually getting multi-threaded reads is when they are both set \nto true.  Any other combination will silently fail.  I think it would be \nconfusing if you set index.threads=true and got no error message but \ndidn't get multi-threaded reads either (or vice versa).\n\nIf you want to prevent any of the scary \"ignoring IEOT extension\" from \never happening then your only option is to turn off the IEOT writing by \ndefault.  The downside is that people have to discover and turn it on if \nthey want the improvements.  This can be achieved by changing the \ndefault for index.threads from \"true\" to \"false.\"\n\ndiff --git a/config.c b/config.c\nindex 2ee29f6f86..86f5c14294 100644\n--- a/config.c\n+++ b/config.c\n@@ -2291,7 +2291,7 @@ int git_config_get_fsmonitor(void)\n\n  int git_config_get_index_threads(void)\n  {\n-       int is_bool, val = 0;\n+       int is_bool, val = 1;\n\n         val = git_env_ulong(\"GIT_TEST_INDEX_THREADS\", 0);\n         if (val)\n\n\nIf you want to provide a way for a concerned user to disable the message \nafter the first time they have seen it, then they can be instructed to \nrun 'git config --global index.threads false'\n\nThere is no way to get multi-threaded reads and NOT get the scary \nmessage with older versions of git.  Multi-threaded reads require the \nIEOT extension to be written into the index and the existence of the \nIEOT extension in the index will always generate the scary warning.\n\n> Jonathan\n> \n"},{"id":"363388","messageId":"70f48153-fedf-4b82-780b-eca08981a8eb@gmail.com","threadId":"49204","inReplyTo":"xmqqh8gkqr2l.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 3/3] index: do not warn about unrecognized extensions","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-14T18:19:31Z","receivedAt":"2018-11-14T18:19:43Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/13/2018 10:24 PM, Junio C Hamano wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n> \n>> We cannot change the past, but for index extensions of the future,\n>> there is a straightforward improvement: silence that message except\n>> when tracing.  This way, the message is still available when\n>> debugging, but in everyday use it does not show up so (once most Git\n>> users have this patch) we can turn on new optional extensions right\n>> away without alarming people.\n> \n> That argument ignores the \"let the users know they are using a stale\n> version when they did use (either by accident or deliberately) a\n> more recent one\" value, though.\n> \n> Even if we consider that this is only for debugging, I am not sure\n> if tracing is the right place to add.  As long as the \"optional\n> extensions can be ignored without affecting the correctness\" rule\n> holds, there is nothing gained by letting these messages shown for\n> debugging purposes\n\nHaving recently written a couple of patches that utilize an optional \nextension - I actually found the warning to be a helpful debugging tool \nand would like to see them enabled via tracing.  It would also be \nhelpful to see the opposite - I'm looking for an optional extension but \nit is missing.\n\nThe most common scenario was when I'd be testing my changes in different \nrepos and forget that I needed to force an updated index to be written \nthat contained the extension I was trying to test.  The \"silently ignore \nthe optional extension\" behavior is good for end users but as a \ndeveloper, I'd like to be able to have it yell at me via tracing. :-)\n\nIMHO - if an end user has to turn on tracing, I view that as a failure \non our part.  No end user should have to understand the inner workings \nof git to be able to use it effectively.\n\nand if there is such a bug (e.g. we introduced\n> an optional extension but the code that wrote an index with an\n> optional extension wrote the non-optional part in such a way that it\n> cannot be correctly handled without the extension that is supposed\n> to be optional) we'd probably want to let users notice without\n> having to explicitly go into a debugging session.  If Googling for\n> \"ignoring FNCY ext\" gives \"discard your index with 'reset HEAD',\n> because an index file with FNCY ext cannot be read without\n> understanding it\", that may prevent damages from happening in the\n> first place.  On the other hand, hiding it behind tracing would mean\n> the user first need to exprience an unknown breakage first and then\n> has to enable tracing among other 47 different things to diagnose\n> and drill down to the root cause.\n> \n> \n"},{"id":"363407","messageId":"20181115000539.GA92137@google.com","threadId":"49204","inReplyTo":"75c91c81-f66f-ab2d-2b29-339deb3a6557@gmail.com","subject":"Re: [PATCH 2/3] ieot: default to not writing IEOT section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-15T00:05:39Z","receivedAt":"2018-11-15T00:05:45Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nBen Peart wrote:\n\n> There is no way to get multi-threaded reads and NOT get the scary message\n> with older versions of git.  Multi-threaded reads require the IEOT extension\n> to be written into the index and the existence of the IEOT extension in the\n> index will always generate the scary warning.\n\nThis is where I think we differ.  I want my local copy of Git to get\nmulti-threaded reads as long as IEOT happens to be there, even if I am\nnot ready to write IEOT myself yet.\n\nI understand that this differs from what you would prefer, so I'd like\nto find some compromise that makes us both happy.  I've tried to\nsuggest one:\n\n   Make explicitly setting index.threads=true imply\n   index.recordOffsetTable=true.  That way, the default behavior is the\n   behavior I prefer, and a client can simply set index.threads=true to\n   get the behavior I think you are describing preferring.\n\nMy preference is instead what I sent in patch 2/3 (for simplicity,\nespecially since the default of index.recordOffsetTable=false would be\nonly temporary), but this would work okay for me.\n\nI'll send this as a patch.  If there is a reason it won't work for\nyou, I would be very happy to learn more about why.\n\nThanks,\nJonathan\n"},{"id":"363408","messageId":"20181115001915.GB92137@google.com","threadId":"49204","inReplyTo":"xmqqzhuctp72.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 1/3] eoie: default to not writing EOIE section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-15T00:19:15Z","receivedAt":"2018-11-15T00:19:20Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n\n>>  1. Using multiple versions of Git on a single machine.  For example,\n>>     some IDEs bundle a particular version of Git, which can be a\n>>     different version from the system copy, or on a Mac, /usr/bin/git\n>>     quickly goes out of sync with the Homebrew git in\n>>     /usr/local/bin/git.\n>\n> Exactly this, especially the latter, is the answer to your\n> question in an earlier message:\n>\n>>> Am I understanding correctly?  Can you give an example of when a user\n>>> would *want* to see this message and what they would do in response?\n>\n> The user may not be even aware of using another version of Git that\n> does not know how to take advantage of the version of Git you have\n> used in the repository, and it can be a mistake the user may want to\n> fix (e.g. by futzing with PATH).\n\nAh, thanks much.  I'll add a hint along those lines (e.g.\n\n warning: ignoring optional IEOT index extension\n hint: This is likely due to the file having been written by a newer\n hint: version of Git than is reading it.  You can upgrade Git to\n hint: take advantage of performance improvements from the updated\n hint: file format.\n hint:\n hint: You can run \"git config advice.unknownIndexExtension true\" to\n hint: suppress this message.\n\nI am still vaguely uncomfortable with this since it seems analogous to\nwarning that the server is advertising an unrecognized capability, but\nI can live with it. :)\n\nPatch coming in a few moments.\n\nJonathan\n"},{"id":"363702","messageId":"20181120060920.GA144753@google.com","threadId":"49204","inReplyTo":"xmqqo9asqrxu.fsf@gitster-ct.c.googlers.com","subject":"[PATCH v2 0/5] Avoid confusing messages from new index extensions","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-20T06:09:20Z","receivedAt":"2018-11-20T06:09:26Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n> Ben Peart <peartben@gmail.com> writes:\n\n>> Why introduce a new setting to disable writing the IEOT extension\n>> instead of just using the existing index.threads setting?\n>\n> But index.threads is about what the reader does, not about the\n> writer who does not even know who will be reading the resulting\n> index, no?\n\nIt affects the writer, too, since it affects the number of blocks, but\nfrom an end user's point of view, I agree.\n\nHere's an updated version of the series.\n\nPatches 1-3 are as before, except that they are rebased to avoid\nconflicting with nd/config-split.\n\nPatch 4 allows enabling the new index extensions with a single config\nsetting, to address the feedback above.\n\nPatch 5 revives the noisiness when encountering an unknown index\nextension, guarded with an advice setting.\n\nSorry for the delay in getting this out.  Thoughts of all kinds\nwelcome, as always.\n\nSincerely,\nJonathan Nieder (5):\n  eoie: default to not writing EOIE section\n  ieot: default to not writing IEOT section\n  index: do not warn about unrecognized extensions\n  index: make index.threads=true enable ieot and eoie\n  index: offer advice for unknown index extensions\n\n Documentation/config/index.txt | 16 ++++++++++\n advice.c                       |  2 ++\n advice.h                       |  1 +\n config.c                       | 17 ++++++-----\n config.h                       |  2 +-\n read-cache.c                   | 54 +++++++++++++++++++++++++++++-----\n t/t1700-split-index.sh         | 11 ++++---\n 7 files changed, 84 insertions(+), 19 deletions(-)\n"},{"id":"363703","messageId":"20181120061147.GB144753@google.com","threadId":"49204","inReplyTo":"20181120060920.GA144753@google.com","subject":"[PATCH 1/5] eoie: default to not writing EOIE section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-20T06:11:47Z","receivedAt":"2018-11-20T06:11:51Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Since 3b1d9e04 (eoie: add End of Index Entry (EOIE) extension,\n2018-10-10) Git defaults to writing the new EOIE section when writing\nout an index file.  Usually that is a good thing because it improves\nthreaded performance, but when a Git repository is shared with older\nversions of Git, it produces a confusing warning:\n\n  $ git status\n  ignoring EOIE extension\n  HEAD detached at 371ed0defa\n  nothing to commit, working tree clean\n\nLet's introduce the new index extension more gently.  First we'll roll\nout the new version of Git that understands it, and then once\nsufficiently many users are using such a version, we can flip the\ndefault to writing it by default.\n\nIntroduce a '[index] recordEndOfIndexEntries' configuration variable\nto allow interested users to benefit from this index extension early.\n\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n---\nRebased.  No other change from v1.\n\nAs Jonathan pointed out, it would be nice to have tests here.  Ben,\nany advice for how I could write some in a followup change?  E.g. does\nDerrick Stolee's tracing based testing trick apply here?\n\n Documentation/config/index.txt |  7 +++++++\n read-cache.c                   | 11 ++++++++++-\n t/t1700-split-index.sh         | 11 +++++++----\n 3 files changed, 24 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/config/index.txt b/Documentation/config/index.txt\nindex 4b94b6bedc..8e138aba7a 100644\n--- a/Documentation/config/index.txt\n+++ b/Documentation/config/index.txt\n@@ -1,3 +1,10 @@\n+index.recordEndOfIndexEntries::\n+\tSpecifies whether the index file should include an \"End Of Index\n+\tEntry\" section. This reduces index load time on multiprocessor\n+\tmachines but produces a message \"ignoring EOIE extension\" when\n+\treading the index using Git versions before 2.20. Defaults to\n+\t'false'.\n+\n index.threads::\n \tSpecifies the number of threads to spawn when loading the index.\n \tThis is meant to reduce index load time on multiprocessor machines.\ndiff --git a/read-cache.c b/read-cache.c\nindex 4ca81286c0..1e9c772603 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2689,6 +2689,15 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n \t\trollback_lock_file(lockfile);\n }\n \n+static int record_eoie(void)\n+{\n+\tint val;\n+\n+\tif (!git_config_get_bool(\"index.recordendofindexentries\", &val))\n+\t\treturn val;\n+\treturn 0;\n+}\n+\n /*\n  * On success, `tempfile` is closed. If it is the temporary file\n  * of a `struct lock_file`, we will therefore effectively perform\n@@ -2936,7 +2945,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \t * read.  Write it out regardless of the strip_extensions parameter as we need it\n \t * when loading the shared index.\n \t */\n-\tif (offset) {\n+\tif (offset && record_eoie()) {\n \t\tstruct strbuf sb = STRBUF_INIT;\n \n \t\twrite_eoie_extension(&sb, &eoie_c, offset);\ndiff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\nindex 2ac47aa0e4..0cbac64e28 100755\n--- a/t/t1700-split-index.sh\n+++ b/t/t1700-split-index.sh\n@@ -25,14 +25,17 @@ test_expect_success 'enable split index' '\n \tgit update-index --split-index &&\n \ttest-tool dump-split-index .git/index >actual &&\n \tindexversion=$(test-tool index-version <.git/index) &&\n+\n+\t# NEEDSWORK: Stop hard-coding checksums.\n \tif test \"$indexversion\" = \"4\"\n \tthen\n-\t\town=3527df833c6c100d3d1d921a9a782d62a8be4b58\n-\t\tbase=746f7ab2ed44fb839efdfbffcf399d0b113fb4cb\n+\t\town=432ef4b63f32193984f339431fd50ca796493569\n+\t\tbase=508851a7f0dfa8691e9f69c7f055865389012491\n \telse\n-\t\town=5e9b60117ece18da410ddecc8b8d43766a0e4204\n-\t\tbase=4370042739b31cd17a5c5cd6043a77c9a00df113\n+\t\town=8299b0bcd1ac364e5f1d7768efb62fa2da79a339\n+\t\tbase=39d890139ee5356c7ef572216cebcd27aa41f9df\n \tfi &&\n+\n \tcat >expect <<-EOF &&\n \town $own\n \tbase $base\n-- \n2.20.0.rc0.387.gc7a69e6b6c\n\n"},{"id":"363704","messageId":"20181120061221.GC144753@google.com","threadId":"49204","inReplyTo":"20181120060920.GA144753@google.com","subject":"[PATCH 2/5] ieot: default to not writing IEOT section","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-20T06:12:22Z","receivedAt":"2018-11-20T06:12:26Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"As with EOIE, popular versions of Git do not support the new IEOT\nextension yet.  When accessing a Git repository written by a more\nmodern version of Git, they correctly ignore the unrecognized section,\nbut in the process they loudly warn\n\n\tignoring IEOT extension\n\nresulting in confusion for users.  Introduce the index extension more\ngently by not writing it yet in this first version with support for\nit.  Soon, once sufficiently many users are running a modern version\nof Git, we can flip the default so users benefit from this index\nextension by default.\n\nIntroduce a '[index] recordOffsetTable' configuration variable to\ncontrol whether the new index extension is written.\n\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n---\nAs with patch 1/5, no change from v1 other than rebasing.\n\n Documentation/config/index.txt |  7 +++++++\n read-cache.c                   | 11 ++++++++++-\n 2 files changed, 17 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/config/index.txt b/Documentation/config/index.txt\nindex 8e138aba7a..de44183235 100644\n--- a/Documentation/config/index.txt\n+++ b/Documentation/config/index.txt\n@@ -5,6 +5,13 @@ index.recordEndOfIndexEntries::\n \treading the index using Git versions before 2.20. Defaults to\n \t'false'.\n \n+index.recordOffsetTable::\n+\tSpecifies whether the index file should include an \"Index Entry\n+\tOffset Table\" section. This reduces index load time on\n+\tmultiprocessor machines but produces a message \"ignoring IEOT\n+\textension\" when reading the index using Git versions before 2.20.\n+\tDefaults to 'false'.\n+\n index.threads::\n \tSpecifies the number of threads to spawn when loading the index.\n \tThis is meant to reduce index load time on multiprocessor machines.\ndiff --git a/read-cache.c b/read-cache.c\nindex 1e9c772603..f3d5638d9e 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2698,6 +2698,15 @@ static int record_eoie(void)\n \treturn 0;\n }\n \n+static int record_ieot(void)\n+{\n+\tint val;\n+\n+\tif (!git_config_get_bool(\"index.recordoffsettable\", &val))\n+\t\treturn val;\n+\treturn 0;\n+}\n+\n /*\n  * On success, `tempfile` is closed. If it is the temporary file\n  * of a `struct lock_file`, we will therefore effectively perform\n@@ -2761,7 +2770,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \telse\n \t\tnr_threads = 1;\n \n-\tif (nr_threads != 1) {\n+\tif (nr_threads != 1 && record_ieot()) {\n \t\tint ieot_blocks, cpus;\n \n \t\t/*\n-- \n2.20.0.rc0.387.gc7a69e6b6c-goog\n\n"},{"id":"363705","messageId":"20181120061251.GD144753@google.com","threadId":"49204","inReplyTo":"20181120060920.GA144753@google.com","subject":"[PATCH 3/5] index: do not warn about unrecognized extensions","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-20T06:12:51Z","receivedAt":"2018-11-20T06:12:56Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Documentation/technical/index-format explains:\n\n     4-byte extension signature. If the first byte is 'A'..'Z' the\n     extension is optional and can be ignored.\n\nThis allows gracefully introducing a new index extension without\nhaving to rely on all readers having support for it.  Mandatory\nextensions start with a lowercase letter and optional ones start with\na capital.  Thus the versions of Git acting on a shared local\nrepository do not have to upgrade in lockstep.\n\nWe almost obey that convention, but there is a problem: when\nencountering an unrecognized optional extension, we write\n\n\tignoring FNCY extension\n\nto stderr, which alarms users.  This means that in practice we have\nhad to introduce index extensions in two steps: first add read\nsupport, and then a while later, start writing by default.  This\ndelays when users can benefit from improvements to the index format.\n\nWe cannot change the past, but for index extensions of the future,\nthere is a straightforward improvement: silence that message except\nwhen tracing.  This way, the message is still available when\ndebugging, but in everyday use it does not show up so (once most Git\nusers have this patch) we can turn on new optional extensions right\naway without alarming people.\n\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n---\nUnchanged.\n\n read-cache.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex f3d5638d9e..83d24357a6 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1726,7 +1726,7 @@ static int read_index_extension(struct index_state *istate,\n \t\tif (*ext < 'A' || 'Z' < *ext)\n \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n \t\t\t\t     ext);\n-\t\tfprintf(stderr, \"ignoring %.4s extension\\n\", ext);\n+\t\ttrace_printf(\"ignoring %.4s extension\\n\", ext);\n \t\tbreak;\n \t}\n \treturn 0;\n-- \n2.20.0.rc0.387.gc7a69e6b6c\n\n"},{"id":"363707","messageId":"20181120061426.GE144753@google.com","threadId":"49204","inReplyTo":"20181120060920.GA144753@google.com","subject":"[PATCH 4/5] index: make index.threads=true enable ieot and eoie","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-20T06:14:26Z","receivedAt":"2018-11-20T06:14:31Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"If a user explicitly sets\n\n\t[index]\n\t\tthreads = true\n\nto read the index using multiple threads, ensure that index writes\ninclude the offset table by default to make that possible.  This\nensures that the user's intent of turning on threading is respected.\n\nIn other words, permit the following configurations:\n\n- index.threads and index.recordOffsetTable unspecified: do not write\n  the offset table yet (to avoid alarming the user with \"ignoring IEOT\n  extension\" messages when an older version of Git accesses the\n  repository) but do make use of multiple threads to read the index if\n  the supporting offset table is present.\n\n  This can also be requested explicitly by setting index.threads=true,\n  0, or >1 and index.recordOffsetTable=false.\n\n- index.threads=false or 1: do not write the offset table, and do not\n  make use of the offset table.\n\n  One can set index.recordOffsetTable=false as well, to be more\n  explicit.\n\n- index.threads=true, 0, or >1 and index.recordOffsetTable unspecified:\n  write the offset table and make use of threads at read time.\n\n  This can also be requested by setting index.threads=true, 0, >1, or\n  unspecified and index.recordOffsetTable=true.\n\nFortunately the complication is temporary: once most Git installations\nhave upgraded to a version with support for the IEOT and EOIE\nextensions, we can flip the defaults for index.recordEndOfIndexEntries\nand index.recordOffsetTable to true and eliminate the settings.\n\nHelped-by: Ben Peart <benpeart@microsoft.com>\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n---\nNew, based on Ben Peart's feedback.  Turned out simpler than I feared\n--- thanks, Ben, for pushing for this.\n\n Documentation/config/index.txt |  6 ++++--\n config.c                       | 17 ++++++++++-------\n config.h                       |  2 +-\n read-cache.c                   | 23 +++++++++++++++++------\n 4 files changed, 32 insertions(+), 16 deletions(-)\n\ndiff --git a/Documentation/config/index.txt b/Documentation/config/index.txt\nindex de44183235..f181503041 100644\n--- a/Documentation/config/index.txt\n+++ b/Documentation/config/index.txt\n@@ -3,14 +3,16 @@ index.recordEndOfIndexEntries::\n \tEntry\" section. This reduces index load time on multiprocessor\n \tmachines but produces a message \"ignoring EOIE extension\" when\n \treading the index using Git versions before 2.20. Defaults to\n-\t'false'.\n+\t'true' if index.threads has been explicitly enabled, 'false'\n+\totherwise.\n \n index.recordOffsetTable::\n \tSpecifies whether the index file should include an \"Index Entry\n \tOffset Table\" section. This reduces index load time on\n \tmultiprocessor machines but produces a message \"ignoring IEOT\n \textension\" when reading the index using Git versions before 2.20.\n-\tDefaults to 'false'.\n+\tDefaults to 'true' if index.threads has been explicitly enabled,\n+\t'false' otherwise.\n \n index.threads::\n \tSpecifies the number of threads to spawn when loading the index.\ndiff --git a/config.c b/config.c\nindex 04286f7717..ff521eb27a 100644\n--- a/config.c\n+++ b/config.c\n@@ -2294,22 +2294,25 @@ int git_config_get_fsmonitor(void)\n \treturn 0;\n }\n \n-int git_config_get_index_threads(void)\n+int git_config_get_index_threads(int *dest)\n {\n-\tint is_bool, val = 0;\n+\tint is_bool, val;\n \n \tval = git_env_ulong(\"GIT_TEST_INDEX_THREADS\", 0);\n-\tif (val)\n-\t\treturn val;\n+\tif (val) {\n+\t\t*dest = val;\n+\t\treturn 0;\n+\t}\n \n \tif (!git_config_get_bool_or_int(\"index.threads\", &is_bool, &val)) {\n \t\tif (is_bool)\n-\t\t\treturn val ? 0 : 1;\n+\t\t\t*dest = val ? 0 : 1;\n \t\telse\n-\t\t\treturn val;\n+\t\t\t*dest = val;\n+\t\treturn 0;\n \t}\n \n-\treturn 0; /* auto */\n+\treturn 1;\n }\n \n NORETURN\ndiff --git a/config.h b/config.h\nindex a06027e69b..ee5d3fa7b4 100644\n--- a/config.h\n+++ b/config.h\n@@ -246,11 +246,11 @@ extern int git_config_get_bool(const char *key, int *dest);\n extern int git_config_get_bool_or_int(const char *key, int *is_bool, int *dest);\n extern int git_config_get_maybe_bool(const char *key, int *dest);\n extern int git_config_get_pathname(const char *key, const char **dest);\n+extern int git_config_get_index_threads(int *dest);\n extern int git_config_get_untracked_cache(void);\n extern int git_config_get_split_index(void);\n extern int git_config_get_max_percent_split_change(void);\n extern int git_config_get_fsmonitor(void);\n-extern int git_config_get_index_threads(void);\n \n /* This dies if the configured or default date is in the future */\n extern int git_config_get_expiry(const char *key, const char **output);\ndiff --git a/read-cache.c b/read-cache.c\nindex 83d24357a6..002ed2c1e4 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -2176,7 +2176,8 @@ int do_read_index(struct index_state *istate, const char *path, int must_exist)\n \n \tsrc_offset = sizeof(*hdr);\n \n-\tnr_threads = git_config_get_index_threads();\n+\tif (git_config_get_index_threads(&nr_threads))\n+\t\tnr_threads = 1;\n \n \t/* TODO: does creating more threads than cores help? */\n \tif (!nr_threads) {\n@@ -2695,7 +2696,13 @@ static int record_eoie(void)\n \n \tif (!git_config_get_bool(\"index.recordendofindexentries\", &val))\n \t\treturn val;\n-\treturn 0;\n+\n+\t/*\n+\t * As a convenience, the end of index entries extension\n+\t * used for threading is written by default if the user\n+\t * explicitly requested threaded index reads.\n+\t */\n+\treturn !git_config_get_index_threads(&val) && val != 1;\n }\n \n static int record_ieot(void)\n@@ -2704,7 +2711,13 @@ static int record_ieot(void)\n \n \tif (!git_config_get_bool(\"index.recordoffsettable\", &val))\n \t\treturn val;\n-\treturn 0;\n+\n+\t/*\n+\t * As a convenience, the offset table used for threading is\n+\t * written by default if the user explicitly requested\n+\t * threaded index reads.\n+\t */\n+\treturn !git_config_get_index_threads(&val) && val != 1;\n }\n \n /*\n@@ -2765,9 +2778,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n \tif (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)\n \t\treturn -1;\n \n-\tif (HAVE_THREADS)\n-\t\tnr_threads = git_config_get_index_threads();\n-\telse\n+\tif (!HAVE_THREADS || git_config_get_index_threads(&nr_threads))\n \t\tnr_threads = 1;\n \n \tif (nr_threads != 1 && record_ieot()) {\n-- \n2.20.0.rc0.387.gc7a69e6b6c\n\n"},{"id":"363708","messageId":"20181120061544.GF144753@google.com","threadId":"49204","inReplyTo":"20181120060920.GA144753@google.com","subject":"[PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-20T06:15:44Z","receivedAt":"2018-11-20T06:15:49Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"It is not unusual for multiple distinct versions of Git to act on a\nsingle repository.  For example, some IDEs bundle a particular version\nof Git, which can be a different version from the system copy of Git,\nor on a Mac, /usr/bin/git quickly goes out of sync with the Homebrew\ngit in /usr/local/bin/git.\n\nWhen a newer version of Git writes an index file that an older version\nof Git does not know how to make full use of, this is a teaching\nopportunity.  The user may not be aware of what version of Git they\nare using.  Print an advice message to help the user to use the most\nfull featured version of Git (e.g. by futzing with their PATH).\n\n  warning: ignoring optional IEOT index extension\n  hint: This is likely due to the file having been written by a newer\n  hint: version of Git than is reading it.  You can upgrade Git to\n  hint: take advantage of performance improvements from the updated\n  hint: file format.\n  hint:\n  hint: You can run \"git config advice.unknownIndexExtension false\"\n  hint: to suppress this message.\n\nThis replaces the message\n\n  ignoring IEOT extension\n\nthat existed previously and did not provide enough detail for a user\nto act on it or suppress it.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n---\nNew, based on Junio's hints about the message removed in patch 3/5.\n\nThat's the end of the series.  Thanks for reading, and thanks again\nfor your help so far.\n\n advice.c     |  2 ++\n advice.h     |  1 +\n read-cache.c | 11 +++++++++++\n 3 files changed, 14 insertions(+)\n\ndiff --git a/advice.c b/advice.c\nindex 5f35656409..91a55046fd 100644\n--- a/advice.c\n+++ b/advice.c\n@@ -24,6 +24,7 @@ int advice_add_embedded_repo = 1;\n int advice_ignored_hook = 1;\n int advice_waiting_for_editor = 1;\n int advice_graft_file_deprecated = 1;\n+int advice_unknown_index_extension = 1;\n int advice_checkout_ambiguous_remote_branch_name = 1;\n \n static int advice_use_color = -1;\n@@ -78,6 +79,7 @@ static struct {\n \t{ \"ignoredHook\", &advice_ignored_hook },\n \t{ \"waitingForEditor\", &advice_waiting_for_editor },\n \t{ \"graftFileDeprecated\", &advice_graft_file_deprecated },\n+\t{ \"unknownIndexExtension\", &advice_unknown_index_extension },\n \t{ \"checkoutAmbiguousRemoteBranchName\", &advice_checkout_ambiguous_remote_branch_name },\n \n \t/* make this an alias for backward compatibility */\ndiff --git a/advice.h b/advice.h\nindex 696bf0e7d2..8da0845cfc 100644\n--- a/advice.h\n+++ b/advice.h\n@@ -24,6 +24,7 @@ extern int advice_add_embedded_repo;\n extern int advice_ignored_hook;\n extern int advice_waiting_for_editor;\n extern int advice_graft_file_deprecated;\n+extern int advice_unknown_index_extension;\n extern int advice_checkout_ambiguous_remote_branch_name;\n \n int git_default_advice_config(const char *var, const char *value);\ndiff --git a/read-cache.c b/read-cache.c\nindex 002ed2c1e4..d1d903e5a1 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -1727,6 +1727,17 @@ static int read_index_extension(struct index_state *istate,\n \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n \t\t\t\t     ext);\n \t\ttrace_printf(\"ignoring %.4s extension\\n\", ext);\n+\t\tif (advice_unknown_index_extension) {\n+\t\t\twarning(_(\"ignoring optional %.4s index extension\"), ext);\n+\t\t\tadvise(_(\"This is likely due to the file having been written by a newer\\n\"\n+\t\t\t\t \"version of Git than is reading it. You can upgrade Git to\\n\"\n+\t\t\t\t \"take advantage of performance improvements from the updated\\n\"\n+\t\t\t\t \"file format.\\n\"\n+\t\t\t\t \"\\n\"\n+\t\t\t\t \"Run \\\"%s\\\"\\n\"\n+\t\t\t\t \"to suppress this message.\"),\n+\t\t\t       \"git config advice.unknownIndexExtension false\");\n+\t\t}\n \t\tbreak;\n \t}\n \treturn 0;\n-- \n2.20.0.rc0.387.gc7a69e6b6c\n\n"},{"id":"363713","messageId":"87sgzwyu94.fsf@evledraar.gmail.com","threadId":"49204","inReplyTo":"20181120061544.GF144753@google.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-11-20T09:26:47Z","receivedAt":"2018-11-20T09:26:53Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Nov 20 2018, Jonathan Nieder wrote:\n\nJust commenting here on the end-state of this since it's easier than\neach patch at a time:\n\nFirst, do we still need to be doing %.4s instead of just %s? It would be\neasier for translators / to understand what's going on if it were just\n%s. I.e. \"this is the extension name\" v.s. \"this is the first 4 bytes of\nwhatever it is...\".\n\n>  \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n>  \t\t\t\t     ext);\n\nMissing _(). Not the fault of this series, but something to fix while\nwe're at it.\n\nAlso not the fault of this series, the \"is this upper case\" test is\nunportable, but this is probably the tip of the iceberg for git not\nworking on EBCDIC systems.\n\nThis message should say something like \"Index uses the mandatory %s\nextension\" to clarify and distinguish it from the below. We don't\nunderstand the upper-case one either, but the important distinction is\nthat one is mandatory, and the other can be dropped. The two messages\nshould make this clear.\n\nAlso, having advice() for that case is even more valuable since we have\na hard error at this point. So something like:\n\n    \"This is likely due to the index having been written by a future\n    version of Git. All-lowercase index extensions are mandatory, as\n    opposed to optional all-uppercase ones which we'll drop with a\n    warning if we see them\".\n\n>  \t\ttrace_printf(\"ignoring %.4s extension\\n\", ext);\n> +\t\tif (advice_unknown_index_extension) {\n> +\t\t\twarning(_(\"ignoring optional %.4s index extension\"), ext);\n\nShould start with upper-case. Good that it says \"optional\".\n\n> +\t\t\tadvise(_(\"This is likely due to the file having been written by a newer\\n\"\n> +\t\t\t\t \"version of Git than is reading it. You can upgrade Git to\\n\"\n> +\t\t\t\t \"take advantage of performance improvements from the updated\\n\"\n> +\t\t\t\t \"file format.\\n\"\n\nLet's not promise performance improvements with this extension in a\nfuture version. We have no idea what the extension is, yeah right now\nit's going to be true for the extension that prompted this patch series,\nbut may not be in the future. So just something like this for the last\nsentence:\n\n    You can try upgrading Git to use this new index format.\n\n> +\t\t\t\t \"\\n\"\n> +\t\t\t\t \"Run \\\"%s\\\"\\n\"\n> +\t\t\t\t \"to suppress this message.\"),\n> +\t\t\t       \"git config advice.unknownIndexExtension false\");\n\nSomewhat of an aside, but if I grep:\n\n    git grep -C10 'git config advice\\..*false' -- '*.[ch]'\n\nThere's a few existing examples of this, but the majority of advice()\nmessages don't say in the message how you can turn these off. Do we\nthink this a message users would especially like to turn off? I have the\nopposite impression, it's a one-off in most cases, although not in the\ncase where an editor has an embedded git.\n\nI think it would make sense to add this sort of thing to the advice()\nAPI, i.e.:\n\n    advice_with_config_hint(_(\"<message>\"), \"unknownIndexExtension\");\n\nWhich would then know how to consistently print this advice about how to\nturn off the warning.\n"},{"id":"363739","messageId":"efa1d7fb-1da3-c093-1cb1-873a2e1c445c@gmail.com","threadId":"49204","inReplyTo":"20181120061147.GB144753@google.com","subject":"Re: [PATCH 1/5] eoie: default to not writing EOIE section","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-20T13:06:16Z","receivedAt":"2018-11-20T13:06:23Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/20/2018 1:11 AM, Jonathan Nieder wrote:\n> Since 3b1d9e04 (eoie: add End of Index Entry (EOIE) extension,\n> 2018-10-10) Git defaults to writing the new EOIE section when writing\n> out an index file.  Usually that is a good thing because it improves\n> threaded performance, but when a Git repository is shared with older\n> versions of Git, it produces a confusing warning:\n> \n>    $ git status\n>    ignoring EOIE extension\n>    HEAD detached at 371ed0defa\n>    nothing to commit, working tree clean\n> \n> Let's introduce the new index extension more gently.  First we'll roll\n> out the new version of Git that understands it, and then once\n> sufficiently many users are using such a version, we can flip the\n> default to writing it by default.\n> \n> Introduce a '[index] recordEndOfIndexEntries' configuration variable\n> to allow interested users to benefit from this index extension early.\n> \n> Signed-off-by: Jonathan Nieder <jrnieder@gmail.com>\n> ---\n> Rebased.  No other change from v1.\n> \n> As Jonathan pointed out, it would be nice to have tests here.  Ben,\n> any advice for how I could write some in a followup change?  E.g. does\n> Derrick Stolee's tracing based testing trick apply here?\n> \n>   Documentation/config/index.txt |  7 +++++++\n>   read-cache.c                   | 11 ++++++++++-\n>   t/t1700-split-index.sh         | 11 +++++++----\n>   3 files changed, 24 insertions(+), 5 deletions(-)\n> \n> diff --git a/Documentation/config/index.txt b/Documentation/config/index.txt\n> index 4b94b6bedc..8e138aba7a 100644\n> --- a/Documentation/config/index.txt\n> +++ b/Documentation/config/index.txt\n> @@ -1,3 +1,10 @@\n> +index.recordEndOfIndexEntries::\n> +\tSpecifies whether the index file should include an \"End Of Index\n> +\tEntry\" section. This reduces index load time on multiprocessor\n> +\tmachines but produces a message \"ignoring EOIE extension\" when\n> +\treading the index using Git versions before 2.20. Defaults to\n> +\t'false'.\n> +\n>   index.threads::\n>   \tSpecifies the number of threads to spawn when loading the index.\n>   \tThis is meant to reduce index load time on multiprocessor machines.\n> diff --git a/read-cache.c b/read-cache.c\n> index 4ca81286c0..1e9c772603 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -2689,6 +2689,15 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n>   \t\trollback_lock_file(lockfile);\n>   }\n>   \n> +static int record_eoie(void)\n> +{\n> +\tint val;\n\nI believe you are going to want to initialize val to 0 here as it is on \nthe stack so is not guaranteed to be zero.\n\n> +\n> +\tif (!git_config_get_bool(\"index.recordendofindexentries\", &val))\n> +\t\treturn val;\n> +\treturn 0;\n> +}\n> +\n>   /*\n>    * On success, `tempfile` is closed. If it is the temporary file\n>    * of a `struct lock_file`, we will therefore effectively perform\n> @@ -2936,7 +2945,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>   \t * read.  Write it out regardless of the strip_extensions parameter as we need it\n>   \t * when loading the shared index.\n>   \t */\n> -\tif (offset) {\n> +\tif (offset && record_eoie()) {\n>   \t\tstruct strbuf sb = STRBUF_INIT;\n>   \n>   \t\twrite_eoie_extension(&sb, &eoie_c, offset);\n> diff --git a/t/t1700-split-index.sh b/t/t1700-split-index.sh\n> index 2ac47aa0e4..0cbac64e28 100755\n> --- a/t/t1700-split-index.sh\n> +++ b/t/t1700-split-index.sh\n> @@ -25,14 +25,17 @@ test_expect_success 'enable split index' '\n>   \tgit update-index --split-index &&\n>   \ttest-tool dump-split-index .git/index >actual &&\n>   \tindexversion=$(test-tool index-version <.git/index) &&\n> +\n> +\t# NEEDSWORK: Stop hard-coding checksums.\n>   \tif test \"$indexversion\" = \"4\"\n>   \tthen\n> -\t\town=3527df833c6c100d3d1d921a9a782d62a8be4b58\n> -\t\tbase=746f7ab2ed44fb839efdfbffcf399d0b113fb4cb\n> +\t\town=432ef4b63f32193984f339431fd50ca796493569\n> +\t\tbase=508851a7f0dfa8691e9f69c7f055865389012491\n>   \telse\n> -\t\town=5e9b60117ece18da410ddecc8b8d43766a0e4204\n> -\t\tbase=4370042739b31cd17a5c5cd6043a77c9a00df113\n> +\t\town=8299b0bcd1ac364e5f1d7768efb62fa2da79a339\n> +\t\tbase=39d890139ee5356c7ef572216cebcd27aa41f9df\n>   \tfi &&\n> +\n>   \tcat >expect <<-EOF &&\n>   \town $own\n>   \tbase $base\n> \n"},{"id":"363740","messageId":"05e7df80-0dfc-c1ec-df14-c196357524f4@gmail.com","threadId":"49204","inReplyTo":"20181120061221.GC144753@google.com","subject":"Re: [PATCH 2/5] ieot: default to not writing IEOT section","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-20T13:07:49Z","receivedAt":"2018-11-20T13:07:54Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/20/2018 1:12 AM, Jonathan Nieder wrote:\n> As with EOIE, popular versions of Git do not support the new IEOT\n> extension yet.  When accessing a Git repository written by a more\n> modern version of Git, they correctly ignore the unrecognized section,\n> but in the process they loudly warn\n> \n> \tignoring IEOT extension\n> \n> resulting in confusion for users.  Introduce the index extension more\n> gently by not writing it yet in this first version with support for\n> it.  Soon, once sufficiently many users are running a modern version\n> of Git, we can flip the default so users benefit from this index\n> extension by default.\n> \n> Introduce a '[index] recordOffsetTable' configuration variable to\n> control whether the new index extension is written.\n> \n> Signed-off-by: Jonathan Nieder <jrnieder@gmail.com>\n> ---\n> As with patch 1/5, no change from v1 other than rebasing.\n> \n>   Documentation/config/index.txt |  7 +++++++\n>   read-cache.c                   | 11 ++++++++++-\n>   2 files changed, 17 insertions(+), 1 deletion(-)\n> \n> diff --git a/Documentation/config/index.txt b/Documentation/config/index.txt\n> index 8e138aba7a..de44183235 100644\n> --- a/Documentation/config/index.txt\n> +++ b/Documentation/config/index.txt\n> @@ -5,6 +5,13 @@ index.recordEndOfIndexEntries::\n>   \treading the index using Git versions before 2.20. Defaults to\n>   \t'false'.\n>   \n> +index.recordOffsetTable::\n> +\tSpecifies whether the index file should include an \"Index Entry\n> +\tOffset Table\" section. This reduces index load time on\n> +\tmultiprocessor machines but produces a message \"ignoring IEOT\n> +\textension\" when reading the index using Git versions before 2.20.\n> +\tDefaults to 'false'.\n> +\n>   index.threads::\n>   \tSpecifies the number of threads to spawn when loading the index.\n>   \tThis is meant to reduce index load time on multiprocessor machines.\n> diff --git a/read-cache.c b/read-cache.c\n> index 1e9c772603..f3d5638d9e 100644\n> --- a/read-cache.c\n> +++ b/read-cache.c\n> @@ -2698,6 +2698,15 @@ static int record_eoie(void)\n>   \treturn 0;\n>   }\n>   \n> +static int record_ieot(void)\n> +{\n> +\tint val;\n> +\n\nInitialize stack val to zero to ensure proper default.\n\n> +\tif (!git_config_get_bool(\"index.recordoffsettable\", &val))\n> +\t\treturn val;\n> +\treturn 0;\n> +}\n> +\n>   /*\n>    * On success, `tempfile` is closed. If it is the temporary file\n>    * of a `struct lock_file`, we will therefore effectively perform\n> @@ -2761,7 +2770,7 @@ static int do_write_index(struct index_state *istate, struct tempfile *tempfile,\n>   \telse\n>   \t\tnr_threads = 1;\n>   \n> -\tif (nr_threads != 1) {\n> +\tif (nr_threads != 1 && record_ieot()) {\n>   \t\tint ieot_blocks, cpus;\n>   \n>   \t\t/*\n> \n"},{"id":"363742","messageId":"20181120132151.GA30222@szeder.dev","threadId":"49204","inReplyTo":"efa1d7fb-1da3-c093-1cb1-873a2e1c445c@gmail.com","subject":"Re: [PATCH 1/5] eoie: default to not writing EOIE section","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2018-11-20T13:21:51Z","receivedAt":"2018-11-20T13:21:56Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Tue, Nov 20, 2018 at 08:06:16AM -0500, Ben Peart wrote:\n> >diff --git a/read-cache.c b/read-cache.c\n> >index 4ca81286c0..1e9c772603 100644\n> >--- a/read-cache.c\n> >+++ b/read-cache.c\n> >@@ -2689,6 +2689,15 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n> >  \t\trollback_lock_file(lockfile);\n> >  }\n> >+static int record_eoie(void)\n> >+{\n> >+\tint val;\n> \n> I believe you are going to want to initialize val to 0 here as it is on the\n> stack so is not guaranteed to be zero.\n\nThe git_config_get_bool() call below will initialize it anyway.\n\n> >+\n> >+\tif (!git_config_get_bool(\"index.recordendofindexentries\", &val))\n> >+\t\treturn val;\n> >+\treturn 0;\n> >+}\n"},{"id":"363743","messageId":"f4c28f3f-f3e0-8a23-ea12-70b4fef5d96c@gmail.com","threadId":"49204","inReplyTo":"20181120061426.GE144753@google.com","subject":"Re: [PATCH 4/5] index: make index.threads=true enable ieot and eoie","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-20T13:24:35Z","receivedAt":"2018-11-20T13:24:41Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/20/2018 1:14 AM, Jonathan Nieder wrote:\n> If a user explicitly sets\n> \n> \t[index]\n> \t\tthreads = true\n> \n> to read the index using multiple threads, ensure that index writes\n> include the offset table by default to make that possible.  This\n> ensures that the user's intent of turning on threading is respected.\n> \n> In other words, permit the following configurations:\n> \n> - index.threads and index.recordOffsetTable unspecified: do not write\n>    the offset table yet (to avoid alarming the user with \"ignoring IEOT\n>    extension\" messages when an older version of Git accesses the\n>    repository) but do make use of multiple threads to read the index if\n>    the supporting offset table is present.\n> \n>    This can also be requested explicitly by setting index.threads=true,\n>    0, or >1 and index.recordOffsetTable=false.\n> \n> - index.threads=false or 1: do not write the offset table, and do not\n>    make use of the offset table.\n> \n>    One can set index.recordOffsetTable=false as well, to be more\n>    explicit.\n> \n> - index.threads=true, 0, or >1 and index.recordOffsetTable unspecified:\n>    write the offset table and make use of threads at read time.\n> \n>    This can also be requested by setting index.threads=true, 0, >1, or\n>    unspecified and index.recordOffsetTable=true.\n> \n> Fortunately the complication is temporary: once most Git installations\n> have upgraded to a version with support for the IEOT and EOIE\n> extensions, we can flip the defaults for index.recordEndOfIndexEntries\n> and index.recordOffsetTable to true and eliminate the settings.\n> \n\nThis looks good.  I think this provides good default behavior while \nenabling fine grained control to those who want/need it.\n\nI'm looking forward to the day when we can turn it back on by default so \nthat people can take advantage of the speed improvements.\n\n"},{"id":"363744","messageId":"cabd2e37-7389-ac74-6626-629eab7da53f@gmail.com","threadId":"49204","inReplyTo":"87sgzwyu94.fsf@evledraar.gmail.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-20T13:30:50Z","receivedAt":"2018-11-20T13:30:56Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/20/2018 4:26 AM, Ævar Arnfjörð Bjarmason wrote:\n> \n> On Tue, Nov 20 2018, Jonathan Nieder wrote:\n> \n> Just commenting here on the end-state of this since it's easier than\n> each patch at a time:\n> \n> First, do we still need to be doing %.4s instead of just %s? It would be\n> easier for translators / to understand what's going on if it were just\n> %s. I.e. \"this is the extension name\" v.s. \"this is the first 4 bytes of\n> whatever it is...\".\n> \n>>   \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n>>   \t\t\t\t     ext);\n> \n> Missing _(). Not the fault of this series, but something to fix while\n> we're at it.\n> \n> Also not the fault of this series, the \"is this upper case\" test is\n> unportable, but this is probably the tip of the iceberg for git not\n> working on EBCDIC systems.\n> \n> This message should say something like \"Index uses the mandatory %s\n> extension\" to clarify and distinguish it from the below. We don't\n> understand the upper-case one either, but the important distinction is\n> that one is mandatory, and the other can be dropped. The two messages\n> should make this clear.\n> \n> Also, having advice() for that case is even more valuable since we have\n> a hard error at this point. So something like:\n> \n>      \"This is likely due to the index having been written by a future\n>      version of Git. All-lowercase index extensions are mandatory, as\n>      opposed to optional all-uppercase ones which we'll drop with a\n>      warning if we see them\".\n> \n\nI agree that we should have different messages for mandatory and \noptional extensions.  I don't think we should try and educate the end \nuser on the implementation detail that git makes lower cases mandatory \nand upper case optional (ie drop the 'All-lowercase...\" part).  They \nwill never see the lower vs upper case difference and can't do anything \nabout it anyway.\n\n>>   \t\ttrace_printf(\"ignoring %.4s extension\\n\", ext);\n>> +\t\tif (advice_unknown_index_extension) {\n>> +\t\t\twarning(_(\"ignoring optional %.4s index extension\"), ext);\n> \n> Should start with upper-case. Good that it says \"optional\".\n> \n>> +\t\t\tadvise(_(\"This is likely due to the file having been written by a newer\\n\"\n>> +\t\t\t\t \"version of Git than is reading it. You can upgrade Git to\\n\"\n>> +\t\t\t\t \"take advantage of performance improvements from the updated\\n\"\n>> +\t\t\t\t \"file format.\\n\"\n> \n> Let's not promise performance improvements with this extension in a\n> future version. We have no idea what the extension is, yeah right now\n> it's going to be true for the extension that prompted this patch series,\n> but may not be in the future. So just something like this for the last\n> sentence:\n> \n>      You can try upgrading Git to use this new index format.\n\nAgree - not all are guaranteed to be perf related.\n\n> \n>> +\t\t\t\t \"\\n\"\n>> +\t\t\t\t \"Run \\\"%s\\\"\\n\"\n>> +\t\t\t\t \"to suppress this message.\"),\n>> +\t\t\t       \"git config advice.unknownIndexExtension false\");\n> \n> Somewhat of an aside, but if I grep:\n> \n>      git grep -C10 'git config advice\\..*false' -- '*.[ch]'\n> \n> There's a few existing examples of this, but the majority of advice()\n> messages don't say in the message how you can turn these off. Do we\n> think this a message users would especially like to turn off? I have the\n> opposite impression, it's a one-off in most cases, although not in the\n> case where an editor has an embedded git.\n> \n> I think it would make sense to add this sort of thing to the advice()\n> API, i.e.:\n> \n>      advice_with_config_hint(_(\"<message>\"), \"unknownIndexExtension\");\n> \n> Which would then know how to consistently print this advice about how to\n> turn off the warning.\n> \n\nI like this.  I personally never knew you could turn of the \"spent xxx \nseconds finding untracked files\" advice until I worked on this patch \nseries. This would help make that feature more discoverable.\n"},{"id":"363745","messageId":"7a3bf106-a8ce-c64a-3015-d8543feee4d9@gmail.com","threadId":"49204","inReplyTo":"20181120061147.GB144753@google.com","subject":"Re: [PATCH 1/5] eoie: default to not writing EOIE section","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-20T15:01:45Z","receivedAt":"2018-11-20T15:01:51Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/20/2018 1:11 AM, Jonathan Nieder wrote:\n> Since 3b1d9e04 (eoie: add End of Index Entry (EOIE) extension,\n> 2018-10-10) Git defaults to writing the new EOIE section when writing\n> out an index file.  Usually that is a good thing because it improves\n> threaded performance, but when a Git repository is shared with older\n> versions of Git, it produces a confusing warning:\n> \n>    $ git status\n>    ignoring EOIE extension\n>    HEAD detached at 371ed0defa\n>    nothing to commit, working tree clean\n> \n> Let's introduce the new index extension more gently.  First we'll roll\n> out the new version of Git that understands it, and then once\n> sufficiently many users are using such a version, we can flip the\n> default to writing it by default.\n> \n> Introduce a '[index] recordEndOfIndexEntries' configuration variable\n> to allow interested users to benefit from this index extension early.\n> \n> Signed-off-by: Jonathan Nieder <jrnieder@gmail.com>\n> ---\n> Rebased.  No other change from v1.\n> \n> As Jonathan pointed out, it would be nice to have tests here.  Ben,\n> any advice for how I could write some in a followup change?  E.g. does\n> Derrick Stolee's tracing based testing trick apply here?\n> \n\nI suppose a 'test-dump-eoie' could be written along the lines of \ntest-dump-fsmonitor or test-dump-untracked-cache.  Unlike those, there \nisn't much state to dump other than the existence of the extension and \nthe offset.  That could be used to test that the new settings are \nworking properly.\n\n"},{"id":"363790","messageId":"xmqqefbf9t4j.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"cabd2e37-7389-ac74-6626-629eab7da53f@gmail.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-21T00:22:36Z","receivedAt":"2018-11-21T00:22:46Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Peart <peartben@gmail.com> writes:\n\n>> This message should say something like \"Index uses the mandatory %s\n>> extension\" to clarify and distinguish it from the below. We don't\n>> understand the upper-case one either, but the important distinction is\n>> that one is mandatory, and the other can be dropped. The two messages\n>> should make this clear.\n>>\n>> Also, having advice() for that case is even more valuable since we have\n>> a hard error at this point. So something like:\n>>\n>>      \"This is likely due to the index having been written by a future\n>>      version of Git. All-lowercase index extensions are mandatory, as\n>>      opposed to optional all-uppercase ones which we'll drop with a\n>>      warning if we see them\".\n>>\n>\n> I agree that we should have different messages for mandatory and\n> optional extensions.  I don't think we should try and educate the end\n> user on the implementation detail that git makes lower cases mandatory\n> and upper case optional (ie drop the 'All-lowercase...\" part).  They\n> will never see the lower vs upper case difference and can't do\n> anything about it anyway.\n\nI agree that the \"warn and continue\" message should say \"optional\"\n(meaning: safe to ignore but you would want to take note) while\n\"cannot continue\" message should say something different.\n\nI do not mind a more verbose error message when we saw unknown but\nrequired extension, but unlike the \"warn and continue\" case, the\nprogram will stop and die with such an error right there, so I am\nnot sure if it is worth allowing to tone it down by putting some\npart of the verbosity behind the advise() mechanism.\n\n>>>   \t\ttrace_printf(\"ignoring %.4s extension\\n\", ext);\n>>> +\t\tif (advice_unknown_index_extension) {\n>>> +\t\t\twarning(_(\"ignoring optional %.4s index extension\"), ext);\n\nSo from that point of view, the distinction between this message and this one\n\n>>>   \t\t\treturn error(\"index uses %.4s extension, which we do not understand\",\n>>>   \t\t\t\t     ext);\n\nis halfway there.  The message needs to anticipate and answer an\nend-user reaction: \"we do not understand\" so what?\n\nI am still puzzled by the insistence of 3/5 and this step that wants\nto kill the coalmine canary.  But I am even more puzzled by the\nfirst two steps that want to disable the two optional extensions.\n\nWhat's so different this time with the new optional extensions?\n\nThe other early optional extensions like cache-tree or resolve-undo\nwere added unconditionally and by definition appeared much earlier\nin git-core than any other Git reimplementations.  verbody who\nrecorded the fact that s/he resolved merge conflicts got REUC, and\nwe would have given warning when an older Git did not understand\nthese extensions [*1*].  We knudged users to more modern Git by\npreparing the old Gits to warn when there are unknown extensions,\neither by upgrading their Git themselves, or by bugging their\ntoolsmiths.  Nobody complained to propose to rip the messages like\nthis round.  This series has a strong smell of pushing back by the\ntoolsmiths who refuse to promptly upgrade to help their users, and\nthat is why I do not feel entirely happy with this series.\n\n\n[Footnote]\n\n *1* A Git that did not understand TREE would have been silent, as\n  it was the first extension and that was the first time we became\n  aware of the need to warn unknown extensions.\n"},{"id":"363792","messageId":"20181121003912.GC149929@google.com","threadId":"49204","inReplyTo":"xmqqefbf9t4j.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-21T00:39:12Z","receivedAt":"2018-11-21T00:39:17Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nJunio C Hamano wrote:\n\n> I am still puzzled by the insistence of 3/5 and this step that wants\n> to kill the coalmine canary.  But I am even more puzzled by the\n> first two steps that want to disable the two optional extensions.\n>\n> What's so different this time with the new optional extensions?\n>\n> The other early optional extensions like cache-tree or resolve-undo\n> were added unconditionally and by definition appeared much earlier\n> in git-core than any other Git reimplementations.  Everbody who\n> recorded the fact that s/he resolved merge conflicts got REUC, and\n> we would have given warning when an older Git did not understand\n> these extensions [*1*].  We knudged users to more modern Git by\n> preparing the old Gits to warn when there are unknown extensions,\n> either by upgrading their Git themselves, or by bugging their\n> toolsmiths.  Nobody complained to propose to rip the messages like\n> this round.  This series has a strong smell of pushing back by the\n> toolsmiths who refuse to promptly upgrade to help their users, and\n> that is why I do not feel entirely happy with this series.\n\nI acknowledge your puzzlement.  I'm not sure what to do about it.\n\nThere are a few significant differences from the REUC case:\n\n 1. This happens whenever the index is refreshed.  REUC, as you\n    mentioned, only affected resolutions of conflicted merges.  So\n    users ran into it less often.\n\n 2. I never ran into the REUC case.  If I had, I would have sent the\n    same patch then.\n\n 3. Time has passed and people's standards may have gone up.\n\nI wish I had been around when the message was added in the first\nplace, so that I could have provided the same feedback about the\nmessage then.  But I do not think that that should be held against me.\nI'm describing a real user problem.\n\nAre the commit messages unclear?  Is there some missing use case that\nthis version of the patch misses?\n\nI don't *think* you intend to say \"sure, you got user reports, but\n(those users are wrong | those users are not real | you are not\ninterpreting those users correctly)\", but that is what I am hearing.\nOn the other hand, I don't want to discourage useful review feedback,\nand I think adding the advise() call was a real improvement.  I'm just\ngetting confused about why I am still not being heard.\n\nJonathan\n"},{"id":"363793","messageId":"20181121004423.GD149929@google.com","threadId":"49204","inReplyTo":"20181121003912.GC149929@google.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-21T00:44:23Z","receivedAt":"2018-11-21T00:44:27Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"onathan Nieder wrote:\n> Junio C Hamano wrote:\n\n>> I am still puzzled by the insistence of 3/5 and this step that wants\n>> to kill the coalmine canary.  But I am even more puzzled by the\n>> first two steps that want to disable the two optional extensions.\n[...]\n> I acknowledge your puzzlement.  I'm not sure what to do about it.\n>\n> There are a few significant differences from the REUC case:\n>\n>  1. This happens whenever the index is refreshed.  REUC, as you\n>     mentioned, only affected resolutions of conflicted merges.  So\n>     users ran into it less often.\n>\n>  2. I never ran into the REUC case.  If I had, I would have sent the\n>     same patch then.\n>\n>  3. Time has passed and people's standards may have gone up.\n>\n> I wish I had been around when the message was added in the first\n> place, so that I could have provided the same feedback about the\n> message then.  But I do not think that that should be held against me.\n> I'm describing a real user problem.\n\nAnd to be clear, it is the first two patches that address the\nimmediate user problem.  Whatever improvements we make to the warning\nmessage today, we cannot retroactively change the other versions of\nGit that users are using that want to access the same repository.\n\nJonathan\n"},{"id":"363794","messageId":"20181121010309.GE149929@google.com","threadId":"49204","inReplyTo":"xmqqefbf9t4j.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-21T01:03:09Z","receivedAt":"2018-11-21T01:03:14Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n\n>              This series has a strong smell of pushing back by the\n> toolsmiths who refuse to promptly upgrade to help their users, and\n> that is why I do not feel entirely happy with this series.\n\nLast reply, I promise. :)\n\nThis sentence might have the key to the misunderstanding.  Let me say\na little more about where this showed up in the internal deployment\nhere, to clarify things a little.\n\nAt Google we deploy snapshots of the \"next\" branch approximately\nweekly so that we can find problems early before they affect a\npublished release.  We rely on the ability to roll back quickly when a\nproblem is discovered, and we might care more about compatibility than\nsome others because of that.\n\nA popular tool within Google has a bundled copy of Git (also a\nsnapshot of the \"next\" branch, but from a few weeks prior) and when we\ndeployed Git with the EOIE and IEOT extensions, users of that tool\nvery quickly reported the mysterious message.\n\nThat said, the maintainers of that tool did not complain at all, so\nhopefully I can allay your worries about toolsmiths pushing back.\nOnce the problem reached my attention (a few days later than I would\nhave liked it to), the Git team at Google knew that we could not roll\nback and were certainly alarmed about what that means about our\nability to cope with other problems should we need to.  But we were\nable to quickly update that popular tool --- no issue.\n\nInstead, we ran into a number of other users running into the same\nproblem, when sharing repositories between machines using sshfs, etc.\nThat, plus the aforementioned inability to roll back Git if we need\nto, meant that this was a serious issue so we quickly addressed it in\nthe internal installation.\n\nIn general, we haven't had much trouble getting people to use Git\n2.19.1 or newer.  So the problem here does not have to do with users\nbeing slow to upgrade.\n\nInstead, it's simply that upgrading Git should not cause the older,\nwidely deployed version of Git to complain about the repositories it\nacts on.  That's a recipe for difficult debugging situations, it can\nlead to people upgrading less quickly and reporting bugs later, and\nall in all it's a bad situation to be in.  I've used tools like\nSubversion that would upgrade repositories so they are unusable by the\nprevious version and experienced all of these problems.\n\nSo I consider it important *to Git upstream* to handle this well in\nthe Git 2.20 release.  We can flip the default soon after, even as\nsoon as 2.21.\n\nMoreover, I am not the only one who ran into this --- e.g. from [1],\n2018-10-19:\n\n  17:10 <peff> jrnieder: Yes, I noticed that annoyance myself. ;)\n  17:11 <newren> Yeah, I saw that message a few times and was slightly\n                 annoyed as well.\n\nNow, a meta point.  Throughout this discussion, I have been hoping for\nsome acknowledgement of the problem --- e.g. an \"I am sympathetic to\nwhat you are trying to do, but <X>\".  I wasn't able to find that, and\nthat is part of what contributed to the feeling of not being heard.\n\nThanks for your patient explanations, and hope that helps,\nJonathan\n\n[1] https://colabti.org/irclogger/irclogger_log/git-devel?date=2018-10-19#l114\n"},{"id":"363803","messageId":"xmqqpnuz83f9.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20181121010309.GE149929@google.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-21T04:23:06Z","receivedAt":"2018-11-21T04:25:54Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> Now, a meta point.  Throughout this discussion, I have been hoping for\n> some acknowledgement of the problem --- e.g. an \"I am sympathetic to\n> what you are trying to do, but <X>\".  I wasn't able to find that, and\n> that is part of what contributed to the feeling of not being heard.\n\nI had little sympathy for what you were trying to do, i.e. killing\nthe coalmine canary that warns users about using older version of\nGit when there clearly is a sign that a newer one is available to\nthem and they have already used it; there was no problem to be\nacknowledged.\n\nNot before \"why is it different this time?\" question gets answered\nanyway.\n\nAnd seeing the same \"let's not enable the new extension\" patch again\nwithout much improved justification contributed greatly to the\nfeeling of not being heard.  The feeling is mutual.\n\n"},{"id":"363807","messageId":"20181121045700.GC245855@google.com","threadId":"49204","inReplyTo":"xmqqpnuz83f9.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-21T04:57:00Z","receivedAt":"2018-11-21T04:57:05Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n\n>> Now, a meta point.  Throughout this discussion, I have been hoping for\n>> some acknowledgement of the problem --- e.g. an \"I am sympathetic to\n>> what you are trying to do, but <X>\".  I wasn't able to find that, and\n>> that is part of what contributed to the feeling of not being heard.\n[...]\n> And seeing the same \"let's not enable the new extension\" patch again\n> without much improved justification contributed greatly to the\n> feeling of not being heard.\n\nThanks for the pointer.  I had not understood before that you were\nunhappy with those commit messages.\n\nThe commit message describes symptoms and the motivation for the\nchange.  I was confused at the original replies to patch 1 and 2 that\nseemed to be more about patch 3; patch 1 and 2 are meaningful without\npatch 3, so it would be odd to include a justification for patch 3 in\ntheir commit message.\n\nThat said, it sounds like their commit messages are not adequate.  I'd\nappreciate help from someone else to improve them.\n\n>                              The feeling is mutual.\n\nI was trying to diagnose what was going wrong with the conversation so\nas to move things forward on a better footing.  It seems I only\nescalated things more. :(\n\nSorry about that, and I hope there's some way to move forward.\n\nWhat is the best way to handle this?  I am feeling somewhat burnt by\nthis review process.  If Ben and I, working together, are able to come\nup with a series that we both like, will you consider it for 2.20?  Is\nthere some other trusted contributor, such as Peff or Duy, that you\nwould trust to represent your wishes so I can pursue their Reviewed-by\nwithout risking getting burnt in the same way again?\n\nJonathan\n"},{"id":"363808","messageId":"xmqq36rv81nr.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20181121003912.GC149929@google.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-21T05:01:12Z","receivedAt":"2018-11-21T05:01:16Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> I don't *think* you intend to say \"sure, you got user reports, but\n> (those users are wrong | those users are not real | you are not\n> interpreting those users correctly)\", but that is what I am hearing.\n\nWhat I have been saying is \"we are sending a wrong message to those\nusers by not clearly saying 'optional' (i.e. it is OK for your Git\nnot to understand this optional bits of information---you do not\nhave to get alarmed immediately) and also not hinting where that\noptional thing comes from (i.e. if users realized they come from the\nfuture, the coalmine canary message will serve its purpose of\nreminding them that a newer Git is not just available but has been\nused already in their repository and help them to rectify the\nsituation sooner)\".\n\nAs the deployed versions of Git will keep sending the wrong message,\nI do not mind applying 1/5 and 2/5, given especially that Ben seems\nto be OK with the plan.  I however do not think 3 thru 5 is ready\nyet with this round---there were some discussions on phrasing in\nthis thread.\n"},{"id":"363809","messageId":"20181121050458.GD245855@google.com","threadId":"49204","inReplyTo":"xmqq36rv81nr.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-11-21T05:04:58Z","receivedAt":"2018-11-21T05:05:04Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n\n>> I don't *think* you intend to say \"sure, you got user reports, but\n>> (those users are wrong | those users are not real | you are not\n>> interpreting those users correctly)\", but that is what I am hearing.\n>\n> What I have been saying is \"we are sending a wrong message to those\n> users by not clearly saying 'optional' (i.e. it is OK for your Git\n> not to understand this optional bits of information---you do not\n> have to get alarmed immediately) and also not hinting where that\n> optional thing comes from (i.e. if users realized they come from the\n> future, the coalmine canary message will serve its purpose of\n> reminding them that a newer Git is not just available but has been\n> used already in their repository and help them to rectify the\n> situation sooner)\".\n>\n> As the deployed versions of Git will keep sending the wrong message,\n> I do not mind applying 1/5 and 2/5, given especially that Ben seems\n> to be OK with the plan.  I however do not think 3 thru 5 is ready\n> yet with this round---there were some discussions on phrasing in\n> this thread.\n\nThanks much --- that helps a lot.\n\nWould you mind taking patch 4/5 as well?  (It's a tweak to the\nconfiguration introduced in patches 1 and 2 that addresses a concern\nBen Peart had.)\n\nAs for patches 3 and 5, I agree.  In particular, patch 5 needs an\ns/performance//, and it seems that the commit messages need some work\nas well.\n\nSorry for getting the conversation in the wrong direction, and I'm\nglad to hear we have a good way forward.\n\nSincerely,\nJonathan\n"},{"id":"363810","messageId":"xmqqtvkb6mfc.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"xmqq36rv81nr.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-21T05:15:35Z","receivedAt":"2018-11-21T05:15:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> As the deployed versions of Git will keep sending the wrong message,\n> I do not mind applying 1/5 and 2/5, given especially that Ben seems\n> to be OK with the plan.  I however do not think 3 thru 5 is ready\n> yet with this round---there were some discussions on phrasing in\n> this thread.\n\nI ran out of time looking at the surrounding code, but I think 1, 2\nand 4 form a set that would give us immediate benefit to fast track\nto the upcoming release.\n\nI do not know if it makes sense to have 3 and 5 separate; I suspect\na single patch that does \"clarify the warning, and allow those who\nhave no choice in which version of Git to choose squelch it\" would\nsuffice.\n\nThe phrasing in 5 received a couple of good concrete suggestions\nalready, so it is not ready in its current form but need a bit of\nwordsmithing.  I also do not think a new \"trace_printf\" would\nparticularly help.  If I stared it a lot longer, I may spot more\nissues in it.\n\nBut what that step does primarily would help long after the upcoming\nrelease and in that sense can wait a bit longer than 1, 2 & 4 (which\nI am hoping can be merged in -rc1).\n"},{"id":"363811","messageId":"xmqqpnuz6lpm.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"xmqqtvkb6mfc.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-21T05:31:01Z","receivedAt":"2018-11-21T05:31:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> I do not know if it makes sense to have 3 and 5 separate; I suspect\n> a single patch that does \"clarify the warning, and allow those who\n> have no choice in which version of Git to choose squelch it\" would\n> suffice.\n\nI actually do not mind two patches for these, but I think the\nseparation presented in the series is wrong (first to kill the\ncanary completely, and then add it as if it were a completely\nseparate advice).  \n\nIt would make more sense, at least to me, if \n\n - the earlier step is to clarify the warning on two points\n   (i.e. this is safe to ignore, but you may want to know that you\n   are using a stale Git when we see evidence that a newer one has\n   already been used here) and then\n\n - the later step is to optionally make it possible to squelch it\n   for those who do not have control over what version of Git they\n   are allowed to run.\n\nBut again, a single patch to do all of that is also fine.\n"},{"id":"363813","messageId":"87o9aizsjz.fsf@evledraar.gmail.com","threadId":"49204","inReplyTo":"20181121010309.GE149929@google.com","subject":"Re: [PATCH 5/5] index: offer advice for unknown index extensions","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-11-21T09:30:24Z","receivedAt":"2018-11-21T09:30:32Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Nov 21 2018, Jonathan Nieder wrote:\n\n> Junio C Hamano wrote:\n>\n>>              This series has a strong smell of pushing back by the\n>> toolsmiths who refuse to promptly upgrade to help their users, and\n>> that is why I do not feel entirely happy with this series.\n>\n> Last reply, I promise. :)\n>\n> This sentence might have the key to the misunderstanding.  Let me say\n> a little more about where this showed up in the internal deployment\n> here, to clarify things a little.\n>\n> At Google we deploy snapshots of the \"next\" branch approximately\n> weekly so that we can find problems early before they affect a\n> published release.  We rely on the ability to roll back quickly when a\n> problem is discovered, and we might care more about compatibility than\n> some others because of that.\n>\n> A popular tool within Google has a bundled copy of Git (also a\n> snapshot of the \"next\" branch, but from a few weeks prior) and when we\n> deployed Git with the EOIE and IEOT extensions, users of that tool\n> very quickly reported the mysterious message.\n>\n> That said, the maintainers of that tool did not complain at all, so\n> hopefully I can allay your worries about toolsmiths pushing back.\n> Once the problem reached my attention (a few days later than I would\n> have liked it to), the Git team at Google knew that we could not roll\n> back and were certainly alarmed about what that means about our\n> ability to cope with other problems should we need to.  But we were\n> able to quickly update that popular tool --- no issue.\n>\n> Instead, we ran into a number of other users running into the same\n> problem, when sharing repositories between machines using sshfs, etc.\n> That, plus the aforementioned inability to roll back Git if we need\n> to, meant that this was a serious issue so we quickly addressed it in\n> the internal installation.\n>\n> In general, we haven't had much trouble getting people to use Git\n> 2.19.1 or newer.  So the problem here does not have to do with users\n> being slow to upgrade.\n>\n> Instead, it's simply that upgrading Git should not cause the older,\n> widely deployed version of Git to complain about the repositories it\n> acts on.  That's a recipe for difficult debugging situations, it can\n> lead to people upgrading less quickly and reporting bugs later, and\n> all in all it's a bad situation to be in.  I've used tools like\n> Subversion that would upgrade repositories so they are unusable by the\n> previous version and experienced all of these problems.\n>\n> So I consider it important *to Git upstream* to handle this well in\n> the Git 2.20 release.  We can flip the default soon after, even as\n> soon as 2.21.\n>\n> Moreover, I am not the only one who ran into this --- e.g. from [1],\n> 2018-10-19:\n>\n>   17:10 <peff> jrnieder: Yes, I noticed that annoyance myself. ;)\n>   17:11 <newren> Yeah, I saw that message a few times and was slightly\n>                  annoyed as well.\n>\n> Now, a meta point.  Throughout this discussion, I have been hoping for\n> some acknowledgement of the problem --- e.g. an \"I am sympathetic to\n> what you are trying to do, but <X>\".  I wasn't able to find that, and\n> that is part of what contributed to the feeling of not being heard.\n>\n> Thanks for your patient explanations, and hope that helps,\n> Jonathan\n\nI think it makes total sense to fix this. I had not spotted this myself\nsince I tend to just roll forward and only use one version of git on one\nsystem, but fixing this makes sense.\n"},{"id":"363842","messageId":"20181121164619.GA13860@sigill.intra.peff.net","threadId":"49204","inReplyTo":"20181120132151.GA30222@szeder.dev","subject":"Re: [PATCH 1/5] eoie: default to not writing EOIE section","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-11-21T16:46:20Z","receivedAt":"2018-11-21T16:46:23Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Nov 20, 2018 at 02:21:51PM +0100, SZEDER Gábor wrote:\n\n> On Tue, Nov 20, 2018 at 08:06:16AM -0500, Ben Peart wrote:\n> > >diff --git a/read-cache.c b/read-cache.c\n> > >index 4ca81286c0..1e9c772603 100644\n> > >--- a/read-cache.c\n> > >+++ b/read-cache.c\n> > >@@ -2689,6 +2689,15 @@ void update_index_if_able(struct index_state *istate, struct lock_file *lockfile\n> > >  \t\trollback_lock_file(lockfile);\n> > >  }\n> > >+static int record_eoie(void)\n> > >+{\n> > >+\tint val;\n> > \n> > I believe you are going to want to initialize val to 0 here as it is on the\n> > stack so is not guaranteed to be zero.\n> \n> The git_config_get_bool() call below will initialize it anyway.\n\nYes, there are two ways to write this. With a conditional to initialize\nand return or to return the default, as we have here:\n\n> > >+\tif (!git_config_get_bool(\"index.recordendofindexentries\", &val))\n> > >+\t\treturn val;\n> > >+\treturn 0;\n\nOr initialize the default ahead of time, and rely on the function not to\nmodify it when the entry is missing:\n\n  int val = 0;\n  git_config_get_bool(\"index.whatever\", &val);\n  return val;\n\nI think either is perfectly fine, but since I also had to look at it\ntwice to make sure it was doing the right thing, I figured it is worth\nmentioning as a possible style/convention thing we may want to decide\non.\n\n-Peff\n"},{"id":"363856","messageId":"xmqqefbe546m.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"20181121164619.GA13860@sigill.intra.peff.net","subject":"Re: [PATCH 1/5] eoie: default to not writing EOIE section","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-22T00:47:13Z","receivedAt":"2018-11-22T00:47:18Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Yes, there are two ways to write this. With a conditional to initialize\n> and return or to return the default, as we have here:\n>\n>> > >+\tif (!git_config_get_bool(\"index.recordendofindexentries\", &val))\n>> > >+\t\treturn val;\n>> > >+\treturn 0;\n>\n> Or initialize the default ahead of time, and rely on the function not to\n> modify it when the entry is missing:\n>\n>   int val = 0;\n>   git_config_get_bool(\"index.whatever\", &val);\n>   return val;\n>\n> I think either is perfectly fine, but since I also had to look at it\n> twice to make sure it was doing the right thing, I figured it is worth\n> mentioning as a possible style/convention thing we may want to decide\n> on.\n\nI too think either is fine, and both rely on the git_config_get_*()\nto modify the value return only when it sees that it is set.\n\nI'd choose the latter when the default value is simple, as the\nreader does not have to even know what the return value from the\ngit_config_get_*() function means to follow what is going on.\n\nOn the other hand, the former (i.e. the original by Jonathan) is\nmore flexible, and it makes it possible to write a piece of code,\nwhich computes a default that is expensive to build only when\nnecessary, in the most natural way.  The readers do need to be aware\nof how the functin signals \"I didn't get anything\" with its return\nvalue, though.\n\nI do not mind standardizing on the latter, though.  A caller with an\nexpensive default can initialize val to an impossible \"sentinel\"\nvalue that signals the fact that git_config_get_*() did not get\nanything, as long as the type has a natural sentinel (like -1 for a\nbool to signal \"unset\"), and code that comes either immediately\nafter git_config_get_*() or much much later in the control flow can\ncheck for the sentinel to see if it needs to compute the expensive\ndefault.\n\n"},{"id":"364096","messageId":"CAGZ79kbaPKaCFGGXnbNchvk=1Q4Q5Hgt2hXOhcGo6pVwquhaEg@mail.gmail.com","threadId":"49204","inReplyTo":"05e7df80-0dfc-c1ec-df14-c196357524f4@gmail.com","subject":"Re: [PATCH 2/5] ieot: default to not writing IEOT section","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-11-26T19:59:16Z","receivedAt":"2018-11-26T19:59:31Z","isPatch":true,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"> > +static int record_ieot(void)\n> > +{\n> > +     int val;\n> > +\n>\n> Initialize stack val to zero to ensure proper default.\n\nI don't think that is needed here, as we only use `val` when\nwe first write to it via git_config_get_bool.\n\nDid you spot this via code review and thought of\ndefensive programming or is there a tool that\nhas a false positive here?\n\n>\n> > +     if (!git_config_get_bool(\"index.recordoffsettable\", &val))\n> > +             return val;\n> > +     return 0;\n> > +}\n"},{"id":"364100","messageId":"7e06ec44-a3a0-fa38-75a7-7b875ae0679e@gmail.com","threadId":"49204","inReplyTo":"CAGZ79kbaPKaCFGGXnbNchvk=1Q4Q5Hgt2hXOhcGo6pVwquhaEg@mail.gmail.com","subject":"Re: [PATCH 2/5] ieot: default to not writing IEOT section","fromName":"Ben Peart","fromEmail":"peartben@gmail.com","sentAt":"2018-11-26T21:47:56Z","receivedAt":"2018-11-26T21:48:03Z","isPatch":true,"sender":{"key":"benpeart@microsoft.com","avatar":"https://avatars.githubusercontent.com/u/15252029?v=4"},"body":"\n\nOn 11/26/2018 2:59 PM, Stefan Beller wrote:\n>>> +static int record_ieot(void)\n>>> +{\n>>> +     int val;\n>>> +\n>>\n>> Initialize stack val to zero to ensure proper default.\n> \n> I don't think that is needed here, as we only use `val` when\n> we first write to it via git_config_get_bool.\n> \n> Did you spot this via code review and thought of\n> defensive programming or is there a tool that\n> has a false positive here?\n> \n\nCode review and defensive programming.  I had to review the code in \ngit_config_get_bool() to see if it always initialized the val even if it \ndidn't find the requested config variable (esp since we don't pass in a \ndefault value for this function like we do others).\n\n>>\n>>> +     if (!git_config_get_bool(\"index.recordoffsettable\", &val))\n>>> +             return val;\n>>> +     return 0;\n>>> +}\n"},{"id":"364102","messageId":"CAGZ79kZLFGtKzQLAXcrXRy8D9NhMwG=7iTCMzYu_h3Ba4G8cYQ@mail.gmail.com","threadId":"49204","inReplyTo":"7e06ec44-a3a0-fa38-75a7-7b875ae0679e@gmail.com","subject":"Re: [PATCH 2/5] ieot: default to not writing IEOT section","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-11-26T22:02:51Z","receivedAt":"2018-11-26T22:03:05Z","isPatch":true,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Mon, Nov 26, 2018 at 1:48 PM Ben Peart <peartben@gmail.com> wrote:\n>\n>\n>\n> On 11/26/2018 2:59 PM, Stefan Beller wrote:\n> >>> +static int record_ieot(void)\n> >>> +{\n> >>> +     int val;\n> >>> +\n> >>\n> >> Initialize stack val to zero to ensure proper default.\n> >\n> > I don't think that is needed here, as we only use `val` when\n> > we first write to it via git_config_get_bool.\n> >\n> > Did you spot this via code review and thought of\n> > defensive programming or is there a tool that\n> > has a false positive here?\n> >\n>\n> Code review and defensive programming.  I had to review the code in\n> git_config_get_bool() to see if it always initialized the val even if it\n> didn't find the requested config variable (esp since we don't pass in a\n> default value for this function like we do others).\n>\n\nAh, sorry to have sent out this email, which I found as one of the\nearliest discussions in my mailbox. The later patches/discussions\nbecame a lot more heated from my cursory skimming and sorted\nout this as well.\n\nIt is interesting to notice that, as I also had to lookup how the config\nmachinery works (once? a couple times?) but now it is so hardcoded\nin my brain to assume that if functions like git_config_* take the\nbranch, we can access the value that the config function was supposed\nto read into.\n\nSorry for the noise,\nStefan\n"},{"id":"364112","messageId":"xmqq8t1fuywv.fsf@gitster-ct.c.googlers.com","threadId":"49204","inReplyTo":"CAGZ79kbaPKaCFGGXnbNchvk=1Q4Q5Hgt2hXOhcGo6pVwquhaEg@mail.gmail.com","subject":"Re: [PATCH 2/5] ieot: default to not writing IEOT section","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-11-27T00:50:08Z","receivedAt":"2018-11-27T00:50:17Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Stefan Beller <sbeller@google.com> writes:\n\n>> > +static int record_ieot(void)\n>> > +{\n>> > +     int val;\n>> > +\n>>\n>> Initialize stack val to zero to ensure proper default.\n>\n> I don't think that is needed here, as we only use `val` when\n> we first write to it via git_config_get_bool.\n\nYup.\n"}]}