{"thread":{"id":"52499","subject":"[PATCH 0/9] [RFC] Changed Paths Bloom Filters","startedAt":"2019-12-20T22:05:26Z","lastAt":"2020-08-05T17:11:13Z","messageCount":159,"participants":["Garima Singh via GitGitGadget","Junio C Hamano","Philip Oakley","Christian Couder","Jeff King","Derrick Stolee","Jakub Narebski","Garima Singh","Emily Shaffer","Jeff King via GitGitGadget","Derrick Stolee via GitGitGadget","SZEDER Gábor","Bryan Turner","Jakub Narębski","Taylor Blau"],"isPatch":true,"patchVersion":1,"patchTotal":9},"messages":[{"id":"388681","messageId":"pull.497.git.1576879520.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":null,"subject":"[PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-20T22:05:11Z","receivedAt":"2019-12-20T22:05:26Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"Hey! \n\nThe commit graph feature brought in a lot of performance improvements across\nmultiple commands. However, file based history continues to be a performance\npain point, especially in large repositories. \n\nAdopting changed path bloom filters has been discussed on the list before,\nand a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\nDerrick Stolee [1]. This series is based on Dr. Stolee's approach [2] and\npresents an updated and more polished RFC version of the feature. \n\nPerformance Gains: We tested the performance of git log -- path on the git\nrepo, the linux repo and some internal large repos, with a variety of paths\nof varying depths.\n\nOn the git and linux repos: We observed a 2x to 5x speed up.\n\nOn a large internal repo with files seated 6-10 levels deep in the tree: We\nobserved 10x to 20x speed ups, with some paths going up to 28 times faster.\n\nFuture Work (not included in the scope of this series):\n\n 1. Supporting multiple path based revision walk\n 2. Adopting it in git blame logic. \n 3. Interactions with line log git log -L\n\nThis series is intended to start the conversation and many of the commit\nmessages include specific call outs for suggestions and thoughts. \n\nCheers! Garima Singh\n\n[1] https://lore.kernel.org/git/20181009193445.21908-1-szeder.dev@gmail.com/\n[2] \nhttps://lore.kernel.org/git/61559c5b-546e-d61b-d2e1-68de692f5972@gmail.com/\n\nGarima Singh (9):\n  commit-graph: add --changed-paths option to write\n  commit-graph: write changed paths bloom filters\n  commit-graph: use MAX_NUM_CHUNKS\n  commit-graph: document bloom filter format\n  commit-graph: write changed path bloom filters to commit-graph file.\n  commit-graph: test commit-graph write --changed-paths\n  commit-graph: reuse existing bloom filters during write.\n  revision.c: use bloom filters to speed up path based revision walks\n  commit-graph: add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag\n\n Documentation/git-commit-graph.txt            |   5 +\n .../technical/commit-graph-format.txt         |  17 ++\n Makefile                                      |   1 +\n bloom.c                                       | 257 +++++++++++++++++\n bloom.h                                       |  51 ++++\n builtin/commit-graph.c                        |   9 +-\n ci/run-build-and-tests.sh                     |   1 +\n commit-graph.c                                | 116 +++++++-\n commit-graph.h                                |   9 +-\n revision.c                                    |  67 ++++-\n revision.h                                    |   5 +\n t/README                                      |   3 +\n t/helper/test-read-graph.c                    |   4 +\n t/t4216-log-bloom.sh                          |  77 ++++++\n t/t5318-commit-graph.sh                       |   2 +\n t/t5324-split-commit-graph.sh                 |   1 +\n t/t5325-commit-graph-bloom.sh                 | 258 ++++++++++++++++++\n 17 files changed, 875 insertions(+), 8 deletions(-)\n create mode 100644 bloom.c\n create mode 100644 bloom.h\n create mode 100755 t/t4216-log-bloom.sh\n create mode 100755 t/t5325-commit-graph-bloom.sh\n\n\nbase-commit: b02fd2accad4d48078671adf38fe5b5976d77304\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-497%2Fgarimasi514%2FcoreGit-bloomFilters-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-497/garimasi514/coreGit-bloomFilters-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/497\n-- \ngitgitgadget\n"},{"id":"388682","messageId":"6bdde5e4f0ceb54546978e3e9cdde00045d45468.1576879520.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"[PATCH 1/9] commit-graph: add --changed-paths option to write","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-20T22:05:12Z","receivedAt":"2019-12-20T22:05:29Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd --changed-paths option to git commit-graph write. This option will\nsoon allow users to compute bloom filters for the paths changed between\na commit and its first significant parent, and write this information\ninto the commit-graph file.\n\nNote: This commit does not change any behavior. It only introduces\nthe option and passes down the appropriate flag to the commit-graph.\n\nRFC Notes:\n1. We named the option --changed-paths to capture what the option does,\n   instead of how it does it. The current implementation does this\n   using bloom filters. We believe using --changed-paths however keeps\n   the implementation open to other data structures.\n   All thoughts and suggestions for the name and this approach are\n   welcome\n\n2. Currently, a subsequent commit in this series will add tests that\n   exercise this option. I plan to split that test commit across the\n   series as appropriate.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n Documentation/git-commit-graph.txt | 5 +++++\n builtin/commit-graph.c             | 9 +++++++--\n commit-graph.h                     | 3 ++-\n 3 files changed, 14 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/git-commit-graph.txt b/Documentation/git-commit-graph.txt\nindex bcd85c1976..1efe6e5c5a 100644\n--- a/Documentation/git-commit-graph.txt\n+++ b/Documentation/git-commit-graph.txt\n@@ -54,6 +54,11 @@ or `--stdin-packs`.)\n With the `--append` option, include all commits that are present in the\n existing commit-graph file.\n +\n+With the `--changed-paths` option, compute and write information about the\n+paths changed between a commit and it's first parent. This operation can\n+take a while on large repositories. It provides significant performance gains\n+for getting file based history logs with `git log`\n++\n With the `--split` option, write the commit-graph as a chain of multiple\n commit-graph files stored in `<dir>/info/commit-graphs`. The new commits\n not already in the commit-graph are added in a new \"tip\" file. This file\ndiff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\nindex e0c6fc4bbf..9bd1e11161 100644\n--- a/builtin/commit-graph.c\n+++ b/builtin/commit-graph.c\n@@ -9,7 +9,7 @@\n \n static char const * const builtin_commit_graph_usage[] = {\n \tN_(\"git commit-graph verify [--object-dir <objdir>] [--shallow] [--[no-]progress]\"),\n-\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n+\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n \tNULL\n };\n \n@@ -19,7 +19,7 @@ static const char * const builtin_commit_graph_verify_usage[] = {\n };\n \n static const char * const builtin_commit_graph_write_usage[] = {\n-\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n+\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n \tNULL\n };\n \n@@ -32,6 +32,7 @@ static struct opts_commit_graph {\n \tint split;\n \tint shallow;\n \tint progress;\n+\tint enable_bloom_filters;\n } opts;\n \n static int graph_verify(int argc, const char **argv)\n@@ -110,6 +111,8 @@ static int graph_write(int argc, const char **argv)\n \t\t\tN_(\"start walk at commits listed by stdin\")),\n \t\tOPT_BOOL(0, \"append\", &opts.append,\n \t\t\tN_(\"include all commits already in the commit-graph file\")),\n+\t\tOPT_BOOL(0, \"changed-paths\", &opts.enable_bloom_filters,\n+\t\t\tN_(\"enable computation for changed paths\")),\n \t\tOPT_BOOL(0, \"progress\", &opts.progress, N_(\"force progress reporting\")),\n \t\tOPT_BOOL(0, \"split\", &opts.split,\n \t\t\tN_(\"allow writing an incremental commit-graph file\")),\n@@ -143,6 +146,8 @@ static int graph_write(int argc, const char **argv)\n \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n \tif (opts.progress)\n \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n+\tif (opts.enable_bloom_filters)\n+\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n \n \tread_replace_refs = 0;\n \ndiff --git a/commit-graph.h b/commit-graph.h\nindex 7f5c933fa2..952a4b83be 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -76,7 +76,8 @@ enum commit_graph_write_flags {\n \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\n \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n \t/* Make sure that each OID in the input is a valid commit OID. */\n-\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3)\n+\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n+\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n };\n \n struct split_commit_graph_opts {\n-- \ngitgitgadget\n\n"},{"id":"388683","messageId":"a15f87fdcbea1a37a20a05135832b42f36f682f1.1576879520.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"[PATCH 3/9] commit-graph: use MAX_NUM_CHUNKS","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-20T22:05:14Z","receivedAt":"2019-12-20T22:05:29Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nThis is a minor cleanup to make it easier to change the\nnumber of chunks being written to the commit-graph in the future.\n\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 5 +++--\n 1 file changed, 3 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 61e60ff98a..8c4941eeaa 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -24,6 +24,7 @@\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n+#define MAX_NUM_CHUNKS 5\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -1381,8 +1382,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tint fd;\n \tstruct hashfile *f;\n \tstruct lock_file lk = LOCK_INIT;\n-\tuint32_t chunk_ids[6];\n-\tuint64_t chunk_offsets[6];\n+\tuint32_t chunk_ids[MAX_NUM_CHUNKS + 1];\n+\tuint64_t chunk_offsets[MAX_NUM_CHUNKS + 1];\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \tstruct strbuf progress_title = STRBUF_INIT;\n \tint num_chunks = 3;\n-- \ngitgitgadget\n\n"},{"id":"388684","messageId":"3182a11f7c07af834ba71dc7861742458754eb91.1576879520.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"[PATCH 4/9] commit-graph: document bloom filter format","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-20T22:05:15Z","receivedAt":"2019-12-20T22:05:30Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nUpdate the technical documentation for commit-graph-format with BIDX\nand BDAT chunk information.\n\nRFC Notes:\n1. [Call for advice] We specifically mention that we are using bloom\n   filters in this technical document. Should this document also be\n   made open to other data structures in the future, with versioning\n   information?\n\n2. [Call for advice] We are also not describing the explicit nature\n   of how we store the bloom filter binary data. Would it be useful\n   to document details about the hash algorithm, the number of hashes\n   and the specific seed values we are using in a separate document,\n   or perhaps in a separate section in this document?\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n Documentation/technical/commit-graph-format.txt | 17 +++++++++++++++++\n 1 file changed, 17 insertions(+)\n\ndiff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\nindex a4f17441ae..6497f19f08 100644\n--- a/Documentation/technical/commit-graph-format.txt\n+++ b/Documentation/technical/commit-graph-format.txt\n@@ -17,6 +17,9 @@ metadata, including:\n - The parents of the commit, stored using positional references within\n   the graph file.\n \n+- The bloom filter of the commit carrying the paths that were changed between\n+  the commit and it's first parent.\n+\n These positional references are stored as unsigned 32-bit integers\n corresponding to the array position within the list of commit OIDs. Due\n to some special constants we use to track parents, we can store at most\n@@ -93,6 +96,20 @@ CHUNK DATA:\n       positions for the parents until reaching a value with the most-significant\n       bit on. The other bits correspond to the position of the last parent.\n \n+  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) [Optional]\n+      For each commit we store the offset of its bloom filter in the BDAT chunk\n+      as follows:\n+      BIDX[i] = number of 8-byte words in all the bloom filters from commit 0 to\n+\t\tcommit i (inclusive)\n+\n+  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n+      * It starts with three 32 bit integers for the\n+\t    - version of the hash algorithm being used\n+\t    - the number of hashes used in the computation\n+\t    - the number of bits per entry\n+\t  * The rest of the chunk is the concatenation of all the computed bloom \n+\t  filters for the commits in lexicographic order.\n+\n   Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n       This list of H-byte hashes describe a set of B commit-graph files that\n       form a commit-graph chain. The graph position for the ith commit in this\n-- \ngitgitgadget\n\n"},{"id":"388685","messageId":"72a2bbf6765a1e99a3a23372f801099a07fe11a5.1576879520.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"[PATCH 8/9] revision.c: use bloom filters to speed up path based revision walks","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-20T22:05:19Z","receivedAt":"2019-12-20T22:05:33Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nIf bloom filters have been written to the commit-graph file, revision walk will\nuse them to speed up revision walks for a particular path.\nNote: The current implementation does this in the case of single pathspec\ncase only.\n\nWe load the bloom filters during the prepare_revision_walk step when dealing\nwith a single pathspec. While comparing trees in rev_compare_trees(), if the\nbloom filter says that the file is not different between the two trees, we\ndon't need to compute the expensive diff. This is where we get our performance\ngains.\n\nPerformance Gains:\nWe tested the performance of `git log --path` on the git repo, the linux and\nsome internal large repos, with a variety of paths of varying depths.\n\nOn the git and linux repos:\nwe observed a 2x to 5x speed up.\n\nOn a large internal repo with files seated 6-10 levels deep in the tree:\nwe observed 10x to 20x speed ups, with some paths going up to 28 times faster.\n\nRFC Notes:\nI plan to collect the folloowing statistics around this usage of bloom filters\nand trace them out using trace2.\n- number of bloom filter queries,\n- number of \"No\" responses (file hasn't changed)\n- number of \"Maybe\" responses (file may have changed)\n- number of \"Commit not parsed\" cases (commit had too many changes to have a\n  bloom filter written out, currently our limit is 512 diffs)\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com\nHelped-by: SZEDER Gábor <szeder.dev@gmail.com>\nHelped-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c              | 20 ++++++++++++\n bloom.h              |  4 +++\n revision.c           | 67 +++++++++++++++++++++++++++++++++++++--\n revision.h           |  5 +++\n t/t4216-log-bloom.sh | 74 ++++++++++++++++++++++++++++++++++++++++++++\n 5 files changed, 168 insertions(+), 2 deletions(-)\n create mode 100755 t/t4216-log-bloom.sh\n\ndiff --git a/bloom.c b/bloom.c\nindex 86b1005802..0c7505d3d6 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -235,3 +235,23 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \treturn filter;\n }\n+\n+int bloom_filter_contains(struct bloom_filter *filter,\n+\t\t\t  struct bloom_key *key,\n+\t\t\t  struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tuint64_t mod = filter->len * BITS_PER_BLOCK;\n+\n+\tif (!mod)\n+\t\treturn 1;\n+\n+\tfor (i = 0; i < settings->num_hashes; i++) {\n+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n+\t\tuint64_t block_pos = hash_mod / BITS_PER_BLOCK;\n+\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n+\t\t\treturn 0;\n+\t}\n+\n+\treturn 1;\n+}\ndiff --git a/bloom.h b/bloom.h\nindex 101d689bbd..9bdacd0a8e 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -44,4 +44,8 @@ void fill_bloom_key(const char *data,\n \t\t    struct bloom_key *key,\n \t\t    struct bloom_filter_settings *settings);\n \n+int bloom_filter_contains(struct bloom_filter *filter,\n+\t\t\t  struct bloom_key *key,\n+\t\t\t  struct bloom_filter_settings *settings);\n+\n #endif\ndiff --git a/revision.c b/revision.c\nindex 39a25e7a5d..01f5330740 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -29,6 +29,7 @@\n #include \"prio-queue.h\"\n #include \"hashmap.h\"\n #include \"utf8.h\"\n+#include \"bloom.h\"\n \n volatile show_early_output_fn_t show_early_output;\n \n@@ -624,11 +625,34 @@ static void file_change(struct diff_options *options,\n \toptions->flags.has_changes = 1;\n }\n \n+static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n+\t\t\t\t\t\t struct commit *commit,\n+\t\t\t\t\t\t struct bloom_key *key,\n+\t\t\t\t\t\t struct bloom_filter_settings *settings)\n+{\n+\tstruct bloom_filter *filter;\n+\n+\tif (!revs->repo->objects->commit_graph)\n+\t\treturn -1;\n+\tif (commit->generation == GENERATION_NUMBER_INFINITY)\n+\t\treturn -1;\n+\tif (!key || !settings)\n+\t\treturn -1;\n+\n+\tfilter = get_bloom_filter(revs->repo, commit, 0);\n+\n+\tif (!filter || !filter->len)\n+\t\treturn 1;\n+\n+\treturn bloom_filter_contains(filter, key, settings);\n+}\n+\n static int rev_compare_tree(struct rev_info *revs,\n-\t\t\t    struct commit *parent, struct commit *commit)\n+\t\t\t    struct commit *parent, struct commit *commit, int nth_parent)\n {\n \tstruct tree *t1 = get_commit_tree(parent);\n \tstruct tree *t2 = get_commit_tree(commit);\n+\tint bloom_ret = 1;\n \n \tif (!t1)\n \t\treturn REV_TREE_NEW;\n@@ -653,6 +677,16 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\treturn REV_TREE_SAME;\n \t}\n \n+\tif (revs->pruning.pathspec.nr == 1 && !nth_parent) {\n+\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs,\n+\t\t\t\t\t\t\t\t  commit,\n+\t\t\t\t\t\t\t\t  revs->bloom_key,\n+\t\t\t\t\t\t\t\t  revs->bloom_filter_settings);\n+\n+\t\tif (bloom_ret == 0)\n+\t\t\treturn REV_TREE_SAME;\n+\t}\n+\n \ttree_difference = REV_TREE_SAME;\n \trevs->pruning.flags.has_changes = 0;\n \tif (diff_tree_oid(&t1->object.oid, &t2->object.oid, \"\",\n@@ -855,7 +889,7 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit)\n \t\t\tdie(\"cannot simplify commit %s (because of %s)\",\n \t\t\t    oid_to_hex(&commit->object.oid),\n \t\t\t    oid_to_hex(&p->object.oid));\n-\t\tswitch (rev_compare_tree(revs, p, commit)) {\n+\t\tswitch (rev_compare_tree(revs, p, commit, nth_parent)) {\n \t\tcase REV_TREE_SAME:\n \t\t\tif (!revs->simplify_history || !relevant_commit(p)) {\n \t\t\t\t/* Even if a merge with an uninteresting\n@@ -3342,6 +3376,33 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n \t}\n }\n \n+static void prepare_to_use_bloom_filter(struct rev_info *revs)\n+{\n+\tstruct pathspec_item *pi;\n+\tconst char *path;\n+\tsize_t len;\n+\n+\tif (!revs->commits)\n+\t    return;\n+\n+\tparse_commit(revs->commits->item);\n+\n+\tif (!revs->repo->objects->commit_graph)\n+\t\treturn;\n+\n+\trevs->bloom_filter_settings = revs->repo->objects->commit_graph->settings;\n+\tif (!revs->bloom_filter_settings)\n+\t\treturn;\n+\n+\tpi = &revs->pruning.pathspec.items[0];\n+\tpath = pi->match;\n+\tlen = strlen(path);\n+\n+\tload_bloom_filters();\n+\trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n+\tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n+}\n+\n int prepare_revision_walk(struct rev_info *revs)\n {\n \tint i;\n@@ -3391,6 +3452,8 @@ int prepare_revision_walk(struct rev_info *revs)\n \t\tsimplify_merges(revs);\n \tif (revs->children.name)\n \t\tset_children(revs);\n+\tif (revs->pruning.pathspec.nr == 1)\n+\t    prepare_to_use_bloom_filter(revs);\n \treturn 0;\n }\n \ndiff --git a/revision.h b/revision.h\nindex a1a804bd3d..65dc11e8f1 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -56,6 +56,8 @@ struct repository;\n struct rev_info;\n struct string_list;\n struct saved_parents;\n+struct bloom_key;\n+struct bloom_filter_settings;\n define_shared_commit_slab(revision_sources, char *);\n \n struct rev_cmdline_info {\n@@ -291,6 +293,9 @@ struct rev_info {\n \tstruct revision_sources *sources;\n \n \tstruct topo_walk_info *topo_walk_info;\n+\n+\tstruct bloom_key *bloom_key;\n+\tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n int ref_excluded(struct string_list *, const char *path);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nnew file mode 100755\nindex 0000000000..d42f077998\n--- /dev/null\n+++ b/t/t4216-log-bloom.sh\n@@ -0,0 +1,74 @@\n+#!/bin/sh\n+\n+test_description='git log for a path with bloom filters'\n+. ./test-lib.sh\n+\n+test_expect_success 'setup repo' '\n+\tgit init &&\n+\tgit config core.commitGraph true &&\n+\tgit config gc.writeCommitGraph false &&\n+\tinfodir=\".git/objects/info\" &&\n+\tgraphdir=\"$infodir/commit-graphs\" &&\n+\ttest_oid_init\n+'\n+\n+test_expect_success 'create 9 commits and repack' '\n+\ttest_commit c1 file1 &&\n+\ttest_commit c2 file2 &&\n+\ttest_commit c3 file3 &&\n+\ttest_commit c4 file1 &&\n+\ttest_commit c5 file2 &&\n+\ttest_commit c6 file3 &&\n+\ttest_commit c7 file1 &&\n+\ttest_commit c8 file2 &&\n+\ttest_commit c9 file3\n+'\n+\n+printf \"c7\\nc4\\nc1\" > expect_file1\n+\n+test_expect_success 'log without bloom filters' '\n+\tgit log --pretty=\"format:%s\"  -- file1 > actual &&\n+\ttest_cmp expect_file1 actual\n+'\n+\n+printf \"c8\\nc7\\nc5\\nc4\\nc2\\nc1\" > expect_file1_file2\n+\n+test_expect_success 'multi-path log without bloom filters' '\n+\tgit log --pretty=\"format:%s\"  -- file1 file2 > actual &&\n+\ttest_cmp expect_file1_file2 actual\n+'\n+\n+graph_read_expect() {\n+\tOPTIONAL=\"\"\n+\tNUM_CHUNKS=5\n+\tif test ! -z $2\n+\tthen\n+\t\tOPTIONAL=\" $2\"\n+\t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n+\tfi\n+\tcat >expect <<- EOF\n+\theader: 43475048 1 1 $NUM_CHUNKS 0\n+\tnum_commits: $1\n+\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data$OPTIONAL\n+\tEOF\n+\ttest-tool read-graph >output &&\n+\ttest_cmp expect output\n+}\n+\n+test_expect_success 'write commit graph with bloom filters' '\n+\tgit commit-graph write --reachable --changed-paths &&\n+\ttest_path_is_file $infodir/commit-graph &&\n+\tgraph_read_expect \"9\"\n+'\n+\n+test_expect_success 'log using bloom filters' '\n+\tgit log --pretty=\"format:%s\" -- file1 > actual &&\n+\ttest_cmp expect_file1 actual\n+'\n+\n+test_expect_success 'multi-path log using bloom filters' '\n+\tgit log --pretty=\"format:%s\"  -- file1 file2 > actual &&\n+\ttest_cmp expect_file1_file2 actual\n+'\n+\n+test_done\n-- \ngitgitgadget\n\n"},{"id":"388686","messageId":"1e2acb37ad710cb0d1c09ed163fdd4473a27335c.1576879520.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"[PATCH 7/9] commit-graph: reuse existing bloom filters during write.","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-20T22:05:18Z","receivedAt":"2019-12-20T22:05:34Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nRead previously computed bloom filters from the commit-graph file if possible\nto avoid recomputing during commit-graph write.\n\nReading from the commit-graph is based on the format in which bloom filters are\nwritten in the commit graph file. See method `fill_filter_from_graph` in bloom.c\n\nFor reading the bloom filter for commit at lexicographic position i:\n1. Read BIDX[i] which essentially gives us the starting index in BDAT for filter\n   of commit i+1 (called the next_index in the code)\n\n2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT for\n   filter of commit i (called the prev_index in the code)\n   For i = 0, prev_index will be 0. The first lexicographic commit's filter will\n   start at BDAT.\n\n3. The length of the filter will be next_index - prev_index, because BIDX[i]\n   gives the cumulative 8-byte words including the ith commit's filter.\n\nWe toggle whether bloom filters should be recomputed based on the compute_if_null\nflag.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c        | 40 ++++++++++++++++++++++++++++++++++++++--\n bloom.h        |  3 ++-\n commit-graph.c |  6 +++---\n 3 files changed, 43 insertions(+), 6 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 08328cc381..86b1005802 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -1,5 +1,7 @@\n #include \"git-compat-util.h\"\n #include \"bloom.h\"\n+#include \"commit.h\"\n+#include \"commit-slab.h\"\n #include \"commit-graph.h\"\n #include \"object-store.h\"\n #include \"diff.h\"\n@@ -119,13 +121,35 @@ static void add_key_to_filter(struct bloom_key *key,\n \t}\n }\n \n+static void fill_filter_from_graph(struct commit_graph *g,\n+\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t   struct commit *c)\n+{\n+\tuint32_t lex_pos, prev_index, next_index;\n+\n+\twhile (c->graph_pos < g->num_commits_in_base)\n+\t\tg = g->base_graph;\n+\n+\tlex_pos = c->graph_pos - g->num_commits_in_base;\n+\n+\tnext_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n+\tif (lex_pos)\n+\t\tprev_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n+\telse\n+\t\tprev_index = 0;\n+\n+\tfilter->len = next_index - prev_index;\n+\tfilter->data = (uint64_t *)(g->chunk_bloom_data + 8 * prev_index + 12);\n+}\n+\n void load_bloom_filters(void)\n {\n \tinit_bloom_filter_slab(&bloom_filters);\n }\n \n struct bloom_filter *get_bloom_filter(struct repository *r,\n-\t\t\t\t      struct commit *c)\n+\t\t\t\t      struct commit *c,\n+\t\t\t\t      int compute_if_null)\n {\n \tstruct bloom_filter *filter;\n \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n@@ -134,6 +158,18 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \tconst char *revs_argv[] = {NULL, \"HEAD\", NULL};\n \n \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n+\n+\tif (!filter->data) {\n+\t\tload_commit_graph_info(r, c);\n+\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH && r->objects->commit_graph->chunk_bloom_indexes) {\n+\t\t\tfill_filter_from_graph(r->objects->commit_graph, filter, c);\n+\t\t\treturn filter;\n+\t\t}\n+\t}\n+\n+\tif (filter->data || !compute_if_null)\n+\t\t\treturn filter;\n+\n \tinit_revisions(&revs, NULL);\n \trevs.diffopt.flags.recursive = 1;\n \n@@ -198,4 +234,4 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n \n \treturn filter;\n-}\n\\ No newline at end of file\n+}\ndiff --git a/bloom.h b/bloom.h\nindex ba8ae70b67..101d689bbd 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -36,7 +36,8 @@ struct bloom_key {\n void load_bloom_filters(void);\n \n struct bloom_filter *get_bloom_filter(struct repository *r,\n-\t\t\t\t      struct commit *c);\n+\t\t\t\t      struct commit *c,\n+\t\t\t\t      int compute_if_null);\n \n void fill_bloom_key(const char *data,\n \t\t    int len,\ndiff --git a/commit-graph.c b/commit-graph.c\nindex def2ade166..0580ce75d5 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1032,7 +1032,7 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n \tuint32_t cur_pos = 0;\n \n \twhile (list < last) {\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n \t\tcur_pos += filter->len;\n \t\thashwrite_be32(f, cur_pos);\n \t\tlist++;\n@@ -1051,7 +1051,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n \thashwrite_be32(f, settings->bits_per_entry);\n \n \twhile (first < last) {\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *first);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *first, 0);\n \t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n \t\tfirst++;\n \t}\n@@ -1218,7 +1218,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tstruct commit *c = ctx->commits.list[i];\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n \t\tctx->total_bloom_filter_size += sizeof(uint64_t) * filter->len;\n \t\tdisplay_progress(progress, i + 1);\n \t}\n-- \ngitgitgadget\n\n"},{"id":"388687","messageId":"e1c315d0a766af147eb4ead41a172f724e90cc34.1576879520.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"[PATCH 9/9] commit-graph: add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-20T22:05:20Z","receivedAt":"2019-12-20T22:05:35Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag to the test setup suite in\norder to toggle writing bloom filters when running any of the git tests. If set\nto true, we will compute and write bloom filters every time a test calls\n`git commit-graph write`.\n\nThe test suite passes when GIT_TEST_COMMIT_GRAPH and\nGIT_COMMIT_GRAPH_BLOOM_FILTERS are enabled.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n builtin/commit-graph.c        | 2 +-\n ci/run-build-and-tests.sh     | 1 +\n commit-graph.h                | 1 +\n t/README                      | 3 +++\n t/t4216-log-bloom.sh          | 3 +++\n t/t5318-commit-graph.sh       | 2 ++\n t/t5324-split-commit-graph.sh | 1 +\n t/t5325-commit-graph-bloom.sh | 3 +++\n 8 files changed, 15 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\nindex 9bd1e11161..97167959b2 100644\n--- a/builtin/commit-graph.c\n+++ b/builtin/commit-graph.c\n@@ -146,7 +146,7 @@ static int graph_write(int argc, const char **argv)\n \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n \tif (opts.progress)\n \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n-\tif (opts.enable_bloom_filters)\n+\tif (opts.enable_bloom_filters || git_env_bool(GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS, 0))\n \t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n \n \tread_replace_refs = 0;\ndiff --git a/ci/run-build-and-tests.sh b/ci/run-build-and-tests.sh\nindex ff0ef7f08e..19d0846d34 100755\n--- a/ci/run-build-and-tests.sh\n+++ b/ci/run-build-and-tests.sh\n@@ -19,6 +19,7 @@ linux-gcc)\n \texport GIT_TEST_OE_SIZE=10\n \texport GIT_TEST_OE_DELTA_SIZE=5\n \texport GIT_TEST_COMMIT_GRAPH=1\n+\texport GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=1\n \texport GIT_TEST_MULTI_PACK_INDEX=1\n \tmake test\n \t;;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 2202ad91ae..d914e6abf1 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -8,6 +8,7 @@\n \n #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n+#define GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS \"GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS\"\n \n struct commit;\n struct bloom_filter_settings;\ndiff --git a/t/README b/t/README\nindex caa125ba9a..399b190437 100644\n--- a/t/README\n+++ b/t/README\n@@ -378,6 +378,9 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=<boolean>, when true, forces commit-graph\n+write to compute and write bloom filters for every 'git commit-graph write'\n+\n GIT_TEST_FSMONITOR=$PWD/t7519/fsmonitor-all exercises the fsmonitor\n code path for utilizing a file system monitor to speed up detecting\n new or changed files.\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex d42f077998..0e092b387c 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -3,6 +3,9 @@\n test_description='git log for a path with bloom filters'\n . ./test-lib.sh\n \n+GIT_TEST_COMMIT_GRAPH=0\n+GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n+\n test_expect_success 'setup repo' '\n \tgit init &&\n \tgit config core.commitGraph true &&\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 3f03de6018..613228bb12 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -3,6 +3,8 @@\n test_description='commit graph'\n . ./test-lib.sh\n \n+GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n+\n test_expect_success 'setup full repo' '\n \tmkdir full &&\n \tcd \"$TRASH_DIRECTORY/full\" &&\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex c24823431f..181ca7e0cb 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -4,6 +4,7 @@ test_description='split commit graph'\n . ./test-lib.sh\n \n GIT_TEST_COMMIT_GRAPH=0\n+GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n \n test_expect_success 'setup repo' '\n \tgit init &&\ndiff --git a/t/t5325-commit-graph-bloom.sh b/t/t5325-commit-graph-bloom.sh\nindex d7ef0e7fb3..a9c9e9fef6 100755\n--- a/t/t5325-commit-graph-bloom.sh\n+++ b/t/t5325-commit-graph-bloom.sh\n@@ -3,6 +3,9 @@\n test_description='commit graph with bloom filters'\n . ./test-lib.sh\n \n+GIT_TEST_COMMIT_GRAPH=0\n+GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n+\n test_expect_success 'setup repo' '\n \tgit init &&\n \tgit config core.commitGraph true &&\n-- \ngitgitgadget\n"},{"id":"388688","messageId":"7648021072ca11153ac65c90f0ebed5973f20e1a.1576879520.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"[PATCH 5/9] commit-graph: write changed path bloom filters to commit-graph file.","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-20T22:05:16Z","receivedAt":"2019-12-20T22:05:36Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nWrite bloom filters to the commit-graph using the format described in\nDocumentation/technical/commit-graph-format.txt\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 81 +++++++++++++++++++++++++++++++++++++++++++++++++-\n commit-graph.h |  5 ++++\n 2 files changed, 85 insertions(+), 1 deletion(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 8c4941eeaa..def2ade166 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -24,7 +24,9 @@\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 5\n+#define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n+#define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n+#define MAX_NUM_CHUNKS 7\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -282,6 +284,32 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n \t\t\t\tchunk_repeated = 1;\n \t\t\telse\n \t\t\t\tgraph->chunk_base_graphs = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_BLOOMINDEXES:\n+\t\t\tif (graph->chunk_bloom_indexes)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_bloom_indexes = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_BLOOMDATA:\n+\t\t\tif (graph->chunk_bloom_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse {\n+\t\t\t\tuint32_t hash_version;\n+\t\t\t\tgraph->chunk_bloom_data = data + chunk_offset;\n+\t\t\t\thash_version = get_be32(data + chunk_offset);\n+\n+\t\t\t\tif (hash_version != 1)\n+\t\t\t\t\tbreak;\n+\n+\t\t\t\tgraph->settings = xmalloc(sizeof(struct bloom_filter_settings));\n+\t\t\t\tgraph->settings->hash_version = hash_version;\n+\t\t\t\tgraph->settings->num_hashes = get_be32(data + chunk_offset + 4);\n+\t\t\t\tgraph->settings->bits_per_entry = get_be32(data + chunk_offset + 8);\n+\t\t\t}\n+\t\t\tbreak;\n \t\t}\n \n \t\tif (chunk_repeated) {\n@@ -996,6 +1024,39 @@ static void write_graph_chunk_extra_edges(struct hashfile *f,\n \t}\n }\n \n+static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n+\t\t\t\t\t    struct write_commit_graph_context *ctx)\n+{\n+\tstruct commit **list = ctx->commits.list;\n+\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n+\tuint32_t cur_pos = 0;\n+\n+\twhile (list < last) {\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tcur_pos += filter->len;\n+\t\thashwrite_be32(f, cur_pos);\n+\t\tlist++;\n+\t}\n+}\n+\n+static void write_graph_chunk_bloom_data(struct hashfile *f,\n+\t\t\t\t\t struct write_commit_graph_context *ctx,\n+\t\t\t\t\t struct bloom_filter_settings *settings)\n+{\n+\tstruct commit **first = ctx->commits.list;\n+\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n+\n+\thashwrite_be32(f, settings->hash_version);\n+\thashwrite_be32(f, settings->num_hashes);\n+\thashwrite_be32(f, settings->bits_per_entry);\n+\n+\twhile (first < last) {\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *first);\n+\t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n+\t\tfirst++;\n+\t}\n+}\n+\n static int oid_compare(const void *_a, const void *_b)\n {\n \tconst struct object_id *a = (const struct object_id *)_a;\n@@ -1388,6 +1449,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tstruct strbuf progress_title = STRBUF_INIT;\n \tint num_chunks = 3;\n \tstruct object_id file_hash;\n+\tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \n \tif (ctx->split) {\n \t\tstruct strbuf tmp_file = STRBUF_INIT;\n@@ -1432,6 +1494,12 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tnum_chunks++;\n \t}\n+\tif (ctx->bloom) {\n+\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMINDEXES;\n+\t\tnum_chunks++;\n+\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMDATA;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_commit_graphs_after > 1) {\n \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BASE;\n \t\tnum_chunks++;\n@@ -1450,6 +1518,13 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t\t\t\t\t\t4 * ctx->num_extra_edges;\n \t\tnum_chunks++;\n \t}\n+\tif (ctx->bloom) {\n+\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] + sizeof(uint32_t) * ctx->commits.nr;\n+\t\tnum_chunks++;\n+\n+\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] + sizeof(uint32_t) * 3 + ctx->total_bloom_filter_size;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_commit_graphs_after > 1) {\n \t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n \t\t\t\t\t\thashsz * (ctx->num_commit_graphs_after - 1);\n@@ -1487,6 +1562,10 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \twrite_graph_chunk_data(f, hashsz, ctx);\n \tif (ctx->num_extra_edges)\n \t\twrite_graph_chunk_extra_edges(f, ctx);\n+\tif (ctx->bloom) {\n+\t\twrite_graph_chunk_bloom_indexes(f, ctx);\n+\t\twrite_graph_chunk_bloom_data(f, ctx, &bloom_settings);\n+\t}\n \tif (ctx->num_commit_graphs_after > 1 &&\n \t    write_graph_chunk_base(f, ctx)) {\n \t\treturn -1;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 952a4b83be..2202ad91ae 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -10,6 +10,7 @@\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n \n struct commit;\n+struct bloom_filter_settings;\n \n char *get_commit_graph_filename(const char *obj_dir);\n int open_commit_graph(const char *graph_file, int *fd, struct stat *st);\n@@ -58,6 +59,10 @@ struct commit_graph {\n \tconst unsigned char *chunk_commit_data;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n+\tconst unsigned char *chunk_bloom_indexes;\n+\tconst unsigned char *chunk_bloom_data;\n+\n+\tstruct bloom_filter_settings *settings;\n };\n \n struct commit_graph *load_commit_graph_one_fd_st(int fd, struct stat *st);\n-- \ngitgitgadget\n\n"},{"id":"388689","messageId":"85bfdfa59c48891343e3eeb740a4b3554405909a.1576879520.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"[PATCH 6/9] commit-graph: test commit-graph write --changed-paths","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-20T22:05:17Z","receivedAt":"2019-12-20T22:05:37Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd tests for the --changed-paths feature when writing\ncommit-graphs.\n\nRFC Notes:\nI plan to split this test across some of the earlier commits\nas appropriate.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n t/helper/test-read-graph.c    |   4 +\n t/t5325-commit-graph-bloom.sh | 255 ++++++++++++++++++++++++++++++++++\n 2 files changed, 259 insertions(+)\n create mode 100755 t/t5325-commit-graph-bloom.sh\n\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex d2884efe0a..aff597c7a3 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -45,6 +45,10 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" commit_metadata\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n+\tif (graph->chunk_bloom_indexes)\n+\t\tprintf(\" bloom_indexes\");\n+\tif (graph->chunk_bloom_data)\n+\t\tprintf(\" bloom_data\");\n \tprintf(\"\\n\");\n \n \tUNLEAK(graph);\ndiff --git a/t/t5325-commit-graph-bloom.sh b/t/t5325-commit-graph-bloom.sh\nnew file mode 100755\nindex 0000000000..d7ef0e7fb3\n--- /dev/null\n+++ b/t/t5325-commit-graph-bloom.sh\n@@ -0,0 +1,255 @@\n+#!/bin/sh\n+\n+test_description='commit graph with bloom filters'\n+. ./test-lib.sh\n+\n+test_expect_success 'setup repo' '\n+\tgit init &&\n+\tgit config core.commitGraph true &&\n+\tgit config gc.writeCommitGraph false &&\n+\tinfodir=\".git/objects/info\" &&\n+\tgraphdir=\"$infodir/commit-graphs\" &&\n+\ttest_oid_init\n+'\n+\n+graph_read_expect() {\n+\tOPTIONAL=\"\"\n+\tNUM_CHUNKS=5\n+\tif test ! -z $2\n+\tthen\n+\t\tOPTIONAL=\" $2\"\n+\t\tNUM_CHUNKS=$((NUM_CHUNKS + $(echo \"$2\" | wc -w)))\n+\tfi\n+\tcat >expect <<- EOF\n+\theader: 43475048 1 1 $NUM_CHUNKS 0\n+\tnum_commits: $1\n+\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tEOF\n+\ttest-tool read-graph >output &&\n+\ttest_cmp expect output\n+}\n+\n+test_expect_success 'create commits and write commit-graph' '\n+\tfor i in $(test_seq 3)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit commit-graph write --reachable --changed-paths &&\n+\ttest_path_is_file $infodir/commit-graph &&\n+\tgraph_read_expect 3\n+'\n+\n+graph_git_two_modes() {\n+\tgit -c core.commitGraph=true $1 >output\n+\tgit -c core.commitGraph=false $1 >expect\n+\ttest_cmp expect output\n+}\n+\n+graph_git_behavior() {\n+\tMSG=$1\n+\tBRANCH=$2\n+\tCOMPARE=$3\n+\ttest_expect_success \"check normal git operations: $MSG\" '\n+\t\tgraph_git_two_modes \"log --oneline $BRANCH\" &&\n+\t\tgraph_git_two_modes \"log --topo-order $BRANCH\" &&\n+\t\tgraph_git_two_modes \"log --graph $COMPARE..$BRANCH\" &&\n+\t\tgraph_git_two_modes \"branch -vv\" &&\n+\t\tgraph_git_two_modes \"merge-base -a $BRANCH $COMPARE\"\n+\t'\n+}\n+\n+graph_git_behavior 'graph exists' commits/3 commits/1\n+\n+verify_chain_files_exist() {\n+\tfor hash in $(cat $1/commit-graph-chain)\n+\tdo\n+\t\ttest_path_is_file $1/graph-$hash.graph || return 1\n+\tdone\n+}\n+\n+test_expect_success 'add more commits, and write a new base graph' '\n+\tgit reset --hard commits/1 &&\n+\tfor i in $(test_seq 4 5)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/2 &&\n+\tfor i in $(test_seq 6 10)\n+\tdo\n+\t\ttest_commit $i &&\n+\t\tgit branch commits/$i || return 1\n+\tdone &&\n+\tgit reset --hard commits/2 &&\n+\tgit merge commits/4 &&\n+\tgit branch merge/1 &&\n+\tgit reset --hard commits/4 &&\n+\tgit merge commits/6 &&\n+\tgit branch merge/2 &&\n+\tgit commit-graph write --reachable --changed-paths &&\n+\tgraph_read_expect 12\n+'\n+\n+test_expect_success 'fork and fail to base a chain on a commit-graph file' '\n+\ttest_when_finished rm -rf fork &&\n+\tgit clone . fork &&\n+\t(\n+\t\tcd fork &&\n+\t\trm .git/objects/info/commit-graph &&\n+\t\techo \"$(pwd)/../.git/objects\" >.git/objects/info/alternates &&\n+\t\ttest_commit new-commit &&\n+\t\tgit commit-graph write --reachable --split --changed-paths &&\n+\t\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\t\tverify_chain_files_exist $graphdir\n+\t)\n+'\n+\n+test_expect_success 'add three more commits, write a tip graph' '\n+\tgit reset --hard commits/3 &&\n+\tgit merge merge/1 &&\n+\tgit merge commits/5 &&\n+\tgit merge merge/2 &&\n+\tgit branch merge/3 &&\n+\tgit commit-graph write --reachable --split --changed-paths &&\n+\ttest_path_is_missing $infodir/commit-graph &&\n+\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\tls $graphdir/graph-*.graph >graph-files &&\n+\ttest_line_count = 2 graph-files &&\n+\tverify_chain_files_exist $graphdir\n+'\n+\n+graph_git_behavior 'split commit-graph: merge 3 vs 2' merge/3 merge/2\n+\n+test_expect_success 'add one commit, write a tip graph' '\n+\ttest_commit 11 &&\n+\tgit branch commits/11 &&\n+\tgit commit-graph write --reachable --split --changed-paths &&\n+\ttest_path_is_missing $infodir/commit-graph &&\n+\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\tls $graphdir/graph-*.graph >graph-files &&\n+\ttest_line_count = 3 graph-files &&\n+\tverify_chain_files_exist $graphdir\n+'\n+\n+graph_git_behavior 'three-layer commit-graph: commit 11 vs 6' commits/11 commits/6\n+\n+test_expect_success 'add one commit, write a merged graph' '\n+\ttest_commit 12 &&\n+\tgit branch commits/12 &&\n+\tgit commit-graph write --reachable --split --changed-paths &&\n+\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\tls $graphdir/graph-*.graph >graph-files &&\n+\ttest_line_count = 2 graph-files &&\n+\tverify_chain_files_exist $graphdir\n+'\n+\n+graph_git_behavior 'merged commit-graph: commit 12 vs 6' commits/12 commits/6\n+\n+test_expect_success 'create fork and chain across alternate' '\n+\tgit clone . fork &&\n+\t(\n+\t\tcd fork &&\n+\t\tgit config core.commitGraph true &&\n+\t\trm -rf $graphdir &&\n+\t\techo \"$(pwd)/../.git/objects\" >.git/objects/info/alternates &&\n+\t\ttest_commit 13 &&\n+\t\tgit branch commits/13 &&\n+\t\tgit commit-graph write --reachable --split --changed-paths &&\n+\t\ttest_path_is_file $graphdir/commit-graph-chain &&\n+\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n+\t\tls $graphdir/graph-*.graph >graph-files &&\n+\t\ttest_line_count = 1 graph-files &&\n+\t\tgit -c core.commitGraph=true  rev-list HEAD >expect &&\n+\t\tgit -c core.commitGraph=false rev-list HEAD >actual &&\n+\t\ttest_cmp expect actual &&\n+\t\ttest_commit 14 &&\n+\t\tgit commit-graph write --reachable --split --changed-paths --object-dir=.git/objects/ &&\n+\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n+\t\tls $graphdir/graph-*.graph >graph-files &&\n+\t\ttest_line_count = 1 graph-files\n+\t)\n+'\n+\n+graph_git_behavior 'alternate: commit 13 vs 6' commits/13 commits/6\n+\n+test_expect_success 'test merge stragety constants' '\n+\tgit clone . merge-2 &&\n+\t(\n+\t\tcd merge-2 &&\n+\t\tgit config core.commitGraph true &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\ttest_commit 14 &&\n+\t\tgit commit-graph write --reachable --split --changed-paths --size-multiple=2 &&\n+\t\ttest_line_count = 3 $graphdir/commit-graph-chain\n+\n+\t) &&\n+\tgit clone . merge-10 &&\n+\t(\n+\t\tcd merge-10 &&\n+\t\tgit config core.commitGraph true &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\ttest_commit 14 &&\n+\t\tgit commit-graph write --reachable --split --changed-paths --size-multiple=10 &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\t\tls $graphdir/graph-*.graph >graph-files &&\n+\t\ttest_line_count = 1 graph-files\n+\t) &&\n+\tgit clone . merge-10-expire &&\n+\t(\n+\t\tcd merge-10-expire &&\n+\t\tgit config core.commitGraph true &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\ttest_commit 15 &&\n+\t\tgit commit-graph write --reachable --split --changed-paths --size-multiple=10 --expire-time=1980-01-01 &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\t\tls $graphdir/graph-*.graph >graph-files &&\n+\t\ttest_line_count = 3 graph-files\n+\t) &&\n+\tgit clone --no-hardlinks . max-commits &&\n+\t(\n+\t\tcd max-commits &&\n+\t\tgit config core.commitGraph true &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\ttest_commit 16 &&\n+\t\ttest_commit 17 &&\n+\t\tgit commit-graph write --reachable --split --changed-paths --max-commits=1 &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n+\t\tls $graphdir/graph-*.graph >graph-files &&\n+\t\ttest_line_count = 1 graph-files\n+\t)\n+'\n+\n+test_expect_success 'remove commit-graph-chain file after flattening' '\n+\tgit clone . flatten &&\n+\t(\n+\t\tcd flatten &&\n+\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n+\t\tgit commit-graph write --reachable &&\n+\t\ttest_path_is_missing $graphdir/commit-graph-chain &&\n+\t\tls $graphdir >graph-files &&\n+\t\ttest_must_be_empty graph-files\n+\t)\n+'\n+\n+graph_git_behavior 'graph exists' merge/octopus commits/12\n+\n+test_expect_success 'split across alternate where alternate is not split' '\n+\tgit commit-graph write --reachable &&\n+\ttest_path_is_file .git/objects/info/commit-graph &&\n+\tcp .git/objects/info/commit-graph . &&\n+\tgit clone --no-hardlinks . alt-split &&\n+\t(\n+\t\tcd alt-split &&\n+\t\trm -f .git/objects/info/commit-graph &&\n+\t\techo \"$(pwd)\"/../.git/objects >.git/objects/info/alternates &&\n+\t\ttest_commit 18 &&\n+\t\tgit commit-graph write --reachable --split --changed-paths &&\n+\t\ttest_line_count = 1 $graphdir/commit-graph-chain\n+\t) &&\n+\ttest_cmp commit-graph .git/objects/info/commit-graph\n+'\n+\n+test_done\n-- \ngitgitgadget\n\n"},{"id":"388690","messageId":"e52c7ad37a306891487bd79a09b040bfb657d723.1576879520.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"[PATCH 2/9] commit-graph: write changed paths bloom filters","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-20T22:05:13Z","receivedAt":"2019-12-20T22:05:38Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nThe changed path bloom filters help determine which paths changed between a\ncommit and its first parent. We already have the \"--changed-paths\" option\nfor the \"git commit-graph write\" subcommand, now actually compute them under\nthat option. The COMMIT_GRAPH_WRITE_BLOOM_FILTERS flag enables this\ncomputation.\n\nRFC Notes: Here are some details about the implementation and I would love\nto know your thoughts and suggestions for improvements here.\n\nFor details on what bloom filters are and how they work, please refer to\nDr. Derrick Stolee's blog post [1].\n[1] https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-bloom-filters/\n\n1. The implementation sticks to the recommended values of 7 and 10 for the\n   number of hashes and the size of each entry, as described in the blog.\n   The implementation while not completely open to it at the moment, is flexible\n   enough to allow for tweaking these settings in the future.\n   Note: The performance gains we have observed so far with these values is\n   significant enough to not that we did not need to tweak these settings.\n   The cover letter of this series has the details and the commit where we have\n   git log use bloom filters.\n\n2. As described in the blog and the linked technical paper therin, we do not need\n   7 independent hashing functions. We use the Murmur3 hashing scheme - seed it\n   twice and then combine those to procure an arbitrary number of hash values.\n\n3. The filters are sized according to the number of changes in the each commit,\n   with minimum size of one 64 bit word.\n\n[Call for advice] We currently cap writing bloom filters for commits with\natmost 512 changed files. In the current implementation, we compute the diff,\nand then just throw it away once we see it has more than 512 changes.\nAny suggestiongs on how to reduce the work we are doing in this case are more\nthan welcome.\n\n[Call for advice] Would the git community like this commit to be split up into\nmore granular commits? This commit could possibly be split out further with the\nbloom.c code in its own commit, to be used by the commit-graph in a subsequent\ncommit. While I prefer it being contained in one commit this way, I am open to\nsuggestions.\n\n[Call for advice] Would a technical document explaining the exact details of\nthe bloom filter implemenation and the hashing calculations be helpful? I will\nbe adding details into Documentation/technical/commit-graph-format.txt, but the\nbloom filter code is an independent subsystem and could be used outside of the\ncommit-graph feature. Is it worth a separate document, or should we apply \"You\nAin't Gonna Need It\" principles?\n\n[Call for advice] I plan to add unit tests for bloom.c, specifically to ensure\nthat the hash algorithm and bloom key calculations are stable across versions.\n\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\n---\n Makefile       |   1 +\n bloom.c        | 201 +++++++++++++++++++++++++++++++++++++++++++++++++\n bloom.h        |  46 +++++++++++\n commit-graph.c |  32 +++++++-\n 4 files changed, 279 insertions(+), 1 deletion(-)\n create mode 100644 bloom.c\n create mode 100644 bloom.h\n\ndiff --git a/Makefile b/Makefile\nindex 42a061d3fb..9d5e26f5d6 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -838,6 +838,7 @@ LIB_OBJS += base85.o\n LIB_OBJS += bisect.o\n LIB_OBJS += blame.o\n LIB_OBJS += blob.o\n+LIB_OBJS += bloom.o\n LIB_OBJS += branch.o\n LIB_OBJS += bulk-checkin.o\n LIB_OBJS += bundle.o\ndiff --git a/bloom.c b/bloom.c\nnew file mode 100644\nindex 0000000000..08328cc381\n--- /dev/null\n+++ b/bloom.c\n@@ -0,0 +1,201 @@\n+#include \"git-compat-util.h\"\n+#include \"bloom.h\"\n+#include \"commit-graph.h\"\n+#include \"object-store.h\"\n+#include \"diff.h\"\n+#include \"diffcore.h\"\n+#include \"revision.h\"\n+#include \"hashmap.h\"\n+\n+#define BITS_PER_BLOCK 64\n+\n+define_commit_slab(bloom_filter_slab, struct bloom_filter);\n+\n+struct bloom_filter_slab bloom_filters;\n+\n+struct pathmap_hash_entry {\n+    struct hashmap_entry entry;\n+    const char path[FLEX_ARRAY];\n+};\n+\n+static uint32_t rotate_right(uint32_t value, int32_t count)\n+{\n+\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n+\tcount &= mask;\n+\treturn ((value >> count) | (value << ((-count) & mask)));\n+}\n+\n+static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n+{\n+\tconst uint32_t c1 = 0xcc9e2d51;\n+\tconst uint32_t c2 = 0x1b873593;\n+\tconst int32_t r1 = 15;\n+\tconst int32_t r2 = 13;\n+\tconst uint32_t m = 5;\n+\tconst uint32_t n = 0xe6546b64;\n+\tint i;\n+\tuint32_t k1 = 0;\n+\tconst char *tail;\n+\n+\tint len4 = len / sizeof(uint32_t);\n+\n+\tconst uint32_t *blocks = (const uint32_t*)data;\n+\n+\tuint32_t k;\n+\tfor (i = 0; i < len4; i++)\n+\t{\n+\t\tk = blocks[i];\n+\t\tk *= c1;\n+\t\tk = rotate_right(k, r1);\n+\t\tk *= c2;\n+\n+\t\tseed ^= k;\n+\t\tseed = rotate_right(seed, r2) * m + n;\n+\t}\n+\n+\ttail = (data + len4 * sizeof(uint32_t));\n+\n+\tswitch (len & (sizeof(uint32_t) - 1))\n+\t{\n+\tcase 3:\n+\t\tk1 ^= ((uint32_t)tail[2]) << 16;\n+\t\t/*-fallthrough*/\n+\tcase 2:\n+\t\tk1 ^= ((uint32_t)tail[1]) << 8;\n+\t\t/*-fallthrough*/\n+\tcase 1:\n+\t\tk1 ^= ((uint32_t)tail[0]) << 0;\n+\t\tk1 *= c1;\n+\t\tk1 = rotate_right(k1, r1);\n+\t\tk1 *= c2;\n+\t\tseed ^= k1;\n+\t\tbreak;\n+\t}\n+\n+\tseed ^= (uint32_t)len;\n+\tseed ^= (seed >> 16);\n+\tseed *= 0x85ebca6b;\n+\tseed ^= (seed >> 13);\n+\tseed *= 0xc2b2ae35;\n+\tseed ^= (seed >> 16);\n+\n+\treturn seed;\n+}\n+\n+static inline uint64_t get_bitmask(uint32_t pos)\n+{\n+\treturn ((uint64_t)1) << (pos & (BITS_PER_BLOCK - 1));\n+}\n+\n+void fill_bloom_key(const char *data,\n+\t\t    int len,\n+\t\t    struct bloom_key *key,\n+\t\t    struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tuint32_t seed0 = 0x293ae76f;\n+\tuint32_t seed1 = 0x7e646e2c;\n+\n+\tuint32_t hash0 = seed_murmur3(seed0, data, len);\n+\tuint32_t hash1 = seed_murmur3(seed1, data, len);\n+\n+\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n+\tfor (i = 0; i < settings->num_hashes; i++)\n+\t\tkey->hashes[i] = hash0 + i * hash1;\n+}\n+\n+static void add_key_to_filter(struct bloom_key *key,\n+\t\t\t      struct bloom_filter *filter,\n+\t\t\t      struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tuint64_t mod = filter->len * BITS_PER_BLOCK;\n+\n+\tfor (i = 0; i < settings->num_hashes; i++) {\n+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n+\t\tuint64_t block_pos = hash_mod / BITS_PER_BLOCK;\n+\n+\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n+\t}\n+}\n+\n+void load_bloom_filters(void)\n+{\n+\tinit_bloom_filter_slab(&bloom_filters);\n+}\n+\n+struct bloom_filter *get_bloom_filter(struct repository *r,\n+\t\t\t\t      struct commit *c)\n+{\n+\tstruct bloom_filter *filter;\n+\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\tint i;\n+\tstruct rev_info revs;\n+\tconst char *revs_argv[] = {NULL, \"HEAD\", NULL};\n+\n+\tfilter = bloom_filter_slab_at(&bloom_filters, c);\n+\tinit_revisions(&revs, NULL);\n+\trevs.diffopt.flags.recursive = 1;\n+\n+\tsetup_revisions(2, revs_argv, &revs, NULL);\n+\n+\tif (c->parents)\n+\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &revs.diffopt);\n+\telse\n+\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &revs.diffopt);\n+\tdiffcore_std(&revs.diffopt);\n+\n+\tif (diff_queued_diff.nr <= 512) {\n+\t\tstruct hashmap pathmap;\n+\t\tstruct pathmap_hash_entry* e;\n+\t\tstruct hashmap_iter iter;\n+\t\thashmap_init(&pathmap, NULL, NULL, 0);\n+\n+\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n+\t\t    const char* path = diff_queued_diff.queue[i]->two->path;\n+\t\t    const char* p = path;\n+\n+\t\t    /*\n+\t\t     * Add each leading directory of the changed file, i.e. for\n+\t\t     * 'dir/subdir/file' add 'dir' and 'dir/subdir' as well, so\n+\t\t     * the Bloom filter could be used to speed up commands like\n+\t\t     * 'git log dir/subdir', too.\n+\t\t     *\n+\t\t     * Note that directories are added without the trailing '/'.\n+\t\t     */\n+\t\t    do {\n+\t\t\t\tchar* last_slash = strrchr(p, '/');\n+\n+\t\t\t\tFLEX_ALLOC_STR(e, path, path);\n+\t\t\t\thashmap_entry_init(&e->entry, strhash(p));\n+\t\t\t\thashmap_add(&pathmap, &e->entry);\n+\n+\t\t\t\tif (!last_slash)\n+\t\t\t\t    last_slash = (char*)p;\n+\t\t\t\t*last_slash = '\\0';\n+\n+\t\t    } while (*p);\n+\n+\t\t    diff_free_filepair(diff_queued_diff.queue[i]);\n+\t\t}\n+\n+\t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_BLOCK - 1) / BITS_PER_BLOCK;\n+\t\tfilter->data = xcalloc(filter->len, sizeof(uint64_t));\n+\n+\t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n+\t\t    struct bloom_key key;\n+\t\t    fill_bloom_key(e->path, strlen(e->path), &key, &settings);\n+\t\t    add_key_to_filter(&key, filter, &settings);\n+\t\t}\n+\n+\t\thashmap_free_entries(&pathmap, struct pathmap_hash_entry, entry);\n+\t} else {\n+\t\tfilter->data = NULL;\n+\t\tfilter->len = 0;\n+\t}\n+\n+\tfree(diff_queued_diff.queue);\n+\tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n+\n+\treturn filter;\n+}\n\\ No newline at end of file\ndiff --git a/bloom.h b/bloom.h\nnew file mode 100644\nindex 0000000000..ba8ae70b67\n--- /dev/null\n+++ b/bloom.h\n@@ -0,0 +1,46 @@\n+#ifndef BLOOM_H\n+#define BLOOM_H\n+\n+struct commit;\n+struct repository;\n+\n+struct bloom_filter_settings {\n+\tuint32_t hash_version;\n+\tuint32_t num_hashes;\n+\tuint32_t bits_per_entry;\n+};\n+\n+#define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n+\n+/*\n+ * A bloom_filter struct represents a data segment to\n+ * use when testing hash values. The 'len' member\n+ * dictates how many uint64_t entries are stored in\n+ * 'data'.\n+ */\n+struct bloom_filter {\n+\tuint64_t *data;\n+\tint len;\n+};\n+\n+/*\n+ * A bloom_key represents the k hash values for a\n+ * given hash input. These can be precomputed and\n+ * stored in a bloom_key for re-use when testing\n+ * against a bloom_filter.\n+ */\n+struct bloom_key {\n+\tuint32_t *hashes;\n+};\n+\n+void load_bloom_filters(void);\n+\n+struct bloom_filter *get_bloom_filter(struct repository *r,\n+\t\t\t\t      struct commit *c);\n+\n+void fill_bloom_key(const char *data,\n+\t\t    int len,\n+\t\t    struct bloom_key *key,\n+\t\t    struct bloom_filter_settings *settings);\n+\n+#endif\ndiff --git a/commit-graph.c b/commit-graph.c\nindex e771394aff..61e60ff98a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -16,6 +16,7 @@\n #include \"hashmap.h\"\n #include \"replace-object.h\"\n #include \"progress.h\"\n+#include \"bloom.h\"\n \n #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n@@ -794,9 +795,11 @@ struct write_commit_graph_context {\n \tunsigned append:1,\n \t\t report_progress:1,\n \t\t split:1,\n-\t\t check_oids:1;\n+\t\t check_oids:1,\n+\t\t bloom:1;\n \n \tconst struct split_commit_graph_opts *split_opts;\n+\tuint32_t total_bloom_filter_size;\n };\n \n static void write_graph_chunk_fanout(struct hashfile *f,\n@@ -1139,6 +1142,28 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \tstop_progress(&ctx->progress);\n }\n \n+static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n+{\n+\tint i;\n+\tstruct progress *progress = NULL;\n+\n+\tload_bloom_filters();\n+\n+\tif (ctx->report_progress)\n+\t\tprogress = start_progress(\n+\t\t\t_(\"Computing commit diff Bloom filters\"),\n+\t\t\tctx->commits.nr);\n+\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n+\t\tctx->total_bloom_filter_size += sizeof(uint64_t) * filter->len;\n+\t\tdisplay_progress(progress, i + 1);\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n static int add_ref_to_list(const char *refname,\n \t\t\t   const struct object_id *oid,\n \t\t\t   int flags, void *cb_data)\n@@ -1791,6 +1816,8 @@ int write_commit_graph(const char *obj_dir,\n \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n \tctx->check_oids = flags & COMMIT_GRAPH_WRITE_CHECK_OIDS ? 1 : 0;\n \tctx->split_opts = split_opts;\n+\tctx->bloom = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n+\tctx->total_bloom_filter_size = 0;\n \n \tif (ctx->split) {\n \t\tstruct commit_graph *g;\n@@ -1885,6 +1912,9 @@ int write_commit_graph(const char *obj_dir,\n \n \tcompute_generation_numbers(ctx);\n \n+\tif (ctx->bloom)\n+\t\tcompute_bloom_filters(ctx);\n+\n \tres = write_commit_graph_file(ctx);\n \n \tif (ctx->split)\n-- \ngitgitgadget\n\n"},{"id":"388691","messageId":"xmqq5zia8x1g.fsf@gitster-ct.c.googlers.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-20T22:14:19Z","receivedAt":"2019-12-20T22:14:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Adopting changed path bloom filters has been discussed on the list before,\n> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n> Derrick Stolee [1]. This series is based on Dr. Stolee's approach [2] and\n> presents an updated and more polished RFC version of the feature. \n\n;-)\n"},{"id":"388711","messageId":"39ef6bc6-4f21-1ba6-ad6e-06cb1a2423ac@iee.email","threadId":"52499","inReplyTo":"e52c7ad37a306891487bd79a09b040bfb657d723.1576879520.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/9] commit-graph: write changed paths bloom filters","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2019-12-21T16:48:44Z","receivedAt":"2019-12-21T16:48:46Z","isPatch":true,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"spelling nit?\n\nOn 20/12/2019 22:05, Garima Singh via GitGitGadget wrote:\n> 1. The implementation sticks to the recommended values of 7 and 10 for the\n>    number of hashes and the size of each entry, as described in the blog.\n>    The implementation while not completely open to it at the moment, is flexible\n>    enough to allow for tweaking these settings in the future.\n>    Note: The performance gains we have observed so far with these values is\n>    significant enough to not that we did not need to tweak these settings.\ns/not/note/ (first occurrence)\n>    The cover letter of this series has the details and the commit where we have\n>    git log use bloom filters.\nPhilip\n"},{"id":"388787","messageId":"CAP8UFD1mOEUngLofTZ2hjsTooR49FktfWHWJGzQ9Y-a=oB-mZw@mail.gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2019-12-22T09:26:20Z","receivedAt":"2019-12-22T09:30:41Z","isPatch":true,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"Hi,\n\nOn Fri, Dec 20, 2019 at 11:07 PM Garima Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> The commit graph feature brought in a lot of performance improvements across\n> multiple commands. However, file based history continues to be a performance\n> pain point, especially in large repositories.\n>\n> Adopting changed path bloom filters has been discussed on the list before,\n> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n> Derrick Stolee [1]. This series is based on Dr. Stolee's approach [2] and\n> presents an updated and more polished RFC version of the feature.\n\nThanks for working on this!\n\n> Performance Gains: We tested the performance of git log -- path on the git\n> repo, the linux repo and some internal large repos, with a variety of paths\n> of varying depths.\n>\n> On the git and linux repos: We observed a 2x to 5x speed up.\n>\n> On a large internal repo with files seated 6-10 levels deep in the tree: We\n> observed 10x to 20x speed ups, with some paths going up to 28 times faster.\n\nVery nice!\n\nI have a question though. Are the performance gains only available\nwith `git log -- path` or are they already available for example when\ndoing a partial clone and/or a sparse checkout?\n\n> Future Work (not included in the scope of this series):\n>\n>  1. Supporting multiple path based revision walk\n>  2. Adopting it in git blame logic.\n>  3. Interactions with line log git log -L\n\nGreat!\n\n> This series is intended to start the conversation and many of the commit\n> messages include specific call outs for suggestions and thoughts.\n\nI think Peff said during the Virtual Contributor Summit that he was\ninterested in using bitmaps to speed up partial clone on the server\nside. Would it make sense to use both bitmaps and bloom filters?\n\nThanks,\nChristian.\n"},{"id":"388788","messageId":"20191222093036.GA3449072@coredump.intra.peff.net","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-12-22T09:30:36Z","receivedAt":"2019-12-22T09:30:41Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Dec 20, 2019 at 10:05:11PM +0000, Garima Singh via GitGitGadget wrote:\n\n> Adopting changed path bloom filters has been discussed on the list before,\n> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n> Derrick Stolee [1]. This series is based on Dr. Stolee's approach [2] and\n> presents an updated and more polished RFC version of the feature.\n\nGreat to see progress here. I probably won't have time to review this\ncarefully before the new year, but I did notice some low-hanging fruit\non the generation side.\n\nSo here are a few patches to reduce the CPU and memory usage. They could\nbe squashed in at the appropriate spots, or perhaps taken as inspiration\nif there are better solutions (especially for the first one).\n\nI think we could go further still, by actually doing a non-recursive\ndiff_tree_oid(), and then recursing into sub-trees ourselves. That would\nsave us having to split apart each path to add the leading paths to the\nhashmap (most of which will be duplicates if the commit touched \"a/b/c\"\nand \"a/b/d\", etc). I doubt it would be that huge a speedup though. We\nhave to keep a list of the touched paths anyway (since the bloom key\nparameters depend on the number of entries), and most of the time is\nalmost certainly spent inflating the trees in the first place. However\nit might be easier to follow the code, and it would make it simpler to\nstop traversing at the 512-entry limit, rather than generating a huge\ndiff only to throw it away.\n\n  [1/3]: commit-graph: examine changed-path objects in pack order\n  [2/3]: commit-graph: free large diffs, too\n  [3/3]: commit-graph: stop using full rev_info for diffs\n\n bloom.c        | 18 +++++++++---------\n commit-graph.c | 34 +++++++++++++++++++++++++++++++++-\n 2 files changed, 42 insertions(+), 10 deletions(-)\n\n-Peff\n"},{"id":"388789","messageId":"20191222093206.GA3460818@coredump.intra.peff.net","threadId":"52499","inReplyTo":"20191222093036.GA3449072@coredump.intra.peff.net","subject":"[PATCH 1/3] commit-graph: examine changed-path objects in pack order","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-12-22T09:32:06Z","receivedAt":"2019-12-22T09:32:09Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"Looking at the diff of commit objects in pack order is much faster than\nin sha1 order, as it gives locality to the access of tree deltas\n(whereas sha1 order is effectively random). Unfortunately the\ncommit-graph code sorts the commits (several times, sometimes as an oid\nand sometimes a pointer-to-commit), and we ultimately traverse in sha1\norder.\n\nInstead, let's remember the position at which we see each commit, and\ntraverse in that order when looking at bloom filters. This drops my time\nfor \"git commit-graph write --changed-paths\" in linux.git from ~4\nminutes to ~1.5 minutes.\n\nProbably the \"--reachable\" code path would want something similar.\n\nOr alternatively, we could use a different data structure (either a\nhash, or maybe even just a bit in \"struct commit\") to keep track of\nwhich oids we've seen, etc instead of sorting. And then we could keep\nthe original order.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n commit-graph.c | 34 +++++++++++++++++++++++++++++++++-\n 1 file changed, 33 insertions(+), 1 deletion(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 0580ce75d5..bf6c663772 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -17,6 +17,7 @@\n #include \"replace-object.h\"\n #include \"progress.h\"\n #include \"bloom.h\"\n+#include \"commit-slab.h\"\n \n #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n@@ -48,6 +49,29 @@\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n+/* Keep track of the order in which commits are added to our list. */\n+define_commit_slab(commit_pos, int);\n+static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n+\n+static void set_commit_pos(struct repository *r, const struct object_id *oid)\n+{\n+\tstatic int32_t max_pos;\n+\tstruct commit *commit = lookup_commit(r, oid);\n+\n+\tif (!commit)\n+\t\treturn; /* should never happen, but be lenient */\n+\n+\t*commit_pos_at(&commit_pos, commit) = max_pos++;\n+}\n+\n+static int commit_pos_cmp(const void *va, const void *vb)\n+{\n+\tconst struct commit *a = *(const struct commit **)va;\n+\tconst struct commit *b = *(const struct commit **)vb;\n+\treturn commit_pos_at(&commit_pos, a) -\n+\t       commit_pos_at(&commit_pos, b);\n+}\n+\n char *get_commit_graph_filename(const char *obj_dir)\n {\n \tchar *filename = xstrfmt(\"%s/info/commit-graph\", obj_dir);\n@@ -1088,6 +1112,8 @@ static int add_packed_commits(const struct object_id *oid,\n \toidcpy(&(ctx->oids.list[ctx->oids.nr]), oid);\n \tctx->oids.nr++;\n \n+\tset_commit_pos(ctx->r, oid);\n+\n \treturn 0;\n }\n \n@@ -1208,6 +1234,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n {\n \tint i;\n \tstruct progress *progress = NULL;\n+\tstruct commit **sorted_by_pos;\n \n \tload_bloom_filters();\n \n@@ -1216,13 +1243,18 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \t\t\t_(\"Computing commit diff Bloom filters\"),\n \t\t\tctx->commits.nr);\n \n+\tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n+\tCOPY_ARRAY(sorted_by_pos, ctx->commits.list, ctx->commits.nr);\n+\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n+\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tstruct commit *c = ctx->commits.list[i];\n+\t\tstruct commit *c = sorted_by_pos[i];\n \t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n \t\tctx->total_bloom_filter_size += sizeof(uint64_t) * filter->len;\n \t\tdisplay_progress(progress, i + 1);\n \t}\n \n+\tfree(sorted_by_pos);\n \tstop_progress(&progress);\n }\n \n-- \n2.24.1.1152.gda0b849012\n\n"},{"id":"388790","messageId":"20191222093216.GB3460818@coredump.intra.peff.net","threadId":"52499","inReplyTo":"20191222093036.GA3449072@coredump.intra.peff.net","subject":"[PATCH 2/3] commit-graph: free large diffs, too","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-12-22T09:32:16Z","receivedAt":"2019-12-22T09:32:19Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"If a diff we compute for --changed-path has more than 512 entries, we\ndon't bother generating a bloom filter for it. But since we don't\niterate over diff_queued_diff, we also don't free the filepairs and\nfilespecs from the diff before clearing the queue. Let's make sure we do\nso.\n\nThis drops the peak heap usage of \"commit-graph write --changed-paths\"\non linux.git from ~8GB to ~4GB.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n bloom.c | 2 ++\n 1 file changed, 2 insertions(+)\n\ndiff --git a/bloom.c b/bloom.c\nindex 0c7505d3d6..d1d3796e11 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -226,6 +226,8 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \t\thashmap_free_entries(&pathmap, struct pathmap_hash_entry, entry);\n \t} else {\n+\t\tfor (i = 0; i < diff_queued_diff.nr; i++)\n+\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n \t\tfilter->data = NULL;\n \t\tfilter->len = 0;\n \t}\n-- \n2.24.1.1152.gda0b849012\n\n"},{"id":"388791","messageId":"20191222093222.GC3460818@coredump.intra.peff.net","threadId":"52499","inReplyTo":"20191222093036.GA3449072@coredump.intra.peff.net","subject":"[PATCH 3/3] commit-graph: stop using full rev_info for diffs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-12-22T09:32:22Z","receivedAt":"2019-12-22T09:32:24Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"When we perform a diff to get the set of changed paths for a commit,\nwe initialize a full \"struct rev_info\" with setup_revisions(). But the\nonly part of it we use is the diff_options struct. Besides being overly\ncomplex, this also leaks memory, as we use the fake argv to\nsetup_revisions() create a pending array which is never cleared.\n\nLet's just use diff_options directly. This reduces the peak heap usage\nof \"git commit-graph write --changed-paths\" on linux.git from ~4GB to\n~1.2GB.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n bloom.c | 16 +++++++---------\n 1 file changed, 7 insertions(+), 9 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex d1d3796e11..ea77631cc2 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -154,8 +154,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \tstruct bloom_filter *filter;\n \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \tint i;\n-\tstruct rev_info revs;\n-\tconst char *revs_argv[] = {NULL, \"HEAD\", NULL};\n+\tstruct diff_options diffopt;\n \n \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n \n@@ -170,16 +169,15 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \tif (filter->data || !compute_if_null)\n \t\t\treturn filter;\n \n-\tinit_revisions(&revs, NULL);\n-\trevs.diffopt.flags.recursive = 1;\n-\n-\tsetup_revisions(2, revs_argv, &revs, NULL);\n+\trepo_diff_setup(r, &diffopt);\n+\tdiffopt.flags.recursive = 1;\n+\tdiff_setup_done(&diffopt);\n \n \tif (c->parents)\n-\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &revs.diffopt);\n+\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &diffopt);\n \telse\n-\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &revs.diffopt);\n-\tdiffcore_std(&revs.diffopt);\n+\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n+\tdiffcore_std(&diffopt);\n \n \tif (diff_queued_diff.nr <= 512) {\n \t\tstruct hashmap pathmap;\n-- \n2.24.1.1152.gda0b849012\n"},{"id":"388792","messageId":"20191222093857.GB3449072@coredump.intra.peff.net","threadId":"52499","inReplyTo":"CAP8UFD1mOEUngLofTZ2hjsTooR49FktfWHWJGzQ9Y-a=oB-mZw@mail.gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-12-22T09:38:57Z","receivedAt":"2019-12-22T09:39:00Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Dec 22, 2019 at 10:26:20AM +0100, Christian Couder wrote:\n\n> I have a question though. Are the performance gains only available\n> with `git log -- path` or are they already available for example when\n> doing a partial clone and/or a sparse checkout?\n\nFrom my quick look at the code, anything that feeds a pathspec to a\nrevision traversal would be helped. I'm not sure if it would help for\npartial/sparse traversals, though. There we actually need to know which\nblobs correspond to the paths in question, not just whether any\nparticular commit touched them.\n\nI also took a brief look at adding support to the custom blame-tree\nimplementation we use at GitHub, and got about a 6x speedup.\n\n> > This series is intended to start the conversation and many of the commit\n> > messages include specific call outs for suggestions and thoughts.\n> \n> I think Peff said during the Virtual Contributor Summit that he was\n> interested in using bitmaps to speed up partial clone on the server\n> side. Would it make sense to use both bitmaps and bloom filters?\n\nI think they're orthogonal. For size-based filters on blobs, you'd just\nuse bitmaps as normal, because you can post-process the result to check\nthe type and size of each object in the list (and I have patches to do\nthis, but they need some polishing and we're not yet running them).\n\nFor path-based filters like a sparse specification, you can't use\nbitmaps at all; you have to do a real traversal. But there you still\ngenerally get all of the commits. I guess if a commit doesn't touch any\npath you're interested in, you could avoid walking into its tree at all,\nwhich might help. I haven't given it much thought yet.\n\n-Peff\n"},{"id":"388918","messageId":"fc30441a-1bb7-73e5-43f6-6e26824e04f6@gmail.com","threadId":"52499","inReplyTo":"20191222093036.GA3449072@coredump.intra.peff.net","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-12-26T14:21:36Z","receivedAt":"2019-12-26T14:21:41Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/22/2019 4:30 AM, Jeff King wrote:\n> On Fri, Dec 20, 2019 at 10:05:11PM +0000, Garima Singh via GitGitGadget wrote:\n> \n>> Adopting changed path bloom filters has been discussed on the list before,\n>> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n>> Derrick Stolee [1]. This series is based on Dr. Stolee's approach [2] and\n>> presents an updated and more polished RFC version of the feature.\n> \n> Great to see progress here. I probably won't have time to review this\n> carefully before the new year, but I did notice some low-hanging fruit\n> on the generation side.\n> \n> So here are a few patches to reduce the CPU and memory usage. They could\n> be squashed in at the appropriate spots, or perhaps taken as inspiration\n> if there are better solutions (especially for the first one).\n> \n> I think we could go further still, by actually doing a non-recursive\n> diff_tree_oid(), and then recursing into sub-trees ourselves. That would\n> save us having to split apart each path to add the leading paths to the\n> hashmap (most of which will be duplicates if the commit touched \"a/b/c\"\n> and \"a/b/d\", etc). I doubt it would be that huge a speedup though. We\n> have to keep a list of the touched paths anyway (since the bloom key\n> parameters depend on the number of entries), and most of the time is\n> almost certainly spent inflating the trees in the first place. However\n> it might be easier to follow the code, and it would make it simpler to\n> stop traversing at the 512-entry limit, rather than generating a huge\n> diff only to throw it away.\n\nThanks for these improvements. This diff machinery is new to us (Garima\nand myself).\n\nHere are some recommendations (to Garima) for how to proceed with these\npatches. Please let me know if anyone disagrees.\n\n>   [1/3]: commit-graph: examine changed-path objects in pack order\n\nThis one is best kept as its own patch, as it shows a clear reason why\nwe want to do the sort-by-position. It would also be a complicated\npatch to include this logic along with the first use of\ncompute_bloom_filters().\n\n>   [2/3]: commit-graph: free large diffs, too\nThis one seems best to squash into \"commit-graph: write changed paths\nbloom filters\" with a Helped-by for Peff.\n\n>   [3/3]: commit-graph: stop using full rev_info for diffs\n\nWhile I appreciate the clear benefit in the commit-message here, it\nmay be best to also squash this one similarly.\n\nOf course, if we create our own diff logic with the short-circuit\ncapability, then perhaps these patches become obsolete. I'll spend\na little time playing with options here.\n\nThanks!\n-Stolee\n"},{"id":"388991","messageId":"8b331ef6-f431-56ef-37a9-1d6e263ea0fe@gmail.com","threadId":"52499","inReplyTo":"20191222093206.GA3460818@coredump.intra.peff.net","subject":"Re: [PATCH 1/3] commit-graph: examine changed-path objects in pack order","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-12-27T14:51:02Z","receivedAt":"2019-12-27T14:51:06Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/22/2019 4:32 AM, Jeff King wrote:\n> Looking at the diff of commit objects in pack order is much faster than\n> in sha1 order, as it gives locality to the access of tree deltas\n> (whereas sha1 order is effectively random). Unfortunately the\n> commit-graph code sorts the commits (several times, sometimes as an oid\n> and sometimes a pointer-to-commit), and we ultimately traverse in sha1\n> order.\n> \n> Instead, let's remember the position at which we see each commit, and\n> traverse in that order when looking at bloom filters. This drops my time\n> for \"git commit-graph write --changed-paths\" in linux.git from ~4\n> minutes to ~1.5 minutes.\n\nI'm doing my own perf tests on these patches, and my copy of linux.git\nhas four packs of varying sizes (corresponding with my rare fetches and\nlack of repacks). My time goes from 3m50s to 3m00s. I was confused at\nfirst, but then realized that I used the \"--reachable\" flag. In that\ncase, we never run set_commit_pos(), so all positions are equal and the\nsort is not helpful.\n\nI thought that inserting some set_commit_pos() calls into close_reachable()\nand add_missing_parents() would give some amount of time-order to the\ncommits as we compute the filters. However, the time did not change at\nall.\n\nI've included the patch below for reference, anyway.\n\nThanks,\n-Stolee\n\n-->8--\n\nFrom e7c63d8db09be81ce213ba7f112bb3d2f537bf4a Mon Sep 17 00:00:00 2001\nFrom: Derrick Stolee <dstolee@microsoft.com>\nDate: Fri, 27 Dec 2019 09:47:49 -0500\nSubject: [PATCH] commit-graph: set commit positions for --reachable\n\nWhen running 'git commit-graph write --changed-paths', we sort the\ncommits by pack-order to save time when computing the changed-paths\nbloom filters. This does not help when finding the commits via the\n--reachable flag.\n\nAdd calls to set_commit_pos() when walking the reachable commits,\nwhich provides an ordering similar to a topological ordering.\n\nUnfortunately, the performance did not improve with this change.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\n---\n commit-graph.c | 6 +++++-\n 1 file changed, 5 insertions(+), 1 deletion(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex bf6c663772..a6c4ab401e 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1126,6 +1126,8 @@ static void add_missing_parents(struct write_commit_graph_context *ctx, struct c\n \t\t\toidcpy(&ctx->oids.list[ctx->oids.nr], &(parent->item->object.oid));\n \t\t\tctx->oids.nr++;\n \t\t\tparent->item->object.flags |= REACHABLE;\n+\n+\t\t\tset_commit_pos(ctx->r, &parent->item->object.oid);\n \t\t}\n \t}\n }\n@@ -1142,8 +1144,10 @@ static void close_reachable(struct write_commit_graph_context *ctx)\n \tfor (i = 0; i < ctx->oids.nr; i++) {\n \t\tdisplay_progress(ctx->progress, i + 1);\n \t\tcommit = lookup_commit(ctx->r, &ctx->oids.list[i]);\n-\t\tif (commit)\n+\t\tif (commit) {\n \t\t\tcommit->object.flags |= REACHABLE;\n+\t\t\tset_commit_pos(ctx->r, &commit->object.oid);\n+\t\t}\n \t}\n \tstop_progress(&ctx->progress);\n \n-- \n2.25.0.rc0\n\n"},{"id":"388992","messageId":"e932bed5-f90c-da51-7d45-54e14aa27734@gmail.com","threadId":"52499","inReplyTo":"20191222093216.GB3460818@coredump.intra.peff.net","subject":"Re: [PATCH 2/3] commit-graph: free large diffs, too","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-12-27T14:52:10Z","receivedAt":"2019-12-27T14:52:13Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/22/2019 4:32 AM, Jeff King wrote:\n> If a diff we compute for --changed-path has more than 512 entries, we\n> don't bother generating a bloom filter for it. But since we don't\n> iterate over diff_queued_diff, we also don't free the filepairs and\n> filespecs from the diff before clearing the queue. Let's make sure we do\n> so.\n> \n> This drops the peak heap usage of \"commit-graph write --changed-paths\"\n> on linux.git from ~8GB to ~4GB.\n\nIn my testing, the heap size went from ~10gb to ~6gb.\n\nThanks,\n-Stolee\n"},{"id":"388993","messageId":"061c6800-1516-ddd9-968d-a1274e93d6a1@gmail.com","threadId":"52499","inReplyTo":"20191222093222.GC3460818@coredump.intra.peff.net","subject":"Re: [PATCH 3/3] commit-graph: stop using full rev_info for diffs","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-12-27T14:53:23Z","receivedAt":"2019-12-27T14:53:26Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/22/2019 4:32 AM, Jeff King wrote:\n> When we perform a diff to get the set of changed paths for a commit,\n> we initialize a full \"struct rev_info\" with setup_revisions(). But the\n> only part of it we use is the diff_options struct. Besides being overly\n> complex, this also leaks memory, as we use the fake argv to\n> setup_revisions() create a pending array which is never cleared.\n> \n> Let's just use diff_options directly. This reduces the peak heap usage\n> of \"git commit-graph write --changed-paths\" on linux.git from ~4GB to\n> ~1.2GB.\n\nIn my testing, this went from ~6gb to ~4gb.\n\nI'm guessing that my memory difference is related to how poorly my\npacks are repacked/redeltified.\n\nThanks,\n-Stolee\n\n"},{"id":"388995","messageId":"e9a4e4ff-5466-dc39-c3f5-c9a8b8f2f11d@gmail.com","threadId":"52499","inReplyTo":"20191222093036.GA3449072@coredump.intra.peff.net","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-12-27T16:11:37Z","receivedAt":"2019-12-27T16:11:45Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/22/2019 4:30 AM, Jeff King wrote:\n> On Fri, Dec 20, 2019 at 10:05:11PM +0000, Garima Singh via GitGitGadget wrote:\n> \n>> Adopting changed path bloom filters has been discussed on the list before,\n>> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n>> Derrick Stolee [1]. This series is based on Dr. Stolee's approach [2] and\n>> presents an updated and more polished RFC version of the feature.\n> \n> Great to see progress here. I probably won't have time to review this\n> carefully before the new year, but I did notice some low-hanging fruit\n> on the generation side.\n> \n> So here are a few patches to reduce the CPU and memory usage. They could\n> be squashed in at the appropriate spots, or perhaps taken as inspiration\n> if there are better solutions (especially for the first one).\n\nI tested these patches with the Linux kernel repo and reported my results\non each patch. However, I wanted to also test on a larger internal repo\n(the AzureDevOps repo), which has ~500 commits with more than 512 changes,\nand generally has larger diffs than the Linux kernel repo.\n\n| Version  | Time   | Memory |\n|----------|--------|--------|\n| Garima   | 16m36s | 27.0gb |\n| Peff 1   | 6m32s  | 28.0gb |\n| Peff 2   | 6m48s  |  5.6gb |\n| Peff 3   | 6m14s  |  4.5gb |\n| Shortcut | 3m47s  |  4.5gb |\n\nFor reference, I found the time and memory information using\n\"/usr/bin/time --verbose\" in a bash script.\n \n> I think we could go further still, by actually doing a non-recursive\n> diff_tree_oid(), and then recursing into sub-trees ourselves. That would\n> save us having to split apart each path to add the leading paths to the\n> hashmap (most of which will be duplicates if the commit touched \"a/b/c\"\n> and \"a/b/d\", etc). I doubt it would be that huge a speedup though. We\n> have to keep a list of the touched paths anyway (since the bloom key\n> parameters depend on the number of entries), and most of the time is\n> almost certainly spent inflating the trees in the first place. However\n> it might be easier to follow the code, and it would make it simpler to\n> stop traversing at the 512-entry limit, rather than generating a huge\n> diff only to throw it away.\n\nBy \"Shortcut\" in the table above, I mean the following patch on top of\nGarima's and Peff's changes. It inserts a max_changes option into struct\ndiff_options to halt the diff early. This seemed like an easier change\nthan creating a new tree diff algorithm wholesale.\n\nThanks,\n-Stolee\n\n-->8--\n\nFrom: Derrick Stolee <dstolee@microsoft.com>\nDate: Fri, 27 Dec 2019 10:13:48 -0500\nSubject: [PATCH] diff: halt tree-diff early after max_changes\n\nWhen computing the changed-paths bloom filters for the commit-graph,\nwe limit the size of the filter by restricting the number of paths\nin the diff. Instead of computing a large diff and then ignoring the\nresult, it is better to halt the diff computation early.\n\nCreate a new \"max_changes\" option in struct diff_options. If non-zero,\nthen halt the diff computation after discovering strictly more changed\npaths. This includes paths corresponding to trees that change.\n\nUse this max_changes option in the bloom filter calculations. This\nreduces the time taken to compute the filters for the Linux kernel\nrepo from 2m50s to 2m35s. For a larger repo with more commits changing\nmany paths, the time reduces from 6 minutes to under 4 minutes.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\n---\n bloom.c     | 4 +++-\n diff.h      | 5 +++++\n tree-diff.c | 5 +++++\n 3 files changed, 13 insertions(+), 1 deletion(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex ea77631cc2..83dde2378b 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -155,6 +155,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \tint i;\n \tstruct diff_options diffopt;\n+\tint max_changes = 512;\n \n \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n \n@@ -171,6 +172,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \trepo_diff_setup(r, &diffopt);\n \tdiffopt.flags.recursive = 1;\n+\tdiffopt.max_changes = max_changes;\n \tdiff_setup_done(&diffopt);\n \n \tif (c->parents)\n@@ -179,7 +181,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n \tdiffcore_std(&diffopt);\n \n-\tif (diff_queued_diff.nr <= 512) {\n+\tif (diff_queued_diff.nr <= max_changes) {\n \t\tstruct hashmap pathmap;\n \t\tstruct pathmap_hash_entry* e;\n \t\tstruct hashmap_iter iter;\ndiff --git a/diff.h b/diff.h\nindex 6febe7e365..9443dc1b00 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -285,6 +285,11 @@ struct diff_options {\n \t/* Number of hexdigits to abbreviate raw format output to. */\n \tint abbrev;\n \n+\t/* If non-zero, then stop computing after this many changes. */\n+\tint max_changes;\n+\t/* For internal use only. */\n+\tint num_changes;\n+\n \tint ita_invisible_in_index;\n /* white-space error highlighting */\n #define WSEH_NEW (1<<12)\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 33ded7f8b3..16a21d9f34 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -434,6 +434,9 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \n+\t\tif (opt->max_changes && opt->num_changes > opt->max_changes)\n+\t\t\tbreak;\n+\n \t\tif (opt->pathspec.nr) {\n \t\t\tskip_uninteresting(&t, base, opt);\n \t\t\tfor (i = 0; i < nparent; i++)\n@@ -518,6 +521,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \n \t\t\t/* t↓ */\n \t\t\tupdate_tree_entry(&t);\n+\t\t\topt->num_changes++;\n \t\t}\n \n \t\t/* t > p[imin] */\n@@ -535,6 +539,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\tskip_emit_tp:\n \t\t\t/* ∀ pi=p[imin]  pi↓ */\n \t\t\tupdate_tp_entries(tp, nparent);\n+\t\t\topt->num_changes++;\n \t\t}\n \t}\n \n-- \n2.25.0.rc0\n\n"},{"id":"389024","messageId":"20191229060308.GA220034@coredump.intra.peff.net","threadId":"52499","inReplyTo":"fc30441a-1bb7-73e5-43f6-6e26824e04f6@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-12-29T06:03:08Z","receivedAt":"2019-12-29T06:03:12Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 26, 2019 at 09:21:36AM -0500, Derrick Stolee wrote:\n\n> Here are some recommendations (to Garima) for how to proceed with these\n> patches. Please let me know if anyone disagrees.\n> \n> >   [1/3]: commit-graph: examine changed-path objects in pack order\n> \n> This one is best kept as its own patch, as it shows a clear reason why\n> we want to do the sort-by-position. It would also be a complicated\n> patch to include this logic along with the first use of\n> compute_bloom_filters().\n\nYeah, I'd agree this one could be a separate patch. It does need more\nwork, though (as you found out, it does not cover --reachable at all).\n\nThe position counter also probably ought to be an unsigned (or even a\nuint32_t, which we usually consider a maximum bound for number of\nobjects).\n\n-Peff\n"},{"id":"389025","messageId":"20191229061246.GB220034@coredump.intra.peff.net","threadId":"52499","inReplyTo":"8b331ef6-f431-56ef-37a9-1d6e263ea0fe@gmail.com","subject":"Re: [PATCH 1/3] commit-graph: examine changed-path objects in pack order","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-12-29T06:12:46Z","receivedAt":"2019-12-29T06:12:48Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Dec 27, 2019 at 09:51:02AM -0500, Derrick Stolee wrote:\n\n> On 12/22/2019 4:32 AM, Jeff King wrote:\n> > Looking at the diff of commit objects in pack order is much faster than\n> > in sha1 order, as it gives locality to the access of tree deltas\n> > (whereas sha1 order is effectively random). Unfortunately the\n> > commit-graph code sorts the commits (several times, sometimes as an oid\n> > and sometimes a pointer-to-commit), and we ultimately traverse in sha1\n> > order.\n> > \n> > Instead, let's remember the position at which we see each commit, and\n> > traverse in that order when looking at bloom filters. This drops my time\n> > for \"git commit-graph write --changed-paths\" in linux.git from ~4\n> > minutes to ~1.5 minutes.\n> \n> I'm doing my own perf tests on these patches, and my copy of linux.git\n> has four packs of varying sizes (corresponding with my rare fetches and\n> lack of repacks). My time goes from 3m50s to 3m00s. I was confused at\n> first, but then realized that I used the \"--reachable\" flag. In that\n> case, we never run set_commit_pos(), so all positions are equal and the\n> sort is not helpful.\n> \n> I thought that inserting some set_commit_pos() calls into close_reachable()\n> and add_missing_parents() would give some amount of time-order to the\n> commits as we compute the filters. However, the time did not change at\n> all.\n> \n> I've included the patch below for reference, anyway.\n\nYeah, I expected that would cover it, too. But instrumenting it to dump\nthe position of each commit (see patch below), and then decorating \"git\nlog\" output with the positions (see script below) shows that we're all\nover the map:\n\n  *   3\n  |\\  \n  | * 2791\n  | * 5476\n  | * 8520\n  | * 12040\n  | * 16036\n  * |   2790\n  |\\ \\  \n  | * | 5475\n  | * | 8519\n  | * | 12039\n  | * | 16035\n  | * | 20517\n  | * | 25527\n  | |/  \n  * |   5474\n  |\\ \\  \n  | * | 8518\n  | * | 12038\n  * | |   8517\n  [...]\n\nI think the root issue is that we never do any date-sorting on the\ncommits. So:\n\n  - we hit each ref tip in lexical order; with tags, this is quite often\n    the opposite of reverse-chronological\n\n  - we traverse breadth-first, but we don't order queue at all. So if we\n    see a merge X, then we'll next process X^1 and X^2, and then X^1^,\n    and then X^2^, and so forth. So we keep digging equally down\n    simultaneous branches, even if one branch is way shorter than the\n    other. Whereas a regular Git traversal will order the queue by\n    commit timestamp, so it tends to be roughly chronological (of course\n    a topo-sort would work too, but that's probably overkill).\n\nI wonder if this would be simpler if \"commit-graph --reachable\" just\nused the regular revision machinery instead of doing its own custom\ntraversal.\n\n-Peff\n"},{"id":"389026","messageId":"20191229062825.GA222211@coredump.intra.peff.net","threadId":"52499","inReplyTo":"20191229061246.GB220034@coredump.intra.peff.net","subject":"Re: [PATCH 1/3] commit-graph: examine changed-path objects in pack order","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-12-29T06:28:25Z","receivedAt":"2019-12-29T06:33:05Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Dec 29, 2019 at 01:12:46AM -0500, Jeff King wrote:\n\n> Yeah, I expected that would cover it, too. But instrumenting it to dump\n> the position of each commit (see patch below), and then decorating \"git\n> log\" output with the positions (see script below) shows that we're all\n> over the map:\n\nI forgot the patch, of course. :)\n\nI just dumped this trace:\n\n---\ndiff --git a/commit-graph.c b/commit-graph.c\nindex a6c4ab401e..1cb77be45f 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -61,6 +61,7 @@ static void set_commit_pos(struct repository *r, const struct object_id *oid)\n \tif (!commit)\n \t\treturn; /* should never happen, but be lenient */\n \n+\ttrace_printf(\"pos %s = %d\", oid_to_hex(oid), max_pos);\n \t*commit_pos_at(&commit_pos, commit) = max_pos++;\n }\n \n\nwith:\n\n  rm -f .git/objects/info/commit-graph\n  GIT_TRACE=$PWD/trace.out git commit-graph write --changed-paths --reachable\n\nand then used:\n\n  cat >foo.pl <<\\EOF\n  #!/usr/bin/perl\n  \n  my %deco = do {\n  \topen(my $fh, '<', 'trace.out');\n  \tmap { /pos (\\S+) = (\\d+)/ ? ($1 => $2) : () } <$fh>\n  };\n  while (<>) {\n  \ts/([0-9a-f]{40})/$deco{$1}/;\n  \tprint;\n  }\n  EOF\n\nlike so:\n\n  git log --graph --format=%H |\n  perl foo.pl |\n  less\n\n-Peff\n"},{"id":"389027","messageId":"20191229062414.GC220034@coredump.intra.peff.net","threadId":"52499","inReplyTo":"e9a4e4ff-5466-dc39-c3f5-c9a8b8f2f11d@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-12-29T06:24:14Z","receivedAt":"2019-12-29T06:33:05Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Dec 27, 2019 at 11:11:37AM -0500, Derrick Stolee wrote:\n\n> > So here are a few patches to reduce the CPU and memory usage. They could\n> > be squashed in at the appropriate spots, or perhaps taken as inspiration\n> > if there are better solutions (especially for the first one).\n> \n> I tested these patches with the Linux kernel repo and reported my results\n> on each patch. However, I wanted to also test on a larger internal repo\n> (the AzureDevOps repo), which has ~500 commits with more than 512 changes,\n> and generally has larger diffs than the Linux kernel repo.\n> \n> | Version  | Time   | Memory |\n> |----------|--------|--------|\n> | Garima   | 16m36s | 27.0gb |\n> | Peff 1   | 6m32s  | 28.0gb |\n> | Peff 2   | 6m48s  |  5.6gb |\n> | Peff 3   | 6m14s  |  4.5gb |\n> | Shortcut | 3m47s  |  4.5gb |\n> \n> For reference, I found the time and memory information using\n> \"/usr/bin/time --verbose\" in a bash script.\n\nThanks for giving it more exercise. My heap profiling was done with\nmassif, which measures the heap directly. Measuring RSS would cover\nthat, but will also include the mmap'd packfiles. That's probably why\nyour linux.git numbers were slightly higher than mine.\n\n(massif is a really great tool if you haven't used it, as it also shows\nwhich allocations were using the memory. But it's part of valgrind, so\nit definitely doesn't run on native Windows. It might work under WSL,\nthough. I'm sure there are also other heap profilers on Windows).\n\n> By \"Shortcut\" in the table above, I mean the following patch on top of\n> Garima's and Peff's changes. It inserts a max_changes option into struct\n> diff_options to halt the diff early. This seemed like an easier change\n> than creating a new tree diff algorithm wholesale.\n\nYeah, I'm not opposed to a diff feature like this.\n\nBut be careful, because...\n\n> diff --git a/diff.h b/diff.h\n> index 6febe7e365..9443dc1b00 100644\n> --- a/diff.h\n> +++ b/diff.h\n> @@ -285,6 +285,11 @@ struct diff_options {\n>  \t/* Number of hexdigits to abbreviate raw format output to. */\n>  \tint abbrev;\n>  \n> +\t/* If non-zero, then stop computing after this many changes. */\n> +\tint max_changes;\n> +\t/* For internal use only. */\n> +\tint num_changes;\n\nThis is holding internal state in diff_options, but the same\ndiff_options is often used for multiple diffs (e.g., \"git log --raw\"\nwould use the same rev_info.diffopt over and over again).\n\nSo it would need to be cleared between diffs. There's a similar feature\nin the \"has_changes\" flag, though it looks like it is cleared manually\nby callers. Yuck.\n\nThis isn't a problem for commit-graph right now, but:\n\n  - it actually could be using a single diff_options, which would be\n    slightly simpler (it doesn't seem to save much CPU, though, because\n    the initialization is relatively cheap)\n\n  - it's a bit of a subtle bug to leave hanging around for the next\n    person who tries to use the feature\n\nI actually wonder if this could be rolled into the has_changes and\ndiff_can_quit_early() feature. This really just a generalization of that\nfeature (which is like setting max_changes to \"1\").\n\n-Peff\n"},{"id":"389059","messageId":"b9bd0c2e-63fc-5658-7a24-b8ab078acd44@gmail.com","threadId":"52499","inReplyTo":"20191229061246.GB220034@coredump.intra.peff.net","subject":"Re: [PATCH 1/3] commit-graph: examine changed-path objects in pack order","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-12-30T14:37:28Z","receivedAt":"2019-12-30T14:37:32Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/29/2019 1:12 AM, Jeff King wrote:\n> On Fri, Dec 27, 2019 at 09:51:02AM -0500, Derrick Stolee wrote:\n> \n>> On 12/22/2019 4:32 AM, Jeff King wrote:\n>>> Looking at the diff of commit objects in pack order is much faster than\n>>> in sha1 order, as it gives locality to the access of tree deltas\n>>> (whereas sha1 order is effectively random). Unfortunately the\n>>> commit-graph code sorts the commits (several times, sometimes as an oid\n>>> and sometimes a pointer-to-commit), and we ultimately traverse in sha1\n>>> order.\n>>>\n>>> Instead, let's remember the position at which we see each commit, and\n>>> traverse in that order when looking at bloom filters. This drops my time\n>>> for \"git commit-graph write --changed-paths\" in linux.git from ~4\n>>> minutes to ~1.5 minutes.\n>>\n>> I'm doing my own perf tests on these patches, and my copy of linux.git\n>> has four packs of varying sizes (corresponding with my rare fetches and\n>> lack of repacks). My time goes from 3m50s to 3m00s. I was confused at\n>> first, but then realized that I used the \"--reachable\" flag. In that\n>> case, we never run set_commit_pos(), so all positions are equal and the\n>> sort is not helpful.\n>>\n>> I thought that inserting some set_commit_pos() calls into close_reachable()\n>> and add_missing_parents() would give some amount of time-order to the\n>> commits as we compute the filters. However, the time did not change at\n>> all.\n>>\n>> I've included the patch below for reference, anyway.\n> \n> Yeah, I expected that would cover it, too. But instrumenting it to dump\n> the position of each commit (see patch below), and then decorating \"git\n> log\" output with the positions (see script below) shows that we're all\n> over the map:\n> \n>   *   3\n>   |\\  \n>   | * 2791\n>   | * 5476\n>   | * 8520\n>   | * 12040\n>   | * 16036\n>   * |   2790\n>   |\\ \\  \n>   | * | 5475\n>   | * | 8519\n>   | * | 12039\n>   | * | 16035\n>   | * | 20517\n>   | * | 25527\n>   | |/  \n>   * |   5474\n>   |\\ \\  \n>   | * | 8518\n>   | * | 12038\n>   * | |   8517\n>   [...]\n\nThis makes a lot of sense why the previous approach did not work. Thanks!\n\n> I think the root issue is that we never do any date-sorting on the\n> commits. So:\n> \n>   - we hit each ref tip in lexical order; with tags, this is quite often\n>     the opposite of reverse-chronological\n> \n>   - we traverse breadth-first, but we don't order queue at all. So if we\n>     see a merge X, then we'll next process X^1 and X^2, and then X^1^,\n>     and then X^2^, and so forth. So we keep digging equally down\n>     simultaneous branches, even if one branch is way shorter than the\n>     other. Whereas a regular Git traversal will order the queue by\n>     commit timestamp, so it tends to be roughly chronological (of course\n>     a topo-sort would work too, but that's probably overkill).\n> \n> I wonder if this would be simpler if \"commit-graph --reachable\" just\n> used the regular revision machinery instead of doing its own custom\n> traversal.\n\nInstead, why not use our already-computed generation numbers? That seems\nto improve the time a bit. (6m30s to 4m50s)\n\n-->8--\n\nFrom: Derrick Stolee <dstolee@microsoft.com>\nDate: Fri, 27 Dec 2019 09:47:49 -0500\nSubject: [PATCH] commit-graph: examine commits by generation number\n\nWhen running 'git commit-graph write --changed-paths', we sort the\ncommits by pack-order to save time when computing the changed-paths\nbloom filters. This does not help when finding the commits via the\n--reachable flag.\n\nIf not using pack-order, then sort by generation number before\nexamining the diff. Commits with similar generation are more likely\nto have many trees in common, making the diff faster.\n\nOn the Linux kernel repository, this change reduced the computation\ntime for 'git commit-graph write --reachable --changed-paths' from\n6m30s to 4m50s.\n\nHelped-by: Jeff King <peff@peff.net>\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\n---\n commit-graph.c | 33 ++++++++++++++++++++++++++++++---\n 1 file changed, 30 insertions(+), 3 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex bf6c663772..fe4ab545f2 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -72,6 +72,25 @@ static int commit_pos_cmp(const void *va, const void *vb)\n \t       commit_pos_at(&commit_pos, b);\n }\n \n+static int commit_gen_cmp(const void *va, const void *vb)\n+{\n+\tconst struct commit *a = *(const struct commit **)va;\n+\tconst struct commit *b = *(const struct commit **)vb;\n+\n+\t/* lower generation commits first */\n+\tif (a->generation < b->generation)\n+\t\treturn -1;\n+\telse if (a->generation > b->generation)\n+\t\treturn 1;\n+\n+\t/* use date as a heuristic when generations are equal */\n+\tif (a->date < b->date)\n+\t\treturn -1;\n+\telse if (a->date > b->date)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n char *get_commit_graph_filename(const char *obj_dir)\n {\n \tchar *filename = xstrfmt(\"%s/info/commit-graph\", obj_dir);\n@@ -849,7 +868,8 @@ struct write_commit_graph_context {\n \t\t report_progress:1,\n \t\t split:1,\n \t\t check_oids:1,\n-\t\t bloom:1;\n+\t\t bloom:1,\n+\t\t order_by_pack:1;\n \n \tconst struct split_commit_graph_opts *split_opts;\n \tuint32_t total_bloom_filter_size;\n@@ -1245,7 +1265,11 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \n \tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n \tCOPY_ARRAY(sorted_by_pos, ctx->commits.list, ctx->commits.nr);\n-\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n+\n+\tif (ctx->order_by_pack)\n+\t\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n+\telse\n+\t\tQSORT(sorted_by_pos, ctx->commits.nr, commit_gen_cmp);\n \n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tstruct commit *c = sorted_by_pos[i];\n@@ -1979,6 +2003,7 @@ int write_commit_graph(const char *obj_dir,\n \t}\n \n \tif (pack_indexes) {\n+\t\tctx->order_by_pack = 1;\n \t\tif ((res = fill_oids_from_packs(ctx, pack_indexes)))\n \t\t\tgoto cleanup;\n \t}\n@@ -1988,8 +2013,10 @@ int write_commit_graph(const char *obj_dir,\n \t\t\tgoto cleanup;\n \t}\n \n-\tif (!pack_indexes && !commit_hex)\n+\tif (!pack_indexes && !commit_hex) {\n+\t\tctx->order_by_pack = 1;\n \t\tfill_oids_from_all_packs(ctx);\n+\t}\n \n \tclose_reachable(ctx);\n \n-- \n2.25.0.rc0\n\n\n\n"},{"id":"389061","messageId":"f0579d4c-b44e-e17f-f395-ae8970765f20@gmail.com","threadId":"52499","inReplyTo":"b9bd0c2e-63fc-5658-7a24-b8ab078acd44@gmail.com","subject":"Re: [PATCH 1/3] commit-graph: examine changed-path objects in pack order","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-12-30T14:51:27Z","receivedAt":"2019-12-30T14:51:30Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/30/2019 9:37 AM, Derrick Stolee wrote:\n> On the Linux kernel repository, this change reduced the computation\n> time for 'git commit-graph write --reachable --changed-paths' from\n> 6m30s to 4m50s.\n\nI apologize, these numbers are based on the AzureDevOps repo, not the\nLinux kernel repo. After re-running with the Linux kernel repo my\ntimes improve from 3m00s to 1m37s.\n\n-Stolee\n\n\n"},{"id":"389070","messageId":"6d20f568-9681-3e66-b892-8f076e20dc63@gmail.com","threadId":"52499","inReplyTo":"20191229062414.GC220034@coredump.intra.peff.net","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-12-30T16:04:53Z","receivedAt":"2019-12-30T16:05:06Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 12/29/2019 1:24 AM, Jeff King wrote:\n> On Fri, Dec 27, 2019 at 11:11:37AM -0500, Derrick Stolee wrote:\n> \n>>> So here are a few patches to reduce the CPU and memory usage. They could\n>>> be squashed in at the appropriate spots, or perhaps taken as inspiration\n>>> if there are better solutions (especially for the first one).\n>>\n>> I tested these patches with the Linux kernel repo and reported my results\n>> on each patch. However, I wanted to also test on a larger internal repo\n>> (the AzureDevOps repo), which has ~500 commits with more than 512 changes,\n>> and generally has larger diffs than the Linux kernel repo.\n>>\n>> | Version  | Time   | Memory |\n>> |----------|--------|--------|\n>> | Garima   | 16m36s | 27.0gb |\n>> | Peff 1   | 6m32s  | 28.0gb |\n>> | Peff 2   | 6m48s  |  5.6gb |\n>> | Peff 3   | 6m14s  |  4.5gb |\n>> | Shortcut | 3m47s  |  4.5gb |\n>>\n>> For reference, I found the time and memory information using\n>> \"/usr/bin/time --verbose\" in a bash script.\n> \n> Thanks for giving it more exercise. My heap profiling was done with\n> massif, which measures the heap directly. Measuring RSS would cover\n> that, but will also include the mmap'd packfiles. That's probably why\n> your linux.git numbers were slightly higher than mine.\n\nThat's interesting. I initially avoided massif because it is so much\nslower than /usr/bin/time. However, the inflated numbers could be\nexplained by that. Also, the distinction between mem_heap and\nmem_heap_extra may have interesting implications. Looking online, it\nseems that large mem_heap_extra implies the heap is fragmented from\nmany small allocations.\n\nHere are my findings on the Linux repo:\n\n| Version  | mem_heap | mem_heap_extra |\n|----------|----------|----------------|\n| Peff 1   |  6,500mb |          913mb |\n| Peff 2   |  3,100mb |          286mb |\n| Peff 3   |    781mb |          235mb |\n\nThese numbers more closely match your numbers (in sum of the two\ncolumns).\n\n> (massif is a really great tool if you haven't used it, as it also shows\n> which allocations were using the memory. But it's part of valgrind, so\n> it definitely doesn't run on native Windows. It might work under WSL,\n> though. I'm sure there are also other heap profilers on Windows).\n\nI am using my Linux machine for my tests. Garima is using her Windows\nmachine.\n\n>> By \"Shortcut\" in the table above, I mean the following patch on top of\n>> Garima's and Peff's changes. It inserts a max_changes option into struct\n>> diff_options to halt the diff early. This seemed like an easier change\n>> than creating a new tree diff algorithm wholesale.\n> \n> Yeah, I'm not opposed to a diff feature like this.\n> \n> But be careful, because...\n> \n>> diff --git a/diff.h b/diff.h\n>> index 6febe7e365..9443dc1b00 100644\n>> --- a/diff.h\n>> +++ b/diff.h\n>> @@ -285,6 +285,11 @@ struct diff_options {\n>>  \t/* Number of hexdigits to abbreviate raw format output to. */\n>>  \tint abbrev;\n>>  \n>> +\t/* If non-zero, then stop computing after this many changes. */\n>> +\tint max_changes;\n>> +\t/* For internal use only. */\n>> +\tint num_changes;\n> \n> This is holding internal state in diff_options, but the same\n> diff_options is often used for multiple diffs (e.g., \"git log --raw\"\n> would use the same rev_info.diffopt over and over again).\n> \n> So it would need to be cleared between diffs. There's a similar feature\n> in the \"has_changes\" flag, though it looks like it is cleared manually\n> by callers. Yuck.\n\nYou're right about this. What if we initialize it in diff_tree_paths()\nbefore it calls the recursive ll_difF_tree_paths()?\n\n> This isn't a problem for commit-graph right now, but:\n> \n>   - it actually could be using a single diff_options, which would be\n>     slightly simpler (it doesn't seem to save much CPU, though, because\n>     the initialization is relatively cheap)\n> \n>   - it's a bit of a subtle bug to leave hanging around for the next\n>     person who tries to use the feature\n> \n> I actually wonder if this could be rolled into the has_changes and\n> diff_can_quit_early() feature. This really just a generalization of that\n> feature (which is like setting max_changes to \"1\").\n\nI thought about this at first, but it only takes a struct diff_options\nright now. It does have an internally-mutated member (flags.has_changes)\nbut it also seems a bit wrong to add a uint32_t of the count in this.\nChanging the prototype could be messy, too.\n\nThere are also multiple callers, and limiting everything to tree-diff.c\nlimits the impact.\n\nThanks,\n-Stolee\n"},{"id":"389076","messageId":"xmqqsgl1hhlw.fsf@gitster-ct.c.googlers.com","threadId":"52499","inReplyTo":"20191229062414.GC220034@coredump.intra.peff.net","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-30T17:02:19Z","receivedAt":"2019-12-30T17:02:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> This is holding internal state in diff_options, but the same\n> diff_options is often used for multiple diffs (e.g., \"git log --raw\"\n> would use the same rev_info.diffopt over and over again).\n>\n> So it would need to be cleared between diffs. There's a similar feature\n> in the \"has_changes\" flag, though it looks like it is cleared manually\n> by callers. Yuck.\n\nDo you mean we want reset_per_invocation_part_of_diff_options()\nhelper or something?\n\n> I actually wonder if this could be rolled into the has_changes and\n> diff_can_quit_early() feature. This really just a generalization of that\n> feature (which is like setting max_changes to \"1\").\n\nYeah, I wondered about the same thing, after seeing the impressive\nnumbers ;-)\n"},{"id":"389136","messageId":"86d0c44f5s.fsf@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2019-12-31T16:45:51Z","receivedAt":"2019-12-31T16:45:59Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Hey! \n>\n> The commit graph feature brought in a lot of performance improvements across\n> multiple commands. However, file based history continues to be a performance\n> pain point, especially in large repositories. \n>\n> Adopting changed path bloom filters has been discussed on the list before,\n> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n> Derrick Stolee [1]. This series is based on Dr. Stolee's approach [2] and\n> presents an updated and more polished RFC version of the feature. \n\nIt is nice to have this picked up for upstream, finally.  The proof of\nconcept works[1][2] were started more than a year ago.\n\nOn the other hand slow and steady adoption of commit-graph serialization\nand then extending it (generation numbers, topological sort, incremental\nupdate) feels like a good approach.\n\n> Performance Gains: We tested the performance of 'git log -- <path>' on the git\n> repo, the linux repo and some internal large repos, with a variety of paths\n> of varying depths.\n>\n> On the git and linux repos: We observed a 2x to 5x speed up.\n>\n> On a large internal repo with files seated 6-10 levels deep in the tree: We\n> observed 10x to 20x speed ups, with some paths going up to 28 times faster.\n\nCould you provide some more statistics about this internal repository,\nsuch as number of files, number of commits, perhaps also number of all\nobjects?  Thanks in advance.\n\nI wonder why such large difference in performance 2-5x vs 10-20x.  Is it\nabout the depth of the file hierarchy?  How would the numbers look for\nfiles seated closer to the root in the same large repository, like 3-5\nlevels deep in the tree?\n\n> Future Work (not included in the scope of this series):\n>\n>  1. Supporting multiple path based revision walk\n\nI wonder if it would ever be possible to support globbing, e.g. '*.c'\n\n>  2. Adopting it in git blame logic.\n\nWhat about 'git log --follow <path>'?\n\n>  3. Interactions with line log git log -L\n>\n> This series is intended to start the conversation and many of the commit\n> messages include specific call outs for suggestions and thoughts. \n>\n> Cheers! Garima Singh\n>\n> [1] https://lore.kernel.org/git/20181009193445.21908-1-szeder.dev@gmail.com/\n> [2] https://lore.kernel.org/git/61559c5b-546e-d61b-d2e1-68de692f5972@gmail.com/\n>\n> Garima Singh (9):\n>   commit-graph: add --changed-paths option to write\n\nThis summary is not easy to understand on first glance.  Maybe:\n\n    commit-graph: add --changed-paths option to the write subcommand\n\nor\n\n    commit-graph: add --changed-paths option to 'git commit-graph write'\n\nwould be better?\n\n>   commit-graph: write changed paths bloom filters\n>   commit-graph: use MAX_NUM_CHUNKS\n>   commit-graph: document bloom filter format\n>   commit-graph: write changed path bloom filters to commit-graph file.\n>   commit-graph: test commit-graph write --changed-paths\n>   commit-graph: reuse existing bloom filters during write.\n>   revision.c: use bloom filters to speed up path based revision walks\n>   commit-graph: add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag\n>\n>  Documentation/git-commit-graph.txt            |   5 +\n>  .../technical/commit-graph-format.txt         |  17 ++\n>  Makefile                                      |   1 +\n>  bloom.c                                       | 257 +++++++++++++++++\n>  bloom.h                                       |  51 ++++\n>  builtin/commit-graph.c                        |   9 +-\n>  ci/run-build-and-tests.sh                     |   1 +\n>  commit-graph.c                                | 116 +++++++-\n>  commit-graph.h                                |   9 +-\n>  revision.c                                    |  67 ++++-\n>  revision.h                                    |   5 +\n>  t/README                                      |   3 +\n>  t/helper/test-read-graph.c                    |   4 +\n>  t/t4216-log-bloom.sh                          |  77 ++++++\n>  t/t5318-commit-graph.sh                       |   2 +\n>  t/t5324-split-commit-graph.sh                 |   1 +\n>  t/t5325-commit-graph-bloom.sh                 | 258 ++++++++++++++++++\n>  17 files changed, 875 insertions(+), 8 deletions(-)\n>  create mode 100644 bloom.c\n>  create mode 100644 bloom.h\n>  create mode 100755 t/t4216-log-bloom.sh\n>  create mode 100755 t/t5325-commit-graph-bloom.sh\n>\n>\n> base-commit: b02fd2accad4d48078671adf38fe5b5976d77304\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-497%2Fgarimasi514%2FcoreGit-bloomFilters-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-497/garimasi514/coreGit-bloomFilters-v1\n> Pull-Request: https://github.com/gitgitgadget/git/pull/497\n"},{"id":"389150","messageId":"865zhv4c2w.fsf@gmail.com","threadId":"52499","inReplyTo":"20191222093857.GB3449072@coredump.intra.peff.net","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-01T12:04:39Z","receivedAt":"2020-01-01T12:04:49Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Sun, Dec 22, 2019 at 10:26:20AM +0100, Christian Couder wrote:\n>\n>> I have a question though. Are the performance gains only available\n>> with `git log -- path` or are they already available for example when\n>> doing a partial clone and/or a sparse checkout?\n>\n> From my quick look at the code, anything that feeds a pathspec to a\n> revision traversal would be helped. I'm not sure if it would help for\n> partial/sparse traversals, though. There we actually need to know which\n> blobs correspond to the paths in question, not just whether any\n> particular commit touched them.\n>\n> I also took a brief look at adding support to the custom blame-tree\n> implementation we use at GitHub, and got about a 6x speedup.\n\nIs there any chance of upstreaming the blame-tree algorithm, perhaps as\na separate mode for git-blame (invoked with `git blame <directory>`?\nOr is the algorithm too GitHub-specific?\n\nBest,\n-- \nJakub Narębski\n"},{"id":"389152","messageId":"86tv5f2ak7.fsf@gmail.com","threadId":"52499","inReplyTo":"6bdde5e4f0ceb54546978e3e9cdde00045d45468.1576879520.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/9] commit-graph: add --changed-paths option to write","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-01T20:20:24Z","receivedAt":"2020-01-01T20:22:47Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Add --changed-paths option to git commit-graph write. This option will\n> soon allow users to compute bloom filters for the paths changed between\n> a commit and its first significant parent, and write this information\n> into the commit-graph file.\n\nA slightly nitpicky comment.\n\nFirst, I think it is \"Bloom filter\", not \"bloom filter\" (from the name\nof the person that discovered them, Burton Howard Bloom).\n\nSecond, I would rather that the commit message started with at least one\nsentence of describing purpose of this new option, not going straight to\nthe technical details (i.e. using Bloom filters).  Or in any other way\ndescribe that this option would make Git store some helper data that\nwould help find out faster if a given path was changed in given commit.\n\n> Note: This commit does not change any behavior. It only introduces\n> the option and passes down the appropriate flag to the commit-graph.\n\nAll right.\n\nPersonally, I don't have strong opinion for or against separating this\nchange into its own patch.\n\n> RFC Notes:\n> 1. We named the option --changed-paths to capture what the option does,\n>    instead of how it does it. The current implementation does this\n>    using bloom filters. We believe using --changed-paths however keeps\n>    the implementation open to other data structures.\n>    All thoughts and suggestions for the name and this approach are\n>    welcome\n\nIt is all right name.  Another option could be for example\n`git commit-graph write --changeset-info`, or something like that.\n\n>\n> 2. Currently, a subsequent commit in this series will add tests that\n>    exercise this option. I plan to split that test commit across the\n>    series as appropriate.\n\nThere is another thing, but one that could be left for the followup\nseries, namely the configuration variables for this behavior.  In the\nfuture it should be possible to switch some configuration variable to\nhave this feature on by default when manually or automatically running\n`git commit-graph write`.\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  Documentation/git-commit-graph.txt | 5 +++++\n>  builtin/commit-graph.c             | 9 +++++++--\n>  commit-graph.h                     | 3 ++-\n>  3 files changed, 14 insertions(+), 3 deletions(-)\n>\n> diff --git a/Documentation/git-commit-graph.txt b/Documentation/git-commit-graph.txt\n> index bcd85c1976..1efe6e5c5a 100644\n> --- a/Documentation/git-commit-graph.txt\n> +++ b/Documentation/git-commit-graph.txt\n\nIt is nice to have option documented.\n\nAll right, the 'write' subcommand has the following synopsis:\n\n  'git commit-graph write' <options> [--object-dir <dir>] [--[no-]progress]\n\nso the is no need to adjust it when adding a new option.\n\n> @@ -54,6 +54,11 @@ or `--stdin-packs`.)\n>  With the `--append` option, include all commits that are present in the\n>  existing commit-graph file.\n>  +\n> +With the `--changed-paths` option, compute and write information about the\n> +paths changed between a commit and it's first parent. This operation can\n> +take a while on large repositories. It provides significant performance gains\n> +for getting file based history logs with `git log`\n       ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n\nThis might be not entirely clear for someone that is not familiar with\nGit jargon.  Perhaps it would better read as \"for getting history of a\ndirectory or a file with `git log <path>`\", or something like that.\n\nSide note: the sentence is missing its finishing full stop.\n\n> ++\n>  With the `--split` option, write the commit-graph as a chain of multiple\n>  commit-graph files stored in `<dir>/info/commit-graphs`. The new commits\n>  not already in the commit-graph are added in a new \"tip\" file. This file\n> diff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\n> index e0c6fc4bbf..9bd1e11161 100644\n> --- a/builtin/commit-graph.c\n> +++ b/builtin/commit-graph.c\n> @@ -9,7 +9,7 @@\n>  \n>  static char const * const builtin_commit_graph_usage[] = {\n>  \tN_(\"git commit-graph verify [--object-dir <objdir>] [--shallow] [--[no-]progress]\"),\n> -\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n> +\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n>  \tNULL\n>  };\n>  \n> @@ -19,7 +19,7 @@ static const char * const builtin_commit_graph_verify_usage[] = {\n>  };\n>  \n>  static const char * const builtin_commit_graph_write_usage[] = {\n> -\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n> +\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n>  \tNULL\n>  };\n\nI was at first wondering why the duplication (not caused by your patch,\nthough), and then realized that it is to have usage for command, and for\nindividual subcommands, separately.\n\n> @@ -32,6 +32,7 @@ static struct opts_commit_graph {\n>  \tint split;\n>  \tint shallow;\n>  \tint progress;\n> +\tint enable_bloom_filters;\n\nWhy the field is called `enable_bloom_filters`, while option is called\n`--changed-paths`?  I know it is not user-visible thing, so it would be\neasy to change if we ever go beyond Bloom filters, though...\n\nSo I am not against keeping it as it is currently.\n\n>  } opts;\n>  \n>  static int graph_verify(int argc, const char **argv)\n> @@ -110,6 +111,8 @@ static int graph_write(int argc, const char **argv)\n>  \t\t\tN_(\"start walk at commits listed by stdin\")),\n>  \t\tOPT_BOOL(0, \"append\", &opts.append,\n>  \t\t\tN_(\"include all commits already in the commit-graph file\")),\n> +\t\tOPT_BOOL(0, \"changed-paths\", &opts.enable_bloom_filters,\n> +\t\t\tN_(\"enable computation for changed paths\")),\n>  \t\tOPT_BOOL(0, \"progress\", &opts.progress, N_(\"force progress reporting\")),\n>  \t\tOPT_BOOL(0, \"split\", &opts.split,\n>  \t\t\tN_(\"allow writing an incremental commit-graph file\")),\n> @@ -143,6 +146,8 @@ static int graph_write(int argc, const char **argv)\n>  \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n>  \tif (opts.progress)\n>  \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n> +\tif (opts.enable_bloom_filters)\n> +\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n\nMinor nitpick: are we all right having this ordering of options (not for\nexample having opt.progress last)?\n\nDisregarding this, it looks all right.\n\n>  \n>  \tread_replace_refs = 0;\n>  \n> diff --git a/commit-graph.h b/commit-graph.h\n> index 7f5c933fa2..952a4b83be 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -76,7 +76,8 @@ enum commit_graph_write_flags {\n>  \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\n>  \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n>  \t/* Make sure that each OID in the input is a valid commit OID. */\n> -\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3)\n> +\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n> +\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n\nI wonder if we should add comment describing the flag, like for the one\nabove...\n\n>  };\n>  \n>  struct split_commit_graph_opts {\n\nBest,\n-- \nJakub Narębski\n"},{"id":"389275","messageId":"86eewczapt.fsf@gmail.com","threadId":"52499","inReplyTo":"e52c7ad37a306891487bd79a09b040bfb657d723.1576879520.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 2/9] commit-graph: write changed paths bloom filters","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-06T18:44:14Z","receivedAt":"2020-01-06T18:44:30Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> The changed path bloom filters help determine which paths changed between a\n> commit and its first parent. We already have the \"--changed-paths\" option\n> for the \"git commit-graph write\" subcommand, now actually compute them under\n> that option. The COMMIT_GRAPH_WRITE_BLOOM_FILTERS flag enables this\n> computation.\n>\n> RFC Notes: Here are some details about the implementation and I would love\n> to know your thoughts and suggestions for improvements here.\n>\n> For details on what bloom filters are and how they work, please refer to\n> Dr. Derrick Stolee's blog post [1].\n> [1] https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-bloom-filters/\n>\n> 1. The implementation sticks to the recommended values of 7 and 10 for the\n>    number of hashes and the size of each entry, as described in the blog.\n\nPlease provide references to original work for this.  Derrick Stolee\nblog post references the following work:\n\n  Flavio Bonomi, Michael Mitzenmacher, Rina Panigrahy, Sushil Singh, George Varghese\n  \"An Improved Construction for Counting Bloom Filters\"\n  http://theory.stanford.edu/~rinap/papers/esa2006b.pdf\n  https://doi.org/10.1007/11841036_61\n\nHowever, we do not use Counting Bloom Filters, but ordinary Bloom\nFilters, if used in untypical way: instead of testing many elements\n(keys) against single filter, we test single element (path) against\nmainy filters.\n\nAlso, I'm not sure that values 10 bits per entry and 7 hash functions\nare recommended; the work states:\n\n  \"For example, when n/m = 10 and k = 7 the false positive probability\n  is just over 0.008.\"\n\nGiven false positive probablity we can calculate best choice for n/m and\nk.\n\nOn the other hand in https://arxiv.org/abs/1912.08258 we have\n\n  \"For efficient memory usage, a Bloom filter with a false-positive\n   probability ϵ should use about − log_2(ϵ) hash functions\n   [Broder2004]. At a false-positive probability of 1%, seven hash\n   functions are thus required.\"\n\nSo k=7 being optimal is somewhat confirmed.\n\n>    The implementation while not completely open to it at the moment, is flexible\n>    enough to allow for tweaking these settings in the future.\n\nAll right.\n\n>    Note: The performance gains we have observed so far with these values is\n>    significant enough to not that we did not need to tweak these settings.\n                           ^^^\ns/not/note/\n\nDid you try to tweak settings, i.e. numbers of bits per entry, number\nof hash functions (which is derivative of the former - at least the\noptimal number), the size of the block, the cutoff threshold value?\nIt is not needed to be in this patch series - fine tuning is probably\nbetter left for later.\n\n>    The cover letter of this series has the details and the commit where we have\n>    git log use bloom filters.\n\nThe second part of this sentence, from \"and the commit...\" is a bit\nunclear.  Did you mean here that the future / subsequent commit in this\npatch series that makes Git actually use Bloom filters in `git log --\n<path>` will have more details in its commit message?\n\n> 2. As described in the blog and the linked technical paper therin, we do not need\n                                                             ^^^^^^\ns/therin/therein/\n\n>    7 independent hashing functions. We use the Murmur3 hashing scheme - seed it\n>    twice and then combine those to procure an arbitrary number of hash values.\n\nThe \"linked technical paper\" in the blog post (which I would prefer to\nhave linked directly to in the commit message) is\n\n  Peter C. Dillinger and Panagiotis Manolios\n  \"Bloom Filters in Probabilistic Verification\"\n  http://www.ccs.neu.edu/home/pete/pub/bloom-filters-verification.pdf\n  https://doi.org/10.1007/978-3-540-30494-4_26\n\nSidenote: it looks like it is a reference from Wikipedia on Bloom filters.\nThis is according to authors the original paper with the _double hashing_\ntechnique.\n\nThey also examine in much detail the optimal number of hash functions.\n\n> 3. The filters are sized according to the number of changes in the each commit,\n>    with minimum size of one 64 bit word.\n\nDo I understand it correctly that the size of filter is 10*(number of\nchanged files) bits, rounded up to nearest multiple of 64?\n\nHow do you count renames and copies?  As two changes?\n\nDo I understand it correctly that commit with no changes in it (which\ncan rarely happen) would have 64-bits i.e. 8-bytes Bloom filter of all\nzeros: 0x0000000000000000?\n\nHow merges are handled?  Does the filter uses all changed files, or just\nchanges compared to first parent?\n\n>\n> [Call for advice] We currently cap writing bloom filters for commits with\n> atmost 512 changed files. In the current implementation, we compute the diff,\n> and then just throw it away once we see it has more than 512 changes.\n> Any suggestiongs on how to reduce the work we are doing in this case are more\n> than welcome.\n\nThis got solved in \"[PATCH] diff: halt tree-diff early after max_changes\"\nhttps://public-inbox.org/git/e9a4e4ff-5466-dc39-c3f5-c9a8b8f2f11d@gmail.com/\n\n> [Call for advice] Would the git community like this commit to be split up into\n> more granular commits? This commit could possibly be split out further with the\n> bloom.c code in its own commit, to be used by the commit-graph in a subsequent\n> commit. While I prefer it being contained in one commit this way, I am open to\n> suggestions.\n\nI think it might be a good idea to split this commit into purely Bloom\nfilter implementation (bloom.c) AND unit tests for Bloom filter itself\n(which would probably involve some new test-tool).\n\nI have not read further messages in the series [edit: they don't], so I\ndon't know if such tests already exist or not.  One could test for\nnegative match, maybe also (for specific choice of hash function) for\npositive and maybe even false positive match, for filter size depending\non the number of changes, for changes cap (maybe), maybe also for\nno-changes scenario.\n\n\nAs for splitting the main part of the series, I would envision it in the\nfollowing way (which is of course only one possibility):\n\n1. Implementation of generic-ish Bloom filter (with elements being\n   strings / paths, and optimized to test single key against many\n   filters, each taking small-ish space, variable size filter, limit on\n   maximum number of elements).\n\n   Technical documentation in comments in bloom.h (description of API)\n   and bloom.c (details of the algorithm, with references).\n\n   TODO: test-tool and unit tests.\n\n2. Using per-commit Bloom filter(s) to store changeset information\n   i.e. changed paths.  This would implement in-memory storage (on slab)\n   and creating Bloom filter out of commit and repository information.\n\n   Perhaps this should also get its own unit tests (that Bloom filter\n   catches changed files, and excluding false positivess catches\n   unchanged files).\n\n3. Storing per-commit Bloom filters in the commit-graph file:\n\n   a.) writing Bloom filters data to commit-graph file, which means\n       designing the chunk(s) format,\n   b.) verifying Bloom filter chunks, at least sanity-checks\n   c.) reading Bloom filters from commit-graph file into memory\n\n   Perhaps also some integration tests that the information is stored\n   and retrieved correctly, and that verifying finds bugs in\n   intentionally corrupted Bloom filter chunks.\n\n4. Using Bloom filters to speed up `git log -- <path>` (and similar\n   commands).\n\n   It would be nice to have some functional tests, and maybe some\n   performance tests, if possible.\n\n\n> [Call for advice] Would a technical document explaining the exact details of\n> the bloom filter implemenation and the hashing calculations be helpful? I will\n> be adding details into Documentation/technical/commit-graph-format.txt, but the\n> bloom filter code is an independent subsystem and could be used outside of the\n> commit-graph feature. Is it worth a separate document, or should we apply \"You\n> Ain't Gonna Need It\" principles?\n\nAs nowadays technical reference documentation is being moved from\nDocumentation/technical/api-*.txt to appropriate header files, maybe the\ndocumentation of Bloom filter API (and some technical documentation and\nreferences) be put in bloom.h?  See for example comments in strbuf.h.\n\n> [Call for advice] I plan to add unit tests for bloom.c, specifically to ensure\n> that the hash algorithm and bloom key calculations are stable across versions.\n\nAh, so the unit tests for bloom.c does not exist, yet...\n\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> ---\n>  Makefile       |   1 +\n>  bloom.c        | 201 +++++++++++++++++++++++++++++++++++++++++++++++++\n>  bloom.h        |  46 +++++++++++\n>  commit-graph.c |  32 +++++++-\n>  4 files changed, 279 insertions(+), 1 deletion(-)\n>  create mode 100644 bloom.c\n>  create mode 100644 bloom.h\n>\n> diff --git a/Makefile b/Makefile\n> index 42a061d3fb..9d5e26f5d6 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -838,6 +838,7 @@ LIB_OBJS += base85.o\n>  LIB_OBJS += bisect.o\n>  LIB_OBJS += blame.o\n>  LIB_OBJS += blob.o\n> +LIB_OBJS += bloom.o\n>  LIB_OBJS += branch.o\n>  LIB_OBJS += bulk-checkin.o\n>  LIB_OBJS += bundle.o\n\nI'll put bloom.h first, to make it easier to review.\n\n> diff --git a/bloom.h b/bloom.h\n> new file mode 100644\n> index 0000000000..ba8ae70b67\n> --- /dev/null\n> +++ b/bloom.h\n> @@ -0,0 +1,46 @@\n> +#ifndef BLOOM_H\n> +#define BLOOM_H\n> +\n> +struct commit;\n> +struct repository;\n\nThis would probably be missing if this patch was split in two:\nintroducing Bloom filter and saving Bloom filter in the repository\nmetadata (in commit-graphh file).\n\n> +\n\nO.K., the names of fields are descriptive enough so that this struct\ndoesn't need detailed description in comment (like the next one).\n\n> +struct bloom_filter_settings {\n> +\tuint32_t hash_version;\n\nDo we need full half-word for hash version?\n\n> +\tuint32_t num_hashes;\n\nDo we need full 32-bits for number of hashes?  The \"Bloom Filters in\nProbabilistic Verification\" paper mentioned in Stolee blog states that\nno one should need number of hashes greater than k=32 - the accuracy is\nso high that it doesn't matter that it is not optimal.\n\n  \"Notice one last thing about Bloom filters in verification, if $m$ is\n   several gigabytes or less and $m/n$ calls for more than about 32 index\n   functions, the accuracy is going to be so high that there is not much\n   reason to use more than 32—for the next several years at least. In\n   response to this, 3SPIN currently limits the user to $k = 32$. The\n   point of this observation is that we do not have to worry about the\n   runtime cost of $k$ being on the order of 64 or 100, because those\n   choices do not really buy us anything over 32.\"\n\nHere 'm' is the number of bits in Bloom filter, and m/n is number of\nbits per element added to filter.\n\n> +\tuint32_t bits_per_entry;\n\nAll right, we wouldn't really want large Bloom filters, as we use one\nfilter per commit to match againts one key, not single Bloom filter to\nmatch againts many keys.\n\n> +};\n> +\n> +#define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n> +\n> +/*\n> + * A bloom_filter struct represents a data segment to\n> + * use when testing hash values. The 'len' member\n> + * dictates how many uint64_t entries are stored in\n> + * 'data'.\n> + */\n> +struct bloom_filter {\n> +\tuint64_t *data;\n> +\tint len;\n> +};\n\nO.K., so it is single variable-sized (in 64-bit increments) Bloom filter\nbit vector (bitmap).\n\n> +\n> +/*\n> + * A bloom_key represents the k hash values for a\n> + * given hash input. These can be precomputed and\n> + * stored in a bloom_key for re-use when testing\n> + * against a bloom_filter.\n> + */\n> +struct bloom_key {\n> +\tuint32_t *hashes;\n> +};\n\nThat is smart.  I wonder however if it wouldn't be a good idea to\n'typedef' a hash function return type.\n\nI repeat myself: in Git case we have one key that we try to match\nagainst many Bloom filters which are never updated, while in an ordinary\ncase many keys are matched against single Bloom filter - in many cases\nupdated (with keys inserted to Bloom filter).\n\nI wonder if somebody from academia have examined such situation.\nI couldn't find a good search query.\n\n\nSidenote: perhaps Xor or Xor+ filters from Graf & Lemire (2019)\nhttps://arxiv.org/abs/1912.08258 would be better solution - they also\nassume unchanging filter.  Though they are a very fresh proposal;\nalso construction time might be important for Git.\nhttps://github.com/FastFilter/xor_singleheader\n\n> +\n> +void load_bloom_filters(void);\n> +\n> +struct bloom_filter *get_bloom_filter(struct repository *r,\n> +\t\t\t\t      struct commit *c);\n\nThose two functions really need API documentation on how they are used,\nif they are to be used in any other role, especially what is their\ncalling convention?  Why load_bloom_filters() doesn't take any\nparameters?\n\nAnyway, if this patch would be split into pure Bloom filter\nimplementation and Git use^W store of Bloom filters, then this would be\nleft for the latter patch.\n\n> +\n> +void fill_bloom_key(const char *data,\n> +\t\t    int len,\n> +\t\t    struct bloom_key *key,\n> +\t\t    struct bloom_filter_settings *settings);\n\nIt is a bit strange that two basic Bloom filter operations, namely\nadding element to Bloom filter (and constructing Bloom filter), and\ntesting whether element is in Bloom filter are not part of a public\nAPI...\n\nThis function should probably be documented, in particular the fact that\n*key is an in/out parameter.  This could also be a good place to\ndocument the mechanism itself (i.e. our implementation of Bloom filter,\nwith references), though it might be better to keep the details of how\nit works in the bloom.c - close to the actual source (while keeping\ndescription of API in bloom.h comments).\n\n> +\n> +#endif\n> diff --git a/bloom.c b/bloom.c\n> new file mode 100644\n> index 0000000000..08328cc381\n> --- /dev/null\n> +++ b/bloom.c\n> @@ -0,0 +1,201 @@\n> +#include \"git-compat-util.h\"\n> +#include \"bloom.h\"\n> +#include \"commit-graph.h\"\n> +#include \"object-store.h\"\n> +#include \"diff.h\"\n> +#include \"diffcore.h\"\n> +#include \"revision.h\"\n> +#include \"hashmap.h\"\n> +\n> +#define BITS_PER_BLOCK 64\n> +\n> +define_commit_slab(bloom_filter_slab, struct bloom_filter);\n> +\n> +struct bloom_filter_slab bloom_filters;\n\nAll right, so the Bloom filter data would be on slab.  This should\nprobably be mentioned in the commit message, like in\nhttps://lore.kernel.org/git/61559c5b-546e-d61b-d2e1-68de692f5972@gmail.com/\n\nSidenote: If I remember correctly one of the unmet prerequisites for\nswitching to generation numbers v2 (corrected commit date with monotonic\noffsets) was moving 'generation' field out of 'struct commit' and on to\nslab (possibly also 'graph_pos'), and certainly having 'corrected_date'\non slab (Inside-Out Object style).  Which probably could be done with\nCoccinelle script...\n\n> +\n> +struct pathmap_hash_entry {\n> +    struct hashmap_entry entry;\n> +    const char path[FLEX_ARRAY];\n> +};\n\nHmmm... I wonder why use hashmap and not string_list.  This is for\nadding path with leading directories to the Bloom filter, isn't it?\n\n> +\n> +static uint32_t rotate_right(uint32_t value, int32_t count)\n> +{\n> +\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n> +\tcount &= mask;\n> +\treturn ((value >> count) | (value << ((-count) & mask)));\n> +}\n\nDoes it actually work with count being negative?  Shouldn't 'count' be\nof unsigned type, and if int32_t is needed, perhaps add an assertion (if\nneeded)?  I think it does not.\n\nIt looks like it is John Regehr [2] safe and compiler-friendly\nimplementation, with explicit 8 in place of CHAR_BIT from <limits.h>,\nwhich should compile to \"rotate\" assembly instruction... it looks like\nit is the case, see https://godbolt.org/z/5JP1Jb (at least for C++\ncompiler).\n\n[2]: https://en.wikipedia.org/wiki/Circular_shift\n\n\nI wonder if this should, in the future, be a part of 'compat/', maybe\neven using compiler intrinsics for \"rotate right\" if available (see\nhttps://stackoverflow.com/a/776523/46058).  But that might be outside of\nthe scope of this patch (perhaps outside of choosing function name).\n\n> +\n\nIt would be nice to have reference to the source of algorithm, or to the\ncode that was borrowed for this in the header comment for the following\nfunction.\n\nI will be comparing the algorithm itself in Wikipedia\nhttps://en.wikipedia.org/wiki/MurmurHash#Algorithm\nand its implementation in C in qLibc library (BSD licensed)\nhttps://github.com/wolkykim/qlibc/blob/03a8ce035391adf88d6d755f9a26967c16a1a567/src/utilities/qhash.c#L258\n\n> +static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n> +{\n> +\tconst uint32_t c1 = 0xcc9e2d51;\n> +\tconst uint32_t c2 = 0x1b873593;\n> +\tconst int32_t r1 = 15;\n> +\tconst int32_t r2 = 13;\n\nThose two: r1 and r1, probably should be both uint32_t type.\n\n> +\tconst uint32_t m = 5;\n> +\tconst uint32_t n = 0xe6546b64;\n> +\tint i;\n> +\tuint32_t k1 = 0;\n> +\tconst char *tail;\n\n*tail should probably be 'uint8_t', not 'char', isn't it?\n\n> +\n> +\tint len4 = len / sizeof(uint32_t);\n\nAll length variables and parameters, i.e. `len`, `len4`, `i`, could\npossibly be `size_t` and not `int` type.\n\n> +\n> +\tconst uint32_t *blocks = (const uint32_t*)data;\n> +\n\nSome implementations copy `seed` (or assume seed=0) to the local\nvariable named `h` or `hash`.\n\n> +\tuint32_t k;\n> +\tfor (i = 0; i < len4; i++)\n> +\t{\n> +\t\tk = blocks[i];\n> +\t\tk *= c1;\n> +\t\tk = rotate_right(k, r1);\n\nShouldn't it be *rotate_left* (ROL), not rotate_right (ROR)???\nThis affects all cases / uses.\n\n> +\t\tk *= c2;\n> +\n> +\t\tseed ^= k;\n> +\t\tseed = rotate_right(seed, r2) * m + n;\n> +\t}\n> +\n> +\ttail = (data + len4 * sizeof(uint32_t));\n> +\n\nWe could have reused variable `k`, like the implementation in qLibc\ndoes, instead of introducing new `k1` variable, but this way it is more\nclean.  Or name it `remainingBytes` instead of `k1`\n\n> +\tswitch (len & (sizeof(uint32_t) - 1))\n> +\t{\n> +\tcase 3:\n> +\t\tk1 ^= ((uint32_t)tail[2]) << 16;\n> +\t\t/*-fallthrough*/\n> +\tcase 2:\n> +\t\tk1 ^= ((uint32_t)tail[1]) << 8;\n> +\t\t/*-fallthrough*/\n> +\tcase 1:\n> +\t\tk1 ^= ((uint32_t)tail[0]) << 0;\n> +\t\tk1 *= c1;\n> +\t\tk1 = rotate_right(k1, r1);\n> +\t\tk1 *= c2;\n> +\t\tseed ^= k1;\n> +\t\tbreak;\n> +\t}\n> +\n> +\tseed ^= (uint32_t)len;\n> +\tseed ^= (seed >> 16);\n> +\tseed *= 0x85ebca6b;\n> +\tseed ^= (seed >> 13);\n> +\tseed *= 0xc2b2ae35;\n> +\tseed ^= (seed >> 16);\n> +\n> +\treturn seed;\n> +}\n> +\n\nIt would be nice to have header comment describing what this function is\nintended to actually do.\n\n> +static inline uint64_t get_bitmask(uint32_t pos)\n> +{\n> +\treturn ((uint64_t)1) << (pos & (BITS_PER_BLOCK - 1));\n> +}\n\nSidenote: I wonder if ewah/bitmap.c implements something similar.\nCertainly possible consolidation, if any possible exists, should be left\nfor the future.\n\n> +\n> +void fill_bloom_key(const char *data,\n> +\t\t    int len,\n> +\t\t    struct bloom_key *key,\n> +\t\t    struct bloom_filter_settings *settings)\n> +{\n> +\tint i;\n> +\tuint32_t seed0 = 0x293ae76f;\n> +\tuint32_t seed1 = 0x7e646e2c;\n\nWhere did those constants came from?  It would be nice to have a\nreference either in header comment (in bloom.h or bloom.c), or in a\ncommit message, or both.\n\nNote that above *constants* are each used only once.\n\n> +\n> +\tuint32_t hash0 = seed_murmur3(seed0, data, len);\n> +\tuint32_t hash1 = seed_murmur3(seed1, data, len);\n\nThose are constant values, so perhaps they should be `const uint32_t`.\n\n> +\n> +\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n> +\tfor (i = 0; i < settings->num_hashes; i++)\n> +\t\tkey->hashes[i] = hash0 + i * hash1;\n\nIt looks like this code implements the double hashing technique given in\nEq. (4) in http://www.ccs.neu.edu/home/pete/pub/bloom-filters-verification.pdf\nthat is \"Bloom Filters in Probabilistic Verification\".\n\nNote that Dillinger and Manolios in this paper propose also _enhanced_\ndouble hashing algorithm (Algorithm 2 on page 11), which has closed form\ngiven by Eq. (6) - with better science-theoretical properties at\nsimilar cost.\n\n\nIt might be a good idea to explicitly state in the header comment that\nall arithmetic is performed with unsigned 32-bit integers, which means\nthat operations are performed modulo 2^32.  Or it might not be needed.\n\n> +}\n> +\n> +static void add_key_to_filter(struct bloom_key *key,\n> +\t\t\t      struct bloom_filter *filter,\n> +\t\t\t      struct bloom_filter_settings *settings)\n> +{\n> +\tint i;\n> +\tuint64_t mod = filter->len * BITS_PER_BLOCK;\n> +\n> +\tfor (i = 0; i < settings->num_hashes; i++) {\n> +\t\tuint64_t hash_mod = key->hashes[i] % mod;\n> +\t\tuint64_t block_pos = hash_mod / BITS_PER_BLOCK;\n\nAll right.  Because Bloom filters for different commits (and the same\nkey) may have different lengths, we can perform modulo operation only\nhere.  `hash_mod` is i-th hash modulo size of filter, and `block_pos` is\nthe block the '1' bit would go into.\n\n> +\n> +\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n\nI'm not quite convinced that get_bitmask() is a good name: this function\nreturns bitmap with hash_mod's bit set to 1.  On the other hand it\ndoesn't matter, because it is static (file-local) helper function.\n\nNever mind then.\n\n> +\t}\n> +}\n> +\n> +void load_bloom_filters(void)\n> +{\n> +\tinit_bloom_filter_slab(&bloom_filters);\n> +}\n\nWhy *load* if all it does is initialize?\n\n> +\n> +struct bloom_filter *get_bloom_filter(struct repository *r,\n> +\t\t\t\t      struct commit *c)\n\nI will not comment on this function; see Jeff King reply and Derrick\nStolee reply.\n\n> +{\n[...]\n> +}\n> \\ No newline at end of file\n\nWhy there is no newline at the end of the file?  Accident?\n\n> diff --git a/commit-graph.c b/commit-graph.c\n> index e771394aff..61e60ff98a 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -16,6 +16,7 @@\n>  #include \"hashmap.h\"\n>  #include \"replace-object.h\"\n>  #include \"progress.h\"\n> +#include \"bloom.h\"\n>  \n>  #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n>  #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n> @@ -794,9 +795,11 @@ struct write_commit_graph_context {\n>  \tunsigned append:1,\n>  \t\t report_progress:1,\n>  \t\t split:1,\n> -\t\t check_oids:1;\n> +\t\t check_oids:1,\n> +\t\t bloom:1;\n\nVery minor nitpick: why `bloom` and not `bloom_filter`?\n\n>  \n>  \tconst struct split_commit_graph_opts *split_opts;\n> +\tuint32_t total_bloom_filter_size;\n\nAll right, I guess size of all Bloom filters would fit in uint32_t, no\nneed for size_t, is it?\n\nShouldn't it be total_bloom_filters_size -- it is not a single Bloom\nfilter, but many (minor nitpick)?\n\n>  };\n>  \n>  static void write_graph_chunk_fanout(struct hashfile *f,\n> @@ -1139,6 +1142,28 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \tstop_progress(&ctx->progress);\n>  }\n>  \n> +static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n> +{\n> +\tint i;\n> +\tstruct progress *progress = NULL;\n> +\n> +\tload_bloom_filters();\n> +\n> +\tif (ctx->report_progress)\n> +\t\tprogress = start_progress(\n> +\t\t\t_(\"Computing commit diff Bloom filters\"),\n> +\t\t\tctx->commits.nr);\n> +\n> +\tfor (i = 0; i < ctx->commits.nr; i++) {\n> +\t\tstruct commit *c = ctx->commits.list[i];\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n> +\t\tctx->total_bloom_filter_size += sizeof(uint64_t) * filter->len;\n\nWouldn't it be more future proof instead of using `sizeof(uint64_t)` to\nuse `sizeof(filter->data[0])` here?  This may be not worth it, and be\nless readable (we have hard-coded use of 64-bits blocks in other places).\n\n> +\t\tdisplay_progress(progress, i + 1);\n> +\t}\n> +\n> +\tstop_progress(&progress);\n> +}\n> +\n>  static int add_ref_to_list(const char *refname,\n>  \t\t\t   const struct object_id *oid,\n>  \t\t\t   int flags, void *cb_data)\n> @@ -1791,6 +1816,8 @@ int write_commit_graph(const char *obj_dir,\n>  \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n>  \tctx->check_oids = flags & COMMIT_GRAPH_WRITE_CHECK_OIDS ? 1 : 0;\n>  \tctx->split_opts = split_opts;\n> +\tctx->bloom = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n\nAll right, this flag was defined in [PATCH 1/9].\n\nThe ordering of setting `ctx` members looks a bit strange.  Now it is\nneither check `flags` firsts, neither keep related stuff together (see\nctx->split vs ctx->split_opts).  This is a very minor nitpick.\n\n> +\tctx->total_bloom_filter_size = 0;\n>  \n>  \tif (ctx->split) {\n>  \t\tstruct commit_graph *g;\n> @@ -1885,6 +1912,9 @@ int write_commit_graph(const char *obj_dir,\n>  \n>  \tcompute_generation_numbers(ctx);\n>  \n> +\tif (ctx->bloom)\n> +\t\tcompute_bloom_filters(ctx);\n> +\n>  \tres = write_commit_graph_file(ctx);\n>  \n>  \tif (ctx->split)\n\nRegards,\n-- \nJakub Narębski\n"},{"id":"389360","messageId":"865zhnzcg5.fsf@gmail.com","threadId":"52499","inReplyTo":"a15f87fdcbea1a37a20a05135832b42f36f682f1.1576879520.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 3/9] commit-graph: use MAX_NUM_CHUNKS","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-07T12:19:06Z","receivedAt":"2020-01-07T12:19:14Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> This is a minor cleanup to make it easier to change the\n> number of chunks being written to the commit-graph in the future.\n\nVery minor nit: in the whole commit message it is not stated explicitly\nwhat MAX_NUM_CHUNKS is for, though it is very easy to guess (from the\nname itself).\n\n>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  commit-graph.c | 5 +++--\n>  1 file changed, 3 insertions(+), 2 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 61e60ff98a..8c4941eeaa 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -24,6 +24,7 @@\n>  #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n>  #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n>  #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n> +#define MAX_NUM_CHUNKS 5\n\nMinor nit: MAX_NUM_CHUNKS or MAX_CHUNKS?\n\n>  \n>  #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n>  \n> @@ -1381,8 +1382,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \tint fd;\n>  \tstruct hashfile *f;\n>  \tstruct lock_file lk = LOCK_INIT;\n> -\tuint32_t chunk_ids[6];\n> -\tuint64_t chunk_offsets[6];\n> +\tuint32_t chunk_ids[MAX_NUM_CHUNKS + 1];\n> +\tuint64_t chunk_offsets[MAX_NUM_CHUNKS + 1];\n\nLooks good.  I guess we won't ever have more chunks than 5:\nOIDF, OIDL, CDAT, EDGE, BASE (and they cannot repeat, and last two are\noptional).\n\n>  \tconst unsigned hashsz = the_hash_algo->rawsz;\n>  \tstruct strbuf progress_title = STRBUF_INIT;\n>  \tint num_chunks = 3;\n\nGood.\n\nLooks good to me.\n-- \nJakub Narębski\n"},{"id":"389375","messageId":"86h817xr1i.fsf@gmail.com","threadId":"52499","inReplyTo":"3182a11f7c07af834ba71dc7861742458754eb91.1576879520.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 4/9] commit-graph: document bloom filter format","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-07T14:46:49Z","receivedAt":"2020-01-07T14:47:08Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Update the technical documentation for commit-graph-format with BIDX\n> and BDAT chunk information.\n>\n> RFC Notes:\n> 1. [Call for advice] We specifically mention that we are using Bloom\n>    filters in this technical document. Should this document also be\n>    made open to other data structures in the future, with versioning\n>    information?\n\nI'm not sure.  In theory we might want to switch to another\nprobabilistic set inclusion query structure, like xor filters or cuckoo\nhashing.\n\nOn one hand side we could use separate chunks (e.g. XIDX, XDAT for xor\nfilters), on the other hand we need only one such structure.  On the\ngripping hand this can be left for the future, if needed.\n\nSidenote: using Bloom filters is somewhat encoded in the name of chunk\n(B from Bloom filter).  I don't have a better poposal for 4-char name\n(XIDX / XDAT for cXange?  CHDX / CHDT for CHange?  FIDX / FDAT for\nchanged Files?... I don't know).\n\n>\n> 2. [Call for advice] We are also not describing the explicit nature\n>    of how we store the bloom filter binary data. Would it be useful\n>    to document details about the hash algorithm, the number of hashes\n>    and the specific seed values we are using in a separate document,\n>    or perhaps in a separate section in this document?\n\nI think it would be best to keep description of the commit graph format\nconcise.  The details about Bloom filter implementation would be better\nput in Documentation/technical/commit-graph.txt in my opinion, together\nwith reasoning behind it (perhaps borrowing from Derrick Stolee blog\npost).\n\nThis could be done as a separate patch.\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  Documentation/technical/commit-graph-format.txt | 17 +++++++++++++++++\n>  1 file changed, 17 insertions(+)\n>\n> diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n> index a4f17441ae..6497f19f08 100644\n> --- a/Documentation/technical/commit-graph-format.txt\n> +++ b/Documentation/technical/commit-graph-format.txt\n> @@ -17,6 +17,9 @@ metadata, including:\n>  - The parents of the commit, stored using positional references within\n>    the graph file.\n>  \n> +- The bloom filter of the commit carrying the paths that were changed between\n> +  the commit and it's first parent.\n\ns/bloom/Bloom/ and s/it's/its/\n\nI am not sure about exact wording, but I could at this time think of a\nbetter but concise way of stating it.\n\n> +\n>  These positional references are stored as unsigned 32-bit integers\n>  corresponding to the array position within the list of commit OIDs. Due\n>  to some special constants we use to track parents, we can store at most\n> @@ -93,6 +96,20 @@ CHUNK DATA:\n>        positions for the parents until reaching a value with the most-significant\n>        bit on. The other bits correspond to the position of the last parent.\n>  \n> +  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) [Optional]\n> +      For each commit we store the offset of its bloom filter in the BDAT chunk\n> +      as follows:\n> +      BIDX[i] = number of 8-byte words in all the bloom filters from commit 0 to\n> +\t\tcommit i (inclusive)\n\nI think it would be better for consistency and ease of reading to follow\nthe example of OID Fanout (OIDF) chunk description:\n\n +  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) (N * 4 bytes) [Optional]\n +      The ith entry, BIDX[i], stores the number of 8-byte word blocks\n +      in all Bloom filters from commit 0 up to commit i (inclusive)\n +      in lexicographical order.\n\nMaybe even add the following to make implementing it easier:\n\n +      Data for Bloom filter for i-th commit spans from BIDX[i-1] to\n +      BIDX[i] (plus header length), where we take BIDX[-1] to be 0.\n\nIs it possible for (BIDX[i] - BIDX[i-1]) to be zero (no Bloom filter),\nfor example for commits with more than 512 changes?  Or is this case\nhandled by 1 8-byte word Bloom filter of all bits sets to '1', i.e.\n0xffffffffffffffff?\n\nHow the case of too many changes is distingushed from the case of no\nchanges (`git commit --allow-empty`, or `git merge --ours`)?  Is the\ncase of no changes uninteresting, i.e. Bloom filter consisting of zero,\nthat is with all bits set to '0'?\n\n> +\n> +  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n> +      * It starts with three 32 bit integers for the\n\nI would say \"It starts with the header consisting of three unsigned\n32-bit integers:\" (but current version is not bad).\n\nI wonder if this metadata should perhaps be put in a separate chunk,\nBMET... (Bloom filter METadata).\n\n> +\t    - version of the hash algorithm being used\n\nThis number not only encodes that the base hash algorithm being used is\n32-bit Murmur3 hash, but that 'k' hashes used in the computation are\ncreated out of Murmur3 hash using double hashing technique, and\nspecifies two specific seed values for this double hashing technique.\n\n[Maybe we should store those two seed values here too?]\n\nIt might be important to say that the currently supported version is\n'1', and if Git encounters unknown hashing algorithm version it should\nnot use Bloom filter data.\n\nUnless we store encoded _name_ of the hash algorithm, e.g. bytes\n'm','u','r','3' for MurmurHash3_32... though it is about more than\na base hash.\n\nDo we need whole 4 bytes for hash version, or is it for ease of use and\nalignment?\n\n> +\t    - the number of hashes used in the computation\n\nAll right.  Perhaps we should test in the future patches that the value\ndifferent from the default of 7 would also work.\n\nAlso 8-bits / 1 byte for number of hashes (hash functions) should be\nenough: as I have written in prevous reply there is no need for k > 32.\n\n> +\t    - the number of bits per entry\n\nThis is important for construction of Bloom filter, but I think it is\nnot necessary to use it -- so it may not be necessary to store it.\n\nWould also fit in a single byte: we don't need exceedingly low false\npositive probability.\n\nWe could use it to estimate the false positive probability, and...\n\n> +\t  * The rest of the chunk is the concatenation of all the computed bloom \n> +\t  filters for the commits in lexicographic order.\n\n  +\t * The rest of the chunk is the concatenation of all the computed Bloom \n  +\t   filters for the commits in lexicographic order.\n\nIt would be, I think, a good idea to make it explicit that BDAT is\npresent iff BIDX is present (iff == if and only if), i.e. that either\nboth or neither of those chunks should be present.\n\n> +\n>    Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n>        This list of H-byte hashes describe a set of B commit-graph files that\n>        form a commit-graph chain. The graph position for the ith commit in this\n\nBest,\n-- \nJakub Narębski\n"},{"id":"389390","messageId":"867e23xnkh.fsf@gmail.com","threadId":"52499","inReplyTo":"7648021072ca11153ac65c90f0ebed5973f20e1a.1576879520.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 5/9] commit-graph: write changed path bloom filters to commit-graph file.","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-07T16:01:50Z","receivedAt":"2020-01-07T16:01:58Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Write bloom filters to the commit-graph using the format described in\n> Documentation/technical/commit-graph-format.txt\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n\nLooks good to me.\n\n> ---\n>  commit-graph.c | 81 +++++++++++++++++++++++++++++++++++++++++++++++++-\n>  commit-graph.h |  5 ++++\n>  2 files changed, 85 insertions(+), 1 deletion(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 8c4941eeaa..def2ade166 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -24,7 +24,9 @@\n>  #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n>  #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n>  #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n> -#define MAX_NUM_CHUNKS 5\n> +#define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n> +#define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n> +#define MAX_NUM_CHUNKS 7\n\nVery minor nitpick: shouldn't we follow the order in the\ncommit-graph-format.txt document (i.e. \"BASE\" as last chunk and last\npreprocessor constant)?\n\n>  \n>  #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n>  \n> @@ -282,6 +284,32 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n>  \t\t\t\tchunk_repeated = 1;\n>  \t\t\telse\n>  \t\t\t\tgraph->chunk_base_graphs = data + chunk_offset;\n> +\t\t\tbreak;\n> +\n> +\t\tcase GRAPH_CHUNKID_BLOOMINDEXES:\n> +\t\t\tif (graph->chunk_bloom_indexes)\n> +\t\t\t\tchunk_repeated = 1;\n> +\t\t\telse\n> +\t\t\t\tgraph->chunk_bloom_indexes = data + chunk_offset;\n> +\t\t\tbreak;\n\nAll right.\n\n> +\n> +\t\tcase GRAPH_CHUNKID_BLOOMDATA:\n> +\t\t\tif (graph->chunk_bloom_data)\n> +\t\t\t\tchunk_repeated = 1;\n> +\t\t\telse {\n> +\t\t\t\tuint32_t hash_version;\n> +\t\t\t\tgraph->chunk_bloom_data = data + chunk_offset;\n> +\t\t\t\thash_version = get_be32(data + chunk_offset);\n\nAll right, now I see why all those header values for BDAT chunk are\ndefined to be 32-bit integers.  For code simplicity.\n\n> +\n> +\t\t\t\tif (hash_version != 1)\n> +\t\t\t\t\tbreak;\n\nWhat does it mean for Git?  Behave as if there were no Bloom filter\ndata?\n\n> +\n> +\t\t\t\tgraph->settings = xmalloc(sizeof(struct bloom_filter_settings));\n> +\t\t\t\tgraph->settings->hash_version = hash_version;\n> +\t\t\t\tgraph->settings->num_hashes = get_be32(data + chunk_offset + 4);\n> +\t\t\t\tgraph->settings->bits_per_entry = get_be32(data + chunk_offset + 8);\n\nAll right, looks O.K.\n\n> +\t\t\t}\n> +\t\t\tbreak;\n>  \t\t}\n>  \n>  \t\tif (chunk_repeated) {\n> @@ -996,6 +1024,39 @@ static void write_graph_chunk_extra_edges(struct hashfile *f,\n>  \t}\n>  }\n>  \n> +static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n> +\t\t\t\t\t    struct write_commit_graph_context *ctx)\n> +{\n> +\tstruct commit **list = ctx->commits.list;\n> +\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n> +\tuint32_t cur_pos = 0;\n> +\n> +\twhile (list < last) {\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tcur_pos += filter->len;\n> +\t\thashwrite_be32(f, cur_pos);\n> +\t\tlist++;\n> +\t}\n\nWhy not follow the write_graph_chunk_oids() example, instead of\nwrite_graph_chunk_data(), that is use simply:\n\n  +\tstruct commit **list = ctx->commits.list;\n  +\tuint32_t cur_pos = 0;\n  +\n  +\tfor (count = 0; count < ctx->commits.nr; count++, list++) {\n  +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n  +\t\tcur_pos += filter->len;\n  +\t\thashwrite_be32(f, cur_pos);\n  +\t}\n\nI guess using here\n\n  +\t\tcur_pos += get_bloom_filter(ctx->r, *list)->len;\n\nwould be too cryptic, and hard to debug?\n\nAlso, wouldn't we need\n\n  +\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n\nbefore hashwrite_be32()?\n\n> +}\n> +\n> +static void write_graph_chunk_bloom_data(struct hashfile *f,\n> +\t\t\t\t\t struct write_commit_graph_context *ctx,\n> +\t\t\t\t\t struct bloom_filter_settings *settings)\n> +{\n> +\tstruct commit **first = ctx->commits.list;\n\nEven if we decide to use `while` loop, like write_graph_chunk_data(),\nand not `for` loop, like write_graph_chunk_oids(), why the change from\n`struct commit **list = ...` to `struct commit **first = ...`?\n\n> +\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n> +\n> +\thashwrite_be32(f, settings->hash_version);\n> +\thashwrite_be32(f, settings->num_hashes);\n> +\thashwrite_be32(f, settings->bits_per_entry);\n\nAll right, simple.\n\n> +\n> +\twhile (first < last) {\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *first);\n\nHmmm... wouldn't this compute Bloom filter second time?\nget_bloom_filter() does work unconditionally.\n\nWouldn't\n\n  +\t\tstruct bloom_filter *filter = bloom_filter_slab_at(&bloom_filters, *first);\n\nbe enough?\n\nOr make get_bloom_filter() use *_peek() to check if Bloom filter for\ngiven commit was already computed, and only if it returns NULL do the\nwork.\n\n> +\t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n\nMight need display_progress() before hashwrote().\n\n> +\t\tfirst++;\n> +\t}\n> +}\n> +\n>  static int oid_compare(const void *_a, const void *_b)\n>  {\n>  \tconst struct object_id *a = (const struct object_id *)_a;\n> @@ -1388,6 +1449,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \tstruct strbuf progress_title = STRBUF_INIT;\n>  \tint num_chunks = 3;\n>  \tstruct object_id file_hash;\n> +\tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n>  \n>  \tif (ctx->split) {\n>  \t\tstruct strbuf tmp_file = STRBUF_INIT;\n> @@ -1432,6 +1494,12 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n>  \t\tnum_chunks++;\n>  \t}\n> +\tif (ctx->bloom) {\n> +\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMINDEXES;\n> +\t\tnum_chunks++;\n> +\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMDATA;\n> +\t\tnum_chunks++;\n> +\t}\n\nLooks all right.\n\n>  \tif (ctx->num_commit_graphs_after > 1) {\n>  \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BASE;\n>  \t\tnum_chunks++;\n> @@ -1450,6 +1518,13 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \t\t\t\t\t\t4 * ctx->num_extra_edges;\n>  \t\tnum_chunks++;\n>  \t}\n> +\tif (ctx->bloom) {\n> +\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] + sizeof(uint32_t) * ctx->commits.nr;\n> +\t\tnum_chunks++;\n> +\n> +\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] + sizeof(uint32_t) * 3 + ctx->total_bloom_filter_size;\n> +\t\tnum_chunks++;\n> +\t}\n\nBetter wrap those long lines, like above:\n\n  +\tif (ctx->bloom) {\n  +\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n  +\t\t\t\t\t\tsizeof(uint32_t) * ctx->commits.nr;\n  +\t\tnum_chunks++;\n  +\n  +\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n  +\t\t\t\t\t\tsizeof(uint32_t) * 3 + ctx->total_bloom_filter_size;\n  +\t\tnum_chunks++;\n  +\t}\n\n>  \tif (ctx->num_commit_graphs_after > 1) {\n>  \t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n>  \t\t\t\t\t\thashsz * (ctx->num_commit_graphs_after - 1);\n> @@ -1487,6 +1562,10 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \twrite_graph_chunk_data(f, hashsz, ctx);\n>  \tif (ctx->num_extra_edges)\n>  \t\twrite_graph_chunk_extra_edges(f, ctx);\n> +\tif (ctx->bloom) {\n> +\t\twrite_graph_chunk_bloom_indexes(f, ctx);\n> +\t\twrite_graph_chunk_bloom_data(f, ctx, &bloom_settings);\n> +\t}\n\nAll right.\n\n>  \tif (ctx->num_commit_graphs_after > 1 &&\n>  \t    write_graph_chunk_base(f, ctx)) {\n>  \t\treturn -1;\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 952a4b83be..2202ad91ae 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -10,6 +10,7 @@\n>  #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n>  \n>  struct commit;\n> +struct bloom_filter_settings;\n>  \n>  char *get_commit_graph_filename(const char *obj_dir);\n>  int open_commit_graph(const char *graph_file, int *fd, struct stat *st);\n> @@ -58,6 +59,10 @@ struct commit_graph {\n>  \tconst unsigned char *chunk_commit_data;\n>  \tconst unsigned char *chunk_extra_edges;\n>  \tconst unsigned char *chunk_base_graphs;\n> +\tconst unsigned char *chunk_bloom_indexes;\n> +\tconst unsigned char *chunk_bloom_data;\n> +\n> +\tstruct bloom_filter_settings *settings;\n\nShould this be part of `struct commit_graph`?  Shouldn't we free() this\ndata, or is it a pointer into xmmap-ped file... no it isn't -- we\nxalloc() it, so we should free() it.\n\nI think it should be done in 'cleanup:' section of write_commit_graph(),\nbut I am not entirely sure.\n\n>  };\n>  \n>  struct commit_graph *load_commit_graph_one_fd_st(int fd, struct stat *st);\n\nBest,\n-- \nJakub Narębski\n"},{"id":"389440","messageId":"86r20awzxe.fsf@gmail.com","threadId":"52499","inReplyTo":"85bfdfa59c48891343e3eeb740a4b3554405909a.1576879520.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 6/9] commit-graph: test commit-graph write --changed-paths","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-08T00:32:29Z","receivedAt":"2020-01-08T00:32:40Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Add tests for the --changed-paths feature when writing\n> commit-graphs.\n\nIt doesn't look however as if this test is actually testing the _Bloom\nfilter_ functionality itself -- because this test looks like\ncopy'n'paste of t/t5324-split-commit-graph.sh, just with\n`--changed-paths` added to the `git commit-graph write` invocation, and\nadded checking via enhanced test-tool that there are Bloom filter chunks\n(\"bloom_indexes\" and \"bloom_data\").\n\nPlease correct me if I am wrong, but this looks like a simple sanity\ncheck for me.\n\n>\n> RFC Notes:\n> I plan to split this test across some of the earlier commits\n> as appropriate.\n\nAbout adding tests to earlier commits in this series:\n\n1. Testing Bloom filter functionality:\n   - creating Bloom filter and adding elements to it\n   - testing Bloom filter functionality\n     - for element in set the answer is \"maybe\"\n     - for element not in set the answer is \"no\" or \"maybe\"\n   - automatic resizing works (6 and 7 elements)\n   - it works for different number of hash functions,\n     and different number of bits per element (maybe?)\n\n2. Testing Bloom filter for commit changeset:\n   - it works for commit with no changes\n   - it works for merge commit with no changes to first parent\n     (`git merge --strategy=ours`)\n   - with number of changes that require filter size change\n   - with maximal number of changes, one changed file less,\n     one changed file more\n   - that for file deeper in hierarchy, path/to/file, all of\n     changed directories (path/to/ and path/) are also added\n\n3. Test writing and reading commit-graph with Bloom filters\n   - that after writing Bloom filters with `--changed-paths`\n     the data is present in commit-graph files\n   - it works correctly with split commit-graph\n   - it doesn't crash if confronted with unknown settings:\n     hash version different than 1, different number of hash\n     functions, different number of bits per element\n\n4. Bloom filter specific `git commit-graph verify` parts\n   - fail if Bloom filter chunks appear multiple times\n   - fail if only one of BIDX or BDAT chunks are present\n   - fail if BIDX is not monotonic, that is if size of Bloom filter\n     for a commit is negative\n   - fail if BDAT size does not agree with BIDX,\n     being either too small, or too large\n   - check if values of number of hash functions\n     and number of bits per element added are sane\n\n5. Using Bloom filters to speed up Git operations\n   - test that with and without Bloom filters (or commit-graph)\n     the following operations work the same:\n     - git log -- <path/to/file>\n     - git log -- <path/to/directory>\n     - git log -- '*.c'  # or other glob pattern\n     - git log -- <file1> <file2>\n     - git log --follow <file>\n     - maybe also `git log --full-history -- <file>`\n   - if possible, add performance tests, see `t/perf`\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  t/helper/test-read-graph.c    |   4 +\n>  t/t5325-commit-graph-bloom.sh | 255 ++++++++++++++++++++++++++++++++++\n>  2 files changed, 259 insertions(+)\n>  create mode 100755 t/t5325-commit-graph-bloom.sh\n>\n> diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\n> index d2884efe0a..aff597c7a3 100644\n> --- a/t/helper/test-read-graph.c\n> +++ b/t/helper/test-read-graph.c\n> @@ -45,6 +45,10 @@ int cmd__read_graph(int argc, const char **argv)\n>  \t\tprintf(\" commit_metadata\");\n>  \tif (graph->chunk_extra_edges)\n>  \t\tprintf(\" extra_edges\");\n> +\tif (graph->chunk_bloom_indexes)\n> +\t\tprintf(\" bloom_indexes\");\n> +\tif (graph->chunk_bloom_data)\n> +\t\tprintf(\" bloom_data\");\n\nAll right, though it is very basic information.\n\n>  \tprintf(\"\\n\");\n>  \n>  \tUNLEAK(graph);\n> diff --git a/t/t5325-commit-graph-bloom.sh b/t/t5325-commit-graph-bloom.sh\n> new file mode 100755\n> index 0000000000..d7ef0e7fb3\n> --- /dev/null\n> +++ b/t/t5325-commit-graph-bloom.sh\n> @@ -0,0 +1,255 @@\n> +#!/bin/sh\n> +\n> +test_description='commit graph with bloom filters'\n> +. ./test-lib.sh\n> +\n> +test_expect_success 'setup repo' '\n> +\tgit init &&\n> +\tgit config core.commitGraph true &&\n> +\tgit config gc.writeCommitGraph false &&\n> +\tinfodir=\".git/objects/info\" &&\n> +\tgraphdir=\"$infodir/commit-graphs\" &&\n> +\ttest_oid_init\n> +'\n> +\n> +graph_read_expect() {\n\nStyle: space between function name and parentheses, i.e.\n\n  +graph_read_expect () {\n\n> +\tOPTIONAL=\"\"\n\nNot used anywhere.\n\n> +\tNUM_CHUNKS=5\n> +\tif test ! -z $2\n\nIt might be good idea to add names to those parameters by setting some\nlocal variables to $1 and $2; or, alternatively add comment describing\nthis function.\n\n> +\tthen\n> +\t\tOPTIONAL=\" $2\"\n> +\t\tNUM_CHUNKS=$((NUM_CHUNKS + $(echo \"$2\" | wc -w)))\n> +\tfi\n> +\tcat >expect <<- EOF\n> +\theader: 43475048 1 1 $NUM_CHUNKS 0\n> +\tnum_commits: $1\n> +\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n> +\tEOF\n> +\ttest-tool read-graph >output &&\n> +\ttest_cmp expect output\n> +}\n\nNo comments below this point...\n\nBest,\n\n  Jakub Narębski\n\n\n> +\n> +test_expect_success 'create commits and write commit-graph' '\n> +\tfor i in $(test_seq 3)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit commit-graph write --reachable --changed-paths &&\n> +\ttest_path_is_file $infodir/commit-graph &&\n> +\tgraph_read_expect 3\n> +'\n> +\n> +graph_git_two_modes() {\n> +\tgit -c core.commitGraph=true $1 >output\n> +\tgit -c core.commitGraph=false $1 >expect\n> +\ttest_cmp expect output\n> +}\n> +\n> +graph_git_behavior() {\n> +\tMSG=$1\n> +\tBRANCH=$2\n> +\tCOMPARE=$3\n> +\ttest_expect_success \"check normal git operations: $MSG\" '\n> +\t\tgraph_git_two_modes \"log --oneline $BRANCH\" &&\n> +\t\tgraph_git_two_modes \"log --topo-order $BRANCH\" &&\n> +\t\tgraph_git_two_modes \"log --graph $COMPARE..$BRANCH\" &&\n> +\t\tgraph_git_two_modes \"branch -vv\" &&\n> +\t\tgraph_git_two_modes \"merge-base -a $BRANCH $COMPARE\"\n> +\t'\n> +}\n> +\n> +graph_git_behavior 'graph exists' commits/3 commits/1\n> +\n> +verify_chain_files_exist() {\n> +\tfor hash in $(cat $1/commit-graph-chain)\n> +\tdo\n> +\t\ttest_path_is_file $1/graph-$hash.graph || return 1\n> +\tdone\n> +}\n> +\n> +test_expect_success 'add more commits, and write a new base graph' '\n> +\tgit reset --hard commits/1 &&\n> +\tfor i in $(test_seq 4 5)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit reset --hard commits/2 &&\n> +\tfor i in $(test_seq 6 10)\n> +\tdo\n> +\t\ttest_commit $i &&\n> +\t\tgit branch commits/$i || return 1\n> +\tdone &&\n> +\tgit reset --hard commits/2 &&\n> +\tgit merge commits/4 &&\n> +\tgit branch merge/1 &&\n> +\tgit reset --hard commits/4 &&\n> +\tgit merge commits/6 &&\n> +\tgit branch merge/2 &&\n> +\tgit commit-graph write --reachable --changed-paths &&\n> +\tgraph_read_expect 12\n> +'\n> +\n> +test_expect_success 'fork and fail to base a chain on a commit-graph file' '\n> +\ttest_when_finished rm -rf fork &&\n> +\tgit clone . fork &&\n> +\t(\n> +\t\tcd fork &&\n> +\t\trm .git/objects/info/commit-graph &&\n> +\t\techo \"$(pwd)/../.git/objects\" >.git/objects/info/alternates &&\n> +\t\ttest_commit new-commit &&\n> +\t\tgit commit-graph write --reachable --split --changed-paths &&\n> +\t\ttest_path_is_file $graphdir/commit-graph-chain &&\n> +\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n> +\t\tverify_chain_files_exist $graphdir\n> +\t)\n> +'\n> +\n> +test_expect_success 'add three more commits, write a tip graph' '\n> +\tgit reset --hard commits/3 &&\n> +\tgit merge merge/1 &&\n> +\tgit merge commits/5 &&\n> +\tgit merge merge/2 &&\n> +\tgit branch merge/3 &&\n> +\tgit commit-graph write --reachable --split --changed-paths &&\n> +\ttest_path_is_missing $infodir/commit-graph &&\n> +\ttest_path_is_file $graphdir/commit-graph-chain &&\n> +\tls $graphdir/graph-*.graph >graph-files &&\n> +\ttest_line_count = 2 graph-files &&\n> +\tverify_chain_files_exist $graphdir\n> +'\n> +\n> +graph_git_behavior 'split commit-graph: merge 3 vs 2' merge/3 merge/2\n> +\n> +test_expect_success 'add one commit, write a tip graph' '\n> +\ttest_commit 11 &&\n> +\tgit branch commits/11 &&\n> +\tgit commit-graph write --reachable --split --changed-paths &&\n> +\ttest_path_is_missing $infodir/commit-graph &&\n> +\ttest_path_is_file $graphdir/commit-graph-chain &&\n> +\tls $graphdir/graph-*.graph >graph-files &&\n> +\ttest_line_count = 3 graph-files &&\n> +\tverify_chain_files_exist $graphdir\n> +'\n> +\n> +graph_git_behavior 'three-layer commit-graph: commit 11 vs 6' commits/11 commits/6\n> +\n> +test_expect_success 'add one commit, write a merged graph' '\n> +\ttest_commit 12 &&\n> +\tgit branch commits/12 &&\n> +\tgit commit-graph write --reachable --split --changed-paths &&\n> +\ttest_path_is_file $graphdir/commit-graph-chain &&\n> +\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n> +\tls $graphdir/graph-*.graph >graph-files &&\n> +\ttest_line_count = 2 graph-files &&\n> +\tverify_chain_files_exist $graphdir\n> +'\n> +\n> +graph_git_behavior 'merged commit-graph: commit 12 vs 6' commits/12 commits/6\n> +\n> +test_expect_success 'create fork and chain across alternate' '\n> +\tgit clone . fork &&\n> +\t(\n> +\t\tcd fork &&\n> +\t\tgit config core.commitGraph true &&\n> +\t\trm -rf $graphdir &&\n> +\t\techo \"$(pwd)/../.git/objects\" >.git/objects/info/alternates &&\n> +\t\ttest_commit 13 &&\n> +\t\tgit branch commits/13 &&\n> +\t\tgit commit-graph write --reachable --split --changed-paths &&\n> +\t\ttest_path_is_file $graphdir/commit-graph-chain &&\n> +\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n> +\t\tls $graphdir/graph-*.graph >graph-files &&\n> +\t\ttest_line_count = 1 graph-files &&\n> +\t\tgit -c core.commitGraph=true  rev-list HEAD >expect &&\n> +\t\tgit -c core.commitGraph=false rev-list HEAD >actual &&\n> +\t\ttest_cmp expect actual &&\n> +\t\ttest_commit 14 &&\n> +\t\tgit commit-graph write --reachable --split --changed-paths --object-dir=.git/objects/ &&\n> +\t\ttest_line_count = 3 $graphdir/commit-graph-chain &&\n> +\t\tls $graphdir/graph-*.graph >graph-files &&\n> +\t\ttest_line_count = 1 graph-files\n> +\t)\n> +'\n> +\n> +graph_git_behavior 'alternate: commit 13 vs 6' commits/13 commits/6\n> +\n> +test_expect_success 'test merge stragety constants' '\n> +\tgit clone . merge-2 &&\n> +\t(\n> +\t\tcd merge-2 &&\n> +\t\tgit config core.commitGraph true &&\n> +\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n> +\t\ttest_commit 14 &&\n> +\t\tgit commit-graph write --reachable --split --changed-paths --size-multiple=2 &&\n> +\t\ttest_line_count = 3 $graphdir/commit-graph-chain\n> +\n> +\t) &&\n> +\tgit clone . merge-10 &&\n> +\t(\n> +\t\tcd merge-10 &&\n> +\t\tgit config core.commitGraph true &&\n> +\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n> +\t\ttest_commit 14 &&\n> +\t\tgit commit-graph write --reachable --split --changed-paths --size-multiple=10 &&\n> +\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n> +\t\tls $graphdir/graph-*.graph >graph-files &&\n> +\t\ttest_line_count = 1 graph-files\n> +\t) &&\n> +\tgit clone . merge-10-expire &&\n> +\t(\n> +\t\tcd merge-10-expire &&\n> +\t\tgit config core.commitGraph true &&\n> +\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n> +\t\ttest_commit 15 &&\n> +\t\tgit commit-graph write --reachable --split --changed-paths --size-multiple=10 --expire-time=1980-01-01 &&\n> +\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n> +\t\tls $graphdir/graph-*.graph >graph-files &&\n> +\t\ttest_line_count = 3 graph-files\n> +\t) &&\n> +\tgit clone --no-hardlinks . max-commits &&\n> +\t(\n> +\t\tcd max-commits &&\n> +\t\tgit config core.commitGraph true &&\n> +\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n> +\t\ttest_commit 16 &&\n> +\t\ttest_commit 17 &&\n> +\t\tgit commit-graph write --reachable --split --changed-paths --max-commits=1 &&\n> +\t\ttest_line_count = 1 $graphdir/commit-graph-chain &&\n> +\t\tls $graphdir/graph-*.graph >graph-files &&\n> +\t\ttest_line_count = 1 graph-files\n> +\t)\n> +'\n> +\n> +test_expect_success 'remove commit-graph-chain file after flattening' '\n> +\tgit clone . flatten &&\n> +\t(\n> +\t\tcd flatten &&\n> +\t\ttest_line_count = 2 $graphdir/commit-graph-chain &&\n> +\t\tgit commit-graph write --reachable &&\n> +\t\ttest_path_is_missing $graphdir/commit-graph-chain &&\n> +\t\tls $graphdir >graph-files &&\n> +\t\ttest_must_be_empty graph-files\n> +\t)\n> +'\n> +\n> +graph_git_behavior 'graph exists' merge/octopus commits/12\n> +\n> +test_expect_success 'split across alternate where alternate is not split' '\n> +\tgit commit-graph write --reachable &&\n> +\ttest_path_is_file .git/objects/info/commit-graph &&\n> +\tcp .git/objects/info/commit-graph . &&\n> +\tgit clone --no-hardlinks . alt-split &&\n> +\t(\n> +\t\tcd alt-split &&\n> +\t\trm -f .git/objects/info/commit-graph &&\n> +\t\techo \"$(pwd)\"/../.git/objects >.git/objects/info/alternates &&\n> +\t\ttest_commit 18 &&\n> +\t\tgit commit-graph write --reachable --split --changed-paths &&\n> +\t\ttest_line_count = 1 $graphdir/commit-graph-chain\n> +\t) &&\n> +\ttest_cmp commit-graph .git/objects/info/commit-graph\n> +'\n> +\n> +test_done\n"},{"id":"389542","messageId":"86r208v3zs.fsf@gmail.com","threadId":"52499","inReplyTo":"1e2acb37ad710cb0d1c09ed163fdd4473a27335c.1576879520.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 7/9] commit-graph: reuse existing bloom filters during write.","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-09T19:12:07Z","receivedAt":"2020-01-09T19:12:22Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\nOverly long lines in the commit message.\n\n> Read previously computed bloom filters from the commit-graph file if possible\n> to avoid recomputing during commit-graph write.\n\nHmmm.  This fixes (somewhat) the problem that I have noticed in previous\npatch that the Bloom filter was computed at least twice, once for BIDX\nchunk, once fo BDAT chunk.\n\nI think the order should be:\n - use Bloom filter on slab, if present\n - fill it from commit graph, if saved there\n - if needed, compute it from scratch (expensive operation!)\n\nIf I understand it correctly, it now does it... though possibly with\nunnecessary memory allocation if commit-graph file does not include\nBloomm filters data, and (re)computing is not requested (see later).\n\nBut I might be wrong here.\n\n>\n> Reading from the commit-graph is based on the format in which bloom filters are\n> written in the commit graph file. See method `fill_filter_from_graph` in bloom.c\n\nThis description reads a bit strange; it looks like it states a truism\n(we read in the format we wrote).  It think it should be rephrased in\ndifferent way for better readability.\n\n>\n> For reading the bloom filter for commit at lexicographic position i:\n\nI think it would better read as:\n\n  To read Bloom filter for a given commit with lexicographic position\n  'i' we need to:\n\n> 1. Read BIDX[i] which essentially gives us the starting index in BDAT for filter\n>    of commit i+1 (called the next_index in the code)\n                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n                    |\n                    \\- I think would be not needed with better var name\n\nI would also add that it gives the position [one past] the end of Bloom\nfilter data for i-th commit.\n\n>\n> 2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT for\n>    filter of commit i (called the prev_index in the code)\n\nMinor nitpick: Full stops are missing.\n\nWhy it is called prev_index and next_index, while it is either\ncurr_index and next_index or prev_index and curr_index, or maybe even\nbetter beg_index and end_index?\n\n>    For i = 0, prev_index will be 0. The first lexicographic commit's filter will\n>    start at BDAT.\n\nI would state it\n\n     For first commit, with i = 0, Bloom filter data starts at the\n     beginning, just past the header in BDAT chunk.\n\n>\n> 3. The length of the filter will be next_index - prev_index, because BIDX[i]\n>    gives the cumulative 8-byte words including the ith commit's filter.\n>\n> We toggle whether bloom filters should be recomputed based on the compute_if_null\n> flag.\n\nAll right.\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  bloom.c        | 40 ++++++++++++++++++++++++++++++++++++++--\n>  bloom.h        |  3 ++-\n>  commit-graph.c |  6 +++---\n>  3 files changed, 43 insertions(+), 6 deletions(-)\n>\n> diff --git a/bloom.c b/bloom.c\n> index 08328cc381..86b1005802 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -1,5 +1,7 @@\n>  #include \"git-compat-util.h\"\n>  #include \"bloom.h\"\n> +#include \"commit.h\"\n> +#include \"commit-slab.h\"\n>  #include \"commit-graph.h\"\n>  #include \"object-store.h\"\n>  #include \"diff.h\"\n> @@ -119,13 +121,35 @@ static void add_key_to_filter(struct bloom_key *key,\n>  \t}\n>  }\n>  \n> +static void fill_filter_from_graph(struct commit_graph *g,\n> +\t\t\t\t   struct bloom_filter *filter,\n> +\t\t\t\t   struct commit *c)\n> +{\n> +\tuint32_t lex_pos, prev_index, next_index;\n> +\n\n> +\twhile (c->graph_pos < g->num_commits_in_base)\n> +\t\tg = g->base_graph;\n> +\n> +\tlex_pos = c->graph_pos - g->num_commits_in_base;\n\nThis part shares common code with load_oid_from_graph(), only without\nsome of error checking; perhaps it might be good to extract it into a\nseparate helper function, e.g. `lex_index(&g, c->graph_pos)`.\n\nMinor nitpick about the consistency of function names: why\nload_oid_from_graph(), but fill_filter_from_graph(), and not\nload_filter_from_graph() / load_bloom_from_graph()?\n\n> +\n> +\tnext_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n> +\tif (lex_pos)\n\nI think using\n\n  +\tif (lex_pos > 0)\n\nor\n\n  +\tif (lex_pos >= 0)\n\nmight be easier to reason about.\n\n> +\t\tprev_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n> +\telse\n> +\t\tprev_index = 0;\n> +\n> +\tfilter->len = next_index - prev_index;\n\nThe above command reads a bit strange: next - prev?  not  next - curr,\nor curr - prev?  Wouldn't it be better to name it begin_index and\nend_index, or beg_index and end_index for brevity?\n\n> +\tfilter->data = (uint64_t *)(g->chunk_bloom_data + 8 * prev_index + 12);\n\nPlease do not use magic constants; use instead something like:\n\n  +\tfilter->data = (uint64_t *)(g->chunk_bloom_data +\n  +\t\t\t\t    sizeof(uint64_t) * prev_index +\n  +\t\t\t\t    BLOOMDATA_CHUNK_HEADER_SIZE);\n\nPerhaps using `3*sizeof(unit32_t)` instead of magic value 12 would be\nenough; but having symbolic name for BDAT chunk header size is better, I\nthink.\n\n> +}\n> +\n>  void load_bloom_filters(void)\n>  {\n>  \tinit_bloom_filter_slab(&bloom_filters);\n>  }\n>  \n>  struct bloom_filter *get_bloom_filter(struct repository *r,\n> -\t\t\t\t      struct commit *c)\n> +\t\t\t\t      struct commit *c,\n> +\t\t\t\t      int compute_if_null)\n\nI'm not sure about `compute_if_null` name...\n\n>  {\n>  \tstruct bloom_filter *filter;\n>  \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n> @@ -134,6 +158,18 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \tconst char *revs_argv[] = {NULL, \"HEAD\", NULL};\n>  \n>  \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n\nNote that the documentation for `slab_at(slab, commit)` is documented in\n`commit-slab.h` as\n\n *   This function locates the data associated with the given commit in\n *   the indegree slab, and returns the pointer to it.  The location to\n *   store the data is allocated as necessary.\n                       ~~~~~~~~~~~~~~~~~~~~~~\n\nShould we worry about this possibly unnecessary allocation (if there is\nno Bloom filter chunk in the commit-graph, and we are not recomputing\nit)?\n\nThere is `slab_peek(slab_commit)` with the following properties:\n\n *   This function is similar to indegree_at(), but it will return NULL\n *   until a call to indegree_at() was made for the commit.\n\n> +\n> +\tif (!filter->data) {\n\nWhat does the Bloom filter for a commit with no changes looks like?\nWhat about for a commit with more than 512 changes?  Do, in either of\nthose cases, filter->len is 0?  If yes, what about filter->data?\n\n> +\t\tload_commit_graph_info(r, c);\n> +\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH && r->objects->commit_graph->chunk_bloom_indexes) {\n\nPlease wrap overly long lines (109 characters seems too long; the\nCodingGuidelines states:\n\n - We try to keep to at most 80 characters per line.\n\n  +\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH &&\n  +\t\t    r->objects->commit_graph->chunk_bloom_indexes) {\n\nYou seem to assume here that in the chain of commit-graph files either\nall of them would have Bloom filters, or all of them would be missing\nBloom filters.  Isn't it however possible for only some of the\ncommit-graph files in chain to include Bloom filter data chunks?\n\nIn such case, the top commit-graph file may have BIDX chunk\n(bloom_indexes), but the commit-graph with the commit 'c' might not have\nit.  Or the top commit-graph file may be missing BIDX chunk, so Git\nwould recompute it even if the commit-graph file for commit 'c' includes\nit.\n\nIf such situation is forbidden, how the restriction is managed?\n\nNote: in any case, this needs to be tested!\n\n> +\t\t\tfill_filter_from_graph(r->objects->commit_graph, filter, c);\n> +\t\t\treturn filter;\n> +\t\t}\n> +\t}\n\nAll right, if it is not in slab, we try to read if from the commit\ngraph.  Looks all right.\n\n> +\n> +\tif (filter->data || !compute_if_null)\n> +\t\t\treturn filter;\n                ^^^^^^^^\n                 |\n                 \\- one tab too many\n\nIf we have found existing filter (on slab or in the commit-graph), or if\nwe won't be recomputing it, return it.  O.K.\n\n> +\n>  \tinit_revisions(&revs, NULL);\n>  \trevs.diffopt.flags.recursive = 1;\n>  \n> @@ -198,4 +234,4 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n>  \n>  \treturn filter;\n> -}\n> \\ No newline at end of file\n> +}\n\nAccidental change.\n\n> diff --git a/bloom.h b/bloom.h\n> index ba8ae70b67..101d689bbd 100644\n> --- a/bloom.h\n> +++ b/bloom.h\n> @@ -36,7 +36,8 @@ struct bloom_key {\n>  void load_bloom_filters(void);\n>  \n>  struct bloom_filter *get_bloom_filter(struct repository *r,\n> -\t\t\t\t      struct commit *c);\n> +\t\t\t\t      struct commit *c,\n> +\t\t\t\t      int compute_if_null);\n>\n\nAll right, this is just update of the function signature.\n\n>  void fill_bloom_key(const char *data,\n>  \t\t    int len,\n> diff --git a/commit-graph.c b/commit-graph.c\n> index def2ade166..0580ce75d5 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -1032,7 +1032,7 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n>  \tuint32_t cur_pos = 0;\n>  \n>  \twhile (list < last) {\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n>  \t\tcur_pos += filter->len;\n>  \t\thashwrite_be32(f, cur_pos);\n>  \t\tlist++;\n> @@ -1051,7 +1051,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n>  \thashwrite_be32(f, settings->bits_per_entry);\n>  \n>  \twhile (first < last) {\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *first);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *first, 0);\n>  \t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n>  \t\tfirst++;\n>  \t}\n\nO.K., so those two do not compute Bloom filters, but they are called\nfrom write_commit_graph_file(), which in turn is called in\nwrite_commit_graph() *after* running compute_bloom_filters().\n\nLooks good to me, then.\n\n> @@ -1218,7 +1218,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>  \n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>  \t\tstruct commit *c = ctx->commits.list[i];\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n>  \t\tctx->total_bloom_filter_size += sizeof(uint64_t) * filter->len;\n>  \t\tdisplay_progress(progress, i + 1);\n>  \t}\n\nAll right, so compute_bloom_filters() ensures that it is actually\ncomputed (if needed).\n\nLooks good to me.\n\nRegards,\n-- \nJakub Narębski\n"},{"id":"389602","messageId":"86d0bqsuqc.fsf@gmail.com","threadId":"52499","inReplyTo":"72a2bbf6765a1e99a3a23372f801099a07fe11a5.1576879520.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 8/9] revision.c: use bloom filters to speed up path based revision walks","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-11T00:27:23Z","receivedAt":"2020-01-11T00:27:37Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> If bloom filters have been written to the commit-graph file, revision walk will\n> use them to speed up revision walks for a particular path.\n\nI'd propose the following change, to make structure of the above\nsentence simpler:\n\n  Revision walk will now use Bloom filters for commits to speed up\n  revision walks for a particular path (for computing history of a\n  path), if they are present in the commit-graph file.\n\n> Note: The current implementation does this in the case of single pathspec\n> case only.\n>\n> We load the bloom filters during the prepare_revision_walk step when dealing\n\nI think it would flow better with the following change:\n\ns/when dealing/, but only when dealing/\n\n> with a single pathspec. While comparing trees in rev_compare_trees(), if the\n> bloom filter says that the file is not different between the two trees, we\n> don't need to compute the expensive diff. This is where we get our performance\n> gains.\n\nMaybe we should also add:\n\n  The other answer we can get from the Bloom filter is \"maybe\".\n\n>\n> Performance Gains:\n> We tested the performance of `git log --path` on the git repo, the linux and\n> some internal large repos, with a variety of paths of varying depths.\n\nI think you meant `git log <path>`, not `git log --path`.\n\n>\n> On the git and linux repos:\n> we observed a 2x to 5x speed up.\n\nIt would be good, I think, to have some specific numbers: starting from\nthis version, for this path, with and without Bloom filters it takes so\nlong (and the improvement in %).\n\nWhile at it, it might be good idea to provide _costs_: how much\nadditional space on disk Bloom filters take for those specific examples,\nand how much extra time (and memory) it takes to compute Bloom filters.\n\nWith actual specific numbers we can estimate when it would start to be\nworth it to create Bloom filter data...\n\n>\n> On a large internal repo with files seated 6-10 levels deep in the tree:\n> we observed 10x to 20x speed ups, with some paths going up to 28 times\n> faster.\n\nIt would be nice to see specific numbers, if showing pathnames is\npossible.  In any case it would be good to have more information: what\npaths give 10x, what give 20x, and what kinds give 28x speedup (what\nis path depth, how many objects, etc.).\n\n>\n> RFC Notes:\n> I plan to collect the folloowing statistics around this usage of bloom filters\n> and trace them out using trace2.\n> - number of bloom filter queries,\n> - number of \"No\" responses (file hasn't changed)\n> - number of \"Maybe\" responses (file may have changed)\n> - number of \"Commit not parsed\" cases (commit had too many changes to have a\n>   bloom filter written out, currently our limit is 512 diffs)\n\nPerhaps also:\n  - histogram of bloom filter sizes in 64-bit blocks\n    (which is rough histogram of number of changes per commit)\n\nThough I think all those statistics are a bit specific to the\nrepository, and how you use Git.\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com\n> Helped-by: SZEDER Gábor <szeder.dev@gmail.com>\n> Helped-by: Jonathan Tan <jonathantanmy@google.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  bloom.c              | 20 ++++++++++++\n>  bloom.h              |  4 +++\n>  revision.c           | 67 +++++++++++++++++++++++++++++++++++++--\n>  revision.h           |  5 +++\n>  t/t4216-log-bloom.sh | 74 ++++++++++++++++++++++++++++++++++++++++++++\n>  5 files changed, 168 insertions(+), 2 deletions(-)\n>  create mode 100755 t/t4216-log-bloom.sh\n>\n> diff --git a/bloom.c b/bloom.c\n> index 86b1005802..0c7505d3d6 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -235,3 +235,23 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \n>  \treturn filter;\n>  }\n> +\n> +int bloom_filter_contains(struct bloom_filter *filter,\n> +\t\t\t  struct bloom_key *key,\n> +\t\t\t  struct bloom_filter_settings *settings)\n> +{\n> +\tint i;\n> +\tuint64_t mod = filter->len * BITS_PER_BLOCK;\n> +\n> +\tif (!mod)\n> +\t\treturn 1;\n\nAh, so filter->len equal to zero denotes too many changes for a Bloom\nfilter, and this conditional is a short-circuit test: always return\n\"maybe\".\n\nI wonder if we should explicitly check for filter->len = 1 and\nfilter->data[0] = 0, which should be empty Bloom filter -- for commit\nwith no changes with respect to first parent, and short-circuit\nreturning 0 (no file would ever belong).\n\n> +\n> +\tfor (i = 0; i < settings->num_hashes; i++) {\n> +\t\tuint64_t hash_mod = key->hashes[i] % mod;\n> +\t\tuint64_t block_pos = hash_mod / BITS_PER_BLOCK;\n\nI have seen this code before... ;-)  The add_key_to_filter() function\nincludes almost identical code, but I am not sure if it is feasible to\neliminate this (slight) code duplication.  Probably not.\n\n> +\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n> +\t\t\treturn 0;\n\nAll right, if at least one hash function does not match, return \"no\"...\n\n> +\t}\n> +\n> +\treturn 1;\n\n... else return \"maybe\".\n\n> +}\n\nI am wondering however if this code should not be moved earlier in the\nseries, so that commit 2/9 in series actually adds fully functional\n[semi-generic] Bloom filter implementation.\n\n> diff --git a/bloom.h b/bloom.h\n> index 101d689bbd..9bdacd0a8e 100644\n> --- a/bloom.h\n> +++ b/bloom.h\n> @@ -44,4 +44,8 @@ void fill_bloom_key(const char *data,\n>  \t\t    struct bloom_key *key,\n>  \t\t    struct bloom_filter_settings *settings);\n>  \n> +int bloom_filter_contains(struct bloom_filter *filter,\n> +\t\t\t  struct bloom_key *key,\n> +\t\t\t  struct bloom_filter_settings *settings);\n> +\n>  #endif\n> diff --git a/revision.c b/revision.c\n> index 39a25e7a5d..01f5330740 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -29,6 +29,7 @@\n>  #include \"prio-queue.h\"\n>  #include \"hashmap.h\"\n>  #include \"utf8.h\"\n> +#include \"bloom.h\"\n>  \n>  volatile show_early_output_fn_t show_early_output;\n>  \n> @@ -624,11 +625,34 @@ static void file_change(struct diff_options *options,\n>  \toptions->flags.has_changes = 1;\n>  }\n>  \n> +static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n> +\t\t\t\t\t\t struct commit *commit,\n> +\t\t\t\t\t\t struct bloom_key *key,\n> +\t\t\t\t\t\t struct bloom_filter_settings *settings)\n\nAll right, this function name is certainly descriptive.  I wonder if it\nwouldn't be better to use a shorter name, like maybe_different(), or\nmaybe_not_treesame(), so that the conditional using this function would\nread naturally:\n\n  if (!maybe_different(revs, commit, revs->bloom_key,\n  \t\t       revs->bloom_filter_settings))\n  \treturn REV_TREE_SAME;\n\nBut this might be just a matter of taste.\n\n> +{\n> +\tstruct bloom_filter *filter;\n> +\n> +\tif (!revs->repo->objects->commit_graph)\n> +\t\treturn -1;\n> +\tif (commit->generation == GENERATION_NUMBER_INFINITY)\n> +\t\treturn -1;\n\nO.K., so we check that there is loaded commit graph, and that given\ncommit is in a commit graph, otherwise we return \"no data\".\n\nI agree with distinguishing between \"no data\" value (which is <0, a\nconvention for denoting errors), and \"maybe\" value.  Currently this\ndistinction is not utilized, but it would help in the case of more than\none path given -- in \"no data\" case there is no need to check other\npaths against non-existing Bloom filter.\n\n> +\tif (!key || !settings)\n> +\t\treturn -1;\n\nWhy the check for non-NULL 'key' and 'settings' is the last check?\n\nBTW. when it is possible for key or settings to be NULL?  Or is it just\ndefensive programming?\n\n> +\n> +\tfilter = get_bloom_filter(revs->repo, commit, 0);\n\nO.K., we won't be recomputing Bloom filter if it is not present either\non slab (as \"inside-out\" auxiliary data for a commit), or in the\ncommit-graph (in Bloom filter chunks).\n\n> +\n> +\tif (!filter || !filter->len)\n> +\t\treturn 1;\n\nSidenote: bloom_filter_contains() would also indirectly check for\nfilter->len being zero.  Though this doesn't cost much.\n\nShouldn't !filter case return value of -1 i.e. \"no data\", rather than\nreturn value of 1 i.e. \"maybe\"?\n\n> +\n> +\treturn bloom_filter_contains(filter, key, settings);\n> +}\n\nAll right.\n\n> +\n>  static int rev_compare_tree(struct rev_info *revs,\n> -\t\t\t    struct commit *parent, struct commit *commit)\n> +\t\t\t    struct commit *parent, struct commit *commit, int nth_parent)\n>  {\n>  \tstruct tree *t1 = get_commit_tree(parent);\n>  \tstruct tree *t2 = get_commit_tree(commit);\n> +\tint bloom_ret = 1;\n>  \n>  \tif (!t1)\n>  \t\treturn REV_TREE_NEW;\n> @@ -653,6 +677,16 @@ static int rev_compare_tree(struct rev_info *revs,\n>  \t\t\treturn REV_TREE_SAME;\n>  \t}\n>  \n> +\tif (revs->pruning.pathspec.nr == 1 && !nth_parent) {\n\nAll right, Bloom filter stores information about changed paths with\nrespect to first-parent changes only, so if we are asking about is not a\nfirst parent (where nth_parent == 0), we cannot use Bloom filter.\n\nCurrently we limit the check to single pathspec only; that is a good\nstart and a good simplification.\n\n> +\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs,\n> +\t\t\t\t\t\t\t\t  commit,\n> +\t\t\t\t\t\t\t\t  revs->bloom_key,\n> +\t\t\t\t\t\t\t\t  revs->bloom_filter_settings);\n> +\n> +\t\tif (bloom_ret == 0)\n> +\t\t\treturn REV_TREE_SAME;\n\nPretty straightforward.\n\n> +\t}\n> +\n>  \ttree_difference = REV_TREE_SAME;\n>  \trevs->pruning.flags.has_changes = 0;\n>  \tif (diff_tree_oid(&t1->object.oid, &t2->object.oid, \"\",\n> @@ -855,7 +889,7 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit)\n>  \t\t\tdie(\"cannot simplify commit %s (because of %s)\",\n>  \t\t\t    oid_to_hex(&commit->object.oid),\n>  \t\t\t    oid_to_hex(&p->object.oid));\n> -\t\tswitch (rev_compare_tree(revs, p, commit)) {\n> +\t\tswitch (rev_compare_tree(revs, p, commit, nth_parent)) {\n\nAll right, we need to pass information about the index of the parent;\nand we have just done that.  Good.\n\n>  \t\tcase REV_TREE_SAME:\n>  \t\t\tif (!revs->simplify_history || !relevant_commit(p)) {\n>  \t\t\t\t/* Even if a merge with an uninteresting\n> @@ -3342,6 +3376,33 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n>  \t}\n>  }\n>  \n> +static void prepare_to_use_bloom_filter(struct rev_info *revs)\n\nAll right, I see that pointers to bloom_key and bloom_filter_settings\nwere added to the rev_info struct.  I understand why the former is here,\nbut the latter seems to be there just as a shortcut (to not owned data),\nwhich is fine but a bit strange.\n\nOr is the latter here to allow for Bloom filter settings to possibly\nchange from commit-graph file in the chain to commit-graph file, and\nthus from commit to commit?\n\n> +{\n> +\tstruct pathspec_item *pi;\n\nMaybe 'pathspec' (or the like) instead of short and cryptic 'pi' would\nbe better a better name... unless 'pi' is used in other places already.\n\n> +\tconst char *path;\n> +\tsize_t len;\n> +\n> +\tif (!revs->commits)\n> +\t    return;\n\nWhen revs->commits may be NULL?  I understand that we need to have this\ncheck because we use revs->commits->item next (sidenote: can revs ever\nbe NULL?).\n\nWould `git log --all -- <path>` use Bloom filters (as it theoretically\ncould)?\n\n> +\n> +\tparse_commit(revs->commits->item);\n\nWhy parsing first commit on the list of starting commits is needed here?\nPlease help me understand this line.\n\nAnd shouldn't we use repo_parse_commit() here?\n\n> +\n> +\tif (!revs->repo->objects->commit_graph)\n> +\t\treturn;\n> +\n> +\trevs->bloom_filter_settings = revs->repo->objects->commit_graph->settings;\n> +\tif (!revs->bloom_filter_settings)\n> +\t\treturn;\n\nAll right, so if there is no commit graph, or the commit graph does not\ninclude Bloom filter data, there is nothing to do.\n\nThough I worry that it would make Git do not use Bloom filter if the top\ncommit-graph in the chain does not include Bloom filter data, while\nother commit-graph files do (and Git could have used that information to\nspeed up the file history query).\n\n> +\n> +\tpi = &revs->pruning.pathspec.items[0];\n> +\tpath = pi->match;\n> +\tlen = strlen(path);\n\n\nWhy not the following, if we do not do any checks for 'pi' value:\n\n  +\tpath = &revs->pruning.pathspec.items[0]->match;\n\nA question: is the path in the `match` field in `struct pathspec`\nnormalized with respect to trailing slash (for directories)?  Bloom\nfilter stores pathnames for directories without trailing slash.\n\nWhat I mean is if, for example, both of those would use Bloom filter\ndata:\n\n  $ git log -- Documentation/\n  $ git log -- Documentation\n\n> +\n> +\tload_bloom_filters();\n> +\trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n> +\tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n\nAll right, looks good.\n\nThough... do we leak revs->bloom_key, as should we worry about it?\n\n> +}\n> +\n>  int prepare_revision_walk(struct rev_info *revs)\n>  {\n>  \tint i;\n> @@ -3391,6 +3452,8 @@ int prepare_revision_walk(struct rev_info *revs)\n>  \t\tsimplify_merges(revs);\n>  \tif (revs->children.name)\n>  \t\tset_children(revs);\n> +\tif (revs->pruning.pathspec.nr == 1)\n> +\t    prepare_to_use_bloom_filter(revs);\n\nLooks good.\n\nMinor nitpick: 4 spaces instead of tab are used for indentation (or, to\nbe more exact a tab followed by 4 space, instead of two tabs):\n\n  +\tif (revs->pruning.pathspec.nr == 1)\n  +\t\tprepare_to_use_bloom_filter(revs);\n\n>  \treturn 0;\n>  }\n>  \n> diff --git a/revision.h b/revision.h\n> index a1a804bd3d..65dc11e8f1 100644\n> --- a/revision.h\n> +++ b/revision.h\n> @@ -56,6 +56,8 @@ struct repository;\n>  struct rev_info;\n>  struct string_list;\n>  struct saved_parents;\n> +struct bloom_key;\n> +struct bloom_filter_settings;\n>  define_shared_commit_slab(revision_sources, char *);\n>  \n>  struct rev_cmdline_info {\n> @@ -291,6 +293,9 @@ struct rev_info {\n>  \tstruct revision_sources *sources;\n>  \n>  \tstruct topo_walk_info *topo_walk_info;\n> +\n> +\tstruct bloom_key *bloom_key;\n> +\tstruct bloom_filter_settings *bloom_filter_settings;\n\nIt might be a good idea to add one-line comment above those newly\nintroduced fields.  The `struct rev_info` has many subsections of fields\ndescribed by such comments (or even block comments), like e.g.\n\n  /* Starting list */\n  /* Parents of shown commits */\n  /* The end-points specified by the end user */\n  /* excluding from --branches, --refs, etc. expansion */\n  /* Traversal flags */\n  /* diff info for patches and for paths limiting */\n\n  /*\n   * Whether the arguments parsed by setup_revisions() included any\n   * \"input\" revisions that might still have yielded an empty pending\n   * list (e.g., patterns like \"--all\" or \"--glob\").\n   */\n\n>  };\n>  \n>  int ref_excluded(struct string_list *, const char *path);\n> diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> new file mode 100755\n> index 0000000000..d42f077998\n> --- /dev/null\n> +++ b/t/t4216-log-bloom.sh\n> @@ -0,0 +1,74 @@\n> +#!/bin/sh\n> +\n> +test_description='git log for a path with bloom filters'\n> +. ./test-lib.sh\n> +\n> +test_expect_success 'setup repo' '\n> +\tgit init &&\n> +\tgit config core.commitGraph true &&\n> +\tgit config gc.writeCommitGraph false &&\n> +\tinfodir=\".git/objects/info\" &&\n> +\tgraphdir=\"$infodir/commit-graphs\" &&\n> +\ttest_oid_init\n\nWhy do you use `test_oid_init`, when you are *not* using `test_oid`?\nI guess it is because t5318-commit-graph.sh uses it, isn't it?\n\nThe `graphdir` shell variable is not used either.\n\n> +'\n> +\n> +test_expect_success 'create 9 commits and repack' '\n> +\ttest_commit c1 file1 &&\n> +\ttest_commit c2 file2 &&\n> +\ttest_commit c3 file3 &&\n> +\ttest_commit c4 file1 &&\n> +\ttest_commit c5 file2 &&\n> +\ttest_commit c6 file3 &&\n> +\ttest_commit c7 file1 &&\n> +\ttest_commit c8 file2 &&\n> +\ttest_commit c9 file3\n> +'\n\nWouldn't it be better for this step to be a part of 'setup repo' step?\nAnyway, the test name says '... and repack', but `git repack` is missing\n(it should be done after last test_commit).\n\nI think it would be good idea to test behavior of Bloom filters with\nrespect to directories, so at least one file should be in a subdirectory\n(maybe even deeper in hierarchy).\n\nWe should also test the behavior with respect to merges, and when we are\nnot following the first parent.  But that might be a separate part of\nthis test.\n\n> +\n> +printf \"c7\\nc4\\nc1\" > expect_file1\n\nDoing things outside test is discouraged.  We can create a separate test\nthat creates those expect_file* files, or it can be a part of 'create\ncommits' test.\n\nAnyway, instead of doing test without Bloom filters (something that\nshould have been tested already by other parts of testsuite), and then\ndoing the same test with Bloom filter, why not compare that the result\nwithout and with Bloom filter is the same.  The t5318-commit-graph.sh\ntest does this with help of graph_git_two_modes() function:\n\n  graph_git_two_modes () {\n  \tgit -c core.commitGraph=true  $1 >output\n  \tgit -c core.commitGraph=false $1 >expect\n  \ttest_cmp expect output\n  }\n\nSidenote: I wonder if it is high time to create t/lib-commit-graph.sh\nhelper with, among others, this common function.\n\n> +\n> +test_expect_success 'log without bloom filters' '\n> +\tgit log --pretty=\"format:%s\"  -- file1 > actual &&\n\nCodingGuidelines:57: Redirection operators should be written with space\nbefore, but no space after them.  (Minor nitpick)\n\n  +\tgit log --pretty=\"format:%s\"  -- file1 >actual &&\n\n> +\ttest_cmp expect_file1 actual\n> +'\n> +\n> +printf \"c8\\nc7\\nc5\\nc4\\nc2\\nc1\" > expect_file1_file2\n\n  +printf \"c8\\nc7\\nc5\\nc4\\nc2\\nc1\" >expect_file1_file2\n\n> +\n> +test_expect_success 'multi-path log without bloom filters' '\n> +\tgit log --pretty=\"format:%s\"  -- file1 file2 > actual &&\n\n  +\tgit log --pretty=\"format:%s\"  -- file1 file2 >actual &&\n\n> +\ttest_cmp expect_file1_file2 actual\n> +'\n> +\n> +graph_read_expect() {\n\nCodingGuidelines:144: We prefer a space between the function name and\nthe parentheses, and no space inside the parentheses.  (Minor nitpick)\n\n  +graph_read_expect () {\n\n> +\tOPTIONAL=\"\"\n> +\tNUM_CHUNKS=5\n> +\tif test ! -z $2\n> +\tthen\n> +\t\tOPTIONAL=\" $2\"\n> +\t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n\nThis should be either\n\n  +\t\tNUM_CHUNKS=$((5 + $(echo \"$2\" | wc -w)))\n\nor more future proof\n\n  +\t\tNUM_CHUNKS=$(($NUM_CHUNKS + $(echo \"$2\" | wc -w)))\n\nWe got away with this bug because we were not using octopus merges, and\nthere were no optional core chunks, that is we never call\ngraph_read_expect with second parameter in this test.\n\n> +\tfi\n> +\tcat >expect <<- EOF\n\nWhy there is space between \"<<-\" and \"EOF\"?\n\n> +\theader: 43475048 1 1 $NUM_CHUNKS 0\n> +\tnum_commits: $1\n> +\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data$OPTIONAL\n> +\tEOF\n> +\ttest-tool read-graph >output &&\n> +\ttest_cmp expect output\n> +}\n> +\n> +test_expect_success 'write commit graph with bloom filters' '\n> +\tgit commit-graph write --reachable --changed-paths &&\n> +\ttest_path_is_file $infodir/commit-graph &&\n> +\tgraph_read_expect \"9\"\n> +'\n\nAll right, this is preparatory step for further tests.\n\n> +\n> +test_expect_success 'log using bloom filters' '\n> +\tgit log --pretty=\"format:%s\" -- file1 > actual &&\n> +\ttest_cmp expect_file1 actual\n> +'\n> +\n> +test_expect_success 'multi-path log using bloom filters' '\n> +\tgit log --pretty=\"format:%s\"  -- file1 file2 > actual &&\n> +\ttest_cmp expect_file1_file2 actual\n> +'\n\nWith graph_git_two_modes() it would be much simpler:\n\n  +test_expect_success 'single-path log' '\n  +\tgit_graph_two_modes \"git log --pretty=format:%s -- file1\"\n  +'\n\n  +test_expect_success 'multi-path log' '\n  +\tgit_graph_two_modes \"git log --pretty=format:%s -- file1 file2\"\n  +'\n\nWhat is missing is:\n1. checking that Git is actually _using_ Bloom filters\n   (which might be difficult to do)\n2. testing that Bloom filters work also for history of subdirectories\n   e.g. \"git log -- subdir/\" and \"git log -- subdir\"; this would of\n   course require adjusting setup step\n3. testing specific behaviors, like \"git log --all -- file1\"\n4. merges with history following second parent\n5. commits with no changes and/or merges with no first-parent changes\n6. commit with more than 512 changed files (marked as slow test,\n   and perhaps created with fast-import interface, like bulk commit\n   creation in test_commit_bulk)\n\n> +\n> +test_done\n\nBest regards,\n-- \nJakub Narębski\n"},{"id":"389617","messageId":"86v9phrcml.fsf@gmail.com","threadId":"52499","inReplyTo":"e1c315d0a766af147eb4ead41a172f724e90cc34.1576879520.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 9/9] commit-graph: add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-11T19:56:02Z","receivedAt":"2020-01-11T19:56:13Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag to the test setup suite in\n> order to toggle writing bloom filters when running any of the git tests. If set\n> to true, we will compute and write bloom filters every time a test calls\n> `git commit-graph write`.\n\nOK, so it works in addition to GIT_TEST_COMMIT_GRAPH.\n\n>\n> The test suite passes when GIT_TEST_COMMIT_GRAPH and\n> GIT_COMMIT_GRAPH_BLOOM_FILTERS are enabled.\n\nGood.  Very good.\n\nNo errors found by Continuous Integration setup either?\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  builtin/commit-graph.c        | 2 +-\n>  ci/run-build-and-tests.sh     | 1 +\n>  commit-graph.h                | 1 +\n>  t/README                      | 3 +++\n>  t/t4216-log-bloom.sh          | 3 +++\n>  t/t5318-commit-graph.sh       | 2 ++\n>  t/t5324-split-commit-graph.sh | 1 +\n>  t/t5325-commit-graph-bloom.sh | 3 +++\n>  8 files changed, 15 insertions(+), 1 deletion(-)\n>\n> diff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\n> index 9bd1e11161..97167959b2 100644\n> --- a/builtin/commit-graph.c\n> +++ b/builtin/commit-graph.c\n> @@ -146,7 +146,7 @@ static int graph_write(int argc, const char **argv)\n>  \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n>  \tif (opts.progress)\n>  \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n> -\tif (opts.enable_bloom_filters)\n> +\tif (opts.enable_bloom_filters || git_env_bool(GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS, 0))\n>  \t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n>\n\nVery minor nitpick: not to make this line long, I would break it at the\nboolean operator, that is write\n\n  -\tif (opts.enable_bloom_filters)\n  +\tif (opts.enable_bloom_filters ||\n  +\t    git_env_bool(GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS, 0))  \n   \t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n\nI agree that this is a good place to put this check, by pretending that\n`--changed-paths` option was given on command line.  Looks good.\n\n>  \tread_replace_refs = 0;\n> diff --git a/ci/run-build-and-tests.sh b/ci/run-build-and-tests.sh\n> index ff0ef7f08e..19d0846d34 100755\n> --- a/ci/run-build-and-tests.sh\n> +++ b/ci/run-build-and-tests.sh\n> @@ -19,6 +19,7 @@ linux-gcc)\n>  \texport GIT_TEST_OE_SIZE=10\n>  \texport GIT_TEST_OE_DELTA_SIZE=5\n>  \texport GIT_TEST_COMMIT_GRAPH=1\n> +\texport GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=1\n>  \texport GIT_TEST_MULTI_PACK_INDEX=1\n>  \tmake test\n>  \t;;\n\nAll right, adding this to CI would certainly exercise this feature.\n\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 2202ad91ae..d914e6abf1 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -8,6 +8,7 @@\n>  \n>  #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n>  #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n> +#define GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS \"GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS\"\n>\n\nAll right (the ordering is a mater of taste).\n\n>  struct commit;\n>  struct bloom_filter_settings;\n> diff --git a/t/README b/t/README\n> index caa125ba9a..399b190437 100644\n> --- a/t/README\n> +++ b/t/README\n> @@ -378,6 +378,9 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n>  be written after every 'git commit' command, and overrides the\n>  'core.commitGraph' setting to true.\n>  \n> +GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=<boolean>, when true, forces commit-graph\n> +write to compute and write bloom filters for every 'git commit-graph write'\n> +\n\nThanks for documenting this.\n\nMissing full stop '.' at the end of the sentence.  (Minor nit).\n\nWe might want to add \", as if '--changed-paths' option was given.\", but\nit is not strictly necessary.\n\n>  GIT_TEST_FSMONITOR=$PWD/t7519/fsmonitor-all exercises the fsmonitor\n>  code path for utilizing a file system monitor to speed up detecting\n>  new or changed files.\n> diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> index d42f077998..0e092b387c 100755\n> --- a/t/t4216-log-bloom.sh\n> +++ b/t/t4216-log-bloom.sh\n> @@ -3,6 +3,9 @@\n>  test_description='git log for a path with bloom filters'\n>  . ./test-lib.sh\n>  \n> +GIT_TEST_COMMIT_GRAPH=0\n> +GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n> +\n\nOK, neither of those setting increase the coverage for this test.\n\nOn the other hand they won't make tests fail (or at least they\nshouldn't, I think).  The t5318-commit-graph.sh doesn't include \nGIT_TEST_COMMIT_GRAPH=0, after all.\n\n>  test_expect_success 'setup repo' '\n>  \tgit init &&\n>  \tgit config core.commitGraph true &&\n> diff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\n> index 3f03de6018..613228bb12 100755\n> --- a/t/t5318-commit-graph.sh\n> +++ b/t/t5318-commit-graph.sh\n> @@ -3,6 +3,8 @@\n>  test_description='commit graph'\n>  . ./test-lib.sh\n>  \n> +GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n> +\n\nAll right, adding Bloom filters to commit-graph file would make test\ncases utilizing graph_read_expect() fail.\n\n>  test_expect_success 'setup full repo' '\n>  \tmkdir full &&\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> index c24823431f..181ca7e0cb 100755\n> --- a/t/t5324-split-commit-graph.sh\n> +++ b/t/t5324-split-commit-graph.sh\n> @@ -4,6 +4,7 @@ test_description='split commit graph'\n>  . ./test-lib.sh\n>  \n>  GIT_TEST_COMMIT_GRAPH=0\n> +GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n\nAll right, same here: adding Bloom filters to commit-graph would make\ntest cases utilizing graph_read_expect() fail.\n\nSidenote: Here without GIT_TEST_COMMIT_GRAPH=0 tests that rely on\nprecise timing of writing commit-graph to create split commit-graph\nwould fail.\n\n>  \n>  test_expect_success 'setup repo' '\n>  \tgit init &&\n> diff --git a/t/t5325-commit-graph-bloom.sh b/t/t5325-commit-graph-bloom.sh\n> index d7ef0e7fb3..a9c9e9fef6 100755\n> --- a/t/t5325-commit-graph-bloom.sh\n> +++ b/t/t5325-commit-graph-bloom.sh\n> @@ -3,6 +3,9 @@\n>  test_description='commit graph with bloom filters'\n>  . ./test-lib.sh\n>  \n> +GIT_TEST_COMMIT_GRAPH=0\n> +GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n> +\n\nThis test also includes some split commit-graph test cases, so the above\nis necessary.\n\nAll right.\n\n>  test_expect_success 'setup repo' '\n>  \tgit init &&\n>  \tgit config core.commitGraph true &&\n\nBest,\n-- \nJakub Narębski\n"},{"id":"389716","messageId":"3aaf02fe-ac83-5694-2c69-e133879a0030@gmail.com","threadId":"52499","inReplyTo":"86d0c44f5s.fsf@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-01-13T16:54:32Z","receivedAt":"2020-01-13T16:54:36Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"On 12/31/2019 11:45 AM, Jakub Narebski wrote:\n> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>> Performance Gains: We tested the performance of 'git log -- <path>' on the git\n>> repo, the linux repo and some internal large repos, with a variety of paths\n>> of varying depths.\n>>\n>> On the git and linux repos: We observed a 2x to 5x speed up.\n>>\n>> On a large internal repo with files seated 6-10 levels deep in the tree: We\n>> observed 10x to 20x speed ups, with some paths going up to 28 times faster.\n> \n> Could you provide some more statistics about this internal repository,\n> such as number of files, number of commits, perhaps also number of all\n> objects?  Thanks in advance.\n> \n> I wonder why such large difference in performance 2-5x vs 10-20x.  Is it\n> about the depth of the file hierarchy?  How would the numbers look for\n> files seated closer to the root in the same large repository, like 3-5\n> levels deep in the tree?\n\nThe internal repository we saw these massive gains on has:\n- 413579 commits. \n- 183303 files distributed across 34482 folders\nThe size on disk is about 17 GiB. \n\nAnd yes, the difference is performance gains is mostly because of how \ndeep the files were in the hierarchy. How often a file has been touched\nalso makes a difference. The performance gains are less dramatic if the \nfile has a very sparse history even if it is a deep file. \n\nThe numbers from the git and linux repos for instance, are for files \ncloser to the root, hence 2x to 5x. \n\nThanks! \nGarima Singh\n"},{"id":"389723","messageId":"6cefadde-7171-e081-ba5b-ba2543d9f22f@gmail.com","threadId":"52499","inReplyTo":"86eewczapt.fsf@gmail.com","subject":"Re: [PATCH 2/9] commit-graph: write changed paths bloom filters","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-01-13T19:48:04Z","receivedAt":"2020-01-13T19:48:08Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 1/6/2020 1:44 PM, Jakub Narebski wrote:\n\n> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes: \n>> 3. The filters are sized according to the number of changes in the each commit,\n>>    with minimum size of one 64 bit word.\n> \n> Do I understand it correctly that the size of filter is 10*(number of\n> changed files) bits, rounded up to nearest multiple of 64?\n>\n\nYes.\n  \n>> +\n>> +struct pathmap_hash_entry {\n>> +    struct hashmap_entry entry;\n>> +    const char path[FLEX_ARRAY];\n>> +};\n> \n> Hmmm... I wonder why use hashmap and not string_list.  This is for\n> adding path with leading directories to the Bloom filter, isn't it?\n> \n\nYes. We do not want to repeat directories in the filter.\n\nThanks!\nGarima Singh\n\n"},{"id":"389746","messageId":"3b7d77a1-aed9-d202-8646-4b964cb965db@gmail.com","threadId":"52499","inReplyTo":"867e23xnkh.fsf@gmail.com","subject":"Re: [PATCH 5/9] commit-graph: write changed path bloom filters to commit-graph file.","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-01-14T15:14:30Z","receivedAt":"2020-01-14T15:14:33Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 1/7/2020 11:01 AM, Jakub Narebski wrote:\n> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>> +\n>> +\t\t\t\tif (hash_version != 1)\n>> +\t\t\t\t\tbreak;\n> \n> What does it mean for Git?  Behave as if there were no Bloom filter\n> data?\n>\n\nYes. We choose to use Bloom filters best effort, and in such cases, we will \njust fall back to the original code path. \n\n>>  \tif (ctx->num_commit_graphs_after > 1 &&\n>>  \t    write_graph_chunk_base(f, ctx)) {\n>>  \t\treturn -1;\n>> diff --git a/commit-graph.h b/commit-graph.h\n>> index 952a4b83be..2202ad91ae 100644\n>> --- a/commit-graph.h\n>> +++ b/commit-graph.h\n>> @@ -10,6 +10,7 @@\n>>  #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n>>  \n>>  struct commit;\n>> +struct bloom_filter_settings;\n>>  \n>>  char *get_commit_graph_filename(const char *obj_dir);\n>>  int open_commit_graph(const char *graph_file, int *fd, struct stat *st);\n>> @@ -58,6 +59,10 @@ struct commit_graph {\n>>  \tconst unsigned char *chunk_commit_data;\n>>  \tconst unsigned char *chunk_extra_edges;\n>>  \tconst unsigned char *chunk_base_graphs;\n>> +\tconst unsigned char *chunk_bloom_indexes;\n>> +\tconst unsigned char *chunk_bloom_data;\n>> +\n>> +\tstruct bloom_filter_settings *settings;\n> \n> Should this be part of `struct commit_graph`?  Shouldn't we free() this\n> data, or is it a pointer into xmmap-ped file... no it isn't -- we\n> xalloc() it, so we should free() it.\n> \n> I think it should be done in 'cleanup:' section of write_commit_graph(),\n> but I am not entirely sure.\n> \n\nThanks for calling this out! This is definitely a bug in how commit-graph.c\nfrees up the graph. The right way to free the graph would be to call\nfree_commit_graph() instead of free(graph) like many places in that file. \n\nCleaning up this entire pattern would be orthogonal to this series, so \nI will follow up with a separate series that cleans it up overall. \n\nFor now, I will free up `bloom_filter_settings` in free_commit_graph(). \n\nCheers! \nGarima Singh\n"},{"id":"389796","messageId":"706f7e8f-9211-36f0-1e83-cc98dbdd3ae1@gmail.com","threadId":"52499","inReplyTo":"86d0bqsuqc.fsf@gmail.com","subject":"Re: [PATCH 8/9] revision.c: use bloom filters to speed up path based revision walks","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-01-15T00:08:46Z","receivedAt":"2020-01-15T00:08:50Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 1/10/2020 7:27 PM, Jakub Narebski wrote:\n> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>  \t\tcase REV_TREE_SAME:\n>>  \t\t\tif (!revs->simplify_history || !relevant_commit(p)) {\n>>  \t\t\t\t/* Even if a merge with an uninteresting\n>> @@ -3342,6 +3376,33 @@ static void expand_topo_walk(struct rev_info *revs, struct commit *commit)\n>>  \t}\n>>  }\n>>  \n>> +static void prepare_to_use_bloom_filter(struct rev_info *revs)\n> \n> All right, I see that pointers to bloom_key and bloom_filter_settings\n> were added to the rev_info struct.  I understand why the former is here,\n> but the latter seems to be there just as a shortcut (to not owned data),\n> which is fine but a bit strange.\n> \n> Or is the latter here to allow for Bloom filter settings to possibly\n> change from commit-graph file in the chain to commit-graph file, and\n> thus from commit to commit?\n> \n\nThe latter. The idea is to keep the implementation open to that \npossibility. Load the bloom filter settings from the commit graph file \nyou are dealing with and use those settings to fill out the bloom_key.\n\n>> +\tconst char *path;\n>> +\tsize_t len;\n>> +\n>> +\tif (!revs->commits)\n>> +\t    return;\n> \n> When revs->commits may be NULL?  I understand that we need to have this\n> check because we use revs->commits->item next (sidenote: can revs ever\n> be NULL?).\n> \n> Would `git log --all -- <path>` use Bloom filters (as it theoretically\n> could)?\n> \n\nI am being defensive about revs->commits being NULL for the reason you \ncalled out. \n\nAnd yes, `git log --all <path>` does use Bloom filters if provided a \nsingle pathspec and when following the first parent.  \n\n>> +\n>> +\tif (!revs->repo->objects->commit_graph)\n>> +\t\treturn;\n>> +\n>> +\trevs->bloom_filter_settings = revs->repo->objects->commit_graph->settings;\n>> +\tif (!revs->bloom_filter_settings)\n>> +\t\treturn;\n> \n> All right, so if there is no commit graph, or the commit graph does not\n> include Bloom filter data, there is nothing to do.\n> \n> Though I worry that it would make Git do not use Bloom filter if the top\n> commit-graph in the chain does not include Bloom filter data, while\n> other commit-graph files do (and Git could have used that information to\n> speed up the file history query).\n> \n\nThanks for bringing this up! I need to test this scenario out more. \n\n>> +\n>> +printf \"c7\\nc4\\nc1\" > expect_file1\n> \n> Doing things outside test is discouraged.  We can create a separate test\n> that creates those expect_file* files, or it can be a part of 'create\n> commits' test.\n> \n> Anyway, instead of doing test without Bloom filters (something that\n> should have been tested already by other parts of testsuite), and then\n> doing the same test with Bloom filter, why not compare that the result\n> without and with Bloom filter is the same.  The t5318-commit-graph.sh\n> test does this with help of graph_git_two_modes() function:\n> \n>   graph_git_two_modes () {\n>   \tgit -c core.commitGraph=true  $1 >output\n>   \tgit -c core.commitGraph=false $1 >expect\n>   \ttest_cmp expect output\n>   }\n> \n> Sidenote: I wonder if it is high time to create t/lib-commit-graph.sh\n> helper with, among others, this common function.\n>\n\nThank you! I have restructured the test to do almost everything you \nhave suggested. \n\nAlso, a note for this v1 RFC series: I am working on proper formal tests \nright now. I didn't want to wait to get these nailed down before sending \nout the RFC series and getting the ball rolling. \n\nI have taken note of all the testing suggestions you have made in your \nreview. They are very helpful and I appreciate it! \n\nCreating t/lib-commit-graph.sh helper would be orthogonal to this series,\nso I will follow up with a separate series that does this. \n\nCheers! \nGarima Singh\n"},{"id":"389797","messageId":"f8e9c4c7-b448-4014-5647-87c833a21dcc@gmail.com","threadId":"52499","inReplyTo":"86v9phrcml.fsf@gmail.com","subject":"Re: [PATCH 9/9] commit-graph: add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-01-15T00:55:49Z","receivedAt":"2020-01-15T00:56:03Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 1/11/2020 2:56 PM, Jakub Narebski wrote:\n> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>\n>> The test suite passes when GIT_TEST_COMMIT_GRAPH and\n>> GIT_COMMIT_GRAPH_BLOOM_FILTERS are enabled.\n> \n> Good.  Very good.\n> \n> No errors found by Continuous Integration setup either?\n> \n\nYes, the CI test pipelines all passed.\n\nCheers! \nGarima Singh\n"},{"id":"390075","messageId":"868sm2ck7w.fsf@gmail.com","threadId":"52499","inReplyTo":"3aaf02fe-ac83-5694-2c69-e133879a0030@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-01-20T13:48:19Z","receivedAt":"2020-01-20T13:48:29Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Garima Singh <garimasigit@gmail.com> writes:\n> On 12/31/2019 11:45 AM, Jakub Narebski wrote:\n>> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>>>\n>>> Performance Gains: We tested the performance of 'git log -- <path>' on the git\n>>> repo, the linux repo and some internal large repos, with a variety of paths\n>>> of varying depths.\n>>>\n>>> On the git and linux repos: We observed a 2x to 5x speed up.\n>>>\n>>> On a large internal repo with files seated 6-10 levels deep in the tree: We\n>>> observed 10x to 20x speed ups, with some paths going up to 28 times faster.\n>> \n>> Could you provide some more statistics about this internal repository,\n>> such as number of files, number of commits, perhaps also number of all\n>> objects?  Thanks in advance.\n>> \n>> I wonder why such large difference in performance 2-5x vs 10-20x.  Is it\n>> about the depth of the file hierarchy?  How would the numbers look for\n>> files seated closer to the root in the same large repository, like 3-5\n>> levels deep in the tree?\n>\n> The internal repository we saw these massive gains on has:\n> - 413579 commits. \n> - 183303 files distributed across 34482 folders\n> The size on disk is about 17 GiB. \n\nThank you for the data.  Such information would be important\nconsideration to help to find out whether enabling Bloom filters in\ngiven repository would be worth it.\n\n> And yes, the difference is performance gains is mostly because of how \n> deep the files were in the hierarchy.\n\nRight, this is understandable.  If files are diep in hierarchy, then we\nhave to unpack more tree objects to find out if the file was changed in\na given commit (provided that finding differences do not terminate early\nthanks to hierarchical structure of tree objects).\n\n>                                      How often a file has been touched\n> also makes a difference. The performance gains are less dramatic if the \n> file has a very sparse history even if it is a deep file.\n\nThis looks a bit strange (or maybe I don't understand something).\n\nBloom filter can answer \"no\" and \"maybe\" to subset inclusion query.\nThis means that if file was *not* changed, with great probability the\nanswer from Bloom filter would be \"no\", and we would skip diff-ing\ntrees (which may terminate early, though).\n\nOn the other hand if file was changed by the commit, and the answer from\na Bloom filter is \"maybe\", then we have to perform diffing to make sure.\n\n>\n> The numbers from the git and linux repos for instance, are for files \n> closer to the root, hence 2x to 5x. \n\nThat is quite nice speedup, anyway (git repository cannot be even\nconsidered large; medium -- maybe).\n\n\nP.S. I wonder if it would be worth to create some synthetical repository\nto test performance gains of Bloom filters, perhaps in t/perf...\n\nBest,\n-- \nJakub Narębski\n"},{"id":"390173","messageId":"f5625b23-d7c4-9a72-4ed6-69893de103b0@gmail.com","threadId":"52499","inReplyTo":"868sm2ck7w.fsf@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-01-21T16:14:48Z","receivedAt":"2020-01-21T16:14:51Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 1/20/2020 8:48 AM, Jakub Narebski wrote: >>                                      How often a file has been touched\n>> also makes a difference. The performance gains are less dramatic if the \n>> file has a very sparse history even if it is a deep file.\n> \n> This looks a bit strange (or maybe I don't understand something).\n> \n> Bloom filter can answer \"no\" and \"maybe\" to subset inclusion query.\n> This means that if file was *not* changed, with great probability the\n> answer from Bloom filter would be \"no\", and we would skip diff-ing\n> trees (which may terminate early, though).\n> \n> On the other hand if file was changed by the commit, and the answer from\n> a Bloom filter is \"maybe\", then we have to perform diffing to make sure.\n>\n\nYes. What I meant by statement however is that the performance gain i.e. \ndifference in performance between using and not using bloom filters, is not \nalways as dramatic if the history is sparse and the trees aren't touched \nas often. So it is largely dependent on the shape of the repo and the shape\nof the commit graph. \n \n>>\n>> The numbers from the git and linux repos for instance, are for files \n>> closer to the root, hence 2x to 5x. \n> \n> That is quite nice speedup, anyway (git repository cannot be even\n> considered large; medium -- maybe).\n> \n\nYeah. Git and Linux served as nice initial test beds. If you have any \nsuggestions for interesting repos it would be worth running performanc \ninvestigations on, do let me know! \n\n> \n> P.S. I wonder if it would be worth to create some synthetical repository\n> to test performance gains of Bloom filters, perhaps in t/perf...\n> \n\nI will look into this after I get v1 out on the mailing list. \nThanks! \n\nCheers\nGarima Singh\n"},{"id":"390215","messageId":"20200121234055.GN181522@google.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2020-01-21T23:40:55Z","receivedAt":"2020-01-21T23:41:05Z","isPatch":true,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"Hi,\n\nOn Fri, Dec 20, 2019 at 10:05:11PM +0000, Garima Singh via GitGitGadget wrote:\n> This series is intended to start the conversation and many of the commit\n> messages include specific call outs for suggestions and thoughts. \n\nSince it's mostly in RFC stage, I'm holding off on line by line comments\nfor now. I read the series (thanks for your patience) and I'll try to\nleave some review thoughts in the diffspec here.\n\n> Garima Singh (9):\n>   [1/9] commit-graph: add --changed-paths option to write\nI wonder if this can be combined with 2; without 2 actually the\ndocumentation is wrong for this one, right? Although I suppose you also\nmentioned 2 perhaps being too long :)\n\n>   [2/9] commit-graph: write changed paths bloom filters\nAs I understand it, this one always regenerates the bloom filter pieces,\nand doesn't write it down in the commit-graph file. How much longer does\nthat take now than before? I don't have a great feel for how often 'git\ncommit-graph' is run, or whether I need to be invoking it manually.\n\n>   [3/9] commit-graph: use MAX_NUM_CHUNKS\n>   [4/9] commit-graph: document bloom filter format\nI suppose I might like to see this commit squashed with 5, but it's a\nnit. I'm thinking it'd be handy to say \"git blame commit-graph\" and see\nsome nice doc about the format expected in the commit-graph file.\n\n>   [5/9] commit-graph: write changed path bloom filters to commit-graph file.\nAh, so here we finally write down the result from 2 to disk in the\ncommit-graph file. And without 7, this gets recalculated every time we\ncall 'git commit-graph' still.\n\nAs for a technical doc around here, I'd really appreciate one. But I'm\nspeaking selfishly - I'd also be happy if I could watch a talk about\nthis design to make sure I understand it right :)\n\n>   [6/9] commit-graph: test commit-graph write --changed-paths\n>   [7/9] commit-graph: reuse existing bloom filters during write.\nI saw an option to give up if there wasn't an existing bloom filter, but\nI didn't see an option here to force recalculating. Is there a scenario\nwhen that would be useful? What's the mitigation path if:\n - I have a commit-graph with v0 of the bloom index piece, but update to\n   Git which uses v1?\n - My commit-graph file is corrupted in a way that the bloom filter\n   results are incorrect and I am missing a blob change (and therefore\n   not finding it during the walk)?\nI think I understand that without this commit, 8 is not much speedup\nbecause we will be recalculating the filter for each commit, rather than\nusing the written-down commit-graph file.\n\n>   [8/9] revision.c: use bloom filters to speed up path based revision walks\n>   [9/9] commit-graph: add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag\n\nSpeaking of making sure I understand the change right, I will also\nsummarize my own understanding in the hopes that I can be corrected and\nothers can learn too ;)\n\n - The general idea is that we can write down a hint at the tree level\n   saying \"something I own did change this commit\"; when we look for an\n   object later, we can skip the commit where it looks like that path\n   didn't change.\n - The change is in two pieces: first, to generate the hints per tree\n   (which happens during commit-graph); second, to use those hints to\n   optimize a rev walk (which happens in revision.c patch 8)\n - When we calculate the hints during commit-graph, we check the diff of\n   each tree compared to its recent ancestor to see if there was a\n   change; if so we calculate a hash for each path and use that as a key\n   for a map from hash to path. After we look through everything changed\n   in the diff, we can add it to a cumulative bloom filter list (one\n   filter per commit) so we have a handy in-memory idea of which paths\n   changed in each commit.\n - When it's time to do the rev walk, we ask for the bloom filter for\n   each commit and check if that commit's map contains the path to the\n   object we're worried about; if so, then it's OK to unpack the tree\n   and check if the path we are interested in actually did get changed\n   during that commit.\n\nThanks.\n - Emily\n\n> \n>  Documentation/git-commit-graph.txt            |   5 +\n>  .../technical/commit-graph-format.txt         |  17 ++\n>  Makefile                                      |   1 +\n>  bloom.c                                       | 257 +++++++++++++++++\n>  bloom.h                                       |  51 ++++\n>  builtin/commit-graph.c                        |   9 +-\n>  ci/run-build-and-tests.sh                     |   1 +\n>  commit-graph.c                                | 116 +++++++-\n>  commit-graph.h                                |   9 +-\n>  revision.c                                    |  67 ++++-\n>  revision.h                                    |   5 +\n>  t/README                                      |   3 +\n>  t/helper/test-read-graph.c                    |   4 +\n>  t/t4216-log-bloom.sh                          |  77 ++++++\n>  t/t5318-commit-graph.sh                       |   2 +\n>  t/t5324-split-commit-graph.sh                 |   1 +\n>  t/t5325-commit-graph-bloom.sh                 | 258 ++++++++++++++++++\n>  17 files changed, 875 insertions(+), 8 deletions(-)\n>  create mode 100644 bloom.c\n>  create mode 100644 bloom.h\n>  create mode 100755 t/t4216-log-bloom.sh\n>  create mode 100755 t/t5325-commit-graph-bloom.sh\n> \n> \n> base-commit: b02fd2accad4d48078671adf38fe5b5976d77304\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-497%2Fgarimasi514%2FcoreGit-bloomFilters-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-497/garimasi514/coreGit-bloomFilters-v1\n> Pull-Request: https://github.com/gitgitgadget/git/pull/497\n> -- \n> gitgitgadget\n"},{"id":"390571","messageId":"00ccf0f0-9598-171e-d868-2ab0ea97cc7b@gmail.com","threadId":"52499","inReplyTo":"20200121234055.GN181522@google.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-01-27T18:24:10Z","receivedAt":"2020-01-27T18:24:15Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"On 1/21/2020 6:40 PM, Emily Shaffer wrote:>> Garima Singh (9):\n>>   [1/9] commit-graph: add --changed-paths option to write\n> I wonder if this can be combined with 2; without 2 actually the\n> documentation is wrong for this one, right? Although I suppose you also\n> mentioned 2 perhaps being too long :)\n> \n\nTrue. The ordering of these commits has been a very subjective discussion. \nLeaving this commit isolated the way it is, does help separate the option\nfrom the bloom filter computation and commit graph write step. \n\nAlso for clarity, I have changed the message for 2/9 to:\n`commit-graph: compute Bloom filters for changed paths`\n\n \n>>   [2/9] commit-graph: write changed paths bloom filters\n> As I understand it, this one always regenerates the bloom filter pieces,\n> and doesn't write it down in the commit-graph file. How much longer does\n> that take now than before? I don't have a great feel for how often 'git\n> commit-graph' is run, or whether I need to be invoking it manually.\n> \n\nYes. Computation and writing the bloom filters to the commit-graph file\n(2/9 and 5/9) are ideally one time operations per commit. The time taken\ndepends on the shape and size of the repo since computation involves \nrunning a full diff. If you look at the discussions and patches exchanged\nby Dr. Stolee and Peff: it has improved greatly since this RFC patch. \nMy next submission will carry these patches and more concrete numbers for\ntime and memory. \n\nAlso, `git commit-graph write` is not run very frequently and is ideally\nincremental. A full rewrite is usually only done in case there is a \ncorruption caught by `git commit-graph verify` in which case, you delete\nand rewrite. These are both manual operations. \nSee the docs here: \nhttps://git-scm.com/docs/git-commit-graph\n\nNote: that fetch.writeCommitGraph is now on by default in which case \nnew computations happen automatically for newly fetched commits. But \nI am not adding the --changed-paths option to that yet. So computing, \nwriting and using bloom filters will still be an opt-in feature \nthat users need to manually run. \n \n>>   [7/9] commit-graph: reuse existing bloom filters during write.\n> I saw an option to give up if there wasn't an existing bloom filter, but\n> I didn't see an option here to force recalculating. \n\nThe way to force recalculation at the moment would be to delete the commit\ngraph file and write it again. \n\nIs there a scenario when that would be useful? \n\nYes, if there are algorithm or hash version changes in the changed paths logic\nwe would need to rewrite. Each of the cases I can think of would involve\ntriggering a recalculation by deleting and rewriting.\n\nWhat's the mitigation path if:\n>  - I have a commit-graph with v0 of the bloom index piece, but update to\n>    Git which uses v1?     \n   \n     Take a look at 5/9, when we are parsing the commit graph: if the code \n     is expected to work with a particular version (hash_version = 1 in the \n     current version), and the commit graph has a different version, we just \n     ignore it. In the future, this is where we could extend this to support \n     multiple versions.\n\n>  - My commit-graph file is corrupted in a way that the bloom filter\n>    results are incorrect and I am missing a blob change (and therefore\n>    not finding it during the walk)?\n         The mitigation for any commit graph corruption is to delete and \n     rewrite. \n\n     If however, we are confident that the bloom filter computation itself\n     is wrong, the immediate mitigations would be to deleting the commit\n     graph file and rewriting without the --changed-paths option; and ofc\n     report the bug so it can be investigated and fixed. :) \n\n> Speaking of making sure I understand the change right, I will also\n> summarize my own understanding in the hopes that I can be corrected and\n> others can learn too ;)\n> \n>  - The general idea is that we can write down a hint at the tree level\n>    saying \"something I own did change this commit\"; when we look for an\n>    object later, we can skip the commit where it looks like that path\n>    didn't change.\n\nThe hint we are storing is a bloom filter which answers \"No\" or \"Maybe\"\nto the question \"Did file A change in commit c\"\nIf the answer is No, we can ignore walking that commit. Else, we fall\nback to the diff algorithm like before to confirm if the file changed or\nnot. \n\n>  - The change is in two pieces: first, to generate the hints per tree\n>    (which happens during commit-graph); second, to use those hints to\n>    optimize a rev walk (which happens in revision.c patch 8)\n\nYes. \n\n>  - When we calculate the hints during commit-graph, we check the diff of\n>    each tree compared to its recent ancestor to see if there was a\n>    change; if so we calculate a hash for each path and use that as a key\n>    for a map from hash to path. After we look through everything changed\n>    in the diff, we can add it to a cumulative bloom filter list (one\n>    filter per commit) so we have a handy in-memory idea of which paths\n>    changed in each commit.\n>  - When it's time to do the rev walk, we ask for the bloom filter for\n>    each commit and check if that commit's map contains the path to the\n>    object we're worried about; if so, then it's OK to unpack the tree\n>    and check if the path we are interested in actually did get changed\n>    during that commit.\n\nEssentially yes. There are a few implementation specifics this description\nis glossing over, but I understand that is the intention. \n\n \n> Thanks.\n>  - Emily\n\nCheers! \nGarima Singh\n"},{"id":"391002","messageId":"86blqhsx2p.fsf@gmail.com","threadId":"52499","inReplyTo":"20200121234055.GN181522@google.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-01T23:32:30Z","receivedAt":"2020-02-01T23:32:40Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Emily Shaffer <emilyshaffer@google.com> writes:\n\n[...]\n> Speaking of making sure I understand the change right, I will also\n> summarize my own understanding in the hopes that I can be corrected and\n> others can learn too ;)\n>\n>  - The general idea is that we can write down a hint at the tree level\n>    saying \"something I own did change this commit\"; when we look for an\n>    object later, we can skip the commit where it looks like that path\n>    didn't change.\n\nOr, to be more exact, we write hint about all the files and directories\nchanged in the commit at the commit level.\n\nSay, for example, that changed files are 'README' and 'subdir/file'.\nWe store hint that 'README', 'subdir/file' and 'subdir' paths have\nchanges in them.\n\n>  - The change is in two pieces: first, to generate the hints per tree\n>    (which happens during commit-graph); second, to use those hints to\n>    optimize a rev walk (which happens in revision.c patch 8)\n\nRight.\n\n>  - When we calculate the hints during commit-graph, we check the diff of\n>    each tree compared to its recent ancestor to see if there was a\n>    change;\n\ns/recent ancestor/first parent/\n\nRight, though commits without any changes with respect to first-parent\n(or null tree in case of root i.e. parentless commit) should be rare.\n\n>           if so we calculate a hash for each path and use that as a key\n>    for a map from hash to path.\n\nYes, and no.  Here we enter the details how Bloom filter is constructed.\nWe don't store paths in Bloom filter -- it would take too much space.\n\nThe hashmap is used as an implementation of mathematical set, a\ntemporary structure used during Bloom filter construction.  We could\nhave used string_list, but then we could waste time trying to add\nintermediate directories multiple times (for example if 'foo/bar' and\n'foo/baz' files changed, we need to add 'foo' path only once to Bloom\nfilter).\n\nYou can think of Bloom filter as a compact (and probabilistic, see\nbelow) representation of set of changed paths.\n\n>                                After we look through everything changed\n>    in the diff, we can add it to a cumulative bloom filter list (one\n>    filter per commit) so we have a handy in-memory idea of which paths\n>    changed in each commit.\n\nYes.\n\n>  - When it's time to do the rev walk, we ask for the bloom filter for\n>    each commit and check if that commit's map contains the path to the\n>    object we're worried about; if so, then it's OK to unpack the tree\n>    and check if the path we are interested in actually did get changed\n>    during that commit.\n\nFrom the point of view of rev walk, we ask for the Bloom filter for each\ncommit walked, and check if the (sub)set of changed paths includes given\npath.\n\nBloom filter can answer \"no\" -- then we can skip the commit simplifying\nhistory, or it can answer \"maybe\" -- then we need to check if file was\nactually changed unpacking the trees (there is around 1% probability\nthat Bloom filter will say \"maybe\" if the path is not actually changed).\n\n\nFrom the point of view of Bloom filter, if the path was actually changed\nthe filter will always answer \"maybe\".  If the path was not changed,\nthen in most cases the filter will answer \"no\" but there is 1% of chance\nthat it will answer \"maybe\".\n\n\nI hope that helps,\n--\nJakub Narębski\n"},{"id":"391006","messageId":"86mua0zv6k.fsf@gmail.com","threadId":"52499","inReplyTo":"f5625b23-d7c4-9a72-4ed6-69893de103b0@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-02T18:43:47Z","receivedAt":"2020-02-02T18:44:00Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Garima Singh <garimasigit@gmail.com> writes:\n> On 1/20/2020 8:48 AM, Jakub Narebski wrote:\n\n>>>                                      How often a file has been touched\n>>> also makes a difference. The performance gains are less dramatic if the \n>>> file has a very sparse history even if it is a deep file.\n>> \n>> This looks a bit strange (or maybe I don't understand something).\n>> \n>> Bloom filter can answer \"no\" and \"maybe\" to subset inclusion query.\n>> This means that if file was *not* changed, with great probability the\n>> answer from Bloom filter would be \"no\", and we would skip diff-ing\n>> trees (which may terminate early, though).\n>> \n>> On the other hand if file was changed by the commit, and the answer from\n>> a Bloom filter is \"maybe\", then we have to perform diffing to make sure.\n>\n> Yes. What I meant by statement however is that the performance gain i.e. \n> difference in performance between using and not using bloom filters, is not \n> always as dramatic if the history is sparse and the trees aren't touched \n> as often. So it is largely dependent on the shape of the repo and the shape\n> of the commit graph. \n\nIt probably depends on the depth of changes in a typical skipped commit,\nI think.\n\nIf we are getting history for the\ncore/java/com/android/ims/internal/uce/presence/IPresenceListener.aidl\nand the change is contained in libs/input/ directory, we have to unpack\nonly two trees, independent on the depth of the file we are asking\nabout.\n\nIt could be also possible, if Git is smart enough about it, to halt\nearly if we are checking only if given path changed rather than\ncalculating a full difftree.  Say, for example, that the change was in\ncore/java/org/apache/http/conn/ssl/SSLSocketFactory.java file, while\nwe were getting the history for the following file\ncore/java/com/android/ims/internal/uce/presence/IPresenceListener.aidl\nAfter unpacking two or three threes we know that the second file was not\nchanged.  But if we compute full diff, we have to unpack 8 trees.\nQuite a difference.\n\n\nAll example paths above came from AOSP repository that was used for\ntesting different proposed generation numbers v2, see\nhttps://github.com/derrickstolee/gen-test/blob/master/clone-repos.sh\n\n>>> The numbers from the git and linux repos for instance, are for files \n>>> closer to the root, hence 2x to 5x. \n>> \n>> That is quite nice speedup, anyway (git repository cannot be even\n>> considered large; medium -- maybe).\n>\n> Yeah. Git and Linux served as nice initial test beds. If you have any \n> suggestions for interesting repos it would be worth running performanc \n> investigations on, do let me know! \n\nIf we want repositories with deep path hierarchy, Java projects with\nmandated directory structures might be a good choice, for example\nAndroid (AOSP):\n\n  git clone https://android.googlesource.com/platform/frameworks/base/ android-base\n\nIt is also quite large repository; in 2019 it had around 874000 commits,\naround the same as the Linux kernel repository.\n\nAnother large repository is Chromium -- though I don't know if it has\ndeep filesystem hierarchy.\n\nYou can use the list of different large and large-ish repositories from\nhttps://github.com/derrickstolee/gen-test/blob/master/clone-repos.sh\nOther repositories with large number of commmits not on that list are\nLLVM Compiler, GCC (GNU Compiler Collection) -- just converted to Git,\nHomebrew, and Ruby on Rails.\n\n>> P.S. I wonder if it would be worth to create some synthetical repository\n>> to test performance gains of Bloom filters, perhaps in t/perf...\n>> \n>\n> I will look into this after I get v1 out on the mailing list. \n> Thanks! \n\nIt would be nice to have, but it can wait.\n\n\nKeep up the good work!\n-- \nJakub Narębski\n"},{"id":"391201","messageId":"bf6b93878af5be81148614087aee6b4435ef0396.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 01/11] commit-graph: use MAX_NUM_CHUNKS","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:20Z","receivedAt":"2020-02-05T22:56:36Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nThis is a minor cleanup to make it easier to change the\nnumber of chunks being written to the commit-graph in the future.\n\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 5 +++--\n 1 file changed, 3 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex b205e65ed1..3c4d411326 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -23,6 +23,7 @@\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n+#define MAX_NUM_CHUNKS 5\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -1356,8 +1357,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tint fd;\n \tstruct hashfile *f;\n \tstruct lock_file lk = LOCK_INIT;\n-\tuint32_t chunk_ids[6];\n-\tuint64_t chunk_offsets[6];\n+\tuint32_t chunk_ids[MAX_NUM_CHUNKS + 1];\n+\tuint64_t chunk_offsets[MAX_NUM_CHUNKS + 1];\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \tstruct strbuf progress_title = STRBUF_INIT;\n \tint num_chunks = 3;\n-- \ngitgitgadget\n\n"},{"id":"391202","messageId":"c17bbcbc66ea77bb480391804d1f2db66ffa0926.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 04/11] commit-graph: compute Bloom filters for changed paths","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:23Z","receivedAt":"2020-02-05T22:56:37Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nCompute Bloom filters for the paths that changed between a commit and its\nfirst parent using the implementation in bloom.c, when the\nCOMMIT_GRAPH_WRITE_CHANGED_PATHS flag is set. This computation is done on a\ncommit-by-commit basis. We will write these Bloom filters to the commit graph\nfile in the next change.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 32 +++++++++++++++++++++++++++++++-\n commit-graph.h |  3 ++-\n 2 files changed, 33 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 3c4d411326..724bfcffc4 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -16,6 +16,7 @@\n #include \"hashmap.h\"\n #include \"replace-object.h\"\n #include \"progress.h\"\n+#include \"bloom.h\"\n \n #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n@@ -795,9 +796,11 @@ struct write_commit_graph_context {\n \tunsigned append:1,\n \t\t report_progress:1,\n \t\t split:1,\n-\t\t check_oids:1;\n+\t\t check_oids:1,\n+\t\t changed_paths:1;\n \n \tconst struct split_commit_graph_opts *split_opts;\n+\tuint32_t total_bloom_filter_data_size;\n };\n \n static void write_graph_chunk_fanout(struct hashfile *f,\n@@ -1140,6 +1143,28 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \tstop_progress(&ctx->progress);\n }\n \n+static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n+{\n+\tint i;\n+\tstruct progress *progress = NULL;\n+\n+\tload_bloom_filters();\n+\n+\tif (ctx->report_progress)\n+\t\tprogress = start_progress(\n+\t\t\t_(\"Computing commit diff Bloom filters\"),\n+\t\t\tctx->commits.nr);\n+\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n+\t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n+\t\tdisplay_progress(progress, i + 1);\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n static int add_ref_to_list(const char *refname,\n \t\t\t   const struct object_id *oid,\n \t\t\t   int flags, void *cb_data)\n@@ -1794,6 +1819,8 @@ int write_commit_graph(const char *obj_dir,\n \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n \tctx->check_oids = flags & COMMIT_GRAPH_WRITE_CHECK_OIDS ? 1 : 0;\n \tctx->split_opts = split_opts;\n+\tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n+\tctx->total_bloom_filter_data_size = 0;\n \n \tif (ctx->split) {\n \t\tstruct commit_graph *g;\n@@ -1888,6 +1915,9 @@ int write_commit_graph(const char *obj_dir,\n \n \tcompute_generation_numbers(ctx);\n \n+\tif (ctx->changed_paths)\n+\t\tcompute_bloom_filters(ctx);\n+\n \tres = write_commit_graph_file(ctx);\n \n \tif (ctx->split)\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 7f5c933fa2..952a4b83be 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -76,7 +76,8 @@ enum commit_graph_write_flags {\n \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\n \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n \t/* Make sure that each OID in the input is a valid commit OID. */\n-\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3)\n+\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n+\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n };\n \n struct split_commit_graph_opts {\n-- \ngitgitgadget\n\n"},{"id":"391203","messageId":"78e8e49c3a1131ffacf660603de60729b3dbadc9.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 05/11] commit-graph: examine changed-path objects in pack order","fromName":"Jeff King via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:24Z","receivedAt":"2020-02-05T22:56:39Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nLooking at the diff of commit objects in pack order is much faster than\nin sha1 order, as it gives locality to the access of tree deltas\n(whereas sha1 order is effectively random). Unfortunately the\ncommit-graph code sorts the commits (several times, sometimes as an oid\nand sometimes a pointer-to-commit), and we ultimately traverse in sha1\norder.\n\nInstead, let's remember the position at which we see each commit, and\ntraverse in that order when looking at bloom filters. This drops my time\nfor \"git commit-graph write --changed-paths\" in linux.git from ~4\nminutes to ~1.5 minutes.\n\nProbably the \"--reachable\" code path would want something similar.\n\nOr alternatively, we could use a different data structure (either a\nhash, or maybe even just a bit in \"struct commit\") to keep track of\nwhich oids we've seen, etc instead of sorting. And then we could keep\nthe original order.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 34 +++++++++++++++++++++++++++++++++-\n 1 file changed, 33 insertions(+), 1 deletion(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 724bfcffc4..e125511a1c 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -17,6 +17,7 @@\n #include \"replace-object.h\"\n #include \"progress.h\"\n #include \"bloom.h\"\n+#include \"commit-slab.h\"\n \n #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n@@ -46,6 +47,29 @@\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n+/* Keep track of the order in which commits are added to our list. */\n+define_commit_slab(commit_pos, int);\n+static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n+\n+static void set_commit_pos(struct repository *r, const struct object_id *oid)\n+{\n+\tstatic int32_t max_pos;\n+\tstruct commit *commit = lookup_commit(r, oid);\n+\n+\tif (!commit)\n+\t\treturn; /* should never happen, but be lenient */\n+\n+\t*commit_pos_at(&commit_pos, commit) = max_pos++;\n+}\n+\n+static int commit_pos_cmp(const void *va, const void *vb)\n+{\n+\tconst struct commit *a = *(const struct commit **)va;\n+\tconst struct commit *b = *(const struct commit **)vb;\n+\treturn commit_pos_at(&commit_pos, a) -\n+\t       commit_pos_at(&commit_pos, b);\n+}\n+\n char *get_commit_graph_filename(const char *obj_dir)\n {\n \tchar *filename = xstrfmt(\"%s/info/commit-graph\", obj_dir);\n@@ -1027,6 +1051,8 @@ static int add_packed_commits(const struct object_id *oid,\n \toidcpy(&(ctx->oids.list[ctx->oids.nr]), oid);\n \tctx->oids.nr++;\n \n+\tset_commit_pos(ctx->r, oid);\n+\n \treturn 0;\n }\n \n@@ -1147,6 +1173,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n {\n \tint i;\n \tstruct progress *progress = NULL;\n+\tstruct commit **sorted_by_pos;\n \n \tload_bloom_filters();\n \n@@ -1155,13 +1182,18 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \t\t\t_(\"Computing commit diff Bloom filters\"),\n \t\t\tctx->commits.nr);\n \n+\tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n+\tCOPY_ARRAY(sorted_by_pos, ctx->commits.list, ctx->commits.nr);\n+\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n+\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tstruct commit *c = ctx->commits.list[i];\n+\t\tstruct commit *c = sorted_by_pos[i];\n \t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n \t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n \t\tdisplay_progress(progress, i + 1);\n \t}\n \n+\tfree(sorted_by_pos);\n \tstop_progress(&progress);\n }\n \n-- \ngitgitgadget\n\n"},{"id":"391204","messageId":"58704d81b6b4fbc54715457246aeed783eb32a99.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 06/11] commit-graph: examine commits by generation number","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:25Z","receivedAt":"2020-02-05T22:56:40Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <dstolee@microsoft.com>\n\nWhen running 'git commit-graph write --changed-paths', we sort the\ncommits by pack-order to save time when computing the changed-paths\nbloom filters. This does not help when finding the commits via the\n--reachable flag.\n\nIf not using pack-order, then sort by generation number before\nexamining the diff. Commits with similar generation are more likely\nto have many trees in common, making the diff faster.\n\nOn the Linux kernel repository, this change reduced the computation\ntime for 'git commit-graph write --reachable --changed-paths' from\n3m00s to 1m37s.\n\nHelped-by: Jeff King <peff@peff.net>\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 33 ++++++++++++++++++++++++++++++---\n 1 file changed, 30 insertions(+), 3 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex e125511a1c..32a315058f 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -70,6 +70,25 @@ static int commit_pos_cmp(const void *va, const void *vb)\n \t       commit_pos_at(&commit_pos, b);\n }\n \n+static int commit_gen_cmp(const void *va, const void *vb)\n+{\n+\tconst struct commit *a = *(const struct commit **)va;\n+\tconst struct commit *b = *(const struct commit **)vb;\n+\n+\t/* lower generation commits first */\n+\tif (a->generation < b->generation)\n+\t\treturn -1;\n+\telse if (a->generation > b->generation)\n+\t\treturn 1;\n+\n+\t/* use date as a heuristic when generations are equal */\n+\tif (a->date < b->date)\n+\t\treturn -1;\n+\telse if (a->date > b->date)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n char *get_commit_graph_filename(const char *obj_dir)\n {\n \tchar *filename = xstrfmt(\"%s/info/commit-graph\", obj_dir);\n@@ -821,7 +840,8 @@ struct write_commit_graph_context {\n \t\t report_progress:1,\n \t\t split:1,\n \t\t check_oids:1,\n-\t\t changed_paths:1;\n+\t\t changed_paths:1,\n+\t\t order_by_pack:1;\n \n \tconst struct split_commit_graph_opts *split_opts;\n \tuint32_t total_bloom_filter_data_size;\n@@ -1184,7 +1204,11 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \n \tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n \tCOPY_ARRAY(sorted_by_pos, ctx->commits.list, ctx->commits.nr);\n-\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n+\n+\tif (ctx->order_by_pack)\n+\t\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n+\telse\n+\t\tQSORT(sorted_by_pos, ctx->commits.nr, commit_gen_cmp);\n \n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tstruct commit *c = sorted_by_pos[i];\n@@ -1902,6 +1926,7 @@ int write_commit_graph(const char *obj_dir,\n \t}\n \n \tif (pack_indexes) {\n+\t\tctx->order_by_pack = 1;\n \t\tif ((res = fill_oids_from_packs(ctx, pack_indexes)))\n \t\t\tgoto cleanup;\n \t}\n@@ -1911,8 +1936,10 @@ int write_commit_graph(const char *obj_dir,\n \t\t\tgoto cleanup;\n \t}\n \n-\tif (!pack_indexes && !commit_hex)\n+\tif (!pack_indexes && !commit_hex) {\n+\t\tctx->order_by_pack = 1;\n \t\tfill_oids_from_all_packs(ctx);\n+\t}\n \n \tclose_reachable(ctx);\n \n-- \ngitgitgadget\n\n"},{"id":"391205","messageId":"02b16d94227470059dcee2781e29ae7ae010f602.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 02/11] bloom: core Bloom filter implementation for changed paths","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:21Z","receivedAt":"2020-02-05T22:56:41Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd the core Bloom filter logic for computing the paths changed between a\ncommit and its first parent. For details on what Bloom filters are and how they\nwork, please refer to Dr. Derrick Stolee's blog post [1]. It provides a concise\nexplaination of the adoption of Bloom filters as described in [2] and [3]\n\n1. We currently use 7 and 10 for the number of hashes and the size of each\n   entry respectively. They served as great starting values, the mathematical\n   details behind this choice are described in [1] and [4]. The implementation\n   while not completely open to it at the moment, is flexible enough to allow\n   for tweaking these settings in the future.\n\n   Note: The performance gains we have observed with these values are\n   significant enough that we did not need to tweak these settings.\n   The performance numbers are included in the cover letter of this series\n   and in the message of a subsequent commit where we use Bloom filters in\n   to speed up `git log -- <path>`.\n\n2. As described in the blog and in [3], we do not need 7 independent hashing\n   functions. We use the Murmur3 hashing scheme. Seed it twice and then\n   combine those to procure an arbitrary number of hash values.\n\n3. The filters are sized according to the number of changes in the each commit,\n   with minimum size of one 64 bit word.\n\n4. We fill the Bloom filters as (const char *data, int len) pairs as\n   \"struct bloom_filter\"s in a commit slab.\n\n5. The seed_murmur3 method is implemented as described in [5]. It hashes the\n   given data using a given seed and produces a uniformly distributed hash\n   value.\n\n[1] https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n\n[2] Flavio Bonomi, Michael Mitzenmacher, Rina Panigrahy, Sushil Singh, George Varghese\n    \"An Improved Construction for Counting Bloom Filters\"\n    http://theory.stanford.edu/~rinap/papers/esa2006b.pdf\n    https://doi.org/10.1007/11841036_61\n\n[3] Peter C. Dillinger and Panagiotis Manolios\n    \"Bloom Filters in Probabilistic Verification\"\n    http://www.ccs.neu.edu/home/pete/pub/Bloom-filters-verification.pdf\n    https://doi.org/10.1007/978-3-540-30494-4_26\n\n[4] Thomas Mueller Graf, Daniel Lemire\n    \"Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters\"\n    https://arxiv.org/abs/1912.08258\n\n[5] https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n\nHelped-by: Jeff King <peff@peff.net>\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n Makefile              |   2 +\n bloom.c               | 228 ++++++++++++++++++++++++++++++++++++++++++\n bloom.h               |  56 +++++++++++\n t/helper/test-bloom.c |  84 ++++++++++++++++\n t/helper/test-tool.c  |   1 +\n t/helper/test-tool.h  |   1 +\n t/t0095-bloom.sh      | 113 +++++++++++++++++++++\n 7 files changed, 485 insertions(+)\n create mode 100644 bloom.c\n create mode 100644 bloom.h\n create mode 100644 t/helper/test-bloom.c\n create mode 100755 t/t0095-bloom.sh\n\ndiff --git a/Makefile b/Makefile\nindex 6134104ae6..afba81f4a8 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -695,6 +695,7 @@ X =\n \n PROGRAMS += $(patsubst %.o,git-%$X,$(PROGRAM_OBJS))\n \n+TEST_BUILTINS_OBJS += test-bloom.o\n TEST_BUILTINS_OBJS += test-chmtime.o\n TEST_BUILTINS_OBJS += test-config.o\n TEST_BUILTINS_OBJS += test-ctype.o\n@@ -840,6 +841,7 @@ LIB_OBJS += base85.o\n LIB_OBJS += bisect.o\n LIB_OBJS += blame.o\n LIB_OBJS += blob.o\n+LIB_OBJS += bloom.o\n LIB_OBJS += branch.o\n LIB_OBJS += bulk-checkin.o\n LIB_OBJS += bundle.o\ndiff --git a/bloom.c b/bloom.c\nnew file mode 100644\nindex 0000000000..6082193a75\n--- /dev/null\n+++ b/bloom.c\n@@ -0,0 +1,228 @@\n+#include \"git-compat-util.h\"\n+#include \"bloom.h\"\n+#include \"commit-graph.h\"\n+#include \"object-store.h\"\n+#include \"diff.h\"\n+#include \"diffcore.h\"\n+#include \"revision.h\"\n+#include \"hashmap.h\"\n+\n+define_commit_slab(bloom_filter_slab, struct bloom_filter);\n+\n+struct bloom_filter_slab bloom_filters;\n+\n+struct pathmap_hash_entry {\n+    struct hashmap_entry entry;\n+    const char path[FLEX_ARRAY];\n+};\n+\n+static uint32_t rotate_right(uint32_t value, int32_t count)\n+{\n+\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n+\tcount &= mask;\n+\treturn ((value >> count) | (value << ((-count) & mask)));\n+}\n+\n+/*\n+ * Calculate a hash value for the given data using the given seed.\n+ * Produces a uniformly distributed hash value.\n+ * Not considered to be cryptographically secure.\n+ * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n+ **/\n+static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n+{\n+\tconst uint32_t c1 = 0xcc9e2d51;\n+\tconst uint32_t c2 = 0x1b873593;\n+\tconst uint32_t r1 = 15;\n+\tconst uint32_t r2 = 13;\n+\tconst uint32_t m = 5;\n+\tconst uint32_t n = 0xe6546b64;\n+\tint i;\n+\tuint32_t k1 = 0;\n+\tconst char *tail;\n+\n+\tint len4 = len / sizeof(uint32_t);\n+\n+\tconst uint32_t *blocks = (const uint32_t*)data;\n+\n+\tuint32_t k;\n+\tfor (i = 0; i < len4; i++)\n+\t{\n+\t\tk = blocks[i];\n+\t\tk *= c1;\n+\t\tk = rotate_right(k, r1);\n+\t\tk *= c2;\n+\n+\t\tseed ^= k;\n+\t\tseed = rotate_right(seed, r2) * m + n;\n+\t}\n+\n+\ttail = (data + len4 * sizeof(uint32_t));\n+\n+\tswitch (len & (sizeof(uint32_t) - 1))\n+\t{\n+\tcase 3:\n+\t\tk1 ^= ((uint32_t)tail[2]) << 16;\n+\t\t/*-fallthrough*/\n+\tcase 2:\n+\t\tk1 ^= ((uint32_t)tail[1]) << 8;\n+\t\t/*-fallthrough*/\n+\tcase 1:\n+\t\tk1 ^= ((uint32_t)tail[0]) << 0;\n+\t\tk1 *= c1;\n+\t\tk1 = rotate_right(k1, r1);\n+\t\tk1 *= c2;\n+\t\tseed ^= k1;\n+\t\tbreak;\n+\t}\n+\n+\tseed ^= (uint32_t)len;\n+\tseed ^= (seed >> 16);\n+\tseed *= 0x85ebca6b;\n+\tseed ^= (seed >> 13);\n+\tseed *= 0xc2b2ae35;\n+\tseed ^= (seed >> 16);\n+\n+\treturn seed;\n+}\n+\n+static inline uint64_t get_bitmask(uint32_t pos)\n+{\n+\treturn ((uint64_t)1) << (pos & (BITS_PER_WORD - 1));\n+}\n+\n+void load_bloom_filters(void)\n+{\n+\tinit_bloom_filter_slab(&bloom_filters);\n+}\n+\n+void fill_bloom_key(const char *data,\n+\t\t\t\t\tint len,\n+\t\t\t\t\tstruct bloom_key *key,\n+\t\t\t\t\tstruct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tconst uint32_t seed0 = 0x293ae76f;\n+\tconst uint32_t seed1 = 0x7e646e2c;\n+\tconst uint32_t hash0 = seed_murmur3(seed0, data, len);\n+\tconst uint32_t hash1 = seed_murmur3(seed1, data, len);\n+\n+\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n+\tfor (i = 0; i < settings->num_hashes; i++)\n+\t\tkey->hashes[i] = hash0 + i * hash1;\n+}\n+\n+void add_key_to_filter(struct bloom_key *key,\n+\t\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t\t   struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tuint64_t mod = filter->len * BITS_PER_WORD;\n+\n+\tfor (i = 0; i < settings->num_hashes; i++) {\n+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n+\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n+\n+\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n+\t}\n+}\n+\n+struct bloom_filter *get_bloom_filter(struct repository *r,\n+\t\t\t\t      struct commit *c)\n+{\n+\tstruct bloom_filter *filter;\n+\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\tint i;\n+\tstruct diff_options diffopt;\n+\n+\tif (!bloom_filters.slab_size)\n+\t\treturn NULL;\n+\n+\tfilter = bloom_filter_slab_at(&bloom_filters, c);\n+\n+\trepo_diff_setup(r, &diffopt);\n+\tdiffopt.flags.recursive = 1;\n+\tdiff_setup_done(&diffopt);\n+\n+\tif (c->parents)\n+\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &diffopt);\n+\telse\n+\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n+\tdiffcore_std(&diffopt);\n+\n+\tif (diff_queued_diff.nr <= 512) {\n+\t\tstruct hashmap pathmap;\n+\t\tstruct pathmap_hash_entry* e;\n+\t\tstruct hashmap_iter iter;\n+\t\thashmap_init(&pathmap, NULL, NULL, 0);\n+\n+\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n+\t\t\tconst char* path = diff_queued_diff.queue[i]->two->path;\n+\t\t\tconst char* p = path;\n+\n+\t\t\t/*\n+\t\t\t* Add each leading directory of the changed file, i.e. for\n+\t\t\t* 'dir/subdir/file' add 'dir' and 'dir/subdir' as well, so\n+\t\t\t* the Bloom filter could be used to speed up commands like\n+\t\t\t* 'git log dir/subdir', too.\n+\t\t\t*\n+\t\t\t* Note that directories are added without the trailing '/'.\n+\t\t\t*/\n+\t\t\tdo {\n+\t\t\t\tchar* last_slash = strrchr(p, '/');\n+\n+\t\t\t\tFLEX_ALLOC_STR(e, path, path);\n+\t\t\t\thashmap_entry_init(&e->entry, strhash(p));\n+\t\t\t\thashmap_add(&pathmap, &e->entry);\n+\n+\t\t\t\tif (!last_slash)\n+\t\t\t\t\tlast_slash = (char*)p;\n+\t\t\t\t*last_slash = '\\0';\n+\n+\t\t\t} while (*p);\n+\n+\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n+\t\t}\n+\n+\t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n+\t\tfilter->data = xcalloc(filter->len, sizeof(uint64_t));\n+\n+\t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n+\t\t\tstruct bloom_key key;\n+\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, &settings);\n+\t\t\tadd_key_to_filter(&key, filter, &settings);\n+\t\t}\n+\n+\t\thashmap_free_entries(&pathmap, struct pathmap_hash_entry, entry);\n+\t} else {\n+\t\tfor (i = 0; i < diff_queued_diff.nr; i++)\n+\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n+\t\tfilter->data = NULL;\n+\t\tfilter->len = 0;\n+\t}\n+\n+\tfree(diff_queued_diff.queue);\n+\tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n+\n+\treturn filter;\n+}\n+\n+int bloom_filter_contains(struct bloom_filter *filter,\n+\t\t\t  struct bloom_key *key,\n+\t\t\t  struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tuint64_t mod = filter->len * BITS_PER_WORD;\n+\n+\tif (!mod)\n+\t\treturn -1;\n+\n+\tfor (i = 0; i < settings->num_hashes; i++) {\n+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n+\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n+\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n+\t\t\treturn 0;\n+\t}\n+\n+\treturn 1;\n+}\ndiff --git a/bloom.h b/bloom.h\nnew file mode 100644\nindex 0000000000..7f40c751f7\n--- /dev/null\n+++ b/bloom.h\n@@ -0,0 +1,56 @@\n+#ifndef BLOOM_H\n+#define BLOOM_H\n+\n+struct commit;\n+struct repository;\n+struct commit_graph;\n+\n+struct bloom_filter_settings {\n+\tuint32_t hash_version;\n+\tuint32_t num_hashes;\n+\tuint32_t bits_per_entry;\n+};\n+\n+#define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n+#define BITS_PER_WORD 64\n+\n+/*\n+ * A bloom_filter struct represents a data segment to\n+ * use when testing hash values. The 'len' member\n+ * dictates how many uint64_t entries are stored in\n+ * 'data'.\n+ */\n+struct bloom_filter {\n+\tuint64_t *data;\n+\tint len;\n+};\n+\n+/*\n+ * A bloom_key represents the k hash values for a\n+ * given hash input. These can be precomputed and\n+ * stored in a bloom_key for re-use when testing\n+ * against a bloom_filter.\n+ */\n+struct bloom_key {\n+\tuint32_t *hashes;\n+};\n+\n+void load_bloom_filters(void);\n+\n+void fill_bloom_key(const char *data,\n+\t\t    int len,\n+\t\t    struct bloom_key *key,\n+\t\t    struct bloom_filter_settings *settings);\n+\n+void add_key_to_filter(struct bloom_key *key,\n+\t\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t\t   struct bloom_filter_settings *settings);\n+\n+struct bloom_filter *get_bloom_filter(struct repository *r,\n+\t\t\t\t      struct commit *c);\n+\n+int bloom_filter_contains(struct bloom_filter *filter,\n+\t\t\t  struct bloom_key *key,\n+\t\t\t  struct bloom_filter_settings *settings);\n+\n+#endif\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nnew file mode 100644\nindex 0000000000..331957011b\n--- /dev/null\n+++ b/t/helper/test-bloom.c\n@@ -0,0 +1,84 @@\n+#include \"test-tool.h\"\n+#include \"git-compat-util.h\"\n+#include \"bloom.h\"\n+#include \"test-tool.h\"\n+#include \"cache.h\"\n+#include \"commit-graph.h\"\n+#include \"commit.h\"\n+#include \"config.h\"\n+#include \"object-store.h\"\n+#include \"object.h\"\n+#include \"repository.h\"\n+#include \"tree.h\"\n+\n+struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\n+static void print_bloom_filter(struct bloom_filter *filter) {\n+\tint i;\n+\n+\tif (!filter) {\n+\t\tprintf(\"No filter.\\n\");\n+\t\treturn;\n+\t}\n+\tprintf(\"Filter_Length:%d\\n\", filter->len);\n+\tprintf(\"Filter_Data:\");\n+\tfor (i = 0; i < filter->len; i++){\n+\t\tprintf(\"%\"PRIx64\"|\", filter->data[i]);\n+\t}\n+\tprintf(\"\\n\");\n+}\n+\n+static void add_string_to_filter(const char *data, struct bloom_filter *filter) {\n+\t\tstruct bloom_key key;\n+\t\tint i;\n+\n+\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n+\t\tprintf(\"Hashes:\");\n+\t\tfor (i = 0; i < settings.num_hashes; i++){\n+\t\t\tprintf(\"%08x|\", key.hashes[i]);\n+\t\t}\n+\t\tprintf(\"\\n\");\n+\t\tadd_key_to_filter(&key, filter, &settings);\n+}\n+\n+static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n+{\n+\tstruct commit *c;\n+\tstruct bloom_filter *filter;\n+\tsetup_git_directory();\n+\tc = lookup_commit(the_repository, commit_oid);\n+\tfilter = get_bloom_filter(the_repository, c);\n+\tprint_bloom_filter(filter);\n+}\n+\n+int cmd__bloom(int argc, const char **argv)\n+{\n+    if (!strcmp(argv[1], \"generate_filter\")) {\n+\t\tstruct bloom_filter filter;\n+\t\tint i = 2;\n+\t\tfilter.len =  (settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n+\t\tfilter.data = xcalloc(filter.len, sizeof(uint64_t));\n+\n+\t\tif (!argv[2]){\n+\t\t\tdie(\"at least one input string expected\");\n+\t\t}\n+\n+\t\twhile (argv[i]) {\n+\t\t\tadd_string_to_filter(argv[i], &filter);\n+\t\t\ti++;\n+\t\t}\n+\n+\t\tprint_bloom_filter(&filter);\n+\t}\n+\n+\tif (!strcmp(argv[1], \"get_filter_for_commit\")) {\n+\t\tstruct object_id oid;\n+\t\tconst char *end;\n+\t\tif (parse_oid_hex(argv[2], &oid, &end))\n+\t\t\tdie(\"cannot parse oid '%s'\", argv[2]);\n+\t\tload_bloom_filters();\n+\t\tget_bloom_filter_for_commit(&oid);\n+\t}\n+\n+\treturn 0;\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex c9a232d238..ca4f4b0066 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -14,6 +14,7 @@ struct test_cmd {\n };\n \n static struct test_cmd cmds[] = {\n+\t{ \"bloom\", cmd__bloom },\n \t{ \"chmtime\", cmd__chmtime },\n \t{ \"config\", cmd__config },\n \t{ \"ctype\", cmd__ctype },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex c8549fd87f..05d2b32451 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -4,6 +4,7 @@\n #define USE_THE_INDEX_COMPATIBILITY_MACROS\n #include \"git-compat-util.h\"\n \n+int cmd__bloom(int argc, const char **argv);\n int cmd__chmtime(int argc, const char **argv);\n int cmd__config(int argc, const char **argv);\n int cmd__ctype(int argc, const char **argv);\ndiff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\nnew file mode 100755\nindex 0000000000..424fe4fc29\n--- /dev/null\n+++ b/t/t0095-bloom.sh\n@@ -0,0 +1,113 @@\n+#!/bin/sh\n+\n+test_description='test bloom.c'\n+. ./test-lib.sh\n+\n+test_expect_success 'get bloom filters for commit with no changes' '\n+\tgit init &&\n+\tgit commit --allow-empty -m \"c0\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:0\n+\tFilter_Data:\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'get bloom filter for commit with 10 changes' '\n+\trm actual &&\n+\trm expect &&\n+\tmkdir smallDir &&\n+\tfor i in $(test_seq 0 9)\n+\tdo\n+\t\techo $i >smallDir/$i\n+\tdone &&\n+\tgit add smallDir &&\n+\tgit commit -m \"commit with 10 changes\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:4\n+\tFilter_Data:508928809087080a|8a7648210804001|4089824400951000|841ab310098051a8|\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success EXPENSIVE 'get bloom filter for commit with 513 changes' '\n+\trm actual &&\n+\trm expect &&\n+\tmkdir bigDir &&\n+\tfor i in $(test_seq 0 512)\n+\tdo\n+\t\techo $i >bigDir/$i\n+\tdone &&\n+\tgit add bigDir &&\n+\tgit commit -m \"commit with 513 changes\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:0\n+\tFilter_Data:\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for empty string' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:5615800c|5b966560|61174ab4|66983008|6c19155c|7199fab0|771ae004|\n+\tFilter_Length:1\n+\tFilter_Data:11000110001110|\n+\tEOF\n+\ttest-tool bloom generate_filter \"\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for whitespace' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:1bf014e6|8a91b50b|f9335530|67d4f555|d676957a|4518359f|b3b9d5c4|\n+\tFilter_Length:1\n+\tFilter_Data:401004080200810|\n+\tEOF\n+\ttest-tool bloom generate_filter \" \" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for a root level folder' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:1a21016f|fff1c06d|e5c27f6b|cb933e69|b163fd67|9734bc65|7d057b63|\n+\tFilter_Length:1\n+\tFilter_Data:aaa800000000|\n+\tEOF\n+\ttest-tool bloom generate_filter \"A\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for a root level file' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:e2d51107|30970605|7e58fb03|cc1af001|19dce4ff|679ed9fd|b560cefb|\n+\tFilter_Length:1\n+\tFilter_Data:a8000000000000aa|\n+\tEOF\n+\ttest-tool bloom generate_filter \"file.txt\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for a deep folder' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:864cf838|27f055cd|c993b362|6b3710f7|0cda6e8c|ae7dcc21|502129b6|\n+\tFilter_Length:1\n+\tFilter_Data:1c0000600003000|\n+\tEOF\n+\ttest-tool bloom generate_filter \"A/B/C/D/E\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for a deep file' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:07cdf850|4af629c7|8e1e5b3e|d1468cb5|146ebe2c|5796efa3|9abf211a|\n+\tFilter_Length:1\n+\tFilter_Data:4020100804010080|\n+\tEOF\n+\ttest-tool bloom generate_filter \"A/B/C/D/E/file.txt\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_done\n-- \ngitgitgadget\n\n"},{"id":"391206","messageId":"a698c04a78cf2988fb822e0aa532989f925e0a9e.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 03/11] diff: halt tree-diff early after max_changes","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:22Z","receivedAt":"2020-02-05T22:56:43Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <dstolee@microsoft.com>\n\nWhen computing the changed-paths bloom filters for the commit-graph,\nwe limit the size of the filter by restricting the number of paths\nin the diff. Instead of computing a large diff and then ignoring the\nresult, it is better to halt the diff computation early.\n\nCreate a new \"max_changes\" option in struct diff_options. If non-zero,\nthen halt the diff computation after discovering strictly more changed\npaths. This includes paths corresponding to trees that change.\n\nUse this max_changes option in the bloom filter calculations. This\nreduces the time taken to compute the filters for the Linux kernel\nrepo from 2m50s to 2m35s. On a large internal repository with ~500\ncommits that perform tree-wide changes, the time reduced from\n6m15s to 3m48s.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c     | 4 +++-\n diff.h      | 5 +++++\n tree-diff.c | 6 ++++++\n 3 files changed, 14 insertions(+), 1 deletion(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 6082193a75..818382c03b 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -134,6 +134,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \tint i;\n \tstruct diff_options diffopt;\n+\tint max_changes = 512;\n \n \tif (!bloom_filters.slab_size)\n \t\treturn NULL;\n@@ -142,6 +143,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \trepo_diff_setup(r, &diffopt);\n \tdiffopt.flags.recursive = 1;\n+\tdiffopt.max_changes = max_changes;\n \tdiff_setup_done(&diffopt);\n \n \tif (c->parents)\n@@ -150,7 +152,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n \tdiffcore_std(&diffopt);\n \n-\tif (diff_queued_diff.nr <= 512) {\n+\tif (diff_queued_diff.nr <= max_changes) {\n \t\tstruct hashmap pathmap;\n \t\tstruct pathmap_hash_entry* e;\n \t\tstruct hashmap_iter iter;\ndiff --git a/diff.h b/diff.h\nindex 6febe7e365..9443dc1b00 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -285,6 +285,11 @@ struct diff_options {\n \t/* Number of hexdigits to abbreviate raw format output to. */\n \tint abbrev;\n \n+\t/* If non-zero, then stop computing after this many changes. */\n+\tint max_changes;\n+\t/* For internal use only. */\n+\tint num_changes;\n+\n \tint ita_invisible_in_index;\n /* white-space error highlighting */\n #define WSEH_NEW (1<<12)\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 33ded7f8b3..f3d303c6e5 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -434,6 +434,9 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \n+\t\tif (opt->max_changes && opt->num_changes > opt->max_changes)\n+\t\t\tbreak;\n+\n \t\tif (opt->pathspec.nr) {\n \t\t\tskip_uninteresting(&t, base, opt);\n \t\t\tfor (i = 0; i < nparent; i++)\n@@ -518,6 +521,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \n \t\t\t/* t↓ */\n \t\t\tupdate_tree_entry(&t);\n+\t\t\topt->num_changes++;\n \t\t}\n \n \t\t/* t > p[imin] */\n@@ -535,6 +539,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\tskip_emit_tp:\n \t\t\t/* ∀ pi=p[imin]  pi↓ */\n \t\t\tupdate_tp_entries(tp, nparent);\n+\t\t\topt->num_changes++;\n \t\t}\n \t}\n \n@@ -552,6 +557,7 @@ struct combine_diff_path *diff_tree_paths(\n \tconst struct object_id **parents_oid, int nparent,\n \tstruct strbuf *base, struct diff_options *opt)\n {\n+\topt->num_changes = 0;\n \tp = ll_diff_tree_paths(p, oid, parents_oid, nparent, base, opt);\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"391208","messageId":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"[PATCH v2 00/11] Changed Paths Bloom Filters","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:19Z","receivedAt":"2020-02-05T22:56:44Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"Hey! \n\nThe commit graph feature brought in a lot of performance improvements across\nmultiple commands. However, file based history continues to be a performance\npain point, especially in large repositories. \n\nAdopting changed path bloom filters has been discussed on the list before,\nand a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\nDerrick Stolee [1]. This series is based on Dr. Stolee's proof of concept in\n[2]\n\nPerformance Gains: We tested the performance of git log -- path on the git\nrepo, the linux repo and some internal large repos, with a variety of paths\nof varying depths.\n\nOn the git and linux repos: We observed a 2x to 5x speed up.\n\nOn a large internal repo with files seated 6-10 levels deep in the tree: We\nobserved 10x to 20x speed ups, with some paths going up to 28 times faster.\n\nFuture Work (not included in the scope of this series):\n\n 1. Supporting multiple path based revision walk\n 2. Adopting it in git blame logic. \n 3. Interactions with line log git log -L\n\n\n----------------------------------------------------------------------------\n\nUpdates since the last submission\n\n * Removed all the RFC callouts, this is a ready for full review version\n * Added unit tests for the bloom filter computation layer\n * Added more evolved functional tests for git log\n * Fixed a lot of the bugs found by the tests\n * Reacted to other miscellaneous feedback on the RFC series. \n\nCheers! Garima Singh\n\n[1] https://lore.kernel.org/git/20181009193445.21908-1-szeder.dev@gmail.com/\n[2] \nhttps://lore.kernel.org/git/61559c5b-546e-d61b-d2e1-68de692f5972@gmail.com/\n\nDerrick Stolee (2):\n  diff: halt tree-diff early after max_changes\n  commit-graph: examine commits by generation number\n\nGarima Singh (8):\n  commit-graph: use MAX_NUM_CHUNKS\n  bloom: core Bloom filter implementation for changed paths\n  commit-graph: compute Bloom filters for changed paths\n  commit-graph: write Bloom filters to commit graph file\n  commit-graph: reuse existing Bloom filters during write.\n  commit-graph: add --changed-paths option to write subcommand\n  revision.c: use Bloom filters to speed up path based revision walks\n  commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag\n\nJeff King (1):\n  commit-graph: examine changed-path objects in pack order\n\n Documentation/git-commit-graph.txt            |   5 +\n .../technical/commit-graph-format.txt         |  24 ++\n Makefile                                      |   2 +\n bloom.c                                       | 277 ++++++++++++++++++\n bloom.h                                       |  58 ++++\n builtin/commit-graph.c                        |  10 +-\n ci/run-build-and-tests.sh                     |   1 +\n commit-graph.c                                | 211 ++++++++++++-\n commit-graph.h                                |   9 +-\n diff.h                                        |   5 +\n revision.c                                    | 124 +++++++-\n revision.h                                    |  11 +\n t/README                                      |   5 +\n t/helper/test-bloom.c                         |  84 ++++++\n t/helper/test-read-graph.c                    |   4 +\n t/helper/test-tool.c                          |   1 +\n t/helper/test-tool.h                          |   1 +\n t/t0095-bloom.sh                              | 113 +++++++\n t/t4216-log-bloom.sh                          | 143 +++++++++\n t/t5318-commit-graph.sh                       |   2 +\n t/t5324-split-commit-graph.sh                 |   1 +\n tree-diff.c                                   |   6 +\n 22 files changed, 1088 insertions(+), 9 deletions(-)\n create mode 100644 bloom.c\n create mode 100644 bloom.h\n create mode 100644 t/helper/test-bloom.c\n create mode 100755 t/t0095-bloom.sh\n create mode 100755 t/t4216-log-bloom.sh\n\n\nbase-commit: 5b0ca878e008e82f91300091e793427205ce3544\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-497%2Fgarimasi514%2FcoreGit-bloomFilters-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-497/garimasi514/coreGit-bloomFilters-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/497\n\nRange-diff vs v1:\n\n  3:  a15f87fdcb =  1:  bf6b93878a commit-graph: use MAX_NUM_CHUNKS\n  2:  e52c7ad37a !  2:  02b16d9422 commit-graph: write changed paths bloom filters\n     @@ -1,65 +1,72 @@\n      Author: Garima Singh <garima.singh@microsoft.com>\n      \n     -    commit-graph: write changed paths bloom filters\n     +    bloom: core Bloom filter implementation for changed paths\n      \n     -    The changed path bloom filters help determine which paths changed between a\n     -    commit and its first parent. We already have the \"--changed-paths\" option\n     -    for the \"git commit-graph write\" subcommand, now actually compute them under\n     -    that option. The COMMIT_GRAPH_WRITE_BLOOM_FILTERS flag enables this\n     -    computation.\n     +    Add the core Bloom filter logic for computing the paths changed between a\n     +    commit and its first parent. For details on what Bloom filters are and how they\n     +    work, please refer to Dr. Derrick Stolee's blog post [1]. It provides a concise\n     +    explaination of the adoption of Bloom filters as described in [2] and [3]\n      \n     -    RFC Notes: Here are some details about the implementation and I would love\n     -    to know your thoughts and suggestions for improvements here.\n     +    1. We currently use 7 and 10 for the number of hashes and the size of each\n     +       entry respectively. They served as great starting values, the mathematical\n     +       details behind this choice are described in [1] and [4]. The implementation\n     +       while not completely open to it at the moment, is flexible enough to allow\n     +       for tweaking these settings in the future.\n      \n     -    For details on what bloom filters are and how they work, please refer to\n     -    Dr. Derrick Stolee's blog post [1].\n     -    [1] https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-bloom-filters/\n     +       Note: The performance gains we have observed with these values are\n     +       significant enough that we did not need to tweak these settings.\n     +       The performance numbers are included in the cover letter of this series\n     +       and in the message of a subsequent commit where we use Bloom filters in\n     +       to speed up `git log -- <path>`.\n      \n     -    1. The implementation sticks to the recommended values of 7 and 10 for the\n     -       number of hashes and the size of each entry, as described in the blog.\n     -       The implementation while not completely open to it at the moment, is flexible\n     -       enough to allow for tweaking these settings in the future.\n     -       Note: The performance gains we have observed so far with these values is\n     -       significant enough to not that we did not need to tweak these settings.\n     -       The cover letter of this series has the details and the commit where we have\n     -       git log use bloom filters.\n     -\n     -    2. As described in the blog and the linked technical paper therin, we do not need\n     -       7 independent hashing functions. We use the Murmur3 hashing scheme - seed it\n     -       twice and then combine those to procure an arbitrary number of hash values.\n     +    2. As described in the blog and in [3], we do not need 7 independent hashing\n     +       functions. We use the Murmur3 hashing scheme. Seed it twice and then\n     +       combine those to procure an arbitrary number of hash values.\n      \n          3. The filters are sized according to the number of changes in the each commit,\n             with minimum size of one 64 bit word.\n      \n     -    [Call for advice] We currently cap writing bloom filters for commits with\n     -    atmost 512 changed files. In the current implementation, we compute the diff,\n     -    and then just throw it away once we see it has more than 512 changes.\n     -    Any suggestiongs on how to reduce the work we are doing in this case are more\n     -    than welcome.\n     +    4. We fill the Bloom filters as (const char *data, int len) pairs as\n     +       \"struct bloom_filter\"s in a commit slab.\n      \n     -    [Call for advice] Would the git community like this commit to be split up into\n     -    more granular commits? This commit could possibly be split out further with the\n     -    bloom.c code in its own commit, to be used by the commit-graph in a subsequent\n     -    commit. While I prefer it being contained in one commit this way, I am open to\n     -    suggestions.\n     +    5. The seed_murmur3 method is implemented as described in [5]. It hashes the\n     +       given data using a given seed and produces a uniformly distributed hash\n     +       value.\n      \n     -    [Call for advice] Would a technical document explaining the exact details of\n     -    the bloom filter implemenation and the hashing calculations be helpful? I will\n     -    be adding details into Documentation/technical/commit-graph-format.txt, but the\n     -    bloom filter code is an independent subsystem and could be used outside of the\n     -    commit-graph feature. Is it worth a separate document, or should we apply \"You\n     -    Ain't Gonna Need It\" principles?\n     +    [1] https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n      \n     -    [Call for advice] I plan to add unit tests for bloom.c, specifically to ensure\n     -    that the hash algorithm and bloom key calculations are stable across versions.\n     +    [2] Flavio Bonomi, Michael Mitzenmacher, Rina Panigrahy, Sushil Singh, George Varghese\n     +        \"An Improved Construction for Counting Bloom Filters\"\n     +        http://theory.stanford.edu/~rinap/papers/esa2006b.pdf\n     +        https://doi.org/10.1007/11841036_61\n      \n     -    Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n     +    [3] Peter C. Dillinger and Panagiotis Manolios\n     +        \"Bloom Filters in Probabilistic Verification\"\n     +        http://www.ccs.neu.edu/home/pete/pub/Bloom-filters-verification.pdf\n     +        https://doi.org/10.1007/978-3-540-30494-4_26\n     +\n     +    [4] Thomas Mueller Graf, Daniel Lemire\n     +        \"Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters\"\n     +        https://arxiv.org/abs/1912.08258\n     +\n     +    [5] https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n     +\n     +    Helped-by: Jeff King <peff@peff.net>\n          Helped-by: Derrick Stolee <dstolee@microsoft.com>\n     +    Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n      \n       diff --git a/Makefile b/Makefile\n       --- a/Makefile\n       +++ b/Makefile\n      @@\n     + \n     + PROGRAMS += $(patsubst %.o,git-%$X,$(PROGRAM_OBJS))\n     + \n     ++TEST_BUILTINS_OBJS += test-bloom.o\n     + TEST_BUILTINS_OBJS += test-chmtime.o\n     + TEST_BUILTINS_OBJS += test-config.o\n     + TEST_BUILTINS_OBJS += test-ctype.o\n     +@@\n       LIB_OBJS += bisect.o\n       LIB_OBJS += blame.o\n       LIB_OBJS += blob.o\n     @@ -82,8 +89,6 @@\n      +#include \"revision.h\"\n      +#include \"hashmap.h\"\n      +\n     -+#define BITS_PER_BLOCK 64\n     -+\n      +define_commit_slab(bloom_filter_slab, struct bloom_filter);\n      +\n      +struct bloom_filter_slab bloom_filters;\n     @@ -100,12 +105,18 @@\n      +\treturn ((value >> count) | (value << ((-count) & mask)));\n      +}\n      +\n     ++/*\n     ++ * Calculate a hash value for the given data using the given seed.\n     ++ * Produces a uniformly distributed hash value.\n     ++ * Not considered to be cryptographically secure.\n     ++ * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n     ++ **/\n      +static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n      +{\n      +\tconst uint32_t c1 = 0xcc9e2d51;\n      +\tconst uint32_t c2 = 0x1b873593;\n     -+\tconst int32_t r1 = 15;\n     -+\tconst int32_t r2 = 13;\n     ++\tconst uint32_t r1 = 15;\n     ++\tconst uint32_t r2 = 13;\n      +\tconst uint32_t m = 5;\n      +\tconst uint32_t n = 0xe6546b64;\n      +\tint i;\n     @@ -159,66 +170,67 @@\n      +\n      +static inline uint64_t get_bitmask(uint32_t pos)\n      +{\n     -+\treturn ((uint64_t)1) << (pos & (BITS_PER_BLOCK - 1));\n     ++\treturn ((uint64_t)1) << (pos & (BITS_PER_WORD - 1));\n     ++}\n     ++\n     ++void load_bloom_filters(void)\n     ++{\n     ++\tinit_bloom_filter_slab(&bloom_filters);\n      +}\n      +\n      +void fill_bloom_key(const char *data,\n     -+\t\t    int len,\n     -+\t\t    struct bloom_key *key,\n     -+\t\t    struct bloom_filter_settings *settings)\n     ++\t\t\t\t\tint len,\n     ++\t\t\t\t\tstruct bloom_key *key,\n     ++\t\t\t\t\tstruct bloom_filter_settings *settings)\n      +{\n      +\tint i;\n     -+\tuint32_t seed0 = 0x293ae76f;\n     -+\tuint32_t seed1 = 0x7e646e2c;\n     -+\n     -+\tuint32_t hash0 = seed_murmur3(seed0, data, len);\n     -+\tuint32_t hash1 = seed_murmur3(seed1, data, len);\n     ++\tconst uint32_t seed0 = 0x293ae76f;\n     ++\tconst uint32_t seed1 = 0x7e646e2c;\n     ++\tconst uint32_t hash0 = seed_murmur3(seed0, data, len);\n     ++\tconst uint32_t hash1 = seed_murmur3(seed1, data, len);\n      +\n      +\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n      +\tfor (i = 0; i < settings->num_hashes; i++)\n      +\t\tkey->hashes[i] = hash0 + i * hash1;\n      +}\n      +\n     -+static void add_key_to_filter(struct bloom_key *key,\n     -+\t\t\t      struct bloom_filter *filter,\n     -+\t\t\t      struct bloom_filter_settings *settings)\n     ++void add_key_to_filter(struct bloom_key *key,\n     ++\t\t\t\t\t   struct bloom_filter *filter,\n     ++\t\t\t\t\t   struct bloom_filter_settings *settings)\n      +{\n      +\tint i;\n     -+\tuint64_t mod = filter->len * BITS_PER_BLOCK;\n     ++\tuint64_t mod = filter->len * BITS_PER_WORD;\n      +\n      +\tfor (i = 0; i < settings->num_hashes; i++) {\n      +\t\tuint64_t hash_mod = key->hashes[i] % mod;\n     -+\t\tuint64_t block_pos = hash_mod / BITS_PER_BLOCK;\n     ++\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n      +\n      +\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n      +\t}\n      +}\n      +\n     -+void load_bloom_filters(void)\n     -+{\n     -+\tinit_bloom_filter_slab(&bloom_filters);\n     -+}\n     -+\n      +struct bloom_filter *get_bloom_filter(struct repository *r,\n      +\t\t\t\t      struct commit *c)\n      +{\n      +\tstruct bloom_filter *filter;\n      +\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n      +\tint i;\n     -+\tstruct rev_info revs;\n     -+\tconst char *revs_argv[] = {NULL, \"HEAD\", NULL};\n     ++\tstruct diff_options diffopt;\n     ++\n     ++\tif (!bloom_filters.slab_size)\n     ++\t\treturn NULL;\n      +\n      +\tfilter = bloom_filter_slab_at(&bloom_filters, c);\n     -+\tinit_revisions(&revs, NULL);\n     -+\trevs.diffopt.flags.recursive = 1;\n      +\n     -+\tsetup_revisions(2, revs_argv, &revs, NULL);\n     ++\trepo_diff_setup(r, &diffopt);\n     ++\tdiffopt.flags.recursive = 1;\n     ++\tdiff_setup_done(&diffopt);\n      +\n      +\tif (c->parents)\n     -+\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &revs.diffopt);\n     ++\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &diffopt);\n      +\telse\n     -+\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &revs.diffopt);\n     -+\tdiffcore_std(&revs.diffopt);\n     ++\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n     ++\tdiffcore_std(&diffopt);\n      +\n      +\tif (diff_queued_diff.nr <= 512) {\n      +\t\tstruct hashmap pathmap;\n     @@ -227,18 +239,18 @@\n      +\t\thashmap_init(&pathmap, NULL, NULL, 0);\n      +\n      +\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n     -+\t\t    const char* path = diff_queued_diff.queue[i]->two->path;\n     -+\t\t    const char* p = path;\n     -+\n     -+\t\t    /*\n     -+\t\t     * Add each leading directory of the changed file, i.e. for\n     -+\t\t     * 'dir/subdir/file' add 'dir' and 'dir/subdir' as well, so\n     -+\t\t     * the Bloom filter could be used to speed up commands like\n     -+\t\t     * 'git log dir/subdir', too.\n     -+\t\t     *\n     -+\t\t     * Note that directories are added without the trailing '/'.\n     -+\t\t     */\n     -+\t\t    do {\n     ++\t\t\tconst char* path = diff_queued_diff.queue[i]->two->path;\n     ++\t\t\tconst char* p = path;\n     ++\n     ++\t\t\t/*\n     ++\t\t\t* Add each leading directory of the changed file, i.e. for\n     ++\t\t\t* 'dir/subdir/file' add 'dir' and 'dir/subdir' as well, so\n     ++\t\t\t* the Bloom filter could be used to speed up commands like\n     ++\t\t\t* 'git log dir/subdir', too.\n     ++\t\t\t*\n     ++\t\t\t* Note that directories are added without the trailing '/'.\n     ++\t\t\t*/\n     ++\t\t\tdo {\n      +\t\t\t\tchar* last_slash = strrchr(p, '/');\n      +\n      +\t\t\t\tFLEX_ALLOC_STR(e, path, path);\n     @@ -246,25 +258,27 @@\n      +\t\t\t\thashmap_add(&pathmap, &e->entry);\n      +\n      +\t\t\t\tif (!last_slash)\n     -+\t\t\t\t    last_slash = (char*)p;\n     ++\t\t\t\t\tlast_slash = (char*)p;\n      +\t\t\t\t*last_slash = '\\0';\n      +\n     -+\t\t    } while (*p);\n     ++\t\t\t} while (*p);\n      +\n     -+\t\t    diff_free_filepair(diff_queued_diff.queue[i]);\n     ++\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n      +\t\t}\n      +\n     -+\t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_BLOCK - 1) / BITS_PER_BLOCK;\n     ++\t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n      +\t\tfilter->data = xcalloc(filter->len, sizeof(uint64_t));\n      +\n      +\t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n     -+\t\t    struct bloom_key key;\n     -+\t\t    fill_bloom_key(e->path, strlen(e->path), &key, &settings);\n     -+\t\t    add_key_to_filter(&key, filter, &settings);\n     ++\t\t\tstruct bloom_key key;\n     ++\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, &settings);\n     ++\t\t\tadd_key_to_filter(&key, filter, &settings);\n      +\t\t}\n      +\n      +\t\thashmap_free_entries(&pathmap, struct pathmap_hash_entry, entry);\n      +\t} else {\n     ++\t\tfor (i = 0; i < diff_queued_diff.nr; i++)\n     ++\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n      +\t\tfilter->data = NULL;\n      +\t\tfilter->len = 0;\n      +\t}\n     @@ -274,7 +288,26 @@\n      +\n      +\treturn filter;\n      +}\n     - \\ No newline at end of file\n     ++\n     ++int bloom_filter_contains(struct bloom_filter *filter,\n     ++\t\t\t  struct bloom_key *key,\n     ++\t\t\t  struct bloom_filter_settings *settings)\n     ++{\n     ++\tint i;\n     ++\tuint64_t mod = filter->len * BITS_PER_WORD;\n     ++\n     ++\tif (!mod)\n     ++\t\treturn -1;\n     ++\n     ++\tfor (i = 0; i < settings->num_hashes; i++) {\n     ++\t\tuint64_t hash_mod = key->hashes[i] % mod;\n     ++\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n     ++\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n     ++\t\t\treturn 0;\n     ++\t}\n     ++\n     ++\treturn 1;\n     ++}\n      \n       diff --git a/bloom.h b/bloom.h\n       new file mode 100644\n     @@ -286,6 +319,7 @@\n      +\n      +struct commit;\n      +struct repository;\n     ++struct commit_graph;\n      +\n      +struct bloom_filter_settings {\n      +\tuint32_t hash_version;\n     @@ -294,6 +328,7 @@\n      +};\n      +\n      +#define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n     ++#define BITS_PER_WORD 64\n      +\n      +/*\n      + * A bloom_filter struct represents a data segment to\n     @@ -318,85 +353,253 @@\n      +\n      +void load_bloom_filters(void);\n      +\n     -+struct bloom_filter *get_bloom_filter(struct repository *r,\n     -+\t\t\t\t      struct commit *c);\n     -+\n      +void fill_bloom_key(const char *data,\n      +\t\t    int len,\n      +\t\t    struct bloom_key *key,\n      +\t\t    struct bloom_filter_settings *settings);\n      +\n     ++void add_key_to_filter(struct bloom_key *key,\n     ++\t\t\t\t\t   struct bloom_filter *filter,\n     ++\t\t\t\t\t   struct bloom_filter_settings *settings);\n     ++\n     ++struct bloom_filter *get_bloom_filter(struct repository *r,\n     ++\t\t\t\t      struct commit *c);\n     ++\n     ++int bloom_filter_contains(struct bloom_filter *filter,\n     ++\t\t\t  struct bloom_key *key,\n     ++\t\t\t  struct bloom_filter_settings *settings);\n     ++\n      +#endif\n      \n     - diff --git a/commit-graph.c b/commit-graph.c\n     - --- a/commit-graph.c\n     - +++ b/commit-graph.c\n     + diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n     + new file mode 100644\n     + --- /dev/null\n     + +++ b/t/helper/test-bloom.c\n      @@\n     - #include \"hashmap.h\"\n     - #include \"replace-object.h\"\n     - #include \"progress.h\"\n     ++#include \"test-tool.h\"\n     ++#include \"git-compat-util.h\"\n      +#include \"bloom.h\"\n     - \n     - #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n     - #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n     -@@\n     - \tunsigned append:1,\n     - \t\t report_progress:1,\n     - \t\t split:1,\n     --\t\t check_oids:1;\n     -+\t\t check_oids:1,\n     -+\t\t bloom:1;\n     - \n     - \tconst struct split_commit_graph_opts *split_opts;\n     -+\tuint32_t total_bloom_filter_size;\n     - };\n     - \n     - static void write_graph_chunk_fanout(struct hashfile *f,\n     -@@\n     - \tstop_progress(&ctx->progress);\n     - }\n     - \n     -+static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n     -+{\n     -+\tint i;\n     -+\tstruct progress *progress = NULL;\n     ++#include \"test-tool.h\"\n     ++#include \"cache.h\"\n     ++#include \"commit-graph.h\"\n     ++#include \"commit.h\"\n     ++#include \"config.h\"\n     ++#include \"object-store.h\"\n     ++#include \"object.h\"\n     ++#include \"repository.h\"\n     ++#include \"tree.h\"\n      +\n     -+\tload_bloom_filters();\n     ++struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n      +\n     -+\tif (ctx->report_progress)\n     -+\t\tprogress = start_progress(\n     -+\t\t\t_(\"Computing commit diff Bloom filters\"),\n     -+\t\t\tctx->commits.nr);\n     ++static void print_bloom_filter(struct bloom_filter *filter) {\n     ++\tint i;\n      +\n     -+\tfor (i = 0; i < ctx->commits.nr; i++) {\n     -+\t\tstruct commit *c = ctx->commits.list[i];\n     -+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n     -+\t\tctx->total_bloom_filter_size += sizeof(uint64_t) * filter->len;\n     -+\t\tdisplay_progress(progress, i + 1);\n     ++\tif (!filter) {\n     ++\t\tprintf(\"No filter.\\n\");\n     ++\t\treturn;\n     ++\t}\n     ++\tprintf(\"Filter_Length:%d\\n\", filter->len);\n     ++\tprintf(\"Filter_Data:\");\n     ++\tfor (i = 0; i < filter->len; i++){\n     ++\t\tprintf(\"%\"PRIx64\"|\", filter->data[i]);\n      +\t}\n     ++\tprintf(\"\\n\");\n     ++}\n     ++\n     ++static void add_string_to_filter(const char *data, struct bloom_filter *filter) {\n     ++\t\tstruct bloom_key key;\n     ++\t\tint i;\n     ++\n     ++\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n     ++\t\tprintf(\"Hashes:\");\n     ++\t\tfor (i = 0; i < settings.num_hashes; i++){\n     ++\t\t\tprintf(\"%08x|\", key.hashes[i]);\n     ++\t\t}\n     ++\t\tprintf(\"\\n\");\n     ++\t\tadd_key_to_filter(&key, filter, &settings);\n     ++}\n      +\n     -+\tstop_progress(&progress);\n     ++static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n     ++{\n     ++\tstruct commit *c;\n     ++\tstruct bloom_filter *filter;\n     ++\tsetup_git_directory();\n     ++\tc = lookup_commit(the_repository, commit_oid);\n     ++\tfilter = get_bloom_filter(the_repository, c);\n     ++\tprint_bloom_filter(filter);\n      +}\n      +\n     - static int add_ref_to_list(const char *refname,\n     - \t\t\t   const struct object_id *oid,\n     - \t\t\t   int flags, void *cb_data)\n     ++int cmd__bloom(int argc, const char **argv)\n     ++{\n     ++    if (!strcmp(argv[1], \"generate_filter\")) {\n     ++\t\tstruct bloom_filter filter;\n     ++\t\tint i = 2;\n     ++\t\tfilter.len =  (settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n     ++\t\tfilter.data = xcalloc(filter.len, sizeof(uint64_t));\n     ++\n     ++\t\tif (!argv[2]){\n     ++\t\t\tdie(\"at least one input string expected\");\n     ++\t\t}\n     ++\n     ++\t\twhile (argv[i]) {\n     ++\t\t\tadd_string_to_filter(argv[i], &filter);\n     ++\t\t\ti++;\n     ++\t\t}\n     ++\n     ++\t\tprint_bloom_filter(&filter);\n     ++\t}\n     ++\n     ++\tif (!strcmp(argv[1], \"get_filter_for_commit\")) {\n     ++\t\tstruct object_id oid;\n     ++\t\tconst char *end;\n     ++\t\tif (parse_oid_hex(argv[2], &oid, &end))\n     ++\t\t\tdie(\"cannot parse oid '%s'\", argv[2]);\n     ++\t\tload_bloom_filters();\n     ++\t\tget_bloom_filter_for_commit(&oid);\n     ++\t}\n     ++\n     ++\treturn 0;\n     ++}\n     +\n     + diff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\n     + --- a/t/helper/test-tool.c\n     + +++ b/t/helper/test-tool.c\n      @@\n     - \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n     - \tctx->check_oids = flags & COMMIT_GRAPH_WRITE_CHECK_OIDS ? 1 : 0;\n     - \tctx->split_opts = split_opts;\n     -+\tctx->bloom = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n     -+\tctx->total_bloom_filter_size = 0;\n     + };\n       \n     - \tif (ctx->split) {\n     - \t\tstruct commit_graph *g;\n     + static struct test_cmd cmds[] = {\n     ++\t{ \"bloom\", cmd__bloom },\n     + \t{ \"chmtime\", cmd__chmtime },\n     + \t{ \"config\", cmd__config },\n     + \t{ \"ctype\", cmd__ctype },\n     +\n     + diff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\n     + --- a/t/helper/test-tool.h\n     + +++ b/t/helper/test-tool.h\n      @@\n     + #define USE_THE_INDEX_COMPATIBILITY_MACROS\n     + #include \"git-compat-util.h\"\n       \n     - \tcompute_generation_numbers(ctx);\n     - \n     -+\tif (ctx->bloom)\n     -+\t\tcompute_bloom_filters(ctx);\n     -+\n     - \tres = write_commit_graph_file(ctx);\n     - \n     - \tif (ctx->split)\n     ++int cmd__bloom(int argc, const char **argv);\n     + int cmd__chmtime(int argc, const char **argv);\n     + int cmd__config(int argc, const char **argv);\n     + int cmd__ctype(int argc, const char **argv);\n     +\n     + diff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\n     + new file mode 100755\n     + --- /dev/null\n     + +++ b/t/t0095-bloom.sh\n     +@@\n     ++#!/bin/sh\n     ++\n     ++test_description='test bloom.c'\n     ++. ./test-lib.sh\n     ++\n     ++test_expect_success 'get bloom filters for commit with no changes' '\n     ++\tgit init &&\n     ++\tgit commit --allow-empty -m \"c0\" &&\n     ++\tcat >expect <<-\\EOF &&\n     ++\tFilter_Length:0\n     ++\tFilter_Data:\n     ++\tEOF\n     ++\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n     ++\ttest_cmp expect actual\n     ++'\n     ++\n     ++test_expect_success 'get bloom filter for commit with 10 changes' '\n     ++\trm actual &&\n     ++\trm expect &&\n     ++\tmkdir smallDir &&\n     ++\tfor i in $(test_seq 0 9)\n     ++\tdo\n     ++\t\techo $i >smallDir/$i\n     ++\tdone &&\n     ++\tgit add smallDir &&\n     ++\tgit commit -m \"commit with 10 changes\" &&\n     ++\tcat >expect <<-\\EOF &&\n     ++\tFilter_Length:4\n     ++\tFilter_Data:508928809087080a|8a7648210804001|4089824400951000|841ab310098051a8|\n     ++\tEOF\n     ++\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n     ++\ttest_cmp expect actual\n     ++'\n     ++\n     ++test_expect_success EXPENSIVE 'get bloom filter for commit with 513 changes' '\n     ++\trm actual &&\n     ++\trm expect &&\n     ++\tmkdir bigDir &&\n     ++\tfor i in $(test_seq 0 512)\n     ++\tdo\n     ++\t\techo $i >bigDir/$i\n     ++\tdone &&\n     ++\tgit add bigDir &&\n     ++\tgit commit -m \"commit with 513 changes\" &&\n     ++\tcat >expect <<-\\EOF &&\n     ++\tFilter_Length:0\n     ++\tFilter_Data:\n     ++\tEOF\n     ++\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n     ++\ttest_cmp expect actual\n     ++'\n     ++\n     ++test_expect_success 'compute bloom key for empty string' '\n     ++\tcat >expect <<-\\EOF &&\n     ++\tHashes:5615800c|5b966560|61174ab4|66983008|6c19155c|7199fab0|771ae004|\n     ++\tFilter_Length:1\n     ++\tFilter_Data:11000110001110|\n     ++\tEOF\n     ++\ttest-tool bloom generate_filter \"\" >actual &&\n     ++\ttest_cmp expect actual\n     ++'\n     ++\n     ++test_expect_success 'compute bloom key for whitespace' '\n     ++\tcat >expect <<-\\EOF &&\n     ++\tHashes:1bf014e6|8a91b50b|f9335530|67d4f555|d676957a|4518359f|b3b9d5c4|\n     ++\tFilter_Length:1\n     ++\tFilter_Data:401004080200810|\n     ++\tEOF\n     ++\ttest-tool bloom generate_filter \" \" >actual &&\n     ++\ttest_cmp expect actual\n     ++'\n     ++\n     ++test_expect_success 'compute bloom key for a root level folder' '\n     ++\tcat >expect <<-\\EOF &&\n     ++\tHashes:1a21016f|fff1c06d|e5c27f6b|cb933e69|b163fd67|9734bc65|7d057b63|\n     ++\tFilter_Length:1\n     ++\tFilter_Data:aaa800000000|\n     ++\tEOF\n     ++\ttest-tool bloom generate_filter \"A\" >actual &&\n     ++\ttest_cmp expect actual\n     ++'\n     ++\n     ++test_expect_success 'compute bloom key for a root level file' '\n     ++\tcat >expect <<-\\EOF &&\n     ++\tHashes:e2d51107|30970605|7e58fb03|cc1af001|19dce4ff|679ed9fd|b560cefb|\n     ++\tFilter_Length:1\n     ++\tFilter_Data:a8000000000000aa|\n     ++\tEOF\n     ++\ttest-tool bloom generate_filter \"file.txt\" >actual &&\n     ++\ttest_cmp expect actual\n     ++'\n     ++\n     ++test_expect_success 'compute bloom key for a deep folder' '\n     ++\tcat >expect <<-\\EOF &&\n     ++\tHashes:864cf838|27f055cd|c993b362|6b3710f7|0cda6e8c|ae7dcc21|502129b6|\n     ++\tFilter_Length:1\n     ++\tFilter_Data:1c0000600003000|\n     ++\tEOF\n     ++\ttest-tool bloom generate_filter \"A/B/C/D/E\" >actual &&\n     ++\ttest_cmp expect actual\n     ++'\n     ++\n     ++test_expect_success 'compute bloom key for a deep file' '\n     ++\tcat >expect <<-\\EOF &&\n     ++\tHashes:07cdf850|4af629c7|8e1e5b3e|d1468cb5|146ebe2c|5796efa3|9abf211a|\n     ++\tFilter_Length:1\n     ++\tFilter_Data:4020100804010080|\n     ++\tEOF\n     ++\ttest-tool bloom generate_filter \"A/B/C/D/E/file.txt\" >actual &&\n     ++\ttest_cmp expect actual\n     ++'\n     ++\n     ++test_done\n  -:  ---------- >  3:  a698c04a78 diff: halt tree-diff early after max_changes\n  -:  ---------- >  4:  c17bbcbc66 commit-graph: compute Bloom filters for changed paths\n  -:  ---------- >  5:  78e8e49c3a commit-graph: examine changed-path objects in pack order\n  -:  ---------- >  6:  58704d81b6 commit-graph: examine commits by generation number\n  5:  7648021072 !  7:  39ee061080 commit-graph: write changed path bloom filters to commit-graph file.\n     @@ -1,23 +1,67 @@\n      Author: Garima Singh <garima.singh@microsoft.com>\n      \n     -    commit-graph: write changed path bloom filters to commit-graph file.\n     +    commit-graph: write Bloom filters to commit graph file\n      \n     -    Write bloom filters to the commit-graph using the format described in\n     -    Documentation/technical/commit-graph-format.txt\n     +    Update the technical documentation for commit-graph-format with the formats for\n     +    the Bloom filter index (BIDX) and Bloom filter data (BDAT) chunks. Write the\n     +    computed Bloom filters information to the commit graph file using this format.\n      \n          Helped-by: Derrick Stolee <dstolee@microsoft.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n      \n     + diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n     + --- a/Documentation/technical/commit-graph-format.txt\n     + +++ b/Documentation/technical/commit-graph-format.txt\n     +@@\n     + - The parents of the commit, stored using positional references within\n     +   the graph file.\n     + \n     ++- The Bloom filter of the commit carrying the paths that were changed between\n     ++  the commit and its first parent.\n     ++\n     + These positional references are stored as unsigned 32-bit integers\n     + corresponding to the array position within the list of commit OIDs. Due\n     + to some special constants we use to track parents, we can store at most\n     +@@\n     +       positions for the parents until reaching a value with the most-significant\n     +       bit on. The other bits correspond to the position of the last parent.\n     + \n     ++  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) (N * 4 bytes) [Optional]\n     ++    * The ith entry, BIDX[i], stores the number of 8-byte word blocks in all\n     ++      Bloom filters from commit 0 to commit i (inclusive) in lexicographic\n     ++      order. The Bloom filter for the i-th commit spans from BIDX[i-1] to\n     ++      BIDX[i] (plus header length), where BIDX[-1] is 0.\n     ++    * The BIDX chunk is ignored if the BDAT chunk is not present.\n     ++\n     ++  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n     ++    * It starts with header consisting of three unsigned 32-bit integers:\n     ++      - Version of the hash algorithm being used. We currently only support\n     ++\tvalue 1 which implies the murmur3 hash implemented exactly as described\n     ++\tin https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n     ++      - The number of times a path is hashed and hence the number of bit positions\n     ++\tthat cumulatively determine whether a file is present in the commit.\n     ++      - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n     ++\tcontains 'n' entries, then the filter size is the minimum number of 64-bit\n     ++\twords that contain n*b bits.\n     ++    * The rest of the chunk is the concatenation of all the computed Bloom\n     ++      filters for the commits in lexicographic order.\n     ++    * The BDAT chunk is present iff BIDX is present.\n     ++\n     +   Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n     +       This list of H-byte hashes describe a set of B commit-graph files that\n     +       form a commit-graph chain. The graph position for the ith commit in this\n     +\n       diff --git a/commit-graph.c b/commit-graph.c\n       --- a/commit-graph.c\n       +++ b/commit-graph.c\n      @@\n     + #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n       #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n       #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n     - #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n     --#define MAX_NUM_CHUNKS 5\n      +#define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n      +#define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n     + #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n     +-#define MAX_NUM_CHUNKS 5\n      +#define MAX_NUM_CHUNKS 7\n       \n       #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n     @@ -46,15 +90,33 @@\n      +\t\t\t\tif (hash_version != 1)\n      +\t\t\t\t\tbreak;\n      +\n     -+\t\t\t\tgraph->settings = xmalloc(sizeof(struct bloom_filter_settings));\n     -+\t\t\t\tgraph->settings->hash_version = hash_version;\n     -+\t\t\t\tgraph->settings->num_hashes = get_be32(data + chunk_offset + 4);\n     -+\t\t\t\tgraph->settings->bits_per_entry = get_be32(data + chunk_offset + 8);\n     ++\t\t\t\tgraph->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n     ++\t\t\t\tgraph->bloom_filter_settings->hash_version = hash_version;\n     ++\t\t\t\tgraph->bloom_filter_settings->num_hashes = get_be32(data + chunk_offset + 4);\n     ++\t\t\t\tgraph->bloom_filter_settings->bits_per_entry = get_be32(data + chunk_offset + 8);\n      +\t\t\t}\n      +\t\t\tbreak;\n       \t\t}\n       \n       \t\tif (chunk_repeated) {\n     +@@\n     + \t\tlast_chunk_offset = chunk_offset;\n     + \t}\n     + \n     ++\t/* We need both the bloom chunks to exist together. Else ignore the data */\n     ++\tif ((graph->chunk_bloom_indexes && !graph->chunk_bloom_data)\n     ++\t\t || (!graph->chunk_bloom_indexes && graph->chunk_bloom_data)) {\n     ++\t\tgraph->chunk_bloom_indexes = NULL;\n     ++\t\tgraph->chunk_bloom_data = NULL;\n     ++\t\tgraph->bloom_filter_settings = NULL;\n     ++\t}\n     ++\n     ++\tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data)\n     ++\t\tload_bloom_filters();\n     ++\n     + \thashcpy(graph->oid.hash, graph->data + graph->data_len - graph->hash_len);\n     + \n     + \tif (verify_commit_graph_lite(graph)) {\n      @@\n       \t}\n       }\n     @@ -65,36 +127,67 @@\n      +\tstruct commit **list = ctx->commits.list;\n      +\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n      +\tuint32_t cur_pos = 0;\n     ++\tstruct progress *progress = NULL;\n     ++\tint i = 0;\n     ++\n     ++\tif (ctx->report_progress)\n     ++\t\tprogress = start_delayed_progress(\n     ++\t\t\t_(\"Writing changed paths Bloom filters index\"),\n     ++\t\t\tctx->commits.nr);\n      +\n      +\twhile (list < last) {\n      +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n      +\t\tcur_pos += filter->len;\n     ++\t\tdisplay_progress(progress, ++i);\n      +\t\thashwrite_be32(f, cur_pos);\n      +\t\tlist++;\n      +\t}\n     ++\n     ++\tstop_progress(&progress);\n      +}\n      +\n      +static void write_graph_chunk_bloom_data(struct hashfile *f,\n      +\t\t\t\t\t struct write_commit_graph_context *ctx,\n      +\t\t\t\t\t struct bloom_filter_settings *settings)\n      +{\n     -+\tstruct commit **first = ctx->commits.list;\n     ++\tstruct commit **list = ctx->commits.list;\n      +\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n     ++\tstruct progress *progress = NULL;\n     ++\tint i = 0;\n     ++\n     ++\tif (ctx->report_progress)\n     ++\t\tprogress = start_delayed_progress(\n     ++\t\t\t_(\"Writing changed paths Bloom filters data\"),\n     ++\t\t\tctx->commits.nr);\n      +\n      +\thashwrite_be32(f, settings->hash_version);\n      +\thashwrite_be32(f, settings->num_hashes);\n      +\thashwrite_be32(f, settings->bits_per_entry);\n      +\n     -+\twhile (first < last) {\n     -+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *first);\n     ++\twhile (list < last) {\n     ++\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n     ++\t\tdisplay_progress(progress, ++i);\n      +\t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n     -+\t\tfirst++;\n     ++\t\tlist++;\n      +\t}\n     ++\n     ++\tstop_progress(&progress);\n      +}\n      +\n       static int oid_compare(const void *_a, const void *_b)\n       {\n       \tconst struct object_id *a = (const struct object_id *)_a;\n     +@@\n     + \tload_bloom_filters();\n     + \n     + \tif (ctx->report_progress)\n     +-\t\tprogress = start_progress(\n     +-\t\t\t_(\"Computing commit diff Bloom filters\"),\n     ++\t\tprogress = start_delayed_progress(\n     ++\t\t\t_(\"Computing changed paths Bloom filters\"),\n     + \t\t\tctx->commits.nr);\n     + \n     + \tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n      @@\n       \tstruct strbuf progress_title = STRBUF_INIT;\n       \tint num_chunks = 3;\n     @@ -107,7 +200,7 @@\n       \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n       \t\tnum_chunks++;\n       \t}\n     -+\tif (ctx->bloom) {\n     ++\tif (ctx->changed_paths) {\n      +\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMINDEXES;\n      +\t\tnum_chunks++;\n      +\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMDATA;\n     @@ -120,11 +213,13 @@\n       \t\t\t\t\t\t4 * ctx->num_extra_edges;\n       \t\tnum_chunks++;\n       \t}\n     -+\tif (ctx->bloom) {\n     -+\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] + sizeof(uint32_t) * ctx->commits.nr;\n     ++\tif (ctx->changed_paths) {\n     ++\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n     ++\t\t\t\t\t\tsizeof(uint32_t) * ctx->commits.nr;\n      +\t\tnum_chunks++;\n      +\n     -+\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] + sizeof(uint32_t) * 3 + ctx->total_bloom_filter_size;\n     ++\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n     ++\t\t\t\t\t\tsizeof(uint32_t) * 3 + ctx->total_bloom_filter_data_size;\n      +\t\tnum_chunks++;\n      +\t}\n       \tif (ctx->num_commit_graphs_after > 1) {\n     @@ -134,7 +229,7 @@\n       \twrite_graph_chunk_data(f, hashsz, ctx);\n       \tif (ctx->num_extra_edges)\n       \t\twrite_graph_chunk_extra_edges(f, ctx);\n     -+\tif (ctx->bloom) {\n     ++\tif (ctx->changed_paths) {\n      +\t\twrite_graph_chunk_bloom_indexes(f, ctx);\n      +\t\twrite_graph_chunk_bloom_data(f, ctx, &bloom_settings);\n      +\t}\n     @@ -160,7 +255,16 @@\n      +\tconst unsigned char *chunk_bloom_indexes;\n      +\tconst unsigned char *chunk_bloom_data;\n      +\n     -+\tstruct bloom_filter_settings *settings;\n     ++\tstruct bloom_filter_settings *bloom_filter_settings;\n       };\n       \n       struct commit_graph *load_commit_graph_one_fd_st(int fd, struct stat *st);\n     +@@\n     + \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n     + \t/* Make sure that each OID in the input is a valid commit OID. */\n     + \tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n     +-\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n     ++\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4),\n     + };\n     + \n     + struct split_commit_graph_opts {\n  7:  1e2acb37ad !  8:  b20c8d2b20 commit-graph: reuse existing bloom filters during write.\n     @@ -1,27 +1,31 @@\n      Author: Garima Singh <garima.singh@microsoft.com>\n      \n     -    commit-graph: reuse existing bloom filters during write.\n     +    commit-graph: reuse existing Bloom filters during write.\n      \n     -    Read previously computed bloom filters from the commit-graph file if possible\n     -    to avoid recomputing during commit-graph write.\n     +    Read previously computed Bloom filters from the commit-graph file if\n     +    possible to avoid recomputing during commit-graph write.\n      \n     -    Reading from the commit-graph is based on the format in which bloom filters are\n     -    written in the commit graph file. See method `fill_filter_from_graph` in bloom.c\n     +    See Documentation/technical/commit-graph-format for the format in which\n     +    the Bloom filter information is written to the commit graph file.\n      \n     -    For reading the bloom filter for commit at lexicographic position i:\n     -    1. Read BIDX[i] which essentially gives us the starting index in BDAT for filter\n     -       of commit i+1 (called the next_index in the code)\n     +    To read Bloom filter for a given commit with lexicographic position\n     +    'i' we need to:\n     +    1. Read BIDX[i] which essentially gives us the starting index in BDAT for\n     +       filter of commit i+1. It is essentially the index past the end\n     +       of the filter of commit i. It is called end_index in the code.\n      \n     -    2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT for\n     -       filter of commit i (called the prev_index in the code)\n     -       For i = 0, prev_index will be 0. The first lexicographic commit's filter will\n     -       start at BDAT.\n     +    2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT\n     +       for filter of commit i. It is called the start_index in the code.\n     +       For the first commit, where i = 0, Bloom filter data starts at the\n     +       beginning, just past the header in the BDAT chunk. Hence, start_index\n     +       will be 0.\n      \n     -    3. The length of the filter will be next_index - prev_index, because BIDX[i]\n     -       gives the cumulative 8-byte words including the ith commit's filter.\n     +    3. The length of the filter will be end_index - start_index, because\n     +       BIDX[i] gives the cumulative 8-byte words including the ith\n     +       commit's filter.\n      \n     -    We toggle whether bloom filters should be recomputed based on the compute_if_null\n     -    flag.\n     +    We toggle whether Bloom filters should be recomputed based on the\n     +    compute_if_null flag.\n      \n          Helped-by: Derrick Stolee <dstolee@microsoft.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n     @@ -41,107 +45,135 @@\n       \t}\n       }\n       \n     -+static void fill_filter_from_graph(struct commit_graph *g,\n     ++static int load_bloom_filter_from_graph(struct commit_graph *g,\n      +\t\t\t\t   struct bloom_filter *filter,\n      +\t\t\t\t   struct commit *c)\n      +{\n     -+\tuint32_t lex_pos, prev_index, next_index;\n     ++\tuint32_t lex_pos, start_index, end_index;\n      +\n      +\twhile (c->graph_pos < g->num_commits_in_base)\n      +\t\tg = g->base_graph;\n      +\n     ++\t/* The commit graph commit 'c' lives in doesn't carry bloom filters. */\n     ++\tif (!g->chunk_bloom_indexes)\n     ++\t\treturn 0;\n     ++\n      +\tlex_pos = c->graph_pos - g->num_commits_in_base;\n      +\n     -+\tnext_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n     ++\tend_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n     ++\n      +\tif (lex_pos)\n     -+\t\tprev_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n     ++\t\tstart_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n      +\telse\n     -+\t\tprev_index = 0;\n     ++\t\tstart_index = 0;\n      +\n     -+\tfilter->len = next_index - prev_index;\n     -+\tfilter->data = (uint64_t *)(g->chunk_bloom_data + 8 * prev_index + 12);\n     ++\tfilter->len = end_index - start_index;\n     ++\tfilter->data = (uint64_t *)(g->chunk_bloom_data +\n     ++\t\t\t\t\tsizeof(uint64_t) * start_index +\n     ++\t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n     ++\n     ++\treturn 1;\n      +}\n      +\n     - void load_bloom_filters(void)\n     - {\n     - \tinit_bloom_filter_slab(&bloom_filters);\n     - }\n     - \n       struct bloom_filter *get_bloom_filter(struct repository *r,\n      -\t\t\t\t      struct commit *c)\n      +\t\t\t\t      struct commit *c,\n     -+\t\t\t\t      int compute_if_null)\n     ++\t\t\t\t      int compute_if_not_present)\n       {\n       \tstruct bloom_filter *filter;\n       \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n      @@\n     - \tconst char *revs_argv[] = {NULL, \"HEAD\", NULL};\n       \n       \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n     -+\n     + \n      +\tif (!filter->data) {\n      +\t\tload_commit_graph_info(r, c);\n     -+\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH && r->objects->commit_graph->chunk_bloom_indexes) {\n     -+\t\t\tfill_filter_from_graph(r->objects->commit_graph, filter, c);\n     -+\t\t\treturn filter;\n     ++\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH &&\n     ++\t\t\tr->objects->commit_graph->chunk_bloom_indexes) {\n     ++\t\t\tif (load_bloom_filter_from_graph(r->objects->commit_graph, filter, c))\n     ++\t\t\t\treturn filter;\n     ++\t\t\telse\n     ++\t\t\t\treturn NULL;\n      +\t\t}\n      +\t}\n      +\n     -+\tif (filter->data || !compute_if_null)\n     -+\t\t\treturn filter;\n     ++\tif (filter->data || !compute_if_not_present)\n     ++\t\treturn filter;\n      +\n     - \tinit_revisions(&revs, NULL);\n     - \trevs.diffopt.flags.recursive = 1;\n     - \n     -@@\n     - \tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n     - \n     - \treturn filter;\n     --}\n     - \\ No newline at end of file\n     -+}\n     + \trepo_diff_setup(r, &diffopt);\n     + \tdiffopt.flags.recursive = 1;\n     + \tdiffopt.max_changes = max_changes;\n      \n       diff --git a/bloom.h b/bloom.h\n       --- a/bloom.h\n       +++ b/bloom.h\n      @@\n     - void load_bloom_filters(void);\n     + \n     + #define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n     + #define BITS_PER_WORD 64\n     ++#define BLOOMDATA_CHUNK_HEADER_SIZE 3*sizeof(uint32_t)\n     + \n     + /*\n     +  * A bloom_filter struct represents a data segment to\n     +@@\n     + \t\t\t\t\t   struct bloom_filter_settings *settings);\n       \n       struct bloom_filter *get_bloom_filter(struct repository *r,\n      -\t\t\t\t      struct commit *c);\n      +\t\t\t\t      struct commit *c,\n     -+\t\t\t\t      int compute_if_null);\n     ++\t\t\t\t      int compute_if_not_present);\n       \n     - void fill_bloom_key(const char *data,\n     - \t\t    int len,\n     + int bloom_filter_contains(struct bloom_filter *filter,\n     + \t\t\t  struct bloom_key *key,\n      \n       diff --git a/commit-graph.c b/commit-graph.c\n       --- a/commit-graph.c\n       +++ b/commit-graph.c\n      @@\n     - \tuint32_t cur_pos = 0;\n     + \t\t\tctx->commits.nr);\n       \n       \twhile (list < last) {\n      -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n      +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n       \t\tcur_pos += filter->len;\n     + \t\tdisplay_progress(progress, ++i);\n       \t\thashwrite_be32(f, cur_pos);\n     - \t\tlist++;\n      @@\n       \thashwrite_be32(f, settings->bits_per_entry);\n       \n     - \twhile (first < last) {\n     --\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *first);\n     -+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *first, 0);\n     + \twhile (list < last) {\n     +-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n     ++\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n     + \t\tdisplay_progress(progress, ++i);\n       \t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n     - \t\tfirst++;\n     - \t}\n     + \t\tlist++;\n      @@\n       \n       \tfor (i = 0; i < ctx->commits.nr; i++) {\n     - \t\tstruct commit *c = ctx->commits.list[i];\n     + \t\tstruct commit *c = sorted_by_pos[i];\n      -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n      +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n     - \t\tctx->total_bloom_filter_size += sizeof(uint64_t) * filter->len;\n     + \t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n       \t\tdisplay_progress(progress, i + 1);\n       \t}\n     +@@\n     + \t\tg->data = NULL;\n     + \t\tclose(g->graph_fd);\n     + \t}\n     ++\tfree(g->bloom_filter_settings);\n     + \tfree(g->filename);\n     + \tfree(g);\n     + }\n     +\n     + diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n     + --- a/t/helper/test-bloom.c\n     + +++ b/t/helper/test-bloom.c\n     +@@\n     + \tstruct bloom_filter *filter;\n     + \tsetup_git_directory();\n     + \tc = lookup_commit(the_repository, commit_oid);\n     +-\tfilter = get_bloom_filter(the_repository, c);\n     ++\tfilter = get_bloom_filter(the_repository, c, 1);\n     + \tprint_bloom_filter(filter);\n     + }\n     + \n  1:  6bdde5e4f0 !  9:  3d7ee0c969 commit-graph: add --changed-paths option to write\n     @@ -1,26 +1,13 @@\n      Author: Garima Singh <garima.singh@microsoft.com>\n      \n     -    commit-graph: add --changed-paths option to write\n     +    commit-graph: add --changed-paths option to write subcommand\n      \n          Add --changed-paths option to git commit-graph write. This option will\n     -    soon allow users to compute bloom filters for the paths changed between\n     -    a commit and its first significant parent, and write this information\n     -    into the commit-graph file.\n     -\n     -    Note: This commit does not change any behavior. It only introduces\n     -    the option and passes down the appropriate flag to the commit-graph.\n     -\n     -    RFC Notes:\n     -    1. We named the option --changed-paths to capture what the option does,\n     -       instead of how it does it. The current implementation does this\n     -       using bloom filters. We believe using --changed-paths however keeps\n     -       the implementation open to other data structures.\n     -       All thoughts and suggestions for the name and this approach are\n     -       welcome\n     -\n     -    2. Currently, a subsequent commit in this series will add tests that\n     -       exercise this option. I plan to split that test commit across the\n     -       series as appropriate.\n     +    allow users to compute information about the paths that have changed\n     +    between a commit and its first parent, and write it into the commit graph\n     +    file. If the option is passed to the write subcommand we set the\n     +    COMMIT_GRAPH_WRITE_BLOOM_FILTERS flag and pass it down to the\n     +    commit-graph logic.\n      \n          Helped-by: Derrick Stolee <dstolee@microsoft.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n     @@ -35,7 +22,7 @@\n      +With the `--changed-paths` option, compute and write information about the\n      +paths changed between a commit and it's first parent. This operation can\n      +take a while on large repositories. It provides significant performance gains\n     -+for getting file based history logs with `git log`\n     ++for getting history of a directory or a file with `git log -- <path>`.\n      ++\n       With the `--split` option, write the commit-graph as a chain of multiple\n       commit-graph files stored in `<dir>/info/commit-graphs`. The new commits\n     @@ -66,7 +53,7 @@\n       \tint split;\n       \tint shallow;\n       \tint progress;\n     -+\tint enable_bloom_filters;\n     ++\tint enable_changed_paths;\n       } opts;\n       \n       static int graph_verify(int argc, const char **argv)\n     @@ -74,7 +61,7 @@\n       \t\t\tN_(\"start walk at commits listed by stdin\")),\n       \t\tOPT_BOOL(0, \"append\", &opts.append,\n       \t\t\tN_(\"include all commits already in the commit-graph file\")),\n     -+\t\tOPT_BOOL(0, \"changed-paths\", &opts.enable_bloom_filters,\n     ++\t\tOPT_BOOL(0, \"changed-paths\", &opts.enable_changed_paths,\n      +\t\t\tN_(\"enable computation for changed paths\")),\n       \t\tOPT_BOOL(0, \"progress\", &opts.progress, N_(\"force progress reporting\")),\n       \t\tOPT_BOOL(0, \"split\", &opts.split,\n     @@ -83,22 +70,8 @@\n       \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n       \tif (opts.progress)\n       \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n     -+\tif (opts.enable_bloom_filters)\n     ++\tif (opts.enable_changed_paths)\n      +\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n       \n       \tread_replace_refs = 0;\n       \n     -\n     - diff --git a/commit-graph.h b/commit-graph.h\n     - --- a/commit-graph.h\n     - +++ b/commit-graph.h\n     -@@\n     - \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\n     - \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n     - \t/* Make sure that each OID in the input is a valid commit OID. */\n     --\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3)\n     -+\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n     -+\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n     - };\n     - \n     - struct split_commit_graph_opts {\n  4:  3182a11f7c <  -:  ---------- commit-graph: document bloom filter format\n  6:  85bfdfa59c <  -:  ---------- commit-graph: test commit-graph write --changed-paths\n  8:  72a2bbf676 ! 10:  77f1c561e8 revision.c: use bloom filters to speed up path based revision walks\n     @@ -1,83 +1,35 @@\n      Author: Garima Singh <garima.singh@microsoft.com>\n      \n     -    revision.c: use bloom filters to speed up path based revision walks\n     +    revision.c: use Bloom filters to speed up path based revision walks\n      \n     -    If bloom filters have been written to the commit-graph file, revision walk will\n     -    use them to speed up revision walks for a particular path.\n     -    Note: The current implementation does this in the case of single pathspec\n     -    case only.\n     +    Revision walk will now use Bloom filters for commits to speed up revision\n     +    walks for a particular path (for computing history for that path), if they\n     +    are present in the commit-graph file.\n      \n     -    We load the bloom filters during the prepare_revision_walk step when dealing\n     -    with a single pathspec. While comparing trees in rev_compare_trees(), if the\n     -    bloom filter says that the file is not different between the two trees, we\n     -    don't need to compute the expensive diff. This is where we get our performance\n     -    gains.\n     +    We load the Bloom filters during the prepare_revision_walk step, but only\n     +    when dealing with a single pathspec. While comparing trees in\n     +    rev_compare_trees(), if the Bloom filter says that the file is not different\n     +    between the two trees, we don't need to compute the expensive diff. This is\n     +    where we get our performance gains. The other response of the Bloom filter\n     +    is `maybe`, in which case we fall back to the full diff calculation to\n     +    determine if the path was changed in the commit.\n      \n          Performance Gains:\n     -    We tested the performance of `git log --path` on the git repo, the linux and\n     -    some internal large repos, with a variety of paths of varying depths.\n     +    We tested the performance of `git log -- <path>` on the git repo, the linux\n     +    and some internal large repos, with a variety of paths of varying depths.\n      \n          On the git and linux repos:\n     -    we observed a 2x to 5x speed up.\n     +    - we observed a 2x to 5x speed up.\n      \n          On a large internal repo with files seated 6-10 levels deep in the tree:\n     -    we observed 10x to 20x speed ups, with some paths going up to 28 times faster.\n     -\n     -    RFC Notes:\n     -    I plan to collect the folloowing statistics around this usage of bloom filters\n     -    and trace them out using trace2.\n     -    - number of bloom filter queries,\n     -    - number of \"No\" responses (file hasn't changed)\n     -    - number of \"Maybe\" responses (file may have changed)\n     -    - number of \"Commit not parsed\" cases (commit had too many changes to have a\n     -      bloom filter written out, currently our limit is 512 diffs)\n     +    - we observed 10x to 20x speed ups, with some paths going up to 28 times\n     +      faster.\n      \n          Helped-by: Derrick Stolee <dstolee@microsoft.com\n          Helped-by: SZEDER Gábor <szeder.dev@gmail.com>\n          Helped-by: Jonathan Tan <jonathantanmy@google.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n      \n     - diff --git a/bloom.c b/bloom.c\n     - --- a/bloom.c\n     - +++ b/bloom.c\n     -@@\n     - \n     - \treturn filter;\n     - }\n     -+\n     -+int bloom_filter_contains(struct bloom_filter *filter,\n     -+\t\t\t  struct bloom_key *key,\n     -+\t\t\t  struct bloom_filter_settings *settings)\n     -+{\n     -+\tint i;\n     -+\tuint64_t mod = filter->len * BITS_PER_BLOCK;\n     -+\n     -+\tif (!mod)\n     -+\t\treturn 1;\n     -+\n     -+\tfor (i = 0; i < settings->num_hashes; i++) {\n     -+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n     -+\t\tuint64_t block_pos = hash_mod / BITS_PER_BLOCK;\n     -+\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n     -+\t\t\treturn 0;\n     -+\t}\n     -+\n     -+\treturn 1;\n     -+}\n     -\n     - diff --git a/bloom.h b/bloom.h\n     - --- a/bloom.h\n     - +++ b/bloom.h\n     -@@\n     - \t\t    struct bloom_key *key,\n     - \t\t    struct bloom_filter_settings *settings);\n     - \n     -+int bloom_filter_contains(struct bloom_filter *filter,\n     -+\t\t\t  struct bloom_key *key,\n     -+\t\t\t  struct bloom_filter_settings *settings);\n     -+\n     - #endif\n     -\n       diff --git a/revision.c b/revision.c\n       --- a/revision.c\n       +++ b/revision.c\n     @@ -86,6 +38,7 @@\n       #include \"hashmap.h\"\n       #include \"utf8.h\"\n      +#include \"bloom.h\"\n     ++#include \"json-writer.h\"\n       \n       volatile show_early_output_fn_t show_early_output;\n       \n     @@ -93,26 +46,106 @@\n       \toptions->flags.has_changes = 1;\n       }\n       \n     ++static int bloom_filter_atexit_registered;\n     ++static unsigned int count_bloom_filter_maybe;\n     ++static unsigned int count_bloom_filter_definitely_not;\n     ++static unsigned int count_bloom_filter_false_positive;\n     ++static unsigned int count_bloom_filter_not_present;\n     ++static unsigned int count_bloom_filter_length_zero;\n     ++\n     ++static void trace2_bloom_filter_statistics_atexit(void)\n     ++{\n     ++\tstruct json_writer jw = JSON_WRITER_INIT;\n     ++\n     ++\tjw_object_begin(&jw, 0);\n     ++\tjw_object_intmax(&jw, \"filter_not_present\", count_bloom_filter_not_present);\n     ++\tjw_object_intmax(&jw, \"zero_length_filter\", count_bloom_filter_length_zero);\n     ++\tjw_object_intmax(&jw, \"maybe\", count_bloom_filter_maybe);\n     ++\tjw_object_intmax(&jw, \"definitely_not\", count_bloom_filter_definitely_not);\n     ++\tjw_end(&jw);\n     ++\n     ++\ttrace2_data_json(\"bloom\", the_repository, \"statistics\", &jw);\n     ++\n     ++\tjw_release(&jw);\n     ++}\n     ++\n     ++static void prepare_to_use_bloom_filter(struct rev_info *revs)\n     ++{\n     ++\tstruct pathspec_item *pi;\n     ++\tchar *path_alloc = NULL;\n     ++\tconst char *path;\n     ++\tint last_index;\n     ++\tint len;\n     ++\n     ++\tif (!revs->commits)\n     ++\t    return;\n     ++\n     ++\trepo_parse_commit(revs->repo, revs->commits->item);\n     ++\n     ++\tif (!revs->repo->objects->commit_graph)\n     ++\t\treturn;\n     ++\n     ++\trevs->bloom_filter_settings = revs->repo->objects->commit_graph->bloom_filter_settings;\n     ++\tif (!revs->bloom_filter_settings)\n     ++\t\treturn;\n     ++\n     ++\tpi = &revs->pruning.pathspec.items[0];\n     ++\tlast_index = pi->len - 1;\n     ++\n     ++\tif (pi->match[last_index] == '/') {\n     ++\t    path_alloc = xstrdup(pi->match);\n     ++\t    path_alloc[last_index] = '\\0';\n     ++\t    path = path_alloc;\n     ++\t} else\n     ++\t    path = pi->match;\n     ++\n     ++\tlen = strlen(path);\n     ++\n     ++\trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n     ++\tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n     ++\n     ++\tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n     ++\t\tatexit(trace2_bloom_filter_statistics_atexit);\n     ++\t\tbloom_filter_atexit_registered = 1;\n     ++\t}\n     ++\n     ++\tfree(path_alloc);\n     ++}\n     ++\n      +static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n     -+\t\t\t\t\t\t struct commit *commit,\n     -+\t\t\t\t\t\t struct bloom_key *key,\n     -+\t\t\t\t\t\t struct bloom_filter_settings *settings)\n     ++\t\t\t\t\t\t struct commit *commit)\n      +{\n      +\tstruct bloom_filter *filter;\n     ++\tint result;\n      +\n      +\tif (!revs->repo->objects->commit_graph)\n      +\t\treturn -1;\n     ++\n      +\tif (commit->generation == GENERATION_NUMBER_INFINITY)\n      +\t\treturn -1;\n     -+\tif (!key || !settings)\n     -+\t\treturn -1;\n      +\n      +\tfilter = get_bloom_filter(revs->repo, commit, 0);\n      +\n     -+\tif (!filter || !filter->len)\n     -+\t\treturn 1;\n     ++\tif (!filter) {\n     ++\t\tcount_bloom_filter_not_present++;\n     ++\t\treturn -1;\n     ++\t}\n      +\n     -+\treturn bloom_filter_contains(filter, key, settings);\n     ++\tif (!filter->len) {\n     ++\t\tcount_bloom_filter_length_zero++;\n     ++\t\treturn -1;\n     ++\t}\n     ++\n     ++\tresult = bloom_filter_contains(filter,\n     ++\t\t\t\t       revs->bloom_key,\n     ++\t\t\t\t       revs->bloom_filter_settings);\n     ++\n     ++\tif (result)\n     ++\t\tcount_bloom_filter_maybe++;\n     ++\telse\n     ++\t\tcount_bloom_filter_definitely_not++;\n     ++\n     ++\treturn result;\n      +}\n      +\n       static int rev_compare_tree(struct rev_info *revs,\n     @@ -129,11 +162,8 @@\n       \t\t\treturn REV_TREE_SAME;\n       \t}\n       \n     -+\tif (revs->pruning.pathspec.nr == 1 && !nth_parent) {\n     -+\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs,\n     -+\t\t\t\t\t\t\t\t  commit,\n     -+\t\t\t\t\t\t\t\t  revs->bloom_key,\n     -+\t\t\t\t\t\t\t\t  revs->bloom_filter_settings);\n     ++\tif (revs->pruning.pathspec.nr == 1 && !revs->reflog_info && !nth_parent) {\n     ++\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n      +\n      +\t\tif (bloom_ret == 0)\n      +\t\t\treturn REV_TREE_SAME;\n     @@ -142,6 +172,16 @@\n       \ttree_difference = REV_TREE_SAME;\n       \trevs->pruning.flags.has_changes = 0;\n       \tif (diff_tree_oid(&t1->object.oid, &t2->object.oid, \"\",\n     + \t\t\t   &revs->pruning) < 0)\n     + \t\treturn REV_TREE_DIFFERENT;\n     ++\n     ++\tif (!nth_parent)\n     ++\t\tif (bloom_ret == 1 && tree_difference == REV_TREE_SAME)\n     ++\t\t\tcount_bloom_filter_false_positive++;\n     ++\n     + \treturn tree_difference;\n     + }\n     + \n      @@\n       \t\t\tdie(\"cannot simplify commit %s (because of %s)\",\n       \t\t\t    oid_to_hex(&commit->object.oid),\n     @@ -152,45 +192,19 @@\n       \t\t\tif (!revs->simplify_history || !relevant_commit(p)) {\n       \t\t\t\t/* Even if a merge with an uninteresting\n      @@\n     + \t\t\t\t       FOR_EACH_OBJECT_PROMISOR_ONLY);\n       \t}\n     - }\n       \n     -+static void prepare_to_use_bloom_filter(struct rev_info *revs)\n     -+{\n     -+\tstruct pathspec_item *pi;\n     -+\tconst char *path;\n     -+\tsize_t len;\n     -+\n     -+\tif (!revs->commits)\n     -+\t    return;\n     -+\n     -+\tparse_commit(revs->commits->item);\n     -+\n     -+\tif (!revs->repo->objects->commit_graph)\n     -+\t\treturn;\n     -+\n     -+\trevs->bloom_filter_settings = revs->repo->objects->commit_graph->settings;\n     -+\tif (!revs->bloom_filter_settings)\n     -+\t\treturn;\n     -+\n     -+\tpi = &revs->pruning.pathspec.items[0];\n     -+\tpath = pi->match;\n     -+\tlen = strlen(path);\n     -+\n     -+\tload_bloom_filters();\n     -+\trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n     -+\tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n     -+}\n     -+\n     - int prepare_revision_walk(struct rev_info *revs)\n     - {\n     - \tint i;\n     ++\tif (revs->pruning.pathspec.nr == 1 && !revs->reflog_info)\n     ++\t\tprepare_to_use_bloom_filter(revs);\n     + \tif (revs->no_walk != REVISION_WALK_NO_WALK_UNSORTED)\n     + \t\tcommit_list_sort_by_date(&revs->commits);\n     + \tif (revs->no_walk)\n      @@\n       \t\tsimplify_merges(revs);\n       \tif (revs->children.name)\n       \t\tset_children(revs);\n     -+\tif (revs->pruning.pathspec.nr == 1)\n     -+\t    prepare_to_use_bloom_filter(revs);\n     ++\n       \treturn 0;\n       }\n       \n     @@ -212,12 +226,33 @@\n       \n       \tstruct topo_walk_info *topo_walk_info;\n      +\n     ++\t/* Commit graph bloom filter fields */\n     ++\t/* The bloom filter key for the pathspec */\n      +\tstruct bloom_key *bloom_key;\n     ++\t/*\n     ++\t * The bloom filter settings used to generate the key.\n     ++\t * This is loaded from the commit-graph being used.\n     ++\t */\n      +\tstruct bloom_filter_settings *bloom_filter_settings;\n       };\n       \n       int ref_excluded(struct string_list *, const char *path);\n      \n     + diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\n     + --- a/t/helper/test-read-graph.c\n     + +++ b/t/helper/test-read-graph.c\n     +@@\n     + \t\tprintf(\" commit_metadata\");\n     + \tif (graph->chunk_extra_edges)\n     + \t\tprintf(\" extra_edges\");\n     ++\tif (graph->chunk_bloom_indexes)\n     ++\t\tprintf(\" bloom_indexes\");\n     ++\tif (graph->chunk_bloom_data)\n     ++\t\tprintf(\" bloom_data\");\n     + \tprintf(\"\\n\");\n     + \n     + \tUNLEAK(graph);\n     +\n       diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n       new file mode 100755\n       --- /dev/null\n     @@ -228,72 +263,138 @@\n      +test_description='git log for a path with bloom filters'\n      +. ./test-lib.sh\n      +\n     -+test_expect_success 'setup repo' '\n     ++test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n      +\tgit init &&\n     -+\tgit config core.commitGraph true &&\n     -+\tgit config gc.writeCommitGraph false &&\n     -+\tinfodir=\".git/objects/info\" &&\n     -+\tgraphdir=\"$infodir/commit-graphs\" &&\n     -+\ttest_oid_init\n     ++\tmkdir A A/B A/B/C &&\n     ++\ttest_commit c1 A/file1 &&\n     ++\ttest_commit c2 A/B/file2 &&\n     ++\ttest_commit c3 A/B/C/file3 &&\n     ++\ttest_commit c4 A/file1 &&\n     ++\ttest_commit c5 A/B/file2 &&\n     ++\ttest_commit c6 A/B/C/file3 &&\n     ++\ttest_commit c7 A/file1 &&\n     ++\ttest_commit c8 A/B/file2 &&\n     ++\ttest_commit c9 A/B/C/file3 &&\n     ++\tgit checkout -b side HEAD~4 &&\n     ++\ttest_commit side-1 file4 &&\n     ++\tgit checkout master &&\n     ++\tgit merge side &&\n     ++\ttest_commit c10 file5 &&\n     ++\tmv file5 file5_renamed &&\n     ++\tgit add file5_renamed &&\n     ++\tgit commit -m \"rename\" &&\n     ++\tgit commit-graph write --reachable --changed-paths\n      +'\n     -+\n     -+test_expect_success 'create 9 commits and repack' '\n     -+\ttest_commit c1 file1 &&\n     -+\ttest_commit c2 file2 &&\n     -+\ttest_commit c3 file3 &&\n     -+\ttest_commit c4 file1 &&\n     -+\ttest_commit c5 file2 &&\n     -+\ttest_commit c6 file3 &&\n     -+\ttest_commit c7 file1 &&\n     -+\ttest_commit c8 file2 &&\n     -+\ttest_commit c9 file3\n     -+'\n     -+\n     -+printf \"c7\\nc4\\nc1\" > expect_file1\n     -+\n     -+test_expect_success 'log without bloom filters' '\n     -+\tgit log --pretty=\"format:%s\"  -- file1 > actual &&\n     -+\ttest_cmp expect_file1 actual\n     -+'\n     -+\n     -+printf \"c8\\nc7\\nc5\\nc4\\nc2\\nc1\" > expect_file1_file2\n     -+\n     -+test_expect_success 'multi-path log without bloom filters' '\n     -+\tgit log --pretty=\"format:%s\"  -- file1 file2 > actual &&\n     -+\ttest_cmp expect_file1_file2 actual\n     -+'\n     -+\n      +graph_read_expect() {\n      +\tOPTIONAL=\"\"\n      +\tNUM_CHUNKS=5\n     -+\tif test ! -z $2\n     -+\tthen\n     -+\t\tOPTIONAL=\" $2\"\n     -+\t\tNUM_CHUNKS=$((3 + $(echo \"$2\" | wc -w)))\n     -+\tfi\n      +\tcat >expect <<- EOF\n      +\theader: 43475048 1 1 $NUM_CHUNKS 0\n      +\tnum_commits: $1\n     -+\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data$OPTIONAL\n     ++\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n      +\tEOF\n      +\ttest-tool read-graph >output &&\n      +\ttest_cmp expect output\n      +}\n      +\n     -+test_expect_success 'write commit graph with bloom filters' '\n     -+\tgit commit-graph write --reachable --changed-paths &&\n     -+\ttest_path_is_file $infodir/commit-graph &&\n     -+\tgraph_read_expect \"9\"\n     ++test_expect_success 'commit-graph write wrote out the bloom chunks' '\n     ++\tgraph_read_expect 13\n     ++'\n     ++\n     ++setup() {\n     ++\trm output\n     ++\trm \"$TRASH_DIRECTORY/trace.perf\"\n     ++\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom\n     ++\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n     ++}\n     ++\n     ++test_bloom_filters_used() {\n     ++\tlog_args=$1\n     ++\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":0,\\\"zero_length_filter\\\":0,\\\"maybe\\\"\"\n     ++\tsetup \"$log_args\"\n     ++\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && test_cmp log_wo_bloom log_w_bloom\n     ++}\n     ++\n     ++test_bloom_filters_not_used() {\n     ++\tlog_args=$1\n     ++\tsetup \"$log_args\"\n     ++\t!(grep -q \"statistics:{\\\"filter_not_present\\\":\" \"$TRASH_DIRECTORY/trace.perf\") && test_cmp log_wo_bloom log_w_bloom\n     ++}\n     ++\n     ++for path in A A/B A/B/C A/file1 A/B/file2 A/B/C/file3 file4 file5_renamed\n     ++do\n     ++\tfor option in \"\" \\\n     ++\t\t      \"--full-history\" \\\n     ++\t\t      \"--full-history --simplify-merges\" \\\n     ++\t\t      \"--simplify-merges\" \\\n     ++\t\t      \"--simplify-by-decoration\" \\\n     ++\t\t      \"--follow\" \\\n     ++\t\t      \"--first-parent\" \\\n     ++\t\t      \"--topo-order\" \\\n     ++\t\t      \"--date-order\" \\\n     ++\t\t      \"--author-date-order\" \\\n     ++\t\t      \"--ancestry-path side..master\"\n     ++\tdo\n     ++\t\ttest_expect_success \"git log option: $option for path: $path\" '\n     ++\t\t\ttest_bloom_filters_used \"$option -- $path\"\n     ++\t\t'\n     ++\tdone\n     ++done\n     ++\n     ++test_expect_success 'git log -- folder works with and without the trailing slash' '\n     ++\ttest_bloom_filters_used \"-- A\" &&\n     ++\ttest_bloom_filters_used \"-- A/\"\n     ++'\n     ++\n     ++test_expect_success 'git log for path that does not exist. ' '\n     ++\ttest_bloom_filters_used \"-- path_does_not_exist\"\n     ++'\n     ++\n     ++test_expect_success 'git log with --walk-reflogs does not use bloom filters' '\n     ++\ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n     ++'\n     ++\n     ++test_expect_success 'git log -- multiple path specs does not use bloom filters' '\n     ++\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n     ++'\n     ++\n     ++test_expect_success 'git log with wildcard that resolves to a single path uses bloom filters' '\n     ++\ttest_bloom_filters_used \"-- *4\" &&\n     ++\ttest_bloom_filters_used \"-- *renamed\"\n      +'\n      +\n     -+test_expect_success 'log using bloom filters' '\n     -+\tgit log --pretty=\"format:%s\" -- file1 > actual &&\n     -+\ttest_cmp expect_file1 actual\n     ++test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses bloom filters' '\n     ++\ttest_bloom_filters_not_used \"-- *\" &&\n     ++\ttest_bloom_filters_not_used \"-- file*\"\n      +'\n      +\n     -+test_expect_success 'multi-path log using bloom filters' '\n     -+\tgit log --pretty=\"format:%s\"  -- file1 file2 > actual &&\n     -+\ttest_cmp expect_file1_file2 actual\n     ++test_expect_success 'setup - add commit-graph to the chain without bloom filters' '\n     ++\ttest_commit c14 A/anotherFile2 &&\n     ++\ttest_commit c15 A/B/anotherFile2 &&\n     ++\ttest_commit c16 A/B/C/anotherFile2 &&\n     ++\tGIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0 git commit-graph write --reachable --split &&\n     ++\ttest_line_count = 2 .git/objects/info/commit-graphs/commit-graph-chain\n     ++'\n     ++\n     ++test_expect_success 'git log does not use bloom filters if the latest graph does not have bloom filters.' '\n     ++\ttest_bloom_filters_not_used \"-- A/B\"\n     ++'\n     ++\n     ++test_expect_success 'setup - add commit-graph to the chain with bloom filters' '\n     ++\ttest_commit c17 A/anotherFile3 &&\n     ++\tgit commit-graph write --reachable --changed-paths --split &&\n     ++\ttest_line_count = 3 .git/objects/info/commit-graphs/commit-graph-chain\n     ++'\n     ++\n     ++test_bloom_filters_used_when_some_filters_are_missing() {\n     ++\tlog_args=$1\n     ++\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":3,\\\"zero_length_filter\\\":0,\\\"maybe\\\":6,\\\"definitely_not\\\":6\"\n     ++\tsetup \"$log_args\"\n     ++\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && test_cmp log_wo_bloom log_w_bloom\n     ++}\n     ++\n     ++test_expect_success 'git log uses bloom filters if they exist in the latest but not all commit graphs in the chain.' '\n     ++\ttest_bloom_filters_used_when_some_filters_are_missing \"-- A/B\"\n      +'\n      +\n      +test_done\n  9:  e1c315d0a7 ! 11:  e1b076a714 commit-graph: add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag\n     @@ -1,14 +1,15 @@\n      Author: Garima Singh <garima.singh@microsoft.com>\n      \n     -    commit-graph: add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag\n     +    commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag\n      \n     -    Add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag to the test setup suite in\n     -    order to toggle writing bloom filters when running any of the git tests. If set\n     -    to true, we will compute and write bloom filters every time a test calls\n     -    `git commit-graph write`.\n     +    Add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag to the test setup suite\n     +    in order to toggle writing Bloom filters when running any of the git tests.\n     +    If set to true, we will compute and write Bloom filters every time a test\n     +    calls `git commit-graph write`, as if the `--changed-paths` option was\n     +    passed in.\n      \n          The test suite passes when GIT_TEST_COMMIT_GRAPH and\n     -    GIT_COMMIT_GRAPH_BLOOM_FILTERS are enabled.\n     +    GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS are enabled.\n      \n          Helped-by: Derrick Stolee <dstolee@microsoft.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n     @@ -20,8 +21,9 @@\n       \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n       \tif (opts.progress)\n       \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n     --\tif (opts.enable_bloom_filters)\n     -+\tif (opts.enable_bloom_filters || git_env_bool(GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS, 0))\n     +-\tif (opts.enable_changed_paths)\n     ++\tif (opts.enable_changed_paths ||\n     ++\t    git_env_bool(GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS, 0))\n       \t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n       \n       \tread_replace_refs = 0;\n     @@ -33,7 +35,7 @@\n       \texport GIT_TEST_OE_SIZE=10\n       \texport GIT_TEST_OE_DELTA_SIZE=5\n       \texport GIT_TEST_COMMIT_GRAPH=1\n     -+\texport GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=1\n     ++\texport GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=1\n       \texport GIT_TEST_MULTI_PACK_INDEX=1\n       \tmake test\n       \t;;\n     @@ -45,7 +47,7 @@\n       \n       #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n       #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n     -+#define GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS \"GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS\"\n     ++#define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n       \n       struct commit;\n       struct bloom_filter_settings;\n     @@ -57,8 +59,10 @@\n       be written after every 'git commit' command, and overrides the\n       'core.commitGraph' setting to true.\n       \n     -+GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=<boolean>, when true, forces commit-graph\n     -+write to compute and write bloom filters for every 'git commit-graph write'\n     ++GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n     ++commit-graph write to compute and write changed path Bloom filters for\n     ++every 'git commit-graph write', as if the `--changed-paths` option was\n     ++passed in.\n      +\n       GIT_TEST_FSMONITOR=$PWD/t7519/fsmonitor-all exercises the fsmonitor\n       code path for utilizing a file system monitor to speed up detecting\n     @@ -72,11 +76,11 @@\n       . ./test-lib.sh\n       \n      +GIT_TEST_COMMIT_GRAPH=0\n     -+GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n     ++GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n      +\n     - test_expect_success 'setup repo' '\n     + test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n       \tgit init &&\n     - \tgit config core.commitGraph true &&\n     + \tmkdir A A/B A/B/C &&\n      \n       diff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\n       --- a/t/t5318-commit-graph.sh\n     @@ -85,7 +89,7 @@\n       test_description='commit graph'\n       . ./test-lib.sh\n       \n     -+GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n     ++GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n      +\n       test_expect_success 'setup full repo' '\n       \tmkdir full &&\n     @@ -98,21 +102,7 @@\n       . ./test-lib.sh\n       \n       GIT_TEST_COMMIT_GRAPH=0\n     -+GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n     ++GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n       \n       test_expect_success 'setup repo' '\n       \tgit init &&\n     -\n     - diff --git a/t/t5325-commit-graph-bloom.sh b/t/t5325-commit-graph-bloom.sh\n     - --- a/t/t5325-commit-graph-bloom.sh\n     - +++ b/t/t5325-commit-graph-bloom.sh\n     -@@\n     - test_description='commit graph with bloom filters'\n     - . ./test-lib.sh\n     - \n     -+GIT_TEST_COMMIT_GRAPH=0\n     -+GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS=0\n     -+\n     - test_expect_success 'setup repo' '\n     - \tgit init &&\n     - \tgit config core.commitGraph true &&\n\n-- \ngitgitgadget\n"},{"id":"391207","messageId":"39ee0610800d7d2d92785d392df941fc5a0b231b.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 07/11] commit-graph: write Bloom filters to commit graph file","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:26Z","receivedAt":"2020-02-05T22:56:46Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nUpdate the technical documentation for commit-graph-format with the formats for\nthe Bloom filter index (BIDX) and Bloom filter data (BDAT) chunks. Write the\ncomputed Bloom filters information to the commit graph file using this format.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n .../technical/commit-graph-format.txt         |  24 ++++\n commit-graph.c                                | 118 +++++++++++++++++-\n commit-graph.h                                |   7 +-\n 3 files changed, 145 insertions(+), 4 deletions(-)\n\ndiff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\nindex a4f17441ae..22e511643d 100644\n--- a/Documentation/technical/commit-graph-format.txt\n+++ b/Documentation/technical/commit-graph-format.txt\n@@ -17,6 +17,9 @@ metadata, including:\n - The parents of the commit, stored using positional references within\n   the graph file.\n \n+- The Bloom filter of the commit carrying the paths that were changed between\n+  the commit and its first parent.\n+\n These positional references are stored as unsigned 32-bit integers\n corresponding to the array position within the list of commit OIDs. Due\n to some special constants we use to track parents, we can store at most\n@@ -93,6 +96,27 @@ CHUNK DATA:\n       positions for the parents until reaching a value with the most-significant\n       bit on. The other bits correspond to the position of the last parent.\n \n+  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) (N * 4 bytes) [Optional]\n+    * The ith entry, BIDX[i], stores the number of 8-byte word blocks in all\n+      Bloom filters from commit 0 to commit i (inclusive) in lexicographic\n+      order. The Bloom filter for the i-th commit spans from BIDX[i-1] to\n+      BIDX[i] (plus header length), where BIDX[-1] is 0.\n+    * The BIDX chunk is ignored if the BDAT chunk is not present.\n+\n+  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n+    * It starts with header consisting of three unsigned 32-bit integers:\n+      - Version of the hash algorithm being used. We currently only support\n+\tvalue 1 which implies the murmur3 hash implemented exactly as described\n+\tin https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n+      - The number of times a path is hashed and hence the number of bit positions\n+\tthat cumulatively determine whether a file is present in the commit.\n+      - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n+\tcontains 'n' entries, then the filter size is the minimum number of 64-bit\n+\twords that contain n*b bits.\n+    * The rest of the chunk is the concatenation of all the computed Bloom\n+      filters for the commits in lexicographic order.\n+    * The BDAT chunk is present iff BIDX is present.\n+\n   Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n       This list of H-byte hashes describe a set of B commit-graph files that\n       form a commit-graph chain. The graph position for the ith commit in this\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 32a315058f..4585b3b702 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -24,8 +24,10 @@\n #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n+#define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n+#define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 5\n+#define MAX_NUM_CHUNKS 7\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -325,6 +327,32 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n \t\t\t\tchunk_repeated = 1;\n \t\t\telse\n \t\t\t\tgraph->chunk_base_graphs = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_BLOOMINDEXES:\n+\t\t\tif (graph->chunk_bloom_indexes)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_bloom_indexes = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_BLOOMDATA:\n+\t\t\tif (graph->chunk_bloom_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse {\n+\t\t\t\tuint32_t hash_version;\n+\t\t\t\tgraph->chunk_bloom_data = data + chunk_offset;\n+\t\t\t\thash_version = get_be32(data + chunk_offset);\n+\n+\t\t\t\tif (hash_version != 1)\n+\t\t\t\t\tbreak;\n+\n+\t\t\t\tgraph->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n+\t\t\t\tgraph->bloom_filter_settings->hash_version = hash_version;\n+\t\t\t\tgraph->bloom_filter_settings->num_hashes = get_be32(data + chunk_offset + 4);\n+\t\t\t\tgraph->bloom_filter_settings->bits_per_entry = get_be32(data + chunk_offset + 8);\n+\t\t\t}\n+\t\t\tbreak;\n \t\t}\n \n \t\tif (chunk_repeated) {\n@@ -343,6 +371,17 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n \t\tlast_chunk_offset = chunk_offset;\n \t}\n \n+\t/* We need both the bloom chunks to exist together. Else ignore the data */\n+\tif ((graph->chunk_bloom_indexes && !graph->chunk_bloom_data)\n+\t\t || (!graph->chunk_bloom_indexes && graph->chunk_bloom_data)) {\n+\t\tgraph->chunk_bloom_indexes = NULL;\n+\t\tgraph->chunk_bloom_data = NULL;\n+\t\tgraph->bloom_filter_settings = NULL;\n+\t}\n+\n+\tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data)\n+\t\tload_bloom_filters();\n+\n \thashcpy(graph->oid.hash, graph->data + graph->data_len - graph->hash_len);\n \n \tif (verify_commit_graph_lite(graph)) {\n@@ -1040,6 +1079,59 @@ static void write_graph_chunk_extra_edges(struct hashfile *f,\n \t}\n }\n \n+static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n+\t\t\t\t\t    struct write_commit_graph_context *ctx)\n+{\n+\tstruct commit **list = ctx->commits.list;\n+\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n+\tuint32_t cur_pos = 0;\n+\tstruct progress *progress = NULL;\n+\tint i = 0;\n+\n+\tif (ctx->report_progress)\n+\t\tprogress = start_delayed_progress(\n+\t\t\t_(\"Writing changed paths Bloom filters index\"),\n+\t\t\tctx->commits.nr);\n+\n+\twhile (list < last) {\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tcur_pos += filter->len;\n+\t\tdisplay_progress(progress, ++i);\n+\t\thashwrite_be32(f, cur_pos);\n+\t\tlist++;\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n+static void write_graph_chunk_bloom_data(struct hashfile *f,\n+\t\t\t\t\t struct write_commit_graph_context *ctx,\n+\t\t\t\t\t struct bloom_filter_settings *settings)\n+{\n+\tstruct commit **list = ctx->commits.list;\n+\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n+\tstruct progress *progress = NULL;\n+\tint i = 0;\n+\n+\tif (ctx->report_progress)\n+\t\tprogress = start_delayed_progress(\n+\t\t\t_(\"Writing changed paths Bloom filters data\"),\n+\t\t\tctx->commits.nr);\n+\n+\thashwrite_be32(f, settings->hash_version);\n+\thashwrite_be32(f, settings->num_hashes);\n+\thashwrite_be32(f, settings->bits_per_entry);\n+\n+\twhile (list < last) {\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tdisplay_progress(progress, ++i);\n+\t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n+\t\tlist++;\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n static int oid_compare(const void *_a, const void *_b)\n {\n \tconst struct object_id *a = (const struct object_id *)_a;\n@@ -1198,8 +1290,8 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \tload_bloom_filters();\n \n \tif (ctx->report_progress)\n-\t\tprogress = start_progress(\n-\t\t\t_(\"Computing commit diff Bloom filters\"),\n+\t\tprogress = start_delayed_progress(\n+\t\t\t_(\"Computing changed paths Bloom filters\"),\n \t\t\tctx->commits.nr);\n \n \tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n@@ -1444,6 +1536,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tstruct strbuf progress_title = STRBUF_INIT;\n \tint num_chunks = 3;\n \tstruct object_id file_hash;\n+\tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \n \tif (ctx->split) {\n \t\tstruct strbuf tmp_file = STRBUF_INIT;\n@@ -1488,6 +1581,12 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tnum_chunks++;\n \t}\n+\tif (ctx->changed_paths) {\n+\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMINDEXES;\n+\t\tnum_chunks++;\n+\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMDATA;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_commit_graphs_after > 1) {\n \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BASE;\n \t\tnum_chunks++;\n@@ -1506,6 +1605,15 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t\t\t\t\t\t4 * ctx->num_extra_edges;\n \t\tnum_chunks++;\n \t}\n+\tif (ctx->changed_paths) {\n+\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n+\t\t\t\t\t\tsizeof(uint32_t) * ctx->commits.nr;\n+\t\tnum_chunks++;\n+\n+\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n+\t\t\t\t\t\tsizeof(uint32_t) * 3 + ctx->total_bloom_filter_data_size;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_commit_graphs_after > 1) {\n \t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n \t\t\t\t\t\thashsz * (ctx->num_commit_graphs_after - 1);\n@@ -1543,6 +1651,10 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \twrite_graph_chunk_data(f, hashsz, ctx);\n \tif (ctx->num_extra_edges)\n \t\twrite_graph_chunk_extra_edges(f, ctx);\n+\tif (ctx->changed_paths) {\n+\t\twrite_graph_chunk_bloom_indexes(f, ctx);\n+\t\twrite_graph_chunk_bloom_data(f, ctx, &bloom_settings);\n+\t}\n \tif (ctx->num_commit_graphs_after > 1 &&\n \t    write_graph_chunk_base(f, ctx)) {\n \t\treturn -1;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 952a4b83be..25fefefb3e 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -10,6 +10,7 @@\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n \n struct commit;\n+struct bloom_filter_settings;\n \n char *get_commit_graph_filename(const char *obj_dir);\n int open_commit_graph(const char *graph_file, int *fd, struct stat *st);\n@@ -58,6 +59,10 @@ struct commit_graph {\n \tconst unsigned char *chunk_commit_data;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n+\tconst unsigned char *chunk_bloom_indexes;\n+\tconst unsigned char *chunk_bloom_data;\n+\n+\tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n struct commit_graph *load_commit_graph_one_fd_st(int fd, struct stat *st);\n@@ -77,7 +82,7 @@ enum commit_graph_write_flags {\n \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n \t/* Make sure that each OID in the input is a valid commit OID. */\n \tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n-\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n+\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4),\n };\n \n struct split_commit_graph_opts {\n-- \ngitgitgadget\n\n"},{"id":"391209","messageId":"b20c8d2b2096bf10fe1a5f37a5181c57873a9676.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 08/11] commit-graph: reuse existing Bloom filters during write.","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:27Z","receivedAt":"2020-02-05T22:56:47Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nRead previously computed Bloom filters from the commit-graph file if\npossible to avoid recomputing during commit-graph write.\n\nSee Documentation/technical/commit-graph-format for the format in which\nthe Bloom filter information is written to the commit graph file.\n\nTo read Bloom filter for a given commit with lexicographic position\n'i' we need to:\n1. Read BIDX[i] which essentially gives us the starting index in BDAT for\n   filter of commit i+1. It is essentially the index past the end\n   of the filter of commit i. It is called end_index in the code.\n\n2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT\n   for filter of commit i. It is called the start_index in the code.\n   For the first commit, where i = 0, Bloom filter data starts at the\n   beginning, just past the header in the BDAT chunk. Hence, start_index\n   will be 0.\n\n3. The length of the filter will be end_index - start_index, because\n   BIDX[i] gives the cumulative 8-byte words including the ith\n   commit's filter.\n\nWe toggle whether Bloom filters should be recomputed based on the\ncompute_if_null flag.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c               | 49 ++++++++++++++++++++++++++++++++++++++++++-\n bloom.h               |  4 +++-\n commit-graph.c        |  7 ++++---\n t/helper/test-bloom.c |  2 +-\n 4 files changed, 56 insertions(+), 6 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 818382c03b..90d84dc713 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -1,5 +1,7 @@\n #include \"git-compat-util.h\"\n #include \"bloom.h\"\n+#include \"commit.h\"\n+#include \"commit-slab.h\"\n #include \"commit-graph.h\"\n #include \"object-store.h\"\n #include \"diff.h\"\n@@ -127,8 +129,39 @@ void add_key_to_filter(struct bloom_key *key,\n \t}\n }\n \n+static int load_bloom_filter_from_graph(struct commit_graph *g,\n+\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t   struct commit *c)\n+{\n+\tuint32_t lex_pos, start_index, end_index;\n+\n+\twhile (c->graph_pos < g->num_commits_in_base)\n+\t\tg = g->base_graph;\n+\n+\t/* The commit graph commit 'c' lives in doesn't carry bloom filters. */\n+\tif (!g->chunk_bloom_indexes)\n+\t\treturn 0;\n+\n+\tlex_pos = c->graph_pos - g->num_commits_in_base;\n+\n+\tend_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n+\n+\tif (lex_pos)\n+\t\tstart_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n+\telse\n+\t\tstart_index = 0;\n+\n+\tfilter->len = end_index - start_index;\n+\tfilter->data = (uint64_t *)(g->chunk_bloom_data +\n+\t\t\t\t\tsizeof(uint64_t) * start_index +\n+\t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n+\n+\treturn 1;\n+}\n+\n struct bloom_filter *get_bloom_filter(struct repository *r,\n-\t\t\t\t      struct commit *c)\n+\t\t\t\t      struct commit *c,\n+\t\t\t\t      int compute_if_not_present)\n {\n \tstruct bloom_filter *filter;\n \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n@@ -141,6 +174,20 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n \n+\tif (!filter->data) {\n+\t\tload_commit_graph_info(r, c);\n+\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH &&\n+\t\t\tr->objects->commit_graph->chunk_bloom_indexes) {\n+\t\t\tif (load_bloom_filter_from_graph(r->objects->commit_graph, filter, c))\n+\t\t\t\treturn filter;\n+\t\t\telse\n+\t\t\t\treturn NULL;\n+\t\t}\n+\t}\n+\n+\tif (filter->data || !compute_if_not_present)\n+\t\treturn filter;\n+\n \trepo_diff_setup(r, &diffopt);\n \tdiffopt.flags.recursive = 1;\n \tdiffopt.max_changes = max_changes;\ndiff --git a/bloom.h b/bloom.h\nindex 7f40c751f7..76f8a9ad0c 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -13,6 +13,7 @@ struct bloom_filter_settings {\n \n #define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n #define BITS_PER_WORD 64\n+#define BLOOMDATA_CHUNK_HEADER_SIZE 3*sizeof(uint32_t)\n \n /*\n  * A bloom_filter struct represents a data segment to\n@@ -47,7 +48,8 @@ void add_key_to_filter(struct bloom_key *key,\n \t\t\t\t\t   struct bloom_filter_settings *settings);\n \n struct bloom_filter *get_bloom_filter(struct repository *r,\n-\t\t\t\t      struct commit *c);\n+\t\t\t\t      struct commit *c,\n+\t\t\t\t      int compute_if_not_present);\n \n int bloom_filter_contains(struct bloom_filter *filter,\n \t\t\t  struct bloom_key *key,\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 4585b3b702..c0e9834bf2 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1094,7 +1094,7 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n \t\t\tctx->commits.nr);\n \n \twhile (list < last) {\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n \t\tcur_pos += filter->len;\n \t\tdisplay_progress(progress, ++i);\n \t\thashwrite_be32(f, cur_pos);\n@@ -1123,7 +1123,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n \thashwrite_be32(f, settings->bits_per_entry);\n \n \twhile (list < last) {\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n \t\tdisplay_progress(progress, ++i);\n \t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n \t\tlist++;\n@@ -1304,7 +1304,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tstruct commit *c = sorted_by_pos[i];\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n \t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n \t\tdisplay_progress(progress, i + 1);\n \t}\n@@ -2314,6 +2314,7 @@ void free_commit_graph(struct commit_graph *g)\n \t\tg->data = NULL;\n \t\tclose(g->graph_fd);\n \t}\n+\tfree(g->bloom_filter_settings);\n \tfree(g->filename);\n \tfree(g);\n }\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 331957011b..9b4be97f75 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -47,7 +47,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n \tstruct bloom_filter *filter;\n \tsetup_git_directory();\n \tc = lookup_commit(the_repository, commit_oid);\n-\tfilter = get_bloom_filter(the_repository, c);\n+\tfilter = get_bloom_filter(the_repository, c, 1);\n \tprint_bloom_filter(filter);\n }\n \n-- \ngitgitgadget\n\n"},{"id":"391210","messageId":"e1b076a714d611e59d3d71c89221e41a3427fae4.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 11/11] commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:30Z","receivedAt":"2020-02-05T22:56:49Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag to the test setup suite\nin order to toggle writing Bloom filters when running any of the git tests.\nIf set to true, we will compute and write Bloom filters every time a test\ncalls `git commit-graph write`, as if the `--changed-paths` option was\npassed in.\n\nThe test suite passes when GIT_TEST_COMMIT_GRAPH and\nGIT_TEST_COMMIT_GRAPH_CHANGED_PATHS are enabled.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n builtin/commit-graph.c        | 3 ++-\n ci/run-build-and-tests.sh     | 1 +\n commit-graph.h                | 1 +\n t/README                      | 5 +++++\n t/t4216-log-bloom.sh          | 3 +++\n t/t5318-commit-graph.sh       | 2 ++\n t/t5324-split-commit-graph.sh | 1 +\n 7 files changed, 15 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\nindex 261dcce091..fc9b234ab0 100644\n--- a/builtin/commit-graph.c\n+++ b/builtin/commit-graph.c\n@@ -146,7 +146,8 @@ static int graph_write(int argc, const char **argv)\n \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n \tif (opts.progress)\n \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n-\tif (opts.enable_changed_paths)\n+\tif (opts.enable_changed_paths ||\n+\t    git_env_bool(GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS, 0))\n \t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n \n \tread_replace_refs = 0;\ndiff --git a/ci/run-build-and-tests.sh b/ci/run-build-and-tests.sh\nindex ff0ef7f08e..7b4857651d 100755\n--- a/ci/run-build-and-tests.sh\n+++ b/ci/run-build-and-tests.sh\n@@ -19,6 +19,7 @@ linux-gcc)\n \texport GIT_TEST_OE_SIZE=10\n \texport GIT_TEST_OE_DELTA_SIZE=5\n \texport GIT_TEST_COMMIT_GRAPH=1\n+\texport GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=1\n \texport GIT_TEST_MULTI_PACK_INDEX=1\n \tmake test\n \t;;\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 25fefefb3e..4c202ff3d7 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -8,6 +8,7 @@\n \n #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n+#define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n \n struct commit;\n struct bloom_filter_settings;\ndiff --git a/t/README b/t/README\nindex caa125ba9a..be2f7d7fd2 100644\n--- a/t/README\n+++ b/t/README\n@@ -378,6 +378,11 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n+commit-graph write to compute and write changed path Bloom filters for\n+every 'git commit-graph write', as if the `--changed-paths` option was\n+passed in.\n+\n GIT_TEST_FSMONITOR=$PWD/t7519/fsmonitor-all exercises the fsmonitor\n code path for utilizing a file system monitor to speed up detecting\n new or changed files.\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 19eca1864b..7acebb3962 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -3,6 +3,9 @@\n test_description='git log for a path with bloom filters'\n . ./test-lib.sh\n \n+GIT_TEST_COMMIT_GRAPH=0\n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n+\n test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n \tgit init &&\n \tmkdir A A/B A/B/C &&\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 3f03de6018..973020be2d 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -3,6 +3,8 @@\n test_description='commit graph'\n . ./test-lib.sh\n \n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n+\n test_expect_success 'setup full repo' '\n \tmkdir full &&\n \tcd \"$TRASH_DIRECTORY/full\" &&\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex c24823431f..9235db4561 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -4,6 +4,7 @@ test_description='split commit graph'\n . ./test-lib.sh\n \n GIT_TEST_COMMIT_GRAPH=0\n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n \n test_expect_success 'setup repo' '\n \tgit init &&\n-- \ngitgitgadget\n"},{"id":"391211","messageId":"77f1c561e8205c0598b57bf572640d21d64757f8.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 10/11] revision.c: use Bloom filters to speed up path based revision walks","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:29Z","receivedAt":"2020-02-05T22:56:49Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nRevision walk will now use Bloom filters for commits to speed up revision\nwalks for a particular path (for computing history for that path), if they\nare present in the commit-graph file.\n\nWe load the Bloom filters during the prepare_revision_walk step, but only\nwhen dealing with a single pathspec. While comparing trees in\nrev_compare_trees(), if the Bloom filter says that the file is not different\nbetween the two trees, we don't need to compute the expensive diff. This is\nwhere we get our performance gains. The other response of the Bloom filter\nis `maybe`, in which case we fall back to the full diff calculation to\ndetermine if the path was changed in the commit.\n\nPerformance Gains:\nWe tested the performance of `git log -- <path>` on the git repo, the linux\nand some internal large repos, with a variety of paths of varying depths.\n\nOn the git and linux repos:\n- we observed a 2x to 5x speed up.\n\nOn a large internal repo with files seated 6-10 levels deep in the tree:\n- we observed 10x to 20x speed ups, with some paths going up to 28 times\n  faster.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com\nHelped-by: SZEDER Gábor <szeder.dev@gmail.com>\nHelped-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n revision.c                 | 124 +++++++++++++++++++++++++++++++-\n revision.h                 |  11 +++\n t/helper/test-read-graph.c |   4 ++\n t/t4216-log-bloom.sh       | 140 +++++++++++++++++++++++++++++++++++++\n 4 files changed, 277 insertions(+), 2 deletions(-)\n create mode 100755 t/t4216-log-bloom.sh\n\ndiff --git a/revision.c b/revision.c\nindex 8136929e23..d1622afa17 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -29,6 +29,8 @@\n #include \"prio-queue.h\"\n #include \"hashmap.h\"\n #include \"utf8.h\"\n+#include \"bloom.h\"\n+#include \"json-writer.h\"\n \n volatile show_early_output_fn_t show_early_output;\n \n@@ -624,11 +626,114 @@ static void file_change(struct diff_options *options,\n \toptions->flags.has_changes = 1;\n }\n \n+static int bloom_filter_atexit_registered;\n+static unsigned int count_bloom_filter_maybe;\n+static unsigned int count_bloom_filter_definitely_not;\n+static unsigned int count_bloom_filter_false_positive;\n+static unsigned int count_bloom_filter_not_present;\n+static unsigned int count_bloom_filter_length_zero;\n+\n+static void trace2_bloom_filter_statistics_atexit(void)\n+{\n+\tstruct json_writer jw = JSON_WRITER_INIT;\n+\n+\tjw_object_begin(&jw, 0);\n+\tjw_object_intmax(&jw, \"filter_not_present\", count_bloom_filter_not_present);\n+\tjw_object_intmax(&jw, \"zero_length_filter\", count_bloom_filter_length_zero);\n+\tjw_object_intmax(&jw, \"maybe\", count_bloom_filter_maybe);\n+\tjw_object_intmax(&jw, \"definitely_not\", count_bloom_filter_definitely_not);\n+\tjw_end(&jw);\n+\n+\ttrace2_data_json(\"bloom\", the_repository, \"statistics\", &jw);\n+\n+\tjw_release(&jw);\n+}\n+\n+static void prepare_to_use_bloom_filter(struct rev_info *revs)\n+{\n+\tstruct pathspec_item *pi;\n+\tchar *path_alloc = NULL;\n+\tconst char *path;\n+\tint last_index;\n+\tint len;\n+\n+\tif (!revs->commits)\n+\t    return;\n+\n+\trepo_parse_commit(revs->repo, revs->commits->item);\n+\n+\tif (!revs->repo->objects->commit_graph)\n+\t\treturn;\n+\n+\trevs->bloom_filter_settings = revs->repo->objects->commit_graph->bloom_filter_settings;\n+\tif (!revs->bloom_filter_settings)\n+\t\treturn;\n+\n+\tpi = &revs->pruning.pathspec.items[0];\n+\tlast_index = pi->len - 1;\n+\n+\tif (pi->match[last_index] == '/') {\n+\t    path_alloc = xstrdup(pi->match);\n+\t    path_alloc[last_index] = '\\0';\n+\t    path = path_alloc;\n+\t} else\n+\t    path = pi->match;\n+\n+\tlen = strlen(path);\n+\n+\trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n+\tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n+\n+\tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n+\t\tatexit(trace2_bloom_filter_statistics_atexit);\n+\t\tbloom_filter_atexit_registered = 1;\n+\t}\n+\n+\tfree(path_alloc);\n+}\n+\n+static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n+\t\t\t\t\t\t struct commit *commit)\n+{\n+\tstruct bloom_filter *filter;\n+\tint result;\n+\n+\tif (!revs->repo->objects->commit_graph)\n+\t\treturn -1;\n+\n+\tif (commit->generation == GENERATION_NUMBER_INFINITY)\n+\t\treturn -1;\n+\n+\tfilter = get_bloom_filter(revs->repo, commit, 0);\n+\n+\tif (!filter) {\n+\t\tcount_bloom_filter_not_present++;\n+\t\treturn -1;\n+\t}\n+\n+\tif (!filter->len) {\n+\t\tcount_bloom_filter_length_zero++;\n+\t\treturn -1;\n+\t}\n+\n+\tresult = bloom_filter_contains(filter,\n+\t\t\t\t       revs->bloom_key,\n+\t\t\t\t       revs->bloom_filter_settings);\n+\n+\tif (result)\n+\t\tcount_bloom_filter_maybe++;\n+\telse\n+\t\tcount_bloom_filter_definitely_not++;\n+\n+\treturn result;\n+}\n+\n static int rev_compare_tree(struct rev_info *revs,\n-\t\t\t    struct commit *parent, struct commit *commit)\n+\t\t\t    struct commit *parent, struct commit *commit, int nth_parent)\n {\n \tstruct tree *t1 = get_commit_tree(parent);\n \tstruct tree *t2 = get_commit_tree(commit);\n+\tint bloom_ret = 1;\n \n \tif (!t1)\n \t\treturn REV_TREE_NEW;\n@@ -653,11 +758,23 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\treturn REV_TREE_SAME;\n \t}\n \n+\tif (revs->pruning.pathspec.nr == 1 && !revs->reflog_info && !nth_parent) {\n+\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n+\n+\t\tif (bloom_ret == 0)\n+\t\t\treturn REV_TREE_SAME;\n+\t}\n+\n \ttree_difference = REV_TREE_SAME;\n \trevs->pruning.flags.has_changes = 0;\n \tif (diff_tree_oid(&t1->object.oid, &t2->object.oid, \"\",\n \t\t\t   &revs->pruning) < 0)\n \t\treturn REV_TREE_DIFFERENT;\n+\n+\tif (!nth_parent)\n+\t\tif (bloom_ret == 1 && tree_difference == REV_TREE_SAME)\n+\t\t\tcount_bloom_filter_false_positive++;\n+\n \treturn tree_difference;\n }\n \n@@ -855,7 +972,7 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit)\n \t\t\tdie(\"cannot simplify commit %s (because of %s)\",\n \t\t\t    oid_to_hex(&commit->object.oid),\n \t\t\t    oid_to_hex(&p->object.oid));\n-\t\tswitch (rev_compare_tree(revs, p, commit)) {\n+\t\tswitch (rev_compare_tree(revs, p, commit, nth_parent)) {\n \t\tcase REV_TREE_SAME:\n \t\t\tif (!revs->simplify_history || !relevant_commit(p)) {\n \t\t\t\t/* Even if a merge with an uninteresting\n@@ -3362,6 +3479,8 @@ int prepare_revision_walk(struct rev_info *revs)\n \t\t\t\t       FOR_EACH_OBJECT_PROMISOR_ONLY);\n \t}\n \n+\tif (revs->pruning.pathspec.nr == 1 && !revs->reflog_info)\n+\t\tprepare_to_use_bloom_filter(revs);\n \tif (revs->no_walk != REVISION_WALK_NO_WALK_UNSORTED)\n \t\tcommit_list_sort_by_date(&revs->commits);\n \tif (revs->no_walk)\n@@ -3379,6 +3498,7 @@ int prepare_revision_walk(struct rev_info *revs)\n \t\tsimplify_merges(revs);\n \tif (revs->children.name)\n \t\tset_children(revs);\n+\n \treturn 0;\n }\n \ndiff --git a/revision.h b/revision.h\nindex 475f048fb6..7c026fe41f 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -56,6 +56,8 @@ struct repository;\n struct rev_info;\n struct string_list;\n struct saved_parents;\n+struct bloom_key;\n+struct bloom_filter_settings;\n define_shared_commit_slab(revision_sources, char *);\n \n struct rev_cmdline_info {\n@@ -291,6 +293,15 @@ struct rev_info {\n \tstruct revision_sources *sources;\n \n \tstruct topo_walk_info *topo_walk_info;\n+\n+\t/* Commit graph bloom filter fields */\n+\t/* The bloom filter key for the pathspec */\n+\tstruct bloom_key *bloom_key;\n+\t/*\n+\t * The bloom filter settings used to generate the key.\n+\t * This is loaded from the commit-graph being used.\n+\t */\n+\tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n int ref_excluded(struct string_list *, const char *path);\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex d2884efe0a..aff597c7a3 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -45,6 +45,10 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" commit_metadata\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n+\tif (graph->chunk_bloom_indexes)\n+\t\tprintf(\" bloom_indexes\");\n+\tif (graph->chunk_bloom_data)\n+\t\tprintf(\" bloom_data\");\n \tprintf(\"\\n\");\n \n \tUNLEAK(graph);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nnew file mode 100755\nindex 0000000000..19eca1864b\n--- /dev/null\n+++ b/t/t4216-log-bloom.sh\n@@ -0,0 +1,140 @@\n+#!/bin/sh\n+\n+test_description='git log for a path with bloom filters'\n+. ./test-lib.sh\n+\n+test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n+\tgit init &&\n+\tmkdir A A/B A/B/C &&\n+\ttest_commit c1 A/file1 &&\n+\ttest_commit c2 A/B/file2 &&\n+\ttest_commit c3 A/B/C/file3 &&\n+\ttest_commit c4 A/file1 &&\n+\ttest_commit c5 A/B/file2 &&\n+\ttest_commit c6 A/B/C/file3 &&\n+\ttest_commit c7 A/file1 &&\n+\ttest_commit c8 A/B/file2 &&\n+\ttest_commit c9 A/B/C/file3 &&\n+\tgit checkout -b side HEAD~4 &&\n+\ttest_commit side-1 file4 &&\n+\tgit checkout master &&\n+\tgit merge side &&\n+\ttest_commit c10 file5 &&\n+\tmv file5 file5_renamed &&\n+\tgit add file5_renamed &&\n+\tgit commit -m \"rename\" &&\n+\tgit commit-graph write --reachable --changed-paths\n+'\n+graph_read_expect() {\n+\tOPTIONAL=\"\"\n+\tNUM_CHUNKS=5\n+\tcat >expect <<- EOF\n+\theader: 43475048 1 1 $NUM_CHUNKS 0\n+\tnum_commits: $1\n+\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tEOF\n+\ttest-tool read-graph >output &&\n+\ttest_cmp expect output\n+}\n+\n+test_expect_success 'commit-graph write wrote out the bloom chunks' '\n+\tgraph_read_expect 13\n+'\n+\n+setup() {\n+\trm output\n+\trm \"$TRASH_DIRECTORY/trace.perf\"\n+\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom\n+\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n+}\n+\n+test_bloom_filters_used() {\n+\tlog_args=$1\n+\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":0,\\\"zero_length_filter\\\":0,\\\"maybe\\\"\"\n+\tsetup \"$log_args\"\n+\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && test_cmp log_wo_bloom log_w_bloom\n+}\n+\n+test_bloom_filters_not_used() {\n+\tlog_args=$1\n+\tsetup \"$log_args\"\n+\t!(grep -q \"statistics:{\\\"filter_not_present\\\":\" \"$TRASH_DIRECTORY/trace.perf\") && test_cmp log_wo_bloom log_w_bloom\n+}\n+\n+for path in A A/B A/B/C A/file1 A/B/file2 A/B/C/file3 file4 file5_renamed\n+do\n+\tfor option in \"\" \\\n+\t\t      \"--full-history\" \\\n+\t\t      \"--full-history --simplify-merges\" \\\n+\t\t      \"--simplify-merges\" \\\n+\t\t      \"--simplify-by-decoration\" \\\n+\t\t      \"--follow\" \\\n+\t\t      \"--first-parent\" \\\n+\t\t      \"--topo-order\" \\\n+\t\t      \"--date-order\" \\\n+\t\t      \"--author-date-order\" \\\n+\t\t      \"--ancestry-path side..master\"\n+\tdo\n+\t\ttest_expect_success \"git log option: $option for path: $path\" '\n+\t\t\ttest_bloom_filters_used \"$option -- $path\"\n+\t\t'\n+\tdone\n+done\n+\n+test_expect_success 'git log -- folder works with and without the trailing slash' '\n+\ttest_bloom_filters_used \"-- A\" &&\n+\ttest_bloom_filters_used \"-- A/\"\n+'\n+\n+test_expect_success 'git log for path that does not exist. ' '\n+\ttest_bloom_filters_used \"-- path_does_not_exist\"\n+'\n+\n+test_expect_success 'git log with --walk-reflogs does not use bloom filters' '\n+\ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n+'\n+\n+test_expect_success 'git log -- multiple path specs does not use bloom filters' '\n+\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n+'\n+\n+test_expect_success 'git log with wildcard that resolves to a single path uses bloom filters' '\n+\ttest_bloom_filters_used \"-- *4\" &&\n+\ttest_bloom_filters_used \"-- *renamed\"\n+'\n+\n+test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses bloom filters' '\n+\ttest_bloom_filters_not_used \"-- *\" &&\n+\ttest_bloom_filters_not_used \"-- file*\"\n+'\n+\n+test_expect_success 'setup - add commit-graph to the chain without bloom filters' '\n+\ttest_commit c14 A/anotherFile2 &&\n+\ttest_commit c15 A/B/anotherFile2 &&\n+\ttest_commit c16 A/B/C/anotherFile2 &&\n+\tGIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0 git commit-graph write --reachable --split &&\n+\ttest_line_count = 2 .git/objects/info/commit-graphs/commit-graph-chain\n+'\n+\n+test_expect_success 'git log does not use bloom filters if the latest graph does not have bloom filters.' '\n+\ttest_bloom_filters_not_used \"-- A/B\"\n+'\n+\n+test_expect_success 'setup - add commit-graph to the chain with bloom filters' '\n+\ttest_commit c17 A/anotherFile3 &&\n+\tgit commit-graph write --reachable --changed-paths --split &&\n+\ttest_line_count = 3 .git/objects/info/commit-graphs/commit-graph-chain\n+'\n+\n+test_bloom_filters_used_when_some_filters_are_missing() {\n+\tlog_args=$1\n+\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":3,\\\"zero_length_filter\\\":0,\\\"maybe\\\":6,\\\"definitely_not\\\":6\"\n+\tsetup \"$log_args\"\n+\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && test_cmp log_wo_bloom log_w_bloom\n+}\n+\n+test_expect_success 'git log uses bloom filters if they exist in the latest but not all commit graphs in the chain.' '\n+\ttest_bloom_filters_used_when_some_filters_are_missing \"-- A/B\"\n+'\n+\n+test_done\n-- \ngitgitgadget\n\n"},{"id":"391212","messageId":"3d7ee0c96955dc15c87d04982d8cdec8b62750b2.1580943390.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v2 09/11] commit-graph: add --changed-paths option to write subcommand","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-02-05T22:56:28Z","receivedAt":"2020-02-05T22:56:49Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd --changed-paths option to git commit-graph write. This option will\nallow users to compute information about the paths that have changed\nbetween a commit and its first parent, and write it into the commit graph\nfile. If the option is passed to the write subcommand we set the\nCOMMIT_GRAPH_WRITE_BLOOM_FILTERS flag and pass it down to the\ncommit-graph logic.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n Documentation/git-commit-graph.txt | 5 +++++\n builtin/commit-graph.c             | 9 +++++++--\n 2 files changed, 12 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-commit-graph.txt b/Documentation/git-commit-graph.txt\nindex bcd85c1976..907d703b30 100644\n--- a/Documentation/git-commit-graph.txt\n+++ b/Documentation/git-commit-graph.txt\n@@ -54,6 +54,11 @@ or `--stdin-packs`.)\n With the `--append` option, include all commits that are present in the\n existing commit-graph file.\n +\n+With the `--changed-paths` option, compute and write information about the\n+paths changed between a commit and it's first parent. This operation can\n+take a while on large repositories. It provides significant performance gains\n+for getting history of a directory or a file with `git log -- <path>`.\n++\n With the `--split` option, write the commit-graph as a chain of multiple\n commit-graph files stored in `<dir>/info/commit-graphs`. The new commits\n not already in the commit-graph are added in a new \"tip\" file. This file\ndiff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\nindex e0c6fc4bbf..261dcce091 100644\n--- a/builtin/commit-graph.c\n+++ b/builtin/commit-graph.c\n@@ -9,7 +9,7 @@\n \n static char const * const builtin_commit_graph_usage[] = {\n \tN_(\"git commit-graph verify [--object-dir <objdir>] [--shallow] [--[no-]progress]\"),\n-\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n+\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n \tNULL\n };\n \n@@ -19,7 +19,7 @@ static const char * const builtin_commit_graph_verify_usage[] = {\n };\n \n static const char * const builtin_commit_graph_write_usage[] = {\n-\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n+\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n \tNULL\n };\n \n@@ -32,6 +32,7 @@ static struct opts_commit_graph {\n \tint split;\n \tint shallow;\n \tint progress;\n+\tint enable_changed_paths;\n } opts;\n \n static int graph_verify(int argc, const char **argv)\n@@ -110,6 +111,8 @@ static int graph_write(int argc, const char **argv)\n \t\t\tN_(\"start walk at commits listed by stdin\")),\n \t\tOPT_BOOL(0, \"append\", &opts.append,\n \t\t\tN_(\"include all commits already in the commit-graph file\")),\n+\t\tOPT_BOOL(0, \"changed-paths\", &opts.enable_changed_paths,\n+\t\t\tN_(\"enable computation for changed paths\")),\n \t\tOPT_BOOL(0, \"progress\", &opts.progress, N_(\"force progress reporting\")),\n \t\tOPT_BOOL(0, \"split\", &opts.split,\n \t\t\tN_(\"allow writing an incremental commit-graph file\")),\n@@ -143,6 +146,8 @@ static int graph_write(int argc, const char **argv)\n \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n \tif (opts.progress)\n \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n+\tif (opts.enable_changed_paths)\n+\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n \n \tread_replace_refs = 0;\n \n-- \ngitgitgadget\n\n"},{"id":"391323","messageId":"20200207135249.GD2868@szeder.dev","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 00/11] Changed Paths Bloom Filters","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-02-07T13:52:50Z","receivedAt":"2020-02-07T13:52:58Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Wed, Feb 05, 2020 at 10:56:19PM +0000, Garima Singh via GitGitGadget wrote:\n> Hey! \n> \n> The commit graph feature brought in a lot of performance improvements across\n> multiple commands. However, file based history continues to be a performance\n> pain point, especially in large repositories. \n> \n> Adopting changed path bloom filters has been discussed on the list before,\n> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n> Derrick Stolee [1]. This series is based on Dr. Stolee's proof of concept in\n> [2]\n> \n> Performance Gains: We tested the performance of git log -- path on the git\n> repo, the linux repo and some internal large repos, with a variety of paths\n> of varying depths.\n> \n> On the git and linux repos: We observed a 2x to 5x speed up.\n> \n> On a large internal repo with files seated 6-10 levels deep in the tree: We\n> observed 10x to 20x speed ups, with some paths going up to 28 times faster.\n> \n> Future Work (not included in the scope of this series):\n> \n>  1. Supporting multiple path based revision walk\n>  2. Adopting it in git blame logic. \n>  3. Interactions with line log git log -L\n> \n> \n> ----------------------------------------------------------------------------\n> \n> Updates since the last submission\n> \n>  * Removed all the RFC callouts, this is a ready for full review version\n\nDon't know when I'll find enough time to properly review the series.\nmaybe someday...\n\n>  * Added unit tests for the bloom filter computation layer\n\nThis fails on big endian, e.g. in Travis CI's s390x build:\n\n  https://travis-ci.org/szeder/git-cooking-topics-for-travis-ci/jobs/647253022#L2210\n\n(The link highlights the failure, but I'm afraid your browser won't\njump there right away; you'll have to click on the print-test-failures\nfold at the bottom, and scroll down a bit...)\n"},{"id":"391330","messageId":"140cf2f4-23d5-09ab-8f23-bbbd397c68f7@gmail.com","threadId":"52499","inReplyTo":"20200207135249.GD2868@szeder.dev","subject":"Re: [PATCH v2 00/11] Changed Paths Bloom Filters","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-07T15:09:09Z","receivedAt":"2020-02-07T15:09:12Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/7/2020 8:52 AM, SZEDER Gábor wrote:\n>>  * Added unit tests for the bloom filter computation layer\n> \n> This fails on big endian, e.g. in Travis CI's s390x build:\n> \n>   https://travis-ci.org/szeder/git-cooking-topics-for-travis-ci/jobs/647253022#L2210\n> \n> (The link highlights the failure, but I'm afraid your browser won't\n> jump there right away; you'll have to click on the print-test-failures\n> fold at the bottom, and scroll down a bit...)\n> \n\nThank you so much for running this pipeline and pointing out the error!\n\nWe will carefully review our interactions with the binary data and \nhopefully solve this in the next version. \n\nCheers!\nGarima Singh\n"},{"id":"391332","messageId":"88c8e5da-72f2-25cc-f55b-f62500c52a24@gmail.com","threadId":"52499","inReplyTo":"140cf2f4-23d5-09ab-8f23-bbbd397c68f7@gmail.com","subject":"Re: [PATCH v2 00/11] Changed Paths Bloom Filters","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-02-07T15:36:58Z","receivedAt":"2020-02-07T15:37:15Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 2/7/2020 10:09 AM, Garima Singh wrote:\n> \n> On 2/7/2020 8:52 AM, SZEDER Gábor wrote:\n>>>  * Added unit tests for the bloom filter computation layer\n>>\n>> This fails on big endian, e.g. in Travis CI's s390x build:\n>>\n>>   https://travis-ci.org/szeder/git-cooking-topics-for-travis-ci/jobs/647253022#L2210\n>>\n>> (The link highlights the failure, but I'm afraid your browser won't\n>> jump there right away; you'll have to click on the print-test-failures\n>> fold at the bottom, and scroll down a bit...)\n>>\n> \n> Thank you so much for running this pipeline and pointing out the error!\n> \n> We will carefully review our interactions with the binary data and \n> hopefully solve this in the next version. \n\nSzeder,\n\nThanks so much for running this test. We don't have access to a big endian\nmachine right now, so could you please apply this patch and re-run your tests?\n\nThe issue is described in the message below, and Garima is working to ensure\nthe handling of the filter data is clarified in the next version.\n\nThis is an issue from WAY back in the original prototype, and it highlights\nthat we've never been writing the data in network-byte order. This is completely\nmy fault.\n\nThanks,\n-Stolee\n\n\n-->8--\n\nFrom c1067db5d618b2dae430dfe373a11c771517da9e Mon Sep 17 00:00:00 2001\nFrom: Derrick Stolee <dstolee@microsoft.com>\nDate: Fri, 7 Feb 2020 10:24:05 -0500\nSubject: [PATCH] fixup! bloom: core Bloom filter implementation for changed\n paths\n\nThe 'data' field of 'struct bloom_filter' can point to a memory location\n(when computing one before writing to the commit-graph) or a memmap()'d\nfile location (when reading from the Bloom data chunk of the commit-graph\nfile). This means that the memory representation may be backwards in\nLittle Endian or Big Endian machines.\n\nAlways write and read bits from 'filter->data' using network order. This\nallows us to avoid loading the data streams from the file into memory\nbuffers.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\n---\n bloom.c               | 6 ++++--\n t/helper/test-bloom.c | 2 +-\n 2 files changed, 5 insertions(+), 3 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 90d84dc713..aa6896584b 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -124,8 +124,9 @@ void add_key_to_filter(struct bloom_key *key,\n \tfor (i = 0; i < settings->num_hashes; i++) {\n \t\tuint64_t hash_mod = key->hashes[i] % mod;\n \t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n+\t\tuint64_t bit = get_bitmask(hash_mod);\n \n-\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n+\t\tfilter->data[block_pos] |= htonll(bit);\n \t}\n }\n \n@@ -269,7 +270,8 @@ int bloom_filter_contains(struct bloom_filter *filter,\n \tfor (i = 0; i < settings->num_hashes; i++) {\n \t\tuint64_t hash_mod = key->hashes[i] % mod;\n \t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n-\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n+\t\tuint64_t bit = get_bitmask(hash_mod);\n+\t\tif (!(filter->data[block_pos] & htonll(bit)))\n \t\t\treturn 0;\n \t}\n \ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 9b4be97f75..09b2bb0a00 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -23,7 +23,7 @@ static void print_bloom_filter(struct bloom_filter *filter) {\n \tprintf(\"Filter_Length:%d\\n\", filter->len);\n \tprintf(\"Filter_Data:\");\n \tfor (i = 0; i < filter->len; i++){\n-\t\tprintf(\"%\"PRIx64\"|\", filter->data[i]);\n+\t\tprintf(\"%\"PRIx64\"|\", ntohll(filter->data[i]));\n \t}\n \tprintf(\"\\n\");\n }\n-- \n2.25.0.vfs.1.1.1.g9906319d24.dirty\n\n\n\n"},{"id":"391335","messageId":"20200207161528.GA18146@szeder.dev","threadId":"52499","inReplyTo":"88c8e5da-72f2-25cc-f55b-f62500c52a24@gmail.com","subject":"Re: [PATCH v2 00/11] Changed Paths Bloom Filters","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-02-07T16:15:28Z","receivedAt":"2020-02-07T16:15:39Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Fri, Feb 07, 2020 at 10:36:58AM -0500, Derrick Stolee wrote:\n> On 2/7/2020 10:09 AM, Garima Singh wrote:\n> > \n> > On 2/7/2020 8:52 AM, SZEDER Gábor wrote:\n> >>>  * Added unit tests for the bloom filter computation layer\n> >>\n> >> This fails on big endian, e.g. in Travis CI's s390x build:\n> >>\n> >>   https://travis-ci.org/szeder/git-cooking-topics-for-travis-ci/jobs/647253022#L2210\n> >>\n> >> (The link highlights the failure, but I'm afraid your browser won't\n> >> jump there right away; you'll have to click on the print-test-failures\n> >> fold at the bottom, and scroll down a bit...)\n> >>\n> > \n> > Thank you so much for running this pipeline and pointing out the error!\n> > \n> > We will carefully review our interactions with the binary data and \n> > hopefully solve this in the next version. \n> \n> Szeder,\n> \n> Thanks so much for running this test. We don't have access to a big endian\n> machine right now, so could you please apply this patch and re-run your tests?\n\nUnfortunately, it still failed:\n\n  https://travis-ci.org/szeder/git-cooking-topics-for-travis-ci/jobs/647395554#L2204\n\n> The issue is described in the message below, and Garima is working to ensure\n> the handling of the filter data is clarified in the next version.\n> \n> This is an issue from WAY back in the original prototype, and it highlights\n> that we've never been writing the data in network-byte order. This is completely\n> my fault.\n> \n> Thanks,\n> -Stolee\n> \n> \n> -->8--\n> \n> From c1067db5d618b2dae430dfe373a11c771517da9e Mon Sep 17 00:00:00 2001\n> From: Derrick Stolee <dstolee@microsoft.com>\n> Date: Fri, 7 Feb 2020 10:24:05 -0500\n> Subject: [PATCH] fixup! bloom: core Bloom filter implementation for changed\n>  paths\n> \n> The 'data' field of 'struct bloom_filter' can point to a memory location\n> (when computing one before writing to the commit-graph) or a memmap()'d\n> file location (when reading from the Bloom data chunk of the commit-graph\n> file). This means that the memory representation may be backwards in\n> Little Endian or Big Endian machines.\n> \n> Always write and read bits from 'filter->data' using network order. This\n> allows us to avoid loading the data streams from the file into memory\n> buffers.\n> \n> Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n> ---\n>  bloom.c               | 6 ++++--\n>  t/helper/test-bloom.c | 2 +-\n>  2 files changed, 5 insertions(+), 3 deletions(-)\n> \n> diff --git a/bloom.c b/bloom.c\n> index 90d84dc713..aa6896584b 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -124,8 +124,9 @@ void add_key_to_filter(struct bloom_key *key,\n>  \tfor (i = 0; i < settings->num_hashes; i++) {\n>  \t\tuint64_t hash_mod = key->hashes[i] % mod;\n>  \t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n> +\t\tuint64_t bit = get_bitmask(hash_mod);\n>  \n> -\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n> +\t\tfilter->data[block_pos] |= htonll(bit);\n>  \t}\n>  }\n>  \n> @@ -269,7 +270,8 @@ int bloom_filter_contains(struct bloom_filter *filter,\n>  \tfor (i = 0; i < settings->num_hashes; i++) {\n>  \t\tuint64_t hash_mod = key->hashes[i] % mod;\n>  \t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n> -\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n> +\t\tuint64_t bit = get_bitmask(hash_mod);\n> +\t\tif (!(filter->data[block_pos] & htonll(bit)))\n>  \t\t\treturn 0;\n>  \t}\n>  \n> diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n> index 9b4be97f75..09b2bb0a00 100644\n> --- a/t/helper/test-bloom.c\n> +++ b/t/helper/test-bloom.c\n> @@ -23,7 +23,7 @@ static void print_bloom_filter(struct bloom_filter *filter) {\n>  \tprintf(\"Filter_Length:%d\\n\", filter->len);\n>  \tprintf(\"Filter_Data:\");\n>  \tfor (i = 0; i < filter->len; i++){\n> -\t\tprintf(\"%\"PRIx64\"|\", filter->data[i]);\n> +\t\tprintf(\"%\"PRIx64\"|\", ntohll(filter->data[i]));\n>  \t}\n>  \tprintf(\"\\n\");\n>  }\n> -- \n> 2.25.0.vfs.1.1.1.g9906319d24.dirty\n> \n> \n> \n"},{"id":"391337","messageId":"49cd87c0-3a5d-09fb-fe08-57b0a7b7a194@gmail.com","threadId":"52499","inReplyTo":"20200207161528.GA18146@szeder.dev","subject":"Re: [PATCH v2 00/11] Changed Paths Bloom Filters","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-02-07T16:33:15Z","receivedAt":"2020-02-07T16:33:21Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 2/7/2020 11:15 AM, SZEDER Gábor wrote:\n> On Fri, Feb 07, 2020 at 10:36:58AM -0500, Derrick Stolee wrote:\n>> On 2/7/2020 10:09 AM, Garima Singh wrote:\n>>>\n>>> On 2/7/2020 8:52 AM, SZEDER Gábor wrote:\n>>>>>  * Added unit tests for the bloom filter computation layer\n>>>>\n>>>> This fails on big endian, e.g. in Travis CI's s390x build:\n>>>>\n>>>>   https://travis-ci.org/szeder/git-cooking-topics-for-travis-ci/jobs/647253022#L2210\n>>>>\n>>>> (The link highlights the failure, but I'm afraid your browser won't\n>>>> jump there right away; you'll have to click on the print-test-failures\n>>>> fold at the bottom, and scroll down a bit...)\n>>>>\n>>>\n>>> Thank you so much for running this pipeline and pointing out the error!\n>>>\n>>> We will carefully review our interactions with the binary data and \n>>> hopefully solve this in the next version. \n>>\n>> Szeder,\n>>\n>> Thanks so much for running this test. We don't have access to a big endian\n>> machine right now, so could you please apply this patch and re-run your tests?\n> \n> Unfortunately, it still failed:\n> \n>   https://travis-ci.org/szeder/git-cooking-topics-for-travis-ci/jobs/647395554#L2204\n\nThanks! Both fail on test 2 of t0095-bloom.sh, which includes this\nexpected output line:\n\n\tFilter_Data:508928809087080a|8a7648210804001|4089824400951000|841ab310098051a8|\n\nWe may not be properly adjusting the output in the test-helper.\n\nI still think the fixup patch I included is a good idea, but Garima\ncontinues to dig into the problem from all angles to understand this\nfailure and the full fix.\n\n-Stolee\n\n"},{"id":"391380","messageId":"86a75swuie.fsf@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 00/11] Changed Paths Bloom Filters","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-08T23:04:41Z","receivedAt":"2020-02-08T23:04:58Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> Hey! \n>\n> The commit graph feature brought in a lot of performance improvements across\n> multiple commands. However, file based history continues to be a performance\n> pain point, especially in large repositories. \n>\n> Adopting changed path Bloom filters has been discussed on the list before,\n> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n> Derrick Stolee [1]. This series is based on Dr. Stolee's proof of\n> concept in [2].\n\nSidenote: I wondered why it did use MurmurHash3 (64-bit version), which\nrequires adding its implementation, instead of reusing FNV-1 hash\n(Fowler–Noll–Vo hash function) used by Git hashmap implementation, see\nhttps://github.com/git/git/blob/228f53135a4a41a37b6be8e4d6e2b6153db4a8ed/hashmap.h#L109\nBeside the fact that everyone is using MurmurHash for Bloom filters ;-)\n\nIt turns out that in various benchmark MurmurHash is faster and also\nslightly better as a hash than FNV-1 or FNV-1b.\n\n\nI wonder then if it would be a good idea (in the future) to make it easy\nto use hashmap with MurmurHash3 instead of FNV-1, or maybe to even make\nit the default for hashing strings.\n\n>\n> Performance Gains: We tested the performance of git log -- path on the git\n> repo, the linux repo and some internal large repos, with a variety of paths\n> of varying depths.\n\nAs I wrote in reply to previous version of this series, a good public\nrepository (and thus being able to use by anyone) to test the Bloom\nfilter performance improvements could be AOSP (Android) base:\n\n  https://android.googlesource.com/platform/frameworks/base/\n\nwhich is a large repository with long path depths (due to Java file\nnaming conventions).\n\n>\n> On the git and linux repos: We observed a 2x to 5x speed up.\n>\n> On a large internal repo with files seated 6-10 levels deep in the tree: We\n> observed 10x to 20x speed ups, with some paths going up to 28 times faster.\n\nVery nice! Good work!\n\nWhat is the cost of this feature, that is how long it takes to generate\nBloom filters, and how much larger commit-graph file gets?  It would be\nnice to know.\n\n>\n> Future Work (not included in the scope of this series):\n>\n>  1. Supporting multiple path based revision walk\n\nShouldn't then tests that were added in v2 mark use of Bloom filters\nwith multiple paths revision walking as _not working *yet*_\n(test_expect_failure), and not expected to not work (test_expect_success\nwith test_bloom_filters_not_used)?\n\n>  2. Adopting it in git blame logic. \n>  3. Interactions with line log git log -L\n>\n>\n> ----------------------------------------------------------------------------\n>\n> Updates since the last submission\n>\n>  * Removed all the RFC callouts, this is a ready for full review version\n>  * Added unit tests for the bloom filter computation layer\n>  * Added more evolved functional tests for git log\n>  * Fixed a lot of the bugs found by the tests\n>  * Reacted to other miscellaneous feedback on the RFC series. \n>\n> Cheers! Garima Singh\n>\n> [1] https://lore.kernel.org/git/20181009193445.21908-1-szeder.dev@gmail.com/\n> [2] https://lore.kernel.org/git/61559c5b-546e-d61b-d2e1-68de692f5972@gmail.com/\n>\n> Derrick Stolee (2):\n>   diff: halt tree-diff early after max_changes\n>   commit-graph: examine commits by generation number\n>\n> Garima Singh (8):\n>   commit-graph: use MAX_NUM_CHUNKS\n>   bloom: core Bloom filter implementation for changed paths\n>   commit-graph: compute Bloom filters for changed paths\n>   commit-graph: write Bloom filters to commit graph file\n>   commit-graph: reuse existing Bloom filters during write.\n>   commit-graph: add --changed-paths option to write subcommand\n>   revision.c: use Bloom filters to speed up path based revision walks\n>   commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag\n>\n> Jeff King (1):\n>   commit-graph: examine changed-path objects in pack order\n\nThe shortlog summary is a fine tool to show contributors to the patch\nseries, but is not as useful to show patch series as a whole: splitting\nof patches and their ordering.\n\nI will review each of patches individually, but now I would like to say\na few things about the series as a whole.\n\n- [PATCH v2 01/11] commit-graph: use MAX_NUM_CHUNKS\n\n  Simple and non-controversial patch, improvement to existing code with\n  the goal of helping future development (including further patches).\n\n- [PATCH v2 02/11] bloom: core Bloom filter implementation for changed paths\n\n  In my opinion this patch could be split into three individual pieces,\n  though one might think it is not worth it.\n\n  a. Add implementation of MurmurHash v3 (64-bit)\n  \n  Include tests based on test-tool (creating file similar to the\n  t/helper/test-hash.c, or enhancing to that file) that the\n  implementation is correct, for example that 'The quick brown fox jumps\n  over the lazy dog' with given seed (for example the default feed of 0)\n  hashes to the same value as other implementations.\n\n  b. Add implementation of Bloom filter\n\n  Include generic Bloom filter tests i.e. that it correctly answers\n  \"yes\" and \"maybe\" (create filter, save it or print it, then use stored\n  filter), and tests specific to our implementation, namely that the\n  size of the filter behaves as it should.\n\n  c. Bloom filter implementation for changed paths\n\n  Here include tests that use 'test-tool bloom get_filter_for_commit',\n  that filter for commit with no changes and for commit with more than\n  512 fies changed works correctly, that directories are added along the\n  paths, etc.\n\n- [PATCH v2 03/11] diff: halt tree-diff early after max_changes\n\n  I think keeping this patch as a separate step makes individual commits\n  easier to understand and review.\n\n- [PATCH v2 04/11] commit-graph: compute Bloom filters for changed paths\n\n  Here we compute Bloom filters for changed paths for each commit in the\n  commit-graph file, without writing it to file; as a side-effect we\n  calculate total Bloom filters data size.\n\n  This doesn't make much sense as a standalone patch, but it is nice,\n  easy to understand incremental step in building the feature.\n\n- [PATCH v2 05/11] commit-graph: examine changed-path objects in pack order\n- [PATCH v2 06/11] commit-graph: examine commits by generation number\n\n  Those two are performance improvements of previous step.  It is good\n  to keep them as separate commits, makes it easier to understand (and\n  easier to catch error via git-bisect, if there would be any)\n\n- [PATCH v2 07/11] commit-graph: write Bloom filters to commit graph file\n\n  This commit includes the documentation of the two new chunks of\n  commit-graph file format.\n\n  I wonder if the 9th patch in this series, namely\n  commit-graph: add --changed-paths option to write subcommand\n  should not precede this commit.  Otherwise we have this new code but\n  no way of testing it.  On the other hand it makes it easier to\n  review.  On the gripping hand, you can't really test that writing\n  works without the ability to parse Bloom filter data out of\n  commit-graph file... which is the next commit.\n\n- [PATCH v2 08/11] commit-graph: reuse existing Bloom filters during write\n\n  This implements reading Bloom filters data from commit-graph file.\n  Is it a good split?  I think it makes it easier to review the single\n  patch, but itt also makes them less standalone.\n\n- [PATCH v2 09/11] commit-graph: add --changed-paths option to write subcommand\n\n  One thing we could test there is that we are writing two new chunks to\n  the commit-graph file (and perhaps checking that they are correctly\n  formatted, and have correct shape).\n\n- [PATCH v2 10/11] revision.c: use Bloom filters to speed up path based revision walks\n\n  This is quite a big and involved patch, which in my opinion could be\n  split in two or three parts:\n\n  a. Add a bare bones implementation, like in v2\n\n  This limits amount of testing we can do; the only thing we can really\n  test is that we get the same results with and without Bloom filters.\n\n  b.1. Add trace2 Bloom filter statistics\n  b.2. Use said trace2 statistics to test use of Bloom filters\n\n- [PATCH v2 11/11] commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag\n\n  This one is for (optional) exhaustive testing of the feature.\n\n\nFeel free to disagree with those ideas.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"391390","messageId":"8636bkvss2.fsf@gmail.com","threadId":"52499","inReplyTo":"bf6b93878af5be81148614087aee6b4435ef0396.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 01/11] commit-graph: use MAX_NUM_CHUNKS","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-09T12:39:41Z","receivedAt":"2020-02-09T12:40:07Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n> Subject: Re: [PATCH v2 01/11] commit-graph: use MAX_NUM_CHUNKS\n>\n> This is a minor cleanup to make it easier to change the\n> number of chunks being written to the commit-graph in the future.\n\nLooks good to me...\n\n...with the very minor possible nitpick that the subject probably should\nbe\n\n  [PATCH v2 01/11] commit-graph: define and use MAX_NUM_CHUNKS\n\nBut this is just a bikeshedding.  Feel free to disregard this.\n\nBest,\n-- \nJakub Narębski\n\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  commit-graph.c | 5 +++--\n>  1 file changed, 3 insertions(+), 2 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index b205e65ed1..3c4d411326 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -23,6 +23,7 @@\n>  #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n>  #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n>  #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n> +#define MAX_NUM_CHUNKS 5\n>  \n>  #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n>  \n> @@ -1356,8 +1357,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \tint fd;\n>  \tstruct hashfile *f;\n>  \tstruct lock_file lk = LOCK_INIT;\n> -\tuint32_t chunk_ids[6];\n> -\tuint64_t chunk_offsets[6];\n> +\tuint32_t chunk_ids[MAX_NUM_CHUNKS + 1];\n> +\tuint64_t chunk_offsets[MAX_NUM_CHUNKS + 1];\n>  \tconst unsigned hashsz = the_hash_algo->rawsz;\n>  \tstruct strbuf progress_title = STRBUF_INIT;\n>  \tint num_chunks = 3;\n"},{"id":"391562","messageId":"ba856e20-0a3c-e2d2-6744-b9abfacdc465@gmail.com","threadId":"52499","inReplyTo":"140cf2f4-23d5-09ab-8f23-bbbd397c68f7@gmail.com","subject":"Re: [PATCH v2 00/11] Changed Paths Bloom Filters","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-11T19:08:53Z","receivedAt":"2020-02-11T19:08:57Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/7/2020 10:09 AM, Garima Singh wrote:\n> \n> On 2/7/2020 8:52 AM, SZEDER Gábor wrote:\n>>>  * Added unit tests for the bloom filter computation layer\n>>\n>> This fails on big endian, e.g. in Travis CI's s390x build:\n>>\n>>   https://travis-ci.org/szeder/git-cooking-topics-for-travis-ci/jobs/647253022#L2210\n>>\n>> (The link highlights the failure, but I'm afraid your browser won't\n>> jump there right away; you'll have to click on the print-test-failures\n>> fold at the bottom, and scroll down a bit...)\n>>\n> \n> Thank you so much for running this pipeline and pointing out the error!\n> \n> We will carefully review our interactions with the binary data and \n> hopefully solve this in the next version. \n> \n> Cheers!\n> Garima Singh\n> \n\nHey! \n\nThe patch below carries the fix for the failure on Big-endian architectures.\nWe now treat bloom filter data as a simple binary stream of 1 byte words \ninstead of 8 byte words. This avoids the Big-endian vs Little-endian \nconfusion on different CPU architectures. \n\nHere is the successful run of SZEDER's Travis CI s390x build. \n\n https://travis-ci.org/szeder/git/jobs/649044879\n\nI will be squashing this patch into the appropriate commits in the series\nin v3, which I will send out after people have had a chance to complete\ntheir review of v2. \n\nA special thanks to SZEDER for helping us test our patches on his CI \npipeline and saving us the overhead of setting up a Big-endian machine!\n\nCheers!\nGarima Singh\n\n-->8--\n\nFrom ee72310dd8c3ad2b810914edb651008f637e7c2a Mon Sep 17 00:00:00 2001\nFrom: Garima Singh <garima.singh@microsoft.com>\nDate: Tue, 11 Feb 2020 13:55:03 -0500\nSubject: [PATCH] Process bloom filter data as 1 byte words\n\nProcess bloom filter data as 1 byte words instead of 8 byte\nwords to avoid the Big-endian vs Little-endian confusion on\ndifferent CPU architectures\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c               |  24 ++++-----\n bloom.h               |   4 +-\n commit-graph.c        |   4 +-\n t/helper/test-bloom.c |   4 +-\n t/t0095-bloom.sh      | 118 +++++++++++++++++++++---------------------\n 5 files changed, 77 insertions(+), 77 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 90d84dc713..6d5d6bb2ef 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -45,12 +45,13 @@ static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n \n \tint len4 = len / sizeof(uint32_t);\n \n-\tconst uint32_t *blocks = (const uint32_t*)data;\n-\n \tuint32_t k;\n-\tfor (i = 0; i < len4; i++)\n-\t{\n-\t\tk = blocks[i];\n+\tfor (i = 0; i < len4; i++) {\t\n+\t\tuint32_t byte1 = (uint32_t)data[4*i];\n+\t\tuint32_t byte2 = ((uint32_t)data[4*i + 1]) << 8;\n+\t\tuint32_t byte3 = ((uint32_t)data[4*i + 2]) << 16;\n+\t\tuint32_t byte4 = ((uint32_t)data[4*i + 3]) << 24;\n+\t\tk = byte1 | byte2 | byte3 | byte4;\n \t\tk *= c1;\n \t\tk = rotate_right(k, r1);\n \t\tk *= c2;\n@@ -61,8 +62,7 @@ static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n \n \ttail = (data + len4 * sizeof(uint32_t));\n \n-\tswitch (len & (sizeof(uint32_t) - 1))\n-\t{\n+\tswitch (len & (sizeof(uint32_t) - 1)) {\n \tcase 3:\n \t\tk1 ^= ((uint32_t)tail[2]) << 16;\n \t\t/*-fallthrough*/\n@@ -88,9 +88,9 @@ static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n \treturn seed;\n }\n \n-static inline uint64_t get_bitmask(uint32_t pos)\n+static inline unsigned char get_bitmask(uint32_t pos)\n {\n-\treturn ((uint64_t)1) << (pos & (BITS_PER_WORD - 1));\n+\treturn ((unsigned char)1) << (pos & (BITS_PER_WORD - 1));\n }\n \n void load_bloom_filters(void)\n@@ -152,8 +152,8 @@ static int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\tstart_index = 0;\n \n \tfilter->len = end_index - start_index;\n-\tfilter->data = (uint64_t *)(g->chunk_bloom_data +\n-\t\t\t\t\tsizeof(uint64_t) * start_index +\n+\tfilter->data = (unsigned char *)(g->chunk_bloom_data +\n+\t\t\t\t\tsizeof(unsigned char) * start_index +\n \t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n \n \treturn 1;\n@@ -234,7 +234,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \t\t}\n \n \t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n-\t\tfilter->data = xcalloc(filter->len, sizeof(uint64_t));\n+\t\tfilter->data = xcalloc(filter->len, sizeof(unsigned char));\n \n \t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n \t\t\tstruct bloom_key key;\ndiff --git a/bloom.h b/bloom.h\nindex 76f8a9ad0c..9604723ce0 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -12,7 +12,7 @@ struct bloom_filter_settings {\n };\n \n #define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n-#define BITS_PER_WORD 64\n+#define BITS_PER_WORD 8\n #define BLOOMDATA_CHUNK_HEADER_SIZE 3*sizeof(uint32_t)\n \n /*\n@@ -22,7 +22,7 @@ struct bloom_filter_settings {\n  * 'data'.\n  */\n struct bloom_filter {\n-\tuint64_t *data;\n+\tunsigned char *data;\n \tint len;\n };\n \ndiff --git a/commit-graph.c b/commit-graph.c\nindex c0e9834bf2..f5f9a23c9a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1125,7 +1125,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n \twhile (list < last) {\n \t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n \t\tdisplay_progress(progress, ++i);\n-\t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n+\t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n \t\tlist++;\n \t}\n \n@@ -1305,7 +1305,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tstruct commit *c = sorted_by_pos[i];\n \t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n-\t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n+\t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n \t\tdisplay_progress(progress, i + 1);\n \t}\n \ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 9b4be97f75..8fa2d8fc25 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -23,7 +23,7 @@ static void print_bloom_filter(struct bloom_filter *filter) {\n \tprintf(\"Filter_Length:%d\\n\", filter->len);\n \tprintf(\"Filter_Data:\");\n \tfor (i = 0; i < filter->len; i++){\n-\t\tprintf(\"%\"PRIx64\"|\", filter->data[i]);\n+\t\tprintf(\"%02x|\", filter->data[i]);\n \t}\n \tprintf(\"\\n\");\n }\n@@ -57,7 +57,7 @@ int cmd__bloom(int argc, const char **argv)\n \t\tstruct bloom_filter filter;\n \t\tint i = 2;\n \t\tfilter.len =  (settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n-\t\tfilter.data = xcalloc(filter.len, sizeof(uint64_t));\n+\t\tfilter.data = xcalloc(filter.len, sizeof(unsigned char));\n \n \t\tif (!argv[2]){\n \t\t\tdie(\"at least one input string expected\");\ndiff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\nindex 424fe4fc29..58273219ff 100755\n--- a/t/t0095-bloom.sh\n+++ b/t/t0095-bloom.sh\n@@ -3,58 +3,11 @@\n test_description='test bloom.c'\n . ./test-lib.sh\n \n-test_expect_success 'get bloom filters for commit with no changes' '\n-\tgit init &&\n-\tgit commit --allow-empty -m \"c0\" &&\n-\tcat >expect <<-\\EOF &&\n-\tFilter_Length:0\n-\tFilter_Data:\n-\tEOF\n-\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n-\ttest_cmp expect actual\n-'\n-\n-test_expect_success 'get bloom filter for commit with 10 changes' '\n-\trm actual &&\n-\trm expect &&\n-\tmkdir smallDir &&\n-\tfor i in $(test_seq 0 9)\n-\tdo\n-\t\techo $i >smallDir/$i\n-\tdone &&\n-\tgit add smallDir &&\n-\tgit commit -m \"commit with 10 changes\" &&\n-\tcat >expect <<-\\EOF &&\n-\tFilter_Length:4\n-\tFilter_Data:508928809087080a|8a7648210804001|4089824400951000|841ab310098051a8|\n-\tEOF\n-\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n-\ttest_cmp expect actual\n-'\n-\n-test_expect_success EXPENSIVE 'get bloom filter for commit with 513 changes' '\n-\trm actual &&\n-\trm expect &&\n-\tmkdir bigDir &&\n-\tfor i in $(test_seq 0 512)\n-\tdo\n-\t\techo $i >bigDir/$i\n-\tdone &&\n-\tgit add bigDir &&\n-\tgit commit -m \"commit with 513 changes\" &&\n-\tcat >expect <<-\\EOF &&\n-\tFilter_Length:0\n-\tFilter_Data:\n-\tEOF\n-\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n-\ttest_cmp expect actual\n-'\n-\n test_expect_success 'compute bloom key for empty string' '\n \tcat >expect <<-\\EOF &&\n \tHashes:5615800c|5b966560|61174ab4|66983008|6c19155c|7199fab0|771ae004|\n-\tFilter_Length:1\n-\tFilter_Data:11000110001110|\n+\tFilter_Length:2\n+\tFilter_Data:11|11|\n \tEOF\n \ttest-tool bloom generate_filter \"\" >actual &&\n \ttest_cmp expect actual\n@@ -63,8 +16,8 @@ test_expect_success 'compute bloom key for empty string' '\n test_expect_success 'compute bloom key for whitespace' '\n \tcat >expect <<-\\EOF &&\n \tHashes:1bf014e6|8a91b50b|f9335530|67d4f555|d676957a|4518359f|b3b9d5c4|\n-\tFilter_Length:1\n-\tFilter_Data:401004080200810|\n+\tFilter_Length:2\n+\tFilter_Data:71|8c|\n \tEOF\n \ttest-tool bloom generate_filter \" \" >actual &&\n \ttest_cmp expect actual\n@@ -73,8 +26,8 @@ test_expect_success 'compute bloom key for whitespace' '\n test_expect_success 'compute bloom key for a root level folder' '\n \tcat >expect <<-\\EOF &&\n \tHashes:1a21016f|fff1c06d|e5c27f6b|cb933e69|b163fd67|9734bc65|7d057b63|\n-\tFilter_Length:1\n-\tFilter_Data:aaa800000000|\n+\tFilter_Length:2\n+\tFilter_Data:a8|aa|\n \tEOF\n \ttest-tool bloom generate_filter \"A\" >actual &&\n \ttest_cmp expect actual\n@@ -83,8 +36,8 @@ test_expect_success 'compute bloom key for a root level folder' '\n test_expect_success 'compute bloom key for a root level file' '\n \tcat >expect <<-\\EOF &&\n \tHashes:e2d51107|30970605|7e58fb03|cc1af001|19dce4ff|679ed9fd|b560cefb|\n-\tFilter_Length:1\n-\tFilter_Data:a8000000000000aa|\n+\tFilter_Length:2\n+\tFilter_Data:aa|a8|\n \tEOF\n \ttest-tool bloom generate_filter \"file.txt\" >actual &&\n \ttest_cmp expect actual\n@@ -93,8 +46,8 @@ test_expect_success 'compute bloom key for a root level file' '\n test_expect_success 'compute bloom key for a deep folder' '\n \tcat >expect <<-\\EOF &&\n \tHashes:864cf838|27f055cd|c993b362|6b3710f7|0cda6e8c|ae7dcc21|502129b6|\n-\tFilter_Length:1\n-\tFilter_Data:1c0000600003000|\n+\tFilter_Length:2\n+\tFilter_Data:c6|31|\n \tEOF\n \ttest-tool bloom generate_filter \"A/B/C/D/E\" >actual &&\n \ttest_cmp expect actual\n@@ -103,11 +56,58 @@ test_expect_success 'compute bloom key for a deep folder' '\n test_expect_success 'compute bloom key for a deep file' '\n \tcat >expect <<-\\EOF &&\n \tHashes:07cdf850|4af629c7|8e1e5b3e|d1468cb5|146ebe2c|5796efa3|9abf211a|\n-\tFilter_Length:1\n-\tFilter_Data:4020100804010080|\n+\tFilter_Length:2\n+\tFilter_Data:a9|54|\n \tEOF\n \ttest-tool bloom generate_filter \"A/B/C/D/E/file.txt\" >actual &&\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'get bloom filters for commit with no changes' '\n+\tgit init &&\n+\tgit commit --allow-empty -m \"c0\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:0\n+\tFilter_Data:\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'get bloom filter for commit with 10 changes' '\n+\trm actual &&\n+\trm expect &&\n+\tmkdir smallDir &&\n+\tfor i in $(test_seq 0 9)\n+\tdo\n+\t\techo $i >smallDir/$i\n+\tdone &&\n+\tgit add smallDir &&\n+\tgit commit -m \"commit with 10 changes\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:25\n+\tFilter_Data:c2|0b|b8|c0|10|88|f0|1d|c1|0c|01|a4|01|28|81|80|01|30|10|d0|92|be|88|10|8a|\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success EXPENSIVE 'get bloom filter for commit with 513 changes' '\n+\trm actual &&\n+\trm expect &&\n+\tmkdir bigDir &&\n+\tfor i in $(test_seq 0 512)\n+\tdo\n+\t\techo $i >bigDir/$i\n+\tdone &&\n+\tgit add bigDir &&\n+\tgit commit -m \"commit with 513 changes\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:0\n+\tFilter_Data:\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \n2.22.0.windows.1\n\n"},{"id":"391831","messageId":"86eeuvwz0s.fsf@gmail.com","threadId":"52499","inReplyTo":"02b16d94227470059dcee2781e29ae7ae010f602.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 02/11] bloom: core Bloom filter implementation for changed paths","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-15T17:17:39Z","receivedAt":"2020-02-15T17:17:53Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Add the core Bloom filter logic for computing the paths changed between a\n> commit and its first parent. For details on what Bloom filters are and how they\n> work, please refer to Dr. Derrick Stolee's blog post [1]. It provides a concise\n> explaination of the adoption of Bloom filters as described in [2] and [3].\n                                                                           ^^- to add\n>\n> 1. We currently use 7 and 10 for the number of hashes and the size of each\n>    entry respectively. They served as great starting values, the mathematical\n>    details behind this choice are described in [1] and [4]. The implementation,\n                                                                                ^^- to add\n>    while not completely open to it at the moment, is flexible enough to allow\n>    for tweaking these settings in the future.\n\nI don't know if it is worth it, but I think it should be size of each\nentry, or in other words number of bits per element in the set, as first\nvalue, and number of hashes as second.\n\nAbout where those values come from.  The idea is that you decide on the\nacceptable number of false positives, for example 1% (or 0.8% given that\nthe values must be integers); that gives you number of bits per element\ni.e. 10, and from there you can find optimal number of hashes i.e. 7.\nThe references mentioned (and Wikipedia article) have those equations.\n\n>\n>    Note: The performance gains we have observed with these values are\n>    significant enough that we did not need to tweak these settings.\n>    The performance numbers are included in the cover letter of this series\n>    and in the message of a subsequent commit where we use Bloom filters in\n>    to speed up `git log -- <path>`.\n\nAll right.\n\n>\n> 2. As described in the blog and in [3], we do not need 7 independent hashing\n>    functions. We use the Murmur3 hashing scheme. Seed it twice and then\n>    combine those to procure an arbitrary number of hash values.\n\nThe technique from [3] is called \"double hashing\" (Algorithm 1 and\nequation (4) on page 10).  Note that in this paper there is also\npresented \"enhanced double hashing\" scheme (Algorithm 2 and equation\n(6)) -- more about it later.\n\nThis is a standard technique from the hashing literature, called open\naddressing with double hashing in hash tables.\n\nThis \"enhanced double hashing\" technique is further analyzed in [6].\n\n[6] Adam Kirsch, Michael Mitzenmacher\n    \"Less Hashing, Same Performance: Building a Better Bloom Filter\"\n    https://www.eecs.harvard.edu/~michaelm/postscripts/esa2006a.pdf\n    https://doi.org/10.5555/1400123.1400125\n\n>\n> 3. The filters are sized according to the number of changes in the each commit,\n>    with minimum size of one 64 bit word.\n\nIf I understand it correctly (but which might not be entirely clear),\nthe filter size in bits is the number of changes^* times 10, rounded up\nto the nearest multiple of 64.\n\n[*] where the number of changes is the number of changed files (new blob\nobjects) _and_ the number of changed directories (new tree objects,\nexcluding root tree object change).\n\n\nThe interesting corner case, which might be worth specifying explicitly,\nis what happens in the case there are _no changes_ with respect to first\nparent (which can happen with either commit created with `git commit\n--allow-empty`, or merge created e.g. with `git merge --strategy=ours`).\nIs this case represented as Bloom filter of length 0, or as a Bloom\nfilter of length  of one 64-bit word which is minimal length composed of\nall 0's (0x0000000000000000)?\n\n>\n> 4. We fill the Bloom filters as (const char *data, int len) pairs as\n>    \"struct bloom_filter\"s in a commit slab.\n\nAll right.\n\n>\n> 5. The seed_murmur3 method is implemented as described in [5]. It hashes the\n>    given data using a given seed and produces a uniformly distributed hash\n>    value.\n\nActually there are two variants of Murmur3 hash, and we should specify\nwhich one we are using.  There is Murmur3_32 which returns 32-bit value,\nand Murmur3_128 which returns 128-bit value (which is different for x86\nand x64 versions).  We use Murmur3_32.\n\nAlso, seed_murmur3 is the name given the function, not the name of the\nmethod i.e. of a non-cryptographic hash function.\n\n\nOne question that one might as is why use Murmur3 hash instead for\nexample already implemented FNV hash from hashmap implementation (FNV\nhash i.e. Fowler–Noll–Vo hash function is another non-cryptographic hash\nfunction).  The answer is of course performance while maintaining good\nenough quality (and for Bloom filter there is no problem of \"hash\nflooding\" denial-of-service like for there is for a hash table -- no\nneed for SipHash or similar).\n\n>\n> [1] https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n\nI would write it in full, similar to subsequent bibliographical entries,\nthat is:\n\n  [1] Derrick Stolee\n      \"Supercharging the Git Commit Graph IV: Bloom Filters\"\n      https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n\nBut that is just a matter of style.\n\n>\n> [2] Flavio Bonomi, Michael Mitzenmacher, Rina Panigrahy, Sushil Singh, George Varghese\n>     \"An Improved Construction for Counting Bloom Filters\"\n>     http://theory.stanford.edu/~rinap/papers/esa2006b.pdf\n>     https://doi.org/10.1007/11841036_61\n>\n> [3] Peter C. Dillinger and Panagiotis Manolios\n>     \"Bloom Filters in Probabilistic Verification\"\n>     http://www.ccs.neu.edu/home/pete/pub/Bloom-filters-verification.pdf\n>     https://doi.org/10.1007/978-3-540-30494-4_26\n\nGood, we should be able to find them even if the URL with PDF stops\nworking for some reason.\n\n>\n> [4] Thomas Mueller Graf, Daniel Lemire\n>     \"Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters\"\n>     https://arxiv.org/abs/1912.08258\n>\n> [5] https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n>\n> Helped-by: Jeff King <peff@peff.net>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  Makefile              |   2 +\n>  bloom.c               | 228 ++++++++++++++++++++++++++++++++++++++++++\n>  bloom.h               |  56 +++++++++++\n>  t/helper/test-bloom.c |  84 ++++++++++++++++\n>  t/helper/test-tool.c  |   1 +\n>  t/helper/test-tool.h  |   1 +\n>  t/t0095-bloom.sh      | 113 +++++++++++++++++++++\n>  7 files changed, 485 insertions(+)\n>  create mode 100644 bloom.c\n>  create mode 100644 bloom.h\n>  create mode 100644 t/helper/test-bloom.c\n>  create mode 100755 t/t0095-bloom.sh\n\nAs I wrote earlier, In my opinion this patch could be split into three\nindividual single-functionality pieces, to make it easier to review and\naid in bisectability if needed.\n\n1. Add implementation of MurmurHash v3 (32-bit result)\n  \nInclude tests based on test-tool (creating file similar to the\nt/helper/test-hash.c, or enhancing to that file) that the implementation\nis correct, for example that 'The quick brown fox jumps over the lazy\ndog' or 'Hello world!' with a given seed (for example the default seed\nof 0) hashes to the same value as other implementations, including the\nreference implementation in https://github.com/aappleby/smhasher\n\n\n2. Add implementation of [variant of] Bloom filter\n\nInclude generic Bloom filter tests i.e. that it correctly answers \"yes\"\nand \"maybe\" (create filter, save it or print it, then use stored\nfilter), and tests specific to our implementation, namely that the size\nof the filter behaves as it should.\n\n\n3. Bloom filter implementation for changed paths\n\nHere include tests that use 'test-tool bloom get_filter_for_commit',\nthat filter for commit with no changes and for commit with more than 512\nchanges works correctly, that directories are added along the files,\netc.\n\n\nThis split would make it easier to distinguish if the problems with\ntests failing on big-endian architectures is caused by different output\nfrom our implementation of Murmur3 hash, different bit sequence in the\nBloom filter, or just different printed output of Bloom filter data.\n\n>\n> diff --git a/Makefile b/Makefile\n> index 6134104ae6..afba81f4a8 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -695,6 +695,7 @@ X =\n>  \n>  PROGRAMS += $(patsubst %.o,git-%$X,$(PROGRAM_OBJS))\n>  \n> +TEST_BUILTINS_OBJS += test-bloom.o\n>  TEST_BUILTINS_OBJS += test-chmtime.o\n>  TEST_BUILTINS_OBJS += test-config.o\n>  TEST_BUILTINS_OBJS += test-ctype.o\n> @@ -840,6 +841,7 @@ LIB_OBJS += base85.o\n>  LIB_OBJS += bisect.o\n>  LIB_OBJS += blame.o\n>  LIB_OBJS += blob.o\n> +LIB_OBJS += bloom.o\n>  LIB_OBJS += branch.o\n>  LIB_OBJS += bulk-checkin.o\n>  LIB_OBJS += bundle.o\n\nAll right.\n\n> diff --git a/bloom.c b/bloom.c\n> new file mode 100644\n> index 0000000000..6082193a75\n> --- /dev/null\n> +++ b/bloom.c\n> @@ -0,0 +1,228 @@\n> +#include \"git-compat-util.h\"\n> +#include \"bloom.h\"\n> +#include \"commit-graph.h\"\n> +#include \"object-store.h\"\n> +#include \"diff.h\"\n> +#include \"diffcore.h\"\n> +#include \"revision.h\"\n> +#include \"hashmap.h\"\n> +\n> +define_commit_slab(bloom_filter_slab, struct bloom_filter);\n> +\n> +struct bloom_filter_slab bloom_filters;\n\nAll right, this is needed to store per-commit Bloom filter data\n(inside-out object style, or in other jargon stored on slab).\n\n> +\n> +struct pathmap_hash_entry {\n> +    struct hashmap_entry entry;\n> +    const char path[FLEX_ARRAY];\n> +};\n\nO.K. this is used to add gather paths to add them all as elements to the\nBloom filter.\n\n> +\n> +static uint32_t rotate_right(uint32_t value, int32_t count)\n> +{\n> +\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n> +\tcount &= mask;\n> +\treturn ((value >> count) | (value << ((-count) & mask)));\n> +}\n\nHmmm... both the algoritm on Wikipedia, and reference implementation use\nrotate *left*, not rotate *right* in the implementation of Murmur3 hash,\nsee\n\n  https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n  https://github.com/aappleby/smhasher/blob/master/src/MurmurHash3.cpp#L23\n\n\ninline uint32_t rotl32 ( uint32_t x, int8_t r )\n{\n  return (x << r) | (x >> (32 - r));\n}\n\n> +\n> +/*\n> + * Calculate a hash value for the given data using the given seed.\n> + * Produces a uniformly distributed hash value.\n> + * Not considered to be cryptographically secure.\n> + * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n> + **/\n    ^^-- why two _trailing_ asterisks?\n\nPerhaps it would be worth it to add that this hash function is intended\nto be fast while being reasonably good (it is distributed randomly\nenough, and it doesn't have too many hash collisions on typical inputs).\nBut this might be too much for a comment.\n\n> +static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n\nA few things: name of the function, type of parameters and ordering of\nparameters.\n\n\nAbout the name: when I first saw seed_murmur3() used, I thought it was\n_setting_ the seed, not that it was returning the 32-bit hash value.\nOther implementations use either murmur3_32, MurmurHash3_x86_32, or\nsomething similar like hashmurmur3_32.  If we were to specify that\n'seed' is one of parameters, then using this word as part of suffix\nwould be better than using seed_ prefix; if we need it at all.\n\nBecause there is 32-bit and 128-bit variants of Murmur3, I think the _32\nsuffix should be a part of function name.\n\nIn short, I think that the name of the function should be murmur3_32, or\nmurmurhash3_32, or possibly murmur3_32_seed, or something like that.\n\n\nAbout types of parameters and the return type of function: I understand\nthat 'data' parameter is of type 'const char *', instead of more generic\n'const uint8_t*' or 'const void *' because of what we will be using the\nhash function for.  On the other hand taking a look at implementation of\nFNV hash function in hashmap.{c,h} we see that the 'str*' variants take\n'const char *' parameter _without_ length, and 'mem*' variants take\n'const void *' parmeter with length of data.\n\nShouldn't 'len' parameter be of 'size_t' type, rather than 'int'?  Both\nthe example implementation in C on Wikipedia page, and implementation in\nC in qLibc use 'size_t'; the implementation of FNV hash in hashmap in\nGit also uses 'size_t' (while admittedly the reference implementation in\nC++ of Austin Appleby uses 'int' type for len parameter).\n\nFor 32-bit output variant of Murmur3 hash, using uint32_t as return type\nis just fine.  The '*hash*' functions from hashmap.{c,h} use 'unsigned\nint' but I think 'uint32_t' is better.\n\n\nAbout names and ordering of parameters: the 'seed' or 'hash_seed'\nparameter should be either first or last; it is a matter of preference.\nWhile example implementation on Wikipedia page, Appleby's reference\nimplementation in C++ have 'seed' as last parameter, memihash_cont()\nfrom hashmap.c in Git has it as first parameter.\n\nIn short: I'm fine with either order (seed parameter first or last), and\neither name (be it 'seed' or 'hash_seed').\n\n> +{\n> +\tconst uint32_t c1 = 0xcc9e2d51;\n> +\tconst uint32_t c2 = 0x1b873593;\n> +\tconst uint32_t r1 = 15;\n> +\tconst uint32_t r2 = 13;\n> +\tconst uint32_t m = 5;\n> +\tconst uint32_t n = 0xe6546b64;\n> +\tint i;\n> +\tuint32_t k1 = 0;\n> +\tconst char *tail;\n> +\n> +\tint len4 = len / sizeof(uint32_t);\n> +\n> +\tconst uint32_t *blocks = (const uint32_t*)data;\n> +\n> +\tuint32_t k;\n> +\tfor (i = 0; i < len4; i++)\n> +\t{\n> +\t\tk = blocks[i];\n\nIMPORTANT: There is a comment around there in the example implementation\nin C on Wikipedia that this operation above is a source of differing\nresults across endianness.  The pseudo-code description of the algorithm\non Wikipedia (above of C code) says that endian swapping is only\nnecessary on big-endian machines (and that it is needed to place the\nmeaningful digits towards the low end of the value, to not be discarded\nby the modulo arithmetic under overflow).\n\nThe original / reference implementation by Austin Appleby in C++ uses\ngetblock32() function for doing the block read... but it doesn't\nactually implement the endian-swapping on big-endian architecture:\n\n  //-----------------------------------------------------------------------------\n  // Block read - if your platform needs to do endian-swapping or can only\n  // handle aligned reads, do the conversion here\n\n  FORCE_INLINE uint32_t getblock32 ( const uint32_t * p, int i )\n  {\n    return p[i];\n  }\n\nReferences:\n-----------\n1. https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n2. https://github.com/aappleby/smhasher/blob/master/src/MurmurHash3.cpp\n\n> +\t\tk *= c1;\n> +\t\tk = rotate_right(k, r1);\n\nIt is  k ROL r1 / ROTL32(k,15) / (k << 15) | (k >> (32 - 15))\n(in other implementations), not rotate_right.\n\n> +\t\tk *= c2;\n> +\n> +\t\tseed ^= k;\n> +\t\tseed = rotate_right(seed, r2) * m + n;\n\nIt is  hash ROL r2 / ROTL32(h1,13) / (h << 13) | (h >> (32 - 13))\n(in other implementations), not rotate_right.\n\nReferences:\n-----------\n1. https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n2. https://github.com/aappleby/smhasher/blob/master/src/MurmurHash3.cpp#L94\n3. https://github.com/wolkykim/qlibc/blob/master/src/utilities/qhash.c#L258\n\n> +\t}\n> +\n> +\ttail = (data + len4 * sizeof(uint32_t));\n\nHmmm... in the pseudocode implementation on Wikipedia this is the place\nwhere one needs to respect endianness:\n\n    with any remainingBytesInKey do\n        remainingBytes ← SwapToLittleEndian(remainingBytesInKey)\n        // Note: Endian swapping is only necessary on big-endian machines.\n        //       The purpose is to place the meaningful digits towards the low end of the value,\n        //       so that these digits have the greatest potential to affect the low range digits\n        //       in the subsequent multiplication.  Consider that locating the meaningful digits\n        //       in the high range would produce a greater effect upon the high digits of the\n        //       multiplication, and notably, that such high digits are likely to be discarded\n        //       by the modulo arithmetic under overflow.  We don't want that.\n\nOn the other hand in the reference Appleby's C++ implementation the\nendian-swapping is [ssumed to be] done only in the loop over data.\nEither should be enough alone, but doing swapping for remaining bytes\nonly would work, it would be a better solution -- you do swap only once,\nat the end.\n\nIt looks like the Crhomium implementation in C by Shane Day (public\ndomain) uses the second solution; well almost, see:\nhttps://chromium.googlesource.com/external/smhasher/+/5b8fd3c31a58b87b80605dca7a64fad6cb3f8a0f/PMurHash.c#189\n\n> +\n> +\tswitch (len & (sizeof(uint32_t) - 1))\n> +\t{\n> +\tcase 3:\n> +\t\tk1 ^= ((uint32_t)tail[2]) << 16;\n> +\t\t/*-fallthrough*/\n> +\tcase 2:\n> +\t\tk1 ^= ((uint32_t)tail[1]) << 8;\n> +\t\t/*-fallthrough*/\n> +\tcase 1:\n> +\t\tk1 ^= ((uint32_t)tail[0]) << 0;\n> +\t\tk1 *= c1;\n> +\t\tk1 = rotate_right(k1, r1);\n\n\n\n> +\t\tk1 *= c2;\n> +\t\tseed ^= k1;\n> +\t\tbreak;\n> +\t}\n\n\n\n> +\n> +\tseed ^= (uint32_t)len;\n> +\tseed ^= (seed >> 16);\n> +\tseed *= 0x85ebca6b;\n> +\tseed ^= (seed >> 13);\n> +\tseed *= 0xc2b2ae35;\n> +\tseed ^= (seed >> 16);\n> +\n> +\treturn seed;\n> +}\n> +\n> +static inline uint64_t get_bitmask(uint32_t pos)\n> +{\n> +\treturn ((uint64_t)1) << (pos & (BITS_PER_WORD - 1));\n> +}\n> +\n> +void load_bloom_filters(void)\n> +{\n> +\tinit_bloom_filter_slab(&bloom_filters);\n> +}\n> +\n> +void fill_bloom_key(const char *data,\n> +\t\t\t\t\tint len,\n> +\t\t\t\t\tstruct bloom_key *key,\n> +\t\t\t\t\tstruct bloom_filter_settings *settings)\n> +{\n> +\tint i;\n> +\tconst uint32_t seed0 = 0x293ae76f;\n> +\tconst uint32_t seed1 = 0x7e646e2c;\n> +\tconst uint32_t hash0 = seed_murmur3(seed0, data, len);\n> +\tconst uint32_t hash1 = seed_murmur3(seed1, data, len);\n> +\n> +\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n> +\tfor (i = 0; i < settings->num_hashes; i++)\n> +\t\tkey->hashes[i] = hash0 + i * hash1;\n> +}\n> +\n> +void add_key_to_filter(struct bloom_key *key,\n> +\t\t\t\t\t   struct bloom_filter *filter,\n> +\t\t\t\t\t   struct bloom_filter_settings *settings)\n> +{\n> +\tint i;\n> +\tuint64_t mod = filter->len * BITS_PER_WORD;\n> +\n> +\tfor (i = 0; i < settings->num_hashes; i++) {\n> +\t\tuint64_t hash_mod = key->hashes[i] % mod;\n> +\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n> +\n> +\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n> +\t}\n> +}\n> +\n> +struct bloom_filter *get_bloom_filter(struct repository *r,\n> +\t\t\t\t      struct commit *c)\n> +{\n> +\tstruct bloom_filter *filter;\n> +\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n> +\tint i;\n> +\tstruct diff_options diffopt;\n> +\n> +\tif (!bloom_filters.slab_size)\n> +\t\treturn NULL;\n> +\n> +\tfilter = bloom_filter_slab_at(&bloom_filters, c);\n> +\n> +\trepo_diff_setup(r, &diffopt);\n> +\tdiffopt.flags.recursive = 1;\n> +\tdiff_setup_done(&diffopt);\n> +\n> +\tif (c->parents)\n> +\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &diffopt);\n> +\telse\n> +\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n> +\tdiffcore_std(&diffopt);\n> +\n> +\tif (diff_queued_diff.nr <= 512) {\n> +\t\tstruct hashmap pathmap;\n> +\t\tstruct pathmap_hash_entry* e;\n> +\t\tstruct hashmap_iter iter;\n> +\t\thashmap_init(&pathmap, NULL, NULL, 0);\n> +\n> +\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n> +\t\t\tconst char* path = diff_queued_diff.queue[i]->two->path;\n> +\t\t\tconst char* p = path;\n> +\n> +\t\t\t/*\n> +\t\t\t* Add each leading directory of the changed file, i.e. for\n> +\t\t\t* 'dir/subdir/file' add 'dir' and 'dir/subdir' as well, so\n> +\t\t\t* the Bloom filter could be used to speed up commands like\n> +\t\t\t* 'git log dir/subdir', too.\n> +\t\t\t*\n> +\t\t\t* Note that directories are added without the trailing '/'.\n> +\t\t\t*/\n> +\t\t\tdo {\n> +\t\t\t\tchar* last_slash = strrchr(p, '/');\n> +\n> +\t\t\t\tFLEX_ALLOC_STR(e, path, path);\n> +\t\t\t\thashmap_entry_init(&e->entry, strhash(p));\n> +\t\t\t\thashmap_add(&pathmap, &e->entry);\n> +\n> +\t\t\t\tif (!last_slash)\n> +\t\t\t\t\tlast_slash = (char*)p;\n> +\t\t\t\t*last_slash = '\\0';\n> +\n> +\t\t\t} while (*p);\n> +\n> +\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n> +\t\t}\n> +\n> +\t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n> +\t\tfilter->data = xcalloc(filter->len, sizeof(uint64_t));\n> +\n> +\t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n> +\t\t\tstruct bloom_key key;\n> +\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, &settings);\n> +\t\t\tadd_key_to_filter(&key, filter, &settings);\n> +\t\t}\n> +\n> +\t\thashmap_free_entries(&pathmap, struct pathmap_hash_entry, entry);\n> +\t} else {\n> +\t\tfor (i = 0; i < diff_queued_diff.nr; i++)\n> +\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n> +\t\tfilter->data = NULL;\n> +\t\tfilter->len = 0;\n> +\t}\n> +\n> +\tfree(diff_queued_diff.queue);\n> +\tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n> +\n> +\treturn filter;\n> +}\n> +\n> +int bloom_filter_contains(struct bloom_filter *filter,\n> +\t\t\t  struct bloom_key *key,\n> +\t\t\t  struct bloom_filter_settings *settings)\n> +{\n> +\tint i;\n> +\tuint64_t mod = filter->len * BITS_PER_WORD;\n> +\n> +\tif (!mod)\n> +\t\treturn -1;\n> +\n> +\tfor (i = 0; i < settings->num_hashes; i++) {\n> +\t\tuint64_t hash_mod = key->hashes[i] % mod;\n> +\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n> +\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n> +\t\t\treturn 0;\n> +\t}\n> +\n> +\treturn 1;\n> +}\n> diff --git a/bloom.h b/bloom.h\n> new file mode 100644\n> index 0000000000..7f40c751f7\n> --- /dev/null\n> +++ b/bloom.h\n> @@ -0,0 +1,56 @@\n> +#ifndef BLOOM_H\n> +#define BLOOM_H\n> +\n> +struct commit;\n> +struct repository;\n> +struct commit_graph;\n> +\n> +struct bloom_filter_settings {\n> +\tuint32_t hash_version;\n> +\tuint32_t num_hashes;\n> +\tuint32_t bits_per_entry;\n> +};\n> +\n> +#define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n> +#define BITS_PER_WORD 64\n> +\n> +/*\n> + * A bloom_filter struct represents a data segment to\n> + * use when testing hash values. The 'len' member\n> + * dictates how many uint64_t entries are stored in\n> + * 'data'.\n> + */\n> +struct bloom_filter {\n> +\tuint64_t *data;\n> +\tint len;\n> +};\n> +\n> +/*\n> + * A bloom_key represents the k hash values for a\n> + * given hash input. These can be precomputed and\n> + * stored in a bloom_key for re-use when testing\n> + * against a bloom_filter.\n> + */\n> +struct bloom_key {\n> +\tuint32_t *hashes;\n> +};\n> +\n> +void load_bloom_filters(void);\n> +\n> +void fill_bloom_key(const char *data,\n> +\t\t    int len,\n> +\t\t    struct bloom_key *key,\n> +\t\t    struct bloom_filter_settings *settings);\n> +\n> +void add_key_to_filter(struct bloom_key *key,\n> +\t\t\t\t\t   struct bloom_filter *filter,\n> +\t\t\t\t\t   struct bloom_filter_settings *settings);\n> +\n> +struct bloom_filter *get_bloom_filter(struct repository *r,\n> +\t\t\t\t      struct commit *c);\n> +\n> +int bloom_filter_contains(struct bloom_filter *filter,\n> +\t\t\t  struct bloom_key *key,\n> +\t\t\t  struct bloom_filter_settings *settings);\n> +\n> +#endif\n> diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n> new file mode 100644\n> index 0000000000..331957011b\n> --- /dev/null\n> +++ b/t/helper/test-bloom.c\n> @@ -0,0 +1,84 @@\n> +#include \"test-tool.h\"\n> +#include \"git-compat-util.h\"\n> +#include \"bloom.h\"\n> +#include \"test-tool.h\"\n> +#include \"cache.h\"\n> +#include \"commit-graph.h\"\n> +#include \"commit.h\"\n> +#include \"config.h\"\n> +#include \"object-store.h\"\n> +#include \"object.h\"\n> +#include \"repository.h\"\n> +#include \"tree.h\"\n> +\n> +struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n> +\n> +static void print_bloom_filter(struct bloom_filter *filter) {\n> +\tint i;\n> +\n> +\tif (!filter) {\n> +\t\tprintf(\"No filter.\\n\");\n> +\t\treturn;\n> +\t}\n> +\tprintf(\"Filter_Length:%d\\n\", filter->len);\n> +\tprintf(\"Filter_Data:\");\n> +\tfor (i = 0; i < filter->len; i++){\n> +\t\tprintf(\"%\"PRIx64\"|\", filter->data[i]);\n> +\t}\n> +\tprintf(\"\\n\");\n> +}\n> +\n> +static void add_string_to_filter(const char *data, struct bloom_filter *filter) {\n> +\t\tstruct bloom_key key;\n> +\t\tint i;\n> +\n> +\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n> +\t\tprintf(\"Hashes:\");\n> +\t\tfor (i = 0; i < settings.num_hashes; i++){\n> +\t\t\tprintf(\"%08x|\", key.hashes[i]);\n> +\t\t}\n> +\t\tprintf(\"\\n\");\n> +\t\tadd_key_to_filter(&key, filter, &settings);\n> +}\n> +\n> +static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n> +{\n> +\tstruct commit *c;\n> +\tstruct bloom_filter *filter;\n> +\tsetup_git_directory();\n> +\tc = lookup_commit(the_repository, commit_oid);\n> +\tfilter = get_bloom_filter(the_repository, c);\n> +\tprint_bloom_filter(filter);\n> +}\n> +\n> +int cmd__bloom(int argc, const char **argv)\n> +{\n> +    if (!strcmp(argv[1], \"generate_filter\")) {\n> +\t\tstruct bloom_filter filter;\n> +\t\tint i = 2;\n> +\t\tfilter.len =  (settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n> +\t\tfilter.data = xcalloc(filter.len, sizeof(uint64_t));\n> +\n> +\t\tif (!argv[2]){\n> +\t\t\tdie(\"at least one input string expected\");\n> +\t\t}\n> +\n> +\t\twhile (argv[i]) {\n> +\t\t\tadd_string_to_filter(argv[i], &filter);\n> +\t\t\ti++;\n> +\t\t}\n> +\n> +\t\tprint_bloom_filter(&filter);\n> +\t}\n> +\n> +\tif (!strcmp(argv[1], \"get_filter_for_commit\")) {\n> +\t\tstruct object_id oid;\n> +\t\tconst char *end;\n> +\t\tif (parse_oid_hex(argv[2], &oid, &end))\n> +\t\t\tdie(\"cannot parse oid '%s'\", argv[2]);\n> +\t\tload_bloom_filters();\n> +\t\tget_bloom_filter_for_commit(&oid);\n> +\t}\n> +\n> +\treturn 0;\n> +}\n> diff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\n> index c9a232d238..ca4f4b0066 100644\n> --- a/t/helper/test-tool.c\n> +++ b/t/helper/test-tool.c\n> @@ -14,6 +14,7 @@ struct test_cmd {\n>  };\n>  \n>  static struct test_cmd cmds[] = {\n> +\t{ \"bloom\", cmd__bloom },\n>  \t{ \"chmtime\", cmd__chmtime },\n>  \t{ \"config\", cmd__config },\n>  \t{ \"ctype\", cmd__ctype },\n> diff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\n> index c8549fd87f..05d2b32451 100644\n> --- a/t/helper/test-tool.h\n> +++ b/t/helper/test-tool.h\n> @@ -4,6 +4,7 @@\n>  #define USE_THE_INDEX_COMPATIBILITY_MACROS\n>  #include \"git-compat-util.h\"\n>  \n> +int cmd__bloom(int argc, const char **argv);\n>  int cmd__chmtime(int argc, const char **argv);\n>  int cmd__config(int argc, const char **argv);\n>  int cmd__ctype(int argc, const char **argv);\n> diff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\n> new file mode 100755\n> index 0000000000..424fe4fc29\n> --- /dev/null\n> +++ b/t/t0095-bloom.sh\n> @@ -0,0 +1,113 @@\n> +#!/bin/sh\n> +\n> +test_description='test bloom.c'\n> +. ./test-lib.sh\n> +\n> +test_expect_success 'get bloom filters for commit with no changes' '\n> +\tgit init &&\n> +\tgit commit --allow-empty -m \"c0\" &&\n> +\tcat >expect <<-\\EOF &&\n> +\tFilter_Length:0\n> +\tFilter_Data:\n> +\tEOF\n> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'get bloom filter for commit with 10 changes' '\n> +\trm actual &&\n> +\trm expect &&\n> +\tmkdir smallDir &&\n> +\tfor i in $(test_seq 0 9)\n> +\tdo\n> +\t\techo $i >smallDir/$i\n> +\tdone &&\n> +\tgit add smallDir &&\n> +\tgit commit -m \"commit with 10 changes\" &&\n> +\tcat >expect <<-\\EOF &&\n> +\tFilter_Length:4\n> +\tFilter_Data:508928809087080a|8a7648210804001|4089824400951000|841ab310098051a8|\n> +\tEOF\n> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success EXPENSIVE 'get bloom filter for commit with 513 changes' '\n> +\trm actual &&\n> +\trm expect &&\n> +\tmkdir bigDir &&\n> +\tfor i in $(test_seq 0 512)\n> +\tdo\n> +\t\techo $i >bigDir/$i\n> +\tdone &&\n> +\tgit add bigDir &&\n> +\tgit commit -m \"commit with 513 changes\" &&\n> +\tcat >expect <<-\\EOF &&\n> +\tFilter_Length:0\n> +\tFilter_Data:\n> +\tEOF\n> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'compute bloom key for empty string' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:5615800c|5b966560|61174ab4|66983008|6c19155c|7199fab0|771ae004|\n> +\tFilter_Length:1\n> +\tFilter_Data:11000110001110|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \"\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'compute bloom key for whitespace' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:1bf014e6|8a91b50b|f9335530|67d4f555|d676957a|4518359f|b3b9d5c4|\n> +\tFilter_Length:1\n> +\tFilter_Data:401004080200810|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \" \" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'compute bloom key for a root level folder' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:1a21016f|fff1c06d|e5c27f6b|cb933e69|b163fd67|9734bc65|7d057b63|\n> +\tFilter_Length:1\n> +\tFilter_Data:aaa800000000|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \"A\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'compute bloom key for a root level file' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:e2d51107|30970605|7e58fb03|cc1af001|19dce4ff|679ed9fd|b560cefb|\n> +\tFilter_Length:1\n> +\tFilter_Data:a8000000000000aa|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \"file.txt\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'compute bloom key for a deep folder' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:864cf838|27f055cd|c993b362|6b3710f7|0cda6e8c|ae7dcc21|502129b6|\n> +\tFilter_Length:1\n> +\tFilter_Data:1c0000600003000|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \"A/B/C/D/E\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'compute bloom key for a deep file' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:07cdf850|4af629c7|8e1e5b3e|d1468cb5|146ebe2c|5796efa3|9abf211a|\n> +\tFilter_Length:1\n> +\tFilter_Data:4020100804010080|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \"A/B/C/D/E/file.txt\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_done\n"},{"id":"391874","messageId":"86tv3qqxyn.fsf@gmail.com","threadId":"52499","inReplyTo":"02b16d94227470059dcee2781e29ae7ae010f602.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 02/11] bloom: core Bloom filter implementation for changed paths","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-16T16:49:20Z","receivedAt":"2020-02-16T16:49:39Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[I'm sorry for accidentally sending unfinished version of this email]\n\n\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Add the core Bloom filter logic for computing the paths changed between a\n> commit and its first parent. For details on what Bloom filters are and how they\n> work, please refer to Dr. Derrick Stolee's blog post [1]. It provides a concise\n> explaination of the adoption of Bloom filters as described in [2] and [3].\n                                                                           ^^- to add\n>\n> 1. We currently use 7 and 10 for the number of hashes and the size of each\n>    entry respectively. They served as great starting values, the mathematical\n>    details behind this choice are described in [1] and [4]. The implementation,\n                                                                                ^^- to add\n>    while not completely open to it at the moment, is flexible enough to allow\n>    for tweaking these settings in the future.\n\nI don't know if it is worth it, but I think it should be size of each\nentry, or in other words number of bits per element in the set, as first\nvalue, and number of hashes as second.\n\nAbout where those values come from.  The idea is that you decide on the\nacceptable number of false positives, for example 1% (or 0.8% given that\nthe values must be integers); that gives you number of bits per element\ni.e. 10, and from there you can find optimal number of hashes i.e. 7.\nThe references mentioned (and Wikipedia article) have those equations.\n\n>\n>    Note: The performance gains we have observed with these values are\n>    significant enough that we did not need to tweak these settings.\n>    The performance numbers are included in the cover letter of this series\n>    and in the message of a subsequent commit where we use Bloom filters in\n>    to speed up `git log -- <path>`.\n\nAll right.\n\n>\n> 2. As described in the blog and in [3], we do not need 7 independent hashing\n>    functions. We use the Murmur3 hashing scheme. Seed it twice and then\n>    combine those to procure an arbitrary number of hash values.\n\nThe technique from [3] is called \"double hashing\" (Algorithm 1 and\nequation (4) on page 10).  Note that in this paper there is also\npresented \"enhanced double hashing\" scheme (Algorithm 2 and equation\n(6)) -- more about it later.\n\nThis is a standard technique from the hashing literature, called open\naddressing with double hashing in hash tables.\n\nThis \"enhanced double hashing\" technique is further analyzed in [6].\n\n[6] Adam Kirsch, Michael Mitzenmacher\n    \"Less Hashing, Same Performance: Building a Better Bloom Filter\"\n    https://www.eecs.harvard.edu/~michaelm/postscripts/esa2006a.pdf\n    https://doi.org/10.5555/1400123.1400125\n\n>\n> 3. The filters are sized according to the number of changes in the each commit,\n>    with minimum size of one 64 bit word.\n\nIf I understand it correctly (but which might not be entirely clear),\nthe filter size in bits is the number of changes^* times 10, rounded up\nto the nearest multiple of 64.\n\n[*] where the number of changes is the number of changed files (new blob\nobjects) _and_ the number of changed directories (new tree objects,\nexcluding root tree object change).\n\n\nThe interesting corner case, which might be worth specifying explicitly,\nis what happens in the case there are _no changes_ with respect to first\nparent (which can happen with either commit created with `git commit\n--allow-empty`, or merge created e.g. with `git merge --strategy=ours`).\nIs this case represented as Bloom filter of length 0, or as a Bloom\nfilter of length  of one 64-bit word which is minimal length composed of\nall 0's (0x0000000000000000)?\n\n>\n> 4. We fill the Bloom filters as (const char *data, int len) pairs as\n>    \"struct bloom_filter\"s in a commit slab.\n\nAll right.\n\n>\n> 5. The seed_murmur3 method is implemented as described in [5]. It hashes the\n>    given data using a given seed and produces a uniformly distributed hash\n>    value.\n\nActually there are two variants of Murmur3 hash, and we should specify\nwhich one we are using.  There is Murmur3_32 which returns 32-bit value,\nand Murmur3_128 which returns 128-bit value (which is different for x86\nand x64 versions).  We use Murmur3_32.\n\nAlso, seed_murmur3 is the name given the function, not the name of the\nmethod i.e. of a non-cryptographic hash function.\n\n\nOne question that one might as is why use Murmur3 hash instead for\nexample already implemented FNV hash from hashmap implementation (FNV\nhash i.e. Fowler–Noll–Vo hash function is another non-cryptographic hash\nfunction).  The answer is of course performance while maintaining good\nenough quality (and for Bloom filter there is no problem of \"hash\nflooding\" denial-of-service like for there is for a hash table -- no\nneed for SipHash or similar).\n\n>\n> [1] https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n\nI would write it in full, similar to subsequent bibliographical entries,\nthat is:\n\n  [1] Derrick Stolee\n      \"Supercharging the Git Commit Graph IV: Bloom Filters\"\n      https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n\nBut that is just a matter of style.\n\n>\n> [2] Flavio Bonomi, Michael Mitzenmacher, Rina Panigrahy, Sushil Singh, George Varghese\n>     \"An Improved Construction for Counting Bloom Filters\"\n>     http://theory.stanford.edu/~rinap/papers/esa2006b.pdf\n>     https://doi.org/10.1007/11841036_61\n>\n> [3] Peter C. Dillinger and Panagiotis Manolios\n>     \"Bloom Filters in Probabilistic Verification\"\n>     http://www.ccs.neu.edu/home/pete/pub/Bloom-filters-verification.pdf\n>     https://doi.org/10.1007/978-3-540-30494-4_26\n\nGood, we should be able to find them even if the URL with PDF stops\nworking for some reason.\n\n>\n> [4] Thomas Mueller Graf, Daniel Lemire\n>     \"Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters\"\n>     https://arxiv.org/abs/1912.08258\n>\n> [5] https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n>\n> Helped-by: Jeff King <peff@peff.net>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  Makefile              |   2 +\n>  bloom.c               | 228 ++++++++++++++++++++++++++++++++++++++++++\n>  bloom.h               |  56 +++++++++++\n>  t/helper/test-bloom.c |  84 ++++++++++++++++\n>  t/helper/test-tool.c  |   1 +\n>  t/helper/test-tool.h  |   1 +\n>  t/t0095-bloom.sh      | 113 +++++++++++++++++++++\n>  7 files changed, 485 insertions(+)\n>  create mode 100644 bloom.c\n>  create mode 100644 bloom.h\n>  create mode 100644 t/helper/test-bloom.c\n>  create mode 100755 t/t0095-bloom.sh\n\nAs I wrote earlier, In my opinion this patch could be split into three\nindividual single-functionality pieces, to make it easier to review and\naid in bisectability if needed.\n\n1. Add implementation of MurmurHash v3 (32-bit result)\n  \nInclude tests based on test-tool (creating file similar to the\nt/helper/test-hash.c, or enhancing to that file) that the implementation\nis correct, for example that 'The quick brown fox jumps over the lazy\ndog' or 'Hello world!' with a given seed (for example the default seed\nof 0) hashes to the same value as other implementations, including the\nreference implementation in https://github.com/aappleby/smhasher\n\n\n2. Add implementation of [variant of] Bloom filter\n\nInclude generic Bloom filter tests i.e. that it correctly answers \"yes\"\nand \"maybe\" (create filter, save it or print it, then use stored\nfilter), and tests specific to our implementation, namely that the size\nof the filter behaves as it should.\n\n\n3. Bloom filter implementation for changed paths\n\nHere include tests that use 'test-tool bloom get_filter_for_commit',\nthat filter for commit with no changes and for commit with more than 512\nchanges works correctly, that directories are added along the files,\netc.\n\n\nThis split would make it easier to distinguish if the problems with\ntests failing on big-endian architectures is caused by different output\nfrom our implementation of Murmur3 hash, different bit sequence in the\nBloom filter, or just different printed output of Bloom filter data.\n\n>\n> diff --git a/Makefile b/Makefile\n> index 6134104ae6..afba81f4a8 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -695,6 +695,7 @@ X =\n>  \n>  PROGRAMS += $(patsubst %.o,git-%$X,$(PROGRAM_OBJS))\n>  \n> +TEST_BUILTINS_OBJS += test-bloom.o\n>  TEST_BUILTINS_OBJS += test-chmtime.o\n>  TEST_BUILTINS_OBJS += test-config.o\n>  TEST_BUILTINS_OBJS += test-ctype.o\n> @@ -840,6 +841,7 @@ LIB_OBJS += base85.o\n>  LIB_OBJS += bisect.o\n>  LIB_OBJS += blame.o\n>  LIB_OBJS += blob.o\n> +LIB_OBJS += bloom.o\n>  LIB_OBJS += branch.o\n>  LIB_OBJS += bulk-checkin.o\n>  LIB_OBJS += bundle.o\n\nAll right.\n\n> diff --git a/bloom.c b/bloom.c\n> new file mode 100644\n> index 0000000000..6082193a75\n> --- /dev/null\n> +++ b/bloom.c\n> @@ -0,0 +1,228 @@\n> +#include \"git-compat-util.h\"\n> +#include \"bloom.h\"\n> +#include \"commit-graph.h\"\n> +#include \"object-store.h\"\n> +#include \"diff.h\"\n> +#include \"diffcore.h\"\n> +#include \"revision.h\"\n> +#include \"hashmap.h\"\n> +\n> +define_commit_slab(bloom_filter_slab, struct bloom_filter);\n> +\n> +struct bloom_filter_slab bloom_filters;\n\nAll right, this is needed to store per-commit Bloom filter data\n(inside-out object style, or in other jargon stored on slab).\n\n> +\n> +struct pathmap_hash_entry {\n> +    struct hashmap_entry entry;\n> +    const char path[FLEX_ARRAY];\n> +};\n\nO.K. this is used to add gather paths to add them all as elements to the\nBloom filter.\n\n> +\n> +static uint32_t rotate_right(uint32_t value, int32_t count)\n> +{\n> +\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n> +\tcount &= mask;\n> +\treturn ((value >> count) | (value << ((-count) & mask)));\n> +}\n\nHmmm... both the algoritm on Wikipedia, and reference implementation use\nrotate *left*, not rotate *right* in the implementation of Murmur3 hash,\nsee\n\n  https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n  https://github.com/aappleby/smhasher/blob/master/src/MurmurHash3.cpp#L23\n\n\ninline uint32_t rotl32 ( uint32_t x, int8_t r )\n{\n  return (x << r) | (x >> (32 - r));\n}\n\n> +\n> +/*\n> + * Calculate a hash value for the given data using the given seed.\n> + * Produces a uniformly distributed hash value.\n> + * Not considered to be cryptographically secure.\n> + * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n> + **/\n    ^^-- why two _trailing_ asterisks?\n\nPerhaps it would be worth it to add that this hash function is intended\nto be fast while being reasonably good (it is distributed randomly\nenough, and it doesn't have too many hash collisions on typical inputs).\nBut this might be too much for a comment.\n\n> +static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n\nA few things: name of the function, type of parameters and ordering of\nparameters.\n\n\nAbout the name: when I first saw seed_murmur3() used, I thought it was\n_setting_ the seed, not that it was returning the 32-bit hash value.\nOther implementations use either murmur3_32, MurmurHash3_x86_32, or\nsomething similar like hashmurmur3_32.  If we were to specify that\n'seed' is one of parameters, then using this word as part of suffix\nwould be better than using seed_ prefix; if we need it at all.\n\nBecause there is 32-bit and 128-bit variants of Murmur3, I think the _32\nsuffix should be a part of function name.\n\nIn short, I think that the name of the function should be murmur3_32, or\nmurmurhash3_32, or possibly murmur3_32_seed, or something like that.\n\n\nAbout types of parameters and the return type of function: I understand\nthat 'data' parameter is of type 'const char *', instead of more generic\n'const uint8_t*' or 'const void *' because of what we will be using the\nhash function for.  On the other hand taking a look at implementation of\nFNV hash function in hashmap.{c,h} we see that the 'str*' variants take\n'const char *' parameter _without_ length, and 'mem*' variants take\n'const void *' parmeter with length of data.\n\nShouldn't 'len' parameter be of 'size_t' type, rather than 'int'?  Both\nthe example implementation in C on Wikipedia page, and implementation in\nC in qLibc use 'size_t'; the implementation of FNV hash in hashmap in\nGit also uses 'size_t' (while admittedly the reference implementation in\nC++ of Austin Appleby uses 'int' type for len parameter).\n\nFor 32-bit output variant of Murmur3 hash, using uint32_t as return type\nis just fine.  The '*hash*' functions from hashmap.{c,h} use 'unsigned\nint' but I think 'uint32_t' is better.\n\n\nAbout names and ordering of parameters: the 'seed' or 'hash_seed'\nparameter should be either first or last; it is a matter of preference.\nWhile example implementation on Wikipedia page, Appleby's reference\nimplementation in C++ have 'seed' as last parameter, memihash_cont()\nfrom hashmap.c in Git has it as first parameter.\n\nIn short: I'm fine with either order (seed parameter first or last), and\neither name (be it 'seed' or 'hash_seed').\n\n> +{\n> +\tconst uint32_t c1 = 0xcc9e2d51;\n> +\tconst uint32_t c2 = 0x1b873593;\n> +\tconst uint32_t r1 = 15;\n> +\tconst uint32_t r2 = 13;\n> +\tconst uint32_t m = 5;\n> +\tconst uint32_t n = 0xe6546b64;\n> +\tint i;\n> +\tuint32_t k1 = 0;\n> +\tconst char *tail;\n> +\n> +\tint len4 = len / sizeof(uint32_t);\n> +\n> +\tconst uint32_t *blocks = (const uint32_t*)data;\n> +\n> +\tuint32_t k;\n> +\tfor (i = 0; i < len4; i++)\n> +\t{\n> +\t\tk = blocks[i];\n\nIMPORTANT: There is a comment around there in the example implementation\nin C on Wikipedia that this operation above is a source of differing\nresults across endianness.  The pseudo-code description of the algorithm\non Wikipedia (above of C code) says that endian swapping is only\nnecessary on big-endian machines (and that it is needed to place the\nmeaningful digits towards the low end of the value, to not be discarded\nby the modulo arithmetic under overflow).\n\nThe original / reference implementation by Austin Appleby in C++ uses\ngetblock32() function for doing the block read... but it doesn't\nactually implement the endian-swapping on big-endian architecture:\n\n  //-----------------------------------------------------------------------------\n  // Block read - if your platform needs to do endian-swapping or can only\n  // handle aligned reads, do the conversion here\n\n  FORCE_INLINE uint32_t getblock32 ( const uint32_t * p, int i )\n  {\n    return p[i];\n  }\n\nReferences:\n-----------\n1. https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n2. https://github.com/aappleby/smhasher/blob/master/src/MurmurHash3.cpp\n\n> +\t\tk *= c1;\n> +\t\tk = rotate_right(k, r1);\n\nIt is  k ROL r1 / ROTL32(k,15) / (k << 15) | (k >> (32 - 15))\n(in other implementations), not rotate_right.\n\n> +\t\tk *= c2;\n> +\n> +\t\tseed ^= k;\n> +\t\tseed = rotate_right(seed, r2) * m + n;\n\nIt is  hash ROL r2 / ROTL32(h1,13) / (h << 13) | (h >> (32 - 13))\n(in other implementations), not rotate_right.\n\nReferences:\n-----------\n1. https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n2. https://github.com/aappleby/smhasher/blob/master/src/MurmurHash3.cpp#L94\n3. https://github.com/wolkykim/qlibc/blob/master/src/utilities/qhash.c#L258\n\n> +\t}\n> +\n> +\ttail = (data + len4 * sizeof(uint32_t));\n\nHmmm... in the pseudocode implementation on Wikipedia this is the place\nwhere one needs to respect endianness:\n\n    with any remainingBytesInKey do\n        remainingBytes ← SwapToLittleEndian(remainingBytesInKey)\n        // Note: Endian swapping is only necessary on big-endian machines.\n        //       The purpose is to place the meaningful digits towards the low end of the value,\n        //       so that these digits have the greatest potential to affect the low range digits\n        //       in the subsequent multiplication.  Consider that locating the meaningful digits\n        //       in the high range would produce a greater effect upon the high digits of the\n        //       multiplication, and notably, that such high digits are likely to be discarded\n        //       by the modulo arithmetic under overflow.  We don't want that.\n\nOn the other hand in the reference Appleby's C++ implementation the\nendian-swapping is [ssumed to be] done only in the loop over data.\nEither should be enough alone, but doing swapping for remaining bytes\nonly would work, it would be a better solution -- you do swap only once,\nat the end.\n\nIt looks like the Chromium implementation in C by Shane Day (public\ndomain) uses the second solution; well almost, see:\nhttps://chromium.googlesource.com/external/smhasher/+/5b8fd3c31a58b87b80605dca7a64fad6cb3f8a0f/PMurHash.c#189\n\n> +\n> +\tswitch (len & (sizeof(uint32_t) - 1))\n> +\t{\n> +\tcase 3:\n> +\t\tk1 ^= ((uint32_t)tail[2]) << 16;\n> +\t\t/*-fallthrough*/\n> +\tcase 2:\n> +\t\tk1 ^= ((uint32_t)tail[1]) << 8;\n> +\t\t/*-fallthrough*/\n> +\tcase 1:\n> +\t\tk1 ^= ((uint32_t)tail[0]) << 0;\n> +\t\tk1 *= c1;\n> +\t\tk1 = rotate_right(k1, r1);\n\nIt is  remainingBytes ROL r1 / ROTL32(k1,15) / (k << 15) | (k >> (32 - 15)) \n(in other implementations), not rotate_right.  The same references as\nbefore.\n\n> +\t\tk1 *= c2;\n> +\t\tseed ^= k1;\n> +\t\tbreak;\n> +\t}\n> +\n> +\tseed ^= (uint32_t)len;\n> +\tseed ^= (seed >> 16);\n> +\tseed *= 0x85ebca6b;\n> +\tseed ^= (seed >> 13);\n> +\tseed *= 0xc2b2ae35;\n> +\tseed ^= (seed >> 16);\n> +\n> +\treturn seed;\n> +}\n\nIn https://public-inbox.org/git/ba856e20-0a3c-e2d2-6744-b9abfacdc465@gmail.com/\nyou posted \"[PATCH] Process bloom filter data as 1 byte words\".\nThis may avoid the Big-endian vs Little-endian confusion,\nthat is wrong results on Big-endian architectures, but\nit also may slow down the algorithm.\n\nThe public domain implementation in PMurHash.c in SMHasher\n(re)implementation in Chromium (see URL above) fall backs to 1-byte\noperations only if it doesn't know the endianness (or if it is neither\nlittle-endian, nor big-endian, i.e. middle-endian or mixed-endian --\nthough I doubt that Git works correctly on mixed-endian anyway).\n\n\nSidenote: it looks like the current implementation if Murmur hash in\nCromium uses MurmurHash3_x86_32, i.e. little-endian unaligned-safe\nimplementation, but prepares data by swapping with StringToLE32\nhttps://github.com/chromium/chromium/blob/master/components/variations/variations_murmur_hash.h\n\n\nAssuming that the terminating NUL (\"\\0\") character of a c-string is not\nincluded in hash calculations, then murmur3_x86_32 hash has the\nfollowing results (all results are for seed equal 0):\n\n''               -> 0x00000000\n' '              -> 0x7ef49b98\n'Hello world!'   -> 0x627b0c2c\n'The quick brown fox jumps over the lazy dog'   -> 0x2e4ff723\n\nC source (from Wikipedia): https://godbolt.org/z/ofa2p8\nC++ source (Appleby's):    https://godbolt.org/z/BoSt6V\n\nThe implementation provided in this patch, with rotate_right (instead of\nrotate_left) gives, on little-endian machine, different results:\n\n''               -> 0x00000000\n' '              -> 0xd1f27e64\n'Hello world!'   -> 0xa0791ad7\n'The quick brown fox jumps over the lazy dog'   -> 0x99f1676c\n\nhttps://github.com/gitgitgadget/git/blob/e1b076a714d611e59d3d71c89221e41a3427fae4/bloom.c#L21\nC source (via GitGitGadget): https://godbolt.org/z/R9s8Tt\n\nSidenote: While Godbolt.org site supports compiling with many different\ncompilers, including GCC, Clang (LLVM), icc (Intel), MSVC (via Wine),\nand cross compiling for different platforms, including x86_64, ARM, MIPS,\nPowerPC, power64 and power64le, AVR, it allows for execution only on\nx86_64 i.e. little-endian.\n\n\nWe could create test similar to the one for SHA-1 and SHA-256 in\nt/t0015-hash.sh but for murmur3, for example:\n\n  test_expect_success 'test basic Murmur3_32 hash values' '\n  \tprintf \" \" | test-tool murmur3_32 0 >actual &&\n        printf \"7ef49b98\" >expected &&\n        test_cmp expected actual &&\n        ...\n  '\n\nor\n\n  test_expect_success 'test basic Murmur3_32 hash values' '\n  \tprintf \" \" | test-tool murmur3_32 0 >actual &&\n        grep \"7ef49b98\" actual &&\n        ...\n  '\n\n> +\n> +static inline uint64_t get_bitmask(uint32_t pos)\n> +{\n> +\treturn ((uint64_t)1) << (pos & (BITS_PER_WORD - 1));\n> +}\n\nAll right, that creates 64-bit wide mask with 1 bit set to 1 for a\n64-bit word within filter data.  I just wonder if the trick with the &\noperation is truly faster than using simpler to understand modulo\nwith compiler optimizations.\n\n   static inline uint64_t get_bitmask(uint32_t pos)\n   {\n   \treturn ((uint64_t)1) << (pos % BITS_PER_WORD);\n   }\n\nAnyway, looks good (beside naming things, but I don't have better\nproposal, and the function is static i.e. file-local anyway).\n\n> +\n> +void load_bloom_filters(void)\n> +{\n> +\tinit_bloom_filter_slab(&bloom_filters);\n> +}\n\n\nActually this function doesn't load anything.  Perhaps it should be\nnamed init_bloom_filters() or init_bloom_filters_storage(), or\nbloom_filters_init()?\n\n> +\n> +void fill_bloom_key(const char *data,\n> +\t\t\t\t\tint len,\n> +\t\t\t\t\tstruct bloom_key *key,\n> +\t\t\t\t\tstruct bloom_filter_settings *settings)\n\nThe last parameter could be of 'const bloom_filter_settings *' type.\n\n> +{\n> +\tint i;\n> +\tconst uint32_t seed0 = 0x293ae76f;\n> +\tconst uint32_t seed1 = 0x7e646e2c;\n\nWhere did those seeds values came from?\n\n> +\tconst uint32_t hash0 = seed_murmur3(seed0, data, len);\n> +\tconst uint32_t hash1 = seed_murmur3(seed1, data, len);\n> +\n> +\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n> +\tfor (i = 0; i < settings->num_hashes; i++)\n> +\t\tkey->hashes[i] = hash0 + i * hash1;\n\nNote that in [3] authors say that double hashing technique has some\nproblems.  For one, we should ensure that hash1 is not zero, and even\nbetter that it is odd (which makes it relatively prime to filter size\nwhich is multiple of 64).  It also suffers from something called\n\"approximate fingerprint collisions\".\n\nThat is why the define \"enhanced double hashing\" technique, which does\nnot suffer from those problems (Algorithm 2, page 11/15).\n\n  +\tfor (i = 0; i < settings->num_hashes; i++) {\n  +\t\tkey->hashes[i] = hash0;\n  +\n  +\t\thash0 = hash0 + hash1;\n  +\t\thash1 = hash1 + i;\n  +\t}\n\nThis can also be written in closed form, based on equation (6)\n\n  +\tfor (i = 0; i < settings->num_hashes; i++)\n  +\t\tkey->hashes[i] = hash0 + i * hash1 + i*(i*i - 1)/6;\n\n\nIn later paper [6] the closed form for \"enhanced double hashing\"\n(p. 188) is slightly modified (or rather they use different variant of\nthis technique):\n\n  +\tfor (i = 0; i < settings->num_hashes; i++)\n  +\t\tkey->hashes[i] = hash0 + i * hash1 + i*i;\n\nThis is a variant of more generic \"enhanced double hashing\", section\n5.2 (Enhanced) Double Hashing Schemes (page 199):\n\n        h_1(u) + i h_2(u) + f(i)    mod m\n\nwith f(i) = i^2 = i*i.\n\nThey have tested that enhanced double hashing with both f(i) equal i*i\nand equal i*i*i, and triple hashing technique, and they have found that\nit performs slightly better than straight double hashing technique\n(Fig. 1, page 212, section 3).\n\n> +}\n> +\n> +void add_key_to_filter(struct bloom_key *key,\n> +\t\t\t\t\t   struct bloom_filter *filter,\n> +\t\t\t\t\t   struct bloom_filter_settings *settings)\n\nHere again the 'settings' argument can be const (as can the 'key'\nparameter).\n\n> +{\n> +\tint i;\n> +\tuint64_t mod = filter->len * BITS_PER_WORD;\n> +\n> +\tfor (i = 0; i < settings->num_hashes; i++) {\n> +\t\tuint64_t hash_mod = key->hashes[i] % mod;\n> +\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n> +\n> +\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n> +\t}\n> +}\n\nAll right, bloom_key is an intermediate representation that is used both\nfor creating Bloom filter, and for querying it.  In the latter case the\nsame path may be tested against Bloom filters for commits with different\nnumber of (blob and tree) changes, and thus against Bloom filters with\ndifferent lengths.  It makes sense for bloom_key to store just values of\nhash functions, without arithmetics modulo filter size.\n\nThough I think it could be a good idea to create add_str_to_filter() as\na wrapper around add_key_to_filter() and fill_bloom_key() functions.\n\n> +\n> +struct bloom_filter *get_bloom_filter(struct repository *r,\n> +\t\t\t\t      struct commit *c)\n> +{\n> +\tstruct bloom_filter *filter;\n> +\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n> +\tint i;\n> +\tstruct diff_options diffopt;\n> +\n> +\tif (!bloom_filters.slab_size)\n> +\t\treturn NULL;\n\nThis is testing that commit slab for per-commit Bloom filters is\ninitialized, isn't it?\n\nFirst, should we write the condition as\n\n\tif (!bloom_filters.slab_size)\n\nor would the following be more readable\n\n\tif (bloom_filters.slab_size == 0)\n\nSecond, should we return NULL, or should we just initialize the slab?\nOr is non-existence of slab treated as a signal that the Bloom filters\nmechanism is turned off?\n\n> +\n> +\tfilter = bloom_filter_slab_at(&bloom_filters, c);\n\nWouldn't it be better to check if the data for commit exists already on\nthe slab, and create the Bloom filter for commit changes only if it does\nnot exists, i.e.:\n\n  +\tfilter = bloom_filter_slab_peek(&bloom_filters, c);\n  +\tif (filter)\n  +\t\treturn filter;\n  +\tfilter = bloom_filter_slab_at(&bloom_filters, c);\n\n> +\n> +\trepo_diff_setup(r, &diffopt);\n> +\tdiffopt.flags.recursive = 1;\n> +\tdiff_setup_done(&diffopt);\n\nI'll punt on checking this.  Looks all right from first glance, and\nfollows calling sequence in https://github.com/git/git/blob/master/diff.h#L26\n\n> +\n> +\tif (c->parents)\n> +\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &diffopt);\n> +\telse\n> +\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n> +\tdiffcore_std(&diffopt);\n\nAll right, that computes first-parent diff (or diff from empty tree of\nthere are no parents).\n\n> +\n> +\tif (diff_queued_diff.nr <= 512) {\n\nFirst, shouldn't this magic value 512 be hidden behind some symbolic\nname (some preprocessor constant), e.g. BLOOM_MAX_CHANGES?  On the other\nhand this value is used only once (except tests), so it might be not\nworth it -- especially coming up with a good name.\n\nSecond, there is a minor issue that diff_queue_struct.nr stores the\nnumber of filepairs, that is the number of changed files, while the\nnumber of elements added to Bloom filter is number of changed blobs and\ntrees.  For example if the following files are changed:\n\n  sub/dir/file1\n  sub/file2\n\nthen diff_queued_diff.nr is 2, but number of elements to be added to\nBloom filter is 4.\n\n  sub/dir/file1\n  sub/file2\n  sub/dir/\n  sub/\n\nI'm not sure if it matters in practice.\n\n> +\t\tstruct hashmap pathmap;\n> +\t\tstruct pathmap_hash_entry* e;\n> +\t\tstruct hashmap_iter iter;\n> +\t\thashmap_init(&pathmap, NULL, NULL, 0);\n\nStylistic issue: I have just noticed that here (and in some other\nplaces), but not in all cases, you declare pointer types with asterisk\ncuddled to type name, not to variable name, which contradicts\nCodingGuidelines:\n\n - When declaring pointers, the star sides with the variable\n   name, i.e. \"char *string\", not \"char* string\" or\n   \"char * string\".  This makes it easier to understand code\n   like \"char *string, c;\".\n\nIn this case it should be\n\n  +\t\tstruct pathmap_hash_entry *e;\n\nIn many other places in this patch it is correct, though.\n\n> +\n> +\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n> +\t\t\tconst char* path = diff_queued_diff.queue[i]->two->path;\n\nIs that correct that we consider only post-image name for storing\nchanges in Bloom filter?  Currently if file was renamed (or deleted), it\nis considered changed, and `git log -- <old-name>` lists commit that\nchanged file name too.\n\n> +\t\t\tconst char* p = path;\n\nIt should be \"const char *\" for both.\n\n> +\n> +\t\t\t/*\n> +\t\t\t* Add each leading directory of the changed file, i.e. for\n> +\t\t\t* 'dir/subdir/file' add 'dir' and 'dir/subdir' as well, so\n> +\t\t\t* the Bloom filter could be used to speed up commands like\n> +\t\t\t* 'git log dir/subdir', too.\n> +\t\t\t*\n> +\t\t\t* Note that directories are added without the trailing '/'.\n> +\t\t\t*/\n> +\t\t\tdo {\n> +\t\t\t\tchar* last_slash = strrchr(p, '/');\n> +\n> +\t\t\t\tFLEX_ALLOC_STR(e, path, path);\n\nHere first 'path' is the field name, i.e. pathmap_hash_entry.path,\nsecond 'path' is the name of local variable, aliased also to 'p'.\n\n> +\t\t\t\thashmap_entry_init(&e->entry, strhash(p));\n\nI don't know why both 'path' and 'p' are used, while both point to the\nsame memory (and thus have the same contents).  It is a bit confusing.\nSee also my previous comment.\n\n> +\t\t\t\thashmap_add(&pathmap, &e->entry);\n> +\n> +\t\t\t\tif (!last_slash)\n> +\t\t\t\t\tlast_slash = (char*)p;\n> +\t\t\t\t*last_slash = '\\0';\n> +\n> +\t\t\t} while (*p);\n\nLooks good.  We overwrite '/' with '\\0', and gather shrinking pathnames\nalong the way.\n\n> +\n> +\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n> +\t\t}\n> +\n> +\t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n\nAll right, this is division by BITS_PER_WORD, rounding up.\n\nSidenote: I see now why hashmap was used, it was to be able to get\nnumber of unique changes (changed blobs and trees) easily.\n\n> +\t\tfilter->data = xcalloc(filter->len, sizeof(uint64_t));\n> +\n> +\t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n> +\t\t\tstruct bloom_key key;\n> +\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, &settings);\n> +\t\t\tadd_key_to_filter(&key, filter, &settings);\n> +\t\t}\n\nAll right.\n\n> +\n> +\t\thashmap_free_entries(&pathmap, struct pathmap_hash_entry, entry);\n> +\t} else {\n> +\t\tfor (i = 0; i < diff_queued_diff.nr; i++)\n> +\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n\nAll right, that frees the memory taken by diff results.\n\n> +\t\tfilter->data = NULL;\n> +\t\tfilter->len = 0;\n\nThis needs to be explicitly stated both in the commit message and in the\nAPI documentation (in comments) that bloom_filter.len == 0 means \"no\ndata\", while \"no changes\" is represented as bloom_filter with len == 1\nand *data == (uint64_t)0;\n\nEDIT: actually \"no changes\" is also represented as bloom_filter with len\nequal 0, as it turns out.\n\nOne possible alternative could be representing \"no data\" value with\nBloom filter of length 1 and all 64 bits set to 1, and \"no changes\"\nrepresented as filter of length 0.  This is not unambiguous choice!\n\n> +\t}\n> +\n> +\tfree(diff_queued_diff.queue);\n> +\tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n> +\n> +\treturn filter;\n> +}\n\nAll right.\n\n> +\n> +int bloom_filter_contains(struct bloom_filter *filter,\n> +\t\t\t  struct bloom_key *key,\n> +\t\t\t  struct bloom_filter_settings *settings)\n\nIt might be good idea to define enum for return values, that is\nNO_DATA = -1, NO = 0, MAYBE = 1.\n\n> +{\n> +\tint i;\n> +\tuint64_t mod = filter->len * BITS_PER_WORD;\n> +\n> +\tif (!mod)\n> +\t\treturn -1;\n\nAll right, it is different way of writing\n\n\tif (filter->len == 0)\n\t\treturn -1;\n\nwhich means \"no data\" (too many elements for Bloom filter to store).\nEDIT: or \"no changes\".\n\n> +\n> +\tfor (i = 0; i < settings->num_hashes; i++) {\n> +\t\tuint64_t hash_mod = key->hashes[i] % mod;\n> +\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n> +\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n> +\t\t\treturn 0;\n\nAll right, if any of hash functions (hash results) doesn't match what is\nstored in filter, then the key cannot be contained in the Bloom filter.\n\n> +\t}\n> +\n> +\treturn 1;\n\nAll right, otherwise the key is probably included in filter, but may be\nfalse positive (with around 1% probability in theory).\n\nThis means that if we get value of 0, we can skip checking the diff; we\nknow commit is TREESAME with respect to the path given.\n\n> +}\n> diff --git a/bloom.h b/bloom.h\n> new file mode 100644\n> index 0000000000..7f40c751f7\n> --- /dev/null\n> +++ b/bloom.h\n> @@ -0,0 +1,56 @@\n> +#ifndef BLOOM_H\n> +#define BLOOM_H\n\nShould we #include the stdint.h header for uint32_t and uint64_t types?\n\n> +\n> +struct commit;\n> +struct repository;\n> +struct commit_graph;\n> +\n\nPerhaps we should add block comment for this struct, like there is one\nfor struct bloom_filter below.\n\n> +struct bloom_filter_settings {\n> +\tuint32_t hash_version;\n> +\tuint32_t num_hashes;\n> +\tuint32_t bits_per_entry;\n\nI guess that the type uint32_t was chosen to make it easier to store\nthis information and later retrieve it from the commit-graph file, isn't\nit?  Otherwise those types are much too large for sensible range of\nvalues (which would all fit in 8-bits byte).\n\n> +};\n> +\n> +#define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n> +#define BITS_PER_WORD 64\n\nSidenote: While CodingGuidelines explicitly says:\n\n - We try to support a wide range of C compilers to compile Git with,\n   including old ones.  You should not use features from newer C\n   standard, even if your compiler groks them.\n\n   There are a few exceptions to this guideline:\n\n   [...]\n\n   . since mid 2017 with cbc0f81d, we have been using designated\n     initializers for struct (e.g. \"struct t v = { .val = 'a' };\").\n\nI don't think however that using designated initializers in\nDEFAULT_BLOOM_FILTER_SETTINGS is needed, as this preprocessor constant\nis just below the definition of struct bloom_filter_settings type.\n\n> +\n> +/*\n> + * A bloom_filter struct represents a data segment to\n> + * use when testing hash values. The 'len' member\n> + * dictates how many uint64_t entries are stored in\n> + * 'data'.\n> + */\n> +struct bloom_filter {\n> +\tuint64_t *data;\n> +\tint len;\n> +};\n\nJust wondering: is there any advantage or disadvantage to putting 'len'\nfield first (i.e. before 'data') versus putting it after (i.e. after\n'data')?  Is there a convention that Git uses?\n\n> +\n> +/*\n> + * A bloom_key represents the k hash values for a\n> + * given hash input. These can be precomputed and\n> + * stored in a bloom_key for re-use when testing\n> + * against a bloom_filter.\n\nWe might want to add that the number of hash values is given by Bloom\nfilter settings, and it is assumed to be the same for all bloom_key\nvariables / objects.\n\n> + */\n> +struct bloom_key {\n> +\tuint32_t *hashes;\n> +};\n> +\n> +void load_bloom_filters(void);\n> +\n> +void fill_bloom_key(const char *data,\n> +\t\t    int len,\n> +\t\t    struct bloom_key *key,\n> +\t\t    struct bloom_filter_settings *settings);\n> +\n> +void add_key_to_filter(struct bloom_key *key,\n> +\t\t\t\t\t   struct bloom_filter *filter,\n> +\t\t\t\t\t   struct bloom_filter_settings *settings);\n> +\n> +struct bloom_filter *get_bloom_filter(struct repository *r,\n> +\t\t\t\t      struct commit *c);\n> +\n> +int bloom_filter_contains(struct bloom_filter *filter,\n> +\t\t\t  struct bloom_key *key,\n> +\t\t\t  struct bloom_filter_settings *settings);\n> +\n> +#endif\n> diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n> new file mode 100644\n> index 0000000000..331957011b\n> --- /dev/null\n> +++ b/t/helper/test-bloom.c\n> @@ -0,0 +1,84 @@\n> +#include \"test-tool.h\"\n> +#include \"git-compat-util.h\"\n> +#include \"bloom.h\"\n> +#include \"test-tool.h\"\n> +#include \"cache.h\"\n> +#include \"commit-graph.h\"\n> +#include \"commit.h\"\n> +#include \"config.h\"\n> +#include \"object-store.h\"\n> +#include \"object.h\"\n> +#include \"repository.h\"\n> +#include \"tree.h\"\n> +\n> +struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n> +\n> +static void print_bloom_filter(struct bloom_filter *filter) {\n> +\tint i;\n> +\n> +\tif (!filter) {\n> +\t\tprintf(\"No filter.\\n\");\n> +\t\treturn;\n> +\t}\n> +\tprintf(\"Filter_Length:%d\\n\", filter->len);\n> +\tprintf(\"Filter_Data:\");\n> +\tfor (i = 0; i < filter->len; i++){\n> +\t\tprintf(\"%\"PRIx64\"|\", filter->data[i]);\n> +\t}\n> +\tprintf(\"\\n\");\n> +}\n> +\n> +static void add_string_to_filter(const char *data, struct bloom_filter *filter) {\n> +\t\tstruct bloom_key key;\n> +\t\tint i;\n> +\n> +\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n> +\t\tprintf(\"Hashes:\");\n> +\t\tfor (i = 0; i < settings.num_hashes; i++){\n> +\t\t\tprintf(\"%08x|\", key.hashes[i]);\n> +\t\t}\n> +\t\tprintf(\"\\n\");\n> +\t\tadd_key_to_filter(&key, filter, &settings);\n> +}\n> +\n> +static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n> +{\n> +\tstruct commit *c;\n> +\tstruct bloom_filter *filter;\n> +\tsetup_git_directory();\n> +\tc = lookup_commit(the_repository, commit_oid);\n> +\tfilter = get_bloom_filter(the_repository, c);\n> +\tprint_bloom_filter(filter);\n> +}\n> +\n> +int cmd__bloom(int argc, const char **argv)\n> +{\n> +    if (!strcmp(argv[1], \"generate_filter\")) {\n> +\t\tstruct bloom_filter filter;\n> +\t\tint i = 2;\n> +\t\tfilter.len =  (settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n> +\t\tfilter.data = xcalloc(filter.len, sizeof(uint64_t));\n> +\n> +\t\tif (!argv[2]){\n> +\t\t\tdie(\"at least one input string expected\");\n> +\t\t}\n> +\n> +\t\twhile (argv[i]) {\n> +\t\t\tadd_string_to_filter(argv[i], &filter);\n> +\t\t\ti++;\n> +\t\t}\n> +\n> +\t\tprint_bloom_filter(&filter);\n> +\t}\n> +\n> +\tif (!strcmp(argv[1], \"get_filter_for_commit\")) {\n> +\t\tstruct object_id oid;\n> +\t\tconst char *end;\n> +\t\tif (parse_oid_hex(argv[2], &oid, &end))\n> +\t\t\tdie(\"cannot parse oid '%s'\", argv[2]);\n> +\t\tload_bloom_filters();\n> +\t\tget_bloom_filter_for_commit(&oid);\n> +\t}\n> +\n> +\treturn 0;\n> +}\n\n\nI won't comment on test-tool code, as I think the Bloom filter and\nMurmur3 hash tests should be structured differently, which would\ncompletely change test-bloom.c code.\n\n> diff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\n> index c9a232d238..ca4f4b0066 100644\n> --- a/t/helper/test-tool.c\n> +++ b/t/helper/test-tool.c\n> @@ -14,6 +14,7 @@ struct test_cmd {\n>  };\n>  \n>  static struct test_cmd cmds[] = {\n> +\t{ \"bloom\", cmd__bloom },\n>  \t{ \"chmtime\", cmd__chmtime },\n>  \t{ \"config\", cmd__config },\n>  \t{ \"ctype\", cmd__ctype },\n\n> diff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\n> index c8549fd87f..05d2b32451 100644\n> --- a/t/helper/test-tool.h\n> +++ b/t/helper/test-tool.h\n> @@ -4,6 +4,7 @@\n>  #define USE_THE_INDEX_COMPATIBILITY_MACROS\n>  #include \"git-compat-util.h\"\n>  \n> +int cmd__bloom(int argc, const char **argv);\n>  int cmd__chmtime(int argc, const char **argv);\n>  int cmd__config(int argc, const char **argv);\n>  int cmd__ctype(int argc, const char **argv);\n\nAll right, looks good.\n\n> diff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\n> new file mode 100755\n> index 0000000000..424fe4fc29\n> --- /dev/null\n> +++ b/t/t0095-bloom.sh\n> @@ -0,0 +1,113 @@\n> +#!/bin/sh\n> +\n> +test_description='test bloom.c'\n\nThis description is a bit lackluster...\n\n> +. ./test-lib.sh\n> +\n> +test_expect_success 'get bloom filters for commit with no changes' '\n> +\tgit init &&\n> +\tgit commit --allow-empty -m \"c0\" &&\n> +\tcat >expect <<-\\EOF &&\n> +\tFilter_Length:0\n> +\tFilter_Data:\n> +\tEOF\n> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n\nA few things.  First, I wonder why we need to provide object ID;\ncouldn't 'test-tool bloom get_filter_for_commit' parse commit-ish\nargument, or would it make it too complicated for no reason?\n\nSecond, why both \"no changes\" (here) and \"no data\" have the same\nrepresentation of filter with length equal 0?  Let's take a look at the\ncode.\n\nFor no changes:\n\n  filter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n                 ^^^^^^^^^^^^^^^^^^^^^^^^^^ == 0  for no changes\n                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n                 \\-- == 0 + BITS_PER_WORD - 1     for no changes\n                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n                 \\-- == 0  for no changes\n  filter->data = xcalloc(filter->len, sizeof(uint64_t));\n                         ^^^^^^^^^^^ == 0  for no changes\n                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n                 \\-- is NULL or unique pointer that can be passed to free()\n\nFor more than 512 changed files:\n\n  filter->data = NULL;\n  filter->len = 0;\n\nNot being able to distinguish between \"no data\" and \"no changes in the\ncommit\" cases means that we would always perform full diff for commit\nwith no changes, unnecessarily.  Fortunately there should be no hit to\nperformance, as in this case we need to simply compare objects IDs of\ntop tree to know that there is no change.\n\nIf it is a design decision we go with, it should be in my opinion at\nleast explained in the commit message explicitly.\n\n> +\n> +test_expect_success 'get bloom filter for commit with 10 changes' '\n> +\trm actual &&\n> +\trm expect &&\n> +\tmkdir smallDir &&\n> +\tfor i in $(test_seq 0 9)\n> +\tdo\n> +\t\techo $i >smallDir/$i\n> +\tdone &&\n> +\tgit add smallDir &&\n> +\tgit commit -m \"commit with 10 changes\" &&\n> +\tcat >expect <<-\\EOF &&\n> +\tFilter_Length:4\n> +\tFilter_Data:508928809087080a|8a7648210804001|4089824400951000|841ab310098051a8|\n> +\tEOF\n> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n\nThis test is in my opinion fragile, as it unnecessarily test the\nimplementation details instead of the functionality provided.  If we\nchange the hashing scheme (for example going from double hashing to some\nvariant of enhanced double hashing), or change the base hash function\n(for example from Murmur3_32 to xxHash_64), or change the number of hash\nfunctions (perhaps because changing of number of bits per element, and\nthus optimal number of hash functions from 7 to 6), or change from\n64-bit word blocks to 32-bit word blocks, the test would have to be\nchanged.\n\nWhat I think would be a good test is something like t/t0011-hashmap.sh.\nFor example test that the Bloom filter size scales correctly could look\nlike this:\n\n   test_bloom() {\n   \techo \"$1\" | test-tool bloom $3 >actual &&\n   \techo \"$2\" >expect &&\n   \ttest_cmp expect actual\n   }\n\n   test_expect_success 'Bloom filter for commit size scales with number of changes' '\n   \tmkdir smallDir &&\n\tfor i in $(test_seq 0 9)\n\tdo\n\t\techo $i >smallDir/$i\n\tdone &&\n\tgit add smallDir &&\n\tgit commit -m \"commit with 10 changes\" &&\n        HEAD=$(git rev-parse HEAD) &&\n        cat | test-tool bloom >actual <<-EOF &&\n        add-commit $HEAD\n        len-commit $HEAD\n        EOF\n        echo \"4\" >expect &&\n        test_cmp expect actual\n   '\n\n> +\n> +test_expect_success EXPENSIVE 'get bloom filter for commit with 513 changes' '\n> +\trm actual &&\n> +\trm expect &&\n> +\tmkdir bigDir &&\n> +\tfor i in $(test_seq 0 512)\n> +\tdo\n> +\t\techo $i >bigDir/$i\n> +\tdone &&\n> +\tgit add bigDir &&\n> +\tgit commit -m \"commit with 513 changes\" &&\n> +\tcat >expect <<-\\EOF &&\n> +\tFilter_Length:0\n> +\tFilter_Data:\n> +\tEOF\n> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n\nAll right, it is good test to have (though perhaps in modified form,\nless fragile one).\n\n> +\n> +test_expect_success 'compute bloom key for empty string' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:5615800c|5b966560|61174ab4|66983008|6c19155c|7199fab0|771ae004|\n> +\tFilter_Length:1\n> +\tFilter_Data:11000110001110|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \"\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n\nThis might be unnecessarily fragile test, but it might be a good test\nfor double hashing or enhanced double hashing technique.  Murmur3 hash\non empty data (empty string) always return seed value, so the result of\n(enhanced) double hashing technique is predictable, given two seed\nvalues.\n\n> +\n> +test_expect_success 'compute bloom key for whitespace' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:1bf014e6|8a91b50b|f9335530|67d4f555|d676957a|4518359f|b3b9d5c4|\n> +\tFilter_Length:1\n> +\tFilter_Data:401004080200810|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \" \" >actual &&\n> +\ttest_cmp expect actual\n> +'\n\nInstead of those two fragile tests (that depend on irrelevant details of\nthe implementation), it would be better to create test similar to those\nin t/t0011-hashmap.sh, for example:\n\n   test_expect_success 'testing Bloom filter querying' '\n   \ttest_bloom \"add abc\n        add abcdef\n        check abc\n        check abcdef\n        check abcdee\n        check abcdefghi\n        len\" \"maybe\n        maybe\n        no\n        no\n        1\"\n   '\n\nOr maybe something like this:\n\n   test_expect_success 'testing Bloom filter querying' '\n   \tcat >commands <<\\-EOF &&\n        add abc\n        add abcdef\n        check abc\n        check abcdef\n        check abcdee\n        check abcdefghi\n        len\n        EOF\n\n   \tcat >expect <<\\-EOF &&\n        maybe\n        maybe\n        no\n        no\n        1\n        EOF\n        \n        test-tool bloom <commands >actual &&\n        test_cmp expect actual\n   '\n\n> +\n> +test_expect_success 'compute bloom key for a root level folder' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:1a21016f|fff1c06d|e5c27f6b|cb933e69|b163fd67|9734bc65|7d057b63|\n> +\tFilter_Length:1\n> +\tFilter_Data:aaa800000000|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \"A\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'compute bloom key for a root level file' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:e2d51107|30970605|7e58fb03|cc1af001|19dce4ff|679ed9fd|b560cefb|\n> +\tFilter_Length:1\n> +\tFilter_Data:a8000000000000aa|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \"file.txt\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'compute bloom key for a deep folder' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:864cf838|27f055cd|c993b362|6b3710f7|0cda6e8c|ae7dcc21|502129b6|\n> +\tFilter_Length:1\n> +\tFilter_Data:1c0000600003000|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \"A/B/C/D/E\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'compute bloom key for a deep file' '\n> +\tcat >expect <<-\\EOF &&\n> +\tHashes:07cdf850|4af629c7|8e1e5b3e|d1468cb5|146ebe2c|5796efa3|9abf211a|\n> +\tFilter_Length:1\n> +\tFilter_Data:4020100804010080|\n> +\tEOF\n> +\ttest-tool bloom generate_filter \"A/B/C/D/E/file.txt\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n\nWhat are those meant to test?  For the Bloom filter itself it doesn't\nmatter if we add \"A/B/C/file.txt\" string to filter, or \"ABC\" string.\n\nWhat we didn't test is that changed _directories_ are also added to the\nBloom filter for a commit.  Such test could look like this:\n\n   test_expect_success 'changed directories are added to Bloom filter' '\n   \tmkdir -p A/B &&\n\techo \"foo\" >A/B/file.txt &&\n\tgit add A/B/file.txt &&\n\tgit commit -m \"add A/B/file.txt\" &&\n        HEAD=$(git rev-parse HEAD) &&\n\n   \tcat >commands <<-EOF &&\n        add-commit $HEAD\n        check A/B/file.txt\n        check A/B\n        check A\n        EOF\n\n   \tcat >expect <<\\-EOF &&\n        maybe\n        maybe\n\tmaybe\n        EOF\n        \n        test-tool bloom <commands >actual &&\n        test_cmp expect actual\n   '\n\n\n> +\n> +test_done\n\nReviewed-by: Jakub Narębski <jnareb@gmail.com>\n\nThanks for working on this.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"391883","messageId":"86h7zqqdze.fsf@gmail.com","threadId":"52499","inReplyTo":"a698c04a78cf2988fb822e0aa532989f925e0a9e.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 03/11] diff: halt tree-diff early after max_changes","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-17T00:00:53Z","receivedAt":"2020-02-17T00:01:09Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Derrick Stolee via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Derrick Stolee <dstolee@microsoft.com>\n>\n> When computing the changed-paths bloom filters for the commit-graph,\n> we limit the size of the filter by restricting the number of paths\n> in the diff. Instead of computing a large diff and then ignoring the\n> result, it is better to halt the diff computation early.\n\nGood idea.\n\n>\n> Create a new \"max_changes\" option in struct diff_options. If non-zero,\n> then halt the diff computation after discovering strictly more changed\n> paths. This includes paths corresponding to trees that change.\n\nAll right; also, it doesn't need to be exact, though it would be good if\nit was.\n\n512 changed paths (changed files) usually generate more than 512\nelements to be added to the Bloom filter (changed directories and\nfiles), anyway.\n\n>\n> Use this max_changes option in the bloom filter calculations. This\n> reduces the time taken to compute the filters for the Linux kernel\n> repo from 2m50s to 2m35s. On a large internal repository with ~500\n> commits that perform tree-wide changes, the time reduced from\n> 6m15s to 3m48s.\n\nI wonder if there is some large open-source project with many commits\nperforming tree-wide changes, that is with many commits with more than\n512 changed files with respect to the first parent.\n\nMaybe https://github.com/whosonfirst-data/whosonfirst-data-venue-us-ny\nfrom \"Top Ten Worst Repositories to host on GitHub - Git Merge 2017\"\ncould be a good repository to test ;-)\n\n>\n> Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n\nLooks good to me, but that is from cursory examination.\nDon't know the area to say anything more.\n\n> ---\n>  bloom.c     | 4 +++-\n>  diff.h      | 5 +++++\n>  tree-diff.c | 6 ++++++\n>  3 files changed, 14 insertions(+), 1 deletion(-)\n>\n> diff --git a/bloom.c b/bloom.c\n> index 6082193a75..818382c03b 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -134,6 +134,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n>  \tint i;\n>  \tstruct diff_options diffopt;\n> +\tint max_changes = 512;\n>  \n>  \tif (!bloom_filters.slab_size)\n>  \t\treturn NULL;\n> @@ -142,6 +143,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \n>  \trepo_diff_setup(r, &diffopt);\n>  \tdiffopt.flags.recursive = 1;\n> +\tdiffopt.max_changes = max_changes;\n>  \tdiff_setup_done(&diffopt);\n>  \n>  \tif (c->parents)\n> @@ -150,7 +152,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n>  \tdiffcore_std(&diffopt);\n>  \n> -\tif (diff_queued_diff.nr <= 512) {\n> +\tif (diff_queued_diff.nr <= max_changes) {\n>  \t\tstruct hashmap pathmap;\n>  \t\tstruct pathmap_hash_entry* e;\n>  \t\tstruct hashmap_iter iter;\n> diff --git a/diff.h b/diff.h\n> index 6febe7e365..9443dc1b00 100644\n> --- a/diff.h\n> +++ b/diff.h\n> @@ -285,6 +285,11 @@ struct diff_options {\n>  \t/* Number of hexdigits to abbreviate raw format output to. */\n>  \tint abbrev;\n>  \n> +\t/* If non-zero, then stop computing after this many changes. */\n> +\tint max_changes;\n> +\t/* For internal use only. */\n> +\tint num_changes;\n> +\n>  \tint ita_invisible_in_index;\n>  /* white-space error highlighting */\n>  #define WSEH_NEW (1<<12)\n> diff --git a/tree-diff.c b/tree-diff.c\n> index 33ded7f8b3..f3d303c6e5 100644\n> --- a/tree-diff.c\n> +++ b/tree-diff.c\n> @@ -434,6 +434,9 @@ static struct combine_diff_path *ll_diff_tree_paths(\n>  \t\tif (diff_can_quit_early(opt))\n>  \t\t\tbreak;\n>  \n> +\t\tif (opt->max_changes && opt->num_changes > opt->max_changes)\n> +\t\t\tbreak;\n> +\n>  \t\tif (opt->pathspec.nr) {\n>  \t\t\tskip_uninteresting(&t, base, opt);\n>  \t\t\tfor (i = 0; i < nparent; i++)\n> @@ -518,6 +521,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n>  \n>  \t\t\t/* t↓ */\n>  \t\t\tupdate_tree_entry(&t);\n> +\t\t\topt->num_changes++;\n>  \t\t}\n>  \n>  \t\t/* t > p[imin] */\n> @@ -535,6 +539,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n>  \t\tskip_emit_tp:\n>  \t\t\t/* ∀ pi=p[imin]  pi↓ */\n>  \t\t\tupdate_tp_entries(tp, nparent);\n> +\t\t\topt->num_changes++;\n>  \t\t}\n>  \t}\n>  \n> @@ -552,6 +557,7 @@ struct combine_diff_path *diff_tree_paths(\n>  \tconst struct object_id **parents_oid, int nparent,\n>  \tstruct strbuf *base, struct diff_options *opt)\n>  {\n> +\topt->num_changes = 0;\n>  \tp = ll_diff_tree_paths(p, oid, parents_oid, nparent, base, opt);\n>  \n>  \t/*\n"},{"id":"391955","messageId":"86k14klvyb.fsf@gmail.com","threadId":"52499","inReplyTo":"c17bbcbc66ea77bb480391804d1f2db66ffa0926.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 04/11] commit-graph: compute Bloom filters for changed paths","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-17T21:56:12Z","receivedAt":"2020-02-17T21:56:24Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n> Subject: [PATCH v2 04/11] commit-graph: compute Bloom filters for changed paths\n>\n> Compute Bloom filters for the paths that changed between a commit and its\n> first parent using the implementation in bloom.c, when the\n> COMMIT_GRAPH_WRITE_CHANGED_PATHS flag is set. This computation is done on a\n> commit-by-commit basis. We will write these Bloom filters to the commit graph\n> file in the next change.\n\nI have no major complaints about the contents of this patch (except lack\nof test, and type of total_bloom_filter_data_size), but the commit\nmessage could have been worded better.\n\nI would write something like this instead:\n\n  Add new COMMIT_GRAPH_WRITE_CHANGED_PATHS flag that makes Git compute\n  Bloom filters that store the information about changed paths (that\n  changed between a commit and its first parent) for each commit in the\n  commit-graph.  This computation is done on a commit-by-commit basis.\n\n  We will write these Bloom filters to the commit-graph file, to store\n  this data on disk, in the next change in this series.\n\nIn my opinion the fact that we compute Bloom filters for each and every\ncommit in the commit-graph file is more important than quite obvious\nfact that we use implementation from bloom.c.\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  commit-graph.c | 32 +++++++++++++++++++++++++++++++-\n>  commit-graph.h |  3 ++-\n>  2 files changed, 33 insertions(+), 2 deletions(-)\n\nIt would be good to have at least sanity check of this feature, perhaps\none that would check that the number of per-commit Bloom filters on slab\nmatches the number of commits in the commit-graph.\n\nIt could look something like this:\n\n  test_expect_success 'create Bloom filters for all commit-graph commits' '\n  \t# create commit-graph with 2 commits\n  \tgit rev-parse HEAD HEAD^ | git commit-graph write --stdin-commits &&\n  \t# generate Bloom filters for commit-graph commits\n  \tcat >commands <<\\-EOF &&\n  \tadd-graph-commits\n  \tfilters-count\n  \tEOF\n  \tNUM_FILTERS=$(git test-tool bloom <commands) %%\n  \ttest \"$NUM_FILTERS\" -eq 2\n  '\n\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 3c4d411326..724bfcffc4 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -16,6 +16,7 @@\n>  #include \"hashmap.h\"\n>  #include \"replace-object.h\"\n>  #include \"progress.h\"\n> +#include \"bloom.h\"\n>  \n>  #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n>  #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n> @@ -795,9 +796,11 @@ struct write_commit_graph_context {\n>  \tunsigned append:1,\n>  \t\t report_progress:1,\n>  \t\t split:1,\n> -\t\t check_oids:1;\n> +\t\t check_oids:1,\n> +\t\t changed_paths:1;\n\nAll right, this flag will be used for handling future `--changed-paths`\noption to the `git commit-graph write`.\n\n>  \n>  \tconst struct split_commit_graph_opts *split_opts;\n> +\tuint32_t total_bloom_filter_data_size;\n\nThis is total size of Bloom filters data, in bytes, that will later be\nused for BDAT chunk size.  However the commit-graph format uses 8 bytes\nfor byte-offset, not 4 bytes.  Why it is uint32_t and not uint64_t then?\n\n>  };\n>  \n>  static void write_graph_chunk_fanout(struct hashfile *f,\n> @@ -1140,6 +1143,28 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>  \tstop_progress(&ctx->progress);\n>  }\n>  \n> +static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n> +{\n> +\tint i;\n> +\tstruct progress *progress = NULL;\n> +\n> +\tload_bloom_filters();\n> +\n> +\tif (ctx->report_progress)\n> +\t\tprogress = start_progress(\n> +\t\t\t_(\"Computing commit diff Bloom filters\"),\n> +\t\t\tctx->commits.nr);\n> +\n\nShouldn't we initialize ctx->total_bloom_filter_data_size to 0 here?  We\ncannot use compute_bloom_filters() to _update_ Bloom filters data, I\nthink -- we don't distinguish here between new and existing data (where\nexisting data size is already included in total Bloom filters size).  At\nleast I don't think so.\n\n\n> +\tfor (i = 0; i < ctx->commits.nr; i++) {\n> +\t\tstruct commit *c = ctx->commits.list[i];\n\nHere we process commit in whatever order commits are in the\ncommits.list, which probably means lexicographical order, in practice\nrandom order.\n\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n> +\t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n> +\t\tdisplay_progress(progress, i + 1);\n> +\t}\n> +\n> +\tstop_progress(&progress);\n> +}\n> +\n>  static int add_ref_to_list(const char *refname,\n>  \t\t\t   const struct object_id *oid,\n>  \t\t\t   int flags, void *cb_data)\n\n> @@ -1794,6 +1819,8 @@ int write_commit_graph(const char *obj_dir,\n>  \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n>  \tctx->check_oids = flags & COMMIT_GRAPH_WRITE_CHECK_OIDS ? 1 : 0;\n>  \tctx->split_opts = split_opts;\n> +\tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n> +\tctx->total_bloom_filter_data_size = 0;\n>  \n>  \tif (ctx->split) {\n>  \t\tstruct commit_graph *g;\n> @@ -1888,6 +1915,9 @@ int write_commit_graph(const char *obj_dir,\n>  \n>  \tcompute_generation_numbers(ctx);\n>  \n> +\tif (ctx->changed_paths)\n> +\t\tcompute_bloom_filters(ctx);\n> +\n\nAll right.\n\n>  \tres = write_commit_graph_file(ctx);\n>  \n>  \tif (ctx->split)\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 7f5c933fa2..952a4b83be 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -76,7 +76,8 @@ enum commit_graph_write_flags {\n>  \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\n>  \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n>  \t/* Make sure that each OID in the input is a valid commit OID. */\n> -\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3)\n> +\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n> +\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n\nAll right.\n\n\nSide note: perhaps we could add trailing comma after new enum entry,\nthat is\n\n  +\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4),\n\nfollowing new CodingGuidelines recommendation\n\n - We try to support a wide range of C compilers to compile Git with,\n   including old ones.  You should not use features from newer C\n   standard, even if your compiler groks them.\n\n   There are a few exceptions to this guideline:\n\n   . since early 2012 with e1327023ea, we have been using an enum\n     definition whose last element is followed by a comma.  This, like\n     an array initializer that ends with a trailing comma, can be used\n     to reduce the patch noise when adding a new identifier at the end.\n\nhttps://github.com/git/git/blob/master/Documentation/CodingGuidelines#L197\n\n>  };\n>  \n>  struct split_commit_graph_opts {\n"},{"id":"391998","messageId":"86k14jkc8s.fsf@gmail.com","threadId":"52499","inReplyTo":"78e8e49c3a1131ffacf660603de60729b3dbadc9.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 05/11] commit-graph: examine changed-path objects in pack order","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-18T17:59:31Z","receivedAt":"2020-02-18T17:59:46Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Jeff King via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Jeff King <peff@peff.net>\n>\n> Looking at the diff of commit objects in pack order is much faster than\n> in sha1 order, as it gives locality to the access of tree deltas\n\nNitpick: should we still say sha1 order?  Git is still using SHA-1 as an\n*oid*, but hopefully soon it will be transitioning to NewHash = SHA-256.\n(No need to change anything.)\n\n> (whereas sha1 order is effectively random). Unfortunately the\n> commit-graph code sorts the commits (several times, sometimes as an oid\n> and sometimes a pointer-to-commit), and we ultimately traverse in sha1\n> order.\n\nActually, commit-graph code needs write_commit_graph_context.commits.list\nto be in lexicographical order to be able to turn position in graph into\nreference to a commit.  The information about the parents of the commit\nare stored using positional references within the graph file.\n\n>\n> Instead, let's remember the position at which we see each commit, and\n> traverse in that order when looking at bloom filters. This drops my time\n> for \"git commit-graph write --changed-paths\" in linux.git from ~4\n> minutes to ~1.5 minutes.\n\nNitpick: with reordering of patches (which I think is otherwise a good\nthing) this patch actually comes before the one adding \"--changed-paths\"\noption to \"git commit-graph write\".  So it 'This would drop my time'\nrather than 'This drops my time...' ;-)\n\n>\n> Probably the \"--reachable\" code path would want something similar.\n\nHas anyone tried doing this?\n\n>\n> Or alternatively, we could use a different data structure (either a\n> hash, or maybe even just a bit in \"struct commit\") to keep track of\n> which oids we've seen, etc instead of sorting. And then we could keep\n> the original order.\n\nI think it is nice to keep those \"what ifs?\" thoughts in the commit\nmessage.  They add some color.\n\n>\n> Signed-off-by: Jeff King <peff@peff.net>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  commit-graph.c | 34 +++++++++++++++++++++++++++++++++-\n>  1 file changed, 33 insertions(+), 1 deletion(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 724bfcffc4..e125511a1c 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -17,6 +17,7 @@\n>  #include \"replace-object.h\"\n>  #include \"progress.h\"\n>  #include \"bloom.h\"\n> +#include \"commit-slab.h\"\n>  \n>  #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n>  #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n> @@ -46,6 +47,29 @@\n>  /* Remember to update object flag allocation in object.h */\n>  #define REACHABLE       (1u<<15)\n>  \n> +/* Keep track of the order in which commits are added to our list. */\n> +define_commit_slab(commit_pos, int);\n> +static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n> +\n> +static void set_commit_pos(struct repository *r, const struct object_id *oid)\n> +{\n> +\tstatic int32_t max_pos;\n> +\tstruct commit *commit = lookup_commit(r, oid);\n> +\n> +\tif (!commit)\n> +\t\treturn; /* should never happen, but be lenient */\n> +\n> +\t*commit_pos_at(&commit_pos, commit) = max_pos++;\n> +}\n\nAll right, that is nice and universal function.\n\n> +\n> +static int commit_pos_cmp(const void *va, const void *vb)\n> +{\n> +\tconst struct commit *a = *(const struct commit **)va;\n> +\tconst struct commit *b = *(const struct commit **)vb;\n> +\treturn commit_pos_at(&commit_pos, a) -\n> +\t       commit_pos_at(&commit_pos, b);\n> +}\n\nHmmm... I wonder what would happen in commit_pos was not set (like\ne.g. commit-graph commits not coming from the packfile).  Let's look up\nthe documenation...\n\ncommit_pos_at() returns a pointer to an int... why are we comparing\npointers and not values?  Shouldn't it be\n\n  +\treturn *commit_pos_at(&commit_pos, a) -\n  +\t       *commit_pos_at(&commit_pos, b);\n\n\nWith commit_pos_at() the location to store the data is allocated as\nnecessary (if data for commit doesn't exists), and because we are using\nxalloc() the *commit_pos_at() is 0-initialized.  This means that if\ncommits didn't come from the packfile, we sort all commits as being\nequal.  Luckily we fix that in next patch.\n\n> +\n>  char *get_commit_graph_filename(const char *obj_dir)\n>  {\n>  \tchar *filename = xstrfmt(\"%s/info/commit-graph\", obj_dir);\n> @@ -1027,6 +1051,8 @@ static int add_packed_commits(const struct object_id *oid,\n>  \toidcpy(&(ctx->oids.list[ctx->oids.nr]), oid);\n>  \tctx->oids.nr++;\n>  \n> +\tset_commit_pos(ctx->r, oid);\n> +\n>  \treturn 0;\n>  }\n>  \n> @@ -1147,6 +1173,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>  {\n>  \tint i;\n>  \tstruct progress *progress = NULL;\n> +\tstruct commit **sorted_by_pos;\n\nIn the next patch in series we would sort commits by generation number\nand creation data; shouldn't this variable name be more generic to\nreflect this, for example just `sorted_commits` or `commits_sorted`?\n\n>  \n>  \tload_bloom_filters();\n>  \n> @@ -1155,13 +1182,18 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>  \t\t\t_(\"Computing commit diff Bloom filters\"),\n>  \t\t\tctx->commits.nr);\n>  \n> +\tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n> +\tCOPY_ARRAY(sorted_by_pos, ctx->commits.list, ctx->commits.nr);\n> +\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n> +\n\nAll right: allocate array, copy data, sort it.\n\nWe need to copy data because (what I think) we need commits in\nlexicographical order to be able to turn the position in graph that\nparents of a commit are stored as into the reference to this commit.\n\n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n> -\t\tstruct commit *c = ctx->commits.list[i];\n> +\t\tstruct commit *c = sorted_by_pos[i];\n\nAll right: use sorted data.\n\n>  \t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n>  \t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n>  \t\tdisplay_progress(progress, i + 1);\n>  \t}\n>  \n> +\tfree(sorted_by_pos);\n\nCan we free the slab data, i.e. call `clear_commit_pos(&commit_pos);`\nhere?  Otherwise we are leaking memory (well, except that finishing\ncommand makes the operating system to free memory for us).\n\n>  \tstop_progress(&progress);\n>  }\n\nBest,\n-- \nJakub Narębski\n"},{"id":"392037","messageId":"865zg3ju2j.fsf@gmail.com","threadId":"52499","inReplyTo":"58704d81b6b4fbc54715457246aeed783eb32a99.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 06/11] commit-graph: examine commits by generation number","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-19T00:32:04Z","receivedAt":"2020-02-19T00:33:13Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Derrick Stolee via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Derrick Stolee <dstolee@microsoft.com>\n>\n> When running 'git commit-graph write --changed-paths', we sort the\n> commits by pack-order to save time when computing the changed-paths\n> bloom filters. This does not help when finding the commits via the\n> --reachable flag.\n\nMinor improvement suggestion: s/--reachable flag/'--reachable' flag/.\n\n>\n> If not using pack-order, then sort by generation number before\n> examining the diff.\n\nAll right, that is good description of what the patch does.\n\n>                     Commits with similar generation are more likely\n> to have many trees in common, making the diff faster.\n\nIs this what causes the performance improvement, that subsequently\nexamined commits are more likely to have more trees in common, which\nmeans that those trees would be hot in cache, making generating diff\nfaster?  Is it what profiling shows?\n\n>\n> On the Linux kernel repository, this change reduced the computation\n> time for 'git commit-graph write --reachable --changed-paths' from\n> 3m00s to 1m37s.\n\nWould using the trick used for packfiles also for '--reachable', which\nwould mean commits examined in recency / reachability order, give\nsimilar, worse or better performance improvements?\n\nWe would want this sorting order as one of possibilities anyway, because\n'--stdin-commits' we could get commits in random order.\n\n>\n> Helped-by: Jeff King <peff@peff.net>\n> Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  commit-graph.c | 33 ++++++++++++++++++++++++++++++---\n>  1 file changed, 30 insertions(+), 3 deletions(-)\n>\n> diff --git a/commit-graph.c b/commit-graph.c\n> index e125511a1c..32a315058f 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -70,6 +70,25 @@ static int commit_pos_cmp(const void *va, const void *vb)\n>  \t       commit_pos_at(&commit_pos, b);\n>  }\n>  \n> +static int commit_gen_cmp(const void *va, const void *vb)\n> +{\n> +\tconst struct commit *a = *(const struct commit **)va;\n> +\tconst struct commit *b = *(const struct commit **)vb;\n> +\n> +\t/* lower generation commits first */\n\nShouldn't higher generation commits come first, in recency-like order?\nOr it doesn't matter if it is sorted in ascending or descending order,\nas long as commits with close generation numbers are examined close\ntogether?\n\n> +\tif (a->generation < b->generation)\n> +\t\treturn -1;\n> +\telse if (a->generation > b->generation)\n> +\t\treturn 1;\n> +\n> +\t/* use date as a heuristic when generations are equal */\n> +\tif (a->date < b->date)\n> +\t\treturn -1;\n> +\telse if (a->date > b->date)\n> +\t\treturn 1;\n> +\treturn 0;\n> +}\n\nI thought we have had such comparison function defined somewhere in Git\nalready, but I think I'm wrong here.\n\n> +\n>  char *get_commit_graph_filename(const char *obj_dir)\n>  {\n>  \tchar *filename = xstrfmt(\"%s/info/commit-graph\", obj_dir);\n> @@ -821,7 +840,8 @@ struct write_commit_graph_context {\n>  \t\t report_progress:1,\n>  \t\t split:1,\n>  \t\t check_oids:1,\n> -\t\t changed_paths:1;\n> +\t\t changed_paths:1,\n> +\t\t order_by_pack:1;\n>  \n>  \tconst struct split_commit_graph_opts *split_opts;\n>  \tuint32_t total_bloom_filter_data_size;\n> @@ -1184,7 +1204,11 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>  \n>  \tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n>  \tCOPY_ARRAY(sorted_by_pos, ctx->commits.list, ctx->commits.nr);\n> -\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n> +\n> +\tif (ctx->order_by_pack)\n> +\t\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n> +\telse\n> +\t\tQSORT(sorted_by_pos, ctx->commits.nr, commit_gen_cmp);\n\nHere 'sorted_b_pos' variable name no longer reflects reality...\n(see comment to the previous patch in the series).\n\n>  \n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>  \t\tstruct commit *c = sorted_by_pos[i];\n> @@ -1902,6 +1926,7 @@ int write_commit_graph(const char *obj_dir,\n>  \t}\n>  \n>  \tif (pack_indexes) {\n> +\t\tctx->order_by_pack = 1;\n>  \t\tif ((res = fill_oids_from_packs(ctx, pack_indexes)))\n>  \t\t\tgoto cleanup;\n>  \t}\n> @@ -1911,8 +1936,10 @@ int write_commit_graph(const char *obj_dir,\n>  \t\t\tgoto cleanup;\n>  \t}\n>  \n> -\tif (!pack_indexes && !commit_hex)\n> +\tif (!pack_indexes && !commit_hex) {\n> +\t\tctx->order_by_pack = 1;\n>  \t\tfill_oids_from_all_packs(ctx);\n> +\t}\n>  \n>  \tclose_reachable(ctx);\n\nAll right, that covers all cases where 'git commit-graph write' writes\nserialized commit-graph based on the commits found in packfiles:\n'--stdin-packs' and default no option case, in that order.\n\nLooks good.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"392061","messageId":"86pneahaop.fsf@gmail.com","threadId":"52499","inReplyTo":"39ee0610800d7d2d92785d392df941fc5a0b231b.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 07/11] commit-graph: write Bloom filters to commit graph file","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-19T15:13:42Z","receivedAt":"2020-02-19T15:13:56Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Update the technical documentation for commit-graph-format with the formats for\n> the Bloom filter index (BIDX) and Bloom filter data (BDAT) chunks. Write the\n> computed Bloom filters information to the commit graph file using this format.\n\nNice description.\n\nThe only minor nitpick is with the formating: it is 80-character wide,\nwhich is a bit wide.\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  .../technical/commit-graph-format.txt         |  24 ++++\n>  commit-graph.c                                | 118 +++++++++++++++++-\n>  commit-graph.h                                |   7 +-\n>  3 files changed, 145 insertions(+), 4 deletions(-)\n>\n> diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n> index a4f17441ae..22e511643d 100644\n> --- a/Documentation/technical/commit-graph-format.txt\n> +++ b/Documentation/technical/commit-graph-format.txt\n> @@ -17,6 +17,9 @@ metadata, including:\n>  - The parents of the commit, stored using positional references within\n>    the graph file.\n>  \n> +- The Bloom filter of the commit carrying the paths that were changed between\n> +  the commit and its first parent.\n> +\n\nAll right.\n\nShould we also state that it is optional (meta)data?  This would be\nfirst optional piece of data stored in commit-graph, I think.\n\n>  These positional references are stored as unsigned 32-bit integers\n>  corresponding to the array position within the list of commit OIDs. Due\n>  to some special constants we use to track parents, we can store at most\n> @@ -93,6 +96,27 @@ CHUNK DATA:\n>        positions for the parents until reaching a value with the most-significant\n>        bit on. The other bits correspond to the position of the last parent.\n>  \n> +  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) (N * 4 bytes) [Optional]\n> +    * The ith entry, BIDX[i], stores the number of 8-byte word blocks in all\n> +      Bloom filters from commit 0 to commit i (inclusive) in lexicographic\n> +      order. The Bloom filter for the i-th commit spans from BIDX[i-1] to\n> +      BIDX[i] (plus header length), where BIDX[-1] is 0.\n> +    * The BIDX chunk is ignored if the BDAT chunk is not present.\n\nAll right.  Looks good.\n\n> +\n> +  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n> +    * It starts with header consisting of three unsigned 32-bit integers:\n> +      - Version of the hash algorithm being used. We currently only support\n> +\tvalue 1 which implies the murmur3 hash implemented exactly as described\n> +\tin https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n\nFirst a minor issue: shouldn't this nested unordered list be indented\nwith a hanging indent formatted with spaces?  That is be formatted like\nthe following:\n\n  +  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n  +    * It starts with header consisting of three unsigned 32-bit integers:\n  +      - Version of the hash algorithm being used. We currently only support\n  +        value 1 which implies the murmur3 hash implemented exactly as\n  +        described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n\nBut the existing formatting with spaces and tabs might be fine as it is,\nthat is it renders as nested list with Asciidoc; it only looks a bit\nweird as patch, not so as text.\n\nSecond, and more important: it is in my opinion not enough information,\nat least if we are assuming that the information in this document should\nbe enough for clean-room reimplementation of Bloom filter functionality\n(for example by JGit).  To generate compatible Bloom filters, one needs\nalso the information on how to create $k$ functionally-independent hash\nfunctions out of murmur3 hash.  We do it currently using double hashing\ntechnique; if that changes then the exact set of bits in the Bloom\nfilter would also change.\n\nThe additional description could look something like the following:\n\n  +    * It starts with header consisting of three unsigned 32-bit integers:\n  +      - Version of the hash algorithm being used. We currently only support\n  +        value 1 which implies the murmur3_32 hash implemented exactly as\n  +        described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n  +        and double hashing technique with 0x293ae76f and 0x7e646e2c seeds\n  +        as described in https://doi.org/10.1007/978-3-540-30494-4_26\n  +        \"Bloom Filters in Probabilistic Verification\"\n\nAlso, it should be explicitly noted that we use murmur3_32, because\nthere is also 128-bit version of murmur3 hash.\n\n> +      - The number of times a path is hashed and hence the number of bit positions\n> +\tthat cumulatively determine whether a file is present in the commit.\n\nAll right, in the original Bloom filter it was the number of different\nhash functions.  With the double hashing technique, it is the number of\ntimes a path is hashed.\n\n> +      - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n> +\tcontains 'n' entries, then the filter size is the minimum number of 64-bit\n> +\twords that contain n*b bits.\n\nAll right, that means empty Bloom filter, representing \"no changes\",\nwith 'n' equal 0 entries, is represented as size 0 filter.  That is, if\nwe read this rule exactly as written.\n\nShould we add the information that size 0 / length 0 filter is\nconsidered \"no data\" case?  Or should we leave it to implementation?\n\nThere are two corner cases:\n- \"no changes\" case, where all queries are answered with \"no\"\n  can be represented as filter of size 0, or as Bloom filter with all\n  bits set to 0\n- \"no data\" case (used when there are more than 512 changed files)\n  where all queries are answered with \"maybe\", currently represented\n  as filter of size 0; can also be represented as Bloom filter with all\n  bits set to 1\n\n> +    * The rest of the chunk is the concatenation of all the computed Bloom\n> +      filters for the commits in lexicographic order.\n\nAll right.\n\n> +    * The BDAT chunk is present iff BIDX is present.\n\nPerhaps we should spell 'iff' in full, that is 'if and only if'?\n\n> +\n>    Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n>        This list of H-byte hashes describe a set of B commit-graph files that\n>        form a commit-graph chain. The graph position for the ith commit in this\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 32a315058f..4585b3b702 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -24,8 +24,10 @@\n>  #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n>  #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n>  #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n> +#define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n> +#define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n>  #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n> -#define MAX_NUM_CHUNKS 5\n> +#define MAX_NUM_CHUNKS 7\n>  \n>  #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n>  \n> @@ -325,6 +327,32 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n>  \t\t\t\tchunk_repeated = 1;\n>  \t\t\telse\n>  \t\t\t\tgraph->chunk_base_graphs = data + chunk_offset;\n> +\t\t\tbreak;\n> +\n> +\t\tcase GRAPH_CHUNKID_BLOOMINDEXES:\n> +\t\t\tif (graph->chunk_bloom_indexes)\n> +\t\t\t\tchunk_repeated = 1;\n> +\t\t\telse\n> +\t\t\t\tgraph->chunk_bloom_indexes = data + chunk_offset;\n> +\t\t\tbreak;\n> +\n> +\t\tcase GRAPH_CHUNKID_BLOOMDATA:\n> +\t\t\tif (graph->chunk_bloom_data)\n> +\t\t\t\tchunk_repeated = 1;\n> +\t\t\telse {\n> +\t\t\t\tuint32_t hash_version;\n> +\t\t\t\tgraph->chunk_bloom_data = data + chunk_offset;\n> +\t\t\t\thash_version = get_be32(data + chunk_offset);\n> +\n> +\t\t\t\tif (hash_version != 1)\n> +\t\t\t\t\tbreak;\n\nShouldn't we mark Bloom filter as not to be used?  Or is it left for\nlater commit?\n\nIn the future it might be good idea to notify the user (perhaps\nprotected with some advice.* option) that there is problem with Bloom\nfilter data, namely that we have encountered unsupported hash version.\n\n> +\n> +\t\t\t\tgraph->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n\nWhy is this structure allocated dynamically?  We are leaking admittedly\na small amount of memory because we never free this xmalloc() result.\n\nIf we need this field being a pointer to struct to have NULL mean no\nsupported Bloom filter data, we could have instead use chunk_bloom_*\nfields instead - we can set at least one of them to NULL.\n\n> +\t\t\t\tgraph->bloom_filter_settings->hash_version = hash_version;\n> +\t\t\t\tgraph->bloom_filter_settings->num_hashes = get_be32(data + chunk_offset + 4);\n> +\t\t\t\tgraph->bloom_filter_settings->bits_per_entry = get_be32(data + chunk_offset + 8);\n\nAll right; these 4 and 8 are sizeof(uint32_t) and 2*sizeof(uint32_t),\nrespectively.\n\n> +\t\t\t}\n> +\t\t\tbreak;\n>  \t\t}\n>  \n>  \t\tif (chunk_repeated) {\n> @@ -343,6 +371,17 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n>  \t\tlast_chunk_offset = chunk_offset;\n>  \t}\n>  \n> +\t/* We need both the bloom chunks to exist together. Else ignore the data */\n> +\tif ((graph->chunk_bloom_indexes && !graph->chunk_bloom_data)\n> +\t\t || (!graph->chunk_bloom_indexes && graph->chunk_bloom_data)) {\n> +\t\tgraph->chunk_bloom_indexes = NULL;\n> +\t\tgraph->chunk_bloom_data = NULL;\n> +\t\tgraph->bloom_filter_settings = NULL;\n> +\t}\n> +\n> +\tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data)\n> +\t\tload_bloom_filters();\n\nWouldn't it be simpler to rely on the fact that both Bloom chunks must\nexists for it to matter, and write it like this:\n\n  +\tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data) {\n  +\t\tload_bloom_filters();\n  +\t} else {\n  +\t\tgraph->chunk_bloom_indexes = NULL;\n  +\t\tgraph->chunk_bloom_data = NULL;\n  +\t\tgraph->bloom_filter_settings = NULL;\n  +\t}\n\n> +\n>  \thashcpy(graph->oid.hash, graph->data + graph->data_len - graph->hash_len);\n>  \n>  \tif (verify_commit_graph_lite(graph)) {\n> @@ -1040,6 +1079,59 @@ static void write_graph_chunk_extra_edges(struct hashfile *f,\n>  \t}\n>  }\n>  \n> +static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n> +\t\t\t\t\t    struct write_commit_graph_context *ctx)\n> +{\n> +\tstruct commit **list = ctx->commits.list;\n> +\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n> +\tuint32_t cur_pos = 0;\n> +\tstruct progress *progress = NULL;\n> +\tint i = 0;\n> +\n> +\tif (ctx->report_progress)\n> +\t\tprogress = start_delayed_progress(\n> +\t\t\t_(\"Writing changed paths Bloom filters index\"),\n> +\t\t\tctx->commits.nr);\n> +\n> +\twhile (list < last) {\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tcur_pos += filter->len;\n> +\t\tdisplay_progress(progress, ++i);\n> +\t\thashwrite_be32(f, cur_pos);\n> +\t\tlist++;\n> +\t}\n> +\n> +\tstop_progress(&progress);\n> +}\n\nAll right, looks good.\n\n> +\n> +static void write_graph_chunk_bloom_data(struct hashfile *f,\n> +\t\t\t\t\t struct write_commit_graph_context *ctx,\n> +\t\t\t\t\t struct bloom_filter_settings *settings)\n> +{\n> +\tstruct commit **list = ctx->commits.list;\n> +\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n> +\tstruct progress *progress = NULL;\n> +\tint i = 0;\n> +\n> +\tif (ctx->report_progress)\n> +\t\tprogress = start_delayed_progress(\n> +\t\t\t_(\"Writing changed paths Bloom filters data\"),\n> +\t\t\tctx->commits.nr);\n> +\n> +\thashwrite_be32(f, settings->hash_version);\n> +\thashwrite_be32(f, settings->num_hashes);\n> +\thashwrite_be32(f, settings->bits_per_entry);\n> +\n> +\twhile (list < last) {\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tdisplay_progress(progress, ++i);\n> +\t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n> +\t\tlist++;\n> +\t}\n> +\n> +\tstop_progress(&progress);\n> +}\n\nAll right, looks good.\n\nSide note: why have while loop here instead of for loop, like in\nprevious patches?  I'm not saying this is a bad idea (especially with\nsame names for same variables).\n\n> +\n>  static int oid_compare(const void *_a, const void *_b)\n>  {\n>  \tconst struct object_id *a = (const struct object_id *)_a;\n> @@ -1198,8 +1290,8 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>  \tload_bloom_filters();\n>  \n>  \tif (ctx->report_progress)\n> -\t\tprogress = start_progress(\n> -\t\t\t_(\"Computing commit diff Bloom filters\"),\n> +\t\tprogress = start_delayed_progress(\n> +\t\t\t_(\"Computing changed paths Bloom filters\"),\n>  \t\t\tctx->commits.nr);\n>\n\nOoops.  This look like a fixup which should be made to the original\nearlier commit instead, isn't it?\n\n>  \tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n> @@ -1444,6 +1536,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \tstruct strbuf progress_title = STRBUF_INIT;\n>  \tint num_chunks = 3;\n>  \tstruct object_id file_hash;\n> +\tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n>  \n>  \tif (ctx->split) {\n>  \t\tstruct strbuf tmp_file = STRBUF_INIT;\n> @@ -1488,6 +1581,12 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n>  \t\tnum_chunks++;\n>  \t}\n> +\tif (ctx->changed_paths) {\n> +\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMINDEXES;\n> +\t\tnum_chunks++;\n> +\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMDATA;\n> +\t\tnum_chunks++;\n> +\t}\n\nAll right, adding chunks and counting them.\n\n>  \tif (ctx->num_commit_graphs_after > 1) {\n>  \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BASE;\n>  \t\tnum_chunks++;\n> @@ -1506,6 +1605,15 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \t\t\t\t\t\t4 * ctx->num_extra_edges;\n>  \t\tnum_chunks++;\n>  \t}\n> +\tif (ctx->changed_paths) {\n> +\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n> +\t\t\t\t\t\tsizeof(uint32_t) * ctx->commits.nr;\n> +\t\tnum_chunks++;\n> +\n> +\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n> +\t\t\t\t\t\tsizeof(uint32_t) * 3 + ctx->total_bloom_filter_data_size;\n> +\t\tnum_chunks++;\n> +\t}\n\nAll right, calculating chunk offsets.\n\n>  \tif (ctx->num_commit_graphs_after > 1) {\n>  \t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n>  \t\t\t\t\t\thashsz * (ctx->num_commit_graphs_after - 1);\n> @@ -1543,6 +1651,10 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n>  \twrite_graph_chunk_data(f, hashsz, ctx);\n>  \tif (ctx->num_extra_edges)\n>  \t\twrite_graph_chunk_extra_edges(f, ctx);\n> +\tif (ctx->changed_paths) {\n> +\t\twrite_graph_chunk_bloom_indexes(f, ctx);\n> +\t\twrite_graph_chunk_bloom_data(f, ctx, &bloom_settings);\n> +\t}\n\nAll right, writing BIDX and BDAT chunks with default settings.\n\nBy the way, in the future, when appending to existing commit-graph file,\nshouldn't we re-use existing settings even if they are different from\ndefault settings?  But that is question for the future...\n\n>  \tif (ctx->num_commit_graphs_after > 1 &&\n>  \t    write_graph_chunk_base(f, ctx)) {\n>  \t\treturn -1;\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 952a4b83be..25fefefb3e 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -10,6 +10,7 @@\n>  #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n>  \n>  struct commit;\n> +struct bloom_filter_settings;\n>  \n>  char *get_commit_graph_filename(const char *obj_dir);\n>  int open_commit_graph(const char *graph_file, int *fd, struct stat *st);\n> @@ -58,6 +59,10 @@ struct commit_graph {\n>  \tconst unsigned char *chunk_commit_data;\n>  \tconst unsigned char *chunk_extra_edges;\n>  \tconst unsigned char *chunk_base_graphs;\n> +\tconst unsigned char *chunk_bloom_indexes;\n> +\tconst unsigned char *chunk_bloom_data;\n\nAll right.\n\n> +\n> +\tstruct bloom_filter_settings *bloom_filter_settings;\n\nWhy it is pointer to struct, instead of being just struct type?\nIs there reason for that?\n\n>  };\n>  \n>  struct commit_graph *load_commit_graph_one_fd_st(int fd, struct stat *st);\n> @@ -77,7 +82,7 @@ enum commit_graph_write_flags {\n>  \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n>  \t/* Make sure that each OID in the input is a valid commit OID. */\n>  \tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n> -\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n> +\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4),\n\nThis looks like accidental change; if we want to use trailing comma in\nenum, this change should be in my opinion done in the commit that added\nCOMMIT_GRAPH_WRITE_BLOOM_FILTERS (as I have written in a comment there).\n\n>  };\n>  \n>  struct split_commit_graph_opts {\n\nThank you for your work on this series.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"392182","messageId":"86r1ypf62y.fsf@gmail.com","threadId":"52499","inReplyTo":"b20c8d2b2096bf10fe1a5f37a5181c57873a9676.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 08/11] commit-graph: reuse existing Bloom filters during write.","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-20T18:48:21Z","receivedAt":"2020-02-20T18:48:35Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Read previously computed Bloom filters from the commit-graph file if\n> possible to avoid recomputing during commit-graph write.\n\nAll right, what is written makes sense for this point in patch series.\n\nBut it my opinion it is more important to state that this commit adds\n\"parsing\" of the Bloom filter data from commit-graph file.  This means\nthat it needs to be calculated only once, then stored in commit-graph,\nready to be re-used.\n\n>\n> See Documentation/technical/commit-graph-format for the format in which\n> the Bloom filter information is written to the commit graph file.\n>\n> To read Bloom filter for a given commit with lexicographic position\n> 'i' we need to:\n> 1. Read BIDX[i] which essentially gives us the starting index in BDAT for\n>    filter of commit i+1. It is essentially the index past the end\n>    of the filter of commit i. It is called end_index in the code.\n>\n> 2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT\n>    for filter of commit i. It is called the start_index in the code.\n>    For the first commit, where i = 0, Bloom filter data starts at the\n>    beginning, just past the header in the BDAT chunk. Hence, start_index\n>    will be 0.\n>\n> 3. The length of the filter will be end_index - start_index, because\n>    BIDX[i] gives the cumulative 8-byte words including the ith\n>    commit's filter.\n>\n> We toggle whether Bloom filters should be recomputed based on the\n> compute_if_null flag.\n\nNitpick: the flag (the parameter) is called compute_if_not_present, not\ncompute_if_null.\n\nAll right, this explanation is nice and clear.\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  bloom.c               | 49 ++++++++++++++++++++++++++++++++++++++++++-\n>  bloom.h               |  4 +++-\n>  commit-graph.c        |  7 ++++---\n>  t/helper/test-bloom.c |  2 +-\n>  4 files changed, 56 insertions(+), 6 deletions(-)\n>\n> diff --git a/bloom.c b/bloom.c\n> index 818382c03b..90d84dc713 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -1,5 +1,7 @@\n>  #include \"git-compat-util.h\"\n>  #include \"bloom.h\"\n> +#include \"commit.h\"\n> +#include \"commit-slab.h\"\n>  #include \"commit-graph.h\"\n>  #include \"object-store.h\"\n>  #include \"diff.h\"\n> @@ -127,8 +129,39 @@ void add_key_to_filter(struct bloom_key *key,\n>  \t}\n>  }\n>  \n> +static int load_bloom_filter_from_graph(struct commit_graph *g,\n> +\t\t\t\t   struct bloom_filter *filter,\n> +\t\t\t\t   struct commit *c)\n> +{\n> +\tuint32_t lex_pos, start_index, end_index;\n> +\n> +\twhile (c->graph_pos < g->num_commits_in_base)\n> +\t\tg = g->base_graph;\n> +\n> +\t/* The commit graph commit 'c' lives in doesn't carry bloom filters. */\n> +\tif (!g->chunk_bloom_indexes)\n> +\t\treturn 0;\n> +\n> +\tlex_pos = c->graph_pos - g->num_commits_in_base;\n\nAll right, this finds lexicographical position of the commit following\nthe chain of incremental commit-graph files, and also check if the\ncommit-graph fragment that contains the commit in question has Bloom\nfilter data included.\n\n> +\n> +\tend_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n> +\n> +\tif (lex_pos)\n\nWouldn't it be better to be more explicit, and write\n\n  +\tif (lex_pos > 0)\n\n\n> +\t\tstart_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n> +\telse\n> +\t\tstart_index = 0;\n\nAll right, here we find start_index and end_index.\n\nIt might be good idea to at least assert() that start_index <= end_index,\nthough that should not happen (that is why I propose for this check to\nbe compiled on only for debug builds).\n\n> +\n> +\tfilter->len = end_index - start_index;\n> +\tfilter->data = (uint64_t *)(g->chunk_bloom_data +\n> +\t\t\t\t\tsizeof(uint64_t) * start_index +\n> +\t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n\nAll right, nice use of constant.\n\n> +\n> +\treturn 1;\n> +}\n> +\n>  struct bloom_filter *get_bloom_filter(struct repository *r,\n> -\t\t\t\t      struct commit *c)\n> +\t\t\t\t      struct commit *c,\n> +\t\t\t\t      int compute_if_not_present)\n>  {\n>  \tstruct bloom_filter *filter;\n>  \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n> @@ -141,6 +174,20 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \n>  \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n>  \n> +\tif (!filter->data) {\n> +\t\tload_commit_graph_info(r, c);\n> +\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH &&\n> +\t\t\tr->objects->commit_graph->chunk_bloom_indexes) {\n\nAll right, the limitation that the top layer of incremental commit graph\nneeds to have Bloom filters enabled for it to be even considered is\nreasonable tradeoff, in my opinion.\n\n> +\t\t\tif (load_bloom_filter_from_graph(r->objects->commit_graph, filter, c))\n> +\t\t\t\treturn filter;\n> +\t\t\telse\n> +\t\t\t\treturn NULL;\n\nIf it should have filter, return it, otherwise return NULL.\n\nI wonder however when it can return NULL (and whether it should compute\nBloom filters if required instead).\n\n> +\t\t}\n> +\t}\n> +\n> +\tif (filter->data || !compute_if_not_present)\n> +\t\treturn filter;\n\nIf we have filter from slab, return it.  All right.\n\nHowever, according to documentation contained in comments in\ncommit-slab.h, bloom_filter_slab_at() will allocate the location to\nstore the data, and return freshly allocated memory... fortunately it\nuses xcalloc() so returned bloom_filter would have ->len == 0 and\n->data == 0.\n\n> +\n>  \trepo_diff_setup(r, &diffopt);\n>  \tdiffopt.flags.recursive = 1;\n>  \tdiffopt.max_changes = max_changes;\n> diff --git a/bloom.h b/bloom.h\n> index 7f40c751f7..76f8a9ad0c 100644\n> --- a/bloom.h\n> +++ b/bloom.h\n> @@ -13,6 +13,7 @@ struct bloom_filter_settings {\n>  \n>  #define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n>  #define BITS_PER_WORD 64\n> +#define BLOOMDATA_CHUNK_HEADER_SIZE 3*sizeof(uint32_t)\n\nAll right.\n\n>  \n>  /*\n>   * A bloom_filter struct represents a data segment to\n> @@ -47,7 +48,8 @@ void add_key_to_filter(struct bloom_key *key,\n>  \t\t\t\t\t   struct bloom_filter_settings *settings);\n>  \n>  struct bloom_filter *get_bloom_filter(struct repository *r,\n> -\t\t\t\t      struct commit *c);\n> +\t\t\t\t      struct commit *c,\n> +\t\t\t\t      int compute_if_not_present);\n>\n\nAll right, adding new parameter (changing function signature).\n\n>  int bloom_filter_contains(struct bloom_filter *filter,\n>  \t\t\t  struct bloom_key *key,\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 4585b3b702..c0e9834bf2 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -1094,7 +1094,7 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n>  \t\t\tctx->commits.nr);\n>  \n>  \twhile (list < last) {\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n>  \t\tcur_pos += filter->len;\n>  \t\tdisplay_progress(progress, ++i);\n>  \t\thashwrite_be32(f, cur_pos);\n> @@ -1123,7 +1123,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n>  \thashwrite_be32(f, settings->bits_per_entry);\n>  \n>  \twhile (list < last) {\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n>  \t\tdisplay_progress(progress, ++i);\n>  \t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n>  \t\tlist++;\n\nAll right, if needed (that is, if '--changed-path' option from the\nfuture commit is provided to 'git commit-graph write'),\ncompute_bloom_filters() would be called befor write_commit_graph_file(),\nwhich in turn runs write_graph_chunk_bloom_index() and *_data().\n\n\nActually, when writing Bloom data chunks (BIDX and BDAT) we could have\nrequested recomputing filters if necessary: slab storage works as\nmemoization, so you would calculate Bloom filter data for each commit in\nthe commit-graph only once.  And write_graph_chunk_bloom_indexes()\nand write_graph_chunk_bloom_data() are called only if ctx->changed_paths\nis true.\n\nSo it would work with\n\n  +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 1);\n\nOnly in the future we would really need to call with compute_if_not_present\nparameter set to falsy value.\n\n> @@ -1304,7 +1304,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>  \n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>  \t\tstruct commit *c = sorted_by_pos[i];\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n>  \t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n>  \t\tdisplay_progress(progress, i + 1);\n>  \t}\n> @@ -2314,6 +2314,7 @@ void free_commit_graph(struct commit_graph *g)\n>  \t\tg->data = NULL;\n>  \t\tclose(g->graph_fd);\n>  \t}\n> +\tfree(g->bloom_filter_settings);\n>  \tfree(g->filename);\n>  \tfree(g);\n\nShouldn't this fixup be added to earlier commit?\n\n>  }\n> diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n> index 331957011b..9b4be97f75 100644\n> --- a/t/helper/test-bloom.c\n> +++ b/t/helper/test-bloom.c\n> @@ -47,7 +47,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n>  \tstruct bloom_filter *filter;\n>  \tsetup_git_directory();\n>  \tc = lookup_commit(the_repository, commit_oid);\n> -\tfilter = get_bloom_filter(the_repository, c);\n> +\tfilter = get_bloom_filter(the_repository, c, 1);\n>  \tprint_bloom_filter(filter);\n>  }\n\nI would like to see some tests, but that needs to wait for patch that\nadds --changed-paths option to the 'write' subcommand.\n\nThings to be tested:\n1. That after reading commit-graph with Bloom filter:\n   - that commit(s) in commit-graph have Bloom filter\n   - that commits outside commit-graph do not have Bloom filter\n2. That incremental commit-graph feature works:\n   - for commits in deeper layer that have Bloom filter chunks\n   - for commits in deeper layer that do not have Bloom filter chunks\n\nBest,\n-- \nJakub Narębski\n"},{"id":"392195","messageId":"86y2sxdmw9.fsf@gmail.com","threadId":"52499","inReplyTo":"3d7ee0c96955dc15c87d04982d8cdec8b62750b2.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 09/11] commit-graph: add --changed-paths option to write subcommand","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-20T20:28:06Z","receivedAt":"2020-02-20T20:28:18Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Add --changed-paths option to git commit-graph write. This option will\n> allow users to compute information about the paths that have changed\n> between a commit and its first parent, and write it into the commit graph\n> file. If the option is passed to the write subcommand we set the\n> COMMIT_GRAPH_WRITE_BLOOM_FILTERS flag and pass it down to the\n> commit-graph logic.\n\nIn the manpage you write that this operation (computing Bloom filters)\ncan take a while on large repositories.  Could you perhaps provide some\nnumbers: how much longer does it take to write commit-graph file with\nand without '--changed-paths' for example for Linux kernel, or some\nother large repository?  Thanks in advance.\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  Documentation/git-commit-graph.txt | 5 +++++\n>  builtin/commit-graph.c             | 9 +++++++--\n>  2 files changed, 12 insertions(+), 2 deletions(-)\n\nWhat is missing is some sanity tests: that bloom index and bloom data\nchunks are not present without '--changed-paths', and that they are\nadded with '--changed-paths'.\n\nIf possible, maybe also check in a separate test that the size of\nbloom_index chunk agrees with the number of commits in the commit graph.\n\n\nAlso, we can now add those tests I have wrote about in my review of\nprevious patch, that is:\n\n1. If you write commit-graph with --changed-paths, and either add some\n   commits later or exclude some commits from the commit graph, then:\n\n   a.) commit(s) in commit-graph have Bloom filter\n   b.) commit(s) not in commit-graph do not have Bloom filter\n\n2. If you write commit-graph without --changed-paths as base layer,\n   and then write next layer with --changed-paths and --split, then:\n\n   a.) commit(s) in top layer have Bloom filter(s)\n   b.) commit(s) in bottom layer don't have Bloom filter(s)\n\n>\n> diff --git a/Documentation/git-commit-graph.txt b/Documentation/git-commit-graph.txt\n> index bcd85c1976..907d703b30 100644\n> --- a/Documentation/git-commit-graph.txt\n> +++ b/Documentation/git-commit-graph.txt\n> @@ -54,6 +54,11 @@ or `--stdin-packs`.)\n>  With the `--append` option, include all commits that are present in the\n>  existing commit-graph file.\n>  +\n> +With the `--changed-paths` option, compute and write information about the\n> +paths changed between a commit and it's first parent. This operation can\n> +take a while on large repositories. It provides significant performance gains\n> +for getting history of a directory or a file with `git log -- <path>`.\n> ++\n\nShould we write about limitation that the topmost layer in the split\ncommit graph needs to be written with '--changed-paths' for Git to use\nthis information?  Or perhaps we should try (in the future) to remove\nthis limitation??\n\n>  With the `--split` option, write the commit-graph as a chain of multiple\n>  commit-graph files stored in `<dir>/info/commit-graphs`. The new commits\n>  not already in the commit-graph are added in a new \"tip\" file. This file\n> diff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\n> index e0c6fc4bbf..261dcce091 100644\n> --- a/builtin/commit-graph.c\n> +++ b/builtin/commit-graph.c\n> @@ -9,7 +9,7 @@\n>  \n>  static char const * const builtin_commit_graph_usage[] = {\n>  \tN_(\"git commit-graph verify [--object-dir <objdir>] [--shallow] [--[no-]progress]\"),\n> -\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n> +\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n>  \tNULL\n>  };\n>  \n> @@ -19,7 +19,7 @@ static const char * const builtin_commit_graph_verify_usage[] = {\n>  };\n>  \n>  static const char * const builtin_commit_graph_write_usage[] = {\n> -\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n> +\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n>  \tNULL\n>  };\n>\n\nAll right.\n\n> @@ -32,6 +32,7 @@ static struct opts_commit_graph {\n>  \tint split;\n>  \tint shallow;\n>  \tint progress;\n> +\tint enable_changed_paths;\n\nBikeshed painting: should this field be called enable_changed_paths or\nsimply changed_paths?\n\n>  } opts;\n>  \n>  static int graph_verify(int argc, const char **argv)\n> @@ -110,6 +111,8 @@ static int graph_write(int argc, const char **argv)\n>  \t\t\tN_(\"start walk at commits listed by stdin\")),\n>  \t\tOPT_BOOL(0, \"append\", &opts.append,\n>  \t\t\tN_(\"include all commits already in the commit-graph file\")),\n> +\t\tOPT_BOOL(0, \"changed-paths\", &opts.enable_changed_paths,\n> +\t\t\tN_(\"enable computation for changed paths\")),\n>  \t\tOPT_BOOL(0, \"progress\", &opts.progress, N_(\"force progress reporting\")),\n>  \t\tOPT_BOOL(0, \"split\", &opts.split,\n>  \t\t\tN_(\"allow writing an incremental commit-graph file\")),\n\nAll right.\n\n> @@ -143,6 +146,8 @@ static int graph_write(int argc, const char **argv)\n>  \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n>  \tif (opts.progress)\n>  \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n> +\tif (opts.enable_changed_paths)\n> +\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n>  \n>  \tread_replace_refs = 0;\n\nAll right.  This actually turns on calculation Bloom filters for changed\npaths, thanks to\n\n \tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n\nthat was added by the \"[PATCH v2 04/11] commit-graph: compute Bloom\nfilters for changed paths\" patch.\n\nThough... should this enabling be split into two separate patches like\nthis?\n\n\nBest,\n-- \nJakub Narębski\n"},{"id":"392206","messageId":"CAGyf7-FzaG3Jb92JTx1QyADAoLhHCREyadVbTM2vZW-wxK4zEg@mail.gmail.com","threadId":"52499","inReplyTo":"3d7ee0c96955dc15c87d04982d8cdec8b62750b2.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 09/11] commit-graph: add --changed-paths option to write subcommand","fromName":"Bryan Turner","fromEmail":"bturner@atlassian.com","sentAt":"2020-02-20T22:10:57Z","receivedAt":"2020-02-20T22:11:12Z","isPatch":true,"sender":{"key":"bturner@atlassian.com","avatar":"https://gravatar.com/avatar/16bcf3167981c1ef7c804e502642366d888a35b0d0b0a4ca01fdc442aa1acb1e?d=mp&s=160"},"body":"On Wed, Feb 5, 2020 at 2:56 PM Garima Singh via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Add --changed-paths option to git commit-graph write. This option will\n> allow users to compute information about the paths that have changed\n> between a commit and its first parent, and write it into the commit graph\n> file. If the option is passed to the write subcommand we set the\n> COMMIT_GRAPH_WRITE_BLOOM_FILTERS flag and pass it down to the\n> commit-graph logic.\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  Documentation/git-commit-graph.txt | 5 +++++\n>  builtin/commit-graph.c             | 9 +++++++--\n>  2 files changed, 12 insertions(+), 2 deletions(-)\n>\n> diff --git a/Documentation/git-commit-graph.txt b/Documentation/git-commit-graph.txt\n> index bcd85c1976..907d703b30 100644\n> --- a/Documentation/git-commit-graph.txt\n> +++ b/Documentation/git-commit-graph.txt\n> @@ -54,6 +54,11 @@ or `--stdin-packs`.)\n>  With the `--append` option, include all commits that are present in the\n>  existing commit-graph file.\n>  +\n> +With the `--changed-paths` option, compute and write information about the\n> +paths changed between a commit and it's first parent. This operation can\n\n\"its first parent\"\n\n(Pardon the grammar nit from the peanut gallery!)\n\n> +take a while on large repositories. It provides significant performance gains\n> +for getting history of a directory or a file with `git log -- <path>`.\n> ++\n>  With the `--split` option, write the commit-graph as a chain of multiple\n>  commit-graph files stored in `<dir>/info/commit-graphs`. The new commits\n>  not already in the commit-graph are added in a new \"tip\" file. This file\n> diff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\n> index e0c6fc4bbf..261dcce091 100644\n> --- a/builtin/commit-graph.c\n> +++ b/builtin/commit-graph.c\n> @@ -9,7 +9,7 @@\n>\n>  static char const * const builtin_commit_graph_usage[] = {\n>         N_(\"git commit-graph verify [--object-dir <objdir>] [--shallow] [--[no-]progress]\"),\n> -       N_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n> +       N_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n>         NULL\n>  };\n>\n> @@ -19,7 +19,7 @@ static const char * const builtin_commit_graph_verify_usage[] = {\n>  };\n>\n>  static const char * const builtin_commit_graph_write_usage[] = {\n> -       N_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n> +       N_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n>         NULL\n>  };\n>\n> @@ -32,6 +32,7 @@ static struct opts_commit_graph {\n>         int split;\n>         int shallow;\n>         int progress;\n> +       int enable_changed_paths;\n>  } opts;\n>\n>  static int graph_verify(int argc, const char **argv)\n> @@ -110,6 +111,8 @@ static int graph_write(int argc, const char **argv)\n>                         N_(\"start walk at commits listed by stdin\")),\n>                 OPT_BOOL(0, \"append\", &opts.append,\n>                         N_(\"include all commits already in the commit-graph file\")),\n> +               OPT_BOOL(0, \"changed-paths\", &opts.enable_changed_paths,\n> +                       N_(\"enable computation for changed paths\")),\n>                 OPT_BOOL(0, \"progress\", &opts.progress, N_(\"force progress reporting\")),\n>                 OPT_BOOL(0, \"split\", &opts.split,\n>                         N_(\"allow writing an incremental commit-graph file\")),\n> @@ -143,6 +146,8 @@ static int graph_write(int argc, const char **argv)\n>                 flags |= COMMIT_GRAPH_WRITE_SPLIT;\n>         if (opts.progress)\n>                 flags |= COMMIT_GRAPH_WRITE_PROGRESS;\n> +       if (opts.enable_changed_paths)\n> +               flags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n>\n>         read_replace_refs = 0;\n>\n> --\n> gitgitgadget\n>\n"},{"id":"392272","messageId":"86o8trdeyh.fsf@gmail.com","threadId":"52499","inReplyTo":"77f1c561e8205c0598b57bf572640d21d64757f8.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 10/11] revision.c: use Bloom filters to speed up path based revision walks","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-21T17:31:50Z","receivedAt":"2020-02-21T17:32:01Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Revision walk will now use Bloom filters for commits to speed up revision\n> walks for a particular path (for computing history for that path), if they\n> are present in the commit-graph file.\n\nWhy do we need to turn this feature off for --walk-reflog?\n\nAnyway, in my opinion this restriction should be stated explicitly in\nthe commit message, if kept. \n\n>\n> We load the Bloom filters during the prepare_revision_walk step, but only\n> when dealing with a single pathspec.\n\nI would add the qualifier \"currently\" here, i.e. s/only/currently only/\nto make it clear that it is the limitation of current implementation,\nand not the inherent implementation of the technique.\n\n>                                      While comparing trees in\n> rev_compare_trees(), if the Bloom filter says that the file is not different\n> between the two trees, we don't need to compute the expensive diff. This is\n> where we get our performance gains. The other response of the Bloom filter\n> is `maybe`, in which case we fall back to the full diff calculation to\n> determine if the path was changed in the commit.\n\nAll right, looks good.\n\nVery minor nitpick: s/`maybe`/'maybe'/ (in my opinion).\n\n>\n> Performance Gains:\n> We tested the performance of `git log -- <path>` on the git repo, the linux\n> and some internal large repos, with a variety of paths of varying depths.\n\nAnother repository that we could test Bloom filters feature would be, as\nI have written before, Android AOSP frameworks core repository\nhttps://android.googlesource.com/platform/frameworks/base/\nbecause being written in Java it has deep path hierarchy, and it also\nhas large number of commits.\n\n>\n> On the git and linux repos:\n> - we observed a 2x to 5x speed up.\n\nIt would be nice to have at least one specific and repeatable example:\nin given repository, starting from given commit or tag, following the\nhistory of given path, what are timing results for doing some specific\ncommand with and without Bloom filters computed and enabled.\n\nOne might also want to know the cost of this speedup: how much disk\nspace does it take (i.e. how large is the commit-graph file with and\nwithout Bloom filters chunks), and how long does it take to compute\n(i.e. how much time writing commit-graph takes with and without using\n--changed-paths options).\n\n>\n> On a large internal repo with files seated 6-10 levels deep in the tree:\n> - we observed 10x to 20x speed ups, with some paths going up to 28 times\n>   faster.\n\nThis is good to know.\n\nIn the future we might want to have procedurally generated synthetic\nrepository, where we would be able to control number of files, depth of\nfilesystem hierarchy, average number of changes per commit, etc. to be\nused for performance testing.  (Just wishful thinking)\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com\n> Helped-by: SZEDER Gábor <szeder.dev@gmail.com>\n> Helped-by: Jonathan Tan <jonathantanmy@google.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  revision.c                 | 124 +++++++++++++++++++++++++++++++-\n>  revision.h                 |  11 +++\n>  t/helper/test-read-graph.c |   4 ++\n>  t/t4216-log-bloom.sh       | 140 +++++++++++++++++++++++++++++++++++++\n>  4 files changed, 277 insertions(+), 2 deletions(-)\n>  create mode 100755 t/t4216-log-bloom.sh\n>\n> diff --git a/revision.c b/revision.c\n> index 8136929e23..d1622afa17 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -29,6 +29,8 @@\n>  #include \"prio-queue.h\"\n>  #include \"hashmap.h\"\n>  #include \"utf8.h\"\n> +#include \"bloom.h\"\n> +#include \"json-writer.h\"\n>  \n>  volatile show_early_output_fn_t show_early_output;\n>  \n> @@ -624,11 +626,114 @@ static void file_change(struct diff_options *options,\n>  \toptions->flags.has_changes = 1;\n>  }\n>  \n> +static int bloom_filter_atexit_registered;\n> +static unsigned int count_bloom_filter_maybe;\n> +static unsigned int count_bloom_filter_definitely_not;\n> +static unsigned int count_bloom_filter_false_positive;\n> +static unsigned int count_bloom_filter_not_present;\n> +static unsigned int count_bloom_filter_length_zero;\n> +\n> +static void trace2_bloom_filter_statistics_atexit(void)\n> +{\n> +\tstruct json_writer jw = JSON_WRITER_INIT;\n> +\n> +\tjw_object_begin(&jw, 0);\n> +\tjw_object_intmax(&jw, \"filter_not_present\", count_bloom_filter_not_present);\n> +\tjw_object_intmax(&jw, \"zero_length_filter\", count_bloom_filter_length_zero);\n> +\tjw_object_intmax(&jw, \"maybe\", count_bloom_filter_maybe);\n> +\tjw_object_intmax(&jw, \"definitely_not\", count_bloom_filter_definitely_not);\n> +\tjw_end(&jw);\n> +\n> +\ttrace2_data_json(\"bloom\", the_repository, \"statistics\", &jw);\n> +\n> +\tjw_release(&jw);\n> +}\n\nI thought that it would be better to put this part together with tests\nthat absolutely require this functionality in a separate subsequent\npatch, but now I am not so sure.  It is nice to have all or almost all\ntests created in a single patch.\n\nLooks good to me, but I don't know much about trace2 API, so take it\nwith a pinch of salt.\n\n> +\n> +static void prepare_to_use_bloom_filter(struct rev_info *revs)\n> +{\n> +\tstruct pathspec_item *pi;\n> +\tchar *path_alloc = NULL;\n> +\tconst char *path;\n> +\tint last_index;\n> +\tint len;\n> +\n> +\tif (!revs->commits)\n> +\t    return;\n\nI see that we need this because in next command we dereference\nrevs->commits to get revs->commits->item.\n\nIf I understand it correctly empty pending list may happen with \"--all\"\nor \"--glob\" options, but somebody with more experience in this area of\ncode is needed to state for sure.\n\nShould we test `git log --all -- <path>`?\n\n> +\n> +\trepo_parse_commit(revs->repo, revs->commits->item);\n\nAre we calling this function for its side-effects?  Wouldn't using\nprepare_commit_graph(revs->repo) here be a better solution?\n\n> +\n> +\tif (!revs->repo->objects->commit_graph)\n> +\t\treturn;\n\nLooks good to me.  If there is no commit graph, then there are no Bloom\nfilters to consult.\n\n> +\n> +\trevs->bloom_filter_settings = revs->repo->objects->commit_graph->bloom_filter_settings;\n\nHmmm... is that why bloom_filter_settings is a pointer to struct, and\nnot struct itself?\n\n> +\tif (!revs->bloom_filter_settings)\n> +\t\treturn;\n\nLooks good to me.  If there is no Bloomm filter in the commit-graph\nfile, then there are no Bloom filters to consult.\n\n> +\n> +\tpi = &revs->pruning.pathspec.items[0];\n> +\tlast_index = pi->len - 1;\n> +\n\nIt might be a good idea to add a comment explaining what is happening\nhere, for example:\n\n  +\t/* remove single trailing slash from path, if needed */\n> +\tif (pi->match[last_index] == '/') {\n> +\t    path_alloc = xstrdup(pi->match);\n> +\t    path_alloc[last_index] = '\\0';\n> +\t    path = path_alloc;\n> +\t} else\n> +\t    path = pi->match;\n> +\n> +\tlen = strlen(path);\n\nWe can avoid computing strlen(path) here, because in first branch of\nthis conditional we have len = last_index, in the second branch we have\nlen = pi->len.\n\n> +\n> +\trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n> +\tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n\nAll right, this is the meat of this function: creating bloom_key for a\npath.  Looks good to me.\n\n> +\n> +\tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n> +\t\tatexit(trace2_bloom_filter_statistics_atexit);\n> +\t\tbloom_filter_atexit_registered = 1;\n> +\t}\n\nOK, here we register trace2 Bloom filter statistics handler, but only\nonce, and only when needed.\n\n> +\n> +\tfree(path_alloc);\n\nOK, path_alloc is either xstrdup-ed string, or NULL, and is no longer\nneeded (after possibly being used to create bloom_key).\n\n> +}\n> +\n> +static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n> +\t\t\t\t\t\t struct commit *commit)\n> +{\n> +\tstruct bloom_filter *filter;\n> +\tint result;\n> +\n> +\tif (!revs->repo->objects->commit_graph)\n> +\t\treturn -1;\n> +\n> +\tif (commit->generation == GENERATION_NUMBER_INFINITY)\n> +\t\treturn -1;\n\nIdle thought: would it be useful to gather for trace2 statistics also\nnumber of commits encountered that were outside commit-graph?\n\n> +\n> +\tfilter = get_bloom_filter(revs->repo, commit, 0);\n> +\n> +\tif (!filter) {\n> +\t\tcount_bloom_filter_not_present++;\n> +\t\treturn -1;\n> +\t}\n> +\n> +\tif (!filter->len) {\n> +\t\tcount_bloom_filter_length_zero++;\n> +\t\treturn -1;\n> +\t}\n> +\n> +\tresult = bloom_filter_contains(filter,\n> +\t\t\t\t       revs->bloom_key,\n> +\t\t\t\t       revs->bloom_filter_settings);\n> +\n> +\tif (result)\n> +\t\tcount_bloom_filter_maybe++;\n> +\telse\n> +\t\tcount_bloom_filter_definitely_not++;\n> +\n> +\treturn result;\n> +}\n\nThe whole check_maybe_different_in_bloom_filter() looks good to me,\nthanks to designing and building a good API.\n\n> +\n>  static int rev_compare_tree(struct rev_info *revs,\n> -\t\t\t    struct commit *parent, struct commit *commit)\n> +\t\t\t    struct commit *parent, struct commit *commit, int nth_parent)\n>  {\n>  \tstruct tree *t1 = get_commit_tree(parent);\n>  \tstruct tree *t2 = get_commit_tree(commit);\n> +\tint bloom_ret = 1;\n\nI don't understand why it is initialized to 1, and not to 0.\n\n>  \n>  \tif (!t1)\n>  \t\treturn REV_TREE_NEW;\n> @@ -653,11 +758,23 @@ static int rev_compare_tree(struct rev_info *revs,\n>  \t\t\treturn REV_TREE_SAME;\n>  \t}\n>  \n> +\tif (revs->pruning.pathspec.nr == 1 && !revs->reflog_info && !nth_parent) {\n\nShouldn't we check upfront here that revs->bloom_key is not NULL?\nI don't think we check this down the callchain...\n\nOr even better replace the first two checks with it, as revs->bloom_key\nis set only if (revs->pruning.pathspec.nr == 1 && !revs->reflog_info),\nsee addition to prepare_revision_walk() below.\n\nOf course the !nth_parent check needs to be kept, as this changes during\nthe revision walk (it is a limitation of current version of Bloom filter\nin that only changes with respect to first parent are stored in filter).\n\n> +\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n> +\n> +\t\tif (bloom_ret == 0)\n> +\t\t\treturn REV_TREE_SAME;\n> +\t}\n\nAll right, if we have single pathspec, and we don't walk reflog (?), and\nwe are interested in first parent, then we query the Bloom filter.\n\nThe Bloom filter can return 'no' or 'maybe'; if it returns 'no' then we\ncan short-circuit and avoid computing the tree diff.\n\n> +\n>  \ttree_difference = REV_TREE_SAME;\n>  \trevs->pruning.flags.has_changes = 0;\n>  \tif (diff_tree_oid(&t1->object.oid, &t2->object.oid, \"\",\n>  \t\t\t   &revs->pruning) < 0)\n>  \t\treturn REV_TREE_DIFFERENT;\n> +\n> +\tif (!nth_parent)\n\nShouldn't this condition be exactly the same as for running\ncheck_maybe_different_in_bloom_filter()?  Otherwise due to initializing\nbloom_ret to 1 we would get wrong statistics, isn't it?\n\n> +\t\tif (bloom_ret == 1 && tree_difference == REV_TREE_SAME)\n> +\t\t\tcount_bloom_filter_false_positive++;\n> +\n\nAll right, looks good.\n\n>  \treturn tree_difference;\n>  }\n>  \n> @@ -855,7 +972,7 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit)\n>  \t\t\tdie(\"cannot simplify commit %s (because of %s)\",\n>  \t\t\t    oid_to_hex(&commit->object.oid),\n>  \t\t\t    oid_to_hex(&p->object.oid));\n> -\t\tswitch (rev_compare_tree(revs, p, commit)) {\n> +\t\tswitch (rev_compare_tree(revs, p, commit, nth_parent)) {\n>  \t\tcase REV_TREE_SAME:\n>  \t\t\tif (!revs->simplify_history || !relevant_commit(p)) {\n>  \t\t\t\t/* Even if a merge with an uninteresting\n\nOK, we are just dding new parameter, with the information needed to\ndecide whether Bloom filters can be used or not.\n\n> @@ -3362,6 +3479,8 @@ int prepare_revision_walk(struct rev_info *revs)\n>  \t\t\t\t       FOR_EACH_OBJECT_PROMISOR_ONLY);\n>  \t}\n>  \n> +\tif (revs->pruning.pathspec.nr == 1 && !revs->reflog_info)\n> +\t\tprepare_to_use_bloom_filter(revs);\n\nWell, the limitation that the technique _currently_ works only with a\nsingle pathspec is stated explicitly, but the fact that it is turned off\nfor some reason for --walk-reflog is not.\n\nOtherwise, looks good to me.\n\n>  \tif (revs->no_walk != REVISION_WALK_NO_WALK_UNSORTED)\n>  \t\tcommit_list_sort_by_date(&revs->commits);\n>  \tif (revs->no_walk)\n> @@ -3379,6 +3498,7 @@ int prepare_revision_walk(struct rev_info *revs)\n>  \t\tsimplify_merges(revs);\n>  \tif (revs->children.name)\n>  \t\tset_children(revs);\n> +\n>  \treturn 0;\n>  }\n\nUnrelated coding style fixup, but we are doing changes in the\nneighborhood.  All right, I can agree to that.\n\n>  \n> diff --git a/revision.h b/revision.h\n> index 475f048fb6..7c026fe41f 100644\n> --- a/revision.h\n> +++ b/revision.h\n> @@ -56,6 +56,8 @@ struct repository;\n>  struct rev_info;\n>  struct string_list;\n>  struct saved_parents;\n> +struct bloom_key;\n> +struct bloom_filter_settings;\n>  define_shared_commit_slab(revision_sources, char *);\n>  \n>  struct rev_cmdline_info {\n> @@ -291,6 +293,15 @@ struct rev_info {\n>  \tstruct revision_sources *sources;\n>  \n>  \tstruct topo_walk_info *topo_walk_info;\n> +\n> +\t/* Commit graph bloom filter fields */\n> +\t/* The bloom filter key for the pathspec */\n> +\tstruct bloom_key *bloom_key;\n> +\t/*\n> +\t * The bloom filter settings used to generate the key.\n> +\t * This is loaded from the commit-graph being used.\n> +\t */\n> +\tstruct bloom_filter_settings *bloom_filter_settings;\n\nIt is nice having those explanatory comments.\n\nSidenote: if I understand it correctly, revs->bloom_key is allocated but\nnever free()d.  On the other hand revs->bloom_filter_settings is a weak\nreference / is set to the value of other pointer, which is allocated and\nfree()d together with commit_graph struct.\n\n>  };\n>  \n>  int ref_excluded(struct string_list *, const char *path);\n> diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\n> index d2884efe0a..aff597c7a3 100644\n> --- a/t/helper/test-read-graph.c\n> +++ b/t/helper/test-read-graph.c\n> @@ -45,6 +45,10 @@ int cmd__read_graph(int argc, const char **argv)\n>  \t\tprintf(\" commit_metadata\");\n>  \tif (graph->chunk_extra_edges)\n>  \t\tprintf(\" extra_edges\");\n> +\tif (graph->chunk_bloom_indexes)\n> +\t\tprintf(\" bloom_indexes\");\n> +\tif (graph->chunk_bloom_data)\n> +\t\tprintf(\" bloom_data\");\n>  \tprintf(\"\\n\");\n\nThis chunk could be moved to the commit adding --changed-paths\noption... on the other hand if all tests are to be added by this patch,\nit can be left as is.\n\n>  \n>  \tUNLEAK(graph);\n> diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> new file mode 100755\n> index 0000000000..19eca1864b\n> --- /dev/null\n> +++ b/t/t4216-log-bloom.sh\n[...]\n\nI'll leave reviewing tests of this feature for the next email.\n\nBest regards,\n-- \nJakub Narębski\n"},{"id":"392273","messageId":"fdcbd793-57c2-f5ea-ccb9-cf34e911b669@gmail.com","threadId":"52499","inReplyTo":"86a75swuie.fsf@gmail.com","subject":"Re: [PATCH v2 00/11] Changed Paths Bloom Filters","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-21T17:41:22Z","receivedAt":"2020-02-21T17:41:30Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/8/2020 6:04 PM, Jakub Narebski wrote:\n> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> Hey! \n>>\n>> The commit graph feature brought in a lot of performance improvements across\n>> multiple commands. However, file based history continues to be a performance\n>> pain point, especially in large repositories. \n>>\n>> Adopting changed path Bloom filters has been discussed on the list before,\n>> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n>> Derrick Stolee [1]. This series is based on Dr. Stolee's proof of\n>> concept in [2].\n> \n> Sidenote: I wondered why it did use MurmurHash3 (64-bit version), which\n> requires adding its implementation, instead of reusing FNV-1 hash\n> (Fowler–Noll–Vo hash function) used by Git hashmap implementation, see\n> https://github.com/git/git/blob/228f53135a4a41a37b6be8e4d6e2b6153db4a8ed/hashmap.h#L109\n> Beside the fact that everyone is using MurmurHash for Bloom filters ;-)\n> \n> It turns out that in various benchmark MurmurHash is faster and also\n> slightly better as a hash than FNV-1 or FNV-1b.\n> \n> \n> I wonder then if it would be a good idea (in the future) to make it easy\n> to use hashmap with MurmurHash3 instead of FNV-1, or maybe to even make\n> it the default for hashing strings.\n> \n\nMaking Murmur3 hash the default for hashing strings is definitely outside the\nscope of this series. Also, if the method signatures for the murmur3 hash \nmatched the existing hash method signatures in hashmap.c, then it would be \nappropriate to place them adjacently, even if no hashmap consumer uses it for \nhashmaps. However, we need the option to start at a custom seed to do our double\nhashing. A change in the future that involves adopting murmur3 in the hashmap\ncode would involve a simple code move before creating the new methods that \navoid a custom seed. So for now, it makes sense that these methods leave in \nbloom.c where they are being used for a very specific purpose. \n\n>>\n>> Performance Gains: We tested the performance of git log -- path on the git\n>> repo, the linux repo and some internal large repos, with a variety of paths\n>> of varying depths.\n> \n> As I wrote in reply to previous version of this series, a good public\n> repository (and thus being able to use by anyone) to test the Bloom\n> filter performance improvements could be AOSP (Android) base:\n> \n>   https://android.googlesource.com/platform/frameworks/base/\n> \n> which is a large repository with long path depths (due to Java file\n> naming conventions).\n> \n\nThank you! I will incorporate these results into the commit messages as \nappropriate in v3. \n\n>>\n>> On the git and linux repos: We observed a 2x to 5x speed up.\n>>\n>> On a large internal repo with files seated 6-10 levels deep in the tree: We\n>> observed 10x to 20x speed ups, with some paths going up to 28 times faster.\n> \n> Very nice! Good work!\n> \n> What is the cost of this feature, that is how long it takes to generate\n> Bloom filters, and how much larger commit-graph file gets?  It would be\n> nice to know.\n> \n\nThe cost of writing is much better now with Peff and Dr. Stolee's improvements. \nI will include these numbers as well in the commit messages as appropriate in \nv3. \n\n>>\n>> Future Work (not included in the scope of this series):\n>>\n>>  1. Supporting multiple path based revision walk\n> \n> Shouldn't then tests that were added in v2 mark use of Bloom filters\n> with multiple paths revision walking as _not working *yet*_\n> (test_expect_failure), and not expected to not work (test_expect_success\n> with test_bloom_filters_not_used)?\n> \n\nMy intent is to ensure that bloom filters are not being used in any of the \nunsupported code paths. I don't have a strong preference about the test \nsemantics as long as I get that coverage :) So I will look into switching it \nto test_expect_failure as you have suggested. \n\n>> Derrick Stolee (2):\n>>   diff: halt tree-diff early after max_changes\n>>   commit-graph: examine commits by generation number\n>>\n>> Garima Singh (8):\n>>   commit-graph: use MAX_NUM_CHUNKS\n>>   bloom: core Bloom filter implementation for changed paths\n>>   commit-graph: compute Bloom filters for changed paths\n>>   commit-graph: write Bloom filters to commit graph file\n>>   commit-graph: reuse existing Bloom filters during write.\n>>   commit-graph: add --changed-paths option to write subcommand\n>>   revision.c: use Bloom filters to speed up path based revision walks\n>>   commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag\n>>\n>> Jeff King (1):\n>>   commit-graph: examine changed-path objects in pack order\n> \n> The shortlog summary is a fine tool to show contributors to the patch\n> series, but is not as useful to show patch series as a whole: splitting\n> of patches and their ordering.\n> \n\nThis is a GitGitGadget specific thing, and it is probably by design. I have \nopened an issue in that repo for any follow up discussions:\n  https://github.com/gitgitgadget/gitgitgadget/issues/203\n\n> - [PATCH v2 02/11] bloom: core Bloom filter implementation for changed paths\n> \n>   In my opinion this patch could be split into three individual pieces,\n>   though one might think it is not worth it.\n> \n\nI have gone back and forth on doing this. I like most of the core Bloom filter\ncomputations being isolated in one patch/commit. But based on the rest of your\nreview, it seems like you are leaning heavily on having this split out. \nSo, I will take a proper stab at doing it for v3. \n\n> - [PATCH v2 07/11] commit-graph: write Bloom filters to commit graph file\n> \n>   This commit includes the documentation of the two new chunks of\n>   commit-graph file format.\n> \n>   I wonder if the 9th patch in this series, namely\n>   commit-graph: add --changed-paths option to write subcommand\n>   should not precede this commit.  Otherwise we have this new code but\n>   no way of testing it.  On the other hand it makes it easier to\n>   review.  On the gripping hand, you can't really test that writing\n>   works without the ability to parse Bloom filter data out of\n>   commit-graph file... which is the next commit.\n> \n\nGetting complete test coverage within a single patch would require 2 or 3 of \nthese patches to be combined. This would lead to a large patch that would be \nmuch more difficult to review.\n \nMy tests in the patches following this one run git commands. Hence the tests \nget introduced when the command line is ready to use all the new code. \n\nThe current ordering of patches works better than adding the --changed-paths \noption before the logic that computes and writes. Otherwise the option will not \nbe doing what it is supposed to do in the patch it was introduced in.\n\n> - [PATCH v2 08/11] commit-graph: reuse existing Bloom filters during write\n> \n>   This implements reading Bloom filters data from commit-graph file.\n>   Is it a good split?  I think it makes it easier to review the single\n>   patch, but itt also makes them less standalone.\n> \n\nAll the logic upto this point works just fine without the ability to read and \nparse precomputed bloom filters. This patch is an enhancement and it also \nseparates out the reading and writing logic. Reusing existing bloom filters \nduring write is the simplest interatcion that involves reading from the commit\ngraph file, and builds the foundation to make the `git log` improvements. \nHence, it warrants its own patch and review. \n\n> - [PATCH v2 10/11] revision.c: use Bloom filters to speed up path based revision walks\n> \n>   This is quite a big and involved patch, which in my opinion could be\n>   split in two or three parts:\n> \n>   a. Add a bare bones implementation, like in v2\n> \n>   This limits amount of testing we can do; the only thing we can really\n>   test is that we get the same results with and without Bloom filters.\n> \n>   b.1. Add trace2 Bloom filter statistics\n>   b.2. Use said trace2 statistics to test use of Bloom filters\n> \n\nSure. I will look into doing this split as well for v3. \n\n> \n> Feel free to disagree with those ideas.\n> \n> Best,\n\nThanks for taking the time for reviewing this series so thoroughly! \nIt is greatly appreciated! \n\nCheers,\nGarima Singh\n\n"},{"id":"392285","messageId":"86r1ynbluo.fsf@gmail.com","threadId":"52499","inReplyTo":"77f1c561e8205c0598b57bf572640d21d64757f8.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 10/11] revision.c: use Bloom filters to speed up path based revision walks","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-21T22:45:51Z","receivedAt":"2020-02-21T22:46:09Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\nThis is a second part of my response, focusing solely on tests of the\nBloom filters feature.\n\n[...]\n> diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\n> index d2884efe0a..aff597c7a3 100644\n> --- a/t/helper/test-read-graph.c\n> +++ b/t/helper/test-read-graph.c\n> @@ -45,6 +45,10 @@ int cmd__read_graph(int argc, const char **argv)\n>  \t\tprintf(\" commit_metadata\");\n>  \tif (graph->chunk_extra_edges)\n>  \t\tprintf(\" extra_edges\");\n> +\tif (graph->chunk_bloom_indexes)\n> +\t\tprintf(\" bloom_indexes\");\n> +\tif (graph->chunk_bloom_data)\n> +\t\tprintf(\" bloom_data\");\n>  \tprintf(\"\\n\");\n>\n\nAll right, that is simple extension of 'test-helper read-graph'.\n\n>  \tUNLEAK(graph);\n> diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> new file mode 100755\n> index 0000000000..19eca1864b\n> --- /dev/null\n> +++ b/t/t4216-log-bloom.sh\n> @@ -0,0 +1,140 @@\n> +#!/bin/sh\n> +\n> +test_description='git log for a path with bloom filters'\n> +. ./test-lib.sh\n> +\n> +test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n> +\tgit init &&\n> +\tmkdir A A/B A/B/C &&\n> +\ttest_commit c1 A/file1 &&\n> +\ttest_commit c2 A/B/file2 &&\n> +\ttest_commit c3 A/B/C/file3 &&\n> +\ttest_commit c4 A/file1 &&\n> +\ttest_commit c5 A/B/file2 &&\n> +\ttest_commit c6 A/B/C/file3 &&\n> +\ttest_commit c7 A/file1 &&\n> +\ttest_commit c8 A/B/file2 &&\n> +\ttest_commit c9 A/B/C/file3 &&\n> +\tgit checkout -b side HEAD~4 &&\n> +\ttest_commit side-1 file4 &&\n> +\tgit checkout master &&\n> +\tgit merge side &&\n> +\ttest_commit c10 file5 &&\n\nUnfortunately this might be not enough for Git's heuristic similarity\nbased rename detection, as it creates 'file5' file with content 'c10'.\n\n[Checking something].  Well, actually it looks like it works, even with\nnot much contents.  I thought you would need to use something like\n\n  +\ttest_write_lines 1 2 3 4 5 6 7 8 9 >file5 &&\n  +\tgit add file5 &&\n  +\tgit commit -m c10 &&\n\nBut it turns out that it is, s far as I have checked, not necessary.\n\n> +\tmv file5 file5_renamed &&\n> +\tgit add file5_renamed &&\n> +\tgit commit -m \"rename\" &&\n> +\tgit commit-graph write --reachable --changed-paths\n> +'\n\nHmmm... there is no test for file that was present in history but got\ndeleted.  Might be important (because of pre-image vs post-image name\nissues).\n\n\nVery minor issue: following the style used in t/test-lib-functions.sh\nand the style guide in CodingGuidelines, it should be\n\n  +graph_read_expect () {\n\nand the same for the following functions.\n\n\nhttps://github.com/git/git/blob/master/Documentation/CodingGuidelines#L144\n\n - We prefer a space between the function name and the parentheses,\n   and no space inside the parentheses. The opening \"{\" should also\n   be on the same line.\n\n\t(incorrect)\n\tmy_function(){\n\t\t...\n\n\t(correct)\n\tmy_function () {\n\t\t...\n\n> +graph_read_expect() {\n> +\tOPTIONAL=\"\"\n> +\tNUM_CHUNKS=5\n> +\tcat >expect <<- EOF\n> +\theader: 43475048 1 1 $NUM_CHUNKS 0\n> +\tnum_commits: $1\n> +\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n\nEither OPTIONAL remains unused, and should be removed, or we leave it\nfor possible future extension, and we write\n\n  +\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data$OPTIONAL\n\nlike in t/t5318-commit-graph.sh.\n\n> +\tEOF\n> +\ttest-tool read-graph >output &&\n> +\ttest_cmp expect output\n\nWhy 'output', and not 'actual'?\n\n> +}\n> +\n> +test_expect_success 'commit-graph write wrote out the bloom chunks' '\n> +\tgraph_read_expect 13\n> +'\n\nAll right, that is sanity-checking 'git commit-graph write --changed-paths'.\n\n> +\n> +setup() {\n\nI wonder if we can come up with a better name... setup_log(),\nsetup_log_bloom(), log_compare()?\n\n> +\trm output\n\nThis shouldn't be here, in this function.  Or perhaps it shouldn't even\nbe used at all; having 'output' doesn't hinder anything.\n\n> +\trm \"$TRASH_DIRECTORY/trace.perf\"\n\nAll right, this cleanup is needed.\n\n> +\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom\n> +\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n\nAll right, we prepare for comparing version without Bloom filters\n(reference) and with Bloom filters, and for checking if Bloom filters\nwere used.\n\n> +}\n\nThis setup() function above is missing the && chain.\n\nIt should then in my opinion read:\n\n  +setup () {\n  +\trm \"$TRASH_DIRECTORY/trace.perf\" &&\n  +\tgit -c core.commitGraph=false log --format=\"%s\" $1 >log_wo_bloom &&\n  +\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" \\\n  +\tgit -c core.commitGraph=true log --format=\"%s\" $1 >log_w_bloom\n  +}\n\nAlso, perhaps we should add at the beginning of this test file, outside\nanu test_expect_success block, the following (see t/*trace2*.sh files):\n\n  # Turn off any inherited trace2 settings for this test.\n  sane_unset GIT_TRACE2 GIT_TRACE2_PERF GIT_TRACE2_EVENT\n  sane_unset GIT_TRACE2_PERF_BRIEF\n  sane_unset GIT_TRACE2_CONFIG_PARAMS\n\n> +\n> +test_bloom_filters_used() {\n> +\tlog_args=$1\n> +\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":0,\\\"zero_length_filter\\\":0,\\\"maybe\\\"\"\n> +\tsetup \"$log_args\"\n\nMissing && chain.\n\n> +\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && test_cmp log_wo_bloom log_w_bloom\n\nWhy no line break after &&?\n\n> +}\n\nUgh, examining JSON output with regexp is in my opinion quite fragile.\nThough I am not sure if requiring Perl and JSON module installed like\nt/t0212-trace2-event.sh is any better.\n\n> +\n> +test_bloom_filters_not_used() {\n> +\tlog_args=$1\n> +\tsetup \"$log_args\"\n> +\t!(grep -q \"statistics:{\\\"filter_not_present\\\":\" \"$TRASH_DIRECTORY/trace.perf\") && test_cmp log_wo_bloom log_w_bloom\n\nWe should also check that \"$TRASH_DIRECTORY/trace.perf\" file exist with\ntest_path_is_file.\n\nAlso, testing that something was not found is a bit fragile, but I don't\nhave any better idea on how to do this test without negating grep exit\nvalue.\n\n> +}\n> +\n> +for path in A A/B A/B/C A/file1 A/B/file2 A/B/C/file3 file4 file5_renamed\n\nNOTE: file5 is missing from this list!\n\nI suspect that adding it might cause the test to fail.\n\n> +do\n> +\tfor option in \"\" \\\n> +\t\t      \"--full-history\" \\\n> +\t\t      \"--full-history --simplify-merges\" \\\n> +\t\t      \"--simplify-merges\" \\\n> +\t\t      \"--simplify-by-decoration\" \\\n> +\t\t      \"--follow\" \\\n> +\t\t      \"--first-parent\" \\\n> +\t\t      \"--topo-order\" \\\n> +\t\t      \"--date-order\" \\\n> +\t\t      \"--author-date-order\" \\\n> +\t\t      \"--ancestry-path side..master\"\n> +\tdo\n> +\t\ttest_expect_success \"git log option: $option for path: $path\" '\n> +\t\t\ttest_bloom_filters_used \"$option -- $path\"\n\nAll right, this tests that Bloom filters were used *and* that the\ncommand run with Bloom filters and without Bloom filters (without\ncommit-graph) produces the same output.\n\n> +\t\t'\n> +\tdone\n> +done\n> +\n> +test_expect_success 'git log -- folder works with and without the trailing slash' '\n> +\ttest_bloom_filters_used \"-- A\" &&\n> +\ttest_bloom_filters_used \"-- A/\"\n> +'\n\nAll right.\n\nI wonder if we should test for insane test case, namely pathname to an\nordinary file that ends with slash:\n\n  +\ttest_bloom_filters_used \"-- file4\" &&\n  +\ttest_bloom_filters_used \"-- file4/\"\n\nThe latter should produce no output, being treated as not existing file.\n\n> +\n> +test_expect_success 'git log for path that does not exist. ' '\n> +\ttest_bloom_filters_used \"-- path_does_not_exist\"\n> +'\n\nAll right.\n\n> +\n> +test_expect_success 'git log with --walk-reflogs does not use bloom filters' '\n> +\ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n> +'\n\nAll right, but why is it so?\n\n> +\n> +test_expect_success 'git log -- multiple path specs does not use bloom filters' '\n> +\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n> +'\n\nAll right, though this is limitation of current code, not limitation of\ntechnique, so _maybe_ it would be better to test_expect_failure that for\nmultiple pathspecs bloom_filters_used...\n\n> +\n> +test_expect_success 'git log with wildcard that resolves to a single path uses bloom filters' '\n> +\ttest_bloom_filters_used \"-- *4\" &&\n> +\ttest_bloom_filters_used \"-- *renamed\"\n> +'\n> +\n> +test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses bloom filters' '\n> +\ttest_bloom_filters_not_used \"-- *\" &&\n> +\ttest_bloom_filters_not_used \"-- file*\"\n> +'\n\nSame here.\n\n> +\n> +test_expect_success 'setup - add commit-graph to the chain without bloom filters' '\n> +\ttest_commit c14 A/anotherFile2 &&\n> +\ttest_commit c15 A/B/anotherFile2 &&\n> +\ttest_commit c16 A/B/C/anotherFile2 &&\n> +\tGIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0 git commit-graph write --reachable --split &&\n> +\ttest_line_count = 2 .git/objects/info/commit-graphs/commit-graph-chain\n> +'\n> +\n> +test_expect_success 'git log does not use bloom filters if the latest graph does not have bloom filters.' '\n> +\ttest_bloom_filters_not_used \"-- A/B\"\n> +'\n\nAll right... though I would try to come up with a shorter test name :-)\n\n> +\n> +test_expect_success 'setup - add commit-graph to the chain with bloom filters' '\n> +\ttest_commit c17 A/anotherFile3 &&\n> +\tgit commit-graph write --reachable --changed-paths --split &&\n> +\ttest_line_count = 3 .git/objects/info/commit-graphs/commit-graph-chain\n> +'\n> +\n> +test_bloom_filters_used_when_some_filters_are_missing() {\n> +\tlog_args=$1\n> +\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":3,\\\"zero_length_filter\\\":0,\\\"maybe\\\":6,\\\"definitely_not\\\":6\"\n\nPerhaps a better solution would be to use (enhanced) 'test-tool bloom'\nto check which commits have Bloom filters and which do not.\n\n> +\tsetup \"$log_args\"\n> +\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && test_cmp log_wo_bloom log_w_bloom\n> +}\n\nWhy broken && chain between setup() and the resr, and why && is not\nfollowed by line break (as before)?\n\n> +\n> +test_expect_success 'git log uses bloom filters if they exist in the latest but not all commit graphs in the chain.' '\n> +\ttest_bloom_filters_used_when_some_filters_are_missing \"-- A/B\"\n> +'\n> +\n> +test_done\n\nAll right... though the description of this test is a bit long.\n\n\nThank you for your work on this series.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"392287","messageId":"86lfovbhvx.fsf@gmail.com","threadId":"52499","inReplyTo":"e1b076a714d611e59d3d71c89221e41a3427fae4.1580943390.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 11/11] commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-22T00:11:30Z","receivedAt":"2020-02-22T00:11:40Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Garima Singh <garima.singh@microsoft.com>\n>\n> Add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag to the test setup suite\n> in order to toggle writing Bloom filters when running any of the git tests.\n> If set to true, we will compute and write Bloom filters every time a test\n> calls `git commit-graph write`, as if the `--changed-paths` option was\n> passed in.\n>\n> The test suite passes when GIT_TEST_COMMIT_GRAPH and\n> GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS are enabled.\n\nAll right.  Nice.\n\n>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  builtin/commit-graph.c        | 3 ++-\n>  ci/run-build-and-tests.sh     | 1 +\n>  commit-graph.h                | 1 +\n>  t/README                      | 5 +++++\n>  t/t4216-log-bloom.sh          | 3 +++\n>  t/t5318-commit-graph.sh       | 2 ++\n>  t/t5324-split-commit-graph.sh | 1 +\n>  7 files changed, 15 insertions(+), 1 deletion(-)\n>\n> diff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\n> index 261dcce091..fc9b234ab0 100644\n> --- a/builtin/commit-graph.c\n> +++ b/builtin/commit-graph.c\n> @@ -146,7 +146,8 @@ static int graph_write(int argc, const char **argv)\n>  \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n>  \tif (opts.progress)\n>  \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n> -\tif (opts.enable_changed_paths)\n> +\tif (opts.enable_changed_paths ||\n> +\t    git_env_bool(GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS, 0))\n>  \t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n>\n\nLooks good to me.\n\n>  \tread_replace_refs = 0;\n> diff --git a/ci/run-build-and-tests.sh b/ci/run-build-and-tests.sh\n> index ff0ef7f08e..7b4857651d 100755\n> --- a/ci/run-build-and-tests.sh\n> +++ b/ci/run-build-and-tests.sh\n> @@ -19,6 +19,7 @@ linux-gcc)\n>  \texport GIT_TEST_OE_SIZE=10\n>  \texport GIT_TEST_OE_DELTA_SIZE=5\n>  \texport GIT_TEST_COMMIT_GRAPH=1\n> +\texport GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=1\n>  \texport GIT_TEST_MULTI_PACK_INDEX=1\n>  \tmake test\n>  \t;;\n\nOK, include in continuous integration.\n\n> diff --git a/commit-graph.h b/commit-graph.h\n> index 25fefefb3e..4c202ff3d7 100644\n> --- a/commit-graph.h\n> +++ b/commit-graph.h\n> @@ -8,6 +8,7 @@\n>  \n>  #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n>  #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n> +#define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n>\n\nLooks good to me.\n\n>  struct commit;\n>  struct bloom_filter_settings;\n> diff --git a/t/README b/t/README\n> index caa125ba9a..be2f7d7fd2 100644\n> --- a/t/README\n> +++ b/t/README\n> @@ -378,6 +378,11 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n>  be written after every 'git commit' command, and overrides the\n>  'core.commitGraph' setting to true.\n>  \n> +GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n> +commit-graph write to compute and write changed path Bloom filters for\n> +every 'git commit-graph write', as if the `--changed-paths` option was\n> +passed in.\n> +\n\nGood, it is documented in README for tests.\n\n>  GIT_TEST_FSMONITOR=$PWD/t7519/fsmonitor-all exercises the fsmonitor\n>  code path for utilizing a file system monitor to speed up detecting\n>  new or changed files.\n> diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> index 19eca1864b..7acebb3962 100755\n> --- a/t/t4216-log-bloom.sh\n> +++ b/t/t4216-log-bloom.sh\n> @@ -3,6 +3,9 @@\n>  test_description='git log for a path with bloom filters'\n>  . ./test-lib.sh\n>  \n> +GIT_TEST_COMMIT_GRAPH=0\n> +GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n> +\n\nAll right, we need to ensure that 'git commit-graph write' is not run\nautomatically, otherwise split / incremental commit-graph tests would\nnot work.\n\nWe also need to ensure that '--changed-paths' is not added\nautomatically, so that we can test that commit-graph does not include\nBloom filters chunks if not requested.\n\n>  test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n>  \tgit init &&\n>  \tmkdir A A/B A/B/C &&\n> diff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\n> index 3f03de6018..973020be2d 100755\n> --- a/t/t5318-commit-graph.sh\n> +++ b/t/t5318-commit-graph.sh\n> @@ -3,6 +3,8 @@\n>  test_description='commit graph'\n>  . ./test-lib.sh\n>  \n> +GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n> +\n\nOK, otherwise it would screw up checking the content of commit-graph\nwith 'test-tool read-graph'.\n\n>  test_expect_success 'setup full repo' '\n>  \tmkdir full &&\n>  \tcd \"$TRASH_DIRECTORY/full\" &&\n> diff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\n> index c24823431f..9235db4561 100755\n> --- a/t/t5324-split-commit-graph.sh\n> +++ b/t/t5324-split-commit-graph.sh\n> @@ -4,6 +4,7 @@ test_description='split commit graph'\n>  . ./test-lib.sh\n>  \n>  GIT_TEST_COMMIT_GRAPH=0\n> +GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n>  \n>  test_expect_success 'setup repo' '\n>  \tgit init &&\n\nSame here.\n\nLooks good to me.\n-- \nJakub Narębski\n"},{"id":"392288","messageId":"a6c08b27-18d7-cf30-c076-3f6451a21519@gmail.com","threadId":"52499","inReplyTo":"86tv3qqxyn.fsf@gmail.com","subject":"Re: [PATCH v2 02/11] bloom: core Bloom filter implementation for changed paths","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-22T00:32:03Z","receivedAt":"2020-02-22T00:32:09Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/16/2020 11:49 AM, Jakub Narebski wrote:\n>> From: Garima Singh <garima.singh@microsoft.com>\n>>\n>> Add the core Bloom filter logic for computing the paths changed between a\n>> commit and its first parent. For details on what Bloom filters are and how they\n>> work, please refer to Dr. Derrick Stolee's blog post [1]. It provides a concise\n>> explaination of the adoption of Bloom filters as described in [2] and [3].\n>                                                                            ^^- to add\n\nNot sure what this means. Can you please clarify. \n\n>> 1. We currently use 7 and 10 for the number of hashes and the size of each\n>>    entry respectively. They served as great starting values, the mathematical\n>>    details behind this choice are described in [1] and [4]. The implementation,\n>                                                                                 ^^- to add\n\nNot sure what this means. Can you please clarify.\n\n>> 3. The filters are sized according to the number of changes in the each commit,\n>>    with minimum size of one 64 bit word.\n> \n> If I understand it correctly (but which might not be entirely clear),\n> the filter size in bits is the number of changes^* times 10, rounded up\n> to the nearest multiple of 64.\n> \n> [*] where the number of changes is the number of changed files (new blob\n> objects) _and_ the number of changed directories (new tree objects,\n> excluding root tree object change).\n> \n\nYes. \n\n> The interesting corner case, which might be worth specifying explicitly,\n> is what happens in the case there are _no changes_ with respect to first\n> parent (which can happen with either commit created with `git commit\n> --allow-empty`, or merge created e.g. with `git merge --strategy=ours`).\n> Is this case represented as Bloom filter of length 0, or as a Bloom\n> filter of length  of one 64-bit word which is minimal length composed of\n> all 0's (0x0000000000000000)?\n> \n\nSee t0095-bloom.sh: The filter for a commit with no changes is of length 0.\nI will call it out specifically in the appropriate commit message as well. \n\n>>\n>> [1] https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n> \n> I would write it in full, similar to subsequent bibliographical entries,\n> that is:\n> \n>   [1] Derrick Stolee\n>       \"Supercharging the Git Commit Graph IV: Bloom Filters\"\n>       https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n> \n> But that is just a matter of style.\n> \n\nSounds good. Will do. \n\n>>\n>> [4] Thomas Mueller Graf, Daniel Lemire\n>>     \"Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters\"\n>>     https://arxiv.org/abs/1912.08258\n>>\n>> [5] https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n>>\n>> Helped-by: Jeff King <peff@peff.net>\n>> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n>> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n>> ---\n>>  Makefile              |   2 +\n>>  bloom.c               | 228 ++++++++++++++++++++++++++++++++++++++++++\n>>  bloom.h               |  56 +++++++++++\n>>  t/helper/test-bloom.c |  84 ++++++++++++++++\n>>  t/helper/test-tool.c  |   1 +\n>>  t/helper/test-tool.h  |   1 +\n>>  t/t0095-bloom.sh      | 113 +++++++++++++++++++++\n>>  7 files changed, 485 insertions(+)\n>>  create mode 100644 bloom.c\n>>  create mode 100644 bloom.h\n>>  create mode 100644 t/helper/test-bloom.c\n>>  create mode 100755 t/t0095-bloom.sh\n> \n> As I wrote earlier, In my opinion this patch could be split into three\n> individual single-functionality pieces, to make it easier to review and\n> aid in bisectability if needed.\n> \n\nDoing this in v3. \n\n\n>> +\n>> +static uint32_t rotate_right(uint32_t value, int32_t count)\n>> +{\n>> +\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n>> +\tcount &= mask;\n>> +\treturn ((value >> count) | (value << ((-count) & mask)));\n>> +}\n> \n> Hmmm... both the algoritm on Wikipedia, and reference implementation use\n> rotate *left*, not rotate *right* in the implementation of Murmur3 hash,\n> see\n> \n>   https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n>   https://github.com/aappleby/smhasher/blob/master/src/MurmurHash3.cpp#L23\n> \n> \n> inline uint32_t rotl32 ( uint32_t x, int8_t r )\n> {\n>   return (x << r) | (x >> (32 - r));\n> }\n> \n\nThanks! Fixed this in v3. More on it later. \n\n>> +\n>> +/*\n>> + * Calculate a hash value for the given data using the given seed.\n>> + * Produces a uniformly distributed hash value.\n>> + * Not considered to be cryptographically secure.\n>> + * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n>> + **/\n>     ^^-- why two _trailing_ asterisks?\n> \n\nOops. Fixed. \n\n>> +static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n> \n> In short, I think that the name of the function should be murmur3_32, or\n> murmurhash3_32, or possibly murmur3_32_seed, or something like that.\n>\n\nRenamed it to murmur3_seeded in v3. The input and output types in the \nsignature make it clear that it is 32-bit version.\n \n>> +{\n>> +\tconst uint32_t c1 = 0xcc9e2d51;\n>> +\tconst uint32_t c2 = 0x1b873593;\n>> +\tconst uint32_t r1 = 15;\n>> +\tconst uint32_t r2 = 13;\n>> +\tconst uint32_t m = 5;\n>> +\tconst uint32_t n = 0xe6546b64;\n>> +\tint i;\n>> +\tuint32_t k1 = 0;\n>> +\tconst char *tail;\n>> +\n>> +\tint len4 = len / sizeof(uint32_t);\n>> +\n>> +\tconst uint32_t *blocks = (const uint32_t*)data;\n>> +\n>> +\tuint32_t k;\n>> +\tfor (i = 0; i < len4; i++)\n>> +\t{\n>> +\t\tk = blocks[i];\n> \n> IMPORTANT: There is a comment around there in the example implementation\n> in C on Wikipedia that this operation above is a source of differing\n> results across endianness.  \n\nThanks! SZEDER found this on his CI pipeline and we have fixed it to \nprocess the data in 1 byte words to avoid hitting any endian-ness issues. \nSee this part of the thread that carries the fix and the related discussion. \n  https://lore.kernel.org/git/ba856e20-0a3c-e2d2-6744-b9abfacdc465@gmail.com/\nI will be squashing those changes in appropriately in v3.  \n \n>> +\t\tk1 *= c2;\n>> +\t\tseed ^= k1;\n>> +\t\tbreak;\n>> +\t}\n>> +\n>> +\tseed ^= (uint32_t)len;\n>> +\tseed ^= (seed >> 16);\n>> +\tseed *= 0x85ebca6b;\n>> +\tseed ^= (seed >> 13);\n>> +\tseed *= 0xc2b2ae35;\n>> +\tseed ^= (seed >> 16);\n>> +\n>> +\treturn seed;\n>> +}\n> \n> In https://public-inbox.org/git/ba856e20-0a3c-e2d2-6744-b9abfacdc465@gmail.com/\n> you posted \"[PATCH] Process bloom filter data as 1 byte words\".\n> This may avoid the Big-endian vs Little-endian confusion,\n> that is wrong results on Big-endian architectures, but\n> it also may slow down the algorithm.\n> \n\nOh cool! You have seen that patch. And yes, we understand that it might add \na little overhead but at this point it is more important to be correct on all\narchitectures instead of micro-optimizing and introducing different \nimplementations for Little-endian and Big-endian. This would make this \nseries overly complicated. Optimizing the hashing techniques would deserve a\nseries of its own, which we can definitely revisit later.\n\n> The public domain implementation in PMurHash.c in SMHasher\n> (re)implementation in Chromium (see URL above) fall backs to 1-byte\n> operations only if it doesn't know the endianness (or if it is neither\n> little-endian, nor big-endian, i.e. middle-endian or mixed-endian --\n> though I doubt that Git works correctly on mixed-endian anyway).\n> \n> \n> Sidenote: it looks like the current implementation if Murmur hash in\n> Cromium uses MurmurHash3_x86_32, i.e. little-endian unaligned-safe\n> implementation, but prepares data by swapping with StringToLE32\n> https://github.com/chromium/chromium/blob/master/components/variations/variations_murmur_hash.h\n> \n> \n> Assuming that the terminating NUL (\"\\0\") character of a c-string is not\n> included in hash calculations, then murmur3_x86_32 hash has the\n> following results (all results are for seed equal 0):\n> \n> ''               -> 0x00000000\n> ' '              -> 0x7ef49b98\n> 'Hello world!'   -> 0x627b0c2c\n> 'The quick brown fox jumps over the lazy dog'   -> 0x2e4ff723\n> \n> C source (from Wikipedia): https://godbolt.org/z/ofa2p8\n> C++ source (Appleby's):    https://godbolt.org/z/BoSt6V\n> \n> The implementation provided in this patch, with rotate_right (instead of\n> rotate_left) gives, on little-endian machine, different results:\n> \n> ''               -> 0x00000000\n> ' '              -> 0xd1f27e64\n> 'Hello world!'   -> 0xa0791ad7\n> 'The quick brown fox jumps over the lazy dog'   -> 0x99f1676c\n> \n> https://github.com/gitgitgadget/git/blob/e1b076a714d611e59d3d71c89221e41a3427fae4/bloom.c#L21\n> C source (via GitGitGadget): https://godbolt.org/z/R9s8Tt\n> \n\nThanks! This is an excellent catch! Fixing the rotate_right to rotate_left, \ngives us the same answers as the two implementations you pointed out. I have\nadded the appropriate unit tests in v3 and they match the values you obtained \nfrom the other implementations. Thanks a lot for the rigor! \n\nWe based our implementation on the pseudo code and not on the sample code \npresented here: https://en.wikipedia.org/wiki/MurmurHash#Algorithm\nWe just didn't parse the ROL instruction correctly. \n\n>> +\n>> +void load_bloom_filters(void)\n>> +{\n>> +\tinit_bloom_filter_slab(&bloom_filters);\n>> +}\n> \n> \n> Actually this function doesn't load anything.  Perhaps it should be\n> named init_bloom_filters() or init_bloom_filters_storage(), or\n> bloom_filters_init()?\n>\n\nChanged to init_bloom_filters() in v3. Thanks! \n \n>> +\n>> +void fill_bloom_key(const char *data,\n>> +\t\t\t\t\tint len,\n>> +\t\t\t\t\tstruct bloom_key *key,\n>> +\t\t\t\t\tstruct bloom_filter_settings *settings)\n> \n> The last parameter could be of 'const bloom_filter_settings *' type.\n> \n\nDone. \n\n>> +{\n>> +\tint i;\n>> +\tconst uint32_t seed0 = 0x293ae76f;\n>> +\tconst uint32_t seed1 = 0x7e646e2c;\n> \n> Where did those seeds values came from?\n> \n\nThose values were chosen randomly. They will be fixed constants for the \ncurrent hashing version. I will add a note calling this out in the \nappropriate commit messages and the Documentation in v3. \n\n>> +\tconst uint32_t hash0 = seed_murmur3(seed0, data, len);\n>> +\tconst uint32_t hash1 = seed_murmur3(seed1, data, len);\n>> +\n>> +\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n>> +\tfor (i = 0; i < settings->num_hashes; i++)\n>> +\t\tkey->hashes[i] = hash0 + i * hash1;\n> \n> Note that in [3] authors say that double hashing technique has some\n> problems.  For one, we should ensure that hash1 is not zero, and even\n> better that it is odd (which makes it relatively prime to filter size\n> which is multiple of 64).  It also suffers from something called\n> \"approximate fingerprint collisions\".\n> \n> That is why the define \"enhanced double hashing\" technique, which does\n> not suffer from those problems (Algorithm 2, page 11/15).\n> \n>   +\tfor (i = 0; i < settings->num_hashes; i++) {\n>   +\t\tkey->hashes[i] = hash0;\n>   +\n>   +\t\thash0 = hash0 + hash1;\n>   +\t\thash1 = hash1 + i;\n>   +\t}\n> \n> This can also be written in closed form, based on equation (6)\n> \n>   +\tfor (i = 0; i < settings->num_hashes; i++)\n>   +\t\tkey->hashes[i] = hash0 + i * hash1 + i*(i*i - 1)/6;\n> \n> \n> In later paper [6] the closed form for \"enhanced double hashing\"\n> (p. 188) is slightly modified (or rather they use different variant of\n> this technique):\n> \n>   +\tfor (i = 0; i < settings->num_hashes; i++)\n>   +\t\tkey->hashes[i] = hash0 + i * hash1 + i*i;\n> \n> This is a variant of more generic \"enhanced double hashing\", section\n> 5.2 (Enhanced) Double Hashing Schemes (page 199):\n> \n>         h_1(u) + i h_2(u) + f(i)    mod m\n> \n> with f(i) = i^2 = i*i.\n> \n> They have tested that enhanced double hashing with both f(i) equal i*i\n> and equal i*i*i, and triple hashing technique, and they have found that\n> it performs slightly better than straight double hashing technique\n> (Fig. 1, page 212, section 3).\n> \n\nThanks for the detailed research here! The hash becoming zero and the \napproximate fingerprint collision are both extremely rare situations. In both\ncases, we would just see git log having to diff more trees than if it didn't \noccur. While these techniques would be great optimizations to do, especially\nif this implementation gets pulled into more generic hashing applications\nin the code, we think that for the purposes of the current series - it is not \nworth it. I say this because Azure Repos has been using this exact hashing \ntechnique for several years now without any glitches. And we think it would\nbe great to rely on this battle tested strategy in atleast the first version\nof this feature. \n\n>> +}\n>> +\n>> +void add_key_to_filter(struct bloom_key *key,\n>> +\t\t\t\t\t   struct bloom_filter *filter,\n>> +\t\t\t\t\t   struct bloom_filter_settings *settings)\n> \n> Here again the 'settings' argument can be const (as can the 'key'\n> parameter).\n> \n\nDone. \n\n>> +\n>> +struct bloom_filter *get_bloom_filter(struct repository *r,\n>> +\t\t\t\t      struct commit *c)\n>> +{\n>> +\tstruct bloom_filter *filter;\n>> +\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n>> +\tint i;\n>> +\tstruct diff_options diffopt;\n>> +\n>> +\tif (!bloom_filters.slab_size)\n>> +\t\treturn NULL;\n> \n> This is testing that commit slab for per-commit Bloom filters is\n> initialized, isn't it?\n> \n> First, should we write the condition as\n> \n> \tif (!bloom_filters.slab_size)\n> \n> or would the following be more readable\n> \n> \tif (bloom_filters.slab_size == 0)\n> \n\nSure. Switched to `if (bloom_filter.slab_size == 0)` in v3. \n\n> Second, should we return NULL, or should we just initialize the slab?\n> Or is non-existence of slab treated as a signal that the Bloom filters\n> mechanism is turned off?\n> \n\nYes. We purposefully choose to return NULL and ignore the mechanism \noverall because we use Bloom filters best effort only. \n\n>> +\n>> +\tif (diff_queued_diff.nr <= 512) {\n>\n> Second, there is a minor issue that diff_queue_struct.nr stores the\n> number of filepairs, that is the number of changed files, while the\n> number of elements added to Bloom filter is number of changed blobs and\n> trees.  For example if the following files are changed:\n> \n>   sub/dir/file1\n>   sub/file2\n> \n> then diff_queued_diff.nr is 2, but number of elements to be added to\n> Bloom filter is 4.\n> \n>   sub/dir/file1\n>   sub/file2\n>   sub/dir/\n>   sub/\n> \n> I'm not sure if it matters in practice.\n> \n\nIt does not matter much in practice, since the directories usually tend\nto collapse across the changes. Still, I will add another limit after \ncreating the hashmap entries to cap at 640 so that we have a maximum of \n100 changes in the bloom filter. \n\nWe plan to make these values configurable later. \n\n>> +\t\tstruct hashmap pathmap;\n>> +\t\tstruct pathmap_hash_entry* e;\n>> +\t\tstruct hashmap_iter iter;\n>> +\t\thashmap_init(&pathmap, NULL, NULL, 0);\n> \n> Stylistic issue: I have just noticed that here (and in some other\n> places), but not in all cases, you declare pointer types with asterisk\n> cuddled to type name, not to variable name, which contradicts\n> CodingGuidelines\n\nThanks for noticing that! Fixed all of these in v3. \n\n>> +\n>> +\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n>> +\t\t\tconst char* path = diff_queued_diff.queue[i]->two->path;\n> \n> Is that correct that we consider only post-image name for storing\n> changes in Bloom filter?  Currently if file was renamed (or deleted), it\n> is considered changed, and `git log -- <old-name>` lists commit that\n> changed file name too.\n>\n\nThe tests in t4216-log-bloom.sh ensure that the output of `git log -- <oldname>` \nremains unchanged for renamed and deleted files, when using bloom filters. \nI realize that I fat fingered over checking the old name, and didn't have an \nexplicit deleted file in the test. I have added them in v3, and the tests pass. \nSo the behavior is preserved and as expected when using Bloom filters. \nThanks for paying close attention! \n \n>> +\t\t\tconst char* p = path;\n> \n> It should be \"const char *\" for both.\n> \n>> +\n>> +\t\t\t/*\n>> +\t\t\t* Add each leading directory of the changed file, i.e. for\n>> +\t\t\t* 'dir/subdir/file' add 'dir' and 'dir/subdir' as well, so\n>> +\t\t\t* the Bloom filter could be used to speed up commands like\n>> +\t\t\t* 'git log dir/subdir', too.\n>> +\t\t\t*\n>> +\t\t\t* Note that directories are added without the trailing '/'.\n>> +\t\t\t*/\n>> +\t\t\tdo {\n>> +\t\t\t\tchar* last_slash = strrchr(p, '/');\n>> +\n>> +\t\t\t\tFLEX_ALLOC_STR(e, path, path);\n> \n> Here first 'path' is the field name, i.e. pathmap_hash_entry.path,\n> second 'path' is the name of local variable, aliased also to 'p'.\n> \n>> +\t\t\t\thashmap_entry_init(&e->entry, strhash(p));\n> \n> I don't know why both 'path' and 'p' are used, while both point to the\n> same memory (and thus have the same contents).  It is a bit confusing.\n> See also my previous comment.\n> \n\nCleaned up in v3. Thanks! \n>> +\t\tfilter->data = NULL;\n>> +\t\tfilter->len = 0;\n> \n> This needs to be explicitly stated both in the commit message and in the\n> API documentation (in comments) that bloom_filter.len == 0 means \"no\n> data\", while \"no changes\" is represented as bloom_filter with len == 1\n> and *data == (uint64_t)0;\n> \n> EDIT: actually \"no changes\" is also represented as bloom_filter with len\n> equal 0, as it turns out.\n> \n> One possible alternative could be representing \"no data\" value with\n> Bloom filter of length 1 and all 64 bits set to 1, and \"no changes\"\n> represented as filter of length 0.  This is not unambiguous choice!\n>\n\nThere is no gain in distinguishing between the absence of a filter and\na commit having no changes. The effect on `git log -- path` is the same in \nboth cases. We fall back to the normal diffing algorithm in revision.c.\nI will make this clearer in the appropriate commit messages and in the \nDocumentation in v3. \n \n>> +}\n>> diff --git a/bloom.h b/bloom.h\n>> new file mode 100644\n>> index 0000000000..7f40c751f7\n>> --- /dev/null\n>> +++ b/bloom.h\n>> @@ -0,0 +1,56 @@\n>> +#ifndef BLOOM_H\n>> +#define BLOOM_H\n> \n> Should we #include the stdint.h header for uint32_t and uint64_t types?\n> \n\ngit-compat-util.h takes care of this. \n\n>> +\n>> +struct commit;\n>> +struct repository;\n>> +struct commit_graph;\n>> +\n> \n> Perhaps we should add block comment for this struct, like there is one\n> for struct bloom_filter below.\n> \n\nDone in v3.\n\n>> +struct bloom_filter_settings {\n>> +\tuint32_t hash_version;\n>> +\tuint32_t num_hashes;\n>> +\tuint32_t bits_per_entry;\n> \n> I guess that the type uint32_t was chosen to make it easier to store\n> this information and later retrieve it from the commit-graph file, isn't\n> it?  Otherwise those types are much too large for sensible range of\n> values (which would all fit in 8-bits byte).\n> \n\nYes.\n\n>> +\n>> +/*\n>> + * A bloom_key represents the k hash values for a\n>> + * given hash input. These can be precomputed and\n>> + * stored in a bloom_key for re-use when testing\n>> + * against a bloom_filter.\n> \n> We might want to add that the number of hash values is given by Bloom\n> filter settings, and it is assumed to be the same for all bloom_key\n> variables / objects.\n> \n\nIncorporated in v3. \n\n>> +. ./test-lib.sh\n>> +\n>> +test_expect_success 'get bloom filters for commit with no changes' '\n>> +\tgit init &&\n>> +\tgit commit --allow-empty -m \"c0\" &&\n>> +\tcat >expect <<-\\EOF &&\n>> +\tFilter_Length:0\n>> +\tFilter_Data:\n>> +\tEOF\n>> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n>> +\ttest_cmp expect actual\n>> +'\n> \n> A few things.  First, I wonder why we need to provide object ID;\n> couldn't 'test-tool bloom get_filter_for_commit' parse commit-ish\n> argument, or would it make it too complicated for no reason?\n> \n\nYes it was overkill for what I need in the test. \n\n>> +\n>> +test_expect_success 'get bloom filter for commit with 10 changes' '\n>> +\trm actual &&\n>> +\trm expect &&\n>> +\tmkdir smallDir &&\n>> +\tfor i in $(test_seq 0 9)\n>> +\tdo\n>> +\t\techo $i >smallDir/$i\n>> +\tdone &&\n>> +\tgit add smallDir &&\n>> +\tgit commit -m \"commit with 10 changes\" &&\n>> +\tcat >expect <<-\\EOF &&\n>> +\tFilter_Length:4\n>> +\tFilter_Data:508928809087080a|8a7648210804001|4089824400951000|841ab310098051a8|\n>> +\tEOF\n>> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n>> +\ttest_cmp expect actual\n>> +'\n> \n> This test is in my opinion fragile, as it unnecessarily test the\n> implementation details instead of the functionality provided.  If we\n> change the hashing scheme (for example going from double hashing to some\n> variant of enhanced double hashing), or change the base hash function\n> (for example from Murmur3_32 to xxHash_64), or change the number of hash\n> functions (perhaps because changing of number of bits per element, and\n> thus optimal number of hash functions from 7 to 6), or change from\n> 64-bit word blocks to 32-bit word blocks, the test would have to be\n> changed.\n> \n\nRegarding this and the rest of you comments on t0095-log-bloom.sh:\n\nI am tweaking it as necessary but the entire point of these tests is to\nbreak for the things you called out. They need to be intricately tied\nto the current hashing strategy and are hence intended to be fragile so \nas to catch any subtle or accidental changes in the hashing computation. \nAny change like the ones you have called out would require a hash version\nchange and all the compatibility reactions that come with it. \n\nI have added more tests around the murmur3_seeded method in v3. Removed\nsome of the redundant ones. \n\nThe other more evolved test cases you call out are covered in the e2e\nintegration tests in t4216-log-bloom.sh\n\n> \n> Reviewed-by: Jakub Narębski <jnareb@gmail.com>\n> \n> Thanks for working on this.\n> \n> Best,\n> \n\nThank you once again for an excellent and in-depth review of this patch! \nYou have helped make this code so much better!\n\nCheers! \nGarima Singh\n"},{"id":"392289","messageId":"189827fa-7fab-5064-7baa-f856b149559b@gmail.com","threadId":"52499","inReplyTo":"86h7zqqdze.fsf@gmail.com","subject":"Re: [PATCH v2 03/11] diff: halt tree-diff early after max_changes","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-22T00:37:30Z","receivedAt":"2020-02-22T00:37:33Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/16/2020 7:00 PM, Jakub Narebski wrote:\n> \"Derrick Stolee via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> From: Derrick Stolee <dstolee@microsoft.com>\n>>\n>> Use this max_changes option in the bloom filter calculations. This\n>> reduces the time taken to compute the filters for the Linux kernel\n>> repo from 2m50s to 2m35s. On a large internal repository with ~500\n>> commits that perform tree-wide changes, the time reduced from\n>> 6m15s to 3m48s.\n> \n> I wonder if there is some large open-source project with many commits\n> performing tree-wide changes, that is with many commits with more than\n> 512 changed files with respect to the first parent.\n> \n> Maybe https://github.com/whosonfirst-data/whosonfirst-data-venue-us-ny\n> from \"Top Ten Worst Repositories to host on GitHub - Git Merge 2017\"\n> could be a good repository to test ;-)\n> \n\nThanks for the suggestion! I will see if any of these repos gives us a \ngood test bed and add the perf improvement numbers in the appropriate\ncommit messages in v3. \n\nCheers! \nGarima Singh\n"},{"id":"392290","messageId":"0f1ab477-fae8-b744-5c48-87995f7fc8eb@gmail.com","threadId":"52499","inReplyTo":"86k14klvyb.fsf@gmail.com","subject":"Re: [PATCH v2 04/11] commit-graph: compute Bloom filters for changed paths","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-22T00:55:56Z","receivedAt":"2020-02-22T00:56:01Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/17/2020 4:56 PM, Jakub Narebski wrote:\n> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> From: Garima Singh <garima.singh@microsoft.com>\n>> Subject: [PATCH v2 04/11] commit-graph: compute Bloom filters for changed paths\n>>\n>> Compute Bloom filters for the paths that changed between a commit and its\n>> first parent using the implementation in bloom.c, when the\n>> COMMIT_GRAPH_WRITE_CHANGED_PATHS flag is set. This computation is done on a\n>> commit-by-commit basis. We will write these Bloom filters to the commit graph\n>> file in the next change.\n> \n> I have no major complaints about the contents of this patch (except lack\n> of test, and type of total_bloom_filter_data_size), but the commit\n> message could have been worded better.\n> \n> I would write something like this instead:\n> \n>   Add new COMMIT_GRAPH_WRITE_CHANGED_PATHS flag that makes Git compute\n>   Bloom filters that store the information about changed paths (that\n>   changed between a commit and its first parent) for each commit in the\n>   commit-graph.  This computation is done on a commit-by-commit basis.\n> \n>   We will write these Bloom filters to the commit-graph file, to store\n>   this data on disk, in the next change in this series.\n> \n> In my opinion the fact that we compute Bloom filters for each and every\n> commit in the commit-graph file is more important than quite obvious\n> fact that we use implementation from bloom.c.\n> \n\nNice! Incorporated in v3. Thanks!\n\n>>\n>> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n>> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n>> ---\n>>  commit-graph.c | 32 +++++++++++++++++++++++++++++++-\n>>  commit-graph.h |  3 ++-\n>>  2 files changed, 33 insertions(+), 2 deletions(-)\n> \n> It would be good to have at least sanity check of this feature, perhaps\n> one that would check that the number of per-commit Bloom filters on slab\n> matches the number of commits in the commit-graph.\n> \n\nThe combination of all the e2e tests in this series with the test\nflag being turned on in the CI, and the performance gains we are seeing\nconfirm that this is happening correctly.\n\n>>  \n>>  \tconst struct split_commit_graph_opts *split_opts;\n>> +\tuint32_t total_bloom_filter_data_size;\n> \n> This is total size of Bloom filters data, in bytes, that will later be\n> used for BDAT chunk size.  However the commit-graph format uses 8 bytes\n> for byte-offset, not 4 bytes.  Why it is uint32_t and not uint64_t then?\n>\n\nChanged to size_t. Thanks for noticing! \n \n>>  };\n>>  \n>>  static void write_graph_chunk_fanout(struct hashfile *f,\n>> @@ -1140,6 +1143,28 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>>  \tstop_progress(&ctx->progress);\n>>  }\n>>  \n>> +static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>> +{\n>> +\tint i;\n>> +\tstruct progress *progress = NULL;\n>> +\n>> +\tload_bloom_filters();\n>> +\n>> +\tif (ctx->report_progress)\n>> +\t\tprogress = start_progress(\n>> +\t\t\t_(\"Computing commit diff Bloom filters\"),\n>> +\t\t\tctx->commits.nr);\n>> +\n> \n> Shouldn't we initialize ctx->total_bloom_filter_data_size to 0 here?  We\n> cannot use compute_bloom_filters() to _update_ Bloom filters data, I\n> think -- we don't distinguish here between new and existing data (where\n> existing data size is already included in total Bloom filters size).  At\n> least I don't think so.\n> \n\nThis line in commit-graph.c takes care of reinitializing the graph context and\nby consequence the bloom filter data size. \n  ctx = xcalloc(1, sizeof(struct write_commit_graph_context));\n  \nSo the total size gets recalculated every time, which is correct. \n\n> \n> Side note: perhaps we could add trailing comma after new enum entry,\n> that is\n> \n>   +\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4),\n> \n> following new CodingGuidelines recommendation\n> \n\nThanks! Fixed in v3.\n\nCheers! \nGarima Singh\n"},{"id":"392291","messageId":"c3a222e7-a782-2c25-e809-8eda4ce418ce@gmail.com","threadId":"52499","inReplyTo":"CAGyf7-FzaG3Jb92JTx1QyADAoLhHCREyadVbTM2vZW-wxK4zEg@mail.gmail.com","subject":"Re: [PATCH v2 09/11] commit-graph: add --changed-paths option to write subcommand","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-22T01:44:35Z","receivedAt":"2020-02-22T01:44:41Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/20/2020 5:10 PM, Bryan Turner wrote:\n> On Wed, Feb 5, 2020 at 2:56 PM Garima Singh via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n>> diff --git a/Documentation/git-commit-graph.txt b/Documentation/git-commit-graph.txt\n>> index bcd85c1976..907d703b30 100644\n>> --- a/Documentation/git-commit-graph.txt\n>> +++ b/Documentation/git-commit-graph.txt\n>> @@ -54,6 +54,11 @@ or `--stdin-packs`.)\n>>  With the `--append` option, include all commits that are present in the\n>>  existing commit-graph file.\n>>  +\n>> +With the `--changed-paths` option, compute and write information about the\n>> +paths changed between a commit and it's first parent. This operation can\n> \n> \"its first parent\"\n> \n> (Pardon the grammar nit from the peanut gallery!)\n> \n\n:)\nThank you! Fixed in v3. \n\nCheers! \nGarima Singh\n"},{"id":"392355","messageId":"86wo8d8lus.fsf@gmail.com","threadId":"52499","inReplyTo":"a6c08b27-18d7-cf30-c076-3f6451a21519@gmail.com","subject":"Re: [PATCH v2 02/11] bloom: core Bloom filter implementation for changed paths","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-23T13:38:35Z","receivedAt":"2020-02-23T13:38:47Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Garima Singh <garimasigit@gmail.com> writes:\n> On 2/16/2020 11:49 AM, Jakub Narebski wrote:\n>>> From: Garima Singh <garima.singh@microsoft.com>\n>>>\n>>> Add the core Bloom filter logic for computing the paths changed between a\n>>> commit and its first parent. For details on what Bloom filters are and how they\n>>> work, please refer to Dr. Derrick Stolee's blog post [1]. It provides a concise\n>>> explaination of the adoption of Bloom filters as described in [2] and [3].\n>>                                                                           ^^- to add\n>\n> Not sure what this means. Can you please clarify. \n>\n>>> 1. We currently use 7 and 10 for the number of hashes and the size of each\n>>>    entry respectively. They served as great starting values, the mathematical\n>>>    details behind this choice are described in [1] and [4]. The implementation,\n>>                                                                                ^^- to add\n>\n> Not sure what this means. Can you please clarify.\n\nI'm sorry for not being clear.  What I wanted to say that in both cases\nthe last line should have ended in either full stop in first case, or\ncomma in second case:\n\n  \"as described in [2] and [3].\"\n\n  \"The implementation,\"\n\nWhat I wrote (trying to put the arrow below final fullstop or comma)\nonly works when one is using with fixed-width font.\n\n>>> 3. The filters are sized according to the number of changes in the each commit,\n>>>    with minimum size of one 64 bit word.\n\n[...]\n>> The interesting corner case, which might be worth specifying explicitly,\n>> is what happens in the case there are _no changes_ with respect to first\n>> parent (which can happen with either commit created with `git commit\n>> --allow-empty`, or merge created e.g. with `git merge --strategy=ours`).\n>> Is this case represented as Bloom filter of length 0, or as a Bloom\n>> filter of length  of one 64-bit word which is minimal length composed of\n>> all 0's (0x0000000000000000)?\n>> \n>\n> See t0095-bloom.sh: The filter for a commit with no changes is of length 0.\n> I will call it out specifically in the appropriate commit message as well. \n\nI have realized this only later that both \"no changes\" and \"no data\"\nuses filter of length 0; which works well because checking the diff if\nthere were no changes is cheap (both tree oids are the same).\n\n>>> ---\n>>>  Makefile              |   2 +\n>>>  bloom.c               | 228 ++++++++++++++++++++++++++++++++++++++++++\n>>>  bloom.h               |  56 +++++++++++\n>>>  t/helper/test-bloom.c |  84 ++++++++++++++++\n>>>  t/helper/test-tool.c  |   1 +\n>>>  t/helper/test-tool.h  |   1 +\n>>>  t/t0095-bloom.sh      | 113 +++++++++++++++++++++\n>>>  7 files changed, 485 insertions(+)\n>>>  create mode 100644 bloom.c\n>>>  create mode 100644 bloom.h\n>>>  create mode 100644 t/helper/test-bloom.c\n>>>  create mode 100755 t/t0095-bloom.sh\n>> \n>> As I wrote earlier, In my opinion this patch could be split into three\n>> individual single-functionality pieces, to make it easier to review and\n>> aid in bisectability if needed.\n>\n> Doing this in v3. \n\nThanks.  Though if it makes (much) more work for you, I can work with\nunsplit patch, no problem.\n\n>>> +\n>>> +static uint32_t rotate_right(uint32_t value, int32_t count)\n>>> +{\n>>> +\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n>>> +\tcount &= mask;\n>>> +\treturn ((value >> count) | (value << ((-count) & mask)));\n>>> +}\n>> \n>> Hmmm... both the algoritm on Wikipedia, and reference implementation use\n>> rotate *left*, not rotate *right* in the implementation of Murmur3 hash,\n>> see\n>> \n>>   https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n>>   https://github.com/aappleby/smhasher/blob/master/src/MurmurHash3.cpp#L23\n>> \n>> \n>> inline uint32_t rotl32 ( uint32_t x, int8_t r )\n>> {\n>>   return (x << r) | (x >> (32 - r));\n>> }\n>\n> Thanks! Fixed this in v3. More on it later. \n\nSidenote: If I understand it correctly Bloom filters functionality is\nincluded in Scalar [1].  What will happen then with all those Bloom\nfilter chunks in commit-graph files with wrong hash functions?\n\n[1]: https://devblogs.microsoft.com/devops/introducing-scalar/\n\n>>> +\n>>> +/*\n>>> + * Calculate a hash value for the given data using the given seed.\n>>> + * Produces a uniformly distributed hash value.\n>>> + * Not considered to be cryptographically secure.\n>>> + * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n>>> + **/\n>>     ^^-- why two _trailing_ asterisks?\n>\n> Oops. Fixed. \n\nOften two _leading_ asterisks are used to mark commit as containing\ndocstring in some specific format, like Doxygen.  Two _trailing_\nasterisks looks like typo.\n\n>>> +static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n>> \n>> In short, I think that the name of the function should be murmur3_32, or\n>> murmurhash3_32, or possibly murmur3_32_seed, or something like that.\n>\n> Renamed it to murmur3_seeded in v3. The input and output types in the \n> signature make it clear that it is 32-bit version.\n\nAll right, I can agree with that.\n\n>>> +{\n>>> +\tconst uint32_t c1 = 0xcc9e2d51;\n>>> +\tconst uint32_t c2 = 0x1b873593;\n>>> +\tconst uint32_t r1 = 15;\n>>> +\tconst uint32_t r2 = 13;\n>>> +\tconst uint32_t m = 5;\n>>> +\tconst uint32_t n = 0xe6546b64;\n>>> +\tint i;\n>>> +\tuint32_t k1 = 0;\n>>> +\tconst char *tail;\n>>> +\n>>> +\tint len4 = len / sizeof(uint32_t);\n>>> +\n>>> +\tconst uint32_t *blocks = (const uint32_t*)data;\n>>> +\n>>> +\tuint32_t k;\n>>> +\tfor (i = 0; i < len4; i++)\n>>> +\t{\n>>> +\t\tk = blocks[i];\n>> \n>> IMPORTANT: There is a comment around there in the example implementation\n>> in C on Wikipedia that this operation above is a source of differing\n>> results across endianness.  \n>\n> Thanks! SZEDER found this on his CI pipeline and we have fixed it to \n> process the data in 1 byte words to avoid hitting any endian-ness issues. \n> See this part of the thread that carries the fix and the related discussion. \n>   https://lore.kernel.org/git/ba856e20-0a3c-e2d2-6744-b9abfacdc465@gmail.com/\n> I will be squashing those changes in appropriately in v3.  \n\n[...]\n>>> +\t\tk1 *= c2;\n>>> +\t\tseed ^= k1;\n>>> +\t\tbreak;\n>>> +\t}\n>>> +\n>>> +\tseed ^= (uint32_t)len;\n>>> +\tseed ^= (seed >> 16);\n>>> +\tseed *= 0x85ebca6b;\n>>> +\tseed ^= (seed >> 13);\n>>> +\tseed *= 0xc2b2ae35;\n>>> +\tseed ^= (seed >> 16);\n>>> +\n>>> +\treturn seed;\n>>> +}\n>> \n>> In https://public-inbox.org/git/ba856e20-0a3c-e2d2-6744-b9abfacdc465@gmail.com/\n>> you posted \"[PATCH] Process bloom filter data as 1 byte words\".\n>> This may avoid the Big-endian vs Little-endian confusion,\n>> that is wrong results on Big-endian architectures, but\n>> it also may slow down the algorithm.\n>\n> Oh cool! You have seen that patch. And yes, we understand that it might add \n> a little overhead but at this point it is more important to be correct on all\n> architectures instead of micro-optimizing and introducing different \n> implementations for Little-endian and Big-endian. This would make this \n> series overly complicated. Optimizing the hashing techniques would deserve a\n> series of its own, which we can definitely revisit later.\n\nRight, \"first make it work, then make it right, and, finally, make it fast.\".\n\nAnyway, could you maybe compare performance of Git for old version\n(operating on 32-bit/4-bytes words) and new version (operating on 1-byte\nwords) file history operation with Bloom filters, to see if it matters\nor not?\n\n>> The public domain implementation in PMurHash.c in SMHasher\n>> (re)implementation in Chromium (see URL above) fall backs to 1-byte\n>> operations only if it doesn't know the endianness (or if it is neither\n>> little-endian, nor big-endian, i.e. middle-endian or mixed-endian --\n>> though I doubt that Git works correctly on mixed-endian anyway).\n>> \n>> \n>> Sidenote: it looks like the current implementation if Murmur hash in\n>> Chromium uses MurmurHash3_x86_32, i.e. little-endian unaligned-safe\n>> implementation, but prepares data by swapping with StringToLE32\n>> https://github.com/chromium/chromium/blob/master/components/variations/variations_murmur_hash.h\n\nThe solution in PMurHash.c in Chromium, and the pseudo-code algorithm on\nWikipedia do endian handling only for remaining bytes (while the\nsolution in Appleby's code [beginnings of], and in current\nabove-mentioned Chromium implementation do the conversion for all\nbytes).  I think that handling it only for remaining bytes (for data\nsizes not being multiply of 32-bits / 4-bytes) is enough; all other\noperations, that is multiply, rotate, xor and addition do not depend on\nendianness.\n\n>> Assuming that the terminating NUL (\"\\0\") character of a c-string is not\n>> included in hash calculations, then murmur3_x86_32 hash has the\n>> following results (all results are for seed equal 0):\n>> \n>> ''               -> 0x00000000\n>> ' '              -> 0x7ef49b98\n>> 'Hello world!'   -> 0x627b0c2c\n>> 'The quick brown fox jumps over the lazy dog'   -> 0x2e4ff723\n>> \n>> C source (from Wikipedia): https://godbolt.org/z/ofa2p8\n>> C++ source (Appleby's):    https://godbolt.org/z/BoSt6V\n>> \n>> The implementation provided in this patch, with rotate_right (instead of\n>> rotate_left) gives, on little-endian machine, different results:\n>> \n>> ''               -> 0x00000000\n>> ' '              -> 0xd1f27e64\n>> 'Hello world!'   -> 0xa0791ad7\n>> 'The quick brown fox jumps over the lazy dog'   -> 0x99f1676c\n>> \n>> https://github.com/gitgitgadget/git/blob/e1b076a714d611e59d3d71c89221e41a3427fae4/bloom.c#L21\n>> C source (via GitGitGadget): https://godbolt.org/z/R9s8Tt\n>> \n>\n> Thanks! This is an excellent catch! Fixing the rotate_right to rotate_left, \n> gives us the same answers as the two implementations you pointed out. I have\n> added the appropriate unit tests in v3 and they match the values you obtained \n> from the other implementations. Thanks a lot for the rigor! \n>\n> We based our implementation on the pseudo code and not on the sample code \n> presented here: https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n> We just didn't parse the ROL instruction correctly. \n\nAll right, that's good.\n\nNote that the pseudo code includes the following:\n\n    with any remainingBytesInKey do\n        remainingBytes ← SwapToLittleEndian(remainingBytesInKey)\n        // Note: Endian swapping is only necessary on big-endian machines.\n        //       The purpose is to place the meaningful digits towards the low end of the value,\n        //       so that these digits have the greatest potential to affect the low range digits\n        //       in the subsequent multiplication.  Consider that locating the meaningful digits\n        //       in the high range would produce a greater effect upon the high digits of the\n        //       multiplication, and notably, that such high digits are likely to be discarded\n        //       by the modulo arithmetic under overflow.  We don't want that.\n\n[...]\n>>> +{\n>>> +\tint i;\n>>> +\tconst uint32_t seed0 = 0x293ae76f;\n>>> +\tconst uint32_t seed1 = 0x7e646e2c;\n>> \n>> Where did those seeds values came from?\n>>\n>\n> Those values were chosen randomly. They will be fixed constants for the \n> current hashing version. I will add a note calling this out in the \n> appropriate commit messages and the Documentation in v3. \n\nNice to know.\n\nI wonder if those seed values should be relatively prime, and whether\nseed1 should be odd (from theoretical point of view).\n\n>>> +\tconst uint32_t hash0 = seed_murmur3(seed0, data, len);\n>>> +\tconst uint32_t hash1 = seed_murmur3(seed1, data, len);\n>>> +\n>>> +\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n>>> +\tfor (i = 0; i < settings->num_hashes; i++)\n>>> +\t\tkey->hashes[i] = hash0 + i * hash1;\n>> \n>> Note that in [3] authors say that double hashing technique has some\n>> problems.  For one, we should ensure that hash1 is not zero, and even\n>> better that it is odd (which makes it relatively prime to filter size\n>> which is multiple of 64).  It also suffers from something called\n>> \"approximate fingerprint collisions\".\n>> \n>> That is why the define \"enhanced double hashing\" technique, which does\n>> not suffer from those problems (Algorithm 2, page 11/15).\n>> \n>>   +\tfor (i = 0; i < settings->num_hashes; i++) {\n>>   +\t\tkey->hashes[i] = hash0;\n>>   +\n>>   +\t\thash0 = hash0 + hash1;\n>>   +\t\thash1 = hash1 + i;\n>>   +\t}\n>> \n>> This can also be written in closed form, based on equation (6)\n>> \n>>   +\tfor (i = 0; i < settings->num_hashes; i++)\n>>   +\t\tkey->hashes[i] = hash0 + i * hash1 + i*(i*i - 1)/6;\n>> \n>> \n>> In later paper [6] the closed form for \"enhanced double hashing\"\n>> (p. 188) is slightly modified (or rather they use different variant of\n>> this technique):\n>> \n>>   +\tfor (i = 0; i < settings->num_hashes; i++)\n>>   +\t\tkey->hashes[i] = hash0 + i * hash1 + i*i;\n>> \n>> This is a variant of more generic \"enhanced double hashing\", section\n>> 5.2 (Enhanced) Double Hashing Schemes (page 199):\n>> \n>>         h_1(u) + i h_2(u) + f(i)    mod m\n>> \n>> with f(i) = i^2 = i*i.\n>> \n>> They have tested that enhanced double hashing with both f(i) equal i*i\n>> and equal i*i*i, and triple hashing technique, and they have found that\n>> it performs slightly better than straight double hashing technique\n>> (Fig. 1, page 212, section 3).\n>> \n>\n> Thanks for the detailed research here! The hash becoming zero and the \n> approximate fingerprint collision are both extremely rare situations. In both\n> cases, we would just see `git log` having to diff more trees than if it didn't \n> occur. While these techniques would be great optimizations to do, especially\n> if this implementation gets pulled into more generic hashing applications\n> in the code, we think that for the purposes of the current series - it is not \n> worth it. I say this because Azure Repos has been using this exact hashing \n> technique for several years now without any glitches. And we think it would\n> be great to rely on this battle tested strategy in at least the first version\n> of this feature. \n\nAll right, that is a good strategy.\n\nI wonder if switching from double hashing to enhanced double hashing\n(for example the variant with i*i added) would bring any noticeable\nperformance improvements in Git operations (due to less false\npositives).\n\n>>> +\n>>> +struct bloom_filter *get_bloom_filter(struct repository *r,\n>>> +\t\t\t\t      struct commit *c)\n>>> +{\n>>> +\tstruct bloom_filter *filter;\n>>> +\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n>>> +\tint i;\n>>> +\tstruct diff_options diffopt;\n>>> +\n>>> +\tif (!bloom_filters.slab_size)\n>>> +\t\treturn NULL;\n>> \n>> This is testing that commit slab for per-commit Bloom filters is\n>> initialized, isn't it?\n>> \n>> First, should we write the condition as\n>> \n>> \tif (!bloom_filters.slab_size)\n>> \n>> or would the following be more readable\n>> \n>> \tif (bloom_filters.slab_size == 0)\n>> \n>\n> Sure. Switched to `if (bloom_filter.slab_size == 0)` in v3. \n\nThough either works, and the former looks more like the test if\nbloom_filters slab are initialized, now that I thought about it a bit.\nYour choice.\n\n>> Second, should we return NULL, or should we just initialize the slab?\n>> Or is non-existence of slab treated as a signal that the Bloom filters\n>> mechanism is turned off?\n>> \n>\n> Yes. We purposefully choose to return NULL and ignore the mechanism \n> overall because we use Bloom filters best effort only. \n\nAll right.\n\n>>> +\n>>> +\tif (diff_queued_diff.nr <= 512) {\n>>\n>> Second, there is a minor issue that diff_queue_struct.nr stores the\n>> number of filepairs, that is the number of changed files, while the\n>> number of elements added to Bloom filter is number of changed blobs and\n>> trees.  For example if the following files are changed:\n>> \n>>   sub/dir/file1\n>>   sub/file2\n>> \n>> then diff_queued_diff.nr is 2, but number of elements to be added to\n>> Bloom filter is 4.\n>> \n>>   sub/dir/file1\n>>   sub/file2\n>>   sub/dir/\n>>   sub/\n>> \n>> I'm not sure if it matters in practice.\n>> \n>\n> It does not matter much in practice, since the directories usually tend\n> to collapse across the changes. Still, I will add another limit after \n> creating the hashmap entries to cap at 640 so that we have a maximum of \n> 100 changes in the bloom filter. \n>\n> We plan to make these values configurable later. \n\nI'm not sure if it is truly necessary; we can treat limit on number of\nchanged paths as \"best effort\" limit on Bloom filter size.\n\nI just wanted to point out the difference.\n\n\nSide note: I wonder if it would be worth it (in the future) to change\nhandling commits with large amount of changes.  I was thinking about\nswitching to soft and hard limit: soft limit would be on the size of the\nBloom filter, that is if number of elements times bits per element is\ngreater that size threshold, we don't increase the size of the filter.\n\nThis would mean that the false positives ratio (the number of files that\nare not present but get answer \"maybe\" instead of \"no\" out of the\nfilter) would increase, so there would be a need for another hard limit\nwhere we decide that it is not worth it, and not store the data for the\nBloom filter -- current \"no data\" case with empty filter with length 0.\nThis hard limit can be imposed on number of changed files, or on number\nof paths added to filter, or on number of bits set to 1 in the filter\n(on popcount), or some combination thereof.\n\n[...]\n>>> +\n>>> +\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n>>> +\t\t\tconst char* path = diff_queued_diff.queue[i]->two->path;\n>> \n>> Is that correct that we consider only post-image name for storing\n>> changes in Bloom filter?  Currently if file was renamed (or deleted), it\n>> is considered changed, and `git log -- <old-name>` lists commit that\n>> changed file name too.\n>\n> The tests in t4216-log-bloom.sh ensure that the output of `git log -- <oldname>` \n> remains unchanged for renamed and deleted files, when using bloom filters. \n> I realize that I fat fingered over checking the old name, and didn't have an \n> explicit deleted file in the test. I have added them in v3, and the tests pass. \n> So the behavior is preserved and as expected when using Bloom filters. \n> Thanks for paying close attention! \n\nIt seems like it shouldn't be working, as we are not adding the old name\nto Bloom filter, but that only means that I misunderstood how\ndiff_tree_oid() works with default options.  It turns out that without\nexplicitly turning on rename detection it shows rename as deletion of\nold name and addition of new name -- so if tracking deletion works\ncorrectly, then tracking renames should work correctly.\n\nSo it is in fact correct, which as you said was confirmed by (improved)\ntests.  I think also that if there was a bug in handling renames in this\ncode it would have been detected when running CI with\nGIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS.\n\n[...]\n>>> +\t\tfilter->data = NULL;\n>>> +\t\tfilter->len = 0;\n>> \n>> This needs to be explicitly stated both in the commit message and in the\n>> API documentation (in comments) that bloom_filter.len == 0 means \"no\n>> data\", while \"no changes\" is represented as bloom_filter with len == 1\n>> and *data == (uint64_t)0;\n>> \n>> EDIT: actually \"no changes\" is also represented as bloom_filter with len\n>> equal 0, as it turns out.\n>> \n>> One possible alternative could be representing \"no data\" value with\n>> Bloom filter of length 1 and all 64 bits set to 1, and \"no changes\"\n>> represented as filter of length 0.  This is not unambiguous choice!\n>>\n>\n> There is no gain in distinguishing between the absence of a filter and\n> a commit having no changes. The effect on `git log -- path` is the same in \n> both cases. We fall back to the normal diffing algorithm in revision.c.\n> I will make this clearer in the appropriate commit messages and in the \n> Documentation in v3. \n\nYou are right, which I have realized only when reviewing subsequent\npatches in the series.\n\nIn the absence of a filter, the \"no data\" case, we need to fall back to\nexamining the diff anyway.\n\nIn the case of commit having no changes, the \"no changes\" case,\ncomputing the diff is cheap because Git can realize that both trees have\nthe same oid.  So we do not lose performance this way, and we avoid\nspecial-casing it (avoiding branching) when computing the Bloom filter,\nif the \"no change\" case was represented by filter of length 1 and all\nzero bits as data.  Comparing tree oids and matching first hash function\nin bloom_key against all zeros Bloom filter should be, I think, of\nsimilar performance.\n\n[...]\n>>> +. ./test-lib.sh\n>>> +\n>>> +test_expect_success 'get bloom filters for commit with no changes' '\n>>> +\tgit init &&\n>>> +\tgit commit --allow-empty -m \"c0\" &&\n>>> +\tcat >expect <<-\\EOF &&\n>>> +\tFilter_Length:0\n>>> +\tFilter_Data:\n>>> +\tEOF\n>>> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n>>> +\ttest_cmp expect actual\n>>> +'\n>> \n>> A few things.  First, I wonder why we need to provide object ID;\n>> couldn't 'test-tool bloom get_filter_for_commit' parse commit-ish\n>> argument, or would it make it too complicated for no reason? \n>\n> Yes it was overkill for what I need in the test. \n\nAll right, I agree with that.\n\n>>> +\n>>> +test_expect_success 'get bloom filter for commit with 10 changes' '\n>>> +\trm actual &&\n>>> +\trm expect &&\n>>> +\tmkdir smallDir &&\n>>> +\tfor i in $(test_seq 0 9)\n>>> +\tdo\n>>> +\t\techo $i >smallDir/$i\n>>> +\tdone &&\n>>> +\tgit add smallDir &&\n>>> +\tgit commit -m \"commit with 10 changes\" &&\n>>> +\tcat >expect <<-\\EOF &&\n>>> +\tFilter_Length:4\n>>> +\tFilter_Data:508928809087080a|8a7648210804001|4089824400951000|841ab310098051a8|\n>>> +\tEOF\n>>> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n>>> +\ttest_cmp expect actual\n>>> +'\n>> \n>> This test is in my opinion fragile, as it unnecessarily test the\n>> implementation details instead of the functionality provided.  If we\n>> change the hashing scheme (for example going from double hashing to some\n>> variant of enhanced double hashing), or change the base hash function\n>> (for example from Murmur3_32 to xxHash_64), or change the number of hash\n>> functions (perhaps because changing of number of bits per element, and\n>> thus optimal number of hash functions from 7 to 6), or change from\n>> 64-bit word blocks to 32-bit word blocks, the test would have to be\n>> changed.\n>\n> Regarding this and the rest of you comments on t0095-log-bloom.sh:\n>\n> I am tweaking it as necessary but the entire point of these tests is to\n> break for the things you called out. They need to be intricately tied\n> to the current hashing strategy and are hence intended to be fragile so \n> as to catch any subtle or accidental changes in the hashing computation. \n> Any change like the ones you have called out would require a hash version\n> change and all the compatibility reactions that come with it. \n\nAll right, if we assume that commit-graph is not something purely local^*,\nand we need iteroperability, then this test is necessary and is\nnecessarily fragile.\n\n*. This may happen because the repository and the commit-graph file in\n   it is on network disk, and accessed by hosts with different\n   endianness.  Or in the future (or possibly now, if one is using\n   Scalar) the commit-graph file can be sent together with packfile\n   during the fetch operation.\n\nOn the other hand testing the functionality of Murmur hash, and of Bloom\nfilter would help finding possible troubles if we decide in the future\nto change the algorithm details (change hash function, and/or move from\ndouble hashing to enhanced double hashing, and/or change how commits\nwith large number of changes are handled, or even switching to xor\nfilters [1]).\n\n[1]: Graf, Thomas Mueller; Lemire, Daniel (2019), \"Xor Filters: Faster\nand Smaller Than Bloom and Cuckoo Filters\", https://arxiv.org/abs/1912.08258\n\n> I have added more tests around the murmur3_seeded method in v3. Removed\n> some of the redundant ones. \n\nThere is another test that might be worth adding (see the comment below\nwhy), namely one test checking that bloom_key is computed as expected.\n\n> The other more evolved test cases you call out are covered in the e2e\n> integration tests in t4216-log-bloom.sh\n\nAll right, but there is another issue to consider.  Good tests should\nnot only catch the breakage, but also help to detect where the bug is.\nThat is one of advantages that unit tests (like the ones I have\nproposed) have over end-to-end functional tests.  They are also often\nfaster.\n\nOn the other hand e2e tests can catch problems with integration, and\nactually check that the user-visible behaviour is as expected.\n\n\nBest,\n--\nJakub Narębski\n\n>> \n>> Reviewed-by: Jakub Narębski <jnareb@gmail.com>\n>> \n>> Thanks for working on this.\n>> \n>> Best, \n>\n> Thank you once again for an excellent and in-depth review of this patch! \n> You have helped make this code so much better!\n"},{"id":"392359","messageId":"86o8tp8axt.fsf@gmail.com","threadId":"52499","inReplyTo":"0f1ab477-fae8-b744-5c48-87995f7fc8eb@gmail.com","subject":"Re: [PATCH v2 04/11] commit-graph: compute Bloom filters for changed paths","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-23T17:34:22Z","receivedAt":"2020-02-23T17:34:31Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Garima Singh <garimasigit@gmail.com> writes:\n> On 2/17/2020 4:56 PM, Jakub Narebski wrote:\n>> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n[...]\n>>> ---\n>>>  commit-graph.c | 32 +++++++++++++++++++++++++++++++-\n>>>  commit-graph.h |  3 ++-\n>>>  2 files changed, 33 insertions(+), 2 deletions(-)\n>> \n>> It would be good to have at least sanity check of this feature, perhaps\n>> one that would check that the number of per-commit Bloom filters on slab\n>> matches the number of commits in the commit-graph.\n>\n> The combination of all the e2e tests in this series with the test\n> flag being turned on in the CI, and the performance gains we are seeing\n> confirm that this is happening correctly.\n\nWell, the advantage of unit tests over e2e functional tests is that they\ncan pinpoint the source of bug much more easily.\n\nThat said, I don't think there is absolute need for unit tests here,\nthough it would be nice to have them.\n\n>>>  \n>>>  \tconst struct split_commit_graph_opts *split_opts;\n>>> +\tuint32_t total_bloom_filter_data_size;\n>> \n>> This is total size of Bloom filters data, in bytes, that will later be\n>> used for BDAT chunk size.  However the commit-graph format uses 8 bytes\n>> for byte-offset, not 4 bytes.  Why it is uint32_t and not uint64_t then?\n>\n> Changed to size_t. Thanks for noticing! \n\nRight, this is a local value (size_t may be different size on different\narchitectures), even though it will be stored indirectly in chunk lookup\ntable as pair of uint64_t offsets.\n\n>>>  };\n>>>  \n>>>  static void write_graph_chunk_fanout(struct hashfile *f,\n>>> @@ -1140,6 +1143,28 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n>>>  \tstop_progress(&ctx->progress);\n>>>  }\n>>>  \n>>> +static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>>> +{\n>>> +\tint i;\n>>> +\tstruct progress *progress = NULL;\n>>> +\n>>> +\tload_bloom_filters();\n>>> +\n>>> +\tif (ctx->report_progress)\n>>> +\t\tprogress = start_progress(\n>>> +\t\t\t_(\"Computing commit diff Bloom filters\"),\n>>> +\t\t\tctx->commits.nr);\n>>> +\n>> \n>> Shouldn't we initialize ctx->total_bloom_filter_data_size to 0 here?  We\n>> cannot use compute_bloom_filters() to _update_ Bloom filters data, I\n>> think -- we don't distinguish here between new and existing data (where\n>> existing data size is already included in total Bloom filters size).  At\n>> least I don't think so.\n>> \n>\n> This line in commit-graph.c takes care of reinitializing the graph context and\n> by consequence the bloom filter data size.\n>\n>   ctx = xcalloc(1, sizeof(struct write_commit_graph_context));\n>   \n> So the total size gets recalculated every time, which is correct. \n\nTrue, I have missed this.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"392414","messageId":"8f7a6168-8555-0a37-9945-1dc41d85c50d@gmail.com","threadId":"52499","inReplyTo":"86wo8d8lus.fsf@gmail.com","subject":"Re: [PATCH v2 02/11] bloom: core Bloom filter implementation for changed paths","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-24T17:34:02Z","receivedAt":"2020-02-24T17:34:09Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/23/2020 8:38 AM, Jakub Narebski wrote:\n> Garima Singh <garimasigit@gmail.com> writes:\n>> On 2/16/2020 11:49 AM, Jakub Narebski wrote:\n>>>> From: Garima Singh <garima.singh@microsoft.com>\n>>>>\n>>>> Add the core Bloom filter logic for computing the paths changed between a\n>>>> commit and its first parent. For details on what Bloom filters are and how they\n>>>> work, please refer to Dr. Derrick Stolee's blog post [1]. It provides a concise\n>>>> explaination of the adoption of Bloom filters as described in [2] and [3].\n>>>                                                                           ^^- to add\n>>\n>> Not sure what this means. Can you please clarify. \n>>\n>>>> 1. We currently use 7 and 10 for the number of hashes and the size of each\n>>>>    entry respectively. They served as great starting values, the mathematical\n>>>>    details behind this choice are described in [1] and [4]. The implementation,\n>>>                                                                                ^^- to add\n>>\n>> Not sure what this means. Can you please clarify.\n> \n> I'm sorry for not being clear.  What I wanted to say that in both cases\n> the last line should have ended in either full stop in first case, or\n> comma in second case:\n> \n>   \"as described in [2] and [3].\"\n> \n>   \"The implementation,\"\n> \n> What I wrote (trying to put the arrow below final fullstop or comma)\n> only works when one is using with fixed-width font.\n> \n\nAah. Cool. Thanks! \n\n>>>> ---\n>>>>  Makefile              |   2 +\n>>>>  bloom.c               | 228 ++++++++++++++++++++++++++++++++++++++++++\n>>>>  bloom.h               |  56 +++++++++++\n>>>>  t/helper/test-bloom.c |  84 ++++++++++++++++\n>>>>  t/helper/test-tool.c  |   1 +\n>>>>  t/helper/test-tool.h  |   1 +\n>>>>  t/t0095-bloom.sh      | 113 +++++++++++++++++++++\n>>>>  7 files changed, 485 insertions(+)\n>>>>  create mode 100644 bloom.c\n>>>>  create mode 100644 bloom.h\n>>>>  create mode 100644 t/helper/test-bloom.c\n>>>>  create mode 100755 t/t0095-bloom.sh\n>>>\n>>> As I wrote earlier, In my opinion this patch could be split into three\n>>> individual single-functionality pieces, to make it easier to review and\n>>> aid in bisectability if needed.\n>>\n>> Doing this in v3. \n> \n> Thanks.  Though if it makes (much) more work for you, I can work with\n> unsplit patch, no problem.\n> \n\nThanks! That's great! Splitting the patches will add some overhead. I will \ntry and do it provided it does not delay getting v3 on the list. \n\n>>>> +\n>>>> +static uint32_t rotate_right(uint32_t value, int32_t count)\n>>>> +{\n>>>> +\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n>>>> +\tcount &= mask;\n>>>> +\treturn ((value >> count) | (value << ((-count) & mask)));\n>>>> +}\n>>>\n>>> Hmmm... both the algoritm on Wikipedia, and reference implementation use\n>>> rotate *left*, not rotate *right* in the implementation of Murmur3 hash,\n>>> see\n>>>\n>>>   https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n>>>   https://github.com/aappleby/smhasher/blob/master/src/MurmurHash3.cpp#L23\n>>>\n>>>\n>>> inline uint32_t rotl32 ( uint32_t x, int8_t r )\n>>> {\n>>>   return (x << r) | (x >> (32 - r));\n>>> }\n>>\n>> Thanks! Fixed this in v3. More on it later. \n> \n> Sidenote: If I understand it correctly Bloom filters functionality is\n> included in Scalar [1].  What will happen then with all those Bloom\n> filter chunks in commit-graph files with wrong hash functions?\n> \n> [1]: https://devblogs.microsoft.com/devops/introducing-scalar/\n> \n\nIt is not included in Scalar. Scalar will write to the commit-graph in \nthe background using the features available in the git version it is working\nwith. It will update to include changed path Bloom filters when they are \navailable in git. We are not taking the Bloom filter into microsoft/git \nuntil the format is approved and accepted by the core git community.\n\n>>>> +{\n>>>> +\tconst uint32_t c1 = 0xcc9e2d51;\n>>>> +\tconst uint32_t c2 = 0x1b873593;\n>>>> +\tconst uint32_t r1 = 15;\n>>>> +\tconst uint32_t r2 = 13;\n>>>> +\tconst uint32_t m = 5;\n>>>> +\tconst uint32_t n = 0xe6546b64;\n>>>> +\tint i;\n>>>> +\tuint32_t k1 = 0;\n>>>> +\tconst char *tail;\n>>>> +\n>>>> +\tint len4 = len / sizeof(uint32_t);\n>>>> +\n>>>> +\tconst uint32_t *blocks = (const uint32_t*)data;\n>>>> +\n>>>> +\tuint32_t k;\n>>>> +\tfor (i = 0; i < len4; i++)\n>>>> +\t{\n>>>> +\t\tk = blocks[i];\n>>>\n>>> IMPORTANT: There is a comment around there in the example implementation\n>>> in C on Wikipedia that this operation above is a source of differing\n>>> results across endianness.  \n>>\n>> Thanks! SZEDER found this on his CI pipeline and we have fixed it to \n>> process the data in 1 byte words to avoid hitting any endian-ness issues. \n>> See this part of the thread that carries the fix and the related discussion. \n>>   https://lore.kernel.org/git/ba856e20-0a3c-e2d2-6744-b9abfacdc465@gmail.com/\n>> I will be squashing those changes in appropriately in v3.  \n> \n> [...]\n>>>> +\t\tk1 *= c2;\n>>>> +\t\tseed ^= k1;\n>>>> +\t\tbreak;\n>>>> +\t}\n>>>> +\n>>>> +\tseed ^= (uint32_t)len;\n>>>> +\tseed ^= (seed >> 16);\n>>>> +\tseed *= 0x85ebca6b;\n>>>> +\tseed ^= (seed >> 13);\n>>>> +\tseed *= 0xc2b2ae35;\n>>>> +\tseed ^= (seed >> 16);\n>>>> +\n>>>> +\treturn seed;\n>>>> +}\n>>>\n>>> In https://public-inbox.org/git/ba856e20-0a3c-e2d2-6744-b9abfacdc465@gmail.com/\n>>> you posted \"[PATCH] Process bloom filter data as 1 byte words\".\n>>> This may avoid the Big-endian vs Little-endian confusion,\n>>> that is wrong results on Big-endian architectures, but\n>>> it also may slow down the algorithm.\n>>\n>> Oh cool! You have seen that patch. And yes, we understand that it might add \n>> a little overhead but at this point it is more important to be correct on all\n>> architectures instead of micro-optimizing and introducing different \n>> implementations for Little-endian and Big-endian. This would make this \n>> series overly complicated. Optimizing the hashing techniques would deserve a\n>> series of its own, which we can definitely revisit later.\n> \n> Right, \"first make it work, then make it right, and, finally, make it fast.\".\n> \n> Anyway, could you maybe compare performance of Git for old version\n> (operating on 32-bit/4-bytes words) and new version (operating on 1-byte\n> words) file history operation with Bloom filters, to see if it matters\n> or not?\n> \n\nWe chose to switch to 1 byte words for correctness, not performance. \nAlso, this specific implementation choice is a very small portion of the \nend to end time spent computing and writing Bloom filters. We run two murmur3 \nhashes per path, which is one path per `git log` query; and one path per change \nafter parsing trees to compute a diff. Measuring performance and micro-optimizing \nis not worth the effort and/or trading in the simplicity here.\n\n\n>>>> +\n>>>> +struct bloom_filter *get_bloom_filter(struct repository *r,\n>>>> +\t\t\t\t      struct commit *c)\n>>>> +{\n>>>> +\tstruct bloom_filter *filter;\n>>>> +\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n>>>> +\tint i;\n>>>> +\tstruct diff_options diffopt;\n>>>> +\n>>>> +\tif (!bloom_filters.slab_size)\n>>>> +\t\treturn NULL;\n>>>\n>>> This is testing that commit slab for per-commit Bloom filters is\n>>> initialized, isn't it?\n>>>\n>>> First, should we write the condition as\n>>>\n>>> \tif (!bloom_filters.slab_size)\n>>>\n>>> or would the following be more readable\n>>>\n>>> \tif (bloom_filters.slab_size == 0)\n>>>\n>>\n>> Sure. Switched to `if (bloom_filter.slab_size == 0)` in v3. \n> \n> Though either works, and the former looks more like the test if\n> bloom_filters slab are initialized, now that I thought about it a bit.\n> Your choice.\n> \n\n:) \n\n\n>>>> +\n>>>> +\tif (diff_queued_diff.nr <= 512) {\n>>>\n>>> Second, there is a minor issue that diff_queue_struct.nr stores the\n>>> number of filepairs, that is the number of changed files, while the\n>>> number of elements added to Bloom filter is number of changed blobs and\n>>> trees.  For example if the following files are changed:\n>>>\n>>>   sub/dir/file1\n>>>   sub/file2\n>>>\n>>> then diff_queued_diff.nr is 2, but number of elements to be added to\n>>> Bloom filter is 4.\n>>>\n>>>   sub/dir/file1\n>>>   sub/file2\n>>>   sub/dir/\n>>>   sub/\n>>>\n>>> I'm not sure if it matters in practice.\n>>>\n>>\n>> It does not matter much in practice, since the directories usually tend\n>> to collapse across the changes. Still, I will add another limit after \n>> creating the hashmap entries to cap at 640 so that we have a maximum of \n>> 100 changes in the bloom filter. \n>>\n>> We plan to make these values configurable later. \n> \n> I'm not sure if it is truly necessary; we can treat limit on number of\n> changed paths as \"best effort\" limit on Bloom filter size.\n> \n> I just wanted to point out the difference.\n> \n\nSure. Not doing this for v3. Glad it got discussed here though!  \n\n> \n> Side note: I wonder if it would be worth it (in the future) to change\n> handling commits with large amount of changes.  I was thinking about\n> switching to soft and hard limit: soft limit would be on the size of the\n> Bloom filter, that is if number of elements times bits per element is\n> greater that size threshold, we don't increase the size of the filter.\n> \n> This would mean that the false positives ratio (the number of files that\n> are not present but get answer \"maybe\" instead of \"no\" out of the\n> filter) would increase, so there would be a need for another hard limit\n> where we decide that it is not worth it, and not store the data for the\n> Bloom filter -- current \"no data\" case with empty filter with length 0.\n> This hard limit can be imposed on number of changed files, or on number\n> of paths added to filter, or on number of bits set to 1 in the filter\n> (on popcount), or some combination thereof.\n> \n> [...]\n\nCould be considered in the future. Doesn't make the cut for the current\nseries though. \n\nThanks\nGarima Singh\n"},{"id":"392419","messageId":"86h7zf97ak.fsf@gmail.com","threadId":"52499","inReplyTo":"8f7a6168-8555-0a37-9945-1dc41d85c50d@gmail.com","subject":"Re: [PATCH v2 02/11] bloom: core Bloom filter implementation for changed paths","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-24T18:20:03Z","receivedAt":"2020-02-24T18:20:10Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Garima Singh <garimasigit@gmail.com> writes:\n\n> On 2/23/2020 8:38 AM, Jakub Narebski wrote:\n>> Garima Singh <garimasigit@gmail.com> writes:\n>>> On 2/16/2020 11:49 AM, Jakub Narebski wrote:\n>>>>> From: Garima Singh <garima.singh@microsoft.com>\n\n[...]\n>>>> IMPORTANT: There is a comment around there in the example implementation\n>>>> in C on Wikipedia that this operation above is a source of differing\n>>>> results across endianness.  \n>>>\n>>> Thanks! SZEDER found this on his CI pipeline and we have fixed it to \n>>> process the data in 1 byte words to avoid hitting any endian-ness issues. \n>>> See this part of the thread that carries the fix and the related discussion. \n>>>   https://lore.kernel.org/git/ba856e20-0a3c-e2d2-6744-b9abfacdc465@gmail.com/\n>>> I will be squashing those changes in appropriately in v3.  \n>> \n>> [...]\n>>>>\n>>>> In https://public-inbox.org/git/ba856e20-0a3c-e2d2-6744-b9abfacdc465@gmail.com/\n>>>> you posted \"[PATCH] Process bloom filter data as 1 byte words\".\n>>>> This may avoid the Big-endian vs Little-endian confusion,\n>>>> that is wrong results on Big-endian architectures, but\n>>>> it also may slow down the algorithm.\n>>>\n>>> Oh cool! You have seen that patch. And yes, we understand that it might add \n>>> a little overhead but at this point it is more important to be correct on all\n>>> architectures instead of micro-optimizing and introducing different \n>>> implementations for Little-endian and Big-endian. This would make this \n>>> series overly complicated. Optimizing the hashing techniques would deserve a\n>>> series of its own, which we can definitely revisit later.\n>> \n>> Right, \"first make it work, then make it right, and, finally, make it fast.\".\n>> \n>> Anyway, could you maybe compare performance of Git for old version\n>> (operating on 32-bit/4-bytes words) and new version (operating on 1-byte\n>> words) file history operation with Bloom filters, to see if it matters\n>> or not?\n>> \n>\n> We chose to switch to 1 byte words for correctness, not performance. \n> Also, this specific implementation choice is a very small portion of the \n> end to end time spent computing and writing Bloom filters. We run two murmur3 \n> hashes per path, which is one path per `git log` query; and one path per change \n> after parsing trees to compute a diff. Measuring performance and micro-optimizing \n> is not worth the effort and/or trading in the simplicity here.\n\nAll right.\n\nI still think that adding to_le32() invocation before the part that\nprocesses remaining bytes (the 'switch' instruction in v2 code), just\nlike in pseudo-code on Wikipedia:\n\n    with any remainingBytesInKey do\n        remainingBytes ← SwapToLittleEndian(remainingBytesInKey)\n\nwould be enough to have correct results regardlless of endianness.\n\nAs I wrote\n\nJN> The solution in PMurHash.c in Chromium [1], and the pseudo-code algorithm on\nJN> Wikipedia do endian handling only for remaining bytes (while the\nJN> beginnings of solution in Appleby's code, and solution in current\nJN> above-mentioned Chromium implementation do the conversion for all\nJN> bytes).  I think that handling it only for remaining bytes (for data\nJN> sizes not being multiply of 32-bits / 4-bytes) is enough; all other\nJN> operations, that is multiply, rotate, xor and addition do not depend on\nJN> endianness.\n\n[1]: https://chromium.googlesource.com/external/smhasher/+/5b8fd3c31a58b87b80605dca7a64fad6cb3f8a0f/PMurHash.c\n\nIf you have access to, or can run code on some big-endian architecture,\nit should be easy enough to check it.\n\n\nAnyway, if you decide on 1-byte at time implementation, please put a\ncomment about 32-bit chunk implementation.\n\n>> Side note: I wonder if it would be worth it (in the future) to change\n>> handling commits with large amount of changes.  I was thinking about\n>> switching to soft and hard limit: soft limit would be on the size of the\n>> Bloom filter, that is if number of elements times bits per element is\n>> greater that size threshold, we don't increase the size of the filter.\n>> \n>> This would mean that the false positives ratio (the number of files that\n>> are not present but get answer \"maybe\" instead of \"no\" out of the\n>> filter) would increase, so there would be a need for another hard limit\n>> where we decide that it is not worth it, and not store the data for the\n>> Bloom filter -- current \"no data\" case with empty filter with length 0.\n>> This hard limit can be imposed on number of changed files, or on number\n>> of paths added to filter, or on number of bits set to 1 in the filter\n>> (on popcount), or some combination thereof.\n>> \n>> [...]\n>\n> Could be considered in the future. Doesn't make the cut for the current\n> series though. \n\nRight.\n\nBest,\n-- \nJakub Narębski\n"},{"id":"392423","messageId":"f7a0d25d-0dfe-6b16-712b-1b2744378982@gmail.com","threadId":"52499","inReplyTo":"86k14jkc8s.fsf@gmail.com","subject":"Re: [PATCH v2 05/11] commit-graph: examine changed-path objects in pack order","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-24T18:29:44Z","receivedAt":"2020-02-24T18:29:50Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/18/2020 12:59 PM, Jakub Narebski wrote:\n> \"Jeff King via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> From: Jeff King <peff@peff.net>\n>>\n>> Looking at the diff of commit objects in pack order is much faster than\n>> in sha1 order, as it gives locality to the access of tree deltas\n> \n> Nitpick: should we still say sha1 order?  Git is still using SHA-1 as an\n> *oid*, but hopefully soon it will be transitioning to NewHash = SHA-256.\n> (No need to change anything.)\n> \n>> (whereas sha1 order is effectively random). Unfortunately the\n>> commit-graph code sorts the commits (several times, sometimes as an oid\n>> and sometimes a pointer-to-commit), and we ultimately traverse in sha1\n>> order.\n> \n> Actually, commit-graph code needs write_commit_graph_context.commits.list\n> to be in lexicographical order to be able to turn position in graph into\n> reference to a commit.  The information about the parents of the commit\n> are stored using positional references within the graph file.\n> \n\nYou are right. Fixing the commit message in v3. \n\n>>\n>> Instead, let's remember the position at which we see each commit, and\n>> traverse in that order when looking at bloom filters. This drops my time\n>> for \"git commit-graph write --changed-paths\" in linux.git from ~4\n>> minutes to ~1.5 minutes.\n> \n> Nitpick: with reordering of patches (which I think is otherwise a good\n> thing) this patch actually comes before the one adding \"--changed-paths\"\n> option to \"git commit-graph write\".  So it 'This would drop my time'\n> rather than 'This drops my time...' ;-)\n> \n\n:) I will fix that up. \n\n>>\n>> Probably the \"--reachable\" code path would want something similar.\n> \n> Has anyone tried doing this?\n> \n\nI will and I will include the perf numbers in the appropriately in v3. \n\n\n>> +\n>>  char *get_commit_graph_filename(const char *obj_dir)\n>>  {\n>>  \tchar *filename = xstrfmt(\"%s/info/commit-graph\", obj_dir);\n>> @@ -1027,6 +1051,8 @@ static int add_packed_commits(const struct object_id *oid,\n>>  \toidcpy(&(ctx->oids.list[ctx->oids.nr]), oid);\n>>  \tctx->oids.nr++;\n>>  \n>> +\tset_commit_pos(ctx->r, oid);\n>> +\n>>  \treturn 0;\n>>  }\n>>  \n>> @@ -1147,6 +1173,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>>  {\n>>  \tint i;\n>>  \tstruct progress *progress = NULL;\n>> +\tstruct commit **sorted_by_pos;\n> \n> In the next patch in series we would sort commits by generation number\n> and creation data; shouldn't this variable name be more generic to\n> reflect this, for example just `sorted_commits` or `commits_sorted`?\n> \n\nGood call. I will clean this up in both commits. \n\nThanks for the review! \nCheers! \nGarima Singh\n"},{"id":"392436","messageId":"b959d5d4-4e9f-1fc9-f06f-40228d56b80c@gmail.com","threadId":"52499","inReplyTo":"865zg3ju2j.fsf@gmail.com","subject":"Re: [PATCH v2 06/11] commit-graph: examine commits by generation number","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-24T20:45:55Z","receivedAt":"2020-02-24T20:46:03Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/18/2020 7:32 PM, Jakub Narebski wrote:\n> \"Derrick Stolee via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> From: Derrick Stolee <dstolee@microsoft.com>\n>>\n>> When running 'git commit-graph write --changed-paths', we sort the\n>> commits by pack-order to save time when computing the changed-paths\n>> bloom filters. This does not help when finding the commits via the\n>> --reachable flag.\n> \n> Minor improvement suggestion: s/--reachable flag/'--reachable' flag/.\n> \n\nSure. \n\n\n>>                     Commits with similar generation are more likely\n>> to have many trees in common, making the diff faster.\n> \n> Is this what causes the performance improvement, that subsequently\n> examined commits are more likely to have more trees in common, which\n> means that those trees would be hot in cache, making generating diff\n> faster?  Is it what profiling shows?\n> \n\nYes. \n\n>>\n>> Helped-by: Jeff King <peff@peff.net>\n>> Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n>> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n>> ---\n>>  commit-graph.c | 33 ++++++++++++++++++++++++++++++---\n>>  1 file changed, 30 insertions(+), 3 deletions(-)\n>>\n>> diff --git a/commit-graph.c b/commit-graph.c\n>> index e125511a1c..32a315058f 100644\n>> --- a/commit-graph.c\n>> +++ b/commit-graph.c\n>> @@ -70,6 +70,25 @@ static int commit_pos_cmp(const void *va, const void *vb)\n>>  \t       commit_pos_at(&commit_pos, b);\n>>  }\n>>  \n>> +static int commit_gen_cmp(const void *va, const void *vb)\n>> +{\n>> +\tconst struct commit *a = *(const struct commit **)va;\n>> +\tconst struct commit *b = *(const struct commit **)vb;\n>> +\n>> +\t/* lower generation commits first */\n> \n> Shouldn't higher generation commits come first, in recency-like order?\n> Or it doesn't matter if it is sorted in ascending or descending order,\n> as long as commits with close generation numbers are examined close\n> together?\n> \n\nThe direction does not matter. Locality is important. \n\n>> +\tif (a->generation < b->generation)\n>> +\t\treturn -1;\n>> +\telse if (a->generation > b->generation)\n>> +\t\treturn 1;\n>> +\n>> +\t/* use date as a heuristic when generations are equal */\n>> +\tif (a->date < b->date)\n>> +\t\treturn -1;\n>> +\telse if (a->date > b->date)\n>> +\t\treturn 1;\n>> +\treturn 0;\n>> +}\n> \n> I thought we have had such comparison function defined somewhere in Git\n> already, but I think I'm wrong here.\n> \n\nIt actually exists in commit.h\nI will just use it here. \nThanks for pointing it out! \n\n>> +\n>>  char *get_commit_graph_filename(const char *obj_dir)\n>>  {\n>>  \tchar *filename = xstrfmt(\"%s/info/commit-graph\", obj_dir);\n>> @@ -821,7 +840,8 @@ struct write_commit_graph_context {\n>>  \t\t report_progress:1,\n>>  \t\t split:1,\n>>  \t\t check_oids:1,\n>> -\t\t changed_paths:1;\n>> +\t\t changed_paths:1,\n>> +\t\t order_by_pack:1;\n>>  \n>>  \tconst struct split_commit_graph_opts *split_opts;\n>>  \tuint32_t total_bloom_filter_data_size;\n>> @@ -1184,7 +1204,11 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>>  \n>>  \tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n>>  \tCOPY_ARRAY(sorted_by_pos, ctx->commits.list, ctx->commits.nr);\n>> -\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n>> +\n>> +\tif (ctx->order_by_pack)\n>> +\t\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n>> +\telse\n>> +\t\tQSORT(sorted_by_pos, ctx->commits.nr, commit_gen_cmp);\n> \n> Here 'sorted_b_pos' variable name no longer reflects reality...\n> (see comment to the previous patch in the series).\n> \n\nYup. Fixing. \n\nThanks!\nGarima Singh\n"},{"id":"392441","messageId":"de3f1f7e-0f2f-6c5d-6290-3ba5d37a0ea5@gmail.com","threadId":"52499","inReplyTo":"86pneahaop.fsf@gmail.com","subject":"Re: [PATCH v2 07/11] commit-graph: write Bloom filters to commit graph file","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-24T21:14:38Z","receivedAt":"2020-02-24T21:14:45Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\n\nOn 2/19/2020 10:13 AM, Jakub Narebski wrote:\n> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> From: Garima Singh <garima.singh@microsoft.com>\n>>\n>> Update the technical documentation for commit-graph-format with the formats for\n>> the Bloom filter index (BIDX) and Bloom filter data (BDAT) chunks. Write the\n>> computed Bloom filters information to the commit graph file using this format.\n> \n> Nice description.\n> \n> The only minor nitpick is with the formating: it is 80-character wide,\n> which is a bit wide.\n> \n\nFixed in v3. Thanks! \n\n>>\n>> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n>> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n>> ---\n>>  .../technical/commit-graph-format.txt         |  24 ++++\n>>  commit-graph.c                                | 118 +++++++++++++++++-\n>>  commit-graph.h                                |   7 +-\n>>  3 files changed, 145 insertions(+), 4 deletions(-)\n>>\n>> diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n>> index a4f17441ae..22e511643d 100644\n>> --- a/Documentation/technical/commit-graph-format.txt\n>> +++ b/Documentation/technical/commit-graph-format.txt\n>> @@ -17,6 +17,9 @@ metadata, including:\n>>  - The parents of the commit, stored using positional references within\n>>    the graph file.\n>>  \n>> +- The Bloom filter of the commit carrying the paths that were changed between\n>> +  the commit and its first parent.\n>> +\n> \n> All right.\n> \n> Should we also state that it is optional (meta)data?  This would be\n> first optional piece of data stored in commit-graph, I think.\n> \n\nHowever the entire commit graph file is non critical metadata since git commands\nwork just fine without it, just slower. The same applies to the changed path\nbloom filters. \n\nBased on the definition of optional you are suggesting, edge data is optional\nbecause not every commit-graph has octopus merges. \n\n>> +\n>> +  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n>> +    * It starts with header consisting of three unsigned 32-bit integers:\n>> +      - Version of the hash algorithm being used. We currently only support\n>> +\tvalue 1 which implies the murmur3 hash implemented exactly as described\n>> +\tin https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n> \n> First a minor issue: shouldn't this nested unordered list be indented\n> with a hanging indent formatted with spaces?  That is be formatted like\n> the following:\n> \n>   +  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n>   +    * It starts with header consisting of three unsigned 32-bit integers:\n>   +      - Version of the hash algorithm being used. We currently only support\n>   +        value 1 which implies the murmur3 hash implemented exactly as\n>   +        described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n> \n> But the existing formatting with spaces and tabs might be fine as it is,\n> that is it renders as nested list with Asciidoc; it only looks a bit\n> weird as patch, not so as text.\n> \n> Second, and more important: it is in my opinion not enough information,\n> at least if we are assuming that the information in this document should\n> be enough for clean-room reimplementation of Bloom filter functionality\n> (for example by JGit).  To generate compatible Bloom filters, one needs\n> also the information on how to create $k$ functionally-independent hash\n> functions out of murmur3 hash.  We do it currently using double hashing\n> technique; if that changes then the exact set of bits in the Bloom\n> filter would also change.\n> \n> The additional description could look something like the following:\n> \n>   +    * It starts with header consisting of three unsigned 32-bit integers:\n>   +      - Version of the hash algorithm being used. We currently only support\n>   +        value 1 which implies the murmur3_32 hash implemented exactly as\n>   +        described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n>   +        and double hashing technique with 0x293ae76f and 0x7e646e2c seeds\n>   +        as described in https://doi.org/10.1007/978-3-540-30494-4_26\n>   +        \"Bloom Filters in Probabilistic Verification\"\n> \n> Also, it should be explicitly noted that we use murmur3_32, because\n> there is also 128-bit version of murmur3 hash.\n> \n\nI will incorporate this in. Thanks! \n\n\n>> +    * The BDAT chunk is present iff BIDX is present.\n> \n> Perhaps we should spell 'iff' in full, that is 'if and only if'?\n> \n\nSure. \n\n>> +\n>>    Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n>>        This list of H-byte hashes describe a set of B commit-graph files that\n>>        form a commit-graph chain. The graph position for the ith commit in this\n>> diff --git a/commit-graph.c b/commit-graph.c\n>> index 32a315058f..4585b3b702 100644\n>> --- a/commit-graph.c\n>> +++ b/commit-graph.c\n>> @@ -24,8 +24,10 @@\n>>  #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n>>  #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n>>  #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n>> +#define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n>> +#define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n>>  #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n>> -#define MAX_NUM_CHUNKS 5\n>> +#define MAX_NUM_CHUNKS 7\n>>  \n>>  #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n>>  \n>> @@ -325,6 +327,32 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n>>  \t\t\t\tchunk_repeated = 1;\n>>  \t\t\telse\n>>  \t\t\t\tgraph->chunk_base_graphs = data + chunk_offset;\n>> +\t\t\tbreak;\n>> +\n>> +\t\tcase GRAPH_CHUNKID_BLOOMINDEXES:\n>> +\t\t\tif (graph->chunk_bloom_indexes)\n>> +\t\t\t\tchunk_repeated = 1;\n>> +\t\t\telse\n>> +\t\t\t\tgraph->chunk_bloom_indexes = data + chunk_offset;\n>> +\t\t\tbreak;\n>> +\n>> +\t\tcase GRAPH_CHUNKID_BLOOMDATA:\n>> +\t\t\tif (graph->chunk_bloom_data)\n>> +\t\t\t\tchunk_repeated = 1;\n>> +\t\t\telse {\n>> +\t\t\t\tuint32_t hash_version;\n>> +\t\t\t\tgraph->chunk_bloom_data = data + chunk_offset;\n>> +\t\t\t\thash_version = get_be32(data + chunk_offset);\n>> +\n>> +\t\t\t\tif (hash_version != 1)\n>> +\t\t\t\t\tbreak;\n> \n> Shouldn't we mark Bloom filter as not to be used?  Or is it left for\n> later commit?\n> \n\nWe take care of this in line 375. \n\n> In the future it might be good idea to notify the user (perhaps\n> protected with some advice.* option) that there is problem with Bloom\n> filter data, namely that we have encountered unsupported hash version.\n> \n>> +\n>> +\t\t\t\tgraph->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n> \n> Why is this structure allocated dynamically?  We are leaking admittedly\n> a small amount of memory because we never free this xmalloc() result.\n> \n> If we need this field being a pointer to struct to have NULL mean no\n> supported Bloom filter data, we could have instead use chunk_bloom_*\n> fields instead - we can set at least one of them to NULL.\n> \n\nI am freeing this up in free_commit_graph but I messed up putting it in the right commit. \nSorry about that. Fixed in v3. \n\nAlso as discussed in https://lore.kernel.org/git/3b7d77a1-aed9-d202-8646-4b964cb965db@gmail.com/\nthere is a bug in commit-graph.c where we should be calling free_commit_graph() instead of \njust free(graph). I will do this in a separate series. \n\n>> +\t\t\t}\n>> +\t\t\tbreak;\n>>  \t\t}\n>>  \n>>  \t\tif (chunk_repeated) {\n>> @@ -343,6 +371,17 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n>>  \t\tlast_chunk_offset = chunk_offset;\n>>  \t}\n>>  \n>> +\t/* We need both the bloom chunks to exist together. Else ignore the data */\n>> +\tif ((graph->chunk_bloom_indexes && !graph->chunk_bloom_data)\n>> +\t\t || (!graph->chunk_bloom_indexes && graph->chunk_bloom_data)) {\n>> +\t\tgraph->chunk_bloom_indexes = NULL;\n>> +\t\tgraph->chunk_bloom_data = NULL;\n>> +\t\tgraph->bloom_filter_settings = NULL;\n>> +\t}\n>> +\n>> +\tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data)\n>> +\t\tload_bloom_filters();\n> \n> Wouldn't it be simpler to rely on the fact that both Bloom chunks must\n> exists for it to matter, and write it like this:\n> \n>   +\tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data) {\n>   +\t\tload_bloom_filters();\n>   +\t} else {\n>   +\t\tgraph->chunk_bloom_indexes = NULL;\n>   +\t\tgraph->chunk_bloom_data = NULL;\n>   +\t\tgraph->bloom_filter_settings = NULL;\n>   +\t}\n> \n\n:) Yes. Fixed in v3. \n\n>> +\n>>  static int oid_compare(const void *_a, const void *_b)\n>>  {\n>>  \tconst struct object_id *a = (const struct object_id *)_a;\n>> @@ -1198,8 +1290,8 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>>  \tload_bloom_filters();\n>>  \n>>  \tif (ctx->report_progress)\n>> -\t\tprogress = start_progress(\n>> -\t\t\t_(\"Computing commit diff Bloom filters\"),\n>> +\t\tprogress = start_delayed_progress(\n>> +\t\t\t_(\"Computing changed paths Bloom filters\"),\n>>  \t\t\tctx->commits.nr);\n>>\n> \n> Ooops.  This look like a fixup which should be made to the original\n> earlier commit instead, isn't it?\n\n\nYes. Should have been in a previous commit. Fixed in v3. \n\n\n>>  };\n>>  \n>>  struct commit_graph *load_commit_graph_one_fd_st(int fd, struct stat *st);\n>> @@ -77,7 +82,7 @@ enum commit_graph_write_flags {\n>>  \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n>>  \t/* Make sure that each OID in the input is a valid commit OID. */\n>>  \tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n>> -\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n>> +\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4),\n> \n> This looks like accidental change; if we want to use trailing comma in\n> enum, this change should be in my opinion done in the commit that added\n> COMMIT_GRAPH_WRITE_BLOOM_FILTERS (as I have written in a comment there).\n> \n\nYes, I noticed the lack of the comma later and forgot to move it to the right\ncommit. Fixed in v3. \n\n> \n> Thank you for your work on this series.\n> \n> Best,\n> \n"},{"id":"392444","messageId":"2ca9f6ab-41c4-37d2-7681-8f973204d6a2@gmail.com","threadId":"52499","inReplyTo":"86r1ypf62y.fsf@gmail.com","subject":"Re: [PATCH v2 08/11] commit-graph: reuse existing Bloom filters during write.","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-24T21:45:00Z","receivedAt":"2020-02-24T21:45:06Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\n\nOn 2/20/2020 1:48 PM, Jakub Narebski wrote:\n> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> From: Garima Singh <garima.singh@microsoft.com>\n>>\n>> Read previously computed Bloom filters from the commit-graph file if\n>> possible to avoid recomputing during commit-graph write.\n> \n> All right, what is written makes sense for this point in patch series.\n> \n> But it my opinion it is more important to state that this commit adds\n> \"parsing\" of the Bloom filter data from commit-graph file.  This means\n> that it needs to be calculated only once, then stored in commit-graph,\n> ready to be re-used.\n> \n\nGood point. Incorporated in v3.\n\n>>\n>> See Documentation/technical/commit-graph-format for the format in which\n>> the Bloom filter information is written to the commit graph file.\n>>\n>> To read Bloom filter for a given commit with lexicographic position\n>> 'i' we need to:\n>> 1. Read BIDX[i] which essentially gives us the starting index in BDAT for\n>>    filter of commit i+1. It is essentially the index past the end\n>>    of the filter of commit i. It is called end_index in the code.\n>>\n>> 2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT\n>>    for filter of commit i. It is called the start_index in the code.\n>>    For the first commit, where i = 0, Bloom filter data starts at the\n>>    beginning, just past the header in the BDAT chunk. Hence, start_index\n>>    will be 0.\n>>\n>> 3. The length of the filter will be end_index - start_index, because\n>>    BIDX[i] gives the cumulative 8-byte words including the ith\n>>    commit's filter.\n>>\n>> We toggle whether Bloom filters should be recomputed based on the\n>> compute_if_null flag.\n> \n> Nitpick: the flag (the parameter) is called compute_if_not_present, not\n> compute_if_null.\n> \nOops. Fixed in v3. \n\n>> +\n>> +\tend_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n>> +\n>> +\tif (lex_pos)\n> \n> Wouldn't it be better to be more explicit, and write\n> \n>   +\tif (lex_pos > 0)\n> \n> \n\nSure. \n\n>> +\t\tstart_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n>> +\telse\n>> +\t\tstart_index = 0;\n> \n> All right, here we find start_index and end_index.\n> \n> It might be good idea to at least assert() that start_index <= end_index,\n> though that should not happen (that is why I propose for this check to\n> be compiled on only for debug builds).\n> \n\nI will look into this. Thanks! \n\n\n>> @@ -1304,7 +1304,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>>  \n>>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>>  \t\tstruct commit *c = sorted_by_pos[i];\n>> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n>> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n>>  \t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n>>  \t\tdisplay_progress(progress, i + 1);\n>>  \t}\n>> @@ -2314,6 +2314,7 @@ void free_commit_graph(struct commit_graph *g)\n>>  \t\tg->data = NULL;\n>>  \t\tclose(g->graph_fd);\n>>  \t}\n>> +\tfree(g->bloom_filter_settings);\n>>  \tfree(g->filename);\n>>  \tfree(g);\n> \n> Shouldn't this fixup be added to earlier commit?\n> \n\nYes. \n\n>>  }\n>> diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n>> index 331957011b..9b4be97f75 100644\n>> --- a/t/helper/test-bloom.c\n>> +++ b/t/helper/test-bloom.c\n>> @@ -47,7 +47,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n>>  \tstruct bloom_filter *filter;\n>>  \tsetup_git_directory();\n>>  \tc = lookup_commit(the_repository, commit_oid);\n>> -\tfilter = get_bloom_filter(the_repository, c);\n>> +\tfilter = get_bloom_filter(the_repository, c, 1);\n>>  \tprint_bloom_filter(filter);\n>>  }\n> \n> I would like to see some tests, but that needs to wait for patch that\n> adds --changed-paths option to the 'write' subcommand.\n> \n> Things to be tested:\n> 1. That after reading commit-graph with Bloom filter:\n>    - that commit(s) in commit-graph have Bloom filter\n>    - that commits outside commit-graph do not have Bloom filter\n> 2. That incremental commit-graph feature works:\n>    - for commits in deeper layer that have Bloom filter chunks\n>    - for commits in deeper layer that do not have Bloom filter chunks\n> \n\nIncluded in later commits. \n\n> Best,\n> \n"},{"id":"392445","messageId":"56150788-c477-5526-2d6d-e9325ccb4da6@gmail.com","threadId":"52499","inReplyTo":"86y2sxdmw9.fsf@gmail.com","subject":"Re: [PATCH v2 09/11] commit-graph: add --changed-paths option to write subcommand","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-24T21:51:48Z","receivedAt":"2020-02-24T21:51:54Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\n\nOn 2/20/2020 3:28 PM, Jakub Narebski wrote:\n> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> From: Garima Singh <garima.singh@microsoft.com>\n>>\n>> Add --changed-paths option to git commit-graph write. This option will\n>> allow users to compute information about the paths that have changed\n>> between a commit and its first parent, and write it into the commit graph\n>> file. If the option is passed to the write subcommand we set the\n>> COMMIT_GRAPH_WRITE_BLOOM_FILTERS flag and pass it down to the\n>> commit-graph logic.\n> \n> In the manpage you write that this operation (computing Bloom filters)\n> can take a while on large repositories.  Could you perhaps provide some\n> numbers: how much longer does it take to write commit-graph file with\n> and without '--changed-paths' for example for Linux kernel, or some\n> other large repository?  Thanks in advance.\n> \n\nYes. Will include numbers as appropriate in v3. \n\n>>\n>> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n>> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n>> ---\n>>  Documentation/git-commit-graph.txt | 5 +++++\n>>  builtin/commit-graph.c             | 9 +++++++--\n>>  2 files changed, 12 insertions(+), 2 deletions(-)\n> \n> What is missing is some sanity tests: that bloom index and bloom data\n> chunks are not present without '--changed-paths', and that they are\n> added with '--changed-paths'.\n> \n> If possible, maybe also check in a separate test that the size of\n> bloom_index chunk agrees with the number of commits in the commit graph.\n> \n> \n> Also, we can now add those tests I have wrote about in my review of\n> previous patch, that is:\n> \n> 1. If you write commit-graph with --changed-paths, and either add some\n>    commits later or exclude some commits from the commit graph, then:\n> \n>    a.) commit(s) in commit-graph have Bloom filter\n>    b.) commit(s) not in commit-graph do not have Bloom filter\n> \n> 2. If you write commit-graph without --changed-paths as base layer,\n>    and then write next layer with --changed-paths and --split, then:\n> \n>    a.) commit(s) in top layer have Bloom filter(s)\n>    b.) commit(s) in bottom layer don't have Bloom filter(s)\n> \n\nI will see what more can be done here. \n\n>>\n>> diff --git a/Documentation/git-commit-graph.txt b/Documentation/git-commit-graph.txt\n>> index bcd85c1976..907d703b30 100644\n>> --- a/Documentation/git-commit-graph.txt\n>> +++ b/Documentation/git-commit-graph.txt\n>> @@ -54,6 +54,11 @@ or `--stdin-packs`.)\n>>  With the `--append` option, include all commits that are present in the\n>>  existing commit-graph file.\n>>  +\n>> +With the `--changed-paths` option, compute and write information about the\n>> +paths changed between a commit and it's first parent. This operation can\n>> +take a while on large repositories. It provides significant performance gains\n>> +for getting history of a directory or a file with `git log -- <path>`.\n>> ++\n> \n> Should we write about limitation that the topmost layer in the split\n> commit graph needs to be written with '--changed-paths' for Git to use\n> this information?  Or perhaps we should try (in the future) to remove\n> this limitation??\n> \n\nGiven that this information is going to be used best effort, it would be \nsuperfluous to describe every case and conditional that decides whether \nthis information is being used.\n>> @@ -143,6 +146,8 @@ static int graph_write(int argc, const char **argv)\n>>  \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n>>  \tif (opts.progress)\n>>  \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n>> +\tif (opts.enable_changed_paths)\n>> +\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n>>  \n>>  \tread_replace_refs = 0;\n> \n> All right.  This actually turns on calculation Bloom filters for changed\n> paths, thanks to\n> \n>  \tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n> \n> that was added by the \"[PATCH v2 04/11] commit-graph: compute Bloom\n> filters for changed paths\" patch.\n> \n> Though... should this enabling be split into two separate patches like\n> this?\n> \n\nThe idea is that in 4/11 We compute only if the flag is set. \nAnd between that patch and this one: we prepare the foundational code \nthat is now ready for that flag to be set via an opt-in by the user. \n\n> \n> Best,\n> \n"},{"id":"392469","messageId":"86sgiyc2ta.fsf@gmail.com","threadId":"52499","inReplyTo":"de3f1f7e-0f2f-6c5d-6290-3ba5d37a0ea5@gmail.com","subject":"Re: [PATCH v2 07/11] commit-graph: write Bloom filters to commit graph file","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-25T11:40:49Z","receivedAt":"2020-02-25T11:40:57Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Garima Singh <garimasigit@gmail.com> writes:\n> On 2/19/2020 10:13 AM, Jakub Narebski wrote:\n>> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n[...]\n>>> diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n>>> index a4f17441ae..22e511643d 100644\n>>> --- a/Documentation/technical/commit-graph-format.txt\n>>> +++ b/Documentation/technical/commit-graph-format.txt\n>>> @@ -17,6 +17,9 @@ metadata, including:\n>>>  - The parents of the commit, stored using positional references within\n>>>    the graph file.\n>>>  \n>>> +- The Bloom filter of the commit carrying the paths that were changed between\n>>> +  the commit and its first parent.\n>>> +\n>> \n>> All right.\n>> \n>> Should we also state that it is optional (meta)data?  This would be\n>> first optional piece of data stored in commit-graph, I think.\n>> \n>\n> However the entire commit graph file is non critical metadata since git commands\n> work just fine without it, just slower. The same applies to the changed path\n> bloom filters. \n>\n> Based on the definition of optional you are suggesting, edge data is optional\n> because not every commit-graph has octopus merges. \n\nWell, edge data (EDGE chunk) is optional in different way from Bloom\nfilter data.  The former depends on the repository (whether there are\noctopus merges used), the latter is opt-in user choice (whether to run\n`git commit-graph write` with the `--changed-paths` option, or in the\nfuture equivalent config option).\n\nTo provide some advise that can be acted upon: perhaps it would be\nbetter to start with \"It can store\", or end with \"if requested\" or\n\"optionally\".  For example the change could look like the following\nsuggestion:\n\n\n The Git commit graph stores a list of commit OIDs and some associated\n metadata, including:\n[...]\n+- The Bloom filter of the commit carrying the paths that were changed between\n+  the commit and its first parent, if requested.\n+\n\nBest,\n-- \nJakub Narębski\n"},{"id":"392470","messageId":"86fteyc1fi.fsf@gmail.com","threadId":"52499","inReplyTo":"56150788-c477-5526-2d6d-e9325ccb4da6@gmail.com","subject":"Re: [PATCH v2 09/11] commit-graph: add --changed-paths option to write subcommand","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2020-02-25T12:10:41Z","receivedAt":"2020-02-25T12:10:48Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Garima Singh <garimasigit@gmail.com> writes:\n> On 2/20/2020 3:28 PM, Jakub Narebski wrote:\n>> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n[...]\n>>> --- a/Documentation/git-commit-graph.txt\n>>> +++ b/Documentation/git-commit-graph.txt\n>>> @@ -54,6 +54,11 @@ or `--stdin-packs`.)\n>>>  With the `--append` option, include all commits that are present in the\n>>>  existing commit-graph file.\n>>>  +\n>>> +With the `--changed-paths` option, compute and write information about the\n>>> +paths changed between a commit and it's first parent. This operation can\n>>> +take a while on large repositories. It provides significant performance gains\n>>> +for getting history of a directory or a file with `git log -- <path>`.\n>>> ++\n>> \n>> Should we write about limitation that the topmost layer in the split\n>> commit graph needs to be written with '--changed-paths' for Git to use\n>> this information?  Or perhaps we should try (in the future) to remove\n>> this limitation?\n>\n> Given that this information is going to be used best effort, it would be \n> superfluous to describe every case and conditional that decides whether \n> this information is being used.\n\nI can somewhat agree with this reasoning.\n\nHowever what I would like to avoid is surprising users.  If one creates\nbase commit-graph with Bloom filters data, but then when creating\nnew layer of commit-graph (updating it incrementally), it may be\nsurprising that `git log -- <path>` is now much slower.\n\nOn the other hand if one would update commit-graph in a non-incremental\nway (rewriting the commit-graph file), loosing the Bloom filter\ninformation and performance of `git log -- <path>` because one forgot to\ninclude `--changed-paths` is not that unexpected.\n\nAnyway, in the future when this mechanism will be controlled by\nappropriate config variable, this whole discussion would become somewhat\nmoot.\n\n\nThought for the future: perhaps `git commit-graph verify` could detect\nthat split graph has Bloom filters only for some layers, and inform the\nuser?  But that is almost certainly out of scope of this patch series.\n\n>>> @@ -143,6 +146,8 @@ static int graph_write(int argc, const char **argv)\n>>>  \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n>>>  \tif (opts.progress)\n>>>  \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n>>> +\tif (opts.enable_changed_paths)\n>>> +\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n>>>  \n>>>  \tread_replace_refs = 0;\n>> \n>> All right.  This actually turns on calculation Bloom filters for changed\n>> paths, thanks to\n>> \n>>  \tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n>> \n>> that was added by the \"[PATCH v2 04/11] commit-graph: compute Bloom\n>> filters for changed paths\" patch.\n>> \n>> Though... should this enabling be split into two separate patches like\n>> this?\n>\n> The idea is that in 4/11 We compute only if the flag is set. \n> And between that patch and this one: we prepare the foundational code \n> that is now ready for that flag to be set via an opt-in by the user. \n\nAll right.\n\nChoosing how to split large change into series is not easy.  One one\nhand one would want for each change to be small and self contained.  On\nthe other hand it would be good if each change was testable (test-tool\ncan help here).\n\nBest,\n-- \nJakub Narębski\n"},{"id":"392472","messageId":"65c3bce0-40b5-6c25-fd4a-11429c7f2196@gmail.com","threadId":"52499","inReplyTo":"86sgiyc2ta.fsf@gmail.com","subject":"Re: [PATCH v2 07/11] commit-graph: write Bloom filters to commit graph file","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-02-25T15:58:49Z","receivedAt":"2020-02-25T15:58:56Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"\nOn 2/25/2020 6:40 AM, Jakub Narebski wrote:\n> Garima Singh <garimasigit@gmail.com> writes:\n>> On 2/19/2020 10:13 AM, Jakub Narebski wrote:\n>>> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> [...]\n>>>> diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n>>>> index a4f17441ae..22e511643d 100644\n>>>> --- a/Documentation/technical/commit-graph-format.txt\n>>>> +++ b/Documentation/technical/commit-graph-format.txt\n>>>> @@ -17,6 +17,9 @@ metadata, including:\n>>>>  - The parents of the commit, stored using positional references within\n>>>>    the graph file.\n>>>>  \n>>>> +- The Bloom filter of the commit carrying the paths that were changed between\n>>>> +  the commit and its first parent.\n>>>> +\n>>>\n>>> All right.\n>>>\n>>> Should we also state that it is optional (meta)data?  This would be\n>>> first optional piece of data stored in commit-graph, I think.\n>>>\n>>\n>> However the entire commit graph file is non critical metadata since git commands\n>> work just fine without it, just slower. The same applies to the changed path\n>> bloom filters. \n>>\n>> Based on the definition of optional you are suggesting, edge data is optional\n>> because not every commit-graph has octopus merges. \n> \n> Well, edge data (EDGE chunk) is optional in different way from Bloom\n> filter data.  The former depends on the repository (whether there are\n> octopus merges used), the latter is opt-in user choice (whether to run\n> `git commit-graph write` with the `--changed-paths` option, or in the\n> future equivalent config option).\n> \n> To provide some advise that can be acted upon: perhaps it would be\n> better to start with \"It can store\", or end with \"if requested\" or\n> \"optionally\".  For example the change could look like the following\n> suggestion:\n> \n> \n>  The Git commit graph stores a list of commit OIDs and some associated\n>  metadata, including:\n> [...]\n> +- The Bloom filter of the commit carrying the paths that were changed between\n> +  the commit and its first parent, if requested.\n> +\n> \n> Best,\n> \n\nSure. That makes sense. Will incorporate in v3. \n\nCheers!\nGarima Singh\n"},{"id":"392887","messageId":"e52a50f5-d050-3fd1-5014-4893375e2d7b@gmail.com","threadId":"52499","inReplyTo":"pull.497.git.1576879520.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/9] [RFC] Changed Paths Bloom Filters","fromName":"Garima Singh","fromEmail":"garimasigit@gmail.com","sentAt":"2020-03-05T19:49:39Z","receivedAt":"2020-03-05T19:49:44Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"My apologies that things have been quite on this series for the past\nweek or so. An unexpected high priority task at work demanded all of \nmy attention and will continue to do so through the end of this week. \n\nHopefully I will be able to pick this up again early next week and \nhave v3 out soon! \n\nCheers!\nGarima Singh\n\nOn 12/20/2019 5:05 PM, Garima Singh via GitGitGadget wrote:\n> Hey! \n> \n> The commit graph feature brought in a lot of performance improvements across\n> multiple commands. However, file based history continues to be a performance\n> pain point, especially in large repositories. \n> \n> Adopting changed path bloom filters has been discussed on the list before,\n> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n> Derrick Stolee [1]. This series is based on Dr. Stolee's approach [2] and\n> presents an updated and more polished RFC version of the feature. \n> \n> Performance Gains: We tested the performance of git log -- path on the git\n> repo, the linux repo and some internal large repos, with a variety of paths\n> of varying depths.\n> \n> On the git and linux repos: We observed a 2x to 5x speed up.\n> \n> On a large internal repo with files seated 6-10 levels deep in the tree: We\n> observed 10x to 20x speed ups, with some paths going up to 28 times faster.\n> \n> Future Work (not included in the scope of this series):\n> \n>  1. Supporting multiple path based revision walk\n>  2. Adopting it in git blame logic. \n>  3. Interactions with line log git log -L\n> \n> This series is intended to start the conversation and many of the commit\n> messages include specific call outs for suggestions and thoughts. \n> \n> Cheers! Garima Singh\n> \n> [1] https://lore.kernel.org/git/20181009193445.21908-1-szeder.dev@gmail.com/\n> [2] \n> https://lore.kernel.org/git/61559c5b-546e-d61b-d2e1-68de692f5972@gmail.com/\n> \n> Garima Singh (9):\n>   commit-graph: add --changed-paths option to write\n>   commit-graph: write changed paths bloom filters\n>   commit-graph: use MAX_NUM_CHUNKS\n>   commit-graph: document bloom filter format\n>   commit-graph: write changed path bloom filters to commit-graph file.\n>   commit-graph: test commit-graph write --changed-paths\n>   commit-graph: reuse existing bloom filters during write.\n>   revision.c: use bloom filters to speed up path based revision walks\n>   commit-graph: add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flag\n> \n>  Documentation/git-commit-graph.txt            |   5 +\n>  .../technical/commit-graph-format.txt         |  17 ++\n>  Makefile                                      |   1 +\n>  bloom.c                                       | 257 +++++++++++++++++\n>  bloom.h                                       |  51 ++++\n>  builtin/commit-graph.c                        |   9 +-\n>  ci/run-build-and-tests.sh                     |   1 +\n>  commit-graph.c                                | 116 +++++++-\n>  commit-graph.h                                |   9 +-\n>  revision.c                                    |  67 ++++-\n>  revision.h                                    |   5 +\n>  t/README                                      |   3 +\n>  t/helper/test-read-graph.c                    |   4 +\n>  t/t4216-log-bloom.sh                          |  77 ++++++\n>  t/t5318-commit-graph.sh                       |   2 +\n>  t/t5324-split-commit-graph.sh                 |   1 +\n>  t/t5325-commit-graph-bloom.sh                 | 258 ++++++++++++++++++\n>  17 files changed, 875 insertions(+), 8 deletions(-)\n>  create mode 100644 bloom.c\n>  create mode 100644 bloom.h\n>  create mode 100755 t/t4216-log-bloom.sh\n>  create mode 100755 t/t5325-commit-graph-bloom.sh\n> \n> \n> base-commit: b02fd2accad4d48078671adf38fe5b5976d77304\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-497%2Fgarimasi514%2FcoreGit-bloomFilters-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-497/garimasi514/coreGit-bloomFilters-v1\n> Pull-Request: https://github.com/gitgitgadget/git/pull/497\n> \n"},{"id":"394295","messageId":"xmqqtv279ffq.fsf@gitster.c.googlers.com","threadId":"52499","inReplyTo":"fdcbd793-57c2-f5ea-ccb9-cf34e911b669@gmail.com","subject":"Re: [PATCH v2 00/11] Changed Paths Bloom Filters","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-03-29T18:36:09Z","receivedAt":"2020-03-29T18:36:17Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Garima Singh <garimasigit@gmail.com> writes:\n\n> On 2/8/2020 6:04 PM, Jakub Narebski wrote:\n>> \"Garima Singh via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>> ...\n> I have gone back and forth on doing this. I like most of the core Bloom filter\n> computations being isolated in one patch/commit. But based on the rest of your\n> review, it seems like you are leaning heavily on having this split out. \n> So, I will take a proper stab at doing it for v3. \n> ...\n> Thanks for taking the time for reviewing this series so thoroughly! \n> It is greatly appreciated! \n\nThanks for a great discussion.  Just a friendly ping to the thread,\nso that something from the discussion thread will stay on the first\npage of mailing list archive's threaded view ;-)\n\n"},{"id":"394302","messageId":"c3ffd9820d50196793833d0faf35f9d5dd70ff6d.1585528298.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 01/16] commit-graph: define and use MAX_NUM_CHUNKS","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:23Z","receivedAt":"2020-03-30T00:31:46Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nThis is a minor cleanup to make it easier to change\nthe number of chunks being written to the commit\ngraph.\n\nReviewed-by: Jakub Narębski <jnareb@gmail.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 5 +++--\n 1 file changed, 3 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex f013a84e294..e4f1a5b2f1a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -23,6 +23,7 @@\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n+#define MAX_NUM_CHUNKS 5\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -1350,8 +1351,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tint fd;\n \tstruct hashfile *f;\n \tstruct lock_file lk = LOCK_INIT;\n-\tuint32_t chunk_ids[6];\n-\tuint64_t chunk_offsets[6];\n+\tuint32_t chunk_ids[MAX_NUM_CHUNKS + 1];\n+\tuint64_t chunk_offsets[MAX_NUM_CHUNKS + 1];\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \tstruct strbuf progress_title = STRBUF_INIT;\n \tint num_chunks = 3;\n-- \ngitgitgadget\n\n"},{"id":"394303","messageId":"a5aa3415c05ee9bc67a9471445a20c71a9834673.1585528298.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 02/16] bloom.c: add the murmur3 hash implementation","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:24Z","receivedAt":"2020-03-30T00:31:48Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nIn preparation for computing changed paths Bloom filters,\nimplement the Murmur3 hash algorithm as described in [1].\nIt hashes the given data using the given seed and produces\na uniformly distributed hash value.\n\n[1] https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nHelped-by: Szeder Gábor <szeder.dev@gmail.com>\nReviewed-by: Jakub Narębski <jnareb@gmail.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n Makefile              |  2 ++\n bloom.c               | 73 +++++++++++++++++++++++++++++++++++++++++++\n bloom.h               | 13 ++++++++\n t/helper/test-bloom.c | 13 ++++++++\n t/helper/test-tool.c  |  1 +\n t/helper/test-tool.h  |  1 +\n t/t0095-bloom.sh      | 30 ++++++++++++++++++\n 7 files changed, 133 insertions(+)\n create mode 100644 bloom.c\n create mode 100644 bloom.h\n create mode 100644 t/helper/test-bloom.c\n create mode 100755 t/t0095-bloom.sh\n\ndiff --git a/Makefile b/Makefile\nindex ef1ff2228f0..491f75e68c5 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -695,6 +695,7 @@ X =\n PROGRAMS += $(patsubst %.o,git-%$X,$(PROGRAM_OBJS))\n \n TEST_BUILTINS_OBJS += test-advise.o\n+TEST_BUILTINS_OBJS += test-bloom.o\n TEST_BUILTINS_OBJS += test-chmtime.o\n TEST_BUILTINS_OBJS += test-config.o\n TEST_BUILTINS_OBJS += test-ctype.o\n@@ -840,6 +841,7 @@ LIB_OBJS += base85.o\n LIB_OBJS += bisect.o\n LIB_OBJS += blame.o\n LIB_OBJS += blob.o\n+LIB_OBJS += bloom.o\n LIB_OBJS += branch.o\n LIB_OBJS += bulk-checkin.o\n LIB_OBJS += bundle.o\ndiff --git a/bloom.c b/bloom.c\nnew file mode 100644\nindex 00000000000..40e87632aeb\n--- /dev/null\n+++ b/bloom.c\n@@ -0,0 +1,73 @@\n+#include \"git-compat-util.h\"\n+#include \"bloom.h\"\n+\n+static uint32_t rotate_left(uint32_t value, int32_t count)\n+{\n+\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n+\tcount &= mask;\n+\treturn ((value << count) | (value >> ((-count) & mask)));\n+}\n+\n+/*\n+ * Calculate the murmur3 32-bit hash value for the given data\n+ * using the given seed.\n+ * Produces a uniformly distributed hash value.\n+ * Not considered to be cryptographically secure.\n+ * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n+ */\n+uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len)\n+{\n+\tconst uint32_t c1 = 0xcc9e2d51;\n+\tconst uint32_t c2 = 0x1b873593;\n+\tconst uint32_t r1 = 15;\n+\tconst uint32_t r2 = 13;\n+\tconst uint32_t m = 5;\n+\tconst uint32_t n = 0xe6546b64;\n+\tint i;\n+\tuint32_t k1 = 0;\n+\tconst char *tail;\n+\n+\tint len4 = len / sizeof(uint32_t);\n+\n+\tuint32_t k;\n+\tfor (i = 0; i < len4; i++) {\n+\t\tuint32_t byte1 = (uint32_t)data[4*i];\n+\t\tuint32_t byte2 = ((uint32_t)data[4*i + 1]) << 8;\n+\t\tuint32_t byte3 = ((uint32_t)data[4*i + 2]) << 16;\n+\t\tuint32_t byte4 = ((uint32_t)data[4*i + 3]) << 24;\n+\t\tk = byte1 | byte2 | byte3 | byte4;\n+\t\tk *= c1;\n+\t\tk = rotate_left(k, r1);\n+\t\tk *= c2;\n+\n+\t\tseed ^= k;\n+\t\tseed = rotate_left(seed, r2) * m + n;\n+\t}\n+\n+\ttail = (data + len4 * sizeof(uint32_t));\n+\n+\tswitch (len & (sizeof(uint32_t) - 1)) {\n+\tcase 3:\n+\t\tk1 ^= ((uint32_t)tail[2]) << 16;\n+\t\t/*-fallthrough*/\n+\tcase 2:\n+\t\tk1 ^= ((uint32_t)tail[1]) << 8;\n+\t\t/*-fallthrough*/\n+\tcase 1:\n+\t\tk1 ^= ((uint32_t)tail[0]) << 0;\n+\t\tk1 *= c1;\n+\t\tk1 = rotate_left(k1, r1);\n+\t\tk1 *= c2;\n+\t\tseed ^= k1;\n+\t\tbreak;\n+\t}\n+\n+\tseed ^= (uint32_t)len;\n+\tseed ^= (seed >> 16);\n+\tseed *= 0x85ebca6b;\n+\tseed ^= (seed >> 13);\n+\tseed *= 0xc2b2ae35;\n+\tseed ^= (seed >> 16);\n+\n+\treturn seed;\n+}\n\\ No newline at end of file\ndiff --git a/bloom.h b/bloom.h\nnew file mode 100644\nindex 00000000000..d0fcc5f0aa6\n--- /dev/null\n+++ b/bloom.h\n@@ -0,0 +1,13 @@\n+#ifndef BLOOM_H\n+#define BLOOM_H\n+\n+/*\n+ * Calculate the murmur3 32-bit hash value for the given data\n+ * using the given seed.\n+ * Produces a uniformly distributed hash value.\n+ * Not considered to be cryptographically secure.\n+ * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n+ */\n+uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len);\n+\n+#endif\n\\ No newline at end of file\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nnew file mode 100644\nindex 00000000000..60ee2043689\n--- /dev/null\n+++ b/t/helper/test-bloom.c\n@@ -0,0 +1,13 @@\n+#include \"git-compat-util.h\"\n+#include \"bloom.h\"\n+#include \"test-tool.h\"\n+\n+int cmd__bloom(int argc, const char **argv)\n+{\n+\tif (!strcmp(argv[1], \"get_murmur3\")) {\n+\t\tuint32_t hashed = murmur3_seeded(0, argv[2], strlen(argv[2]));\n+\t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n+\t}\n+\n+\treturn 0;\n+}\n\\ No newline at end of file\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex 31eedcd241f..6e26bd65c97 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -15,6 +15,7 @@ struct test_cmd {\n \n static struct test_cmd cmds[] = {\n \t{ \"advise\", cmd__advise_if_enabled },\n+\t{ \"bloom\", cmd__bloom },\n \t{ \"chmtime\", cmd__chmtime },\n \t{ \"config\", cmd__config },\n \t{ \"ctype\", cmd__ctype },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex 4eb5e6609e1..dceeef1d5c2 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -5,6 +5,7 @@\n #include \"git-compat-util.h\"\n \n int cmd__advise_if_enabled(int argc, const char **argv);\n+int cmd__bloom(int argc, const char **argv);\n int cmd__chmtime(int argc, const char **argv);\n int cmd__config(int argc, const char **argv);\n int cmd__ctype(int argc, const char **argv);\ndiff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\nnew file mode 100755\nindex 00000000000..2dad8c4a94e\n--- /dev/null\n+++ b/t/t0095-bloom.sh\n@@ -0,0 +1,30 @@\n+#!/bin/sh\n+\n+test_description='Testing the various Bloom filter computations in bloom.c'\n+. ./test-lib.sh\n+\n+test_expect_success 'compute unseeded murmur3 hash for empty string' '\n+\tcat >expect <<-\\EOF &&\n+\tMurmur3 Hash with seed=0:0x00000000\n+\tEOF\n+\ttest-tool bloom get_murmur3 \"\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute unseeded murmur3 hash for test string 1' '\n+\tcat >expect <<-\\EOF &&\n+\tMurmur3 Hash with seed=0:0x627b0c2c\n+\tEOF\n+\ttest-tool bloom get_murmur3 \"Hello world!\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute unseeded murmur3 hash for test string 2' '\n+\tcat >expect <<-\\EOF &&\n+\tMurmur3 Hash with seed=0:0x2e4ff723\n+\tEOF\n+\ttest-tool bloom get_murmur3 \"The quick brown fox jumps over the lazy dog\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_done\n\\ No newline at end of file\n-- \ngitgitgadget\n\n"},{"id":"394304","messageId":"a7702c1afde1ce9d4c628a831927b4e75bff8515.1585528298.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 03/16] bloom.c: introduce core Bloom filter constructs","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:25Z","receivedAt":"2020-03-30T00:31:49Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nIntroduce the constructs for Bloom filters, Bloom filter keys\nand Bloom filter settings.\nFor details on what Bloom filters are and how they work, refer\nto Dr. Derrick Stolee's blog post [1]. It provides a concise\nexplanation of the adoption of Bloom filters as described in\n[2] and [3].\n\nImplementation specifics:\n1. We currently use 7 and 10 for the number of hashes and the\n   size of each entry respectively. They served as great starting\n   values, the mathematical details behind this choice are\n   described in [1] and [4]. The implementation, while not\n   completely open to it at the moment, is flexible enough to allow\n   for tweaking these settings in the future.\n\n   Note: The performance gains we have observed with these values\n   are significant enough that we did not need to tweak these\n   settings. The performance numbers are included in the cover letter\n   of this series and in the commit message of the subsequent commit\n   where we use Bloom filters to speed up `git log -- path`.\n\n2. As described in [1] and [3], we do not need 7 independent hashing\n   functions. We use the Murmur3 hashing scheme, seed it twice and\n   then combine those to procure an arbitrary number of hash values.\n\n3. The filters will be sized according to the number of changes in\n   each commit, in multiples of 8 bit words.\n\n[1] Derrick Stolee\n      \"Supercharging the Git Commit Graph IV: Bloom Filters\"\n      https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n\n[2] Flavio Bonomi, Michael Mitzenmacher, Rina Panigrahy, Sushil Singh, George Varghese\n    \"An Improved Construction for Counting Bloom Filters\"\n    http://theory.stanford.edu/~rinap/papers/esa2006b.pdf\n    https://doi.org/10.1007/11841036_61\n\n[3] Peter C. Dillinger and Panagiotis Manolios\n    \"Bloom Filters in Probabilistic Verification\"\n    http://www.ccs.neu.edu/home/pete/pub/Bloom-filters-verification.pdf\n    https://doi.org/10.1007/978-3-540-30494-4_26\n\n[4] Thomas Mueller Graf, Daniel Lemire\n    \"Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters\"\n    https://arxiv.org/abs/1912.08258\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nReviewed-by: Jakub Narębski <jnareb@gmail.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c               | 38 +++++++++++++++++++++++++-\n bloom.h               | 63 +++++++++++++++++++++++++++++++++++++++++++\n t/helper/test-bloom.c | 48 +++++++++++++++++++++++++++++++++\n t/t0095-bloom.sh      | 40 +++++++++++++++++++++++++++\n 4 files changed, 188 insertions(+), 1 deletion(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 40e87632aeb..888b67f1ea6 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -8,6 +8,11 @@ static uint32_t rotate_left(uint32_t value, int32_t count)\n \treturn ((value << count) | (value >> ((-count) & mask)));\n }\n \n+static inline unsigned char get_bitmask(uint32_t pos)\n+{\n+\treturn ((unsigned char)1) << (pos & (BITS_PER_WORD - 1));\n+}\n+\n /*\n  * Calculate the murmur3 32-bit hash value for the given data\n  * using the given seed.\n@@ -70,4 +75,35 @@ uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len)\n \tseed ^= (seed >> 16);\n \n \treturn seed;\n-}\n\\ No newline at end of file\n+}\n+\n+void fill_bloom_key(const char *data,\n+\t\t\t\t\tsize_t len,\n+\t\t\t\t\tstruct bloom_key *key,\n+\t\t\t\t\tconst struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tconst uint32_t seed0 = 0x293ae76f;\n+\tconst uint32_t seed1 = 0x7e646e2c;\n+\tconst uint32_t hash0 = murmur3_seeded(seed0, data, len);\n+\tconst uint32_t hash1 = murmur3_seeded(seed1, data, len);\n+\n+\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n+\tfor (i = 0; i < settings->num_hashes; i++)\n+\t\tkey->hashes[i] = hash0 + i * hash1;\n+}\n+\n+void add_key_to_filter(const struct bloom_key *key,\n+\t\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t\t   const struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tuint64_t mod = filter->len * BITS_PER_WORD;\n+\n+\tfor (i = 0; i < settings->num_hashes; i++) {\n+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n+\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n+\n+\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n+\t}\n+}\ndiff --git a/bloom.h b/bloom.h\nindex d0fcc5f0aa6..b9ce422ca2d 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -1,6 +1,60 @@\n #ifndef BLOOM_H\n #define BLOOM_H\n \n+struct bloom_filter_settings {\n+\t/*\n+\t * The version of the hashing technique being used.\n+\t * We currently only support version = 1 which is\n+\t * the seeded murmur3 hashing technique implemented\n+\t * in bloom.c.\n+\t */\n+\tuint32_t hash_version;\n+\n+\t/*\n+\t * The number of times a path is hashed, i.e. the\n+\t * number of bit positions tht cumulatively\n+\t * determine whether a path is present in the\n+\t * Bloom filter.\n+\t */\n+\tuint32_t num_hashes;\n+\n+\t/*\n+\t * The minimum number of bits per entry in the Bloom\n+\t * filter. If the filter contains 'n' entries, then\n+\t * filter size is the minimum number of 8-bit words\n+\t * that contain n*b bits.\n+\t */\n+\tuint32_t bits_per_entry;\n+};\n+\n+#define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n+#define BITS_PER_WORD 8\n+\n+/*\n+ * A bloom_filter struct represents a data segment to\n+ * use when testing hash values. The 'len' member\n+ * dictates how many entries are stored in\n+ * 'data'.\n+ */\n+struct bloom_filter {\n+\tunsigned char *data;\n+\tsize_t len;\n+};\n+\n+/*\n+ * A bloom_key represents the k hash values for a\n+ * given string. These can be precomputed and\n+ * stored in a bloom_key for re-use when testing\n+ * against a bloom_filter. The number of hashes is\n+ * given by the Bloom filter settings and is the same\n+ * for all Bloom filters and keys interacting with\n+ * the loaded version of the commit graph file and\n+ * the Bloom data chunks.\n+ */\n+struct bloom_key {\n+\tuint32_t *hashes;\n+};\n+\n /*\n  * Calculate the murmur3 32-bit hash value for the given data\n  * using the given seed.\n@@ -10,4 +64,13 @@\n  */\n uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len);\n \n+void fill_bloom_key(const char *data,\n+\t\t    size_t len,\n+\t\t    struct bloom_key *key,\n+\t\t    const struct bloom_filter_settings *settings);\n+\n+void add_key_to_filter(const struct bloom_key *key,\n+\t\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t\t   const struct bloom_filter_settings *settings);\n+\n #endif\n\\ No newline at end of file\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 60ee2043689..20460cde775 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -2,6 +2,36 @@\n #include \"bloom.h\"\n #include \"test-tool.h\"\n \n+struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\n+static void add_string_to_filter(const char *data, struct bloom_filter *filter) {\n+\t\tstruct bloom_key key;\n+\t\tint i;\n+\n+\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n+\t\tprintf(\"Hashes:\");\n+\t\tfor (i = 0; i < settings.num_hashes; i++){\n+\t\t\tprintf(\"0x%08x|\", key.hashes[i]);\n+\t\t}\n+\t\tprintf(\"\\n\");\n+\t\tadd_key_to_filter(&key, filter, &settings);\n+}\n+\n+static void print_bloom_filter(struct bloom_filter *filter) {\n+\tint i;\n+\n+\tif (!filter) {\n+\t\tprintf(\"No filter.\\n\");\n+\t\treturn;\n+\t}\n+\tprintf(\"Filter_Length:%d\\n\", (int)filter->len);\n+\tprintf(\"Filter_Data:\");\n+\tfor (i = 0; i < filter->len; i++){\n+\t\tprintf(\"%02x|\", filter->data[i]);\n+\t}\n+\tprintf(\"\\n\");\n+}\n+\n int cmd__bloom(int argc, const char **argv)\n {\n \tif (!strcmp(argv[1], \"get_murmur3\")) {\n@@ -9,5 +39,23 @@ int cmd__bloom(int argc, const char **argv)\n \t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n \t}\n \n+    if (!strcmp(argv[1], \"generate_filter\")) {\n+\t\tstruct bloom_filter filter;\n+\t\tint i = 2;\n+\t\tfilter.len =  (settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n+\t\tfilter.data = xcalloc(filter.len, sizeof(unsigned char));\n+\n+\t\tif (!argv[2]){\n+\t\t\tdie(\"at least one input string expected\");\n+\t\t}\n+\n+\t\twhile (argv[i]) {\n+\t\t\tadd_string_to_filter(argv[i], &filter);\n+\t\t\ti++;\n+\t\t}\n+\n+\t\tprint_bloom_filter(&filter);\n+\t}\n+\n \treturn 0;\n }\n\\ No newline at end of file\ndiff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\nindex 2dad8c4a94e..36a086c7c60 100755\n--- a/t/t0095-bloom.sh\n+++ b/t/t0095-bloom.sh\n@@ -27,4 +27,44 @@ test_expect_success 'compute unseeded murmur3 hash for test string 2' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'compute bloom key for empty string' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:0x5615800c|0x5b966560|0x61174ab4|0x66983008|0x6c19155c|0x7199fab0|0x771ae004|\n+\tFilter_Length:2\n+\tFilter_Data:11|11|\n+\tEOF\n+\ttest-tool bloom generate_filter \"\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for whitespace' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:0xf178874c|0x5f3d6eb6|0xcd025620|0x3ac73d8a|0xa88c24f4|0x16510c5e|0x8415f3c8|\n+\tFilter_Length:2\n+\tFilter_Data:51|55|\n+\tEOF\n+\ttest-tool bloom generate_filter \" \" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for test string 1' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:0xb270de9b|0x1bb6f26e|0x84fd0641|0xee431a14|0x57892de7|0xc0cf41ba|0x2a15558d|\n+\tFilter_Length:2\n+\tFilter_Data:92|6c|\n+\tEOF\n+\ttest-tool bloom generate_filter \"Hello world!\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for test string 2' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:0x20ab385b|0xf5237fe2|0xc99bc769|0x9e140ef0|0x728c5677|0x47049dfe|0x1b7ce585|\n+\tFilter_Length:2\n+\tFilter_Data:a5|4a|\n+\tEOF\n+\ttest-tool bloom generate_filter \"file.txt\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_done\n\\ No newline at end of file\n-- \ngitgitgadget\n\n"},{"id":"394308","messageId":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v2.git.1580943390.gitgitgadget@gmail.com","subject":"[PATCH v3 00/16] Changed Paths Bloom Filters","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:22Z","receivedAt":"2020-03-30T00:31:49Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"Hey! \n\nThe commit graph feature brought in a lot of performance improvements across\nmultiple commands. However, file based history continues to be a performance\npain point, especially in large repositories. \n\nAdopting changed path Bloom filters has been discussed on the list before,\nand a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\nDerrick Stolee [1]. This series is based on Dr. Stolee's proof of concept in\n[2]\n\nWith the changes in this series, git users will be able to choose to write\nBloom filters to the commit-graph using the following command:\n\n'git commit-graph write --changed-paths'\n\nSubsequent 'git log -- path' commands will use these computed Bloom filters\nto decided which commits are worth exploring further to produce the history\nof the provided path. \n\nCost of computing and writing Bloom filters\n===========================================\n\nComputing and writing Bloom filters to the commit graph for the first time\nimplies computing the diffs and the resulting Bloom filters for all the\ncommits in the repository. This adds a non trivial amount of time to run\ntime. Every subsequent run is incremental i.e. we reuse the previously\ncomputed Bloom filters. So this is a one time cost. \n\nTime taken by 'git commit-graph write' with and w/o --changed-paths, speed\nup in 'git log -- path' with computed Bloom filters (see a):- \n\n-------------------------------------------------------------------------\n| Repo        | w/o --changed-paths | with --changed-paths | Speed up   |\n-------------------------------------------------------------------------\n| git [3]     | 0.9 seconds         | 7 seconds            | 2x to 6x   |\n| linux [4]   | 16 seconds          | 1 minute 8 seconds   | 2x to 6x   | \n| android [5] | 9 seconds           | 48 seconds           | 2x to 6x   |\n| AzDo(see b) | 1 minute            | 5 minutes 2 seconds  | 10x to 30x |\n-------------------------------------------------------------------------\n\na) We tested the performance of git log -- path with randomly chosen paths\nof varying depths in each repo. The speed up depends on how deep the files\nare in the hierarchy and how often a file has been touched in the commit\nhistory.\n\nb) This internal repository has about 420k commits, 183k files distributed\nacross 34k folders, the size on disk is about 17 GiB. The most massive gains\non this repository were for files 6-12 levels deep in the tree. \n\nc) These numbers were collected on a Windows machine, except for the linux\nrepo which was tested on a Linux machine. \n\nFuture Work (not included in the scope of this series)\n======================================================\n\n 1. Supporting multiple path based revision walk\n 2. Adopting it in git blame logic. \n 3. Interactions with line log git log -L\n\nCheers! Garima Singh\n\n[1] https://lore.kernel.org/git/20181009193445.21908-1-szeder.dev@gmail.com/\n\n[2] \nhttps://lore.kernel.org/git/61559c5b-546e-d61b-d2e1-68de692f5972@gmail.com/\n\n[3] https://github.com/git/git\n\n[4] https://github.com/torvalds/linux\n\n[5] https://android.googlesource.com/platform/frameworks/base/\n\njeffhost@microsoft.com, me@ttaylorr.com, peff@peff.net, \ngarimasigit@gmail.com,jnareb@gmail.com, christian.couder@gmail.com, \nemilyshaffer@gmail.com,gitster@pobox.com\n\nDerrick Stolee (1):\n  diff: halt tree-diff early after max_changes\n\nGarima Singh (14):\n  commit-graph: define and use MAX_NUM_CHUNKS\n  bloom.c: add the murmur3 hash implementation\n  bloom.c: introduce core Bloom filter constructs\n  bloom.c: core Bloom filter implementation for changed paths.\n  commit-graph: compute Bloom filters for changed paths\n  commit-graph: examine commits by generation number\n  diff: skip batch object download when possible\n  commit-graph: write Bloom filters to commit graph file\n  commit-graph: reuse existing Bloom filters during write\n  commit-graph: add --changed-paths option to write subcommand\n  revision.c: use Bloom filters to speed up path based revision walks\n  revision.c: add trace2 stats around Bloom filter usage\n  t4216: add end to end tests for git log with Bloom filters\n  commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag\n\nJeff King (1):\n  commit-graph: examine changed-path objects in pack order\n\n Documentation/git-commit-graph.txt            |   5 +\n .../technical/commit-graph-format.txt         |  30 ++\n Makefile                                      |   2 +\n bloom.c                                       | 276 ++++++++++++++++++\n bloom.h                                       |  90 ++++++\n builtin/commit-graph.c                        |  10 +-\n ci/run-build-and-tests.sh                     |   1 +\n commit-graph.c                                | 213 +++++++++++++-\n commit-graph.h                                |   9 +-\n diff.c                                        |   8 +-\n diff.h                                        |   6 +\n revision.c                                    | 126 +++++++-\n revision.h                                    |  11 +\n t/README                                      |   5 +\n t/helper/test-bloom.c                         |  81 +++++\n t/helper/test-read-graph.c                    |   4 +\n t/helper/test-tool.c                          |   1 +\n t/helper/test-tool.h                          |   1 +\n t/t0095-bloom.sh                              | 117 ++++++++\n t/t4216-log-bloom.sh                          | 155 ++++++++++\n t/t5318-commit-graph.sh                       |   2 +\n t/t5324-split-commit-graph.sh                 |   1 +\n tree-diff.c                                   |   6 +\n 23 files changed, 1148 insertions(+), 12 deletions(-)\n create mode 100644 bloom.c\n create mode 100644 bloom.h\n create mode 100644 t/helper/test-bloom.c\n create mode 100755 t/t0095-bloom.sh\n create mode 100755 t/t4216-log-bloom.sh\n\n\nbase-commit: 3bab5d56259722843359702bc27111475437ad2a\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-497%2Fgarimasi514%2FcoreGit-bloomFilters-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-497/garimasi514/coreGit-bloomFilters-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/497\n\nRange-diff vs v2:\n\n  1:  bf6b93878af !  1:  c3ffd9820d5 commit-graph: use MAX_NUM_CHUNKS\n     @@ -1,10 +1,12 @@\n      Author: Garima Singh <garima.singh@microsoft.com>\n      \n     -    commit-graph: use MAX_NUM_CHUNKS\n     +    commit-graph: define and use MAX_NUM_CHUNKS\n      \n     -    This is a minor cleanup to make it easier to change the\n     -    number of chunks being written to the commit-graph in the future.\n     +    This is a minor cleanup to make it easier to change\n     +    the number of chunks being written to the commit\n     +    graph.\n      \n     +    Reviewed-by: Jakub Narębski <jnareb@gmail.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n      \n       diff --git a/commit-graph.c b/commit-graph.c\n  -:  ----------- >  2:  a5aa3415c05 bloom.c: add the murmur3 hash implementation\n  -:  ----------- >  3:  a7702c1afde bloom.c: introduce core Bloom filter constructs\n  2:  02b16d94227 !  4:  8304c297520 bloom: core Bloom filter implementation for changed paths\n     @@ -1,89 +1,33 @@\n      Author: Garima Singh <garima.singh@microsoft.com>\n      \n     -    bloom: core Bloom filter implementation for changed paths\n     +    bloom.c: core Bloom filter implementation for changed paths.\n      \n     -    Add the core Bloom filter logic for computing the paths changed between a\n     -    commit and its first parent. For details on what Bloom filters are and how they\n     -    work, please refer to Dr. Derrick Stolee's blog post [1]. It provides a concise\n     -    explaination of the adoption of Bloom filters as described in [2] and [3]\n     +    Add the core implementation for computing Bloom filters for\n     +    the paths changed between a commit and it's first parent.\n      \n     -    1. We currently use 7 and 10 for the number of hashes and the size of each\n     -       entry respectively. They served as great starting values, the mathematical\n     -       details behind this choice are described in [1] and [4]. The implementation\n     -       while not completely open to it at the moment, is flexible enough to allow\n     -       for tweaking these settings in the future.\n     +    We fill the Bloom filters as (const char *data, int len) pairs\n     +    as `struct bloom_filters\" within a commit slab.\n      \n     -       Note: The performance gains we have observed with these values are\n     -       significant enough that we did not need to tweak these settings.\n     -       The performance numbers are included in the cover letter of this series\n     -       and in the message of a subsequent commit where we use Bloom filters in\n     -       to speed up `git log -- <path>`.\n     -\n     -    2. As described in the blog and in [3], we do not need 7 independent hashing\n     -       functions. We use the Murmur3 hashing scheme. Seed it twice and then\n     -       combine those to procure an arbitrary number of hash values.\n     -\n     -    3. The filters are sized according to the number of changes in the each commit,\n     -       with minimum size of one 64 bit word.\n     -\n     -    4. We fill the Bloom filters as (const char *data, int len) pairs as\n     -       \"struct bloom_filter\"s in a commit slab.\n     -\n     -    5. The seed_murmur3 method is implemented as described in [5]. It hashes the\n     -       given data using a given seed and produces a uniformly distributed hash\n     -       value.\n     -\n     -    [1] https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n     -\n     -    [2] Flavio Bonomi, Michael Mitzenmacher, Rina Panigrahy, Sushil Singh, George Varghese\n     -        \"An Improved Construction for Counting Bloom Filters\"\n     -        http://theory.stanford.edu/~rinap/papers/esa2006b.pdf\n     -        https://doi.org/10.1007/11841036_61\n     -\n     -    [3] Peter C. Dillinger and Panagiotis Manolios\n     -        \"Bloom Filters in Probabilistic Verification\"\n     -        http://www.ccs.neu.edu/home/pete/pub/Bloom-filters-verification.pdf\n     -        https://doi.org/10.1007/978-3-540-30494-4_26\n     -\n     -    [4] Thomas Mueller Graf, Daniel Lemire\n     -        \"Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters\"\n     -        https://arxiv.org/abs/1912.08258\n     -\n     -    [5] https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n     +    Filters for commits with no changes and more than 512 changes,\n     +    is represented with a filter of length zero. There is no gain\n     +    in distinguishing between a computed filter of length zero for\n     +    a commit with no changes, and an uncomputed filter for new commits\n     +    or for commits with more than 512 changes. The effect on\n     +    `git log -- path` is the same in both cases. We will fall back to\n     +    the normal diffing algorithm when we can't benefit from the\n     +    existence of Bloom filters.\n      \n          Helped-by: Jeff King <peff@peff.net>\n          Helped-by: Derrick Stolee <dstolee@microsoft.com>\n     +    Reviewed-by: Jakub Narębski <jnareb@gmail.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n      \n     - diff --git a/Makefile b/Makefile\n     - --- a/Makefile\n     - +++ b/Makefile\n     -@@\n     - \n     - PROGRAMS += $(patsubst %.o,git-%$X,$(PROGRAM_OBJS))\n     - \n     -+TEST_BUILTINS_OBJS += test-bloom.o\n     - TEST_BUILTINS_OBJS += test-chmtime.o\n     - TEST_BUILTINS_OBJS += test-config.o\n     - TEST_BUILTINS_OBJS += test-ctype.o\n     -@@\n     - LIB_OBJS += bisect.o\n     - LIB_OBJS += blame.o\n     - LIB_OBJS += blob.o\n     -+LIB_OBJS += bloom.o\n     - LIB_OBJS += branch.o\n     - LIB_OBJS += bulk-checkin.o\n     - LIB_OBJS += bundle.o\n     -\n       diff --git a/bloom.c b/bloom.c\n     - new file mode 100644\n     - --- /dev/null\n     + --- a/bloom.c\n       +++ b/bloom.c\n      @@\n     -+#include \"git-compat-util.h\"\n     -+#include \"bloom.h\"\n     -+#include \"commit-graph.h\"\n     -+#include \"object-store.h\"\n     + #include \"git-compat-util.h\"\n     + #include \"bloom.h\"\n      +#include \"diff.h\"\n      +#include \"diffcore.h\"\n      +#include \"revision.h\"\n     @@ -97,118 +41,19 @@\n      +    struct hashmap_entry entry;\n      +    const char path[FLEX_ARRAY];\n      +};\n     + \n     + static uint32_t rotate_left(uint32_t value, int32_t count)\n     + {\n     +@@\n     + \t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n     + \t}\n     + }\n      +\n     -+static uint32_t rotate_right(uint32_t value, int32_t count)\n     -+{\n     -+\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n     -+\tcount &= mask;\n     -+\treturn ((value >> count) | (value << ((-count) & mask)));\n     -+}\n     -+\n     -+/*\n     -+ * Calculate a hash value for the given data using the given seed.\n     -+ * Produces a uniformly distributed hash value.\n     -+ * Not considered to be cryptographically secure.\n     -+ * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n     -+ **/\n     -+static uint32_t seed_murmur3(uint32_t seed, const char *data, int len)\n     -+{\n     -+\tconst uint32_t c1 = 0xcc9e2d51;\n     -+\tconst uint32_t c2 = 0x1b873593;\n     -+\tconst uint32_t r1 = 15;\n     -+\tconst uint32_t r2 = 13;\n     -+\tconst uint32_t m = 5;\n     -+\tconst uint32_t n = 0xe6546b64;\n     -+\tint i;\n     -+\tuint32_t k1 = 0;\n     -+\tconst char *tail;\n     -+\n     -+\tint len4 = len / sizeof(uint32_t);\n     -+\n     -+\tconst uint32_t *blocks = (const uint32_t*)data;\n     -+\n     -+\tuint32_t k;\n     -+\tfor (i = 0; i < len4; i++)\n     -+\t{\n     -+\t\tk = blocks[i];\n     -+\t\tk *= c1;\n     -+\t\tk = rotate_right(k, r1);\n     -+\t\tk *= c2;\n     -+\n     -+\t\tseed ^= k;\n     -+\t\tseed = rotate_right(seed, r2) * m + n;\n     -+\t}\n     -+\n     -+\ttail = (data + len4 * sizeof(uint32_t));\n     -+\n     -+\tswitch (len & (sizeof(uint32_t) - 1))\n     -+\t{\n     -+\tcase 3:\n     -+\t\tk1 ^= ((uint32_t)tail[2]) << 16;\n     -+\t\t/*-fallthrough*/\n     -+\tcase 2:\n     -+\t\tk1 ^= ((uint32_t)tail[1]) << 8;\n     -+\t\t/*-fallthrough*/\n     -+\tcase 1:\n     -+\t\tk1 ^= ((uint32_t)tail[0]) << 0;\n     -+\t\tk1 *= c1;\n     -+\t\tk1 = rotate_right(k1, r1);\n     -+\t\tk1 *= c2;\n     -+\t\tseed ^= k1;\n     -+\t\tbreak;\n     -+\t}\n     -+\n     -+\tseed ^= (uint32_t)len;\n     -+\tseed ^= (seed >> 16);\n     -+\tseed *= 0x85ebca6b;\n     -+\tseed ^= (seed >> 13);\n     -+\tseed *= 0xc2b2ae35;\n     -+\tseed ^= (seed >> 16);\n     -+\n     -+\treturn seed;\n     -+}\n     -+\n     -+static inline uint64_t get_bitmask(uint32_t pos)\n     -+{\n     -+\treturn ((uint64_t)1) << (pos & (BITS_PER_WORD - 1));\n     -+}\n     -+\n     -+void load_bloom_filters(void)\n     ++void init_bloom_filters(void)\n      +{\n      +\tinit_bloom_filter_slab(&bloom_filters);\n      +}\n      +\n     -+void fill_bloom_key(const char *data,\n     -+\t\t\t\t\tint len,\n     -+\t\t\t\t\tstruct bloom_key *key,\n     -+\t\t\t\t\tstruct bloom_filter_settings *settings)\n     -+{\n     -+\tint i;\n     -+\tconst uint32_t seed0 = 0x293ae76f;\n     -+\tconst uint32_t seed1 = 0x7e646e2c;\n     -+\tconst uint32_t hash0 = seed_murmur3(seed0, data, len);\n     -+\tconst uint32_t hash1 = seed_murmur3(seed1, data, len);\n     -+\n     -+\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n     -+\tfor (i = 0; i < settings->num_hashes; i++)\n     -+\t\tkey->hashes[i] = hash0 + i * hash1;\n     -+}\n     -+\n     -+void add_key_to_filter(struct bloom_key *key,\n     -+\t\t\t\t\t   struct bloom_filter *filter,\n     -+\t\t\t\t\t   struct bloom_filter_settings *settings)\n     -+{\n     -+\tint i;\n     -+\tuint64_t mod = filter->len * BITS_PER_WORD;\n     -+\n     -+\tfor (i = 0; i < settings->num_hashes; i++) {\n     -+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n     -+\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n     -+\n     -+\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n     -+\t}\n     -+}\n     -+\n      +struct bloom_filter *get_bloom_filter(struct repository *r,\n      +\t\t\t\t      struct commit *c)\n      +{\n     @@ -217,7 +62,7 @@\n      +\tint i;\n      +\tstruct diff_options diffopt;\n      +\n     -+\tif (!bloom_filters.slab_size)\n     ++\tif (bloom_filters.slab_size == 0)\n      +\t\treturn NULL;\n      +\n      +\tfilter = bloom_filter_slab_at(&bloom_filters, c);\n     @@ -234,13 +79,12 @@\n      +\n      +\tif (diff_queued_diff.nr <= 512) {\n      +\t\tstruct hashmap pathmap;\n     -+\t\tstruct pathmap_hash_entry* e;\n     ++\t\tstruct pathmap_hash_entry *e;\n      +\t\tstruct hashmap_iter iter;\n      +\t\thashmap_init(&pathmap, NULL, NULL, 0);\n      +\n      +\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n     -+\t\t\tconst char* path = diff_queued_diff.queue[i]->two->path;\n     -+\t\t\tconst char* p = path;\n     ++\t\t\tconst char *path = diff_queued_diff.queue[i]->two->path;\n      +\n      +\t\t\t/*\n      +\t\t\t* Add each leading directory of the changed file, i.e. for\n     @@ -251,23 +95,23 @@\n      +\t\t\t* Note that directories are added without the trailing '/'.\n      +\t\t\t*/\n      +\t\t\tdo {\n     -+\t\t\t\tchar* last_slash = strrchr(p, '/');\n     ++\t\t\t\tchar *last_slash = strrchr(path, '/');\n      +\n      +\t\t\t\tFLEX_ALLOC_STR(e, path, path);\n     -+\t\t\t\thashmap_entry_init(&e->entry, strhash(p));\n     ++\t\t\t\thashmap_entry_init(&e->entry, strhash(path));\n      +\t\t\t\thashmap_add(&pathmap, &e->entry);\n      +\n      +\t\t\t\tif (!last_slash)\n     -+\t\t\t\t\tlast_slash = (char*)p;\n     ++\t\t\t\t\tlast_slash = (char*)path;\n      +\t\t\t\t*last_slash = '\\0';\n      +\n     -+\t\t\t} while (*p);\n     ++\t\t\t} while (*path);\n      +\n      +\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n      +\t\t}\n      +\n      +\t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n     -+\t\tfilter->data = xcalloc(filter->len, sizeof(uint64_t));\n     ++\t\tfilter->data = xcalloc(filter->len, sizeof(unsigned char));\n      +\n      +\t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n      +\t\t\tstruct bloom_key key;\n     @@ -287,138 +131,48 @@\n      +\tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n      +\n      +\treturn filter;\n     -+}\n     -+\n     -+int bloom_filter_contains(struct bloom_filter *filter,\n     -+\t\t\t  struct bloom_key *key,\n     -+\t\t\t  struct bloom_filter_settings *settings)\n     -+{\n     -+\tint i;\n     -+\tuint64_t mod = filter->len * BITS_PER_WORD;\n     -+\n     -+\tif (!mod)\n     -+\t\treturn -1;\n     -+\n     -+\tfor (i = 0; i < settings->num_hashes; i++) {\n     -+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n     -+\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n     -+\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n     -+\t\t\treturn 0;\n     -+\t}\n     -+\n     -+\treturn 1;\n      +}\n      \n       diff --git a/bloom.h b/bloom.h\n     - new file mode 100644\n     - --- /dev/null\n     + --- a/bloom.h\n       +++ b/bloom.h\n      @@\n     -+#ifndef BLOOM_H\n     -+#define BLOOM_H\n     -+\n     + #ifndef BLOOM_H\n     + #define BLOOM_H\n     + \n      +struct commit;\n      +struct repository;\n     -+struct commit_graph;\n     -+\n     -+struct bloom_filter_settings {\n     -+\tuint32_t hash_version;\n     -+\tuint32_t num_hashes;\n     -+\tuint32_t bits_per_entry;\n     -+};\n     -+\n     -+#define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n     -+#define BITS_PER_WORD 64\n      +\n     -+/*\n     -+ * A bloom_filter struct represents a data segment to\n     -+ * use when testing hash values. The 'len' member\n     -+ * dictates how many uint64_t entries are stored in\n     -+ * 'data'.\n     -+ */\n     -+struct bloom_filter {\n     -+\tuint64_t *data;\n     -+\tint len;\n     -+};\n     -+\n     -+/*\n     -+ * A bloom_key represents the k hash values for a\n     -+ * given hash input. These can be precomputed and\n     -+ * stored in a bloom_key for re-use when testing\n     -+ * against a bloom_filter.\n     -+ */\n     -+struct bloom_key {\n     -+\tuint32_t *hashes;\n     -+};\n     -+\n     -+void load_bloom_filters(void);\n     -+\n     -+void fill_bloom_key(const char *data,\n     -+\t\t    int len,\n     -+\t\t    struct bloom_key *key,\n     -+\t\t    struct bloom_filter_settings *settings);\n     -+\n     -+void add_key_to_filter(struct bloom_key *key,\n     -+\t\t\t\t\t   struct bloom_filter *filter,\n     -+\t\t\t\t\t   struct bloom_filter_settings *settings);\n     + struct bloom_filter_settings {\n     + \t/*\n     + \t * The version of the hashing technique being used.\n     +@@\n     + \t\t\t\t\t   struct bloom_filter *filter,\n     + \t\t\t\t\t   const struct bloom_filter_settings *settings);\n     + \n     ++void init_bloom_filters(void);\n      +\n      +struct bloom_filter *get_bloom_filter(struct repository *r,\n      +\t\t\t\t      struct commit *c);\n      +\n     -+int bloom_filter_contains(struct bloom_filter *filter,\n     -+\t\t\t  struct bloom_key *key,\n     -+\t\t\t  struct bloom_filter_settings *settings);\n     -+\n     -+#endif\n     + #endif\n     + \\ No newline at end of file\n      \n       diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n     - new file mode 100644\n     - --- /dev/null\n     + --- a/t/helper/test-bloom.c\n       +++ b/t/helper/test-bloom.c\n      @@\n     -+#include \"test-tool.h\"\n     -+#include \"git-compat-util.h\"\n     -+#include \"bloom.h\"\n     -+#include \"test-tool.h\"\n     -+#include \"cache.h\"\n     -+#include \"commit-graph.h\"\n     + #include \"git-compat-util.h\"\n     + #include \"bloom.h\"\n     + #include \"test-tool.h\"\n      +#include \"commit.h\"\n     -+#include \"config.h\"\n     -+#include \"object-store.h\"\n     -+#include \"object.h\"\n     -+#include \"repository.h\"\n     -+#include \"tree.h\"\n     -+\n     -+struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n     -+\n     -+static void print_bloom_filter(struct bloom_filter *filter) {\n     -+\tint i;\n     -+\n     -+\tif (!filter) {\n     -+\t\tprintf(\"No filter.\\n\");\n     -+\t\treturn;\n     -+\t}\n     -+\tprintf(\"Filter_Length:%d\\n\", filter->len);\n     -+\tprintf(\"Filter_Data:\");\n     -+\tfor (i = 0; i < filter->len; i++){\n     -+\t\tprintf(\"%\"PRIx64\"|\", filter->data[i]);\n     -+\t}\n     -+\tprintf(\"\\n\");\n     -+}\n     -+\n     -+static void add_string_to_filter(const char *data, struct bloom_filter *filter) {\n     -+\t\tstruct bloom_key key;\n     -+\t\tint i;\n     -+\n     -+\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n     -+\t\tprintf(\"Hashes:\");\n     -+\t\tfor (i = 0; i < settings.num_hashes; i++){\n     -+\t\t\tprintf(\"%08x|\", key.hashes[i]);\n     -+\t\t}\n     -+\t\tprintf(\"\\n\");\n     -+\t\tadd_key_to_filter(&key, filter, &settings);\n     -+}\n     -+\n     + \n     + struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n     + \n     +@@\n     + \tprintf(\"\\n\");\n     + }\n     + \n      +static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n      +{\n      +\tstruct commit *c;\n     @@ -429,72 +183,33 @@\n      +\tprint_bloom_filter(filter);\n      +}\n      +\n     -+int cmd__bloom(int argc, const char **argv)\n     -+{\n     -+    if (!strcmp(argv[1], \"generate_filter\")) {\n     -+\t\tstruct bloom_filter filter;\n     -+\t\tint i = 2;\n     -+\t\tfilter.len =  (settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n     -+\t\tfilter.data = xcalloc(filter.len, sizeof(uint64_t));\n     -+\n     -+\t\tif (!argv[2]){\n     -+\t\t\tdie(\"at least one input string expected\");\n     -+\t\t}\n     -+\n     -+\t\twhile (argv[i]) {\n     -+\t\t\tadd_string_to_filter(argv[i], &filter);\n     -+\t\t\ti++;\n     -+\t\t}\n     -+\n     -+\t\tprint_bloom_filter(&filter);\n     -+\t}\n     -+\n     -+\tif (!strcmp(argv[1], \"get_filter_for_commit\")) {\n     + int cmd__bloom(int argc, const char **argv)\n     + {\n     + \tif (!strcmp(argv[1], \"get_murmur3\")) {\n     +@@\n     + \t\tprint_bloom_filter(&filter);\n     + \t}\n     + \n     ++    if (!strcmp(argv[1], \"get_filter_for_commit\")) {\n      +\t\tstruct object_id oid;\n      +\t\tconst char *end;\n      +\t\tif (parse_oid_hex(argv[2], &oid, &end))\n      +\t\t\tdie(\"cannot parse oid '%s'\", argv[2]);\n     -+\t\tload_bloom_filters();\n     ++\t\tinit_bloom_filters();\n      +\t\tget_bloom_filter_for_commit(&oid);\n      +\t}\n      +\n     -+\treturn 0;\n     -+}\n     -\n     - diff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\n     - --- a/t/helper/test-tool.c\n     - +++ b/t/helper/test-tool.c\n     -@@\n     - };\n     - \n     - static struct test_cmd cmds[] = {\n     -+\t{ \"bloom\", cmd__bloom },\n     - \t{ \"chmtime\", cmd__chmtime },\n     - \t{ \"config\", cmd__config },\n     - \t{ \"ctype\", cmd__ctype },\n     -\n     - diff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\n     - --- a/t/helper/test-tool.h\n     - +++ b/t/helper/test-tool.h\n     -@@\n     - #define USE_THE_INDEX_COMPATIBILITY_MACROS\n     - #include \"git-compat-util.h\"\n     - \n     -+int cmd__bloom(int argc, const char **argv);\n     - int cmd__chmtime(int argc, const char **argv);\n     - int cmd__config(int argc, const char **argv);\n     - int cmd__ctype(int argc, const char **argv);\n     + \treturn 0;\n     + }\n     + \\ No newline at end of file\n      \n       diff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\n     - new file mode 100755\n     - --- /dev/null\n     + --- a/t/t0095-bloom.sh\n       +++ b/t/t0095-bloom.sh\n      @@\n     -+#!/bin/sh\n     -+\n     -+test_description='test bloom.c'\n     -+. ./test-lib.sh\n     -+\n     + \ttest_cmp expect actual\n     + '\n     + \n      +test_expect_success 'get bloom filters for commit with no changes' '\n      +\tgit init &&\n      +\tgit commit --allow-empty -m \"c0\" &&\n     @@ -517,8 +232,8 @@\n      +\tgit add smallDir &&\n      +\tgit commit -m \"commit with 10 changes\" &&\n      +\tcat >expect <<-\\EOF &&\n     -+\tFilter_Length:4\n     -+\tFilter_Data:508928809087080a|8a7648210804001|4089824400951000|841ab310098051a8|\n     ++\tFilter_Length:25\n     ++\tFilter_Data:82|a0|65|47|0c|92|90|c0|a1|40|02|a0|e2|40|e0|04|0a|9a|66|cf|80|19|85|42|23|\n      +\tEOF\n      +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n      +\ttest_cmp expect actual\n     @@ -542,64 +257,5 @@\n      +\ttest_cmp expect actual\n      +'\n      +\n     -+test_expect_success 'compute bloom key for empty string' '\n     -+\tcat >expect <<-\\EOF &&\n     -+\tHashes:5615800c|5b966560|61174ab4|66983008|6c19155c|7199fab0|771ae004|\n     -+\tFilter_Length:1\n     -+\tFilter_Data:11000110001110|\n     -+\tEOF\n     -+\ttest-tool bloom generate_filter \"\" >actual &&\n     -+\ttest_cmp expect actual\n     -+'\n     -+\n     -+test_expect_success 'compute bloom key for whitespace' '\n     -+\tcat >expect <<-\\EOF &&\n     -+\tHashes:1bf014e6|8a91b50b|f9335530|67d4f555|d676957a|4518359f|b3b9d5c4|\n     -+\tFilter_Length:1\n     -+\tFilter_Data:401004080200810|\n     -+\tEOF\n     -+\ttest-tool bloom generate_filter \" \" >actual &&\n     -+\ttest_cmp expect actual\n     -+'\n     -+\n     -+test_expect_success 'compute bloom key for a root level folder' '\n     -+\tcat >expect <<-\\EOF &&\n     -+\tHashes:1a21016f|fff1c06d|e5c27f6b|cb933e69|b163fd67|9734bc65|7d057b63|\n     -+\tFilter_Length:1\n     -+\tFilter_Data:aaa800000000|\n     -+\tEOF\n     -+\ttest-tool bloom generate_filter \"A\" >actual &&\n     -+\ttest_cmp expect actual\n     -+'\n     -+\n     -+test_expect_success 'compute bloom key for a root level file' '\n     -+\tcat >expect <<-\\EOF &&\n     -+\tHashes:e2d51107|30970605|7e58fb03|cc1af001|19dce4ff|679ed9fd|b560cefb|\n     -+\tFilter_Length:1\n     -+\tFilter_Data:a8000000000000aa|\n     -+\tEOF\n     -+\ttest-tool bloom generate_filter \"file.txt\" >actual &&\n     -+\ttest_cmp expect actual\n     -+'\n     -+\n     -+test_expect_success 'compute bloom key for a deep folder' '\n     -+\tcat >expect <<-\\EOF &&\n     -+\tHashes:864cf838|27f055cd|c993b362|6b3710f7|0cda6e8c|ae7dcc21|502129b6|\n     -+\tFilter_Length:1\n     -+\tFilter_Data:1c0000600003000|\n     -+\tEOF\n     -+\ttest-tool bloom generate_filter \"A/B/C/D/E\" >actual &&\n     -+\ttest_cmp expect actual\n     -+'\n     -+\n     -+test_expect_success 'compute bloom key for a deep file' '\n     -+\tcat >expect <<-\\EOF &&\n     -+\tHashes:07cdf850|4af629c7|8e1e5b3e|d1468cb5|146ebe2c|5796efa3|9abf211a|\n     -+\tFilter_Length:1\n     -+\tFilter_Data:4020100804010080|\n     -+\tEOF\n     -+\ttest-tool bloom generate_filter \"A/B/C/D/E/file.txt\" >actual &&\n     -+\ttest_cmp expect actual\n     -+'\n     -+\n     -+test_done\n     + test_done\n     + \\ No newline at end of file\n  3:  a698c04a78c !  5:  2d4c0b2da38 diff: halt tree-diff early after max_changes\n     @@ -29,7 +29,7 @@\n       \tstruct diff_options diffopt;\n      +\tint max_changes = 512;\n       \n     - \tif (!bloom_filters.slab_size)\n     + \tif (bloom_filters.slab_size == 0)\n       \t\treturn NULL;\n      @@\n       \n     @@ -46,7 +46,7 @@\n      -\tif (diff_queued_diff.nr <= 512) {\n      +\tif (diff_queued_diff.nr <= max_changes) {\n       \t\tstruct hashmap pathmap;\n     - \t\tstruct pathmap_hash_entry* e;\n     + \t\tstruct pathmap_hash_entry *e;\n       \t\tstruct hashmap_iter iter;\n      \n       diff --git a/diff.h b/diff.h\n  4:  c17bbcbc66e !  6:  c38b9b386ef commit-graph: compute Bloom filters for changed paths\n     @@ -2,11 +2,13 @@\n      \n          commit-graph: compute Bloom filters for changed paths\n      \n     -    Compute Bloom filters for the paths that changed between a commit and its\n     -    first parent using the implementation in bloom.c, when the\n     -    COMMIT_GRAPH_WRITE_CHANGED_PATHS flag is set. This computation is done on a\n     -    commit-by-commit basis. We will write these Bloom filters to the commit graph\n     -    file in the next change.\n     +    Add new COMMIT_GRAPH_WRITE_CHANGED_PATHS flag that makes Git compute\n     +    Bloom filters for the paths that changed between a commit and it's\n     +    first parent, for each commit in the commit-graph.  This computation\n     +    is done on a commit-by-commit basis.\n     +\n     +    We will write these Bloom filters to the commit-graph file, to store\n     +    this data on disk, in the next change in this series.\n      \n          Helped-by: Derrick Stolee <dstolee@microsoft.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n     @@ -31,7 +33,7 @@\n      +\t\t changed_paths:1;\n       \n       \tconst struct split_commit_graph_opts *split_opts;\n     -+\tuint32_t total_bloom_filter_data_size;\n     ++\tsize_t total_bloom_filter_data_size;\n       };\n       \n       static void write_graph_chunk_fanout(struct hashfile *f,\n     @@ -44,17 +46,17 @@\n      +\tint i;\n      +\tstruct progress *progress = NULL;\n      +\n     -+\tload_bloom_filters();\n     ++\tinit_bloom_filters();\n      +\n      +\tif (ctx->report_progress)\n     -+\t\tprogress = start_progress(\n     -+\t\t\t_(\"Computing commit diff Bloom filters\"),\n     ++\t\tprogress = start_delayed_progress(\n     ++\t\t\t_(\"Computing commit changed paths Bloom filters\"),\n      +\t\t\tctx->commits.nr);\n      +\n      +\tfor (i = 0; i < ctx->commits.nr; i++) {\n      +\t\tstruct commit *c = ctx->commits.list[i];\n      +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n     -+\t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n     ++\t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n      +\t\tdisplay_progress(progress, i + 1);\n      +\t}\n      +\n     @@ -93,7 +95,7 @@\n       \t/* Make sure that each OID in the input is a valid commit OID. */\n      -\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3)\n      +\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n     -+\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n     ++\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4),\n       };\n       \n       struct split_commit_graph_opts {\n  5:  78e8e49c3a1 !  7:  d24c85c54ef commit-graph: examine changed-path objects in pack order\n     @@ -39,6 +39,7 @@\n       /* Remember to update object flag allocation in object.h */\n       #define REACHABLE       (1u<<15)\n       \n     +-char *get_commit_graph_filename(struct object_directory *odb)\n      +/* Keep track of the order in which commits are added to our list. */\n      +define_commit_slab(commit_pos, int);\n      +static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n     @@ -55,16 +56,20 @@\n      +}\n      +\n      +static int commit_pos_cmp(const void *va, const void *vb)\n     -+{\n     + {\n     +-\treturn xstrfmt(\"%s/info/commit-graph\", odb->path);\n      +\tconst struct commit *a = *(const struct commit **)va;\n      +\tconst struct commit *b = *(const struct commit **)vb;\n      +\treturn commit_pos_at(&commit_pos, a) -\n      +\t       commit_pos_at(&commit_pos, b);\n      +}\n      +\n     - char *get_commit_graph_filename(const char *obj_dir)\n     - {\n     - \tchar *filename = xstrfmt(\"%s/info/commit-graph\", obj_dir);\n     ++char *get_commit_graph_filename(struct object_directory *obj_dir)\n     ++{\n     ++\treturn xstrfmt(\"%s/info/commit-graph\", obj_dir->path);\n     + }\n     + \n     + static char *get_split_graph_filename(struct object_directory *odb,\n      @@\n       \toidcpy(&(ctx->oids.list[ctx->oids.nr]), oid);\n       \tctx->oids.nr++;\n     @@ -78,27 +83,27 @@\n       {\n       \tint i;\n       \tstruct progress *progress = NULL;\n     -+\tstruct commit **sorted_by_pos;\n     ++\tstruct commit **sorted_commits;\n       \n     - \tload_bloom_filters();\n     + \tinit_bloom_filters();\n       \n      @@\n     - \t\t\t_(\"Computing commit diff Bloom filters\"),\n     + \t\t\t_(\"Computing commit changed paths Bloom filters\"),\n       \t\t\tctx->commits.nr);\n       \n     -+\tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n     -+\tCOPY_ARRAY(sorted_by_pos, ctx->commits.list, ctx->commits.nr);\n     -+\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n     ++\tALLOC_ARRAY(sorted_commits, ctx->commits.nr);\n     ++\tCOPY_ARRAY(sorted_commits, ctx->commits.list, ctx->commits.nr);\n     ++\tQSORT(sorted_commits, ctx->commits.nr, commit_pos_cmp);\n      +\n       \tfor (i = 0; i < ctx->commits.nr; i++) {\n      -\t\tstruct commit *c = ctx->commits.list[i];\n     -+\t\tstruct commit *c = sorted_by_pos[i];\n     ++\t\tstruct commit *c = sorted_commits[i];\n       \t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n     - \t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n     + \t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n       \t\tdisplay_progress(progress, i + 1);\n       \t}\n       \n     -+\tfree(sorted_by_pos);\n     ++\tfree(sorted_commits);\n       \tstop_progress(&progress);\n       }\n       \n  6:  58704d81b6b !  8:  5ed16f35fed commit-graph: examine commits by generation number\n     @@ -1,11 +1,11 @@\n     -Author: Derrick Stolee <dstolee@microsoft.com>\n     +Author: Garima Singh <garima.singh@microsoft.com>\n      \n          commit-graph: examine commits by generation number\n      \n          When running 'git commit-graph write --changed-paths', we sort the\n          commits by pack-order to save time when computing the changed-paths\n          bloom filters. This does not help when finding the commits via the\n     -    --reachable flag.\n     +    '--reachable' flag.\n      \n          If not using pack-order, then sort by generation number before\n          examining the diff. Commits with similar generation are more likely\n     @@ -45,9 +45,9 @@\n      +\treturn 0;\n      +}\n      +\n     - char *get_commit_graph_filename(const char *obj_dir)\n     + char *get_commit_graph_filename(struct object_directory *obj_dir)\n       {\n     - \tchar *filename = xstrfmt(\"%s/info/commit-graph\", obj_dir);\n     + \treturn xstrfmt(\"%s/info/commit-graph\", obj_dir->path);\n      @@\n       \t\t report_progress:1,\n       \t\t split:1,\n     @@ -57,20 +57,20 @@\n      +\t\t order_by_pack:1;\n       \n       \tconst struct split_commit_graph_opts *split_opts;\n     - \tuint32_t total_bloom_filter_data_size;\n     + \tsize_t total_bloom_filter_data_size;\n      @@\n       \n     - \tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n     - \tCOPY_ARRAY(sorted_by_pos, ctx->commits.list, ctx->commits.nr);\n     --\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n     + \tALLOC_ARRAY(sorted_commits, ctx->commits.nr);\n     + \tCOPY_ARRAY(sorted_commits, ctx->commits.list, ctx->commits.nr);\n     +-\tQSORT(sorted_commits, ctx->commits.nr, commit_pos_cmp);\n      +\n      +\tif (ctx->order_by_pack)\n     -+\t\tQSORT(sorted_by_pos, ctx->commits.nr, commit_pos_cmp);\n     ++\t\tQSORT(sorted_commits, ctx->commits.nr, commit_pos_cmp);\n      +\telse\n     -+\t\tQSORT(sorted_by_pos, ctx->commits.nr, commit_gen_cmp);\n     ++\t\tQSORT(sorted_commits, ctx->commits.nr, commit_gen_cmp);\n       \n       \tfor (i = 0; i < ctx->commits.nr; i++) {\n     - \t\tstruct commit *c = sorted_by_pos[i];\n     + \t\tstruct commit *c = sorted_commits[i];\n      @@\n       \t}\n       \n  -:  ----------- >  9:  55824cda89c diff: skip batch object download when possible\n  7:  39ee0610800 ! 10:  1e4663523de commit-graph: write Bloom filters to commit graph file\n     @@ -2,9 +2,10 @@\n      \n          commit-graph: write Bloom filters to commit graph file\n      \n     -    Update the technical documentation for commit-graph-format with the formats for\n     -    the Bloom filter index (BIDX) and Bloom filter data (BDAT) chunks. Write the\n     -    computed Bloom filters information to the commit graph file using this format.\n     +    Update the technical documentation for commit-graph-format with\n     +    the formats for the Bloom filter index (BIDX) and Bloom filter\n     +    data (BDAT) chunks. Write the computed Bloom filters information\n     +    to the commit graph file using this format.\n      \n          Helped-by: Derrick Stolee <dstolee@microsoft.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n     @@ -17,7 +18,7 @@\n         the graph file.\n       \n      +- The Bloom filter of the commit carrying the paths that were changed between\n     -+  the commit and its first parent.\n     ++  the commit and its first parent, if requested.\n      +\n       These positional references are stored as unsigned 32-bit integers\n       corresponding to the array position within the list of commit OIDs. Due\n     @@ -36,16 +37,22 @@\n      +  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n      +    * It starts with header consisting of three unsigned 32-bit integers:\n      +      - Version of the hash algorithm being used. We currently only support\n     -+\tvalue 1 which implies the murmur3 hash implemented exactly as described\n     -+\tin https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n     ++\tvalue 1 which corresponds to the 32-bit version of the murmur3 hash\n     ++\timplemented exactly as described in\n     ++\thttps://en.wikipedia.org/wiki/MurmurHash#Algorithm and the double\n     ++\thashing technique using seed values 0x293ae76f and 0x7e646e2 as\n     ++\tdescribed in https://doi.org/10.1007/978-3-540-30494-4_26 \"Bloom Filters\n     ++\tin Probabilistic Verification\"\n      +      - The number of times a path is hashed and hence the number of bit positions\n     -+\tthat cumulatively determine whether a file is present in the commit.\n     ++\t      that cumulatively determine whether a file is present in the commit.\n      +      - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n     -+\tcontains 'n' entries, then the filter size is the minimum number of 64-bit\n     -+\twords that contain n*b bits.\n     ++\t      contains 'n' entries, then the filter size is the minimum number of 64-bit\n     ++\t      words that contain n*b bits.\n      +    * The rest of the chunk is the concatenation of all the computed Bloom\n      +      filters for the commits in lexicographic order.\n     -+    * The BDAT chunk is present iff BIDX is present.\n     ++    * Note: Commits with no changes or more than 512 changes have Bloom filters\n     ++      of length zero.\n     ++    * The BDAT chunk is present if and only if BIDX is present.\n      +\n         Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n             This list of H-byte hashes describe a set of B commit-graph files that\n     @@ -103,16 +110,14 @@\n       \t\tlast_chunk_offset = chunk_offset;\n       \t}\n       \n     -+\t/* We need both the bloom chunks to exist together. Else ignore the data */\n     -+\tif ((graph->chunk_bloom_indexes && !graph->chunk_bloom_data)\n     -+\t\t || (!graph->chunk_bloom_indexes && graph->chunk_bloom_data)) {\n     ++\tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data) {\n     ++\t\tinit_bloom_filters();\n     ++\t} else {\n     ++\t\t/* We need both the bloom chunks to exist together. Else ignore the data */\n      +\t\tgraph->chunk_bloom_indexes = NULL;\n      +\t\tgraph->chunk_bloom_data = NULL;\n      +\t\tgraph->bloom_filter_settings = NULL;\n      +\t}\n     -+\n     -+\tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data)\n     -+\t\tload_bloom_filters();\n      +\n       \thashcpy(graph->oid.hash, graph->data + graph->data_len - graph->hash_len);\n       \n     @@ -148,7 +153,7 @@\n      +\n      +static void write_graph_chunk_bloom_data(struct hashfile *f,\n      +\t\t\t\t\t struct write_commit_graph_context *ctx,\n     -+\t\t\t\t\t struct bloom_filter_settings *settings)\n     ++\t\t\t\t\t const struct bloom_filter_settings *settings)\n      +{\n      +\tstruct commit **list = ctx->commits.list;\n      +\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n     @@ -167,7 +172,7 @@\n      +\twhile (list < last) {\n      +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n      +\t\tdisplay_progress(progress, ++i);\n     -+\t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n     ++\t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n      +\t\tlist++;\n      +\t}\n      +\n     @@ -177,22 +182,11 @@\n       static int oid_compare(const void *_a, const void *_b)\n       {\n       \tconst struct object_id *a = (const struct object_id *)_a;\n     -@@\n     - \tload_bloom_filters();\n     - \n     - \tif (ctx->report_progress)\n     --\t\tprogress = start_progress(\n     --\t\t\t_(\"Computing commit diff Bloom filters\"),\n     -+\t\tprogress = start_delayed_progress(\n     -+\t\t\t_(\"Computing changed paths Bloom filters\"),\n     - \t\t\tctx->commits.nr);\n     - \n     - \tALLOC_ARRAY(sorted_by_pos, ctx->commits.nr);\n      @@\n       \tstruct strbuf progress_title = STRBUF_INIT;\n       \tint num_chunks = 3;\n       \tstruct object_id file_hash;\n     -+\tstruct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n     ++\tconst struct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n       \n       \tif (ctx->split) {\n       \t\tstruct strbuf tmp_file = STRBUF_INIT;\n     @@ -236,6 +230,14 @@\n       \tif (ctx->num_commit_graphs_after > 1 &&\n       \t    write_graph_chunk_base(f, ctx)) {\n       \t\treturn -1;\n     +@@\n     + \t\tclose(g->graph_fd);\n     + \t}\n     + \tfree(g->filename);\n     ++\tfree(g->bloom_filter_settings);\n     + \tfree(g);\n     + }\n     + \n      \n       diff --git a/commit-graph.h b/commit-graph.h\n       --- a/commit-graph.h\n     @@ -246,7 +248,7 @@\n       struct commit;\n      +struct bloom_filter_settings;\n       \n     - char *get_commit_graph_filename(const char *obj_dir);\n     + char *get_commit_graph_filename(struct object_directory *odb);\n       int open_commit_graph(const char *graph_file, int *fd, struct stat *st);\n      @@\n       \tconst unsigned char *chunk_commit_data;\n     @@ -258,13 +260,4 @@\n      +\tstruct bloom_filter_settings *bloom_filter_settings;\n       };\n       \n     - struct commit_graph *load_commit_graph_one_fd_st(int fd, struct stat *st);\n     -@@\n     - \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n     - \t/* Make sure that each OID in the input is a valid commit OID. */\n     - \tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n     --\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4)\n     -+\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4),\n     - };\n     - \n     - struct split_commit_graph_opts {\n     + struct commit_graph *load_commit_graph_one_fd_st(int fd, struct stat *st,\n  8:  b20c8d2b209 ! 11:  68395d4051b commit-graph: reuse existing Bloom filters during write.\n     @@ -1,9 +1,10 @@\n      Author: Garima Singh <garima.singh@microsoft.com>\n      \n     -    commit-graph: reuse existing Bloom filters during write.\n     +    commit-graph: reuse existing Bloom filters during write\n      \n     -    Read previously computed Bloom filters from the commit-graph file if\n     -    possible to avoid recomputing during commit-graph write.\n     +    Add logic to\n     +    a) parse Bloom filter information from the commit graph file and,\n     +    b) re-use existing Bloom filters.\n      \n          See Documentation/technical/commit-graph-format for the format in which\n          the Bloom filter information is written to the commit graph file.\n     @@ -25,7 +26,7 @@\n             commit's filter.\n      \n          We toggle whether Bloom filters should be recomputed based on the\n     -    compute_if_null flag.\n     +    compute_if_not_present flag.\n      \n          Helped-by: Derrick Stolee <dstolee@microsoft.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n     @@ -34,15 +35,16 @@\n       --- a/bloom.c\n       +++ b/bloom.c\n      @@\n     - #include \"git-compat-util.h\"\n     - #include \"bloom.h\"\n     + #include \"diffcore.h\"\n     + #include \"revision.h\"\n     + #include \"hashmap.h\"\n     ++#include \"commit-graph.h\"\n      +#include \"commit.h\"\n     -+#include \"commit-slab.h\"\n     - #include \"commit-graph.h\"\n     - #include \"object-store.h\"\n     - #include \"diff.h\"\n     + \n     + define_commit_slab(bloom_filter_slab, struct bloom_filter);\n     + \n      @@\n     - \t}\n     + \treturn ((unsigned char)1) << (pos & (BITS_PER_WORD - 1));\n       }\n       \n      +static int load_bloom_filter_from_graph(struct commit_graph *g,\n     @@ -62,23 +64,29 @@\n      +\n      +\tend_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n      +\n     -+\tif (lex_pos)\n     ++\tif (lex_pos > 0)\n      +\t\tstart_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n      +\telse\n      +\t\tstart_index = 0;\n      +\n      +\tfilter->len = end_index - start_index;\n     -+\tfilter->data = (uint64_t *)(g->chunk_bloom_data +\n     -+\t\t\t\t\tsizeof(uint64_t) * start_index +\n     ++\tfilter->data = (unsigned char *)(g->chunk_bloom_data +\n     ++\t\t\t\t\tsizeof(unsigned char) * start_index +\n      +\t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n      +\n      +\treturn 1;\n      +}\n      +\n     + /*\n     +  * Calculate the murmur3 32-bit hash value for the given data\n     +  * using the given seed.\n     +@@\n     + }\n     + \n       struct bloom_filter *get_bloom_filter(struct repository *r,\n      -\t\t\t\t      struct commit *c)\n      +\t\t\t\t      struct commit *c,\n     -+\t\t\t\t      int compute_if_not_present)\n     ++\t\t\t\t\t  int compute_if_not_present)\n       {\n       \tstruct bloom_filter *filter;\n       \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n     @@ -102,7 +110,7 @@\n      +\n       \trepo_diff_setup(r, &diffopt);\n       \tdiffopt.flags.recursive = 1;\n     - \tdiffopt.max_changes = max_changes;\n     + \tdiffopt.detect_rename = 0;\n      \n       diff --git a/bloom.h b/bloom.h\n       --- a/bloom.h\n     @@ -110,21 +118,21 @@\n      @@\n       \n       #define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n     - #define BITS_PER_WORD 64\n     -+#define BLOOMDATA_CHUNK_HEADER_SIZE 3*sizeof(uint32_t)\n     + #define BITS_PER_WORD 8\n     ++#define BLOOMDATA_CHUNK_HEADER_SIZE 3 * sizeof(uint32_t)\n       \n       /*\n        * A bloom_filter struct represents a data segment to\n      @@\n     - \t\t\t\t\t   struct bloom_filter_settings *settings);\n     + void init_bloom_filters(void);\n       \n       struct bloom_filter *get_bloom_filter(struct repository *r,\n      -\t\t\t\t      struct commit *c);\n      +\t\t\t\t      struct commit *c,\n      +\t\t\t\t      int compute_if_not_present);\n       \n     - int bloom_filter_contains(struct bloom_filter *filter,\n     - \t\t\t  struct bloom_key *key,\n     + #endif\n     + \\ No newline at end of file\n      \n       diff --git a/commit-graph.c b/commit-graph.c\n       --- a/commit-graph.c\n     @@ -145,25 +153,17 @@\n      -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n      +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n       \t\tdisplay_progress(progress, ++i);\n     - \t\thashwrite(f, filter->data, filter->len * sizeof(uint64_t));\n     + \t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n       \t\tlist++;\n      @@\n       \n       \tfor (i = 0; i < ctx->commits.nr; i++) {\n     - \t\tstruct commit *c = sorted_by_pos[i];\n     + \t\tstruct commit *c = sorted_commits[i];\n      -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n      +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n     - \t\tctx->total_bloom_filter_data_size += sizeof(uint64_t) * filter->len;\n     + \t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n       \t\tdisplay_progress(progress, i + 1);\n       \t}\n     -@@\n     - \t\tg->data = NULL;\n     - \t\tclose(g->graph_fd);\n     - \t}\n     -+\tfree(g->bloom_filter_settings);\n     - \tfree(g->filename);\n     - \tfree(g);\n     - }\n      \n       diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n       --- a/t/helper/test-bloom.c\n  9:  3d7ee0c9695 ! 12:  7e450e45236 commit-graph: add --changed-paths option to write subcommand\n     @@ -56,7 +56,7 @@\n      +\tint enable_changed_paths;\n       } opts;\n       \n     - static int graph_verify(int argc, const char **argv)\n     + static struct object_directory *find_odb(struct repository *r,\n      @@\n       \t\t\tN_(\"start walk at commits listed by stdin\")),\n       \t\tOPT_BOOL(0, \"append\", &opts.append,\n     @@ -74,4 +74,4 @@\n      +\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n       \n       \tread_replace_refs = 0;\n     - \n     + \todb = find_odb(the_repository, opts.obj_dir);\n 10:  77f1c561e82 ! 13:  b18af58aa3e revision.c: use Bloom filters to speed up path based revision walks\n     @@ -2,17 +2,27 @@\n      \n          revision.c: use Bloom filters to speed up path based revision walks\n      \n     -    Revision walk will now use Bloom filters for commits to speed up revision\n     -    walks for a particular path (for computing history for that path), if they\n     -    are present in the commit-graph file.\n     +    Revision walk will now use Bloom filters for commits to speed up\n     +    revision walks for a particular path (for computing history for\n     +    that path), if they are present in the commit-graph file.\n      \n     -    We load the Bloom filters during the prepare_revision_walk step, but only\n     -    when dealing with a single pathspec. While comparing trees in\n     -    rev_compare_trees(), if the Bloom filter says that the file is not different\n     -    between the two trees, we don't need to compute the expensive diff. This is\n     -    where we get our performance gains. The other response of the Bloom filter\n     -    is `maybe`, in which case we fall back to the full diff calculation to\n     -    determine if the path was changed in the commit.\n     +    We load the Bloom filters during the prepare_revision_walk step,\n     +    currently only when dealing with a single pathspec. Extending\n     +    it to work with multiple pathspecs can be explored and built on\n     +    top of this series in the future.\n     +\n     +    While comparing trees in rev_compare_trees(), if the Bloom filter\n     +    says that the file is not different between the two trees, we don't\n     +    need to compute the expensive diff. This is where we get our\n     +    performance gains. The other response of the Bloom filter is '`:maybe',\n     +    in which case we fall back to the full diff calculation to determine\n     +    if the path was changed in the commit.\n     +\n     +    We do not try to use Bloom filters when the '--walk-reflogs' option\n     +    is specified. The '--walk-reflogs' option does not walk the commit\n     +    ancestry chain like the rest of the options. Incorporating the\n     +    performance gains when walking reflog entries would add more\n     +    complexity, and can be explored in a later series.\n      \n          Performance Gains:\n          We tested the performance of `git log -- <path>` on the git repo, the linux\n     @@ -30,6 +40,49 @@\n          Helped-by: Jonathan Tan <jonathantanmy@google.com>\n          Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n      \n     + diff --git a/bloom.c b/bloom.c\n     + --- a/bloom.c\n     + +++ b/bloom.c\n     +@@\n     + \n     + \treturn filter;\n     + }\n     ++\n     ++int bloom_filter_contains(const struct bloom_filter *filter,\n     ++\t\t\t  const struct bloom_key *key,\n     ++\t\t\t  const struct bloom_filter_settings *settings)\n     ++{\n     ++\tint i;\n     ++\tuint64_t mod = filter->len * BITS_PER_WORD;\n     ++\n     ++\tif (!mod)\n     ++\t\treturn -1;\n     ++\n     ++\tfor (i = 0; i < settings->num_hashes; i++) {\n     ++\t\tuint64_t hash_mod = key->hashes[i] % mod;\n     ++\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n     ++\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n     ++\t\t\treturn 0;\n     ++\t}\n     ++\n     ++\treturn 1;\n     ++}\n     + \\ No newline at end of file\n     +\n     + diff --git a/bloom.h b/bloom.h\n     + --- a/bloom.h\n     + +++ b/bloom.h\n     +@@\n     + \t\t\t\t      struct commit *c,\n     + \t\t\t\t      int compute_if_not_present);\n     + \n     ++int bloom_filter_contains(const struct bloom_filter *filter,\n     ++\t\t\t  const struct bloom_key *key,\n     ++\t\t\t  const struct bloom_filter_settings *settings);\n     ++\n     + #endif\n     + \\ No newline at end of file\n     +\n       diff --git a/revision.c b/revision.c\n       --- a/revision.c\n       +++ b/revision.c\n     @@ -38,7 +91,6 @@\n       #include \"hashmap.h\"\n       #include \"utf8.h\"\n      +#include \"bloom.h\"\n     -+#include \"json-writer.h\"\n       \n       volatile show_early_output_fn_t show_early_output;\n       \n     @@ -46,29 +98,6 @@\n       \toptions->flags.has_changes = 1;\n       }\n       \n     -+static int bloom_filter_atexit_registered;\n     -+static unsigned int count_bloom_filter_maybe;\n     -+static unsigned int count_bloom_filter_definitely_not;\n     -+static unsigned int count_bloom_filter_false_positive;\n     -+static unsigned int count_bloom_filter_not_present;\n     -+static unsigned int count_bloom_filter_length_zero;\n     -+\n     -+static void trace2_bloom_filter_statistics_atexit(void)\n     -+{\n     -+\tstruct json_writer jw = JSON_WRITER_INIT;\n     -+\n     -+\tjw_object_begin(&jw, 0);\n     -+\tjw_object_intmax(&jw, \"filter_not_present\", count_bloom_filter_not_present);\n     -+\tjw_object_intmax(&jw, \"zero_length_filter\", count_bloom_filter_length_zero);\n     -+\tjw_object_intmax(&jw, \"maybe\", count_bloom_filter_maybe);\n     -+\tjw_object_intmax(&jw, \"definitely_not\", count_bloom_filter_definitely_not);\n     -+\tjw_end(&jw);\n     -+\n     -+\ttrace2_data_json(\"bloom\", the_repository, \"statistics\", &jw);\n     -+\n     -+\tjw_release(&jw);\n     -+}\n     -+\n      +static void prepare_to_use_bloom_filter(struct rev_info *revs)\n      +{\n      +\tstruct pathspec_item *pi;\n     @@ -92,6 +121,7 @@\n      +\tpi = &revs->pruning.pathspec.items[0];\n      +\tlast_index = pi->len - 1;\n      +\n     ++\t/* remove single trailing slash from path, if needed */\n      +\tif (pi->match[last_index] == '/') {\n      +\t    path_alloc = xstrdup(pi->match);\n      +\t    path_alloc[last_index] = '\\0';\n     @@ -104,11 +134,6 @@\n      +\trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n      +\tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n      +\n     -+\tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n     -+\t\tatexit(trace2_bloom_filter_statistics_atexit);\n     -+\t\tbloom_filter_atexit_registered = 1;\n     -+\t}\n     -+\n      +\tfree(path_alloc);\n      +}\n      +\n     @@ -127,12 +152,10 @@\n      +\tfilter = get_bloom_filter(revs->repo, commit, 0);\n      +\n      +\tif (!filter) {\n     -+\t\tcount_bloom_filter_not_present++;\n      +\t\treturn -1;\n      +\t}\n      +\n      +\tif (!filter->len) {\n     -+\t\tcount_bloom_filter_length_zero++;\n      +\t\treturn -1;\n      +\t}\n      +\n     @@ -140,11 +163,6 @@\n      +\t\t\t\t       revs->bloom_key,\n      +\t\t\t\t       revs->bloom_filter_settings);\n      +\n     -+\tif (result)\n     -+\t\tcount_bloom_filter_maybe++;\n     -+\telse\n     -+\t\tcount_bloom_filter_definitely_not++;\n     -+\n      +\treturn result;\n      +}\n      +\n     @@ -162,7 +180,7 @@\n       \t\t\treturn REV_TREE_SAME;\n       \t}\n       \n     -+\tif (revs->pruning.pathspec.nr == 1 && !revs->reflog_info && !nth_parent) {\n     ++\tif (revs->bloom_key && !nth_parent) {\n      +\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n      +\n      +\t\tif (bloom_ret == 0)\n     @@ -174,10 +192,6 @@\n       \tif (diff_tree_oid(&t1->object.oid, &t2->object.oid, \"\",\n       \t\t\t   &revs->pruning) < 0)\n       \t\treturn REV_TREE_DIFFERENT;\n     -+\n     -+\tif (!nth_parent)\n     -+\t\tif (bloom_ret == 1 && tree_difference == REV_TREE_SAME)\n     -+\t\t\tcount_bloom_filter_false_positive++;\n      +\n       \treturn tree_difference;\n       }\n     @@ -237,164 +251,3 @@\n       };\n       \n       int ref_excluded(struct string_list *, const char *path);\n     -\n     - diff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\n     - --- a/t/helper/test-read-graph.c\n     - +++ b/t/helper/test-read-graph.c\n     -@@\n     - \t\tprintf(\" commit_metadata\");\n     - \tif (graph->chunk_extra_edges)\n     - \t\tprintf(\" extra_edges\");\n     -+\tif (graph->chunk_bloom_indexes)\n     -+\t\tprintf(\" bloom_indexes\");\n     -+\tif (graph->chunk_bloom_data)\n     -+\t\tprintf(\" bloom_data\");\n     - \tprintf(\"\\n\");\n     - \n     - \tUNLEAK(graph);\n     -\n     - diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n     - new file mode 100755\n     - --- /dev/null\n     - +++ b/t/t4216-log-bloom.sh\n     -@@\n     -+#!/bin/sh\n     -+\n     -+test_description='git log for a path with bloom filters'\n     -+. ./test-lib.sh\n     -+\n     -+test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n     -+\tgit init &&\n     -+\tmkdir A A/B A/B/C &&\n     -+\ttest_commit c1 A/file1 &&\n     -+\ttest_commit c2 A/B/file2 &&\n     -+\ttest_commit c3 A/B/C/file3 &&\n     -+\ttest_commit c4 A/file1 &&\n     -+\ttest_commit c5 A/B/file2 &&\n     -+\ttest_commit c6 A/B/C/file3 &&\n     -+\ttest_commit c7 A/file1 &&\n     -+\ttest_commit c8 A/B/file2 &&\n     -+\ttest_commit c9 A/B/C/file3 &&\n     -+\tgit checkout -b side HEAD~4 &&\n     -+\ttest_commit side-1 file4 &&\n     -+\tgit checkout master &&\n     -+\tgit merge side &&\n     -+\ttest_commit c10 file5 &&\n     -+\tmv file5 file5_renamed &&\n     -+\tgit add file5_renamed &&\n     -+\tgit commit -m \"rename\" &&\n     -+\tgit commit-graph write --reachable --changed-paths\n     -+'\n     -+graph_read_expect() {\n     -+\tOPTIONAL=\"\"\n     -+\tNUM_CHUNKS=5\n     -+\tcat >expect <<- EOF\n     -+\theader: 43475048 1 1 $NUM_CHUNKS 0\n     -+\tnum_commits: $1\n     -+\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n     -+\tEOF\n     -+\ttest-tool read-graph >output &&\n     -+\ttest_cmp expect output\n     -+}\n     -+\n     -+test_expect_success 'commit-graph write wrote out the bloom chunks' '\n     -+\tgraph_read_expect 13\n     -+'\n     -+\n     -+setup() {\n     -+\trm output\n     -+\trm \"$TRASH_DIRECTORY/trace.perf\"\n     -+\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom\n     -+\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n     -+}\n     -+\n     -+test_bloom_filters_used() {\n     -+\tlog_args=$1\n     -+\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":0,\\\"zero_length_filter\\\":0,\\\"maybe\\\"\"\n     -+\tsetup \"$log_args\"\n     -+\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && test_cmp log_wo_bloom log_w_bloom\n     -+}\n     -+\n     -+test_bloom_filters_not_used() {\n     -+\tlog_args=$1\n     -+\tsetup \"$log_args\"\n     -+\t!(grep -q \"statistics:{\\\"filter_not_present\\\":\" \"$TRASH_DIRECTORY/trace.perf\") && test_cmp log_wo_bloom log_w_bloom\n     -+}\n     -+\n     -+for path in A A/B A/B/C A/file1 A/B/file2 A/B/C/file3 file4 file5_renamed\n     -+do\n     -+\tfor option in \"\" \\\n     -+\t\t      \"--full-history\" \\\n     -+\t\t      \"--full-history --simplify-merges\" \\\n     -+\t\t      \"--simplify-merges\" \\\n     -+\t\t      \"--simplify-by-decoration\" \\\n     -+\t\t      \"--follow\" \\\n     -+\t\t      \"--first-parent\" \\\n     -+\t\t      \"--topo-order\" \\\n     -+\t\t      \"--date-order\" \\\n     -+\t\t      \"--author-date-order\" \\\n     -+\t\t      \"--ancestry-path side..master\"\n     -+\tdo\n     -+\t\ttest_expect_success \"git log option: $option for path: $path\" '\n     -+\t\t\ttest_bloom_filters_used \"$option -- $path\"\n     -+\t\t'\n     -+\tdone\n     -+done\n     -+\n     -+test_expect_success 'git log -- folder works with and without the trailing slash' '\n     -+\ttest_bloom_filters_used \"-- A\" &&\n     -+\ttest_bloom_filters_used \"-- A/\"\n     -+'\n     -+\n     -+test_expect_success 'git log for path that does not exist. ' '\n     -+\ttest_bloom_filters_used \"-- path_does_not_exist\"\n     -+'\n     -+\n     -+test_expect_success 'git log with --walk-reflogs does not use bloom filters' '\n     -+\ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n     -+'\n     -+\n     -+test_expect_success 'git log -- multiple path specs does not use bloom filters' '\n     -+\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n     -+'\n     -+\n     -+test_expect_success 'git log with wildcard that resolves to a single path uses bloom filters' '\n     -+\ttest_bloom_filters_used \"-- *4\" &&\n     -+\ttest_bloom_filters_used \"-- *renamed\"\n     -+'\n     -+\n     -+test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses bloom filters' '\n     -+\ttest_bloom_filters_not_used \"-- *\" &&\n     -+\ttest_bloom_filters_not_used \"-- file*\"\n     -+'\n     -+\n     -+test_expect_success 'setup - add commit-graph to the chain without bloom filters' '\n     -+\ttest_commit c14 A/anotherFile2 &&\n     -+\ttest_commit c15 A/B/anotherFile2 &&\n     -+\ttest_commit c16 A/B/C/anotherFile2 &&\n     -+\tGIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0 git commit-graph write --reachable --split &&\n     -+\ttest_line_count = 2 .git/objects/info/commit-graphs/commit-graph-chain\n     -+'\n     -+\n     -+test_expect_success 'git log does not use bloom filters if the latest graph does not have bloom filters.' '\n     -+\ttest_bloom_filters_not_used \"-- A/B\"\n     -+'\n     -+\n     -+test_expect_success 'setup - add commit-graph to the chain with bloom filters' '\n     -+\ttest_commit c17 A/anotherFile3 &&\n     -+\tgit commit-graph write --reachable --changed-paths --split &&\n     -+\ttest_line_count = 3 .git/objects/info/commit-graphs/commit-graph-chain\n     -+'\n     -+\n     -+test_bloom_filters_used_when_some_filters_are_missing() {\n     -+\tlog_args=$1\n     -+\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":3,\\\"zero_length_filter\\\":0,\\\"maybe\\\":6,\\\"definitely_not\\\":6\"\n     -+\tsetup \"$log_args\"\n     -+\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && test_cmp log_wo_bloom log_w_bloom\n     -+}\n     -+\n     -+test_expect_success 'git log uses bloom filters if they exist in the latest but not all commit graphs in the chain.' '\n     -+\ttest_bloom_filters_used_when_some_filters_are_missing \"-- A/B\"\n     -+'\n     -+\n     -+test_done\n  -:  ----------- > 14:  b5eb280178f revision.c: add trace2 stats around Bloom filter usage\n  -:  ----------- > 15:  3019ef72881 t4216: add end to end tests for git log with Bloom filters\n 11:  e1b076a714d ! 16:  213abb5d895 commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag\n     @@ -37,8 +37,8 @@\n       \texport GIT_TEST_COMMIT_GRAPH=1\n      +\texport GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=1\n       \texport GIT_TEST_MULTI_PACK_INDEX=1\n     + \texport GIT_TEST_ADD_I_USE_BUILTIN=1\n       \tmake test\n     - \t;;\n      \n       diff --git a/commit-graph.h b/commit-graph.h\n       --- a/commit-graph.h\n     @@ -68,20 +68,6 @@\n       code path for utilizing a file system monitor to speed up detecting\n       new or changed files.\n      \n     - diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n     - --- a/t/t4216-log-bloom.sh\n     - +++ b/t/t4216-log-bloom.sh\n     -@@\n     - test_description='git log for a path with bloom filters'\n     - . ./test-lib.sh\n     - \n     -+GIT_TEST_COMMIT_GRAPH=0\n     -+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n     -+\n     - test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n     - \tgit init &&\n     - \tmkdir A A/B A/B/C &&\n     -\n       diff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\n       --- a/t/t5318-commit-graph.sh\n       +++ b/t/t5318-commit-graph.sh\n\n-- \ngitgitgadget\n"},{"id":"394305","messageId":"c38b9b386ef246cd3144ec5ede991c3cad32f3d9.1585528298.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 06/16] commit-graph: compute Bloom filters for changed paths","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:28Z","receivedAt":"2020-03-30T00:31:50Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd new COMMIT_GRAPH_WRITE_CHANGED_PATHS flag that makes Git compute\nBloom filters for the paths that changed between a commit and it's\nfirst parent, for each commit in the commit-graph.  This computation\nis done on a commit-by-commit basis.\n\nWe will write these Bloom filters to the commit-graph file, to store\nthis data on disk, in the next change in this series.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 32 +++++++++++++++++++++++++++++++-\n commit-graph.h |  3 ++-\n 2 files changed, 33 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex e4f1a5b2f1a..862a00d67ed 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -16,6 +16,7 @@\n #include \"hashmap.h\"\n #include \"replace-object.h\"\n #include \"progress.h\"\n+#include \"bloom.h\"\n \n #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n@@ -789,9 +790,11 @@ struct write_commit_graph_context {\n \tunsigned append:1,\n \t\t report_progress:1,\n \t\t split:1,\n-\t\t check_oids:1;\n+\t\t check_oids:1,\n+\t\t changed_paths:1;\n \n \tconst struct split_commit_graph_opts *split_opts;\n+\tsize_t total_bloom_filter_data_size;\n };\n \n static void write_graph_chunk_fanout(struct hashfile *f,\n@@ -1134,6 +1137,28 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \tstop_progress(&ctx->progress);\n }\n \n+static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n+{\n+\tint i;\n+\tstruct progress *progress = NULL;\n+\n+\tinit_bloom_filters();\n+\n+\tif (ctx->report_progress)\n+\t\tprogress = start_delayed_progress(\n+\t\t\t_(\"Computing commit changed paths Bloom filters\"),\n+\t\t\tctx->commits.nr);\n+\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n+\t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n+\t\tdisplay_progress(progress, i + 1);\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n static int add_ref_to_list(const char *refname,\n \t\t\t   const struct object_id *oid,\n \t\t\t   int flags, void *cb_data)\n@@ -1776,6 +1801,8 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n \tctx->check_oids = flags & COMMIT_GRAPH_WRITE_CHECK_OIDS ? 1 : 0;\n \tctx->split_opts = split_opts;\n+\tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n+\tctx->total_bloom_filter_data_size = 0;\n \n \tif (ctx->split) {\n \t\tstruct commit_graph *g;\n@@ -1870,6 +1897,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \tcompute_generation_numbers(ctx);\n \n+\tif (ctx->changed_paths)\n+\t\tcompute_bloom_filters(ctx);\n+\n \tres = write_commit_graph_file(ctx);\n \n \tif (ctx->split)\ndiff --git a/commit-graph.h b/commit-graph.h\nindex e87a6f63600..86be81219da 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -79,7 +79,8 @@ enum commit_graph_write_flags {\n \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\n \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n \t/* Make sure that each OID in the input is a valid commit OID. */\n-\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3)\n+\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n+\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4),\n };\n \n struct split_commit_graph_opts {\n-- \ngitgitgadget\n\n"},{"id":"394306","messageId":"2d4c0b2da38632424c8bd31ccb2037e0676c3c74.1585528298.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 05/16] diff: halt tree-diff early after max_changes","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:27Z","receivedAt":"2020-03-30T00:31:51Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <dstolee@microsoft.com>\n\nWhen computing the changed-paths bloom filters for the commit-graph,\nwe limit the size of the filter by restricting the number of paths\nin the diff. Instead of computing a large diff and then ignoring the\nresult, it is better to halt the diff computation early.\n\nCreate a new \"max_changes\" option in struct diff_options. If non-zero,\nthen halt the diff computation after discovering strictly more changed\npaths. This includes paths corresponding to trees that change.\n\nUse this max_changes option in the bloom filter calculations. This\nreduces the time taken to compute the filters for the Linux kernel\nrepo from 2m50s to 2m35s. On a large internal repository with ~500\ncommits that perform tree-wide changes, the time reduced from\n6m15s to 3m48s.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c     | 4 +++-\n diff.h      | 5 +++++\n tree-diff.c | 6 ++++++\n 3 files changed, 14 insertions(+), 1 deletion(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 881a9841ede..a16eee92331 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -133,6 +133,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \tint i;\n \tstruct diff_options diffopt;\n+\tint max_changes = 512;\n \n \tif (bloom_filters.slab_size == 0)\n \t\treturn NULL;\n@@ -141,6 +142,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \trepo_diff_setup(r, &diffopt);\n \tdiffopt.flags.recursive = 1;\n+\tdiffopt.max_changes = max_changes;\n \tdiff_setup_done(&diffopt);\n \n \tif (c->parents)\n@@ -149,7 +151,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n \tdiffcore_std(&diffopt);\n \n-\tif (diff_queued_diff.nr <= 512) {\n+\tif (diff_queued_diff.nr <= max_changes) {\n \t\tstruct hashmap pathmap;\n \t\tstruct pathmap_hash_entry *e;\n \t\tstruct hashmap_iter iter;\ndiff --git a/diff.h b/diff.h\nindex 6febe7e3656..9443dc1b003 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -285,6 +285,11 @@ struct diff_options {\n \t/* Number of hexdigits to abbreviate raw format output to. */\n \tint abbrev;\n \n+\t/* If non-zero, then stop computing after this many changes. */\n+\tint max_changes;\n+\t/* For internal use only. */\n+\tint num_changes;\n+\n \tint ita_invisible_in_index;\n /* white-space error highlighting */\n #define WSEH_NEW (1<<12)\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 33ded7f8b3e..f3d303c6e54 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -434,6 +434,9 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \n+\t\tif (opt->max_changes && opt->num_changes > opt->max_changes)\n+\t\t\tbreak;\n+\n \t\tif (opt->pathspec.nr) {\n \t\t\tskip_uninteresting(&t, base, opt);\n \t\t\tfor (i = 0; i < nparent; i++)\n@@ -518,6 +521,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \n \t\t\t/* t↓ */\n \t\t\tupdate_tree_entry(&t);\n+\t\t\topt->num_changes++;\n \t\t}\n \n \t\t/* t > p[imin] */\n@@ -535,6 +539,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\tskip_emit_tp:\n \t\t\t/* ∀ pi=p[imin]  pi↓ */\n \t\t\tupdate_tp_entries(tp, nparent);\n+\t\t\topt->num_changes++;\n \t\t}\n \t}\n \n@@ -552,6 +557,7 @@ struct combine_diff_path *diff_tree_paths(\n \tconst struct object_id **parents_oid, int nparent,\n \tstruct strbuf *base, struct diff_options *opt)\n {\n+\topt->num_changes = 0;\n \tp = ll_diff_tree_paths(p, oid, parents_oid, nparent, base, opt);\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"394307","messageId":"8304c2975207ee847c6709abd71efee918fc4142.1585528298.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 04/16] bloom.c: core Bloom filter implementation for changed paths.","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:26Z","receivedAt":"2020-03-30T00:31:52Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd the core implementation for computing Bloom filters for\nthe paths changed between a commit and it's first parent.\n\nWe fill the Bloom filters as (const char *data, int len) pairs\nas `struct bloom_filters\" within a commit slab.\n\nFilters for commits with no changes and more than 512 changes,\nis represented with a filter of length zero. There is no gain\nin distinguishing between a computed filter of length zero for\na commit with no changes, and an uncomputed filter for new commits\nor for commits with more than 512 changes. The effect on\n`git log -- path` is the same in both cases. We will fall back to\nthe normal diffing algorithm when we can't benefit from the\nexistence of Bloom filters.\n\nHelped-by: Jeff King <peff@peff.net>\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nReviewed-by: Jakub Narębski <jnareb@gmail.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c               | 97 +++++++++++++++++++++++++++++++++++++++++++\n bloom.h               |  8 ++++\n t/helper/test-bloom.c | 20 +++++++++\n t/t0095-bloom.sh      | 47 +++++++++++++++++++++\n 4 files changed, 172 insertions(+)\n\ndiff --git a/bloom.c b/bloom.c\nindex 888b67f1ea6..881a9841ede 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -1,5 +1,18 @@\n #include \"git-compat-util.h\"\n #include \"bloom.h\"\n+#include \"diff.h\"\n+#include \"diffcore.h\"\n+#include \"revision.h\"\n+#include \"hashmap.h\"\n+\n+define_commit_slab(bloom_filter_slab, struct bloom_filter);\n+\n+struct bloom_filter_slab bloom_filters;\n+\n+struct pathmap_hash_entry {\n+    struct hashmap_entry entry;\n+    const char path[FLEX_ARRAY];\n+};\n \n static uint32_t rotate_left(uint32_t value, int32_t count)\n {\n@@ -107,3 +120,87 @@ void add_key_to_filter(const struct bloom_key *key,\n \t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n \t}\n }\n+\n+void init_bloom_filters(void)\n+{\n+\tinit_bloom_filter_slab(&bloom_filters);\n+}\n+\n+struct bloom_filter *get_bloom_filter(struct repository *r,\n+\t\t\t\t      struct commit *c)\n+{\n+\tstruct bloom_filter *filter;\n+\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\tint i;\n+\tstruct diff_options diffopt;\n+\n+\tif (bloom_filters.slab_size == 0)\n+\t\treturn NULL;\n+\n+\tfilter = bloom_filter_slab_at(&bloom_filters, c);\n+\n+\trepo_diff_setup(r, &diffopt);\n+\tdiffopt.flags.recursive = 1;\n+\tdiff_setup_done(&diffopt);\n+\n+\tif (c->parents)\n+\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &diffopt);\n+\telse\n+\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n+\tdiffcore_std(&diffopt);\n+\n+\tif (diff_queued_diff.nr <= 512) {\n+\t\tstruct hashmap pathmap;\n+\t\tstruct pathmap_hash_entry *e;\n+\t\tstruct hashmap_iter iter;\n+\t\thashmap_init(&pathmap, NULL, NULL, 0);\n+\n+\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n+\t\t\tconst char *path = diff_queued_diff.queue[i]->two->path;\n+\n+\t\t\t/*\n+\t\t\t* Add each leading directory of the changed file, i.e. for\n+\t\t\t* 'dir/subdir/file' add 'dir' and 'dir/subdir' as well, so\n+\t\t\t* the Bloom filter could be used to speed up commands like\n+\t\t\t* 'git log dir/subdir', too.\n+\t\t\t*\n+\t\t\t* Note that directories are added without the trailing '/'.\n+\t\t\t*/\n+\t\t\tdo {\n+\t\t\t\tchar *last_slash = strrchr(path, '/');\n+\n+\t\t\t\tFLEX_ALLOC_STR(e, path, path);\n+\t\t\t\thashmap_entry_init(&e->entry, strhash(path));\n+\t\t\t\thashmap_add(&pathmap, &e->entry);\n+\n+\t\t\t\tif (!last_slash)\n+\t\t\t\t\tlast_slash = (char*)path;\n+\t\t\t\t*last_slash = '\\0';\n+\n+\t\t\t} while (*path);\n+\n+\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n+\t\t}\n+\n+\t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n+\t\tfilter->data = xcalloc(filter->len, sizeof(unsigned char));\n+\n+\t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n+\t\t\tstruct bloom_key key;\n+\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, &settings);\n+\t\t\tadd_key_to_filter(&key, filter, &settings);\n+\t\t}\n+\n+\t\thashmap_free_entries(&pathmap, struct pathmap_hash_entry, entry);\n+\t} else {\n+\t\tfor (i = 0; i < diff_queued_diff.nr; i++)\n+\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n+\t\tfilter->data = NULL;\n+\t\tfilter->len = 0;\n+\t}\n+\n+\tfree(diff_queued_diff.queue);\n+\tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n+\n+\treturn filter;\n+}\ndiff --git a/bloom.h b/bloom.h\nindex b9ce422ca2d..85ab8e9423d 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -1,6 +1,9 @@\n #ifndef BLOOM_H\n #define BLOOM_H\n \n+struct commit;\n+struct repository;\n+\n struct bloom_filter_settings {\n \t/*\n \t * The version of the hashing technique being used.\n@@ -73,4 +76,9 @@ void add_key_to_filter(const struct bloom_key *key,\n \t\t\t\t\t   struct bloom_filter *filter,\n \t\t\t\t\t   const struct bloom_filter_settings *settings);\n \n+void init_bloom_filters(void);\n+\n+struct bloom_filter *get_bloom_filter(struct repository *r,\n+\t\t\t\t      struct commit *c);\n+\n #endif\n\\ No newline at end of file\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 20460cde775..f18d1b722e1 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -1,6 +1,7 @@\n #include \"git-compat-util.h\"\n #include \"bloom.h\"\n #include \"test-tool.h\"\n+#include \"commit.h\"\n \n struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \n@@ -32,6 +33,16 @@ static void print_bloom_filter(struct bloom_filter *filter) {\n \tprintf(\"\\n\");\n }\n \n+static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n+{\n+\tstruct commit *c;\n+\tstruct bloom_filter *filter;\n+\tsetup_git_directory();\n+\tc = lookup_commit(the_repository, commit_oid);\n+\tfilter = get_bloom_filter(the_repository, c);\n+\tprint_bloom_filter(filter);\n+}\n+\n int cmd__bloom(int argc, const char **argv)\n {\n \tif (!strcmp(argv[1], \"get_murmur3\")) {\n@@ -57,5 +68,14 @@ int cmd__bloom(int argc, const char **argv)\n \t\tprint_bloom_filter(&filter);\n \t}\n \n+    if (!strcmp(argv[1], \"get_filter_for_commit\")) {\n+\t\tstruct object_id oid;\n+\t\tconst char *end;\n+\t\tif (parse_oid_hex(argv[2], &oid, &end))\n+\t\t\tdie(\"cannot parse oid '%s'\", argv[2]);\n+\t\tinit_bloom_filters();\n+\t\tget_bloom_filter_for_commit(&oid);\n+\t}\n+\n \treturn 0;\n }\n\\ No newline at end of file\ndiff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\nindex 36a086c7c60..8f9eef116dc 100755\n--- a/t/t0095-bloom.sh\n+++ b/t/t0095-bloom.sh\n@@ -67,4 +67,51 @@ test_expect_success 'compute bloom key for test string 2' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'get bloom filters for commit with no changes' '\n+\tgit init &&\n+\tgit commit --allow-empty -m \"c0\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:0\n+\tFilter_Data:\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'get bloom filter for commit with 10 changes' '\n+\trm actual &&\n+\trm expect &&\n+\tmkdir smallDir &&\n+\tfor i in $(test_seq 0 9)\n+\tdo\n+\t\techo $i >smallDir/$i\n+\tdone &&\n+\tgit add smallDir &&\n+\tgit commit -m \"commit with 10 changes\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:25\n+\tFilter_Data:82|a0|65|47|0c|92|90|c0|a1|40|02|a0|e2|40|e0|04|0a|9a|66|cf|80|19|85|42|23|\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success EXPENSIVE 'get bloom filter for commit with 513 changes' '\n+\trm actual &&\n+\trm expect &&\n+\tmkdir bigDir &&\n+\tfor i in $(test_seq 0 512)\n+\tdo\n+\t\techo $i >bigDir/$i\n+\tdone &&\n+\tgit add bigDir &&\n+\tgit commit -m \"commit with 513 changes\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:0\n+\tFilter_Data:\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_done\n\\ No newline at end of file\n-- \ngitgitgadget\n\n"},{"id":"394309","messageId":"d24c85c54ef841eb2a62d95937c1bf9884ba690a.1585528298.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 07/16] commit-graph: examine changed-path objects in pack order","fromName":"Jeff King via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:29Z","receivedAt":"2020-03-30T00:31:54Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nLooking at the diff of commit objects in pack order is much faster than\nin sha1 order, as it gives locality to the access of tree deltas\n(whereas sha1 order is effectively random). Unfortunately the\ncommit-graph code sorts the commits (several times, sometimes as an oid\nand sometimes a pointer-to-commit), and we ultimately traverse in sha1\norder.\n\nInstead, let's remember the position at which we see each commit, and\ntraverse in that order when looking at bloom filters. This drops my time\nfor \"git commit-graph write --changed-paths\" in linux.git from ~4\nminutes to ~1.5 minutes.\n\nProbably the \"--reachable\" code path would want something similar.\n\nOr alternatively, we could use a different data structure (either a\nhash, or maybe even just a bit in \"struct commit\") to keep track of\nwhich oids we've seen, etc instead of sorting. And then we could keep\nthe original order.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 38 +++++++++++++++++++++++++++++++++++---\n 1 file changed, 35 insertions(+), 3 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 862a00d67ed..31b06f878ce 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -17,6 +17,7 @@\n #include \"replace-object.h\"\n #include \"progress.h\"\n #include \"bloom.h\"\n+#include \"commit-slab.h\"\n \n #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n@@ -46,9 +47,32 @@\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n-char *get_commit_graph_filename(struct object_directory *odb)\n+/* Keep track of the order in which commits are added to our list. */\n+define_commit_slab(commit_pos, int);\n+static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n+\n+static void set_commit_pos(struct repository *r, const struct object_id *oid)\n+{\n+\tstatic int32_t max_pos;\n+\tstruct commit *commit = lookup_commit(r, oid);\n+\n+\tif (!commit)\n+\t\treturn; /* should never happen, but be lenient */\n+\n+\t*commit_pos_at(&commit_pos, commit) = max_pos++;\n+}\n+\n+static int commit_pos_cmp(const void *va, const void *vb)\n {\n-\treturn xstrfmt(\"%s/info/commit-graph\", odb->path);\n+\tconst struct commit *a = *(const struct commit **)va;\n+\tconst struct commit *b = *(const struct commit **)vb;\n+\treturn commit_pos_at(&commit_pos, a) -\n+\t       commit_pos_at(&commit_pos, b);\n+}\n+\n+char *get_commit_graph_filename(struct object_directory *obj_dir)\n+{\n+\treturn xstrfmt(\"%s/info/commit-graph\", obj_dir->path);\n }\n \n static char *get_split_graph_filename(struct object_directory *odb,\n@@ -1021,6 +1045,8 @@ static int add_packed_commits(const struct object_id *oid,\n \toidcpy(&(ctx->oids.list[ctx->oids.nr]), oid);\n \tctx->oids.nr++;\n \n+\tset_commit_pos(ctx->r, oid);\n+\n \treturn 0;\n }\n \n@@ -1141,6 +1167,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n {\n \tint i;\n \tstruct progress *progress = NULL;\n+\tstruct commit **sorted_commits;\n \n \tinit_bloom_filters();\n \n@@ -1149,13 +1176,18 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \t\t\t_(\"Computing commit changed paths Bloom filters\"),\n \t\t\tctx->commits.nr);\n \n+\tALLOC_ARRAY(sorted_commits, ctx->commits.nr);\n+\tCOPY_ARRAY(sorted_commits, ctx->commits.list, ctx->commits.nr);\n+\tQSORT(sorted_commits, ctx->commits.nr, commit_pos_cmp);\n+\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tstruct commit *c = ctx->commits.list[i];\n+\t\tstruct commit *c = sorted_commits[i];\n \t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n \t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n \t\tdisplay_progress(progress, i + 1);\n \t}\n \n+\tfree(sorted_commits);\n \tstop_progress(&progress);\n }\n \n-- \ngitgitgadget\n\n"},{"id":"394311","messageId":"1e4663523ded1e97cc43a5987776eb2a7d6db451.1585528299.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 10/16] commit-graph: write Bloom filters to commit graph file","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:32Z","receivedAt":"2020-03-30T00:31:56Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nUpdate the technical documentation for commit-graph-format with\nthe formats for the Bloom filter index (BIDX) and Bloom filter\ndata (BDAT) chunks. Write the computed Bloom filters information\nto the commit graph file using this format.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n .../technical/commit-graph-format.txt         |  30 +++++\n commit-graph.c                                | 113 +++++++++++++++++-\n commit-graph.h                                |   5 +\n 3 files changed, 147 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\nindex a4f17441aed..de56f9f1efd 100644\n--- a/Documentation/technical/commit-graph-format.txt\n+++ b/Documentation/technical/commit-graph-format.txt\n@@ -17,6 +17,9 @@ metadata, including:\n - The parents of the commit, stored using positional references within\n   the graph file.\n \n+- The Bloom filter of the commit carrying the paths that were changed between\n+  the commit and its first parent, if requested.\n+\n These positional references are stored as unsigned 32-bit integers\n corresponding to the array position within the list of commit OIDs. Due\n to some special constants we use to track parents, we can store at most\n@@ -93,6 +96,33 @@ CHUNK DATA:\n       positions for the parents until reaching a value with the most-significant\n       bit on. The other bits correspond to the position of the last parent.\n \n+  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) (N * 4 bytes) [Optional]\n+    * The ith entry, BIDX[i], stores the number of 8-byte word blocks in all\n+      Bloom filters from commit 0 to commit i (inclusive) in lexicographic\n+      order. The Bloom filter for the i-th commit spans from BIDX[i-1] to\n+      BIDX[i] (plus header length), where BIDX[-1] is 0.\n+    * The BIDX chunk is ignored if the BDAT chunk is not present.\n+\n+  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n+    * It starts with header consisting of three unsigned 32-bit integers:\n+      - Version of the hash algorithm being used. We currently only support\n+\tvalue 1 which corresponds to the 32-bit version of the murmur3 hash\n+\timplemented exactly as described in\n+\thttps://en.wikipedia.org/wiki/MurmurHash#Algorithm and the double\n+\thashing technique using seed values 0x293ae76f and 0x7e646e2 as\n+\tdescribed in https://doi.org/10.1007/978-3-540-30494-4_26 \"Bloom Filters\n+\tin Probabilistic Verification\"\n+      - The number of times a path is hashed and hence the number of bit positions\n+\t      that cumulatively determine whether a file is present in the commit.\n+      - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n+\t      contains 'n' entries, then the filter size is the minimum number of 64-bit\n+\t      words that contain n*b bits.\n+    * The rest of the chunk is the concatenation of all the computed Bloom\n+      filters for the commits in lexicographic order.\n+    * Note: Commits with no changes or more than 512 changes have Bloom filters\n+      of length zero.\n+    * The BDAT chunk is present if and only if BIDX is present.\n+\n   Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n       This list of H-byte hashes describe a set of B commit-graph files that\n       form a commit-graph chain. The graph position for the ith commit in this\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 732c81fa1b2..a8b6b5cca5d 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -24,8 +24,10 @@\n #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n+#define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n+#define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 5\n+#define MAX_NUM_CHUNKS 7\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -319,6 +321,32 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n \t\t\t\tchunk_repeated = 1;\n \t\t\telse\n \t\t\t\tgraph->chunk_base_graphs = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_BLOOMINDEXES:\n+\t\t\tif (graph->chunk_bloom_indexes)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_bloom_indexes = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_BLOOMDATA:\n+\t\t\tif (graph->chunk_bloom_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse {\n+\t\t\t\tuint32_t hash_version;\n+\t\t\t\tgraph->chunk_bloom_data = data + chunk_offset;\n+\t\t\t\thash_version = get_be32(data + chunk_offset);\n+\n+\t\t\t\tif (hash_version != 1)\n+\t\t\t\t\tbreak;\n+\n+\t\t\t\tgraph->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n+\t\t\t\tgraph->bloom_filter_settings->hash_version = hash_version;\n+\t\t\t\tgraph->bloom_filter_settings->num_hashes = get_be32(data + chunk_offset + 4);\n+\t\t\t\tgraph->bloom_filter_settings->bits_per_entry = get_be32(data + chunk_offset + 8);\n+\t\t\t}\n+\t\t\tbreak;\n \t\t}\n \n \t\tif (chunk_repeated) {\n@@ -337,6 +365,15 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n \t\tlast_chunk_offset = chunk_offset;\n \t}\n \n+\tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data) {\n+\t\tinit_bloom_filters();\n+\t} else {\n+\t\t/* We need both the bloom chunks to exist together. Else ignore the data */\n+\t\tgraph->chunk_bloom_indexes = NULL;\n+\t\tgraph->chunk_bloom_data = NULL;\n+\t\tgraph->bloom_filter_settings = NULL;\n+\t}\n+\n \thashcpy(graph->oid.hash, graph->data + graph->data_len - graph->hash_len);\n \n \tif (verify_commit_graph_lite(graph)) {\n@@ -1034,6 +1071,59 @@ static void write_graph_chunk_extra_edges(struct hashfile *f,\n \t}\n }\n \n+static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n+\t\t\t\t\t    struct write_commit_graph_context *ctx)\n+{\n+\tstruct commit **list = ctx->commits.list;\n+\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n+\tuint32_t cur_pos = 0;\n+\tstruct progress *progress = NULL;\n+\tint i = 0;\n+\n+\tif (ctx->report_progress)\n+\t\tprogress = start_delayed_progress(\n+\t\t\t_(\"Writing changed paths Bloom filters index\"),\n+\t\t\tctx->commits.nr);\n+\n+\twhile (list < last) {\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tcur_pos += filter->len;\n+\t\tdisplay_progress(progress, ++i);\n+\t\thashwrite_be32(f, cur_pos);\n+\t\tlist++;\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n+static void write_graph_chunk_bloom_data(struct hashfile *f,\n+\t\t\t\t\t struct write_commit_graph_context *ctx,\n+\t\t\t\t\t const struct bloom_filter_settings *settings)\n+{\n+\tstruct commit **list = ctx->commits.list;\n+\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n+\tstruct progress *progress = NULL;\n+\tint i = 0;\n+\n+\tif (ctx->report_progress)\n+\t\tprogress = start_delayed_progress(\n+\t\t\t_(\"Writing changed paths Bloom filters data\"),\n+\t\t\tctx->commits.nr);\n+\n+\thashwrite_be32(f, settings->hash_version);\n+\thashwrite_be32(f, settings->num_hashes);\n+\thashwrite_be32(f, settings->bits_per_entry);\n+\n+\twhile (list < last) {\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tdisplay_progress(progress, ++i);\n+\t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n+\t\tlist++;\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n static int oid_compare(const void *_a, const void *_b)\n {\n \tconst struct object_id *a = (const struct object_id *)_a;\n@@ -1438,6 +1528,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tstruct strbuf progress_title = STRBUF_INIT;\n \tint num_chunks = 3;\n \tstruct object_id file_hash;\n+\tconst struct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \n \tif (ctx->split) {\n \t\tstruct strbuf tmp_file = STRBUF_INIT;\n@@ -1482,6 +1573,12 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tnum_chunks++;\n \t}\n+\tif (ctx->changed_paths) {\n+\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMINDEXES;\n+\t\tnum_chunks++;\n+\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMDATA;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_commit_graphs_after > 1) {\n \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BASE;\n \t\tnum_chunks++;\n@@ -1500,6 +1597,15 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t\t\t\t\t\t4 * ctx->num_extra_edges;\n \t\tnum_chunks++;\n \t}\n+\tif (ctx->changed_paths) {\n+\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n+\t\t\t\t\t\tsizeof(uint32_t) * ctx->commits.nr;\n+\t\tnum_chunks++;\n+\n+\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n+\t\t\t\t\t\tsizeof(uint32_t) * 3 + ctx->total_bloom_filter_data_size;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_commit_graphs_after > 1) {\n \t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n \t\t\t\t\t\thashsz * (ctx->num_commit_graphs_after - 1);\n@@ -1537,6 +1643,10 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \twrite_graph_chunk_data(f, hashsz, ctx);\n \tif (ctx->num_extra_edges)\n \t\twrite_graph_chunk_extra_edges(f, ctx);\n+\tif (ctx->changed_paths) {\n+\t\twrite_graph_chunk_bloom_indexes(f, ctx);\n+\t\twrite_graph_chunk_bloom_data(f, ctx, &bloom_settings);\n+\t}\n \tif (ctx->num_commit_graphs_after > 1 &&\n \t    write_graph_chunk_base(f, ctx)) {\n \t\treturn -1;\n@@ -2184,6 +2294,7 @@ void free_commit_graph(struct commit_graph *g)\n \t\tclose(g->graph_fd);\n \t}\n \tfree(g->filename);\n+\tfree(g->bloom_filter_settings);\n \tfree(g);\n }\n \ndiff --git a/commit-graph.h b/commit-graph.h\nindex 86be81219da..8e7a8e0e5b2 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -11,6 +11,7 @@\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n \n struct commit;\n+struct bloom_filter_settings;\n \n char *get_commit_graph_filename(struct object_directory *odb);\n int open_commit_graph(const char *graph_file, int *fd, struct stat *st);\n@@ -59,6 +60,10 @@ struct commit_graph {\n \tconst unsigned char *chunk_commit_data;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n+\tconst unsigned char *chunk_bloom_indexes;\n+\tconst unsigned char *chunk_bloom_data;\n+\n+\tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n struct commit_graph *load_commit_graph_one_fd_st(int fd, struct stat *st,\n-- \ngitgitgadget\n\n"},{"id":"394314","messageId":"b18af58aa3ed7644955bc027a9e6f710a1d36530.1585528299.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 13/16] revision.c: use Bloom filters to speed up path based revision walks","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:35Z","receivedAt":"2020-03-30T00:31:56Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nRevision walk will now use Bloom filters for commits to speed up\nrevision walks for a particular path (for computing history for\nthat path), if they are present in the commit-graph file.\n\nWe load the Bloom filters during the prepare_revision_walk step,\ncurrently only when dealing with a single pathspec. Extending\nit to work with multiple pathspecs can be explored and built on\ntop of this series in the future.\n\nWhile comparing trees in rev_compare_trees(), if the Bloom filter\nsays that the file is not different between the two trees, we don't\nneed to compute the expensive diff. This is where we get our\nperformance gains. The other response of the Bloom filter is '`:maybe',\nin which case we fall back to the full diff calculation to determine\nif the path was changed in the commit.\n\nWe do not try to use Bloom filters when the '--walk-reflogs' option\nis specified. The '--walk-reflogs' option does not walk the commit\nancestry chain like the rest of the options. Incorporating the\nperformance gains when walking reflog entries would add more\ncomplexity, and can be explored in a later series.\n\nPerformance Gains:\nWe tested the performance of `git log -- <path>` on the git repo, the linux\nand some internal large repos, with a variety of paths of varying depths.\n\nOn the git and linux repos:\n- we observed a 2x to 5x speed up.\n\nOn a large internal repo with files seated 6-10 levels deep in the tree:\n- we observed 10x to 20x speed ups, with some paths going up to 28 times\n  faster.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com\nHelped-by: SZEDER Gábor <szeder.dev@gmail.com>\nHelped-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c    | 20 +++++++++++++\n bloom.h    |  4 +++\n revision.c | 85 ++++++++++++++++++++++++++++++++++++++++++++++++++++--\n revision.h | 11 +++++++\n 4 files changed, 118 insertions(+), 2 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 151d598ce7b..dd9bab9bbd6 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -254,3 +254,23 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \treturn filter;\n }\n+\n+int bloom_filter_contains(const struct bloom_filter *filter,\n+\t\t\t  const struct bloom_key *key,\n+\t\t\t  const struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tuint64_t mod = filter->len * BITS_PER_WORD;\n+\n+\tif (!mod)\n+\t\treturn -1;\n+\n+\tfor (i = 0; i < settings->num_hashes; i++) {\n+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n+\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n+\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n+\t\t\treturn 0;\n+\t}\n+\n+\treturn 1;\n+}\n\\ No newline at end of file\ndiff --git a/bloom.h b/bloom.h\nindex 760d7122374..b935186425d 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -83,4 +83,8 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \t\t\t\t      struct commit *c,\n \t\t\t\t      int compute_if_not_present);\n \n+int bloom_filter_contains(const struct bloom_filter *filter,\n+\t\t\t  const struct bloom_key *key,\n+\t\t\t  const struct bloom_filter_settings *settings);\n+\n #endif\n\\ No newline at end of file\ndiff --git a/revision.c b/revision.c\nindex 8136929e236..d3fcb7c6ff6 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -29,6 +29,7 @@\n #include \"prio-queue.h\"\n #include \"hashmap.h\"\n #include \"utf8.h\"\n+#include \"bloom.h\"\n \n volatile show_early_output_fn_t show_early_output;\n \n@@ -624,11 +625,80 @@ static void file_change(struct diff_options *options,\n \toptions->flags.has_changes = 1;\n }\n \n+static void prepare_to_use_bloom_filter(struct rev_info *revs)\n+{\n+\tstruct pathspec_item *pi;\n+\tchar *path_alloc = NULL;\n+\tconst char *path;\n+\tint last_index;\n+\tint len;\n+\n+\tif (!revs->commits)\n+\t    return;\n+\n+\trepo_parse_commit(revs->repo, revs->commits->item);\n+\n+\tif (!revs->repo->objects->commit_graph)\n+\t\treturn;\n+\n+\trevs->bloom_filter_settings = revs->repo->objects->commit_graph->bloom_filter_settings;\n+\tif (!revs->bloom_filter_settings)\n+\t\treturn;\n+\n+\tpi = &revs->pruning.pathspec.items[0];\n+\tlast_index = pi->len - 1;\n+\n+\t/* remove single trailing slash from path, if needed */\n+\tif (pi->match[last_index] == '/') {\n+\t    path_alloc = xstrdup(pi->match);\n+\t    path_alloc[last_index] = '\\0';\n+\t    path = path_alloc;\n+\t} else\n+\t    path = pi->match;\n+\n+\tlen = strlen(path);\n+\n+\trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n+\tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n+\n+\tfree(path_alloc);\n+}\n+\n+static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n+\t\t\t\t\t\t struct commit *commit)\n+{\n+\tstruct bloom_filter *filter;\n+\tint result;\n+\n+\tif (!revs->repo->objects->commit_graph)\n+\t\treturn -1;\n+\n+\tif (commit->generation == GENERATION_NUMBER_INFINITY)\n+\t\treturn -1;\n+\n+\tfilter = get_bloom_filter(revs->repo, commit, 0);\n+\n+\tif (!filter) {\n+\t\treturn -1;\n+\t}\n+\n+\tif (!filter->len) {\n+\t\treturn -1;\n+\t}\n+\n+\tresult = bloom_filter_contains(filter,\n+\t\t\t\t       revs->bloom_key,\n+\t\t\t\t       revs->bloom_filter_settings);\n+\n+\treturn result;\n+}\n+\n static int rev_compare_tree(struct rev_info *revs,\n-\t\t\t    struct commit *parent, struct commit *commit)\n+\t\t\t    struct commit *parent, struct commit *commit, int nth_parent)\n {\n \tstruct tree *t1 = get_commit_tree(parent);\n \tstruct tree *t2 = get_commit_tree(commit);\n+\tint bloom_ret = 1;\n \n \tif (!t1)\n \t\treturn REV_TREE_NEW;\n@@ -653,11 +723,19 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\treturn REV_TREE_SAME;\n \t}\n \n+\tif (revs->bloom_key && !nth_parent) {\n+\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n+\n+\t\tif (bloom_ret == 0)\n+\t\t\treturn REV_TREE_SAME;\n+\t}\n+\n \ttree_difference = REV_TREE_SAME;\n \trevs->pruning.flags.has_changes = 0;\n \tif (diff_tree_oid(&t1->object.oid, &t2->object.oid, \"\",\n \t\t\t   &revs->pruning) < 0)\n \t\treturn REV_TREE_DIFFERENT;\n+\n \treturn tree_difference;\n }\n \n@@ -855,7 +933,7 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit)\n \t\t\tdie(\"cannot simplify commit %s (because of %s)\",\n \t\t\t    oid_to_hex(&commit->object.oid),\n \t\t\t    oid_to_hex(&p->object.oid));\n-\t\tswitch (rev_compare_tree(revs, p, commit)) {\n+\t\tswitch (rev_compare_tree(revs, p, commit, nth_parent)) {\n \t\tcase REV_TREE_SAME:\n \t\t\tif (!revs->simplify_history || !relevant_commit(p)) {\n \t\t\t\t/* Even if a merge with an uninteresting\n@@ -3362,6 +3440,8 @@ int prepare_revision_walk(struct rev_info *revs)\n \t\t\t\t       FOR_EACH_OBJECT_PROMISOR_ONLY);\n \t}\n \n+\tif (revs->pruning.pathspec.nr == 1 && !revs->reflog_info)\n+\t\tprepare_to_use_bloom_filter(revs);\n \tif (revs->no_walk != REVISION_WALK_NO_WALK_UNSORTED)\n \t\tcommit_list_sort_by_date(&revs->commits);\n \tif (revs->no_walk)\n@@ -3379,6 +3459,7 @@ int prepare_revision_walk(struct rev_info *revs)\n \t\tsimplify_merges(revs);\n \tif (revs->children.name)\n \t\tset_children(revs);\n+\n \treturn 0;\n }\n \ndiff --git a/revision.h b/revision.h\nindex 475f048fb61..7c026fe41fc 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -56,6 +56,8 @@ struct repository;\n struct rev_info;\n struct string_list;\n struct saved_parents;\n+struct bloom_key;\n+struct bloom_filter_settings;\n define_shared_commit_slab(revision_sources, char *);\n \n struct rev_cmdline_info {\n@@ -291,6 +293,15 @@ struct rev_info {\n \tstruct revision_sources *sources;\n \n \tstruct topo_walk_info *topo_walk_info;\n+\n+\t/* Commit graph bloom filter fields */\n+\t/* The bloom filter key for the pathspec */\n+\tstruct bloom_key *bloom_key;\n+\t/*\n+\t * The bloom filter settings used to generate the key.\n+\t * This is loaded from the commit-graph being used.\n+\t */\n+\tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n int ref_excluded(struct string_list *, const char *path);\n-- \ngitgitgadget\n\n"},{"id":"394310","messageId":"b5eb280178ff314e2f28d0b9e22910a1464b10d6.1585528299.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 14/16] revision.c: add trace2 stats around Bloom filter usage","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:36Z","receivedAt":"2020-03-30T00:31:57Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd trace2 statistics around Bloom filter usage and behavior\nfor 'git log -- path' commands that are hoping to benefit from\nthe presence of computed changed paths Bloom filters.\n\nThese statistics are great for performance analysis work and\nfor formal testing, which we will see in the commit following\nthis one.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com\nHelped-by: SZEDER Gábor <szeder.dev@gmail.com>\nHelped-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n revision.c | 41 +++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 41 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex d3fcb7c6ff6..2b06ee739c8 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -30,6 +30,7 @@\n #include \"hashmap.h\"\n #include \"utf8.h\"\n #include \"bloom.h\"\n+#include \"json-writer.h\"\n \n volatile show_early_output_fn_t show_early_output;\n \n@@ -625,6 +626,30 @@ static void file_change(struct diff_options *options,\n \toptions->flags.has_changes = 1;\n }\n \n+static int bloom_filter_atexit_registered;\n+static unsigned int count_bloom_filter_maybe;\n+static unsigned int count_bloom_filter_definitely_not;\n+static unsigned int count_bloom_filter_false_positive;\n+static unsigned int count_bloom_filter_not_present;\n+static unsigned int count_bloom_filter_length_zero;\n+\n+static void trace2_bloom_filter_statistics_atexit(void)\n+{\n+\tstruct json_writer jw = JSON_WRITER_INIT;\n+\n+\tjw_object_begin(&jw, 0);\n+\tjw_object_intmax(&jw, \"filter_not_present\", count_bloom_filter_not_present);\n+\tjw_object_intmax(&jw, \"zero_length_filter\", count_bloom_filter_length_zero);\n+\tjw_object_intmax(&jw, \"maybe\", count_bloom_filter_maybe);\n+\tjw_object_intmax(&jw, \"definitely_not\", count_bloom_filter_definitely_not);\n+\tjw_object_intmax(&jw, \"false_positive\", count_bloom_filter_false_positive);\n+\tjw_end(&jw);\n+\n+\ttrace2_data_json(\"bloom\", the_repository, \"statistics\", &jw);\n+\n+\tjw_release(&jw);\n+}\n+\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n@@ -661,6 +686,11 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n \tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n \n+\tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n+\t\tatexit(trace2_bloom_filter_statistics_atexit);\n+\t\tbloom_filter_atexit_registered = 1;\n+\t}\n+\n \tfree(path_alloc);\n }\n \n@@ -679,10 +709,12 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \tfilter = get_bloom_filter(revs->repo, commit, 0);\n \n \tif (!filter) {\n+\t\tcount_bloom_filter_not_present++;\n \t\treturn -1;\n \t}\n \n \tif (!filter->len) {\n+\t\tcount_bloom_filter_length_zero++;\n \t\treturn -1;\n \t}\n \n@@ -690,6 +722,11 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\t\t\t       revs->bloom_key,\n \t\t\t\t       revs->bloom_filter_settings);\n \n+\tif (result)\n+\t\tcount_bloom_filter_maybe++;\n+\telse\n+\t\tcount_bloom_filter_definitely_not++;\n+\n \treturn result;\n }\n \n@@ -736,6 +773,10 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\t   &revs->pruning) < 0)\n \t\treturn REV_TREE_DIFFERENT;\n \n+\tif (!nth_parent)\n+\t\tif (bloom_ret == 1 && tree_difference == REV_TREE_SAME)\n+\t\t\tcount_bloom_filter_false_positive++;\n+\n \treturn tree_difference;\n }\n \n-- \ngitgitgadget\n\n"},{"id":"394312","messageId":"3019ef72881289589d5b76c8c0c8177a2c464972.1585528299.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 15/16] t4216: add end to end tests for git log with Bloom filters","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:37Z","receivedAt":"2020-03-30T00:31:59Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nThese tests exercises writing commit graph with Bloom filters\nand exercises 'git log -- path' with all the applicable\noptions. They check that the output is the same with and\nwithout Bloom filters, confirm Bloom filters were used by\nchecking if trace2 statistics were logged correctly.\n\nAlso confirms cases where Bloom filters are not used:\n1. Multiple path specs,\n2. --walk-reflogs (see patch titled 'revision.c: use Bloom filters...'\n   for details,\n3. If the latest commit graph does not have Bloom filters\n\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n t/helper/test-read-graph.c |   4 +\n t/t4216-log-bloom.sh       | 155 +++++++++++++++++++++++++++++++++++++\n 2 files changed, 159 insertions(+)\n create mode 100755 t/t4216-log-bloom.sh\n\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex f8a461767ca..4223ff32fb6 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -45,6 +45,10 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" commit_metadata\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n+\tif (graph->chunk_bloom_indexes)\n+\t\tprintf(\" bloom_indexes\");\n+\tif (graph->chunk_bloom_data)\n+\t\tprintf(\" bloom_data\");\n \tprintf(\"\\n\");\n \n \tUNLEAK(graph);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nnew file mode 100755\nindex 00000000000..38accd272df\n--- /dev/null\n+++ b/t/t4216-log-bloom.sh\n@@ -0,0 +1,155 @@\n+#!/bin/sh\n+\n+test_description='git log for a path with Bloom filters'\n+. ./test-lib.sh\n+\n+GIT_TEST_COMMIT_GRAPH=0\n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n+\n+test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n+\tgit init &&\n+\tmkdir A A/B A/B/C &&\n+\ttest_commit c1 A/file1 &&\n+\ttest_commit c2 A/B/file2 &&\n+\ttest_commit c3 A/B/C/file3 &&\n+\ttest_commit c4 A/file1 &&\n+\ttest_commit c5 A/B/file2 &&\n+\ttest_commit c6 A/B/C/file3 &&\n+\ttest_commit c7 A/file1 &&\n+\ttest_commit c8 A/B/file2 &&\n+\ttest_commit c9 A/B/C/file3 &&\n+\ttest_commit c10 file_to_be_deleted &&\n+\tgit checkout -b side HEAD~4 &&\n+\ttest_commit side-1 file4 &&\n+\tgit checkout master &&\n+\tgit merge side &&\n+\ttest_commit c11 file5 &&\n+\tmv file5 file5_renamed &&\n+\tgit add file5_renamed &&\n+\tgit commit -m \"rename\" &&\n+\trm file_to_be_deleted &&\n+\tgit add . &&\n+\tgit commit -m \"file removed\" &&\n+\tgit commit-graph write --reachable --changed-paths\n+'\n+graph_read_expect () {\n+\tNUM_CHUNKS=5\n+\tcat >expect <<- EOF\n+\theader: 43475048 1 1 $NUM_CHUNKS 0\n+\tnum_commits: $1\n+\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tEOF\n+\ttest-tool read-graph >actual &&\n+\ttest_cmp expect actual\n+}\n+\n+test_expect_success 'commit-graph write wrote out the bloom chunks' '\n+\tgraph_read_expect 15\n+'\n+\n+# Turn off any inherited trace2 settings for this test.\n+sane_unset GIT_TRACE2 GIT_TRACE2_PERF GIT_TRACE2_EVENT\n+sane_unset GIT_TRACE2_PERF_BRIEF\n+sane_unset GIT_TRACE2_CONFIG_PARAMS\n+\n+setup () {\n+\trm \"$TRASH_DIRECTORY/trace.perf\"\n+\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n+\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n+}\n+\n+test_bloom_filters_used () {\n+\tlog_args=$1\n+\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":0,\\\"zero_length_filter\\\":0,\\\"maybe\\\"\"\n+\tsetup \"$log_args\" &&\n+\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && \n+\ttest_cmp log_wo_bloom log_w_bloom &&\n+    test_path_is_file \"$TRASH_DIRECTORY/trace.perf\"\n+}\n+\n+test_bloom_filters_not_used () {\n+\tlog_args=$1\n+\tsetup \"$log_args\" &&\n+\t!(grep -q \"statistics:{\\\"filter_not_present\\\":\" \"$TRASH_DIRECTORY/trace.perf\") && \n+\ttest_cmp log_wo_bloom log_w_bloom\n+}\n+\n+for path in A A/B A/B/C A/file1 A/B/file2 A/B/C/file3 file4 file5 file5_renamed file_to_be_deleted\n+do\n+\tfor option in \"\" \\\n+              \"--all\" \\\n+\t\t      \"--full-history\" \\\n+\t\t      \"--full-history --simplify-merges\" \\\n+\t\t      \"--simplify-merges\" \\\n+\t\t      \"--simplify-by-decoration\" \\\n+\t\t      \"--follow\" \\\n+\t\t      \"--first-parent\" \\\n+\t\t      \"--topo-order\" \\\n+\t\t      \"--date-order\" \\\n+\t\t      \"--author-date-order\" \\\n+\t\t      \"--ancestry-path side..master\"\n+\tdo\n+\t\ttest_expect_success \"git log option: $option for path: $path\" '\n+\t\t\ttest_bloom_filters_used \"$option -- $path\"\n+\t\t'\n+\tdone\n+done\n+\n+test_expect_success 'git log -- folder works with and without the trailing slash' '\n+\ttest_bloom_filters_used \"-- A\" &&\n+\ttest_bloom_filters_used \"-- A/\"\n+'\n+\n+test_expect_success 'git log for path that does not exist. ' '\n+\ttest_bloom_filters_used \"-- path_does_not_exist\"\n+'\n+\n+test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n+\ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n+'\n+\n+test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n+\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n+'\n+\n+test_expect_success 'git log with wildcard that resolves to a single path uses Bloom filters' '\n+\ttest_bloom_filters_used \"-- *4\" &&\n+\ttest_bloom_filters_used \"-- *renamed\"\n+'\n+\n+test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n+\ttest_bloom_filters_not_used \"-- *\" &&\n+\ttest_bloom_filters_not_used \"-- file*\"\n+'\n+\n+test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n+\ttest_commit c14 A/anotherFile2 &&\n+\ttest_commit c15 A/B/anotherFile2 &&\n+\ttest_commit c16 A/B/C/anotherFile2 &&\n+\tGIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0 git commit-graph write --reachable --split &&\n+\ttest_line_count = 2 .git/objects/info/commit-graphs/commit-graph-chain\n+'\n+\n+test_expect_success 'Do not use Bloom filters if the latest graph does not have Bloom filters.' '\n+\ttest_bloom_filters_not_used \"-- A/B\"\n+'\n+\n+test_expect_success 'setup - add commit-graph to the chain with Bloom filters' '\n+\ttest_commit c17 A/anotherFile3 &&\n+\tgit commit-graph write --reachable --changed-paths --split &&\n+\ttest_line_count = 3 .git/objects/info/commit-graphs/commit-graph-chain\n+'\n+\n+test_bloom_filters_used_when_some_filters_are_missing () {\n+\tlog_args=$1\n+\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":3,\\\"zero_length_filter\\\":0,\\\"maybe\\\":8,\\\"definitely_not\\\":6\"\n+\tsetup \"$log_args\" &&\n+\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && \n+\ttest_cmp log_wo_bloom log_w_bloom\n+}\n+\n+test_expect_success 'Use Bloom filters if they exist in the latest but not all commit graphs in the chain.' '\n+\ttest_bloom_filters_used_when_some_filters_are_missing \"-- A/B\"\n+'\n+\n+test_done\n\\ No newline at end of file\n-- \ngitgitgadget\n\n"},{"id":"394313","messageId":"213abb5d8951aac010c2c6be0f7fa45f911e5106.1585528299.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 16/16] commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:38Z","receivedAt":"2020-03-30T00:31:59Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag to the test setup suite\nin order to toggle writing Bloom filters when running any of the git tests.\nIf set to true, we will compute and write Bloom filters every time a test\ncalls `git commit-graph write`, as if the `--changed-paths` option was\npassed in.\n\nThe test suite passes when GIT_TEST_COMMIT_GRAPH and\nGIT_TEST_COMMIT_GRAPH_CHANGED_PATHS are enabled.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n builtin/commit-graph.c        | 3 ++-\n ci/run-build-and-tests.sh     | 1 +\n commit-graph.h                | 1 +\n t/README                      | 5 +++++\n t/t5318-commit-graph.sh       | 2 ++\n t/t5324-split-commit-graph.sh | 1 +\n 6 files changed, 12 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\nindex cacb5d04a80..59009837dc9 100644\n--- a/builtin/commit-graph.c\n+++ b/builtin/commit-graph.c\n@@ -171,7 +171,8 @@ static int graph_write(int argc, const char **argv)\n \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n \tif (opts.progress)\n \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n-\tif (opts.enable_changed_paths)\n+\tif (opts.enable_changed_paths ||\n+\t    git_env_bool(GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS, 0))\n \t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n \n \tread_replace_refs = 0;\ndiff --git a/ci/run-build-and-tests.sh b/ci/run-build-and-tests.sh\nindex 4df54c4efea..17e25aade96 100755\n--- a/ci/run-build-and-tests.sh\n+++ b/ci/run-build-and-tests.sh\n@@ -19,6 +19,7 @@ linux-gcc)\n \texport GIT_TEST_OE_SIZE=10\n \texport GIT_TEST_OE_DELTA_SIZE=5\n \texport GIT_TEST_COMMIT_GRAPH=1\n+\texport GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=1\n \texport GIT_TEST_MULTI_PACK_INDEX=1\n \texport GIT_TEST_ADD_I_USE_BUILTIN=1\n \tmake test\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 8e7a8e0e5b2..8655d064c14 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -9,6 +9,7 @@\n \n #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n+#define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n \n struct commit;\n struct bloom_filter_settings;\ndiff --git a/t/README b/t/README\nindex 369e3a9ded8..4f53da53a15 100644\n--- a/t/README\n+++ b/t/README\n@@ -378,6 +378,11 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n+commit-graph write to compute and write changed path Bloom filters for\n+every 'git commit-graph write', as if the `--changed-paths` option was\n+passed in.\n+\n GIT_TEST_FSMONITOR=$PWD/t7519/fsmonitor-all exercises the fsmonitor\n code path for utilizing a file system monitor to speed up detecting\n new or changed files.\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 9bf920ae171..18304a65e4d 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -3,6 +3,8 @@\n test_description='commit graph'\n . ./test-lib.sh\n \n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n+\n test_expect_success 'setup full repo' '\n \tmkdir full &&\n \tcd \"$TRASH_DIRECTORY/full\" &&\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 53b2e6b4555..d3f1f2c4a71 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -4,6 +4,7 @@ test_description='split commit graph'\n . ./test-lib.sh\n \n GIT_TEST_COMMIT_GRAPH=0\n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n \n test_expect_success 'setup repo' '\n \tgit init &&\n-- \ngitgitgadget\n"},{"id":"394317","messageId":"7e450e452363d32927a4afd979b5c544ac541fcc.1585528299.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 12/16] commit-graph: add --changed-paths option to write subcommand","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:34Z","receivedAt":"2020-03-30T00:32:00Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd --changed-paths option to git commit-graph write. This option will\nallow users to compute information about the paths that have changed\nbetween a commit and its first parent, and write it into the commit graph\nfile. If the option is passed to the write subcommand we set the\nCOMMIT_GRAPH_WRITE_BLOOM_FILTERS flag and pass it down to the\ncommit-graph logic.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n Documentation/git-commit-graph.txt | 5 +++++\n builtin/commit-graph.c             | 9 +++++++--\n 2 files changed, 12 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-commit-graph.txt b/Documentation/git-commit-graph.txt\nindex 28d1fee5053..f4b13c005b8 100644\n--- a/Documentation/git-commit-graph.txt\n+++ b/Documentation/git-commit-graph.txt\n@@ -57,6 +57,11 @@ or `--stdin-packs`.)\n With the `--append` option, include all commits that are present in the\n existing commit-graph file.\n +\n+With the `--changed-paths` option, compute and write information about the\n+paths changed between a commit and it's first parent. This operation can\n+take a while on large repositories. It provides significant performance gains\n+for getting history of a directory or a file with `git log -- <path>`.\n++\n With the `--split` option, write the commit-graph as a chain of multiple\n commit-graph files stored in `<dir>/info/commit-graphs`. The new commits\n not already in the commit-graph are added in a new \"tip\" file. This file\ndiff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\nindex d1ab6625f63..cacb5d04a80 100644\n--- a/builtin/commit-graph.c\n+++ b/builtin/commit-graph.c\n@@ -9,7 +9,7 @@\n \n static char const * const builtin_commit_graph_usage[] = {\n \tN_(\"git commit-graph verify [--object-dir <objdir>] [--shallow] [--[no-]progress]\"),\n-\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n+\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n \tNULL\n };\n \n@@ -19,7 +19,7 @@ static const char * const builtin_commit_graph_verify_usage[] = {\n };\n \n static const char * const builtin_commit_graph_write_usage[] = {\n-\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n+\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n \tNULL\n };\n \n@@ -32,6 +32,7 @@ static struct opts_commit_graph {\n \tint split;\n \tint shallow;\n \tint progress;\n+\tint enable_changed_paths;\n } opts;\n \n static struct object_directory *find_odb(struct repository *r,\n@@ -135,6 +136,8 @@ static int graph_write(int argc, const char **argv)\n \t\t\tN_(\"start walk at commits listed by stdin\")),\n \t\tOPT_BOOL(0, \"append\", &opts.append,\n \t\t\tN_(\"include all commits already in the commit-graph file\")),\n+\t\tOPT_BOOL(0, \"changed-paths\", &opts.enable_changed_paths,\n+\t\t\tN_(\"enable computation for changed paths\")),\n \t\tOPT_BOOL(0, \"progress\", &opts.progress, N_(\"force progress reporting\")),\n \t\tOPT_BOOL(0, \"split\", &opts.split,\n \t\t\tN_(\"allow writing an incremental commit-graph file\")),\n@@ -168,6 +171,8 @@ static int graph_write(int argc, const char **argv)\n \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n \tif (opts.progress)\n \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n+\tif (opts.enable_changed_paths)\n+\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n \n \tread_replace_refs = 0;\n \todb = find_odb(the_repository, opts.obj_dir);\n-- \ngitgitgadget\n\n"},{"id":"394315","messageId":"55824cda89c1dca7756c8c2d831d6e115f4a9ddb.1585528298.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 09/16] diff: skip batch object download when possible","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:31Z","receivedAt":"2020-03-30T00:32:02Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nWhen computing changed-path Bloom filters or performing a name-only\ndiff, we do not need the blob contents before completing the diff\nvalues. Thus, we do not need to download a pack containing the blobs\nwe do not have on-disk before completing our diff calculation.\n\nThis prevents downloading every blob in a partial clone when computing\nchanged path Bloom filters. It also prevents over-aggressive downloads\nduring \"git log --raw\" commands.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c | 1 +\n diff.c  | 8 +++++++-\n diff.h  | 1 +\n 3 files changed, 9 insertions(+), 1 deletion(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex a16eee92331..dbcf594baec 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -142,6 +142,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \trepo_diff_setup(r, &diffopt);\n \tdiffopt.flags.recursive = 1;\n+\tdiffopt.detect_rename = 0;\n \tdiffopt.max_changes = max_changes;\n \tdiff_setup_done(&diffopt);\n \ndiff --git a/diff.c b/diff.c\nindex 1010d806f50..63376adb011 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -4633,6 +4633,10 @@ void diff_setup_done(struct diff_options *options)\n \tif (!options->use_color || external_diff())\n \t\toptions->color_moved = 0;\n \n+\tif (!(options->output_format & ~(DIFF_FORMAT_NAME | DIFF_FORMAT_RAW)) &&\n+\t    !options->detect_rename)\n+\t\t\toptions->skip_batch_download_objects = 1;\n+\n \tFREE_AND_NULL(options->parseopts);\n }\n \n@@ -6507,7 +6511,9 @@ static void add_if_missing(struct repository *r,\n \n void diffcore_std(struct diff_options *options)\n {\n-\tif (options->repo == the_repository && has_promisor_remote()) {\n+\tif (!options->skip_batch_download_objects &&\n+\t    options->repo == the_repository &&\n+\t\thas_promisor_remote()) {\n \t\t/*\n \t\t * Prefetch the diff pairs that are about to be flushed.\n \t\t */\ndiff --git a/diff.h b/diff.h\nindex 9443dc1b003..e9f104309c4 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -281,6 +281,7 @@ struct diff_options {\n \tint show_rename_progress;\n \tint dirstat_permille;\n \tint setup;\n+\tint skip_batch_download_objects;\n \n \t/* Number of hexdigits to abbreviate raw format output to. */\n \tint abbrev;\n-- \ngitgitgadget\n\n"},{"id":"394318","messageId":"5ed16f35fed43018ce441adfff55b85967d3918c.1585528298.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 08/16] commit-graph: examine commits by generation number","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:30Z","receivedAt":"2020-03-30T00:32:02Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nWhen running 'git commit-graph write --changed-paths', we sort the\ncommits by pack-order to save time when computing the changed-paths\nbloom filters. This does not help when finding the commits via the\n'--reachable' flag.\n\nIf not using pack-order, then sort by generation number before\nexamining the diff. Commits with similar generation are more likely\nto have many trees in common, making the diff faster.\n\nOn the Linux kernel repository, this change reduced the computation\ntime for 'git commit-graph write --reachable --changed-paths' from\n3m00s to 1m37s.\n\nHelped-by: Jeff King <peff@peff.net>\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 33 ++++++++++++++++++++++++++++++---\n 1 file changed, 30 insertions(+), 3 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 31b06f878ce..732c81fa1b2 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -70,6 +70,25 @@ static int commit_pos_cmp(const void *va, const void *vb)\n \t       commit_pos_at(&commit_pos, b);\n }\n \n+static int commit_gen_cmp(const void *va, const void *vb)\n+{\n+\tconst struct commit *a = *(const struct commit **)va;\n+\tconst struct commit *b = *(const struct commit **)vb;\n+\n+\t/* lower generation commits first */\n+\tif (a->generation < b->generation)\n+\t\treturn -1;\n+\telse if (a->generation > b->generation)\n+\t\treturn 1;\n+\n+\t/* use date as a heuristic when generations are equal */\n+\tif (a->date < b->date)\n+\t\treturn -1;\n+\telse if (a->date > b->date)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n char *get_commit_graph_filename(struct object_directory *obj_dir)\n {\n \treturn xstrfmt(\"%s/info/commit-graph\", obj_dir->path);\n@@ -815,7 +834,8 @@ struct write_commit_graph_context {\n \t\t report_progress:1,\n \t\t split:1,\n \t\t check_oids:1,\n-\t\t changed_paths:1;\n+\t\t changed_paths:1,\n+\t\t order_by_pack:1;\n \n \tconst struct split_commit_graph_opts *split_opts;\n \tsize_t total_bloom_filter_data_size;\n@@ -1178,7 +1198,11 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \n \tALLOC_ARRAY(sorted_commits, ctx->commits.nr);\n \tCOPY_ARRAY(sorted_commits, ctx->commits.list, ctx->commits.nr);\n-\tQSORT(sorted_commits, ctx->commits.nr, commit_pos_cmp);\n+\n+\tif (ctx->order_by_pack)\n+\t\tQSORT(sorted_commits, ctx->commits.nr, commit_pos_cmp);\n+\telse\n+\t\tQSORT(sorted_commits, ctx->commits.nr, commit_gen_cmp);\n \n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tstruct commit *c = sorted_commits[i];\n@@ -1884,6 +1908,7 @@ int write_commit_graph(struct object_directory *odb,\n \t}\n \n \tif (pack_indexes) {\n+\t\tctx->order_by_pack = 1;\n \t\tif ((res = fill_oids_from_packs(ctx, pack_indexes)))\n \t\t\tgoto cleanup;\n \t}\n@@ -1893,8 +1918,10 @@ int write_commit_graph(struct object_directory *odb,\n \t\t\tgoto cleanup;\n \t}\n \n-\tif (!pack_indexes && !commit_hex)\n+\tif (!pack_indexes && !commit_hex) {\n+\t\tctx->order_by_pack = 1;\n \t\tfill_oids_from_all_packs(ctx);\n+\t}\n \n \tclose_reachable(ctx);\n \n-- \ngitgitgadget\n\n"},{"id":"394316","messageId":"68395d4051b7b2863c1dc3d758e7066f2b16a81d.1585528299.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v3 11/16] commit-graph: reuse existing Bloom filters during write","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-03-30T00:31:33Z","receivedAt":"2020-03-30T00:32:03Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd logic to\na) parse Bloom filter information from the commit graph file and,\nb) re-use existing Bloom filters.\n\nSee Documentation/technical/commit-graph-format for the format in which\nthe Bloom filter information is written to the commit graph file.\n\nTo read Bloom filter for a given commit with lexicographic position\n'i' we need to:\n1. Read BIDX[i] which essentially gives us the starting index in BDAT for\n   filter of commit i+1. It is essentially the index past the end\n   of the filter of commit i. It is called end_index in the code.\n\n2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT\n   for filter of commit i. It is called the start_index in the code.\n   For the first commit, where i = 0, Bloom filter data starts at the\n   beginning, just past the header in the BDAT chunk. Hence, start_index\n   will be 0.\n\n3. The length of the filter will be end_index - start_index, because\n   BIDX[i] gives the cumulative 8-byte words including the ith\n   commit's filter.\n\nWe toggle whether Bloom filters should be recomputed based on the\ncompute_if_not_present flag.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c               | 49 ++++++++++++++++++++++++++++++++++++++++++-\n bloom.h               |  4 +++-\n commit-graph.c        |  6 +++---\n t/helper/test-bloom.c |  2 +-\n 4 files changed, 55 insertions(+), 6 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex dbcf594baec..151d598ce7b 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -4,6 +4,8 @@\n #include \"diffcore.h\"\n #include \"revision.h\"\n #include \"hashmap.h\"\n+#include \"commit-graph.h\"\n+#include \"commit.h\"\n \n define_commit_slab(bloom_filter_slab, struct bloom_filter);\n \n@@ -26,6 +28,36 @@ static inline unsigned char get_bitmask(uint32_t pos)\n \treturn ((unsigned char)1) << (pos & (BITS_PER_WORD - 1));\n }\n \n+static int load_bloom_filter_from_graph(struct commit_graph *g,\n+\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t   struct commit *c)\n+{\n+\tuint32_t lex_pos, start_index, end_index;\n+\n+\twhile (c->graph_pos < g->num_commits_in_base)\n+\t\tg = g->base_graph;\n+\n+\t/* The commit graph commit 'c' lives in doesn't carry bloom filters. */\n+\tif (!g->chunk_bloom_indexes)\n+\t\treturn 0;\n+\n+\tlex_pos = c->graph_pos - g->num_commits_in_base;\n+\n+\tend_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n+\n+\tif (lex_pos > 0)\n+\t\tstart_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n+\telse\n+\t\tstart_index = 0;\n+\n+\tfilter->len = end_index - start_index;\n+\tfilter->data = (unsigned char *)(g->chunk_bloom_data +\n+\t\t\t\t\tsizeof(unsigned char) * start_index +\n+\t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n+\n+\treturn 1;\n+}\n+\n /*\n  * Calculate the murmur3 32-bit hash value for the given data\n  * using the given seed.\n@@ -127,7 +159,8 @@ void init_bloom_filters(void)\n }\n \n struct bloom_filter *get_bloom_filter(struct repository *r,\n-\t\t\t\t      struct commit *c)\n+\t\t\t\t      struct commit *c,\n+\t\t\t\t\t  int compute_if_not_present)\n {\n \tstruct bloom_filter *filter;\n \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n@@ -140,6 +173,20 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n \n+\tif (!filter->data) {\n+\t\tload_commit_graph_info(r, c);\n+\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH &&\n+\t\t\tr->objects->commit_graph->chunk_bloom_indexes) {\n+\t\t\tif (load_bloom_filter_from_graph(r->objects->commit_graph, filter, c))\n+\t\t\t\treturn filter;\n+\t\t\telse\n+\t\t\t\treturn NULL;\n+\t\t}\n+\t}\n+\n+\tif (filter->data || !compute_if_not_present)\n+\t\treturn filter;\n+\n \trepo_diff_setup(r, &diffopt);\n \tdiffopt.flags.recursive = 1;\n \tdiffopt.detect_rename = 0;\ndiff --git a/bloom.h b/bloom.h\nindex 85ab8e9423d..760d7122374 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -32,6 +32,7 @@ struct bloom_filter_settings {\n \n #define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n #define BITS_PER_WORD 8\n+#define BLOOMDATA_CHUNK_HEADER_SIZE 3 * sizeof(uint32_t)\n \n /*\n  * A bloom_filter struct represents a data segment to\n@@ -79,6 +80,7 @@ void add_key_to_filter(const struct bloom_key *key,\n void init_bloom_filters(void);\n \n struct bloom_filter *get_bloom_filter(struct repository *r,\n-\t\t\t\t      struct commit *c);\n+\t\t\t\t      struct commit *c,\n+\t\t\t\t      int compute_if_not_present);\n \n #endif\n\\ No newline at end of file\ndiff --git a/commit-graph.c b/commit-graph.c\nindex a8b6b5cca5d..77668629e27 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1086,7 +1086,7 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n \t\t\tctx->commits.nr);\n \n \twhile (list < last) {\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n \t\tcur_pos += filter->len;\n \t\tdisplay_progress(progress, ++i);\n \t\thashwrite_be32(f, cur_pos);\n@@ -1115,7 +1115,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n \thashwrite_be32(f, settings->bits_per_entry);\n \n \twhile (list < last) {\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n \t\tdisplay_progress(progress, ++i);\n \t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n \t\tlist++;\n@@ -1296,7 +1296,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tstruct commit *c = sorted_commits[i];\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n \t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n \t\tdisplay_progress(progress, i + 1);\n \t}\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex f18d1b722e1..ce412664ba9 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -39,7 +39,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n \tstruct bloom_filter *filter;\n \tsetup_git_directory();\n \tc = lookup_commit(the_repository, commit_oid);\n-\tfilter = get_bloom_filter(the_repository, c);\n+\tfilter = get_bloom_filter(the_repository, c, 1);\n \tprint_bloom_filter(filter);\n }\n \n-- \ngitgitgadget\n\n"},{"id":"394850","messageId":"c3ffd9820d50196793833d0faf35f9d5dd70ff6d.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 01/15] commit-graph: define and use MAX_NUM_CHUNKS","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:41Z","receivedAt":"2020-04-06T17:00:03Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nThis is a minor cleanup to make it easier to change\nthe number of chunks being written to the commit\ngraph.\n\nReviewed-by: Jakub Narębski <jnareb@gmail.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 5 +++--\n 1 file changed, 3 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex f013a84e294..e4f1a5b2f1a 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -23,6 +23,7 @@\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n+#define MAX_NUM_CHUNKS 5\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -1350,8 +1351,8 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tint fd;\n \tstruct hashfile *f;\n \tstruct lock_file lk = LOCK_INIT;\n-\tuint32_t chunk_ids[6];\n-\tuint64_t chunk_offsets[6];\n+\tuint32_t chunk_ids[MAX_NUM_CHUNKS + 1];\n+\tuint64_t chunk_offsets[MAX_NUM_CHUNKS + 1];\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \tstruct strbuf progress_title = STRBUF_INIT;\n \tint num_chunks = 3;\n-- \ngitgitgadget\n\n"},{"id":"394851","messageId":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v3.git.1585528298.gitgitgadget@gmail.com","subject":"[PATCH v4 00/15] Changed Paths Bloom Filters","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:40Z","receivedAt":"2020-04-06T17:00:03Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"Hey! \n\nThe commit graph feature brought in a lot of performance improvements across\nmultiple commands. However, file based history continues to be a performance\npain point, especially in large repositories. \n\nAdopting changed path Bloom filters has been discussed on the list before,\nand a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\nDerrick Stolee [1]. This series is based on Dr. Stolee's proof of concept in\n[2]\n\nWith the changes in this series, git users will be able to choose to write\nBloom filters to the commit-graph using the following command:\n\n'git commit-graph write --changed-paths'\n\nSubsequent 'git log -- path' commands will use these computed Bloom filters\nto decided which commits are worth exploring further to produce the history\nof the provided path. \n\nCost of computing and writing Bloom filters\n===========================================\n\nComputing and writing Bloom filters to the commit graph for the first time\nimplies computing the diffs and the resulting Bloom filters for all the\ncommits in the repository. This adds a non trivial amount of time to run\ntime. Every subsequent run is incremental i.e. we reuse the previously\ncomputed Bloom filters. So this is a one time cost. \n\nTime taken by 'git commit-graph write' with and w/o --changed-paths, speed\nup in 'git log -- path' with computed Bloom filters (see a):- \n\n-------------------------------------------------------------------------\n| Repo        | w/o --changed-paths | with --changed-paths | Speed up   |\n-------------------------------------------------------------------------\n| git [3]     | 0.9 seconds         | 7 seconds            | 2x to 6x   |\n| linux [4]   | 16 seconds          | 1 minute 8 seconds   | 2x to 6x   | \n| android [5] | 9 seconds           | 48 seconds           | 2x to 6x   |\n| AzDo(see b) | 1 minute            | 5 minutes 2 seconds  | 10x to 30x |\n-------------------------------------------------------------------------\n\na) We tested the performance of git log -- path with randomly chosen paths\nof varying depths in each repo. The speed up depends on how deep the files\nare in the hierarchy and how often a file has been touched in the commit\nhistory.\n\nb) This internal repository has about 420k commits, 183k files distributed\nacross 34k folders, the size on disk is about 17 GiB. The most massive gains\non this repository were for files 6-12 levels deep in the tree. \n\nc) These numbers were collected on a Windows machine, except for the linux\nrepo which was tested on a Linux machine. \n\nFuture Work (not included in the scope of this series)\n======================================================\n\n 1. Supporting multiple path based revision walk\n 2. Adopting it in git blame logic. \n 3. Interactions with line log git log -L\n\nCheers! Garima Singh\n\n[1] https://lore.kernel.org/git/20181009193445.21908-1-szeder.dev@gmail.com/\n\n[2] \nhttps://lore.kernel.org/git/61559c5b-546e-d61b-d2e1-68de692f5972@gmail.com/\n\n[3] https://github.com/git/git\n\n[4] https://github.com/torvalds/linux\n\n[5] https://android.googlesource.com/platform/frameworks/base/\n\njeffhost@microsoft.com, me@ttaylorr.com, peff@peff.net, \ngarimasigit@gmail.com,jnareb@gmail.com, christian.couder@gmail.com, \nemilyshaffer@gmail.com,gitster@pobox.com\n\nDerrick Stolee (1):\n  diff: halt tree-diff early after max_changes\n\nGarima Singh (13):\n  commit-graph: define and use MAX_NUM_CHUNKS\n  bloom.c: add the murmur3 hash implementation\n  bloom.c: introduce core Bloom filter constructs\n  bloom.c: core Bloom filter implementation for changed paths.\n  commit-graph: compute Bloom filters for changed paths\n  commit-graph: examine commits by generation number\n  commit-graph: write Bloom filters to commit graph file\n  commit-graph: reuse existing Bloom filters during write\n  commit-graph: add --changed-paths option to write subcommand\n  revision.c: use Bloom filters to speed up path based revision walks\n  revision.c: add trace2 stats around Bloom filter usage\n  t4216: add end to end tests for git log with Bloom filters\n  commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag\n\nJeff King (1):\n  commit-graph: examine changed-path objects in pack order\n\n Documentation/git-commit-graph.txt            |   5 +\n .../technical/commit-graph-format.txt         |  30 ++\n Makefile                                      |   2 +\n bloom.c                                       | 275 ++++++++++++++++++\n bloom.h                                       |  90 ++++++\n builtin/commit-graph.c                        |  10 +-\n ci/run-build-and-tests.sh                     |   1 +\n commit-graph.c                                | 213 +++++++++++++-\n commit-graph.h                                |   9 +-\n diff.h                                        |   5 +\n revision.c                                    | 126 +++++++-\n revision.h                                    |  11 +\n t/README                                      |   5 +\n t/helper/test-bloom.c                         |  81 ++++++\n t/helper/test-read-graph.c                    |   4 +\n t/helper/test-tool.c                          |   1 +\n t/helper/test-tool.h                          |   1 +\n t/t0095-bloom.sh                              | 117 ++++++++\n t/t4216-log-bloom.sh                          | 155 ++++++++++\n t/t5318-commit-graph.sh                       |   2 +\n t/t5324-split-commit-graph.sh                 |   1 +\n tree-diff.c                                   |   6 +\n 22 files changed, 1139 insertions(+), 11 deletions(-)\n create mode 100644 bloom.c\n create mode 100644 bloom.h\n create mode 100644 t/helper/test-bloom.c\n create mode 100755 t/t0095-bloom.sh\n create mode 100755 t/t4216-log-bloom.sh\n\n\nbase-commit: 3bab5d56259722843359702bc27111475437ad2a\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-497%2Fgarimasi514%2FcoreGit-bloomFilters-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-497/garimasi514/coreGit-bloomFilters-v4\nPull-Request: https://github.com/gitgitgadget/git/pull/497\n\nRange-diff vs v3:\n\n  1:  c3ffd9820d5 =  1:  c3ffd9820d5 commit-graph: define and use MAX_NUM_CHUNKS\n  2:  a5aa3415c05 =  2:  a5aa3415c05 bloom.c: add the murmur3 hash implementation\n  3:  a7702c1afde =  3:  a7702c1afde bloom.c: introduce core Bloom filter constructs\n  4:  8304c297520 =  4:  8304c297520 bloom.c: core Bloom filter implementation for changed paths.\n  5:  2d4c0b2da38 =  5:  2d4c0b2da38 diff: halt tree-diff early after max_changes\n  6:  c38b9b386ef =  6:  c38b9b386ef commit-graph: compute Bloom filters for changed paths\n  7:  d24c85c54ef =  7:  d24c85c54ef commit-graph: examine changed-path objects in pack order\n  8:  5ed16f35fed =  8:  5ed16f35fed commit-graph: examine commits by generation number\n  9:  55824cda89c <  -:  ----------- diff: skip batch object download when possible\n 10:  1e4663523de =  9:  ff6b96aad1e commit-graph: write Bloom filters to commit graph file\n 11:  68395d4051b ! 10:  cc8022bdf82 commit-graph: reuse existing Bloom filters during write\n     @@ bloom.c: struct bloom_filter *get_bloom_filter(struct repository *r,\n      +\n       \trepo_diff_setup(r, &diffopt);\n       \tdiffopt.flags.recursive = 1;\n     - \tdiffopt.detect_rename = 0;\n     + \tdiffopt.max_changes = max_changes;\n      \n       ## bloom.h ##\n      @@ bloom.h: struct bloom_filter_settings {\n 12:  7e450e45236 = 11:  c8b86c383ab commit-graph: add --changed-paths option to write subcommand\n 13:  b18af58aa3e = 12:  617f549ef25 revision.c: use Bloom filters to speed up path based revision walks\n 14:  b5eb280178f = 13:  6beaede7159 revision.c: add trace2 stats around Bloom filter usage\n 15:  3019ef72881 = 14:  b899df5c98e t4216: add end to end tests for git log with Bloom filters\n 16:  213abb5d895 = 15:  5656e8590e9 commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag\n\n-- \ngitgitgadget\n"},{"id":"394854","messageId":"a7702c1afde1ce9d4c628a831927b4e75bff8515.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 03/15] bloom.c: introduce core Bloom filter constructs","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:43Z","receivedAt":"2020-04-06T17:00:03Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nIntroduce the constructs for Bloom filters, Bloom filter keys\nand Bloom filter settings.\nFor details on what Bloom filters are and how they work, refer\nto Dr. Derrick Stolee's blog post [1]. It provides a concise\nexplanation of the adoption of Bloom filters as described in\n[2] and [3].\n\nImplementation specifics:\n1. We currently use 7 and 10 for the number of hashes and the\n   size of each entry respectively. They served as great starting\n   values, the mathematical details behind this choice are\n   described in [1] and [4]. The implementation, while not\n   completely open to it at the moment, is flexible enough to allow\n   for tweaking these settings in the future.\n\n   Note: The performance gains we have observed with these values\n   are significant enough that we did not need to tweak these\n   settings. The performance numbers are included in the cover letter\n   of this series and in the commit message of the subsequent commit\n   where we use Bloom filters to speed up `git log -- path`.\n\n2. As described in [1] and [3], we do not need 7 independent hashing\n   functions. We use the Murmur3 hashing scheme, seed it twice and\n   then combine those to procure an arbitrary number of hash values.\n\n3. The filters will be sized according to the number of changes in\n   each commit, in multiples of 8 bit words.\n\n[1] Derrick Stolee\n      \"Supercharging the Git Commit Graph IV: Bloom Filters\"\n      https://devblogs.microsoft.com/devops/super-charging-the-git-commit-graph-iv-Bloom-filters/\n\n[2] Flavio Bonomi, Michael Mitzenmacher, Rina Panigrahy, Sushil Singh, George Varghese\n    \"An Improved Construction for Counting Bloom Filters\"\n    http://theory.stanford.edu/~rinap/papers/esa2006b.pdf\n    https://doi.org/10.1007/11841036_61\n\n[3] Peter C. Dillinger and Panagiotis Manolios\n    \"Bloom Filters in Probabilistic Verification\"\n    http://www.ccs.neu.edu/home/pete/pub/Bloom-filters-verification.pdf\n    https://doi.org/10.1007/978-3-540-30494-4_26\n\n[4] Thomas Mueller Graf, Daniel Lemire\n    \"Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters\"\n    https://arxiv.org/abs/1912.08258\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nReviewed-by: Jakub Narębski <jnareb@gmail.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c               | 38 +++++++++++++++++++++++++-\n bloom.h               | 63 +++++++++++++++++++++++++++++++++++++++++++\n t/helper/test-bloom.c | 48 +++++++++++++++++++++++++++++++++\n t/t0095-bloom.sh      | 40 +++++++++++++++++++++++++++\n 4 files changed, 188 insertions(+), 1 deletion(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 40e87632aeb..888b67f1ea6 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -8,6 +8,11 @@ static uint32_t rotate_left(uint32_t value, int32_t count)\n \treturn ((value << count) | (value >> ((-count) & mask)));\n }\n \n+static inline unsigned char get_bitmask(uint32_t pos)\n+{\n+\treturn ((unsigned char)1) << (pos & (BITS_PER_WORD - 1));\n+}\n+\n /*\n  * Calculate the murmur3 32-bit hash value for the given data\n  * using the given seed.\n@@ -70,4 +75,35 @@ uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len)\n \tseed ^= (seed >> 16);\n \n \treturn seed;\n-}\n\\ No newline at end of file\n+}\n+\n+void fill_bloom_key(const char *data,\n+\t\t\t\t\tsize_t len,\n+\t\t\t\t\tstruct bloom_key *key,\n+\t\t\t\t\tconst struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tconst uint32_t seed0 = 0x293ae76f;\n+\tconst uint32_t seed1 = 0x7e646e2c;\n+\tconst uint32_t hash0 = murmur3_seeded(seed0, data, len);\n+\tconst uint32_t hash1 = murmur3_seeded(seed1, data, len);\n+\n+\tkey->hashes = (uint32_t *)xcalloc(settings->num_hashes, sizeof(uint32_t));\n+\tfor (i = 0; i < settings->num_hashes; i++)\n+\t\tkey->hashes[i] = hash0 + i * hash1;\n+}\n+\n+void add_key_to_filter(const struct bloom_key *key,\n+\t\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t\t   const struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tuint64_t mod = filter->len * BITS_PER_WORD;\n+\n+\tfor (i = 0; i < settings->num_hashes; i++) {\n+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n+\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n+\n+\t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n+\t}\n+}\ndiff --git a/bloom.h b/bloom.h\nindex d0fcc5f0aa6..b9ce422ca2d 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -1,6 +1,60 @@\n #ifndef BLOOM_H\n #define BLOOM_H\n \n+struct bloom_filter_settings {\n+\t/*\n+\t * The version of the hashing technique being used.\n+\t * We currently only support version = 1 which is\n+\t * the seeded murmur3 hashing technique implemented\n+\t * in bloom.c.\n+\t */\n+\tuint32_t hash_version;\n+\n+\t/*\n+\t * The number of times a path is hashed, i.e. the\n+\t * number of bit positions tht cumulatively\n+\t * determine whether a path is present in the\n+\t * Bloom filter.\n+\t */\n+\tuint32_t num_hashes;\n+\n+\t/*\n+\t * The minimum number of bits per entry in the Bloom\n+\t * filter. If the filter contains 'n' entries, then\n+\t * filter size is the minimum number of 8-bit words\n+\t * that contain n*b bits.\n+\t */\n+\tuint32_t bits_per_entry;\n+};\n+\n+#define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n+#define BITS_PER_WORD 8\n+\n+/*\n+ * A bloom_filter struct represents a data segment to\n+ * use when testing hash values. The 'len' member\n+ * dictates how many entries are stored in\n+ * 'data'.\n+ */\n+struct bloom_filter {\n+\tunsigned char *data;\n+\tsize_t len;\n+};\n+\n+/*\n+ * A bloom_key represents the k hash values for a\n+ * given string. These can be precomputed and\n+ * stored in a bloom_key for re-use when testing\n+ * against a bloom_filter. The number of hashes is\n+ * given by the Bloom filter settings and is the same\n+ * for all Bloom filters and keys interacting with\n+ * the loaded version of the commit graph file and\n+ * the Bloom data chunks.\n+ */\n+struct bloom_key {\n+\tuint32_t *hashes;\n+};\n+\n /*\n  * Calculate the murmur3 32-bit hash value for the given data\n  * using the given seed.\n@@ -10,4 +64,13 @@\n  */\n uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len);\n \n+void fill_bloom_key(const char *data,\n+\t\t    size_t len,\n+\t\t    struct bloom_key *key,\n+\t\t    const struct bloom_filter_settings *settings);\n+\n+void add_key_to_filter(const struct bloom_key *key,\n+\t\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t\t   const struct bloom_filter_settings *settings);\n+\n #endif\n\\ No newline at end of file\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 60ee2043689..20460cde775 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -2,6 +2,36 @@\n #include \"bloom.h\"\n #include \"test-tool.h\"\n \n+struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\n+static void add_string_to_filter(const char *data, struct bloom_filter *filter) {\n+\t\tstruct bloom_key key;\n+\t\tint i;\n+\n+\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n+\t\tprintf(\"Hashes:\");\n+\t\tfor (i = 0; i < settings.num_hashes; i++){\n+\t\t\tprintf(\"0x%08x|\", key.hashes[i]);\n+\t\t}\n+\t\tprintf(\"\\n\");\n+\t\tadd_key_to_filter(&key, filter, &settings);\n+}\n+\n+static void print_bloom_filter(struct bloom_filter *filter) {\n+\tint i;\n+\n+\tif (!filter) {\n+\t\tprintf(\"No filter.\\n\");\n+\t\treturn;\n+\t}\n+\tprintf(\"Filter_Length:%d\\n\", (int)filter->len);\n+\tprintf(\"Filter_Data:\");\n+\tfor (i = 0; i < filter->len; i++){\n+\t\tprintf(\"%02x|\", filter->data[i]);\n+\t}\n+\tprintf(\"\\n\");\n+}\n+\n int cmd__bloom(int argc, const char **argv)\n {\n \tif (!strcmp(argv[1], \"get_murmur3\")) {\n@@ -9,5 +39,23 @@ int cmd__bloom(int argc, const char **argv)\n \t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n \t}\n \n+    if (!strcmp(argv[1], \"generate_filter\")) {\n+\t\tstruct bloom_filter filter;\n+\t\tint i = 2;\n+\t\tfilter.len =  (settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n+\t\tfilter.data = xcalloc(filter.len, sizeof(unsigned char));\n+\n+\t\tif (!argv[2]){\n+\t\t\tdie(\"at least one input string expected\");\n+\t\t}\n+\n+\t\twhile (argv[i]) {\n+\t\t\tadd_string_to_filter(argv[i], &filter);\n+\t\t\ti++;\n+\t\t}\n+\n+\t\tprint_bloom_filter(&filter);\n+\t}\n+\n \treturn 0;\n }\n\\ No newline at end of file\ndiff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\nindex 2dad8c4a94e..36a086c7c60 100755\n--- a/t/t0095-bloom.sh\n+++ b/t/t0095-bloom.sh\n@@ -27,4 +27,44 @@ test_expect_success 'compute unseeded murmur3 hash for test string 2' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'compute bloom key for empty string' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:0x5615800c|0x5b966560|0x61174ab4|0x66983008|0x6c19155c|0x7199fab0|0x771ae004|\n+\tFilter_Length:2\n+\tFilter_Data:11|11|\n+\tEOF\n+\ttest-tool bloom generate_filter \"\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for whitespace' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:0xf178874c|0x5f3d6eb6|0xcd025620|0x3ac73d8a|0xa88c24f4|0x16510c5e|0x8415f3c8|\n+\tFilter_Length:2\n+\tFilter_Data:51|55|\n+\tEOF\n+\ttest-tool bloom generate_filter \" \" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for test string 1' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:0xb270de9b|0x1bb6f26e|0x84fd0641|0xee431a14|0x57892de7|0xc0cf41ba|0x2a15558d|\n+\tFilter_Length:2\n+\tFilter_Data:92|6c|\n+\tEOF\n+\ttest-tool bloom generate_filter \"Hello world!\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute bloom key for test string 2' '\n+\tcat >expect <<-\\EOF &&\n+\tHashes:0x20ab385b|0xf5237fe2|0xc99bc769|0x9e140ef0|0x728c5677|0x47049dfe|0x1b7ce585|\n+\tFilter_Length:2\n+\tFilter_Data:a5|4a|\n+\tEOF\n+\ttest-tool bloom generate_filter \"file.txt\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_done\n\\ No newline at end of file\n-- \ngitgitgadget\n\n"},{"id":"394852","messageId":"a5aa3415c05ee9bc67a9471445a20c71a9834673.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 02/15] bloom.c: add the murmur3 hash implementation","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:42Z","receivedAt":"2020-04-06T17:00:04Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nIn preparation for computing changed paths Bloom filters,\nimplement the Murmur3 hash algorithm as described in [1].\nIt hashes the given data using the given seed and produces\na uniformly distributed hash value.\n\n[1] https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nHelped-by: Szeder Gábor <szeder.dev@gmail.com>\nReviewed-by: Jakub Narębski <jnareb@gmail.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n Makefile              |  2 ++\n bloom.c               | 73 +++++++++++++++++++++++++++++++++++++++++++\n bloom.h               | 13 ++++++++\n t/helper/test-bloom.c | 13 ++++++++\n t/helper/test-tool.c  |  1 +\n t/helper/test-tool.h  |  1 +\n t/t0095-bloom.sh      | 30 ++++++++++++++++++\n 7 files changed, 133 insertions(+)\n create mode 100644 bloom.c\n create mode 100644 bloom.h\n create mode 100644 t/helper/test-bloom.c\n create mode 100755 t/t0095-bloom.sh\n\ndiff --git a/Makefile b/Makefile\nindex ef1ff2228f0..491f75e68c5 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -695,6 +695,7 @@ X =\n PROGRAMS += $(patsubst %.o,git-%$X,$(PROGRAM_OBJS))\n \n TEST_BUILTINS_OBJS += test-advise.o\n+TEST_BUILTINS_OBJS += test-bloom.o\n TEST_BUILTINS_OBJS += test-chmtime.o\n TEST_BUILTINS_OBJS += test-config.o\n TEST_BUILTINS_OBJS += test-ctype.o\n@@ -840,6 +841,7 @@ LIB_OBJS += base85.o\n LIB_OBJS += bisect.o\n LIB_OBJS += blame.o\n LIB_OBJS += blob.o\n+LIB_OBJS += bloom.o\n LIB_OBJS += branch.o\n LIB_OBJS += bulk-checkin.o\n LIB_OBJS += bundle.o\ndiff --git a/bloom.c b/bloom.c\nnew file mode 100644\nindex 00000000000..40e87632aeb\n--- /dev/null\n+++ b/bloom.c\n@@ -0,0 +1,73 @@\n+#include \"git-compat-util.h\"\n+#include \"bloom.h\"\n+\n+static uint32_t rotate_left(uint32_t value, int32_t count)\n+{\n+\tuint32_t mask = 8 * sizeof(uint32_t) - 1;\n+\tcount &= mask;\n+\treturn ((value << count) | (value >> ((-count) & mask)));\n+}\n+\n+/*\n+ * Calculate the murmur3 32-bit hash value for the given data\n+ * using the given seed.\n+ * Produces a uniformly distributed hash value.\n+ * Not considered to be cryptographically secure.\n+ * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n+ */\n+uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len)\n+{\n+\tconst uint32_t c1 = 0xcc9e2d51;\n+\tconst uint32_t c2 = 0x1b873593;\n+\tconst uint32_t r1 = 15;\n+\tconst uint32_t r2 = 13;\n+\tconst uint32_t m = 5;\n+\tconst uint32_t n = 0xe6546b64;\n+\tint i;\n+\tuint32_t k1 = 0;\n+\tconst char *tail;\n+\n+\tint len4 = len / sizeof(uint32_t);\n+\n+\tuint32_t k;\n+\tfor (i = 0; i < len4; i++) {\n+\t\tuint32_t byte1 = (uint32_t)data[4*i];\n+\t\tuint32_t byte2 = ((uint32_t)data[4*i + 1]) << 8;\n+\t\tuint32_t byte3 = ((uint32_t)data[4*i + 2]) << 16;\n+\t\tuint32_t byte4 = ((uint32_t)data[4*i + 3]) << 24;\n+\t\tk = byte1 | byte2 | byte3 | byte4;\n+\t\tk *= c1;\n+\t\tk = rotate_left(k, r1);\n+\t\tk *= c2;\n+\n+\t\tseed ^= k;\n+\t\tseed = rotate_left(seed, r2) * m + n;\n+\t}\n+\n+\ttail = (data + len4 * sizeof(uint32_t));\n+\n+\tswitch (len & (sizeof(uint32_t) - 1)) {\n+\tcase 3:\n+\t\tk1 ^= ((uint32_t)tail[2]) << 16;\n+\t\t/*-fallthrough*/\n+\tcase 2:\n+\t\tk1 ^= ((uint32_t)tail[1]) << 8;\n+\t\t/*-fallthrough*/\n+\tcase 1:\n+\t\tk1 ^= ((uint32_t)tail[0]) << 0;\n+\t\tk1 *= c1;\n+\t\tk1 = rotate_left(k1, r1);\n+\t\tk1 *= c2;\n+\t\tseed ^= k1;\n+\t\tbreak;\n+\t}\n+\n+\tseed ^= (uint32_t)len;\n+\tseed ^= (seed >> 16);\n+\tseed *= 0x85ebca6b;\n+\tseed ^= (seed >> 13);\n+\tseed *= 0xc2b2ae35;\n+\tseed ^= (seed >> 16);\n+\n+\treturn seed;\n+}\n\\ No newline at end of file\ndiff --git a/bloom.h b/bloom.h\nnew file mode 100644\nindex 00000000000..d0fcc5f0aa6\n--- /dev/null\n+++ b/bloom.h\n@@ -0,0 +1,13 @@\n+#ifndef BLOOM_H\n+#define BLOOM_H\n+\n+/*\n+ * Calculate the murmur3 32-bit hash value for the given data\n+ * using the given seed.\n+ * Produces a uniformly distributed hash value.\n+ * Not considered to be cryptographically secure.\n+ * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n+ */\n+uint32_t murmur3_seeded(uint32_t seed, const char *data, size_t len);\n+\n+#endif\n\\ No newline at end of file\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nnew file mode 100644\nindex 00000000000..60ee2043689\n--- /dev/null\n+++ b/t/helper/test-bloom.c\n@@ -0,0 +1,13 @@\n+#include \"git-compat-util.h\"\n+#include \"bloom.h\"\n+#include \"test-tool.h\"\n+\n+int cmd__bloom(int argc, const char **argv)\n+{\n+\tif (!strcmp(argv[1], \"get_murmur3\")) {\n+\t\tuint32_t hashed = murmur3_seeded(0, argv[2], strlen(argv[2]));\n+\t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n+\t}\n+\n+\treturn 0;\n+}\n\\ No newline at end of file\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex 31eedcd241f..6e26bd65c97 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -15,6 +15,7 @@ struct test_cmd {\n \n static struct test_cmd cmds[] = {\n \t{ \"advise\", cmd__advise_if_enabled },\n+\t{ \"bloom\", cmd__bloom },\n \t{ \"chmtime\", cmd__chmtime },\n \t{ \"config\", cmd__config },\n \t{ \"ctype\", cmd__ctype },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex 4eb5e6609e1..dceeef1d5c2 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -5,6 +5,7 @@\n #include \"git-compat-util.h\"\n \n int cmd__advise_if_enabled(int argc, const char **argv);\n+int cmd__bloom(int argc, const char **argv);\n int cmd__chmtime(int argc, const char **argv);\n int cmd__config(int argc, const char **argv);\n int cmd__ctype(int argc, const char **argv);\ndiff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\nnew file mode 100755\nindex 00000000000..2dad8c4a94e\n--- /dev/null\n+++ b/t/t0095-bloom.sh\n@@ -0,0 +1,30 @@\n+#!/bin/sh\n+\n+test_description='Testing the various Bloom filter computations in bloom.c'\n+. ./test-lib.sh\n+\n+test_expect_success 'compute unseeded murmur3 hash for empty string' '\n+\tcat >expect <<-\\EOF &&\n+\tMurmur3 Hash with seed=0:0x00000000\n+\tEOF\n+\ttest-tool bloom get_murmur3 \"\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute unseeded murmur3 hash for test string 1' '\n+\tcat >expect <<-\\EOF &&\n+\tMurmur3 Hash with seed=0:0x627b0c2c\n+\tEOF\n+\ttest-tool bloom get_murmur3 \"Hello world!\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'compute unseeded murmur3 hash for test string 2' '\n+\tcat >expect <<-\\EOF &&\n+\tMurmur3 Hash with seed=0:0x2e4ff723\n+\tEOF\n+\ttest-tool bloom get_murmur3 \"The quick brown fox jumps over the lazy dog\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_done\n\\ No newline at end of file\n-- \ngitgitgadget\n\n"},{"id":"394853","messageId":"2d4c0b2da38632424c8bd31ccb2037e0676c3c74.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 05/15] diff: halt tree-diff early after max_changes","fromName":"Derrick Stolee via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:45Z","receivedAt":"2020-04-06T17:00:05Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"From: Derrick Stolee <dstolee@microsoft.com>\n\nWhen computing the changed-paths bloom filters for the commit-graph,\nwe limit the size of the filter by restricting the number of paths\nin the diff. Instead of computing a large diff and then ignoring the\nresult, it is better to halt the diff computation early.\n\nCreate a new \"max_changes\" option in struct diff_options. If non-zero,\nthen halt the diff computation after discovering strictly more changed\npaths. This includes paths corresponding to trees that change.\n\nUse this max_changes option in the bloom filter calculations. This\nreduces the time taken to compute the filters for the Linux kernel\nrepo from 2m50s to 2m35s. On a large internal repository with ~500\ncommits that perform tree-wide changes, the time reduced from\n6m15s to 3m48s.\n\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c     | 4 +++-\n diff.h      | 5 +++++\n tree-diff.c | 6 ++++++\n 3 files changed, 14 insertions(+), 1 deletion(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 881a9841ede..a16eee92331 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -133,6 +133,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \tint i;\n \tstruct diff_options diffopt;\n+\tint max_changes = 512;\n \n \tif (bloom_filters.slab_size == 0)\n \t\treturn NULL;\n@@ -141,6 +142,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \trepo_diff_setup(r, &diffopt);\n \tdiffopt.flags.recursive = 1;\n+\tdiffopt.max_changes = max_changes;\n \tdiff_setup_done(&diffopt);\n \n \tif (c->parents)\n@@ -149,7 +151,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n \tdiffcore_std(&diffopt);\n \n-\tif (diff_queued_diff.nr <= 512) {\n+\tif (diff_queued_diff.nr <= max_changes) {\n \t\tstruct hashmap pathmap;\n \t\tstruct pathmap_hash_entry *e;\n \t\tstruct hashmap_iter iter;\ndiff --git a/diff.h b/diff.h\nindex 6febe7e3656..9443dc1b003 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -285,6 +285,11 @@ struct diff_options {\n \t/* Number of hexdigits to abbreviate raw format output to. */\n \tint abbrev;\n \n+\t/* If non-zero, then stop computing after this many changes. */\n+\tint max_changes;\n+\t/* For internal use only. */\n+\tint num_changes;\n+\n \tint ita_invisible_in_index;\n /* white-space error highlighting */\n #define WSEH_NEW (1<<12)\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 33ded7f8b3e..f3d303c6e54 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -434,6 +434,9 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \n+\t\tif (opt->max_changes && opt->num_changes > opt->max_changes)\n+\t\t\tbreak;\n+\n \t\tif (opt->pathspec.nr) {\n \t\t\tskip_uninteresting(&t, base, opt);\n \t\t\tfor (i = 0; i < nparent; i++)\n@@ -518,6 +521,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \n \t\t\t/* t↓ */\n \t\t\tupdate_tree_entry(&t);\n+\t\t\topt->num_changes++;\n \t\t}\n \n \t\t/* t > p[imin] */\n@@ -535,6 +539,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\tskip_emit_tp:\n \t\t\t/* ∀ pi=p[imin]  pi↓ */\n \t\t\tupdate_tp_entries(tp, nparent);\n+\t\t\topt->num_changes++;\n \t\t}\n \t}\n \n@@ -552,6 +557,7 @@ struct combine_diff_path *diff_tree_paths(\n \tconst struct object_id **parents_oid, int nparent,\n \tstruct strbuf *base, struct diff_options *opt)\n {\n+\topt->num_changes = 0;\n \tp = ll_diff_tree_paths(p, oid, parents_oid, nparent, base, opt);\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"394855","messageId":"8304c2975207ee847c6709abd71efee918fc4142.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 04/15] bloom.c: core Bloom filter implementation for changed paths.","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:44Z","receivedAt":"2020-04-06T17:00:05Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd the core implementation for computing Bloom filters for\nthe paths changed between a commit and it's first parent.\n\nWe fill the Bloom filters as (const char *data, int len) pairs\nas `struct bloom_filters\" within a commit slab.\n\nFilters for commits with no changes and more than 512 changes,\nis represented with a filter of length zero. There is no gain\nin distinguishing between a computed filter of length zero for\na commit with no changes, and an uncomputed filter for new commits\nor for commits with more than 512 changes. The effect on\n`git log -- path` is the same in both cases. We will fall back to\nthe normal diffing algorithm when we can't benefit from the\nexistence of Bloom filters.\n\nHelped-by: Jeff King <peff@peff.net>\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nReviewed-by: Jakub Narębski <jnareb@gmail.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c               | 97 +++++++++++++++++++++++++++++++++++++++++++\n bloom.h               |  8 ++++\n t/helper/test-bloom.c | 20 +++++++++\n t/t0095-bloom.sh      | 47 +++++++++++++++++++++\n 4 files changed, 172 insertions(+)\n\ndiff --git a/bloom.c b/bloom.c\nindex 888b67f1ea6..881a9841ede 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -1,5 +1,18 @@\n #include \"git-compat-util.h\"\n #include \"bloom.h\"\n+#include \"diff.h\"\n+#include \"diffcore.h\"\n+#include \"revision.h\"\n+#include \"hashmap.h\"\n+\n+define_commit_slab(bloom_filter_slab, struct bloom_filter);\n+\n+struct bloom_filter_slab bloom_filters;\n+\n+struct pathmap_hash_entry {\n+    struct hashmap_entry entry;\n+    const char path[FLEX_ARRAY];\n+};\n \n static uint32_t rotate_left(uint32_t value, int32_t count)\n {\n@@ -107,3 +120,87 @@ void add_key_to_filter(const struct bloom_key *key,\n \t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n \t}\n }\n+\n+void init_bloom_filters(void)\n+{\n+\tinit_bloom_filter_slab(&bloom_filters);\n+}\n+\n+struct bloom_filter *get_bloom_filter(struct repository *r,\n+\t\t\t\t      struct commit *c)\n+{\n+\tstruct bloom_filter *filter;\n+\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n+\tint i;\n+\tstruct diff_options diffopt;\n+\n+\tif (bloom_filters.slab_size == 0)\n+\t\treturn NULL;\n+\n+\tfilter = bloom_filter_slab_at(&bloom_filters, c);\n+\n+\trepo_diff_setup(r, &diffopt);\n+\tdiffopt.flags.recursive = 1;\n+\tdiff_setup_done(&diffopt);\n+\n+\tif (c->parents)\n+\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &diffopt);\n+\telse\n+\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n+\tdiffcore_std(&diffopt);\n+\n+\tif (diff_queued_diff.nr <= 512) {\n+\t\tstruct hashmap pathmap;\n+\t\tstruct pathmap_hash_entry *e;\n+\t\tstruct hashmap_iter iter;\n+\t\thashmap_init(&pathmap, NULL, NULL, 0);\n+\n+\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n+\t\t\tconst char *path = diff_queued_diff.queue[i]->two->path;\n+\n+\t\t\t/*\n+\t\t\t* Add each leading directory of the changed file, i.e. for\n+\t\t\t* 'dir/subdir/file' add 'dir' and 'dir/subdir' as well, so\n+\t\t\t* the Bloom filter could be used to speed up commands like\n+\t\t\t* 'git log dir/subdir', too.\n+\t\t\t*\n+\t\t\t* Note that directories are added without the trailing '/'.\n+\t\t\t*/\n+\t\t\tdo {\n+\t\t\t\tchar *last_slash = strrchr(path, '/');\n+\n+\t\t\t\tFLEX_ALLOC_STR(e, path, path);\n+\t\t\t\thashmap_entry_init(&e->entry, strhash(path));\n+\t\t\t\thashmap_add(&pathmap, &e->entry);\n+\n+\t\t\t\tif (!last_slash)\n+\t\t\t\t\tlast_slash = (char*)path;\n+\t\t\t\t*last_slash = '\\0';\n+\n+\t\t\t} while (*path);\n+\n+\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n+\t\t}\n+\n+\t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n+\t\tfilter->data = xcalloc(filter->len, sizeof(unsigned char));\n+\n+\t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n+\t\t\tstruct bloom_key key;\n+\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, &settings);\n+\t\t\tadd_key_to_filter(&key, filter, &settings);\n+\t\t}\n+\n+\t\thashmap_free_entries(&pathmap, struct pathmap_hash_entry, entry);\n+\t} else {\n+\t\tfor (i = 0; i < diff_queued_diff.nr; i++)\n+\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n+\t\tfilter->data = NULL;\n+\t\tfilter->len = 0;\n+\t}\n+\n+\tfree(diff_queued_diff.queue);\n+\tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n+\n+\treturn filter;\n+}\ndiff --git a/bloom.h b/bloom.h\nindex b9ce422ca2d..85ab8e9423d 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -1,6 +1,9 @@\n #ifndef BLOOM_H\n #define BLOOM_H\n \n+struct commit;\n+struct repository;\n+\n struct bloom_filter_settings {\n \t/*\n \t * The version of the hashing technique being used.\n@@ -73,4 +76,9 @@ void add_key_to_filter(const struct bloom_key *key,\n \t\t\t\t\t   struct bloom_filter *filter,\n \t\t\t\t\t   const struct bloom_filter_settings *settings);\n \n+void init_bloom_filters(void);\n+\n+struct bloom_filter *get_bloom_filter(struct repository *r,\n+\t\t\t\t      struct commit *c);\n+\n #endif\n\\ No newline at end of file\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 20460cde775..f18d1b722e1 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -1,6 +1,7 @@\n #include \"git-compat-util.h\"\n #include \"bloom.h\"\n #include \"test-tool.h\"\n+#include \"commit.h\"\n \n struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \n@@ -32,6 +33,16 @@ static void print_bloom_filter(struct bloom_filter *filter) {\n \tprintf(\"\\n\");\n }\n \n+static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n+{\n+\tstruct commit *c;\n+\tstruct bloom_filter *filter;\n+\tsetup_git_directory();\n+\tc = lookup_commit(the_repository, commit_oid);\n+\tfilter = get_bloom_filter(the_repository, c);\n+\tprint_bloom_filter(filter);\n+}\n+\n int cmd__bloom(int argc, const char **argv)\n {\n \tif (!strcmp(argv[1], \"get_murmur3\")) {\n@@ -57,5 +68,14 @@ int cmd__bloom(int argc, const char **argv)\n \t\tprint_bloom_filter(&filter);\n \t}\n \n+    if (!strcmp(argv[1], \"get_filter_for_commit\")) {\n+\t\tstruct object_id oid;\n+\t\tconst char *end;\n+\t\tif (parse_oid_hex(argv[2], &oid, &end))\n+\t\t\tdie(\"cannot parse oid '%s'\", argv[2]);\n+\t\tinit_bloom_filters();\n+\t\tget_bloom_filter_for_commit(&oid);\n+\t}\n+\n \treturn 0;\n }\n\\ No newline at end of file\ndiff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\nindex 36a086c7c60..8f9eef116dc 100755\n--- a/t/t0095-bloom.sh\n+++ b/t/t0095-bloom.sh\n@@ -67,4 +67,51 @@ test_expect_success 'compute bloom key for test string 2' '\n \ttest_cmp expect actual\n '\n \n+test_expect_success 'get bloom filters for commit with no changes' '\n+\tgit init &&\n+\tgit commit --allow-empty -m \"c0\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:0\n+\tFilter_Data:\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'get bloom filter for commit with 10 changes' '\n+\trm actual &&\n+\trm expect &&\n+\tmkdir smallDir &&\n+\tfor i in $(test_seq 0 9)\n+\tdo\n+\t\techo $i >smallDir/$i\n+\tdone &&\n+\tgit add smallDir &&\n+\tgit commit -m \"commit with 10 changes\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:25\n+\tFilter_Data:82|a0|65|47|0c|92|90|c0|a1|40|02|a0|e2|40|e0|04|0a|9a|66|cf|80|19|85|42|23|\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success EXPENSIVE 'get bloom filter for commit with 513 changes' '\n+\trm actual &&\n+\trm expect &&\n+\tmkdir bigDir &&\n+\tfor i in $(test_seq 0 512)\n+\tdo\n+\t\techo $i >bigDir/$i\n+\tdone &&\n+\tgit add bigDir &&\n+\tgit commit -m \"commit with 513 changes\" &&\n+\tcat >expect <<-\\EOF &&\n+\tFilter_Length:0\n+\tFilter_Data:\n+\tEOF\n+\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_done\n\\ No newline at end of file\n-- \ngitgitgadget\n\n"},{"id":"394857","messageId":"c38b9b386ef246cd3144ec5ede991c3cad32f3d9.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 06/15] commit-graph: compute Bloom filters for changed paths","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:46Z","receivedAt":"2020-04-06T17:00:06Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd new COMMIT_GRAPH_WRITE_CHANGED_PATHS flag that makes Git compute\nBloom filters for the paths that changed between a commit and it's\nfirst parent, for each commit in the commit-graph.  This computation\nis done on a commit-by-commit basis.\n\nWe will write these Bloom filters to the commit-graph file, to store\nthis data on disk, in the next change in this series.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 32 +++++++++++++++++++++++++++++++-\n commit-graph.h |  3 ++-\n 2 files changed, 33 insertions(+), 2 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex e4f1a5b2f1a..862a00d67ed 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -16,6 +16,7 @@\n #include \"hashmap.h\"\n #include \"replace-object.h\"\n #include \"progress.h\"\n+#include \"bloom.h\"\n \n #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n@@ -789,9 +790,11 @@ struct write_commit_graph_context {\n \tunsigned append:1,\n \t\t report_progress:1,\n \t\t split:1,\n-\t\t check_oids:1;\n+\t\t check_oids:1,\n+\t\t changed_paths:1;\n \n \tconst struct split_commit_graph_opts *split_opts;\n+\tsize_t total_bloom_filter_data_size;\n };\n \n static void write_graph_chunk_fanout(struct hashfile *f,\n@@ -1134,6 +1137,28 @@ static void compute_generation_numbers(struct write_commit_graph_context *ctx)\n \tstop_progress(&ctx->progress);\n }\n \n+static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n+{\n+\tint i;\n+\tstruct progress *progress = NULL;\n+\n+\tinit_bloom_filters();\n+\n+\tif (ctx->report_progress)\n+\t\tprogress = start_delayed_progress(\n+\t\t\t_(\"Computing commit changed paths Bloom filters\"),\n+\t\t\tctx->commits.nr);\n+\n+\tfor (i = 0; i < ctx->commits.nr; i++) {\n+\t\tstruct commit *c = ctx->commits.list[i];\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n+\t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n+\t\tdisplay_progress(progress, i + 1);\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n static int add_ref_to_list(const char *refname,\n \t\t\t   const struct object_id *oid,\n \t\t\t   int flags, void *cb_data)\n@@ -1776,6 +1801,8 @@ int write_commit_graph(struct object_directory *odb,\n \tctx->split = flags & COMMIT_GRAPH_WRITE_SPLIT ? 1 : 0;\n \tctx->check_oids = flags & COMMIT_GRAPH_WRITE_CHECK_OIDS ? 1 : 0;\n \tctx->split_opts = split_opts;\n+\tctx->changed_paths = flags & COMMIT_GRAPH_WRITE_BLOOM_FILTERS ? 1 : 0;\n+\tctx->total_bloom_filter_data_size = 0;\n \n \tif (ctx->split) {\n \t\tstruct commit_graph *g;\n@@ -1870,6 +1897,9 @@ int write_commit_graph(struct object_directory *odb,\n \n \tcompute_generation_numbers(ctx);\n \n+\tif (ctx->changed_paths)\n+\t\tcompute_bloom_filters(ctx);\n+\n \tres = write_commit_graph_file(ctx);\n \n \tif (ctx->split)\ndiff --git a/commit-graph.h b/commit-graph.h\nindex e87a6f63600..86be81219da 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -79,7 +79,8 @@ enum commit_graph_write_flags {\n \tCOMMIT_GRAPH_WRITE_PROGRESS   = (1 << 1),\n \tCOMMIT_GRAPH_WRITE_SPLIT      = (1 << 2),\n \t/* Make sure that each OID in the input is a valid commit OID. */\n-\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3)\n+\tCOMMIT_GRAPH_WRITE_CHECK_OIDS = (1 << 3),\n+\tCOMMIT_GRAPH_WRITE_BLOOM_FILTERS = (1 << 4),\n };\n \n struct split_commit_graph_opts {\n-- \ngitgitgadget\n\n"},{"id":"394865","messageId":"d24c85c54ef841eb2a62d95937c1bf9884ba690a.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 07/15] commit-graph: examine changed-path objects in pack order","fromName":"Jeff King via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:47Z","receivedAt":"2020-04-06T17:00:08Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nLooking at the diff of commit objects in pack order is much faster than\nin sha1 order, as it gives locality to the access of tree deltas\n(whereas sha1 order is effectively random). Unfortunately the\ncommit-graph code sorts the commits (several times, sometimes as an oid\nand sometimes a pointer-to-commit), and we ultimately traverse in sha1\norder.\n\nInstead, let's remember the position at which we see each commit, and\ntraverse in that order when looking at bloom filters. This drops my time\nfor \"git commit-graph write --changed-paths\" in linux.git from ~4\nminutes to ~1.5 minutes.\n\nProbably the \"--reachable\" code path would want something similar.\n\nOr alternatively, we could use a different data structure (either a\nhash, or maybe even just a bit in \"struct commit\") to keep track of\nwhich oids we've seen, etc instead of sorting. And then we could keep\nthe original order.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 38 +++++++++++++++++++++++++++++++++++---\n 1 file changed, 35 insertions(+), 3 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 862a00d67ed..31b06f878ce 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -17,6 +17,7 @@\n #include \"replace-object.h\"\n #include \"progress.h\"\n #include \"bloom.h\"\n+#include \"commit-slab.h\"\n \n #define GRAPH_SIGNATURE 0x43475048 /* \"CGPH\" */\n #define GRAPH_CHUNKID_OIDFANOUT 0x4f494446 /* \"OIDF\" */\n@@ -46,9 +47,32 @@\n /* Remember to update object flag allocation in object.h */\n #define REACHABLE       (1u<<15)\n \n-char *get_commit_graph_filename(struct object_directory *odb)\n+/* Keep track of the order in which commits are added to our list. */\n+define_commit_slab(commit_pos, int);\n+static struct commit_pos commit_pos = COMMIT_SLAB_INIT(1, commit_pos);\n+\n+static void set_commit_pos(struct repository *r, const struct object_id *oid)\n+{\n+\tstatic int32_t max_pos;\n+\tstruct commit *commit = lookup_commit(r, oid);\n+\n+\tif (!commit)\n+\t\treturn; /* should never happen, but be lenient */\n+\n+\t*commit_pos_at(&commit_pos, commit) = max_pos++;\n+}\n+\n+static int commit_pos_cmp(const void *va, const void *vb)\n {\n-\treturn xstrfmt(\"%s/info/commit-graph\", odb->path);\n+\tconst struct commit *a = *(const struct commit **)va;\n+\tconst struct commit *b = *(const struct commit **)vb;\n+\treturn commit_pos_at(&commit_pos, a) -\n+\t       commit_pos_at(&commit_pos, b);\n+}\n+\n+char *get_commit_graph_filename(struct object_directory *obj_dir)\n+{\n+\treturn xstrfmt(\"%s/info/commit-graph\", obj_dir->path);\n }\n \n static char *get_split_graph_filename(struct object_directory *odb,\n@@ -1021,6 +1045,8 @@ static int add_packed_commits(const struct object_id *oid,\n \toidcpy(&(ctx->oids.list[ctx->oids.nr]), oid);\n \tctx->oids.nr++;\n \n+\tset_commit_pos(ctx->r, oid);\n+\n \treturn 0;\n }\n \n@@ -1141,6 +1167,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n {\n \tint i;\n \tstruct progress *progress = NULL;\n+\tstruct commit **sorted_commits;\n \n \tinit_bloom_filters();\n \n@@ -1149,13 +1176,18 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \t\t\t_(\"Computing commit changed paths Bloom filters\"),\n \t\t\tctx->commits.nr);\n \n+\tALLOC_ARRAY(sorted_commits, ctx->commits.nr);\n+\tCOPY_ARRAY(sorted_commits, ctx->commits.list, ctx->commits.nr);\n+\tQSORT(sorted_commits, ctx->commits.nr, commit_pos_cmp);\n+\n \tfor (i = 0; i < ctx->commits.nr; i++) {\n-\t\tstruct commit *c = ctx->commits.list[i];\n+\t\tstruct commit *c = sorted_commits[i];\n \t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n \t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n \t\tdisplay_progress(progress, i + 1);\n \t}\n \n+\tfree(sorted_commits);\n \tstop_progress(&progress);\n }\n \n-- \ngitgitgadget\n\n"},{"id":"394856","messageId":"cc8022bdf82d0ada326ad546fdd7bb7801fc3675.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 10/15] commit-graph: reuse existing Bloom filters during write","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:50Z","receivedAt":"2020-04-06T17:00:09Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd logic to\na) parse Bloom filter information from the commit graph file and,\nb) re-use existing Bloom filters.\n\nSee Documentation/technical/commit-graph-format for the format in which\nthe Bloom filter information is written to the commit graph file.\n\nTo read Bloom filter for a given commit with lexicographic position\n'i' we need to:\n1. Read BIDX[i] which essentially gives us the starting index in BDAT for\n   filter of commit i+1. It is essentially the index past the end\n   of the filter of commit i. It is called end_index in the code.\n\n2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT\n   for filter of commit i. It is called the start_index in the code.\n   For the first commit, where i = 0, Bloom filter data starts at the\n   beginning, just past the header in the BDAT chunk. Hence, start_index\n   will be 0.\n\n3. The length of the filter will be end_index - start_index, because\n   BIDX[i] gives the cumulative 8-byte words including the ith\n   commit's filter.\n\nWe toggle whether Bloom filters should be recomputed based on the\ncompute_if_not_present flag.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c               | 49 ++++++++++++++++++++++++++++++++++++++++++-\n bloom.h               |  4 +++-\n commit-graph.c        |  6 +++---\n t/helper/test-bloom.c |  2 +-\n 4 files changed, 55 insertions(+), 6 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex a16eee92331..0f714dd76ae 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -4,6 +4,8 @@\n #include \"diffcore.h\"\n #include \"revision.h\"\n #include \"hashmap.h\"\n+#include \"commit-graph.h\"\n+#include \"commit.h\"\n \n define_commit_slab(bloom_filter_slab, struct bloom_filter);\n \n@@ -26,6 +28,36 @@ static inline unsigned char get_bitmask(uint32_t pos)\n \treturn ((unsigned char)1) << (pos & (BITS_PER_WORD - 1));\n }\n \n+static int load_bloom_filter_from_graph(struct commit_graph *g,\n+\t\t\t\t   struct bloom_filter *filter,\n+\t\t\t\t   struct commit *c)\n+{\n+\tuint32_t lex_pos, start_index, end_index;\n+\n+\twhile (c->graph_pos < g->num_commits_in_base)\n+\t\tg = g->base_graph;\n+\n+\t/* The commit graph commit 'c' lives in doesn't carry bloom filters. */\n+\tif (!g->chunk_bloom_indexes)\n+\t\treturn 0;\n+\n+\tlex_pos = c->graph_pos - g->num_commits_in_base;\n+\n+\tend_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n+\n+\tif (lex_pos > 0)\n+\t\tstart_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n+\telse\n+\t\tstart_index = 0;\n+\n+\tfilter->len = end_index - start_index;\n+\tfilter->data = (unsigned char *)(g->chunk_bloom_data +\n+\t\t\t\t\tsizeof(unsigned char) * start_index +\n+\t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n+\n+\treturn 1;\n+}\n+\n /*\n  * Calculate the murmur3 32-bit hash value for the given data\n  * using the given seed.\n@@ -127,7 +159,8 @@ void init_bloom_filters(void)\n }\n \n struct bloom_filter *get_bloom_filter(struct repository *r,\n-\t\t\t\t      struct commit *c)\n+\t\t\t\t      struct commit *c,\n+\t\t\t\t\t  int compute_if_not_present)\n {\n \tstruct bloom_filter *filter;\n \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n@@ -140,6 +173,20 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n \n+\tif (!filter->data) {\n+\t\tload_commit_graph_info(r, c);\n+\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH &&\n+\t\t\tr->objects->commit_graph->chunk_bloom_indexes) {\n+\t\t\tif (load_bloom_filter_from_graph(r->objects->commit_graph, filter, c))\n+\t\t\t\treturn filter;\n+\t\t\telse\n+\t\t\t\treturn NULL;\n+\t\t}\n+\t}\n+\n+\tif (filter->data || !compute_if_not_present)\n+\t\treturn filter;\n+\n \trepo_diff_setup(r, &diffopt);\n \tdiffopt.flags.recursive = 1;\n \tdiffopt.max_changes = max_changes;\ndiff --git a/bloom.h b/bloom.h\nindex 85ab8e9423d..760d7122374 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -32,6 +32,7 @@ struct bloom_filter_settings {\n \n #define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n #define BITS_PER_WORD 8\n+#define BLOOMDATA_CHUNK_HEADER_SIZE 3 * sizeof(uint32_t)\n \n /*\n  * A bloom_filter struct represents a data segment to\n@@ -79,6 +80,7 @@ void add_key_to_filter(const struct bloom_key *key,\n void init_bloom_filters(void);\n \n struct bloom_filter *get_bloom_filter(struct repository *r,\n-\t\t\t\t      struct commit *c);\n+\t\t\t\t      struct commit *c,\n+\t\t\t\t      int compute_if_not_present);\n \n #endif\n\\ No newline at end of file\ndiff --git a/commit-graph.c b/commit-graph.c\nindex a8b6b5cca5d..77668629e27 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1086,7 +1086,7 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n \t\t\tctx->commits.nr);\n \n \twhile (list < last) {\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n \t\tcur_pos += filter->len;\n \t\tdisplay_progress(progress, ++i);\n \t\thashwrite_be32(f, cur_pos);\n@@ -1115,7 +1115,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n \thashwrite_be32(f, settings->bits_per_entry);\n \n \twhile (list < last) {\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n \t\tdisplay_progress(progress, ++i);\n \t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n \t\tlist++;\n@@ -1296,7 +1296,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tstruct commit *c = sorted_commits[i];\n-\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n \t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n \t\tdisplay_progress(progress, i + 1);\n \t}\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex f18d1b722e1..ce412664ba9 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -39,7 +39,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n \tstruct bloom_filter *filter;\n \tsetup_git_directory();\n \tc = lookup_commit(the_repository, commit_oid);\n-\tfilter = get_bloom_filter(the_repository, c);\n+\tfilter = get_bloom_filter(the_repository, c, 1);\n \tprint_bloom_filter(filter);\n }\n \n-- \ngitgitgadget\n\n"},{"id":"394858","messageId":"ff6b96aad1e2317d3ed36c2c8b419905dea84a83.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 09/15] commit-graph: write Bloom filters to commit graph file","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:49Z","receivedAt":"2020-04-06T17:00:10Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nUpdate the technical documentation for commit-graph-format with\nthe formats for the Bloom filter index (BIDX) and Bloom filter\ndata (BDAT) chunks. Write the computed Bloom filters information\nto the commit graph file using this format.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n .../technical/commit-graph-format.txt         |  30 +++++\n commit-graph.c                                | 113 +++++++++++++++++-\n commit-graph.h                                |   5 +\n 3 files changed, 147 insertions(+), 1 deletion(-)\n\ndiff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\nindex a4f17441aed..de56f9f1efd 100644\n--- a/Documentation/technical/commit-graph-format.txt\n+++ b/Documentation/technical/commit-graph-format.txt\n@@ -17,6 +17,9 @@ metadata, including:\n - The parents of the commit, stored using positional references within\n   the graph file.\n \n+- The Bloom filter of the commit carrying the paths that were changed between\n+  the commit and its first parent, if requested.\n+\n These positional references are stored as unsigned 32-bit integers\n corresponding to the array position within the list of commit OIDs. Due\n to some special constants we use to track parents, we can store at most\n@@ -93,6 +96,33 @@ CHUNK DATA:\n       positions for the parents until reaching a value with the most-significant\n       bit on. The other bits correspond to the position of the last parent.\n \n+  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) (N * 4 bytes) [Optional]\n+    * The ith entry, BIDX[i], stores the number of 8-byte word blocks in all\n+      Bloom filters from commit 0 to commit i (inclusive) in lexicographic\n+      order. The Bloom filter for the i-th commit spans from BIDX[i-1] to\n+      BIDX[i] (plus header length), where BIDX[-1] is 0.\n+    * The BIDX chunk is ignored if the BDAT chunk is not present.\n+\n+  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n+    * It starts with header consisting of three unsigned 32-bit integers:\n+      - Version of the hash algorithm being used. We currently only support\n+\tvalue 1 which corresponds to the 32-bit version of the murmur3 hash\n+\timplemented exactly as described in\n+\thttps://en.wikipedia.org/wiki/MurmurHash#Algorithm and the double\n+\thashing technique using seed values 0x293ae76f and 0x7e646e2 as\n+\tdescribed in https://doi.org/10.1007/978-3-540-30494-4_26 \"Bloom Filters\n+\tin Probabilistic Verification\"\n+      - The number of times a path is hashed and hence the number of bit positions\n+\t      that cumulatively determine whether a file is present in the commit.\n+      - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n+\t      contains 'n' entries, then the filter size is the minimum number of 64-bit\n+\t      words that contain n*b bits.\n+    * The rest of the chunk is the concatenation of all the computed Bloom\n+      filters for the commits in lexicographic order.\n+    * Note: Commits with no changes or more than 512 changes have Bloom filters\n+      of length zero.\n+    * The BDAT chunk is present if and only if BIDX is present.\n+\n   Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n       This list of H-byte hashes describe a set of B commit-graph files that\n       form a commit-graph chain. The graph position for the ith commit in this\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 732c81fa1b2..a8b6b5cca5d 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -24,8 +24,10 @@\n #define GRAPH_CHUNKID_OIDLOOKUP 0x4f49444c /* \"OIDL\" */\n #define GRAPH_CHUNKID_DATA 0x43444154 /* \"CDAT\" */\n #define GRAPH_CHUNKID_EXTRAEDGES 0x45444745 /* \"EDGE\" */\n+#define GRAPH_CHUNKID_BLOOMINDEXES 0x42494458 /* \"BIDX\" */\n+#define GRAPH_CHUNKID_BLOOMDATA 0x42444154 /* \"BDAT\" */\n #define GRAPH_CHUNKID_BASE 0x42415345 /* \"BASE\" */\n-#define MAX_NUM_CHUNKS 5\n+#define MAX_NUM_CHUNKS 7\n \n #define GRAPH_DATA_WIDTH (the_hash_algo->rawsz + 16)\n \n@@ -319,6 +321,32 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n \t\t\t\tchunk_repeated = 1;\n \t\t\telse\n \t\t\t\tgraph->chunk_base_graphs = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_BLOOMINDEXES:\n+\t\t\tif (graph->chunk_bloom_indexes)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse\n+\t\t\t\tgraph->chunk_bloom_indexes = data + chunk_offset;\n+\t\t\tbreak;\n+\n+\t\tcase GRAPH_CHUNKID_BLOOMDATA:\n+\t\t\tif (graph->chunk_bloom_data)\n+\t\t\t\tchunk_repeated = 1;\n+\t\t\telse {\n+\t\t\t\tuint32_t hash_version;\n+\t\t\t\tgraph->chunk_bloom_data = data + chunk_offset;\n+\t\t\t\thash_version = get_be32(data + chunk_offset);\n+\n+\t\t\t\tif (hash_version != 1)\n+\t\t\t\t\tbreak;\n+\n+\t\t\t\tgraph->bloom_filter_settings = xmalloc(sizeof(struct bloom_filter_settings));\n+\t\t\t\tgraph->bloom_filter_settings->hash_version = hash_version;\n+\t\t\t\tgraph->bloom_filter_settings->num_hashes = get_be32(data + chunk_offset + 4);\n+\t\t\t\tgraph->bloom_filter_settings->bits_per_entry = get_be32(data + chunk_offset + 8);\n+\t\t\t}\n+\t\t\tbreak;\n \t\t}\n \n \t\tif (chunk_repeated) {\n@@ -337,6 +365,15 @@ struct commit_graph *parse_commit_graph(void *graph_map, int fd,\n \t\tlast_chunk_offset = chunk_offset;\n \t}\n \n+\tif (graph->chunk_bloom_indexes && graph->chunk_bloom_data) {\n+\t\tinit_bloom_filters();\n+\t} else {\n+\t\t/* We need both the bloom chunks to exist together. Else ignore the data */\n+\t\tgraph->chunk_bloom_indexes = NULL;\n+\t\tgraph->chunk_bloom_data = NULL;\n+\t\tgraph->bloom_filter_settings = NULL;\n+\t}\n+\n \thashcpy(graph->oid.hash, graph->data + graph->data_len - graph->hash_len);\n \n \tif (verify_commit_graph_lite(graph)) {\n@@ -1034,6 +1071,59 @@ static void write_graph_chunk_extra_edges(struct hashfile *f,\n \t}\n }\n \n+static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n+\t\t\t\t\t    struct write_commit_graph_context *ctx)\n+{\n+\tstruct commit **list = ctx->commits.list;\n+\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n+\tuint32_t cur_pos = 0;\n+\tstruct progress *progress = NULL;\n+\tint i = 0;\n+\n+\tif (ctx->report_progress)\n+\t\tprogress = start_delayed_progress(\n+\t\t\t_(\"Writing changed paths Bloom filters index\"),\n+\t\t\tctx->commits.nr);\n+\n+\twhile (list < last) {\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tcur_pos += filter->len;\n+\t\tdisplay_progress(progress, ++i);\n+\t\thashwrite_be32(f, cur_pos);\n+\t\tlist++;\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n+static void write_graph_chunk_bloom_data(struct hashfile *f,\n+\t\t\t\t\t struct write_commit_graph_context *ctx,\n+\t\t\t\t\t const struct bloom_filter_settings *settings)\n+{\n+\tstruct commit **list = ctx->commits.list;\n+\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n+\tstruct progress *progress = NULL;\n+\tint i = 0;\n+\n+\tif (ctx->report_progress)\n+\t\tprogress = start_delayed_progress(\n+\t\t\t_(\"Writing changed paths Bloom filters data\"),\n+\t\t\tctx->commits.nr);\n+\n+\thashwrite_be32(f, settings->hash_version);\n+\thashwrite_be32(f, settings->num_hashes);\n+\thashwrite_be32(f, settings->bits_per_entry);\n+\n+\twhile (list < last) {\n+\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n+\t\tdisplay_progress(progress, ++i);\n+\t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n+\t\tlist++;\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n static int oid_compare(const void *_a, const void *_b)\n {\n \tconst struct object_id *a = (const struct object_id *)_a;\n@@ -1438,6 +1528,7 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \tstruct strbuf progress_title = STRBUF_INIT;\n \tint num_chunks = 3;\n \tstruct object_id file_hash;\n+\tconst struct bloom_filter_settings bloom_settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n \n \tif (ctx->split) {\n \t\tstruct strbuf tmp_file = STRBUF_INIT;\n@@ -1482,6 +1573,12 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_EXTRAEDGES;\n \t\tnum_chunks++;\n \t}\n+\tif (ctx->changed_paths) {\n+\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMINDEXES;\n+\t\tnum_chunks++;\n+\t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BLOOMDATA;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_commit_graphs_after > 1) {\n \t\tchunk_ids[num_chunks] = GRAPH_CHUNKID_BASE;\n \t\tnum_chunks++;\n@@ -1500,6 +1597,15 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \t\t\t\t\t\t4 * ctx->num_extra_edges;\n \t\tnum_chunks++;\n \t}\n+\tif (ctx->changed_paths) {\n+\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n+\t\t\t\t\t\tsizeof(uint32_t) * ctx->commits.nr;\n+\t\tnum_chunks++;\n+\n+\t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n+\t\t\t\t\t\tsizeof(uint32_t) * 3 + ctx->total_bloom_filter_data_size;\n+\t\tnum_chunks++;\n+\t}\n \tif (ctx->num_commit_graphs_after > 1) {\n \t\tchunk_offsets[num_chunks + 1] = chunk_offsets[num_chunks] +\n \t\t\t\t\t\thashsz * (ctx->num_commit_graphs_after - 1);\n@@ -1537,6 +1643,10 @@ static int write_commit_graph_file(struct write_commit_graph_context *ctx)\n \twrite_graph_chunk_data(f, hashsz, ctx);\n \tif (ctx->num_extra_edges)\n \t\twrite_graph_chunk_extra_edges(f, ctx);\n+\tif (ctx->changed_paths) {\n+\t\twrite_graph_chunk_bloom_indexes(f, ctx);\n+\t\twrite_graph_chunk_bloom_data(f, ctx, &bloom_settings);\n+\t}\n \tif (ctx->num_commit_graphs_after > 1 &&\n \t    write_graph_chunk_base(f, ctx)) {\n \t\treturn -1;\n@@ -2184,6 +2294,7 @@ void free_commit_graph(struct commit_graph *g)\n \t\tclose(g->graph_fd);\n \t}\n \tfree(g->filename);\n+\tfree(g->bloom_filter_settings);\n \tfree(g);\n }\n \ndiff --git a/commit-graph.h b/commit-graph.h\nindex 86be81219da..8e7a8e0e5b2 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -11,6 +11,7 @@\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n \n struct commit;\n+struct bloom_filter_settings;\n \n char *get_commit_graph_filename(struct object_directory *odb);\n int open_commit_graph(const char *graph_file, int *fd, struct stat *st);\n@@ -59,6 +60,10 @@ struct commit_graph {\n \tconst unsigned char *chunk_commit_data;\n \tconst unsigned char *chunk_extra_edges;\n \tconst unsigned char *chunk_base_graphs;\n+\tconst unsigned char *chunk_bloom_indexes;\n+\tconst unsigned char *chunk_bloom_data;\n+\n+\tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n struct commit_graph *load_commit_graph_one_fd_st(int fd, struct stat *st,\n-- \ngitgitgadget\n\n"},{"id":"394861","messageId":"5ed16f35fed43018ce441adfff55b85967d3918c.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 08/15] commit-graph: examine commits by generation number","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:48Z","receivedAt":"2020-04-06T17:00:10Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nWhen running 'git commit-graph write --changed-paths', we sort the\ncommits by pack-order to save time when computing the changed-paths\nbloom filters. This does not help when finding the commits via the\n'--reachable' flag.\n\nIf not using pack-order, then sort by generation number before\nexamining the diff. Commits with similar generation are more likely\nto have many trees in common, making the diff faster.\n\nOn the Linux kernel repository, this change reduced the computation\ntime for 'git commit-graph write --reachable --changed-paths' from\n3m00s to 1m37s.\n\nHelped-by: Jeff King <peff@peff.net>\nSigned-off-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n commit-graph.c | 33 ++++++++++++++++++++++++++++++---\n 1 file changed, 30 insertions(+), 3 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex 31b06f878ce..732c81fa1b2 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -70,6 +70,25 @@ static int commit_pos_cmp(const void *va, const void *vb)\n \t       commit_pos_at(&commit_pos, b);\n }\n \n+static int commit_gen_cmp(const void *va, const void *vb)\n+{\n+\tconst struct commit *a = *(const struct commit **)va;\n+\tconst struct commit *b = *(const struct commit **)vb;\n+\n+\t/* lower generation commits first */\n+\tif (a->generation < b->generation)\n+\t\treturn -1;\n+\telse if (a->generation > b->generation)\n+\t\treturn 1;\n+\n+\t/* use date as a heuristic when generations are equal */\n+\tif (a->date < b->date)\n+\t\treturn -1;\n+\telse if (a->date > b->date)\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n char *get_commit_graph_filename(struct object_directory *obj_dir)\n {\n \treturn xstrfmt(\"%s/info/commit-graph\", obj_dir->path);\n@@ -815,7 +834,8 @@ struct write_commit_graph_context {\n \t\t report_progress:1,\n \t\t split:1,\n \t\t check_oids:1,\n-\t\t changed_paths:1;\n+\t\t changed_paths:1,\n+\t\t order_by_pack:1;\n \n \tconst struct split_commit_graph_opts *split_opts;\n \tsize_t total_bloom_filter_data_size;\n@@ -1178,7 +1198,11 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n \n \tALLOC_ARRAY(sorted_commits, ctx->commits.nr);\n \tCOPY_ARRAY(sorted_commits, ctx->commits.list, ctx->commits.nr);\n-\tQSORT(sorted_commits, ctx->commits.nr, commit_pos_cmp);\n+\n+\tif (ctx->order_by_pack)\n+\t\tQSORT(sorted_commits, ctx->commits.nr, commit_pos_cmp);\n+\telse\n+\t\tQSORT(sorted_commits, ctx->commits.nr, commit_gen_cmp);\n \n \tfor (i = 0; i < ctx->commits.nr; i++) {\n \t\tstruct commit *c = sorted_commits[i];\n@@ -1884,6 +1908,7 @@ int write_commit_graph(struct object_directory *odb,\n \t}\n \n \tif (pack_indexes) {\n+\t\tctx->order_by_pack = 1;\n \t\tif ((res = fill_oids_from_packs(ctx, pack_indexes)))\n \t\t\tgoto cleanup;\n \t}\n@@ -1893,8 +1918,10 @@ int write_commit_graph(struct object_directory *odb,\n \t\t\tgoto cleanup;\n \t}\n \n-\tif (!pack_indexes && !commit_hex)\n+\tif (!pack_indexes && !commit_hex) {\n+\t\tctx->order_by_pack = 1;\n \t\tfill_oids_from_all_packs(ctx);\n+\t}\n \n \tclose_reachable(ctx);\n \n-- \ngitgitgadget\n\n"},{"id":"394859","messageId":"6beaede715972d7726dda8a58bbd4b920813e194.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 13/15] revision.c: add trace2 stats around Bloom filter usage","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:53Z","receivedAt":"2020-04-06T17:00:12Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd trace2 statistics around Bloom filter usage and behavior\nfor 'git log -- path' commands that are hoping to benefit from\nthe presence of computed changed paths Bloom filters.\n\nThese statistics are great for performance analysis work and\nfor formal testing, which we will see in the commit following\nthis one.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com\nHelped-by: SZEDER Gábor <szeder.dev@gmail.com>\nHelped-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n revision.c | 41 +++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 41 insertions(+)\n\ndiff --git a/revision.c b/revision.c\nindex d3fcb7c6ff6..2b06ee739c8 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -30,6 +30,7 @@\n #include \"hashmap.h\"\n #include \"utf8.h\"\n #include \"bloom.h\"\n+#include \"json-writer.h\"\n \n volatile show_early_output_fn_t show_early_output;\n \n@@ -625,6 +626,30 @@ static void file_change(struct diff_options *options,\n \toptions->flags.has_changes = 1;\n }\n \n+static int bloom_filter_atexit_registered;\n+static unsigned int count_bloom_filter_maybe;\n+static unsigned int count_bloom_filter_definitely_not;\n+static unsigned int count_bloom_filter_false_positive;\n+static unsigned int count_bloom_filter_not_present;\n+static unsigned int count_bloom_filter_length_zero;\n+\n+static void trace2_bloom_filter_statistics_atexit(void)\n+{\n+\tstruct json_writer jw = JSON_WRITER_INIT;\n+\n+\tjw_object_begin(&jw, 0);\n+\tjw_object_intmax(&jw, \"filter_not_present\", count_bloom_filter_not_present);\n+\tjw_object_intmax(&jw, \"zero_length_filter\", count_bloom_filter_length_zero);\n+\tjw_object_intmax(&jw, \"maybe\", count_bloom_filter_maybe);\n+\tjw_object_intmax(&jw, \"definitely_not\", count_bloom_filter_definitely_not);\n+\tjw_object_intmax(&jw, \"false_positive\", count_bloom_filter_false_positive);\n+\tjw_end(&jw);\n+\n+\ttrace2_data_json(\"bloom\", the_repository, \"statistics\", &jw);\n+\n+\tjw_release(&jw);\n+}\n+\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n@@ -661,6 +686,11 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n \tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n \n+\tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n+\t\tatexit(trace2_bloom_filter_statistics_atexit);\n+\t\tbloom_filter_atexit_registered = 1;\n+\t}\n+\n \tfree(path_alloc);\n }\n \n@@ -679,10 +709,12 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \tfilter = get_bloom_filter(revs->repo, commit, 0);\n \n \tif (!filter) {\n+\t\tcount_bloom_filter_not_present++;\n \t\treturn -1;\n \t}\n \n \tif (!filter->len) {\n+\t\tcount_bloom_filter_length_zero++;\n \t\treturn -1;\n \t}\n \n@@ -690,6 +722,11 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\t\t\t       revs->bloom_key,\n \t\t\t\t       revs->bloom_filter_settings);\n \n+\tif (result)\n+\t\tcount_bloom_filter_maybe++;\n+\telse\n+\t\tcount_bloom_filter_definitely_not++;\n+\n \treturn result;\n }\n \n@@ -736,6 +773,10 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\t   &revs->pruning) < 0)\n \t\treturn REV_TREE_DIFFERENT;\n \n+\tif (!nth_parent)\n+\t\tif (bloom_ret == 1 && tree_difference == REV_TREE_SAME)\n+\t\t\tcount_bloom_filter_false_positive++;\n+\n \treturn tree_difference;\n }\n \n-- \ngitgitgadget\n\n"},{"id":"394862","messageId":"b899df5c98e302c23cc50ee2b2d187dfde4e1e7f.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 14/15] t4216: add end to end tests for git log with Bloom filters","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:54Z","receivedAt":"2020-04-06T17:00:12Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nThese tests exercises writing commit graph with Bloom filters\nand exercises 'git log -- path' with all the applicable\noptions. They check that the output is the same with and\nwithout Bloom filters, confirm Bloom filters were used by\nchecking if trace2 statistics were logged correctly.\n\nAlso confirms cases where Bloom filters are not used:\n1. Multiple path specs,\n2. --walk-reflogs (see patch titled 'revision.c: use Bloom filters...'\n   for details,\n3. If the latest commit graph does not have Bloom filters\n\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n t/helper/test-read-graph.c |   4 +\n t/t4216-log-bloom.sh       | 155 +++++++++++++++++++++++++++++++++++++\n 2 files changed, 159 insertions(+)\n create mode 100755 t/t4216-log-bloom.sh\n\ndiff --git a/t/helper/test-read-graph.c b/t/helper/test-read-graph.c\nindex f8a461767ca..4223ff32fb6 100644\n--- a/t/helper/test-read-graph.c\n+++ b/t/helper/test-read-graph.c\n@@ -45,6 +45,10 @@ int cmd__read_graph(int argc, const char **argv)\n \t\tprintf(\" commit_metadata\");\n \tif (graph->chunk_extra_edges)\n \t\tprintf(\" extra_edges\");\n+\tif (graph->chunk_bloom_indexes)\n+\t\tprintf(\" bloom_indexes\");\n+\tif (graph->chunk_bloom_data)\n+\t\tprintf(\" bloom_data\");\n \tprintf(\"\\n\");\n \n \tUNLEAK(graph);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nnew file mode 100755\nindex 00000000000..38accd272df\n--- /dev/null\n+++ b/t/t4216-log-bloom.sh\n@@ -0,0 +1,155 @@\n+#!/bin/sh\n+\n+test_description='git log for a path with Bloom filters'\n+. ./test-lib.sh\n+\n+GIT_TEST_COMMIT_GRAPH=0\n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n+\n+test_expect_success 'setup test - repo, commits, commit graph, log outputs' '\n+\tgit init &&\n+\tmkdir A A/B A/B/C &&\n+\ttest_commit c1 A/file1 &&\n+\ttest_commit c2 A/B/file2 &&\n+\ttest_commit c3 A/B/C/file3 &&\n+\ttest_commit c4 A/file1 &&\n+\ttest_commit c5 A/B/file2 &&\n+\ttest_commit c6 A/B/C/file3 &&\n+\ttest_commit c7 A/file1 &&\n+\ttest_commit c8 A/B/file2 &&\n+\ttest_commit c9 A/B/C/file3 &&\n+\ttest_commit c10 file_to_be_deleted &&\n+\tgit checkout -b side HEAD~4 &&\n+\ttest_commit side-1 file4 &&\n+\tgit checkout master &&\n+\tgit merge side &&\n+\ttest_commit c11 file5 &&\n+\tmv file5 file5_renamed &&\n+\tgit add file5_renamed &&\n+\tgit commit -m \"rename\" &&\n+\trm file_to_be_deleted &&\n+\tgit add . &&\n+\tgit commit -m \"file removed\" &&\n+\tgit commit-graph write --reachable --changed-paths\n+'\n+graph_read_expect () {\n+\tNUM_CHUNKS=5\n+\tcat >expect <<- EOF\n+\theader: 43475048 1 1 $NUM_CHUNKS 0\n+\tnum_commits: $1\n+\tchunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n+\tEOF\n+\ttest-tool read-graph >actual &&\n+\ttest_cmp expect actual\n+}\n+\n+test_expect_success 'commit-graph write wrote out the bloom chunks' '\n+\tgraph_read_expect 15\n+'\n+\n+# Turn off any inherited trace2 settings for this test.\n+sane_unset GIT_TRACE2 GIT_TRACE2_PERF GIT_TRACE2_EVENT\n+sane_unset GIT_TRACE2_PERF_BRIEF\n+sane_unset GIT_TRACE2_CONFIG_PARAMS\n+\n+setup () {\n+\trm \"$TRASH_DIRECTORY/trace.perf\"\n+\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n+\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n+}\n+\n+test_bloom_filters_used () {\n+\tlog_args=$1\n+\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":0,\\\"zero_length_filter\\\":0,\\\"maybe\\\"\"\n+\tsetup \"$log_args\" &&\n+\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && \n+\ttest_cmp log_wo_bloom log_w_bloom &&\n+    test_path_is_file \"$TRASH_DIRECTORY/trace.perf\"\n+}\n+\n+test_bloom_filters_not_used () {\n+\tlog_args=$1\n+\tsetup \"$log_args\" &&\n+\t!(grep -q \"statistics:{\\\"filter_not_present\\\":\" \"$TRASH_DIRECTORY/trace.perf\") && \n+\ttest_cmp log_wo_bloom log_w_bloom\n+}\n+\n+for path in A A/B A/B/C A/file1 A/B/file2 A/B/C/file3 file4 file5 file5_renamed file_to_be_deleted\n+do\n+\tfor option in \"\" \\\n+              \"--all\" \\\n+\t\t      \"--full-history\" \\\n+\t\t      \"--full-history --simplify-merges\" \\\n+\t\t      \"--simplify-merges\" \\\n+\t\t      \"--simplify-by-decoration\" \\\n+\t\t      \"--follow\" \\\n+\t\t      \"--first-parent\" \\\n+\t\t      \"--topo-order\" \\\n+\t\t      \"--date-order\" \\\n+\t\t      \"--author-date-order\" \\\n+\t\t      \"--ancestry-path side..master\"\n+\tdo\n+\t\ttest_expect_success \"git log option: $option for path: $path\" '\n+\t\t\ttest_bloom_filters_used \"$option -- $path\"\n+\t\t'\n+\tdone\n+done\n+\n+test_expect_success 'git log -- folder works with and without the trailing slash' '\n+\ttest_bloom_filters_used \"-- A\" &&\n+\ttest_bloom_filters_used \"-- A/\"\n+'\n+\n+test_expect_success 'git log for path that does not exist. ' '\n+\ttest_bloom_filters_used \"-- path_does_not_exist\"\n+'\n+\n+test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n+\ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n+'\n+\n+test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n+\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n+'\n+\n+test_expect_success 'git log with wildcard that resolves to a single path uses Bloom filters' '\n+\ttest_bloom_filters_used \"-- *4\" &&\n+\ttest_bloom_filters_used \"-- *renamed\"\n+'\n+\n+test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n+\ttest_bloom_filters_not_used \"-- *\" &&\n+\ttest_bloom_filters_not_used \"-- file*\"\n+'\n+\n+test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n+\ttest_commit c14 A/anotherFile2 &&\n+\ttest_commit c15 A/B/anotherFile2 &&\n+\ttest_commit c16 A/B/C/anotherFile2 &&\n+\tGIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0 git commit-graph write --reachable --split &&\n+\ttest_line_count = 2 .git/objects/info/commit-graphs/commit-graph-chain\n+'\n+\n+test_expect_success 'Do not use Bloom filters if the latest graph does not have Bloom filters.' '\n+\ttest_bloom_filters_not_used \"-- A/B\"\n+'\n+\n+test_expect_success 'setup - add commit-graph to the chain with Bloom filters' '\n+\ttest_commit c17 A/anotherFile3 &&\n+\tgit commit-graph write --reachable --changed-paths --split &&\n+\ttest_line_count = 3 .git/objects/info/commit-graphs/commit-graph-chain\n+'\n+\n+test_bloom_filters_used_when_some_filters_are_missing () {\n+\tlog_args=$1\n+\tbloom_trace_prefix=\"statistics:{\\\"filter_not_present\\\":3,\\\"zero_length_filter\\\":0,\\\"maybe\\\":8,\\\"definitely_not\\\":6\"\n+\tsetup \"$log_args\" &&\n+\tgrep -q \"$bloom_trace_prefix\" \"$TRASH_DIRECTORY/trace.perf\" && \n+\ttest_cmp log_wo_bloom log_w_bloom\n+}\n+\n+test_expect_success 'Use Bloom filters if they exist in the latest but not all commit graphs in the chain.' '\n+\ttest_bloom_filters_used_when_some_filters_are_missing \"-- A/B\"\n+'\n+\n+test_done\n\\ No newline at end of file\n-- \ngitgitgadget\n\n"},{"id":"394860","messageId":"617f549ef259424658a84dd67a98685328f6b850.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 12/15] revision.c: use Bloom filters to speed up path based revision walks","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:52Z","receivedAt":"2020-04-06T17:00:14Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nRevision walk will now use Bloom filters for commits to speed up\nrevision walks for a particular path (for computing history for\nthat path), if they are present in the commit-graph file.\n\nWe load the Bloom filters during the prepare_revision_walk step,\ncurrently only when dealing with a single pathspec. Extending\nit to work with multiple pathspecs can be explored and built on\ntop of this series in the future.\n\nWhile comparing trees in rev_compare_trees(), if the Bloom filter\nsays that the file is not different between the two trees, we don't\nneed to compute the expensive diff. This is where we get our\nperformance gains. The other response of the Bloom filter is '`:maybe',\nin which case we fall back to the full diff calculation to determine\nif the path was changed in the commit.\n\nWe do not try to use Bloom filters when the '--walk-reflogs' option\nis specified. The '--walk-reflogs' option does not walk the commit\nancestry chain like the rest of the options. Incorporating the\nperformance gains when walking reflog entries would add more\ncomplexity, and can be explored in a later series.\n\nPerformance Gains:\nWe tested the performance of `git log -- <path>` on the git repo, the linux\nand some internal large repos, with a variety of paths of varying depths.\n\nOn the git and linux repos:\n- we observed a 2x to 5x speed up.\n\nOn a large internal repo with files seated 6-10 levels deep in the tree:\n- we observed 10x to 20x speed ups, with some paths going up to 28 times\n  faster.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com\nHelped-by: SZEDER Gábor <szeder.dev@gmail.com>\nHelped-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n bloom.c    | 20 +++++++++++++\n bloom.h    |  4 +++\n revision.c | 85 ++++++++++++++++++++++++++++++++++++++++++++++++++++--\n revision.h | 11 +++++++\n 4 files changed, 118 insertions(+), 2 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 0f714dd76ae..c5b461d1cfe 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -253,3 +253,23 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \n \treturn filter;\n }\n+\n+int bloom_filter_contains(const struct bloom_filter *filter,\n+\t\t\t  const struct bloom_key *key,\n+\t\t\t  const struct bloom_filter_settings *settings)\n+{\n+\tint i;\n+\tuint64_t mod = filter->len * BITS_PER_WORD;\n+\n+\tif (!mod)\n+\t\treturn -1;\n+\n+\tfor (i = 0; i < settings->num_hashes; i++) {\n+\t\tuint64_t hash_mod = key->hashes[i] % mod;\n+\t\tuint64_t block_pos = hash_mod / BITS_PER_WORD;\n+\t\tif (!(filter->data[block_pos] & get_bitmask(hash_mod)))\n+\t\t\treturn 0;\n+\t}\n+\n+\treturn 1;\n+}\n\\ No newline at end of file\ndiff --git a/bloom.h b/bloom.h\nindex 760d7122374..b935186425d 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -83,4 +83,8 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \t\t\t\t      struct commit *c,\n \t\t\t\t      int compute_if_not_present);\n \n+int bloom_filter_contains(const struct bloom_filter *filter,\n+\t\t\t  const struct bloom_key *key,\n+\t\t\t  const struct bloom_filter_settings *settings);\n+\n #endif\n\\ No newline at end of file\ndiff --git a/revision.c b/revision.c\nindex 8136929e236..d3fcb7c6ff6 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -29,6 +29,7 @@\n #include \"prio-queue.h\"\n #include \"hashmap.h\"\n #include \"utf8.h\"\n+#include \"bloom.h\"\n \n volatile show_early_output_fn_t show_early_output;\n \n@@ -624,11 +625,80 @@ static void file_change(struct diff_options *options,\n \toptions->flags.has_changes = 1;\n }\n \n+static void prepare_to_use_bloom_filter(struct rev_info *revs)\n+{\n+\tstruct pathspec_item *pi;\n+\tchar *path_alloc = NULL;\n+\tconst char *path;\n+\tint last_index;\n+\tint len;\n+\n+\tif (!revs->commits)\n+\t    return;\n+\n+\trepo_parse_commit(revs->repo, revs->commits->item);\n+\n+\tif (!revs->repo->objects->commit_graph)\n+\t\treturn;\n+\n+\trevs->bloom_filter_settings = revs->repo->objects->commit_graph->bloom_filter_settings;\n+\tif (!revs->bloom_filter_settings)\n+\t\treturn;\n+\n+\tpi = &revs->pruning.pathspec.items[0];\n+\tlast_index = pi->len - 1;\n+\n+\t/* remove single trailing slash from path, if needed */\n+\tif (pi->match[last_index] == '/') {\n+\t    path_alloc = xstrdup(pi->match);\n+\t    path_alloc[last_index] = '\\0';\n+\t    path = path_alloc;\n+\t} else\n+\t    path = pi->match;\n+\n+\tlen = strlen(path);\n+\n+\trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n+\tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n+\n+\tfree(path_alloc);\n+}\n+\n+static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n+\t\t\t\t\t\t struct commit *commit)\n+{\n+\tstruct bloom_filter *filter;\n+\tint result;\n+\n+\tif (!revs->repo->objects->commit_graph)\n+\t\treturn -1;\n+\n+\tif (commit->generation == GENERATION_NUMBER_INFINITY)\n+\t\treturn -1;\n+\n+\tfilter = get_bloom_filter(revs->repo, commit, 0);\n+\n+\tif (!filter) {\n+\t\treturn -1;\n+\t}\n+\n+\tif (!filter->len) {\n+\t\treturn -1;\n+\t}\n+\n+\tresult = bloom_filter_contains(filter,\n+\t\t\t\t       revs->bloom_key,\n+\t\t\t\t       revs->bloom_filter_settings);\n+\n+\treturn result;\n+}\n+\n static int rev_compare_tree(struct rev_info *revs,\n-\t\t\t    struct commit *parent, struct commit *commit)\n+\t\t\t    struct commit *parent, struct commit *commit, int nth_parent)\n {\n \tstruct tree *t1 = get_commit_tree(parent);\n \tstruct tree *t2 = get_commit_tree(commit);\n+\tint bloom_ret = 1;\n \n \tif (!t1)\n \t\treturn REV_TREE_NEW;\n@@ -653,11 +723,19 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\treturn REV_TREE_SAME;\n \t}\n \n+\tif (revs->bloom_key && !nth_parent) {\n+\t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n+\n+\t\tif (bloom_ret == 0)\n+\t\t\treturn REV_TREE_SAME;\n+\t}\n+\n \ttree_difference = REV_TREE_SAME;\n \trevs->pruning.flags.has_changes = 0;\n \tif (diff_tree_oid(&t1->object.oid, &t2->object.oid, \"\",\n \t\t\t   &revs->pruning) < 0)\n \t\treturn REV_TREE_DIFFERENT;\n+\n \treturn tree_difference;\n }\n \n@@ -855,7 +933,7 @@ static void try_to_simplify_commit(struct rev_info *revs, struct commit *commit)\n \t\t\tdie(\"cannot simplify commit %s (because of %s)\",\n \t\t\t    oid_to_hex(&commit->object.oid),\n \t\t\t    oid_to_hex(&p->object.oid));\n-\t\tswitch (rev_compare_tree(revs, p, commit)) {\n+\t\tswitch (rev_compare_tree(revs, p, commit, nth_parent)) {\n \t\tcase REV_TREE_SAME:\n \t\t\tif (!revs->simplify_history || !relevant_commit(p)) {\n \t\t\t\t/* Even if a merge with an uninteresting\n@@ -3362,6 +3440,8 @@ int prepare_revision_walk(struct rev_info *revs)\n \t\t\t\t       FOR_EACH_OBJECT_PROMISOR_ONLY);\n \t}\n \n+\tif (revs->pruning.pathspec.nr == 1 && !revs->reflog_info)\n+\t\tprepare_to_use_bloom_filter(revs);\n \tif (revs->no_walk != REVISION_WALK_NO_WALK_UNSORTED)\n \t\tcommit_list_sort_by_date(&revs->commits);\n \tif (revs->no_walk)\n@@ -3379,6 +3459,7 @@ int prepare_revision_walk(struct rev_info *revs)\n \t\tsimplify_merges(revs);\n \tif (revs->children.name)\n \t\tset_children(revs);\n+\n \treturn 0;\n }\n \ndiff --git a/revision.h b/revision.h\nindex 475f048fb61..7c026fe41fc 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -56,6 +56,8 @@ struct repository;\n struct rev_info;\n struct string_list;\n struct saved_parents;\n+struct bloom_key;\n+struct bloom_filter_settings;\n define_shared_commit_slab(revision_sources, char *);\n \n struct rev_cmdline_info {\n@@ -291,6 +293,15 @@ struct rev_info {\n \tstruct revision_sources *sources;\n \n \tstruct topo_walk_info *topo_walk_info;\n+\n+\t/* Commit graph bloom filter fields */\n+\t/* The bloom filter key for the pathspec */\n+\tstruct bloom_key *bloom_key;\n+\t/*\n+\t * The bloom filter settings used to generate the key.\n+\t * This is loaded from the commit-graph being used.\n+\t */\n+\tstruct bloom_filter_settings *bloom_filter_settings;\n };\n \n int ref_excluded(struct string_list *, const char *path);\n-- \ngitgitgadget\n\n"},{"id":"394863","messageId":"c8b86c383abdbbd31ba307eb7e79942ddde1b711.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 11/15] commit-graph: add --changed-paths option to write subcommand","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:51Z","receivedAt":"2020-04-06T17:00:15Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd --changed-paths option to git commit-graph write. This option will\nallow users to compute information about the paths that have changed\nbetween a commit and its first parent, and write it into the commit graph\nfile. If the option is passed to the write subcommand we set the\nCOMMIT_GRAPH_WRITE_BLOOM_FILTERS flag and pass it down to the\ncommit-graph logic.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n Documentation/git-commit-graph.txt | 5 +++++\n builtin/commit-graph.c             | 9 +++++++--\n 2 files changed, 12 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/git-commit-graph.txt b/Documentation/git-commit-graph.txt\nindex 28d1fee5053..f4b13c005b8 100644\n--- a/Documentation/git-commit-graph.txt\n+++ b/Documentation/git-commit-graph.txt\n@@ -57,6 +57,11 @@ or `--stdin-packs`.)\n With the `--append` option, include all commits that are present in the\n existing commit-graph file.\n +\n+With the `--changed-paths` option, compute and write information about the\n+paths changed between a commit and it's first parent. This operation can\n+take a while on large repositories. It provides significant performance gains\n+for getting history of a directory or a file with `git log -- <path>`.\n++\n With the `--split` option, write the commit-graph as a chain of multiple\n commit-graph files stored in `<dir>/info/commit-graphs`. The new commits\n not already in the commit-graph are added in a new \"tip\" file. This file\ndiff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\nindex d1ab6625f63..cacb5d04a80 100644\n--- a/builtin/commit-graph.c\n+++ b/builtin/commit-graph.c\n@@ -9,7 +9,7 @@\n \n static char const * const builtin_commit_graph_usage[] = {\n \tN_(\"git commit-graph verify [--object-dir <objdir>] [--shallow] [--[no-]progress]\"),\n-\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n+\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n \tNULL\n };\n \n@@ -19,7 +19,7 @@ static const char * const builtin_commit_graph_verify_usage[] = {\n };\n \n static const char * const builtin_commit_graph_write_usage[] = {\n-\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n+\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n \tNULL\n };\n \n@@ -32,6 +32,7 @@ static struct opts_commit_graph {\n \tint split;\n \tint shallow;\n \tint progress;\n+\tint enable_changed_paths;\n } opts;\n \n static struct object_directory *find_odb(struct repository *r,\n@@ -135,6 +136,8 @@ static int graph_write(int argc, const char **argv)\n \t\t\tN_(\"start walk at commits listed by stdin\")),\n \t\tOPT_BOOL(0, \"append\", &opts.append,\n \t\t\tN_(\"include all commits already in the commit-graph file\")),\n+\t\tOPT_BOOL(0, \"changed-paths\", &opts.enable_changed_paths,\n+\t\t\tN_(\"enable computation for changed paths\")),\n \t\tOPT_BOOL(0, \"progress\", &opts.progress, N_(\"force progress reporting\")),\n \t\tOPT_BOOL(0, \"split\", &opts.split,\n \t\t\tN_(\"allow writing an incremental commit-graph file\")),\n@@ -168,6 +171,8 @@ static int graph_write(int argc, const char **argv)\n \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n \tif (opts.progress)\n \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n+\tif (opts.enable_changed_paths)\n+\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n \n \tread_replace_refs = 0;\n \todb = find_odb(the_repository, opts.obj_dir);\n-- \ngitgitgadget\n\n"},{"id":"394864","messageId":"5656e8590e9d4b42800e9ce53b10c45c63f03a1a.1586192395.git.gitgitgadget@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"[PATCH v4 15/15] commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag","fromName":"Garima Singh via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2020-04-06T16:59:55Z","receivedAt":"2020-04-06T17:00:16Z","isPatch":true,"sender":{"key":"garimasigit@gmail.com","avatar":null},"body":"From: Garima Singh <garima.singh@microsoft.com>\n\nAdd GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flag to the test setup suite\nin order to toggle writing Bloom filters when running any of the git tests.\nIf set to true, we will compute and write Bloom filters every time a test\ncalls `git commit-graph write`, as if the `--changed-paths` option was\npassed in.\n\nThe test suite passes when GIT_TEST_COMMIT_GRAPH and\nGIT_TEST_COMMIT_GRAPH_CHANGED_PATHS are enabled.\n\nHelped-by: Derrick Stolee <dstolee@microsoft.com>\nSigned-off-by: Garima Singh <garima.singh@microsoft.com>\n---\n builtin/commit-graph.c        | 3 ++-\n ci/run-build-and-tests.sh     | 1 +\n commit-graph.h                | 1 +\n t/README                      | 5 +++++\n t/t5318-commit-graph.sh       | 2 ++\n t/t5324-split-commit-graph.sh | 1 +\n 6 files changed, 12 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\nindex cacb5d04a80..59009837dc9 100644\n--- a/builtin/commit-graph.c\n+++ b/builtin/commit-graph.c\n@@ -171,7 +171,8 @@ static int graph_write(int argc, const char **argv)\n \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n \tif (opts.progress)\n \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n-\tif (opts.enable_changed_paths)\n+\tif (opts.enable_changed_paths ||\n+\t    git_env_bool(GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS, 0))\n \t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n \n \tread_replace_refs = 0;\ndiff --git a/ci/run-build-and-tests.sh b/ci/run-build-and-tests.sh\nindex 4df54c4efea..17e25aade96 100755\n--- a/ci/run-build-and-tests.sh\n+++ b/ci/run-build-and-tests.sh\n@@ -19,6 +19,7 @@ linux-gcc)\n \texport GIT_TEST_OE_SIZE=10\n \texport GIT_TEST_OE_DELTA_SIZE=5\n \texport GIT_TEST_COMMIT_GRAPH=1\n+\texport GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=1\n \texport GIT_TEST_MULTI_PACK_INDEX=1\n \texport GIT_TEST_ADD_I_USE_BUILTIN=1\n \tmake test\ndiff --git a/commit-graph.h b/commit-graph.h\nindex 8e7a8e0e5b2..8655d064c14 100644\n--- a/commit-graph.h\n+++ b/commit-graph.h\n@@ -9,6 +9,7 @@\n \n #define GIT_TEST_COMMIT_GRAPH \"GIT_TEST_COMMIT_GRAPH\"\n #define GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD \"GIT_TEST_COMMIT_GRAPH_DIE_ON_LOAD\"\n+#define GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS \"GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS\"\n \n struct commit;\n struct bloom_filter_settings;\ndiff --git a/t/README b/t/README\nindex 369e3a9ded8..4f53da53a15 100644\n--- a/t/README\n+++ b/t/README\n@@ -378,6 +378,11 @@ GIT_TEST_COMMIT_GRAPH=<boolean>, when true, forces the commit-graph to\n be written after every 'git commit' command, and overrides the\n 'core.commitGraph' setting to true.\n \n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=<boolean>, when true, forces\n+commit-graph write to compute and write changed path Bloom filters for\n+every 'git commit-graph write', as if the `--changed-paths` option was\n+passed in.\n+\n GIT_TEST_FSMONITOR=$PWD/t7519/fsmonitor-all exercises the fsmonitor\n code path for utilizing a file system monitor to speed up detecting\n new or changed files.\ndiff --git a/t/t5318-commit-graph.sh b/t/t5318-commit-graph.sh\nindex 9bf920ae171..18304a65e4d 100755\n--- a/t/t5318-commit-graph.sh\n+++ b/t/t5318-commit-graph.sh\n@@ -3,6 +3,8 @@\n test_description='commit graph'\n . ./test-lib.sh\n \n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n+\n test_expect_success 'setup full repo' '\n \tmkdir full &&\n \tcd \"$TRASH_DIRECTORY/full\" &&\ndiff --git a/t/t5324-split-commit-graph.sh b/t/t5324-split-commit-graph.sh\nindex 53b2e6b4555..d3f1f2c4a71 100755\n--- a/t/t5324-split-commit-graph.sh\n+++ b/t/t5324-split-commit-graph.sh\n@@ -4,6 +4,7 @@ test_description='split commit graph'\n . ./test-lib.sh\n \n GIT_TEST_COMMIT_GRAPH=0\n+GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS=0\n \n test_expect_success 'setup repo' '\n \tgit init &&\n-- \ngitgitgadget\n"},{"id":"395069","messageId":"ced7a793-47d9-8254-93fe-acdda8f12334@gmail.com","threadId":"52499","inReplyTo":"pull.497.v4.git.1586192395.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 00/15] Changed Paths Bloom Filters","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-04-08T15:51:14Z","receivedAt":"2020-04-08T15:51:21Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 4/6/2020 12:59 PM, Garima Singh via GitGitGadget wrote:\n> Hey! \n> \n> The commit graph feature brought in a lot of performance improvements across\n> multiple commands. However, file based history continues to be a performance\n> pain point, especially in large repositories. \n> \n> Adopting changed path Bloom filters has been discussed on the list before,\n> and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n> Derrick Stolee [1]. This series is based on Dr. Stolee's proof of concept in\n> [2]\n> \n> With the changes in this series, git users will be able to choose to write\n> Bloom filters to the commit-graph using the following command:\n> \n> 'git commit-graph write --changed-paths'\n> \n> Subsequent 'git log -- path' commands will use these computed Bloom filters\n> to decided which commits are worth exploring further to produce the history\n> of the provided path. \n\nI noticed Jakub was not CC'd on this email. Jakub: do you plan to re-review\nthe new version? Or are you satisfied with the resolutions to your comments?\n\nIs anyone else planning to review this series?\n\nI'm just wondering when we should take this series to cook in 'next' and\nstart building things on top of it, such as \"git blame\" or \"git log -L\"\nimprovements. While it cooks, any bugs or issues could be resolved with\npatches on top of this version. That would be my preference, anyway.\n\nWhat do you think, Junio?\n\nThanks,\n-Stolee\n\n"},{"id":"395089","messageId":"xmqq7dyppyv8.fsf@gitster.c.googlers.com","threadId":"52499","inReplyTo":"ced7a793-47d9-8254-93fe-acdda8f12334@gmail.com","subject":"Re: [PATCH v4 00/15] Changed Paths Bloom Filters","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-04-08T19:21:31Z","receivedAt":"2020-04-08T19:21:39Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> I noticed Jakub was not CC'd on this email. Jakub: do you plan to re-review\n> the new version? Or are you satisfied with the resolutions to your comments?\n> ...\n> What do you think, Junio?\n\nI was hoping that after Jakub's review, the new round was ready for\n'next' to be extended further by building on top as needed.  Of\ncourse the path-limited revision walk is one of the most important\npart of the entire system, so I'd welcome reviews from others, too.\n\nThanks.\n"},{"id":"395095","messageId":"CANQwDwf-Ok5oToyRoMXxTWOmPN0BhR=FVm3P-AoQxBqbozmc7A@mail.gmail.com","threadId":"52499","inReplyTo":"ced7a793-47d9-8254-93fe-acdda8f12334@gmail.com","subject":"Re: [PATCH v4 00/15] Changed Paths Bloom Filters","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2020-04-08T20:05:01Z","receivedAt":"2020-04-08T20:05:41Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Wed, 8 Apr 2020 at 17:51, Derrick Stolee <stolee@gmail.com> wrote:\n>\n> On 4/6/2020 12:59 PM, Garima Singh via GitGitGadget wrote:\n> > Hey!\n> >\n> > The commit graph feature brought in a lot of performance improvements across\n> > multiple commands. However, file based history continues to be a performance\n> > pain point, especially in large repositories.\n> >\n> > Adopting changed path Bloom filters has been discussed on the list before,\n> > and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n> > Derrick Stolee [1]. This series is based on Dr. Stolee's proof of concept in\n> > [2]\n> >\n> > With the changes in this series, git users will be able to choose to write\n> > Bloom filters to the commit-graph using the following command:\n> >\n> > 'git commit-graph write --changed-paths'\n> >\n> > Subsequent 'git log -- path' commands will use these computed Bloom filters\n> > to decided which commits are worth exploring further to produce the history\n> > of the provided path.\n>\n> I noticed Jakub was not CC'd on this email. Jakub: do you plan to re-review\n> the new version? Or are you satisfied with the resolutions to your comments?\n\nI am planning to re-review v4 of this series when I would have time,\nwhich means probably after Easter.\n\nI think if it handles endianness issues correctly, it should be ready.\n-- \nJakub Narębski\n"},{"id":"395305","messageId":"20200412203410.GA50412@syl.local","threadId":"52499","inReplyTo":"ced7a793-47d9-8254-93fe-acdda8f12334@gmail.com","subject":"Re: [PATCH v4 00/15] Changed Paths Bloom Filters","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2020-04-12T20:34:10Z","receivedAt":"2020-04-12T20:34:15Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"Hi Stolee,\n\nOn Wed, Apr 08, 2020 at 11:51:14AM -0400, Derrick Stolee wrote:\n> On 4/6/2020 12:59 PM, Garima Singh via GitGitGadget wrote:\n> > Hey!\n> >\n> > The commit graph feature brought in a lot of performance improvements across\n> > multiple commands. However, file based history continues to be a performance\n> > pain point, especially in large repositories.\n> >\n> > Adopting changed path Bloom filters has been discussed on the list before,\n> > and a prototype version was worked on by SZEDER Gábor, Jonathan Tan and Dr.\n> > Derrick Stolee [1]. This series is based on Dr. Stolee's proof of concept in\n> > [2]\n> >\n> > With the changes in this series, git users will be able to choose to write\n> > Bloom filters to the commit-graph using the following command:\n> >\n> > 'git commit-graph write --changed-paths'\n> >\n> > Subsequent 'git log -- path' commands will use these computed Bloom filters\n> > to decided which commits are worth exploring further to produce the history\n> > of the provided path.\n>\n> I noticed Jakub was not CC'd on this email. Jakub: do you plan to re-review\n> the new version? Or are you satisfied with the resolutions to your comments?\n>\n> Is anyone else planning to review this series?\n\nI feel horribly that I've had this patch series sitting in my review\nbacklog for months and haven't gotten to it yet, especially because I\nhave such an interest in these patches and know that much care was taken\nto prepare them.\n\nI read through these patches over some coffee today at a cursory level.\nThe high-level approach makes sense to me, and the implementation looks\nsolid. I think that anything that does come up (see below) can be\naddressed in 'next' rather than waiting longer on this series.\n\nFor what it's worth, I'm planning on starting to test this series in\nsome of our testing repositories at GitHub, and I'll report back on our\nexperience with some notes (and patches) should anything come up.\n\n> I'm just wondering when we should take this series to cook in 'next' and\n> start building things on top of it, such as \"git blame\" or \"git log -L\"\n> improvements. While it cooks, any bugs or issues could be resolved with\n> patches on top of this version. That would be my preference, anyway.\n\nThat would be my preference, too.\n\nI noticed a few small things (mostly a couple of typos and other very\nminor details). But, I'd much rather build on top of this series once it\nhas landed in 'next' than go to a fifth re-roll since there are many\npatches involved.\n\nI also noticed that you have already sent some patches in a separate\nseries that are based on this one, which would apply cleanly if this\nseries is merged into next.\n\nI figure that this will also be helpful as I send some patches about\nextra 'commit-graph write' options out of GitHub's fork, since they will\ninevitably create merge conflicts if we both are targeting 'next'. So,\nI figure that this approach will ease some maintainer burden ;-).\n\n>\n> What do you think, Junio?\n>\n> Thanks,\n> -Stolee\n\nThanks,\nTaylor\n"},{"id":"398876","messageId":"20200529085721.GA25128@szeder.dev","threadId":"52499","inReplyTo":"ff6b96aad1e2317d3ed36c2c8b419905dea84a83.1586192395.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 09/15] commit-graph: write Bloom filters to commit graph file","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-05-29T08:57:21Z","receivedAt":"2020-05-29T08:57:32Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Apr 06, 2020 at 04:59:49PM +0000, Garima Singh via GitGitGadget wrote:\n> From: Garima Singh <garima.singh@microsoft.com>\n> \n> Update the technical documentation for commit-graph-format with\n> the formats for the Bloom filter index (BIDX) and Bloom filter\n> data (BDAT) chunks. Write the computed Bloom filters information\n> to the commit graph file using this format.\n> \n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  .../technical/commit-graph-format.txt         |  30 +++++\n>  commit-graph.c                                | 113 +++++++++++++++++-\n>  commit-graph.h                                |   5 +\n>  3 files changed, 147 insertions(+), 1 deletion(-)\n> \n> diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n> index a4f17441aed..de56f9f1efd 100644\n> --- a/Documentation/technical/commit-graph-format.txt\n> +++ b/Documentation/technical/commit-graph-format.txt\n> @@ -17,6 +17,9 @@ metadata, including:\n>  - The parents of the commit, stored using positional references within\n>    the graph file.\n>  \n> +- The Bloom filter of the commit carrying the paths that were changed between\n> +  the commit and its first parent, if requested.\n> +\n>  These positional references are stored as unsigned 32-bit integers\n>  corresponding to the array position within the list of commit OIDs. Due\n>  to some special constants we use to track parents, we can store at most\n> @@ -93,6 +96,33 @@ CHUNK DATA:\n>        positions for the parents until reaching a value with the most-significant\n>        bit on. The other bits correspond to the position of the last parent.\n>  \n> +  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) (N * 4 bytes) [Optional]\n> +    * The ith entry, BIDX[i], stores the number of 8-byte word blocks in all\n\nThis is inconsistent with the implementation: according to the code in\none of the previous patches these entries are simple byte offsets, not\n8-byte word offsets, i.e. the combined size of all modified path\nBloom filters can be at most 2^32 bytes.\n\nThe commit-graph file can contain information about at most 2^31-1\ncommits.  This means that with that many commits each commit can have\na merely 2 byte Bloom filter on average.  When using 7 hashes we'd\nneed 10 bits per path, so in two bytes we could store only a single\npath.\n\nClearly, using 4 byte index entries significantly lowers the max\nnumber of commits that can be stored with modified path Bloom filters.\nIMO every new chunk must support at least 2^31-1 commits.\n\n> +      Bloom filters from commit 0 to commit i (inclusive) in lexicographic\n> +      order. The Bloom filter for the i-th commit spans from BIDX[i-1] to\n> +      BIDX[i] (plus header length), where BIDX[-1] is 0.\n> +    * The BIDX chunk is ignored if the BDAT chunk is not present.\n> +\n> +  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n> +    * It starts with header consisting of three unsigned 32-bit integers:\n> +      - Version of the hash algorithm being used. We currently only support\n> +\tvalue 1 which corresponds to the 32-bit version of the murmur3 hash\n> +\timplemented exactly as described in\n> +\thttps://en.wikipedia.org/wiki/MurmurHash#Algorithm and the double\n> +\thashing technique using seed values 0x293ae76f and 0x7e646e2 as\n> +\tdescribed in https://doi.org/10.1007/978-3-540-30494-4_26 \"Bloom Filters\n> +\tin Probabilistic Verification\"\n\nHow should double hashing compute the k hashes, i.e. using 64 bit or\n32 bit unsigned integer arithmetic?\n\nI'm puzzled that you link to this paper and still use double hashing.\n\nTwo of the contributions of that paper are that it points out some\nshortcomings of the double hashing scheme and provides a better\nalternative in the form of enhanced double hashing, which can cut the\nfalse positive rate in half.\n\nHowever, that paper considers the hashing scheme only in the context\nof one big Bloom filter.  I've found that when it comes to many small\nBloom filters then the k hashes produced by any double hashing variant\nare not independent enough, and \"standard\" double hashing fares the\nworst among them.  There are real repositories out there where double\nhashing has over an order of magnitude higher average false positive\nrate than enhanced double hashing.  Though that's not to say that\nenhanced double hashing is good...\n\nFor details on these issues see\n\n  https://public-inbox.org/git/20200529085038.26008-16-szeder.dev@gmail.com\n\n> +      - The number of times a path is hashed and hence the number of bit positions\n> +\t      that cumulatively determine whether a file is present in the commit.\n> +      - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n> +\t      contains 'n' entries, then the filter size is the minimum number of 64-bit\n> +\t      words that contain n*b bits.\n\nSince the ideal number of bits per element depends only on the number\nof hashes per path (k / ln(2) ≈ k * 10 / 7), why is this value stored\nin the commit-graph?\n\n> +    * The rest of the chunk is the concatenation of all the computed Bloom\n> +      filters for the commits in lexicographic order.\n> +    * Note: Commits with no changes or more than 512 changes have Bloom filters\n> +      of length zero.\n\nWhat does this \"Note:\" prefix mean in the file format specification?\n\nCan an implementation use a one byte Bloom filter with no bits set for\na commit with no changes?  Can an implementation still store a Bloom\nfilter for commits that modify more than 512 paths?\n\n> +    * The BDAT chunk is present if and only if BIDX is present.\n> +\n>    Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n>        This list of H-byte hashes describe a set of B commit-graph files that\n>        form a commit-graph chain. The graph position for the ith commit in this\n> diff --git a/commit-graph.c b/commit-graph.c\n> index 732c81fa1b2..a8b6b5cca5d 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n\n> @@ -1034,6 +1071,59 @@ static void write_graph_chunk_extra_edges(struct hashfile *f,\n>  \t}\n>  }\n>  \n> +static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n> +\t\t\t\t\t    struct write_commit_graph_context *ctx)\n> +{\n> +\tstruct commit **list = ctx->commits.list;\n> +\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n> +\tuint32_t cur_pos = 0;\n> +\tstruct progress *progress = NULL;\n> +\tint i = 0;\n> +\n> +\tif (ctx->report_progress)\n> +\t\tprogress = start_delayed_progress(\n> +\t\t\t_(\"Writing changed paths Bloom filters index\"),\n> +\t\t\tctx->commits.nr);\n> +\n> +\twhile (list < last) {\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tcur_pos += filter->len;\n\nGiven a sufficiently large number of commits with large enough Bloom\nfilters this will silently overflow.\n\n> +\t\tdisplay_progress(progress, ++i);\n> +\t\thashwrite_be32(f, cur_pos);\n> +\t\tlist++;\n> +\t}\n> +\n> +\tstop_progress(&progress);\n> +}\n"},{"id":"398879","messageId":"72cff41c-bb2e-5f87-5db6-d4e9ead25a47@gmail.com","threadId":"52499","inReplyTo":"20200529085721.GA25128@szeder.dev","subject":"Re: [PATCH v4 09/15] commit-graph: write Bloom filters to commit graph file","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-05-29T13:35:17Z","receivedAt":"2020-05-29T13:35:21Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 5/29/2020 4:57 AM, SZEDER Gábor wrote:\n> On Mon, Apr 06, 2020 at 04:59:49PM +0000, Garima Singh via GitGitGadget wrote:\n>> From: Garima Singh <garima.singh@microsoft.com>\n>>\n>> Update the technical documentation for commit-graph-format with\n>> the formats for the Bloom filter index (BIDX) and Bloom filter\n>> data (BDAT) chunks. Write the computed Bloom filters information\n>> to the commit graph file using this format.\n>>\n>> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n>> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n>> ---\n>>  .../technical/commit-graph-format.txt         |  30 +++++\n>>  commit-graph.c                                | 113 +++++++++++++++++-\n>>  commit-graph.h                                |   5 +\n>>  3 files changed, 147 insertions(+), 1 deletion(-)\n>>\n>> diff --git a/Documentation/technical/commit-graph-format.txt b/Documentation/technical/commit-graph-format.txt\n>> index a4f17441aed..de56f9f1efd 100644\n>> --- a/Documentation/technical/commit-graph-format.txt\n>> +++ b/Documentation/technical/commit-graph-format.txt\n>> @@ -17,6 +17,9 @@ metadata, including:\n>>  - The parents of the commit, stored using positional references within\n>>    the graph file.\n>>  \n>> +- The Bloom filter of the commit carrying the paths that were changed between\n>> +  the commit and its first parent, if requested.\n>> +\n>>  These positional references are stored as unsigned 32-bit integers\n>>  corresponding to the array position within the list of commit OIDs. Due\n>>  to some special constants we use to track parents, we can store at most\n>> @@ -93,6 +96,33 @@ CHUNK DATA:\n>>        positions for the parents until reaching a value with the most-significant\n>>        bit on. The other bits correspond to the position of the last parent.\n>>  \n>> +  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) (N * 4 bytes) [Optional]\n>> +    * The ith entry, BIDX[i], stores the number of 8-byte word blocks in all\n> \n> This is inconsistent with the implementation: according to the code in\n> one of the previous patches these entries are simple byte offsets, not\n> 8-byte word offsets, i.e. the combined size of all modified path\n> Bloom filters can be at most 2^32 bytes.\n\nThe documentation was fixed in 88093289cdc (Documentation: changed-path Bloom\nfilters use byte words, 2020-05-11).\n\n> The commit-graph file can contain information about at most 2^31-1\n> commits.  This means that with that many commits each commit can have\n> a merely 2 byte Bloom filter on average.  When using 7 hashes we'd\n> need 10 bits per path, so in two bytes we could store only a single\n> path.\n> \n> Clearly, using 4 byte index entries significantly lowers the max\n> number of commits that can be stored with modified path Bloom filters.\n\nThis is a good point, and certainly the reason for 8-byte multiples.\n\n> IMO every new chunk must support at least 2^31-1 commits.\n\nI'm not sure this is a valid requirement. Even extremely large repositories\n(that are created by actual use, not synthetic) are on the scale of 2^24\ncommits.\n\nYou are right that we should make the commit-graph write process more robust\nto reaching these limits. You point out that we have a new limit when these\nfilters are enabled.\n\nFor reference, the Windows OS repo has ~4.25 million commits and the\ncommit-graph file with changed-path Bloom filters is around 520mb. That's\nthe whole file size, and without the filters it's around 240mb, so the\nfilters are taking <300mb ~ 2^29 and we would need to grow the repo by 8x\nto hit this limit. That's not an unreasonable amount of growth, but is\nalso far enough away that we can handle it in time.\n\nThe incremental commit-graph can actually save us here (and is similar to\nhow we solved a scale issue in Azure Repos around the multi-pack-index):\nwe can refuse to merge layers of an incremental commit-graph if the\nchanged-path filters would exceed the size limit. Of course, the _first_\nwrite of such a commit-graph would need to be aware of this limit and\nplan for it in advance, but that's also a theoretical issue.\n\nI'm tracking some follow-up work [1] for the changed-path filters,\nincluding a way to limit the number of filters computed in one\n\"git commit-graph write\" process. I'll make note of your concerns here,\ntoo.\n\n[1] https://github.com/microsoft/git/issues/272\n\n>> +      Bloom filters from commit 0 to commit i (inclusive) in lexicographic\n>> +      order. The Bloom filter for the i-th commit spans from BIDX[i-1] to\n>> +      BIDX[i] (plus header length), where BIDX[-1] is 0.\n>> +    * The BIDX chunk is ignored if the BDAT chunk is not present.\n>> +\n>> +  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n>> +    * It starts with header consisting of three unsigned 32-bit integers:\n>> +      - Version of the hash algorithm being used. We currently only support\n>> +\tvalue 1 which corresponds to the 32-bit version of the murmur3 hash\n>> +\timplemented exactly as described in\n>> +\thttps://en.wikipedia.org/wiki/MurmurHash#Algorithm and the double\n>> +\thashing technique using seed values 0x293ae76f and 0x7e646e2 as\n>> +\tdescribed in https://doi.org/10.1007/978-3-540-30494-4_26 \"Bloom Filters\n>> +\tin Probabilistic Verification\"\n> \n> How should double hashing compute the k hashes, i.e. using 64 bit or\n> 32 bit unsigned integer arithmetic?\n> \n> I'm puzzled that you link to this paper and still use double hashing.\n> \n> Two of the contributions of that paper are that it points out some\n> shortcomings of the double hashing scheme and provides a better\n> alternative in the form of enhanced double hashing, which can cut the\n> false positive rate in half.\n> \n> However, that paper considers the hashing scheme only in the context\n> of one big Bloom filter.  I've found that when it comes to many small\n> Bloom filters then the k hashes produced by any double hashing variant\n> are not independent enough, and \"standard\" double hashing fares the\n> worst among them.  There are real repositories out there where double\n> hashing has over an order of magnitude higher average false positive\n> rate than enhanced double hashing.  Though that's not to say that\n> enhanced double hashing is good...\n> \n> For details on these issues see\n> \n>   https://public-inbox.org/git/20200529085038.26008-16-szeder.dev@gmail.com\n\nThat message includes very detailed experimental analysis, which is nice.\nWe will need to do some concrete side-by-side comparisons to see if there\nactually is a meaningful difference. (You may have already done this.)\n\n>> +      - The number of times a path is hashed and hence the number of bit positions\n>> +\t      that cumulatively determine whether a file is present in the commit.\n>> +      - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n>> +\t      contains 'n' entries, then the filter size is the minimum number of 64-bit\n>> +\t      words that contain n*b bits.\n> \n> Since the ideal number of bits per element depends only on the number\n> of hashes per path (k / ln(2) ≈ k * 10 / 7), why is this value stored\n> in the commit-graph?\n\nThe ideal number depends also on what false-positive rate you want. In a\nhypothetical future where we want to allow customization here, we want\nthe filters to be consistently sized across all filters.\n\n>> +    * The rest of the chunk is the concatenation of all the computed Bloom\n>> +      filters for the commits in lexicographic order.\n>> +    * Note: Commits with no changes or more than 512 changes have Bloom filters\n>> +      of length zero.\n> \n> What does this \"Note:\" prefix mean in the file format specification?\n> \n> Can an implementation use a one byte Bloom filter with no bits set for\n> a commit with no changes?  Can an implementation still store a Bloom\n> filter for commits that modify more than 512 paths?\n\nThis is currently due to a hard-coded value in the implementation. It's not a\nrequirement of the file format.\n\n>> +    * The BDAT chunk is present if and only if BIDX is present.\n>> +\n>>    Base Graphs List (ID: {'B', 'A', 'S', 'E'}) [Optional]\n>>        This list of H-byte hashes describe a set of B commit-graph files that\n>>        form a commit-graph chain. The graph position for the ith commit in this\n>> diff --git a/commit-graph.c b/commit-graph.c\n>> index 732c81fa1b2..a8b6b5cca5d 100644\n>> --- a/commit-graph.c\n>> +++ b/commit-graph.c\n> \n>> @@ -1034,6 +1071,59 @@ static void write_graph_chunk_extra_edges(struct hashfile *f,\n>>  \t}\n>>  }\n>>  \n>> +static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n>> +\t\t\t\t\t    struct write_commit_graph_context *ctx)\n>> +{\n>> +\tstruct commit **list = ctx->commits.list;\n>> +\tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n>> +\tuint32_t cur_pos = 0;\n>> +\tstruct progress *progress = NULL;\n>> +\tint i = 0;\n>> +\n>> +\tif (ctx->report_progress)\n>> +\t\tprogress = start_delayed_progress(\n>> +\t\t\t_(\"Writing changed paths Bloom filters index\"),\n>> +\t\t\tctx->commits.nr);\n>> +\n>> +\twhile (list < last) {\n>> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n>> +\t\tcur_pos += filter->len;\n> \n> Given a sufficiently large number of commits with large enough Bloom\n> filters this will silently overflow.\n\nWorth fixing, but we are not in a rush. I noted it in my GitHub issue.\n\nThanks,\n-Stolee\n\n"},{"id":"398956","messageId":"20200531172349.GA9990@szeder.dev","threadId":"52499","inReplyTo":"72cff41c-bb2e-5f87-5db6-d4e9ead25a47@gmail.com","subject":"Re: [PATCH v4 09/15] commit-graph: write Bloom filters to commit graph file","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-05-31T17:23:49Z","receivedAt":"2020-05-31T17:23:59Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Fri, May 29, 2020 at 09:35:17AM -0400, Derrick Stolee wrote:\n> >> +  Bloom Filter Index (ID: {'B', 'I', 'D', 'X'}) (N * 4 bytes) [Optional]\n> >> +    * The ith entry, BIDX[i], stores the number of 8-byte word blocks in all\n> > \n> > This is inconsistent with the implementation: according to the code in\n> > one of the previous patches these entries are simple byte offsets, not\n> > 8-byte word offsets, i.e. the combined size of all modified path\n> > Bloom filters can be at most 2^32 bytes.\n> \n> The documentation was fixed in 88093289cdc (Documentation: changed-path Bloom\n> filters use byte words, 2020-05-11).\n\nOh, good.  I'm waaay behind the curve and haven't seen this fix.  Even\nbetter, now I also noticed that two bugs I was about to report have\nbeen fixed already (though both fixes have minor flaws).\n\nOk, so at least the specs are consistent with the implementation.  I'm\nnot sure this was done in the right direction, though, because too\nsmall Bloom filters do hurt performance.\n\n> > Clearly, using 4 byte index entries significantly lowers the max\n> > number of commits that can be stored with modified path Bloom filters.\n> \n> This is a good point, and certainly the reason for 8-byte multiples.\n\nNote that Bloom filters with power-of-two number of bits have higher\nfalse positive probabilities when using some form of double hashing.\nWhen going for 8 byte blocks all commits modifying <= 12 paths\n(assuming 7 hashes per path) will have power-of-2 sized Bloom filters\n(64 or 128 bits), and that is a lot of commits.\n\n> The incremental commit-graph can actually save us here\n\nOh, I haven't thought of that\n\n> >> +      Bloom filters from commit 0 to commit i (inclusive) in lexicographic\n> >> +      order. The Bloom filter for the i-th commit spans from BIDX[i-1] to\n> >> +      BIDX[i] (plus header length), where BIDX[-1] is 0.\n> >> +    * The BIDX chunk is ignored if the BDAT chunk is not present.\n> >> +\n> >> +  Bloom Filter Data (ID: {'B', 'D', 'A', 'T'}) [Optional]\n> >> +    * It starts with header consisting of three unsigned 32-bit integers:\n> >> +      - Version of the hash algorithm being used. We currently only support\n> >> +\tvalue 1 which corresponds to the 32-bit version of the murmur3 hash\n> >> +\timplemented exactly as described in\n> >> +\thttps://en.wikipedia.org/wiki/MurmurHash#Algorithm and the double\n> >> +\thashing technique using seed values 0x293ae76f and 0x7e646e2 as\n> >> +\tdescribed in https://doi.org/10.1007/978-3-540-30494-4_26 \"Bloom Filters\n> >> +\tin Probabilistic Verification\"\n> > \n> > How should double hashing compute the k hashes, i.e. using 64 bit or\n> > 32 bit unsigned integer arithmetic?\n\nNote that this should be clarified in the specs.\n\n> >> +      - The number of times a path is hashed and hence the number of bit positions\n> >> +\t      that cumulatively determine whether a file is present in the commit.\n> >> +      - The minimum number of bits 'b' per entry in the Bloom filter. If the filter\n> >> +\t      contains 'n' entries, then the filter size is the minimum number of 64-bit\n> >> +\t      words that contain n*b bits.\n> > \n> > Since the ideal number of bits per element depends only on the number\n> > of hashes per path (k / ln(2) ≈ k * 10 / 7), why is this value stored\n> > in the commit-graph?\n> \n> The ideal number depends also on what false-positive rate you want.\n\nWell, yes, but indirectly:  according to Wikipedia :) the optimal\nnumber of hashes per element depends only on the desired false\nprobability, and the optimal number of bits per element depends only\non the number of hashes per element.\n\nSo storing the min number of bits per entry seems to be redundant.\n\n> In a\n> hypothetical future where we want to allow customization here, we want\n> the filters to be consistently sized across all filters.\n\nWouldn't customizing through the number of hashes be sufficient?\n\n> >> +    * Note: Commits with no changes or more than 512 changes have Bloom filters\n> >> +      of length zero.\n> > \n> > What does this \"Note:\" prefix mean in the file format specification?\n> > \n> > Can an implementation use a one byte Bloom filter with no bits set for\n> > a commit with no changes?  Can an implementation still store a Bloom\n> > filter for commits that modify more than 512 paths?\n> \n> This is currently due to a hard-coded value in the implementation. It's not a\n> requirement of the file format.\n\nShould an implementation detail like that be part of the specs?  It\nsure caused a bit of confusion here.\n\n"},{"id":"399267","messageId":"20200607222106.GB2898@szeder.dev","threadId":"52499","inReplyTo":"c8b86c383abdbbd31ba307eb7e79942ddde1b711.1586192395.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 11/15] commit-graph: add --changed-paths option to write subcommand","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-06-07T22:21:06Z","receivedAt":"2020-06-07T22:21:14Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Apr 06, 2020 at 04:59:51PM +0000, Garima Singh via GitGitGadget wrote:\n> From: Garima Singh <garima.singh@microsoft.com>\n> \n> Add --changed-paths option to git commit-graph write. This option will\n> allow users to compute information about the paths that have changed\n> between a commit and its first parent, and write it into the commit graph\n> file. If the option is passed to the write subcommand we set the\n> COMMIT_GRAPH_WRITE_BLOOM_FILTERS flag and pass it down to the\n> commit-graph logic.\n> \n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  Documentation/git-commit-graph.txt | 5 +++++\n>  builtin/commit-graph.c             | 9 +++++++--\n>  2 files changed, 12 insertions(+), 2 deletions(-)\n> \n> diff --git a/Documentation/git-commit-graph.txt b/Documentation/git-commit-graph.txt\n> index 28d1fee5053..f4b13c005b8 100644\n> --- a/Documentation/git-commit-graph.txt\n> +++ b/Documentation/git-commit-graph.txt\n> @@ -57,6 +57,11 @@ or `--stdin-packs`.)\n>  With the `--append` option, include all commits that are present in the\n>  existing commit-graph file.\n>  +\n> +With the `--changed-paths` option, compute and write information about the\n> +paths changed between a commit and it's first parent. This operation can\n> +take a while on large repositories. It provides significant performance gains\n> +for getting history of a directory or a file with `git log -- <path>`.\n\nSo 'git commit-graph write' only computes and writes changed path\nBloom filters if this option is specified.  Though not mentioned in\nthe documentation or in the commit message, the negated\n'--no-changed-paths' is supported as well, and it removes Bloom\nfilters from the commit-graph file.  All this is quite reasonable.\n\nHowever, the most important question is what happens when the\ncommit-graph file already contains Bloom filters and neither of these\noptions are specified on the command line.  This isn't mentioned in\nthe docs or in the commit message, either, but as it is implemented in\nthis patch (i.e. COMMIT_GRAPH_WRITE_BLOOM_FILTERS is not passed from\nthe builtin to the commit-graph logic) all those existing Bloom\nfilters are removed from the commit-graph.  Considering how expensive\nit was to compute those Bloom filters this might not be the most\ndesirable behaviour.\n\nThis is important, because 'git commit-graph write' is not the only\ncommand that writes the commit-graph file.  'git gc' does that by\ndefault, too, and will wipe out any modified path Bloom filters while\ndoing so.  Worse, the user doesn't even have to invoke 'git gc'\nmanually, because a lot of git commands invoke 'git gc --auto'.\n\n  $ git commit-graph write --reachable --changed-paths\n  $ ~/src/git/t/helper/test-tool read-graph |grep ^chunks\n  chunks: oid_fanout oid_lookup commit_metadata bloom_indexes bloom_data\n  $ git gc --quiet \n  $ ~/src/git/t/helper/test-tool read-graph |grep ^chunks\n  chunks: oid_fanout oid_lookup commit_metadata\n\nConsequently, if users want to use modified path Bloom filters, then\nthey should avoid gc, both manual and auto, or they'll have to\nre-generate the Bloom filters every once in a while.  That is\ndefinitely not the desired behaviour.\n\n\nNow compare this e.g. to the behaviour of 'git update-index\n--split-index' and '--untracked-cache': both of these options turn on\nfeatures that improve performance and write extra stuff to the index,\nand after they did so all subsequent git commands updating the index\nwill keep writing that extra stuff, including 'git update-index'\nitself even without those options, until it's finally invoked with the\ncorresponding '--no-...' option.  I particularly like how\n'--[no-]untracked-cache' and 'core.untrackedCache' work together and\nwarn when the given command line option goes against the configured\nvalue, and I think the command line options and configuration\nvariables controlling modified path Bloom filters should behave\nsimilarly.\n\n>  With the `--split` option, write the commit-graph as a chain of multiple\n>  commit-graph files stored in `<dir>/info/commit-graphs`. The new commits\n>  not already in the commit-graph are added in a new \"tip\" file. This file\n> diff --git a/builtin/commit-graph.c b/builtin/commit-graph.c\n> index d1ab6625f63..cacb5d04a80 100644\n> --- a/builtin/commit-graph.c\n> +++ b/builtin/commit-graph.c\n> @@ -9,7 +9,7 @@\n>  \n>  static char const * const builtin_commit_graph_usage[] = {\n>  \tN_(\"git commit-graph verify [--object-dir <objdir>] [--shallow] [--[no-]progress]\"),\n> -\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n> +\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n>  \tNULL\n>  };\n>  \n> @@ -19,7 +19,7 @@ static const char * const builtin_commit_graph_verify_usage[] = {\n>  };\n>  \n>  static const char * const builtin_commit_graph_write_usage[] = {\n> -\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--[no-]progress] <split options>\"),\n> +\tN_(\"git commit-graph write [--object-dir <objdir>] [--append|--split] [--reachable|--stdin-packs|--stdin-commits] [--changed-paths] [--[no-]progress] <split options>\"),\n>  \tNULL\n>  };\n>  \n> @@ -32,6 +32,7 @@ static struct opts_commit_graph {\n>  \tint split;\n>  \tint shallow;\n>  \tint progress;\n> +\tint enable_changed_paths;\n>  } opts;\n>  \n>  static struct object_directory *find_odb(struct repository *r,\n> @@ -135,6 +136,8 @@ static int graph_write(int argc, const char **argv)\n>  \t\t\tN_(\"start walk at commits listed by stdin\")),\n>  \t\tOPT_BOOL(0, \"append\", &opts.append,\n>  \t\t\tN_(\"include all commits already in the commit-graph file\")),\n> +\t\tOPT_BOOL(0, \"changed-paths\", &opts.enable_changed_paths,\n> +\t\t\tN_(\"enable computation for changed paths\")),\n>  \t\tOPT_BOOL(0, \"progress\", &opts.progress, N_(\"force progress reporting\")),\n>  \t\tOPT_BOOL(0, \"split\", &opts.split,\n>  \t\t\tN_(\"allow writing an incremental commit-graph file\")),\n> @@ -168,6 +171,8 @@ static int graph_write(int argc, const char **argv)\n>  \t\tflags |= COMMIT_GRAPH_WRITE_SPLIT;\n>  \tif (opts.progress)\n>  \t\tflags |= COMMIT_GRAPH_WRITE_PROGRESS;\n> +\tif (opts.enable_changed_paths)\n> +\t\tflags |= COMMIT_GRAPH_WRITE_BLOOM_FILTERS;\n>  \n>  \tread_replace_refs = 0;\n>  \todb = find_odb(the_repository, opts.obj_dir);\n> -- \n> gitgitgadget\n> \n"},{"id":"400123","messageId":"20200619140230.GB22200@szeder.dev","threadId":"52499","inReplyTo":"cc8022bdf82d0ada326ad546fdd7bb7801fc3675.1586192395.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 10/15] commit-graph: reuse existing Bloom filters during write","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-06-19T14:02:30Z","receivedAt":"2020-06-19T14:02:38Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Apr 06, 2020 at 04:59:50PM +0000, Garima Singh via GitGitGadget wrote:\n> From: Garima Singh <garima.singh@microsoft.com>\n> \n> Add logic to\n> a) parse Bloom filter information from the commit graph file and,\n> b) re-use existing Bloom filters.\n> \n> See Documentation/technical/commit-graph-format for the format in which\n> the Bloom filter information is written to the commit graph file.\n> \n> To read Bloom filter for a given commit with lexicographic position\n> 'i' we need to:\n> 1. Read BIDX[i] which essentially gives us the starting index in BDAT for\n>    filter of commit i+1. It is essentially the index past the end\n>    of the filter of commit i. It is called end_index in the code.\n> \n> 2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT\n>    for filter of commit i. It is called the start_index in the code.\n>    For the first commit, where i = 0, Bloom filter data starts at the\n>    beginning, just past the header in the BDAT chunk. Hence, start_index\n>    will be 0.\n> \n> 3. The length of the filter will be end_index - start_index, because\n>    BIDX[i] gives the cumulative 8-byte words including the ith\n>    commit's filter.\n> \n> We toggle whether Bloom filters should be recomputed based on the\n> compute_if_not_present flag.\n\nA very important question is not discussed here: when should we\nrecompute Bloom filters?\n\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  bloom.c               | 49 ++++++++++++++++++++++++++++++++++++++++++-\n>  bloom.h               |  4 +++-\n>  commit-graph.c        |  6 +++---\n>  t/helper/test-bloom.c |  2 +-\n>  4 files changed, 55 insertions(+), 6 deletions(-)\n> \n> diff --git a/bloom.c b/bloom.c\n> index a16eee92331..0f714dd76ae 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -4,6 +4,8 @@\n>  #include \"diffcore.h\"\n>  #include \"revision.h\"\n>  #include \"hashmap.h\"\n> +#include \"commit-graph.h\"\n> +#include \"commit.h\"\n>  \n>  define_commit_slab(bloom_filter_slab, struct bloom_filter);\n>  \n> @@ -26,6 +28,36 @@ static inline unsigned char get_bitmask(uint32_t pos)\n>  \treturn ((unsigned char)1) << (pos & (BITS_PER_WORD - 1));\n>  }\n>  \n> +static int load_bloom_filter_from_graph(struct commit_graph *g,\n> +\t\t\t\t   struct bloom_filter *filter,\n> +\t\t\t\t   struct commit *c)\n> +{\n> +\tuint32_t lex_pos, start_index, end_index;\n> +\n> +\twhile (c->graph_pos < g->num_commits_in_base)\n> +\t\tg = g->base_graph;\n> +\n> +\t/* The commit graph commit 'c' lives in doesn't carry bloom filters. */\n> +\tif (!g->chunk_bloom_indexes)\n> +\t\treturn 0;\n> +\n> +\tlex_pos = c->graph_pos - g->num_commits_in_base;\n> +\n> +\tend_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n> +\n> +\tif (lex_pos > 0)\n> +\t\tstart_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n> +\telse\n> +\t\tstart_index = 0;\n> +\n> +\tfilter->len = end_index - start_index;\n> +\tfilter->data = (unsigned char *)(g->chunk_bloom_data +\n> +\t\t\t\t\tsizeof(unsigned char) * start_index +\n> +\t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n> +\n> +\treturn 1;\n> +}\n> +\n>  /*\n>   * Calculate the murmur3 32-bit hash value for the given data\n>   * using the given seed.\n> @@ -127,7 +159,8 @@ void init_bloom_filters(void)\n>  }\n>  \n>  struct bloom_filter *get_bloom_filter(struct repository *r,\n> -\t\t\t\t      struct commit *c)\n> +\t\t\t\t      struct commit *c,\n> +\t\t\t\t\t  int compute_if_not_present)\n>  {\n>  \tstruct bloom_filter *filter;\n>  \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n\nThis line in the hunk context sets the default parameters with which\nthis process will compute any new changed path Bloom filters.\n\nNote that this is not the settings instance that eventually gets\nwritten to the header of the Bloom filters chunk:\nwrite_commit_graph_file() has its own 'struct bloom_filter_settings'\ninstance, and that's the one that goes into the chunk header.\n\n> @@ -140,6 +173,20 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \n>  \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n>  \n> +\tif (!filter->data) {\n> +\t\tload_commit_graph_info(r, c);\n> +\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH &&\n> +\t\t\tr->objects->commit_graph->chunk_bloom_indexes) {\n> +\t\t\tif (load_bloom_filter_from_graph(r->objects->commit_graph, filter, c))\n> +\t\t\t\treturn filter;\n> +\t\t\telse\n> +\t\t\t\treturn NULL;\n> +\t\t}\n> +\t}\n\nAnd in the above conditions we try to load the existing Bloom filter\nfor the given commit and return it as-is for reuse if it already\nexists, or go on to compute a new Bloom filter with the parameters set\nat the beginning of the function.\n\nUnfortunately, the parameters used to compute the now reused Bloom\nfilters are not checked anywhere.  In fact this writing process\nentirely ignores all parameters in the header of the existing Bloom\nfilters chunk, and simply replaces them with the default parameters\nhard-coded in write_commit_graph_file().  Consequently, we can end up\nwith Bloom filters computed with different parameters in the same\ncommit-graph file, which, in turn, can result in commits omitted from\nthe output of pathspec-limited revision walks.\n\nThe makeshift (there is no way to override those hard-coded defaults)\ntests below demonstrate this issue.\n\nThis issue raises a good couple of questions:\n\n  - What should we do when updating a commit-graph that was written\n    with different Bloom filter parameters than our hardcoded\n    defaults?\n\n    Reusing the exising Bloom filters is clearly wrong.  Throwing away\n    all existing Bloom filters and recomputing them with our defaults\n    parameters doesn't seem to be good option, because that's a\n    considerable amount of work, and the user might have a reason to\n    chose those parameters.\n\n  - What should we do when updating a commit-graph that was written\n    with different Bloom filter parameters than specified by the user\n    on the command line or in the config?\n\n    Wipe out the old Bloom filters and recompute with new parameters,\n    spending considerable time in bigger repositories?  Or stop with a\n    warning about the different parameters (maybe it's just a typo),\n    and require '--force'?\n    \n    Dunno, and we don't have such options and configuration yet\n    anyway.\n\n  - What about split commit-graphs?\n\n    When split commit-graphs were introduced there was not a single\n    chunk that had its own header.  Now the Bloom filters chunk does\n    have a header, which leads to other questions:\n\n    - Should that Bloom filters header be included in every split\n      commit-graph?\n\n      Not sure, but I suppose that having a header in each split\n      commit-graph file would make loading and parsing that chunk a\n      bit simpler, because all of them should be parsed the same way.\n      Anyway, I think the specs should be explicit about it.  But...:\n\n    - Should we allow different parameters in the Bloom filter chunks\n      in each split commit-graph?\n\n      The point of split commit-graphs is to avoid the overhead of\n      re-writing the whole commit-graph file every time new commits\n      are added, and it's crucial that both writing and merging split\n      commit-graph files are cheap.  However, split commit-graph files\n      using different Bloom filter parameters can't be merged without\n      recomputing those Bloom filters, making merging quite expensive.\n\n      So I don't think that it's a good idea to allow different Bloom\n      filter parameters in split commit-graphs.  But then perhaps it\n      would be better not to have a Bloom filter chunk header in all\n      split commit-graph files after all.\n\n    In any case, the last test below shows that the Bloom filter\n    parameters are only read from the header of the most recent split\n    commit-graph file.\n\n\n  ---  >8  ---\n\n#!/bin/sh\n\ntest_description='test'\n\n. ./test-lib.sh\n\ntest_expect_success 'yuckiest setup ever!' '\n\t(\n\t\tcd \"$GIT_BUILD_DIR\" &&\n\n\t\t# The number of hashes per path cannot be configured\n\t\t# at runtime, so build a dedicated git binary that\n\t\t# writes Bloom filters using only 6 hashes per path.\n\t\tsed -i -e \"/DEFAULT_BLOOM_FILTER_SETTINGS/ s/7/6/\" bloom.h &&\n\t\tmake -j4 git &&\n\t\tcp git git6 &&\n\n\t\t# Revert, rebuild.\n\t\tsed -i -e \"/DEFAULT_BLOOM_FILTER_SETTINGS/ s/6/7/\" bloom.h &&\n\t\tmake -j4 git\n\t) &&\n\tgit6=\"$GIT_BUILD_DIR\"/git6\n'\n\ntest_expect_success 'setup' '\n\t# We need a filename whose 7th hash maps to a different bit\n\t# position than any of its first 6 hashes in a 2-byte Bloom\n\t# filter.\n\tfile=File &&\n\n\ttest_tick &&\n\tgit commit --allow-empty -m initial &&\n\techo 1 >$file &&\n\tgit add $file &&\n\tgit commit -m one $file &&\n\techo 2 >$file &&\n\tgit commit -m two $file &&\n\n\tgit log --oneline -- $file >expect\n'\n\ntest_expect_success 'can read Bloom filters with different parameters' '\n\ttest_when_finished \"rm -rfv .git/objects/info/commit-graph*\" &&\n\n\t# Write a commit-graph with Bloom filters using only 6 hashes\n\t# per path.\n\t\"$git6\" commit-graph write --reachable --changed-paths &&\n\n\t# Try pathspec-limited revision walk with the git binary writing\n\t# Bloom filters using 7 hashes: it still works, because no matter\n\t# how many hashes it would use when writing the commit-graph, the\n\t# reader part respects the nr of hashes stored in the\n\t# commit-graph file.  So far so good.\n\tgit log --oneline $file >actual &&\n\ttest_cmp expect actual\n'\n\ntest_expect_failure 'commit-graph write does not reuse Bloom filters with different parameters' '\n\ttest_when_finished \"rm -rfv .git/objects/info/commit-graph*\" &&\n\n\t# Write a commit-graph with Bloom filters using only 6 hashes\n\t# per path for a subset of commits.\n\tgit rev-parse HEAD^ |\n\t\"$git6\" commit-graph write --stdin-commits --changed-paths &&\n\n\t# Add the rest of the commits to the commit-graph containing Bloom\n\t# filters using 6 hashes with a git version that writes Bloom\n\t# filters using 7 hashes.\n\t# Does it reuse the existing Bloom filters with 6 hashes?\n\tgit commit-graph write --reachable --changed-paths &&\n\n\t# Yes, it does, because these report different filter data,\n\t# even though both commits modified the same file.\n\ttest-tool bloom get_filter_for_commit $(git rev-parse HEAD^) &&\n\ttest-tool bloom get_filter_for_commit $(git rev-parse HEAD) &&\n\n\t# Furthermore, it updated the Bloom filter chunk header as well,\n\t# which now stores that all Bloom filters use 7 hashes.\n\t# Consequently, the first commit whose Bloom filter was written\n\t# with only 6 hashes falls victim of a false negative, and is\n\t# omitted from the output.\n\tgit log --oneline $file >actual &&\n\ttest_cmp expect actual\n'\n\ntest_expect_failure 'split commit-graphs and Bloom filters with different parameters' '\n\ttest_when_finished \"rm -rfv .git/objects/info/commit-graph*\" &&\n\n\tgit rev-parse HEAD^ |\n\t\"$git6\" commit-graph write --stdin-commits --changed-paths --split &&\n\n\tgit commit-graph write --reachable --changed-paths --split=no-merge &&\n\n\t# To make sure that I test what I want, i.e. two commit-graphs\n\t# with one commit in each.  (Though \"test-tool read-graph\" is\n\t# utterly oblivious to split commit graphs...)\n\ttest_line_count = 2 .git/objects/info/commit-graphs/commit-graph-chain &&\n\tverbose test \"$(test-tool read-graph |sed -n -e \"s/^num_commits: //p\")\" = 1 &&\n\n\ttest-tool bloom get_filter_for_commit $(git rev-parse HEAD^) &&\n\ttest-tool bloom get_filter_for_commit $(git rev-parse HEAD) &&\n\n\tgit log --oneline $file >actual &&\n\ttest_cmp expect actual\n'\n\ntest_done\n\n  ---  8<  ---\n\n> +\tif (filter->data || !compute_if_not_present)\n> +\t\treturn filter;\n> +\n>  \trepo_diff_setup(r, &diffopt);\n>  \tdiffopt.flags.recursive = 1;\n>  \tdiffopt.max_changes = max_changes;\n> diff --git a/bloom.h b/bloom.h\n> index 85ab8e9423d..760d7122374 100644\n> --- a/bloom.h\n> +++ b/bloom.h\n> @@ -32,6 +32,7 @@ struct bloom_filter_settings {\n>  \n>  #define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n>  #define BITS_PER_WORD 8\n> +#define BLOOMDATA_CHUNK_HEADER_SIZE 3 * sizeof(uint32_t)\n>  \n>  /*\n>   * A bloom_filter struct represents a data segment to\n> @@ -79,6 +80,7 @@ void add_key_to_filter(const struct bloom_key *key,\n>  void init_bloom_filters(void);\n>  \n>  struct bloom_filter *get_bloom_filter(struct repository *r,\n> -\t\t\t\t      struct commit *c);\n> +\t\t\t\t      struct commit *c,\n> +\t\t\t\t      int compute_if_not_present);\n>  \n>  #endif\n> \\ No newline at end of file\n> diff --git a/commit-graph.c b/commit-graph.c\n> index a8b6b5cca5d..77668629e27 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -1086,7 +1086,7 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n>  \t\t\tctx->commits.nr);\n>  \n>  \twhile (list < last) {\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n>  \t\tcur_pos += filter->len;\n>  \t\tdisplay_progress(progress, ++i);\n>  \t\thashwrite_be32(f, cur_pos);\n> @@ -1115,7 +1115,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n>  \thashwrite_be32(f, settings->bits_per_entry);\n>  \n>  \twhile (list < last) {\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n>  \t\tdisplay_progress(progress, ++i);\n>  \t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n>  \t\tlist++;\n> @@ -1296,7 +1296,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>  \n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>  \t\tstruct commit *c = sorted_commits[i];\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n>  \t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n>  \t\tdisplay_progress(progress, i + 1);\n>  \t}\n> diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n> index f18d1b722e1..ce412664ba9 100644\n> --- a/t/helper/test-bloom.c\n> +++ b/t/helper/test-bloom.c\n> @@ -39,7 +39,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n>  \tstruct bloom_filter *filter;\n>  \tsetup_git_directory();\n>  \tc = lookup_commit(the_repository, commit_oid);\n> -\tfilter = get_bloom_filter(the_repository, c);\n> +\tfilter = get_bloom_filter(the_repository, c, 1);\n>  \tprint_bloom_filter(filter);\n>  }\n>  \n> -- \n> gitgitgadget\n\n"},{"id":"400208","messageId":"xmqqmu4yhn3p.fsf@gitster.c.googlers.com","threadId":"52499","inReplyTo":"20200619140230.GB22200@szeder.dev","subject":"Re: [PATCH v4 10/15] commit-graph: reuse existing Bloom filters during write","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-06-19T19:28:10Z","receivedAt":"2020-06-19T19:28:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"SZEDER Gábor <szeder.dev@gmail.com> writes:\n\n> Note that this is not the settings instance that eventually gets\n> written to the header of the Bloom filters chunk:\n> write_commit_graph_file() has its own 'struct bloom_filter_settings'\n> instance, and that's the one that goes into the chunk header.\n> ...\n> Unfortunately, the parameters used to compute the now reused Bloom\n> filters are not checked anywhere.  In fact this writing process\n> entirely ignores all parameters in the header of the existing Bloom\n> filters chunk, and simply replaces them with the default parameters\n> hard-coded in write_commit_graph_file().  Consequently, we can end up\n> with Bloom filters computed with different parameters in the same\n> commit-graph file, which, in turn, can result in commits omitted from\n> the output of pathspec-limited revision walks.\n\nYeah, the whole design seems quite broken and as you said later,\nmixing other ingredients like split file would only make things\nworse X-<.\n\n> The makeshift (there is no way to override those hard-coded defaults)\n> tests below demonstrate this issue.\n>\n> This issue raises a good couple of questions:\n>\n>   - What should we do when updating a commit-graph that was written\n>     with different Bloom filter parameters than our hardcoded\n>     defaults?\n>\n>     Reusing the exising Bloom filters is clearly wrong.  Throwing away\n>     all existing Bloom filters and recomputing them with our defaults\n>     parameters doesn't seem to be good option, because that's a\n>     considerable amount of work, and the user might have a reason to\n>     chose those parameters.\n>\n>   - What should we do when updating a commit-graph that was written\n>     with different Bloom filter parameters than specified by the user\n>     on the command line or in the config?\n>\n>     Wipe out the old Bloom filters and recompute with new parameters,\n>     spending considerable time in bigger repositories?  Or stop with a\n>     warning about the different parameters (maybe it's just a typo),\n>     and require '--force'?\n>     \n>     Dunno, and we don't have such options and configuration yet\n>     anyway.\n>\n>   - What about split commit-graphs?\n>\n>     When split commit-graphs were introduced there was not a single\n>     chunk that had its own header.  Now the Bloom filters chunk does\n>     have a header, which leads to other questions:\n>\n>     - Should that Bloom filters header be included in every split\n>       commit-graph?\n>\n>       Not sure, but I suppose that having a header in each split\n>       commit-graph file would make loading and parsing that chunk a\n>       bit simpler, because all of them should be parsed the same way.\n>       Anyway, I think the specs should be explicit about it.  But...:\n>\n>     - Should we allow different parameters in the Bloom filter chunks\n>       in each split commit-graph?\n>\n>       The point of split commit-graphs is to avoid the overhead of\n>       re-writing the whole commit-graph file every time new commits\n>       are added, and it's crucial that both writing and merging split\n>       commit-graph files are cheap.  However, split commit-graph files\n>       using different Bloom filter parameters can't be merged without\n>       recomputing those Bloom filters, making merging quite expensive.\n>\n>       So I don't think that it's a good idea to allow different Bloom\n>       filter parameters in split commit-graphs.  But then perhaps it\n>       would be better not to have a Bloom filter chunk header in all\n>       split commit-graph files after all.\n>\n>     In any case, the last test below shows that the Bloom filter\n>     parameters are only read from the header of the most recent split\n>     commit-graph file.\n>\n>\n>   ---  >8  ---\n>\n> #!/bin/sh\n>\n> test_description='test'\n>\n> . ./test-lib.sh\n>\n> test_expect_success 'yuckiest setup ever!' '\n> \t(\n> \t\tcd \"$GIT_BUILD_DIR\" &&\n>\n> \t\t# The number of hashes per path cannot be configured\n> \t\t# at runtime, so build a dedicated git binary that\n> \t\t# writes Bloom filters using only 6 hashes per path.\n> \t\tsed -i -e \"/DEFAULT_BLOOM_FILTER_SETTINGS/ s/7/6/\" bloom.h &&\n> \t\tmake -j4 git &&\n> \t\tcp git git6 &&\n>\n> \t\t# Revert, rebuild.\n> \t\tsed -i -e \"/DEFAULT_BLOOM_FILTER_SETTINGS/ s/6/7/\" bloom.h &&\n> \t\tmake -j4 git\n> \t) &&\n> \tgit6=\"$GIT_BUILD_DIR\"/git6\n> '\n>\n> test_expect_success 'setup' '\n> \t# We need a filename whose 7th hash maps to a different bit\n> \t# position than any of its first 6 hashes in a 2-byte Bloom\n> \t# filter.\n> \tfile=File &&\n>\n> \ttest_tick &&\n> \tgit commit --allow-empty -m initial &&\n> \techo 1 >$file &&\n> \tgit add $file &&\n> \tgit commit -m one $file &&\n> \techo 2 >$file &&\n> \tgit commit -m two $file &&\n>\n> \tgit log --oneline -- $file >expect\n> '\n>\n> test_expect_success 'can read Bloom filters with different parameters' '\n> \ttest_when_finished \"rm -rfv .git/objects/info/commit-graph*\" &&\n>\n> \t# Write a commit-graph with Bloom filters using only 6 hashes\n> \t# per path.\n> \t\"$git6\" commit-graph write --reachable --changed-paths &&\n>\n> \t# Try pathspec-limited revision walk with the git binary writing\n> \t# Bloom filters using 7 hashes: it still works, because no matter\n> \t# how many hashes it would use when writing the commit-graph, the\n> \t# reader part respects the nr of hashes stored in the\n> \t# commit-graph file.  So far so good.\n> \tgit log --oneline $file >actual &&\n> \ttest_cmp expect actual\n> '\n>\n> test_expect_failure 'commit-graph write does not reuse Bloom filters with different parameters' '\n> \ttest_when_finished \"rm -rfv .git/objects/info/commit-graph*\" &&\n>\n> \t# Write a commit-graph with Bloom filters using only 6 hashes\n> \t# per path for a subset of commits.\n> \tgit rev-parse HEAD^ |\n> \t\"$git6\" commit-graph write --stdin-commits --changed-paths &&\n>\n> \t# Add the rest of the commits to the commit-graph containing Bloom\n> \t# filters using 6 hashes with a git version that writes Bloom\n> \t# filters using 7 hashes.\n> \t# Does it reuse the existing Bloom filters with 6 hashes?\n> \tgit commit-graph write --reachable --changed-paths &&\n>\n> \t# Yes, it does, because these report different filter data,\n> \t# even though both commits modified the same file.\n> \ttest-tool bloom get_filter_for_commit $(git rev-parse HEAD^) &&\n> \ttest-tool bloom get_filter_for_commit $(git rev-parse HEAD) &&\n>\n> \t# Furthermore, it updated the Bloom filter chunk header as well,\n> \t# which now stores that all Bloom filters use 7 hashes.\n> \t# Consequently, the first commit whose Bloom filter was written\n> \t# with only 6 hashes falls victim of a false negative, and is\n> \t# omitted from the output.\n> \tgit log --oneline $file >actual &&\n> \ttest_cmp expect actual\n> '\n>\n> test_expect_failure 'split commit-graphs and Bloom filters with different parameters' '\n> \ttest_when_finished \"rm -rfv .git/objects/info/commit-graph*\" &&\n>\n> \tgit rev-parse HEAD^ |\n> \t\"$git6\" commit-graph write --stdin-commits --changed-paths --split &&\n>\n> \tgit commit-graph write --reachable --changed-paths --split=no-merge &&\n>\n> \t# To make sure that I test what I want, i.e. two commit-graphs\n> \t# with one commit in each.  (Though \"test-tool read-graph\" is\n> \t# utterly oblivious to split commit graphs...)\n> \ttest_line_count = 2 .git/objects/info/commit-graphs/commit-graph-chain &&\n> \tverbose test \"$(test-tool read-graph |sed -n -e \"s/^num_commits: //p\")\" = 1 &&\n>\n> \ttest-tool bloom get_filter_for_commit $(git rev-parse HEAD^) &&\n> \ttest-tool bloom get_filter_for_commit $(git rev-parse HEAD) &&\n>\n> \tgit log --oneline $file >actual &&\n> \ttest_cmp expect actual\n> '\n>\n> test_done\n>\n>   ---  8<  ---\n>\n>> +\tif (filter->data || !compute_if_not_present)\n>> +\t\treturn filter;\n>> +\n>>  \trepo_diff_setup(r, &diffopt);\n>>  \tdiffopt.flags.recursive = 1;\n>>  \tdiffopt.max_changes = max_changes;\n>> diff --git a/bloom.h b/bloom.h\n>> index 85ab8e9423d..760d7122374 100644\n>> --- a/bloom.h\n>> +++ b/bloom.h\n>> @@ -32,6 +32,7 @@ struct bloom_filter_settings {\n>>  \n>>  #define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n>>  #define BITS_PER_WORD 8\n>> +#define BLOOMDATA_CHUNK_HEADER_SIZE 3 * sizeof(uint32_t)\n>>  \n>>  /*\n>>   * A bloom_filter struct represents a data segment to\n>> @@ -79,6 +80,7 @@ void add_key_to_filter(const struct bloom_key *key,\n>>  void init_bloom_filters(void);\n>>  \n>>  struct bloom_filter *get_bloom_filter(struct repository *r,\n>> -\t\t\t\t      struct commit *c);\n>> +\t\t\t\t      struct commit *c,\n>> +\t\t\t\t      int compute_if_not_present);\n>>  \n>>  #endif\n>> \\ No newline at end of file\n>> diff --git a/commit-graph.c b/commit-graph.c\n>> index a8b6b5cca5d..77668629e27 100644\n>> --- a/commit-graph.c\n>> +++ b/commit-graph.c\n>> @@ -1086,7 +1086,7 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n>>  \t\t\tctx->commits.nr);\n>>  \n>>  \twhile (list < last) {\n>> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n>> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n>>  \t\tcur_pos += filter->len;\n>>  \t\tdisplay_progress(progress, ++i);\n>>  \t\thashwrite_be32(f, cur_pos);\n>> @@ -1115,7 +1115,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n>>  \thashwrite_be32(f, settings->bits_per_entry);\n>>  \n>>  \twhile (list < last) {\n>> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n>> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n>>  \t\tdisplay_progress(progress, ++i);\n>>  \t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n>>  \t\tlist++;\n>> @@ -1296,7 +1296,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>>  \n>>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>>  \t\tstruct commit *c = sorted_commits[i];\n>> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n>> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n>>  \t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n>>  \t\tdisplay_progress(progress, i + 1);\n>>  \t}\n>> diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n>> index f18d1b722e1..ce412664ba9 100644\n>> --- a/t/helper/test-bloom.c\n>> +++ b/t/helper/test-bloom.c\n>> @@ -39,7 +39,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n>>  \tstruct bloom_filter *filter;\n>>  \tsetup_git_directory();\n>>  \tc = lookup_commit(the_repository, commit_oid);\n>> -\tfilter = get_bloom_filter(the_repository, c);\n>> +\tfilter = get_bloom_filter(the_repository, c, 1);\n>>  \tprint_bloom_filter(filter);\n>>  }\n>>  \n>> -- \n>> gitgitgadget\n"},{"id":"400655","messageId":"20200626063450.GL2898@szeder.dev","threadId":"52499","inReplyTo":"617f549ef259424658a84dd67a98685328f6b850.1586192395.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 12/15] revision.c: use Bloom filters to speed up path based revision walks","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-06-26T06:34:50Z","receivedAt":"2020-06-26T06:34:55Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Apr 06, 2020 at 04:59:52PM +0000, Garima Singh via GitGitGadget wrote:\n> +static void prepare_to_use_bloom_filter(struct rev_info *revs)\n> +{\n> +\tstruct pathspec_item *pi;\n> +\tchar *path_alloc = NULL;\n> +\tconst char *path;\n> +\tint last_index;\n> +\tint len;\n> +\n> +\tif (!revs->commits)\n> +\t    return;\n> +\n> +\trepo_parse_commit(revs->repo, revs->commits->item);\n> +\n> +\tif (!revs->repo->objects->commit_graph)\n> +\t\treturn;\n> +\n> +\trevs->bloom_filter_settings = revs->repo->objects->commit_graph->bloom_filter_settings;\n> +\tif (!revs->bloom_filter_settings)\n> +\t\treturn;\n> +\n> +\tpi = &revs->pruning.pathspec.items[0];\n> +\tlast_index = pi->len - 1;\n> +\n> +\t/* remove single trailing slash from path, if needed */\n> +\tif (pi->match[last_index] == '/') {\n> +\t    path_alloc = xstrdup(pi->match);\n> +\t    path_alloc[last_index] = '\\0';\n> +\t    path = path_alloc;\n\nfill_bloom_key() takes a length parameter, so there is no need to\nduplicate the path to be able to shorten it by one character to remove\nthat trailing '/'.\n\n> +\t} else\n> +\t    path = pi->match;\n> +\n> +\tlen = strlen(path);\n\n'struct pathspec_item's 'len' field already contains the length of the\npath, so there is no need for this strlen().\n\n> +\n> +\trevs->bloom_key = xmalloc(sizeof(struct bloom_key));\n> +\tfill_bloom_key(path, len, revs->bloom_key, revs->bloom_filter_settings);\n> +\n> +\tfree(path_alloc);\n> +}\n\n> @@ -3362,6 +3440,8 @@ int prepare_revision_walk(struct rev_info *revs)\n>  \t\t\t\t       FOR_EACH_OBJECT_PROMISOR_ONLY);\n>  \t}\n>  \n> +\tif (revs->pruning.pathspec.nr == 1 && !revs->reflog_info)\n> +\t\tprepare_to_use_bloom_filter(revs);\n>  \tif (revs->no_walk != REVISION_WALK_NO_WALK_UNSORTED)\n>  \t\tcommit_list_sort_by_date(&revs->commits);\n>  \tif (revs->no_walk)\n                return 0;\n        if (revs->limited) {\n                if (limit_list(revs) < 0)\n                        return -1;\n\nI extended the hunk context a bit to show that\nprepare_to_use_bloom_filter() is called before limit_list().  This is\nimportant, because specifying exclude revs and pathspecs, i.e.  'git\nlog ^v1.2.3 -- dir/file' does perform a lot of diffs in limit_list(),\nand this way we can take advantage of Bloom filters even in this case.\n\n> @@ -3379,6 +3459,7 @@ int prepare_revision_walk(struct rev_info *revs)\n>  \t\tsimplify_merges(revs);\n>  \tif (revs->children.name)\n>  \t\tset_children(revs);\n> +\n>  \treturn 0;\n>  }\n>  \n> diff --git a/revision.h b/revision.h\n> index 475f048fb61..7c026fe41fc 100644\n> --- a/revision.h\n> +++ b/revision.h\n> @@ -56,6 +56,8 @@ struct repository;\n>  struct rev_info;\n>  struct string_list;\n>  struct saved_parents;\n> +struct bloom_key;\n> +struct bloom_filter_settings;\n>  define_shared_commit_slab(revision_sources, char *);\n>  \n>  struct rev_cmdline_info {\n> @@ -291,6 +293,15 @@ struct rev_info {\n>  \tstruct revision_sources *sources;\n>  \n>  \tstruct topo_walk_info *topo_walk_info;\n> +\n> +\t/* Commit graph bloom filter fields */\n> +\t/* The bloom filter key for the pathspec */\n> +\tstruct bloom_key *bloom_key;\n> +\t/*\n> +\t * The bloom filter settings used to generate the key.\n> +\t * This is loaded from the commit-graph being used.\n> +\t */\n> +\tstruct bloom_filter_settings *bloom_filter_settings;\n>  };\n>  \n>  int ref_excluded(struct string_list *, const char *path);\n> -- \n> gitgitgadget\n> \n"},{"id":"400728","messageId":"20200627155336.GB11341@szeder.dev","threadId":"52499","inReplyTo":"8304c2975207ee847c6709abd71efee918fc4142.1586192395.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 04/15] bloom.c: core Bloom filter implementation for changed paths.","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-06-27T15:53:36Z","receivedAt":"2020-06-27T15:53:42Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Apr 06, 2020 at 04:59:44PM +0000, Garima Singh via GitGitGadget wrote:\n> From: Garima Singh <garima.singh@microsoft.com>\n> \n> Add the core implementation for computing Bloom filters for\n> the paths changed between a commit and it's first parent.\n> \n> We fill the Bloom filters as (const char *data, int len) pairs\n> as `struct bloom_filters\" within a commit slab.\n> \n> Filters for commits with no changes and more than 512 changes,\n> is represented with a filter of length zero. There is no gain\n> in distinguishing between a computed filter of length zero for\n> a commit with no changes, and an uncomputed filter for new commits\n> or for commits with more than 512 changes. The effect on\n> `git log -- path` is the same in both cases. We will fall back to\n> the normal diffing algorithm when we can't benefit from the\n> existence of Bloom filters.\n> \n> Helped-by: Jeff King <peff@peff.net>\n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Reviewed-by: Jakub Narębski <jnareb@gmail.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  bloom.c               | 97 +++++++++++++++++++++++++++++++++++++++++++\n>  bloom.h               |  8 ++++\n>  t/helper/test-bloom.c | 20 +++++++++\n>  t/t0095-bloom.sh      | 47 +++++++++++++++++++++\n>  4 files changed, 172 insertions(+)\n> \n> diff --git a/bloom.c b/bloom.c\n> index 888b67f1ea6..881a9841ede 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -1,5 +1,18 @@\n>  #include \"git-compat-util.h\"\n>  #include \"bloom.h\"\n> +#include \"diff.h\"\n> +#include \"diffcore.h\"\n> +#include \"revision.h\"\n> +#include \"hashmap.h\"\n> +\n> +define_commit_slab(bloom_filter_slab, struct bloom_filter);\n\nSo here we define a commit slab for modified path Bloom filters, ...\n\n> +struct bloom_filter_slab bloom_filters;\n> +\n> +struct pathmap_hash_entry {\n> +    struct hashmap_entry entry;\n> +    const char path[FLEX_ARRAY];\n> +};\n>  \n>  static uint32_t rotate_left(uint32_t value, int32_t count)\n>  {\n> @@ -107,3 +120,87 @@ void add_key_to_filter(const struct bloom_key *key,\n>  \t\tfilter->data[block_pos] |= get_bitmask(hash_mod);\n>  \t}\n>  }\n> +\n> +void init_bloom_filters(void)\n> +{\n> +\tinit_bloom_filter_slab(&bloom_filters);\n\n... here initialize the slab ...\n\n> +}\n> +\n> +struct bloom_filter *get_bloom_filter(struct repository *r,\n> +\t\t\t\t      struct commit *c)\n> +{\n> +\tstruct bloom_filter *filter;\n> +\tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n> +\tint i;\n> +\tstruct diff_options diffopt;\n> +\n> +\tif (bloom_filters.slab_size == 0)\n> +\t\treturn NULL;\n> +\n> +\tfilter = bloom_filter_slab_at(&bloom_filters, c);\n\n... allocate an entry in the slab ...\n\n> +\n> +\trepo_diff_setup(r, &diffopt);\n> +\tdiffopt.flags.recursive = 1;\n> +\tdiff_setup_done(&diffopt);\n> +\n> +\tif (c->parents)\n> +\t\tdiff_tree_oid(&c->parents->item->object.oid, &c->object.oid, \"\", &diffopt);\n> +\telse\n> +\t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n> +\tdiffcore_std(&diffopt);\n> +\n> +\tif (diff_queued_diff.nr <= 512) {\n> +\t\tstruct hashmap pathmap;\n> +\t\tstruct pathmap_hash_entry *e;\n> +\t\tstruct hashmap_iter iter;\n> +\t\thashmap_init(&pathmap, NULL, NULL, 0);\n> +\n> +\t\tfor (i = 0; i < diff_queued_diff.nr; i++) {\n> +\t\t\tconst char *path = diff_queued_diff.queue[i]->two->path;\n> +\n> +\t\t\t/*\n> +\t\t\t* Add each leading directory of the changed file, i.e. for\n> +\t\t\t* 'dir/subdir/file' add 'dir' and 'dir/subdir' as well, so\n> +\t\t\t* the Bloom filter could be used to speed up commands like\n> +\t\t\t* 'git log dir/subdir', too.\n> +\t\t\t*\n> +\t\t\t* Note that directories are added without the trailing '/'.\n> +\t\t\t*/\n> +\t\t\tdo {\n> +\t\t\t\tchar *last_slash = strrchr(path, '/');\n> +\n> +\t\t\t\tFLEX_ALLOC_STR(e, path, path);\n> +\t\t\t\thashmap_entry_init(&e->entry, strhash(path));\n> +\t\t\t\thashmap_add(&pathmap, &e->entry);\n> +\n> +\t\t\t\tif (!last_slash)\n> +\t\t\t\t\tlast_slash = (char*)path;\n> +\t\t\t\t*last_slash = '\\0';\n> +\n> +\t\t\t} while (*path);\n> +\n> +\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n> +\t\t}\n> +\n> +\t\tfilter->len = (hashmap_get_size(&pathmap) * settings.bits_per_entry + BITS_PER_WORD - 1) / BITS_PER_WORD;\n> +\t\tfilter->data = xcalloc(filter->len, sizeof(unsigned char));\n\n... and here we fill the slab with data, including a memory allocation\nfor each slab entry.\n\nWhat is missing in this patch or in any of the followup patches is a\nplace where we clear the slab and the additional memory attached to\nit.\n\n> +\n> +\t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n> +\t\t\tstruct bloom_key key;\n> +\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, &settings);\n> +\t\t\tadd_key_to_filter(&key, filter, &settings);\n> +\t\t}\n> +\n> +\t\thashmap_free_entries(&pathmap, struct pathmap_hash_entry, entry);\n> +\t} else {\n> +\t\tfor (i = 0; i < diff_queued_diff.nr; i++)\n> +\t\t\tdiff_free_filepair(diff_queued_diff.queue[i]);\n> +\t\tfilter->data = NULL;\n> +\t\tfilter->len = 0;\n> +\t}\n> +\n> +\tfree(diff_queued_diff.queue);\n> +\tDIFF_QUEUE_CLEAR(&diff_queued_diff);\n> +\n> +\treturn filter;\n> +}\n> diff --git a/bloom.h b/bloom.h\n> index b9ce422ca2d..85ab8e9423d 100644\n> --- a/bloom.h\n> +++ b/bloom.h\n> @@ -1,6 +1,9 @@\n>  #ifndef BLOOM_H\n>  #define BLOOM_H\n>  \n> +struct commit;\n> +struct repository;\n> +\n>  struct bloom_filter_settings {\n>  \t/*\n>  \t * The version of the hashing technique being used.\n> @@ -73,4 +76,9 @@ void add_key_to_filter(const struct bloom_key *key,\n>  \t\t\t\t\t   struct bloom_filter *filter,\n>  \t\t\t\t\t   const struct bloom_filter_settings *settings);\n>  \n> +void init_bloom_filters(void);\n> +\n> +struct bloom_filter *get_bloom_filter(struct repository *r,\n> +\t\t\t\t      struct commit *c);\n> +\n>  #endif\n> \\ No newline at end of file\n> diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n> index 20460cde775..f18d1b722e1 100644\n> --- a/t/helper/test-bloom.c\n> +++ b/t/helper/test-bloom.c\n> @@ -1,6 +1,7 @@\n>  #include \"git-compat-util.h\"\n>  #include \"bloom.h\"\n>  #include \"test-tool.h\"\n> +#include \"commit.h\"\n>  \n>  struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n>  \n> @@ -32,6 +33,16 @@ static void print_bloom_filter(struct bloom_filter *filter) {\n>  \tprintf(\"\\n\");\n>  }\n>  \n> +static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n> +{\n> +\tstruct commit *c;\n> +\tstruct bloom_filter *filter;\n> +\tsetup_git_directory();\n> +\tc = lookup_commit(the_repository, commit_oid);\n> +\tfilter = get_bloom_filter(the_repository, c);\n> +\tprint_bloom_filter(filter);\n> +}\n> +\n>  int cmd__bloom(int argc, const char **argv)\n>  {\n>  \tif (!strcmp(argv[1], \"get_murmur3\")) {\n> @@ -57,5 +68,14 @@ int cmd__bloom(int argc, const char **argv)\n>  \t\tprint_bloom_filter(&filter);\n>  \t}\n>  \n> +    if (!strcmp(argv[1], \"get_filter_for_commit\")) {\n> +\t\tstruct object_id oid;\n> +\t\tconst char *end;\n> +\t\tif (parse_oid_hex(argv[2], &oid, &end))\n> +\t\t\tdie(\"cannot parse oid '%s'\", argv[2]);\n> +\t\tinit_bloom_filters();\n> +\t\tget_bloom_filter_for_commit(&oid);\n> +\t}\n> +\n>  \treturn 0;\n>  }\n> \\ No newline at end of file\n> diff --git a/t/t0095-bloom.sh b/t/t0095-bloom.sh\n> index 36a086c7c60..8f9eef116dc 100755\n> --- a/t/t0095-bloom.sh\n> +++ b/t/t0095-bloom.sh\n> @@ -67,4 +67,51 @@ test_expect_success 'compute bloom key for test string 2' '\n>  \ttest_cmp expect actual\n>  '\n>  \n> +test_expect_success 'get bloom filters for commit with no changes' '\n> +\tgit init &&\n> +\tgit commit --allow-empty -m \"c0\" &&\n> +\tcat >expect <<-\\EOF &&\n> +\tFilter_Length:0\n> +\tFilter_Data:\n> +\tEOF\n> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'get bloom filter for commit with 10 changes' '\n> +\trm actual &&\n> +\trm expect &&\n> +\tmkdir smallDir &&\n> +\tfor i in $(test_seq 0 9)\n> +\tdo\n> +\t\techo $i >smallDir/$i\n> +\tdone &&\n> +\tgit add smallDir &&\n> +\tgit commit -m \"commit with 10 changes\" &&\n> +\tcat >expect <<-\\EOF &&\n> +\tFilter_Length:25\n> +\tFilter_Data:82|a0|65|47|0c|92|90|c0|a1|40|02|a0|e2|40|e0|04|0a|9a|66|cf|80|19|85|42|23|\n> +\tEOF\n> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success EXPENSIVE 'get bloom filter for commit with 513 changes' '\n> +\trm actual &&\n> +\trm expect &&\n> +\tmkdir bigDir &&\n> +\tfor i in $(test_seq 0 512)\n> +\tdo\n> +\t\techo $i >bigDir/$i\n> +\tdone &&\n> +\tgit add bigDir &&\n> +\tgit commit -m \"commit with 513 changes\" &&\n> +\tcat >expect <<-\\EOF &&\n> +\tFilter_Length:0\n> +\tFilter_Data:\n> +\tEOF\n> +\ttest-tool bloom get_filter_for_commit \"$(git rev-parse HEAD)\" >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n>  test_done\n> \\ No newline at end of file\n> -- \n> gitgitgadget\n> \n"},{"id":"401252","messageId":"20200709170003.3020-1-szeder.dev@gmail.com","threadId":"52499","inReplyTo":"ff6b96aad1e2317d3ed36c2c8b419905dea84a83.1586192395.git.gitgitgadget@gmail.com","subject":"[PATCH] commit-graph: fix \"Writing out commit graph\" progress counter","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-07-09T17:00:03Z","receivedAt":"2020-07-09T17:00:13Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"76ffbca71a (commit-graph: write Bloom filters to commit graph file,\n2020-04-06) added two delayed progress lines to writing the Bloom\nfilter index and data chunk.  This is wrong, because a single common\nprogress is used while writing all chunks, which is not updated while\nwriting these two new chunks, resulting in incomplete-looking \"done\"\nlines:\n\n  Expanding reachable commits in commit graph: 888679, done.\n  Computing commit changed paths Bloom filters: 100% (888678/888678), done.\n  Writing out commit graph in 6 passes:  66% (3554712/5332068), done.\n\nUse the common 'struct progress' instance while writing the Bloom\nfilter chunks as well.\n\nSigned-off-by: SZEDER Gábor <szeder.dev@gmail.com>\n---\n commit-graph.c | 22 ++--------------------\n 1 file changed, 2 insertions(+), 20 deletions(-)\n\ndiff --git a/commit-graph.c b/commit-graph.c\nindex aaf3327ede..65cf32637c 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1086,23 +1086,14 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n \tstruct commit **list = ctx->commits.list;\n \tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n \tuint32_t cur_pos = 0;\n-\tstruct progress *progress = NULL;\n-\tint i = 0;\n-\n-\tif (ctx->report_progress)\n-\t\tprogress = start_delayed_progress(\n-\t\t\t_(\"Writing changed paths Bloom filters index\"),\n-\t\t\tctx->commits.nr);\n \n \twhile (list < last) {\n \t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n \t\tcur_pos += filter->len;\n-\t\tdisplay_progress(progress, ++i);\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n \t\thashwrite_be32(f, cur_pos);\n \t\tlist++;\n \t}\n-\n-\tstop_progress(&progress);\n }\n \n static void write_graph_chunk_bloom_data(struct hashfile *f,\n@@ -1111,13 +1102,6 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n {\n \tstruct commit **list = ctx->commits.list;\n \tstruct commit **last = ctx->commits.list + ctx->commits.nr;\n-\tstruct progress *progress = NULL;\n-\tint i = 0;\n-\n-\tif (ctx->report_progress)\n-\t\tprogress = start_delayed_progress(\n-\t\t\t_(\"Writing changed paths Bloom filters data\"),\n-\t\t\tctx->commits.nr);\n \n \thashwrite_be32(f, settings->hash_version);\n \thashwrite_be32(f, settings->num_hashes);\n@@ -1125,12 +1109,10 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n \n \twhile (list < last) {\n \t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n-\t\tdisplay_progress(progress, ++i);\n+\t\tdisplay_progress(ctx->progress, ++ctx->progress_cnt);\n \t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n \t\tlist++;\n \t}\n-\n-\tstop_progress(&progress);\n }\n \n static int oid_compare(const void *_a, const void *_b)\n-- \n2.27.0.547.g4ba2d26563\n\n"},{"id":"401256","messageId":"1b683731-776b-0058-5744-094091c7db4d@gmail.com","threadId":"52499","inReplyTo":"20200709170003.3020-1-szeder.dev@gmail.com","subject":"Re: [PATCH] commit-graph: fix \"Writing out commit graph\" progress counter","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-07-09T18:01:57Z","receivedAt":"2020-07-09T18:02:00Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/9/2020 1:00 PM, SZEDER Gábor wrote:\n> 76ffbca71a (commit-graph: write Bloom filters to commit graph file,\n> 2020-04-06) added two delayed progress lines to writing the Bloom\n> filter index and data chunk.  This is wrong, because a single common\n> progress is used while writing all chunks, which is not updated while\n> writing these two new chunks, resulting in incomplete-looking \"done\"\n> lines:\n> \n>   Expanding reachable commits in commit graph: 888679, done.\n>   Computing commit changed paths Bloom filters: 100% (888678/888678), done.\n>   Writing out commit graph in 6 passes:  66% (3554712/5332068), done.\n> \n> Use the common 'struct progress' instance while writing the Bloom\n> filter chunks as well.\n\nThanks for finding this. It's a clearly correct way to go,\nand is one of the things that did not get updated properly\nbetween the old prototype when applying it on the new code\nthat included this ctx->progress pattern.\n\nJunio: head's up that this will conflict with the final patch\nin ds/maintenance. I'll remove my edits to these methods in\nmy v2 to make that merge a bit easier.\n\nThanks,\n-Stolee\n"},{"id":"401258","messageId":"3484f1fa-3231-f457-9951-a3aa0a3f7c7d@gmail.com","threadId":"52499","inReplyTo":"1b683731-776b-0058-5744-094091c7db4d@gmail.com","subject":"Re: [PATCH] commit-graph: fix \"Writing out commit graph\" progress counter","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-07-09T18:20:21Z","receivedAt":"2020-07-09T18:20:25Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/9/2020 2:01 PM, Derrick Stolee wrote:\n> Junio: head's up that this will conflict with the final patch\n> in ds/maintenance. I'll remove my edits to these methods in\n> my v2 to make that merge a bit easier.\nOr, I'm getting confused, because I changed start_progress()\ncalls in midx.c, not commit-graph.c. Please ignore my scattered\nbrain.\n\nThanks,\n-Stolee\n"},{"id":"402166","messageId":"20200727213312.GP2898@szeder.dev","threadId":"52499","inReplyTo":"cc8022bdf82d0ada326ad546fdd7bb7801fc3675.1586192395.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 10/15] commit-graph: reuse existing Bloom filters during write","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-07-27T21:33:12Z","receivedAt":"2020-07-27T21:33:19Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Apr 06, 2020 at 04:59:50PM +0000, Garima Singh via GitGitGadget wrote:\n> From: Garima Singh <garima.singh@microsoft.com>\n> \n> Add logic to\n> a) parse Bloom filter information from the commit graph file and,\n> b) re-use existing Bloom filters.\n> \n> See Documentation/technical/commit-graph-format for the format in which\n> the Bloom filter information is written to the commit graph file.\n> \n> To read Bloom filter for a given commit with lexicographic position\n> 'i' we need to:\n> 1. Read BIDX[i] which essentially gives us the starting index in BDAT for\n>    filter of commit i+1. It is essentially the index past the end\n>    of the filter of commit i. It is called end_index in the code.\n> \n> 2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT\n>    for filter of commit i. It is called the start_index in the code.\n>    For the first commit, where i = 0, Bloom filter data starts at the\n>    beginning, just past the header in the BDAT chunk. Hence, start_index\n>    will be 0.\n> \n> 3. The length of the filter will be end_index - start_index, because\n>    BIDX[i] gives the cumulative 8-byte words including the ith\n>    commit's filter.\n> \n> We toggle whether Bloom filters should be recomputed based on the\n> compute_if_not_present flag.\n> \n> Helped-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  bloom.c               | 49 ++++++++++++++++++++++++++++++++++++++++++-\n>  bloom.h               |  4 +++-\n>  commit-graph.c        |  6 +++---\n>  t/helper/test-bloom.c |  2 +-\n>  4 files changed, 55 insertions(+), 6 deletions(-)\n> \n> diff --git a/bloom.c b/bloom.c\n> index a16eee92331..0f714dd76ae 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -4,6 +4,8 @@\n>  #include \"diffcore.h\"\n>  #include \"revision.h\"\n>  #include \"hashmap.h\"\n> +#include \"commit-graph.h\"\n> +#include \"commit.h\"\n>  \n>  define_commit_slab(bloom_filter_slab, struct bloom_filter);\n>  \n> @@ -26,6 +28,36 @@ static inline unsigned char get_bitmask(uint32_t pos)\n>  \treturn ((unsigned char)1) << (pos & (BITS_PER_WORD - 1));\n>  }\n>  \n> +static int load_bloom_filter_from_graph(struct commit_graph *g,\n> +\t\t\t\t   struct bloom_filter *filter,\n> +\t\t\t\t   struct commit *c)\n> +{\n> +\tuint32_t lex_pos, start_index, end_index;\n> +\n> +\twhile (c->graph_pos < g->num_commits_in_base)\n> +\t\tg = g->base_graph;\n> +\n> +\t/* The commit graph commit 'c' lives in doesn't carry bloom filters. */\n> +\tif (!g->chunk_bloom_indexes)\n> +\t\treturn 0;\n> +\n> +\tlex_pos = c->graph_pos - g->num_commits_in_base;\n> +\n> +\tend_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);\n\nLet's suppose that we encounter a bogus commit-graph file.  This would\nthen segfault if 'lex_pos' were to point past the end of file, i.e.\npast the mmap()-ed memory region.\n\n> +\n> +\tif (lex_pos > 0)\n> +\t\tstart_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));\n> +\telse\n> +\t\tstart_index = 0;\n> +\n> +\tfilter->len = end_index - start_index;\n> +\tfilter->data = (unsigned char *)(g->chunk_bloom_data +\n> +\t\t\t\t\tsizeof(unsigned char) * start_index +\n> +\t\t\t\t\tBLOOMDATA_CHUNK_HEADER_SIZE);\n\nAnd this could lead to segfault later when accessing the Bloom filter\ndata if 'start_index' or 'end_index' were to point past EOF or\nend_index < start_index.\n\nIMO all indices and offsets read from the commit-graph file must be\nchecked to ensure that they fit in the corresponding chunk, like I did\nin my modified path Bloom filters implementation.  However, I'm not\nsure how it's best to handle an out-of-bounds offset...  Simply\nerroring out in case of a bogus commit-graph file is the\nstraightforward possibility, of course, but since the commit-graph is\nonly an optimization, it would be better user experience to warn and\nignore it and finish the operation without the commit-graph (albeit\nslower).  But is it even possible to ignore the commit-graph, say, in\nthe middle of a 'git rev-list --topo-order HEAD'?\n\n> +\treturn 1;\n> +}\n> +\n>  /*\n>   * Calculate the murmur3 32-bit hash value for the given data\n>   * using the given seed.\n> @@ -127,7 +159,8 @@ void init_bloom_filters(void)\n>  }\n>  \n>  struct bloom_filter *get_bloom_filter(struct repository *r,\n> -\t\t\t\t      struct commit *c)\n> +\t\t\t\t      struct commit *c,\n> +\t\t\t\t\t  int compute_if_not_present)\n>  {\n>  \tstruct bloom_filter *filter;\n>  \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n> @@ -140,6 +173,20 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \n>  \tfilter = bloom_filter_slab_at(&bloom_filters, c);\n>  \n> +\tif (!filter->data) {\n> +\t\tload_commit_graph_info(r, c);\n> +\t\tif (c->graph_pos != COMMIT_NOT_FROM_GRAPH &&\n> +\t\t\tr->objects->commit_graph->chunk_bloom_indexes) {\n> +\t\t\tif (load_bloom_filter_from_graph(r->objects->commit_graph, filter, c))\n> +\t\t\t\treturn filter;\n> +\t\t\telse\n> +\t\t\t\treturn NULL;\n> +\t\t}\n> +\t}\n> +\n> +\tif (filter->data || !compute_if_not_present)\n> +\t\treturn filter;\n> +\n>  \trepo_diff_setup(r, &diffopt);\n>  \tdiffopt.flags.recursive = 1;\n>  \tdiffopt.max_changes = max_changes;\n> diff --git a/bloom.h b/bloom.h\n> index 85ab8e9423d..760d7122374 100644\n> --- a/bloom.h\n> +++ b/bloom.h\n> @@ -32,6 +32,7 @@ struct bloom_filter_settings {\n>  \n>  #define DEFAULT_BLOOM_FILTER_SETTINGS { 1, 7, 10 }\n>  #define BITS_PER_WORD 8\n> +#define BLOOMDATA_CHUNK_HEADER_SIZE 3 * sizeof(uint32_t)\n>  \n>  /*\n>   * A bloom_filter struct represents a data segment to\n> @@ -79,6 +80,7 @@ void add_key_to_filter(const struct bloom_key *key,\n>  void init_bloom_filters(void);\n>  \n>  struct bloom_filter *get_bloom_filter(struct repository *r,\n> -\t\t\t\t      struct commit *c);\n> +\t\t\t\t      struct commit *c,\n> +\t\t\t\t      int compute_if_not_present);\n>  \n>  #endif\n> \\ No newline at end of file\n> diff --git a/commit-graph.c b/commit-graph.c\n> index a8b6b5cca5d..77668629e27 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -1086,7 +1086,7 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,\n>  \t\t\tctx->commits.nr);\n>  \n>  \twhile (list < last) {\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n>  \t\tcur_pos += filter->len;\n>  \t\tdisplay_progress(progress, ++i);\n>  \t\thashwrite_be32(f, cur_pos);\n> @@ -1115,7 +1115,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,\n>  \thashwrite_be32(f, settings->bits_per_entry);\n>  \n>  \twhile (list < last) {\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);\n>  \t\tdisplay_progress(progress, ++i);\n>  \t\thashwrite(f, filter->data, filter->len * sizeof(unsigned char));\n>  \t\tlist++;\n> @@ -1296,7 +1296,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)\n>  \n>  \tfor (i = 0; i < ctx->commits.nr; i++) {\n>  \t\tstruct commit *c = sorted_commits[i];\n> -\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c);\n> +\t\tstruct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);\n>  \t\tctx->total_bloom_filter_data_size += sizeof(unsigned char) * filter->len;\n>  \t\tdisplay_progress(progress, i + 1);\n>  \t}\n> diff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\n> index f18d1b722e1..ce412664ba9 100644\n> --- a/t/helper/test-bloom.c\n> +++ b/t/helper/test-bloom.c\n> @@ -39,7 +39,7 @@ static void get_bloom_filter_for_commit(const struct object_id *commit_oid)\n>  \tstruct bloom_filter *filter;\n>  \tsetup_git_directory();\n>  \tc = lookup_commit(the_repository, commit_oid);\n> -\tfilter = get_bloom_filter(the_repository, c);\n> +\tfilter = get_bloom_filter(the_repository, c, 1);\n>  \tprint_bloom_filter(filter);\n>  }\n>  \n> -- \n> gitgitgadget\n> \n"},{"id":"402816","messageId":"20200804144724.GA25052@szeder.dev","threadId":"52499","inReplyTo":"2d4c0b2da38632424c8bd31ccb2037e0676c3c74.1586192395.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 05/15] diff: halt tree-diff early after max_changes","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-08-04T14:47:24Z","receivedAt":"2020-08-04T14:47:48Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Mon, Apr 06, 2020 at 04:59:45PM +0000, Derrick Stolee via GitGitGadget wrote:\n> From: Derrick Stolee <dstolee@microsoft.com>\n> \n> When computing the changed-paths bloom filters for the commit-graph,\n> we limit the size of the filter by restricting the number of paths\n> in the diff. Instead of computing a large diff and then ignoring the\n> result, it is better to halt the diff computation early.\n> \n> Create a new \"max_changes\" option in struct diff_options. If non-zero,\n> then halt the diff computation after discovering strictly more changed\n> paths. This includes paths corresponding to trees that change.\n> \n> Use this max_changes option in the bloom filter calculations. This\n> reduces the time taken to compute the filters for the Linux kernel\n> repo from 2m50s to 2m35s. On a large internal repository with ~500\n> commits that perform tree-wide changes, the time reduced from\n> 6m15s to 3m48s.\n> \n> Signed-off-by: Derrick Stolee <dstolee@microsoft.com>\n> Signed-off-by: Garima Singh <garima.singh@microsoft.com>\n> ---\n>  bloom.c     | 4 +++-\n>  diff.h      | 5 +++++\n>  tree-diff.c | 6 ++++++\n>  3 files changed, 14 insertions(+), 1 deletion(-)\n> \n> diff --git a/bloom.c b/bloom.c\n> index 881a9841ede..a16eee92331 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -133,6 +133,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \tstruct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n>  \tint i;\n>  \tstruct diff_options diffopt;\n> +\tint max_changes = 512;\n>  \n>  \tif (bloom_filters.slab_size == 0)\n>  \t\treturn NULL;\n> @@ -141,6 +142,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \n>  \trepo_diff_setup(r, &diffopt);\n>  \tdiffopt.flags.recursive = 1;\n> +\tdiffopt.max_changes = max_changes;\n>  \tdiff_setup_done(&diffopt);\n>  \n>  \tif (c->parents)\n> @@ -149,7 +151,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n>  \tdiffcore_std(&diffopt);\n>  \n> -\tif (diff_queued_diff.nr <= 512) {\n> +\tif (diff_queued_diff.nr <= max_changes) {\n>  \t\tstruct hashmap pathmap;\n>  \t\tstruct pathmap_hash_entry *e;\n>  \t\tstruct hashmap_iter iter;\n> diff --git a/diff.h b/diff.h\n> index 6febe7e3656..9443dc1b003 100644\n> --- a/diff.h\n> +++ b/diff.h\n> @@ -285,6 +285,11 @@ struct diff_options {\n>  \t/* Number of hexdigits to abbreviate raw format output to. */\n>  \tint abbrev;\n>  \n> +\t/* If non-zero, then stop computing after this many changes. */\n> +\tint max_changes;\n> +\t/* For internal use only. */\n> +\tint num_changes;\n\n\"For internal use only\", understood.\n\n> +\n>  \tint ita_invisible_in_index;\n>  /* white-space error highlighting */\n>  #define WSEH_NEW (1<<12)\n> diff --git a/tree-diff.c b/tree-diff.c\n> index 33ded7f8b3e..f3d303c6e54 100644\n> --- a/tree-diff.c\n> +++ b/tree-diff.c\n> @@ -434,6 +434,9 @@ static struct combine_diff_path *ll_diff_tree_paths(\n>  \t\tif (diff_can_quit_early(opt))\n>  \t\t\tbreak;\n>  \n> +\t\tif (opt->max_changes && opt->num_changes > opt->max_changes)\n> +\t\t\tbreak;\n> +\n>  \t\tif (opt->pathspec.nr) {\n>  \t\t\tskip_uninteresting(&t, base, opt);\n>  \t\t\tfor (i = 0; i < nparent; i++)\n> @@ -518,6 +521,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n>  \n>  \t\t\t/* t↓ */\n>  \t\t\tupdate_tree_entry(&t);\n> +\t\t\topt->num_changes++;\n>  \t\t}\n>  \n>  \t\t/* t > p[imin] */\n> @@ -535,6 +539,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n>  \t\tskip_emit_tp:\n>  \t\t\t/* ∀ pi=p[imin]  pi↓ */\n>  \t\t\tupdate_tp_entries(tp, nparent);\n> +\t\t\topt->num_changes++;\n>  \t\t}\n>  \t}\n\nThis counter is basically broken, its value is wrong for over 98% of\ncommits, and, worse, its value remains 0 for over 85% of commits in\nthe repositories I usually use to test modified path Bloom filters.\nConsequently, a relatively large number of commits modifying more than\n512 paths get Bloom filters.\n\nThe makeshift tests in the patch below demonstrate these issues as\nmost of them fail, most notably those two tests that demonstrate that\nmodifying existing paths are not counted at all.\n\n\n  ---  >8  ---\n\ndiff --git a/bloom.c b/bloom.c\nindex 9b86aa3f59..3db0fde734 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -203,7 +203,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \trepo_diff_setup(r, &diffopt);\n \tdiffopt.flags.recursive = 1;\n \tdiffopt.detect_rename = 0;\n-\tdiffopt.max_changes = max_changes;\n+\tdiffopt.max_changes = 0;\n \tdiff_setup_done(&diffopt);\n \n \t/* ensure commit is parsed so we have parent information */\n@@ -214,6 +214,7 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \telse\n \t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n \tdiffcore_std(&diffopt);\n+\tprintf(\"%s  %d\\n\", oid_to_hex(&c->object.oid), diffopt.num_changes);\n \n \tif (diffopt.num_changes <= max_changes) {\n \t\tstruct hashmap pathmap;\ndiff --git a/t/t9999-test.sh b/t/t9999-test.sh\nnew file mode 100755\nindex 0000000000..8d2bd9f03f\n--- /dev/null\n+++ b/t/t9999-test.sh\n@@ -0,0 +1,142 @@\n+#!/bin/sh\n+\n+test_description='test'\n+\n+. ./test-lib.sh\n+\n+test_expect_success 'setup' '\n+\ttest_tick &&\n+\n+\techo 1 >file &&\n+\tmkdir -p dir/subdir &&\n+\techo 1 >dir/subdir/file1 &&\n+\techo 1 >dir/subdir/file2 &&\n+\tgit add file dir &&\n+\tgit commit -m setup &&\n+\n+\techo 2 >file &&\n+\tgit commit -a -m \"modify one path in root\" &&\n+\tmod_one_path=$(git rev-parse HEAD) &&\n+\n+\techo 2 >dir/subdir/file1 &&\n+\techo 2 >dir/subdir/file2 &&\n+\tgit commit -a -m \"modify two file two dirs deep\" &&\n+\tmod_four_paths=$(git rev-parse HEAD) &&\n+\n+\t>new-file &&\n+\tgit add new-file &&\n+\tgit commit -m \"add new file in root\" &&\n+\tnew_file_in_root=$(git rev-parse HEAD) &&\n+\n+\tgit rm new-file &&\n+\tgit commit -m \"delete file in root\" &&\n+\tdelete_file_in_root=$(git rev-parse HEAD) &&\n+\n+\t>dir/new-file &&\n+\tgit add dir/new-file &&\n+\tgit commit -m \"add new file in dir\" &&\n+\tnew_file_in_dir=$(git rev-parse HEAD) &&\n+\n+\tgit rm dir/new-file &&\n+\tgit commit -m \"delete file in dir\" &&\n+\tdelete_file_in_dir=$(git rev-parse HEAD) &&\n+\n+\techo 1 >d-f &&\n+\tgit add d-f &&\n+\tgit commit -m foo &&\n+\tgit rm d-f &&\n+\tmkdir d-f &&\n+\techo 2 >d-f/file &&\n+\tgit add d-f &&\n+\tgit commit -m \"replace file with dir\" &&\n+\tfile_to_dir=$(git rev-parse HEAD) &&\n+\n+\t>d-f.c &&\n+\tgit add d-f.c &&\n+\tgit commit -m \"add a file that sorts between d-f and d-f/\" &&\n+\tgit rm -r d-f &&\n+\techo 3 >d-f &&\n+\tgit add d-f &&\n+\tgit commit -m \"replace dir with file\" &&\n+\tdir_to_file=$(git rev-parse HEAD) &&\n+\n+\tbin_sha1=$(git rev-parse HEAD:dir/subdir | hex2oct) &&\n+\t# leading zero in mode: the content of the tree remains the same,\n+\t# but its oid does change!\n+\tprintf \"040000 subdir\\0$bin_sha1\" >rawtree &&\n+\ttree1=$(git hash-object -t tree -w rawtree) &&\n+\tgit cat-file -p HEAD^{tree} >out &&\n+\ttree2=$(sed -e \"s/$(git rev-parse HEAD:dir/)/$tree1/\" out |git mktree) &&\n+\tdifferent_but_same_tree=$(git commit-tree \\\n+\t\t-m \"leading zeros in mode\" \\\n+\t\t-p $(git rev-parse HEAD) $tree2) &&\n+\tgit update-ref HEAD $different_but_same_tree &&\n+\n+\tgit commit-graph write --reachable --changed-paths >out &&\n+\tcat out  # debug\n+'\n+\n+test_expect_success 'modify one path in root' '\n+\tgit diff --name-status $mod_one_path^ $mod_one_path &&\n+\techo \"$mod_one_path  1\" >expect &&\n+\tgrep \"$mod_one_path\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'modify two file two dirs deep' '\n+\tgit diff --name-status $mod_four_paths^ $mod_four_paths &&\n+\techo \"$mod_four_paths  4\" >expect &&\n+\tgrep \"$mod_four_paths\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'add new file in root' '\n+\tgit diff --name-status $new_file_in_root^ $new_file_in_root &&\n+\techo \"$new_file_in_root  1\" >expect &&\n+\tgrep \"$new_file_in_root\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'delete file in root' '\n+\tgit diff --name-status $delete_file_in_root^ $delete_file_in_root &&\n+\techo \"$delete_file_in_root  1\" >expect &&\n+\tgrep \"$delete_file_in_root\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'add new file in dir' '\n+\tgit diff --name-status $new_file_in_dir^ $new_file_in_dir &&\n+\techo \"$new_file_in_dir  2\" >expect &&\n+\tgrep \"$new_file_in_dir\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'delete file in dir' '\n+\tgit diff --name-status $delete_file_in_dir^ $delete_file_in_dir &&\n+\techo \"$delete_file_in_dir  2\" >expect &&\n+\tgrep \"$delete_file_in_dir\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'replace file with dir' '\n+\tgit diff --name-status $file_to_dir^ $file_to_dir &&\n+\techo \"$file_to_dir  2\" >expect &&\n+\tgrep \"$file_to_dir\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'replace dir with file' '\n+\tgit diff --name-status $dir_to_file^ $dir_to_file &&\n+\techo \"$dir_to_file  2\" >expect &&\n+\tgrep \"$dir_to_file\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'leading zeros in mode' '\n+\tgit diff --name-status $different_but_same_tree^ $different_but_same_tree &&\n+\techo \"$different_but_same_tree  0\" >expect &&\n+\tgrep \"$different_but_same_tree\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_done\n\n  ---  >8  ---\n\n\n> @@ -552,6 +557,7 @@ struct combine_diff_path *diff_tree_paths(\n>  \tconst struct object_id **parents_oid, int nparent,\n>  \tstruct strbuf *base, struct diff_options *opt)\n>  {\n> +\topt->num_changes = 0;\n>  \tp = ll_diff_tree_paths(p, oid, parents_oid, nparent, base, opt);\n>  \n>  \t/*\n> -- \n> gitgitgadget\n> \n"},{"id":"402819","messageId":"a08c26bb-54ec-13af-e503-fccd68727cf3@gmail.com","threadId":"52499","inReplyTo":"20200804144724.GA25052@szeder.dev","subject":"Re: [PATCH v4 05/15] diff: halt tree-diff early after max_changes","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-04T16:25:45Z","receivedAt":"2020-08-04T16:25:54Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/4/2020 10:47 AM, SZEDER Gábor wrote:\n> On Mon, Apr 06, 2020 at 04:59:45PM +0000, Derrick Stolee via GitGitGadget wrote:\n> This counter is basically broken, its value is wrong for over 98% of\n> commits, and, worse, its value remains 0 for over 85% of commits in\n> the repositories I usually use to test modified path Bloom filters.\n> Consequently, a relatively large number of commits modifying more than\n> 512 paths get Bloom filters.\n\nThanks for finding this! The counter is only really tested in one\nplace, and that test only considers _file adds_, which is a problem.\n\nIf I understand this correctly, the bug is a performance-only bug\n(since this is a performance-only feature), but it is an important\none to fix.\n\nThere is certainly some dark magic happening in this tree-diff logic,\nso instead of trying to get an accurate count we should just use the\nmagic global diff_queued_diff to track the current list of file changes.\n\nNote: diff_queued_diff does not track the directory changes, so it\nis an under-count for the total changes to track in the Bloom filter.\nThis is later corrected by the block that adds these leading directory\nchanges.\n\n> The makeshift tests in the patch below demonstrate these issues as\n> most of them fail, most notably those two tests that demonstrate that\n> modifying existing paths are not counted at all.\n\nI adapted your diff along with ripping out 'num_changes' in favor\nof diff_queued_diff.nr. This required modifying some of your expected\nvalues in the test script (losing the leading directories in the\ncount).\n\nI'll work with Taylor to create a fix, and include proper testing\nof the logic here. We'll stick it in the v2 of his max-changed-paths\nseries [1]. He already has some helpful logging that can help create\ntests that ensure this logic is performing as expected.\n\nWe plan to have that fix available by later today or early tomorrow.\nWill you be available to help validate it?\n\n[1] https://lore.kernel.org/git/cover.1596480582.git.me@ttaylorr.com/\n\nThanks,\n-Stolee\n\n  --- >8 ---\n\ndiff --git a/bloom.c b/bloom.c\nindex 1a573226e7..b8d6cb9240 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -218,8 +218,9 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n \telse\n \t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n \tdiffcore_std(&diffopt);\n+\tprintf(\"%s  %d\\n\", oid_to_hex(&c->object.oid), diff_queued_diff.nr);\n \n-\tif (diffopt.num_changes <= max_changes) {\n+\tif (diff_queued_diff.nr <= max_changes) {\n \t\tstruct hashmap pathmap;\n \t\tstruct pathmap_hash_entry *e;\n \t\tstruct hashmap_iter iter;\ndiff --git a/diff.h b/diff.h\nindex e0c0af6286b..1d32b718857 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -287,8 +287,6 @@ struct diff_options {\n \n \t/* If non-zero, then stop computing after this many changes. */\n \tint max_changes;\n-\t/* For internal use only. */\n-\tint num_changes;\n \n \tint ita_invisible_in_index;\n /* white-space error highlighting */\ndiff --git a/t/t9999-test.sh b/t/t9999-test.sh\nnew file mode 100755\nindex 00000000000..1f35aa8e2c5\n--- /dev/null\n+++ b/t/t9999-test.sh\n@@ -0,0 +1,142 @@\n+#!/bin/sh\n+\n+test_description='test'\n+\n+. ./test-lib.sh\n+\n+test_expect_success 'setup' '\n+\ttest_tick &&\n+\n+\techo 1 >file &&\n+\tmkdir -p dir/subdir &&\n+\techo 1 >dir/subdir/file1 &&\n+\techo 1 >dir/subdir/file2 &&\n+\tgit add file dir &&\n+\tgit commit -m setup &&\n+\n+\techo 2 >file &&\n+\tgit commit -a -m \"modify one path in root\" &&\n+\tmod_one_path=$(git rev-parse HEAD) &&\n+\n+\techo 2 >dir/subdir/file1 &&\n+\techo 2 >dir/subdir/file2 &&\n+\tgit commit -a -m \"modify two file two dirs deep\" &&\n+\tmod_four_paths=$(git rev-parse HEAD) &&\n+\n+\t>new-file &&\n+\tgit add new-file &&\n+\tgit commit -m \"add new file in root\" &&\n+\tnew_file_in_root=$(git rev-parse HEAD) &&\n+\n+\tgit rm new-file &&\n+\tgit commit -m \"delete file in root\" &&\n+\tdelete_file_in_root=$(git rev-parse HEAD) &&\n+\n+\t>dir/new-file &&\n+\tgit add dir/new-file &&\n+\tgit commit -m \"add new file in dir\" &&\n+\tnew_file_in_dir=$(git rev-parse HEAD) &&\n+\n+\tgit rm dir/new-file &&\n+\tgit commit -m \"delete file in dir\" &&\n+\tdelete_file_in_dir=$(git rev-parse HEAD) &&\n+\n+\techo 1 >d-f &&\n+\tgit add d-f &&\n+\tgit commit -m foo &&\n+\tgit rm d-f &&\n+\tmkdir d-f &&\n+\techo 2 >d-f/file &&\n+\tgit add d-f &&\n+\tgit commit -m \"replace file with dir\" &&\n+\tfile_to_dir=$(git rev-parse HEAD) &&\n+\n+\t>d-f.c &&\n+\tgit add d-f.c &&\n+\tgit commit -m \"add a file that sorts between d-f and d-f/\" &&\n+\tgit rm -r d-f &&\n+\techo 3 >d-f &&\n+\tgit add d-f &&\n+\tgit commit -m \"replace dir with file\" &&\n+\tdir_to_file=$(git rev-parse HEAD) &&\n+\n+\tbin_sha1=$(git rev-parse HEAD:dir/subdir | hex2oct) &&\n+\t# leading zero in mode: the content of the tree remains the same,\n+\t# but its oid does change!\n+\tprintf \"040000 subdir\\0$bin_sha1\" >rawtree &&\n+\ttree1=$(git hash-object -t tree -w rawtree) &&\n+\tgit cat-file -p HEAD^{tree} >out &&\n+\ttree2=$(sed -e \"s/$(git rev-parse HEAD:dir/)/$tree1/\" out |git mktree) &&\n+\tdifferent_but_same_tree=$(git commit-tree \\\n+\t\t-m \"leading zeros in mode\" \\\n+\t\t-p $(git rev-parse HEAD) $tree2) &&\n+\tgit update-ref HEAD $different_but_same_tree &&\n+\n+\tgit commit-graph write --reachable --changed-paths >out &&\n+\tcat out  # debug\n+'\n+\n+test_expect_success 'modify one path in root' '\n+\tgit diff --name-status $mod_one_path^ $mod_one_path &&\n+\techo \"$mod_one_path  1\" >expect &&\n+\tgrep \"$mod_one_path\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'modify two file two dirs deep' '\n+\tgit diff --name-status $mod_four_paths^ $mod_four_paths &&\n+\techo \"$mod_four_paths  2\" >expect &&\n+\tgrep \"$mod_four_paths\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'add new file in root' '\n+\tgit diff --name-status $new_file_in_root^ $new_file_in_root &&\n+\techo \"$new_file_in_root  1\" >expect &&\n+\tgrep \"$new_file_in_root\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'delete file in root' '\n+\tgit diff --name-status $delete_file_in_root^ $delete_file_in_root &&\n+\techo \"$delete_file_in_root  1\" >expect &&\n+\tgrep \"$delete_file_in_root\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'add new file in dir' '\n+\tgit diff --name-status $new_file_in_dir^ $new_file_in_dir &&\n+\techo \"$new_file_in_dir  1\" >expect &&\n+\tgrep \"$new_file_in_dir\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'delete file in dir' '\n+\tgit diff --name-status $delete_file_in_dir^ $delete_file_in_dir &&\n+\techo \"$delete_file_in_dir  1\" >expect &&\n+\tgrep \"$delete_file_in_dir\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'replace file with dir' '\n+\tgit diff --name-status $file_to_dir^ $file_to_dir &&\n+\techo \"$file_to_dir  2\" >expect &&\n+\tgrep \"$file_to_dir\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'replace dir with file' '\n+\tgit diff --name-status $dir_to_file^ $dir_to_file &&\n+\techo \"$dir_to_file  2\" >expect &&\n+\tgrep \"$dir_to_file\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'leading zeros in mode' '\n+\tgit diff --name-status $different_but_same_tree^ $different_but_same_tree &&\n+\techo \"$different_but_same_tree  0\" >expect &&\n+\tgrep \"$different_but_same_tree\" out >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_done\ndiff --git a/tree-diff.c b/tree-diff.c\nindex 6ebad1a46f3..7cebbb327e2 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -434,7 +434,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\tif (diff_can_quit_early(opt))\n \t\t\tbreak;\n \n-\t\tif (opt->max_changes && opt->num_changes > opt->max_changes)\n+\t\tif (opt->max_changes && diff_queued_diff.nr > opt->max_changes)\n \t\t\tbreak;\n \n \t\tif (opt->pathspec.nr) {\n@@ -521,7 +521,6 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \n \t\t\t/* t↓ */\n \t\t\tupdate_tree_entry(&t);\n-\t\t\topt->num_changes++;\n \t\t}\n \n \t\t/* t > p[imin] */\n@@ -539,7 +538,6 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\tskip_emit_tp:\n \t\t\t/* ∀ pi=p[imin]  pi↓ */\n \t\t\tupdate_tp_entries(tp, nparent);\n-\t\t\topt->num_changes++;\n \t\t}\n \t}\n \n@@ -557,7 +555,6 @@ struct combine_diff_path *diff_tree_paths(\n \tconst struct object_id **parents_oid, int nparent,\n \tstruct strbuf *base, struct diff_options *opt)\n {\n-\topt->num_changes = 0;\n \tp = ll_diff_tree_paths(p, oid, parents_oid, nparent, base, opt);\n \n \t/*\n"},{"id":"402823","messageId":"20200804170040.GB25052@szeder.dev","threadId":"52499","inReplyTo":"a08c26bb-54ec-13af-e503-fccd68727cf3@gmail.com","subject":"Re: [PATCH v4 05/15] diff: halt tree-diff early after max_changes","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2020-08-04T17:00:40Z","receivedAt":"2020-08-04T17:00:50Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Tue, Aug 04, 2020 at 12:25:45PM -0400, Derrick Stolee wrote:\n> On 8/4/2020 10:47 AM, SZEDER Gábor wrote:\n> > On Mon, Apr 06, 2020 at 04:59:45PM +0000, Derrick Stolee via GitGitGadget wrote:\n> > This counter is basically broken, its value is wrong for over 98% of\n> > commits, and, worse, its value remains 0 for over 85% of commits in\n> > the repositories I usually use to test modified path Bloom filters.\n> > Consequently, a relatively large number of commits modifying more than\n> > 512 paths get Bloom filters.\n> \n> Thanks for finding this! The counter is only really tested in one\n> place, and that test only considers _file adds_, which is a problem.\n> \n> If I understand this correctly, the bug is a performance-only bug\n> (since this is a performance-only feature), but it is an important\n> one to fix.\n\nOr a performance-only feature in a performance-only feature, because\nthose additional modified path Bloom filters can improve the runtime\nof pathspec-limited revision walks (assuming that the false positive\nrate is low enough).\n\n> There is certainly some dark magic happening in this tree-diff logic,\n> so instead of trying to get an accurate count we should just use the\n> magic global diff_queued_diff to track the current list of file changes.\n> \n> Note: diff_queued_diff does not track the directory changes, so it\n> is an under-count for the total changes to track in the Bloom filter.\n> This is later corrected by the block that adds these leading directory\n> changes.\n> \n> > The makeshift tests in the patch below demonstrate these issues as\n> > most of them fail, most notably those two tests that demonstrate that\n> > modifying existing paths are not counted at all.\n> \n> I adapted your diff along with ripping out 'num_changes' in favor\n> of diff_queued_diff.nr. This required modifying some of your expected\n> values in the test script (losing the leading directories in the\n> count).\n> \n> I'll work with Taylor to create a fix, and include proper testing\n> of the logic here. We'll stick it in the v2 of his max-changed-paths\n> series [1]. He already has some helpful logging that can help create\n> tests that ensure this logic is performing as expected.\n\nDon't forget to include a check of the hashmap's size, to make sure.\n\nFWIW, the patch below does result in the correct count (read: the same\nas in my implemenation) for all but 4 commits in those repositories I\nuse for testing, without adding any memory allocations and extra\nstrcmp() calls.\n\n  ---  >8  ---\n\ndiff --git a/cache.h b/cache.h\nindex 0f0485ecfe..3fc7e1b427 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -1574,6 +1574,7 @@ int repo_interpret_branch_name(struct repository *r,\n int validate_headref(const char *ref);\n \n int base_name_compare(const char *name1, int len1, int mode1, const char *name2, int len2, int mode2);\n+int base_name_compare_df(const char *name1, int len1, int mode1, const char *name2, int len2, int mode2, int *df);\n int df_name_compare(const char *name1, int len1, int mode1, const char *name2, int len2, int mode2);\n int name_compare(const char *name1, size_t len1, const char *name2, size_t len2);\n int cache_name_stage_compare(const char *name1, int len1, int stage1, const char *name2, int len2, int stage2);\ndiff --git a/read-cache.c b/read-cache.c\nindex aa427c5c17..041af19e60 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -460,13 +460,16 @@ int ie_modified(struct index_state *istate,\n \treturn 0;\n }\n \n-int base_name_compare(const char *name1, int len1, int mode1,\n-\t\t      const char *name2, int len2, int mode2)\n+int base_name_compare_df(const char *name1, int len1, int mode1,\n+\t\t\t const char *name2, int len2, int mode2,\n+\t\t\t int *df)\n {\n \tunsigned char c1, c2;\n \tint len = len1 < len2 ? len1 : len2;\n \tint cmp;\n \n+\t*df = 0;\n+\n \tcmp = memcmp(name1, name2, len);\n \tif (cmp)\n \t\treturn cmp;\n@@ -476,7 +479,21 @@ int base_name_compare(const char *name1, int len1, int mode1,\n \t\tc1 = '/';\n \tif (!c2 && S_ISDIR(mode2))\n \t\tc2 = '/';\n-\treturn (c1 < c2) ? -1 : (c1 > c2) ? 1 : 0;\n+\tif (c1 == c2)\n+\t\treturn 0;\t/* TODO: is this even possible? */\n+\tif ((c1 == '/' && !c2) ||\n+\t    (!c1 && c2 == '/'))\n+\t\t*df = 1;\n+\treturn (c1 < c2) ? -1 : 1;\n+}\n+\n+int base_name_compare(const char *name1, int len1, int mode1,\n+\t\t      const char *name2, int len2, int mode2)\n+{\n+\tint unused;\n+\treturn base_name_compare_df(name1, len1, mode1,\n+\t\t\t\t    name2, len2, mode2,\n+\t\t\t\t    &unused);\n }\n \n /*\ndiff --git a/t/t9999-test.sh b/t/t9999-test.sh\nindex 8d2bd9f03f..4f08590b45 100755\n--- a/t/t9999-test.sh\n+++ b/t/t9999-test.sh\n@@ -125,7 +125,7 @@ test_expect_success 'replace file with dir' '\n \ttest_cmp expect actual\n '\n \n-test_expect_success 'replace dir with file' '\n+test_expect_failure 'replace dir with file' '\n \tgit diff --name-status $dir_to_file^ $dir_to_file &&\n \techo \"$dir_to_file  2\" >expect &&\n \tgrep \"$dir_to_file\" out >actual &&\ndiff --git a/tree-diff.c b/tree-diff.c\nindex f3d303c6e5..e27f9c805e 100644\n--- a/tree-diff.c\n+++ b/tree-diff.c\n@@ -46,11 +46,14 @@ static int ll_diff_tree_oid(const struct object_id *old_oid,\n  *      Due to this convention, if trees are scanned in sorted order, all\n  *      non-empty descriptors will be processed first.\n  */\n-static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n+static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2,\n+\t\t\t      int *df)\n {\n \tstruct name_entry *e1, *e2;\n \tint cmp;\n \n+\t*df = 0;\n+\n \t/* empty descriptors sort after valid tree entries */\n \tif (!t1->size)\n \t\treturn t2->size ? 1 : 0;\n@@ -59,8 +62,9 @@ static int tree_entry_pathcmp(struct tree_desc *t1, struct tree_desc *t2)\n \n \te1 = &t1->entry;\n \te2 = &t2->entry;\n-\tcmp = base_name_compare(e1->path, tree_entry_len(e1), e1->mode,\n-\t\t\t\te2->path, tree_entry_len(e2), e2->mode);\n+\tcmp = base_name_compare_df(e1->path, tree_entry_len(e1), e1->mode,\n+\t\t\t\t   e2->path, tree_entry_len(e2), e2->mode,\n+\t\t\t\t   df);\n \treturn cmp;\n }\n \n@@ -410,7 +414,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n {\n \tstruct tree_desc t, *tp;\n \tvoid *ttree, **tptree;\n-\tint i;\n+\tint i, df;\n \n \tFAST_ARRAY_ALLOC(tp, nparent);\n \tFAST_ARRAY_ALLOC(tptree, nparent);\n@@ -463,7 +467,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\ttp[0].entry.mode &= ~S_IFXMIN_NEQ;\n \n \t\tfor (i = 1; i < nparent; ++i) {\n-\t\t\tcmp = tree_entry_pathcmp(&tp[i], &tp[imin]);\n+\t\t\tcmp = tree_entry_pathcmp(&tp[i], &tp[imin], &df);\n \t\t\tif (cmp < 0) {\n \t\t\t\timin = i;\n \t\t\t\ttp[i].entry.mode &= ~S_IFXMIN_NEQ;\n@@ -483,10 +487,12 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \n \n \t\t/* compare t vs p[imin] */\n-\t\tcmp = tree_entry_pathcmp(&t, &tp[imin]);\n+\t\tcmp = tree_entry_pathcmp(&t, &tp[imin], &df);\n \n \t\t/* t = p[imin] */\n \t\tif (cmp == 0) {\n+\t\t\tint prev_num_changes = opt->num_changes;\n+\n \t\t\t/* are either pi > p[imin] or diff(t,pi) != ø ? */\n \t\t\tif (!opt->flags.find_copies_harder) {\n \t\t\t\tfor (i = 0; i < nparent; ++i) {\n@@ -506,6 +512,9 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\t\t/* D += {δ(t,pi) if pi=p[imin];  \"+a\" if pi > p[imin]} */\n \t\t\tp = emit_path(p, base, opt, nparent,\n \t\t\t\t\t&t, tp, imin);\n+\t\t\tif (!(opt->num_changes == prev_num_changes &&\n+\t\t\t      S_ISDIR(t.entry.mode)))\n+\t\t\t\topt->num_changes++;\n \n \t\tskip_emit_t_tp:\n \t\t\t/* t↓,  ∀ pi=p[imin]  pi↓ */\n@@ -518,10 +527,11 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \t\t\t/* D += \"+t\" */\n \t\t\tp = emit_path(p, base, opt, nparent,\n \t\t\t\t\t&t, /*tp=*/NULL, -1);\n+\t\t\tif (!df)\n+\t\t\t\topt->num_changes++;\n \n \t\t\t/* t↓ */\n \t\t\tupdate_tree_entry(&t);\n-\t\t\topt->num_changes++;\n \t\t}\n \n \t\t/* t > p[imin] */\n@@ -535,11 +545,12 @@ static struct combine_diff_path *ll_diff_tree_paths(\n \n \t\t\tp = emit_path(p, base, opt, nparent,\n \t\t\t\t\t/*t=*/NULL, tp, imin);\n+\t\t\tif (!df)\n+\t\t\t\topt->num_changes++;\n \n \t\tskip_emit_tp:\n \t\t\t/* ∀ pi=p[imin]  pi↓ */\n \t\t\tupdate_tp_entries(tp, nparent);\n-\t\t\topt->num_changes++;\n \t\t}\n \t}\n \n  ---  >8  ---\n\n\nHaving said that, the best (i.e faster and accurate) solution to this\nissue is probably:\n\n  - Update the callchain between diff_tree_oid() and the diff callback\n    functions to allow the callbacks to break diffing with a non-zero\n    error code.\n\n  - Fill Bloom filters using the approach presented in:\n\n      https://public-inbox.org/git/20200529085038.26008-21-szeder.dev@gmail.com/\n\n    but modify the callbacks to return non-zero when too many paths\n    have been processed.\n\n  - Drop this counter entirely, as there are no other users.\n\n> We plan to have that fix available by later today or early tomorrow.\n> Will you be available to help validate it?\n> \n> [1] https://lore.kernel.org/git/cover.1596480582.git.me@ttaylorr.com/\n> \n> Thanks,\n> -Stolee\n> \n>   --- >8 ---\n> \n> diff --git a/bloom.c b/bloom.c\n> index 1a573226e7..b8d6cb9240 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -218,8 +218,9 @@ struct bloom_filter *get_bloom_filter(struct repository *r,\n>  \telse\n>  \t\tdiff_tree_oid(NULL, &c->object.oid, \"\", &diffopt);\n>  \tdiffcore_std(&diffopt);\n> +\tprintf(\"%s  %d\\n\", oid_to_hex(&c->object.oid), diff_queued_diff.nr);\n>  \n> -\tif (diffopt.num_changes <= max_changes) {\n> +\tif (diff_queued_diff.nr <= max_changes) {\n>  \t\tstruct hashmap pathmap;\n>  \t\tstruct pathmap_hash_entry *e;\n>  \t\tstruct hashmap_iter iter;\n> diff --git a/diff.h b/diff.h\n> index e0c0af6286b..1d32b718857 100644\n> --- a/diff.h\n> +++ b/diff.h\n> @@ -287,8 +287,6 @@ struct diff_options {\n>  \n>  \t/* If non-zero, then stop computing after this many changes. */\n>  \tint max_changes;\n> -\t/* For internal use only. */\n> -\tint num_changes;\n>  \n>  \tint ita_invisible_in_index;\n>  /* white-space error highlighting */\n> diff --git a/t/t9999-test.sh b/t/t9999-test.sh\n> new file mode 100755\n> index 00000000000..1f35aa8e2c5\n> --- /dev/null\n> +++ b/t/t9999-test.sh\n> @@ -0,0 +1,142 @@\n> +#!/bin/sh\n> +\n> +test_description='test'\n> +\n> +. ./test-lib.sh\n> +\n> +test_expect_success 'setup' '\n> +\ttest_tick &&\n> +\n> +\techo 1 >file &&\n> +\tmkdir -p dir/subdir &&\n> +\techo 1 >dir/subdir/file1 &&\n> +\techo 1 >dir/subdir/file2 &&\n> +\tgit add file dir &&\n> +\tgit commit -m setup &&\n> +\n> +\techo 2 >file &&\n> +\tgit commit -a -m \"modify one path in root\" &&\n> +\tmod_one_path=$(git rev-parse HEAD) &&\n> +\n> +\techo 2 >dir/subdir/file1 &&\n> +\techo 2 >dir/subdir/file2 &&\n> +\tgit commit -a -m \"modify two file two dirs deep\" &&\n> +\tmod_four_paths=$(git rev-parse HEAD) &&\n> +\n> +\t>new-file &&\n> +\tgit add new-file &&\n> +\tgit commit -m \"add new file in root\" &&\n> +\tnew_file_in_root=$(git rev-parse HEAD) &&\n> +\n> +\tgit rm new-file &&\n> +\tgit commit -m \"delete file in root\" &&\n> +\tdelete_file_in_root=$(git rev-parse HEAD) &&\n> +\n> +\t>dir/new-file &&\n> +\tgit add dir/new-file &&\n> +\tgit commit -m \"add new file in dir\" &&\n> +\tnew_file_in_dir=$(git rev-parse HEAD) &&\n> +\n> +\tgit rm dir/new-file &&\n> +\tgit commit -m \"delete file in dir\" &&\n> +\tdelete_file_in_dir=$(git rev-parse HEAD) &&\n> +\n> +\techo 1 >d-f &&\n> +\tgit add d-f &&\n> +\tgit commit -m foo &&\n> +\tgit rm d-f &&\n> +\tmkdir d-f &&\n> +\techo 2 >d-f/file &&\n> +\tgit add d-f &&\n> +\tgit commit -m \"replace file with dir\" &&\n> +\tfile_to_dir=$(git rev-parse HEAD) &&\n> +\n> +\t>d-f.c &&\n> +\tgit add d-f.c &&\n> +\tgit commit -m \"add a file that sorts between d-f and d-f/\" &&\n> +\tgit rm -r d-f &&\n> +\techo 3 >d-f &&\n> +\tgit add d-f &&\n> +\tgit commit -m \"replace dir with file\" &&\n> +\tdir_to_file=$(git rev-parse HEAD) &&\n> +\n> +\tbin_sha1=$(git rev-parse HEAD:dir/subdir | hex2oct) &&\n> +\t# leading zero in mode: the content of the tree remains the same,\n> +\t# but its oid does change!\n> +\tprintf \"040000 subdir\\0$bin_sha1\" >rawtree &&\n> +\ttree1=$(git hash-object -t tree -w rawtree) &&\n> +\tgit cat-file -p HEAD^{tree} >out &&\n> +\ttree2=$(sed -e \"s/$(git rev-parse HEAD:dir/)/$tree1/\" out |git mktree) &&\n> +\tdifferent_but_same_tree=$(git commit-tree \\\n> +\t\t-m \"leading zeros in mode\" \\\n> +\t\t-p $(git rev-parse HEAD) $tree2) &&\n> +\tgit update-ref HEAD $different_but_same_tree &&\n> +\n> +\tgit commit-graph write --reachable --changed-paths >out &&\n> +\tcat out  # debug\n> +'\n> +\n> +test_expect_success 'modify one path in root' '\n> +\tgit diff --name-status $mod_one_path^ $mod_one_path &&\n> +\techo \"$mod_one_path  1\" >expect &&\n> +\tgrep \"$mod_one_path\" out >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'modify two file two dirs deep' '\n> +\tgit diff --name-status $mod_four_paths^ $mod_four_paths &&\n> +\techo \"$mod_four_paths  2\" >expect &&\n> +\tgrep \"$mod_four_paths\" out >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'add new file in root' '\n> +\tgit diff --name-status $new_file_in_root^ $new_file_in_root &&\n> +\techo \"$new_file_in_root  1\" >expect &&\n> +\tgrep \"$new_file_in_root\" out >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'delete file in root' '\n> +\tgit diff --name-status $delete_file_in_root^ $delete_file_in_root &&\n> +\techo \"$delete_file_in_root  1\" >expect &&\n> +\tgrep \"$delete_file_in_root\" out >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'add new file in dir' '\n> +\tgit diff --name-status $new_file_in_dir^ $new_file_in_dir &&\n> +\techo \"$new_file_in_dir  1\" >expect &&\n> +\tgrep \"$new_file_in_dir\" out >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'delete file in dir' '\n> +\tgit diff --name-status $delete_file_in_dir^ $delete_file_in_dir &&\n> +\techo \"$delete_file_in_dir  1\" >expect &&\n> +\tgrep \"$delete_file_in_dir\" out >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'replace file with dir' '\n> +\tgit diff --name-status $file_to_dir^ $file_to_dir &&\n> +\techo \"$file_to_dir  2\" >expect &&\n> +\tgrep \"$file_to_dir\" out >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'replace dir with file' '\n> +\tgit diff --name-status $dir_to_file^ $dir_to_file &&\n> +\techo \"$dir_to_file  2\" >expect &&\n> +\tgrep \"$dir_to_file\" out >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_expect_success 'leading zeros in mode' '\n> +\tgit diff --name-status $different_but_same_tree^ $different_but_same_tree &&\n> +\techo \"$different_but_same_tree  0\" >expect &&\n> +\tgrep \"$different_but_same_tree\" out >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n> +test_done\n> diff --git a/tree-diff.c b/tree-diff.c\n> index 6ebad1a46f3..7cebbb327e2 100644\n> --- a/tree-diff.c\n> +++ b/tree-diff.c\n> @@ -434,7 +434,7 @@ static struct combine_diff_path *ll_diff_tree_paths(\n>  \t\tif (diff_can_quit_early(opt))\n>  \t\t\tbreak;\n>  \n> -\t\tif (opt->max_changes && opt->num_changes > opt->max_changes)\n> +\t\tif (opt->max_changes && diff_queued_diff.nr > opt->max_changes)\n>  \t\t\tbreak;\n>  \n>  \t\tif (opt->pathspec.nr) {\n> @@ -521,7 +521,6 @@ static struct combine_diff_path *ll_diff_tree_paths(\n>  \n>  \t\t\t/* t↓ */\n>  \t\t\tupdate_tree_entry(&t);\n> -\t\t\topt->num_changes++;\n>  \t\t}\n>  \n>  \t\t/* t > p[imin] */\n> @@ -539,7 +538,6 @@ static struct combine_diff_path *ll_diff_tree_paths(\n>  \t\tskip_emit_tp:\n>  \t\t\t/* ∀ pi=p[imin]  pi↓ */\n>  \t\t\tupdate_tp_entries(tp, nparent);\n> -\t\t\topt->num_changes++;\n>  \t\t}\n>  \t}\n>  \n> @@ -557,7 +555,6 @@ struct combine_diff_path *diff_tree_paths(\n>  \tconst struct object_id **parents_oid, int nparent,\n>  \tstruct strbuf *base, struct diff_options *opt)\n>  {\n> -\topt->num_changes = 0;\n>  \tp = ll_diff_tree_paths(p, oid, parents_oid, nparent, base, opt);\n>  \n>  \t/*\n"},{"id":"402825","messageId":"cfa4e35e-e227-5862-6978-53e699a4a1d2@gmail.com","threadId":"52499","inReplyTo":"20200804170040.GB25052@szeder.dev","subject":"Re: [PATCH v4 05/15] diff: halt tree-diff early after max_changes","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-04T17:31:06Z","receivedAt":"2020-08-04T17:31:14Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/4/2020 1:00 PM, SZEDER Gábor wrote:\n> On Tue, Aug 04, 2020 at 12:25:45PM -0400, Derrick Stolee wrote:\n>> On 8/4/2020 10:47 AM, SZEDER Gábor wrote:\n>>> On Mon, Apr 06, 2020 at 04:59:45PM +0000, Derrick Stolee via GitGitGadget wrote:\n>>> This counter is basically broken, its value is wrong for over 98% of\n>>> commits, and, worse, its value remains 0 for over 85% of commits in\n>>> the repositories I usually use to test modified path Bloom filters.\n>>> Consequently, a relatively large number of commits modifying more than\n>>> 512 paths get Bloom filters.\n>>\n>> Thanks for finding this! The counter is only really tested in one\n>> place, and that test only considers _file adds_, which is a problem.\n>>\n>> If I understand this correctly, the bug is a performance-only bug\n>> (since this is a performance-only feature), but it is an important\n>> one to fix.\n> \n> Or a performance-only feature in a performance-only feature, because\n> those additional modified path Bloom filters can improve the runtime\n> of pathspec-limited revision walks (assuming that the false positive\n> rate is low enough).\n> \n>> There is certainly some dark magic happening in this tree-diff logic,\n>> so instead of trying to get an accurate count we should just use the\n>> magic global diff_queued_diff to track the current list of file changes.\n>>\n>> Note: diff_queued_diff does not track the directory changes, so it\n>> is an under-count for the total changes to track in the Bloom filter.\n>> This is later corrected by the block that adds these leading directory\n>> changes.\n>>\n>>> The makeshift tests in the patch below demonstrate these issues as\n>>> most of them fail, most notably those two tests that demonstrate that\n>>> modifying existing paths are not counted at all.\n>>\n>> I adapted your diff along with ripping out 'num_changes' in favor\n>> of diff_queued_diff.nr. This required modifying some of your expected\n>> values in the test script (losing the leading directories in the\n>> count).\n>>\n>> I'll work with Taylor to create a fix, and include proper testing\n>> of the logic here. We'll stick it in the v2 of his max-changed-paths\n>> series [1]. He already has some helpful logging that can help create\n>> tests that ensure this logic is performing as expected.\n> \n> Don't forget to include a check of the hashmap's size, to make sure.\n\nYes, thanks for the pointer. That check is currently not in there,\nsince the code assumes the hashmap's size will match num_changes.\nHopefully, the tests I intend to write around this would have caught\nsuch an omission.\n \n> FWIW, the patch below does result in the correct count (read: the same\n> as in my implemenation) for all but 4 commits in those repositories I\n> use for testing, without adding any memory allocations and extra\n> strcmp() calls.\n\n...\n\n> Having said that, the best (i.e faster and accurate) solution to this\n> issue is probably:\n> \n>   - Update the callchain between diff_tree_oid() and the diff callback\n>     functions to allow the callbacks to break diffing with a non-zero\n>     error code.\n\nIt looks like this part would not be too difficult. The pathchange\ncallback is called by emit_path() which returns a struct combine_diff_path\npointer. This could return NULL to signal an early termination, but\nwe need to update all callers of the following methods to handle NULL\nresponses:\n\n * emit_path()\n * ll_diff_tree_paths()\n * diff_tree_paths()\n\nOf some interest: diff_tree_paths() returns a struct combine_diff_path\npointer, but no callers seem to consume it.\n\n>   - Fill Bloom filters using the approach presented in:\n> \n>       https://public-inbox.org/git/20200529085038.26008-21-szeder.dev@gmail.com/\n> \n>     but modify the callbacks to return non-zero when too many paths\n>     have been processed.\n\nThanks for the pointer to that specific patch. You do a good job of\ndescribing your thought process, including why you used the callback\napproach instead of the diff queue approach. The main reason seemed to\nbe memory overhead from populating the entire diff queue before\nchecking the limit.\n\nHowever, if we are using the diff queue as the short-circuit, then\nperhaps that memory overhead isn't as much of a problem?\n\nYou admit yourself, that\n\n  This patch implements a more efficient, but more complex, approach:\n\nThe logic around matching prefixes definitely seems complex and\nhard to test, especially around the file/directory changes with the\nsort order problems that have plagued similar prefix checks recently.\nI'm not doubting your implementation, just saying that the complexity\nis worth considering before jumping to that solution too quickly.\n\nTo sum up, I intend to start with a fix that uses the diff queue\ncount as a limit, then try the callback approach to see if there are\nmeasurable improvements in performance.\n\n>   - Drop this counter entirely, as there are no other users.\n\nWith the callback approach, \"this counter\" is both num_changes and\nmax_changes, since the callback would perform all of the short-circuit\nlogic.\n\nThanks,\n-Stolee\n"},{"id":"402928","messageId":"58ae3044-fa07-8e88-414a-fa9b410957f9@gmail.com","threadId":"52499","inReplyTo":"cfa4e35e-e227-5862-6978-53e699a4a1d2@gmail.com","subject":"Re: [PATCH v4 05/15] diff: halt tree-diff early after max_changes","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2020-08-05T17:08:52Z","receivedAt":"2020-08-05T17:11:13Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/4/2020 1:31 PM, Derrick Stolee wrote:\n> On 8/4/2020 1:00 PM, SZEDER Gábor wrote:\n\n>> Having said that, the best (i.e faster and accurate) solution to this\n>> issue is probably:\n>>\n>>   - Update the callchain between diff_tree_oid() and the diff callback\n>>     functions to allow the callbacks to break diffing with a non-zero\n>>     error code.\n> \n> It looks like this part would not be too difficult. \n\nOh, my hubris! I gave this a shot for some time this morning. This\nwill definitely take some work to do right. Just changing the callbacks\nto return 'int' is a wide-sweeping change, but the place where they are\ncalled already has an 'int' return that means something different.\n\nI'm not saying this is impossible. It just takes more attention and care\nthan I can currently devote, given my other works in progress right now.\n\n>>   - Fill Bloom filters using the approach presented in:\n>>\n>>       https://public-inbox.org/git/20200529085038.26008-21-szeder.dev@gmail.com/\n>>\n>>     but modify the callbacks to return non-zero when too many paths\n>>     have been processed.\n> \n> Thanks for the pointer to that specific patch. You do a good job of\n> describing your thought process, including why you used the callback\n> approach instead of the diff queue approach. The main reason seemed to\n> be memory overhead from populating the entire diff queue before\n> checking the limit.\n> \n> However, if we are using the diff queue as the short-circuit, then\n> perhaps that memory overhead isn't as much of a problem?\n> \n> You admit yourself, that\n> \n>   This patch implements a more efficient, but more complex, approach:\n> \n> The logic around matching prefixes definitely seems complex and\n> hard to test, especially around the file/directory changes with the\n> sort order problems that have plagued similar prefix checks recently.\n> I'm not doubting your implementation, just saying that the complexity\n> is worth considering before jumping to that solution too quickly.\n> \n> To sum up, I intend to start with a fix that uses the diff queue\n> count as a limit, then try the callback approach to see if there are\n> measurable improvements in performance.\n\nThat fix is now available [1].\n\n[1] https://lore.kernel.org/git/d1c4bbcaa9627068d5d9fbd0e4a2e8c8834a4bd3.1596646576.git.me@ttaylorr.com/\n\nAgain, the callback approach seems promising. The complexity is\nstopping me from trying to apply it on top of the current\nimplementation, while I should be focusing on other things. I completely\nbelieve that that approach is faster and more memory-efficient. I would\nlove to test and review a patch that takes that approach here.\n\nThanks,\n-Stolee\n"}]}