{"thread":{"id":"63691","subject":"[PATCH 0/2] bloom: use bloom filter given multiple pathspec","startedAt":"2025-06-25T12:56:06Z","lastAt":"2025-07-15T15:09:39Z","messageCount":72,"participants":["Lidong Yan","Junio C Hamano","SZEDER Gábor","Patrick Steinhardt","Derrick Stolee"],"isPatch":true,"patchVersion":1,"patchTotal":2},"messages":[{"id":"520693","messageId":"20250625125541.3048632-1-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":null,"subject":"[PATCH 0/2] bloom: use bloom filter given multiple pathspec","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-06-25T12:55:39Z","receivedAt":"2025-06-25T12:56:06Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"git won't use bloom filter for multiple pathspec, which makes the command\n  git log -- file1 file2\nsignificantly slower than\n  git log -- file1 && git log -- file2\n\nThis issue is raised by Kai Koponen at\n  https://lore.kernel.org/git/CADYQcGqaMC=4jgbmnF9Q11oC11jfrqyvH8EuiRRHytpMXd4wYA@mail.gmail.com/\n\nTo fix this, revs->bloom_keys[] needs to become an array of bloom_keys[],\none for each literal pathspec element. For convenience, first commit\ncreates a new struct bloom_keyvec to hold all bloom keys for a single\npathspec. The second commit add for loop to check if any pathspec's keyvec\nis contained in a commit's bloom filter, along with code that initialize\ndestory and test multiple pathspec bloom keyvecs.\n\nWith this change, testing on Kai's example shows that\n  git rev-list -10 3730814f2f2bf24550920c39a16841583de2dac1 -- src/clean.bash src/Make.dist\nruns as fast as\n  git rev-list -10 3730814f2f2bf24550920c39a16841583de2dac1 -- src/Make.dist && \\\n  git rev-list -10 3730814f2f2bf24550920c39a16841583de2dac1 -- src/clean.bash\n\nLidong Yan (2):\n  bloom: replace struct bloom_key * with struct bloom_keyvec\n  bloom: enable multiple pathspec bloom keys\n\n bloom.c              |  47 +++++++++++++++++\n bloom.h              |  14 +++++\n revision.c           | 121 ++++++++++++++++++++++++-------------------\n revision.h           |   5 +-\n t/t4216-log-bloom.sh |  10 ++--\n 5 files changed, 137 insertions(+), 60 deletions(-)\n\n-- \n2.50.0.108.g6ae0c543ae\n\n"},{"id":"520694","messageId":"20250625125541.3048632-2-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250625125541.3048632-1-502024330056@smail.nju.edu.cn","subject":"[PATCH 1/2] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-06-25T12:55:40Z","receivedAt":"2025-06-25T12:56:12Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"struct rev_info uses bloom_keys and bloom_nr to store the bloom keys\ncorresponding to a single pathspec. To allow struct rev_info to store\nthe bloom keys for multiple pathspecs, a new data structure bloom_keyvec\nis introduced. Each struct bloom_keyvec corresponds to a single pathspec.\nIn rev_info, an array bloom_keyvecs is used to store multiple bloom_keyvec\ninstances, along with its length bloom_keyvec_nr.\n\nNew bloom_keyvec_* functions are added to initialize, free, and access\nelements in a keyvec. bloom_filter_contains_vec() is added to check\nif all key in struct bloom_keyvec is contained in a bloom filter.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n bloom.c    | 47 +++++++++++++++++++++++++++++++++++++++++++++++\n bloom.h    | 14 ++++++++++++++\n revision.c | 31 ++++++++++++++++---------------\n revision.h |  5 +++--\n 4 files changed, 80 insertions(+), 17 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 0c8d2cebf9..497bb44567 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -280,6 +280,36 @@ void deinit_bloom_filters(void)\n \tdeep_clear_bloom_filter_slab(&bloom_filters, free_one_bloom_filter);\n }\n \n+void bloom_keyvec_init(struct bloom_keyvec *v, size_t initial_size)\n+{\n+\tALLOC_ARRAY(v->keys, initial_size);\n+\tv->nr = initial_size;\n+}\n+\n+void bloom_keyvec_clear(struct bloom_keyvec *v)\n+{\n+\tsize_t i;\n+\tif (!v->keys)\n+\t\treturn;\n+\n+\tfor (i = 0; i < v->nr; i++)\n+\t\tclear_bloom_key(&v->keys[i]);\n+\n+\tFREE_AND_NULL(v->keys);\n+\tv->nr = 0;\n+}\n+\n+struct bloom_key *bloom_keyvec_at(const struct bloom_keyvec *v, size_t i)\n+{\n+\tassert(i < v->nr);\n+\treturn &v->keys[i];\n+}\n+\n+bool bloom_keyvec_empty(const struct bloom_keyvec *v)\n+{\n+\treturn v->nr == 0 || !v->keys;\n+}\n+\n static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n \t\t       const struct hashmap_entry *eptr,\n \t\t       const struct hashmap_entry *entry_or_key,\n@@ -540,3 +570,20 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \n \treturn 1;\n }\n+\n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *keys,\n+\t\t\t      const struct bloom_filter_settings *settings)\n+{\n+\tconst struct bloom_key *key;\n+\tint i, ret;\n+\n+\tfor (i = 0; i < keys->nr; i++) {\n+\t\tkey = &keys->keys[i];\n+\t\tret = bloom_filter_contains(filter, key, settings);\n+\t\tif (ret <= 0)\n+\t\t\treturn ret;\n+\t}\n+\n+\treturn 1;\n+}\n\\ No newline at end of file\ndiff --git a/bloom.h b/bloom.h\nindex 6e46489a20..d556af7310 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -74,6 +74,11 @@ struct bloom_key {\n \tuint32_t *hashes;\n };\n \n+struct bloom_keyvec {\n+\tstruct bloom_key *keys;\n+\tsize_t nr;\n+};\n+\n int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n@@ -100,6 +105,11 @@ void add_key_to_filter(const struct bloom_key *key,\n void init_bloom_filters(void);\n void deinit_bloom_filters(void);\n \n+void bloom_keyvec_init(struct bloom_keyvec *v, size_t initial_size);\n+void bloom_keyvec_clear(struct bloom_keyvec *v);\n+struct bloom_key *bloom_keyvec_at(const struct bloom_keyvec *v, size_t i);\n+bool bloom_keyvec_empty(const struct bloom_keyvec *v);\n+\n enum bloom_filter_computed {\n \tBLOOM_NOT_COMPUTED = (1 << 0),\n \tBLOOM_COMPUTED     = (1 << 1),\n@@ -137,4 +147,8 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n \t\t\t  const struct bloom_filter_settings *settings);\n \n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *keys,\n+\t\t\t      const struct bloom_filter_settings *settings);\n+\n #endif\ndiff --git a/revision.c b/revision.c\nindex afee111196..cf7dc3b3fa 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -688,6 +688,7 @@ static int forbid_bloom_filters(struct pathspec *spec)\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n+\tstruct bloom_keyvec *bloom_keyvec;\n \tchar *path_alloc = NULL;\n \tconst char *path, *p;\n \tsize_t len;\n@@ -736,10 +737,12 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \t\tp++;\n \t}\n \n-\trevs->bloom_keys_nr = path_component_nr;\n-\tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n+\trevs->bloom_keyvecs_nr = 1;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n+\tbloom_keyvec = &revs->bloom_keyvecs[0];\n+\tbloom_keyvec_init(bloom_keyvec, path_component_nr);\n \n-\tfill_bloom_key(path, len, &revs->bloom_keys[0],\n+\tfill_bloom_key(path, len, bloom_keyvec_at(bloom_keyvec, 0),\n \t\t       revs->bloom_filter_settings);\n \tpath_component_nr = 1;\n \n@@ -747,7 +750,8 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \twhile (p > path) {\n \t\tif (*p == '/')\n \t\t\tfill_bloom_key(path, p - path,\n-\t\t\t\t       &revs->bloom_keys[path_component_nr++],\n+\t\t\t\t       bloom_keyvec_at(bloom_keyvec,\n+\t\t\t\t\t\t       path_component_nr++),\n \t\t\t\t       revs->bloom_filter_settings);\n \t\tp--;\n \t}\n@@ -779,11 +783,8 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\treturn -1;\n \t}\n \n-\tfor (j = 0; result && j < revs->bloom_keys_nr; j++) {\n-\t\tresult = bloom_filter_contains(filter,\n-\t\t\t\t\t       &revs->bloom_keys[j],\n-\t\t\t\t\t       revs->bloom_filter_settings);\n-\t}\n+\tresult = bloom_filter_contains_vec(filter, &revs->bloom_keyvecs[0],\n+\t\t\t\t\t   revs->bloom_filter_settings);\n \n \tif (result)\n \t\tcount_bloom_filter_maybe++;\n@@ -823,7 +824,7 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\treturn REV_TREE_SAME;\n \t}\n \n-\tif (revs->bloom_keys_nr && !nth_parent) {\n+\tif (revs->bloom_keyvecs_nr && !nth_parent) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \n \t\tif (bloom_ret == 0)\n@@ -850,7 +851,7 @@ static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit,\n \tif (!t1)\n \t\treturn 0;\n \n-\tif (!nth_parent && revs->bloom_keys_nr) {\n+\tif (!nth_parent && revs->bloom_keyvecs_nr) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \t\tif (!bloom_ret)\n \t\t\treturn 1;\n@@ -3230,10 +3231,10 @@ void release_revisions(struct rev_info *revs)\n \tline_log_free(revs);\n \toidset_clear(&revs->missing_commits);\n \n-\tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n-\t\tclear_bloom_key(&revs->bloom_keys[i]);\n-\tFREE_AND_NULL(revs->bloom_keys);\n-\trevs->bloom_keys_nr = 0;\n+\tfor (int i = 0; i < revs->bloom_keyvecs_nr; i++)\n+\t\tbloom_keyvec_clear(&revs->bloom_keyvecs[i]);\n+\tFREE_AND_NULL(revs->bloom_keyvecs);\n+\trevs->bloom_keyvecs_nr = 0;\n }\n \n static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\ndiff --git a/revision.h b/revision.h\nindex 6d369cdad6..88167b710c 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -63,6 +63,7 @@ struct rev_info;\n struct string_list;\n struct saved_parents;\n struct bloom_key;\n+struct bloom_keyvec;\n struct bloom_filter_settings;\n struct option;\n struct parse_opt_ctx_t;\n@@ -360,8 +361,8 @@ struct rev_info {\n \n \t/* Commit graph bloom filter fields */\n \t/* The bloom filter key(s) for the pathspec */\n-\tstruct bloom_key *bloom_keys;\n-\tint bloom_keys_nr;\n+\tstruct bloom_keyvec *bloom_keyvecs;\n+\tint bloom_keyvecs_nr;\n \n \t/*\n \t * The bloom filter settings used to generate the key.\n-- \n2.50.0.108.g6ae0c543ae\n\n"},{"id":"520695","messageId":"20250625125541.3048632-3-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250625125541.3048632-1-502024330056@smail.nju.edu.cn","subject":"[PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-06-25T12:55:41Z","receivedAt":"2025-06-25T12:56:16Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Remove `if (spec->nr > 1)` to enable bloom filter given multiple\npathspec. Wrapped for loop around code in prepare_to_use_bloom_filter()\nto initialize each pathspec's struct bloom_keyvec. Add for loop\nin check_maybe_different_in_bloom_filter() to find if at least one\npathspec's bloom_keyvec is contained in bloom filter.\n\nAdd new function release_revisions_bloom_keyvecs() to free all bloom\nkeyvec owned by rev_info.\n\nModify t/t4216 to test if bloom filter is still used given multiple\npathspec.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n revision.c           | 122 ++++++++++++++++++++++++-------------------\n t/t4216-log-bloom.sh |  10 ++--\n 2 files changed, 73 insertions(+), 59 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex cf7dc3b3fa..8818f017f3 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -675,8 +675,6 @@ static int forbid_bloom_filters(struct pathspec *spec)\n {\n \tif (spec->has_wildcard)\n \t\treturn 1;\n-\tif (spec->nr > 1)\n-\t\treturn 1;\n \tif (spec->magic & ~PATHSPEC_LITERAL)\n \t\treturn 1;\n \tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n@@ -685,6 +683,8 @@ static int forbid_bloom_filters(struct pathspec *spec)\n \treturn 0;\n }\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n+\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n@@ -692,7 +692,7 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tchar *path_alloc = NULL;\n \tconst char *path, *p;\n \tsize_t len;\n-\tint path_component_nr = 1;\n+\tint path_component_nr;\n \n \tif (!revs->commits)\n \t\treturn;\n@@ -709,51 +709,53 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n-\tpi = &revs->pruning.pathspec.items[0];\n-\n-\t/* remove single trailing slash from path, if needed */\n-\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n-\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n-\t\tpath = path_alloc;\n-\t} else\n-\t\tpath = pi->match;\n-\n-\tlen = strlen(path);\n-\tif (!len) {\n-\t\trevs->bloom_filter_settings = NULL;\n-\t\tfree(path_alloc);\n-\t\treturn;\n-\t}\n-\n-\tp = path;\n-\twhile (*p) {\n-\t\t/*\n-\t\t * At this point, the path is normalized to use Unix-style\n-\t\t * path separators. This is required due to how the\n-\t\t * changed-path Bloom filters store the paths.\n-\t\t */\n-\t\tif (*p == '/')\n-\t\t\tpath_component_nr++;\n-\t\tp++;\n-\t}\n-\n-\trevs->bloom_keyvecs_nr = 1;\n-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n-\tbloom_keyvec = &revs->bloom_keyvecs[0];\n-\tbloom_keyvec_init(bloom_keyvec, path_component_nr);\n+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n+\t\tpi = &revs->pruning.pathspec.items[i];\n+\t\tpath_component_nr = 1;\n+\n+\t\t/* remove single trailing slash from path, if needed */\n+\t\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n+\t\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n+\t\t\tpath = path_alloc;\n+\t\t} else\n+\t\t\tpath = pi->match;\n+\n+\t\tlen = strlen(path);\n+\t\tif (!len)\n+\t\t\tgoto fail;\n+\n+\t\tp = path;\n+\t\twhile (*p) {\n+\t\t\t/*\n+\t\t\t * At this point, the path is normalized to use\n+\t\t\t * Unix-style path separators. This is required due to\n+\t\t\t * how the changed-path Bloom filters store the paths.\n+\t\t\t */\n+\t\t\tif (*p == '/')\n+\t\t\t\tpath_component_nr++;\n+\t\t\tp++;\n+\t\t}\n \n-\tfill_bloom_key(path, len, bloom_keyvec_at(bloom_keyvec, 0),\n-\t\t       revs->bloom_filter_settings);\n-\tpath_component_nr = 1;\n+\t\tbloom_keyvec = &revs->bloom_keyvecs[i];\n+\t\tbloom_keyvec_init(bloom_keyvec, path_component_nr);\n+\n+\t\tfill_bloom_key(path, len, bloom_keyvec_at(bloom_keyvec, 0),\n+\t\t\t       revs->bloom_filter_settings);\n+\t\tpath_component_nr = 1;\n+\n+\t\tp = path + len - 1;\n+\t\twhile (p > path) {\n+\t\t\tif (*p == '/')\n+\t\t\t\tfill_bloom_key(path, p - path,\n+\t\t\t\t\t       bloom_keyvec_at(bloom_keyvec,\n+\t\t\t\t\t\t\t       path_component_nr++),\n+\t\t\t\t\t       revs->bloom_filter_settings);\n+\t\t\tp--;\n+\t\t}\n \n-\tp = path + len - 1;\n-\twhile (p > path) {\n-\t\tif (*p == '/')\n-\t\t\tfill_bloom_key(path, p - path,\n-\t\t\t\t       bloom_keyvec_at(bloom_keyvec,\n-\t\t\t\t\t\t       path_component_nr++),\n-\t\t\t\t       revs->bloom_filter_settings);\n-\t\tp--;\n+\t\tFREE_AND_NULL(path_alloc);\n \t}\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n@@ -761,14 +763,19 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \t\tbloom_filter_atexit_registered = 1;\n \t}\n \n+\treturn;\n+\n+fail:\n+\trevs->bloom_filter_settings = NULL;\n \tfree(path_alloc);\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\t\t\t\t\t struct commit *commit)\n {\n \tstruct bloom_filter *filter;\n-\tint result = 1, j;\n+\tint result = 0;\n \n \tif (!revs->repo->objects->commit_graph)\n \t\treturn -1;\n@@ -783,8 +790,11 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\treturn -1;\n \t}\n \n-\tresult = bloom_filter_contains_vec(filter, &revs->bloom_keyvecs[0],\n-\t\t\t\t\t   revs->bloom_filter_settings);\n+\tfor (int i = 0; !result && i < revs->bloom_keyvecs_nr; i++) {\n+\t\tresult = bloom_filter_contains_vec(filter,\n+\t\t\t\t\t\t   &revs->bloom_keyvecs[i],\n+\t\t\t\t\t\t   revs->bloom_filter_settings);\n+\t}\n \n \tif (result)\n \t\tcount_bloom_filter_maybe++;\n@@ -3202,6 +3212,14 @@ static void release_revisions_mailmap(struct string_list *mailmap)\n \n static void release_revisions_topo_walk_info(struct topo_walk_info *info);\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs)\n+{\n+\tfor (int i = 0; i < revs->bloom_keyvecs_nr; i++)\n+\t\tbloom_keyvec_clear(&revs->bloom_keyvecs[i]);\n+\tFREE_AND_NULL(revs->bloom_keyvecs);\n+\trevs->bloom_keyvecs_nr = 0;\n+}\n+\n static void free_void_commit_list(void *list)\n {\n \tfree_commit_list(list);\n@@ -3230,11 +3248,7 @@ void release_revisions(struct rev_info *revs)\n \tclear_decoration(&revs->treesame, free);\n \tline_log_free(revs);\n \toidset_clear(&revs->missing_commits);\n-\n-\tfor (int i = 0; i < revs->bloom_keyvecs_nr; i++)\n-\t\tbloom_keyvec_clear(&revs->bloom_keyvecs[i]);\n-\tFREE_AND_NULL(revs->bloom_keyvecs);\n-\trevs->bloom_keyvecs_nr = 0;\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 8910d53cac..46d1900a21 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -138,8 +138,8 @@ test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n '\n \n-test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n+test_expect_success 'git log -- multiple path specs use Bloom filters' '\n+\ttest_bloom_filters_used \"-- file4 A/file1\"\n '\n \n test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n@@ -151,9 +151,9 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n \ttest_bloom_filters_used \"-- *renamed\"\n '\n \n-test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- *\" &&\n-\ttest_bloom_filters_not_used \"-- file*\"\n+test_expect_success 'git log with wildcard that resolves to a multiple paths uses Bloom filters' '\n+\ttest_bloom_filters_used \"-- *\" &&\n+\ttest_bloom_filters_used \"-- file*\"\n '\n \n test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n-- \n2.50.0.108.g6ae0c543ae\n\n"},{"id":"520725","messageId":"xmqq7c0zviat.fsf@gitster.g","threadId":"63691","inReplyTo":"20250625125541.3048632-1-502024330056@smail.nju.edu.cn","subject":"Re: [PATCH 0/2] bloom: use bloom filter given multiple pathspec","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-06-25T17:32:10Z","receivedAt":"2025-06-25T17:32:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lidong Yan <yldhome2d2@gmail.com> writes:\n\n> git won't use bloom filter for multiple pathspec, which makes the command\n\nLet's get the terminology straight.  A pathspec consists of one or\nmore pathspec elements (or pathspec items).\n\nAlso, \"git won't\" is overly general.  The series title shares the\nsame issue (\"given multiple pathspec\" does not even hint that this\nis about revision traversal---you are not making filter used with\npathspec with more than one element in other code paths).\n\nPerhaps like:\n\n    The revision traversal limited by pathspec has optimization when\n    the pathspec has only one element, it does not use any pathspec\n    magic (other than literal), and there is no wildcard.\n\n    While it is much harder to lift the latter two limitations,\n    supporting a pathspec with multiple elements is relatively easy.\n    Just make sure we hash each of them separately and ask the bloom\n    filter about them, and if we see none of them can possibly be\n    affected by the commit, we can skip without tree comparison.\n\nor something along that line?\n\n>   git log -- file1 file2\n> significantly slower than\n>   git log -- file1 && git log -- file2\n>\n> This issue is raised by Kai Koponen at\n>   https://lore.kernel.org/git/CADYQcGqaMC=4jgbmnF9Q11oC11jfrqyvH8EuiRRHytpMXd4wYA@mail.gmail.com/\n>\n> To fix this, revs->bloom_keys[] needs to become an array of bloom_keys[],\n> one for each literal pathspec element. For convenience, first commit\n> creates a new struct bloom_keyvec to hold all bloom keys for a single\n> pathspec. The second commit add for loop to check if any pathspec's keyvec\n> is contained in a commit's bloom filter, along with code that initialize\n> destory and test multiple pathspec bloom keyvecs.\n\nIt is nice to outline an approach to the solution one day, and see\nan almost exact implementation of it appear on the list a few days\nlater.  I wish all the development would go like this ;-)\n\n>  bloom.c              |  47 +++++++++++++++++\n>  bloom.h              |  14 +++++\n>  revision.c           | 121 ++++++++++++++++++++++++-------------------\n>  revision.h           |   5 +-\n>  t/t4216-log-bloom.sh |  10 ++--\n\nCan we have a set of real tests to make sure that the updated filter\ncode still identifies commits that touch the files without false\nnegatives?  False positives are OK as we will follow them with real\ntree comparison to determine what exactly got changed, but false\nnegatives are absolute no-no.\n\nTesting to see that the filter code path is activated is much less\ninteresting than the filter code path still functions correctly with\nthese changes presented here.  I have a feeling that with the\nchanges to the test in this series, you wouldn't even find a bug\nwhere you simply added subpaths for all pathspec elements into a\nsingle array and use the original \"bloom has to say 'possibly yes'\nto all array elements\" logic (which would incorrectly require that\nboth file1 and file2 must be modified).\n\nThanks.\n"},{"id":"520726","messageId":"xmqqtt43u36t.fsf@gitster.g","threadId":"63691","inReplyTo":"20250625125541.3048632-2-502024330056@smail.nju.edu.cn","subject":"Re: [PATCH 1/2] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-06-25T17:43:54Z","receivedAt":"2025-06-25T17:43:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lidong Yan <yldhome2d2@gmail.com> writes:\n\n> +void bloom_keyvec_init(struct bloom_keyvec *v, size_t initial_size)\n> +{\n> +\tALLOC_ARRAY(v->keys, initial_size);\n> +\tv->nr = initial_size;\n> +}\n\nHmph, does this ever grow once initialized?  When I outlined the\nsolution in the earlier discussion, I was wondering a structure that\nlooks more like\n\n\tstruct bloom_keyvec {\n\t\tsize_t count;\n\t\tstruct bloom_key key[FLEX_ARRAY];\n\t};\n\nAlso when your primary use of an array is to use one element at a\ntime (as opposed to the entire array as a single \"collection\"), name\nit singular, so that key[4] is more naturally understood as \"4-th\nkey\", not keys[4].\n\n> +void bloom_keyvec_clear(struct bloom_keyvec *v)\n> +{\n> +\tsize_t i;\n> +\tif (!v->keys)\n> +\t\treturn;\n> +\n> +\tfor (i = 0; i < v->nr; i++)\n> +\t\tclear_bloom_key(&v->keys[i]);\n\nBy doing\n\n\tfor (size_t nr; nr < v->nr; nr++)\n\nyou can\n\n - lose the separate local variable definition at the beginning;\n - avoid confusing \"i\", which hints to be an \"int\", to be of type \"size_t\"\n - limit the scope of \"nr\" a bit tigher.\n\nIf you make keyvec an fixed flex-array, the below would become\nunnecessary (and the check for NULL-ness of .keys[] array).\n\n> +\n> +\tFREE_AND_NULL(v->keys);\n> +\tv->nr = 0;\n> +}\n\n> +struct bloom_key *bloom_keyvec_at(const struct bloom_keyvec *v, size_t i)\n> +{\n\nDitto about abusing the name 'i'.\n\n> +\t\t\treturn ret;\n> +\t}\n> +\n> +\treturn 1;\n> +}\n> \\ No newline at end of file\n\nTell your editor to be more careful, perhaps?\n"},{"id":"520741","messageId":"691CA448-881F-45BF-9D38-190F189DBB4E@smail.nju.edu.cn","threadId":"63691","inReplyTo":"xmqq7c0zviat.fsf@gitster.g","subject":"Re: [PATCH 0/2] bloom: use bloom filter given multiple pathspec","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-06-26T03:34:38Z","receivedAt":"2025-06-26T03:35:05Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n> \n> Lidong Yan <yldhome2d2@gmail.com> writes:\n> \n>> git won't use bloom filter for multiple pathspec, which makes the command\n> \n> Let's get the terminology straight.  A pathspec consists of one or\n> more pathspec elements (or pathspec items).\n\nThanks for the clarification. I will be more precise with the terminology in v2.\n\n> Also, \"git won't\" is overly general.  The series title shares the\n> same issue (\"given multiple pathspec\" does not even hint that this\n> is about revision traversal---you are not making filter used with\n> pathspec with more than one element in other code paths).\n> \n> Perhaps like:\n> \n>    The revision traversal limited by pathspec has optimization when\n>    the pathspec has only one element, it does not use any pathspec\n>    magic (other than literal), and there is no wildcard.\n> \n>    While it is much harder to lift the latter two limitations,\n>    supporting a pathspec with multiple elements is relatively easy.\n>    Just make sure we hash each of them separately and ask the bloom\n>    filter about them, and if we see none of them can possibly be\n>    affected by the commit, we can skip without tree comparison.\n> \n> or something along that line?\n> \n\nWhat you wrote makes perfect sense to me, I’ll just copy and paste\nthose paragraphs into my cover letter. And the title would be\n  \"bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal\"\n\n> Can we have a set of real tests to make sure that the updated filter\n> code still identifies commits that touch the files without false\n> negatives?  False positives are OK as we will follow them with real\n> tree comparison to determine what exactly got changed, but false\n> negatives are absolute no-no.\n> \n> Testing to see that the filter code path is activated is much less\n> interesting than the filter code path still functions correctly with\n> these changes presented here.  I have a feeling that with the\n> changes to the test in this series, you wouldn't even find a bug\n> where you simply added subpaths for all pathspec elements into a\n> single array and use the original \"bloom has to say 'possibly yes'\n> to all array elements\" logic (which would incorrectly require that\n> both file1 and file2 must be modified).\n\nI assume that t4216/test_bloom_filters_used has already verified that\nusing bloom filters with multiple pathspec elements produces the same\nresults as when bloom filters are not used. But I would love to add more\ntest cases to check no false negative happened.\n\nThanks,\nLidong"},{"id":"520742","messageId":"4E0C5F0C-16BA-4B7B-BC5B-632CFEA871E3@smail.nju.edu.cn","threadId":"63691","inReplyTo":"xmqqtt43u36t.fsf@gitster.g","subject":"Re: [PATCH 1/2] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-06-26T03:44:40Z","receivedAt":"2025-06-26T03:45:05Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes\n> \n> Lidong Yan <yldhome2d2@gmail.com> writes:\n> \n>> +void bloom_keyvec_init(struct bloom_keyvec *v, size_t initial_size)\n>> +{\n>> + ALLOC_ARRAY(v->keys, initial_size);\n>> + v->nr = initial_size;\n>> +}\n> \n> Hmph, does this ever grow once initialized?  When I outlined the\n> solution in the earlier discussion, I was wondering a structure that\n> looks more like\n> \n> struct bloom_keyvec {\n> size_t count;\n> struct bloom_key key[FLEX_ARRAY];\n> };\n\nI understand. If we don't need to resize the array dynamically, we should\nallocate the array elements and the array length together to gain benefits\nsuch as spatial locality.\n\n> Also when your primary use of an array is to use one element at a\n> time (as opposed to the entire array as a single \"collection\"), name\n> it singular, so that key[4] is more naturally understood as \"4-th\n> key\", not keys[4].\n\nGot it.\n\n> \n>> +void bloom_keyvec_clear(struct bloom_keyvec *v)\n>> +{\n>> + size_t i;\n>> + if (!v->keys)\n>> + return;\n>> +\n>> + for (i = 0; i < v->nr; i++)\n>> + clear_bloom_key(&v->keys[i]);\n> \n> By doing\n> \n> for (size_t nr; nr < v->nr; nr++)\n> \n> you can\n> \n> - lose the separate local variable definition at the beginning;\n> - avoid confusing \"i\", which hints to be an \"int\", to be of type \"size_t\"\n> - limit the scope of \"nr\" a bit tigher.\n\nGot it.\n\n> \n> If you make keyvec an fixed flex-array, the below would become\n> unnecessary (and the check for NULL-ness of .keys[] array).\n> \n>> +\n>> + FREE_AND_NULL(v->keys);\n>> + v->nr = 0;\n>> +}\n\nAnother benefit to use flex-array.\n\n> \n>> +struct bloom_key *bloom_keyvec_at(const struct bloom_keyvec *v, size_t i)\n>> +{\n> \n> Ditto about abusing the name 'i'.\n> \n>> + return ret;\n>> + }\n>> +\n>> + return 1;\n>> +}\n>> \\ No newline at end of file\n> \n> Tell your editor to be more careful, perhaps?\n\nSeems git clang-format doesn’t add newline for me, I will try\n.editconfig next time.\n\nThanks,\nLidong\n\n"},{"id":"520760","messageId":"xmqqldper3lt.fsf@gitster.g","threadId":"63691","inReplyTo":"691CA448-881F-45BF-9D38-190F189DBB4E@smail.nju.edu.cn","subject":"Re: [PATCH 0/2] bloom: use bloom filter given multiple pathspec","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-06-26T14:15:26Z","receivedAt":"2025-06-26T14:15:30Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lidong Yan <502024330056@smail.nju.edu.cn> writes:\n\n> I assume that t4216/test_bloom_filters_used has already verified that\n> using bloom filters with multiple pathspec elements produces the same\n> results as when bloom filters are not used. But I would love to add more\n> test cases to check no false negative happened.\n\nIf we are sure we have enough test coverage, then it is great. If\nnot, it is superb if you can add some test coverage there.\n\nThanks.\n\n"},{"id":"520792","messageId":"20250627062154.1121530-1-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250625125541.3048632-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v2 0/2] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-06-27T06:21:52Z","receivedAt":"2025-06-27T06:22:08Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"The revision traversal limited by pathspec has optimization when\nthe pathspec has only one element, it does not use any pathspec\nmagic (other than literal), and there is no wildcard. The absence\nof optimization for multiple pathspec elements in revision traversal\ncause an issue raised by Kai Koponen at\n  https://lore.kernel.org/git/CADYQcGqaMC=4jgbmnF9Q11oC11jfrqyvH8EuiRRHytpMXd4wYA@mail.gmail.com/\n\nWhile it is much harder to lift the latter two limitations,\nsupporting a pathspec with multiple elements is relatively easy.\nJust make sure we hash each of them separately and ask the bloom\nfilter about them, and if we see none of them can possibly be\naffected by the commit, we can skip without tree comparison.\n\nFirst commit creates a new data structure `struct bloom_keyvec` to hold\nall bloom keys for a single pathspec item. Second commit add for loop\nto check if any pathspec item's keyvec is contained in a commit's bloom\nfilter.\n\nWith this change, testing on Kai's example shows that\n  git rev-list -10 3730814f2f2bf24550920c39a16841583de2dac1 -- src/clean.bash src/Make.dist\nruns as fast as\n  git rev-list -10 3730814f2f2bf24550920c39a16841583de2dac1 -- src/Make.dist && \\\n  git rev-list -10 3730814f2f2bf24550920c39a16841583de2dac1 -- src/clean.bash\n\nLidong Yan (2):\n  bloom: replace struct bloom_key * with struct bloom_keyvec\n  bloom: optimize multiple pathspec items in revision traversal\n\n bloom.c              |  31 +++++++++++\n bloom.h              |  20 +++++++\n revision.c           | 121 ++++++++++++++++++++++++-------------------\n revision.h           |   6 +--\n t/t4216-log-bloom.sh |  10 ++--\n 5 files changed, 127 insertions(+), 61 deletions(-)\n\n-- \n2.50.0.108.g6ae0c543ae\n\n"},{"id":"520793","messageId":"20250627062154.1121530-2-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250625125541.3048632-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v2 1/2] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-06-27T06:21:53Z","receivedAt":"2025-06-27T06:22:11Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"The revision traversal limited by pathspec has optimization when\nthe pathspec has only one element. To support optimization for\nmultiple pathspec items, we need to modify the data structures\nin struct rev_info.\n\nstruct rev_info uses bloom_keys and bloom_nr to store the bloom keys\ncorresponding to a single pathspec item. To allow struct rev_info\nto store bloom keys for multiple pathspec items, a new data structure\n`struct bloom_keyvec` is introduced. Each `struct bloom_keyvec`\ncorresponds to a single pathspec item.\n\nIn `struct rev_info`, replace bloom_keys and bloom_nr with bloom_keyvecs\nand bloom_keyvec_nr. This commit still optimize one pathspec item, thus\nbloom_keyvec_nr can only be 0 or 1.\n\nNew *_bloom_keyvec functions are added to create and destroy a keyvec.\nbloom_filter_contains_vec() is added to check if all key in keyvec is\ncontained in a bloom filter. fill_bloom_keyvec_key() is added to\ninitialize a key in keyvec.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n bloom.c    | 31 +++++++++++++++++++++++++++++++\n bloom.h    | 20 ++++++++++++++++++++\n revision.c | 36 ++++++++++++++++++------------------\n revision.h |  6 +++---\n 4 files changed, 72 insertions(+), 21 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 0c8d2cebf9..8259cfce51 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -280,6 +280,25 @@ void deinit_bloom_filters(void)\n \tdeep_clear_bloom_filter_slab(&bloom_filters, free_one_bloom_filter);\n }\n \n+struct bloom_keyvec *create_bloom_keyvec(size_t count)\n+{\n+\tstruct bloom_keyvec *vec;\n+\tsize_t sz = sizeof(struct bloom_keyvec);\n+\tsz += count * sizeof(struct bloom_key);\n+\tvec = (struct bloom_keyvec *)xcalloc(1, sz);\n+\tvec->count = count;\n+\treturn vec;\n+}\n+\n+void destroy_bloom_keyvec(struct bloom_keyvec *vec)\n+{\n+\tif (!vec)\n+\t\treturn;\n+\tfor (size_t nr = 0; nr < vec->count; nr++)\n+\t\tclear_bloom_key(&vec->key[nr]);\n+\tfree(vec);\n+}\n+\n static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n \t\t       const struct hashmap_entry *eptr,\n \t\t       const struct hashmap_entry *entry_or_key,\n@@ -540,3 +559,15 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \n \treturn 1;\n }\n+\n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *vec,\n+\t\t\t      const struct bloom_filter_settings *settings)\n+{\n+\tint ret = 1;\n+\n+\tfor (size_t nr = 0; ret > 0 && nr < vec->count; nr++)\n+\t\tret = bloom_filter_contains(filter, &vec->key[nr], settings);\n+\n+\treturn ret;\n+}\ndiff --git a/bloom.h b/bloom.h\nindex 6e46489a20..9e4e832c8c 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -74,6 +74,11 @@ struct bloom_key {\n \tuint32_t *hashes;\n };\n \n+struct bloom_keyvec {\n+\tsize_t count;\n+\tstruct bloom_key key[FLEX_ARRAY];\n+};\n+\n int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n@@ -100,6 +105,17 @@ void add_key_to_filter(const struct bloom_key *key,\n void init_bloom_filters(void);\n void deinit_bloom_filters(void);\n \n+struct bloom_keyvec *create_bloom_keyvec(size_t count);\n+void destroy_bloom_keyvec(struct bloom_keyvec *vec);\n+\n+static inline void fill_bloom_keyvec_key(const char *data, size_t len,\n+\t\t\t\t\t struct bloom_keyvec *vec, size_t nr,\n+\t\t\t\t\t const struct bloom_filter_settings *settings)\n+{\n+\tassert(nr < vec->count);\n+\tfill_bloom_key(data, len, &vec->key[nr], settings);\n+}\n+\n enum bloom_filter_computed {\n \tBLOOM_NOT_COMPUTED = (1 << 0),\n \tBLOOM_COMPUTED     = (1 << 1),\n@@ -137,4 +153,8 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n \t\t\t  const struct bloom_filter_settings *settings);\n \n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *v,\n+\t\t\t      const struct bloom_filter_settings *settings);\n+\n #endif\ndiff --git a/revision.c b/revision.c\nindex afee111196..3aa544c137 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -688,6 +688,7 @@ static int forbid_bloom_filters(struct pathspec *spec)\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n+\tstruct bloom_keyvec *bloom_keyvec;\n \tchar *path_alloc = NULL;\n \tconst char *path, *p;\n \tsize_t len;\n@@ -736,19 +737,21 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \t\tp++;\n \t}\n \n-\trevs->bloom_keys_nr = path_component_nr;\n-\tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n+\trevs->bloom_keyvecs_nr = 1;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n+\tbloom_keyvec = create_bloom_keyvec(path_component_nr);\n+\trevs->bloom_keyvecs[0] = bloom_keyvec;\n \n-\tfill_bloom_key(path, len, &revs->bloom_keys[0],\n-\t\t       revs->bloom_filter_settings);\n+\tfill_bloom_keyvec_key(path, len, bloom_keyvec, 0,\n+\t\t\t      revs->bloom_filter_settings);\n \tpath_component_nr = 1;\n \n \tp = path + len - 1;\n \twhile (p > path) {\n \t\tif (*p == '/')\n-\t\t\tfill_bloom_key(path, p - path,\n-\t\t\t\t       &revs->bloom_keys[path_component_nr++],\n-\t\t\t\t       revs->bloom_filter_settings);\n+\t\t\tfill_bloom_keyvec_key(path, p - path, bloom_keyvec,\n+\t\t\t\t\t      path_component_nr++,\n+\t\t\t\t\t      revs->bloom_filter_settings);\n \t\tp--;\n \t}\n \n@@ -779,11 +782,8 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\treturn -1;\n \t}\n \n-\tfor (j = 0; result && j < revs->bloom_keys_nr; j++) {\n-\t\tresult = bloom_filter_contains(filter,\n-\t\t\t\t\t       &revs->bloom_keys[j],\n-\t\t\t\t\t       revs->bloom_filter_settings);\n-\t}\n+\tresult = bloom_filter_contains_vec(filter, revs->bloom_keyvecs[0],\n+\t\t\t\t\t   revs->bloom_filter_settings);\n \n \tif (result)\n \t\tcount_bloom_filter_maybe++;\n@@ -823,7 +823,7 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\treturn REV_TREE_SAME;\n \t}\n \n-\tif (revs->bloom_keys_nr && !nth_parent) {\n+\tif (revs->bloom_keyvecs_nr && !nth_parent) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \n \t\tif (bloom_ret == 0)\n@@ -850,7 +850,7 @@ static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit,\n \tif (!t1)\n \t\treturn 0;\n \n-\tif (!nth_parent && revs->bloom_keys_nr) {\n+\tif (!nth_parent && revs->bloom_keyvecs_nr) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \t\tif (!bloom_ret)\n \t\t\treturn 1;\n@@ -3230,10 +3230,10 @@ void release_revisions(struct rev_info *revs)\n \tline_log_free(revs);\n \toidset_clear(&revs->missing_commits);\n \n-\tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n-\t\tclear_bloom_key(&revs->bloom_keys[i]);\n-\tFREE_AND_NULL(revs->bloom_keys);\n-\trevs->bloom_keys_nr = 0;\n+\tfor (int i = 0; i < revs->bloom_keyvecs_nr; i++)\n+\t\tdestroy_bloom_keyvec(revs->bloom_keyvecs[i]);\n+\tFREE_AND_NULL(revs->bloom_keyvecs);\n+\trevs->bloom_keyvecs_nr = 0;\n }\n \n static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\ndiff --git a/revision.h b/revision.h\nindex 6d369cdad6..ac843f58d0 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -62,7 +62,7 @@ struct repository;\n struct rev_info;\n struct string_list;\n struct saved_parents;\n-struct bloom_key;\n+struct bloom_keyvec;\n struct bloom_filter_settings;\n struct option;\n struct parse_opt_ctx_t;\n@@ -360,8 +360,8 @@ struct rev_info {\n \n \t/* Commit graph bloom filter fields */\n \t/* The bloom filter key(s) for the pathspec */\n-\tstruct bloom_key *bloom_keys;\n-\tint bloom_keys_nr;\n+\tstruct bloom_keyvec **bloom_keyvecs;\n+\tint bloom_keyvecs_nr;\n \n \t/*\n \t * The bloom filter settings used to generate the key.\n-- \n2.50.0.108.g6ae0c543ae\n\n"},{"id":"520794","messageId":"20250627062154.1121530-3-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250625125541.3048632-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v2 2/2] bloom: optimize multiple pathspec items in revision traversal","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-06-27T06:21:54Z","receivedAt":"2025-06-27T06:22:13Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Remove `if (spec->nr > 1)` to enable bloom filter given multiple\npathspec items. Add for loop in prepare_to_use_bloom_filter()\nto initialize each pathspec item's struct bloom_keyvec. Add for\nloop in check_maybe_different_in_bloom_filter() to find if at least one\nbloom_keyvec is contained in bloom filter.\n\nAdd new function release_revisions_bloom_keyvecs() to free all bloom\nkeyvec owned by rev_info.\n\nModify t/t4216 to ensure consistent results between the optimization\nfor multiple pathspec items using bloom filters and the case without\nbloom filter optimization.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n revision.c           | 121 ++++++++++++++++++++++++-------------------\n t/t4216-log-bloom.sh |  10 ++--\n 2 files changed, 73 insertions(+), 58 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 3aa544c137..5606f6c7f6 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -675,8 +675,6 @@ static int forbid_bloom_filters(struct pathspec *spec)\n {\n \tif (spec->has_wildcard)\n \t\treturn 1;\n-\tif (spec->nr > 1)\n-\t\treturn 1;\n \tif (spec->magic & ~PATHSPEC_LITERAL)\n \t\treturn 1;\n \tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n@@ -685,6 +683,8 @@ static int forbid_bloom_filters(struct pathspec *spec)\n \treturn 0;\n }\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n+\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n@@ -692,7 +692,7 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tchar *path_alloc = NULL;\n \tconst char *path, *p;\n \tsize_t len;\n-\tint path_component_nr = 1;\n+\tint path_component_nr;\n \n \tif (!revs->commits)\n \t\treturn;\n@@ -709,50 +709,53 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n-\tpi = &revs->pruning.pathspec.items[0];\n-\n-\t/* remove single trailing slash from path, if needed */\n-\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n-\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n-\t\tpath = path_alloc;\n-\t} else\n-\t\tpath = pi->match;\n-\n-\tlen = strlen(path);\n-\tif (!len) {\n-\t\trevs->bloom_filter_settings = NULL;\n-\t\tfree(path_alloc);\n-\t\treturn;\n-\t}\n-\n-\tp = path;\n-\twhile (*p) {\n-\t\t/*\n-\t\t * At this point, the path is normalized to use Unix-style\n-\t\t * path separators. This is required due to how the\n-\t\t * changed-path Bloom filters store the paths.\n-\t\t */\n-\t\tif (*p == '/')\n-\t\t\tpath_component_nr++;\n-\t\tp++;\n-\t}\n-\n-\trevs->bloom_keyvecs_nr = 1;\n-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n-\tbloom_keyvec = create_bloom_keyvec(path_component_nr);\n-\trevs->bloom_keyvecs[0] = bloom_keyvec;\n+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n+\t\tpi = &revs->pruning.pathspec.items[i];\n+\t\tpath_component_nr = 1;\n+\n+\t\t/* remove single trailing slash from path, if needed */\n+\t\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n+\t\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n+\t\t\tpath = path_alloc;\n+\t\t} else\n+\t\t\tpath = pi->match;\n+\n+\t\tlen = strlen(path);\n+\t\tif (!len)\n+\t\t\tgoto fail;\n+\n+\t\tp = path;\n+\t\twhile (*p) {\n+\t\t\t/*\n+\t\t\t * At this point, the path is normalized to use\n+\t\t\t * Unix-style path separators. This is required due to\n+\t\t\t * how the changed-path Bloom filters store the paths.\n+\t\t\t */\n+\t\t\tif (*p == '/')\n+\t\t\t\tpath_component_nr++;\n+\t\t\tp++;\n+\t\t}\n \n-\tfill_bloom_keyvec_key(path, len, bloom_keyvec, 0,\n-\t\t\t      revs->bloom_filter_settings);\n-\tpath_component_nr = 1;\n+\t\tbloom_keyvec = create_bloom_keyvec(path_component_nr);\n+\t\trevs->bloom_keyvecs[i] = bloom_keyvec;\n+\n+\t\tfill_bloom_keyvec_key(path, len, bloom_keyvec, 0,\n+\t\t\t       revs->bloom_filter_settings);\n+\t\tpath_component_nr = 1;\n+\n+\t\tp = path + len - 1;\n+\t\twhile (p > path) {\n+\t\t\tif (*p == '/')\n+\t\t\t\tfill_bloom_keyvec_key(path, p - path,\n+\t\t\t\t\t       bloom_keyvec,\n+\t\t\t\t\t\t   path_component_nr++,\n+\t\t\t\t\t       revs->bloom_filter_settings);\n+\t\t\tp--;\n+\t\t}\n \n-\tp = path + len - 1;\n-\twhile (p > path) {\n-\t\tif (*p == '/')\n-\t\t\tfill_bloom_keyvec_key(path, p - path, bloom_keyvec,\n-\t\t\t\t\t      path_component_nr++,\n-\t\t\t\t\t      revs->bloom_filter_settings);\n-\t\tp--;\n+\t\tFREE_AND_NULL(path_alloc);\n \t}\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n@@ -760,14 +763,19 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \t\tbloom_filter_atexit_registered = 1;\n \t}\n \n+\treturn;\n+\n+fail:\n+\trevs->bloom_filter_settings = NULL;\n \tfree(path_alloc);\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\t\t\t\t\t struct commit *commit)\n {\n \tstruct bloom_filter *filter;\n-\tint result = 1, j;\n+\tint result = 0;\n \n \tif (!revs->repo->objects->commit_graph)\n \t\treturn -1;\n@@ -782,8 +790,11 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\treturn -1;\n \t}\n \n-\tresult = bloom_filter_contains_vec(filter, revs->bloom_keyvecs[0],\n-\t\t\t\t\t   revs->bloom_filter_settings);\n+\tfor (size_t nr = 0; !result && nr < revs->bloom_keyvecs_nr; nr++) {\n+\t\tresult = bloom_filter_contains_vec(filter,\n+\t\t\t\t\t\t   revs->bloom_keyvecs[nr],\n+\t\t\t\t\t\t   revs->bloom_filter_settings);\n+\t}\n \n \tif (result)\n \t\tcount_bloom_filter_maybe++;\n@@ -3201,6 +3212,14 @@ static void release_revisions_mailmap(struct string_list *mailmap)\n \n static void release_revisions_topo_walk_info(struct topo_walk_info *info);\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs)\n+{\n+\tfor (size_t nr = 0; nr < revs->bloom_keyvecs_nr; nr++)\n+\t\tdestroy_bloom_keyvec(revs->bloom_keyvecs[nr]);\n+\tFREE_AND_NULL(revs->bloom_keyvecs);\n+\trevs->bloom_keyvecs_nr = 0;\n+}\n+\n static void free_void_commit_list(void *list)\n {\n \tfree_commit_list(list);\n@@ -3229,11 +3248,7 @@ void release_revisions(struct rev_info *revs)\n \tclear_decoration(&revs->treesame, free);\n \tline_log_free(revs);\n \toidset_clear(&revs->missing_commits);\n-\n-\tfor (int i = 0; i < revs->bloom_keyvecs_nr; i++)\n-\t\tdestroy_bloom_keyvec(revs->bloom_keyvecs[i]);\n-\tFREE_AND_NULL(revs->bloom_keyvecs);\n-\trevs->bloom_keyvecs_nr = 0;\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 8910d53cac..46d1900a21 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -138,8 +138,8 @@ test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n '\n \n-test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n+test_expect_success 'git log -- multiple path specs use Bloom filters' '\n+\ttest_bloom_filters_used \"-- file4 A/file1\"\n '\n \n test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n@@ -151,9 +151,9 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n \ttest_bloom_filters_used \"-- *renamed\"\n '\n \n-test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- *\" &&\n-\ttest_bloom_filters_not_used \"-- file*\"\n+test_expect_success 'git log with wildcard that resolves to a multiple paths uses Bloom filters' '\n+\ttest_bloom_filters_used \"-- *\" &&\n+\ttest_bloom_filters_used \"-- file*\"\n '\n \n test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n-- \n2.50.0.108.g6ae0c543ae\n\n"},{"id":"520807","messageId":"xmqqy0td8fa9.fsf@gitster.g","threadId":"63691","inReplyTo":"20250625125541.3048632-3-502024330056@smail.nju.edu.cn","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-06-27T13:50:22Z","receivedAt":"2025-06-27T13:50:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lidong Yan <yldhome2d2@gmail.com> writes:\n\n> Remove `if (spec->nr > 1)` to enable bloom filter given multiple\n> pathspec. Wrapped for loop around code in prepare_to_use_bloom_filter()\n> to initialize each pathspec's struct bloom_keyvec. Add for loop\n> in check_maybe_different_in_bloom_filter() to find if at least one\n> pathspec's bloom_keyvec is contained in bloom filter.\n\nOy.  That's too dense enumeration but I suspect are all \"what the\npatch does\" that can be read from the diff.  The first sentence\ngives \"why\", which is excellent.\n\nYou'd need to check in forbid_bloom_filters() that none of the\npathspec items have magic (other than literal), not just the first\none, no?\n\nTotally outside the topic, but I wonder if we can further optimize\nby adding an early rejection using .nowildcard_len?  Instead of\nallowing a wildcarded \"dir/*\" pathspec element from disabling the\nBloom filter altogether, we could say \"dir/ is not possibly altered,\nso there may be dir/A, dir/B, etc., in the directory, nothing that\nwould match dir/* wildcard would have been modified\", couldn't we?\n\n> diff --git a/revision.c b/revision.c\n> index cf7dc3b3fa..8818f017f3 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -675,8 +675,6 @@ static int forbid_bloom_filters(struct pathspec *spec)\n>  {\n>  \tif (spec->has_wildcard)\n>  \t\treturn 1;\n> -\tif (spec->nr > 1)\n> -\t\treturn 1;\n>  \tif (spec->magic & ~PATHSPEC_LITERAL)\n>  \t\treturn 1;\n>  \tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n\nThis last check only looks at the first item.  It used to be OK\nbecause we didn't look at a pathspec with more than one element, but\nnow shouldn't we care?\n\nThanks.\n"},{"id":"520808","messageId":"4A56A595-55B1-4EEB-9B9E-3E9F7A9A74D4@smail.nju.edu.cn","threadId":"63691","inReplyTo":"xmqqy0td8fa9.fsf@gitster.g","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-06-27T14:24:10Z","receivedAt":"2025-06-27T14:24:39Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n> Lidong Yan <yldhome2d2@gmail.com> writes:\n> \n>> Remove `if (spec->nr > 1)` to enable bloom filter given multiple\n>> pathspec. Wrapped for loop around code in prepare_to_use_bloom_filter()\n>> to initialize each pathspec's struct bloom_keyvec. Add for loop\n>> in check_maybe_different_in_bloom_filter() to find if at least one\n>> pathspec's bloom_keyvec is contained in bloom filter.\n> \n> Oy.  That's too dense enumeration but I suspect are all \"what the\n> patch does\" that can be read from the diff.  The first sentence\n> gives \"why\", which is excellent.\n\nGot it. I will simply this log message in next version.\n\n> You'd need to check in forbid_bloom_filters() that none of the\n> pathspec items have magic (other than literal), not just the first\n> one, no?\n\nYeah, I never notice that. I would add checks in forbid_bloom_filters().\nAnd add test to ensure we don’t use bloom filter if any pathspec item is\nnot literal.\n\n> Totally outside the topic, but I wonder if we can further optimize\n> by adding an early rejection using .nowildcard_len?  Instead of\n> allowing a wildcarded \"dir/*\" pathspec element from disabling the\n> Bloom filter altogether, we could say \"dir/ is not possibly altered,\n> so there may be dir/A, dir/B, etc., in the directory, nothing that\n> would match dir/* wildcard would have been modified\", couldn't we?\n> \n\nI think it's feasible. In that case, we would need to add a condition\n.nowildcard_len > 0 to forbid_bloom_filter. I'm happy to write a new\npatch to address this issue.\n\n>> diff --git a/revision.c b/revision.c\n>> index cf7dc3b3fa..8818f017f3 100644\n>> --- a/revision.c\n>> +++ b/revision.c\n>> @@ -675,8 +675,6 @@ static int forbid_bloom_filters(struct pathspec *spec)\n>> {\n>> if (spec->has_wildcard)\n>> return 1;\n>> - if (spec->nr > 1)\n>> - return 1;\n>> if (spec->magic & ~PATHSPEC_LITERAL)\n>> return 1;\n>> if (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n> \n> This last check only looks at the first item.  It used to be OK\n> because we didn't look at a pathspec with more than one element, but\n> now shouldn't we care?\n\nDefinitely, I would fix that in 3rd version.\n\nThanks,\nLidong\n\n"},{"id":"520815","messageId":"xmqqbjq983aq.fsf@gitster.g","threadId":"63691","inReplyTo":"4A56A595-55B1-4EEB-9B9E-3E9F7A9A74D4@smail.nju.edu.cn","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-06-27T18:09:17Z","receivedAt":"2025-06-27T18:09:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lidong Yan <502024330056@smail.nju.edu.cn> writes:\n\n>> You'd need to check in forbid_bloom_filters() that none of the\n>> pathspec items have magic (other than literal), not just the first\n>> one, no?\n>\n> Yeah, I never notice that. I would add checks in forbid_bloom_filters().\n> And add test to ensure we don’t use bloom filter if any pathspec item is\n> not literal.\n\nSounds great.  I was wondering why your tests did not catch it.\n\n>> Totally outside the topic, but I wonder if we can further optimize\n>> by adding an early rejection using .nowildcard_len?  Instead of\n>> allowing a wildcarded \"dir/*\" pathspec element from disabling the\n>> Bloom filter altogether, we could say \"dir/ is not possibly altered,\n>> so there may be dir/A, dir/B, etc., in the directory, nothing that\n>> would match dir/* wildcard would have been modified\", couldn't we?\n>\n> I think it's feasible. In that case, we would need to add a condition\n> .nowildcard_len > 0 to forbid_bloom_filter. I'm happy to write a new\n> patch to address this issue.\n\nLet's leave it outside the topic and concentrate on the problem at\nhand first.\n\n"},{"id":"520827","messageId":"xmqqqzz47wd3.fsf@gitster.g","threadId":"63691","inReplyTo":"20250625125541.3048632-3-502024330056@smail.nju.edu.cn","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-06-27T20:39:04Z","receivedAt":"2025-06-27T20:39:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lidong Yan <yldhome2d2@gmail.com> writes:\n\nThis is a tangent, but I have to say that whoever wrote the original\ntest does not understand shells very well.  When you have files A\nand B in your working tree, to your $command, the following two does\nnot make any difference:\n\n\t$command ?\n\t$command A B\n\nIn fact it cannot even tell which form was used when composing the\ncommand line.  So this original test ...\n\n> -test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n> -\ttest_bloom_filters_not_used \"-- *\" &&\n> -\ttest_bloom_filters_not_used \"-- file*\"\n\n... is misleading to say the least.\n\n> +test_expect_success 'git log with wildcard that resolves to a multiple paths uses Bloom filters' '\n> +\ttest_bloom_filters_used \"-- *\" &&\n> +\ttest_bloom_filters_used \"-- file*\"\n>  '\n\nI think you should just retitle this to say\n\n\tgit log with multiple literal paths use Bloom filter\n\nor something.\n\nAlso the setup helper test_bloom_filters_{not_,}used helpers call is\nwritten in a way to make it impossible to pass a real wildcard and\nsee how \"$git log\" would behave, because it does this:\n\n\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n\nIt probably should use 'eval' so that the caller can pass a quoted\nwildcard, perhaps like\n\n    eval git -c core.commitgraph=false \\\n\t     log --pretty=format:%s \"$1\" >log_wo_bloom &&\n\nThen a test we can add to see how wildcards prevent Bloom from\nkicking in would look like\n\n\ttest_bloom_filters_used \"-- file*\" &&\n\ttest_bloom_filters_not_used \"-- file4 file\\*\" &&\n\nThe former lets the shell expand file* when the above \"eval\"\nevaluates its (concatenated) strings, while the latter leaves the\nbackslash before the asterisk in the strings fed to \"eval\", so the\n\"log\" will see a pathspec with wildcard.\n\nIf we were to fix that setup() thing, we of course need to be\na bit careful about existing tests.\n\nThanks.\n"},{"id":"520834","messageId":"1D8CE39A-D6F5-4AB7-8613-6C8DF1302907@smail.nju.edu.cn","threadId":"63691","inReplyTo":"xmqqqzz47wd3.fsf@gitster.g","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-06-28T02:54:09Z","receivedAt":"2025-06-28T02:54:40Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n> Also the setup helper test_bloom_filters_{not_,}used helpers call is\n> written in a way to make it impossible to pass a real wildcard and\n> see how \"$git log\" would behave, because it does this:\n> \n> git -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n\nYeah, if $1 contains * and because $1 is not quotes, * would trigger file name\nexpansion.\n\n> It probably should use 'eval' so that the caller can pass a quoted\n> wildcard, perhaps like\n> \n>    eval git -c core.commitgraph=false \\\n>     log --pretty=format:%s \"$1\" >log_wo_bloom &&\n> \n> Then a test we can add to see how wildcards prevent Bloom from\n> kicking in would look like\n> \n> test_bloom_filters_used \"-- file*\" &&\n> test_bloom_filters_not_used \"-- file4 file\\*\" &&\n> \n> The former lets the shell expand file* when the above \"eval\"\n> evaluates its (concatenated) strings, while the latter leaves the\n> backslash before the asterisk in the strings fed to \"eval\", so the\n> \"log\" will see a pathspec with wildcard.\n\nWow, this solution is very clever.\n\n> If we were to fix that setup() thing, we of course need to be\n> a bit careful about existing tests.\n\nThough the uses of test_bloom_filters_(not_)used are not too much,\nI think replace\n  git -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\nwith\n  git -c core.commitGraph=false log --pretty=\"format:%s” “$*\" >log_wo_bloom &&\n\nis not better than add the “eval …” solution, I will just use eval\n\nThanks,\nLidong\n\n"},{"id":"520835","messageId":"20250628042140.1097910-1-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250627062154.1121530-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v3 0/2] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-06-28T04:21:38Z","receivedAt":"2025-06-28T04:21:47Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"This series enables bloom filter optimization for multiple pathspec\nelements. Compared to v2, v3 fixed bugs in forbid_bloom_filter() and\nadd one more test case in t/t4216-log-bloom.sh.\n\nLidong Yan (2):\n  bloom: replace struct bloom_key * with struct bloom_keyvec\n  bloom: optimize multiple pathspec items in revision traversal\n\n bloom.c              |  31 +++++++++++\n bloom.h              |  20 +++++++\n revision.c           | 126 ++++++++++++++++++++++++-------------------\n revision.h           |   6 +--\n t/t4216-log-bloom.sh |  23 ++++----\n 5 files changed, 139 insertions(+), 67 deletions(-)\n\n-- \n2.50.0.108.g6ae0c543ae\n\n"},{"id":"520836","messageId":"20250628042140.1097910-2-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250627062154.1121530-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v3 1/2] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-06-28T04:21:39Z","receivedAt":"2025-06-28T04:21:49Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"The revision traversal limited by pathspec has optimization when\nthe pathspec has only one element. To support optimization for\nmultiple pathspec items, we need to modify the data structures\nin struct rev_info.\n\nstruct rev_info uses bloom_keys and bloom_nr to store the bloom keys\ncorresponding to a single pathspec item. To allow struct rev_info\nto store bloom keys for multiple pathspec items, a new data structure\n`struct bloom_keyvec` is introduced. Each `struct bloom_keyvec`\ncorresponds to a single pathspec item.\n\nIn `struct rev_info`, replace bloom_keys and bloom_nr with bloom_keyvecs\nand bloom_keyvec_nr. This commit still optimize one pathspec item, thus\nbloom_keyvec_nr can only be 0 or 1.\n\nNew *_bloom_keyvec functions are added to create and destroy a keyvec.\nbloom_filter_contains_vec() is added to check if all key in keyvec is\ncontained in a bloom filter. fill_bloom_keyvec_key() is added to\ninitialize a key in keyvec.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n bloom.c    | 31 +++++++++++++++++++++++++++++++\n bloom.h    | 20 ++++++++++++++++++++\n revision.c | 36 ++++++++++++++++++------------------\n revision.h |  6 +++---\n 4 files changed, 72 insertions(+), 21 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 0c8d2cebf9..8259cfce51 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -280,6 +280,25 @@ void deinit_bloom_filters(void)\n \tdeep_clear_bloom_filter_slab(&bloom_filters, free_one_bloom_filter);\n }\n \n+struct bloom_keyvec *create_bloom_keyvec(size_t count)\n+{\n+\tstruct bloom_keyvec *vec;\n+\tsize_t sz = sizeof(struct bloom_keyvec);\n+\tsz += count * sizeof(struct bloom_key);\n+\tvec = (struct bloom_keyvec *)xcalloc(1, sz);\n+\tvec->count = count;\n+\treturn vec;\n+}\n+\n+void destroy_bloom_keyvec(struct bloom_keyvec *vec)\n+{\n+\tif (!vec)\n+\t\treturn;\n+\tfor (size_t nr = 0; nr < vec->count; nr++)\n+\t\tclear_bloom_key(&vec->key[nr]);\n+\tfree(vec);\n+}\n+\n static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n \t\t       const struct hashmap_entry *eptr,\n \t\t       const struct hashmap_entry *entry_or_key,\n@@ -540,3 +559,15 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \n \treturn 1;\n }\n+\n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *vec,\n+\t\t\t      const struct bloom_filter_settings *settings)\n+{\n+\tint ret = 1;\n+\n+\tfor (size_t nr = 0; ret > 0 && nr < vec->count; nr++)\n+\t\tret = bloom_filter_contains(filter, &vec->key[nr], settings);\n+\n+\treturn ret;\n+}\ndiff --git a/bloom.h b/bloom.h\nindex 6e46489a20..9e4e832c8c 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -74,6 +74,11 @@ struct bloom_key {\n \tuint32_t *hashes;\n };\n \n+struct bloom_keyvec {\n+\tsize_t count;\n+\tstruct bloom_key key[FLEX_ARRAY];\n+};\n+\n int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n@@ -100,6 +105,17 @@ void add_key_to_filter(const struct bloom_key *key,\n void init_bloom_filters(void);\n void deinit_bloom_filters(void);\n \n+struct bloom_keyvec *create_bloom_keyvec(size_t count);\n+void destroy_bloom_keyvec(struct bloom_keyvec *vec);\n+\n+static inline void fill_bloom_keyvec_key(const char *data, size_t len,\n+\t\t\t\t\t struct bloom_keyvec *vec, size_t nr,\n+\t\t\t\t\t const struct bloom_filter_settings *settings)\n+{\n+\tassert(nr < vec->count);\n+\tfill_bloom_key(data, len, &vec->key[nr], settings);\n+}\n+\n enum bloom_filter_computed {\n \tBLOOM_NOT_COMPUTED = (1 << 0),\n \tBLOOM_COMPUTED     = (1 << 1),\n@@ -137,4 +153,8 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n \t\t\t  const struct bloom_filter_settings *settings);\n \n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *v,\n+\t\t\t      const struct bloom_filter_settings *settings);\n+\n #endif\ndiff --git a/revision.c b/revision.c\nindex afee111196..3aa544c137 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -688,6 +688,7 @@ static int forbid_bloom_filters(struct pathspec *spec)\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n+\tstruct bloom_keyvec *bloom_keyvec;\n \tchar *path_alloc = NULL;\n \tconst char *path, *p;\n \tsize_t len;\n@@ -736,19 +737,21 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \t\tp++;\n \t}\n \n-\trevs->bloom_keys_nr = path_component_nr;\n-\tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n+\trevs->bloom_keyvecs_nr = 1;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n+\tbloom_keyvec = create_bloom_keyvec(path_component_nr);\n+\trevs->bloom_keyvecs[0] = bloom_keyvec;\n \n-\tfill_bloom_key(path, len, &revs->bloom_keys[0],\n-\t\t       revs->bloom_filter_settings);\n+\tfill_bloom_keyvec_key(path, len, bloom_keyvec, 0,\n+\t\t\t      revs->bloom_filter_settings);\n \tpath_component_nr = 1;\n \n \tp = path + len - 1;\n \twhile (p > path) {\n \t\tif (*p == '/')\n-\t\t\tfill_bloom_key(path, p - path,\n-\t\t\t\t       &revs->bloom_keys[path_component_nr++],\n-\t\t\t\t       revs->bloom_filter_settings);\n+\t\t\tfill_bloom_keyvec_key(path, p - path, bloom_keyvec,\n+\t\t\t\t\t      path_component_nr++,\n+\t\t\t\t\t      revs->bloom_filter_settings);\n \t\tp--;\n \t}\n \n@@ -779,11 +782,8 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\treturn -1;\n \t}\n \n-\tfor (j = 0; result && j < revs->bloom_keys_nr; j++) {\n-\t\tresult = bloom_filter_contains(filter,\n-\t\t\t\t\t       &revs->bloom_keys[j],\n-\t\t\t\t\t       revs->bloom_filter_settings);\n-\t}\n+\tresult = bloom_filter_contains_vec(filter, revs->bloom_keyvecs[0],\n+\t\t\t\t\t   revs->bloom_filter_settings);\n \n \tif (result)\n \t\tcount_bloom_filter_maybe++;\n@@ -823,7 +823,7 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\treturn REV_TREE_SAME;\n \t}\n \n-\tif (revs->bloom_keys_nr && !nth_parent) {\n+\tif (revs->bloom_keyvecs_nr && !nth_parent) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \n \t\tif (bloom_ret == 0)\n@@ -850,7 +850,7 @@ static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit,\n \tif (!t1)\n \t\treturn 0;\n \n-\tif (!nth_parent && revs->bloom_keys_nr) {\n+\tif (!nth_parent && revs->bloom_keyvecs_nr) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \t\tif (!bloom_ret)\n \t\t\treturn 1;\n@@ -3230,10 +3230,10 @@ void release_revisions(struct rev_info *revs)\n \tline_log_free(revs);\n \toidset_clear(&revs->missing_commits);\n \n-\tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n-\t\tclear_bloom_key(&revs->bloom_keys[i]);\n-\tFREE_AND_NULL(revs->bloom_keys);\n-\trevs->bloom_keys_nr = 0;\n+\tfor (int i = 0; i < revs->bloom_keyvecs_nr; i++)\n+\t\tdestroy_bloom_keyvec(revs->bloom_keyvecs[i]);\n+\tFREE_AND_NULL(revs->bloom_keyvecs);\n+\trevs->bloom_keyvecs_nr = 0;\n }\n \n static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\ndiff --git a/revision.h b/revision.h\nindex 6d369cdad6..ac843f58d0 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -62,7 +62,7 @@ struct repository;\n struct rev_info;\n struct string_list;\n struct saved_parents;\n-struct bloom_key;\n+struct bloom_keyvec;\n struct bloom_filter_settings;\n struct option;\n struct parse_opt_ctx_t;\n@@ -360,8 +360,8 @@ struct rev_info {\n \n \t/* Commit graph bloom filter fields */\n \t/* The bloom filter key(s) for the pathspec */\n-\tstruct bloom_key *bloom_keys;\n-\tint bloom_keys_nr;\n+\tstruct bloom_keyvec **bloom_keyvecs;\n+\tint bloom_keyvecs_nr;\n \n \t/*\n \t * The bloom filter settings used to generate the key.\n-- \n2.50.0.108.g6ae0c543ae\n\n"},{"id":"520837","messageId":"20250628042140.1097910-3-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250627062154.1121530-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v3 2/2] bloom: optimize multiple pathspec items in revision traversal","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-06-28T04:21:40Z","receivedAt":"2025-06-28T04:21:51Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"To enable optimize multiple pathspec items in revision traversal,\nreturn 0 if all pathspec item is literal in forbid_bloom_filters().\nAdd code to initialize and check each pathspec item's bloom_keyvec.\n\nAdd new function release_revisions_bloom_keyvecs() to free all bloom\nkeyvec owned by rev_info.\n\nAdd new test cases in t/t4216-log-bloom.sh to ensure\n  - consistent results between the optimization for multiple pathspec\n    items using bloom filter and the case without bloom filter\n    optimization.\n  - does not use bloom filter if any pathspec item is not literal.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n revision.c           | 126 ++++++++++++++++++++++++-------------------\n t/t4216-log-bloom.sh |  23 ++++----\n 2 files changed, 85 insertions(+), 64 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 3aa544c137..8d73395f26 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -675,16 +675,17 @@ static int forbid_bloom_filters(struct pathspec *spec)\n {\n \tif (spec->has_wildcard)\n \t\treturn 1;\n-\tif (spec->nr > 1)\n-\t\treturn 1;\n \tif (spec->magic & ~PATHSPEC_LITERAL)\n \t\treturn 1;\n-\tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n-\t\treturn 1;\n+\tfor (size_t nr = 0; nr < spec->nr; nr++)\n+\t\tif (spec->items[nr].magic & ~PATHSPEC_LITERAL)\n+\t\t\treturn 1;\n \n \treturn 0;\n }\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n+\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n@@ -692,7 +693,7 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tchar *path_alloc = NULL;\n \tconst char *path, *p;\n \tsize_t len;\n-\tint path_component_nr = 1;\n+\tint path_component_nr;\n \n \tif (!revs->commits)\n \t\treturn;\n@@ -709,50 +710,53 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n-\tpi = &revs->pruning.pathspec.items[0];\n-\n-\t/* remove single trailing slash from path, if needed */\n-\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n-\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n-\t\tpath = path_alloc;\n-\t} else\n-\t\tpath = pi->match;\n-\n-\tlen = strlen(path);\n-\tif (!len) {\n-\t\trevs->bloom_filter_settings = NULL;\n-\t\tfree(path_alloc);\n-\t\treturn;\n-\t}\n-\n-\tp = path;\n-\twhile (*p) {\n-\t\t/*\n-\t\t * At this point, the path is normalized to use Unix-style\n-\t\t * path separators. This is required due to how the\n-\t\t * changed-path Bloom filters store the paths.\n-\t\t */\n-\t\tif (*p == '/')\n-\t\t\tpath_component_nr++;\n-\t\tp++;\n-\t}\n-\n-\trevs->bloom_keyvecs_nr = 1;\n-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n-\tbloom_keyvec = create_bloom_keyvec(path_component_nr);\n-\trevs->bloom_keyvecs[0] = bloom_keyvec;\n+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n+\t\tpi = &revs->pruning.pathspec.items[i];\n+\t\tpath_component_nr = 1;\n+\n+\t\t/* remove single trailing slash from path, if needed */\n+\t\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n+\t\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n+\t\t\tpath = path_alloc;\n+\t\t} else\n+\t\t\tpath = pi->match;\n+\n+\t\tlen = strlen(path);\n+\t\tif (!len)\n+\t\t\tgoto fail;\n+\n+\t\tp = path;\n+\t\twhile (*p) {\n+\t\t\t/*\n+\t\t\t * At this point, the path is normalized to use\n+\t\t\t * Unix-style path separators. This is required due to\n+\t\t\t * how the changed-path Bloom filters store the paths.\n+\t\t\t */\n+\t\t\tif (*p == '/')\n+\t\t\t\tpath_component_nr++;\n+\t\t\tp++;\n+\t\t}\n \n-\tfill_bloom_keyvec_key(path, len, bloom_keyvec, 0,\n-\t\t\t      revs->bloom_filter_settings);\n-\tpath_component_nr = 1;\n+\t\tbloom_keyvec = create_bloom_keyvec(path_component_nr);\n+\t\trevs->bloom_keyvecs[i] = bloom_keyvec;\n+\n+\t\tfill_bloom_keyvec_key(path, len, bloom_keyvec, 0,\n+\t\t\t       revs->bloom_filter_settings);\n+\t\tpath_component_nr = 1;\n+\n+\t\tp = path + len - 1;\n+\t\twhile (p > path) {\n+\t\t\tif (*p == '/')\n+\t\t\t\tfill_bloom_keyvec_key(path, p - path,\n+\t\t\t\t\t       bloom_keyvec,\n+\t\t\t\t\t\t   path_component_nr++,\n+\t\t\t\t\t       revs->bloom_filter_settings);\n+\t\t\tp--;\n+\t\t}\n \n-\tp = path + len - 1;\n-\twhile (p > path) {\n-\t\tif (*p == '/')\n-\t\t\tfill_bloom_keyvec_key(path, p - path, bloom_keyvec,\n-\t\t\t\t\t      path_component_nr++,\n-\t\t\t\t\t      revs->bloom_filter_settings);\n-\t\tp--;\n+\t\tFREE_AND_NULL(path_alloc);\n \t}\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n@@ -760,14 +764,19 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \t\tbloom_filter_atexit_registered = 1;\n \t}\n \n+\treturn;\n+\n+fail:\n+\trevs->bloom_filter_settings = NULL;\n \tfree(path_alloc);\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\t\t\t\t\t struct commit *commit)\n {\n \tstruct bloom_filter *filter;\n-\tint result = 1, j;\n+\tint result = 0;\n \n \tif (!revs->repo->objects->commit_graph)\n \t\treturn -1;\n@@ -782,8 +791,11 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\treturn -1;\n \t}\n \n-\tresult = bloom_filter_contains_vec(filter, revs->bloom_keyvecs[0],\n-\t\t\t\t\t   revs->bloom_filter_settings);\n+\tfor (size_t nr = 0; !result && nr < revs->bloom_keyvecs_nr; nr++) {\n+\t\tresult = bloom_filter_contains_vec(filter,\n+\t\t\t\t\t\t   revs->bloom_keyvecs[nr],\n+\t\t\t\t\t\t   revs->bloom_filter_settings);\n+\t}\n \n \tif (result)\n \t\tcount_bloom_filter_maybe++;\n@@ -3201,6 +3213,14 @@ static void release_revisions_mailmap(struct string_list *mailmap)\n \n static void release_revisions_topo_walk_info(struct topo_walk_info *info);\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs)\n+{\n+\tfor (size_t nr = 0; nr < revs->bloom_keyvecs_nr; nr++)\n+\t\tdestroy_bloom_keyvec(revs->bloom_keyvecs[nr]);\n+\tFREE_AND_NULL(revs->bloom_keyvecs);\n+\trevs->bloom_keyvecs_nr = 0;\n+}\n+\n static void free_void_commit_list(void *list)\n {\n \tfree_commit_list(list);\n@@ -3229,11 +3249,7 @@ void release_revisions(struct rev_info *revs)\n \tclear_decoration(&revs->treesame, free);\n \tline_log_free(revs);\n \toidset_clear(&revs->missing_commits);\n-\n-\tfor (int i = 0; i < revs->bloom_keyvecs_nr; i++)\n-\t\tdestroy_bloom_keyvec(revs->bloom_keyvecs[i]);\n-\tFREE_AND_NULL(revs->bloom_keyvecs);\n-\trevs->bloom_keyvecs_nr = 0;\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 8910d53cac..639868ac56 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -66,8 +66,9 @@ sane_unset GIT_TRACE2_CONFIG_PARAMS\n \n setup () {\n \trm -f \"$TRASH_DIRECTORY/trace.perf\" &&\n-\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n-\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n+\teval git -c core.commitGraph=false log --pretty=\"format:%s\" \"$1\" >log_wo_bloom &&\n+\teval \"GIT_TRACE2_PERF=\\\"$TRASH_DIRECTORY/trace.perf\\\"\" \\\n+\t\tgit -c core.commitGraph=true log --pretty=\"format:%s\" \"$1\" >log_w_bloom\n }\n \n test_bloom_filters_used () {\n@@ -138,10 +139,6 @@ test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n '\n \n-test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n-'\n-\n test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"-- .\"\n '\n@@ -151,9 +148,17 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n \ttest_bloom_filters_used \"-- *renamed\"\n '\n \n-test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- *\" &&\n-\ttest_bloom_filters_not_used \"-- file*\"\n+test_expect_success 'git log with multiple literal paths uses Bloom filter' '\n+\ttest_bloom_filters_used \"-- file4 A/file1\" &&\n+\ttest_bloom_filters_used \"-- *\" &&\n+\ttest_bloom_filters_used \"-- file*\"\n+'\n+\n+test_expect_success 'git log with path contains a wildcard does not use Bloom filter' '\n+\ttest_bloom_filters_not_used \"-- file\\*\" &&\n+\ttest_bloom_filters_not_used \"-- A/\\* file4\" &&\n+\ttest_bloom_filters_not_used \"-- file4 A/\\*\" &&\n+\ttest_bloom_filters_not_used \"-- * A/\\*\"\n '\n \n test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n-- \n2.50.0.108.g6ae0c543ae\n\n"},{"id":"520988","messageId":"C8E0D62E-11B1-4921-AD4C-2905F10E07B6@gmail.com","threadId":"63691","inReplyTo":"xmqqy0td8fa9.fsf@gitster.g","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-01T05:52:26Z","receivedAt":"2025-07-01T05:52:37Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n> Totally outside the topic, but I wonder if we can further optimize\n> by adding an early rejection using .nowildcard_len?  Instead of\n> allowing a wildcarded \"dir/*\" pathspec element from disabling the\n> Bloom filter altogether, we could say \"dir/ is not possibly altered,\n> so there may be dir/A, dir/B, etc., in the directory, nothing that\n> would match dir/* wildcard would have been modified\", couldn't we?\n\nI think, except for PATHSPEC_EXCLUDE, all other pathspec magic flags\ncould potentially be optimized using .nowildcard_len by restricting checks to\njust the dir/ part of each pathspec item.\n\nHere;s are all possible pathspec magic\n#define PATHSPEC_FROMTOP\t(1<<0)\n#define PATHSPEC_MAXDEPTH\t(1<<1)\n#define PATHSPEC_LITERAL\t(1<<2)\n#define PATHSPEC_GLOB\t\t(1<<3)\n#define PATHSPEC_ICASE\t\t(1<<4)\n#define PATHSPEC_EXCLUDE\t(1<<5)\n#define PATHSPEC_ATTR\t\t(1<<6)\n\n\n"},{"id":"520998","messageId":"aGOhY2YuJZNG8ovj@szeder.dev","threadId":"63691","inReplyTo":"xmqqy0td8fa9.fsf@gitster.g","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2025-07-01T08:50:43Z","receivedAt":"2025-07-01T08:50:58Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Fri, Jun 27, 2025 at 06:50:22AM -0700, Junio C Hamano wrote:\n> Totally outside the topic, but I wonder if we can further optimize\n> by adding an early rejection using .nowildcard_len?  Instead of\n> allowing a wildcarded \"dir/*\" pathspec element from disabling the\n> Bloom filter altogether, we could say \"dir/ is not possibly altered,\n> so there may be dir/A, dir/B, etc., in the directory, nothing that\n> would match dir/* wildcard would have been modified\", couldn't we?\n\nIndeed, that's what I demonstrated back in:\n\n  https://public-inbox.org/git/20200529085038.26008-35-szeder.dev@gmail.com/\n\n"},{"id":"521017","messageId":"BBAAC895-B24B-47DB-87DA-2276B645830A@gmail.com","threadId":"63691","inReplyTo":"aGOhY2YuJZNG8ovj@szeder.dev","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-01T11:40:01Z","receivedAt":"2025-07-01T11:40:14Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"SZEDER Gábor <szeder.dev@gmail.com> writes:\n> \n> On Fri, Jun 27, 2025 at 06:50:22AM -0700, Junio C Hamano wrote:\n>> Totally outside the topic, but I wonder if we can further optimize\n>> by adding an early rejection using .nowildcard_len?  Instead of\n>> allowing a wildcarded \"dir/*\" pathspec element from disabling the\n>> Bloom filter altogether, we could say \"dir/ is not possibly altered,\n>> so there may be dir/A, dir/B, etc., in the directory, nothing that\n>> would match dir/* wildcard would have been modified\", couldn't we?\n> \n> Indeed, that's what I demonstrated back in:\n> \n>  https://public-inbox.org/git/20200529085038.26008-35-szeder.dev@gmail.com/\n\nThat's interesting. Though I find the bloom part of code changed and I can't\nreuse your patch.\n\nHave you ever considered to optimize other kind of pathspec magic? I am\nnot perfectly sure whether it is feasible to use the same trick (using bloom filter on dir/ path)\non patchspec magic except PATHSPEC_EXCLUDE.\n\nThanks,\nLidong\n\n"},{"id":"521067","messageId":"xmqqo6u4kkg0.fsf@gitster.g","threadId":"63691","inReplyTo":"C8E0D62E-11B1-4921-AD4C-2905F10E07B6@gmail.com","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-01T15:19:27Z","receivedAt":"2025-07-01T15:19:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lidong Yan <yldhome2d2@gmail.com> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>> Totally outside the topic, but I wonder if we can further optimize\n>> by adding an early rejection using .nowildcard_len?  Instead of\n>> allowing a wildcarded \"dir/*\" pathspec element from disabling the\n>> Bloom filter altogether, we could say \"dir/ is not possibly altered,\n>> so there may be dir/A, dir/B, etc., in the directory, nothing that\n>> would match dir/* wildcard would have been modified\", couldn't we?\n>\n> I think, except for PATHSPEC_EXCLUDE, all other pathspec magic flags\n> could potentially be optimized using .nowildcard_len by restricting checks to\n> just the dir/ part of each pathspec item.\n\nA good observation.\n\nI do not know about icase; though.  Asking about \"Dir/Path\" and\ngetting \"Dir/ or Dir/Path cannot possibly be in the set of paths\nthat were modified\" from the changed-path Bloom filter would not\nhelp us optimize the tree comparison out, when we do not want to\nmiss modifications for \"dir/path\".\n\n> Here;s are all possible pathspec magic\n> #define PATHSPEC_FROMTOP\t(1<<0)\n> #define PATHSPEC_MAXDEPTH\t(1<<1)\n> #define PATHSPEC_LITERAL\t(1<<2)\n> #define PATHSPEC_GLOB\t\t(1<<3)\n> #define PATHSPEC_ICASE\t\t(1<<4)\n> #define PATHSPEC_EXCLUDE\t(1<<5)\n> #define PATHSPEC_ATTR\t\t(1<<6)\n\n"},{"id":"521070","messageId":"xmqqfrfflxw8.fsf@gitster.g","threadId":"63691","inReplyTo":"aGOhY2YuJZNG8ovj@szeder.dev","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-01T15:43:35Z","receivedAt":"2025-07-01T15:43:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"SZEDER Gábor <szeder.dev@gmail.com> writes:\n\n> On Fri, Jun 27, 2025 at 06:50:22AM -0700, Junio C Hamano wrote:\n>> Totally outside the topic, but I wonder if we can further optimize\n>> by adding an early rejection using .nowildcard_len?  Instead of\n>> allowing a wildcarded \"dir/*\" pathspec element from disabling the\n>> Bloom filter altogether, we could say \"dir/ is not possibly altered,\n>> so there may be dir/A, dir/B, etc., in the directory, nothing that\n>> would match dir/* wildcard would have been modified\", couldn't we?\n>\n> Indeed, that's what I demonstrated back in:\n>\n>   https://public-inbox.org/git/20200529085038.26008-35-szeder.dev@gmail.com/\n\nHeh, I am always impressed that some people seem to have infinitely\nlong scrollback buffer ;-)  It is curious why nobody else noticed\nand advocated for your patch back then.\n\n"},{"id":"521148","messageId":"D5CB9B7A-C7B2-4F5A-B358-8F46A4E18CDB@gmail.com","threadId":"63691","inReplyTo":"xmqqo6u4kkg0.fsf@gitster.g","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-02T07:14:26Z","receivedAt":"2025-07-02T07:14:45Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n> \n> I do not know about icase; though.  Asking about \"Dir/Path\" and\n> getting \"Dir/ or Dir/Path cannot possibly be in the set of paths\n> that were modified\" from the changed-path Bloom filter would not\n> help us optimize the tree comparison out, when we do not want to\n> miss modifications for \"dir/path\".\n\nMake sense, both PATHSPEC_EXCLUDE and PATHSPEC_ICASE shouldn’t be\noptimized by bloom filter.\n\nI found that my [PATCH v3 2/2] contains two unaligned parameters. Should I reroll\nthis patch and introduce the nowildcard_len change in a separate commit?\n\n> \n>> Here;s are all possible pathspec magic\n>> #define PATHSPEC_FROMTOP (1<<0)\n>> #define PATHSPEC_MAXDEPTH (1<<1)\n>> #define PATHSPEC_LITERAL (1<<2)\n>> #define PATHSPEC_GLOB (1<<3)\n>> #define PATHSPEC_ICASE (1<<4)\n>> #define PATHSPEC_EXCLUDE (1<<5)\n>> #define PATHSPEC_ATTR (1<<6)\n\n"},{"id":"521176","messageId":"aGVLZ9VUf2M1sWhL@pks.im","threadId":"63691","inReplyTo":"20250628042140.1097910-2-502024330056@smail.nju.edu.cn","subject":"Re: [PATCH v3 1/2] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-07-02T15:08:23Z","receivedAt":"2025-07-02T15:08:30Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sat, Jun 28, 2025 at 12:21:39PM +0800, Lidong Yan wrote:\n> diff --git a/bloom.h b/bloom.h\n> index 6e46489a20..9e4e832c8c 100644\n> --- a/bloom.h\n> +++ b/bloom.h\n> @@ -74,6 +74,11 @@ struct bloom_key {\n>  \tuint32_t *hashes;\n>  };\n>  \n> +struct bloom_keyvec {\n> +\tsize_t count;\n> +\tstruct bloom_key key[FLEX_ARRAY];\n> +};\n> +\n\nA short comment would help readers understand what the intent of this\ndata structure is.\n\n>  int load_bloom_filter_from_graph(struct commit_graph *g,\n>  \t\t\t\t struct bloom_filter *filter,\n>  \t\t\t\t uint32_t graph_pos);\n> @@ -100,6 +105,17 @@ void add_key_to_filter(const struct bloom_key *key,\n>  void init_bloom_filters(void);\n>  void deinit_bloom_filters(void);\n>  \n> +struct bloom_keyvec *create_bloom_keyvec(size_t count);\n> +void destroy_bloom_keyvec(struct bloom_keyvec *vec);\n\nThese functions are named very unusually for us -- the first version of\nthis patch series was following our coding guidelines, but this version\nhere isn't anymore.\n\n - The primary data structure that a subsystem 'S' deals with is called\n   `struct S`. Functions that operate on `struct S` are named\n   `S_<verb>()` and should generally receive a pointer to `struct S` as\n   first parameter. E.g.\n\nSecond, the functions should probably be called `*_new()` and `*_free()`\ninstead of `create_*()` and `destroy_*()`.\n\n> +static inline void fill_bloom_keyvec_key(const char *data, size_t len,\n> +\t\t\t\t\t struct bloom_keyvec *vec, size_t nr,\n> +\t\t\t\t\t const struct bloom_filter_settings *settings)\n> +{\n> +\tassert(nr < vec->count);\n> +\tfill_bloom_key(data, len, &vec->key[nr], settings);\n> +}\n> +\n\nSimilarly, this should probably be called `bloom_keyvec_fill_key()`.\n\n>  enum bloom_filter_computed {\n>  \tBLOOM_NOT_COMPUTED = (1 << 0),\n>  \tBLOOM_COMPUTED     = (1 << 1),\n> @@ -137,4 +153,8 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n>  \t\t\t  const struct bloom_key *key,\n>  \t\t\t  const struct bloom_filter_settings *settings);\n>  \n> +int bloom_filter_contains_vec(const struct bloom_filter *filter,\n> +\t\t\t      const struct bloom_keyvec *v,\n> +\t\t\t      const struct bloom_filter_settings *settings);\n> +\n>  #endif\n\nThis one looks alright though.\n\n> diff --git a/revision.c b/revision.c\n> index afee111196..3aa544c137 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -779,11 +782,8 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n>  \t\treturn -1;\n>  \t}\n>  \n> -\tfor (j = 0; result && j < revs->bloom_keys_nr; j++) {\n> -\t\tresult = bloom_filter_contains(filter,\n> -\t\t\t\t\t       &revs->bloom_keys[j],\n> -\t\t\t\t\t       revs->bloom_filter_settings);\n> -\t}\n> +\tresult = bloom_filter_contains_vec(filter, revs->bloom_keyvecs[0],\n> +\t\t\t\t\t   revs->bloom_filter_settings);\n>  \n>  \tif (result)\n>  \t\tcount_bloom_filter_maybe++;\n\nThis conversion feels wrong to me. Why don't we end up iterating through\n`revs->bloom_keyvecs_nr` here?  We do indeed change it back in the next\npatch to use a for loop.\n\n> @@ -3230,10 +3230,10 @@ void release_revisions(struct rev_info *revs)\n>  \tline_log_free(revs);\n>  \toidset_clear(&revs->missing_commits);\n>  \n> -\tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n> -\t\tclear_bloom_key(&revs->bloom_keys[i]);\n> -\tFREE_AND_NULL(revs->bloom_keys);\n> -\trevs->bloom_keys_nr = 0;\n> +\tfor (int i = 0; i < revs->bloom_keyvecs_nr; i++)\n\nIt's puzzling that the number of keys is declared as `int`. It's not an\nissue introduced by you, but can we maybe fix it while at it?\n\nPatrick\n"},{"id":"521180","messageId":"xmqq1pqyfvb2.fsf@gitster.g","threadId":"63691","inReplyTo":"D5CB9B7A-C7B2-4F5A-B358-8F46A4E18CDB@gmail.com","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-02T15:48:17Z","receivedAt":"2025-07-02T15:48:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lidong Yan <yldhome2d2@gmail.com> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>> \n>> I do not know about icase; though.  Asking about \"Dir/Path\" and\n>> getting \"Dir/ or Dir/Path cannot possibly be in the set of paths\n>> that were modified\" from the changed-path Bloom filter would not\n>> help us optimize the tree comparison out, when we do not want to\n>> miss modifications for \"dir/path\".\n>\n> Make sense, both PATHSPEC_EXCLUDE and PATHSPEC_ICASE shouldn’t be\n> optimized by bloom filter.\n\nBefore concluding so, we may want to double check how Bloom filters\nare built on case insensitive systems, though.  If we normalize the\nstring by downcasing before murmuring the string, the resulting\nBloom filter may have more false positives for those who want to\n(ab)use it to optimize case sensitive queries (without affecting\ncorrectness), but case insensitive queries would be helped.  I do\nnot think we support (or want to support) a repository that spans\nacross two filesystems with different case sensitivity, so those who\nworked on our changed-path Bloom filter subsystem may have already\nplaced such an optimization, based on the case sensitivity recorded\nin the repository (core.ignorecase).\n\n> I found that my [PATCH v3 2/2] contains two unaligned parameters. Should I reroll\n> this patch and introduce the nowildcard_len change in a separate commit?\n\nUpdating a patch with a fix to obvious known problems is good.\n\nExtending the scope of the series should be left out for a new\nseparate commit.  It may even be a better idea to hold it while\nthe current set of patches are still being polished, and then sent\nout as a new series after the dust settles (even if you internally\ndeveloped that part as a direct extension to the current effort).\n\nThanks.\n"},{"id":"521181","messageId":"9B2AC9DD-1462-4B6D-B2D5-B2FEEA70B4C3@smail.nju.edu.cn","threadId":"63691","inReplyTo":"aGVLZ9VUf2M1sWhL@pks.im","subject":"Re: [PATCH v3 1/2] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-07-02T15:49:04Z","receivedAt":"2025-07-02T15:49:32Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n> \n> On Sat, Jun 28, 2025 at 12:21:39PM +0800, Lidong Yan wrote:\n>> diff --git a/bloom.h b/bloom.h\n>> index 6e46489a20..9e4e832c8c 100644\n>> --- a/bloom.h\n>> +++ b/bloom.h\n>> @@ -74,6 +74,11 @@ struct bloom_key {\n>> uint32_t *hashes;\n>> };\n>> \n>> +struct bloom_keyvec {\n>> + size_t count;\n>> + struct bloom_key key[FLEX_ARRAY];\n>> +};\n>> +\n> \n> A short comment would help readers understand what the intent of this\n> data structure is.\n\nUnderstood, will add comment for struct bloom_keyvec in v4.\n\n> \n>> int load_bloom_filter_from_graph(struct commit_graph *g,\n>> struct bloom_filter *filter,\n>> uint32_t graph_pos);\n>> @@ -100,6 +105,17 @@ void add_key_to_filter(const struct bloom_key *key,\n>> void init_bloom_filters(void);\n>> void deinit_bloom_filters(void);\n>> \n>> +struct bloom_keyvec *create_bloom_keyvec(size_t count);\n>> +void destroy_bloom_keyvec(struct bloom_keyvec *vec);\n> \n> These functions are named very unusually for us -- the first version of\n> this patch series was following our coding guidelines, but this version\n> here isn't anymore.\n> \n> - The primary data structure that a subsystem 'S' deals with is called\n>   `struct S`. Functions that operate on `struct S` are named\n>   `S_<verb>()` and should generally receive a pointer to `struct S` as\n>   first parameter. E.g.\n> \n> Second, the functions should probably be called `*_new()` and `*_free()`\n> instead of `create_*()` and `destroy_*()`.\n\nThough I think create and destroy doesn’t match the 'operate on struct S’\ndefinition, I will rename these function to *_verb in v4.\n\n> \n>> +static inline void fill_bloom_keyvec_key(const char *data, size_t len,\n>> + struct bloom_keyvec *vec, size_t nr,\n>> + const struct bloom_filter_settings *settings)\n>> +{\n>> + assert(nr < vec->count);\n>> + fill_bloom_key(data, len, &vec->key[nr], settings);\n>> +}\n>> +\n> \n> Similarly, this should probably be called `bloom_keyvec_fill_key()`.\n\nI initially wanted the new function to have a name similar to fill_bloom_key,\nbut perhaps I should consider renaming it.\n\n> \n>> enum bloom_filter_computed {\n>> BLOOM_NOT_COMPUTED = (1 << 0),\n>> BLOOM_COMPUTED     = (1 << 1),\n>> @@ -137,4 +153,8 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n>>  const struct bloom_key *key,\n>>  const struct bloom_filter_settings *settings);\n>> \n>> +int bloom_filter_contains_vec(const struct bloom_filter *filter,\n>> +      const struct bloom_keyvec *v,\n>> +      const struct bloom_filter_settings *settings);\n>> +\n>> #endif\n> \n> This one looks alright though.\n> \n>> diff --git a/revision.c b/revision.c\n>> index afee111196..3aa544c137 100644\n>> --- a/revision.c\n>> +++ b/revision.c\n>> @@ -779,11 +782,8 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n>> return -1;\n>> }\n>> \n>> - for (j = 0; result && j < revs->bloom_keys_nr; j++) {\n>> - result = bloom_filter_contains(filter,\n>> -       &revs->bloom_keys[j],\n>> -       revs->bloom_filter_settings);\n>> - }\n>> + result = bloom_filter_contains_vec(filter, revs->bloom_keyvecs[0],\n>> +   revs->bloom_filter_settings);\n>> \n>> if (result)\n>> count_bloom_filter_maybe++;\n> \n> This conversion feels wrong to me. Why don't we end up iterating through\n> `revs->bloom_keyvecs_nr` here?  We do indeed change it back in the next\n> patch to use a for loop.\n\nMy original intention was to include all the loop-related logic in [PATCH 2/2].\nHowever, the lack of a loop here does make the code look error-prone. I will\nadd the loop in v4.\n\n> \n>> @@ -3230,10 +3230,10 @@ void release_revisions(struct rev_info *revs)\n>> line_log_free(revs);\n>> oidset_clear(&revs->missing_commits);\n>> \n>> - for (int i = 0; i < revs->bloom_keys_nr; i++)\n>> - clear_bloom_key(&revs->bloom_keys[i]);\n>> - FREE_AND_NULL(revs->bloom_keys);\n>> - revs->bloom_keys_nr = 0;\n>> + for (int i = 0; i < revs->bloom_keyvecs_nr; i++)\n> \n> It's puzzling that the number of keys is declared as `int`. It's not an\n> issue introduced by you, but can we maybe fix it while at it?\n\nYes, of course.\n\nThanks for your review,\nLidong\n\n"},{"id":"521193","messageId":"xmqqy0t6curr.fsf@gitster.g","threadId":"63691","inReplyTo":"aGVLZ9VUf2M1sWhL@pks.im","subject":"Re: [PATCH v3 1/2] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-02T18:28:08Z","receivedAt":"2025-07-02T18:28:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>> +static inline void fill_bloom_keyvec_key(const char *data, size_t len,\n>> +\t\t\t\t\t struct bloom_keyvec *vec, size_t nr,\n>> +\t\t\t\t\t const struct bloom_filter_settings *settings)\n>> +{\n>> +\tassert(nr < vec->count);\n>> +\tfill_bloom_key(data, len, &vec->key[nr], settings);\n>> +}\n>> +\n>\n> Similarly, this should probably be called `bloom_keyvec_fill_key()`.\n\nIf so, a preliminary clean-up patch in front of the series is in\norder, as the <bloom.h> header file, without these patches, already\nhas the follwoing external API functions and structures declared,\nthat do not follow your naming rules at all (I have removed the ones\nthat begin with \"bloom_\" from the below):\n\n    int load_bloom_filter_from_graph()\n    uint32_t murmur3_seeded_v2();\n    void fill_bloom_key();\n    void clear_bloom_key(s);\n    void add_key_to_filter();\n    void init_bloom_filters(void);\n    void deinit_bloom_filters(void);\n    struct bloom_filter *get_or_compute_bloom_filter();\n    struct bloom_filter *get_bloom_filter();\n\nIt is very dubious that murmur3_seeded_v2() is exposed (nobody would\nknow it is for Bloom filter subsystem from that name); as far as I\ncan tell, it is only needed for t/helper testing, and makes me\nwonder if we can come up with a better division between the\nproduction code and t/helper/ code around there.\n\nThanks.\n\n"},{"id":"521222","messageId":"B02F0A96-0D2F-41E8-A3CE-A840024092ED@smail.nju.edu.cn","threadId":"63691","inReplyTo":"xmqqy0t6curr.fsf@gitster.g","subject":"Re: [PATCH v3 1/2] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-07-03T01:41:59Z","receivedAt":"2025-07-03T01:42:38Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> It is very dubious that murmur3_seeded_v2() is exposed (nobody would\n> know it is for Bloom filter subsystem from that name); as far as I\n> can tell, it is only needed for t/helper testing, and makes me\n> wonder if we can come up with a better division between the\n> production code and t/helper/ code around there.\n> \n> Thanks.\n\n\nMaybe we can do something like this:\n    struct bloom_filter_settings settings;\n    struct bloom_key key;\n    uint32_t hash0;\n\n    settings->num_hashes = 1;\n    settings->hash_version = 2;\n    fill_bloom_key(argv[2], strlen(argv[2]), &key, &setting);\n    hash0 = key->hash[0];\n    clear_bloom_key(&key);\n\n    return hash0;\n\nIn t/helper, so that we don’t need to export murmur3_seeded_v2()."},{"id":"521223","messageId":"2E8CA6E5-0A2C-4470-A1C0-BE7D72B36DD8@smail.nju.edu.cn","threadId":"63691","inReplyTo":"xmqq1pqyfvb2.fsf@gitster.g","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-07-03T01:52:51Z","receivedAt":"2025-07-03T01:53:08Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n> \n> Before concluding so, we may want to double check how Bloom filters\n> are built on case insensitive systems, though.  If we normalize the\n> string by downcasing before murmuring the string, the resulting\n> Bloom filter may have more false positives for those who want to\n> (ab)use it to optimize case sensitive queries (without affecting\n> correctness), but case insensitive queries would be helped.  I do\n> not think we support (or want to support) a repository that spans\n> across two filesystems with different case sensitivity, so those who\n> worked on our changed-path Bloom filter subsystem may have already\n> placed such an optimization, based on the case sensitivity recorded\n> in the repository (core.ignorecase).\n\nI understand. I should check whether commit graph file's change path\nbloom filter is case sensitive. If the change path bloom filter is case insensitive,\nwe could optimize PATHSPEC_ICASE as well.\n\n> \n> Updating a patch with a fix to obvious known problems is good.\n> \n> Extending the scope of the series should be left out for a new\n> separate commit.  It may even be a better idea to hold it while\n> the current set of patches are still being polished, and then sent\n> out as a new series after the dust settles (even if you internally\n> developed that part as a direct extension to the current effort).\n\nGot it. So for now I should just polish this current patch. \n\nThanks,\nLidong\n\n"},{"id":"521305","messageId":"20250704111437.2660251-1-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250628042140.1097910-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v4 0/4] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-04T11:14:33Z","receivedAt":"2025-07-04T11:14:49Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"This series enables bloom filter optimization for multiple pathspec\nelements. Compared to v3, v4 rename *_bloom_key and *_bloom_keyvec methods\nto bloom_key_* and bloom_keyvec_*, respectively, to follow the git code\nguidelines. Also, it adds a new test helper to return murmur3 hash.\n\nLidong Yan (4):\n  bloom: add test helper to return murmur3 hash\n  bloom: rename function operates on bloom_key\n  bloom: replace struct bloom_key * with struct bloom_keyvec\n  bloom: optimize multiple pathspec items in revision traversal\n\n blame.c               |   2 +-\n bloom.c               |  52 +++++++++++++++--\n bloom.h               |  41 +++++++++----\n line-log.c            |   4 +-\n revision.c            | 131 +++++++++++++++++++++++-------------------\n revision.h            |   6 +-\n t/helper/test-bloom.c |   8 +--\n t/t4216-log-bloom.sh  |  23 +++++---\n 8 files changed, 174 insertions(+), 93 deletions(-)\n\n-- \n2.50.0.107.g33b6ec8c79\n\n"},{"id":"521306","messageId":"20250704111437.2660251-2-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250704111437.2660251-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v4 1/4] bloom: add test helper to return murmur3 hash","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-04T11:14:34Z","receivedAt":"2025-07-04T11:14:54Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"In bloom.h, murmur3_seeded_v2() is exported for the use of test murmur3\nhash. To clarify that murmur3_seeded_v2() is exported solely for testing\npurposes, a new helper function test_murmur3_seeded() was added instead\nof exporting murmur3_seeded_v2() directly.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n bloom.c               | 13 ++++++++++++-\n bloom.h               | 12 +++---------\n t/helper/test-bloom.c |  4 ++--\n 3 files changed, 17 insertions(+), 12 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 0c8d2cebf9..946c5e8c98 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -107,7 +107,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n  * Not considered to be cryptographically secure.\n  * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n  */\n-uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len)\n+static uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len)\n {\n \tconst uint32_t c1 = 0xcc9e2d51;\n \tconst uint32_t c2 = 0x1b873593;\n@@ -540,3 +540,14 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \n \treturn 1;\n }\n+\n+uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n+\t\t\t\t   int version)\n+{\n+\tassert(version == 1 || version == 2);\n+\n+\tif (version == 2)\n+\t\treturn murmur3_seeded_v2(seed, data, len);\n+\telse\n+\t\treturn murmur3_seeded_v1(seed, data, len);\n+}\ndiff --git a/bloom.h b/bloom.h\nindex 6e46489a20..a9ded1822f 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -78,15 +78,6 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n \n-/*\n- * Calculate the murmur3 32-bit hash value for the given data\n- * using the given seed.\n- * Produces a uniformly distributed hash value.\n- * Not considered to be cryptographically secure.\n- * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n- */\n-uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len);\n-\n void fill_bloom_key(const char *data,\n \t\t    size_t len,\n \t\t    struct bloom_key *key,\n@@ -137,4 +128,7 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n \t\t\t  const struct bloom_filter_settings *settings);\n \n+uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n+\t\t\t\t   int version);\n+\n #endif\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 9aa2c5a592..6a24b6e0a6 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -61,13 +61,13 @@ int cmd__bloom(int argc, const char **argv)\n \t\tuint32_t hashed;\n \t\tif (argc < 3)\n \t\t\tusage(bloom_usage);\n-\t\thashed = murmur3_seeded_v2(0, argv[2], strlen(argv[2]));\n+\t\thashed = test_bloom_murmur3_seeded(0, argv[2], strlen(argv[2]), 2);\n \t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n \t}\n \n \tif (!strcmp(argv[1], \"get_murmur3_seven_highbit\")) {\n \t\tuint32_t hashed;\n-\t\thashed = murmur3_seeded_v2(0, \"\\x99\\xaa\\xbb\\xcc\\xdd\\xee\\xff\", 7);\n+\t\thashed = test_bloom_murmur3_seeded(0, \"\\x99\\xaa\\xbb\\xcc\\xdd\\xee\\xff\", 7, 2);\n \t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n \t}\n \n-- \n2.50.0.107.g33b6ec8c79\n\n"},{"id":"521307","messageId":"20250704111437.2660251-3-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250704111437.2660251-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v4 2/4] bloom: rename function operates on bloom_key","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-04T11:14:35Z","receivedAt":"2025-07-04T11:14:58Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"git code style requires that functions operating on a struct S\nshould be named in the form S_verb. However, the functions operating\non struct bloom_key do not follow this convention. Therefore,\nfill_bloom_key() and clear_bloom_key() are renamed to bloom_key_fill()\nand bloom_key_clear(), respectively.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n blame.c               | 2 +-\n bloom.c               | 8 ++++----\n bloom.h               | 4 ++--\n line-log.c            | 4 ++--\n revision.c            | 6 +++---\n t/helper/test-bloom.c | 4 ++--\n 6 files changed, 14 insertions(+), 14 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex 57daa45e89..459043a511 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -1310,7 +1310,7 @@ static void add_bloom_key(struct blame_bloom_data *bd,\n \t}\n \n \tbd->keys[bd->nr] = xmalloc(sizeof(struct bloom_key));\n-\tfill_bloom_key(path, strlen(path), bd->keys[bd->nr], bd->settings);\n+\tbloom_key_fill(path, strlen(path), bd->keys[bd->nr], bd->settings);\n \tbd->nr++;\n }\n \ndiff --git a/bloom.c b/bloom.c\nindex 946c5e8c98..35ff36c31c 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -221,7 +221,7 @@ static uint32_t murmur3_seeded_v1(uint32_t seed, const char *data, size_t len)\n \treturn seed;\n }\n \n-void fill_bloom_key(const char *data,\n+void bloom_key_fill(const char *data,\n \t\t    size_t len,\n \t\t    struct bloom_key *key,\n \t\t    const struct bloom_filter_settings *settings)\n@@ -243,7 +243,7 @@ void fill_bloom_key(const char *data,\n \t\tkey->hashes[i] = hash0 + i * hash1;\n }\n \n-void clear_bloom_key(struct bloom_key *key)\n+void bloom_key_clear(struct bloom_key *key)\n {\n \tFREE_AND_NULL(key->hashes);\n }\n@@ -500,9 +500,9 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \n \t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n \t\t\tstruct bloom_key key;\n-\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, settings);\n+\t\t\tbloom_key_fill(e->path, strlen(e->path), &key, settings);\n \t\t\tadd_key_to_filter(&key, filter, settings);\n-\t\t\tclear_bloom_key(&key);\n+\t\t\tbloom_key_clear(&key);\n \t\t}\n \n \tcleanup:\ndiff --git a/bloom.h b/bloom.h\nindex a9ded1822f..edf14fef3e 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -78,11 +78,11 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n \n-void fill_bloom_key(const char *data,\n+void bloom_key_fill(const char *data,\n \t\t    size_t len,\n \t\t    struct bloom_key *key,\n \t\t    const struct bloom_filter_settings *settings);\n-void clear_bloom_key(struct bloom_key *key);\n+void bloom_key_clear(struct bloom_key *key);\n \n void add_key_to_filter(const struct bloom_key *key,\n \t\t       struct bloom_filter *filter,\ndiff --git a/line-log.c b/line-log.c\nindex 628e3fe3ae..a2aaf869a3 100644\n--- a/line-log.c\n+++ b/line-log.c\n@@ -1172,12 +1172,12 @@ static int bloom_filter_check(struct rev_info *rev,\n \t\treturn 0;\n \n \twhile (!result && range) {\n-\t\tfill_bloom_key(range->path, strlen(range->path), &key, rev->bloom_filter_settings);\n+\t\tbloom_key_fill(range->path, strlen(range->path), &key, rev->bloom_filter_settings);\n \n \t\tif (bloom_filter_contains(filter, &key, rev->bloom_filter_settings))\n \t\t\tresult = 1;\n \n-\t\tclear_bloom_key(&key);\n+\t\tbloom_key_clear(&key);\n \t\trange = range->next;\n \t}\n \ndiff --git a/revision.c b/revision.c\nindex afee111196..49fc650ac7 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -739,14 +739,14 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \trevs->bloom_keys_nr = path_component_nr;\n \tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n \n-\tfill_bloom_key(path, len, &revs->bloom_keys[0],\n+\tbloom_key_fill(path, len, &revs->bloom_keys[0],\n \t\t       revs->bloom_filter_settings);\n \tpath_component_nr = 1;\n \n \tp = path + len - 1;\n \twhile (p > path) {\n \t\tif (*p == '/')\n-\t\t\tfill_bloom_key(path, p - path,\n+\t\t\tbloom_key_fill(path, p - path,\n \t\t\t\t       &revs->bloom_keys[path_component_nr++],\n \t\t\t\t       revs->bloom_filter_settings);\n \t\tp--;\n@@ -3231,7 +3231,7 @@ void release_revisions(struct rev_info *revs)\n \toidset_clear(&revs->missing_commits);\n \n \tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n-\t\tclear_bloom_key(&revs->bloom_keys[i]);\n+\t\tbloom_key_clear(&revs->bloom_keys[i]);\n \tFREE_AND_NULL(revs->bloom_keys);\n \trevs->bloom_keys_nr = 0;\n }\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 6a24b6e0a6..585a107802 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -12,13 +12,13 @@ static struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n static void add_string_to_filter(const char *data, struct bloom_filter *filter) {\n \t\tstruct bloom_key key;\n \n-\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n+\t\tbloom_key_fill(data, strlen(data), &key, &settings);\n \t\tprintf(\"Hashes:\");\n \t\tfor (size_t i = 0; i < settings.num_hashes; i++)\n \t\t\tprintf(\"0x%08x|\", key.hashes[i]);\n \t\tprintf(\"\\n\");\n \t\tadd_key_to_filter(&key, filter, &settings);\n-\t\tclear_bloom_key(&key);\n+\t\tbloom_key_clear(&key);\n }\n \n static void print_bloom_filter(struct bloom_filter *filter) {\n-- \n2.50.0.107.g33b6ec8c79\n\n"},{"id":"521308","messageId":"20250704111437.2660251-4-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250704111437.2660251-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v4 3/4] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-04T11:14:36Z","receivedAt":"2025-07-04T11:15:05Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"The revision traversal limited by pathspec has optimization when\nthe pathspec has only one element. To support optimization for\nmultiple pathspec items, we need to modify the data structures\nin struct rev_info.\n\nstruct rev_info uses bloom_keys and bloom_nr to store the bloom keys\ncorresponding to a single pathspec item. To allow struct rev_info\nto store bloom keys for multiple pathspec items, a new data structure\n`struct bloom_keyvec` is introduced. Each `struct bloom_keyvec`\ncorresponds to a single pathspec item.\n\nIn `struct rev_info`, replace bloom_keys and bloom_nr with bloom_keyvecs\nand bloom_keyvec_nr. This commit still optimize one pathspec item, thus\nbloom_keyvec_nr can only be 0 or 1.\n\nNew bloom_keyvec_* functions are added to create and destroy a keyvec.\nbloom_filter_contains_vec() is added to check if all key in keyvec is\ncontained in a bloom filter. bloom_keyvec_fill_key() is added to\ninitialize a key in keyvec.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n bloom.c    | 31 +++++++++++++++++++++++++++++++\n bloom.h    | 25 +++++++++++++++++++++++++\n revision.c | 39 +++++++++++++++++++++------------------\n revision.h |  6 +++---\n 4 files changed, 80 insertions(+), 21 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 35ff36c31c..877bda0ef3 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -280,6 +280,25 @@ void deinit_bloom_filters(void)\n \tdeep_clear_bloom_filter_slab(&bloom_filters, free_one_bloom_filter);\n }\n \n+struct bloom_keyvec *bloom_keyvec_new(size_t count)\n+{\n+\tstruct bloom_keyvec *vec;\n+\tsize_t sz = sizeof(struct bloom_keyvec);\n+\tsz += count * sizeof(struct bloom_key);\n+\tvec = (struct bloom_keyvec *)xcalloc(1, sz);\n+\tvec->count = count;\n+\treturn vec;\n+}\n+\n+void bloom_keyvec_free(struct bloom_keyvec *vec)\n+{\n+\tif (!vec)\n+\t\treturn;\n+\tfor (size_t nr = 0; nr < vec->count; nr++)\n+\t\tbloom_key_clear(&vec->key[nr]);\n+\tfree(vec);\n+}\n+\n static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n \t\t       const struct hashmap_entry *eptr,\n \t\t       const struct hashmap_entry *entry_or_key,\n@@ -541,6 +560,18 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \treturn 1;\n }\n \n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *vec,\n+\t\t\t      const struct bloom_filter_settings *settings)\n+{\n+\tint ret = 1;\n+\n+\tfor (size_t nr = 0; ret > 0 && nr < vec->count; nr++)\n+\t\tret = bloom_filter_contains(filter, &vec->key[nr], settings);\n+\n+\treturn ret;\n+}\n+\n uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n \t\t\t\t   int version)\n {\ndiff --git a/bloom.h b/bloom.h\nindex edf14fef3e..3669074f3a 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -74,6 +74,16 @@ struct bloom_key {\n \tuint32_t *hashes;\n };\n \n+/*\n+ * A bloom_keyvec is a vector of bloom_keys, which\n+ * can be used to store multiple keys for a single\n+ * pathspec item.\n+ */\n+struct bloom_keyvec {\n+\tsize_t count;\n+\tstruct bloom_key key[FLEX_ARRAY];\n+};\n+\n int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n@@ -84,6 +94,17 @@ void bloom_key_fill(const char *data,\n \t\t    const struct bloom_filter_settings *settings);\n void bloom_key_clear(struct bloom_key *key);\n \n+struct bloom_keyvec *bloom_keyvec_new(size_t count);\n+void bloom_keyvec_free(struct bloom_keyvec *vec);\n+\n+static inline void bloom_keyvec_fill_key(const char *data, size_t len,\n+\t\t\t\t\t struct bloom_keyvec *vec, size_t nr,\n+\t\t\t\t\t const struct bloom_filter_settings *settings)\n+{\n+\tassert(nr < vec->count);\n+\tbloom_key_fill(data, len, &vec->key[nr], settings);\n+}\n+\n void add_key_to_filter(const struct bloom_key *key,\n \t\t       struct bloom_filter *filter,\n \t\t       const struct bloom_filter_settings *settings);\n@@ -128,6 +149,10 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n \t\t\t  const struct bloom_filter_settings *settings);\n \n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *v,\n+\t\t\t      const struct bloom_filter_settings *settings);\n+\n uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n \t\t\t\t   int version);\n \ndiff --git a/revision.c b/revision.c\nindex 49fc650ac7..7cbb49617d 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -688,6 +688,7 @@ static int forbid_bloom_filters(struct pathspec *spec)\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n+\tstruct bloom_keyvec *bloom_keyvec;\n \tchar *path_alloc = NULL;\n \tconst char *path, *p;\n \tsize_t len;\n@@ -736,19 +737,21 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \t\tp++;\n \t}\n \n-\trevs->bloom_keys_nr = path_component_nr;\n-\tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n+\trevs->bloom_keyvecs_nr = 1;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n+\tbloom_keyvec = bloom_keyvec_new(path_component_nr);\n+\trevs->bloom_keyvecs[0] = bloom_keyvec;\n \n-\tbloom_key_fill(path, len, &revs->bloom_keys[0],\n-\t\t       revs->bloom_filter_settings);\n+\tbloom_keyvec_fill_key(path, len, bloom_keyvec, 0,\n+\t\t\t      revs->bloom_filter_settings);\n \tpath_component_nr = 1;\n \n \tp = path + len - 1;\n \twhile (p > path) {\n \t\tif (*p == '/')\n-\t\t\tbloom_key_fill(path, p - path,\n-\t\t\t\t       &revs->bloom_keys[path_component_nr++],\n-\t\t\t\t       revs->bloom_filter_settings);\n+\t\t\tbloom_keyvec_fill_key(path, p - path, bloom_keyvec,\n+\t\t\t\t\t      path_component_nr++,\n+\t\t\t\t\t      revs->bloom_filter_settings);\n \t\tp--;\n \t}\n \n@@ -764,7 +767,7 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\t\t\t\t\t struct commit *commit)\n {\n \tstruct bloom_filter *filter;\n-\tint result = 1, j;\n+\tint result = 0;\n \n \tif (!revs->repo->objects->commit_graph)\n \t\treturn -1;\n@@ -779,10 +782,10 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\treturn -1;\n \t}\n \n-\tfor (j = 0; result && j < revs->bloom_keys_nr; j++) {\n-\t\tresult = bloom_filter_contains(filter,\n-\t\t\t\t\t       &revs->bloom_keys[j],\n-\t\t\t\t\t       revs->bloom_filter_settings);\n+\tfor (size_t nr = 0; !result && nr < revs->bloom_keyvecs_nr; nr++) {\n+\t\tresult = bloom_filter_contains_vec(filter,\n+\t\t\t\t\t\t   revs->bloom_keyvecs[nr],\n+\t\t\t\t\t\t   revs->bloom_filter_settings);\n \t}\n \n \tif (result)\n@@ -823,7 +826,7 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\treturn REV_TREE_SAME;\n \t}\n \n-\tif (revs->bloom_keys_nr && !nth_parent) {\n+\tif (revs->bloom_keyvecs_nr && !nth_parent) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \n \t\tif (bloom_ret == 0)\n@@ -850,7 +853,7 @@ static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit,\n \tif (!t1)\n \t\treturn 0;\n \n-\tif (!nth_parent && revs->bloom_keys_nr) {\n+\tif (!nth_parent && revs->bloom_keyvecs_nr) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \t\tif (!bloom_ret)\n \t\t\treturn 1;\n@@ -3230,10 +3233,10 @@ void release_revisions(struct rev_info *revs)\n \tline_log_free(revs);\n \toidset_clear(&revs->missing_commits);\n \n-\tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n-\t\tbloom_key_clear(&revs->bloom_keys[i]);\n-\tFREE_AND_NULL(revs->bloom_keys);\n-\trevs->bloom_keys_nr = 0;\n+\tfor (size_t i = 0; i < revs->bloom_keyvecs_nr; i++)\n+\t\tbloom_keyvec_free(revs->bloom_keyvecs[i]);\n+\tFREE_AND_NULL(revs->bloom_keyvecs);\n+\trevs->bloom_keyvecs_nr = 0;\n }\n \n static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\ndiff --git a/revision.h b/revision.h\nindex 6d369cdad6..ac843f58d0 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -62,7 +62,7 @@ struct repository;\n struct rev_info;\n struct string_list;\n struct saved_parents;\n-struct bloom_key;\n+struct bloom_keyvec;\n struct bloom_filter_settings;\n struct option;\n struct parse_opt_ctx_t;\n@@ -360,8 +360,8 @@ struct rev_info {\n \n \t/* Commit graph bloom filter fields */\n \t/* The bloom filter key(s) for the pathspec */\n-\tstruct bloom_key *bloom_keys;\n-\tint bloom_keys_nr;\n+\tstruct bloom_keyvec **bloom_keyvecs;\n+\tint bloom_keyvecs_nr;\n \n \t/*\n \t * The bloom filter settings used to generate the key.\n-- \n2.50.0.107.g33b6ec8c79\n\n"},{"id":"521309","messageId":"20250704111437.2660251-5-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250704111437.2660251-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v4 4/4] bloom: optimize multiple pathspec items in revision traversal","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-04T11:14:37Z","receivedAt":"2025-07-04T11:15:09Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"To enable optimize multiple pathspec items in revision traversal,\nreturn 0 if all pathspec item is literal in forbid_bloom_filters().\nAdd code to initialize and check each pathspec item's bloom_keyvec.\n\nAdd new function release_revisions_bloom_keyvecs() to free all bloom\nkeyvec owned by rev_info.\n\nAdd new test cases in t/t4216-log-bloom.sh to ensure\n  - consistent results between the optimization for multiple pathspec\n    items using bloom filter and the case without bloom filter\n    optimization.\n  - does not use bloom filter if any pathspec item is not literal.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n revision.c           | 118 ++++++++++++++++++++++++-------------------\n t/t4216-log-bloom.sh |  23 +++++----\n 2 files changed, 79 insertions(+), 62 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 7cbb49617d..9a77c0d0bc 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -675,16 +675,17 @@ static int forbid_bloom_filters(struct pathspec *spec)\n {\n \tif (spec->has_wildcard)\n \t\treturn 1;\n-\tif (spec->nr > 1)\n-\t\treturn 1;\n \tif (spec->magic & ~PATHSPEC_LITERAL)\n \t\treturn 1;\n-\tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n-\t\treturn 1;\n+\tfor (size_t nr = 0; nr < spec->nr; nr++)\n+\t\tif (spec->items[nr].magic & ~PATHSPEC_LITERAL)\n+\t\t\treturn 1;\n \n \treturn 0;\n }\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n+\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n@@ -692,7 +693,7 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tchar *path_alloc = NULL;\n \tconst char *path, *p;\n \tsize_t len;\n-\tint path_component_nr = 1;\n+\tint path_component_nr;\n \n \tif (!revs->commits)\n \t\treturn;\n@@ -709,50 +710,52 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n-\tpi = &revs->pruning.pathspec.items[0];\n-\n-\t/* remove single trailing slash from path, if needed */\n-\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n-\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n-\t\tpath = path_alloc;\n-\t} else\n-\t\tpath = pi->match;\n-\n-\tlen = strlen(path);\n-\tif (!len) {\n-\t\trevs->bloom_filter_settings = NULL;\n-\t\tfree(path_alloc);\n-\t\treturn;\n-\t}\n-\n-\tp = path;\n-\twhile (*p) {\n-\t\t/*\n-\t\t * At this point, the path is normalized to use Unix-style\n-\t\t * path separators. This is required due to how the\n-\t\t * changed-path Bloom filters store the paths.\n-\t\t */\n-\t\tif (*p == '/')\n-\t\t\tpath_component_nr++;\n-\t\tp++;\n-\t}\n-\n-\trevs->bloom_keyvecs_nr = 1;\n-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n-\tbloom_keyvec = bloom_keyvec_new(path_component_nr);\n-\trevs->bloom_keyvecs[0] = bloom_keyvec;\n-\n-\tbloom_keyvec_fill_key(path, len, bloom_keyvec, 0,\n-\t\t\t      revs->bloom_filter_settings);\n-\tpath_component_nr = 1;\n+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n+\t\tpi = &revs->pruning.pathspec.items[i];\n+\t\tpath_component_nr = 1;\n+\n+\t\t/* remove single trailing slash from path, if needed */\n+\t\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n+\t\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n+\t\t\tpath = path_alloc;\n+\t\t} else\n+\t\t\tpath = pi->match;\n+\n+\t\tlen = strlen(path);\n+\t\tif (!len)\n+\t\t\tgoto fail;\n+\n+\t\tp = path;\n+\t\twhile (*p) {\n+\t\t\t/*\n+\t\t\t * At this point, the path is normalized to use\n+\t\t\t * Unix-style path separators. This is required due to\n+\t\t\t * how the changed-path Bloom filters store the paths.\n+\t\t\t */\n+\t\t\tif (*p == '/')\n+\t\t\t\tpath_component_nr++;\n+\t\t\tp++;\n+\t\t}\n \n-\tp = path + len - 1;\n-\twhile (p > path) {\n-\t\tif (*p == '/')\n-\t\t\tbloom_keyvec_fill_key(path, p - path, bloom_keyvec,\n-\t\t\t\t\t      path_component_nr++,\n-\t\t\t\t\t      revs->bloom_filter_settings);\n-\t\tp--;\n+\t\tbloom_keyvec = bloom_keyvec_new(path_component_nr);\n+\t\trevs->bloom_keyvecs[i] = bloom_keyvec;\n+\n+\t\tbloom_keyvec_fill_key(path, len, bloom_keyvec, 0,\n+\t\t\t\t      revs->bloom_filter_settings);\n+\t\tpath_component_nr = 1;\n+\n+\t\tp = path + len - 1;\n+\t\twhile (p > path) {\n+\t\t\tif (*p == '/')\n+\t\t\t\tbloom_keyvec_fill_key(path, p - path,\n+\t\t\t\t\t\t      bloom_keyvec,\n+\t\t\t\t\t\t      path_component_nr++,\n+\t\t\t\t\t\t      revs->bloom_filter_settings);\n+\t\t\tp--;\n+\t\t}\n+\t\tFREE_AND_NULL(path_alloc);\n \t}\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n@@ -760,7 +763,12 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \t\tbloom_filter_atexit_registered = 1;\n \t}\n \n+\treturn;\n+\n+fail:\n+\trevs->bloom_filter_settings = NULL;\n \tfree(path_alloc);\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n@@ -3204,6 +3212,14 @@ static void release_revisions_mailmap(struct string_list *mailmap)\n \n static void release_revisions_topo_walk_info(struct topo_walk_info *info);\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs)\n+{\n+\tfor (size_t nr = 0; nr < revs->bloom_keyvecs_nr; nr++)\n+\t\tbloom_keyvec_free(revs->bloom_keyvecs[nr]);\n+\tFREE_AND_NULL(revs->bloom_keyvecs);\n+\trevs->bloom_keyvecs_nr = 0;\n+}\n+\n static void free_void_commit_list(void *list)\n {\n \tfree_commit_list(list);\n@@ -3232,11 +3248,7 @@ void release_revisions(struct rev_info *revs)\n \tclear_decoration(&revs->treesame, free);\n \tline_log_free(revs);\n \toidset_clear(&revs->missing_commits);\n-\n-\tfor (size_t i = 0; i < revs->bloom_keyvecs_nr; i++)\n-\t\tbloom_keyvec_free(revs->bloom_keyvecs[i]);\n-\tFREE_AND_NULL(revs->bloom_keyvecs);\n-\trevs->bloom_keyvecs_nr = 0;\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 8910d53cac..639868ac56 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -66,8 +66,9 @@ sane_unset GIT_TRACE2_CONFIG_PARAMS\n \n setup () {\n \trm -f \"$TRASH_DIRECTORY/trace.perf\" &&\n-\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n-\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n+\teval git -c core.commitGraph=false log --pretty=\"format:%s\" \"$1\" >log_wo_bloom &&\n+\teval \"GIT_TRACE2_PERF=\\\"$TRASH_DIRECTORY/trace.perf\\\"\" \\\n+\t\tgit -c core.commitGraph=true log --pretty=\"format:%s\" \"$1\" >log_w_bloom\n }\n \n test_bloom_filters_used () {\n@@ -138,10 +139,6 @@ test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n '\n \n-test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n-'\n-\n test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"-- .\"\n '\n@@ -151,9 +148,17 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n \ttest_bloom_filters_used \"-- *renamed\"\n '\n \n-test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- *\" &&\n-\ttest_bloom_filters_not_used \"-- file*\"\n+test_expect_success 'git log with multiple literal paths uses Bloom filter' '\n+\ttest_bloom_filters_used \"-- file4 A/file1\" &&\n+\ttest_bloom_filters_used \"-- *\" &&\n+\ttest_bloom_filters_used \"-- file*\"\n+'\n+\n+test_expect_success 'git log with path contains a wildcard does not use Bloom filter' '\n+\ttest_bloom_filters_not_used \"-- file\\*\" &&\n+\ttest_bloom_filters_not_used \"-- A/\\* file4\" &&\n+\ttest_bloom_filters_not_used \"-- file4 A/\\*\" &&\n+\ttest_bloom_filters_not_used \"-- * A/\\*\"\n '\n \n test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n-- \n2.50.0.107.g33b6ec8c79\n\n"},{"id":"521311","messageId":"1BD174FD-887A-4002-955C-A67E7DBFEFCA@gmail.com","threadId":"63691","inReplyTo":"xmqq1pqyfvb2.fsf@gitster.g","subject":"Re: [PATCH 2/2] bloom: enable multiple pathspec bloom keys","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-04T12:09:23Z","receivedAt":"2025-07-04T12:09:37Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n> \n> Before concluding so, we may want to double check how Bloom filters\n> are built on case insensitive systems, though.  If we normalize the\n> string by downcasing before murmuring the string, the resulting\n> Bloom filter may have more false positives for those who want to\n> (ab)use it to optimize case sensitive queries (without affecting\n> correctness), but case insensitive queries would be helped.  I do\n> not think we support (or want to support) a repository that spans\n> across two filesystems with different case sensitivity, so those who\n> worked on our changed-path Bloom filter subsystem may have already\n> placed such an optimization, based on the case sensitivity recorded\n> in the repository (core.ignorecase).\n\nIn bloom.c:get_or_compute_bloom_filter(), the computation of a bloom filter\nlooks like:\n    diff_tree_oid(c’s parent or NULL, &c->object.oid, \"\", &diffopt);\n    diffcore_std(&diffopt);\n    struct hashmap path_hashmap;\n\n    for (path : diff_queue_diff) {\n        Add all parts of path to path_hashmap;\n    }\n\n    for_each(path_hashmap) {\n        Add path to filter\n    }\n\nAll these steps do not check config.ignoreCase, so I believe the Bloom filter we\nbuild in the commit graph is case-sensitive.\n\nTo demonstrate this assumption—and since I happen to be a Mac user (where\nconfig.ignoreCase is true by default)—I ran the following commands under the\nllvm-project repository:\n\n$ git commit-graph write --split --reachable --changed-paths\n$ time git log -5 -t -- README.md > /dev/null\nreal\t0m0.089s\nuser\t0m0.067s\nsys\t0m0.021s\n$ time git log -5 -t -- ':(icase)README.md' > /dev/null\nreal\t0m0.281s\nuser\t0m0.239s\nsys\t0m0.041s\n$ time git log -5 -t -- ‘rEADME.md’ > /dev/null\nreal\t0m0.458s\nuser\t0m0.394s\nsys\t0m0.061s\n\nAnd I think it proves that changed-path Bloom filter doesn’t optimize icase\npathspec item in case insensitive file system."},{"id":"521428","messageId":"65dc80f9-a91c-463b-9c6b-cb20d293432b@gmail.com","threadId":"63691","inReplyTo":"20250704111437.2660251-4-502024330056@smail.nju.edu.cn","subject":"Re: [PATCH v4 3/4] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2025-07-07T11:35:42Z","receivedAt":"2025-07-07T11:35:45Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/4/2025 7:14 AM, Lidong Yan wrote:\n> The revision traversal limited by pathspec has optimization when\n> the pathspec has only one element. To support optimization for\n> multiple pathspec items, we need to modify the data structures\n> in struct rev_info.\n\nYou are correct that the revision-walking abandons bloom filters\nwhen there are multiple pathspecs, and fixing this is a valuable\neffort.\n\nThe need for this change is subtle and could use some extra context\nto be sure reviewers understand:\n\nThis change is writing over some code that was created in\nc525ce95b46 (commit-graph: check all leading directories in changed\npath Bloom filters, 2020-07-01) to allow storing multiple bloom\nkeys during the revision walk. The multiple keys are focusing on\nmultiple path components of the literal pathspec. \nThe point is that after the initialization, the bloom key array is\nused directly as a filter for reporting TREESAME commits: a commit\nis automatically reported as TREESAME to its first parent if any\nbloom key results in a \"No, definitely not changed\" result with\nthat commit's bloom filter.\n\nThe reason we need a new data structure is that we need to adjust\nthe conditionals.\n\nBEFORE: \"NOT TREESAME if there EXISTS a bloom key that reports NO\"\n\nAFTER: \"NOT TREESAME if FOR EVERY pathspec there EXISTS a bloom key\n        that reports NO.\"\n\nThis \"FOR EVERY\" condition makes it impossible to use a flat array\nof bloom keys for multiple pathspecs, justifying this change.\n\nWhat is further confusing here is that we already have logic that\ndeals with arrays of bloom keys, so I expected that the vector was\nthe single structure storing a list of those arrays. Instead, the\nvector is replacing the array itself. This is made clear by using\nthe vector immediately in the existing implementation.\n> +struct bloom_keyvec *bloom_keyvec_new(size_t count)\n> +{\n> +\tstruct bloom_keyvec *vec;\n> +\tsize_t sz = sizeof(struct bloom_keyvec);\n> +\tsz += count * sizeof(struct bloom_key);\n> +\tvec = (struct bloom_keyvec *)xcalloc(1, sz);\nYou could use CALLOC_ARRAY() to simplify this and drop\nthe 'sz' variable.\n\n> +\tvec->count = count;\n> +\treturn vec;\n> +}\n> +\n> +void bloom_keyvec_free(struct bloom_keyvec *vec)\n> +{\n> +\tif (!vec)\n> +\t\treturn;\n> +\tfor (size_t nr = 0; nr < vec->count; nr++)\n> +\t\tbloom_key_clear(&vec->key[nr]);\n> +\tfree(vec);\n> +}\n> +\n>  static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n>  \t\t       const struct hashmap_entry *eptr,\n>  \t\t       const struct hashmap_entry *entry_or_key,\n> @@ -541,6 +560,18 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n>  \treturn 1;\n>  }\n>  \n> +int bloom_filter_contains_vec(const struct bloom_filter *filter,\n> +\t\t\t      const struct bloom_keyvec *vec,\n> +\t\t\t      const struct bloom_filter_settings *settings)\n> +{\n> +\tint ret = 1;\n> +\n> +\tfor (size_t nr = 0; ret > 0 && nr < vec->count; nr++)\n> +\t\tret = bloom_filter_contains(filter, &vec->key[nr], settings);\n> +\n> +\treturn ret;\n> +}\n\nThis implementation is where the subtle detail comes in. Might be worth\na comment to say \"if any key in this list is not contained in the filter,\nthen the filter doesn't match this vector.\"\n\n"},{"id":"521429","messageId":"ea144a72-0975-4ac9-b2e4-ae0f7fcb6837@gmail.com","threadId":"63691","inReplyTo":"20250704111437.2660251-5-502024330056@smail.nju.edu.cn","subject":"Re: [PATCH v4 4/4] bloom: optimize multiple pathspec items in revision traversal","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2025-07-07T11:43:19Z","receivedAt":"2025-07-07T11:43:21Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/4/2025 7:14 AM, Lidong Yan wrote:\n> To enable optimize multiple pathspec items in revision traversal,\n> return 0 if all pathspec item is literal in forbid_bloom_filters().\n> Add code to initialize and check each pathspec item's bloom_keyvec.\n> \n> Add new function release_revisions_bloom_keyvecs() to free all bloom\n> keyvec owned by rev_info.\n> \n> Add new test cases in t/t4216-log-bloom.sh to ensure\n>   - consistent results between the optimization for multiple pathspec\n>     items using bloom filter and the case without bloom filter\n>     optimization.\n>   - does not use bloom filter if any pathspec item is not literal.\n\nThis would be a great time to add some performance statistics when\nusing this feature with multiple pathspecs on some standard repos (git\nand the Linux kernel repo are two good examples).\n\nWe don't have a great performance script for this, since each test\nrepo will have different paths to use for comparisons, but you can\nuse 'hyperfine' to assemble your own comparisons before and after\nthis change and report them here (and in your cover letter). \n> Signed-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n> ---\n>  revision.c           | 118 ++++++++++++++++++++++++-------------------\n>  t/t4216-log-bloom.sh |  23 +++++----\n>  2 files changed, 79 insertions(+), 62 deletions(-)\n> \n> diff --git a/revision.c b/revision.c\n> index 7cbb49617d..9a77c0d0bc 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -675,16 +675,17 @@ static int forbid_bloom_filters(struct pathspec *spec)\n>  {\n>  \tif (spec->has_wildcard)\n>  \t\treturn 1;\n> -\tif (spec->nr > 1)\n> -\t\treturn 1;\n>  \tif (spec->magic & ~PATHSPEC_LITERAL)\n>  \t\treturn 1;\n> -\tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n> -\t\treturn 1;\n> +\tfor (size_t nr = 0; nr < spec->nr; nr++)\n> +\t\tif (spec->items[nr].magic & ~PATHSPEC_LITERAL)\n> +\t\t\treturn 1;\n\nThis is a good check: if any is non-literal, then we can't use bloom\nfilters.\n\n> +static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n> +\n>  static void prepare_to_use_bloom_filter(struct rev_info *revs)\n>  {\n>  \tstruct pathspec_item *pi;\n> @@ -692,7 +693,7 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n>  \tchar *path_alloc = NULL;\n>  \tconst char *path, *p;\n>  \tsize_t len;\n> -\tint path_component_nr = 1;\n> +\tint path_component_nr;\n\nWe can move this into the interior of the loop, right?\n\n>  \tif (!revs->commits)\n>  \t\treturn;\n> @@ -709,50 +710,52 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n>  \tif (!revs->pruning.pathspec.nr)\n>  \t\treturn;\n>  \n> -\tpi = &revs->pruning.pathspec.items[0];\n> -\n> -\t/* remove single trailing slash from path, if needed */\n> -\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n> -\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n> -\t\tpath = path_alloc;\n> -\t} else\n> -\t\tpath = pi->match;\n> -\n> -\tlen = strlen(path);\n> -\tif (!len) {\n> -\t\trevs->bloom_filter_settings = NULL;\n> -\t\tfree(path_alloc);\n> -\t\treturn;\n> -\t}\n> -\n> -\tp = path;\n> -\twhile (*p) {\n> -\t\t/*\n> -\t\t * At this point, the path is normalized to use Unix-style\n> -\t\t * path separators. This is required due to how the\n> -\t\t * changed-path Bloom filters store the paths.\n> -\t\t */\n> -\t\tif (*p == '/')\n> -\t\t\tpath_component_nr++;\n> -\t\tp++;\n> -\t}\n> -\n> -\trevs->bloom_keyvecs_nr = 1;\n> -\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n> -\tbloom_keyvec = bloom_keyvec_new(path_component_nr);\n> -\trevs->bloom_keyvecs[0] = bloom_keyvec;\n> -\n> -\tbloom_keyvec_fill_key(path, len, bloom_keyvec, 0,\n> -\t\t\t      revs->bloom_filter_settings);\n> -\tpath_component_nr = 1;\n\nThe size of this diff is unfortunate. I wonder if there could first\nbe an extraction of this logic to operate on a single pathspec and\nbloom_keyvec in a way that would be an obvious code move, then this\npatch could call that method in a loop now that we have an array of\nbloom_keyvecs.\n\nI think it would make a cleaner patch and a cleaner final result.\n\n> +\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n> +\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n> +\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n> +\t\tpi = &revs->pruning.pathspec.items[i];\n> +\t\tpath_component_nr = 1;\n> +\n> +\t\t/* remove single trailing slash from path, if needed */\n> +\t\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n> +\t\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n> +\t\t\tpath = path_alloc;\n> +\t\t} else\n> +\t\t\tpath = pi->match;\n> +\n> +\t\tlen = strlen(path);\n> +\t\tif (!len)\n> +\t\t\tgoto fail;\n> +\n> +\t\tp = path;\n> +\t\twhile (*p) {\n> +\t\t\t/*\n> +\t\t\t * At this point, the path is normalized to use\n> +\t\t\t * Unix-style path separators. This is required due to\n> +\t\t\t * how the changed-path Bloom filters store the paths.\n> +\t\t\t */\n> +\t\t\tif (*p == '/')\n> +\t\t\t\tpath_component_nr++;\n> +\t\t\tp++;\n> +\t\t}\n>  \n> -\tp = path + len - 1;\n> -\twhile (p > path) {\n> -\t\tif (*p == '/')\n> -\t\t\tbloom_keyvec_fill_key(path, p - path, bloom_keyvec,\n> -\t\t\t\t\t      path_component_nr++,\n> -\t\t\t\t\t      revs->bloom_filter_settings);\n> -\t\tp--;\n> +\t\tbloom_keyvec = bloom_keyvec_new(path_component_nr);\n> +\t\trevs->bloom_keyvecs[i] = bloom_keyvec;\n> +\n> +\t\tbloom_keyvec_fill_key(path, len, bloom_keyvec, 0,\n> +\t\t\t\t      revs->bloom_filter_settings);\n> +\t\tpath_component_nr = 1;\n> +\n> +\t\tp = path + len - 1;\n> +\t\twhile (p > path) {\n> +\t\t\tif (*p == '/')\n> +\t\t\t\tbloom_keyvec_fill_key(path, p - path,\n> +\t\t\t\t\t\t      bloom_keyvec,\n> +\t\t\t\t\t\t      path_component_nr++,\n> +\t\t\t\t\t\t      revs->bloom_filter_settings);\n> +\t\t\tp--;\n> +\t\t}\n> +\t\tFREE_AND_NULL(path_alloc);\n\n...\n\n> -test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n> -\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n> -'\n> -\n\nLove to see this.\n\n>  test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n>  \ttest_bloom_filters_not_used \"-- .\"\n>  '\n> @@ -151,9 +148,17 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n>  \ttest_bloom_filters_used \"-- *renamed\"\n>  '\n>  \n> -test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n> -\ttest_bloom_filters_not_used \"-- *\" &&\n> -\ttest_bloom_filters_not_used \"-- file*\"\n> +test_expect_success 'git log with multiple literal paths uses Bloom filter' '\n> +\ttest_bloom_filters_used \"-- file4 A/file1\" &&\n> +\ttest_bloom_filters_used \"-- *\" &&\n> +\ttest_bloom_filters_used \"-- file*\"\n> +'\n> +\n> +test_expect_success 'git log with path contains a wildcard does not use Bloom filter' '\n> +\ttest_bloom_filters_not_used \"-- file\\*\" &&\n> +\ttest_bloom_filters_not_used \"-- A/\\* file4\" &&\n> +\ttest_bloom_filters_not_used \"-- file4 A/\\*\" &&\n> +\ttest_bloom_filters_not_used \"-- * A/\\*\"\n>  '\n\nAnd these new test cases are great.\n\nThanks for this work. I'm happy to see the feature be added and my suggestions\nare purely cosmetic as the proof is in your tests.\n\nThanks,\n-Stolee \n"},{"id":"521439","messageId":"5DB7714D-4009-47C4-A8F7-1C375C6D29AF@smail.nju.edu.cn","threadId":"63691","inReplyTo":"65dc80f9-a91c-463b-9c6b-cb20d293432b@gmail.com","subject":"Re: [PATCH v4 3/4] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-07-07T14:14:22Z","receivedAt":"2025-07-07T14:15:03Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Derrick Stolee <stolee@gmail.com> wrote:\n> \n> BEFORE: \"NOT TREESAME if there EXISTS a bloom key that reports NO\"\n> \n> AFTER: \"NOT TREESAME if FOR EVERY pathspec there EXISTS a bloom key\n>        that reports NO.\"\n> \n> This \"FOR EVERY\" condition makes it impossible to use a flat array\n> of bloom keys for multiple pathspecs, justifying this change.\n> \n> What is further confusing here is that we already have logic that\n> deals with arrays of bloom keys, so I expected that the vector was\n> the single structure storing a list of those arrays. Instead, the\n> vector is replacing the array itself. This is made clear by using\n> the vector immediately in the existing implementation.\n\nI think the problem here is that it clearly enough in my comments.\nWhen writing the code, I thought about converting the original\none-dimensional array struct bloom_key *keys into a two-dimensional\narray struct bloom_key **. Then I came up with the idea that this\ntwo-dimensional array could be designed as struct bloom_keyvec *keyvecs,\nwhich might be clearer. Each struct bloom_keyvec would represent all the\nbloom_key elements for a single pathspec item.\n\n>> +struct bloom_keyvec *bloom_keyvec_new(size_t count)\n>> +{\n>> + struct bloom_keyvec *vec;\n>> + size_t sz = sizeof(struct bloom_keyvec);\n>> + sz += count * sizeof(struct bloom_key);\n>> + vec = (struct bloom_keyvec *)xcalloc(1, sz);\n> You could use CALLOC_ARRAY() to simplify this and drop\n> the 'sz' variable.\n\nYou are right. I think you suggest to write struct bloom_keyvec like:\n\n———————-\n|        count        |\n———————\n|         *keys       |  ——>     key0 | key1 | key2 | … |\n———————\n\nAnd I am doing here makes struct bloom_keyvec looks like\n\n———————\n|       count        |\n———————\n|        key[0]      |\n———————\n|        key[1]      |\n———————\n|        …            |\n\nAlthough bloom_keyvec_new() appears more complex, the advantage\nis that bloom_keyvec_destroy() no longer needs to free keys manually.\nAnd if I understand correctly, junio had suggested to use the second way\n[here](https://lore.kernel.org/git/xmqqtt43u36t.fsf@gitster.g/).\n\n> \n>> + vec->count = count;\n>> + return vec;\n>> +}\n>> +\n>> +void bloom_keyvec_free(struct bloom_keyvec *vec)\n>> +{\n>> + if (!vec)\n>> + return;\n>> + for (size_t nr = 0; nr < vec->count; nr++)\n>> + bloom_key_clear(&vec->key[nr]);\n>> + free(vec);\n>> +}\n>> +\n>> static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n>>       const struct hashmap_entry *eptr,\n>>       const struct hashmap_entry *entry_or_key,\n>> @@ -541,6 +560,18 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n>> return 1;\n>> }\n>> \n>> +int bloom_filter_contains_vec(const struct bloom_filter *filter,\n>> +      const struct bloom_keyvec *vec,\n>> +      const struct bloom_filter_settings *settings)\n>> +{\n>> + int ret = 1;\n>> +\n>> + for (size_t nr = 0; ret > 0 && nr < vec->count; nr++)\n>> + ret = bloom_filter_contains(filter, &vec->key[nr], settings);\n>> +\n>> + return ret;\n>> +}\n> \n> This implementation is where the subtle detail comes in. Might be worth\n> a comment to say \"if any key in this list is not contained in the filter,\n> then the filter doesn't match this vector.\"\n\nI will add comment in the next version\n\nThank you for your review,\nLidong\n\n"},{"id":"521440","messageId":"23232D29-CEFA-409D-90E3-F894EC37C1BE@smail.nju.edu.cn","threadId":"63691","inReplyTo":"ea144a72-0975-4ac9-b2e4-ae0f7fcb6837@gmail.com","subject":"Re: [PATCH v4 4/4] bloom: optimize multiple pathspec items in revision traversal","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-07-07T14:18:45Z","receivedAt":"2025-07-07T14:19:14Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Derrick Stolee <stolee@gmail.com> wrote:\n> \n> On 7/4/2025 7:14 AM, Lidong Yan wrote:\n>> To enable optimize multiple pathspec items in revision traversal,\n>> return 0 if all pathspec item is literal in forbid_bloom_filters().\n>> Add code to initialize and check each pathspec item's bloom_keyvec.\n>> \n>> Add new function release_revisions_bloom_keyvecs() to free all bloom\n>> keyvec owned by rev_info.\n>> \n>> Add new test cases in t/t4216-log-bloom.sh to ensure\n>>  - consistent results between the optimization for multiple pathspec\n>>    items using bloom filter and the case without bloom filter\n>>    optimization.\n>>  - does not use bloom filter if any pathspec item is not literal.\n> \n> This would be a great time to add some performance statistics when\n> using this feature with multiple pathspecs on some standard repos (git\n> and the Linux kernel repo are two good examples).\n> \n> We don't have a great performance script for this, since each test\n> repo will have different paths to use for comparisons, but you can\n> use 'hyperfine' to assemble your own comparisons before and after\n> this change and report them here (and in your cover letter).\n\nGot it, I will add statistics in the next version.\n\n>> + int path_component_nr;\n> \n> We can move this into the interior of the loop, right?\n\nYes, move this into loop looks better.\n\n> \n> The size of this diff is unfortunate. I wonder if there could first\n> be an extraction of this logic to operate on a single pathspec and\n> bloom_keyvec in a way that would be an obvious code move, then this\n> patch could call that method in a loop now that we have an array of\n> bloom_keyvecs.\n\nI will try to figure out a way to make diff cleaner.\n\nThanks,\nLidong\n\n"},{"id":"521443","messageId":"xmqqikk458y7.fsf@gitster.g","threadId":"63691","inReplyTo":"ea144a72-0975-4ac9-b2e4-ae0f7fcb6837@gmail.com","subject":"Re: [PATCH v4 4/4] bloom: optimize multiple pathspec items in revision traversal","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-07T15:14:56Z","receivedAt":"2025-07-07T15:14:57Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> On 7/4/2025 7:14 AM, Lidong Yan wrote:\n>> To enable optimize multiple pathspec items in revision traversal,\n>> return 0 if all pathspec item is literal in forbid_bloom_filters().\n>> Add code to initialize and check each pathspec item's bloom_keyvec.\n>> \n>> Add new function release_revisions_bloom_keyvecs() to free all bloom\n>> keyvec owned by rev_info.\n>> \n>> Add new test cases in t/t4216-log-bloom.sh to ensure\n>>   - consistent results between the optimization for multiple pathspec\n>>     items using bloom filter and the case without bloom filter\n>>     optimization.\n>>   - does not use bloom filter if any pathspec item is not literal.\n>\n> This would be a great time to add some performance statistics when\n> using this feature with multiple pathspecs on some standard repos (git\n> and the Linux kernel repo are two good examples).\n>\n> We don't have a great performance script for this, since each test\n> repo will have different paths to use for comparisons, but you can\n> use 'hyperfine' to assemble your own comparisons before and after\n> this change and report them here (and in your cover letter). \n\nExcellent suggestion.  Very much appreciated.\n\n> The size of this diff is unfortunate. I wonder if there could first\n> be an extraction of this logic to operate on a single pathspec and\n> bloom_keyvec in a way that would be an obvious code move, then this\n> patch could call that method in a loop now that we have an array of\n> bloom_keyvecs.\n>\n> I think it would make a cleaner patch and a cleaner final result.\n\nHmph.  Didn't think of that extra \"preliminary clean-up\" step\nmyself, but I think you have a point.  That would likely make\nthe resulting series easier to follow.\n\nThanks.\n"},{"id":"521740","messageId":"20250710084829.2171855-1-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250704111437.2660251-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v5 0/4] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-10T08:48:25Z","receivedAt":"2025-07-10T08:48:43Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"The revision traversal limited by pathspec has optimization when\nthe pathspec has only one element, it does not use any pathspec\nmagic (other than literal), and there is no wildcard. The absence\nof optimization for multiple pathspec elements in revision traversal\ncause an issue raised by Kai Koponen at\n  https://lore.kernel.org/git/CADYQcGqaMC=4jgbmnF9Q11oC11jfrqyvH8EuiRRHytpMXd4wYA@mail.gmail.com/\n\nWhile it is much harder to lift the latter two limitations,\nsupporting a pathspec with multiple elements is relatively easy.\nJust make sure we hash each of them separately and ask the bloom\nfilter about them, and if we see none of them can possibly be\naffected by the commit, we can skip without tree comparison.\n\nThe difference from v4 is:\n  - for the bloom_key_* functions, we now pass struct bloom_key *\n    as the first parameter.\n  - bloom_keyvec_fill_key() and bloom_keyvec_new() have been merged\n    into a single function, so that each key is filled during the\n    creation of the bloom_keyvec.\n\nBelow is a comparison of the time taken to run git log on Git and\nLLVM repositories before and after applying this patch.\n\nSetup commit-graph:\n  $ cd ~/src/git && git commit-graph write --split --reachable --changed-paths\n  $ cd ~/src/llvm && git commit-graph write --split --reachable --changed-paths\n\nRun git log on Git repository\n  $ cd ~/src/git\n  $ hash -p /usr/bin/git git # my system git binary is 2.43.0\n  $ time git log -100 -- commit.c commit-graph.c >/dev/null\n  real\t0m0.055s\n  user\t0m0.040s\n  sys\t  0m0.015s\n  $ hash -p ~/bin/git/bin/git git\n  $ time git log -100 -- commit.c commit-graph.c >/dev/null\n  real\t0m0.039s\n  user\t0m0.020s\n  sys\t  0m0.020s\n\nRun git log in LLVM repository\n  $ cd ~/src/llvm\n  $ hash -p /usr/bin/git git # my system git binary is 2.43.0\n  $ time git log -100 -- llvm/lib/Support/CommandLine.cpp llvm/lib/Support/CommandLine.h >/dev/null\n  real\t0m2.365s\n  user\t0m2.252s\n  sys\t  0m0.113s\n  $ hash -p ~/bin/git/bin/git git\n  $ time git log -100 -- llvm/lib/Support/CommandLine.cpp llvm/lib/Support/CommandLine.h >/dev/null\n  real\t0m0.240s\n  user\t0m0.158s\n  sys\t  0m0.064s\n\nLidong Yan (4):\n  bloom: add test helper to return murmur3 hash\n  bloom: rename function operates on bloom_key\n  bloom: replace struct bloom_key * with struct bloom_keyvec\n  bloom: optimize multiple pathspec items in revision traversal\n\n blame.c               |  2 +-\n bloom.c               | 84 +++++++++++++++++++++++++++++++++---\n bloom.h               | 54 +++++++++++++++++------\n line-log.c            |  5 ++-\n revision.c            | 99 +++++++++++++++++++------------------------\n revision.h            |  6 +--\n t/helper/test-bloom.c |  8 ++--\n t/t4216-log-bloom.sh  | 23 ++++++----\n 8 files changed, 187 insertions(+), 94 deletions(-)\n\nRange-diff against v4:\n1:  d6883e9d6c = 1:  31c048dcdb bloom: add test helper to return murmur3 hash\n2:  f114556c0f ! 2:  d2603c1752 bloom: rename function operates on bloom_key\n    @@ blame.c: static void add_bloom_key(struct blame_bloom_data *bd,\n      \n      \tbd->keys[bd->nr] = xmalloc(sizeof(struct bloom_key));\n     -\tfill_bloom_key(path, strlen(path), bd->keys[bd->nr], bd->settings);\n    -+\tbloom_key_fill(path, strlen(path), bd->keys[bd->nr], bd->settings);\n    ++\tbloom_key_fill(bd->keys[bd->nr], path, strlen(path), bd->settings);\n      \tbd->nr++;\n      }\n      \n    @@ bloom.c: static uint32_t murmur3_seeded_v1(uint32_t seed, const char *data, size\n      }\n      \n     -void fill_bloom_key(const char *data,\n    -+void bloom_key_fill(const char *data,\n    - \t\t    size_t len,\n    - \t\t    struct bloom_key *key,\n    +-\t\t    size_t len,\n    +-\t\t    struct bloom_key *key,\n    ++void bloom_key_fill(struct bloom_key *key, const char *data, size_t len,\n      \t\t    const struct bloom_filter_settings *settings)\n    + {\n    + \tint i;\n     @@ bloom.c: void fill_bloom_key(const char *data,\n      \t\tkey->hashes[i] = hash0 + i * hash1;\n      }\n    @@ bloom.c: struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n      \t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n      \t\t\tstruct bloom_key key;\n     -\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, settings);\n    -+\t\t\tbloom_key_fill(e->path, strlen(e->path), &key, settings);\n    ++\t\t\tbloom_key_fill(&key, e->path, strlen(e->path), settings);\n      \t\t\tadd_key_to_filter(&key, filter, settings);\n     -\t\t\tclear_bloom_key(&key);\n     +\t\t\tbloom_key_clear(&key);\n    @@ bloom.h: int load_bloom_filter_from_graph(struct commit_graph *g,\n      \t\t\t\t uint32_t graph_pos);\n      \n     -void fill_bloom_key(const char *data,\n    -+void bloom_key_fill(const char *data,\n    - \t\t    size_t len,\n    - \t\t    struct bloom_key *key,\n    +-\t\t    size_t len,\n    +-\t\t    struct bloom_key *key,\n    ++void bloom_key_fill(struct bloom_key *key, const char *data, size_t len,\n      \t\t    const struct bloom_filter_settings *settings);\n     -void clear_bloom_key(struct bloom_key *key);\n     +void bloom_key_clear(struct bloom_key *key);\n    @@ line-log.c: static int bloom_filter_check(struct rev_info *rev,\n      \n      \twhile (!result && range) {\n     -\t\tfill_bloom_key(range->path, strlen(range->path), &key, rev->bloom_filter_settings);\n    -+\t\tbloom_key_fill(range->path, strlen(range->path), &key, rev->bloom_filter_settings);\n    ++\t\tbloom_key_fill(&key, range->path, strlen(range->path),\n    ++\t\t\t       rev->bloom_filter_settings);\n      \n      \t\tif (bloom_filter_contains(filter, &key, rev->bloom_filter_settings))\n      \t\t\tresult = 1;\n    @@ revision.c: static void prepare_to_use_bloom_filter(struct rev_info *revs)\n      \tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n      \n     -\tfill_bloom_key(path, len, &revs->bloom_keys[0],\n    -+\tbloom_key_fill(path, len, &revs->bloom_keys[0],\n    ++\tbloom_key_fill(&revs->bloom_keys[0], path, len,\n      \t\t       revs->bloom_filter_settings);\n      \tpath_component_nr = 1;\n      \n    @@ revision.c: static void prepare_to_use_bloom_filter(struct rev_info *revs)\n      \twhile (p > path) {\n      \t\tif (*p == '/')\n     -\t\t\tfill_bloom_key(path, p - path,\n    -+\t\t\tbloom_key_fill(path, p - path,\n    - \t\t\t\t       &revs->bloom_keys[path_component_nr++],\n    +-\t\t\t\t       &revs->bloom_keys[path_component_nr++],\n    ++\t\t\tbloom_key_fill(&revs->bloom_keys[path_component_nr++],\n    ++\t\t\t\t       path, p - path,\n      \t\t\t\t       revs->bloom_filter_settings);\n      \t\tp--;\n    + \t}\n     @@ revision.c: void release_revisions(struct rev_info *revs)\n      \toidset_clear(&revs->missing_commits);\n      \n    @@ t/helper/test-bloom.c: static struct bloom_filter_settings settings = DEFAULT_BL\n      \t\tstruct bloom_key key;\n      \n     -\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n    -+\t\tbloom_key_fill(data, strlen(data), &key, &settings);\n    ++\t\tbloom_key_fill(&key, data, strlen(data), &settings);\n      \t\tprintf(\"Hashes:\");\n      \t\tfor (size_t i = 0; i < settings.num_hashes; i++)\n      \t\t\tprintf(\"0x%08x|\", key.hashes[i]);\n3:  c042907b92 ! 3:  60a3b16bbb bloom: replace struct bloom_key * with struct bloom_keyvec\n    @@ Metadata\n      ## Commit message ##\n         bloom: replace struct bloom_key * with struct bloom_keyvec\n     \n    -    The revision traversal limited by pathspec has optimization when\n    -    the pathspec has only one element. To support optimization for\n    -    multiple pathspec items, we need to modify the data structures\n    -    in struct rev_info.\n    +    Previously, we stored bloom keys in a flat array and marked a commit\n    +    as NOT TREESAME if any key reported \"definitely not changed\".\n     \n    -    struct rev_info uses bloom_keys and bloom_nr to store the bloom keys\n    -    corresponding to a single pathspec item. To allow struct rev_info\n    -    to store bloom keys for multiple pathspec items, a new data structure\n    -    `struct bloom_keyvec` is introduced. Each `struct bloom_keyvec`\n    -    corresponds to a single pathspec item.\n    +    To support multiple pathspec items, we now require that for each\n    +    pathspec item, there exists a bloom key reporting \"definitely not\n    +    changed\".\n     \n    -    In `struct rev_info`, replace bloom_keys and bloom_nr with bloom_keyvecs\n    -    and bloom_keyvec_nr. This commit still optimize one pathspec item, thus\n    -    bloom_keyvec_nr can only be 0 or 1.\n    +    This \"for every\" condition makes a flat array insufficient, so we\n    +    introduce a new structure to group keys by a single pathspec item.\n    +    `struct bloom_keyvec` is introduced to replace `struct bloom_key *`\n    +    and `bloom_key_nr`. And because we want to support multiple pathspec\n    +    items, we added a bloom_keyvec * and a bloom_keyvec_nr field to\n    +    `struct rev_info` to represent an array of bloom_keyvecs. This commit\n    +    still optimize only one pathspec item, thus bloom_keyvec_nr can only\n    +    be 0 or 1.\n     \n         New bloom_keyvec_* functions are added to create and destroy a keyvec.\n         bloom_filter_contains_vec() is added to check if all key in keyvec is\n    -    contained in a bloom filter. bloom_keyvec_fill_key() is added to\n    -    initialize a key in keyvec.\n    +    contained in a bloom filter.\n     \n         Signed-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n     \n    @@ bloom.c: void deinit_bloom_filters(void)\n      \tdeep_clear_bloom_filter_slab(&bloom_filters, free_one_bloom_filter);\n      }\n      \n    -+struct bloom_keyvec *bloom_keyvec_new(size_t count)\n    ++struct bloom_keyvec *bloom_keyvec_new(const char *path, size_t len,\n    ++\t\t\t\t      const struct bloom_filter_settings *settings)\n     +{\n     +\tstruct bloom_keyvec *vec;\n    -+\tsize_t sz = sizeof(struct bloom_keyvec);\n    -+\tsz += count * sizeof(struct bloom_key);\n    ++\tconst char *p;\n    ++\tsize_t sz;\n    ++\tsize_t nr = 1;\n    ++\n    ++\tp = path;\n    ++\twhile (*p) {\n    ++\t\t/*\n    ++\t\t * At this point, the path is normalized to use Unix-style\n    ++\t\t * path separators. This is required due to how the\n    ++\t\t * changed-path Bloom filters store the paths.\n    ++\t\t */\n    ++\t\tif (*p == '/')\n    ++\t\t\tnr++;\n    ++\t\tp++;\n    ++\t}\n    ++\n    ++\tsz = sizeof(struct bloom_keyvec);\n    ++\tsz += nr * sizeof(struct bloom_key);\n     +\tvec = (struct bloom_keyvec *)xcalloc(1, sz);\n    -+\tvec->count = count;\n    ++\tif (!vec)\n    ++\t\treturn NULL;\n    ++\tvec->count = nr;\n    ++\n    ++\tbloom_key_fill(&vec->key[0], path, len, settings);\n    ++\tnr = 1;\n    ++\tp = path + len - 1;\n    ++\twhile (p > path) {\n    ++\t\tif (*p == '/') {\n    ++\t\t\tbloom_key_fill(&vec->key[nr++], path, p - path, settings);\n    ++\t\t}\n    ++\t\tp--;\n    ++\t}\n    ++\tassert(nr == vec->count);\n     +\treturn vec;\n     +}\n     +\n    @@ bloom.h: struct bloom_key {\n      int load_bloom_filter_from_graph(struct commit_graph *g,\n      \t\t\t\t struct bloom_filter *filter,\n      \t\t\t\t uint32_t graph_pos);\n    -@@ bloom.h: void bloom_key_fill(const char *data,\n    +@@ bloom.h: void bloom_key_fill(struct bloom_key *key, const char *data, size_t len,\n      \t\t    const struct bloom_filter_settings *settings);\n      void bloom_key_clear(struct bloom_key *key);\n      \n    -+struct bloom_keyvec *bloom_keyvec_new(size_t count);\n    ++/*\n    ++ * bloom_keyvec_fill - Allocate and populate a bloom_keyvec with keys for the\n    ++ * given path.\n    ++ *\n    ++ * This function splits the input path by '/' and generates a bloom key for each\n    ++ * prefix, in reverse order of specificity. For example, given the input\n    ++ * \"a/b/c\", it will generate bloom keys for:\n    ++ *   - \"a/b/c\"\n    ++ *   - \"a/b\"\n    ++ *   - \"a\"\n    ++ *\n    ++ * The resulting keys are stored in a newly allocated bloom_keyvec.\n    ++ */\n    ++struct bloom_keyvec *bloom_keyvec_new(const char *path, size_t len,\n    ++\t\t\t\t      const struct bloom_filter_settings *settings);\n     +void bloom_keyvec_free(struct bloom_keyvec *vec);\n    -+\n    -+static inline void bloom_keyvec_fill_key(const char *data, size_t len,\n    -+\t\t\t\t\t struct bloom_keyvec *vec, size_t nr,\n    -+\t\t\t\t\t const struct bloom_filter_settings *settings)\n    -+{\n    -+\tassert(nr < vec->count);\n    -+\tbloom_key_fill(data, len, &vec->key[nr], settings);\n    -+}\n     +\n      void add_key_to_filter(const struct bloom_key *key,\n      \t\t       struct bloom_filter *filter,\n    @@ bloom.h: int bloom_filter_contains(const struct bloom_filter *filter,\n      \t\t\t  const struct bloom_key *key,\n      \t\t\t  const struct bloom_filter_settings *settings);\n      \n    ++/*\n    ++ * bloom_filter_contains_vec - Check if all keys in a key vector are in the\n    ++ * Bloom filter.\n    ++ *\n    ++ * Returns 1 if **all** keys in the vector are present in the filter,\n    ++ * 0 if **any** key is not present.\n    ++ */\n     +int bloom_filter_contains_vec(const struct bloom_filter *filter,\n     +\t\t\t      const struct bloom_keyvec *v,\n     +\t\t\t      const struct bloom_filter_settings *settings);\n    @@ bloom.h: int bloom_filter_contains(const struct bloom_filter *filter,\n     \n      ## revision.c ##\n     @@ revision.c: static int forbid_bloom_filters(struct pathspec *spec)\n    + \treturn 0;\n    + }\n    + \n    ++static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n    ++\n      static void prepare_to_use_bloom_filter(struct rev_info *revs)\n      {\n      \tstruct pathspec_item *pi;\n    @@ revision.c: static int forbid_bloom_filters(struct pathspec *spec)\n      \tchar *path_alloc = NULL;\n      \tconst char *path, *p;\n      \tsize_t len;\n    +-\tint path_component_nr = 1;\n    + \n    + \tif (!revs->commits)\n    + \t\treturn;\n     @@ revision.c: static void prepare_to_use_bloom_filter(struct rev_info *revs)\n    - \t\tp++;\n    - \t}\n    + \tif (!revs->pruning.pathspec.nr)\n    + \t\treturn;\n      \n    --\trevs->bloom_keys_nr = path_component_nr;\n    --\tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n     +\trevs->bloom_keyvecs_nr = 1;\n     +\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n    -+\tbloom_keyvec = bloom_keyvec_new(path_component_nr);\n    -+\trevs->bloom_keyvecs[0] = bloom_keyvec;\n    + \tpi = &revs->pruning.pathspec.items[0];\n    + \n    + \t/* remove single trailing slash from path, if needed */\n    +@@ revision.c: static void prepare_to_use_bloom_filter(struct rev_info *revs)\n    + \t\tpath = pi->match;\n      \n    --\tbloom_key_fill(path, len, &revs->bloom_keys[0],\n    + \tlen = strlen(path);\n    +-\tif (!len) {\n    +-\t\trevs->bloom_filter_settings = NULL;\n    +-\t\tfree(path_alloc);\n    +-\t\treturn;\n    +-\t}\n    +-\n    +-\tp = path;\n    +-\twhile (*p) {\n    +-\t\t/*\n    +-\t\t * At this point, the path is normalized to use Unix-style\n    +-\t\t * path separators. This is required due to how the\n    +-\t\t * changed-path Bloom filters store the paths.\n    +-\t\t */\n    +-\t\tif (*p == '/')\n    +-\t\t\tpath_component_nr++;\n    +-\t\tp++;\n    +-\t}\n    +-\n    +-\trevs->bloom_keys_nr = path_component_nr;\n    +-\tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n    ++\tif (!len)\n    ++\t\tgoto fail;\n    + \n    +-\tbloom_key_fill(&revs->bloom_keys[0], path, len,\n     -\t\t       revs->bloom_filter_settings);\n    -+\tbloom_keyvec_fill_key(path, len, bloom_keyvec, 0,\n    -+\t\t\t      revs->bloom_filter_settings);\n    - \tpath_component_nr = 1;\n    - \n    - \tp = path + len - 1;\n    - \twhile (p > path) {\n    - \t\tif (*p == '/')\n    --\t\t\tbloom_key_fill(path, p - path,\n    --\t\t\t\t       &revs->bloom_keys[path_component_nr++],\n    +-\tpath_component_nr = 1;\n    +-\n    +-\tp = path + len - 1;\n    +-\twhile (p > path) {\n    +-\t\tif (*p == '/')\n    +-\t\t\tbloom_key_fill(&revs->bloom_keys[path_component_nr++],\n    +-\t\t\t\t       path, p - path,\n     -\t\t\t\t       revs->bloom_filter_settings);\n    -+\t\t\tbloom_keyvec_fill_key(path, p - path, bloom_keyvec,\n    -+\t\t\t\t\t      path_component_nr++,\n    -+\t\t\t\t\t      revs->bloom_filter_settings);\n    - \t\tp--;\n    +-\t\tp--;\n    +-\t}\n    ++\trevs->bloom_keyvecs[0] =\n    ++\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n    + \n    + \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n    + \t\tatexit(trace2_bloom_filter_statistics_atexit);\n    + \t\tbloom_filter_atexit_registered = 1;\n      \t}\n      \n    -@@ revision.c: static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n    ++\treturn;\n    ++\n    ++fail:\n    ++\trevs->bloom_filter_settings = NULL;\n    + \tfree(path_alloc);\n    ++\trelease_revisions_bloom_keyvecs(revs);\n    + }\n    + \n    + static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n      \t\t\t\t\t\t struct commit *commit)\n      {\n      \tstruct bloom_filter *filter;\n    @@ revision.c: static int rev_same_tree_as_empty(struct rev_info *revs, struct comm\n      \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n      \t\tif (!bloom_ret)\n      \t\t\treturn 1;\n    +@@ revision.c: static void release_revisions_mailmap(struct string_list *mailmap)\n    + \n    + static void release_revisions_topo_walk_info(struct topo_walk_info *info);\n    + \n    ++static void release_revisions_bloom_keyvecs(struct rev_info *revs)\n    ++{\n    ++\tfor (size_t nr = 0; nr < revs->bloom_keyvecs_nr; nr++)\n    ++\t\tbloom_keyvec_free(revs->bloom_keyvecs[nr]);\n    ++\tFREE_AND_NULL(revs->bloom_keyvecs);\n    ++\trevs->bloom_keyvecs_nr = 0;\n    ++}\n    ++\n    + static void free_void_commit_list(void *list)\n    + {\n    + \tfree_commit_list(list);\n     @@ revision.c: void release_revisions(struct rev_info *revs)\n    + \tclear_decoration(&revs->treesame, free);\n      \tline_log_free(revs);\n      \toidset_clear(&revs->missing_commits);\n    - \n    +-\n     -\tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n     -\t\tbloom_key_clear(&revs->bloom_keys[i]);\n     -\tFREE_AND_NULL(revs->bloom_keys);\n     -\trevs->bloom_keys_nr = 0;\n    -+\tfor (size_t i = 0; i < revs->bloom_keyvecs_nr; i++)\n    -+\t\tbloom_keyvec_free(revs->bloom_keyvecs[i]);\n    -+\tFREE_AND_NULL(revs->bloom_keyvecs);\n    -+\trevs->bloom_keyvecs_nr = 0;\n    ++\trelease_revisions_bloom_keyvecs(revs);\n      }\n      \n      static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\n4:  de4d554b55 ! 4:  198a7da17c bloom: optimize multiple pathspec items in revision traversal\n    @@ Commit message\n     \n         To enable optimize multiple pathspec items in revision traversal,\n         return 0 if all pathspec item is literal in forbid_bloom_filters().\n    -    Add code to initialize and check each pathspec item's bloom_keyvec.\n    -\n    -    Add new function release_revisions_bloom_keyvecs() to free all bloom\n    -    keyvec owned by rev_info.\n    +    Add for loops to initialize and check each pathspec item's bloom_keyvec\n    +    when optimization is possible.\n     \n         Add new test cases in t/t4216-log-bloom.sh to ensure\n           - consistent results between the optimization for multiple pathspec\n    @@ revision.c: static int forbid_bloom_filters(struct pathspec *spec)\n      \n      \treturn 0;\n      }\n    - \n    -+static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n    -+\n    - static void prepare_to_use_bloom_filter(struct rev_info *revs)\n    - {\n    - \tstruct pathspec_item *pi;\n    -@@ revision.c: static void prepare_to_use_bloom_filter(struct rev_info *revs)\n    - \tchar *path_alloc = NULL;\n    - \tconst char *path, *p;\n    - \tsize_t len;\n    --\tint path_component_nr = 1;\n    -+\tint path_component_nr;\n    - \n    - \tif (!revs->commits)\n    - \t\treturn;\n     @@ revision.c: static void prepare_to_use_bloom_filter(struct rev_info *revs)\n      \tif (!revs->pruning.pathspec.nr)\n      \t\treturn;\n      \n    +-\trevs->bloom_keyvecs_nr = 1;\n    +-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n     -\tpi = &revs->pruning.pathspec.items[0];\n    --\n    ++\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n    ++\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n    ++\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n    ++\t\tpi = &revs->pruning.pathspec.items[i];\n    + \n     -\t/* remove single trailing slash from path, if needed */\n     -\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n     -\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n     -\t\tpath = path_alloc;\n     -\t} else\n     -\t\tpath = pi->match;\n    --\n    --\tlen = strlen(path);\n    --\tif (!len) {\n    --\t\trevs->bloom_filter_settings = NULL;\n    --\t\tfree(path_alloc);\n    --\t\treturn;\n    --\t}\n    --\n    --\tp = path;\n    --\twhile (*p) {\n    --\t\t/*\n    --\t\t * At this point, the path is normalized to use Unix-style\n    --\t\t * path separators. This is required due to how the\n    --\t\t * changed-path Bloom filters store the paths.\n    --\t\t */\n    --\t\tif (*p == '/')\n    --\t\t\tpath_component_nr++;\n    --\t\tp++;\n    --\t}\n    --\n    --\trevs->bloom_keyvecs_nr = 1;\n    --\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n    --\tbloom_keyvec = bloom_keyvec_new(path_component_nr);\n    --\trevs->bloom_keyvecs[0] = bloom_keyvec;\n    --\n    --\tbloom_keyvec_fill_key(path, len, bloom_keyvec, 0,\n    --\t\t\t      revs->bloom_filter_settings);\n    --\tpath_component_nr = 1;\n    -+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n    -+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n    -+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n    -+\t\tpi = &revs->pruning.pathspec.items[i];\n    -+\t\tpath_component_nr = 1;\n    -+\n     +\t\t/* remove single trailing slash from path, if needed */\n     +\t\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n     +\t\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n     +\t\t\tpath = path_alloc;\n     +\t\t} else\n     +\t\t\tpath = pi->match;\n    -+\n    + \n    +-\tlen = strlen(path);\n    +-\tif (!len)\n    +-\t\tgoto fail;\n     +\t\tlen = strlen(path);\n     +\t\tif (!len)\n     +\t\t\tgoto fail;\n    -+\n    -+\t\tp = path;\n    -+\t\twhile (*p) {\n    -+\t\t\t/*\n    -+\t\t\t * At this point, the path is normalized to use\n    -+\t\t\t * Unix-style path separators. This is required due to\n    -+\t\t\t * how the changed-path Bloom filters store the paths.\n    -+\t\t\t */\n    -+\t\t\tif (*p == '/')\n    -+\t\t\t\tpath_component_nr++;\n    -+\t\t\tp++;\n    -+\t\t}\n      \n    --\tp = path + len - 1;\n    --\twhile (p > path) {\n    --\t\tif (*p == '/')\n    --\t\t\tbloom_keyvec_fill_key(path, p - path, bloom_keyvec,\n    --\t\t\t\t\t      path_component_nr++,\n    --\t\t\t\t\t      revs->bloom_filter_settings);\n    --\t\tp--;\n    -+\t\tbloom_keyvec = bloom_keyvec_new(path_component_nr);\n    -+\t\trevs->bloom_keyvecs[i] = bloom_keyvec;\n    -+\n    -+\t\tbloom_keyvec_fill_key(path, len, bloom_keyvec, 0,\n    -+\t\t\t\t      revs->bloom_filter_settings);\n    -+\t\tpath_component_nr = 1;\n    -+\n    -+\t\tp = path + len - 1;\n    -+\t\twhile (p > path) {\n    -+\t\t\tif (*p == '/')\n    -+\t\t\t\tbloom_keyvec_fill_key(path, p - path,\n    -+\t\t\t\t\t\t      bloom_keyvec,\n    -+\t\t\t\t\t\t      path_component_nr++,\n    -+\t\t\t\t\t\t      revs->bloom_filter_settings);\n    -+\t\t\tp--;\n    -+\t\t}\n    +-\trevs->bloom_keyvecs[0] =\n    +-\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n    ++\t\trevs->bloom_keyvecs[i] =\n    ++\t\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n     +\t\tFREE_AND_NULL(path_alloc);\n    - \t}\n    ++\t}\n      \n      \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n    -@@ revision.c: static void prepare_to_use_bloom_filter(struct rev_info *revs)\n    - \t\tbloom_filter_atexit_registered = 1;\n    - \t}\n    - \n    -+\treturn;\n    -+\n    -+fail:\n    -+\trevs->bloom_filter_settings = NULL;\n    - \tfree(path_alloc);\n    -+\trelease_revisions_bloom_keyvecs(revs);\n    - }\n    - \n    - static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n    -@@ revision.c: static void release_revisions_mailmap(struct string_list *mailmap)\n    - \n    - static void release_revisions_topo_walk_info(struct topo_walk_info *info);\n    - \n    -+static void release_revisions_bloom_keyvecs(struct rev_info *revs)\n    -+{\n    -+\tfor (size_t nr = 0; nr < revs->bloom_keyvecs_nr; nr++)\n    -+\t\tbloom_keyvec_free(revs->bloom_keyvecs[nr]);\n    -+\tFREE_AND_NULL(revs->bloom_keyvecs);\n    -+\trevs->bloom_keyvecs_nr = 0;\n    -+}\n    -+\n    - static void free_void_commit_list(void *list)\n    - {\n    - \tfree_commit_list(list);\n    -@@ revision.c: void release_revisions(struct rev_info *revs)\n    - \tclear_decoration(&revs->treesame, free);\n    - \tline_log_free(revs);\n    - \toidset_clear(&revs->missing_commits);\n    --\n    --\tfor (size_t i = 0; i < revs->bloom_keyvecs_nr; i++)\n    --\t\tbloom_keyvec_free(revs->bloom_keyvecs[i]);\n    --\tFREE_AND_NULL(revs->bloom_keyvecs);\n    --\trevs->bloom_keyvecs_nr = 0;\n    -+\trelease_revisions_bloom_keyvecs(revs);\n    - }\n    - \n    - static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\n    + \t\tatexit(trace2_bloom_filter_statistics_atexit);\n     \n      ## t/t4216-log-bloom.sh ##\n     @@ t/t4216-log-bloom.sh: sane_unset GIT_TRACE2_CONFIG_PARAMS\n-- \n2.50.0.107.g33b6ec8c79\n\n"},{"id":"521741","messageId":"20250710084829.2171855-2-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250710084829.2171855-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v5 1/4] bloom: add test helper to return murmur3 hash","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-10T08:48:26Z","receivedAt":"2025-07-10T08:48:48Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"In bloom.h, murmur3_seeded_v2() is exported for the use of test murmur3\nhash. To clarify that murmur3_seeded_v2() is exported solely for testing\npurposes, a new helper function test_murmur3_seeded() was added instead\nof exporting murmur3_seeded_v2() directly.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n bloom.c               | 13 ++++++++++++-\n bloom.h               | 12 +++---------\n t/helper/test-bloom.c |  4 ++--\n 3 files changed, 17 insertions(+), 12 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 0c8d2cebf9..946c5e8c98 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -107,7 +107,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n  * Not considered to be cryptographically secure.\n  * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n  */\n-uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len)\n+static uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len)\n {\n \tconst uint32_t c1 = 0xcc9e2d51;\n \tconst uint32_t c2 = 0x1b873593;\n@@ -540,3 +540,14 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \n \treturn 1;\n }\n+\n+uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n+\t\t\t\t   int version)\n+{\n+\tassert(version == 1 || version == 2);\n+\n+\tif (version == 2)\n+\t\treturn murmur3_seeded_v2(seed, data, len);\n+\telse\n+\t\treturn murmur3_seeded_v1(seed, data, len);\n+}\ndiff --git a/bloom.h b/bloom.h\nindex 6e46489a20..a9ded1822f 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -78,15 +78,6 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n \n-/*\n- * Calculate the murmur3 32-bit hash value for the given data\n- * using the given seed.\n- * Produces a uniformly distributed hash value.\n- * Not considered to be cryptographically secure.\n- * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n- */\n-uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len);\n-\n void fill_bloom_key(const char *data,\n \t\t    size_t len,\n \t\t    struct bloom_key *key,\n@@ -137,4 +128,7 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n \t\t\t  const struct bloom_filter_settings *settings);\n \n+uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n+\t\t\t\t   int version);\n+\n #endif\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 9aa2c5a592..6a24b6e0a6 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -61,13 +61,13 @@ int cmd__bloom(int argc, const char **argv)\n \t\tuint32_t hashed;\n \t\tif (argc < 3)\n \t\t\tusage(bloom_usage);\n-\t\thashed = murmur3_seeded_v2(0, argv[2], strlen(argv[2]));\n+\t\thashed = test_bloom_murmur3_seeded(0, argv[2], strlen(argv[2]), 2);\n \t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n \t}\n \n \tif (!strcmp(argv[1], \"get_murmur3_seven_highbit\")) {\n \t\tuint32_t hashed;\n-\t\thashed = murmur3_seeded_v2(0, \"\\x99\\xaa\\xbb\\xcc\\xdd\\xee\\xff\", 7);\n+\t\thashed = test_bloom_murmur3_seeded(0, \"\\x99\\xaa\\xbb\\xcc\\xdd\\xee\\xff\", 7, 2);\n \t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n \t}\n \n-- \n2.50.0.107.g33b6ec8c79\n\n"},{"id":"521742","messageId":"20250710084829.2171855-3-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250710084829.2171855-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v5 2/4] bloom: rename function operates on bloom_key","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-10T08:48:27Z","receivedAt":"2025-07-10T08:48:51Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"git code style requires that functions operating on a struct S\nshould be named in the form S_verb. However, the functions operating\non struct bloom_key do not follow this convention. Therefore,\nfill_bloom_key() and clear_bloom_key() are renamed to bloom_key_fill()\nand bloom_key_clear(), respectively.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n blame.c               |  2 +-\n bloom.c               | 10 ++++------\n bloom.h               |  6 ++----\n line-log.c            |  5 +++--\n revision.c            |  8 ++++----\n t/helper/test-bloom.c |  4 ++--\n 6 files changed, 16 insertions(+), 19 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex 57daa45e89..811c6d8f9f 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -1310,7 +1310,7 @@ static void add_bloom_key(struct blame_bloom_data *bd,\n \t}\n \n \tbd->keys[bd->nr] = xmalloc(sizeof(struct bloom_key));\n-\tfill_bloom_key(path, strlen(path), bd->keys[bd->nr], bd->settings);\n+\tbloom_key_fill(bd->keys[bd->nr], path, strlen(path), bd->settings);\n \tbd->nr++;\n }\n \ndiff --git a/bloom.c b/bloom.c\nindex 946c5e8c98..5523d198c8 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -221,9 +221,7 @@ static uint32_t murmur3_seeded_v1(uint32_t seed, const char *data, size_t len)\n \treturn seed;\n }\n \n-void fill_bloom_key(const char *data,\n-\t\t    size_t len,\n-\t\t    struct bloom_key *key,\n+void bloom_key_fill(struct bloom_key *key, const char *data, size_t len,\n \t\t    const struct bloom_filter_settings *settings)\n {\n \tint i;\n@@ -243,7 +241,7 @@ void fill_bloom_key(const char *data,\n \t\tkey->hashes[i] = hash0 + i * hash1;\n }\n \n-void clear_bloom_key(struct bloom_key *key)\n+void bloom_key_clear(struct bloom_key *key)\n {\n \tFREE_AND_NULL(key->hashes);\n }\n@@ -500,9 +498,9 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \n \t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n \t\t\tstruct bloom_key key;\n-\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, settings);\n+\t\t\tbloom_key_fill(&key, e->path, strlen(e->path), settings);\n \t\t\tadd_key_to_filter(&key, filter, settings);\n-\t\t\tclear_bloom_key(&key);\n+\t\t\tbloom_key_clear(&key);\n \t\t}\n \n \tcleanup:\ndiff --git a/bloom.h b/bloom.h\nindex a9ded1822f..603bc1f90f 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -78,11 +78,9 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n \n-void fill_bloom_key(const char *data,\n-\t\t    size_t len,\n-\t\t    struct bloom_key *key,\n+void bloom_key_fill(struct bloom_key *key, const char *data, size_t len,\n \t\t    const struct bloom_filter_settings *settings);\n-void clear_bloom_key(struct bloom_key *key);\n+void bloom_key_clear(struct bloom_key *key);\n \n void add_key_to_filter(const struct bloom_key *key,\n \t\t       struct bloom_filter *filter,\ndiff --git a/line-log.c b/line-log.c\nindex 628e3fe3ae..07f2154e84 100644\n--- a/line-log.c\n+++ b/line-log.c\n@@ -1172,12 +1172,13 @@ static int bloom_filter_check(struct rev_info *rev,\n \t\treturn 0;\n \n \twhile (!result && range) {\n-\t\tfill_bloom_key(range->path, strlen(range->path), &key, rev->bloom_filter_settings);\n+\t\tbloom_key_fill(&key, range->path, strlen(range->path),\n+\t\t\t       rev->bloom_filter_settings);\n \n \t\tif (bloom_filter_contains(filter, &key, rev->bloom_filter_settings))\n \t\t\tresult = 1;\n \n-\t\tclear_bloom_key(&key);\n+\t\tbloom_key_clear(&key);\n \t\trange = range->next;\n \t}\n \ndiff --git a/revision.c b/revision.c\nindex afee111196..a7eadff0a5 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -739,15 +739,15 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \trevs->bloom_keys_nr = path_component_nr;\n \tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n \n-\tfill_bloom_key(path, len, &revs->bloom_keys[0],\n+\tbloom_key_fill(&revs->bloom_keys[0], path, len,\n \t\t       revs->bloom_filter_settings);\n \tpath_component_nr = 1;\n \n \tp = path + len - 1;\n \twhile (p > path) {\n \t\tif (*p == '/')\n-\t\t\tfill_bloom_key(path, p - path,\n-\t\t\t\t       &revs->bloom_keys[path_component_nr++],\n+\t\t\tbloom_key_fill(&revs->bloom_keys[path_component_nr++],\n+\t\t\t\t       path, p - path,\n \t\t\t\t       revs->bloom_filter_settings);\n \t\tp--;\n \t}\n@@ -3231,7 +3231,7 @@ void release_revisions(struct rev_info *revs)\n \toidset_clear(&revs->missing_commits);\n \n \tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n-\t\tclear_bloom_key(&revs->bloom_keys[i]);\n+\t\tbloom_key_clear(&revs->bloom_keys[i]);\n \tFREE_AND_NULL(revs->bloom_keys);\n \trevs->bloom_keys_nr = 0;\n }\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 6a24b6e0a6..3283544bd3 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -12,13 +12,13 @@ static struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n static void add_string_to_filter(const char *data, struct bloom_filter *filter) {\n \t\tstruct bloom_key key;\n \n-\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n+\t\tbloom_key_fill(&key, data, strlen(data), &settings);\n \t\tprintf(\"Hashes:\");\n \t\tfor (size_t i = 0; i < settings.num_hashes; i++)\n \t\t\tprintf(\"0x%08x|\", key.hashes[i]);\n \t\tprintf(\"\\n\");\n \t\tadd_key_to_filter(&key, filter, &settings);\n-\t\tclear_bloom_key(&key);\n+\t\tbloom_key_clear(&key);\n }\n \n static void print_bloom_filter(struct bloom_filter *filter) {\n-- \n2.50.0.107.g33b6ec8c79\n\n"},{"id":"521743","messageId":"20250710084829.2171855-4-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250710084829.2171855-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v5 3/4] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-10T08:48:28Z","receivedAt":"2025-07-10T08:48:55Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Previously, we stored bloom keys in a flat array and marked a commit\nas NOT TREESAME if any key reported \"definitely not changed\".\n\nTo support multiple pathspec items, we now require that for each\npathspec item, there exists a bloom key reporting \"definitely not\nchanged\".\n\nThis \"for every\" condition makes a flat array insufficient, so we\nintroduce a new structure to group keys by a single pathspec item.\n`struct bloom_keyvec` is introduced to replace `struct bloom_key *`\nand `bloom_key_nr`. And because we want to support multiple pathspec\nitems, we added a bloom_keyvec * and a bloom_keyvec_nr field to\n`struct rev_info` to represent an array of bloom_keyvecs. This commit\nstill optimize only one pathspec item, thus bloom_keyvec_nr can only\nbe 0 or 1.\n\nNew bloom_keyvec_* functions are added to create and destroy a keyvec.\nbloom_filter_contains_vec() is added to check if all key in keyvec is\ncontained in a bloom filter.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n bloom.c    | 61 ++++++++++++++++++++++++++++++++++++++++++++\n bloom.h    | 38 +++++++++++++++++++++++++++\n revision.c | 75 ++++++++++++++++++++++--------------------------------\n revision.h |  6 ++---\n 4 files changed, 132 insertions(+), 48 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 5523d198c8..b86015f6d1 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -278,6 +278,55 @@ void deinit_bloom_filters(void)\n \tdeep_clear_bloom_filter_slab(&bloom_filters, free_one_bloom_filter);\n }\n \n+struct bloom_keyvec *bloom_keyvec_new(const char *path, size_t len,\n+\t\t\t\t      const struct bloom_filter_settings *settings)\n+{\n+\tstruct bloom_keyvec *vec;\n+\tconst char *p;\n+\tsize_t sz;\n+\tsize_t nr = 1;\n+\n+\tp = path;\n+\twhile (*p) {\n+\t\t/*\n+\t\t * At this point, the path is normalized to use Unix-style\n+\t\t * path separators. This is required due to how the\n+\t\t * changed-path Bloom filters store the paths.\n+\t\t */\n+\t\tif (*p == '/')\n+\t\t\tnr++;\n+\t\tp++;\n+\t}\n+\n+\tsz = sizeof(struct bloom_keyvec);\n+\tsz += nr * sizeof(struct bloom_key);\n+\tvec = (struct bloom_keyvec *)xcalloc(1, sz);\n+\tif (!vec)\n+\t\treturn NULL;\n+\tvec->count = nr;\n+\n+\tbloom_key_fill(&vec->key[0], path, len, settings);\n+\tnr = 1;\n+\tp = path + len - 1;\n+\twhile (p > path) {\n+\t\tif (*p == '/') {\n+\t\t\tbloom_key_fill(&vec->key[nr++], path, p - path, settings);\n+\t\t}\n+\t\tp--;\n+\t}\n+\tassert(nr == vec->count);\n+\treturn vec;\n+}\n+\n+void bloom_keyvec_free(struct bloom_keyvec *vec)\n+{\n+\tif (!vec)\n+\t\treturn;\n+\tfor (size_t nr = 0; nr < vec->count; nr++)\n+\t\tbloom_key_clear(&vec->key[nr]);\n+\tfree(vec);\n+}\n+\n static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n \t\t       const struct hashmap_entry *eptr,\n \t\t       const struct hashmap_entry *entry_or_key,\n@@ -539,6 +588,18 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \treturn 1;\n }\n \n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *vec,\n+\t\t\t      const struct bloom_filter_settings *settings)\n+{\n+\tint ret = 1;\n+\n+\tfor (size_t nr = 0; ret > 0 && nr < vec->count; nr++)\n+\t\tret = bloom_filter_contains(filter, &vec->key[nr], settings);\n+\n+\treturn ret;\n+}\n+\n uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n \t\t\t\t   int version)\n {\ndiff --git a/bloom.h b/bloom.h\nindex 603bc1f90f..3602d32054 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -74,6 +74,16 @@ struct bloom_key {\n \tuint32_t *hashes;\n };\n \n+/*\n+ * A bloom_keyvec is a vector of bloom_keys, which\n+ * can be used to store multiple keys for a single\n+ * pathspec item.\n+ */\n+struct bloom_keyvec {\n+\tsize_t count;\n+\tstruct bloom_key key[FLEX_ARRAY];\n+};\n+\n int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n@@ -82,6 +92,23 @@ void bloom_key_fill(struct bloom_key *key, const char *data, size_t len,\n \t\t    const struct bloom_filter_settings *settings);\n void bloom_key_clear(struct bloom_key *key);\n \n+/*\n+ * bloom_keyvec_fill - Allocate and populate a bloom_keyvec with keys for the\n+ * given path.\n+ *\n+ * This function splits the input path by '/' and generates a bloom key for each\n+ * prefix, in reverse order of specificity. For example, given the input\n+ * \"a/b/c\", it will generate bloom keys for:\n+ *   - \"a/b/c\"\n+ *   - \"a/b\"\n+ *   - \"a\"\n+ *\n+ * The resulting keys are stored in a newly allocated bloom_keyvec.\n+ */\n+struct bloom_keyvec *bloom_keyvec_new(const char *path, size_t len,\n+\t\t\t\t      const struct bloom_filter_settings *settings);\n+void bloom_keyvec_free(struct bloom_keyvec *vec);\n+\n void add_key_to_filter(const struct bloom_key *key,\n \t\t       struct bloom_filter *filter,\n \t\t       const struct bloom_filter_settings *settings);\n@@ -126,6 +153,17 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n \t\t\t  const struct bloom_filter_settings *settings);\n \n+/*\n+ * bloom_filter_contains_vec - Check if all keys in a key vector are in the\n+ * Bloom filter.\n+ *\n+ * Returns 1 if **all** keys in the vector are present in the filter,\n+ * 0 if **any** key is not present.\n+ */\n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *v,\n+\t\t\t      const struct bloom_filter_settings *settings);\n+\n uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n \t\t\t\t   int version);\n \ndiff --git a/revision.c b/revision.c\nindex a7eadff0a5..22bcfab7f9 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -685,13 +685,15 @@ static int forbid_bloom_filters(struct pathspec *spec)\n \treturn 0;\n }\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n+\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n+\tstruct bloom_keyvec *bloom_keyvec;\n \tchar *path_alloc = NULL;\n \tconst char *path, *p;\n \tsize_t len;\n-\tint path_component_nr = 1;\n \n \tif (!revs->commits)\n \t\treturn;\n@@ -708,6 +710,8 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n+\trevs->bloom_keyvecs_nr = 1;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n \tpi = &revs->pruning.pathspec.items[0];\n \n \t/* remove single trailing slash from path, if needed */\n@@ -718,53 +722,30 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \t\tpath = pi->match;\n \n \tlen = strlen(path);\n-\tif (!len) {\n-\t\trevs->bloom_filter_settings = NULL;\n-\t\tfree(path_alloc);\n-\t\treturn;\n-\t}\n-\n-\tp = path;\n-\twhile (*p) {\n-\t\t/*\n-\t\t * At this point, the path is normalized to use Unix-style\n-\t\t * path separators. This is required due to how the\n-\t\t * changed-path Bloom filters store the paths.\n-\t\t */\n-\t\tif (*p == '/')\n-\t\t\tpath_component_nr++;\n-\t\tp++;\n-\t}\n-\n-\trevs->bloom_keys_nr = path_component_nr;\n-\tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n+\tif (!len)\n+\t\tgoto fail;\n \n-\tbloom_key_fill(&revs->bloom_keys[0], path, len,\n-\t\t       revs->bloom_filter_settings);\n-\tpath_component_nr = 1;\n-\n-\tp = path + len - 1;\n-\twhile (p > path) {\n-\t\tif (*p == '/')\n-\t\t\tbloom_key_fill(&revs->bloom_keys[path_component_nr++],\n-\t\t\t\t       path, p - path,\n-\t\t\t\t       revs->bloom_filter_settings);\n-\t\tp--;\n-\t}\n+\trevs->bloom_keyvecs[0] =\n+\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n \t\tatexit(trace2_bloom_filter_statistics_atexit);\n \t\tbloom_filter_atexit_registered = 1;\n \t}\n \n+\treturn;\n+\n+fail:\n+\trevs->bloom_filter_settings = NULL;\n \tfree(path_alloc);\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\t\t\t\t\t struct commit *commit)\n {\n \tstruct bloom_filter *filter;\n-\tint result = 1, j;\n+\tint result = 0;\n \n \tif (!revs->repo->objects->commit_graph)\n \t\treturn -1;\n@@ -779,10 +760,10 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\treturn -1;\n \t}\n \n-\tfor (j = 0; result && j < revs->bloom_keys_nr; j++) {\n-\t\tresult = bloom_filter_contains(filter,\n-\t\t\t\t\t       &revs->bloom_keys[j],\n-\t\t\t\t\t       revs->bloom_filter_settings);\n+\tfor (size_t nr = 0; !result && nr < revs->bloom_keyvecs_nr; nr++) {\n+\t\tresult = bloom_filter_contains_vec(filter,\n+\t\t\t\t\t\t   revs->bloom_keyvecs[nr],\n+\t\t\t\t\t\t   revs->bloom_filter_settings);\n \t}\n \n \tif (result)\n@@ -823,7 +804,7 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\treturn REV_TREE_SAME;\n \t}\n \n-\tif (revs->bloom_keys_nr && !nth_parent) {\n+\tif (revs->bloom_keyvecs_nr && !nth_parent) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \n \t\tif (bloom_ret == 0)\n@@ -850,7 +831,7 @@ static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit,\n \tif (!t1)\n \t\treturn 0;\n \n-\tif (!nth_parent && revs->bloom_keys_nr) {\n+\tif (!nth_parent && revs->bloom_keyvecs_nr) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \t\tif (!bloom_ret)\n \t\t\treturn 1;\n@@ -3201,6 +3182,14 @@ static void release_revisions_mailmap(struct string_list *mailmap)\n \n static void release_revisions_topo_walk_info(struct topo_walk_info *info);\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs)\n+{\n+\tfor (size_t nr = 0; nr < revs->bloom_keyvecs_nr; nr++)\n+\t\tbloom_keyvec_free(revs->bloom_keyvecs[nr]);\n+\tFREE_AND_NULL(revs->bloom_keyvecs);\n+\trevs->bloom_keyvecs_nr = 0;\n+}\n+\n static void free_void_commit_list(void *list)\n {\n \tfree_commit_list(list);\n@@ -3229,11 +3218,7 @@ void release_revisions(struct rev_info *revs)\n \tclear_decoration(&revs->treesame, free);\n \tline_log_free(revs);\n \toidset_clear(&revs->missing_commits);\n-\n-\tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n-\t\tbloom_key_clear(&revs->bloom_keys[i]);\n-\tFREE_AND_NULL(revs->bloom_keys);\n-\trevs->bloom_keys_nr = 0;\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\ndiff --git a/revision.h b/revision.h\nindex 6d369cdad6..ac843f58d0 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -62,7 +62,7 @@ struct repository;\n struct rev_info;\n struct string_list;\n struct saved_parents;\n-struct bloom_key;\n+struct bloom_keyvec;\n struct bloom_filter_settings;\n struct option;\n struct parse_opt_ctx_t;\n@@ -360,8 +360,8 @@ struct rev_info {\n \n \t/* Commit graph bloom filter fields */\n \t/* The bloom filter key(s) for the pathspec */\n-\tstruct bloom_key *bloom_keys;\n-\tint bloom_keys_nr;\n+\tstruct bloom_keyvec **bloom_keyvecs;\n+\tint bloom_keyvecs_nr;\n \n \t/*\n \t * The bloom filter settings used to generate the key.\n-- \n2.50.0.107.g33b6ec8c79\n\n"},{"id":"521744","messageId":"20250710084829.2171855-5-502024330056@smail.nju.edu.cn","threadId":"63691","inReplyTo":"20250710084829.2171855-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v5 4/4] bloom: optimize multiple pathspec items in revision traversal","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-10T08:48:29Z","receivedAt":"2025-07-10T08:48:58Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"To enable optimize multiple pathspec items in revision traversal,\nreturn 0 if all pathspec item is literal in forbid_bloom_filters().\nAdd for loops to initialize and check each pathspec item's bloom_keyvec\nwhen optimization is possible.\n\nAdd new test cases in t/t4216-log-bloom.sh to ensure\n  - consistent results between the optimization for multiple pathspec\n    items using bloom filter and the case without bloom filter\n    optimization.\n  - does not use bloom filter if any pathspec item is not literal.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n revision.c           | 38 ++++++++++++++++++++------------------\n t/t4216-log-bloom.sh | 23 ++++++++++++++---------\n 2 files changed, 34 insertions(+), 27 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 22bcfab7f9..f25a61bb6c 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -675,12 +675,11 @@ static int forbid_bloom_filters(struct pathspec *spec)\n {\n \tif (spec->has_wildcard)\n \t\treturn 1;\n-\tif (spec->nr > 1)\n-\t\treturn 1;\n \tif (spec->magic & ~PATHSPEC_LITERAL)\n \t\treturn 1;\n-\tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n-\t\treturn 1;\n+\tfor (size_t nr = 0; nr < spec->nr; nr++)\n+\t\tif (spec->items[nr].magic & ~PATHSPEC_LITERAL)\n+\t\t\treturn 1;\n \n \treturn 0;\n }\n@@ -710,23 +709,26 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n-\trevs->bloom_keyvecs_nr = 1;\n-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n-\tpi = &revs->pruning.pathspec.items[0];\n+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n+\t\tpi = &revs->pruning.pathspec.items[i];\n \n-\t/* remove single trailing slash from path, if needed */\n-\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n-\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n-\t\tpath = path_alloc;\n-\t} else\n-\t\tpath = pi->match;\n+\t\t/* remove single trailing slash from path, if needed */\n+\t\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n+\t\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n+\t\t\tpath = path_alloc;\n+\t\t} else\n+\t\t\tpath = pi->match;\n \n-\tlen = strlen(path);\n-\tif (!len)\n-\t\tgoto fail;\n+\t\tlen = strlen(path);\n+\t\tif (!len)\n+\t\t\tgoto fail;\n \n-\trevs->bloom_keyvecs[0] =\n-\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n+\t\trevs->bloom_keyvecs[i] =\n+\t\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n+\t\tFREE_AND_NULL(path_alloc);\n+\t}\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n \t\tatexit(trace2_bloom_filter_statistics_atexit);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 8910d53cac..639868ac56 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -66,8 +66,9 @@ sane_unset GIT_TRACE2_CONFIG_PARAMS\n \n setup () {\n \trm -f \"$TRASH_DIRECTORY/trace.perf\" &&\n-\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n-\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n+\teval git -c core.commitGraph=false log --pretty=\"format:%s\" \"$1\" >log_wo_bloom &&\n+\teval \"GIT_TRACE2_PERF=\\\"$TRASH_DIRECTORY/trace.perf\\\"\" \\\n+\t\tgit -c core.commitGraph=true log --pretty=\"format:%s\" \"$1\" >log_w_bloom\n }\n \n test_bloom_filters_used () {\n@@ -138,10 +139,6 @@ test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n '\n \n-test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n-'\n-\n test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"-- .\"\n '\n@@ -151,9 +148,17 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n \ttest_bloom_filters_used \"-- *renamed\"\n '\n \n-test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- *\" &&\n-\ttest_bloom_filters_not_used \"-- file*\"\n+test_expect_success 'git log with multiple literal paths uses Bloom filter' '\n+\ttest_bloom_filters_used \"-- file4 A/file1\" &&\n+\ttest_bloom_filters_used \"-- *\" &&\n+\ttest_bloom_filters_used \"-- file*\"\n+'\n+\n+test_expect_success 'git log with path contains a wildcard does not use Bloom filter' '\n+\ttest_bloom_filters_not_used \"-- file\\*\" &&\n+\ttest_bloom_filters_not_used \"-- A/\\* file4\" &&\n+\ttest_bloom_filters_not_used \"-- file4 A/\\*\" &&\n+\ttest_bloom_filters_not_used \"-- * A/\\*\"\n '\n \n test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n-- \n2.50.0.107.g33b6ec8c79\n\n"},{"id":"521751","messageId":"afb68948-218b-4b56-9faa-29578ef9c73c@gmail.com","threadId":"63691","inReplyTo":"20250710084829.2171855-1-502024330056@smail.nju.edu.cn","subject":"Re: [PATCH v5 0/4] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2025-07-10T13:49:50Z","receivedAt":"2025-07-10T13:49:53Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/10/2025 4:48 AM, Lidong Yan wrote:\n\n> The difference from v4 is:\n>   - for the bloom_key_* functions, we now pass struct bloom_key *\n>     as the first parameter.\n>   - bloom_keyvec_fill_key() and bloom_keyvec_new() have been merged\n>     into a single function, so that each key is filled during the\n>     creation of the bloom_keyvec.\n> \n> Below is a comparison of the time taken to run git log on Git and\n> LLVM repositories before and after applying this patch.\n> \n> Setup commit-graph:\n>   $ cd ~/src/git && git commit-graph write --split --reachable --changed-paths\n>   $ cd ~/src/llvm && git commit-graph write --split --reachable --changed-paths\n> \n> Run git log on Git repository\n>   $ cd ~/src/git\n>   $ hash -p /usr/bin/git git # my system git binary is 2.43.0\n>   $ time git log -100 -- commit.c commit-graph.c >/dev/null\n>   real\t0m0.055s\n>   user\t0m0.040s\n>   sys\t  0m0.015s\n>   $ hash -p ~/bin/git/bin/git git\n>   $ time git log -100 -- commit.c commit-graph.c >/dev/null\n>   real\t0m0.039s\n>   user\t0m0.020s\n>   sys\t  0m0.020s\n> \n> Run git log in LLVM repository\n>   $ cd ~/src/llvm\n>   $ hash -p /usr/bin/git git # my system git binary is 2.43.0\n>   $ time git log -100 -- llvm/lib/Support/CommandLine.cpp llvm/lib/Support/CommandLine.h >/dev/null\n>   real\t0m2.365s\n>   user\t0m2.252s\n>   sys\t  0m0.113s\n>   $ hash -p ~/bin/git/bin/git git\n>   $ time git log -100 -- llvm/lib/Support/CommandLine.cpp llvm/lib/Support/CommandLine.h >/dev/null\n>   real\t0m0.240s\n>   user\t0m0.158s\n>   sys\t  0m0.064s\n\nThanks for these stats, though I'd recommend trying hyperfine [1]\nwhich with this setup would let you run these experiments as the\nfollowing:\n\n$ $ hyperfine --warmup=3 \\\n> -n 'old' '~/_git/git-sparse-checkout-clean/git log -100 -- commit.c commit-graph.c' \\\n> -n 'new' '~/_git/git/git log -100 -- commit.c commit-graph.c'\nBenchmark 1: old\n  Time (mean ± σ):      73.1 ms ±   2.9 ms    [User: 48.8 ms, System: 23.9 ms]\n  Range (min … max):    69.9 ms …  84.5 ms    42 runs\n \nBenchmark 2: new\n  Time (mean ± σ):      55.1 ms ±   2.9 ms    [User: 30.5 ms, System: 24.4 ms]\n  Range (min … max):    51.1 ms …  61.2 ms    52 runs\n \nSummary\n  'new' ran\n    1.33 ± 0.09 times faster than 'old'\n\nAnd for LLVM:\n\n$ hyperfine --warmup=3 \\\n> -n 'old' '~/_git/git-sparse-checkout-clean/git log -100 -- llvm/lib/Support/CommandLine.cpp llvm/lib/Support/CommandLine.h' \\\n> -n 'new' '~/_git/git/git log -100 -- llvm/lib/Support/CommandLine.cpp llvm/lib/Support/CommandLine.h'\n\nBenchmark 1: old\n  Time (mean ± σ):      1.974 s ±  0.006 s    [User: 1.877 s, System: 0.097 s]\n  Range (min … max):    1.960 s …  1.983 s    10 runs\n \nBenchmark 2: new\n  Time (mean ± σ):     262.9 ms ±   2.4 ms    [User: 214.2 ms, System: 48.4 ms]\n  Range (min … max):   257.7 ms … 266.2 ms    11 runs\n \nSummary\n  'new' ran\n    7.51 ± 0.07 times faster than 'old'\n\n[1] https://github.com/sharkdp/hyperfine\n\nFinally, putting these performance numbers in the commit message will\nmake the results permanently findable in the repo history.\n\n> 3:  c042907b92 ! 3:  60a3b16bbb bloom: replace struct bloom_key * with struct bloom_keyvec\n\n>     -+struct bloom_keyvec *bloom_keyvec_new(size_t count)\n>     ++struct bloom_keyvec *bloom_keyvec_new(const char *path, size_t len,\n>     ++\t\t\t\t      const struct bloom_filter_settings *settings)\n>      +{\n>      +\tstruct bloom_keyvec *vec;\n>     -+\tsize_t sz = sizeof(struct bloom_keyvec);\n>     -+\tsz += count * sizeof(struct bloom_key);\n>     ++\tconst char *p;\n>     ++\tsize_t sz;\n>     ++\tsize_t nr = 1;\n>     ++\n>     ++\tp = path;\n>     ++\twhile (*p) {\n>     ++\t\t/*\n>     ++\t\t * At this point, the path is normalized to use Unix-style\n>     ++\t\t * path separators. This is required due to how the\n>     ++\t\t * changed-path Bloom filters store the paths.\n>     ++\t\t */\n>     ++\t\tif (*p == '/')\n>     ++\t\t\tnr++;\n>     ++\t\tp++;\n>     ++\t}\n>     ++\n>     ++\tsz = sizeof(struct bloom_keyvec);\n>     ++\tsz += nr * sizeof(struct bloom_key);\n>      +\tvec = (struct bloom_keyvec *)xcalloc(1, sz);\n>     -+\tvec->count = count;\n>     ++\tif (!vec)\n>     ++\t\treturn NULL;\n>     ++\tvec->count = nr;\n>     ++\n>     ++\tbloom_key_fill(&vec->key[0], path, len, settings);\n>     ++\tnr = 1;\n>     ++\tp = path + len - 1;\n>     ++\twhile (p > path) {\n>     ++\t\tif (*p == '/') {\n>     ++\t\t\tbloom_key_fill(&vec->key[nr++], path, p - path, settings);\n>     ++\t\t}\n>     ++\t\tp--;\n>     ++\t}\n>     ++\tassert(nr == vec->count);\n>      +\treturn vec;\n>      +}\n\nThese additions to bloom_keyvec_new() certainly help simplify some\nof the code movement I was talking about, but there is more that\ncan be done to simplify patch 4. I'll send a couple example patches\nin reply to patch 4 with what I mean.\n\nOverall, the patch series is correct. My complaints are stylistic\nmore than anything.\n\nThanks,\n-Stolee\n\n"},{"id":"521752","messageId":"3c59af48-23fe-4cc7-87e9-1de94f509a2b@gmail.com","threadId":"63691","inReplyTo":"20250710084829.2171855-5-502024330056@smail.nju.edu.cn","subject":"[PATCH v5.1 3.5/4] revision: make helper for pathspec to bloom key","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2025-07-10T13:51:13Z","receivedAt":"2025-07-10T13:51:15Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/10/2025 4:48 AM, Lidong Yan wrote:\n> @@ -710,23 +709,26 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n>  \tif (!revs->pruning.pathspec.nr)\n>  \t\treturn;\n>  \n> -\trevs->bloom_keyvecs_nr = 1;\n> -\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n> -\tpi = &revs->pruning.pathspec.items[0];\n> +\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n> +\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n> +\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n> +\t\tpi = &revs->pruning.pathspec.items[i];\n>  \n> -\t/* remove single trailing slash from path, if needed */\n> -\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n> -\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n> -\t\tpath = path_alloc;\n> -\t} else\n> -\t\tpath = pi->match;\n> +\t\t/* remove single trailing slash from path, if needed */\n> +\t\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n> +\t\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n> +\t\t\tpath = path_alloc;\n> +\t\t} else\n> +\t\t\tpath = pi->match;\n>  \n> -\tlen = strlen(path);\n> -\tif (!len)\n> -\t\tgoto fail;\n> +\t\tlen = strlen(path);\n> +\t\tif (!len)\n> +\t\t\tgoto fail;\n>  \n> -\trevs->bloom_keyvecs[0] =\n> -\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n> +\t\trevs->bloom_keyvecs[i] =\n> +\t\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n> +\t\tFREE_AND_NULL(path_alloc);\n> +\t}\n\nThis diff is still bigger than I was hoping, so I'm sending a couple\nof patches that simplify this code movement. Feel free to ignore\nthem as being too nit-picky.\n\n--- >8 ---\n\nFrom 69fa36dc615e140ae842b536f7da792beaebb272 Mon Sep 17 00:00:00 2001\nFrom: Derrick Stolee <stolee@gmail.com>\nDate: Thu, 10 Jul 2025 08:06:29 -0400\nSubject: [PATCH v5.1 3.5/4] revision: make helper for pathspec to bloom key\n\nWhen preparing to use bloom filters in a revision walk, Git populates a\nboom_keyvec with an array of bloom keys for the components of a path.\nBefore we create the ability to map multiple pathspecs to multiple\nbloom_keyvecs, extract the conversion from a pathspec to a bloom_keyvec\ninto its own helper method. This simplifies the state that persists in\nprepare_to_use_bloom_filter() as well as makes the next change much\nsimpler.\n\nSigned-off-by: Derrick Stolee <stolee@gmail.com>\n---\n revision.c | 50 +++++++++++++++++++++++++++++++-------------------\n 1 file changed, 31 insertions(+), 19 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 22bcfab7f93..4c09b594c55 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -687,14 +687,37 @@ static int forbid_bloom_filters(struct pathspec *spec)\n \n static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n \n-static void prepare_to_use_bloom_filter(struct rev_info *revs)\n+static int convert_pathspec_to_filter(const struct pathspec_item *pi,\n+\t\t\t\t      struct bloom_keyvec **bloom_keyvec,\n+\t\t\t\t      const struct bloom_filter_settings *settings)\n {\n-\tstruct pathspec_item *pi;\n-\tstruct bloom_keyvec *bloom_keyvec;\n-\tchar *path_alloc = NULL;\n-\tconst char *path, *p;\n \tsize_t len;\n+\tconst char *path;\n+\tchar *path_alloc = NULL;\n+\tint res = 0;\n+\n+\t/* remove single trailing slash from path, if needed */\n+\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n+\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n+\t\tpath = path_alloc;\n+\t} else\n+\t\tpath = pi->match;\n+\n+\tlen = strlen(path);\n+\tif (!len) {\n+\t\tres = -1;\n+\t\tgoto cleanup;\n+\t}\n+\n+\t*bloom_keyvec = bloom_keyvec_new(path, len, settings);\n \n+cleanup:\n+\tFREE_AND_NULL(path_alloc);\n+\treturn res;\n+}\n+\n+static void prepare_to_use_bloom_filter(struct rev_info *revs)\n+{\n \tif (!revs->commits)\n \t\treturn;\n \n@@ -712,22 +735,12 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \n \trevs->bloom_keyvecs_nr = 1;\n \tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n-\tpi = &revs->pruning.pathspec.items[0];\n \n-\t/* remove single trailing slash from path, if needed */\n-\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n-\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n-\t\tpath = path_alloc;\n-\t} else\n-\t\tpath = pi->match;\n-\n-\tlen = strlen(path);\n-\tif (!len)\n+\tif (convert_pathspec_to_filter(&revs->pruning.pathspec.items[0],\n+\t\t\t\t       &revs->bloom_keyvecs[0],\n+\t\t\t\t       revs->bloom_filter_settings))\n \t\tgoto fail;\n \n-\trevs->bloom_keyvecs[0] =\n-\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n-\n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n \t\tatexit(trace2_bloom_filter_statistics_atexit);\n \t\tbloom_filter_atexit_registered = 1;\n@@ -737,7 +750,6 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \n fail:\n \trevs->bloom_filter_settings = NULL;\n-\tfree(path_alloc);\n \trelease_revisions_bloom_keyvecs(revs);\n }\n \n-- \n2.47.2.vfs.0.2\n\n\n"},{"id":"521753","messageId":"2619038e-05f5-4af8-bb20-e4e01138f839@gmail.com","threadId":"63691","inReplyTo":"20250710084829.2171855-5-502024330056@smail.nju.edu.cn","subject":"[PATCH v5.1 4/4] bloom: optimize multiple pathspec items in revision","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2025-07-10T13:55:15Z","receivedAt":"2025-07-10T13:55:17Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/10/2025 4:48 AM, Lidong Yan wrote:\n> @@ -710,23 +709,26 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n>  \tif (!revs->pruning.pathspec.nr)\n>  \t\treturn;\n>  \n> -\trevs->bloom_keyvecs_nr = 1;\n> -\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n> -\tpi = &revs->pruning.pathspec.items[0];\n> +\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n> +\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n> +\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n> +\t\tpi = &revs->pruning.pathspec.items[i];\n>  \n> -\t/* remove single trailing slash from path, if needed */\n> -\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n> -\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n> -\t\tpath = path_alloc;\n> -\t} else\n> -\t\tpath = pi->match;\n> +\t\t/* remove single trailing slash from path, if needed */\n> +\t\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n> +\t\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n> +\t\t\tpath = path_alloc;\n> +\t\t} else\n> +\t\t\tpath = pi->match;\n>  \n> -\tlen = strlen(path);\n> -\tif (!len)\n> -\t\tgoto fail;\n> +\t\tlen = strlen(path);\n> +\t\tif (!len)\n> +\t\t\tgoto fail;\n>  \n> -\trevs->bloom_keyvecs[0] =\n> -\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n> +\t\trevs->bloom_keyvecs[i] =\n> +\t\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n> +\t\tFREE_AND_NULL(path_alloc);\n> +\t}\n\nFocus on the change to this diff when the patch below is applied on\ntop of the 3.5/4 I sent earlier, resulting in this diff:\n\n@@ -733,13 +732,14 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n-\trevs->bloom_keyvecs_nr = 1;\n-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n-\n-\tif (convert_pathspec_to_filter(&revs->pruning.pathspec.items[0],\n-\t\t\t\t       &revs->bloom_keyvecs[0],\n-\t\t\t\t       revs->bloom_filter_settings))\n-\t\tgoto fail;\n+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n+\t\tif (convert_pathspec_to_filter(&revs->pruning.pathspec.items[i],\n+\t\t\t\t\t       &revs->bloom_keyvecs[i],\n+\t\t\t\t\t       revs->bloom_filter_settings))\n+\t\t\tgoto fail;\n+\t}\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n \t\tatexit(trace2_bloom_filter_statistics_atexit);\n\nAlso, I've included the hyperfine performance output in the commit\nmessage.\n\nThanks,\n-Stolee\n\n\n--- >8 ---\n\nFrom fe255b1acfbe90fa8e4c335435ae18ee95e6243c Mon Sep 17 00:00:00 2001\nFrom: Lidong Yan <502024330056@smail.nju.edu.cn>\nDate: Thu, 10 Jul 2025 08:04:34 -0400\nSubject: [PATCH v5.1 4/4] bloom: optimize multiple pathspec items in revision\n traversal\n\nTo enable optimize multiple pathspec items in revision traversal,\nreturn 0 if all pathspec item is literal in forbid_bloom_filters().\nAdd for loops to initialize and check each pathspec item's bloom_keyvec\nwhen optimization is possible.\n\nAdd new test cases in t/t4216-log-bloom.sh to ensure\n  - consistent results between the optimization for multiple pathspec\n    items using bloom filter and the case without bloom filter\n    optimization.\n  - does not use bloom filter if any pathspec item is not literal.\n\nWith these optimizations, we get some improvements for multi-pathspec runs\nof 'git log'. First, in the Git repository we see these modest results:\n\nBenchmark 1: old\n  Time (mean ± σ):      73.1 ms ±   2.9 ms\n  Range (min … max):    69.9 ms …  84.5 ms    42 runs\n\nBenchmark 2: new\n  Time (mean ± σ):      55.1 ms ±   2.9 ms\n  Range (min … max):    51.1 ms …  61.2 ms    52 runs\n\nSummary\n  'new' ran\n    1.33 ± 0.09 times faster than 'old'\n\nBut in a larger repo, such as the LLVM project repo below, we get even\nbetter results:\n\nBenchmark 1: old\n  Time (mean ± σ):      1.974 s ±  0.006 s\n  Range (min … max):    1.960 s …  1.983 s    10 runs\n\nBenchmark 2: new\n  Time (mean ± σ):     262.9 ms ±   2.4 ms\n  Range (min … max):   257.7 ms … 266.2 ms    11 runs\n\nSummary\n  'new' ran\n    7.51 ± 0.07 times faster than 'old'\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\nSigned-off-by: Derrick Stolee <stolee@gmail.com>\n---\n revision.c           | 22 +++++++++++-----------\n t/t4216-log-bloom.sh | 23 ++++++++++++++---------\n 2 files changed, 25 insertions(+), 20 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 4c09b594c55..ca8c1dde8ca 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -675,12 +675,11 @@ static int forbid_bloom_filters(struct pathspec *spec)\n {\n \tif (spec->has_wildcard)\n \t\treturn 1;\n-\tif (spec->nr > 1)\n-\t\treturn 1;\n \tif (spec->magic & ~PATHSPEC_LITERAL)\n \t\treturn 1;\n-\tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n-\t\treturn 1;\n+\tfor (size_t nr = 0; nr < spec->nr; nr++)\n+\t\tif (spec->items[nr].magic & ~PATHSPEC_LITERAL)\n+\t\t\treturn 1;\n \n \treturn 0;\n }\n@@ -733,13 +732,14 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n-\trevs->bloom_keyvecs_nr = 1;\n-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n-\n-\tif (convert_pathspec_to_filter(&revs->pruning.pathspec.items[0],\n-\t\t\t\t       &revs->bloom_keyvecs[0],\n-\t\t\t\t       revs->bloom_filter_settings))\n-\t\tgoto fail;\n+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n+\t\tif (convert_pathspec_to_filter(&revs->pruning.pathspec.items[i],\n+\t\t\t\t\t       &revs->bloom_keyvecs[i],\n+\t\t\t\t\t       revs->bloom_filter_settings))\n+\t\t\tgoto fail;\n+\t}\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n \t\tatexit(trace2_bloom_filter_statistics_atexit);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 8910d53cac1..639868ac562 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -66,8 +66,9 @@ sane_unset GIT_TRACE2_CONFIG_PARAMS\n \n setup () {\n \trm -f \"$TRASH_DIRECTORY/trace.perf\" &&\n-\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n-\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n+\teval git -c core.commitGraph=false log --pretty=\"format:%s\" \"$1\" >log_wo_bloom &&\n+\teval \"GIT_TRACE2_PERF=\\\"$TRASH_DIRECTORY/trace.perf\\\"\" \\\n+\t\tgit -c core.commitGraph=true log --pretty=\"format:%s\" \"$1\" >log_w_bloom\n }\n \n test_bloom_filters_used () {\n@@ -138,10 +139,6 @@ test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n '\n \n-test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n-'\n-\n test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"-- .\"\n '\n@@ -151,9 +148,17 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n \ttest_bloom_filters_used \"-- *renamed\"\n '\n \n-test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- *\" &&\n-\ttest_bloom_filters_not_used \"-- file*\"\n+test_expect_success 'git log with multiple literal paths uses Bloom filter' '\n+\ttest_bloom_filters_used \"-- file4 A/file1\" &&\n+\ttest_bloom_filters_used \"-- *\" &&\n+\ttest_bloom_filters_used \"-- file*\"\n+'\n+\n+test_expect_success 'git log with path contains a wildcard does not use Bloom filter' '\n+\ttest_bloom_filters_not_used \"-- file\\*\" &&\n+\ttest_bloom_filters_not_used \"-- A/\\* file4\" &&\n+\ttest_bloom_filters_not_used \"-- file4 A/\\*\" &&\n+\ttest_bloom_filters_not_used \"-- * A/\\*\"\n '\n \n test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n-- \n2.47.2.vfs.0.2\n\n\n"},{"id":"521758","messageId":"7885EBB2-0D99-4456-A704-86362219AC17@smail.nju.edu.cn","threadId":"63691","inReplyTo":"3c59af48-23fe-4cc7-87e9-1de94f509a2b@gmail.com","subject":"Re: [PATCH v5.1 3.5/4] revision: make helper for pathspec to bloom key","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-07-10T15:42:18Z","receivedAt":"2025-07-10T15:43:06Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Derrick Stolee <stolee@gmail.com> wrote:\n> \n> This diff is still bigger than I was hoping, so I'm sending a couple\n> of patches that simplify this code movement. Feel free to ignore\n> them as being too nit-picky.\n> \n> --- >8 ---\n> \n> From 69fa36dc615e140ae842b536f7da792beaebb272 Mon Sep 17 00:00:00 2001\n> From: Derrick Stolee <stolee@gmail.com>\n> Date: Thu, 10 Jul 2025 08:06:29 -0400\n> Subject: [PATCH v5.1 3.5/4] revision: make helper for pathspec to bloom key\n> \n> When preparing to use bloom filters in a revision walk, Git populates a\n> boom_keyvec with an array of bloom keys for the components of a path.\n> Before we create the ability to map multiple pathspecs to multiple\n> bloom_keyvecs, extract the conversion from a pathspec to a bloom_keyvec\n> into its own helper method. This simplifies the state that persists in\n> prepare_to_use_bloom_filter() as well as makes the next change much\n> simpler.\n> \n> Signed-off-by: Derrick Stolee <stolee@gmail.com>\n> ---\n> revision.c | 50 +++++++++++++++++++++++++++++++-------------------\n> 1 file changed, 31 insertions(+), 19 deletions(-)\n> \n> diff --git a/revision.c b/revision.c\n> index 22bcfab7f93..4c09b594c55 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -687,14 +687,37 @@ static int forbid_bloom_filters(struct pathspec *spec)\n> \n> static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n> \n> -static void prepare_to_use_bloom_filter(struct rev_info *revs)\n> +static int convert_pathspec_to_filter(const struct pathspec_item *pi,\n> +      struct bloom_keyvec **bloom_keyvec,\n> +      const struct bloom_filter_settings *settings)\n> {\n> - struct pathspec_item *pi;\n> - struct bloom_keyvec *bloom_keyvec;\n> - char *path_alloc = NULL;\n> - const char *path, *p;\n> size_t len;\n> + const char *path;\n> + char *path_alloc = NULL;\n> + int res = 0;\n> +\n> + /* remove single trailing slash from path, if needed */\n> + if (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n> + path_alloc = xmemdupz(pi->match, pi->len - 1);\n> + path = path_alloc;\n> + } else\n> + path = pi->match;\n> +\n> + len = strlen(path);\n> + if (!len) {\n> + res = -1;\n> + goto cleanup;\n> + }\n> +\n> + *bloom_keyvec = bloom_keyvec_new(path, len, settings);\n> \n> +cleanup:\n> + FREE_AND_NULL(path_alloc);\n\nI think we don’t need to NULL path_alloc here, but it doesn’t hurt.\n\n> + return res;\n> +}\n> +\n> +static void prepare_to_use_bloom_filter(struct rev_info *revs)\n> +{\n> if (!revs->commits)\n> return;\n> \n> @@ -712,22 +735,12 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n> \n> revs->bloom_keyvecs_nr = 1;\n> CALLOC_ARRAY(revs->bloom_keyvecs, 1);\n> - pi = &revs->pruning.pathspec.items[0];\n> \n> - /* remove single trailing slash from path, if needed */\n> - if (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n> - path_alloc = xmemdupz(pi->match, pi->len - 1);\n> - path = path_alloc;\n> - } else\n> - path = pi->match;\n> -\n> - len = strlen(path);\n> - if (!len)\n> + if (convert_pathspec_to_filter(&revs->pruning.pathspec.items[0],\n> +       &revs->bloom_keyvecs[0],\n> +       revs->bloom_filter_settings))\n> goto fail;\n> \n> - revs->bloom_keyvecs[0] =\n> - bloom_keyvec_new(path, len, revs->bloom_filter_settings);\n> -\n> if (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n> atexit(trace2_bloom_filter_statistics_atexit);\n> bloom_filter_atexit_registered = 1;\n> @@ -737,7 +750,6 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n> \n> fail:\n> revs->bloom_filter_settings = NULL;\n> - free(path_alloc);\n> release_revisions_bloom_keyvecs(revs);\n> }\n> \n> -- \n> 2.47.2.vfs.0.2\n\nThis looks perfect to me. I would squash patch 3.5 to patch 3 in v6.\n\nThanks for your patch,\nLidong\n\n"},{"id":"521760","messageId":"EF8BBC5E-52A7-46D4-8B7B-9EFB2B726852@smail.nju.edu.cn","threadId":"63691","inReplyTo":"2619038e-05f5-4af8-bb20-e4e01138f839@gmail.com","subject":"Re: [PATCH v5.1 4/4] bloom: optimize multiple pathspec items in revision","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-07-10T15:49:40Z","receivedAt":"2025-07-10T15:50:51Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Derrick Stolee <stolee@gmail.com> wrote:\n> \n> --- >8 ---\n> \n> From fe255b1acfbe90fa8e4c335435ae18ee95e6243c Mon Sep 17 00:00:00 2001\n> From: Lidong Yan <502024330056@smail.nju.edu.cn>\n> Date: Thu, 10 Jul 2025 08:04:34 -0400\n> Subject: [PATCH v5.1 4/4] bloom: optimize multiple pathspec items in revision\n> traversal\n> \n> To enable optimize multiple pathspec items in revision traversal,\n> return 0 if all pathspec item is literal in forbid_bloom_filters().\n> Add for loops to initialize and check each pathspec item's bloom_keyvec\n> when optimization is possible.\n> \n> Add new test cases in t/t4216-log-bloom.sh to ensure\n>  - consistent results between the optimization for multiple pathspec\n>    items using bloom filter and the case without bloom filter\n>    optimization.\n>  - does not use bloom filter if any pathspec item is not literal.\n> \n> With these optimizations, we get some improvements for multi-pathspec runs\n> of 'git log'. First, in the Git repository we see these modest results:\n> \n> Benchmark 1: old\n>  Time (mean ± σ):      73.1 ms ±   2.9 ms\n>  Range (min … max):    69.9 ms …  84.5 ms    42 runs\n> \n> Benchmark 2: new\n>  Time (mean ± σ):      55.1 ms ±   2.9 ms\n>  Range (min … max):    51.1 ms …  61.2 ms    52 runs\n> \n> Summary\n>  'new' ran\n>    1.33 ± 0.09 times faster than 'old'\n> \n> But in a larger repo, such as the LLVM project repo below, we get even\n> better results:\n> \n> Benchmark 1: old\n>  Time (mean ± σ):      1.974 s ±  0.006 s\n>  Range (min … max):    1.960 s …  1.983 s    10 runs\n> \n> Benchmark 2: new\n>  Time (mean ± σ):     262.9 ms ±   2.4 ms\n>  Range (min … max):   257.7 ms … 266.2 ms    11 runs\n> \n> Summary\n>  'new' ran\n>    7.51 ± 0.07 times faster than 'old'\n\nHyperfine do looks better. I will put this into commit message and\ncover letter in v6.\n\n> \n> Signed-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n> Signed-off-by: Derrick Stolee <stolee@gmail.com>\n> ---\n> revision.c           | 22 +++++++++++-----------\n> t/t4216-log-bloom.sh | 23 ++++++++++++++---------\n> 2 files changed, 25 insertions(+), 20 deletions(-)\n> \n> diff --git a/revision.c b/revision.c\n> index 4c09b594c55..ca8c1dde8ca 100644\n> --- a/revision.c\n> +++ b/revision.c\n> @@ -675,12 +675,11 @@ static int forbid_bloom_filters(struct pathspec *spec)\n> {\n> if (spec->has_wildcard)\n> return 1;\n> - if (spec->nr > 1)\n> - return 1;\n> if (spec->magic & ~PATHSPEC_LITERAL)\n> return 1;\n> - if (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n> - return 1;\n> + for (size_t nr = 0; nr < spec->nr; nr++)\n> + if (spec->items[nr].magic & ~PATHSPEC_LITERAL)\n> + return 1;\n> \n> return 0;\n> }\n> @@ -733,13 +732,14 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n> if (!revs->pruning.pathspec.nr)\n> return;\n> \n> - revs->bloom_keyvecs_nr = 1;\n> - CALLOC_ARRAY(revs->bloom_keyvecs, 1);\n> -\n> - if (convert_pathspec_to_filter(&revs->pruning.pathspec.items[0],\n> -       &revs->bloom_keyvecs[0],\n> -       revs->bloom_filter_settings))\n> - goto fail;\n> + revs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n> + CALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n> + for (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n> + if (convert_pathspec_to_filter(&revs->pruning.pathspec.items[i],\n> +       &revs->bloom_keyvecs[i],\n> +       revs->bloom_filter_settings))\n> + goto fail;\n> + }\n> \n> if (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n> atexit(trace2_bloom_filter_statistics_atexit);\n> diff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\n> index 8910d53cac1..639868ac562 100755\n> --- a/t/t4216-log-bloom.sh\n> +++ b/t/t4216-log-bloom.sh\n> @@ -66,8 +66,9 @@ sane_unset GIT_TRACE2_CONFIG_PARAMS\n> \n> setup () {\n> rm -f \"$TRASH_DIRECTORY/trace.perf\" &&\n> - git -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n> - GIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n> + eval git -c core.commitGraph=false log --pretty=\"format:%s\" \"$1\" >log_wo_bloom &&\n> + eval \"GIT_TRACE2_PERF=\\\"$TRASH_DIRECTORY/trace.perf\\\"\" \\\n> + git -c core.commitGraph=true log --pretty=\"format:%s\" \"$1\" >log_w_bloom\n> }\n> \n> test_bloom_filters_used () {\n> @@ -138,10 +139,6 @@ test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n> test_bloom_filters_not_used \"--walk-reflogs -- A\"\n> '\n> \n> -test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n> - test_bloom_filters_not_used \"-- file4 A/file1\"\n> -'\n> -\n> test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n> test_bloom_filters_not_used \"-- .\"\n> '\n> @@ -151,9 +148,17 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n> test_bloom_filters_used \"-- *renamed\"\n> '\n> \n> -test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n> - test_bloom_filters_not_used \"-- *\" &&\n> - test_bloom_filters_not_used \"-- file*\"\n> +test_expect_success 'git log with multiple literal paths uses Bloom filter' '\n> + test_bloom_filters_used \"-- file4 A/file1\" &&\n> + test_bloom_filters_used \"-- *\" &&\n> + test_bloom_filters_used \"-- file*\"\n> +'\n> +\n> +test_expect_success 'git log with path contains a wildcard does not use Bloom filter' '\n> + test_bloom_filters_not_used \"-- file\\*\" &&\n> + test_bloom_filters_not_used \"-- A/\\* file4\" &&\n> + test_bloom_filters_not_used \"-- file4 A/\\*\" &&\n> + test_bloom_filters_not_used \"-- * A/\\*\"\n> '\n> \n> test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n> -- \n> 2.47.2.vfs.0.2\n> \n\nLooks great, I will apply this above patch 3.\n\nThanks,\nLidong\n\n"},{"id":"521763","messageId":"xmqqv7o06mw2.fsf@gitster.g","threadId":"63691","inReplyTo":"20250710084829.2171855-4-502024330056@smail.nju.edu.cn","subject":"Re: [PATCH v5 3/4] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-10T16:17:33Z","receivedAt":"2025-07-10T16:17:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lidong Yan <yldhome2d2@gmail.com> writes:\n\n>  static void prepare_to_use_bloom_filter(struct rev_info *revs)\n>  {\n>  \tstruct pathspec_item *pi;\n> +\tstruct bloom_keyvec *bloom_keyvec;\n\nThis new variable is no longer used, since the code to create a new\nkeyvec is in a helper function and its return value is directly\nstored in the array of keyvecs.\n\n>  \tchar *path_alloc = NULL;\n>  \tconst char *path, *p;\n\nAnd the \"p\" variable no longer is used, because the logic it used to\ncreate a new keyvec is moved elsewhere.\n\n"},{"id":"521815","messageId":"1B012532-E1B3-43CE-871B-B850D86419B1@smail.nju.edu.cn","threadId":"63691","inReplyTo":"xmqqv7o06mw2.fsf@gitster.g","subject":"Re: [PATCH v5 3/4] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-07-11T12:46:27Z","receivedAt":"2025-07-11T12:47:23Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> write:\n> \n> Lidong Yan <yldhome2d2@gmail.com> writes:\n> \n>> static void prepare_to_use_bloom_filter(struct rev_info *revs)\n>> {\n>> struct pathspec_item *pi;\n>> + struct bloom_keyvec *bloom_keyvec;\n> \n> This new variable is no longer used, since the code to create a new\n> keyvec is in a helper function and its return value is directly\n> stored in the array of keyvecs.\n> \n>> char *path_alloc = NULL;\n>> const char *path, *p;\n> \n> And the \"p\" variable no longer is used, because the logic it used to\n> create a new keyvec is moved elsewhere.\n\nWill fix in v6, Thanks,\nLidong\n\n"},{"id":"521820","messageId":"xmqqy0su3gys.fsf@gitster.g","threadId":"63691","inReplyTo":"1B012532-E1B3-43CE-871B-B850D86419B1@smail.nju.edu.cn","subject":"Re: [PATCH v5 3/4] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-11T15:06:03Z","receivedAt":"2025-07-11T15:06:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Lidong Yan <502024330056@smail.nju.edu.cn> writes:\n\n> Junio C Hamano <gitster@pobox.com> write:\n>> \n>> Lidong Yan <yldhome2d2@gmail.com> writes:\n>> \n>>> static void prepare_to_use_bloom_filter(struct rev_info *revs)\n>>> {\n>>> struct pathspec_item *pi;\n>>> + struct bloom_keyvec *bloom_keyvec;\n>> \n>> This new variable is no longer used, since the code to create a new\n>> keyvec is in a helper function and its return value is directly\n>> stored in the array of keyvecs.\n>> \n>>> char *path_alloc = NULL;\n>>> const char *path, *p;\n>> \n>> And the \"p\" variable no longer is used, because the logic it used to\n>> create a new keyvec is moved elsewhere.\n>\n> Will fix in v6, Thanks,\n\nFWIW, what I queued have these two already removed from v5.\n\nThanks.\n\n"},{"id":"521857","messageId":"20250712093517.17907-1-yldhome2d2@gmail.com","threadId":"63691","inReplyTo":"20250710084829.2171855-1-502024330056@smail.nju.edu.cn","subject":"[PATCH v6 0/5] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-12T09:35:12Z","receivedAt":"2025-07-12T09:35:43Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"The revision traversal limited by pathspec has optimization when\nthe pathspec has only one element, it does not use any pathspec\nmagic (other than literal), and there is no wildcard. The absence\nof optimization for multiple pathspec elements in revision traversal\ncause an issue raised by Kai Koponen at\n  https://lore.kernel.org/git/CADYQcGqaMC=4jgbmnF9Q11oC11jfrqyvH8EuiRRHytpMXd4wYA@mail.gmail.com/\n\nWhile it is much harder to lift the latter two limitations,\nsupporting a pathspec with multiple elements is relatively easy.\nJust make sure we hash each of them separately and ask the bloom\nfilter about them, and if we see none of them can possibly be\naffected by the commit, we can skip without tree comparison.\n\nThe difference from v5 is:\n  - extract convert pathspec item to bloom_keyvec logic to\n    a separate function, which simplifies the prepare_to_use_bloom_filter()\n    function.\n  - fix few bugs in v5.\n\nBelow is a comparison of the time taken to run git log on Git and\nLLVM repositories before and after applying this patch. These statistics\nare given by Derrick Stolee at\n  https://lore.kernel.org/git/afb68948-218b-4b56-9faa-29578ef9c73c@gmail.com/\n\nSetup commit-graph:\n  $ cd ~/src/git && git commit-graph write --split --reachable --changed-paths\n  $ cd ~/src/llvm && git commit-graph write --split --reachable --changed-paths\n\nRunning hyperfine [1] on Git repository:\n\n  $ hyperfine --warmup=3 \\\n  > -n 'old' '~/_git/git-sparse-checkout-clean/git log -100 -- commit.c commit-graph.c' \\\n  > -n 'new' '~/_git/git/git log -100 -- commit.c commit-graph.c'\n\nBenchmark 1: old\n  Time (mean ± σ):      73.1 ms ±   2.9 ms    [User: 48.8 ms, System: 23.9 ms]\n  Range (min … max):    69.9 ms …  84.5 ms    42 runs\n\nBenchmark 2: new\n  Time (mean ± σ):      55.1 ms ±   2.9 ms    [User: 30.5 ms, System: 24.4 ms]\n  Range (min … max):    51.1 ms …  61.2 ms    52 runs\n\nSummary\n  'new' ran\n    1.33 ± 0.09 times faster than 'old'\n\nAnd for LLVM:\n\n  $ hyperfine --warmup=3 \\\n  > -n 'old' '~/_git/git-sparse-checkout-clean/git log -100 -- llvm/lib/Support/CommandLine.cpp llvm/lib/Support/CommandLine.h' \\\n  > -n 'new' '~/_git/git/git log -100 -- llvm/lib/Support/CommandLine.cpp llvm/lib/Support/CommandLine.h'\n\nBenchmark 1: old\n  Time (mean ± σ):      1.974 s ±  0.006 s    [User: 1.877 s, System: 0.097 s]\n  Range (min … max):    1.960 s …  1.983 s    10 runs\n\nBenchmark 2: new\n  Time (mean ± σ):     262.9 ms ±   2.4 ms    [User: 214.2 ms, System: 48.4 ms]\n  Range (min … max):   257.7 ms … 266.2 ms    11 runs\n\nSummary\n  'new' ran\n    7.51 ± 0.07 times faster than 'old'\n\n[1] https://github.com/sharkdp/hyperfine\n\nLidong Yan (5):\n  bloom: add test helper to return murmur3 hash\n  bloom: rename function operates on bloom_key\n  bloom: replace struct bloom_key * with struct bloom_keyvec\n  revision: make helper for pathspec to bloom keyvec\n  To enable optimize multiple pathspec items in revision traversal,\n    return 0 if all pathspec item is literal in forbid_bloom_filters().\n    Add for loops to initialize and check each pathspec item's\n    bloom_keyvec when optimization is possible.\n\n blame.c               |   2 +-\n bloom.c               |  84 ++++++++++++++++++++++++++---\n bloom.h               |  54 ++++++++++++++-----\n line-log.c            |   5 +-\n revision.c            | 122 +++++++++++++++++++++---------------------\n revision.h            |   6 +--\n t/helper/test-bloom.c |   8 +--\n t/t4216-log-bloom.sh  |  23 ++++----\n 8 files changed, 204 insertions(+), 100 deletions(-)\n\nRange-diff against v5:\n1:  4d8f60e5ff = 1:  f5ab19063d bloom: add test helper to return murmur3 hash\n2:  acee03e397 = 2:  51a180daa6 bloom: rename function operates on bloom_key\n3:  d7690bd02c ! 3:  e17249ab4b bloom: replace struct bloom_key * with struct bloom_keyvec\n    @@ bloom.h: void bloom_key_fill(struct bloom_key *key, const char *data, size_t len\n      void bloom_key_clear(struct bloom_key *key);\n      \n     +/*\n    -+ * bloom_keyvec_fill - Allocate and populate a bloom_keyvec with keys for the\n    ++ * bloom_keyvec_new - Allocate and populate a bloom_keyvec with keys for the\n     + * given path.\n     + *\n     + * This function splits the input path by '/' and generates a bloom key for each\n    @@ revision.c: static int forbid_bloom_filters(struct pathspec *spec)\n      static void prepare_to_use_bloom_filter(struct rev_info *revs)\n      {\n      \tstruct pathspec_item *pi;\n    -+\tstruct bloom_keyvec *bloom_keyvec;\n      \tchar *path_alloc = NULL;\n    - \tconst char *path, *p;\n    +-\tconst char *path, *p;\n    ++\tconst char *path;\n      \tsize_t len;\n     -\tint path_component_nr = 1;\n      \n-:  ---------- > 4:  b3c1f5bcd1 revision: make helper for pathspec to bloom keyvec\n4:  e577aa1bfd ! 5:  785bd43674 bloom: optimize multiple pathspec items in revision traversal\n    @@ Metadata\n     Author: Lidong Yan <yldhome2d2@gmail.com>\n     \n      ## Commit message ##\n    -    bloom: optimize multiple pathspec items in revision traversal\n    -\n         To enable optimize multiple pathspec items in revision traversal,\n         return 0 if all pathspec item is literal in forbid_bloom_filters().\n         Add for loops to initialize and check each pathspec item's bloom_keyvec\n         when optimization is possible.\n     \n         Add new test cases in t/t4216-log-bloom.sh to ensure\n    -      - consistent results between the optimization for multiple pathspec\n    -        items using bloom filter and the case without bloom filter\n    -        optimization.\n    -      - does not use bloom filter if any pathspec item is not literal.\n    +     - consistent results between the optimization for multiple pathspec\n    +       items using bloom filter and the case without bloom filter\n    +       optimization.\n    +     - does not use bloom filter if any pathspec item is not literal.\n    +\n    +    With these optimizations, we get some improvements for multi-pathspec runs\n    +    of 'git log'. First, in the Git repository we see these modest results:\n    +\n    +    Benchmark 1: old\n    +     Time (mean ± σ):      73.1 ms ±   2.9 ms\n    +     Range (min … max):    69.9 ms …  84.5 ms    42 runs\n    +\n    +    Benchmark 2: new\n    +     Time (mean ± σ):      55.1 ms ±   2.9 ms\n    +     Range (min … max):    51.1 ms …  61.2 ms    52 runs\n    +\n    +    Summary\n    +     'new' ran\n    +       1.33 ± 0.09 times faster than 'old'\n    +\n    +    But in a larger repo, such as the LLVM project repo below, we get even\n    +    better results:\n    +\n    +    Benchmark 1: old\n    +     Time (mean ± σ):      1.974 s ±  0.006 s\n    +     Range (min … max):    1.960 s …  1.983 s    10 runs\n    +\n    +    Benchmark 2: new\n    +     Time (mean ± σ):     262.9 ms ±   2.4 ms\n    +     Range (min … max):   257.7 ms … 266.2 ms    11 runs\n    +\n    +    Summary\n    +     'new' ran\n    +       7.51 ± 0.07 times faster than 'old'\n     \n         Signed-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n    +    Signed-off-by: Derrick Stolee <stolee@gmail.com>\n     \n      ## revision.c ##\n     @@ revision.c: static int forbid_bloom_filters(struct pathspec *spec)\n    @@ revision.c: static void prepare_to_use_bloom_filter(struct rev_info *revs)\n      \n     -\trevs->bloom_keyvecs_nr = 1;\n     -\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n    --\tpi = &revs->pruning.pathspec.items[0];\n     +\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n     +\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n    -+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n    -+\t\tpi = &revs->pruning.pathspec.items[i];\n      \n    --\t/* remove single trailing slash from path, if needed */\n    --\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n    --\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n    --\t\tpath = path_alloc;\n    --\t} else\n    --\t\tpath = pi->match;\n    -+\t\t/* remove single trailing slash from path, if needed */\n    -+\t\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n    -+\t\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n    -+\t\t\tpath = path_alloc;\n    -+\t\t} else\n    -+\t\t\tpath = pi->match;\n    - \n    --\tlen = strlen(path);\n    --\tif (!len)\n    +-\tif (convert_pathspec_to_bloom_keyvec(&revs->bloom_keyvecs[0],\n    +-\t\t\t\t\t     &revs->pruning.pathspec.items[0],\n    +-\t\t\t\t\t     revs->bloom_filter_settings))\n     -\t\tgoto fail;\n    -+\t\tlen = strlen(path);\n    -+\t\tif (!len)\n    ++\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n    ++\t\tif (convert_pathspec_to_bloom_keyvec(&revs->bloom_keyvecs[i],\n    ++\t\t\t\t\t\t     &revs->pruning.pathspec.items[i],\n    ++\t\t\t\t\t\t     revs->bloom_filter_settings))\n     +\t\t\tgoto fail;\n    - \n    --\trevs->bloom_keyvecs[0] =\n    --\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n    -+\t\trevs->bloom_keyvecs[i] =\n    -+\t\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n    -+\t\tFREE_AND_NULL(path_alloc);\n     +\t}\n      \n      \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n-- \n2.39.5 (Apple Git-154)\n\n"},{"id":"521858","messageId":"20250712093517.17907-2-yldhome2d2@gmail.com","threadId":"63691","inReplyTo":"20250712093517.17907-1-yldhome2d2@gmail.com","subject":"[PATCH v6 1/5] bloom: add test helper to return murmur3 hash","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-12T09:35:13Z","receivedAt":"2025-07-12T09:35:53Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"In bloom.h, murmur3_seeded_v2() is exported for the use of test murmur3\nhash. To clarify that murmur3_seeded_v2() is exported solely for testing\npurposes, a new helper function test_murmur3_seeded() was added instead\nof exporting murmur3_seeded_v2() directly.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n bloom.c               | 13 ++++++++++++-\n bloom.h               | 12 +++---------\n t/helper/test-bloom.c |  4 ++--\n 3 files changed, 17 insertions(+), 12 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 0c8d2cebf9..946c5e8c98 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -107,7 +107,7 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n  * Not considered to be cryptographically secure.\n  * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n  */\n-uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len)\n+static uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len)\n {\n \tconst uint32_t c1 = 0xcc9e2d51;\n \tconst uint32_t c2 = 0x1b873593;\n@@ -540,3 +540,14 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \n \treturn 1;\n }\n+\n+uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n+\t\t\t\t   int version)\n+{\n+\tassert(version == 1 || version == 2);\n+\n+\tif (version == 2)\n+\t\treturn murmur3_seeded_v2(seed, data, len);\n+\telse\n+\t\treturn murmur3_seeded_v1(seed, data, len);\n+}\ndiff --git a/bloom.h b/bloom.h\nindex 6e46489a20..a9ded1822f 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -78,15 +78,6 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n \n-/*\n- * Calculate the murmur3 32-bit hash value for the given data\n- * using the given seed.\n- * Produces a uniformly distributed hash value.\n- * Not considered to be cryptographically secure.\n- * Implemented as described in https://en.wikipedia.org/wiki/MurmurHash#Algorithm\n- */\n-uint32_t murmur3_seeded_v2(uint32_t seed, const char *data, size_t len);\n-\n void fill_bloom_key(const char *data,\n \t\t    size_t len,\n \t\t    struct bloom_key *key,\n@@ -137,4 +128,7 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n \t\t\t  const struct bloom_filter_settings *settings);\n \n+uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n+\t\t\t\t   int version);\n+\n #endif\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 9aa2c5a592..6a24b6e0a6 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -61,13 +61,13 @@ int cmd__bloom(int argc, const char **argv)\n \t\tuint32_t hashed;\n \t\tif (argc < 3)\n \t\t\tusage(bloom_usage);\n-\t\thashed = murmur3_seeded_v2(0, argv[2], strlen(argv[2]));\n+\t\thashed = test_bloom_murmur3_seeded(0, argv[2], strlen(argv[2]), 2);\n \t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n \t}\n \n \tif (!strcmp(argv[1], \"get_murmur3_seven_highbit\")) {\n \t\tuint32_t hashed;\n-\t\thashed = murmur3_seeded_v2(0, \"\\x99\\xaa\\xbb\\xcc\\xdd\\xee\\xff\", 7);\n+\t\thashed = test_bloom_murmur3_seeded(0, \"\\x99\\xaa\\xbb\\xcc\\xdd\\xee\\xff\", 7, 2);\n \t\tprintf(\"Murmur3 Hash with seed=0:0x%08x\\n\", hashed);\n \t}\n \n-- \n2.39.5 (Apple Git-154)\n\n"},{"id":"521859","messageId":"20250712093517.17907-3-yldhome2d2@gmail.com","threadId":"63691","inReplyTo":"20250712093517.17907-1-yldhome2d2@gmail.com","subject":"[PATCH v6 2/5] bloom: rename function operates on bloom_key","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-12T09:35:14Z","receivedAt":"2025-07-12T09:35:59Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"git code style requires that functions operating on a struct S\nshould be named in the form S_verb. However, the functions operating\non struct bloom_key do not follow this convention. Therefore,\nfill_bloom_key() and clear_bloom_key() are renamed to bloom_key_fill()\nand bloom_key_clear(), respectively.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n blame.c               |  2 +-\n bloom.c               | 10 ++++------\n bloom.h               |  6 ++----\n line-log.c            |  5 +++--\n revision.c            |  8 ++++----\n t/helper/test-bloom.c |  4 ++--\n 6 files changed, 16 insertions(+), 19 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex 57daa45e89..811c6d8f9f 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -1310,7 +1310,7 @@ static void add_bloom_key(struct blame_bloom_data *bd,\n \t}\n \n \tbd->keys[bd->nr] = xmalloc(sizeof(struct bloom_key));\n-\tfill_bloom_key(path, strlen(path), bd->keys[bd->nr], bd->settings);\n+\tbloom_key_fill(bd->keys[bd->nr], path, strlen(path), bd->settings);\n \tbd->nr++;\n }\n \ndiff --git a/bloom.c b/bloom.c\nindex 946c5e8c98..5523d198c8 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -221,9 +221,7 @@ static uint32_t murmur3_seeded_v1(uint32_t seed, const char *data, size_t len)\n \treturn seed;\n }\n \n-void fill_bloom_key(const char *data,\n-\t\t    size_t len,\n-\t\t    struct bloom_key *key,\n+void bloom_key_fill(struct bloom_key *key, const char *data, size_t len,\n \t\t    const struct bloom_filter_settings *settings)\n {\n \tint i;\n@@ -243,7 +241,7 @@ void fill_bloom_key(const char *data,\n \t\tkey->hashes[i] = hash0 + i * hash1;\n }\n \n-void clear_bloom_key(struct bloom_key *key)\n+void bloom_key_clear(struct bloom_key *key)\n {\n \tFREE_AND_NULL(key->hashes);\n }\n@@ -500,9 +498,9 @@ struct bloom_filter *get_or_compute_bloom_filter(struct repository *r,\n \n \t\thashmap_for_each_entry(&pathmap, &iter, e, entry) {\n \t\t\tstruct bloom_key key;\n-\t\t\tfill_bloom_key(e->path, strlen(e->path), &key, settings);\n+\t\t\tbloom_key_fill(&key, e->path, strlen(e->path), settings);\n \t\t\tadd_key_to_filter(&key, filter, settings);\n-\t\t\tclear_bloom_key(&key);\n+\t\t\tbloom_key_clear(&key);\n \t\t}\n \n \tcleanup:\ndiff --git a/bloom.h b/bloom.h\nindex a9ded1822f..603bc1f90f 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -78,11 +78,9 @@ int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n \n-void fill_bloom_key(const char *data,\n-\t\t    size_t len,\n-\t\t    struct bloom_key *key,\n+void bloom_key_fill(struct bloom_key *key, const char *data, size_t len,\n \t\t    const struct bloom_filter_settings *settings);\n-void clear_bloom_key(struct bloom_key *key);\n+void bloom_key_clear(struct bloom_key *key);\n \n void add_key_to_filter(const struct bloom_key *key,\n \t\t       struct bloom_filter *filter,\ndiff --git a/line-log.c b/line-log.c\nindex 628e3fe3ae..07f2154e84 100644\n--- a/line-log.c\n+++ b/line-log.c\n@@ -1172,12 +1172,13 @@ static int bloom_filter_check(struct rev_info *rev,\n \t\treturn 0;\n \n \twhile (!result && range) {\n-\t\tfill_bloom_key(range->path, strlen(range->path), &key, rev->bloom_filter_settings);\n+\t\tbloom_key_fill(&key, range->path, strlen(range->path),\n+\t\t\t       rev->bloom_filter_settings);\n \n \t\tif (bloom_filter_contains(filter, &key, rev->bloom_filter_settings))\n \t\t\tresult = 1;\n \n-\t\tclear_bloom_key(&key);\n+\t\tbloom_key_clear(&key);\n \t\trange = range->next;\n \t}\n \ndiff --git a/revision.c b/revision.c\nindex afee111196..a7eadff0a5 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -739,15 +739,15 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \trevs->bloom_keys_nr = path_component_nr;\n \tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n \n-\tfill_bloom_key(path, len, &revs->bloom_keys[0],\n+\tbloom_key_fill(&revs->bloom_keys[0], path, len,\n \t\t       revs->bloom_filter_settings);\n \tpath_component_nr = 1;\n \n \tp = path + len - 1;\n \twhile (p > path) {\n \t\tif (*p == '/')\n-\t\t\tfill_bloom_key(path, p - path,\n-\t\t\t\t       &revs->bloom_keys[path_component_nr++],\n+\t\t\tbloom_key_fill(&revs->bloom_keys[path_component_nr++],\n+\t\t\t\t       path, p - path,\n \t\t\t\t       revs->bloom_filter_settings);\n \t\tp--;\n \t}\n@@ -3231,7 +3231,7 @@ void release_revisions(struct rev_info *revs)\n \toidset_clear(&revs->missing_commits);\n \n \tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n-\t\tclear_bloom_key(&revs->bloom_keys[i]);\n+\t\tbloom_key_clear(&revs->bloom_keys[i]);\n \tFREE_AND_NULL(revs->bloom_keys);\n \trevs->bloom_keys_nr = 0;\n }\ndiff --git a/t/helper/test-bloom.c b/t/helper/test-bloom.c\nindex 6a24b6e0a6..3283544bd3 100644\n--- a/t/helper/test-bloom.c\n+++ b/t/helper/test-bloom.c\n@@ -12,13 +12,13 @@ static struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;\n static void add_string_to_filter(const char *data, struct bloom_filter *filter) {\n \t\tstruct bloom_key key;\n \n-\t\tfill_bloom_key(data, strlen(data), &key, &settings);\n+\t\tbloom_key_fill(&key, data, strlen(data), &settings);\n \t\tprintf(\"Hashes:\");\n \t\tfor (size_t i = 0; i < settings.num_hashes; i++)\n \t\t\tprintf(\"0x%08x|\", key.hashes[i]);\n \t\tprintf(\"\\n\");\n \t\tadd_key_to_filter(&key, filter, &settings);\n-\t\tclear_bloom_key(&key);\n+\t\tbloom_key_clear(&key);\n }\n \n static void print_bloom_filter(struct bloom_filter *filter) {\n-- \n2.39.5 (Apple Git-154)\n\n"},{"id":"521860","messageId":"20250712093517.17907-4-yldhome2d2@gmail.com","threadId":"63691","inReplyTo":"20250712093517.17907-1-yldhome2d2@gmail.com","subject":"[PATCH v6 3/5] bloom: replace struct bloom_key * with struct bloom_keyvec","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-12T09:35:15Z","receivedAt":"2025-07-12T09:36:04Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Previously, we stored bloom keys in a flat array and marked a commit\nas NOT TREESAME if any key reported \"definitely not changed\".\n\nTo support multiple pathspec items, we now require that for each\npathspec item, there exists a bloom key reporting \"definitely not\nchanged\".\n\nThis \"for every\" condition makes a flat array insufficient, so we\nintroduce a new structure to group keys by a single pathspec item.\n`struct bloom_keyvec` is introduced to replace `struct bloom_key *`\nand `bloom_key_nr`. And because we want to support multiple pathspec\nitems, we added a bloom_keyvec * and a bloom_keyvec_nr field to\n`struct rev_info` to represent an array of bloom_keyvecs. This commit\nstill optimize only one pathspec item, thus bloom_keyvec_nr can only\nbe 0 or 1.\n\nNew bloom_keyvec_* functions are added to create and destroy a keyvec.\nbloom_filter_contains_vec() is added to check if all key in keyvec is\ncontained in a bloom filter.\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n bloom.c    | 61 +++++++++++++++++++++++++++++++++++++++++++\n bloom.h    | 38 +++++++++++++++++++++++++++\n revision.c | 76 +++++++++++++++++++++---------------------------------\n revision.h |  6 ++---\n 4 files changed, 132 insertions(+), 49 deletions(-)\n\ndiff --git a/bloom.c b/bloom.c\nindex 5523d198c8..b86015f6d1 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -278,6 +278,55 @@ void deinit_bloom_filters(void)\n \tdeep_clear_bloom_filter_slab(&bloom_filters, free_one_bloom_filter);\n }\n \n+struct bloom_keyvec *bloom_keyvec_new(const char *path, size_t len,\n+\t\t\t\t      const struct bloom_filter_settings *settings)\n+{\n+\tstruct bloom_keyvec *vec;\n+\tconst char *p;\n+\tsize_t sz;\n+\tsize_t nr = 1;\n+\n+\tp = path;\n+\twhile (*p) {\n+\t\t/*\n+\t\t * At this point, the path is normalized to use Unix-style\n+\t\t * path separators. This is required due to how the\n+\t\t * changed-path Bloom filters store the paths.\n+\t\t */\n+\t\tif (*p == '/')\n+\t\t\tnr++;\n+\t\tp++;\n+\t}\n+\n+\tsz = sizeof(struct bloom_keyvec);\n+\tsz += nr * sizeof(struct bloom_key);\n+\tvec = (struct bloom_keyvec *)xcalloc(1, sz);\n+\tif (!vec)\n+\t\treturn NULL;\n+\tvec->count = nr;\n+\n+\tbloom_key_fill(&vec->key[0], path, len, settings);\n+\tnr = 1;\n+\tp = path + len - 1;\n+\twhile (p > path) {\n+\t\tif (*p == '/') {\n+\t\t\tbloom_key_fill(&vec->key[nr++], path, p - path, settings);\n+\t\t}\n+\t\tp--;\n+\t}\n+\tassert(nr == vec->count);\n+\treturn vec;\n+}\n+\n+void bloom_keyvec_free(struct bloom_keyvec *vec)\n+{\n+\tif (!vec)\n+\t\treturn;\n+\tfor (size_t nr = 0; nr < vec->count; nr++)\n+\t\tbloom_key_clear(&vec->key[nr]);\n+\tfree(vec);\n+}\n+\n static int pathmap_cmp(const void *hashmap_cmp_fn_data UNUSED,\n \t\t       const struct hashmap_entry *eptr,\n \t\t       const struct hashmap_entry *entry_or_key,\n@@ -539,6 +588,18 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \treturn 1;\n }\n \n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *vec,\n+\t\t\t      const struct bloom_filter_settings *settings)\n+{\n+\tint ret = 1;\n+\n+\tfor (size_t nr = 0; ret > 0 && nr < vec->count; nr++)\n+\t\tret = bloom_filter_contains(filter, &vec->key[nr], settings);\n+\n+\treturn ret;\n+}\n+\n uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n \t\t\t\t   int version)\n {\ndiff --git a/bloom.h b/bloom.h\nindex 603bc1f90f..92ab2100d3 100644\n--- a/bloom.h\n+++ b/bloom.h\n@@ -74,6 +74,16 @@ struct bloom_key {\n \tuint32_t *hashes;\n };\n \n+/*\n+ * A bloom_keyvec is a vector of bloom_keys, which\n+ * can be used to store multiple keys for a single\n+ * pathspec item.\n+ */\n+struct bloom_keyvec {\n+\tsize_t count;\n+\tstruct bloom_key key[FLEX_ARRAY];\n+};\n+\n int load_bloom_filter_from_graph(struct commit_graph *g,\n \t\t\t\t struct bloom_filter *filter,\n \t\t\t\t uint32_t graph_pos);\n@@ -82,6 +92,23 @@ void bloom_key_fill(struct bloom_key *key, const char *data, size_t len,\n \t\t    const struct bloom_filter_settings *settings);\n void bloom_key_clear(struct bloom_key *key);\n \n+/*\n+ * bloom_keyvec_new - Allocate and populate a bloom_keyvec with keys for the\n+ * given path.\n+ *\n+ * This function splits the input path by '/' and generates a bloom key for each\n+ * prefix, in reverse order of specificity. For example, given the input\n+ * \"a/b/c\", it will generate bloom keys for:\n+ *   - \"a/b/c\"\n+ *   - \"a/b\"\n+ *   - \"a\"\n+ *\n+ * The resulting keys are stored in a newly allocated bloom_keyvec.\n+ */\n+struct bloom_keyvec *bloom_keyvec_new(const char *path, size_t len,\n+\t\t\t\t      const struct bloom_filter_settings *settings);\n+void bloom_keyvec_free(struct bloom_keyvec *vec);\n+\n void add_key_to_filter(const struct bloom_key *key,\n \t\t       struct bloom_filter *filter,\n \t\t       const struct bloom_filter_settings *settings);\n@@ -126,6 +153,17 @@ int bloom_filter_contains(const struct bloom_filter *filter,\n \t\t\t  const struct bloom_key *key,\n \t\t\t  const struct bloom_filter_settings *settings);\n \n+/*\n+ * bloom_filter_contains_vec - Check if all keys in a key vector are in the\n+ * Bloom filter.\n+ *\n+ * Returns 1 if **all** keys in the vector are present in the filter,\n+ * 0 if **any** key is not present.\n+ */\n+int bloom_filter_contains_vec(const struct bloom_filter *filter,\n+\t\t\t      const struct bloom_keyvec *v,\n+\t\t\t      const struct bloom_filter_settings *settings);\n+\n uint32_t test_bloom_murmur3_seeded(uint32_t seed, const char *data, size_t len,\n \t\t\t\t   int version);\n \ndiff --git a/revision.c b/revision.c\nindex a7eadff0a5..e4e0c83b0c 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -685,13 +685,14 @@ static int forbid_bloom_filters(struct pathspec *spec)\n \treturn 0;\n }\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n+\n static void prepare_to_use_bloom_filter(struct rev_info *revs)\n {\n \tstruct pathspec_item *pi;\n \tchar *path_alloc = NULL;\n-\tconst char *path, *p;\n+\tconst char *path;\n \tsize_t len;\n-\tint path_component_nr = 1;\n \n \tif (!revs->commits)\n \t\treturn;\n@@ -708,6 +709,8 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n+\trevs->bloom_keyvecs_nr = 1;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n \tpi = &revs->pruning.pathspec.items[0];\n \n \t/* remove single trailing slash from path, if needed */\n@@ -718,53 +721,30 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \t\tpath = pi->match;\n \n \tlen = strlen(path);\n-\tif (!len) {\n-\t\trevs->bloom_filter_settings = NULL;\n-\t\tfree(path_alloc);\n-\t\treturn;\n-\t}\n-\n-\tp = path;\n-\twhile (*p) {\n-\t\t/*\n-\t\t * At this point, the path is normalized to use Unix-style\n-\t\t * path separators. This is required due to how the\n-\t\t * changed-path Bloom filters store the paths.\n-\t\t */\n-\t\tif (*p == '/')\n-\t\t\tpath_component_nr++;\n-\t\tp++;\n-\t}\n-\n-\trevs->bloom_keys_nr = path_component_nr;\n-\tALLOC_ARRAY(revs->bloom_keys, revs->bloom_keys_nr);\n+\tif (!len)\n+\t\tgoto fail;\n \n-\tbloom_key_fill(&revs->bloom_keys[0], path, len,\n-\t\t       revs->bloom_filter_settings);\n-\tpath_component_nr = 1;\n-\n-\tp = path + len - 1;\n-\twhile (p > path) {\n-\t\tif (*p == '/')\n-\t\t\tbloom_key_fill(&revs->bloom_keys[path_component_nr++],\n-\t\t\t\t       path, p - path,\n-\t\t\t\t       revs->bloom_filter_settings);\n-\t\tp--;\n-\t}\n+\trevs->bloom_keyvecs[0] =\n+\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n \t\tatexit(trace2_bloom_filter_statistics_atexit);\n \t\tbloom_filter_atexit_registered = 1;\n \t}\n \n+\treturn;\n+\n+fail:\n+\trevs->bloom_filter_settings = NULL;\n \tfree(path_alloc);\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\t\t\t\t\t struct commit *commit)\n {\n \tstruct bloom_filter *filter;\n-\tint result = 1, j;\n+\tint result = 0;\n \n \tif (!revs->repo->objects->commit_graph)\n \t\treturn -1;\n@@ -779,10 +759,10 @@ static int check_maybe_different_in_bloom_filter(struct rev_info *revs,\n \t\treturn -1;\n \t}\n \n-\tfor (j = 0; result && j < revs->bloom_keys_nr; j++) {\n-\t\tresult = bloom_filter_contains(filter,\n-\t\t\t\t\t       &revs->bloom_keys[j],\n-\t\t\t\t\t       revs->bloom_filter_settings);\n+\tfor (size_t nr = 0; !result && nr < revs->bloom_keyvecs_nr; nr++) {\n+\t\tresult = bloom_filter_contains_vec(filter,\n+\t\t\t\t\t\t   revs->bloom_keyvecs[nr],\n+\t\t\t\t\t\t   revs->bloom_filter_settings);\n \t}\n \n \tif (result)\n@@ -823,7 +803,7 @@ static int rev_compare_tree(struct rev_info *revs,\n \t\t\treturn REV_TREE_SAME;\n \t}\n \n-\tif (revs->bloom_keys_nr && !nth_parent) {\n+\tif (revs->bloom_keyvecs_nr && !nth_parent) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \n \t\tif (bloom_ret == 0)\n@@ -850,7 +830,7 @@ static int rev_same_tree_as_empty(struct rev_info *revs, struct commit *commit,\n \tif (!t1)\n \t\treturn 0;\n \n-\tif (!nth_parent && revs->bloom_keys_nr) {\n+\tif (!nth_parent && revs->bloom_keyvecs_nr) {\n \t\tbloom_ret = check_maybe_different_in_bloom_filter(revs, commit);\n \t\tif (!bloom_ret)\n \t\t\treturn 1;\n@@ -3201,6 +3181,14 @@ static void release_revisions_mailmap(struct string_list *mailmap)\n \n static void release_revisions_topo_walk_info(struct topo_walk_info *info);\n \n+static void release_revisions_bloom_keyvecs(struct rev_info *revs)\n+{\n+\tfor (size_t nr = 0; nr < revs->bloom_keyvecs_nr; nr++)\n+\t\tbloom_keyvec_free(revs->bloom_keyvecs[nr]);\n+\tFREE_AND_NULL(revs->bloom_keyvecs);\n+\trevs->bloom_keyvecs_nr = 0;\n+}\n+\n static void free_void_commit_list(void *list)\n {\n \tfree_commit_list(list);\n@@ -3229,11 +3217,7 @@ void release_revisions(struct rev_info *revs)\n \tclear_decoration(&revs->treesame, free);\n \tline_log_free(revs);\n \toidset_clear(&revs->missing_commits);\n-\n-\tfor (int i = 0; i < revs->bloom_keys_nr; i++)\n-\t\tbloom_key_clear(&revs->bloom_keys[i]);\n-\tFREE_AND_NULL(revs->bloom_keys);\n-\trevs->bloom_keys_nr = 0;\n+\trelease_revisions_bloom_keyvecs(revs);\n }\n \n static void add_child(struct rev_info *revs, struct commit *parent, struct commit *child)\ndiff --git a/revision.h b/revision.h\nindex 6d369cdad6..ac843f58d0 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -62,7 +62,7 @@ struct repository;\n struct rev_info;\n struct string_list;\n struct saved_parents;\n-struct bloom_key;\n+struct bloom_keyvec;\n struct bloom_filter_settings;\n struct option;\n struct parse_opt_ctx_t;\n@@ -360,8 +360,8 @@ struct rev_info {\n \n \t/* Commit graph bloom filter fields */\n \t/* The bloom filter key(s) for the pathspec */\n-\tstruct bloom_key *bloom_keys;\n-\tint bloom_keys_nr;\n+\tstruct bloom_keyvec **bloom_keyvecs;\n+\tint bloom_keyvecs_nr;\n \n \t/*\n \t * The bloom filter settings used to generate the key.\n-- \n2.39.5 (Apple Git-154)\n\n"},{"id":"521861","messageId":"20250712093517.17907-5-yldhome2d2@gmail.com","threadId":"63691","inReplyTo":"20250712093517.17907-1-yldhome2d2@gmail.com","subject":"[PATCH v6 4/5] revision: make helper for pathspec to bloom keyvec","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-12T09:35:16Z","receivedAt":"2025-07-12T09:36:09Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"When preparing to use bloom filters in a revision walk, Git populates a\nboom_keyvec with an array of bloom keys for the components of a path.\nBefore we create the ability to map multiple pathspecs to multiple\nbloom_keyvecs, extract the conversion from a pathspec to a bloom_keyvec\ninto its own helper method. This simplifies the state that persists in\nprepare_to_use_bloom_filter() as well as makes the future change much\nsimpler.\n\nSigned-off-by: Derrick Stolee <stolee@gmail.com>\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n revision.c | 45 +++++++++++++++++++++++++++++----------------\n 1 file changed, 29 insertions(+), 16 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex e4e0c83b0c..1614c6ce0d 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -687,13 +687,37 @@ static int forbid_bloom_filters(struct pathspec *spec)\n \n static void release_revisions_bloom_keyvecs(struct rev_info *revs);\n \n-static void prepare_to_use_bloom_filter(struct rev_info *revs)\n+static int convert_pathspec_to_bloom_keyvec(struct bloom_keyvec **out,\n+\t\t\t\t\t    const struct pathspec_item *pi,\n+\t\t\t\t\t    const struct bloom_filter_settings *settings)\n {\n-\tstruct pathspec_item *pi;\n \tchar *path_alloc = NULL;\n \tconst char *path;\n \tsize_t len;\n+\tint res = 0;\n+\n+\t/* remove single trailing slash from path, if needed */\n+\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n+\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n+\t\tpath = path_alloc;\n+\t} else\n+\t\tpath = pi->match;\n+\n+\tlen = strlen(path);\n+\tif (!len) {\n+\t\tres = -1;\n+\t\tgoto cleanup;\n+\t}\n \n+\t*out = bloom_keyvec_new(path, len, settings);\n+\n+cleanup:\n+\tfree(path_alloc);\n+\treturn res;\n+}\n+\n+static void prepare_to_use_bloom_filter(struct rev_info *revs)\n+{\n \tif (!revs->commits)\n \t\treturn;\n \n@@ -711,22 +735,12 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \n \trevs->bloom_keyvecs_nr = 1;\n \tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n-\tpi = &revs->pruning.pathspec.items[0];\n \n-\t/* remove single trailing slash from path, if needed */\n-\tif (pi->len > 0 && pi->match[pi->len - 1] == '/') {\n-\t\tpath_alloc = xmemdupz(pi->match, pi->len - 1);\n-\t\tpath = path_alloc;\n-\t} else\n-\t\tpath = pi->match;\n-\n-\tlen = strlen(path);\n-\tif (!len)\n+\tif (convert_pathspec_to_bloom_keyvec(&revs->bloom_keyvecs[0],\n+\t\t\t\t\t     &revs->pruning.pathspec.items[0],\n+\t\t\t\t\t     revs->bloom_filter_settings))\n \t\tgoto fail;\n \n-\trevs->bloom_keyvecs[0] =\n-\t\tbloom_keyvec_new(path, len, revs->bloom_filter_settings);\n-\n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n \t\tatexit(trace2_bloom_filter_statistics_atexit);\n \t\tbloom_filter_atexit_registered = 1;\n@@ -736,7 +750,6 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \n fail:\n \trevs->bloom_filter_settings = NULL;\n-\tfree(path_alloc);\n \trelease_revisions_bloom_keyvecs(revs);\n }\n \n-- \n2.39.5 (Apple Git-154)\n\n"},{"id":"521862","messageId":"20250712093517.17907-6-yldhome2d2@gmail.com","threadId":"63691","inReplyTo":"20250712093517.17907-1-yldhome2d2@gmail.com","subject":"[PATCH v6 5/5] To enable optimize multiple pathspec items in revision traversal, return 0 if all pathspec item is literal in forbid_bloom_filters(). Add for loops to initialize and check each pathspec item's bloom_keyvec when optimization is possible.","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-12T09:35:17Z","receivedAt":"2025-07-12T09:36:15Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Add new test cases in t/t4216-log-bloom.sh to ensure\n - consistent results between the optimization for multiple pathspec\n   items using bloom filter and the case without bloom filter\n   optimization.\n - does not use bloom filter if any pathspec item is not literal.\n\nWith these optimizations, we get some improvements for multi-pathspec runs\nof 'git log'. First, in the Git repository we see these modest results:\n\nBenchmark 1: old\n Time (mean ± σ):      73.1 ms ±   2.9 ms\n Range (min … max):    69.9 ms …  84.5 ms    42 runs\n\nBenchmark 2: new\n Time (mean ± σ):      55.1 ms ±   2.9 ms\n Range (min … max):    51.1 ms …  61.2 ms    52 runs\n\nSummary\n 'new' ran\n   1.33 ± 0.09 times faster than 'old'\n\nBut in a larger repo, such as the LLVM project repo below, we get even\nbetter results:\n\nBenchmark 1: old\n Time (mean ± σ):      1.974 s ±  0.006 s\n Range (min … max):    1.960 s …  1.983 s    10 runs\n\nBenchmark 2: new\n Time (mean ± σ):     262.9 ms ±   2.4 ms\n Range (min … max):   257.7 ms … 266.2 ms    11 runs\n\nSummary\n 'new' ran\n   7.51 ± 0.07 times faster than 'old'\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\nSigned-off-by: Derrick Stolee <stolee@gmail.com>\n---\n revision.c           | 21 +++++++++++----------\n t/t4216-log-bloom.sh | 23 ++++++++++++++---------\n 2 files changed, 25 insertions(+), 19 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 1614c6ce0d..cf7198c0ea 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -675,12 +675,11 @@ static int forbid_bloom_filters(struct pathspec *spec)\n {\n \tif (spec->has_wildcard)\n \t\treturn 1;\n-\tif (spec->nr > 1)\n-\t\treturn 1;\n \tif (spec->magic & ~PATHSPEC_LITERAL)\n \t\treturn 1;\n-\tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n-\t\treturn 1;\n+\tfor (size_t nr = 0; nr < spec->nr; nr++)\n+\t\tif (spec->items[nr].magic & ~PATHSPEC_LITERAL)\n+\t\t\treturn 1;\n \n \treturn 0;\n }\n@@ -733,13 +732,15 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n-\trevs->bloom_keyvecs_nr = 1;\n-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n \n-\tif (convert_pathspec_to_bloom_keyvec(&revs->bloom_keyvecs[0],\n-\t\t\t\t\t     &revs->pruning.pathspec.items[0],\n-\t\t\t\t\t     revs->bloom_filter_settings))\n-\t\tgoto fail;\n+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n+\t\tif (convert_pathspec_to_bloom_keyvec(&revs->bloom_keyvecs[i],\n+\t\t\t\t\t\t     &revs->pruning.pathspec.items[i],\n+\t\t\t\t\t\t     revs->bloom_filter_settings))\n+\t\t\tgoto fail;\n+\t}\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n \t\tatexit(trace2_bloom_filter_statistics_atexit);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 8910d53cac..639868ac56 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -66,8 +66,9 @@ sane_unset GIT_TRACE2_CONFIG_PARAMS\n \n setup () {\n \trm -f \"$TRASH_DIRECTORY/trace.perf\" &&\n-\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n-\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n+\teval git -c core.commitGraph=false log --pretty=\"format:%s\" \"$1\" >log_wo_bloom &&\n+\teval \"GIT_TRACE2_PERF=\\\"$TRASH_DIRECTORY/trace.perf\\\"\" \\\n+\t\tgit -c core.commitGraph=true log --pretty=\"format:%s\" \"$1\" >log_w_bloom\n }\n \n test_bloom_filters_used () {\n@@ -138,10 +139,6 @@ test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n '\n \n-test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n-'\n-\n test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"-- .\"\n '\n@@ -151,9 +148,17 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n \ttest_bloom_filters_used \"-- *renamed\"\n '\n \n-test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- *\" &&\n-\ttest_bloom_filters_not_used \"-- file*\"\n+test_expect_success 'git log with multiple literal paths uses Bloom filter' '\n+\ttest_bloom_filters_used \"-- file4 A/file1\" &&\n+\ttest_bloom_filters_used \"-- *\" &&\n+\ttest_bloom_filters_used \"-- file*\"\n+'\n+\n+test_expect_success 'git log with path contains a wildcard does not use Bloom filter' '\n+\ttest_bloom_filters_not_used \"-- file\\*\" &&\n+\ttest_bloom_filters_not_used \"-- A/\\* file4\" &&\n+\ttest_bloom_filters_not_used \"-- file4 A/\\*\" &&\n+\ttest_bloom_filters_not_used \"-- * A/\\*\"\n '\n \n test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n-- \n2.39.5 (Apple Git-154)\n\n"},{"id":"521865","messageId":"A25E64EE-CABB-498D-8B34-27588B349FAC@gmail.com","threadId":"63691","inReplyTo":"20250712093517.17907-6-yldhome2d2@gmail.com","subject":"Re: [PATCH v6 5/5] To enable optimize multiple pathspec items in revision traversal, return 0 if all pathspec item is literal in forbid_bloom_filters(). Add for loops to initialize and check each pathspec item's bloom_keyvec when optimization is possible.","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-12T09:47:53Z","receivedAt":"2025-07-12T09:48:08Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Lidong Yan <yldhome2d2@gmail.com> writes:\n> \n> Add new test cases in t/t4216-log-bloom.sh to ensure\n> - consistent results between the optimization for multiple pathspec\n>   items using bloom filter and the case without bloom filter\n>   optimization.\n> - does not use bloom filter if any pathspec item is not literal.\n> \n> With these optimizations, we get some improvements for multi-pathspec runs\n> of 'git log'. First, in the Git repository we see these modest results:\n\nSorry, seems like I wrote bad commit message, I will resend patch 5/5 soon.\n\n"},{"id":"521866","messageId":"20250712095129.24642-1-yldhome2d2@gmail.com","threadId":"63691","inReplyTo":"A25E64EE-CABB-498D-8B34-27588B349FAC@gmail.com","subject":"[PATCH v6 5/5] bloom: optimize multiple pathspec items in revision","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-12T09:51:29Z","receivedAt":"2025-07-12T09:51:38Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"To enable optimize multiple pathspec items in revision traversal,\nreturn 0 if all pathspec item is literal in forbid_bloom_filters().\nAdd for loops to initialize and check each pathspec item's bloom_keyvec\nwhen optimization is possible.\n\nAdd new test cases in t/t4216-log-bloom.sh to ensure\n - consistent results between the optimization for multiple pathspec\n   items using bloom filter and the case without bloom filter\n   optimization.\n - does not use bloom filter if any pathspec item is not literal.\n\nWith these optimizations, we get some improvements for multi-pathspec runs\nof 'git log'. First, in the Git repository we see these modest results:\n\nBenchmark 1: old\n Time (mean ± σ):      73.1 ms ±   2.9 ms\n Range (min … max):    69.9 ms …  84.5 ms    42 runs\n\nBenchmark 2: new\n Time (mean ± σ):      55.1 ms ±   2.9 ms\n Range (min … max):    51.1 ms …  61.2 ms    52 runs\n\nSummary\n 'new' ran\n   1.33 ± 0.09 times faster than 'old'\n\nBut in a larger repo, such as the LLVM project repo below, we get even\nbetter results:\n\nBenchmark 1: old\n Time (mean ± σ):      1.974 s ±  0.006 s\n Range (min … max):    1.960 s …  1.983 s    10 runs\n\nBenchmark 2: new\n Time (mean ± σ):     262.9 ms ±   2.4 ms\n Range (min … max):   257.7 ms … 266.2 ms    11 runs\n\nSummary\n 'new' ran\n   7.51 ± 0.07 times faster than 'old'\n\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\nSigned-off-by: Derrick Stolee <stolee@gmail.com>\n---\n revision.c           | 21 +++++++++++----------\n t/t4216-log-bloom.sh | 23 ++++++++++++++---------\n 2 files changed, 25 insertions(+), 19 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 1614c6ce0d..cf7198c0ea 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -675,12 +675,11 @@ static int forbid_bloom_filters(struct pathspec *spec)\n {\n \tif (spec->has_wildcard)\n \t\treturn 1;\n-\tif (spec->nr > 1)\n-\t\treturn 1;\n \tif (spec->magic & ~PATHSPEC_LITERAL)\n \t\treturn 1;\n-\tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n-\t\treturn 1;\n+\tfor (size_t nr = 0; nr < spec->nr; nr++)\n+\t\tif (spec->items[nr].magic & ~PATHSPEC_LITERAL)\n+\t\t\treturn 1;\n \n \treturn 0;\n }\n@@ -733,13 +732,15 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n-\trevs->bloom_keyvecs_nr = 1;\n-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n \n-\tif (convert_pathspec_to_bloom_keyvec(&revs->bloom_keyvecs[0],\n-\t\t\t\t\t     &revs->pruning.pathspec.items[0],\n-\t\t\t\t\t     revs->bloom_filter_settings))\n-\t\tgoto fail;\n+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n+\t\tif (convert_pathspec_to_bloom_keyvec(&revs->bloom_keyvecs[i],\n+\t\t\t\t\t\t     &revs->pruning.pathspec.items[i],\n+\t\t\t\t\t\t     revs->bloom_filter_settings))\n+\t\t\tgoto fail;\n+\t}\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n \t\tatexit(trace2_bloom_filter_statistics_atexit);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 8910d53cac..639868ac56 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -66,8 +66,9 @@ sane_unset GIT_TRACE2_CONFIG_PARAMS\n \n setup () {\n \trm -f \"$TRASH_DIRECTORY/trace.perf\" &&\n-\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n-\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n+\teval git -c core.commitGraph=false log --pretty=\"format:%s\" \"$1\" >log_wo_bloom &&\n+\teval \"GIT_TRACE2_PERF=\\\"$TRASH_DIRECTORY/trace.perf\\\"\" \\\n+\t\tgit -c core.commitGraph=true log --pretty=\"format:%s\" \"$1\" >log_w_bloom\n }\n \n test_bloom_filters_used () {\n@@ -138,10 +139,6 @@ test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n '\n \n-test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n-'\n-\n test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"-- .\"\n '\n@@ -151,9 +148,17 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n \ttest_bloom_filters_used \"-- *renamed\"\n '\n \n-test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- *\" &&\n-\ttest_bloom_filters_not_used \"-- file*\"\n+test_expect_success 'git log with multiple literal paths uses Bloom filter' '\n+\ttest_bloom_filters_used \"-- file4 A/file1\" &&\n+\ttest_bloom_filters_used \"-- *\" &&\n+\ttest_bloom_filters_used \"-- file*\"\n+'\n+\n+test_expect_success 'git log with path contains a wildcard does not use Bloom filter' '\n+\ttest_bloom_filters_not_used \"-- file\\*\" &&\n+\ttest_bloom_filters_not_used \"-- A/\\* file4\" &&\n+\ttest_bloom_filters_not_used \"-- file4 A/\\*\" &&\n+\ttest_bloom_filters_not_used \"-- * A/\\*\"\n '\n \n test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n-- \n2.39.5 (Apple Git-154)\n\n"},{"id":"521911","messageId":"30afce8c-c932-4c51-9a27-e63385608514@gmail.com","threadId":"63691","inReplyTo":"20250712095129.24642-1-yldhome2d2@gmail.com","subject":"Re: [PATCH v6 5/5] bloom: optimize multiple pathspec items in revision","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2025-07-14T16:51:56Z","receivedAt":"2025-07-14T16:51:58Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/12/2025 5:51 AM, Lidong Yan wrote:\n> To enable optimize multiple pathspec items in revision traversal,\n> return 0 if all pathspec item is literal in forbid_bloom_filters().\n> Add for loops to initialize and check each pathspec item's bloom_keyvec\n> when optimization is possible.\n\nThe patch itself is good.\n\n> Signed-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n> Signed-off-by: Derrick Stolee <stolee@gmail.com>\n\nHere, I'll just point out that your sign-off should follow mine\nbecause you were the last to touch the patch. In this way, the\nsign-off gives a kind of timestamp to who made the most-recent\nchanges (and that those changes have that person's sign-off,\nand may not have been vetted by previous signers).\n\nThanks,\n-Stolee\n\n"},{"id":"521913","messageId":"0969e176-b9c7-464d-8e97-cf5cd4a06347@gmail.com","threadId":"63691","inReplyTo":"20250712093517.17907-1-yldhome2d2@gmail.com","subject":"Re: [PATCH v6 0/5] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2025-07-14T16:53:35Z","receivedAt":"2025-07-14T16:53:37Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/12/2025 5:35 AM, Lidong Yan wrote:\n\n> The difference from v5 is:\n>   - extract convert pathspec item to bloom_keyvec logic to\n>     a separate function, which simplifies the prepare_to_use_bloom_filter()\n>     function.\n>   - fix few bugs in v5.\n\nThanks for making these changes. Including your fixed patch 5, this\nversion looks ready to me.\n\nI wouldn't say \"fix a few bugs\" but instead \"fix some compile-time\nlinting complaints when using DEVELOPER=1\" to be clear that the\nfunctionality hasn't changed but the code is cleaner.\n\nThanks,\n-Stolee\n\n"},{"id":"521915","messageId":"xmqqbjpmu2oz.fsf@gitster.g","threadId":"63691","inReplyTo":"30afce8c-c932-4c51-9a27-e63385608514@gmail.com","subject":"Re: [PATCH v6 5/5] bloom: optimize multiple pathspec items in revision","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-14T17:01:16Z","receivedAt":"2025-07-14T17:01:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> On 7/12/2025 5:51 AM, Lidong Yan wrote:\n>> To enable optimize multiple pathspec items in revision traversal,\n>> return 0 if all pathspec item is literal in forbid_bloom_filters().\n>> Add for loops to initialize and check each pathspec item's bloom_keyvec\n>> when optimization is possible.\n>\n> The patch itself is good.\n>\n>> Signed-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n>> Signed-off-by: Derrick Stolee <stolee@gmail.com>\n>\n> Here, I'll just point out that your sign-off should follow mine\n> because you were the last to touch the patch. In this way, the\n> sign-off gives a kind of timestamp to who made the most-recent\n> changes (and that those changes have that person's sign-off,\n> and may not have been vetted by previous signers).\n\nThanks for pointing it out.  Also perhaps a single-liner attribution\nto clarify who did what, e.g.\n\n\tSigned-off-by: Derrick\n\t[ly: did this and that to derrick's code to adjust]\n\tSigned-off-by: Lidong\n\nwould be more helpful.\n\n"},{"id":"521916","messageId":"xmqq7c0au2nq.fsf@gitster.g","threadId":"63691","inReplyTo":"0969e176-b9c7-464d-8e97-cf5cd4a06347@gmail.com","subject":"Re: [PATCH v6 0/5] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-14T17:02:01Z","receivedAt":"2025-07-14T17:02:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> On 7/12/2025 5:35 AM, Lidong Yan wrote:\n>\n>> The difference from v5 is:\n>>   - extract convert pathspec item to bloom_keyvec logic to\n>>     a separate function, which simplifies the prepare_to_use_bloom_filter()\n>>     function.\n>>   - fix few bugs in v5.\n>\n> Thanks for making these changes. Including your fixed patch 5, this\n> version looks ready to me.\n>\n> I wouldn't say \"fix a few bugs\" but instead \"fix some compile-time\n> linting complaints when using DEVELOPER=1\" to be clear that the\n> functionality hasn't changed but the code is cleaner.\n\nThanks both, for polishing and reviewing.\n\n"},{"id":"521931","messageId":"B090DCBA-7306-4BA9-A5BA-DA81D1ABB29C@smail.nju.edu.cn","threadId":"63691","inReplyTo":"0969e176-b9c7-464d-8e97-cf5cd4a06347@gmail.com","subject":"Re: [PATCH v6 0/5] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-07-15T01:34:14Z","receivedAt":"2025-07-15T01:34:53Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Derrick Stolee <stolee@gmail.com> wrote:\n> \n> On 7/12/2025 5:35 AM, Lidong Yan wrote:\n> \n>> The difference from v5 is:\n>>  - extract convert pathspec item to bloom_keyvec logic to\n>>    a separate function, which simplifies the prepare_to_use_bloom_filter()\n>>    function.\n>>  - fix few bugs in v5.\n> \n> Thanks for making these changes. Including your fixed patch 5, this\n> version looks ready to me.\n> \n> I wouldn't say \"fix a few bugs\" but instead \"fix some compile-time\n> linting complaints when using DEVELOPER=1\" to be clear that the\n> functionality hasn't changed but the code is cleaner.\n\nI just learned that `make DEVELOPER=1` treats warnings as errors.\nSince this is just a cover letter issue, I feel it might not be worth rerolling\nthe patch again.\n\nThanks,\nLidong"},{"id":"521932","messageId":"55BF9B3C-F9B1-4ADB-9CBC-0D8EA45BA264@smail.nju.edu.cn","threadId":"63691","inReplyTo":"xmqqbjpmu2oz.fsf@gitster.g","subject":"Re: [PATCH v6 5/5] bloom: optimize multiple pathspec items in revision","fromName":"Lidong Yan","fromEmail":"502024330056@smail.nju.edu.cn","sentAt":"2025-07-15T01:37:30Z","receivedAt":"2025-07-15T01:38:08Z","isPatch":true,"sender":{"key":"502024330056@smail.nju.edu.cn","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n> \n> Derrick Stolee <stolee@gmail.com> writes:\n> \n>> On 7/12/2025 5:51 AM, Lidong Yan wrote:\n>>> To enable optimize multiple pathspec items in revision traversal,\n>>> return 0 if all pathspec item is literal in forbid_bloom_filters().\n>>> Add for loops to initialize and check each pathspec item's bloom_keyvec\n>>> when optimization is possible.\n>> \n>> The patch itself is good.\n>> \n>>> Signed-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n>>> Signed-off-by: Derrick Stolee <stolee@gmail.com>\n>> \n>> Here, I'll just point out that your sign-off should follow mine\n>> because you were the last to touch the patch. In this way, the\n>> sign-off gives a kind of timestamp to who made the most-recent\n>> changes (and that those changes have that person's sign-off,\n>> and may not have been vetted by previous signers).\n> \n> Thanks for pointing it out.  Also perhaps a single-liner attribution\n> to clarify who did what, e.g.\n> \n> Signed-off-by: Derrick\n> [ly: did this and that to derrick's code to adjust]\n> Signed-off-by: Lidong\n\nI will fix the order of the sign-offs, add the attributes, and resend\nPATCH v6 5/5.\n\nThanks,\nLidong\n"},{"id":"521934","messageId":"bab82a6f-e704-45a5-b422-75dec2b86d90@gmail.com","threadId":"63691","inReplyTo":"B090DCBA-7306-4BA9-A5BA-DA81D1ABB29C@smail.nju.edu.cn","subject":"Re: [PATCH v6 0/5] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2025-07-15T02:48:33Z","receivedAt":"2025-07-15T02:48:35Z","isPatch":true,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 7/14/2025 9:34 PM, Lidong Yan wrote:\n> Derrick Stolee <stolee@gmail.com> wrote:\n>>\n>> On 7/12/2025 5:35 AM, Lidong Yan wrote:\n>>\n>>> The difference from v5 is:\n>>>  - extract convert pathspec item to bloom_keyvec logic to\n>>>    a separate function, which simplifies the prepare_to_use_bloom_filter()\n>>>    function.\n>>>  - fix few bugs in v5.\n>>\n>> Thanks for making these changes. Including your fixed patch 5, this\n>> version looks ready to me.\n>>\n>> I wouldn't say \"fix a few bugs\" but instead \"fix some compile-time\n>> linting complaints when using DEVELOPER=1\" to be clear that the\n>> functionality hasn't changed but the code is cleaner.\n> \n> I just learned that `make DEVELOPER=1` treats warnings as errors.\n> Since this is just a cover letter issue, I feel it might not be worth rerolling\n> the patch again.\n\nNo need to reroll anything, I think. Junio's got the right\nfixups in place.\n\nThis was just a comment to help you next time.\n\nThanks,\n-Stolee\n\n"},{"id":"521935","messageId":"20250715025622.98646-1-yldhome2d2@gmail.com","threadId":"63691","inReplyTo":"55BF9B3C-F9B1-4ADB-9CBC-0D8EA45BA264@smail.nju.edu.cn","subject":"[RESEND][PATCH v6 5/5] bloom: optimize multiple pathspec items in revision","fromName":"Lidong Yan","fromEmail":"yldhome2d2@gmail.com","sentAt":"2025-07-15T02:56:22Z","receivedAt":"2025-07-15T02:56:50Z","isPatch":true,"sender":{"key":"yldhome2d2@gmail.com","avatar":"https://avatars.githubusercontent.com/u/77328395?v=4"},"body":"To enable optimize multiple pathspec items in revision traversal,\nreturn 0 if all pathspec item is literal in forbid_bloom_filters().\nAdd for loops to initialize and check each pathspec item's bloom_keyvec\nwhen optimization is possible.\n\nAdd new test cases in t/t4216-log-bloom.sh to ensure\n - consistent results between the optimization for multiple pathspec\n   items using bloom filter and the case without bloom filter\n   optimization.\n - does not use bloom filter if any pathspec item is not literal.\n\nWith these optimizations, we get some improvements for multi-pathspec runs\nof 'git log'. First, in the Git repository we see these modest results:\n\nBenchmark 1: old\n Time (mean ± σ):      73.1 ms ±   2.9 ms\n Range (min … max):    69.9 ms …  84.5 ms    42 runs\n\nBenchmark 2: new\n Time (mean ± σ):      55.1 ms ±   2.9 ms\n Range (min … max):    51.1 ms …  61.2 ms    52 runs\n\nSummary\n 'new' ran\n   1.33 ± 0.09 times faster than 'old'\n\nBut in a larger repo, such as the LLVM project repo below, we get even\nbetter results:\n\nBenchmark 1: old\n Time (mean ± σ):      1.974 s ±  0.006 s\n Range (min … max):    1.960 s …  1.983 s    10 runs\n\nBenchmark 2: new\n Time (mean ± σ):     262.9 ms ±   2.4 ms\n Range (min … max):   257.7 ms … 266.2 ms    11 runs\n\nSummary\n 'new' ran\n   7.51 ± 0.07 times faster than 'old'\n\nSigned-off-by: Derrick Stolee <stolee@gmail.com>\n[ly: rename convert_pathspec_to_filter() to convert_pathspec_to_bloom_keyvec()]\nSigned-off-by: Lidong Yan <502024330056@smail.nju.edu.cn>\n---\n revision.c           | 21 +++++++++++----------\n t/t4216-log-bloom.sh | 23 ++++++++++++++---------\n 2 files changed, 25 insertions(+), 19 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 1614c6ce0d..cf7198c0ea 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -675,12 +675,11 @@ static int forbid_bloom_filters(struct pathspec *spec)\n {\n \tif (spec->has_wildcard)\n \t\treturn 1;\n-\tif (spec->nr > 1)\n-\t\treturn 1;\n \tif (spec->magic & ~PATHSPEC_LITERAL)\n \t\treturn 1;\n-\tif (spec->nr && (spec->items[0].magic & ~PATHSPEC_LITERAL))\n-\t\treturn 1;\n+\tfor (size_t nr = 0; nr < spec->nr; nr++)\n+\t\tif (spec->items[nr].magic & ~PATHSPEC_LITERAL)\n+\t\t\treturn 1;\n \n \treturn 0;\n }\n@@ -733,13 +732,15 @@ static void prepare_to_use_bloom_filter(struct rev_info *revs)\n \tif (!revs->pruning.pathspec.nr)\n \t\treturn;\n \n-\trevs->bloom_keyvecs_nr = 1;\n-\tCALLOC_ARRAY(revs->bloom_keyvecs, 1);\n+\trevs->bloom_keyvecs_nr = revs->pruning.pathspec.nr;\n+\tCALLOC_ARRAY(revs->bloom_keyvecs, revs->bloom_keyvecs_nr);\n \n-\tif (convert_pathspec_to_bloom_keyvec(&revs->bloom_keyvecs[0],\n-\t\t\t\t\t     &revs->pruning.pathspec.items[0],\n-\t\t\t\t\t     revs->bloom_filter_settings))\n-\t\tgoto fail;\n+\tfor (int i = 0; i < revs->pruning.pathspec.nr; i++) {\n+\t\tif (convert_pathspec_to_bloom_keyvec(&revs->bloom_keyvecs[i],\n+\t\t\t\t\t\t     &revs->pruning.pathspec.items[i],\n+\t\t\t\t\t\t     revs->bloom_filter_settings))\n+\t\t\tgoto fail;\n+\t}\n \n \tif (trace2_is_enabled() && !bloom_filter_atexit_registered) {\n \t\tatexit(trace2_bloom_filter_statistics_atexit);\ndiff --git a/t/t4216-log-bloom.sh b/t/t4216-log-bloom.sh\nindex 8910d53cac..639868ac56 100755\n--- a/t/t4216-log-bloom.sh\n+++ b/t/t4216-log-bloom.sh\n@@ -66,8 +66,9 @@ sane_unset GIT_TRACE2_CONFIG_PARAMS\n \n setup () {\n \trm -f \"$TRASH_DIRECTORY/trace.perf\" &&\n-\tgit -c core.commitGraph=false log --pretty=\"format:%s\" $1 >log_wo_bloom &&\n-\tGIT_TRACE2_PERF=\"$TRASH_DIRECTORY/trace.perf\" git -c core.commitGraph=true log --pretty=\"format:%s\" $1 >log_w_bloom\n+\teval git -c core.commitGraph=false log --pretty=\"format:%s\" \"$1\" >log_wo_bloom &&\n+\teval \"GIT_TRACE2_PERF=\\\"$TRASH_DIRECTORY/trace.perf\\\"\" \\\n+\t\tgit -c core.commitGraph=true log --pretty=\"format:%s\" \"$1\" >log_w_bloom\n }\n \n test_bloom_filters_used () {\n@@ -138,10 +139,6 @@ test_expect_success 'git log with --walk-reflogs does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"--walk-reflogs -- A\"\n '\n \n-test_expect_success 'git log -- multiple path specs does not use Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- file4 A/file1\"\n-'\n-\n test_expect_success 'git log -- \".\" pathspec at root does not use Bloom filters' '\n \ttest_bloom_filters_not_used \"-- .\"\n '\n@@ -151,9 +148,17 @@ test_expect_success 'git log with wildcard that resolves to a single path uses B\n \ttest_bloom_filters_used \"-- *renamed\"\n '\n \n-test_expect_success 'git log with wildcard that resolves to a multiple paths does not uses Bloom filters' '\n-\ttest_bloom_filters_not_used \"-- *\" &&\n-\ttest_bloom_filters_not_used \"-- file*\"\n+test_expect_success 'git log with multiple literal paths uses Bloom filter' '\n+\ttest_bloom_filters_used \"-- file4 A/file1\" &&\n+\ttest_bloom_filters_used \"-- *\" &&\n+\ttest_bloom_filters_used \"-- file*\"\n+'\n+\n+test_expect_success 'git log with path contains a wildcard does not use Bloom filter' '\n+\ttest_bloom_filters_not_used \"-- file\\*\" &&\n+\ttest_bloom_filters_not_used \"-- A/\\* file4\" &&\n+\ttest_bloom_filters_not_used \"-- file4 A/\\*\" &&\n+\ttest_bloom_filters_not_used \"-- * A/\\*\"\n '\n \n test_expect_success 'setup - add commit-graph to the chain without Bloom filters' '\n-- \n2.39.5 (Apple Git-154)\n\n"},{"id":"522006","messageId":"xmqqv7ntij80.fsf@gitster.g","threadId":"63691","inReplyTo":"bab82a6f-e704-45a5-b422-75dec2b86d90@gmail.com","subject":"Re: [PATCH v6 0/5] bloom: enable bloom filter optimization for multiple pathspec elements in revision traversal","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-07-15T15:09:35Z","receivedAt":"2025-07-15T15:09:39Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> No need to reroll anything, I think. Junio's got the right\n> fixups in place.\n>\n> This was just a comment to help you next time.\n\nThanks.\n"}]}