git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 7/9] commit-graph: reuse existing bloom filters during write.

From
Jakub Narebski <jnareb@gmail.com>
Date
Jan 9, 2020, 19:12 UTC
Message-ID
<86r208v3zs.fsf@gmail.com>
In-Reply-To
<1e2acb37ad710cb0d1c09ed163fdd4473a27335c.1576879520.git.gitgitgadget@gmail.com>
"Garima Singh via GitGitGadget" <gitgitgadget@gmail.com> writes:
> From: Garima Singh <garima.singh@microsoft.com>
>
Overly long lines in the commit message.
> Read previously computed bloom filters from the commit-graph file if possible
> to avoid recomputing during commit-graph write.

Hmmm. This fixes (somewhat) the problem that I have noticed in previous patch that the Bloom filter was computed at least twice, once for BIDX chunk, once fo BDAT chunk.

I think the order should be:
 - use Bloom filter on slab, if present
 - fill it from commit graph, if saved there
 - if needed, compute it from scratch (expensive operation!)

If I understand it correctly, it now does it... though possibly with unnecessary memory allocation if commit-graph file does not include Bloomm filters data, and (re)computing is not requested (see later).

But I might be wrong here.
>
> Reading from the commit-graph is based on the format in which bloom filters are
> written in the commit graph file. See method `fill_filter_from_graph` in bloom.c

This description reads a bit strange; it looks like it states a truism (we read in the format we wrote). It think it should be rephrased in different way for better readability.

>
> For reading the bloom filter for commit at lexicographic position i:
I think it would better read as:
  To read Bloom filter for a given commit with lexicographic position
  'i' we need to:
> 1. Read BIDX[i] which essentially gives us the starting index in BDAT for filter
>    of commit i+1 (called the next_index in the code)
                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                    |
                    \- I think would be not needed with better var name

I would also add that it gives the position [one past] the end of Bloom filter data for i-th commit.

>
> 2. For i>0, read BIDX[i-1] which will give us the starting index in BDAT for
>    filter of commit i (called the prev_index in the code)
Minor nitpick: Full stops are missing.

Why it is called prev_index and next_index, while it is either curr_index and next_index or prev_index and curr_index, or maybe even better beg_index and end_index?

>    For i = 0, prev_index will be 0. The first lexicographic commit's filter will
>    start at BDAT.
I would state it
     For first commit, with i = 0, Bloom filter data starts at the
     beginning, just past the header in BDAT chunk.
Show 6 quoted lines
>
> 3. The length of the filter will be next_index - prev_index, because BIDX[i]
>    gives the cumulative 8-byte words including the ith commit's filter.
>
> We toggle whether bloom filters should be recomputed based on the compute_if_null
> flag.
All right.
Show 31 quoted lines
>
> Helped-by: Derrick Stolee <dstolee@microsoft.com>
> Signed-off-by: Garima Singh <garima.singh@microsoft.com>
> ---
>  bloom.c        | 40 ++++++++++++++++++++++++++++++++++++++--
>  bloom.h        |  3 ++-
>  commit-graph.c |  6 +++---
>  3 files changed, 43 insertions(+), 6 deletions(-)
>
> diff --git a/bloom.c b/bloom.c
> index 08328cc381..86b1005802 100644
> --- a/bloom.c
> +++ b/bloom.c
> @@ -1,5 +1,7 @@
>  #include "git-compat-util.h"
>  #include "bloom.h"
> +#include "commit.h"
> +#include "commit-slab.h"
>  #include "commit-graph.h"
>  #include "object-store.h"
>  #include "diff.h"
> @@ -119,13 +121,35 @@ static void add_key_to_filter(struct bloom_key *key,
>  	}
>  }
>  
> +static void fill_filter_from_graph(struct commit_graph *g,
> +				   struct bloom_filter *filter,
> +				   struct commit *c)
> +{
> +	uint32_t lex_pos, prev_index, next_index;
> +
> +	while (c->graph_pos < g->num_commits_in_base)
> +		g = g->base_graph;
> +
> +	lex_pos = c->graph_pos - g->num_commits_in_base;

This part shares common code with load_oid_from_graph(), only without some of error checking; perhaps it might be good to extract it into a separate helper function, e.g. `lex_index(&g, c->graph_pos)`.

Minor nitpick about the consistency of function names: why load_oid_from_graph(), but fill_filter_from_graph(), and not load_filter_from_graph() / load_bloom_from_graph()?

> +
> +	next_index = get_be32(g->chunk_bloom_indexes + 4 * lex_pos);
> +	if (lex_pos)
I think using
  +	if (lex_pos > 0)
or
  +	if (lex_pos >= 0)
might be easier to reason about.
Show 5 quoted lines
> +		prev_index = get_be32(g->chunk_bloom_indexes + 4 * (lex_pos - 1));
> +	else
> +		prev_index = 0;
> +
> +	filter->len = next_index - prev_index;

The above command reads a bit strange: next - prev? not next - curr, or curr - prev? Wouldn't it be better to name it begin_index and end_index, or beg_index and end_index for brevity?

> +	filter->data = (uint64_t *)(g->chunk_bloom_data + 8 * prev_index + 12);
Please do not use magic constants; use instead something like:
  +	filter->data = (uint64_t *)(g->chunk_bloom_data +
  +				    sizeof(uint64_t) * prev_index +
  +				    BLOOMDATA_CHUNK_HEADER_SIZE);

Perhaps using `3*sizeof(unit32_t)` instead of magic value 12 would be enough; but having symbolic name for BDAT chunk header size is better, I think.

Show 11 quoted lines
> +}
> +
>  void load_bloom_filters(void)
>  {
>  	init_bloom_filter_slab(&bloom_filters);
>  }
>  
>  struct bloom_filter *get_bloom_filter(struct repository *r,
> -				      struct commit *c)
> +				      struct commit *c,
> +				      int compute_if_null)
I'm not sure about `compute_if_null` name...
Show 7 quoted lines
>  {
>  	struct bloom_filter *filter;
>  	struct bloom_filter_settings settings = DEFAULT_BLOOM_FILTER_SETTINGS;
> @@ -134,6 +158,18 @@ struct bloom_filter *get_bloom_filter(struct repository *r,
>  	const char *revs_argv[] = {NULL, "HEAD", NULL};
>  
>  	filter = bloom_filter_slab_at(&bloom_filters, c);

Note that the documentation for `slab_at(slab, commit)` is documented in `commit-slab.h` as

 *   This function locates the data associated with the given commit in
 *   the indegree slab, and returns the pointer to it.  The location to
 *   store the data is allocated as necessary.
                       ~~~~~~~~~~~~~~~~~~~~~~

Should we worry about this possibly unnecessary allocation (if there is no Bloom filter chunk in the commit-graph, and we are not recomputing it)?

There is `slab_peek(slab_commit)` with the following properties:
 *   This function is similar to indegree_at(), but it will return NULL
 *   until a call to indegree_at() was made for the commit.
> +
> +	if (!filter->data) {

What does the Bloom filter for a commit with no changes looks like? What about for a commit with more than 512 changes? Do, in either of those cases, filter->len is 0? If yes, what about filter->data?

> +		load_commit_graph_info(r, c);
> +		if (c->graph_pos != COMMIT_NOT_FROM_GRAPH && r->objects->commit_graph->chunk_bloom_indexes) {

Please wrap overly long lines (109 characters seems too long; the CodingGuidelines states:

 - We try to keep to at most 80 characters per line.
  +		if (c->graph_pos != COMMIT_NOT_FROM_GRAPH &&
  +		    r->objects->commit_graph->chunk_bloom_indexes) {

You seem to assume here that in the chain of commit-graph files either all of them would have Bloom filters, or all of them would be missing Bloom filters. Isn't it however possible for only some of the commit-graph files in chain to include Bloom filter data chunks?

In such case, the top commit-graph file may have BIDX chunk (bloom_indexes), but the commit-graph with the commit 'c' might not have it. Or the top commit-graph file may be missing BIDX chunk, so Git would recompute it even if the commit-graph file for commit 'c' includes it.

If such situation is forbidden, how the restriction is managed?
Note: in any case, this needs to be tested!
> +			fill_filter_from_graph(r->objects->commit_graph, filter, c);
> +			return filter;
> +		}
> +	}

All right, if it is not in slab, we try to read if from the commit graph. Looks all right.

> +
> +	if (filter->data || !compute_if_null)
> +			return filter;
                ^^^^^^^^
                 |
                 \- one tab too many

If we have found existing filter (on slab or in the commit-graph), or if we won't be recomputing it, return it. O.K.

Show 11 quoted lines
> +
>  	init_revisions(&revs, NULL);
>  	revs.diffopt.flags.recursive = 1;
>  
> @@ -198,4 +234,4 @@ struct bloom_filter *get_bloom_filter(struct repository *r,
>  	DIFF_QUEUE_CLEAR(&diff_queued_diff);
>  
>  	return filter;
> -}
> \ No newline at end of file
> +}
Accidental change.
Show 12 quoted lines
> diff --git a/bloom.h b/bloom.h
> index ba8ae70b67..101d689bbd 100644
> --- a/bloom.h
> +++ b/bloom.h
> @@ -36,7 +36,8 @@ struct bloom_key {
>  void load_bloom_filters(void);
>  
>  struct bloom_filter *get_bloom_filter(struct repository *r,
> -				      struct commit *c);
> +				      struct commit *c,
> +				      int compute_if_null);
>
All right, this is just update of the function signature.
Show 24 quoted lines
>  void fill_bloom_key(const char *data,
>  		    int len,
> diff --git a/commit-graph.c b/commit-graph.c
> index def2ade166..0580ce75d5 100644
> --- a/commit-graph.c
> +++ b/commit-graph.c
> @@ -1032,7 +1032,7 @@ static void write_graph_chunk_bloom_indexes(struct hashfile *f,
>  	uint32_t cur_pos = 0;
>  
>  	while (list < last) {
> -		struct bloom_filter *filter = get_bloom_filter(ctx->r, *list);
> +		struct bloom_filter *filter = get_bloom_filter(ctx->r, *list, 0);
>  		cur_pos += filter->len;
>  		hashwrite_be32(f, cur_pos);
>  		list++;
> @@ -1051,7 +1051,7 @@ static void write_graph_chunk_bloom_data(struct hashfile *f,
>  	hashwrite_be32(f, settings->bits_per_entry);
>  
>  	while (first < last) {
> -		struct bloom_filter *filter = get_bloom_filter(ctx->r, *first);
> +		struct bloom_filter *filter = get_bloom_filter(ctx->r, *first, 0);
>  		hashwrite(f, filter->data, filter->len * sizeof(uint64_t));
>  		first++;
>  	}

O.K., so those two do not compute Bloom filters, but they are called from write_commit_graph_file(), which in turn is called in write_commit_graph() *after* running compute_bloom_filters().

Looks good to me, then.
Show 9 quoted lines
> @@ -1218,7 +1218,7 @@ static void compute_bloom_filters(struct write_commit_graph_context *ctx)
>  
>  	for (i = 0; i < ctx->commits.nr; i++) {
>  		struct commit *c = ctx->commits.list[i];
> -		struct bloom_filter *filter = get_bloom_filter(ctx->r, c);
> +		struct bloom_filter *filter = get_bloom_filter(ctx->r, c, 1);
>  		ctx->total_bloom_filter_size += sizeof(uint64_t) * filter->len;
>  		display_progress(progress, i + 1);
>  	}

All right, so compute_bloom_filters() ensures that it is actually computed (if needed).

Looks good to me.
Regards,
-- 
Jakub Narębski
Previous: Garima Singh via GitGitGadgetNext: Garima Singh via GitGitGadget
Message 12 of 159 in “[RFC] Changed Paths Bloom Filters”
  1. 0/9 [RFC] Changed Paths Bloom FiltersGarima Singh via GitGitGadget, Dec 20, 2019
  2. 1/9 commit-graph: add --changed-paths option to writeGarima Singh via GitGitGadget, Dec 20, 2019
  3. Jakub NarebskiJan 1, 2020
  4. 3/9 commit-graph: use MAX_NUM_CHUNKSGarima Singh via GitGitGadget, Dec 20, 2019
  5. Jakub NarebskiJan 7, 2020
  6. 4/9 commit-graph: document bloom filter formatGarima Singh via GitGitGadget, Dec 20, 2019
  7. Jakub NarebskiJan 7, 2020
  8. 8/9 revision.c: use bloom filters to speed up path based revision walksGarima Singh via GitGitGadget, Dec 20, 2019
  9. Jakub NarebskiJan 11, 2020
  10. Garima SinghJan 15, 2020
  11. 7/9 commit-graph: reuse existing bloom filters during write.Garima Singh via GitGitGadget, Dec 20, 2019
  12. Jakub NarebskiJan 9, 2020
  13. 9/9 commit-graph: add GIT_TEST_COMMIT_GRAPH_BLOOM_FILTERS test flagGarima Singh via GitGitGadget, Dec 20, 2019
  14. Jakub NarebskiJan 11, 2020
  15. Garima SinghJan 15, 2020
  16. 5/9 commit-graph: write changed path bloom filters to commit-graph file.Garima Singh via GitGitGadget, Dec 20, 2019
  17. Jakub NarebskiJan 7, 2020
  18. Garima SinghJan 14, 2020
  19. 6/9 commit-graph: test commit-graph write --changed-pathsGarima Singh via GitGitGadget, Dec 20, 2019
  20. Jakub NarebskiJan 8, 2020
  21. 2/9 commit-graph: write changed paths bloom filtersGarima Singh via GitGitGadget, Dec 20, 2019
  22. Philip OakleyDec 21, 2019
  23. Jakub NarebskiJan 6, 2020
  24. Garima SinghJan 13, 2020
  25. Junio C HamanoDec 20, 2019
  26. Christian CouderDec 22, 2019
  27. Jeff KingDec 22, 2019
  28. Jakub NarebskiJan 1, 2020
  29. Jeff KingDec 22, 2019
  30. 1/3 commit-graph: examine changed-path objects in pack orderJeff King, Dec 22, 2019
  31. Derrick StoleeDec 27, 2019
  32. Jeff KingDec 29, 2019
  33. Jeff KingDec 29, 2019
  34. Derrick StoleeDec 30, 2019
  35. Derrick StoleeDec 30, 2019
  36. 2/3 commit-graph: free large diffs, tooJeff King, Dec 22, 2019
  37. Derrick StoleeDec 27, 2019
  38. 3/3 commit-graph: stop using full rev_info for diffsJeff King, Dec 22, 2019
  39. Derrick StoleeDec 27, 2019
  40. Derrick StoleeDec 26, 2019
  41. Jeff KingDec 29, 2019
  42. Derrick StoleeDec 27, 2019
  43. Jeff KingDec 29, 2019
  44. Derrick StoleeDec 30, 2019
  45. Junio C HamanoDec 30, 2019
  46. Jakub NarebskiDec 31, 2019
  47. Garima SinghJan 13, 2020
  48. Jakub NarebskiJan 20, 2020
  49. Garima SinghJan 21, 2020
  50. Jakub NarebskiFeb 2, 2020
  51. Emily ShafferJan 21, 2020
  52. Garima SinghJan 27, 2020
  53. Jakub NarebskiFeb 1, 2020
  54. 00/11 Changed Paths Bloom FiltersGarima Singh via GitGitGadget, Feb 5, 2020
  55. 01/11 commit-graph: use MAX_NUM_CHUNKSGarima Singh via GitGitGadget, Feb 5, 2020
  56. Jakub NarebskiFeb 9, 2020
  57. 04/11 commit-graph: compute Bloom filters for changed pathsGarima Singh via GitGitGadget, Feb 5, 2020
  58. Jakub NarebskiFeb 17, 2020
  59. Garima SinghFeb 22, 2020
  60. Jakub NarebskiFeb 23, 2020
  61. 05/11 commit-graph: examine changed-path objects in pack orderJeff King via GitGitGadget, Feb 5, 2020
  62. Jakub NarebskiFeb 18, 2020
  63. Garima SinghFeb 24, 2020
  64. 06/11 commit-graph: examine commits by generation numberDerrick Stolee via GitGitGadget, Feb 5, 2020
  65. Jakub NarebskiFeb 19, 2020
  66. Garima SinghFeb 24, 2020
  67. 02/11 bloom: core Bloom filter implementation for changed pathsGarima Singh via GitGitGadget, Feb 5, 2020
  68. Jakub NarebskiFeb 15, 2020
  69. Jakub NarebskiFeb 16, 2020
  70. Garima SinghFeb 22, 2020
  71. Jakub NarebskiFeb 23, 2020
  72. Garima SinghFeb 24, 2020
  73. Jakub NarebskiFeb 24, 2020
  74. 03/11 diff: halt tree-diff early after max_changesDerrick Stolee via GitGitGadget, Feb 5, 2020
  75. Jakub NarebskiFeb 17, 2020
  76. Garima SinghFeb 22, 2020
  77. 07/11 commit-graph: write Bloom filters to commit graph fileGarima Singh via GitGitGadget, Feb 5, 2020
  78. Jakub NarebskiFeb 19, 2020
  79. Garima SinghFeb 24, 2020
  80. Jakub NarebskiFeb 25, 2020
  81. Garima SinghFeb 25, 2020
  82. 08/11 commit-graph: reuse existing Bloom filters during write.Garima Singh via GitGitGadget, Feb 5, 2020
  83. Jakub NarebskiFeb 20, 2020
  84. Garima SinghFeb 24, 2020
  85. 11/11 commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flagGarima Singh via GitGitGadget, Feb 5, 2020
  86. Jakub NarebskiFeb 22, 2020
  87. 10/11 revision.c: use Bloom filters to speed up path based revision walksGarima Singh via GitGitGadget, Feb 5, 2020
  88. Jakub NarebskiFeb 21, 2020
  89. Jakub NarebskiFeb 21, 2020
  90. 09/11 commit-graph: add --changed-paths option to write subcommandGarima Singh via GitGitGadget, Feb 5, 2020
  91. Jakub NarebskiFeb 20, 2020
  92. Garima SinghFeb 24, 2020
  93. Jakub NarebskiFeb 25, 2020
  94. Bryan TurnerFeb 20, 2020
  95. Garima SinghFeb 22, 2020
  96. SZEDER GáborFeb 7, 2020
  97. Garima SinghFeb 7, 2020
  98. Derrick StoleeFeb 7, 2020
  99. SZEDER GáborFeb 7, 2020
  100. Derrick StoleeFeb 7, 2020
  101. Garima SinghFeb 11, 2020
  102. Jakub NarebskiFeb 8, 2020
  103. Garima SinghFeb 21, 2020
  104. Junio C HamanoMar 29, 2020
  105. 00/16 Changed Paths Bloom FiltersGarima Singh via GitGitGadget, Mar 30, 2020
  106. 01/16 commit-graph: define and use MAX_NUM_CHUNKSGarima Singh via GitGitGadget, Mar 30, 2020
  107. 02/16 bloom.c: add the murmur3 hash implementationGarima Singh via GitGitGadget, Mar 30, 2020
  108. 03/16 bloom.c: introduce core Bloom filter constructsGarima Singh via GitGitGadget, Mar 30, 2020
  109. 06/16 commit-graph: compute Bloom filters for changed pathsGarima Singh via GitGitGadget, Mar 30, 2020
  110. 05/16 diff: halt tree-diff early after max_changesDerrick Stolee via GitGitGadget, Mar 30, 2020
  111. 04/16 bloom.c: core Bloom filter implementation for changed paths.Garima Singh via GitGitGadget, Mar 30, 2020
  112. 07/16 commit-graph: examine changed-path objects in pack orderJeff King via GitGitGadget, Mar 30, 2020
  113. 10/16 commit-graph: write Bloom filters to commit graph fileGarima Singh via GitGitGadget, Mar 30, 2020
  114. 13/16 revision.c: use Bloom filters to speed up path based revision walksGarima Singh via GitGitGadget, Mar 30, 2020
  115. 14/16 revision.c: add trace2 stats around Bloom filter usageGarima Singh via GitGitGadget, Mar 30, 2020
  116. 15/16 t4216: add end to end tests for git log with Bloom filtersGarima Singh via GitGitGadget, Mar 30, 2020
  117. 16/16 commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flagGarima Singh via GitGitGadget, Mar 30, 2020
  118. 12/16 commit-graph: add --changed-paths option to write subcommandGarima Singh via GitGitGadget, Mar 30, 2020
  119. 09/16 diff: skip batch object download when possibleGarima Singh via GitGitGadget, Mar 30, 2020
  120. 08/16 commit-graph: examine commits by generation numberGarima Singh via GitGitGadget, Mar 30, 2020
  121. 11/16 commit-graph: reuse existing Bloom filters during writeGarima Singh via GitGitGadget, Mar 30, 2020
  122. 00/15 Changed Paths Bloom FiltersGarima Singh via GitGitGadget, Apr 6, 2020
  123. 01/15 commit-graph: define and use MAX_NUM_CHUNKSGarima Singh via GitGitGadget, Apr 6, 2020
  124. 03/15 bloom.c: introduce core Bloom filter constructsGarima Singh via GitGitGadget, Apr 6, 2020
  125. 02/15 bloom.c: add the murmur3 hash implementationGarima Singh via GitGitGadget, Apr 6, 2020
  126. 05/15 diff: halt tree-diff early after max_changesDerrick Stolee via GitGitGadget, Apr 6, 2020
  127. SZEDER GáborAug 4, 2020
  128. Derrick StoleeAug 4, 2020
  129. SZEDER GáborAug 4, 2020
  130. Derrick StoleeAug 4, 2020
  131. Derrick StoleeAug 5, 2020
  132. 04/15 bloom.c: core Bloom filter implementation for changed paths.Garima Singh via GitGitGadget, Apr 6, 2020
  133. SZEDER GáborJun 27, 2020
  134. 06/15 commit-graph: compute Bloom filters for changed pathsGarima Singh via GitGitGadget, Apr 6, 2020
  135. 07/15 commit-graph: examine changed-path objects in pack orderJeff King via GitGitGadget, Apr 6, 2020
  136. 10/15 commit-graph: reuse existing Bloom filters during writeGarima Singh via GitGitGadget, Apr 6, 2020
  137. SZEDER GáborJun 19, 2020
  138. Junio C HamanoJun 19, 2020
  139. SZEDER GáborJul 27, 2020
  140. 09/15 commit-graph: write Bloom filters to commit graph fileGarima Singh via GitGitGadget, Apr 6, 2020
  141. SZEDER GáborMay 29, 2020
  142. Derrick StoleeMay 29, 2020
  143. SZEDER GáborMay 31, 2020
  144. commit-graph: fix "Writing out commit graph" progress counterSZEDER Gábor, Jul 9, 2020
  145. Derrick StoleeJul 9, 2020
  146. Derrick StoleeJul 9, 2020
  147. 08/15 commit-graph: examine commits by generation numberGarima Singh via GitGitGadget, Apr 6, 2020
  148. 13/15 revision.c: add trace2 stats around Bloom filter usageGarima Singh via GitGitGadget, Apr 6, 2020
  149. 14/15 t4216: add end to end tests for git log with Bloom filtersGarima Singh via GitGitGadget, Apr 6, 2020
  150. 12/15 revision.c: use Bloom filters to speed up path based revision walksGarima Singh via GitGitGadget, Apr 6, 2020
  151. SZEDER GáborJun 26, 2020
  152. 11/15 commit-graph: add --changed-paths option to write subcommandGarima Singh via GitGitGadget, Apr 6, 2020
  153. SZEDER GáborJun 7, 2020
  154. 15/15 commit-graph: add GIT_TEST_COMMIT_GRAPH_CHANGED_PATHS test flagGarima Singh via GitGitGadget, Apr 6, 2020
  155. Derrick StoleeApr 8, 2020
  156. Junio C HamanoApr 8, 2020
  157. Jakub NarębskiApr 8, 2020
  158. Taylor BlauApr 12, 2020
  159. Garima SinghMar 5, 2020

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.