Re: [PATCH 2/9] pack-bitmap: handle name-hash lookups in incremental bitmaps
- From
Taylor Blau <me@ttaylorr.com>
- Date
- Nov 13, 2025, 02:55 UTC
- Message-ID
- <aRVIh9R8Pnuk+yS0@nand.local>
- In-Reply-To
- <20251112080151.GB979063@coredump.intra.peff.net>
On Wed, Nov 12, 2025 at 03:01:51AM -0500, Jeff King wrote:
Show 8 quoted lines
> As always with the midx and bitmap code, I am left unsure of which > ordering it is correct to use (pseudo-pack order, or lexical oid order, > or how each splits across incremental files). I _think_ this is right > because it's matching the ordering that is already used for a single > midx. But clearly this area is under-tested, since even when we did not > go off the end of the array we were probably passing back junk > name-hashes (either from the .bitmap file's trailing checksum, or > zero-padding at the end of the mapped page).
Yeah, this is the right order. "index_pos" is a good hint that this is in lexical order. bitmap_writer_finish() has some oid_pos() lookups that use index directly without sorting, so bitmap_writer_finish() expects this array in lexical order.
Commit c528e17966 (pack-bitmap: write multi-pack bitmaps, 2021-08-31) has a comment in (what is now) midx-write.c explaining this assumption in bitmap_writer_finish(), but it should probably be documented explicitly in pack-bitmap.h.
> So it might be worth adding more tests here, but I know this incremental > bitmap code is a big work in progress. So I contented myself with the > reproduction above, and anything else can go onto the incremental todo > pile. :)
Yeah, I agree. The only hash-cache test that I could think of is from t5326, which tests that we can propagate existing name-hash values from a pack bitmap in to a MIDX one. We probably need an equivalent for when writing an incremental MIDX/bitmap too. #leftoverbits
Show 16 quoted lines
> pack-bitmap.c | 27 +++++++++++++++++++++++----
> 1 file changed, 23 insertions(+), 4 deletions(-)
>
> diff --git a/pack-bitmap.c b/pack-bitmap.c
> index 291e1a9cf4..710b86a451 100644
> --- a/pack-bitmap.c
> +++ b/pack-bitmap.c
> @@ -213,6 +213,26 @@ static uint32_t bitmap_num_objects(struct bitmap_index *index)
> return index->pack->num_objects;
> }
>
> +static uint32_t bitmap_name_hash(struct bitmap_index *index, uint32_t pos)
> +{
> + if (bitmap_is_midx(index)) {
> + while (index && pos < index->midx->num_objects_in_base)
> + index = index->base;Looks good. It's too bad that we have to reimplement something very similar to midx_for_object(), but I agree with what you wrote in the patch message and this faithfully captures that. It might be worth doing something like:
while (index && pos < index->midx->num_objects_in_base) {
ASSERT(bitmap_is_midx(index));
index = index->base;
}, which should never trigger, but is a good sanity check. Definitely not worth re-rolling IMHO.
Show 7 quoted lines
> +
> + if (!index)
> + BUG("NULL base bitmap for object position: %"PRIu32, pos);
> +
> + pos -= index->midx->num_objects_in_base;
> + if (pos >= index->midx->num_objects)
> + BUG("out-of-bounds midx bitmap object at %"PRIu32, pos);midx_for_object() spells this portion slightly differently, but what you have here is still good.
Show 6 quoted lines
> + } > + > + if (!index->hashes) > + return 0; > + > + return get_be32(index->hashes + pos);
We *could* double check that that offset is within bounds of index->map_size, and I think that is ultimately worth doing at some point. But I think that stopping where you did makes sense, since it does the minimal thing to fix this bug.
Thanks, Taylor