git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 2/9] pack-bitmap: handle name-hash lookups in incremental bitmaps

From
Jeff King <peff@peff.net>
Date
Nov 18, 2025, 08:59 UTC
Message-ID
<20251118085949.GD4164207@coredump.intra.peff.net>
In-Reply-To
<aRVIh9R8Pnuk+yS0@nand.local>
On Wed, Nov 12, 2025 at 09:55:03PM -0500, Taylor Blau wrote:
Show 14 quoted lines
> On Wed, Nov 12, 2025 at 03:01:51AM -0500, Jeff King wrote:
> > As always with the midx and bitmap code, I am left unsure of which
> > ordering it is correct to use (pseudo-pack order, or lexical oid order,
> > or how each splits across incremental files). I _think_ this is right
> > because it's matching the ordering that is already used for a single
> > midx. But clearly this area is under-tested, since even when we did not
> > go off the end of the array we were probably passing back junk
> > name-hashes (either from the .bitmap file's trailing checksum, or
> > zero-padding at the end of the mapped page).
> 
> Yeah, this is the right order. "index_pos" is a good hint that this is
> in lexical order. bitmap_writer_finish() has some oid_pos() lookups that
> use index directly without sorting, so bitmap_writer_finish() expects
> this array in lexical order.

OK, that matches my analysis. I guess I was just a little surprised that the name hash is in lexical index order, and not pack order. But it definitely is according to the documentation and the implementation. I guess in the end it doesn't really matter that much either way, as you tend to reverse the pack/bit position into a lexical index position anyway to get the oid. So there is no situation where you don't have both anyway.

> Commit c528e17966 (pack-bitmap: write multi-pack bitmaps, 2021-08-31)
> has a comment in (what is now) midx-write.c explaining this assumption
> in bitmap_writer_finish(), but it should probably be documented
> explicitly in pack-bitmap.h.

Maybe, but I think I may just have been overly paranoid that I got it wrong.

Show 18 quoted lines
> > +static uint32_t bitmap_name_hash(struct bitmap_index *index, uint32_t pos)
> > +{
> > +	if (bitmap_is_midx(index)) {
> > +		while (index && pos < index->midx->num_objects_in_base)
> > +			index = index->base;
> 
> Looks good. It's too bad that we have to reimplement something very
> similar to midx_for_object(), but I agree with what you wrote in the
> patch message and this faithfully captures that. It might be worth doing
> something like:
> 
>     while (index && pos < index->midx->num_objects_in_base) {
>         ASSERT(bitmap_is_midx(index));
>         index = index->base;
>     }
> 
> , which should never trigger, but is a good sanity check. Definitely not
> worth re-rolling IMHO.

Yeah, I wondered the same thing while writing it. It would be a pretty horrid bug to have mixed entries in the linked list. But that is also what assertions are there for. ;) I added it for v2.

Show 9 quoted lines
> > +		if (!index)
> > +			BUG("NULL base bitmap for object position: %"PRIu32, pos);
> > +
> > +		pos -= index->midx->num_objects_in_base;
> > +		if (pos >= index->midx->num_objects)
> > +			BUG("out-of-bounds midx bitmap object at %"PRIu32, pos);
> 
> midx_for_object() spells this portion slightly differently, but what you
> have here is still good.

Yes, there it's a die(). But elsewhere, like in pack_pos_to_midx() and its reverse, the same situation is a BUG(). It's not clear to me we get a bit or index position that is out of bounds here (is it truly a bug or programming error, or might we get it from a corrupt on-disk file). So I think it's mostly academic, at least until somebody can generate a real corrupted case.

Show 9 quoted lines
> > +	if (!index->hashes)
> > +		return 0;
> > +
> > +	return get_be32(index->hashes + pos);
> 
> We *could* double check that that offset is within bounds of
> index->map_size, and I think that is ultimately worth doing at some
> point. But I think that stopping where you did makes sense, since it
> does the minimal thing to fix this bug.

I don't think we need to. When we open the bitmap, we check that its hash-cache size matches our expectation based on the number of objects covered by the bitmap (using bitmap_num_objects(), so either from the pack's count or the midx slice's count).

-Peff
Previous: Taylor BlauNext: Jeff King
Message 6 of 64 in “asan bonanza”
  1. 0/9 asan bonanzaJeff King, Nov 12, 2025
  2. 1/9 compat/mmap: mark unused argument in git_munmap()Jeff King, Nov 12, 2025
  3. 2/9 pack-bitmap: handle name-hash lookups in incremental bitmapsJeff King, Nov 12, 2025
  4. Patrick SteinhardtNov 12, 2025
  5. Taylor BlauNov 13, 2025
  6. Jeff KingNov 18, 2025
  7. 3/9 Makefile: turn on NO_MMAP when building with ASanJeff King, Nov 12, 2025
  8. Collin FunkNov 12, 2025
  9. Jeff KingNov 12, 2025
  10. Collin FunkNov 12, 2025
  11. Patrick SteinhardtNov 12, 2025
  12. Taylor BlauNov 13, 2025
  13. Patrick SteinhardtNov 13, 2025
  14. Jeff KingNov 18, 2025
  15. Junio C HamanoNov 13, 2025
  16. Patrick SteinhardtNov 14, 2025
  17. Jeff KingNov 15, 2025
  18. 4/9 cache-tree: avoid strtol() on non-string bufferJeff King, Nov 12, 2025
  19. Patrick SteinhardtNov 12, 2025
  20. Taylor BlauNov 13, 2025
  21. Jeff KingNov 18, 2025
  22. Jeff KingNov 18, 2025
  23. 5/9 fsck: assert newline presence in fsck_ident()Jeff King, Nov 12, 2025
  24. 6/9 fsck: avoid strcspn() in fsck_ident()Jeff King, Nov 12, 2025
  25. 7/9 fsck: remove redundant date timestamp checkJeff King, Nov 12, 2025
  26. 8/9 fsck: avoid parse_timestamp() on buffer that isn't NUL-terminatedJeff King, Nov 12, 2025
  27. Patrick SteinhardtNov 12, 2025
  28. Junio C HamanoNov 12, 2025
  29. Jeff KingNov 15, 2025
  30. 9/9 t: enable ASan's strict_string_checks optionJeff King, Nov 12, 2025
  31. Taylor BlauNov 13, 2025
  32. 0/9 asan bonanzaJeff King, Nov 18, 2025
  33. 1/9 compat/mmap: mark unused argument in git_munmap()Jeff King, Nov 18, 2025
  34. 2/9 pack-bitmap: handle name-hash lookups in incremental bitmapsJeff King, Nov 18, 2025
  35. 3/9 Makefile: turn on NO_MMAP when building with ASanJeff King, Nov 18, 2025
  36. 4/9 cache-tree: avoid strtol() on non-string bufferJeff King, Nov 18, 2025
  37. Phillip WoodNov 18, 2025
  38. Junio C HamanoNov 23, 2025
  39. Phillip WoodNov 23, 2025
  40. Junio C HamanoNov 23, 2025
  41. Jeff KingNov 24, 2025
  42. Junio C HamanoNov 24, 2025
  43. Jeff KingNov 26, 2025
  44. Junio C HamanoNov 26, 2025
  45. 0/4 more robust functions for parsing int from bufJeff King, Nov 30, 2025
  46. 1/4 parse: prefer bool to int for boolean returnsJeff King, Nov 30, 2025
  47. Patrick SteinhardtDec 4, 2025
  48. 2/4 parse: add functions for parsing from non-string buffersJeff King, Nov 30, 2025
  49. my complaints with clarJeff King, Nov 30, 2025
  50. Phillip WoodDec 1, 2025
  51. Patrick SteinhardtDec 4, 2025
  52. Jeff KingDec 5, 2025
  53. Patrick SteinhardtDec 4, 2025
  54. Phillip WoodDec 5, 2025
  55. Junio C HamanoJan 20, 2026
  56. Jeff KingJan 21, 2026
  57. 3/4 cache-tree: use parse_int_from_buf()Jeff King, Nov 30, 2025
  58. 4/4 fsck: use parse_unsigned_from_buf() for parsing timestampJeff King, Nov 30, 2025
  59. 5/9 fsck: assert newline presence in fsck_ident()Jeff King, Nov 18, 2025
  60. 6/9 fsck: avoid strcspn() in fsck_ident()Jeff King, Nov 18, 2025
  61. 7/9 fsck: remove redundant date timestamp checkJeff King, Nov 18, 2025
  62. 8/9 fsck: avoid parse_timestamp() on buffer that isn't NUL-terminatedJeff King, Nov 18, 2025
  63. 9/9 t: enable ASan's strict_string_checks optionJeff King, Nov 18, 2025
  64. Junio C HamanoNov 23, 2025

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.