git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 09/16] documentation: add documentation for the bitmap format

From
Shawn Pearce <spearce@spearce.org>
Date
Jun 27, 2013, 16:07 UTC
Message-ID
<CAJo=hJvOq=CATrDeYAwi+jgkPpqjywWhuKeC1TVYeCXr6NVM6w@mail.gmail.com>
In-Reply-To
<20130627024521.GA6936@sigill.intra.peff.net>
On Wed, Jun 26, 2013 at 7:45 PM, Jeff King <peff@peff.net> wrote:
Show 6 quoted lines
>
> In particular, it seems like the slowness we saw with the v1 bitmap
> format is not what Shawn and Colby have experienced. So it's possible
> that our test setup is bad or different. Or maybe the C v1 reading
> implementation had some problems that are fixable. It's hard to say
> because we haven't shown any code that can be timed and compared.

Right, the format and implementation in JGit can do "Counting objects" in 87ms for the Linux kernel history. But I think we are comparing apples to steaks here, Vincent is (rightfully) concerned about process startup performance, whereas our timings were assuming the process was already running.

It would help everyone to understand the issues involved if we are at least looking at the same file format.

> And the pack-order versus idx-order for the bitmaps is still up in the
> air. Do we have numbers on the on-disk sizes of the resulting EWAHs?

I did not see any presented in this thread, and I am very interested in this aspect of the series. The path hash cache should be taking about 9.9M of disk space, but I recall reading the bitmap file is 8M. I don't understand.

Colby and I were very concerned about the size of the EWAH compressed bitmaps because we wanted hundreds of them for a large history like the kernel, and we wanted to minimize the amount of memory consumed by the bitmap index when loaded into the process.

In the JGit implementation our copy of Linus' tree has 3.1M objects, an 81.5 MiB idx file, and a 3.8 MiB bitmap file. We were trying to keep the overhead below 10% of the idx file, and I think we have succeeded on that. With 3.1M objects the v2 bitmap proposed in this thread needs at least 11.8M, or 14+% overhead just for the path hash cache.

The path hash cache may still be required, Colby and I have been debating the merits of having the data available for delta compression vs. the increase in memory required to hold it.

> The
> pack-order ones should be more amenable to run-length encoding,
> especially as you get further down into history (the tip ones would
> mostly be 1's, no matter how you order them).

This is also true for the type bitmaps, but especially so in the JGit file ordering where we always write all trees before any blobs. The type bitmaps are very compact and basically amount to defining a single range in the file. This takes only a few words in the EWAH compressed format.

Previous: Jeff KingNext: Jeff King
Message 31 of 64 in “Speed up Counting Objects with bitmap data”
  1. 00/16 Speed up Counting Objects with bitmap dataVicent Marti, Jun 24, 2013
  2. 01/16 list-objects: mark tree as unparsed when we free its bufferVicent Marti, Jun 24, 2013
  3. 02/16 sha1_file: refactor into `find_pack_object_pos`Vicent Marti, Jun 24, 2013
  4. Thomas RastJun 25, 2013
  5. 03/16 pack-objects: use a faster hash tableVicent Marti, Jun 24, 2013
  6. Thomas RastJun 25, 2013
  7. Jeff KingJun 26, 2013
  8. Jeff KingJun 26, 2013
  9. Ramkumar RamachandraJun 25, 2013
  10. Junio C HamanoJun 25, 2013
  11. Vicent MartíJun 25, 2013
  12. 04/16 pack-objects: make `pack_name_hash` globalVicent Marti, Jun 24, 2013
  13. 05/16 revision: allow setting custom limiter functionVicent Marti, Jun 24, 2013
  14. 06/16 sha1_file: export `git_open_noatime`Vicent Marti, Jun 24, 2013
  15. 07/16 compat: add endinanness helpersVicent Marti, Jun 24, 2013
  16. Peter KreftingJun 25, 2013
  17. Vicent MartíJun 25, 2013
  18. Peter KreftingJun 27, 2013
  19. 08/16 ewah: compressed bitmap implementationVicent Marti, Jun 24, 2013
  20. Junio C HamanoJun 25, 2013
  21. Junio C HamanoJun 25, 2013
  22. Thomas RastJun 25, 2013
  23. 09/16 documentation: add documentation for the bitmap formatVicent Marti, Jun 24, 2013
  24. Shawn PearceJun 25, 2013
  25. Vicent MartíJun 25, 2013
  26. Junio C HamanoJun 25, 2013
  27. Vicent MartíJun 25, 2013
  28. Shawn PearceJun 27, 2013
  29. Vicent MartíJun 27, 2013
  30. Jeff KingJun 27, 2013
  31. Shawn PearceJun 27, 2013
  32. Jeff KingJun 27, 2013
  33. Colby RangerJul 1, 2013
  34. Shawn PearceJul 1, 2013
  35. Jeff KingJul 7, 2013
  36. Shawn PearceJul 7, 2013
  37. Jeff KingJun 26, 2013
  38. Colby RangerJun 26, 2013
  39. Colby RangerJun 26, 2013
  40. Colby RangerJun 27, 2013
  41. Shawn PearceJun 27, 2013
  42. Shawn PearceJun 27, 2013
  43. Thomas RastJun 25, 2013
  44. Vicent MartíJun 25, 2013
  45. Thomas RastJun 26, 2013
  46. Thomas RastJun 26, 2013
  47. 10/16 pack-objects: use bitmaps when packing objectsVicent Marti, Jun 24, 2013
  48. Ramkumar RamachandraJun 25, 2013
  49. Thomas RastJun 25, 2013
  50. Junio C HamanoJun 25, 2013
  51. Vicent MartíJun 25, 2013
  52. 11/16 rev-list: add bitmap mode to speed up listsVicent Marti, Jun 24, 2013
  53. Thomas RastJun 25, 2013
  54. Vicent MartíJun 26, 2013
  55. Thomas RastJun 26, 2013
  56. Jeff KingJun 26, 2013
  57. 12/16 pack-objects: implement bitmap writingVicent Marti, Jun 24, 2013
  58. 13/16 repack: consider bitmaps when performing repacksVicent Marti, Jun 24, 2013
  59. Junio C HamanoJun 25, 2013
  60. Vicent MartíJun 25, 2013
  61. 14/16 sha1_file: implement `nth_packed_object_info`Vicent Marti, Jun 24, 2013
  62. 15/16 write-bitmap: implement new git command to write bitmapsVicent Marti, Jun 24, 2013
  63. 16/16 rev-list: Optimize --count using bitmaps tooVicent Marti, Jun 24, 2013
  64. Thomas RastJun 25, 2013

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.