git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 09/16] documentation: add documentation for the bitmap format

From
Shawn Pearce <spearce@spearce.org>
Date
Jun 27, 2013, 01:29 UTC
Message-ID
<CAJo=hJuH98sT5WNykxQ5JX+yKxOH-5p3CCRGa-WLAYtMGAj6oA@mail.gmail.com>
In-Reply-To
<20130626051117.GB26755@sigill.intra.peff.net>
On Tue, Jun 25, 2013 at 11:11 PM, Jeff King <peff@peff.net> wrote:
Show 22 quoted lines
> On Tue, Jun 25, 2013 at 09:33:11PM +0200, Vicent Martí wrote:
>
>> > One way we side-stepped the size inflation problem in JGit was to only
>> > use the bitmap index information when sending data on the wire to a
>> > client. Here delta reuse plays a significant factor in building the
>> > pack, and we don't have to be as accurate on matching deltas. During
>> > the equivalent of `git repack` bitmaps are not used, allowing the
>> > traditional graph enumeration algorithm to generate path hash
>> > information.
>>
>> OH BOY HERE WE GO. This is worth its own thread, lots to discuss here.
>> I think peff will have a patchset regarding this to upstream soon,
>> we'll get back to it later.
>
> We do the same thing (only use bitmaps during on-the-wire fetches).  But
> there a few problems with assuming delta reuse.
>
> For us (GitHub), the foremost one is that we pack many "forks" of a
> repository together into a single packfile. That means when you clone
> torvalds/linux, an object you want may be stored in the on-disk pack
> with a delta against an object that you are not going to get. So we have
> to throw out that delta and find a new one.

Gerrit Code Review ran into the same problem a few years ago with the refs/changes namespace. Objects reachable from a branch were often delta compressed against dropped code review revisions, making for some slow transfers. We fixed this by creating a pack of everything reachable from refs/heads/* and then another pack of the other stuff.

I would encourage you to do what you suggest...
Show 10 quoted lines
> I'm dealing with that by adding an option to respect "islands" during
> packing, where an island is a set of common objects (we split it by
> fork, since we expect those objects to be fetched together, but you
> could use other criteria). The rule is that an object cannot delta
> against another object that is not in all of its islands. So everybody
> can delta against shared history, but objects in your fork can only
> delta against other objects in the fork.  You are guaranteed to be able
> to reuse such deltas during a full clone of a fork, and the on-disk pack
> size does not suffer all that much (because there is usually a good
> alternate delta base within your reachable history).

Yes, exactly. I want to do the same thing on our servers, as we have many forks of some popular open source repositories that are also not small (Linux kernel, WebKit). Unfortunately Google has not had the time to develop the necessary support into JGit.

Show 10 quoted lines
> So with that series, we can get good reuse for clones. But there are
> still two cases worth considering:
>
>   1. When you fetch a subset of the commits, git marks only the edges as
>      preferred bases, and does not walk the full object graph down to
>      the roots. So any object you want that is delta'd against something
>      older will not get reused. If you have reachability bitmaps, I
>      don't think there is any reason that we cannot use the entire
>      object graph (starting at the "have" tips, of course) as preferred
>      bases.

In JGit we use the reachability bitmap to provide proof a client has an object. Even if its not in the edges. This allows us much better delta reuse, as often frequently deltas will be available pointing to something behind the edge, but that the client certainly has given the edges we know about.

We also use the reachability bitmap to provide proof a client does not need an object. We found a reduction in number of objects transferred because the "want AND NOT have" subtracted out a number of objects not in the edge. Apparently merges, reverts and cherry-picks happen often enough in the repositories we host that this particular optimization helps reduce data transfer, and work at both server and client ends of the connection. Its a nice freebie the bitmap algorithm gives us.

>   2. The server is not necessarily fully packed. In an active repo, you
>      may have a large "base" pack with bitmaps, with several recently
>      pushed packs on top. You still need to delta the recently pushed
>      objects against the base objects.

Yes, this is unfortunate. One way we avoid this in JGit is to keep everything in pack files, rather than exploding loose. The reachability bitmap often proves the client has the delta base the pusher used to make the object, allowing us to reuse the delta. It may not be the absolute best delta in the world, but reuse is faster than inflate()+delta()+deflate(), and the delta is probably "good enough" until the server can do a real GC in the background.

We combine small packs from pushes together by almost literally just concat'ing the packs together and creating a new .idx. Newer pushed data is put in front of the older data, the pack is clustered by "commit, tree, blob" ordering, duplicates are removed, and its written back to disk. Typically we complete this "pack concat" operation mere seconds after a push finishes, so readers have very few packs to deal with.

> I don't have measurements on how much the deltas suffer in those two
> cases. I know they suffered quite badly for clones without the name
> hashes in our alternates repos, but that part should go away with my
> patch series.

JGit doesn't poke objects into the object table (or even the object list) when a bitmap is used. We spool the bits out of the bitmap in bitmap order and write them to the wire in that order. Its way faster, but depends on the bitmap being in pack-ordering. So clones are crazy fast even though we don't have the path-hash table.

JGit also has another optimization where we figure out based on the bitmap if the client needs *everything* in this pack. Which given a pack created only for refs/heads/* is the common case for a clone. If the client is getting all objects we essentially just do a sendfile() for the region starting at offset 12 through end-20. It can't be a sendfile() syscall because it has to be computed into the trailer SHA-1 the client sees, but its a crazy tight IO copy loop with no Git smarts beyond the SHA-1 updating.

Like I said, there are a ton of optimizations you guys missed. And we think they make a bigger difference than screwing around with little-endian format to favor x86 CPUs.

Previous: Shawn PearceNext: Thomas Rast
Message 42 of 64 in “Speed up Counting Objects with bitmap data”
  1. 00/16 Speed up Counting Objects with bitmap dataVicent Marti, Jun 24, 2013
  2. 01/16 list-objects: mark tree as unparsed when we free its bufferVicent Marti, Jun 24, 2013
  3. 02/16 sha1_file: refactor into `find_pack_object_pos`Vicent Marti, Jun 24, 2013
  4. Thomas RastJun 25, 2013
  5. 03/16 pack-objects: use a faster hash tableVicent Marti, Jun 24, 2013
  6. Thomas RastJun 25, 2013
  7. Jeff KingJun 26, 2013
  8. Jeff KingJun 26, 2013
  9. Ramkumar RamachandraJun 25, 2013
  10. Junio C HamanoJun 25, 2013
  11. Vicent MartíJun 25, 2013
  12. 04/16 pack-objects: make `pack_name_hash` globalVicent Marti, Jun 24, 2013
  13. 05/16 revision: allow setting custom limiter functionVicent Marti, Jun 24, 2013
  14. 06/16 sha1_file: export `git_open_noatime`Vicent Marti, Jun 24, 2013
  15. 07/16 compat: add endinanness helpersVicent Marti, Jun 24, 2013
  16. Peter KreftingJun 25, 2013
  17. Vicent MartíJun 25, 2013
  18. Peter KreftingJun 27, 2013
  19. 08/16 ewah: compressed bitmap implementationVicent Marti, Jun 24, 2013
  20. Junio C HamanoJun 25, 2013
  21. Junio C HamanoJun 25, 2013
  22. Thomas RastJun 25, 2013
  23. 09/16 documentation: add documentation for the bitmap formatVicent Marti, Jun 24, 2013
  24. Shawn PearceJun 25, 2013
  25. Vicent MartíJun 25, 2013
  26. Junio C HamanoJun 25, 2013
  27. Vicent MartíJun 25, 2013
  28. Shawn PearceJun 27, 2013
  29. Vicent MartíJun 27, 2013
  30. Jeff KingJun 27, 2013
  31. Shawn PearceJun 27, 2013
  32. Jeff KingJun 27, 2013
  33. Colby RangerJul 1, 2013
  34. Shawn PearceJul 1, 2013
  35. Jeff KingJul 7, 2013
  36. Shawn PearceJul 7, 2013
  37. Jeff KingJun 26, 2013
  38. Colby RangerJun 26, 2013
  39. Colby RangerJun 26, 2013
  40. Colby RangerJun 27, 2013
  41. Shawn PearceJun 27, 2013
  42. Shawn PearceJun 27, 2013
  43. Thomas RastJun 25, 2013
  44. Vicent MartíJun 25, 2013
  45. Thomas RastJun 26, 2013
  46. Thomas RastJun 26, 2013
  47. 10/16 pack-objects: use bitmaps when packing objectsVicent Marti, Jun 24, 2013
  48. Ramkumar RamachandraJun 25, 2013
  49. Thomas RastJun 25, 2013
  50. Junio C HamanoJun 25, 2013
  51. Vicent MartíJun 25, 2013
  52. 11/16 rev-list: add bitmap mode to speed up listsVicent Marti, Jun 24, 2013
  53. Thomas RastJun 25, 2013
  54. Vicent MartíJun 26, 2013
  55. Thomas RastJun 26, 2013
  56. Jeff KingJun 26, 2013
  57. 12/16 pack-objects: implement bitmap writingVicent Marti, Jun 24, 2013
  58. 13/16 repack: consider bitmaps when performing repacksVicent Marti, Jun 24, 2013
  59. Junio C HamanoJun 25, 2013
  60. Vicent MartíJun 25, 2013
  61. 14/16 sha1_file: implement `nth_packed_object_info`Vicent Marti, Jun 24, 2013
  62. 15/16 write-bitmap: implement new git command to write bitmapsVicent Marti, Jun 24, 2013
  63. 16/16 rev-list: Optimize --count using bitmaps tooVicent Marti, Jun 24, 2013
  64. Thomas RastJun 25, 2013

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.