git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC] Design for http-pull on repo with packs

From
Junio C Hamano <junkio@cox.net>
Date
Jul 11, 2005, 23:30 UTC
Message-ID
<7vu0j0ncnr.fsf@assigned-by-dhcp.cox.net>
In-Reply-To
<42D2960E.3050008@gmail.com>
Dan Holmsand <holmsand@gmail.com> writes:
Show 11 quoted lines
> I did a little experiment. I cloned Linus' current tree, and git
> repacked everything (that's 63M + 3.3M worth of pack files). Then I
> got something like 25 or so of Jeff's branches. That's 6.9M of object
> files, and 1.4M packed. Total size: 70M for the entire
> .git/objects/pack directory.
>
> Repacking all of that to a single pack file gives, somewhat
> surprisingly, a pack size of 62M (+ 1.3M index). In other words, the
> cost of getting all those branches, and all of the new stuff from
> Linus, turns out to be *negative* (probably due to some strange
> deltification coincidence).

We do _not_ want to optimize for initial slurps into empty repositories. Quite the opposite. We want to optimize for allowing quick updates of reasonably up-to-date developer repos. If initial slurps are _also_ efficient then that is an added bonus; that is something the baseline big pack (60M Linus pack) would give us already. So repacking everything into a single pack nightly is _not_ what we want to do, even though that would give the maximum compression ;-). I know you understand this, but just stating the second of the above paragraphs would give casual readers a wrong impression.

> I think that this shows that (at least in this case), having many
> branches isn't particularly wasteful (1.4M in this case with one
> incremental pack).
> And that fewer packs beats many packs quite handily.

You are correct. For somebody like Jeff, having the Linus baseline pack with one pack of all of his head (incremental that excludes what is already in the Linus baseline pack) would help pullers.

Show 5 quoted lines
> The big problem, however, comes when Jeff (or anyone else) decides to
> repack. Then, if you fetch both his repo and Linus', you might end up
> with several really big pack files, that mostly overlap. That could
> easily mean storing most objects many times, if you don't do some
> smart selective un/repacking when fetching.

Indeed. Overlapping packs is a possibility, but my gut feeling is that it would not be too bad, if things are arranged so that packs are expanded-and-then-repacked _very_ rarely if ever. Instead, at least for your public repository, if you only repack incrementally I think you would be OK.

Previous: Tony LuckNext: Dan Holmsand
Message 8 of 9 in “[RFC] Design for http-pull on repo with packs”
  1. Daniel BarkalowJul 10, 2005
  2. Dan HolmsandJul 10, 2005
  3. Daniel BarkalowJul 10, 2005
  4. Dan HolmsandJul 10, 2005
  5. Junio C HamanoJul 11, 2005
  6. Dan HolmsandJul 11, 2005
  7. Tony LuckJul 11, 2005
  8. Junio C HamanoJul 11, 2005
  9. Dan HolmsandJul 12, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.