git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Balanced packing strategy

From
Junio C Hamano <junkio@cox.net>
Date
Nov 13, 2005, 23:13 UTC
Message-ID
<7vbr0o87lj.fsf@assigned-by-dhcp.cox.net>
In-Reply-To
<200511132106.29841.Josef.Weidendorfer@gmx.de>
Petr Baudis <pasky@suse.cz> writes:
Show 5 quoted lines
> This has the property that the second half of given pack is covered by
> objects with precision lower by one. This is a relatively high overload
> (this can be balanced by only keeping the last third or whatever), but
> it designed to reduce the overhead of fetching packs over dumb
> transport.

I have a feeling that you would be better off if instead do the repacking on the server side to prepare multiple packs, each of which has all the necessary objects to bring people who was up-to-date at various timerange ago, to arrange that you would need only one patch fetch with individual objects near the tip.

This obviously needs smarter client-side support.

Suppose we are somewhere after releasing v1.8 and inching towards v1.9:

In your proposal, the object ranges each pack contains would look like this:

 v1.0..v1.5 --------
 v1.5..v1.6        -----
 v1.6..v1.7            -------
 v1.7..v1.8                  -----
 individual objects               ....

That is, there are slight overlaps but you would do multiple packs if you are really behind.

Instead, you could do this:
 v1.0..v1.8 ----------------------
 v1.5..v1.8         --------------
 v1.6..v1.8             ----------
 v1.7..v1.8                   ----
 individual objects               ....

Everybody starts from the tip, fetching individual objects, and when the last repack boundary (the time we released 1.8) is reached, the dumb protocol downloader now faces a choice. The indices are fairly small, so you fetch all of them and see how many objects you are lacking from each pack. If you were up-to-date very long time ago, say at v1.2, you would obviously need to fetch the longest pack. If you were up-to-date recently, say after v1.6 was released, you need to fetch smaller pack.

Given the self containedness requirements, any path that is touched once in a period needs at least one full copy of it in each pack (all other revisions could be deltified), and I suspect in practice the oldest pack (v1.0..v1.5 pack in your scheme) would not save much space by not having v1.5..v1.8 history. We could tweak things further to do something like this:

 v1.0..v1.8 ------------------
 v1.5..v1.8         ----------
 v1.6..v1.8             ------
 v1.7..v1.8                   ----
 individual objects               ....

to also account for a fact that the recent ones cover shorter time range and not many paths are touched.

Previous: Josef Weidendorfer
Message 18 of 18 in “Remove unneeded packs”
  1. Marcel HoltmannNov 12, 2005
  2. Andreas EricssonNov 12, 2005
  3. Marcel HoltmannNov 12, 2005
  4. Lukas SandströmNov 12, 2005
  5. Marcel HoltmannNov 12, 2005
  6. Junio C HamanoNov 13, 2005
  7. Lukas SandströmNov 13, 2005
  8. Sergey VlasovNov 13, 2005
  9. Lukas SandströmNov 13, 2005
  10. Sergey VlasovNov 13, 2005
  11. Lukas SandströmNov 13, 2005
  12. Craig SchlenterNov 12, 2005
  13. Balanced packing strategyPetr Baudis, Nov 12, 2005
  14. Craig SchlenterNov 12, 2005
  15. Junio C HamanoNov 13, 2005
  16. Petr BaudisNov 13, 2005
  17. Josef WeidendorferNov 13, 2005
  18. Junio C HamanoNov 13, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.