git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: With big repos and slower connections, git clone can be hard to work with

From
Emily Shaffer <nasamuffin@google.com>
Date
Jun 10, 2024, 19:04 UTC
Message-ID
<CAJoAoZkP58ZM4J3ejemyiqkkbEaQdphoyGj_LmX9-xb_eMgb4A@mail.gmail.com>
In-Reply-To
<20240608084323.GB2390433@coredump.intra.peff.net>
On Sat, Jun 8, 2024 at 1:43 AM Jeff King <peff@peff.net> wrote:
Show 17 quoted lines
>
> On Sat, Jun 08, 2024 at 02:46:38AM +0200, ellie wrote:
>
> > The deepening worked perfectly, thank you so much! I hope a resume will
> > still be considered however, if even just to help out newcomers.
>
> Because the packfile to send the user is created on the fly, making a
> clone fully resumable is tricky (a second clone may get an equivalent
> but slightly different pack due to new objects entering the repo, or
> even raciness between threads).
>
> One strategy people have worked on is for servers to point clients at
> static packfiles (which _do_ remain byte-for-byte identical, and can be
> resumed) to get some of the objects. But it requires some scheme on the
> server side to decide when and how to create those packfiles. So while
> there is support inside Git itself for this idea (both on the server and
> client side), I don't know of any servers where it is in active use.

We use packfile offloading heavily at Google (any repositories hosted at *.googlesource.com, as well as our internal-facing hosting). It works quite well for us scaling large projects like Android and Chrome; we've been using it for some time now and are happy with it.

However, one thing that's missing is the resumable download Ellie is describing. With a clone which has been turned into a packfile fetch from a different data store, it *should* be resumable. But the client currently lacks the ability to do that. (This just came up for us internally the other day, and we ended up moving an internal bug to https://git.g-issues.gerritcodereview.com/issues/345241684.) After a resumed clone like this, you may not necessarily have latest - for example, you may lose connection with 90% of the clone finished, then not get connection back for some days, after which point upstream has moved as Peff described elsewhere in this thread. But it would still probably be cheaper to resume that 10% of packfile fetch from the offloaded data store, then do an incremental fetch back to the server to get the couple days of updates on top, as compared to starting over from zero with the server.

It seems to me that packfile URIs and bundle URIs are similar enough that we could work out similar logic for both, no? Or maybe there's something I'm missing about the way bundle offloading differs from packfiles.

 - Emily
>
> -Peff
>
Previous: Patrick SteinhardtNext: Junio C Hamano
Message 15 of 43 in “With big repos and slower connections, git clone can be hard to work with”
  1. ellieJun 7, 2024
  2. rsbecker@nexbridge.comJun 7, 2024
  3. ellieJun 8, 2024
  4. rsbecker@nexbridge.comJun 8, 2024
  5. ellieJun 8, 2024
  6. Jeff KingJun 8, 2024
  7. ellieJun 8, 2024
  8. ellieJun 8, 2024
  9. Jeff KingJun 8, 2024
  10. Jeff KingJun 8, 2024
  11. ellieJun 8, 2024
  12. Junio C HamanoJun 8, 2024
  13. ellieJun 8, 2024
  14. Patrick SteinhardtJun 10, 2024
  15. Emily ShafferJun 10, 2024
  16. Junio C HamanoJun 10, 2024
  17. ellieJun 10, 2024
  18. Toon claesJun 13, 2024
  19. Jeff KingJun 11, 2024
  20. Junio C HamanoJun 11, 2024
  21. Sitaram ChamartyJun 29, 2024
  22. Jeff KingJun 11, 2024
  23. Ivan FradeJun 11, 2024
  24. ellieJul 7, 2024
  25. rsbecker@nexbridge.comJul 8, 2024
  26. ellieJul 8, 2024
  27. rsbecker@nexbridge.comJul 8, 2024
  28. ellieJul 8, 2024
  29. Konstantin KhomoutovJul 8, 2024
  30. rsbecker@nexbridge.comJul 8, 2024
  31. ellieJul 8, 2024
  32. rsbecker@nexbridge.comJul 8, 2024
  33. ellieJul 8, 2024
  34. rsbecker@nexbridge.comJul 8, 2024
  35. ellieJul 8, 2024
  36. rsbecker@nexbridge.comJul 8, 2024
  37. Emanuel CziraiJul 8, 2024
  38. Konstantin KhomoutovJul 8, 2024
  39. rsbecker@nexbridge.comJul 8, 2024
  40. ellieJul 14, 2024
  41. ellieJul 24, 2024
  42. EllieSep 8, 2025
  43. EllieSep 30, 2024

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.