git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Resumable git clone?

From
Josh Triplett <josh@joshtriplett.org>
Date
Mar 2, 2016, 16:41 UTC
Message-ID
<20160302164118.GA13732@x>
In-Reply-To
<xmqq4mcp5lij.fsf@gitster.mtv.corp.google.com>
On Wed, Mar 02, 2016 at 12:31:16AM -0800, Junio C Hamano wrote:
Show 12 quoted lines
> Josh Triplett <josh@joshtriplett.org> writes:
> > I think several simpler optimizations seem
> > preferable, such as binary object names, and abbreviating complete
> > object sets ("I have these commits/trees and everything they need
> > recursively; I also have this stack of random objects.").
> 
> Given the way pack stream is organized (i.e. commits first and then
> trees and blobs that belong to the same delta chain together), and
> our assumed goal being to salvage objects from an interrupted
> transfer of a packfile, you are unlikely to ever see "I have these
> commits/trees and everything they need" that are salvaged from such
> a failed transfer.  So I doubt such an optimization is worth doing.

True for the resumable clone case. For that optimization, I was thinking of the "pull during the merge window" case that Al Viro was also interested in optimizing.

> Besides it is very expensive to compute (the computation is done on
> the client side, so the cycles burned and the time the user has to
> wait is of much less concern, though); you'd essentially be doing
> "git fsck" to find the "dangling" objects.

Trading client-side computation for bandwidth can potentially be worthwhile if you have plenty of local compute but a slow and metered link.

Show 17 quoted lines
> The list of what would be transferred needs to come in full from the
> server end, as the list names objects that the receiving end may not
> have seen, but the response by the client could be encoded much
> tightly.  For the full list of N objects from the server, we can
> think of your response to be a bitstream of N bits, each on-bit in
> which signals an unwanted object in the list.  You can optimize this
> transfer by RLE compressing the bitstream, for example.
> 
> As git-over-HTTP is stateless, however, you cannot assume that the
> server side remembers what it sent to the client (instead, the
> client side needs to re-post what it heard from the server in the
> previous exchange to allow the server side to use it after
> validating).  So "objects at these indices in your list" kind of
> optimization may not work very well in that environment.  I'd
> imagine that an exchange of "Here are the list of objects", "Give me
> these objects" done naively in full 40-hex object names would work
> OK there, though.

Good point. Between statelessness and Duy's point about the client list usually being smaller than the server list, perhaps it would make sense to not have the server send a list at all, and just have the client send its own list.

- Josh Triplett
Previous: Duy NguyenNext: Josh Triplett
Message 10 of 24 in “Resumable git clone?”
  1. Josh TriplettMar 2, 2016
  2. Stefan BellerMar 2, 2016
  3. Al ViroMar 2, 2016
  4. Junio C HamanoMar 2, 2016
  5. Duy NguyenMar 2, 2016
  6. Duy NguyenMar 2, 2016
  7. Josh TriplettMar 2, 2016
  8. Junio C HamanoMar 2, 2016
  9. Duy NguyenMar 2, 2016
  10. Josh TriplettMar 2, 2016
  11. Josh TriplettMar 2, 2016
  12. Duy NguyenMar 2, 2016
  13. Jeff KingMar 2, 2016
  14. Bhavik BavishiMar 2, 2016
  15. Josh TriplettMar 2, 2016
  16. Duy NguyenMar 2, 2016
  17. Duy NguyenMar 2, 2016
  18. Junio C HamanoMar 2, 2016
  19. Konstantin RyabitsevMar 2, 2016
  20. Josh TriplettMar 2, 2016
  21. Junio C HamanoMar 2, 2016
  22. Philip OakleyMar 24, 2016
  23. Junio C HamanoMar 24, 2016
  24. Philip OakleyMar 24, 2016

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.