git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Inefficiency of partial shallow clone vs shallow clone + "old-style" sparse checkout

From
Jeff King <peff@peff.net>
Date
Apr 1, 2020, 11:44 UTC
Message-ID
<20200401114446.GA1589184@coredump.intra.peff.net>
In-Reply-To
<8268671585700012@iva3-58091f505f14.qloud-c.yandex.net>
On Wed, Apr 01, 2020 at 04:49:20AM +0300, Konstantin Tokarev wrote:
Show 11 quoted lines
> > Less efficient use of network bandwidth is one thing, but shallow clones are
> > also more CPU-intensive with the "counting objects" phase on the server. Your
> > link shares the following end-to-end timings:
> >
> > * Shallow-clone: 234s
> > * Partial clone: 286s
> > * Both(???): 1023s
> >
> > The data implies that by asking for both you actually got a full clone (4.1 GB).
> 
> No, this is still a partial clone, full clone takes more than 6 GB

I think that 4GB number is just because of the bug, though. With the fix I showed earlier, doing clones of linux.git from a local repo yields:

  type       objects (in passes)      bytes  time
  ----       -----------------------  -----  ----
  shallow      71447 (  71447+  n/a)  188MB   23s
  blob:none  5260567 (5193557+67010)  870MB   99s
  both         71447 (   4437+67010)  188MB   37s

The object counts and sizes make sense. blob:none is still going to get the whole history of commits and trees, which are substantial. The sizes for "shallow" and "both" are the same, because the checkout is going to grab all of the blobs from the tip commit, which were included in the original "shallow" anyway. It does take longer, because they come in a second followup fetch (though I'm surprised it's so _much_ slower).

So to me that implies that shallow is strictly better than partial if you're just going to check out the full tip commit. But doing both together opens up the possibility of narrowing the sparse checkout. Doing:

  $ git clone --no-local --no-checkout --filter=blob:none --depth=1 \
      /path/to/linux sparse
  $ cd sparse
  $ git sparse-checkout set arch
fetches 20795 objects (4437+16357+1), consuming only 27MB.
-Peff
Previous: Konstantin TokarevNext: Jeff King
Message 10 of 13 in “Inefficiency of partial shallow clone vs shallow clone + "old-style" sparse checkout”
  1. Konstantin TokarevMar 27, 2020
  2. Jeff KingMar 28, 2020
  3. Derrick StoleeMar 28, 2020
  4. Taylor BlauMar 31, 2020
  5. Jeff KingApr 1, 2020
  6. Konstantin TokarevMar 31, 2020
  7. Konstantin TokarevMar 31, 2020
  8. Derrick StoleeApr 1, 2020
  9. Konstantin TokarevApr 1, 2020
  10. Jeff KingApr 1, 2020
  11. clone: use "quick" lookup while following tagsJeff King, Apr 1, 2020
  12. Konstantin TokarevApr 1, 2020
  13. Jeff KingApr 1, 2020

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.