git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Partial-clone cause big performance impact on server

From
Derrick Stolee <derrickstolee@github.com>
Date
Aug 12, 2022, 12:21 UTC
Message-ID
<16633d89-6ccd-859d-8533-9861ad831c45@github.com>
In-Reply-To
<bfa3de4485614badb4a27d8cfba99968@xiaomi.com>
On 8/11/22 4:09 AM, 程洋 wrote:> Hi.
>      We observed big disk space save by partial-clone and require all of our users (2000+) to clone repository with partial-clone (filter=blob:none)
>      However at busy time, we found it's extremely slow for user to fetch. Here is what we did.
> 
>     1. ask all users to fetch with filter=blob:none. And it's remarkable. Now our download size per user decrease from 460G to 180G.

I hope this includes the blob download during the initial checkout, because otherwise you have a very strange shape to make your commits and trees take up 180 GB.

>     2. But at busy time, everyone's fetch become slow. (at idle hours, it takes us 5 minutes to clone a big repositories, but it takes more than 1 hour to clone the same repositories at busy hours)
>     3. with GIT_TRACE_PACKET=1. We found on big repositories (200K+refs, 6m+ objects). Git will sends 40k want.

You only have six million objects in the repo and yet have that size? It must be some very large blobs.

>     4. And we then track our server(which is gerrit with jgit). We found the server is couting objects. Then we check those 40k objects, most of them are blobs rather than commit. (which means they're not in bitmap)

Are you seeing any commits in these requests? If the Git client is asking for blobs, then they should not be mixed with commit wants. What kind of operation are you doing to see these mixed wants?

If the request was only blobs, then the server should not need a "Counting objects" phase. It should jump immediately to preparing the objects (which will likely require parsing deltas, and that can be expensive). I don't know if JGit is doing something different, though.

>     5. We believe that's the root cause of our problem. Git sends too many "want SHA1" which are not in bitmap, cause the server to count objects  frequently, which then slow down the server.
> 
> What we want is, download the things we need to checkout to specific commit. But if one commit contain so many objects (like us , 40k+). It takes more time to counting than downloading.

One thing that the microsoft/git fork uses in its "git-gvfs-helper" tool (which speaks the GVFS Protocol as a replacement for partial clone when using Azure Repos as a server) is a batched download of missing objects [1]. The initial limit is 4000 objects at a time, but that helps keep each request small enough that it is less likely to fail for scale reasons alone.

[1] https://github.com/microsoft/git/blob/vfs-2.37.1/gvfs-helper.c#L3510-L3520

It might be interesting to create such batch-downloads for these partial clone blob-fetches.

Thanks, -Stolee

Previous: 程洋Next: Jeff King
Message 7 of 40 in “Partial-clone cause big performance impact on server”
  1. 程洋Aug 11, 2022
  2. Jonathan TanAug 11, 2022
  3. 回复: [External Mail]Re: Partial-clone cause big performance impact on server程洋, Aug 13, 2022
  4. 回复: [External Mail]Re: Partial-clone cause big performance impact on server程洋, Aug 13, 2022
  5. ZheNing HuAug 15, 2022
  6. 程洋Aug 15, 2022
  7. Derrick StoleeAug 12, 2022
  8. Jeff KingAug 14, 2022
  9. Derrick StoleeAug 15, 2022
  10. 程洋Aug 15, 2022
  11. 程洋Aug 17, 2022
  12. Derrick StoleeAug 17, 2022
  13. Jeff KingAug 18, 2022
  14. 程洋Sep 1, 2022
  15. Jeff KingSep 1, 2022
  16. 程洋Sep 5, 2022
  17. Jeff KingSep 6, 2022
  18. 0/3 speeding up on-demand fetch for blobs in partial cloneJeff King, Sep 6, 2022
  19. 1/3 parse_object(): allow skipping hash checkJeff King, Sep 6, 2022
  20. Derrick StoleeSep 7, 2022
  21. Jeff KingSep 7, 2022
  22. 2/3 upload-pack: skip parse-object re-hashing of "want" objectsJeff King, Sep 6, 2022
  23. Derrick StoleeSep 7, 2022
  24. Derrick StoleeSep 7, 2022
  25. Jeff KingSep 7, 2022
  26. Junio C HamanoSep 7, 2022
  27. Jeff KingSep 7, 2022
  28. [BUG] t1800: Fails for error text comparisonrsbecker@nexbridge.com, Sep 7, 2022
  29. Junio C HamanoSep 7, 2022
  30. rsbecker@nexbridge.comSep 7, 2022
  31. Jeff KingSep 7, 2022
  32. Junio C HamanoSep 7, 2022
  33. Jeff KingSep 8, 2022
  34. Junio C HamanoSep 8, 2022
  35. 3/3 parse_object(): check commit-graph when skip_hash setJeff King, Sep 6, 2022
  36. Derrick StoleeSep 7, 2022
  37. Junio C HamanoSep 7, 2022
  38. 程洋Sep 8, 2022
  39. Jeff KingSep 8, 2022
  40. Derrick StoleeSep 7, 2022

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.