git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Performance issue: initial git clone causes massive repack

From
Nicolas Pitre <nico@cam.org>
Date
Apr 6, 2009, 04:06 UTC
Message-ID
<alpine.LFD.2.00.0904052336260.6741@xanadu.home>
In-Reply-To
<20090405T230552Z@curie.orbis-terrarum.net>
On Sun, 5 Apr 2009, Robin H. Johnson wrote:
Show 22 quoted lines
> On Sun, Apr 05, 2009 at 03:57:14PM -0400, Jeff King wrote:
> > > During an initial clone, I see that git-upload-pack invokes
> > > pack-objects, despite the ENTIRE repository already being packed - no
> > > loose objects whatsoever. git-upload-pack then seems to buffer in
> > > memory.
> > We need to run pack-objects even if the repo is fully packed because we
> > don't know what's _in_ the existing pack (or packs). In particular we
> > want to:
> >   - combine multiple packs into a single pack; this is more efficient on
> >     the network, because you can find more deltas, and I believe is
> >     required because the protocol sends only a single pack.
> > 
> >   - cull any objects which are not actually part of the reachability
> >     chain from the refs we are sending
> > 
> > If no work needs to be done for either case, then pack-objects should
> > basically just figure that out and then send the existing pack (the
> > expensive bit is doing deltas, and we don't consider objects in the same
> > pack for deltas, as we know we have already considered that during the
> > last repack). It does mmap the whole pack, so you will see your virtual
> > memory jump, but nothing should require the whole pack being in memory
> > at once.

Actually the pack is mapped with a (configurable) window. See the core.packedGitWindowSize and core.packedGitLimit config options for details.

> While my current pack setup has multiple packs of not more than 100MiB
> each, that was simply for ease of resume with rsync+http tests. Even
> when I already had a single pack, with every object reachable,
> pack-objects was redoing the packing.
In that case it shouldn't have.
Show 19 quoted lines
> > pack-objects streams the output to upload-pack, which should only ever
> > have an 8K buffer of it in memory at any given time.
> > 
> > At least that is how it is all supposed to work, according to my
> > understanding. So if you are seeing very high memory usage, I wonder if
> > there is a bug in pack-objects or upload-pack that can be fixed.
> > 
> > Maybe somebody more knowledgeable than me about packing can comment.
> Looking at the source, I agree that it should be buffering, however top and ps
> seem to disagree. 3GiB VSZ and 2.5GiB RSS here now.
> 
> %CPU %MEM     VSZ     RSS STAT START   TIME COMMAND
>  0.0  0.0  140932    1040 Ss   16:09   0:00 \_ git-upload-pack /code/gentoo/gentoo-git/gentoo-x86.git 
> 32.2  0.0       0       0 Z    16:09   1:50     \_ [git-upload-pack] <defunct>
> 80.8 44.2 3018484 2545700 Sl   16:09   4:36     \_ git pack-objects --stdout --progress --delta-base-offset 
> 
> Also, I did another trace, using some other hardware, in a LAN setting, and
> noticed that git-upload-pack/pack-objects only seems to start output to the
> network after it reaches 100% in 'remote: Compressing objects:'.

That's to be expected. Delta compression matches objects which are not in the stream order at all. Therefore it is not possible to start outputting pack data until this pass is done. Still, this pass should not be invoked if your repository is already fully packed into one pack. Can you confirm this is actually the case?

> Relatedly, throwing more RAM (6GiB total, vs. the previous 2GiB) at 
> the server in this case cut the 200 wallclock minutes before any 
> sending too place down to 5 minutes.

Well... here's a wild guess. In the source repository serving clone requests, please do:

	git config pack.deltaCacheSize 1
	git config pack.deltaCacheLimit 0
and try cloning again with a fully packed repository.
Show 8 quoted lines
> > > For the initial clone, can the git-upload-pack algorithm please send
> > > existing packs, and only generate a pack containing the non-packed
> > > items?
> > 
> > I believe that would require a change to the protocol to allow multiple
> > packs. However, it may be possible to munge the pack header in such a
> > way that you basically concatenate multiple packs. You would still want
> > to peek in the big pack to try deltas from the non-packed items, though.

As explained already, even if the protocol requires a single pack to be created, it is still made up of unmodified data segments from existing packs as much as possible. So you should see it more or less as the concatenation of those packs already, plus some munging over the edges.

Show 7 quoted lines
> > I think all of this falls into the realm of the GSOC pack caching project.
> > There have been other discussions on the list, so you might want to look
> > through those for something useful.
> Yes, both changing the protocol, and recognizing that existing packs may be
> suitable to send could be considered as part of the caching project, as they
> fall under the aegis of making good use of what's stored in the cache already
> to send.

The caching pack project is to address a different issue: mainly to bypass the object enumeration cost. In other words, it could allow for skipping the "Counting objects" pass, and a tiny bit more. At least in theory that's about the main difference. This has many drawbacks as well though.

Nicolas
Previous: Nicolas PitreNext: Robin H. Johnson
Message 94 of 97 in “Performance issue: initial git clone causes massive repack”
  1. Robin H. JohnsonApr 4, 2009
  2. Nicolas SebrechtApr 5, 2009
  3. Robin H. JohnsonApr 5, 2009
  4. Nicolas SebrechtApr 5, 2009
  5. Nicolas SebrechtApr 5, 2009
  6. Robin H. JohnsonApr 5, 2009
  7. Nicolas SebrechtApr 5, 2009
  8. Shawn O. PearceApr 5, 2009
  9. Robin H. JohnsonApr 5, 2009
  10. Robin H. JohnsonApr 5, 2009
  11. Shawn O. PearceApr 5, 2009
  12. david@lang.hmApr 5, 2009
  13. Sverre RabbelierApr 5, 2009
  14. Nicolas PitreApr 6, 2009
  15. Björn SteinbrinkApr 7, 2009
  16. Jakub NarebskiApr 7, 2009
  17. Nicolas PitreApr 7, 2009
  18. Jakub NarebskiApr 7, 2009
  19. Jon SmirlApr 7, 2009
  20. Nicolas PitreApr 7, 2009
  21. Björn SteinbrinkApr 7, 2009
  22. Nicolas PitreApr 7, 2009
  23. Björn SteinbrinkApr 7, 2009
  24. Nicolas PitreApr 7, 2009
  25. Björn SteinbrinkApr 7, 2009
  26. Nicolas PitreApr 8, 2009
  27. Robin H. JohnsonApr 10, 2009
  28. Nicolas PitreApr 11, 2009
  29. Mike HommeyApr 11, 2009
  30. Johannes SchindelinApr 14, 2009
  31. Nicolas PitreApr 14, 2009
  32. Robin H. JohnsonApr 14, 2009
  33. Nicolas PitreApr 14, 2009
  34. Nguyen Thai Ngoc DuyApr 15, 2009
  35. Robin H. JohnsonApr 15, 2009
  36. Junio C HamanoApr 15, 2009
  37. Nicolas PitreApr 15, 2009
  38. Sam VilainApr 22, 2009
  39. Mike RalphsonApr 22, 2009
  40. Pieter de BieApr 22, 2009
  41. Johannes SchindelinApr 22, 2009
  42. Shawn O. PearceApr 22, 2009
  43. Andreas EricssonApr 22, 2009
  44. Johannes SchindelinApr 22, 2009
  45. Christian CouderApr 23, 2009
  46. Nicolas PitreApr 22, 2009
  47. Sam VilainApr 22, 2009
  48. Björn SteinbrinkApr 22, 2009
  49. Nicolas PitreApr 22, 2009
  50. Johannes SchindelinApr 22, 2009
  51. Nicolas PitreApr 23, 2009
  52. Johannes SchindelinApr 14, 2009
  53. Jeff KingApr 7, 2009
  54. Björn SteinbrinkApr 7, 2009
  55. process_{tree,blob}: Remove useless xstrdup callsBjörn Steinbrink, Apr 8, 2009
  56. Linus TorvaldsApr 10, 2009
  57. Linus TorvaldsApr 11, 2009
  58. Linus TorvaldsApr 11, 2009
  59. Nicolas PitreApr 11, 2009
  60. Björn SteinbrinkApr 11, 2009
  61. Björn SteinbrinkApr 11, 2009
  62. Linus TorvaldsApr 11, 2009
  63. Linus TorvaldsApr 11, 2009
  64. Björn SteinbrinkApr 11, 2009
  65. Björn SteinbrinkApr 11, 2009
  66. Linus TorvaldsApr 11, 2009
  67. Björn SteinbrinkApr 11, 2009
  68. Linus TorvaldsApr 11, 2009
  69. Björn SteinbrinkApr 11, 2009
  70. Linus TorvaldsApr 11, 2009
  71. Nicolas SebrechtApr 5, 2009
  72. david@lang.hmApr 5, 2009
  73. Robin RosenbergApr 5, 2009
  74. Nicolas PitreApr 6, 2009
  75. Junio C HamanoApr 6, 2009
  76. Nicolas PitreApr 6, 2009
  77. Jon SmirlApr 6, 2009
  78. Nicolas PitreApr 6, 2009
  79. Jon SmirlApr 6, 2009
  80. Shawn O. PearceApr 6, 2009
  81. Nicolas PitreApr 6, 2009
  82. Jon SmirlApr 6, 2009
  83. Nicolas PitreApr 6, 2009
  84. Matthieu MoyApr 6, 2009
  85. Nicolas PitreApr 6, 2009
  86. Robin H. JohnsonApr 6, 2009
  87. Nicolas PitreApr 6, 2009
  88. Martin LanghoffApr 7, 2009
  89. Jeff KingApr 5, 2009
  90. Robin H. JohnsonApr 5, 2009
  91. Robin H. JohnsonApr 5, 2009
  92. Nguyen Thai Ngoc DuyApr 6, 2009
  93. Nicolas PitreApr 6, 2009
  94. Nicolas PitreApr 6, 2009
  95. Robin H. JohnsonApr 6, 2009
  96. Mark LevedahlApr 11, 2009
  97. Robin H. JohnsonApr 6, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.