git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Performance issue: initial git clone causes massive repack

From
Nicolas Pitre <nico@cam.org>
Date
Apr 6, 2009, 15:14 UTC
Message-ID
<alpine.LFD.2.00.0904061042300.6741@xanadu.home>
In-Reply-To
<9e4733910904060737k3d1c082fk785cd98cdeb6d73d@mail.gmail.com>
On Mon, 6 Apr 2009, Jon Smirl wrote:
Show 17 quoted lines
> On Mon, Apr 6, 2009 at 10:19 AM, Nicolas Pitre <nico@cam.org> wrote:
> > On Mon, 6 Apr 2009, Jon Smirl wrote:
> >
> >> First thing an initial clone does is copy all of the pack files from
> >> the server to the client without even looking at them.
> >
> > This is a no go for reasons already stated many times.  There are
> > security implications (those packs might contain stuff that you didn't
> > intend to be publically accessible) and there might be efficiency
> > reasons as well (you might have a shared object store with lots of stuff
> > unrelated to the particular clone).
> 
> How do you deal with dense history packs? These packs take many hours
> to make (on a server class machine) and can be half the size of a
> regular pack. Shouldn't there be a way to copy these packs intact on
> an initial clone? It's ok if these packs are specially marked as being
> ok to copy.
[sigh]
Let me explain it all again.

There is basically two ways to create a new pack: the intelligent way, and the bruteforce way.

When creating a new pack the intelligent way, what we do is to enumerate all the needed object and look them up in the object store. When a particular object is found, we create a record for that object and note in which pack it is located, at what offset in that pack, how much space it occupies in its compressed form within that pack, , and if whether it is a delta or not. When that object is indeed a delta (the majority of objects usually are) then we also keep a pointer on the record for the base object for that delta.

Next, for all objects in delta form which base object is also part of the object enumeration and obviously part of the same pack, we simply flag those objects as directly reusable without any further processing. This means that, when those objects are about to be stored in the new pack, their raw data is simply copied straight from the original pack using the offset and size noted above. In other words, those objects are simply never redeltified nor redeflated at all, and all the work that was previously done to find the best delta match is preserved with no extra cost.

Of course, when your repository is tightly packed into a single pack, then all enumerated objects fall into the reusable category and therefore a copy of the original pack is indeed sent over the wire. One exception is with older git clients which don't support the delta base offset encoding, in which case the delta reference encoding is substituted on the fly with almost no cost (this is btw another reason why a dumb copy of existing pack may not work universally either). But in the common case, you might see the above as just the same as if git did copy the pack file because it really only reads some data from a pack and immediately writes that data out.

The bruteforce repacking is different because it simply doesn't concern itself with existing deltas at all. It instead start everything from scratch and perform the whole delta search all over for all objects. This is what takes lots of resources and CPU cycles, and as you may guess, is never used for fetch/clone requests.

Nicolas
Previous: Shawn O. PearceNext: Jon Smirl
Message 81 of 97 in “Performance issue: initial git clone causes massive repack”
  1. Robin H. JohnsonApr 4, 2009
  2. Nicolas SebrechtApr 5, 2009
  3. Robin H. JohnsonApr 5, 2009
  4. Nicolas SebrechtApr 5, 2009
  5. Nicolas SebrechtApr 5, 2009
  6. Robin H. JohnsonApr 5, 2009
  7. Nicolas SebrechtApr 5, 2009
  8. Shawn O. PearceApr 5, 2009
  9. Robin H. JohnsonApr 5, 2009
  10. Robin H. JohnsonApr 5, 2009
  11. Shawn O. PearceApr 5, 2009
  12. david@lang.hmApr 5, 2009
  13. Sverre RabbelierApr 5, 2009
  14. Nicolas PitreApr 6, 2009
  15. Björn SteinbrinkApr 7, 2009
  16. Jakub NarebskiApr 7, 2009
  17. Nicolas PitreApr 7, 2009
  18. Jakub NarebskiApr 7, 2009
  19. Jon SmirlApr 7, 2009
  20. Nicolas PitreApr 7, 2009
  21. Björn SteinbrinkApr 7, 2009
  22. Nicolas PitreApr 7, 2009
  23. Björn SteinbrinkApr 7, 2009
  24. Nicolas PitreApr 7, 2009
  25. Björn SteinbrinkApr 7, 2009
  26. Nicolas PitreApr 8, 2009
  27. Robin H. JohnsonApr 10, 2009
  28. Nicolas PitreApr 11, 2009
  29. Mike HommeyApr 11, 2009
  30. Johannes SchindelinApr 14, 2009
  31. Nicolas PitreApr 14, 2009
  32. Robin H. JohnsonApr 14, 2009
  33. Nicolas PitreApr 14, 2009
  34. Nguyen Thai Ngoc DuyApr 15, 2009
  35. Robin H. JohnsonApr 15, 2009
  36. Junio C HamanoApr 15, 2009
  37. Nicolas PitreApr 15, 2009
  38. Sam VilainApr 22, 2009
  39. Mike RalphsonApr 22, 2009
  40. Pieter de BieApr 22, 2009
  41. Johannes SchindelinApr 22, 2009
  42. Shawn O. PearceApr 22, 2009
  43. Andreas EricssonApr 22, 2009
  44. Johannes SchindelinApr 22, 2009
  45. Christian CouderApr 23, 2009
  46. Nicolas PitreApr 22, 2009
  47. Sam VilainApr 22, 2009
  48. Björn SteinbrinkApr 22, 2009
  49. Nicolas PitreApr 22, 2009
  50. Johannes SchindelinApr 22, 2009
  51. Nicolas PitreApr 23, 2009
  52. Johannes SchindelinApr 14, 2009
  53. Jeff KingApr 7, 2009
  54. Björn SteinbrinkApr 7, 2009
  55. process_{tree,blob}: Remove useless xstrdup callsBjörn Steinbrink, Apr 8, 2009
  56. Linus TorvaldsApr 10, 2009
  57. Linus TorvaldsApr 11, 2009
  58. Linus TorvaldsApr 11, 2009
  59. Nicolas PitreApr 11, 2009
  60. Björn SteinbrinkApr 11, 2009
  61. Björn SteinbrinkApr 11, 2009
  62. Linus TorvaldsApr 11, 2009
  63. Linus TorvaldsApr 11, 2009
  64. Björn SteinbrinkApr 11, 2009
  65. Björn SteinbrinkApr 11, 2009
  66. Linus TorvaldsApr 11, 2009
  67. Björn SteinbrinkApr 11, 2009
  68. Linus TorvaldsApr 11, 2009
  69. Björn SteinbrinkApr 11, 2009
  70. Linus TorvaldsApr 11, 2009
  71. Nicolas SebrechtApr 5, 2009
  72. david@lang.hmApr 5, 2009
  73. Robin RosenbergApr 5, 2009
  74. Nicolas PitreApr 6, 2009
  75. Junio C HamanoApr 6, 2009
  76. Nicolas PitreApr 6, 2009
  77. Jon SmirlApr 6, 2009
  78. Nicolas PitreApr 6, 2009
  79. Jon SmirlApr 6, 2009
  80. Shawn O. PearceApr 6, 2009
  81. Nicolas PitreApr 6, 2009
  82. Jon SmirlApr 6, 2009
  83. Nicolas PitreApr 6, 2009
  84. Matthieu MoyApr 6, 2009
  85. Nicolas PitreApr 6, 2009
  86. Robin H. JohnsonApr 6, 2009
  87. Nicolas PitreApr 6, 2009
  88. Martin LanghoffApr 7, 2009
  89. Jeff KingApr 5, 2009
  90. Robin H. JohnsonApr 5, 2009
  91. Robin H. JohnsonApr 5, 2009
  92. Nguyen Thai Ngoc DuyApr 6, 2009
  93. Nicolas PitreApr 6, 2009
  94. Nicolas PitreApr 6, 2009
  95. Robin H. JohnsonApr 6, 2009
  96. Mark LevedahlApr 11, 2009
  97. Robin H. JohnsonApr 6, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.