git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Something is broken in repack

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Dec 11, 2007, 19:17 UTC
Message-ID
<alpine.LFD.0.9999.0712111055590.25032@woody.linux-foundation.org>
In-Reply-To
<9e4733910712111043h6a361996x740f4dba3d742da5@mail.gmail.com>
On Tue, 11 Dec 2007, Jon Smirl wrote:
Show 12 quoted lines
> >
> > So if you want to use more threads, that _forces_ you to have a bigger
> > memory footprint, simply because you have more "live" objects that you
> > work on. Normally, that isn't much of a problem, since most source files
> > are small, but if you have a few deep delta chains on big files, both the
> > delta chain itself is going to use memory (you may have limited the size
> > of the cache, but it's still needed for the actual delta generation, so
> > it's not like the memory usage went away).
> 
> This makes sense. Those runs that blew up to 4.5GB were a combination
> of this effect and fragmentation in the gcc allocator. Google
> allocator appears to be much better at controlling fragmentation.

Yes. I think we do have some case where we simply keep a lot of objects around, and if we are talking reasonably large deltas, we'll have the whole delta-chain in memory just to unpack one single object.

The delta cache size limits kick in only when we explicitly cache old delta results (in case they will be re-used, which is rather common), it doesn't affect the normal "I'm using this data right now" case at all.

And then fragmentation makes it much much worse. Since the allocation patterns aren't nice (they are pretty random and depend on just the sizes of the objects), and the lifetimes aren't always nicely nested _either_ (they become more so when you disable the cache entirely, but that's just death for performance), I'm not surprised that there can be memory allocators that end up having some issues.

> Is there a reasonable scheme to force the chains to only be loaded
> once and then shared between worker threads? The memory blow up
> appears to be directly correlated with chain length.

The worker threads explicitly avoid touching the same objects, and no, you definitely don't want to explode the chains globally once, because the whole point is that we do fit 15 years worth of history into 300MB of pack-file thanks to having a very dense representation. The "loaded once" part is the mmap'ing of the pack-file into memory, but if you were to actually then try to expand the chains, you'd be talking about many *many* more gigabytes of memory than you already see used ;)

So what you actually want to do is to just re-use already packed delta chains directly, which is what we normally do. But you are explicitly looking at the "--no-reuse-delta" (aka "git repack -f") case, which is why it then blows up.

I'm sure we can find places to improve. But I would like to re-iterate the statement that you're kind of doing a "don't do that then" case which is really - by design - meant to be done once and never again, and is using resources - again, pretty much by design - wildly inappropriately just to get an initial packing done.

> That may account for the threaded version needing an extra 20 minutes
> CPU time.  An extra 12% of CPU seems like too much overhead for
> threading. Just letting a couple of those long chain compressions be
> done twice

Well, Nico pointed out that those things should all be thread-private data, so no, the race isn't there (unless there's some other bug there).

> I agree, this problem only occurs when people import giant
> repositories. But every time someone hits these problems they declare
> git to be screwed up and proceed to thrash it in their blogs.

Sure. I'd love to do global packing without paying the cost, but it really was a design decision. Thanks to doing off-line packing ("let it run overnight on some beefy machine") we can get better results. It's expensive, yes. But it was pretty much meant to be expensive. It's a very efficient compression algorithm, after all, and you're turning it up to eleven ;)

I also suspect that the gcc archive makes things more interesting thanks to having some rather large files. The ChangeLog is probably the worst case (large file with *lots* of edits), but I suspect the *.po files aren't wonderful either.

			Linus
Previous: Nicolas PitreNext: Junio C Hamano
Message 76 of 82 in “Something is broken in repack”
  1. Jon SmirlDec 7, 2007
  2. Linus TorvaldsDec 8, 2007
  3. pack-objects: fix delta cache size accountingNicolas Pitre, Dec 8, 2007
  4. Nicolas PitreDec 8, 2007
  5. Jon SmirlDec 8, 2007
  6. Nicolas PitreDec 8, 2007
  7. Jon SmirlDec 8, 2007
  8. David BrownDec 8, 2007
  9. Jon SmirlDec 8, 2007
  10. Nicolas PitreDec 8, 2007
  11. Jon SmirlDec 8, 2007
  12. Nicolas PitreDec 8, 2007
  13. Harvey HarrisonDec 8, 2007
  14. Jon SmirlDec 8, 2007
  15. Harvey HarrisonDec 8, 2007
  16. Junio C HamanoDec 8, 2007
  17. Junio C HamanoDec 9, 2007
  18. Jon SmirlDec 9, 2007
  19. Jon SmirlDec 9, 2007
  20. Nicolas PitreDec 10, 2007
  21. Nicolas PitreDec 10, 2007
  22. David BrownDec 8, 2007
  23. Nicolas PitreDec 10, 2007
  24. Jon SmirlDec 10, 2007
  25. Morten WelinderDec 10, 2007
  26. Jon SmirlDec 11, 2007
  27. Junio C HamanoDec 11, 2007
  28. Nicolas PitreDec 11, 2007
  29. David KastrupDec 11, 2007
  30. Pierre HabouzitDec 11, 2007
  31. David KastrupDec 11, 2007
  32. Nicolas PitreDec 11, 2007
  33. Jon SmirlDec 11, 2007
  34. Jon SmirlDec 11, 2007
  35. Jon SmirlDec 11, 2007
  36. Andreas EricssonDec 11, 2007
  37. Nicolas PitreDec 11, 2007
  38. Nicolas PitreDec 11, 2007
  39. Jon SmirlDec 11, 2007
  40. Nicolas PitreDec 11, 2007
  41. Jon SmirlDec 11, 2007
  42. Nicolas PitreDec 12, 2007
  43. David KastrupDec 12, 2007
  44. Wolfram GlogerDec 14, 2007
  45. Nicolas PitreDec 12, 2007
  46. Paolo BonziniDec 12, 2007
  47. Linus TorvaldsDec 12, 2007
  48. David MillerDec 12, 2007
  49. Linus TorvaldsDec 12, 2007
  50. Jon SmirlDec 12, 2007
  51. Wolfram GlogerDec 14, 2007
  52. David KastrupDec 14, 2007
  53. Wolfram GlogerDec 14, 2007
  54. Nguyen Thai Ngoc DuyDec 13, 2007
  55. Paolo BonziniDec 13, 2007
  56. Paolo BonziniDec 13, 2007
  57. Johannes SixtDec 13, 2007
  58. Jakub NarebskiDec 14, 2007
  59. Paolo BonziniDec 14, 2007
  60. Nguyen Thai Ngoc DuyDec 14, 2007
  61. Paolo BonziniDec 14, 2007
  62. Harvey HarrisonDec 14, 2007
  63. Jakub NarebskiDec 14, 2007
  64. Nguyen Thai Ngoc DuyDec 14, 2007
  65. Nicolas PitreDec 14, 2007
  66. Nicolas PitreDec 12, 2007
  67. Andreas EricssonDec 13, 2007
  68. Wolfram GlogerDec 14, 2007
  69. Linus TorvaldsDec 11, 2007
  70. Nicolas PitreDec 11, 2007
  71. David MillerDec 11, 2007
  72. Nicolas PitreDec 11, 2007
  73. Andreas EricssonDec 11, 2007
  74. Jon SmirlDec 11, 2007
  75. Nicolas PitreDec 11, 2007
  76. Linus TorvaldsDec 11, 2007
  77. Junio C HamanoDec 11, 2007
  78. Andreas EricssonDec 11, 2007
  79. Daniel BerlinDec 11, 2007
  80. Nicolas PitreDec 11, 2007
  81. SeanDec 11, 2007
  82. Jon SmirlDec 11, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.