git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: pack operation is thrashing my server

From
Nicolas Pitre <nico@cam.org>
Date
Aug 13, 2008, 17:04 UTC
Message-ID
<alpine.LFD.1.10.0808131228270.4352@xanadu.home>
In-Reply-To
<20080813155016.GD3782@spearce.org>
On Wed, 13 Aug 2008, Shawn O. Pearce wrote:
Show 6 quoted lines
> Nicolas Pitre <nico@cam.org> wrote:
> > Well, we are talking about 50MB which is not that bad.
> 
> I think we're closer to 100MB here due to the extra overheads
> I just alluded to above, and which weren't in your 104 byte
> per object figure.
Sure.  That should still be workable on a machine with 256MB of RAM.
Show 12 quoted lines
> > However there is a point where we should be realistic and just admit 
> > that you need a sufficiently big machine if you have huge repositories 
> > to deal with.  Git should be fine serving pull requests with relatively 
> > little memory usage, but anything else such as the initial repack simply 
> > require enough RAM to be effective.
> 
> Yea.  But it would also be nice to be able to just concat packs
> together.  Especially if the repository in question is an open source
> one and everything published is already known to be in the wild,
> as say it is also available over dumb HTTP.  Yea, I know people
> like the 'security feature' of the packer not including objects
> which aren't reachable.

It is not only that, even if it is a point I consider important. If you end up with 10 packs, it is likely that a base object in each of those packs could simply be a delta against a single common base object, and therefore the amount of data to transfer might be up to 10 times higher than necessary.

> But how many times has Linus published something to his linux-2.6
> tree that he didn't mean to publish and had to rewind?  I think
> that may be "never".  Yet how many times per day does his tree get
> cloned from scratch?

That's not a good argument. Linus is a very disciplined git users, probably more than average. We should not use that example to paper over technical issues.

> This is also true for many internal corporate repositories.
> Users probably have full read access to the object database anyway,
> and maybe even have direct write access to it.  Doing the object
> enumeration there is pointless as a security measure.
It is good for network bandwidth efficiency as I mentioned.
> I'm too busy to write a pack concat implementation proposal, so
> I'll just shutup now.  But it wouldn't be hard if someone wanted
> to improve at least the initial clone serving case.

A much better solution would consist of finding just _why_ object enumeration is so slow. This is indeed my biggest grip with git performance at the moment.

|nico@xanadu:linux-2.6> time git rev-list --objects --all > /dev/null
|
|real    0m21.742s
|user    0m21.379s
|sys     0m0.360s

That's way too long for 1030198 objects (roughly 48k objects/sec). And it gets even worse with the gcc repository:

|nico@xanadu:gcc> time git rev-list --objects --all > /dev/null
|
|real    1m51.591s
|user    1m50.757s
|sys     0m0.810s
That's for 1267993 objects, or about 11400 objects/sec.
Clearly something is not scaling here.
Nicolas
Previous: Shawn O. PearceNext: Shawn O. Pearce
Message 37 of 80 in “pack operation is thrashing my server”
  1. Ken PrattAug 10, 2008
  2. Martin LanghoffAug 10, 2008
  3. Ken PrattAug 10, 2008
  4. Martin LanghoffAug 10, 2008
  5. Ken PrattAug 10, 2008
  6. Shawn O. PearceAug 11, 2008
  7. Ken PrattAug 11, 2008
  8. Shawn O. PearceAug 11, 2008
  9. Avery PennarunAug 11, 2008
  10. Shawn O. PearceAug 11, 2008
  11. Ken PrattAug 11, 2008
  12. Andi KleenAug 11, 2008
  13. Ken PrattAug 11, 2008
  14. Nicolas PitreAug 13, 2008
  15. Andi KleenAug 13, 2008
  16. Shawn O. PearceAug 13, 2008
  17. Shawn O. PearceAug 11, 2008
  18. Ken PrattAug 11, 2008
  19. Shawn O. PearceAug 11, 2008
  20. Andi KleenAug 11, 2008
  21. Geert BoschAug 13, 2008
  22. Shawn O. PearceAug 13, 2008
  23. Geert BoschAug 13, 2008
  24. Nicolas PitreAug 13, 2008
  25. Jakub NarebskiAug 13, 2008
  26. Shawn O. PearceAug 13, 2008
  27. David TweedAug 13, 2008
  28. Martin LanghoffAug 13, 2008
  29. David TweedAug 14, 2008
  30. Johan HerlandAug 13, 2008
  31. Ken PrattAug 13, 2008
  32. Nicolas PitreAug 13, 2008
  33. Nicolas PitreAug 13, 2008
  34. Shawn O. PearceAug 13, 2008
  35. Nicolas PitreAug 13, 2008
  36. Shawn O. PearceAug 13, 2008
  37. Nicolas PitreAug 13, 2008
  38. Shawn O. PearceAug 13, 2008
  39. Andreas EricssonAug 14, 2008
  40. Thomas RastAug 14, 2008
  41. Andreas EricssonAug 14, 2008
  42. Shawn O. PearceAug 14, 2008
  43. Nicolas PitreAug 15, 2008
  44. Nicolas PitreAug 14, 2008
  45. Linus TorvaldsAug 14, 2008
  46. Linus TorvaldsAug 14, 2008
  47. Nicolas PitreAug 14, 2008
  48. Linus TorvaldsAug 14, 2008
  49. Andi KleenAug 14, 2008
  50. Linus TorvaldsAug 15, 2008
  51. Nicolas PitreAug 14, 2008
  52. Linus TorvaldsAug 14, 2008
  53. Björn SteinbrinkAug 14, 2008
  54. Linus TorvaldsAug 15, 2008
  55. Linus TorvaldsAug 15, 2008
  56. Björn SteinbrinkAug 16, 2008
  57. Linus TorvaldsAug 16, 2008
  58. Junio C HamanoSep 7, 2008
  59. Linus TorvaldsSep 7, 2008
  60. Junio C HamanoSep 7, 2008
  61. Nicolas PitreSep 7, 2008
  62. Junio C HamanoSep 7, 2008
  63. Jon SmirlSep 7, 2008
  64. Linus TorvaldsSep 7, 2008
  65. Jon SmirlSep 7, 2008
  66. Linus TorvaldsSep 7, 2008
  67. Jon SmirlSep 7, 2008
  68. Nicolas PitreSep 7, 2008
  69. Jon SmirlSep 7, 2008
  70. Nicolas PitreSep 8, 2008
  71. Jon SmirlSep 8, 2008
  72. Jon SmirlSep 8, 2008
  73. Andreas EricssonSep 7, 2008
  74. Mike HommeySep 7, 2008
  75. Nicolas PitreAug 14, 2008
  76. Linus TorvaldsAug 14, 2008
  77. Geert BoschAug 13, 2008
  78. Dana HowAug 13, 2008
  79. Nicolas PitreAug 13, 2008
  80. Jakub NarebskiAug 13, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.