git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: pack operation is thrashing my server

From
Shawn O. Pearce <spearce@spearce.org>
Date
Aug 13, 2008, 15:50 UTC
Message-ID
<20080813155016.GD3782@spearce.org>
In-Reply-To
<alpine.LFD.1.10.0808131123221.4352@xanadu.home>
Nicolas Pitre <nico@cam.org> wrote:
Show 7 quoted lines
> On Wed, 13 Aug 2008, Shawn O. Pearce wrote:
> > 
> > Where little memory systems get into trouble with already packed
> > repositories is enumerating the objects to include in the pack.
> 
> I'm counting something like 104 bytes on a 64-bit machine for
> struct object_entry.
Don't forget that we need not just struct object_entry, but
also the struct commit/tree/blob, their hash tables, and the
struct object_entry* in the sorted object list table, and
the pack reverse index table.  It does add up.
 
> > Have 500k objects and its suddenly something quite real in terms
> > of memory usage.
> 
> Well, we are talking about 50MB which is not that bad.

I think we're closer to 100MB here due to the extra overheads I just alluded to above, and which weren't in your 104 byte per object figure.

Show 5 quoted lines
> However there is a point where we should be realistic and just admit 
> that you need a sufficiently big machine if you have huge repositories 
> to deal with.  Git should be fine serving pull requests with relatively 
> little memory usage, but anything else such as the initial repack simply 
> require enough RAM to be effective.

Yea. But it would also be nice to be able to just concat packs together. Especially if the repository in question is an open source one and everything published is already known to be in the wild, as say it is also available over dumb HTTP. Yea, I know people like the 'security feature' of the packer not including objects which aren't reachable.

But how many times has Linus published something to his linux-2.6 tree that he didn't mean to publish and had to rewind? I think that may be "never". Yet how many times per day does his tree get cloned from scratch?

This is also true for many internal corporate repositories. Users probably have full read access to the object database anyway, and maybe even have direct write access to it. Doing the object enumeration there is pointless as a security measure.

I'm too busy to write a pack concat implementation proposal, so I'll just shutup now. But it wouldn't be hard if someone wanted to improve at least the initial clone serving case.

-- 
Shawn.
Previous: Nicolas PitreNext: Nicolas Pitre
Message 36 of 80 in “pack operation is thrashing my server”
  1. Ken PrattAug 10, 2008
  2. Martin LanghoffAug 10, 2008
  3. Ken PrattAug 10, 2008
  4. Martin LanghoffAug 10, 2008
  5. Ken PrattAug 10, 2008
  6. Shawn O. PearceAug 11, 2008
  7. Ken PrattAug 11, 2008
  8. Shawn O. PearceAug 11, 2008
  9. Avery PennarunAug 11, 2008
  10. Shawn O. PearceAug 11, 2008
  11. Ken PrattAug 11, 2008
  12. Andi KleenAug 11, 2008
  13. Ken PrattAug 11, 2008
  14. Nicolas PitreAug 13, 2008
  15. Andi KleenAug 13, 2008
  16. Shawn O. PearceAug 13, 2008
  17. Shawn O. PearceAug 11, 2008
  18. Ken PrattAug 11, 2008
  19. Shawn O. PearceAug 11, 2008
  20. Andi KleenAug 11, 2008
  21. Geert BoschAug 13, 2008
  22. Shawn O. PearceAug 13, 2008
  23. Geert BoschAug 13, 2008
  24. Nicolas PitreAug 13, 2008
  25. Jakub NarebskiAug 13, 2008
  26. Shawn O. PearceAug 13, 2008
  27. David TweedAug 13, 2008
  28. Martin LanghoffAug 13, 2008
  29. David TweedAug 14, 2008
  30. Johan HerlandAug 13, 2008
  31. Ken PrattAug 13, 2008
  32. Nicolas PitreAug 13, 2008
  33. Nicolas PitreAug 13, 2008
  34. Shawn O. PearceAug 13, 2008
  35. Nicolas PitreAug 13, 2008
  36. Shawn O. PearceAug 13, 2008
  37. Nicolas PitreAug 13, 2008
  38. Shawn O. PearceAug 13, 2008
  39. Andreas EricssonAug 14, 2008
  40. Thomas RastAug 14, 2008
  41. Andreas EricssonAug 14, 2008
  42. Shawn O. PearceAug 14, 2008
  43. Nicolas PitreAug 15, 2008
  44. Nicolas PitreAug 14, 2008
  45. Linus TorvaldsAug 14, 2008
  46. Linus TorvaldsAug 14, 2008
  47. Nicolas PitreAug 14, 2008
  48. Linus TorvaldsAug 14, 2008
  49. Andi KleenAug 14, 2008
  50. Linus TorvaldsAug 15, 2008
  51. Nicolas PitreAug 14, 2008
  52. Linus TorvaldsAug 14, 2008
  53. Björn SteinbrinkAug 14, 2008
  54. Linus TorvaldsAug 15, 2008
  55. Linus TorvaldsAug 15, 2008
  56. Björn SteinbrinkAug 16, 2008
  57. Linus TorvaldsAug 16, 2008
  58. Junio C HamanoSep 7, 2008
  59. Linus TorvaldsSep 7, 2008
  60. Junio C HamanoSep 7, 2008
  61. Nicolas PitreSep 7, 2008
  62. Junio C HamanoSep 7, 2008
  63. Jon SmirlSep 7, 2008
  64. Linus TorvaldsSep 7, 2008
  65. Jon SmirlSep 7, 2008
  66. Linus TorvaldsSep 7, 2008
  67. Jon SmirlSep 7, 2008
  68. Nicolas PitreSep 7, 2008
  69. Jon SmirlSep 7, 2008
  70. Nicolas PitreSep 8, 2008
  71. Jon SmirlSep 8, 2008
  72. Jon SmirlSep 8, 2008
  73. Andreas EricssonSep 7, 2008
  74. Mike HommeySep 7, 2008
  75. Nicolas PitreAug 14, 2008
  76. Linus TorvaldsAug 14, 2008
  77. Geert BoschAug 13, 2008
  78. Dana HowAug 13, 2008
  79. Nicolas PitreAug 13, 2008
  80. Jakub NarebskiAug 13, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.