git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: pack operation is thrashing my server

From
Nicolas Pitre <nico@cam.org>
Date
Aug 13, 2008, 17:26 UTC
Message-ID
<alpine.LFD.1.10.0808131304560.4352@xanadu.home>
In-Reply-To
<3E057C8D-FF72-47A2-BBA8-27A22AD67167@adacore.com>
On Wed, 13 Aug 2008, Geert Bosch wrote:
Show 13 quoted lines
> On Aug 13, 2008, at 10:35, Nicolas Pitre wrote:
> > On Tue, 12 Aug 2008, Geert Bosch wrote:
> > 
> > > I've always felt that keeping largish objects (say anything >1MB)
> > > loose makes perfect sense. These objects are accessed infrequently,
> > > often binary or otherwise poor candidates for the delta algorithm.
> > 
> > Or, as I suggested in the past, they can be grouped into a separate
> > pack, or even occupy a pack of their own.
> 
> This is fine, as long as we're not trying to create deltas
> of the large objects, or do other things that requires keeping
> the inflated data in memory.
First, there is the delta attribute:
|commit a74db82e15cd8a2c53a4a83e9a36dc7bf7a4c750
|Author: Junio C Hamano <junkio@cox.net>
|Date:   Sat May 19 00:39:31 2007 -0700
|
|    Teach "delta" attribute to pack-objects.
|
|    This teaches pack-objects to use .gitattributes mechanism so
|    that the user can specify certain blobs are not worth spending
|    CPU cycles to attempt deltification.
|
|    The name of the attrbute is "delta", and when it is set to
|    false, like this:
|
|        == .gitattributes ==
|        *.jpg   -delta
|
|    they are always stored in the plain-compressed base object
|    representation.
This could probably be extended to take a size limit argument as well.
Show 11 quoted lines
> > As soon as you have more than
> > one revision of such largish objects then you lose again by keeping them
> > loose.
> 
> Yes, you lose potentially in terms of disk space, but you avoid the
> large memory footprint during pack generation. For very large blobs,
> it is best to degenerate to having each revision of each file on
> its own (whether we call it a single-file pack, loose object or whatever).
> That way, the large file can stay immutable on disk, and will only
> need to be accessed during checkout. GIT will then scale with good
> performance until we run out of disk space.

Loose objects, though, will always be selected for potential delta generation. Packed objects, deltified or not, are always streamed as is when serving pull requests. And by default delta compression is not (re)attempted between objects which are part of the same pack, the reason being that if they were not deltified on the first packing attempt then there is no point trying again when streaming them over the net. So you always benefit from having your large objects packed with the rest. This, plus the delta prevention mechanism above should cover most cases.

Show 5 quoted lines
> > You'll have memory usage issues whenever such objects are accessed,
> > loose or not.
> Why? The only time we'd need to access their contents for checkout
> or when pushing across the network. These should all be steaming operations
> with small memory footprint.

Pushing across the network, or repacking without -f, is streamed. Checking out currently isn't (although it probably could). Repacking with -f definitely isn't and probably shouldn't because of complexity issues.

Show 11 quoted lines
> > However, once those big objects are packed once, they can
> > be repacked (or streamed over the net) without really "accessing" them.
> > Packed object data is simply copied into a new pack in that case which
> > is less of an issue on memory usage, irrespective of the original pack
> > size.
> Agreed, but still, at least very large objects. If I have a 600MB
> file in my repository, it should just not get in the way. If it gets
> copied around during each repack, that just wastes I/O time for no
> good reason. Even worse, it causes incremental backups or filesystem
> checkpoints to become way more expensive. Just leaving large files
> alone as immutable objects on disk avoids all these issues.

Pack them in a pack of their own and stick a .keep file along with it. At that point they will never be rewritten.

Nicolas
Previous: Dana HowNext: Jakub Narebski
Message 79 of 80 in “pack operation is thrashing my server”
  1. Ken PrattAug 10, 2008
  2. Martin LanghoffAug 10, 2008
  3. Ken PrattAug 10, 2008
  4. Martin LanghoffAug 10, 2008
  5. Ken PrattAug 10, 2008
  6. Shawn O. PearceAug 11, 2008
  7. Ken PrattAug 11, 2008
  8. Shawn O. PearceAug 11, 2008
  9. Avery PennarunAug 11, 2008
  10. Shawn O. PearceAug 11, 2008
  11. Ken PrattAug 11, 2008
  12. Andi KleenAug 11, 2008
  13. Ken PrattAug 11, 2008
  14. Nicolas PitreAug 13, 2008
  15. Andi KleenAug 13, 2008
  16. Shawn O. PearceAug 13, 2008
  17. Shawn O. PearceAug 11, 2008
  18. Ken PrattAug 11, 2008
  19. Shawn O. PearceAug 11, 2008
  20. Andi KleenAug 11, 2008
  21. Geert BoschAug 13, 2008
  22. Shawn O. PearceAug 13, 2008
  23. Geert BoschAug 13, 2008
  24. Nicolas PitreAug 13, 2008
  25. Jakub NarebskiAug 13, 2008
  26. Shawn O. PearceAug 13, 2008
  27. David TweedAug 13, 2008
  28. Martin LanghoffAug 13, 2008
  29. David TweedAug 14, 2008
  30. Johan HerlandAug 13, 2008
  31. Ken PrattAug 13, 2008
  32. Nicolas PitreAug 13, 2008
  33. Nicolas PitreAug 13, 2008
  34. Shawn O. PearceAug 13, 2008
  35. Nicolas PitreAug 13, 2008
  36. Shawn O. PearceAug 13, 2008
  37. Nicolas PitreAug 13, 2008
  38. Shawn O. PearceAug 13, 2008
  39. Andreas EricssonAug 14, 2008
  40. Thomas RastAug 14, 2008
  41. Andreas EricssonAug 14, 2008
  42. Shawn O. PearceAug 14, 2008
  43. Nicolas PitreAug 15, 2008
  44. Nicolas PitreAug 14, 2008
  45. Linus TorvaldsAug 14, 2008
  46. Linus TorvaldsAug 14, 2008
  47. Nicolas PitreAug 14, 2008
  48. Linus TorvaldsAug 14, 2008
  49. Andi KleenAug 14, 2008
  50. Linus TorvaldsAug 15, 2008
  51. Nicolas PitreAug 14, 2008
  52. Linus TorvaldsAug 14, 2008
  53. Björn SteinbrinkAug 14, 2008
  54. Linus TorvaldsAug 15, 2008
  55. Linus TorvaldsAug 15, 2008
  56. Björn SteinbrinkAug 16, 2008
  57. Linus TorvaldsAug 16, 2008
  58. Junio C HamanoSep 7, 2008
  59. Linus TorvaldsSep 7, 2008
  60. Junio C HamanoSep 7, 2008
  61. Nicolas PitreSep 7, 2008
  62. Junio C HamanoSep 7, 2008
  63. Jon SmirlSep 7, 2008
  64. Linus TorvaldsSep 7, 2008
  65. Jon SmirlSep 7, 2008
  66. Linus TorvaldsSep 7, 2008
  67. Jon SmirlSep 7, 2008
  68. Nicolas PitreSep 7, 2008
  69. Jon SmirlSep 7, 2008
  70. Nicolas PitreSep 8, 2008
  71. Jon SmirlSep 8, 2008
  72. Jon SmirlSep 8, 2008
  73. Andreas EricssonSep 7, 2008
  74. Mike HommeySep 7, 2008
  75. Nicolas PitreAug 14, 2008
  76. Linus TorvaldsAug 14, 2008
  77. Geert BoschAug 13, 2008
  78. Dana HowAug 13, 2008
  79. Nicolas PitreAug 13, 2008
  80. Jakub NarebskiAug 13, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.