git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Figured out how to get Mozilla into git

From
Linus Torvalds <torvalds@osdl.org>
Date
Jun 9, 2006, 15:01 UTC
Message-ID
<Pine.LNX.4.64.0606090745390.5498@g5.osdl.org>
In-Reply-To
<e6b798$td3$1@sea.gmane.org>
On Fri, 9 Jun 2006, Jakub Narebski wrote:
Show 9 quoted lines
> Jon Smirl wrote:
> 
> >> git-repack -a -d but it OOMs on my 2GB+2GBswap machine :(
> > 
> > We are all having problems getting this to run on 32 bit machines with
> > the 3-4GB process size limitations.
> 
> Is that expected (for 10GB repository if I remember correctly), or is there
> some way to avoid this OOM?
Well, to some degree, the VM limitations are inevitable with huge packs.

The original idea for packs was to avoid making one huge pack, partly because it was expected to be really really slow to generate (so incremental repacking was a much better strategy), but partly simply because trying to map one huge pack is really hard to do.

For various reasons, we ended up mostly using a single pack most of the time: it's the most efficient model when the project is reasonably sized, and it turns out that with the delta re-use, repacking even moderately large projects like the kernel doesn't actually take all that long.

But the fact that we ended up mostly using a single pack for the kernel, for example, doesn't mean that the fundamental reasons that git supports multiple packs would somehow have gone away. At some point, the project gets large enough that one single pack simply isn't reasonable.

So a single 2GB pack is already very much pushing it. It's really really hard to map in a 2GB file on a 32-bit platform: your VM is usually fragmented enough that it simply isn't practical. In fact, I think the limit for _practical_ usage of single packs is probably somewhere in the half-gig region, unless you just have 64-bit machines.

And yes, I realize that the "single pack" thing actually ends up having become a fact for cloning, for example. Originally, cloning would unpack on the receiving end, and leave the repacking to happen there, but that obviously sucked. So now when we clone, we always get a single pack. That can absolutely be a problem.

I don't know what the right solution is. Single packs _are_ very useful, especially after a clone. So it's possible that we should just make the pack-reading code be able to map partial packs. But the point is that there are certainly ways we can fix this - it's not _really_ fundamental.

It's going to complicate it a bit (damn, how I hate 32-bit VM limitations), but the good news is that the whole git model of "everything is an individual object" means that it's a very _local_ decision: it will probably be painful to re-do some of the pack reading code and have a LRU of pack _fragments_ instead of a LRU of packs, but it's only going to affect a small part of git, and everything else will never even see it.

So large packs are not really a fundamental problem, but right now we have some practical issues with them.

(It's not _just_ packs: running out of memory is also because of git-rev-list --objects being pretty memory hungry. I've improved the memory usage several times by over 50%, but people keep trying larger projects. It used to be that I considered the kernel a large history, now we're talking about things that have ten times the number of objects).

Martin - do you have some place to make that big mozilla repo available? It would be a good test-case..

			Linus
Previous: Jakub NarebskiNext: Nicolas Pitre
Message 6 of 67 in “Figured out how to get Mozilla into git”
  1. Jon SmirlJun 9, 2006
  2. Nicolas PitreJun 9, 2006
  3. Martin LanghoffJun 9, 2006
  4. Jon SmirlJun 9, 2006
  5. Jakub NarebskiJun 9, 2006
  6. Linus TorvaldsJun 9, 2006
  7. Nicolas PitreJun 9, 2006
  8. Linus TorvaldsJun 9, 2006
  9. Nicolas PitreJun 9, 2006
  10. Linus TorvaldsJun 9, 2006
  11. Jakub NarebskiJun 9, 2006
  12. Jon SmirlJun 9, 2006
  13. Linus TorvaldsJun 9, 2006
  14. Jon SmirlJun 9, 2006
  15. Linus TorvaldsJun 9, 2006
  16. Jon SmirlJun 9, 2006
  17. Linus TorvaldsJun 9, 2006
  18. Linus TorvaldsJun 9, 2006
  19. Greg KHJun 9, 2006
  20. Martin LanghoffJun 9, 2006
  21. Linus TorvaldsJun 9, 2006
  22. Jon SmirlJun 10, 2006
  23. Linus TorvaldsJun 10, 2006
  24. Jon SmirlJun 10, 2006
  25. Jon SmirlJun 10, 2006
  26. Jakub NarebskiJun 9, 2006
  27. Nicolas PitreJun 9, 2006
  28. Jon SmirlJun 9, 2006
  29. Martin LanghoffJun 10, 2006
  30. Martin LanghoffJun 10, 2006
  31. Linus TorvaldsJun 10, 2006
  32. Linus TorvaldsJun 10, 2006
  33. Jon SmirlJun 10, 2006
  34. Linus TorvaldsJun 10, 2006
  35. Jon SmirlJun 10, 2006
  36. Carl WorthJun 10, 2006
  37. Linus TorvaldsJun 10, 2006
  38. Jakub NarebskiJun 10, 2006
  39. Junio C HamanoJun 10, 2006
  40. Rogan DawesJun 10, 2006
  41. Junio C HamanoJun 10, 2006
  42. Rogan DawesJun 10, 2006
  43. Jakub NarebskiJun 10, 2006
  44. Nicolas PitreJun 10, 2006
  45. Linus TorvaldsJun 10, 2006
  46. Jon SmirlJun 10, 2006
  47. Rogan DawesJun 10, 2006
  48. Linus TorvaldsJun 10, 2006
  49. Jon SmirlJun 10, 2006
  50. Martin LanghoffJun 10, 2006
  51. Junio C HamanoJun 10, 2006
  52. Linus TorvaldsJun 10, 2006
  53. Linus TorvaldsJun 10, 2006
  54. Jon SmirlJun 10, 2006
  55. Junio C HamanoJun 10, 2006
  56. Jon SmirlJun 10, 2006
  57. Timo HirvonenJun 10, 2006
  58. Petr BaudisJun 10, 2006
  59. Lars JohannsenJun 10, 2006
  60. Nicolas PitreJun 11, 2006
  61. Linus TorvaldsJun 18, 2006
  62. Martin LanghoffJun 18, 2006
  63. Linus TorvaldsJun 18, 2006
  64. Broken PPC sha1.. (Re: Figured out how to get Mozilla into git)Linus Torvalds, Jun 18, 2006
  65. Fix PPC SHA1 routine for large input buffersPaul Mackerras, Jun 18, 2006
  66. Linus TorvaldsJun 19, 2006
  67. Pavel RoskinJun 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.