git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC] Submodules in GIT

From
Linus Torvalds <torvalds@osdl.org>
Date
Dec 2, 2006, 21:22 UTC
Message-ID
<Pine.LNX.4.64.0612021252380.3476@woody.osdl.org>
In-Reply-To
<Pine.LNX.4.64.0612021242080.3476@woody.osdl.org>
On Sat, 2 Dec 2006, Linus Torvalds wrote:
> 
> And watch the memory usage.

Btw, just in case you don't understand _why_ this is true, the fact is, in a git repository, quite fundamentally, because we don't have "backlinks" at any stage at all, we don't know - and fundamentally _cannot_ know - whether we're goign to see the same object in the future.

So operations like "git-rev-list --objects" (or, these days, more commonly anything that just does the equivalent of that internally using the library interfaces - ie "git pack-objects" and friends) VERY FUNDAMENTALLY have to hold on to the object flags for the whole lifetime of the whole operation.

And you should realize that this is really really fundamental. You can't fix it with "smarter memory management". You can't fix it with "garbage collection". This is _not_ a result of the fact that we use C and malloc, and we don't free those objects, like some people sometimes seem to believe.

So garbage collection will never help this kind of situation. It flows _directly_ from the fact that our objects are immutable: because they are immutable, they don't have any backpointers, because we cannot (and must not) add backpointers to an old existing object when a new object is created that points to it.

So this really isn't a memory management issue. You could somewhat work around it by adding a "caching layer" on top of git, and allow that caching layer to modify their cache of old objects (so that they can contain back-pointers), but for 99% of all users that would actually make performance MUCH WORSE, and it would also be a serious problem for coherency issues (one of the things that immutable objects cause is that there are basically never any race conditions, while a "caching layer" like this would have some serious issues about serialization).

So: the very fundamental nature and choices that were made in git also 
means that when you have something like git-pack-objects that wants to 
walk the whole repo, you will end up with something that remembers EVERY 
SINGLE OBJECT it walked. 

And while I've worked very hard to make the memory footprint of individual objects as small as possible, and this means that this all works fine even for fairly large databases (especially since very few operations actually do this "traverse the whole friggin tree" thing), it does mean that there's a very fundamental limit to scalability. You can't just make a whole repository a hundred times bigger - because the operations that traverse the whole thing will require a hundred times more memory!

Now, in "real" projects, this is not a problem. I can pretty much _guarantee_ that memory sizes and hardware will grow faster than projects grow. I'm not AT ALL worried about the fact that in ten years, the linux kernel repository will likely be two or three times the size it is now. Because I'm absolutely convinced that in ten years, the machines we have now will be obsolete.

So on any "individual project" basis, the fact that memory requirements scale roughly as O(n) in the total repository size is simply not a problem. In fact, O(n) is pretty damn good, especially since the constant is pretty small (basically 28 bytes per object - and 20 of those bytes are the SHA1 that you simply cannot avoid).

But it does mean that supermodules really should NOT be so seamless that doing a "git clone" on a supermodule does one _large_ clone. Because it's simply going to be better to:

 - when you clone the supermodule, track the commits you need on all 
   submodules (this _may_ be a reason in itself for the "link" object, 
   just so that you can traverse the supermodule object dependencies and 
   know what subobject you are looking at even _without_ having to look at 
   the path you got there from)
 - clone submodules one-by-one, using the list of objects you gathered.

Maybe there are other solutions, but quite frankly, I doubt it. Yes, you'll end up "traversing" exactly as many objects either way, but the "globe subobjects one by one" is going to be a _hell_ of a lot more memory-efficient, and quite frankly, "memory usage" == "performance" under many loads (notably, any load that uses too much memory will _suck_ performance-wise, either because of swapping or simply because it will throw out caches that "many small invocations" would not have thrown out).

So I guarantee that it's going to be better to do five clones of five small repositories over one clone of one big one. If only because you need less memory to do the five smaller clones.

Previous: Linus TorvaldsNext: Josef Weidendorfer
Message 74 of 160 in “Re: [RFC] Submodules in GIT”
  1. Andy ParkinsNov 28, 2006
  2. Jakub NarebskiNov 28, 2006
  3. Andy ParkinsNov 28, 2006
  4. Shawn PearceNov 28, 2006
  5. Andy ParkinsNov 28, 2006
  6. Shawn PearceNov 28, 2006
  7. Jon LoeligerNov 28, 2006
  8. Martin WaitzNov 29, 2006
  9. sfNov 30, 2006
  10. Steven GrimmNov 28, 2006
  11. Shawn PearceNov 28, 2006
  12. Martin WaitzNov 29, 2006
  13. Andy ParkinsNov 29, 2006
  14. Andreas EricssonNov 30, 2006
  15. Andy ParkinsNov 30, 2006
  16. Martin WaitzNov 30, 2006
  17. Andreas EricssonNov 30, 2006
  18. sfDec 1, 2006
  19. Martin WaitzDec 1, 2006
  20. sfDec 1, 2006
  21. Martin WaitzDec 1, 2006
  22. Stephan FederDec 1, 2006
  23. Martin WaitzDec 1, 2006
  24. Stephan FederDec 1, 2006
  25. Martin WaitzDec 1, 2006
  26. Uwe Kleine-KoenigDec 5, 2006
  27. Andreas EricssonDec 5, 2006
  28. Jakub NarebskiDec 5, 2006
  29. Uwe Kleine-KoenigDec 5, 2006
  30. Andreas EricssonDec 5, 2006
  31. Sven VerdoolaegeDec 5, 2006
  32. Andy ParkinsDec 1, 2006
  33. Martin WaitzDec 1, 2006
  34. sfDec 1, 2006
  35. Martin WaitzDec 1, 2006
  36. sfDec 1, 2006
  37. Martin WaitzDec 1, 2006
  38. Andreas EricssonDec 1, 2006
  39. Martin WaitzDec 1, 2006
  40. Andreas EricssonDec 1, 2006
  41. Martin WaitzDec 1, 2006
  42. Andreas EricssonDec 1, 2006
  43. Linus TorvaldsDec 1, 2006
  44. sfDec 1, 2006
  45. Andreas EricssonDec 1, 2006
  46. Linus TorvaldsDec 1, 2006
  47. Martin WaitzDec 1, 2006
  48. Alan ChandlerDec 1, 2006
  49. Josef WeidendorferDec 1, 2006
  50. Martin WaitzDec 1, 2006
  51. Josef WeidendorferDec 1, 2006
  52. Martin WaitzDec 1, 2006
  53. Josef WeidendorferDec 1, 2006
  54. Martin WaitzDec 2, 2006
  55. Josef WeidendorferDec 3, 2006
  56. Martin WaitzDec 3, 2006
  57. Linus TorvaldsDec 1, 2006
  58. sfDec 1, 2006
  59. Josef WeidendorferDec 1, 2006
  60. Linus TorvaldsDec 1, 2006
  61. Josef WeidendorferDec 1, 2006
  62. Linus TorvaldsDec 2, 2006
  63. Andy ParkinsDec 2, 2006
  64. Josef WeidendorferDec 2, 2006
  65. Linus TorvaldsDec 2, 2006
  66. Martin WaitzDec 2, 2006
  67. Linus TorvaldsDec 2, 2006
  68. Martin WaitzDec 2, 2006
  69. Josef WeidendorferDec 3, 2006
  70. Martin WaitzDec 2, 2006
  71. Linus TorvaldsDec 2, 2006
  72. Martin WaitzDec 2, 2006
  73. Linus TorvaldsDec 2, 2006
  74. Linus TorvaldsDec 2, 2006
  75. Thoughts about memory requirements in traversals [Was: Re: [RFC] Submodules in GIT]Josef Weidendorfer, Dec 3, 2006
  76. Linus TorvaldsDec 3, 2006
  77. Shawn PearceDec 3, 2006
  78. Josef WeidendorferDec 3, 2006
  79. Jakub NarebskiDec 3, 2006
  80. Josef WeidendorferDec 3, 2006
  81. Martin WaitzDec 3, 2006
  82. sfDec 1, 2006
  83. Torgil SvenssonDec 2, 2006
  84. Linus TorvaldsDec 2, 2006
  85. Torgil SvenssonDec 3, 2006
  86. Linus TorvaldsDec 3, 2006
  87. Torgil SvenssonDec 4, 2006
  88. Linus TorvaldsDec 4, 2006
  89. Torgil SvenssonDec 4, 2006
  90. Andreas EricssonDec 5, 2006
  91. Jakub NarebskiDec 5, 2006
  92. Andreas EricssonDec 5, 2006
  93. Jakub NarebskiDec 5, 2006
  94. Andy ParkinsDec 3, 2006
  95. Daniel BarkalowDec 5, 2006
  96. sfDec 5, 2006
  97. R. Steve McKownDec 9, 2006
  98. Torgil SvenssonDec 10, 2006
  99. Torgil SvenssonDec 14, 2006
  100. Josef WeidendorferDec 14, 2006
  101. Torgil SvenssonDec 15, 2006
  102. Josef WeidendorferDec 15, 2006
  103. Torgil SvenssonDec 15, 2006
  104. Torgil SvenssonDec 16, 2006
  105. Torgil SvenssonDec 16, 2006
  106. Jakub NarebskiDec 16, 2006
  107. Torgil SvenssonDec 16, 2006
  108. Jakub NarebskiDec 16, 2006
  109. Junio C HamanoDec 16, 2006
  110. Torgil SvenssonDec 16, 2006
  111. Torgil SvenssonDec 16, 2006
  112. Jakub NarebskiDec 16, 2006
  113. Torgil SvenssonDec 17, 2006
  114. Linus TorvaldsDec 16, 2006
  115. Linus TorvaldsDec 16, 2006
  116. Torgil SvenssonDec 16, 2006
  117. Martin WaitzDec 2, 2006
  118. Josef WeidendorferDec 1, 2006
  119. Martin WaitzDec 1, 2006
  120. Linus TorvaldsDec 1, 2006
  121. Josef WeidendorferDec 2, 2006
  122. Linus TorvaldsDec 2, 2006
  123. Andy ParkinsDec 2, 2006
  124. Michael K. EdwardsDec 4, 2006
  125. Sam VilainDec 5, 2006
  126. Sven VerdoolaegeDec 3, 2006
  127. Linus TorvaldsDec 3, 2006
  128. Jakub NarebskiDec 3, 2006
  129. Josef WeidendorferDec 4, 2006
  130. sfDec 1, 2006
  131. Jon LoeligerDec 8, 2006
  132. Sven VerdoolaegeDec 8, 2006
  133. Andreas EricssonDec 12, 2006
  134. Martin WaitzDec 1, 2006
  135. Martin WaitzDec 1, 2006
  136. Andreas EricssonDec 1, 2006
  137. Martin WaitzDec 1, 2006
  138. Stephan FederDec 1, 2006
  139. Martin WaitzDec 1, 2006
  140. Stephan FederDec 1, 2006
  141. Martin WaitzDec 1, 2006
  142. Stephan FederDec 1, 2006
  143. Martin WaitzDec 1, 2006
  144. sfDec 1, 2006
  145. Martin WaitzDec 2, 2006
  146. Andy ParkinsDec 1, 2006
  147. Martin WaitzDec 1, 2006
  148. Andy ParkinsDec 1, 2006
  149. Martin WaitzDec 1, 2006
  150. Andy ParkinsDec 1, 2006
  151. Martin WaitzDec 1, 2006
  152. Andy ParkinsDec 2, 2006
  153. Josef WeidendorferDec 2, 2006
  154. Martin WaitzDec 2, 2006
  155. Josef WeidendorferDec 3, 2006
  156. Martin WaitzDec 2, 2006
  157. Jakub NarebskiDec 2, 2006
  158. Jakub NarebskiDec 2, 2006
  159. Jakub NarebskiDec 2, 2006
  160. Andy ParkinsDec 3, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.