git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Why Git is so fast (was: Re: Eric Sink's blog - notes on git, dscms and a "whole product" approach)

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
May 1, 2009, 21:37 UTC
Message-ID
<alpine.LFD.2.00.0905011420580.5379@localhost.localdomain>
In-Reply-To
<20090501190854.GA13770@coredump.intra.peff.net>
On Fri, 1 May 2009, Jeff King wrote:
> 
> Thanks for the analysis; what you said makes sense to me. However, there
> is at least one case of somebody complaining that git doesn't scale as
> well as perforce for their load:

So we definitely do have scaling issues, there's no question about that. I just don't think they are about enterprise network servers vs the more workstation-oriented OSS world..

I think they're likely about the whole git mentality of looking at the big picture, and then getting swamped by just how _huge_ that picture can be if somebody just put the whole world in a single repository..

With perforce, repository maintenance is such a central issue that the whole p4 mentality seems to _encourage_ everybody to put everything into basically one single p4 repository. And afaik, p4 basically works mostly like CVS, ie it really ends up being pretty much oriented to a "one file at a time" model.

Which is nice in that you can have a million files, and then only check out a few of them - you'll never even _see_ the impact of the other 999,995 files.

And git obviously doesn't have that kind of model at all. Git fundamnetally never really looks at less than the whole repo. Even if you limit things a bit (ie check out just a portion, or have the history go back just a bit), git ends up still always caring about the whole thing, and carrying the knowledge around.

So git scales really badly if you force it to look at everything as one _huge_ repository. I don't think that part is really fixable, although we can probably improve on it.

And yes, then there's the "big file" issues. I really don't know what to do about huge files. We suck at them, I know. There are work-arounds (like not deltaing big objects at all), but they aren't necessarily that great either.

I bet we could probably improve git large-file behavior for many common cases. Do we have a good test-case of some particular suckiness that is actually relevant enough that people might decide to look at it (and by "people", I do mean myself too - but I'd need to be somewhat motivated by it. A usage case that we suck at and that is available and relevant).

			Linus
Previous: Daniel BarkalowNext: david@lang.hm
Message 34 of 39 in “Eric Sink's blog - notes on git, dscms and a "whole product" approach”
  1. Martin LanghoffApr 27, 2009
  2. Cross-Platform Version Control (was: Eric Sink's blog - notes on git, dscms and a "whole product" approach)Jakub Narebski, Apr 28, 2009
  3. Robin RosenbergApr 28, 2009
  4. Martin LanghoffApr 29, 2009
  5. Jeff KingApr 29, 2009
  6. Markus HeidelbergApr 29, 2009
  7. Jakub NarebskiApr 29, 2009
  8. Martin LanghoffApr 29, 2009
  9. Jakub NarebskiApr 28, 2009
  10. Sitaram ChamartyApr 29, 2009
  11. Why Git is so fast (was: Re: Eric Sink's blog - notes on git, dscms and a "whole product" approach)Jakub Narebski, Apr 30, 2009
  12. Michael WittenApr 30, 2009
  13. Jakub NarebskiApr 30, 2009
  14. Shawn O. PearceApr 30, 2009
  15. Kjetil BarvikApr 30, 2009
  16. Shawn O. PearceApr 30, 2009
  17. Kjetil BarvikApr 30, 2009
  18. Steven NoonanMay 1, 2009
  19. James PickensMay 1, 2009
  20. Kjetil BarvikMay 1, 2009
  21. Mike HommeyMay 1, 2009
  22. Kjetil BarvikMay 1, 2009
  23. Tony FinchMay 1, 2009
  24. Dmitry PotapovMay 1, 2009
  25. Mike HommeyMay 1, 2009
  26. Dmitry PotapovMay 1, 2009
  27. Shawn O. PearceApr 30, 2009
  28. Jeff KingApr 30, 2009
  29. Linus TorvaldsMay 1, 2009
  30. Jeff KingMay 1, 2009
  31. david@lang.hmMay 1, 2009
  32. Nicolas PitreMay 1, 2009
  33. Daniel BarkalowMay 1, 2009
  34. Linus TorvaldsMay 1, 2009
  35. david@lang.hmMay 1, 2009
  36. Nicolas PitreApr 30, 2009
  37. Alex RiesenApr 30, 2009
  38. Andreas EricssonMay 4, 2009
  39. Jakub NarebskiApr 30, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.