From: Linus Torvalds Date: Thu, 28 Apr 2005 15:24:54 GMT Subject: Re: Finding file revisions Message-ID: In-Reply-To: <1114627268.20916.8.camel@tglx.tec.linutronix.de> On Wed, 27 Apr 2005, Thomas Gleixner wrote: > > On Wed, 2005-04-27 at 10:34 -0700, Linus Torvalds wrote: > > > > With more history, "rev-list" should do basically the right thing: it will > > be constant-time for _recent_ commits, and it is linear time in how far > > back you want to go. Which seems quite reasonable. > > Which is quite horrible, if you have a 500k+ blobs repo. It's _not_ linear in blobs. It doesn't care at all about them, in fact. It's linear in how many revisions you go backwards. And I claim that you can't do any better than that, without doing _really_ bad things. > I know you are database allergic, but there a database is the correct > solution. I disagree. I'm not database allergic, I just don't believe in the notion that databases solve all the worlds problems. > Having stored all the relations of those file/tree/commit > blobs in a database it takes <20ms to have a list of all those file > blobs in historical order with some context information retrieved. .. and such an SCM will _suck_ for anything else. You just made creating a commit etc much slower. You now have to update per-file information that you never updated before, and look at information that git simply doesn't _care_ about. Right now, when we create a new version, it's pretty much instantaneous. Exactly becaue we do not look at a _single_ file, and we don't care how they changed from the "previous" version. We just write out the knowledge about what the files are now. Doing a database of file changes would absolutely _suck_. Anybody who thinks that databases are magically faster than not using a database doesn't understand basic physics. Things don't go faster just because you call it a database. Things go faster by _doing_less_. Normally, a database does less by keeping indexes etc around, and the indexes require less work than the data itself. But git _does_ all of that already. Git very much _is_ a database, it's just a specialized one. I dare you to show me wrong. I don't _care_ of you can show the revision history of a single file in 20ms. The easiest way to do that is with a delta format, where the file information basically is single-file in the first place, and you just open the file and print out the results. Guess what? We've had that. It's called RCS/SCCS/CVS, and it's a piece of total and absolute crap. Exactly because single-file revisions simply do not matter. If you want to use a database, go wild. But use it as a _cache_. Then you can build up the database of file revisions "after the fact", and always know that your database is not the real data, it's just an index, and can be thrown away and regenerated at will. That way you don't add overhead to the stuff that actually matters, and that git does a lot better than a general-purpose database could ever do. Linus