Re: dumb transports not being welcomed..
- From
Linus Torvalds <torvalds@osdl.org>
- Date
- Sep 14, 2005, 00:57 UTC
- Message-ID
- <Pine.LNX.4.58.0509131742240.26803@g5.osdl.org>
- In-Reply-To
- <Pine.LNX.4.63.0509140152160.24606@wgmdd8.biozentrum.uni-wuerzburg.de>
On Wed, 14 Sep 2005, Johannes Schindelin wrote:
Show 5 quoted lines
> > IMHO the culprit is git-rev-list, which takes ages and ages for big > repositories (beware: this could be my Darwin client which might be > incapable to stop the rev enumeration in time; but if that can be done > unintentionally, this can be intentionally, too!).
Packed too?
git-rev-list will take a long time if the tree is unpacked and not in the cache. It's all disk seeks. That's _especially_ true of a full clone (which will walk the whole way down).
But I have tons of memory in my machines, and I haven't looked at how badly it does if you don't have that. I know that master.kernel.org is certainly not having any trouble at all with me pulling from lots of trees.. Maybe git-rev-list uses up lots of your memory.
I'm seeing 14 seconds of CPU-time for a _full_ kernel history, with "--objects". Yes, it's not exactly cheap, and maybe I should optimize it (it's all in the "--objects" handling and probably a large portion of it is because trees actually pack very well indeed, so it's actually unpacking a lot of trees), but considering that that is preparing the metadata for pulling down a hundred megs of stuff..
That said, I do think that --objects handling is _very_ CPU-hungry. The offender is this old commit of mine:
4311d328fee11fbd80862e3c5de06a26a0e80046 Author: Linus Torvalds <torvalds@g5.osdl.org> Date: Sat Jul 23 10:01:49 2005 -0700
Be more aggressive about marking trees uninteresting ...
which is much better about avoiding objects in old trees, but it does so at the expense of being _horribly_ CPU-inefficient. It will walk through every tree of every commit that we decided was uninteresting.
You can try to just undo that one commit - it will make pack-files have a few extraneous objects, but I think it will make a huge difference in the CPU cost of "small pulls" (it won't matter at all for the "git clone" case: for that case we just always have to walk the whole object tree).
Linus