Re: Git is not scalable with too many refs/*
- From
Shawn Pearce <spearce@spearce.org>
- Date
- Jun 9, 2011, 15:56 UTC
- Message-ID
- <BANLkTimk06eibz99AO_0BwzoL6FWb5pR8A@mail.gmail.com>
- In-Reply-To
- <4DF0EC32.40001@gmail.com>
On Thu, Jun 9, 2011 at 08:52, A Large Angry SCM <gitzilla@gmail.com> wrote:
Show 8 quoted lines
> On 06/09/2011 11:23 AM, Shawn Pearce wrote: >> Having a reference to every commit in the repository is horrifically >> slow. We run into this with Gerrit Code Review and I need to find >> another solution. Git just wasn't meant to process repositories like >> this. > > Assuming a very large number of refs, what is it that makes git so > horrifically slow? Is there a design or implementation lesson here?
A few things.
Git does a sequential scan of all references when it first needs to access references for an operation. This requires reading the entire packed-refs file, and the recursive scan of the "refs/" subdirectory for any loose refs that might override the packed-refs file.
A lot of operations toss every commit that a reference points at into the revision walker's LRU queue. If you have a tag pointing to every commit, then the entire project history enters the LRU queue at once, up front. That queue is managed with O(N^2) insertion time. And the entire queue has to be filled before anything can be output.
-- Shawn.