Re: [PATCH 1/2] blame: large-scale performance rewrite
- From
Shawn Pearce <spearce@spearce.org>
- Date
- Apr 26, 2014, 00:53 UTC
- Message-ID
- <CAJo=hJukmej1rJXuVoECwd7AxmSue8Wmv4rBmCHEYcWBWNarSw@mail.gmail.com>
- In-Reply-To
- <1398470210-28746-1-git-send-email-dak@gnu.org>
On Fri, Apr 25, 2014 at 4:56 PM, David Kastrup <dak@gnu.org> wrote:
Show 11 quoted lines
> The previous implementation used a single sorted linear list of blame > entries for organizing all partial or completed work. Every subtask had > to scan the whole list, with most entries not being relevant to the > task. The resulting run-time was quadratic to the number of separate > chunks. > > This change gives every subtask its own data to work with. Subtasks are > organized into "struct origin" chains hanging off particular commits. > Commits are organized into a priority queue, processing them in commit > date order in order to keep most of the work affecting a particular blob > collated even in the presence of an extensive merge history.
Without reading the code, this sounds like how JGit runs blame.
> For large files with a diversified history, a speedup by a factor of 3 > or more is not unusual.
And JGit was already usually slower than git-core. Now it will be even slower! :-)