Re: filter-branch performance
- From
Roberto Tyley <roberto.tyley@gmail.com>
- Date
- Dec 10, 2014, 23:44 UTC
- Message-ID
- <CAFY1eda-utVReuQnotSUDPV4-=hiMupbNdLZrYnEiaDryXQboQ@mail.gmail.com>
- In-Reply-To
- <xmqqfvcnjxry.fsf@gitster.dls.corp.google.com>
On 10 December 2014 at 16:05, Junio C Hamano <gitster@pobox.com> wrote:
Show 18 quoted lines
> Roberto Tyley <roberto.tyley@gmail.com> writes: > >> The BFG is generally faster than filter-branch for 3 reasons: >> >> 1. No forking - everything stays in the JVM process >> 2. Embarrassingly parallel algorithm makes good use of multi-core machines >> 3. Memoization means no Git object (file or folder) is cleaned more than once >> >> In the case of your problem, only the first factor will be noticeably >> helpful. Unfortunately commits do need to be cleaned sequentially, as >> their hashes depend on the hashes of their parents, and filter-branch >> doesn't clean /commits/ more than once, the way it does with files or >> folders - so the last 2 reasons in the list won't be significant. > > Just this part. If your history is bushy, you should be able to > rewrite histories of merged branches in parallel up to the point > they are merged---rewriting of the merge commit of course has to > wait until all the branches have been rewritten, though.
That's true, and the bfg does take advantage of that parallelism, so as well as point 1, point 2 will provide some benefit if history is bushy enough :)