git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: cherry-pick is slow

From
Junio C Hamano <gitster@pobox.com>
Date
May 15, 2012, 21:03 UTC
Message-ID
<7vwr4dcg2b.fsf@alter.siamese.dyndns.org>
In-Reply-To
<7v1umldw3i.fsf@alter.siamese.dyndns.org>
Junio C Hamano <gitster@pobox.com> writes:
Show 11 quoted lines
> Unfortunately, I do not think that the actual implementation of
> "cherry-pick" matches that expectation, as it is a full three-way merge.
>
> I am somewhat curious to see what the performance characteristics would be
> if the same commit is replayed using
>
> 	git format-patch -1 --stdout $commit | git apply --index --3way
>
> pipeline.  Depending on the number of paths in the whole tree vs the
> number of paths the $commit touches, I wouldn't be surprised if it is
> faster.

An unscientific datapoint shows that with a project as small as the kernel, the difference is noticeable.

For example, v3.4-rc7-22-g3911ff3 (random tip of the day) touches two paths, and cherry-picking it on top of v3.3 goes like this:

    $ git checkout v3.3 && EDITOR=: /usr/bin/time git cherry-pick 3911ff3
     Author: Jiri Kosina <jkosina@suse.cz>
     2 files changed, 2 insertions(+)
    1.08user 0.20system 0:01.28elapsed 99%CPU (0avgtext+0avgdata 469728maxresident)k
    0inputs+7536outputs (0major+52604minor)pagefaults 0swaps
as opposed to an alternative that touches only these two paths:
    $ git checkout v3.3 && EDITOR=: /usr/bin/time sh -c '
	git format-patch --stdout -1 3911ff3 | git am -3'
    Applying: genirq: export handle_edge_irq() and irq_to_desc()
    0.36user 0.16system 0:00.46elapsed 112%CPU (0avgtext+0avgdata 254720maxresident)k
    0inputs+14872outputs (0major+55145minor)pagefaults 0swaps

Of course, there are vast differences between v3.3 and 3911ff3^1; 11k+ paths touched, countless paths created and deleted.

I _think_ most of the overhead comes from having to match the large trees in unpack_trees() even though none of the changes between the base versions matters for this" cherry-pick".

Both reads the flat index into the core in its entirety and futzing with the index file format would not affect this comparison, even though it could improve the performance of "am", if done right, as it could limit its updates to only two paths. In the merge case, we pretty much rebuild the resulting index from scratch by walking the entire tree in unpack_trees(), so there won't be much benefit.

Perhaps we might want to rethink the way we run merges?
Previous: Junio C HamanoNext: Jeff King
Message 8 of 9 in “cherry-pick is slow”
  1. Dmitry RisenbergMay 12, 2012
  2. Junio C HamanoMay 13, 2012
  3. Dmitry RisenbergMay 13, 2012
  4. Jeff KingMay 14, 2012
  5. Jeff KingMay 15, 2012
  6. Paweł SikoraMay 15, 2012
  7. Junio C HamanoMay 15, 2012
  8. Junio C HamanoMay 15, 2012
  9. Jeff KingMay 19, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.