git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] Adding a cache of commit to patch-id pairs to speed up git-cherry

From
Geoffrey Irving <irving@naml.us>
Date
Jun 7, 2008, 23:50 UTC
Message-ID
<7f9d599f0806071650p6d650b7dwfd69c753850f3e25@mail.gmail.com>
In-Reply-To
<alpine.DEB.1.00.0806021913340.13507@racer.site.net>

On Mon, Jun 2, 2008 at 11:15 AM, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:

Show 37 quoted lines
> Hi,
>
> On Mon, 2 Jun 2008, Geoffrey Irving wrote:
>
>> On Mon, Jun 2, 2008 at 9:18 AM, Johannes Schindelin
>> <Johannes.Schindelin@gmx.de> wrote:
>>
>> > On Mon, 2 Jun 2008, Geoffrey Irving wrote:
>> >
>> >> On Mon, Jun 2, 2008 at 8:37 AM, Johannes Schindelin
>> >> <Johannes.Schindelin@gmx.de> wrote:
>> >>
>> >> > Another issue that just hit me: this cache is append-only, so if it
>> >> > grows too large, you have no other option than to scratch and
>> >> > recreate it. Maybe this needs porcelain support, too?  (git gc?)
>> >>
>> >> If so, the correct operation is to go through the hash and remove
>> >> entries that refer to commits that no longer exist.  I can add this
>> >> if you want.  Hopefully somewhere along the way git-gc constructs an
>> >> easy to traverse list of extant commits, and this will be
>> >> straightforward.
>> >
>> > I don't know... if you have created a cached patch-id for every commit
>> > (by mistake, for example) and do not need it anymore, it might make
>> > git-cherry substantially faster to just scrap the cache.
>>
>> Well, ideally hash maps are O(1), but it could be a difference between a
>> "compare 40 bytes" constant and a "read a 4k block into memory"
>> constant, so in practice yes.  Scrapping it entirely will also make the
>> implementation much simpler.
>>
>> It seems a little sad to wipe all that effort each time, but
>> regenerating the cache is likely to be less expensive than a git-gc, so
>> it shouldn't change any amortized complexities.
>
> Well, how about only scrapping the cache if it is older than, say, 2
> weeks, and is larger than, say, 200kB?  That should help.

That heuristic is insufficient, since it doesn't do anything in the normal case where a new entry appears every few days (e.g., when syncing between two branches with cherry-pick).

I don't know what the best alternative is, so I left garbage collection out of the patch I just submitted. We can add it once we decide what to do. I'm not sure it's a serious problem: if you "accidentally" added entries for all commits in the git tree, the file is still under 1M.

Geoffrey
Previous: Johannes SchindelinNext: Johannes Schindelin
Message 11 of 16 in “Adding a cache of commit to patch-id pairs to speed up git-cherry”
  1. Adding a cache of commit to patch-id pairs to speed up git-cherryGeoffrey Irving, Jun 2, 2008
  2. Johannes SchindelinJun 2, 2008
  3. Jeff KingJun 2, 2008
  4. Geoffrey IrvingJun 2, 2008
  5. Johannes SchindelinJun 2, 2008
  6. Geoffrey IrvingJun 2, 2008
  7. Shawn O. PearceJun 2, 2008
  8. Johannes SchindelinJun 2, 2008
  9. Geoffrey IrvingJun 2, 2008
  10. Johannes SchindelinJun 2, 2008
  11. Geoffrey IrvingJun 7, 2008
  12. Johannes SchindelinJun 8, 2008
  13. Geoffrey IrvingJun 2, 2008
  14. Johannes SchindelinJun 2, 2008
  15. Geoffrey IrvingJun 2, 2008
  16. Johannes SchindelinJun 2, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.