git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] Adding a cache of commit to patch-id pairs to speed up git-cherry

From
Shawn O. Pearce <spearce@spearce.org>
Date
Jun 2, 2008, 15:56 UTC
Message-ID
<20080602155644.GL12896@spearce.org>
In-Reply-To
<7f9d599f0806020849g567461b2kecd65dbd35d3dc3b@mail.gmail.com>
Geoffrey Irving <irving@naml.us> wrote:
Show 9 quoted lines
> On Mon, Jun 2, 2008 at 8:37 AM, Johannes Schindelin
> > Another issue that just hit me: this cache is append-only, so if it grows
> > too large, you have no other option than to scratch and recreate it.
> > Maybe this needs porcelain support, too?  (git gc?)
> 
> If so, the correct operation is to go through the hash and remove
> entries that refer to commits that no longer exist.  I can add this if
> you want.  Hopefully somewhere along the way git-gc constructs an easy
> to traverse list of extant commits, and this will be straightforward.

git-gc doesn't make such a list. Down deep with git-pack-objects (which is called by git-repack, which is called by git-gc) yes, we do make the list of commits that we can find as reachable, and thus should stay in the repository. But that is really low-level plumbing. Wedging a SHA1->SHA1 hashmap gc task down into that is not a good idea.

Instead you'll need to implement something that does `git rev-list --all -g` (or the internal equivilant) and then remove any entries in your hashmap that aren't in that result set. That's not going to be very cheap.

Given how small entries are (what, 40 bytes?) I'd only want to bother with that collection process if the estimated potential wasted space was over 1M (26,000 entries) or some reasonable threshold like that.

E.g. we could just set the GC for this to be once every 26,000 additions, and only during git-gc. Yea, you might waste about 1M worth of space before we clean up. Big deal, I'll bet you have more than that in loose unreachable objects laying around from git-rebase -i usage. ;-)

-- 
Shawn.
Previous: Geoffrey IrvingNext: Johannes Schindelin
Message 7 of 16 in “Adding a cache of commit to patch-id pairs to speed up git-cherry”
  1. Adding a cache of commit to patch-id pairs to speed up git-cherryGeoffrey Irving, Jun 2, 2008
  2. Johannes SchindelinJun 2, 2008
  3. Jeff KingJun 2, 2008
  4. Geoffrey IrvingJun 2, 2008
  5. Johannes SchindelinJun 2, 2008
  6. Geoffrey IrvingJun 2, 2008
  7. Shawn O. PearceJun 2, 2008
  8. Johannes SchindelinJun 2, 2008
  9. Geoffrey IrvingJun 2, 2008
  10. Johannes SchindelinJun 2, 2008
  11. Geoffrey IrvingJun 7, 2008
  12. Johannes SchindelinJun 8, 2008
  13. Geoffrey IrvingJun 2, 2008
  14. Johannes SchindelinJun 2, 2008
  15. Geoffrey IrvingJun 2, 2008
  16. Johannes SchindelinJun 2, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.