git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: gc --aggressive

From
Nicolas Pitre <nico@fluxnic.net>
Date
May 1, 2012, 17:17 UTC
Message-ID
<alpine.LFD.2.02.1205011259200.21030@xanadu.home>
In-Reply-To
<20120501162806.GA15614@sigill.intra.peff.net>
On Tue, 1 May 2012, Jeff King wrote:
Show 10 quoted lines
> On Sun, Apr 29, 2012 at 09:53:31AM -0400, Nicolas Pitre wrote:
> 
> > But my remark was related to the fact that you need to double the 
> > affected resources to gain marginal improvements at some point.  This is 
> > true about computing hardware too: eventually you need way more gates 
> > and spend much more $$$ to gain some performance, and the added 
> > performance is never linear with the spending.
> 
> Right, I agree with that. The trick is just finding the right spot on
> that curve for each repo to maximize the reward/effort ratio.

Absolutely, at least for the default settings. However this is not what --aggressive is meant to be.

Show 21 quoted lines
> > >   1. Should we bump our default window size? The numbers above show that
> > >      typical repos would benefit from jumping to 20 or even 40.
> > 
> > I think this might be a good indication that the number of objects is a 
> > bad metric to size the window, as I mentioned previously.
> > 
> > Given that you have the test repos already, could you re-run it with 
> > --window=1000 and play with --window-memory instead?  I would be curious 
> > to see if this provides more predictable results.
> 
> It doesn't help. The git.git repo does well with about a 1m window
> limit. linux-2.6 is somewhere between 1m and 2m. But the phpmyadmin repo
> wants more like 16m. So it runs into the same issue as using object
> counts.
> 
> But it's much, much worse than that. Here are the actual numbers (same
> format as before; left-hand column is either window size (if no unit) or
> window-memory limit (if k/m unit), followed by resulting pack size, its
> percentage of baseline --window=10 pack, the user CPU time and finally
> its percentage of the baseline):
> [...]

Ouch! Well... so much for good theory. I'm still really surprised and disappointed as I didn't expect such damage at all.

However, this is possibly a good baseline to determine a default value for window-memory though. Given your number, we clearly see that good packing can be achieved with relatively little memory and therefore it might be a good idea not to leave this parameter unbounded by default in order to catch potential pathological cases. Maybe 64M would be a good default value? Having a repack process eating up more than 16GB of RAM because its RAM usage is unbounded is certainly not nice.

Show 13 quoted lines
> > Maybe we could look at the size reduction within the delta search loop.  
> > If the reduction quickly diminishes as tested objects are further away 
> > from the target one then the window doesn't have to be very large, 
> > whereas if the reduction remains more or less constant then it might be 
> > worth searching further.  That could be used to dynamically size the 
> > window at run time.
> 
> I really like the idea of dynamically sizing the window based on what we
> find. If it works. I don't think there's any reason you couldn't have 50
> absolutely terrible delta candidates followed by one really amazing
> delta candidate. But maybe in practice the window tends to get
> progressively worse due to the heuristics, and outliers are unlikely. I
> guess we'd have to experiment.

Yes. The idea is to continue searching if results are not progressively becoming worse fast enough. Coming up with a good way to infer that is far from obvious though.

Nicolas
Previous: Nicolas PitreNext: Jeff King
Message 17 of 26 in “gc --aggressive”
  1. Jay SoffianApr 17, 2012
  2. Jay SoffianApr 17, 2012
  3. Matthieu MoyApr 17, 2012
  4. Jeff KingApr 17, 2012
  5. Jeff KingApr 28, 2012
  6. Nicolas PitreApr 28, 2012
  7. Jeff KingApr 29, 2012
  8. Nicolas PitreApr 29, 2012
  9. Jeff KingMay 1, 2012
  10. Jeff KingMay 1, 2012
  11. Nicolas PitreMay 1, 2012
  12. Junio C HamanoMay 1, 2012
  13. Nicolas PitreMay 1, 2012
  14. Jeff KingMay 1, 2012
  15. Jeff KingMay 1, 2012
  16. Nicolas PitreMay 1, 2012
  17. Nicolas PitreMay 1, 2012
  18. Jeff KingMay 1, 2012
  19. Nicolas PitreMay 1, 2012
  20. Nicolas PitreApr 28, 2012
  21. Jeff KingApr 17, 2012
  22. Junio C HamanoApr 17, 2012
  23. Jeff KingApr 17, 2012
  24. Junio C HamanoApr 17, 2012
  25. Nicolas PitreApr 28, 2012
  26. Andreas EricssonApr 18, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.