git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: dangling commits and blobs: is this normal?

From
Nicolas Pitre <nico@cam.org>
Date
Apr 23, 2009, 18:51 UTC
Message-ID
<alpine.LFD.2.00.0904231438280.6741@xanadu.home>
In-Reply-To
<064C9132-2E72-4665-A44D-A2F4194DAC2B@adacore.com>
On Thu, 23 Apr 2009, Geert Bosch wrote:
Show 20 quoted lines
> 
> On Apr 22, 2009, at 16:05, Jeff King wrote:
> > The other tradeoff, mentioned by Matthieu, is not about speed, but about
> > rollover of files on disk. I think he would be in favor of a less
> > optimal pack setup if it meant rewriting the largest packfile less
> > frequently.
> > 
> > However, it may be reasonable to suggest that he just not manually "gc"
> > then. If he is not generating enough commits to warrant an auto-gc, then
> > he is probably not losing much by having loose objects. And if he is,
> > then auto-gc is already taking care of it.
> 
> For large repositories with lots of large files, git spends too much
> time copying large packs for relatively little gain. This is obvious when
> you include a few dozen large objects in any repository.
> Currently, there is no limit to the number of times this data may
> be copied. In particular, the average amount of I/O needed for
> changes of size X depends linearly on the size of the total repository.
> So, the mere presence of a couple of large objects has an large distributed
> overhead.

You can put a limit on the number of times this data is copied, and even set the limit to zero. Just add a .keep file to your .pack file and that data will remain in stone. Any further repack will consider only those newly added objects you may have.

Show 5 quoted lines
> Wouldn't it be better to have a maximum of N packs, named
> pack_0 .. pack_(N - 1),  in the repository with each pack_i being
> between 2^i and 2^(i+1)-1 bytes large? We could even dispense
> completely with loose objects and instead have each git operation
> create a single new pack.

I suggested that already for large enough objects. For small objects this makes no sense as you may accumulate too many of them and each one would need to be opened in order to find if it contains the desired object whereas currently you need a simple directory lookup.

> number of packs in the way described is useful and will lead to significant
> speedups, especially during large imports that currently require frequent
> repacking of the entire repository.
Others commented on that issue already.
Nicolas
Previous: Matthias AndreeNext: John Dlugosz
Message 20 of 21 in “dangling commits and blobs: is this normal?”
  1. John DlugoszApr 21, 2009
  2. Jeff KingApr 22, 2009
  3. Brandon CaseyApr 22, 2009
  4. Nicolas PitreApr 22, 2009
  5. Matthieu MoyApr 22, 2009
  6. Jeff KingApr 22, 2009
  7. Brandon CaseyApr 22, 2009
  8. Jeff KingApr 22, 2009
  9. Nicolas PitreApr 22, 2009
  10. Matthieu MoyApr 23, 2009
  11. Nicolas PitreApr 22, 2009
  12. Brandon CaseyApr 22, 2009
  13. Nicolas PitreApr 22, 2009
  14. Jeff KingApr 22, 2009
  15. Nicolas PitreApr 22, 2009
  16. Geert BoschApr 23, 2009
  17. Shawn O. PearceApr 23, 2009
  18. Geert BoschApr 23, 2009
  19. Matthias AndreeApr 23, 2009
  20. Nicolas PitreApr 23, 2009
  21. John DlugoszApr 22, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.