git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Keeping unreachable objects in a separate pack instead of loose?

From
Nicolas Pitre <nico@fluxnic.net>
Date
Jun 12, 2012, 18:25 UTC
Message-ID
<alpine.LFD.2.02.1206121359260.23555@xanadu.home>
In-Reply-To
<20120612175438.GB16522@sigill.intra.peff.net>
On Tue, 12 Jun 2012, Jeff King wrote:
Show 11 quoted lines
> On Tue, Jun 12, 2012 at 01:49:26PM -0400, Nicolas Pitre wrote:
> 
> > > Then those objects will remain in the cruft pack. Which is why, as I
> > > said, it is not generally safe to just delete a cruft pack.
> > 
> > ... and my reply was about the needed changes to still make cruft packs 
> > always crufty even if some of its content suddenly becomes useful again.
> 
> I think we are somehow missing each other's point, then. My point is
> that you do not _need_ to make the cruft packs 100% cruft. You can
> tolerate the duplicated objects until they are pruned.

Absolutely. Duplicated objectes are fine, and I was in fact suggesting to actively duplicate any needed object when it is to be found in a cruft pack only.

Show 5 quoted lines
> Earlier in the thread, I outlined another scheme by which you could
> repack and avoid the duplicates. It does not require changes to git's
> object lookup process, because it would involve manually feeding the
> list of cruft objects to pack-objects (which will pack what you ask it,
> regardless of whether the objects are in other packs).

That might be hard to achieve good delta compression though, as the main key to sort those objects is their path name, and with unreferenced objects you might not necessarily have that information. The ability to reuse pack data might mitigate this though.

Show 13 quoted lines
> > > However, when you do a full repack, those objects will be copied into 
> > > the new pack (because they are referenced). Which is why I am claiming 
> > > that it is safe to remove cruft packs at that point.
> > 
> > Yes, but then there is no point marking such packs as cruft if at any 
> > moment they can become useful again.
> 
> How do you know to keep the packs around and expire them after 2 weeks
> if they are not marked in some way? Otherwise you would delete them as
> part of a "git gc", pushing the reachable objects into the new pack and
> the unreachable objects into a new cruft pack. IOW, you need some way of
> keeping the expiration date on the unreachable objects, or they will
> keep getting "refreshed" by each gc.

My feeling is that we should make a step backward and consider if this is actually the right problem to solve. I don't remember why I might have been opposed to a reflog for deleted branches as you say I did, but that is certainly a feature that could prove to be useful.

Then having a repository that can be used as an alternate for other repositories without knowing about it is also a problem that needs fixing and not only because of this object expiry issue. This is not easy to fix though.

Then, the creation of unreferenced objects from successive 'git add' shouldn't create that many objects in the first place. They currently never get the chance to be packed to start with.

So the problem is really about 'git gc' creating more data on disk which is counter productive for a garbage collecting task. Maybe the trick is simply not to delete any of the old pack which content was repacked into a single new pack and let them age before deleting them, rather than exploding a bunch of loose objects. But then we're back to the same issue I wanted to get away from i.e. identifying real cruft packs and making them safely deletable.

Oh well...
Nicolas
Previous: Jeff KingNext: Ted Ts'o
Message 39 of 48 in “Keeping unreachable objects in a separate pack instead of loose?”
  1. Theodore Ts'oJun 10, 2012
  2. Hallvard B FurusethJun 10, 2012
  3. Thomas RastJun 11, 2012
  4. Ted Ts'oJun 11, 2012
  5. Jeff KingJun 11, 2012
  6. Nicolas PitreJun 11, 2012
  7. Ted Ts'oJun 11, 2012
  8. Jeff KingJun 11, 2012
  9. Ted Ts'oJun 11, 2012
  10. Jeff KingJun 11, 2012
  11. Jeff KingJun 11, 2012
  12. Ted Ts'oJun 11, 2012
  13. Jeff KingJun 11, 2012
  14. Hallvard Breien FurusethJun 11, 2012
  15. Jeff KingJun 11, 2012
  16. Hallvard Breien FurusethJun 11, 2012
  17. Ted Ts'oJun 11, 2012
  18. Jeff KingJun 11, 2012
  19. Ted Ts'oJun 11, 2012
  20. Jeff KingJun 11, 2012
  21. Ted Ts'oJun 11, 2012
  22. Jeff KingJun 11, 2012
  23. Nicolas PitreJun 12, 2012
  24. Jeff KingJun 12, 2012
  25. Nicolas PitreJun 12, 2012
  26. Jeff KingJun 12, 2012
  27. Shawn PearceJun 12, 2012
  28. Jeff KingJun 12, 2012
  29. Nicolas PitreJun 12, 2012
  30. Andreas SchwabJun 12, 2012
  31. Jeff KingJun 12, 2012
  32. Nicolas PitreJun 12, 2012
  33. Jeff KingJun 12, 2012
  34. Nicolas PitreJun 12, 2012
  35. Jeff KingJun 12, 2012
  36. Nicolas PitreJun 12, 2012
  37. Nicolas PitreJun 12, 2012
  38. Jeff KingJun 12, 2012
  39. Nicolas PitreJun 12, 2012
  40. Ted Ts'oJun 12, 2012
  41. Nicolas PitreJun 12, 2012
  42. Ted Ts'oJun 12, 2012
  43. Nicolas PitreJun 12, 2012
  44. Ted Ts'oJun 12, 2012
  45. Jeff KingJun 12, 2012
  46. Martin FickJun 13, 2012
  47. Johan HerlandJun 13, 2012
  48. Junio C HamanoJun 11, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.