git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v3 12/14] builtin/pack-objects: use `packfile_store_for_each_object()`

From
Patrick Steinhardt <ps@pks.im>
Date
Jan 26, 2026, 08:53 UTC
Message-ID
<aXcrftLpfcG4S5AX@pks.im>
In-Reply-To
<aXO/YLzRlDXD5IPY@nand.local>
On Fri, Jan 23, 2026 at 01:35:12PM -0500, Taylor Blau wrote:
Show 52 quoted lines
> On Fri, Jan 23, 2026 at 10:43:16AM +0100, Patrick Steinhardt wrote:
> > On Thu, Jan 22, 2026 at 08:21:55PM -0500, Taylor Blau wrote:
> > > On Wed, Jan 21, 2026 at 01:50:28PM +0100, Patrick Steinhardt wrote:
> > > >  static int add_object_in_unpacked_pack(const struct object_id *oid,
> > > > -				       struct packed_git *pack,
> > > > -				       uint32_t pos,
> > > > +				       struct object_info *oi,
> > > >  				       void *data UNUSED)
> > > >  {
> > > >  	if (cruft) {
> > > > -		off_t offset;
> > > > -		time_t mtime;
> > > > -
> > > > -		if (pack->is_cruft) {
> > > > -			if (load_pack_mtimes(pack) < 0)
> > > > -				die(_("could not load cruft pack .mtimes"));
> > > > -			mtime = nth_packed_mtime(pack, pos);
> > > > -		} else {
> > > > -			mtime = pack->mtime;
> > > > -		}
> > > > -		offset = nth_packed_object_offset(pack, pos);
> > > > -
> > > > -		add_cruft_object_entry(oid, OBJ_NONE, pack, offset,
> > > > -				       NULL, mtime);
> > >
> > > OK, here's where we see the existing logic for determining the mtime of
> > > an object in the GC sense. I see there's a subsequent patch that also
> > > makes use of the object_info->mtimep field, and my guess is (not having
> > > completely read that patch yet) that having the same notion of mtime
> > > between the two callsites is desirable.
> > >
> > > I still wonder whether imposing that notion of mtime at the object_info
> > > layer is the right choice. I wonder if it would make more sense to allow
> > > the caller to have a "statp" pointer filled out (or alternatively stick
> > > a "struct stat" in both the packed union type as well as the loose one,
> > > though the latter doesn't yet exist).
> >
> > The problem with filling out a `struct stat` though is that it will only
> > apply to backends that actually have a path to stat. There may be other
> > backends that don't. You could of course pretend that there was a file
> > and fill in the `st_mtime` field. But I don't really see the benefit
> > over having a standalone mtime field.
> 
> I understand what you're saying, but I don't think that this is unique
> to stat. Looking through the object_info struct, there are a handful of
> fields on the request side that are coupled to the objects themselves,
> not their representation, such as typep, sizep, and contentp.
> 
> But there are a handful of fields in the request section which are *not*
> properties of the objects themselves, but rather properties of the way
> those objects are represented as part of the backend-specific
> implementation.

Yes, the specific interpretation will change for some fields. But in general, most of the fields still apply to all backends.

Show 6 quoted lines
> For example, disk_sizep suggests that all objects are stored on disk and
> have a clear notion of how much space they occupy. I could imagine a
> backend implementation where perhaps the contents of objects are divvied
> up into smaller chunks and deduplicated across many objects. I don't
> think there is a clear answer to how much "disk size" an object occupies
> in that case.

This would still apply to other backends though. It's true that "disk" size is a bit of a misnomer now, and that it should probably rather be renamed to "storage" size. But overall, no matter the backend, you will still eventually end up storing the object data somewhere, and that takes up space.

Show 6 quoted lines
> delta_base_oid is another field that I'd argue is not a property of the
> object itself, but rather its representation. Of course, objects stored
> in packfiles may or may not be stored as a delta against some other
> object, and thus being able to ask what that object is makes sense. But
> loose objects don't have the same property as a result of how they are
> stored.

Yup. This field is specific to the packed backend indeed and ideally shouldn't be part of the `struct object_info`. In the best case it could be lifted into `struct object_info::u`, but I'm not sure whether that's easily possible.

> To me this seems like an example where implementation-specific details
> are already leaking through the object_info struct. So in that sense I
> don't think that adding a "struct stat" here is meaningfully changing
> anything.

The thing is that I'm trying to clean up all the different messes that we have. So instead of adding _more_ leakiness, I'd rather prefer to remove some of it.

Show 7 quoted lines
> But I think the proposed mtimep field is a special case not only for the
> reasons stated above, but because an object's mtime has multiple
> interpretations already. For example, if I'm asking about an object's
> mtime, and that object happens to be stored in a cruft pack, which mtime
> am I referring to? Packed objects inherit their mtime from the mtime of
> the *.pack itself, but cruft objects have an additional interpretation
> which is read from the *.mtimes file corresponding to the cruft pack.
This is a good question indeed though.
Show 5 quoted lines
> I don't love bolting another leaky abstraction onto the object_info
> interface, but my broader concern is that the information here is not
> just leaky but ambiguous. By adding a statp pointer, I think the
> information is less ambiguous since the GC-specific interpretation of
> mtime is done at a layer above stat(2).

Fair point. For the current backend, mtime can be ambiguous as the same object may be stored multiple times: either as a loose object, or as part of any of the packfiles.

In the context of `odb_for_each_object()` that info is not ambigous though: we would yield the same object multiple times, and every time we yield it we will may have a different mtime. And this is working as expected for the two callsites:

  - In "reachable.c" we use the mtimep field in the context of recent
    objects. So if at least one of the objects has a new-enough mtime we
    would eventually see it.
 - In "builtin/pack-objects.c" we use basically the same logic as we
   use after my patch seires, where we use either the cruft time or the
   pack time. So things work as expected over there, too.

So I'd claim that this is working sensibly for `odb_for_each_object()`, and there is no ambiguity involved. It's the caller that has to disambiguite, and that's already happening.

But things are a bit different if you invoke `odb_read_object_info()` directly, as we have no way to disambiguate there. We only want to yield _a_ representation of an object, so the mtime will be derived from whatever data structure the object was found in first. This could be helped with better documentation.

Show 51 quoted lines
> > > Then the caller could do something like:
> > >
> > > static time_t object_info_gc_mtime(const struct object_info *oi)
> > > {
> > >     if (!oi->statp)
> > >         BUG("oops!");
> > >
> > >     switch (oi->whence) {
> > >     case OI_CACHED:
> > >         return 0;
> > >     case OI_LOOSE:
> > >         return oi->statp->st_mtime;
> > >     case OI_PACKED:
> > >         struct packed_git *p = oi->u.packed.pack;
> > >         if (p->is_cruft) {
> > >             uint32_t pack_pos;
> > >
> > >             if (load_pack_mtimes(p) < 0)
> > >                 die(_("could not load cruft pack .mtimes for '%s'"),
> > >                     pack_basename(p));
> > >             if (offset_to_pack_pos(p, oi->u.packed.offset, &pack_pos) < 0)
> > >                 die(_("could not find offset for object '%s' in cruft pack '%s'"),
> > >                     oid_to_hex(&oi->oid),
> > >                     pack_basename(p));
> > >
> > >             return nth_packed_mtime(p, pack_pos_to_index(p, pack_pos));
> > >         } else {
> > >             return p->mtime; /* or oi->statp->st_mtime */
> > >         }
> > >     default:
> > >         BUG("unknown oi->whence: %d", oi->whence);
> > >     }
> > > }
> > >
> > > I like the above because it encapsulates the GC-specific interpretation
> > > of an object's mtime outside of the object_info layer, while adding
> > > information (namely statp) that is generic enough to be potentially
> > > useful to other callers who may not be interested in the GC-specific
> > > interpretation.
> >
> > This isn't achieving the goal of making the logic pluggable though, as
> > you now have backend-specific logic outside of the backends. Also, isn't
> > the end result basically the same as what I have proposed, except that
> > my version _is_ fully pluggable because the logic is entirely contained
> > in the backend?
> 
> Yes, the end result is the same, both your patch and what I wrote here
> implement the same GC-specific definition of an object's "mtime". I am
> not following the argument about pluggability, though. The concern I
> have above is that we are pushing domain-specific logic into the object
> storage backend, not the other way around.

To expand on the pluggability bit: every time you add a new backend you'll have to extend the above logic to understand how it represents the mtime. That by itself might be doable, but let's for example consider a backend that is a black box to us (like a shared library that may plug in arbitrary storage logic). In that case you would not even be able to derive the information unless you have a generic layer that lets you convey it to the caller.

So overall I agree with you that there are nuances here, and that the mtimep pointer _can_ be used incorrectly. But I still think that the concept is generic enough across backends, and the refactored logic still works as extended. I'll try to expand the docs and commit message a bit to cover this discussion.

Thanks!
Patrick
Previous: Taylor BlauNext: Jeff King
Message 98 of 120 in “odb: introduce `odb_for_each_object()`”
  1. 00/14 odb: introduce `odb_for_each_object()`Patrick Steinhardt, Jan 15, 2026
  2. 01/14 odb: rename `FOR_EACH_OBJECT_*` flagsPatrick Steinhardt, Jan 15, 2026
  3. Justin ToblerJan 15, 2026
  4. 02/14 odb: fix flags parameter to be unsignedPatrick Steinhardt, Jan 15, 2026
  5. 03/14 object-file: extract function to read object info from pathPatrick Steinhardt, Jan 15, 2026
  6. Justin ToblerJan 15, 2026
  7. Patrick SteinhardtJan 16, 2026
  8. Karthik NayakJan 20, 2026
  9. 04/14 object-file: introduce function to iterate through objectsPatrick Steinhardt, Jan 15, 2026
  10. Justin ToblerJan 15, 2026
  11. Patrick SteinhardtJan 16, 2026
  12. Karthik NayakJan 20, 2026
  13. 05/14 packfile: extract function to iterate through objects of a storePatrick Steinhardt, Jan 15, 2026
  14. 06/14 packfile: introduce function to iterate through objectsPatrick Steinhardt, Jan 15, 2026
  15. 07/14 odb: introduce `odb_for_each_object()`Patrick Steinhardt, Jan 15, 2026
  16. Justin ToblerJan 15, 2026
  17. Patrick SteinhardtJan 16, 2026
  18. Justin ToblerJan 16, 2026
  19. Patrick SteinhardtJan 19, 2026
  20. Karthik NayakJan 20, 2026
  21. Patrick SteinhardtJan 21, 2026
  22. 08/14 builtin/fsck: refactor to use `odb_for_each_object()`Patrick Steinhardt, Jan 15, 2026
  23. Justin ToblerJan 15, 2026
  24. 09/14 treewide: enumerate promisor objects via `odb_for_each_object()`Patrick Steinhardt, Jan 15, 2026
  25. 10/14 treewide: drop uses of `for_each_{loose,packed}_object()`Patrick Steinhardt, Jan 15, 2026
  26. Justin ToblerJan 15, 2026
  27. Patrick SteinhardtJan 16, 2026
  28. Justin ToblerJan 16, 2026
  29. Patrick SteinhardtJan 19, 2026
  30. 11/14 odb: introduce mtime fields for object info requestsPatrick Steinhardt, Jan 15, 2026
  31. 12/14 builtin/pack-objects: use `packfile_store_for_each_object()`Patrick Steinhardt, Jan 15, 2026
  32. 13/14 reachable: convert to use `odb_for_each_object()`Patrick Steinhardt, Jan 15, 2026
  33. 14/14 odb: drop unused `for_each_{loose,packed}_object()` functionsPatrick Steinhardt, Jan 15, 2026
  34. Junio C HamanoJan 15, 2026
  35. Patrick SteinhardtJan 16, 2026
  36. Junio C HamanoJan 16, 2026
  37. 00/14 odb: introduce `odb_for_each_object()`Patrick Steinhardt, Jan 20, 2026
  38. 01/14 odb: rename `FOR_EACH_OBJECT_*` flagsPatrick Steinhardt, Jan 20, 2026
  39. 02/14 odb: fix flags parameter to be unsignedPatrick Steinhardt, Jan 20, 2026
  40. 03/14 object-file: extract function to read object info from pathPatrick Steinhardt, Jan 20, 2026
  41. 04/14 object-file: introduce function to iterate through objectsPatrick Steinhardt, Jan 20, 2026
  42. 05/14 packfile: extract function to iterate through objects of a storePatrick Steinhardt, Jan 20, 2026
  43. 06/14 packfile: introduce function to iterate through objectsPatrick Steinhardt, Jan 20, 2026
  44. 07/14 odb: introduce `odb_for_each_object()`Patrick Steinhardt, Jan 20, 2026
  45. 08/14 builtin/fsck: refactor to use `odb_for_each_object()`Patrick Steinhardt, Jan 20, 2026
  46. 09/14 treewide: enumerate promisor objects via `odb_for_each_object()`Patrick Steinhardt, Jan 20, 2026
  47. 10/14 treewide: drop uses of `for_each_{loose,packed}_object()`Patrick Steinhardt, Jan 20, 2026
  48. 11/14 odb: introduce mtime fields for object info requestsPatrick Steinhardt, Jan 20, 2026
  49. 12/14 builtin/pack-objects: use `packfile_store_for_each_object()`Patrick Steinhardt, Jan 20, 2026
  50. 13/14 reachable: convert to use `odb_for_each_object()`Patrick Steinhardt, Jan 20, 2026
  51. 14/14 odb: drop unused `for_each_{loose,packed}_object()` functionsPatrick Steinhardt, Jan 20, 2026
  52. 00/14 odb: introduce `odb_for_each_object()`Patrick Steinhardt, Jan 21, 2026
  53. 01/14 odb: rename `FOR_EACH_OBJECT_*` flagsPatrick Steinhardt, Jan 21, 2026
  54. 02/14 odb: fix flags parameter to be unsignedPatrick Steinhardt, Jan 21, 2026
  55. Jeff KingJan 21, 2026
  56. Taylor BlauJan 22, 2026
  57. Junio C HamanoJan 22, 2026
  58. Jeff KingJan 22, 2026
  59. Patrick SteinhardtJan 23, 2026
  60. Junio C HamanoJan 26, 2026
  61. Patrick SteinhardtJan 22, 2026
  62. Taylor BlauJan 22, 2026
  63. 03/14 object-file: extract function to read object info from pathPatrick Steinhardt, Jan 21, 2026
  64. Taylor BlauJan 22, 2026
  65. Patrick SteinhardtJan 22, 2026
  66. Taylor BlauJan 22, 2026
  67. 04/14 object-file: introduce function to iterate through objectsPatrick Steinhardt, Jan 21, 2026
  68. Taylor BlauJan 22, 2026
  69. Patrick SteinhardtJan 22, 2026
  70. Taylor BlauJan 23, 2026
  71. 05/14 packfile: extract function to iterate through objects of a storePatrick Steinhardt, Jan 21, 2026
  72. Taylor BlauJan 22, 2026
  73. 06/14 packfile: introduce function to iterate through objectsPatrick Steinhardt, Jan 21, 2026
  74. Taylor BlauJan 23, 2026
  75. Patrick SteinhardtJan 23, 2026
  76. Chris TorekJan 23, 2026
  77. Junio C HamanoJan 23, 2026
  78. Taylor BlauJan 23, 2026
  79. 07/14 odb: introduce `odb_for_each_object()`Patrick Steinhardt, Jan 21, 2026
  80. Taylor BlauJan 23, 2026
  81. 08/14 builtin/fsck: refactor to use `odb_for_each_object()`Patrick Steinhardt, Jan 21, 2026
  82. Taylor BlauJan 23, 2026
  83. Patrick SteinhardtJan 23, 2026
  84. 09/14 treewide: enumerate promisor objects via `odb_for_each_object()`Patrick Steinhardt, Jan 21, 2026
  85. Taylor BlauJan 23, 2026
  86. 10/14 treewide: drop uses of `for_each_{loose,packed}_object()`Patrick Steinhardt, Jan 21, 2026
  87. Taylor BlauJan 23, 2026
  88. Patrick SteinhardtJan 23, 2026
  89. 11/14 odb: introduce mtime fields for object info requestsPatrick Steinhardt, Jan 21, 2026
  90. Taylor BlauJan 23, 2026
  91. Patrick SteinhardtJan 23, 2026
  92. Taylor BlauJan 23, 2026
  93. Patrick SteinhardtJan 26, 2026
  94. 12/14 builtin/pack-objects: use `packfile_store_for_each_object()`Patrick Steinhardt, Jan 21, 2026
  95. Taylor BlauJan 23, 2026
  96. Patrick SteinhardtJan 23, 2026
  97. Taylor BlauJan 23, 2026
  98. Patrick SteinhardtJan 26, 2026
  99. Jeff KingJan 29, 2026
  100. Patrick SteinhardtJan 30, 2026
  101. 13/14 reachable: convert to use `odb_for_each_object()`Patrick Steinhardt, Jan 21, 2026
  102. 14/14 odb: drop unused `for_each_{loose,packed}_object()` functionsPatrick Steinhardt, Jan 21, 2026
  103. Taylor BlauJan 22, 2026
  104. Junio C HamanoJan 22, 2026
  105. 00/14 odb: introduce `odb_for_each_object()`Patrick Steinhardt, Jan 26, 2026
  106. 01/14 odb: rename `FOR_EACH_OBJECT_*` flagsPatrick Steinhardt, Jan 26, 2026
  107. 02/14 odb: fix flags parameter to be unsignedPatrick Steinhardt, Jan 26, 2026
  108. 03/14 object-file: extract function to read object info from pathPatrick Steinhardt, Jan 26, 2026
  109. 04/14 object-file: introduce function to iterate through objectsPatrick Steinhardt, Jan 26, 2026
  110. 05/14 packfile: extract function to iterate through objects of a storePatrick Steinhardt, Jan 26, 2026
  111. 06/14 packfile: introduce function to iterate through objectsPatrick Steinhardt, Jan 26, 2026
  112. 07/14 odb: introduce `odb_for_each_object()`Patrick Steinhardt, Jan 26, 2026
  113. 08/14 builtin/fsck: refactor to use `odb_for_each_object()`Patrick Steinhardt, Jan 26, 2026
  114. 09/14 treewide: enumerate promisor objects via `odb_for_each_object()`Patrick Steinhardt, Jan 26, 2026
  115. 10/14 treewide: drop uses of `for_each_{loose,packed}_object()`Patrick Steinhardt, Jan 26, 2026
  116. 11/14 odb: introduce mtime fields for object info requestsPatrick Steinhardt, Jan 26, 2026
  117. 12/14 builtin/pack-objects: use `packfile_store_for_each_object()`Patrick Steinhardt, Jan 26, 2026
  118. 13/14 reachable: convert to use `odb_for_each_object()`Patrick Steinhardt, Jan 26, 2026
  119. 14/14 odb: drop unused `for_each_{loose,packed}_object()` functionsPatrick Steinhardt, Jan 26, 2026
  120. Junio C HamanoFeb 20, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.