git/list[1] front-page[2] threads[3] people[4] search[5] about
wed 2026-10-07 17:27 UTC

[PATCH 09/17] odb/source: make `read_object_info()` function pluggable

From
Patrick Steinhardt <ps@pks.im>
Date
Feb 23, 2026, 16:18 UTC
Message-ID
<20260223-b4-pks-odb-source-pluggable-v1-9-253bac1db598@pks.im>
In-Reply-To
<20260223-b4-pks-odb-source-pluggable-v1-0-253bac1db598@pks.im>

Introduce a new callback function in `struct odb_source` to make the function pluggable.

Note that this function is a bit less straight-forward to convert compared to the other functions. The reason here is that the logic to read an object is:

  1. We try to read the object. If it exists we return it.
  2. If the object does not exist we reprepare the object database
     source.
  3. We then try reading the object info a second time in case the
     reprepare caused it to appear.

The second read is only supposed to happen for the packfile store though, as reading loose objects is not impacted by repreparing the object database.

Ideally, we'd just move this whole logic into the ODB source. But that's not easily possible because we try to avoid the reprepare unless really required, which is after we have found out that no other ODB source contains the object, either. So the logic spans across multiple ODB sources, and consequently we cannot move it into an individual source.

Instead, introduce a new flag `OBJECT_INFO_SECOND_READ` that tells the backend that we already tried to look up the object once, and that this time around the ODB source should try to find any new objects that may have surfaced due to an on-disk change.

With this flag, the "files" backend can trivially skip trying to re-read the object as a loose object. Furthermore, as we know that we only try the second read via the packfile store, we can skip repreparing loose objects and only reprepare the packfile store.

Signed-off-by: Patrick Steinhardt <ps@pks.im>
---
 object-file.c      | 12 ++++++++-
 odb.c              | 22 +++++++--------
 odb.h              | 24 -----------------
 odb/source-files.c | 15 +++++++++++
 odb/source.h       | 78 ++++++++++++++++++++++++++++++++++++++++++++++++++++++
 packfile.c         | 10 ++++++-
 6 files changed, 123 insertions(+), 38 deletions(-)
diff --git a/object-file.c b/object-file.c
index 25c1146849..eefde72c7d 100644
--- a/object-file.c
+++ b/object-file.c
@@ -543,9 +543,19 @@ static int read_object_info_from_path(struct odb_source *source,
 int odb_source_loose_read_object_info(struct odb_source *source,
 				      const struct object_id *oid,
 				      struct object_info *oi,
-				      unsigned flags)
+				      enum object_info_flags flags)
 {
 	static struct strbuf buf = STRBUF_INIT;
+
+	/*
+	 * The second read shouldn't cause new loose objects to show up, unless
+	 * there was a race condition with a secondary process. We don't care
+	 * about this case though, so we simply skip reading loose objects a
+	 * second time.
+	 */
+	if (flags & OBJECT_INFO_SECOND_READ)
+		return -1;
+
 	odb_loose_path(source, &buf, oid);
 	return read_object_info_from_path(source, buf.buf, oid, oi, flags);
 }
diff --git a/odb.c b/odb.c
index f7487eb0df..c0b8cd062b 100644
--- a/odb.c
+++ b/odb.c
@@ -688,22 +688,20 @@ static int do_oid_object_info_extended(struct object_database *odb,
 	while (1) {
 		struct odb_source *source;
 
-		/* Most likely it's a loose object. */
-		for (source = odb->sources; source; source = source->next) {
-			struct odb_source_files *files = odb_source_files_downcast(source);
-			if (!packfile_store_read_object_info(files->packed, real, oi, flags) ||
-			    !odb_source_loose_read_object_info(source, real, oi, flags))
+		for (source = odb->sources; source; source = source->next)
+			if (!odb_source_read_object_info(source, real, oi, flags))
 				return 0;
-		}
 
-		/* Not a loose object; someone else may have just packed it. */
+		/*
+		 * When the object hasn't been found we try a second read and
+		 * tell the sources so. This may cause them to invalidate
+		 * caches or reload on-disk state.
+		 */
 		if (!(flags & OBJECT_INFO_QUICK)) {
-			odb_reprepare(odb->repo->objects);
-			for (source = odb->sources; source; source = source->next) {
-				struct odb_source_files *files = odb_source_files_downcast(source);
-				if (!packfile_store_read_object_info(files->packed, real, oi, flags))
+			for (source = odb->sources; source; source = source->next)
+				if (!odb_source_read_object_info(source, real, oi,
+								 flags | OBJECT_INFO_SECOND_READ))
 					return 0;
-			}
 		}
 
 		/*
diff --git a/odb.h b/odb.h
index e13b5b7c44..70ffb033f9 100644
--- a/odb.h
+++ b/odb.h
@@ -339,30 +339,6 @@ struct object_info {
  */
 #define OBJECT_INFO_INIT { 0 }
 
-/* Flags that can be passed to `odb_read_object_info_extended()`. */
-enum object_info_flags {
-	/* Invoke lookup_replace_object() on the given hash. */
-	OBJECT_INFO_LOOKUP_REPLACE = (1 << 0),
-
-	/* Do not reprepare object sources when the first lookup has failed. */
-	OBJECT_INFO_QUICK = (1 << 1),
-
-	/*
-	 * Do not attempt to fetch the object if missing (even if fetch_is_missing is
-	 * nonzero).
-	 */
-	OBJECT_INFO_SKIP_FETCH_OBJECT = (1 << 2),
-
-	/* Die if object corruption (not just an object being missing) was detected. */
-	OBJECT_INFO_DIE_IF_CORRUPT = (1 << 3),
-
-	/*
-	 * This is meant for bulk prefetching of missing blobs in a partial
-	 * clone. Implies OBJECT_INFO_SKIP_FETCH_OBJECT and OBJECT_INFO_QUICK.
-	 */
-	OBJECT_INFO_FOR_PREFETCH = (OBJECT_INFO_SKIP_FETCH_OBJECT | OBJECT_INFO_QUICK),
-};
-
 /*
  * Read object info from the object database and populate the `object_info`
  * structure. Returns 0 on success, a negative error code otherwise.
diff --git a/odb/source-files.c b/odb/source-files.c
index 20a24f524a..f2969a1214 100644
--- a/odb/source-files.c
+++ b/odb/source-files.c
@@ -41,6 +41,20 @@ static void odb_source_files_reprepare(struct odb_source *source)
 	packfile_store_reprepare(files->packed);
 }
 
+static int odb_source_files_read_object_info(struct odb_source *source,
+					     const struct object_id *oid,
+					     struct object_info *oi,
+					     enum object_info_flags flags)
+{
+	struct odb_source_files *files = odb_source_files_downcast(source);
+
+	if (!packfile_store_read_object_info(files->packed, oid, oi, flags) ||
+	    !odb_source_loose_read_object_info(source, oid, oi, flags))
+		return 0;
+
+	return -1;
+}
+
 struct odb_source_files *odb_source_files_new(struct object_database *odb,
 					      const char *path,
 					      bool local)
@@ -55,6 +69,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,
 	files->base.free = odb_source_files_free;
 	files->base.close = odb_source_files_close;
 	files->base.reprepare = odb_source_files_reprepare;
+	files->base.read_object_info = odb_source_files_read_object_info;
 
 	/*
 	 * Ideally, we would only ever store absolute paths in the source. This
diff --git a/odb/source.h b/odb/source.h
index 7af4900ab4..45563de61e 100644
--- a/odb/source.h
+++ b/odb/source.h
@@ -13,6 +13,45 @@ enum odb_source_type {
 	ODB_SOURCE_FILES,
 };
 
+/* Flags that can be passed to `odb_read_object_info_extended()`. */
+enum object_info_flags {
+	/* Invoke lookup_replace_object() on the given hash. */
+	OBJECT_INFO_LOOKUP_REPLACE = (1 << 0),
+
+	/* Do not reprepare object sources when the first lookup has failed. */
+	OBJECT_INFO_QUICK = (1 << 1),
+
+	/*
+	 * Do not attempt to fetch the object if missing (even if fetch_is_missing is
+	 * nonzero).
+	 */
+	OBJECT_INFO_SKIP_FETCH_OBJECT = (1 << 2),
+
+	/* Die if object corruption (not just an object being missing) was detected. */
+	OBJECT_INFO_DIE_IF_CORRUPT = (1 << 3),
+
+	/*
+	 * We have already tried reading the object, but it couldn't be found
+	 * via any of the attached sources, and are now doing a second read.
+	 * This second read asks the individual sources to also evaluate
+	 * whether any on-disk state may have changed that may have caused the
+	 * object to appear.
+	 *
+	 * This flag is for internal use, only. The second read only occurs
+	 * when `OBJECT_INFO_QUICK` was not passed.
+	 */
+	OBJECT_INFO_SECOND_READ = (1 << 4),
+
+	/*
+	 * This is meant for bulk prefetching of missing blobs in a partial
+	 * clone. Implies OBJECT_INFO_SKIP_FETCH_OBJECT and OBJECT_INFO_QUICK.
+	 */
+	OBJECT_INFO_FOR_PREFETCH = (OBJECT_INFO_SKIP_FETCH_OBJECT | OBJECT_INFO_QUICK),
+};
+
+struct object_id;
+struct object_info;
+
 /*
  * The source is the part of the object database that stores the actual
  * objects. It thus encapsulates the logic to read and write the specific
@@ -73,6 +112,33 @@ struct odb_source {
 	 * example just been repacked so that new objects will become visible.
 	 */
 	void (*reprepare)(struct odb_source *source);
+
+	/*
+	 * This callback is expected to read object information from the object
+	 * database source. The object info will be partially populated with
+	 * pointers for each bit of information that was requested by the
+	 * caller.
+	 *
+	 * The flags field is a combination of `OBJECT_INFO` flags. Only the
+	 * following fields need to be handled by the backend:
+	 *
+	 *   - `OBJECT_INFO_QUICK` indicates it is fine to use caches without
+	 *     re-verifying the data.
+	 *
+	 *   - `OBJECT_INFO_SECOND_READ` indicates that the initial object
+	 *     lookup has failed and that the object sources should check
+	 *     whether any of its on-disk state has changed that may have
+	 *     caused the object to appear. Sources are free to ignore the
+	 *     second read in case they know that the first read would have
+	 *     already surfaced the object without reloading any on-disk state.
+	 *
+	 * The callback is expected to return a negative error code in case
+	 * reading the object has failed, 0 otherwise.
+	 */
+	int (*read_object_info)(struct odb_source *source,
+				const struct object_id *oid,
+				struct object_info *oi,
+				enum object_info_flags flags);
 };
 
 /*
@@ -132,4 +198,16 @@ static inline void odb_source_reprepare(struct odb_source *source)
 	source->reprepare(source);
 }
 
+/*
+ * Read an object from the object database source identified by its object ID.
+ * Returns 0 on success, a negative error code otherwise.
+ */
+static inline int odb_source_read_object_info(struct odb_source *source,
+					      const struct object_id *oid,
+					      struct object_info *oi,
+					      enum object_info_flags flags)
+{
+	return source->read_object_info(source, oid, oi, flags);
+}
+
 #endif
diff --git a/packfile.c b/packfile.c
index da1c0dfa39..71db10e7c6 100644
--- a/packfile.c
+++ b/packfile.c
@@ -2181,11 +2181,19 @@ int packfile_store_freshen_object(struct packfile_store *store,
 int packfile_store_read_object_info(struct packfile_store *store,
 				    const struct object_id *oid,
 				    struct object_info *oi,
-				    enum object_info_flags flags UNUSED)
+				    enum object_info_flags flags)
 {
 	struct pack_entry e;
 	int ret;
 
+	/*
+	 * In case the first read didn't surface the object, we have to reload
+	 * packfiles. This may cause us to discover new packfiles that have
+	 * been added since the last time we have prepared the packfile store.
+	 */
+	if (flags & OBJECT_INFO_SECOND_READ)
+		packfile_store_reprepare(store);
+
 	if (!find_pack_entry(store, oid, &e))
 		return 1;
 
-- 
2.53.0.536.g309c995771.dirty
Previous: Patrick SteinhardtNext: Patrick Steinhardt
Message 10 of 77 in “odb: make object database sources pluggable”
  1. 00/17 odb: make object database sources pluggablePatrick Steinhardt, Feb 23, 2026
  2. 01/17 odb: split `struct odb_source` into separate headerPatrick Steinhardt, Feb 23, 2026
  3. 02/17 odb: introduce "files" sourcePatrick Steinhardt, Feb 23, 2026
  4. 03/17 odb: embed base source in the "files" backendPatrick Steinhardt, Feb 23, 2026
  5. 04/17 odb: move reparenting logic into respective subsystemsPatrick Steinhardt, Feb 23, 2026
  6. 05/17 odb/source: introduce source type for robustnessPatrick Steinhardt, Feb 23, 2026
  7. 06/17 odb/source: make `free()` function pluggablePatrick Steinhardt, Feb 23, 2026
  8. 07/17 odb/source: make `reprepare()` function pluggablePatrick Steinhardt, Feb 23, 2026
  9. 08/17 odb/source: make `close()` function pluggablePatrick Steinhardt, Feb 23, 2026
  10. 09/17 odb/source: make `read_object_info()` function pluggablePatrick Steinhardt, Feb 23, 2026
  11. 10/17 odb/source: make `read_object_stream()` function pluggablePatrick Steinhardt, Feb 23, 2026
  12. 11/17 odb/source: make `for_each_object()` function pluggablePatrick Steinhardt, Feb 23, 2026
  13. 12/17 odb/source: make `freshen_object()` function pluggablePatrick Steinhardt, Feb 23, 2026
  14. 13/17 odb/source: make `write_object()` function pluggablePatrick Steinhardt, Feb 23, 2026
  15. 14/17 odb/source: make `write_object_stream()` function pluggablePatrick Steinhardt, Feb 23, 2026
  16. 15/17 odb/source: make `read_alternates()` function pluggablePatrick Steinhardt, Feb 23, 2026
  17. 16/17 odb/source: make `write_alternate()` function pluggablePatrick Steinhardt, Feb 23, 2026
  18. 17/17 odb/source: make `begin_transaction()` function pluggablePatrick Steinhardt, Feb 23, 2026
  19. Patrick SteinhardtFeb 23, 2026
  20. Junio C HamanoFeb 23, 2026
  21. Patrick SteinhardtFeb 24, 2026
  22. Justin ToblerMar 4, 2026
  23. Justin ToblerMar 4, 2026
  24. Justin ToblerMar 4, 2026
  25. Justin ToblerMar 4, 2026
  26. Justin ToblerMar 4, 2026
  27. Justin ToblerMar 4, 2026
  28. Justin ToblerMar 4, 2026
  29. Justin ToblerMar 4, 2026
  30. Justin ToblerMar 4, 2026
  31. Justin ToblerMar 4, 2026
  32. Justin ToblerMar 4, 2026
  33. Karthik NayakMar 5, 2026
  34. Karthik NayakMar 5, 2026
  35. Karthik NayakMar 5, 2026
  36. Karthik NayakMar 5, 2026
  37. Karthik NayakMar 5, 2026
  38. Karthik NayakMar 5, 2026
  39. Karthik NayakMar 5, 2026
  40. Patrick SteinhardtMar 5, 2026
  41. Karthik NayakMar 5, 2026
  42. Patrick SteinhardtMar 5, 2026
  43. Patrick SteinhardtMar 5, 2026
  44. Patrick SteinhardtMar 5, 2026
  45. Patrick SteinhardtMar 5, 2026
  46. Patrick SteinhardtMar 5, 2026
  47. Patrick SteinhardtMar 5, 2026
  48. Patrick SteinhardtMar 5, 2026
  49. Patrick SteinhardtMar 5, 2026
  50. Patrick SteinhardtMar 5, 2026
  51. Patrick SteinhardtMar 5, 2026
  52. Patrick SteinhardtMar 5, 2026
  53. Patrick SteinhardtMar 5, 2026
  54. 00/17 odb: make object database sources pluggablePatrick Steinhardt, Mar 5, 2026
  55. 01/17 odb: split `struct odb_source` into separate headerPatrick Steinhardt, Mar 5, 2026
  56. 02/17 odb: introduce "files" sourcePatrick Steinhardt, Mar 5, 2026
  57. 03/17 odb: embed base source in the "files" backendPatrick Steinhardt, Mar 5, 2026
  58. 04/17 odb: move reparenting logic into respective subsystemsPatrick Steinhardt, Mar 5, 2026
  59. 05/17 odb/source: introduce source type for robustnessPatrick Steinhardt, Mar 5, 2026
  60. 06/17 odb/source: make `free()` function pluggablePatrick Steinhardt, Mar 5, 2026
  61. 07/17 odb/source: make `reprepare()` function pluggablePatrick Steinhardt, Mar 5, 2026
  62. 08/17 odb/source: make `close()` function pluggablePatrick Steinhardt, Mar 5, 2026
  63. 09/17 odb/source: make `read_object_info()` function pluggablePatrick Steinhardt, Mar 5, 2026
  64. 10/17 odb/source: make `read_object_stream()` function pluggablePatrick Steinhardt, Mar 5, 2026
  65. 11/17 odb/source: make `for_each_object()` function pluggablePatrick Steinhardt, Mar 5, 2026
  66. 12/17 odb/source: make `freshen_object()` function pluggablePatrick Steinhardt, Mar 5, 2026
  67. 13/17 odb/source: make `write_object()` function pluggablePatrick Steinhardt, Mar 5, 2026
  68. 14/17 odb/source: make `write_object_stream()` function pluggablePatrick Steinhardt, Mar 5, 2026
  69. 15/17 odb/source: make `read_alternates()` function pluggablePatrick Steinhardt, Mar 5, 2026
  70. 16/17 odb/source: make `write_alternate()` function pluggablePatrick Steinhardt, Mar 5, 2026
  71. 17/17 odb/source: make `begin_transaction()` function pluggablePatrick Steinhardt, Mar 5, 2026
  72. Justin ToblerMar 5, 2026
  73. Justin ToblerMar 5, 2026
  74. Justin ToblerMar 5, 2026
  75. Justin ToblerMar 5, 2026
  76. Junio C HamanoMar 5, 2026
  77. Patrick SteinhardtMar 10, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.