{"thread":{"id":"65423","subject":"[PATCH 00/16] odb: introduce \"inmemory\" source","startedAt":"2026-04-03T06:02:17Z","lastAt":"2026-04-14T08:45:18Z","messageCount":85,"participants":["Patrick Steinhardt","Junio C Hamano","Justin Tobler","Karthik Nayak"],"isPatch":true,"patchVersion":1,"patchTotal":16},"messages":[{"id":"540817","messageId":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":null,"subject":"[PATCH 00/16] odb: introduce \"inmemory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:47Z","receivedAt":"2026-04-03T06:02:17Z","isPatch":true,"body":"Hi,\n\nthis patch series introduces the second object database source type,\nwhich is the \"inmemory\" source.\n\nThis source may seem somewhat odd at first: it always starts out empty,\nand any object written into it will only exist in memory until the\nprocess exits. But the source already serves a purpose in our codebase,\nwhere some commands, for example git-blame(1), write an in-memory\nworktree commit.\n\nFurthermore, I think that going forward it can serve more purposes as we\nnow have an easy way to write and read objects that will not get\npersisted. I could see that this may be useful when for example\nre-merging diffs. But eventually, once we have the object storage format\nextension wired up, callers might even want to manually set up an\nin-memory database as the primary ODB for write operations so that no\ndata will be persisted in an arbitrary write.\n\nLast but not least, this patch series also serves the purpose of\neventually getting rid of the `struct object_info::whence` member.\nInstead, we'll simply yield the ODB source a specific object has been\nread from, together with some backend-specific data, which gives\nstrictly more information compared to the status quo.\n\nThe series is based on cf2139f8e1 (The 24th batch, 2026-04-01) with\nps/odb-cleanup at 109bcb7d1d (odb: drop unneeded headers and forward\ndecls, 2026-04-01) merged into it.\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (16):\n      odb: introduce \"inmemory\" source\n      odb/source-inmemory: implement `free()` callback\n      odb: fix unnecessary call to `find_cached_object()`\n      odb/source-inmemory: implement `read_object_info()` callback\n      odb/source-inmemory: implement `read_object_stream()` callback\n      odb/source-inmemory: implement `write_object()` callback\n      odb/source-inmemory: implement `write_object_stream()` callback\n      cbtree: allow using arbitrary wrapper structures for nodes\n      oidtree: add ability to store data\n      odb/source-inmemory: convert to use oidtree\n      odb/source-inmemory: implement `for_each_object()` callback\n      odb/source-inmemory: implement `find_abbrev_len()` callback\n      odb/source-inmemory: implement `count_objects()` callback\n      odb/source-inmemory: implement `freshen_object()` callback\n      odb/source-inmemory: stub out remaining functions\n      odb: generic inmemory source\n\n Makefile                 |   1 +\n cbtree.c                 |  25 +++-\n cbtree.h                 |  11 +-\n loose.c                  |   2 +-\n meson.build              |   1 +\n object-file.c            |   3 +-\n odb.c                    |  82 ++---------\n odb.h                    |   4 +-\n odb/source-inmemory.c    | 375 +++++++++++++++++++++++++++++++++++++++++++++++\n odb/source-inmemory.h    |  33 +++++\n odb/source.h             |   3 +\n oidtree.c                |  66 ++++++---\n oidtree.h                |  12 +-\n t/unit-tests/u-oidtree.c |  26 +++-\n 14 files changed, 529 insertions(+), 115 deletions(-)\n\n\n---\nbase-commit: 3d05c3e2906489caa9f12f0af18dc233a6b8032c\nchange-id: 20260401-b4-pks-odb-source-inmemory-7b17c83d9e43\n\n"},{"id":"540818","messageId":"20260403-b4-pks-odb-source-inmemory-v1-1-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 01/16] odb: introduce \"inmemory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:48Z","receivedAt":"2026-04-03T06:02:20Z","isPatch":true,"body":"Next to our typical object database sources, each object database also\nhas an implicit source of \"cached\" objects. These cached objects only\nexist in memory and some use cases:\n\n  - They contain evergreen objects that we expect to always exist, like\n    for example the empty tree.\n\n  - They can be used to store temporary objects that we don't want to\n    persist to disk.\n\nOverall, their use is somewhat restricted though. For example, we don't\nprovide the ability to use it as a temporary object database source that\nallows the user to write objects, but discard them after Git exists. So\nwhile these cached objects behave almost like a source, they aren't used\nas one.\n\nThis is about to change over the following commits, where we will turn\ncached objects into a new \"inmemory\" source. This will allow us to use\nit exactly the same as any other source by providing the same common\ninterface as the \"files\" source.\n\nFor now, the inmemory source only hosts the cached objects and doesn't\nprovide any logic yet. This will change with subsequent commits, where\nwe move respective functionality into the source.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n Makefile              |  1 +\n meson.build           |  1 +\n odb.c                 | 21 +++++++++++++--------\n odb.h                 |  4 ++--\n odb/source-inmemory.c | 12 ++++++++++++\n odb/source-inmemory.h | 35 +++++++++++++++++++++++++++++++++++\n odb/source.h          |  3 +++\n 7 files changed, 67 insertions(+), 10 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex dbf0022054..175391e6f8 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1218,6 +1218,7 @@ LIB_OBJS += object.o\n LIB_OBJS += odb.o\n LIB_OBJS += odb/source.o\n LIB_OBJS += odb/source-files.o\n+LIB_OBJS += odb/source-inmemory.o\n LIB_OBJS += odb/streaming.o\n LIB_OBJS += oid-array.o\n LIB_OBJS += oidmap.o\ndiff --git a/meson.build b/meson.build\nindex 8309942d18..8f55d2650e 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -404,6 +404,7 @@ libgit_sources = [\n   'odb.c',\n   'odb/source.c',\n   'odb/source-files.c',\n+  'odb/source-inmemory.c',\n   'odb/streaming.c',\n   'oid-array.c',\n   'oidmap.c',\ndiff --git a/odb.c b/odb.c\nindex 9b28fe25ef..95b21e2cfd 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -14,6 +14,7 @@\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/source-inmemory.h\"\n #include \"packfile.h\"\n #include \"path.h\"\n #include \"promisor-remote.h\"\n@@ -53,9 +54,9 @@ static const struct cached_object *find_cached_object(struct object_database *ob\n \t\t.type = OBJ_TREE,\n \t\t.buf = \"\",\n \t};\n-\tconst struct cached_object_entry *co = object_store->cached_objects;\n+\tconst struct cached_object_entry *co = object_store->inmemory_objects->objects;\n \n-\tfor (size_t i = 0; i < object_store->cached_object_nr; i++, co++)\n+\tfor (size_t i = 0; i < object_store->inmemory_objects->objects_nr; i++, co++)\n \t\tif (oideq(&co->oid, oid))\n \t\t\treturn &co->value;\n \n@@ -792,9 +793,10 @@ int odb_pretend_object(struct object_database *odb,\n \t    find_cached_object(odb, oid))\n \t\treturn 0;\n \n-\tALLOC_GROW(odb->cached_objects,\n-\t\t   odb->cached_object_nr + 1, odb->cached_object_alloc);\n-\tco = &odb->cached_objects[odb->cached_object_nr++];\n+\tALLOC_GROW(odb->inmemory_objects->objects,\n+\t\t   odb->inmemory_objects->objects_nr + 1,\n+\t\t   odb->inmemory_objects->objects_alloc);\n+\tco = &odb->inmemory_objects->objects[odb->inmemory_objects->objects_nr++];\n \tco->value.size = len;\n \tco->value.type = type;\n \tco_buf = xmalloc(len);\n@@ -1083,6 +1085,7 @@ struct object_database *odb_new(struct repository *repo,\n \to->sources = odb_source_new(o, primary_source, true);\n \to->sources_tail = &o->sources->next;\n \to->alternate_db = xstrdup_or_null(secondary_sources);\n+\to->inmemory_objects = odb_source_inmemory_new(o);\n \n \tfree(to_free);\n \n@@ -1123,9 +1126,11 @@ void odb_free(struct object_database *o)\n \todb_close(o);\n \todb_free_sources(o);\n \n-\tfor (size_t i = 0; i < o->cached_object_nr; i++)\n-\t\tfree((char *) o->cached_objects[i].value.buf);\n-\tfree(o->cached_objects);\n+\tfor (size_t i = 0; i < o->inmemory_objects->objects_nr; i++)\n+\t\tfree((char *) o->inmemory_objects->objects[i].value.buf);\n+\tfree(o->inmemory_objects->objects);\n+\tfree(o->inmemory_objects->base.path);\n+\tfree(o->inmemory_objects);\n \n \tstring_list_clear(&o->submodule_source_paths, 0);\n \ndiff --git a/odb.h b/odb.h\nindex 3a711f6547..3d20270a05 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -8,6 +8,7 @@\n #include \"thread-utils.h\"\n \n struct cached_object_entry;\n+struct odb_source_inmemory;\n struct packed_git;\n struct repository;\n struct strbuf;\n@@ -98,8 +99,7 @@ struct object_database {\n \t * to write them into the object store (e.g. a browse-only\n \t * application).\n \t */\n-\tstruct cached_object_entry *cached_objects;\n-\tsize_t cached_object_nr, cached_object_alloc;\n+\tstruct odb_source_inmemory *inmemory_objects;\n \n \t/*\n \t * A fast, rough count of the number of objects in the repository.\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nnew file mode 100644\nindex 0000000000..c7ac5c24f0\n--- /dev/null\n+++ b/odb/source-inmemory.c\n@@ -0,0 +1,12 @@\n+#include \"git-compat-util.h\"\n+#include \"odb/source-inmemory.h\"\n+\n+struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n+{\n+\tstruct odb_source_inmemory *source;\n+\n+\tCALLOC_ARRAY(source, 1);\n+\todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n+\n+\treturn source;\n+}\ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nnew file mode 100644\nindex 0000000000..95477bf36d\n--- /dev/null\n+++ b/odb/source-inmemory.h\n@@ -0,0 +1,35 @@\n+#ifndef ODB_SOURCE_INMEMORY_H\n+#define ODB_SOURCE_INMEMORY_H\n+\n+#include \"odb/source.h\"\n+\n+struct cached_object_entry;\n+\n+/*\n+ * An inmemory source that you can write objects to that shall be made\n+ * available for reading, but that shouldn't ever be persisted to disk. Note\n+ * that any objects written to this source will be stored in memory, so the\n+ * number of objects you can store is limited by available system memory.\n+ */\n+struct odb_source_inmemory {\n+\tstruct odb_source base;\n+\n+\tstruct cached_object_entry *objects;\n+\tsize_t objects_nr, objects_alloc;\n+};\n+\n+/* Create a new in-memory object database source. */\n+struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb);\n+\n+/*\n+ * Cast the given object database source to the inmemory backend. This will\n+ * cause a BUG in case the source doesn't use this backend.\n+ */\n+static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n+{\n+\tif (source->type != ODB_SOURCE_INMEMORY)\n+\t\tBUG(\"trying to downcast source of type '%d' to inmemory\", source->type);\n+\treturn container_of(source, struct odb_source_inmemory, base);\n+}\n+\n+#endif\ndiff --git a/odb/source.h b/odb/source.h\nindex f706e0608a..cd14f9e046 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -13,6 +13,9 @@ enum odb_source_type {\n \n \t/* The \"files\" backend that uses loose objects and packfiles. */\n \tODB_SOURCE_FILES,\n+\n+\t/* The \"inmemory\" backend that stores objects in memory. */\n+\tODB_SOURCE_INMEMORY,\n };\n \n struct object_id;\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540820","messageId":"20260403-b4-pks-odb-source-inmemory-v1-2-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 02/16] odb/source-inmemory: implement `free()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:49Z","receivedAt":"2026-04-03T06:02:23Z","isPatch":true,"body":"Implement the `free()` callback function for the \"inmemory\" source.\n\nNote that this requires us to define `struct cached_object_entry` in\n\"odb/source-inmemory.h\", as it is accessed in both \"odb.c\" and\n\"odb/source-inmemory.c\" now. This will be fixed in subsequent commits\nthough.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                 | 25 ++++---------------------\n odb/source-inmemory.c | 12 ++++++++++++\n odb/source-inmemory.h |  9 ++++++++-\n 3 files changed, 24 insertions(+), 22 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 95b21e2cfd..d321242353 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -32,21 +32,6 @@\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct odb_source *, 1, fspathhash, fspatheq)\n \n-/*\n- * This is meant to hold a *small* number of objects that you would\n- * want odb_read_object() to be able to return, but yet you do not want\n- * to write them into the object store (e.g. a browse-only\n- * application).\n- */\n-struct cached_object_entry {\n-\tstruct object_id oid;\n-\tstruct cached_object {\n-\t\tenum object_type type;\n-\t\tconst void *buf;\n-\t\tunsigned long size;\n-\t} value;\n-};\n-\n static const struct cached_object *find_cached_object(struct object_database *object_store,\n \t\t\t\t\t\t      const struct object_id *oid)\n {\n@@ -1109,6 +1094,10 @@ static void odb_free_sources(struct object_database *o)\n \t\todb_source_free(o->sources);\n \t\to->sources = next;\n \t}\n+\n+\todb_source_free(&o->inmemory_objects->base);\n+\to->inmemory_objects = NULL;\n+\n \tkh_destroy_odb_path_map(o->source_by_path);\n \to->source_by_path = NULL;\n }\n@@ -1126,12 +1115,6 @@ void odb_free(struct object_database *o)\n \todb_close(o);\n \todb_free_sources(o);\n \n-\tfor (size_t i = 0; i < o->inmemory_objects->objects_nr; i++)\n-\t\tfree((char *) o->inmemory_objects->objects[i].value.buf);\n-\tfree(o->inmemory_objects->objects);\n-\tfree(o->inmemory_objects->base.path);\n-\tfree(o->inmemory_objects);\n-\n \tstring_list_clear(&o->submodule_source_paths, 0);\n \n \tfree(o);\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex c7ac5c24f0..ccbb622eae 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,6 +1,16 @@\n #include \"git-compat-util.h\"\n #include \"odb/source-inmemory.h\"\n \n+static void odb_source_inmemory_free(struct odb_source *source)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tfor (size_t i = 0; i < inmemory->objects_nr; i++)\n+\t\tfree((char *) inmemory->objects[i].value.buf);\n+\tfree(inmemory->objects);\n+\tfree(inmemory->base.path);\n+\tfree(inmemory);\n+}\n+\n struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n {\n \tstruct odb_source_inmemory *source;\n@@ -8,5 +18,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tCALLOC_ARRAY(source, 1);\n \todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n \n+\tsource->base.free = odb_source_inmemory_free;\n+\n \treturn source;\n }\ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex 95477bf36d..14dc06f7c3 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -3,7 +3,14 @@\n \n #include \"odb/source.h\"\n \n-struct cached_object_entry;\n+struct cached_object_entry {\n+\tstruct object_id oid;\n+\tstruct cached_object {\n+\t\tenum object_type type;\n+\t\tconst void *buf;\n+\t\tunsigned long size;\n+\t} value;\n+};\n \n /*\n  * An inmemory source that you can write objects to that shall be made\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540819","messageId":"20260403-b4-pks-odb-source-inmemory-v1-3-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 03/16] odb: fix unnecessary call to `find_cached_object()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:50Z","receivedAt":"2026-04-03T06:02:26Z","isPatch":true,"body":"The function `odb_pretend_object()` writes an object into the in-memory\nobject database source. The effect of this is that the object will now\nbecome readable, but it won't ever be persisted to disk.\n\nBefore storing the object, we first verify whether the object already\nexists. This is done by calling `odb_has_object()` to check all sources,\nfollowed by `find_cached_object()` to check whether we have already\nstored the object in our in-memory source.\n\nThis is unnecessary though, as `odb_has_object()` already checks the\nin-memory source transitively via:\n\n  - `odb_has_object()`\n  - `odb_read_object_info_extended()`\n  - `do_oid_object_info_extended()`\n  - `find_cached_object()`\n\nDrop the explicit call to `find_cached_object()`.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c | 3 +--\n 1 file changed, 1 insertion(+), 2 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex d321242353..21cdedc31c 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -774,8 +774,7 @@ int odb_pretend_object(struct object_database *odb,\n \tchar *co_buf;\n \n \thash_object_file(odb->repo->hash_algo, buf, len, type, oid);\n-\tif (odb_has_object(odb, oid, 0) ||\n-\t    find_cached_object(odb, oid))\n+\tif (odb_has_object(odb, oid, 0))\n \t\treturn 0;\n \n \tALLOC_GROW(odb->inmemory_objects->objects,\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540821","messageId":"20260403-b4-pks-odb-source-inmemory-v1-4-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 04/16] odb/source-inmemory: implement `read_object_info()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:51Z","receivedAt":"2026-04-03T06:02:28Z","isPatch":true,"body":"Implement the `read_object_info()` callback function for the inmemory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                 | 39 +------------------------------------\n odb/source-inmemory.c | 53 +++++++++++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 54 insertions(+), 38 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 21cdedc31c..b8e7356951 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -32,25 +32,6 @@\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct odb_source *, 1, fspathhash, fspatheq)\n \n-static const struct cached_object *find_cached_object(struct object_database *object_store,\n-\t\t\t\t\t\t      const struct object_id *oid)\n-{\n-\tstatic const struct cached_object empty_tree = {\n-\t\t.type = OBJ_TREE,\n-\t\t.buf = \"\",\n-\t};\n-\tconst struct cached_object_entry *co = object_store->inmemory_objects->objects;\n-\n-\tfor (size_t i = 0; i < object_store->inmemory_objects->objects_nr; i++, co++)\n-\t\tif (oideq(&co->oid, oid))\n-\t\t\treturn &co->value;\n-\n-\tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n-\t\treturn &empty_tree;\n-\n-\treturn NULL;\n-}\n-\n int odb_mkstemp(struct object_database *odb,\n \t\tstruct strbuf *temp_filename, const char *pattern)\n {\n@@ -570,7 +551,6 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \t\t\t\t       const struct object_id *oid,\n \t\t\t\t       struct object_info *oi, unsigned flags)\n {\n-\tconst struct cached_object *co;\n \tconst struct object_id *real = oid;\n \tint already_retried = 0;\n \n@@ -580,25 +560,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \tif (is_null_oid(real))\n \t\treturn -1;\n \n-\tco = find_cached_object(odb, real);\n-\tif (co) {\n-\t\tif (oi) {\n-\t\t\tif (oi->typep)\n-\t\t\t\t*(oi->typep) = co->type;\n-\t\t\tif (oi->sizep)\n-\t\t\t\t*(oi->sizep) = co->size;\n-\t\t\tif (oi->disk_sizep)\n-\t\t\t\t*(oi->disk_sizep) = 0;\n-\t\t\tif (oi->delta_base_oid)\n-\t\t\t\toidclr(oi->delta_base_oid, odb->repo->hash_algo);\n-\t\t\tif (oi->contentp)\n-\t\t\t\t*oi->contentp = xmemdupz(co->buf, co->size);\n-\t\t\tif (oi->mtimep)\n-\t\t\t\t*oi->mtimep = 0;\n-\t\t\toi->whence = OI_CACHED;\n-\t\t}\n+\tif (!odb_source_read_object_info(&odb->inmemory_objects->base, oid, oi, flags))\n \t\treturn 0;\n-\t}\n \n \todb_prepare_alternates(odb);\n \ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex ccbb622eae..12c80f9b34 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,5 +1,57 @@\n #include \"git-compat-util.h\"\n+#include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n+#include \"repository.h\"\n+\n+static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n+\t\t\t\t\t\t      const struct object_id *oid)\n+{\n+\tstatic const struct cached_object empty_tree = {\n+\t\t.type = OBJ_TREE,\n+\t\t.buf = \"\",\n+\t};\n+\tconst struct cached_object_entry *co = source->objects;\n+\n+\tfor (size_t i = 0; i < source->objects_nr; i++, co++)\n+\t\tif (oideq(&co->oid, oid))\n+\t\t\treturn &co->value;\n+\n+\tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n+\t\treturn &empty_tree;\n+\n+\treturn NULL;\n+}\n+\n+static int odb_source_inmemory_read_object_info(struct odb_source *source,\n+\t\t\t\t\t\tconst struct object_id *oid,\n+\t\t\t\t\t\tstruct object_info *oi,\n+\t\t\t\t\t\tenum object_info_flags flags UNUSED)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tconst struct cached_object *object;\n+\n+\tobject = find_cached_object(inmemory, oid);\n+\tif (!object)\n+\t\treturn -1;\n+\n+\tif (oi) {\n+\t\tif (oi->typep)\n+\t\t\t*(oi->typep) = object->type;\n+\t\tif (oi->sizep)\n+\t\t\t*(oi->sizep) = object->size;\n+\t\tif (oi->disk_sizep)\n+\t\t\t*(oi->disk_sizep) = 0;\n+\t\tif (oi->delta_base_oid)\n+\t\t\toidclr(oi->delta_base_oid, source->odb->repo->hash_algo);\n+\t\tif (oi->contentp)\n+\t\t\t*oi->contentp = xmemdupz(object->buf, object->size);\n+\t\tif (oi->mtimep)\n+\t\t\t*oi->mtimep = 0;\n+\t\toi->whence = OI_CACHED;\n+\t}\n+\n+\treturn 0;\n+}\n \n static void odb_source_inmemory_free(struct odb_source *source)\n {\n@@ -19,6 +71,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n \n \tsource->base.free = odb_source_inmemory_free;\n+\tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \n \treturn source;\n }\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540822","messageId":"20260403-b4-pks-odb-source-inmemory-v1-5-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 05/16] odb/source-inmemory: implement `read_object_stream()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:52Z","receivedAt":"2026-04-03T06:02:31Z","isPatch":true,"body":"Implement the `read_object_stream()` callback function for the inmemory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 50 ++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 50 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 12c80f9b34..4a68169430 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,6 +1,7 @@\n #include \"git-compat-util.h\"\n #include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n+#include \"odb/streaming.h\"\n #include \"repository.h\"\n \n static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n@@ -53,6 +54,54 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \treturn 0;\n }\n \n+struct odb_read_stream_inmemory {\n+\tstruct odb_read_stream base;\n+\tconst void *buf;\n+\tsize_t offset;\n+};\n+\n+static ssize_t odb_read_stream_inmemory_read(struct odb_read_stream *stream,\n+\t\t\t\t\t     char *buf, size_t buf_len)\n+{\n+\tstruct odb_read_stream_inmemory *inmemory =\n+\t\tcontainer_of(stream, struct odb_read_stream_inmemory, base);\n+\tsize_t bytes = buf_len;\n+\n+\tif (buf_len > inmemory->base.size - inmemory->offset)\n+\t\tbytes = inmemory->base.size - inmemory->offset;\n+\tmemcpy(buf, inmemory->buf, bytes);\n+\n+\treturn bytes;\n+}\n+\n+static int odb_read_stream_inmemory_close(struct odb_read_stream *stream UNUSED)\n+{\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t\t\t  struct odb_source *source,\n+\t\t\t\t\t\t  const struct object_id *oid)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tstruct odb_read_stream_inmemory *stream;\n+\tconst struct cached_object *object;\n+\n+\tobject = find_cached_object(inmemory, oid);\n+\tif (!object)\n+\t\treturn -1;\n+\n+\tCALLOC_ARRAY(stream, 1);\n+\tstream->base.read = odb_read_stream_inmemory_read;\n+\tstream->base.close = odb_read_stream_inmemory_close;\n+\tstream->base.size = object->size;\n+\tstream->base.type = object->type;\n+\tstream->buf = object->buf;\n+\n+\t*out = &stream->base;\n+\treturn 0;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n@@ -72,6 +121,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \n \tsource->base.free = odb_source_inmemory_free;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n+\tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \n \treturn source;\n }\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540824","messageId":"20260403-b4-pks-odb-source-inmemory-v1-6-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 06/16] odb/source-inmemory: implement `write_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:53Z","receivedAt":"2026-04-03T06:02:34Z","isPatch":true,"body":"Implement the `write_object()` callback function for the inmemory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                 | 16 ++--------------\n odb/source-inmemory.c | 22 ++++++++++++++++++++++\n 2 files changed, 24 insertions(+), 14 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex b8e7356951..34228c0cd5 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -733,24 +733,12 @@ int odb_pretend_object(struct object_database *odb,\n \t\t       void *buf, unsigned long len, enum object_type type,\n \t\t       struct object_id *oid)\n {\n-\tstruct cached_object_entry *co;\n-\tchar *co_buf;\n-\n \thash_object_file(odb->repo->hash_algo, buf, len, type, oid);\n \tif (odb_has_object(odb, oid, 0))\n \t\treturn 0;\n \n-\tALLOC_GROW(odb->inmemory_objects->objects,\n-\t\t   odb->inmemory_objects->objects_nr + 1,\n-\t\t   odb->inmemory_objects->objects_alloc);\n-\tco = &odb->inmemory_objects->objects[odb->inmemory_objects->objects_nr++];\n-\tco->value.size = len;\n-\tco->value.type = type;\n-\tco_buf = xmalloc(len);\n-\tmemcpy(co_buf, buf, len);\n-\tco->value.buf = co_buf;\n-\toidcpy(&co->oid, oid);\n-\treturn 0;\n+\treturn odb_source_write_object(&odb->inmemory_objects->base,\n+\t\t\t\t       buf, len, type, oid, NULL, 0);\n }\n \n void *odb_read_object(struct object_database *odb,\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 4a68169430..d2fc4c4054 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -102,6 +102,27 @@ static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n \treturn 0;\n }\n \n+static int odb_source_inmemory_write_object(struct odb_source *source,\n+\t\t\t\t\t    const void *buf, unsigned long len,\n+\t\t\t\t\t    enum object_type type,\n+\t\t\t\t\t    struct object_id *oid,\n+\t\t\t\t\t    struct object_id *compat_oid UNUSED,\n+\t\t\t\t\t    enum odb_write_object_flags flags UNUSED)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tstruct cached_object_entry *object;\n+\n+\tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n+\t\t   inmemory->objects_alloc);\n+\tobject = &inmemory->objects[inmemory->objects_nr++];\n+\tobject->value.size = len;\n+\tobject->value.type = type;\n+\tobject->value.buf = xmemdupz(buf, len);\n+\toidcpy(&object->oid, oid);\n+\n+\treturn 0;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n@@ -122,6 +143,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.free = odb_source_inmemory_free;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n+\tsource->base.write_object = odb_source_inmemory_write_object;\n \n \treturn source;\n }\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540823","messageId":"20260403-b4-pks-odb-source-inmemory-v1-7-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 07/16] odb/source-inmemory: implement `write_object_stream()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:54Z","receivedAt":"2026-04-03T06:02:36Z","isPatch":true,"body":"Implement the `write_object_stream()` callback function for the inmemory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 40 ++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 40 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex d2fc4c4054..890e2a8c7c 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -123,6 +123,45 @@ static int odb_source_inmemory_write_object(struct odb_source *source,\n \treturn 0;\n }\n \n+static int odb_source_inmemory_write_object_stream(struct odb_source *source,\n+\t\t\t\t\t\t   struct odb_write_stream *stream,\n+\t\t\t\t\t\t   size_t len,\n+\t\t\t\t\t\t   struct object_id *oid)\n+{\n+\tsize_t total_read = 0;\n+\tchar *data;\n+\tint ret;\n+\n+\tCALLOC_ARRAY(data, len);\n+\twhile (!stream->is_finished) {\n+\t\tunsigned long bytes_read;\n+\t\tconst void *in;\n+\n+\t\tin = stream->read(stream, &bytes_read);\n+\t\tif (total_read + bytes_read > len) {\n+\t\t\tret = error(\"object stream yielded more bytes than expected\");\n+\t\t\tgoto out;\n+\t\t}\n+\n+\t\tmemcpy(data, in, bytes_read);\n+\t\ttotal_read += bytes_read;\n+\t}\n+\n+\tif (total_read != len) {\n+\t\tret = error(\"object stream yielded less bytes than expected\");\n+\t\tgoto out;\n+\t}\n+\n+\tret = odb_source_inmemory_write_object(source, data, len, OBJ_BLOB, oid,\n+\t\t\t\t\t       NULL, 0);\n+\tif (ret < 0)\n+\t\tgoto out;\n+\n+out:\n+\tfree(data);\n+\treturn ret;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n@@ -144,6 +183,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n+\tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n \treturn source;\n }\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540825","messageId":"20260403-b4-pks-odb-source-inmemory-v1-8-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 08/16] cbtree: allow using arbitrary wrapper structures for nodes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:55Z","receivedAt":"2026-04-03T06:02:39Z","isPatch":true,"body":"The cbtree subsystem allows the user to store arbitrary data in a\nprefix-free set of strings. This is used by us to store object IDs in a\nway that we can easily iterate through them in lexicograph order, and so\nthat we can easily perform lookups with shortened object IDs.\n\nIn its current form, it is not easily possible to store arbitrary data\nwith the tree nodes. There are a couple of approaches such a caller\ncould try to use, but none of them really work:\n\n  - One may embed the `struct cb_node` in a custom structure. This does\n    not work though as `struct cb_node` contains a flex array, and\n    embedding such a struct in another struct is forbidden.\n\n  - One may use a `union` over `struct cb_node` and ones own data type,\n    which _is_ allowed even if the struct contains a flex array. This\n    does not work though, as the compiler may align members of the\n    struct so that the node key would not immediately start where the\n    flex array starts.\n\n  - One may allocate `struct cb_node` such that it has room for both its\n    key and the custom data. This has the downside though that if the\n    custom data is itself a pointer to allocated memory, then the leak\n    checker will not consider the pointer to be alive anymore.\n\nRefactor the cbtree to drop the flex array and instead take in an\nexplicit offset for where to find the key, which allows the caller to\nembed `struct cb_node` is a wrapper struct.\n\nNote that this change has the downside that we now have a bit of padding\nin our structure, which grows the size from 60 to 64 bytes on a 64 bit\nsystem. On the other hand though, it allows us to get rid of the memory\ncopies that we previously had to do to ensure proper alignment. This\nseems like a reasonable tradeoff.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n cbtree.c  | 25 ++++++++++++++++++-------\n cbtree.h  | 11 ++++++-----\n oidtree.c | 33 ++++++++++++++-------------------\n 3 files changed, 38 insertions(+), 31 deletions(-)\n\ndiff --git a/cbtree.c b/cbtree.c\nindex 4ab794bddc..8f5edbb80a 100644\n--- a/cbtree.c\n+++ b/cbtree.c\n@@ -7,6 +7,11 @@\n #include \"git-compat-util.h\"\n #include \"cbtree.h\"\n \n+static inline uint8_t *cb_node_key(struct cb_tree *t, struct cb_node *node)\n+{\n+\treturn (uint8_t *) node + t->key_offset;\n+}\n+\n static struct cb_node *cb_node_of(const void *p)\n {\n \treturn (struct cb_node *)((uintptr_t)p - 1);\n@@ -33,6 +38,7 @@ struct cb_node *cb_insert(struct cb_tree *t, struct cb_node *node, size_t klen)\n \tuint8_t c;\n \tint newdirection;\n \tstruct cb_node **wherep, *p;\n+\tuint8_t *node_key, *p_key;\n \n \tassert(!((uintptr_t)node & 1)); /* allocations must be aligned */\n \n@@ -41,23 +47,26 @@ struct cb_node *cb_insert(struct cb_tree *t, struct cb_node *node, size_t klen)\n \t\treturn NULL;\t/* success */\n \t}\n \n+\tnode_key = cb_node_key(t, node);\n+\n \t/* see if a node already exists */\n-\tp = cb_internal_best_match(t->root, node->k, klen);\n+\tp = cb_internal_best_match(t->root, node_key, klen);\n+\tp_key = cb_node_key(t, p);\n \n \t/* find first differing byte */\n \tfor (newbyte = 0; newbyte < klen; newbyte++) {\n-\t\tif (p->k[newbyte] != node->k[newbyte])\n+\t\tif (p_key[newbyte] != node_key[newbyte])\n \t\t\tgoto different_byte_found;\n \t}\n \treturn p;\t/* element exists, let user deal with it */\n \n different_byte_found:\n-\tnewotherbits = p->k[newbyte] ^ node->k[newbyte];\n+\tnewotherbits = p_key[newbyte] ^ node_key[newbyte];\n \tnewotherbits |= newotherbits >> 1;\n \tnewotherbits |= newotherbits >> 2;\n \tnewotherbits |= newotherbits >> 4;\n \tnewotherbits = (newotherbits & ~(newotherbits >> 1)) ^ 255;\n-\tc = p->k[newbyte];\n+\tc = p_key[newbyte];\n \tnewdirection = (1 + (newotherbits | c)) >> 8;\n \n \tnode->byte = newbyte;\n@@ -78,7 +87,7 @@ struct cb_node *cb_insert(struct cb_tree *t, struct cb_node *node, size_t klen)\n \t\t\tbreak;\n \t\tif (q->byte == newbyte && q->otherbits > newotherbits)\n \t\t\tbreak;\n-\t\tc = q->byte < klen ? node->k[q->byte] : 0;\n+\t\tc = q->byte < klen ? node_key[q->byte] : 0;\n \t\tdirection = (1 + (q->otherbits | c)) >> 8;\n \t\twherep = q->child + direction;\n \t}\n@@ -93,7 +102,7 @@ struct cb_node *cb_lookup(struct cb_tree *t, const uint8_t *k, size_t klen)\n {\n \tstruct cb_node *p = cb_internal_best_match(t->root, k, klen);\n \n-\treturn p && !memcmp(p->k, k, klen) ? p : NULL;\n+\treturn p && !memcmp(cb_node_key(t, p), k, klen) ? p : NULL;\n }\n \n static int cb_descend(struct cb_node *p, cb_iter fn, void *arg)\n@@ -115,6 +124,7 @@ int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n \tstruct cb_node *p = t->root;\n \tstruct cb_node *top = p;\n \tsize_t i = 0;\n+\tuint8_t *p_key;\n \n \tif (!p)\n \t\treturn 0; /* empty tree */\n@@ -130,8 +140,9 @@ int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n \t\t\ttop = p;\n \t}\n \n+\tp_key = cb_node_key(t, p);\n \tfor (i = 0; i < klen; i++) {\n-\t\tif (p->k[i] != kpfx[i])\n+\t\tif (p_key[i] != kpfx[i])\n \t\t\treturn 0; /* \"best\" match failed */\n \t}\n \ndiff --git a/cbtree.h b/cbtree.h\nindex c374b1b3db..3ce0d6b287 100644\n--- a/cbtree.h\n+++ b/cbtree.h\n@@ -23,18 +23,19 @@ struct cb_node {\n \t */\n \tuint32_t byte;\n \tuint8_t otherbits;\n-\tuint8_t k[FLEX_ARRAY]; /* arbitrary data, unaligned */\n };\n \n struct cb_tree {\n \tstruct cb_node *root;\n+\tptrdiff_t key_offset;\n };\n \n-#define CBTREE_INIT { 0 }\n-\n-static inline void cb_init(struct cb_tree *t)\n+static inline void cb_init(struct cb_tree *t,\n+\t\t\t   ptrdiff_t key_offset)\n {\n-\tstruct cb_tree blank = CBTREE_INIT;\n+\tstruct cb_tree blank = {\n+\t\t.key_offset = key_offset,\n+\t};\n \tmemcpy(t, &blank, sizeof(*t));\n }\n \ndiff --git a/oidtree.c b/oidtree.c\nindex ab9fe7ec7a..117649753f 100644\n--- a/oidtree.c\n+++ b/oidtree.c\n@@ -6,9 +6,14 @@\n #include \"oidtree.h\"\n #include \"hash.h\"\n \n+struct oidtree_node {\n+\tstruct cb_node base;\n+\tstruct object_id key;\n+};\n+\n void oidtree_init(struct oidtree *ot)\n {\n-\tcb_init(&ot->tree);\n+\tcb_init(&ot->tree, offsetof(struct oidtree_node, key));\n \tmem_pool_init(&ot->mem_pool, 0);\n }\n \n@@ -22,20 +27,13 @@ void oidtree_clear(struct oidtree *ot)\n \n void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n {\n-\tstruct cb_node *on;\n-\tstruct object_id k;\n+\tstruct oidtree_node *on;\n \n \tif (!oid->algo)\n \t\tBUG(\"oidtree_insert requires oid->algo\");\n \n-\ton = mem_pool_alloc(&ot->mem_pool, sizeof(*on) + sizeof(*oid));\n-\n-\t/*\n-\t * Clear the padding and copy the result in separate steps to\n-\t * respect the 4-byte alignment needed by struct object_id.\n-\t */\n-\toidcpy(&k, oid);\n-\tmemcpy(on->k, &k, sizeof(k));\n+\ton = mem_pool_alloc(&ot->mem_pool, sizeof(*on));\n+\toidcpy(&on->key, oid);\n \n \t/*\n \t * n.b. Current callers won't get us duplicates, here.  If a\n@@ -43,7 +41,7 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n \t * that won't be freed until oidtree_clear.  Currently it's not\n \t * worth maintaining a free list\n \t */\n-\tcb_insert(&ot->tree, on, sizeof(*oid));\n+\tcb_insert(&ot->tree, &on->base, sizeof(*oid));\n }\n \n bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n@@ -73,21 +71,18 @@ struct oidtree_each_data {\n \n static int iter(struct cb_node *n, void *cb_data)\n {\n+\tstruct oidtree_node *node = container_of(n, struct oidtree_node, base);\n \tstruct oidtree_each_data *data = cb_data;\n-\tstruct object_id k;\n-\n-\t/* Copy to provide 4-byte alignment needed by struct object_id. */\n-\tmemcpy(&k, n->k, sizeof(k));\n \n-\tif (data->algo != GIT_HASH_UNKNOWN && data->algo != k.algo)\n+\tif (data->algo != GIT_HASH_UNKNOWN && data->algo != node->key.algo)\n \t\treturn 0;\n \n \tif (data->last_nibble_at) {\n-\t\tif ((k.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n+\t\tif ((node->key.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n \t\t\treturn 0;\n \t}\n \n-\treturn data->cb(&k, data->cb_data);\n+\treturn data->cb(&node->key, data->cb_data);\n }\n \n int oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540826","messageId":"20260403-b4-pks-odb-source-inmemory-v1-9-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 09/16] oidtree: add ability to store data","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:56Z","receivedAt":"2026-04-03T06:02:41Z","isPatch":true,"body":"The oidtree data structure is currently only used to store object IDs,\nwithout any associated data. So consequently, it can only really be used\nto track which object IDs exist, and we can use the tree structure to\nefficiently operate on OID prefixes.\n\nBut there are valid use cases where we want to both:\n\n  - Store object IDs in a sorted order.\n\n  - Associated arbitrary data with them.\n\nRefactor the oidtree interface so that it allows us to store arbitrary\npayloads within the respective nodes. This will be used in the next\ncommit.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c                  |  2 +-\n object-file.c            |  3 ++-\n oidtree.c                | 37 ++++++++++++++++++++++++++++++++-----\n oidtree.h                | 12 ++++++++++--\n t/unit-tests/u-oidtree.c | 26 +++++++++++++++++++++++---\n 5 files changed, 68 insertions(+), 12 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 07333be696..f7a3dd1a72 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -57,7 +57,7 @@ static int insert_loose_map(struct odb_source *source,\n \tinserted |= insert_oid_pair(map->to_compat, oid, compat_oid);\n \tinserted |= insert_oid_pair(map->to_storage, compat_oid, oid);\n \tif (inserted)\n-\t\toidtree_insert(files->loose->cache, compat_oid);\n+\t\toidtree_insert(files->loose->cache, compat_oid, NULL);\n \n \treturn inserted;\n }\ndiff --git a/object-file.c b/object-file.c\nindex 98a4678ca4..c0805f0ebb 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1850,6 +1850,7 @@ static int for_each_object_wrapper_cb(const struct object_id *oid,\n }\n \n static int for_each_prefixed_object_wrapper_cb(const struct object_id *oid,\n+\t\t\t\t\t       void *node_data UNUSED,\n \t\t\t\t\t       void *cb_data)\n {\n \tstruct for_each_object_wrapper_data *data = cb_data;\n@@ -1995,7 +1996,7 @@ static int append_loose_object(const struct object_id *oid,\n \t\t\t       const char *path UNUSED,\n \t\t\t       void *data)\n {\n-\toidtree_insert(data, oid);\n+\toidtree_insert(data, oid, NULL);\n \treturn 0;\n }\n \ndiff --git a/oidtree.c b/oidtree.c\nindex 117649753f..e43f18026e 100644\n--- a/oidtree.c\n+++ b/oidtree.c\n@@ -9,6 +9,7 @@\n struct oidtree_node {\n \tstruct cb_node base;\n \tstruct object_id key;\n+\tvoid *data;\n };\n \n void oidtree_init(struct oidtree *ot)\n@@ -25,15 +26,22 @@ void oidtree_clear(struct oidtree *ot)\n \t}\n }\n \n-void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n+struct oidtree_data {\n+\tstruct object_id oid;\n+};\n+\n+void oidtree_insert(struct oidtree *ot, const struct object_id *oid,\n+\t\t    void *data)\n {\n \tstruct oidtree_node *on;\n+\tstruct cb_node *node;\n \n \tif (!oid->algo)\n \t\tBUG(\"oidtree_insert requires oid->algo\");\n \n \ton = mem_pool_alloc(&ot->mem_pool, sizeof(*on));\n \toidcpy(&on->key, oid);\n+\ton->data = data;\n \n \t/*\n \t * n.b. Current callers won't get us duplicates, here.  If a\n@@ -41,13 +49,19 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n \t * that won't be freed until oidtree_clear.  Currently it's not\n \t * worth maintaining a free list\n \t */\n-\tcb_insert(&ot->tree, &on->base, sizeof(*oid));\n+\tnode = cb_insert(&ot->tree, &on->base, sizeof(*oid));\n+\tif (node) {\n+\t\tstruct oidtree_node *preexisting = container_of(node, struct oidtree_node, base);\n+\t\tpreexisting->data = data;\n+\t}\n }\n \n-bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n+static struct oidtree_node *oidtree_lookup(struct oidtree *ot,\n+\t\t\t\t\t   const struct object_id *oid)\n {\n \tstruct object_id k;\n \tsize_t klen = sizeof(k);\n+\tstruct cb_node *node;\n \n \toidcpy(&k, oid);\n \n@@ -58,7 +72,20 @@ bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n \tklen += BUILD_ASSERT_OR_ZERO(offsetof(struct object_id, hash) <\n \t\t\t\toffsetof(struct object_id, algo));\n \n-\treturn !!cb_lookup(&ot->tree, (const uint8_t *)&k, klen);\n+\tnode = cb_lookup(&ot->tree, (const uint8_t *)&k, klen);\n+\treturn node ? container_of(node, struct oidtree_node, base) : NULL;\n+}\n+\n+bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n+{\n+\tstruct oidtree_node *node = oidtree_lookup(ot, oid);\n+\treturn node ? 1 : 0;\n+}\n+\n+void *oidtree_get(struct oidtree *ot, const struct object_id *oid)\n+{\n+\tstruct oidtree_node *node = oidtree_lookup(ot, oid);\n+\treturn node ? node->data : NULL;\n }\n \n struct oidtree_each_data {\n@@ -82,7 +109,7 @@ static int iter(struct cb_node *n, void *cb_data)\n \t\t\treturn 0;\n \t}\n \n-\treturn data->cb(&node->key, data->cb_data);\n+\treturn data->cb(&node->key, node->data, data->cb_data);\n }\n \n int oidtree_each(struct oidtree *ot, const struct object_id *prefix,\ndiff --git a/oidtree.h b/oidtree.h\nindex 2b7bad2e60..baa5a436ea 100644\n--- a/oidtree.h\n+++ b/oidtree.h\n@@ -29,18 +29,26 @@ void oidtree_init(struct oidtree *ot);\n  */\n void oidtree_clear(struct oidtree *ot);\n \n-/* Insert the object ID into the tree. */\n-void oidtree_insert(struct oidtree *ot, const struct object_id *oid);\n+/*\n+ * Insert the object ID into the tree and store the given pointer alongside\n+ * with it. The data pointer of any preexisting entry will be overwritten.\n+ */\n+void oidtree_insert(struct oidtree *ot, const struct object_id *oid,\n+\t\t    void *data);\n \n /* Check whether the tree contains the given object ID. */\n bool oidtree_contains(struct oidtree *ot, const struct object_id *oid);\n \n+/* Get the payload stored with the given object ID. */\n+void *oidtree_get(struct oidtree *ot, const struct object_id *oid);\n+\n /*\n  * Callback function used for `oidtree_each()`. Returning a non-zero exit code\n  * will cause iteration to stop. The exit code will be propagated to the caller\n  * of `oidtree_each()`.\n  */\n typedef int (*oidtree_each_cb)(const struct object_id *oid,\n+\t\t\t       void *node_data,\n \t\t\t       void *cb_data);\n \n /*\ndiff --git a/t/unit-tests/u-oidtree.c b/t/unit-tests/u-oidtree.c\nindex d4d05c7dc3..f0d5ebb733 100644\n--- a/t/unit-tests/u-oidtree.c\n+++ b/t/unit-tests/u-oidtree.c\n@@ -19,7 +19,7 @@ static int fill_tree_loc(struct oidtree *ot, const char *hexes[], size_t n)\n \tfor (size_t i = 0; i < n; i++) {\n \t\tstruct object_id oid;\n \t\tcl_parse_any_oid(hexes[i], &oid);\n-\t\toidtree_insert(ot, &oid);\n+\t\toidtree_insert(ot, &oid, NULL);\n \t}\n \treturn 0;\n }\n@@ -38,9 +38,9 @@ struct expected_hex_iter {\n \tconst char *query;\n };\n \n-static int check_each_cb(const struct object_id *oid, void *data)\n+static int check_each_cb(const struct object_id *oid, void *node_data UNUSED, void *cb_data)\n {\n-\tstruct expected_hex_iter *hex_iter = data;\n+\tstruct expected_hex_iter *hex_iter = cb_data;\n \tstruct object_id expected;\n \n \tcl_assert(hex_iter->i < hex_iter->expected_hexes.nr);\n@@ -105,3 +105,23 @@ void test_oidtree__each(void)\n \tcheck_each(&ot, \"32100\", \"321\", NULL);\n \tcheck_each(&ot, \"32\", \"320\", \"321\", NULL);\n }\n+\n+void test_oidtree__insert_overwrites_data(void)\n+{\n+\tstruct object_id oid;\n+\tstruct oidtree ot;\n+\tint a, b;\n+\n+\tcl_parse_any_oid(\"1\", &oid);\n+\n+\toidtree_init(&ot);\n+\n+\toidtree_insert(&ot, &oid, NULL);\n+\tcl_assert_equal_p(oidtree_get(&ot, &oid), NULL);\n+\toidtree_insert(&ot, &oid, &a);\n+\tcl_assert_equal_p(oidtree_get(&ot, &oid), &a);\n+\toidtree_insert(&ot, &oid, &b);\n+\tcl_assert_equal_p(oidtree_get(&ot, &oid), &b);\n+\n+\toidtree_clear(&ot);\n+}\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540827","messageId":"20260403-b4-pks-odb-source-inmemory-v1-10-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 10/16] odb/source-inmemory: convert to use oidtree","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:57Z","receivedAt":"2026-04-03T06:02:44Z","isPatch":true,"body":"The inmemory source stores its objects in a simple array that we grow as\nneeded. This has a couple of downsides:\n\n  - The object lookup is O(n). This doesn't matter in practice because\n    we only store a small number of objects.\n\n  - We don't have an easy way to iterate over all objects in\n    lexicographic order.\n\n  - We don't have an easy way to compute unique object ID prefixes.\n\nRefactor the code to use an oidtree instead. This is the same data\nstructure used by our loose object source, and thus it means we get a\nbunch of functionality for free.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 72 +++++++++++++++++++++++++++++++++++++--------------\n odb/source-inmemory.h | 13 ++--------\n 2 files changed, 54 insertions(+), 31 deletions(-)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 890e2a8c7c..22bae6927e 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -2,20 +2,29 @@\n #include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n #include \"odb/streaming.h\"\n+#include \"oidtree.h\"\n #include \"repository.h\"\n \n-static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n-\t\t\t\t\t\t      const struct object_id *oid)\n+struct inmemory_object {\n+\tenum object_type type;\n+\tconst void *buf;\n+\tunsigned long size;\n+};\n+\n+static const struct inmemory_object *find_cached_object(struct odb_source_inmemory *source,\n+\t\t\t\t\t\t\tconst struct object_id *oid)\n {\n-\tstatic const struct cached_object empty_tree = {\n+\tstatic const struct inmemory_object empty_tree = {\n \t\t.type = OBJ_TREE,\n \t\t.buf = \"\",\n \t};\n-\tconst struct cached_object_entry *co = source->objects;\n+\tconst struct inmemory_object *object;\n \n-\tfor (size_t i = 0; i < source->objects_nr; i++, co++)\n-\t\tif (oideq(&co->oid, oid))\n-\t\t\treturn &co->value;\n+\tif (source->objects) {\n+\t\tobject = oidtree_get(source->objects, oid);\n+\t\tif (object)\n+\t\t\treturn object;\n+\t}\n \n \tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n \t\treturn &empty_tree;\n@@ -29,7 +38,7 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \t\t\t\t\t\tenum object_info_flags flags UNUSED)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n-\tconst struct cached_object *object;\n+\tconst struct inmemory_object *object;\n \n \tobject = find_cached_object(inmemory, oid);\n \tif (!object)\n@@ -85,7 +94,7 @@ static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n \tstruct odb_read_stream_inmemory *stream;\n-\tconst struct cached_object *object;\n+\tconst struct inmemory_object *object;\n \n \tobject = find_cached_object(inmemory, oid);\n \tif (!object)\n@@ -110,15 +119,21 @@ static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    enum odb_write_object_flags flags UNUSED)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n-\tstruct cached_object_entry *object;\n+\tstruct inmemory_object *object;\n \n-\tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n-\t\t   inmemory->objects_alloc);\n-\tobject = &inmemory->objects[inmemory->objects_nr++];\n-\tobject->value.size = len;\n-\tobject->value.type = type;\n-\tobject->value.buf = xmemdupz(buf, len);\n-\toidcpy(&object->oid, oid);\n+\tif (!inmemory->objects) {\n+\t\tCALLOC_ARRAY(inmemory->objects, 1);\n+\t\toidtree_init(inmemory->objects);\n+\t} else if (oidtree_contains(inmemory->objects, oid)) {\n+\t\treturn 0;\n+\t}\n+\n+\tCALLOC_ARRAY(object, 1);\n+\tobject->size = len;\n+\tobject->type = type;\n+\tobject->buf = xmemdupz(buf, len);\n+\n+\toidtree_insert(inmemory->objects, oid, object);\n \n \treturn 0;\n }\n@@ -162,12 +177,29 @@ static int odb_source_inmemory_write_object_stream(struct odb_source *source,\n \treturn ret;\n }\n \n+static int inmemory_object_free(const struct object_id *oid UNUSED,\n+\t\t\t\tvoid *node_data,\n+\t\t\t\tvoid *cb_data UNUSED)\n+{\n+\tstruct inmemory_object *object = node_data;\n+\tfree((void *) object->buf);\n+\tfree(object);\n+\treturn 0;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n-\tfor (size_t i = 0; i < inmemory->objects_nr; i++)\n-\t\tfree((char *) inmemory->objects[i].value.buf);\n-\tfree(inmemory->objects);\n+\n+\tif (inmemory->objects) {\n+\t\tstruct object_id null_oid = { 0 };\n+\n+\t\toidtree_each(inmemory->objects, &null_oid, 0,\n+\t\t\t     inmemory_object_free, NULL);\n+\t\toidtree_clear(inmemory->objects);\n+\t\tfree(inmemory->objects);\n+\t}\n+\n \tfree(inmemory->base.path);\n \tfree(inmemory);\n }\ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex 14dc06f7c3..02cf586b63 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -3,14 +3,7 @@\n \n #include \"odb/source.h\"\n \n-struct cached_object_entry {\n-\tstruct object_id oid;\n-\tstruct cached_object {\n-\t\tenum object_type type;\n-\t\tconst void *buf;\n-\t\tunsigned long size;\n-\t} value;\n-};\n+struct oidtree;\n \n /*\n  * An inmemory source that you can write objects to that shall be made\n@@ -20,9 +13,7 @@ struct cached_object_entry {\n  */\n struct odb_source_inmemory {\n \tstruct odb_source base;\n-\n-\tstruct cached_object_entry *objects;\n-\tsize_t objects_nr, objects_alloc;\n+\tstruct oidtree *objects;\n };\n \n /* Create a new in-memory object database source. */\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540828","messageId":"20260403-b4-pks-odb-source-inmemory-v1-11-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 11/16] odb/source-inmemory: implement `for_each_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:58Z","receivedAt":"2026-04-03T06:02:47Z","isPatch":true,"body":"Implement the `for_each_object()` callback function for the inmemory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 86 +++++++++++++++++++++++++++++++++++++++++----------\n 1 file changed, 70 insertions(+), 16 deletions(-)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 22bae6927e..0ac20df323 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -32,6 +32,28 @@ static const struct inmemory_object *find_cached_object(struct odb_source_inmemo\n \treturn NULL;\n }\n \n+static void populate_object_info(struct odb_source_inmemory *source,\n+\t\t\t\t struct object_info *oi,\n+\t\t\t\t const struct inmemory_object *object)\n+{\n+\tif (!oi)\n+\t\treturn;\n+\n+\tif (oi->typep)\n+\t\t*(oi->typep) = object->type;\n+\tif (oi->sizep)\n+\t\t*(oi->sizep) = object->size;\n+\tif (oi->disk_sizep)\n+\t\t*(oi->disk_sizep) = 0;\n+\tif (oi->delta_base_oid)\n+\t\toidclr(oi->delta_base_oid, source->base.odb->repo->hash_algo);\n+\tif (oi->contentp)\n+\t\t*oi->contentp = xmemdupz(object->buf, object->size);\n+\tif (oi->mtimep)\n+\t\t*oi->mtimep = 0;\n+\toi->whence = OI_CACHED;\n+}\n+\n static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \t\t\t\t\t\tconst struct object_id *oid,\n \t\t\t\t\t\tstruct object_info *oi,\n@@ -44,22 +66,7 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \tif (!object)\n \t\treturn -1;\n \n-\tif (oi) {\n-\t\tif (oi->typep)\n-\t\t\t*(oi->typep) = object->type;\n-\t\tif (oi->sizep)\n-\t\t\t*(oi->sizep) = object->size;\n-\t\tif (oi->disk_sizep)\n-\t\t\t*(oi->disk_sizep) = 0;\n-\t\tif (oi->delta_base_oid)\n-\t\t\toidclr(oi->delta_base_oid, source->odb->repo->hash_algo);\n-\t\tif (oi->contentp)\n-\t\t\t*oi->contentp = xmemdupz(object->buf, object->size);\n-\t\tif (oi->mtimep)\n-\t\t\t*oi->mtimep = 0;\n-\t\toi->whence = OI_CACHED;\n-\t}\n-\n+\tpopulate_object_info(inmemory, oi, object);\n \treturn 0;\n }\n \n@@ -111,6 +118,52 @@ static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n \treturn 0;\n }\n \n+struct odb_source_inmemory_for_each_object_data {\n+\tstruct odb_source_inmemory *inmemory;\n+\tconst struct object_info *request;\n+\todb_for_each_object_cb cb;\n+\tvoid *cb_data;\n+};\n+\n+static int odb_source_inmemory_for_each_object_cb(const struct object_id *oid,\n+\t\t\t\t\t\t  void *node_data, void *cb_data)\n+{\n+\tstruct odb_source_inmemory_for_each_object_data *data = cb_data;\n+\tstruct inmemory_object *object = node_data;\n+\n+\tif (data->request) {\n+\t\tstruct object_info oi = *data->request;\n+\t\tpopulate_object_info(data->inmemory, &oi, object);\n+\t\treturn data->cb(oid, &oi, data->cb_data);\n+\t} else {\n+\t\treturn data->cb(oid, NULL, data->cb_data);\n+\t}\n+}\n+\n+static int odb_source_inmemory_for_each_object(struct odb_source *source,\n+\t\t\t\t\t       const struct object_info *request,\n+\t\t\t\t\t       odb_for_each_object_cb cb,\n+\t\t\t\t\t       void *cb_data,\n+\t\t\t\t\t       const struct odb_for_each_object_options *opts)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tstruct odb_source_inmemory_for_each_object_data payload = {\n+\t\t.inmemory = inmemory,\n+\t\t.request = request,\n+\t\t.cb = cb,\n+\t\t.cb_data = cb_data,\n+\t};\n+\tstruct object_id null_oid = { 0 };\n+\n+\tif ((opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY) ||\n+\t    (opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY && !source->local))\n+\t\treturn 0;\n+\n+\treturn oidtree_each(inmemory->objects,\n+\t\t\t    opts->prefix ? opts->prefix : &null_oid, opts->prefix_hex_len,\n+\t\t\t    odb_source_inmemory_for_each_object_cb, &payload);\n+}\n+\n static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    const void *buf, unsigned long len,\n \t\t\t\t\t    enum object_type type,\n@@ -214,6 +267,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.free = odb_source_inmemory_free;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n+\tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540830","messageId":"20260403-b4-pks-odb-source-inmemory-v1-12-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 12/16] odb/source-inmemory: implement `find_abbrev_len()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:01:59Z","receivedAt":"2026-04-03T06:02:50Z","isPatch":true,"body":"Implement the `find_abbrev_len()` callback function for the inmemory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 39 +++++++++++++++++++++++++++++++++++++++\n 1 file changed, 39 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 0ac20df323..16182bded3 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -164,6 +164,44 @@ static int odb_source_inmemory_for_each_object(struct odb_source *source,\n \t\t\t    odb_source_inmemory_for_each_object_cb, &payload);\n }\n \n+struct find_abbrev_len_data {\n+\tconst struct object_id *oid;\n+\tunsigned len;\n+};\n+\n+static int find_abbrev_len_cb(const struct object_id *oid,\n+\t\t\t      struct object_info *oi UNUSED,\n+\t\t\t      void *cb_data)\n+{\n+\tstruct find_abbrev_len_data *data = cb_data;\n+\tunsigned len = oid_common_prefix_hexlen(oid, data->oid);\n+\tif (len != hash_algos[oid->algo].hexsz && len >= data->len)\n+\t\tdata->len = len + 1;\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t\t       const struct object_id *oid,\n+\t\t\t\t\t       unsigned min_len,\n+\t\t\t\t\t       unsigned *out)\n+{\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = oid,\n+\t\t.prefix_hex_len = min_len,\n+\t};\n+\tstruct find_abbrev_len_data data = {\n+\t\t.oid = oid,\n+\t\t.len = min_len,\n+\t};\n+\tint ret;\n+\n+\tret = odb_source_inmemory_for_each_object(source, NULL, find_abbrev_len_cb,\n+\t\t\t\t\t\t  &data, &opts);\n+\t*out = data.len;\n+\n+\treturn ret;\n+}\n+\n static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    const void *buf, unsigned long len,\n \t\t\t\t\t    enum object_type type,\n@@ -268,6 +306,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n+\tsource->base.find_abbrev_len = odb_source_inmemory_find_abbrev_len;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540829","messageId":"20260403-b4-pks-odb-source-inmemory-v1-13-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 13/16] odb/source-inmemory: implement `count_objects()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:02:00Z","receivedAt":"2026-04-03T06:02:53Z","isPatch":true,"body":"Implement the `count_objects()` callback function for the inmemory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 20 ++++++++++++++++++++\n 1 file changed, 20 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 16182bded3..bd89a7ef14 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -202,6 +202,25 @@ static int odb_source_inmemory_find_abbrev_len(struct odb_source *source,\n \treturn ret;\n }\n \n+static int count_objects_cb(const struct object_id *oid UNUSED,\n+\t\t\t    struct object_info *oi UNUSED,\n+\t\t\t    void *cb_data)\n+{\n+\tunsigned long *counter = cb_data;\n+\t(*counter)++;\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_count_objects(struct odb_source *source,\n+\t\t\t\t\t     enum odb_count_objects_flags flags UNUSED,\n+\t\t\t\t\t     unsigned long *out)\n+{\n+\tstruct odb_for_each_object_options opts = { 0 };\n+\t*out = 0;\n+\treturn odb_source_inmemory_for_each_object(source, NULL, count_objects_cb,\n+\t\t\t\t\t\t   out, &opts);\n+}\n+\n static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    const void *buf, unsigned long len,\n \t\t\t\t\t    enum object_type type,\n@@ -307,6 +326,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n \tsource->base.find_abbrev_len = odb_source_inmemory_find_abbrev_len;\n+\tsource->base.count_objects = odb_source_inmemory_count_objects;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540831","messageId":"20260403-b4-pks-odb-source-inmemory-v1-14-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 14/16] odb/source-inmemory: implement `freshen_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:02:01Z","receivedAt":"2026-04-03T06:02:57Z","isPatch":true,"body":"Implement the `freshen_object()` callback function for the inmemory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 10 ++++++++++\n 1 file changed, 10 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex bd89a7ef14..c5249d04bc 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -287,6 +287,15 @@ static int odb_source_inmemory_write_object_stream(struct odb_source *source,\n \treturn ret;\n }\n \n+static int odb_source_inmemory_freshen_object(struct odb_source *source,\n+\t\t\t\t\t      const struct object_id *oid)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tif (find_cached_object(inmemory, oid))\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n static int inmemory_object_free(const struct object_id *oid UNUSED,\n \t\t\t\tvoid *node_data,\n \t\t\t\tvoid *cb_data UNUSED)\n@@ -329,6 +338,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.count_objects = odb_source_inmemory_count_objects;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n+\tsource->base.freshen_object = odb_source_inmemory_freshen_object;\n \n \treturn source;\n }\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540833","messageId":"20260403-b4-pks-odb-source-inmemory-v1-15-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 15/16] odb/source-inmemory: stub out remaining functions","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:02:02Z","receivedAt":"2026-04-03T06:03:00Z","isPatch":true,"body":"Stub out remaining functions that we either don't need or that are\nbasically no-ops.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 31 +++++++++++++++++++++++++++++++\n 1 file changed, 31 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex c5249d04bc..53009be032 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -296,6 +296,32 @@ static int odb_source_inmemory_freshen_object(struct odb_source *source,\n \treturn 0;\n }\n \n+static int odb_source_inmemory_begin_transaction(struct odb_source *source UNUSED,\n+\t\t\t\t\t\t struct odb_transaction **out UNUSED)\n+{\n+\treturn error(\"inmemory source does not support transactions\");\n+}\n+\n+static int odb_source_inmemory_read_alternates(struct odb_source *source UNUSED,\n+\t\t\t\t\t       struct strvec *out UNUSED)\n+{\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_write_alternate(struct odb_source *source UNUSED,\n+\t\t\t\t\t       const char *alternate UNUSED)\n+{\n+\treturn error(\"inmemory source does not support alternates\");\n+}\n+\n+static void odb_source_inmemory_close(struct odb_source *source UNUSED)\n+{\n+}\n+\n+static void odb_source_inmemory_reprepare(struct odb_source *source UNUSED)\n+{\n+}\n+\n static int inmemory_object_free(const struct object_id *oid UNUSED,\n \t\t\t\tvoid *node_data,\n \t\t\t\tvoid *cb_data UNUSED)\n@@ -331,6 +357,8 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n \n \tsource->base.free = odb_source_inmemory_free;\n+\tsource->base.close = odb_source_inmemory_close;\n+\tsource->base.reprepare = odb_source_inmemory_reprepare;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n@@ -339,6 +367,9 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \tsource->base.freshen_object = odb_source_inmemory_freshen_object;\n+\tsource->base.begin_transaction = odb_source_inmemory_begin_transaction;\n+\tsource->base.read_alternates = odb_source_inmemory_read_alternates;\n+\tsource->base.write_alternate = odb_source_inmemory_write_alternate;\n \n \treturn source;\n }\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540832","messageId":"20260403-b4-pks-odb-source-inmemory-v1-16-8b8d1abaa25e@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH 16/16] odb: generic inmemory source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-03T06:02:03Z","receivedAt":"2026-04-03T06:03:04Z","isPatch":true,"body":"Make the in-memory source generic.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c | 8 ++++----\n odb.h | 2 +-\n 2 files changed, 5 insertions(+), 5 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 34228c0cd5..70c59fef91 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -560,7 +560,7 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \tif (is_null_oid(real))\n \t\treturn -1;\n \n-\tif (!odb_source_read_object_info(&odb->inmemory_objects->base, oid, oi, flags))\n+\tif (!odb_source_read_object_info(odb->inmemory_objects, oid, oi, flags))\n \t\treturn 0;\n \n \todb_prepare_alternates(odb);\n@@ -737,7 +737,7 @@ int odb_pretend_object(struct object_database *odb,\n \tif (odb_has_object(odb, oid, 0))\n \t\treturn 0;\n \n-\treturn odb_source_write_object(&odb->inmemory_objects->base,\n+\treturn odb_source_write_object(odb->inmemory_objects,\n \t\t\t\t       buf, len, type, oid, NULL, 0);\n }\n \n@@ -1020,7 +1020,7 @@ struct object_database *odb_new(struct repository *repo,\n \to->sources = odb_source_new(o, primary_source, true);\n \to->sources_tail = &o->sources->next;\n \to->alternate_db = xstrdup_or_null(secondary_sources);\n-\to->inmemory_objects = odb_source_inmemory_new(o);\n+\to->inmemory_objects = &odb_source_inmemory_new(o)->base;\n \n \tfree(to_free);\n \n@@ -1045,7 +1045,7 @@ static void odb_free_sources(struct object_database *o)\n \t\to->sources = next;\n \t}\n \n-\todb_source_free(&o->inmemory_objects->base);\n+\todb_source_free(o->inmemory_objects);\n \to->inmemory_objects = NULL;\n \n \tkh_destroy_odb_path_map(o->source_by_path);\ndiff --git a/odb.h b/odb.h\nindex 3d20270a05..e3211ad8d4 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -99,7 +99,7 @@ struct object_database {\n \t * to write them into the object store (e.g. a browse-only\n \t * application).\n \t */\n-\tstruct odb_source_inmemory *inmemory_objects;\n+\tstruct odb_source *inmemory_objects;\n \n \t/*\n \t * A fast, rough count of the number of objects in the repository.\n\n-- \n2.53.0.1323.g189a785ab5.dirty\n\n"},{"id":"540856","messageId":"xmqqa4vknjab.fsf@gitster.g","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"Re: [PATCH 00/16] odb: introduce \"inmemory\" source","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-03T15:41:16Z","receivedAt":"2026-04-03T15:41:19Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> this patch series introduces the second object database source type,\n> which is the \"inmemory\" source.\n\nI cannot read the word without a hyphen, i.e.e.g., \"in-memory\".\n\n> This source may seem somewhat odd at first: it always starts out empty,\n> and any object written into it will only exist in memory until the\n> process exits. But the source already serves a purpose in our codebase,\n> where some commands, for example git-blame(1), write an in-memory\n> worktree commit.\n\nIntermediate tree and blob objects you need while making an octopus\nmerge may also benefit from this feature, to stay only in-core\nwithout having to get written out to the outside world.\n\nI understand that this is not meant to be used as a \"we create them\nonly for ourselves and they are available only to us while we work,\nbut once we are satisfied we make them available to others\", which\nis much better done by creating an on-disk ephemeral object store,\nwrite such objects in them, and then decide at the end of the\nprocess between discarding the ephemeral store and moving objects\nfrom there to the main object store.\n\n> Furthermore, I think that going forward it can serve more purposes as we\n> now have an easy way to write and read objects that will not get\n> persisted. I could see that this may be useful when for example\n> re-merging diffs. But eventually, once we have the object storage format\n> extension wired up, callers might even want to manually set up an\n> in-memory database as the primary ODB for write operations so that no\n> data will be persisted in an arbitrary write.\n\n;-)\n\n> Last but not least, this patch series also serves the purpose of\n> eventually getting rid of the `struct object_info::whence` member.\n> Instead, we'll simply yield the ODB source a specific object has been\n> read from, together with some backend-specific data, which gives\n> strictly more information compared to the status quo.\n>\n> The series is based on cf2139f8e1 (The 24th batch, 2026-04-01) with\n> ps/odb-cleanup at 109bcb7d1d (odb: drop unneeded headers and forward\n> decls, 2026-04-01) merged into it.\n>\n> Thanks!\n>\n> Patrick\n>\n> ---\n> Patrick Steinhardt (16):\n>       odb: introduce \"inmemory\" source\n>       odb/source-inmemory: implement `free()` callback\n>       odb: fix unnecessary call to `find_cached_object()`\n>       odb/source-inmemory: implement `read_object_info()` callback\n>       odb/source-inmemory: implement `read_object_stream()` callback\n>       odb/source-inmemory: implement `write_object()` callback\n>       odb/source-inmemory: implement `write_object_stream()` callback\n>       cbtree: allow using arbitrary wrapper structures for nodes\n>       oidtree: add ability to store data\n>       odb/source-inmemory: convert to use oidtree\n>       odb/source-inmemory: implement `for_each_object()` callback\n>       odb/source-inmemory: implement `find_abbrev_len()` callback\n>       odb/source-inmemory: implement `count_objects()` callback\n>       odb/source-inmemory: implement `freshen_object()` callback\n>       odb/source-inmemory: stub out remaining functions\n>       odb: generic inmemory source\n>\n>  Makefile                 |   1 +\n>  cbtree.c                 |  25 +++-\n>  cbtree.h                 |  11 +-\n>  loose.c                  |   2 +-\n>  meson.build              |   1 +\n>  object-file.c            |   3 +-\n>  odb.c                    |  82 ++---------\n>  odb.h                    |   4 +-\n>  odb/source-inmemory.c    | 375 +++++++++++++++++++++++++++++++++++++++++++++++\n>  odb/source-inmemory.h    |  33 +++++\n>  odb/source.h             |   3 +\n>  oidtree.c                |  66 ++++++---\n>  oidtree.h                |  12 +-\n>  t/unit-tests/u-oidtree.c |  26 +++-\n>  14 files changed, 529 insertions(+), 115 deletions(-)\n>\n>\n> ---\n> base-commit: 3d05c3e2906489caa9f12f0af18dc233a6b8032c\n> change-id: 20260401-b4-pks-odb-source-inmemory-7b17c83d9e43\n"},{"id":"540867","messageId":"xmqqzf3jlmnv.fsf@gitster.g","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-7-8b8d1abaa25e@pks.im","subject":"Re: [PATCH 07/16] odb/source-inmemory: implement `write_object_stream()` callback","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-03T22:11:16Z","receivedAt":"2026-04-03T22:11:18Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Implement the `write_object_stream()` callback function for the inmemory\n> source.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  odb/source-inmemory.c | 40 ++++++++++++++++++++++++++++++++++++++++\n>  1 file changed, 40 insertions(+)\n\nAs the signature of the .read() method drastically changes in\nanother topic in flight,\n\n  https://lore.kernel.org/git/20260402213220.2651523-4-jltobler@gmail.com/\n\nthis needs a bit of inter-topic coordination.\n\n> diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> index d2fc4c4054..890e2a8c7c 100644\n> --- a/odb/source-inmemory.c\n> +++ b/odb/source-inmemory.c\n> @@ -123,6 +123,45 @@ static int odb_source_inmemory_write_object(struct odb_source *source,\n>  \treturn 0;\n>  }\n>  \n> +static int odb_source_inmemory_write_object_stream(struct odb_source *source,\n> +\t\t\t\t\t\t   struct odb_write_stream *stream,\n> +\t\t\t\t\t\t   size_t len,\n> +\t\t\t\t\t\t   struct object_id *oid)\n> +{\n> +\tsize_t total_read = 0;\n> +\tchar *data;\n> +\tint ret;\n> +\n> +\tCALLOC_ARRAY(data, len);\n> +\twhile (!stream->is_finished) {\n> +\t\tunsigned long bytes_read;\n> +\t\tconst void *in;\n> +\n> +\t\tin = stream->read(stream, &bytes_read);\n> +\t\tif (total_read + bytes_read > len) {\n> +\t\t\tret = error(\"object stream yielded more bytes than expected\");\n> +\t\t\tgoto out;\n> +\t\t}\n> +\n> +\t\tmemcpy(data, in, bytes_read);\n> +\t\ttotal_read += bytes_read;\n> +\t}\n> +\n> +\tif (total_read != len) {\n> +\t\tret = error(\"object stream yielded less bytes than expected\");\n> +\t\tgoto out;\n> +\t}\n> +\n> +\tret = odb_source_inmemory_write_object(source, data, len, OBJ_BLOB, oid,\n> +\t\t\t\t\t       NULL, 0);\n> +\tif (ret < 0)\n> +\t\tgoto out;\n> +\n> +out:\n> +\tfree(data);\n> +\treturn ret;\n> +}\n> +\n>  static void odb_source_inmemory_free(struct odb_source *source)\n>  {\n>  \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n> @@ -144,6 +183,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n>  \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n>  \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n>  \tsource->base.write_object = odb_source_inmemory_write_object;\n> +\tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n>  \n>  \treturn source;\n>  }\n"},{"id":"541120","messageId":"adYQPmnajLmVr-vh@pks.im","threadId":"65423","inReplyTo":"xmqqa4vknjab.fsf@gitster.g","subject":"Re: [PATCH 00/16] odb: introduce \"inmemory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-08T08:22:22Z","receivedAt":"2026-04-08T08:22:34Z","isPatch":true,"body":"On Fri, Apr 03, 2026 at 08:41:16AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > this patch series introduces the second object database source type,\n> > which is the \"inmemory\" source.\n> \n> I cannot read the word without a hyphen, i.e.e.g., \"in-memory\".\n\nFair. I think I'll keep it as `odb_source_inmemory` in the sources,\nwhich I find easier ot parse than `odb_source_in_memory`, but will adapt\nto \"in-memory\" in prose. I already did this for most of the part, but\nnot in the cover letter indeed.\n\n> > This source may seem somewhat odd at first: it always starts out empty,\n> > and any object written into it will only exist in memory until the\n> > process exits. But the source already serves a purpose in our codebase,\n> > where some commands, for example git-blame(1), write an in-memory\n> > worktree commit.\n> \n> Intermediate tree and blob objects you need while making an octopus\n> merge may also benefit from this feature, to stay only in-core\n> without having to get written out to the outside world.\n> \n> I understand that this is not meant to be used as a \"we create them\n> only for ourselves and they are available only to us while we work,\n> but once we are satisfied we make them available to others\", which\n> is much better done by creating an on-disk ephemeral object store,\n> write such objects in them, and then decide at the end of the\n> process between discarding the ephemeral store and moving objects\n> from there to the main object store.\n\nYeah, I think there's a bunch of use cases where this could be useful\ngoing forward.\n\nPatrick\n"},{"id":"541121","messageId":"adYQUNdsmbVgZ3AT@pks.im","threadId":"65423","inReplyTo":"xmqqzf3jlmnv.fsf@gitster.g","subject":"Re: [PATCH 07/16] odb/source-inmemory: implement `write_object_stream()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-08T08:22:40Z","receivedAt":"2026-04-08T08:22:44Z","isPatch":true,"body":"On Fri, Apr 03, 2026 at 03:11:16PM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > Implement the `write_object_stream()` callback function for the inmemory\n> > source.\n> >\n> > Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> > ---\n> >  odb/source-inmemory.c | 40 ++++++++++++++++++++++++++++++++++++++++\n> >  1 file changed, 40 insertions(+)\n> \n> As the signature of the .read() method drastically changes in\n> another topic in flight,\n> \n>   https://lore.kernel.org/git/20260402213220.2651523-4-jltobler@gmail.com/\n> \n> this needs a bit of inter-topic coordination.\n\nFair. I think Justin's patch series is close to landing, so I'll rebase\nmy patch series on top of his. Thanks for flagging.\n\nPatrick\n"},{"id":"541180","messageId":"ada_W-IWfNKUKnVK@denethor","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-1-8b8d1abaa25e@pks.im","subject":"Re: [PATCH 01/16] odb: introduce \"inmemory\" source","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-08T21:00:48Z","receivedAt":"2026-04-08T21:00:53Z","isPatch":true,"body":"On 26/04/03 08:01AM, Patrick Steinhardt wrote:\n> Next to our typical object database sources, each object database also\n> has an implicit source of \"cached\" objects. These cached objects only\n> exist in memory and some use cases:\n> \n>   - They contain evergreen objects that we expect to always exist, like\n>     for example the empty tree.\n> \n>   - They can be used to store temporary objects that we don't want to\n>     persist to disk.\n> \n> Overall, their use is somewhat restricted though. For example, we don't\n> provide the ability to use it as a temporary object database source that\n> allows the user to write objects, but discard them after Git exists. So\n> while these cached objects behave almost like a source, they aren't used\n> as one.\n\nI find the wording of the second bullet point and paragraph above a\nlittle confusing. Are there existing uses where new objects are written\nto only the cache?\n\n> This is about to change over the following commits, where we will turn\n> cached objects into a new \"inmemory\" source. This will allow us to use\n> it exactly the same as any other source by providing the same common\n> interface as the \"files\" source.\n\nTreating the object cache just like any other ODB source seems like a\ngood direction.\n\n> For now, the inmemory source only hosts the cached objects and doesn't\n> provide any logic yet. This will change with subsequent commits, where\n> we move respective functionality into the source.\n> \n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  Makefile              |  1 +\n>  meson.build           |  1 +\n>  odb.c                 | 21 +++++++++++++--------\n>  odb.h                 |  4 ++--\n>  odb/source-inmemory.c | 12 ++++++++++++\n>  odb/source-inmemory.h | 35 +++++++++++++++++++++++++++++++++++\n>  odb/source.h          |  3 +++\n>  7 files changed, 67 insertions(+), 10 deletions(-)\n> \n[snip]\n> @@ -1123,9 +1126,11 @@ void odb_free(struct object_database *o)\n>  \todb_close(o);\n>  \todb_free_sources(o);\n>  \n> -\tfor (size_t i = 0; i < o->cached_object_nr; i++)\n> -\t\tfree((char *) o->cached_objects[i].value.buf);\n> -\tfree(o->cached_objects);\n> +\tfor (size_t i = 0; i < o->inmemory_objects->objects_nr; i++)\n> +\t\tfree((char *) o->inmemory_objects->objects[i].value.buf);\n> +\tfree(o->inmemory_objects->objects);\n> +\tfree(o->inmemory_objects->base.path);\n> +\tfree(o->inmemory_objects);\n\nShould we have some sort of `odb_source_inmemory_release()`?\n\n>  \n>  \tstring_list_clear(&o->submodule_source_paths, 0);\n>  \n> diff --git a/odb.h b/odb.h\n> index 3a711f6547..3d20270a05 100644\n> --- a/odb.h\n> +++ b/odb.h\n> @@ -8,6 +8,7 @@\n>  #include \"thread-utils.h\"\n>  \n>  struct cached_object_entry;\n> +struct odb_source_inmemory;\n>  struct packed_git;\n>  struct repository;\n>  struct strbuf;\n> @@ -98,8 +99,7 @@ struct object_database {\n>  \t * to write them into the object store (e.g. a browse-only\n>  \t * application).\n>  \t */\n> -\tstruct cached_object_entry *cached_objects;\n> -\tsize_t cached_object_nr, cached_object_alloc;\n> +\tstruct odb_source_inmemory *inmemory_objects;\n\nWe store an inmemory ODB source instead of the cache object info\ndirectly. Makes sense. \n\n>  \n>  \t/*\n>  \t * A fast, rough count of the number of objects in the repository.\n> diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> new file mode 100644\n> index 0000000000..c7ac5c24f0\n> --- /dev/null\n> +++ b/odb/source-inmemory.c\n> @@ -0,0 +1,12 @@\n> +#include \"git-compat-util.h\"\n> +#include \"odb/source-inmemory.h\"\n> +\n> +struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n> +{\n> +\tstruct odb_source_inmemory *source;\n> +\n> +\tCALLOC_ARRAY(source, 1);\n> +\todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n\nhuh, so we set the path for the `struct odb_source` to \"source\". In the\ncontext of an inmemory source, a path doesn't make much sense. I suspect\nthough that storing a path is likely only useful the context of the\nfiles ODB source. Is there reason for us to still keep this around in\nthe generic ODB source?\n\n> +\n> +\treturn source;\n> +}\n> diff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\n> new file mode 100644\n> index 0000000000..95477bf36d\n> --- /dev/null\n> +++ b/odb/source-inmemory.h\n> @@ -0,0 +1,35 @@\n> +#ifndef ODB_SOURCE_INMEMORY_H\n> +#define ODB_SOURCE_INMEMORY_H\n> +\n> +#include \"odb/source.h\"\n> +\n> +struct cached_object_entry;\n> +\n> +/*\n> + * An inmemory source that you can write objects to that shall be made\n> + * available for reading, but that shouldn't ever be persisted to disk. Note\n> + * that any objects written to this source will be stored in memory, so the\n> + * number of objects you can store is limited by available system memory.\n> + */\n> +struct odb_source_inmemory {\n> +\tstruct odb_source base;\n> +\n> +\tstruct cached_object_entry *objects;\n> +\tsize_t objects_nr, objects_alloc;\n> +};\n\nThis new ODB source now just contains the object cache info. Looks good.\n\n-Justin\n"},{"id":"541181","messageId":"adbCaYrt2mJcMPyK@denethor","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-2-8b8d1abaa25e@pks.im","subject":"Re: [PATCH 02/16] odb/source-inmemory: implement `free()` callback","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-08T21:05:10Z","receivedAt":"2026-04-08T21:05:12Z","isPatch":true,"body":"On 26/04/03 08:01AM, Patrick Steinhardt wrote:\n> @@ -1126,12 +1115,6 @@ void odb_free(struct object_database *o)\n>  \todb_close(o);\n>  \todb_free_sources(o);\n>  \n> -\tfor (size_t i = 0; i < o->inmemory_objects->objects_nr; i++)\n> -\t\tfree((char *) o->inmemory_objects->objects[i].value.buf);\n> -\tfree(o->inmemory_objects->objects);\n> -\tfree(o->inmemory_objects->base.path);\n> -\tfree(o->inmemory_objects);\n\nAh ok, this addresses a comment in the previous patch.\n\n> -\n>  \tstring_list_clear(&o->submodule_source_paths, 0);\n>  \n>  \tfree(o);\n> diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> index c7ac5c24f0..ccbb622eae 100644\n> --- a/odb/source-inmemory.c\n> +++ b/odb/source-inmemory.c\n> @@ -1,6 +1,16 @@\n>  #include \"git-compat-util.h\"\n>  #include \"odb/source-inmemory.h\"\n>  \n> +static void odb_source_inmemory_free(struct odb_source *source)\n> +{\n> +\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n> +\tfor (size_t i = 0; i < inmemory->objects_nr; i++)\n> +\t\tfree((char *) inmemory->objects[i].value.buf);\n> +\tfree(inmemory->objects);\n> +\tfree(inmemory->base.path);\n> +\tfree(inmemory);\n> +}\n> +\n>  struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n>  {\n>  \tstruct odb_source_inmemory *source;\n> @@ -8,5 +18,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n>  \tCALLOC_ARRAY(source, 1);\n>  \todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n>  \n> +\tsource->base.free = odb_source_inmemory_free;\n\nWe wire up a function to specifically handle freeing the inmemory ODB\nsource. Looks good.\n\n-Justin\n"},{"id":"541182","messageId":"adbDpPQgvPfctxQS@denethor","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-3-8b8d1abaa25e@pks.im","subject":"Re: [PATCH 03/16] odb: fix unnecessary call to `find_cached_object()`","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-08T21:13:45Z","receivedAt":"2026-04-08T21:13:47Z","isPatch":true,"body":"On 26/04/03 08:01AM, Patrick Steinhardt wrote:\n> diff --git a/odb.c b/odb.c\n> index d321242353..21cdedc31c 100644\n> --- a/odb.c\n> +++ b/odb.c\n> @@ -774,8 +774,7 @@ int odb_pretend_object(struct object_database *odb,\n>  \tchar *co_buf;\n>  \n>  \thash_object_file(odb->repo->hash_algo, buf, len, type, oid);\n> -\tif (odb_has_object(odb, oid, 0) ||\n> -\t    find_cached_object(odb, oid))\n> +\tif (odb_has_object(odb, oid, 0))\n\nNice, odb_has_object() does indeed already check the object cache so\nthat makes the explicit find_cached_object() redundant.\n\nIf a future where temporary objects could be written to the inmemory ODB\nsource, would there ever be a reason for odb_has_object() to\ndifferentiate between inmemory and real objects?\n\n-Justin\n"},{"id":"541183","messageId":"adbG1gIAALhMINlv@denethor","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-5-8b8d1abaa25e@pks.im","subject":"Re: [PATCH 05/16] odb/source-inmemory: implement `read_object_stream()` callback","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-04-08T21:24:13Z","receivedAt":"2026-04-08T21:24:18Z","isPatch":true,"body":"On 26/04/03 08:01AM, Patrick Steinhardt wrote:\n> Implement the `read_object_stream()` callback function for the inmemory\n> source.\n\nHmmm, if the whole object is already in memory, outside providing a\ncomplete ODB source interface, is there really much reason for streaming\nthe object in practice?\n\nThe patch itself looks good though.\n\n-Justin\n"},{"id":"541187","messageId":"xmqq5x61xgvv.fsf@gitster.g","threadId":"65423","inReplyTo":"adYQPmnajLmVr-vh@pks.im","subject":"Re: [PATCH 00/16] odb: introduce \"inmemory\" source","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-08T21:48:52Z","receivedAt":"2026-04-08T21:48:54Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> On Fri, Apr 03, 2026 at 08:41:16AM -0700, Junio C Hamano wrote:\n>> Patrick Steinhardt <ps@pks.im> writes:\n>> \n>> > this patch series introduces the second object database source type,\n>> > which is the \"inmemory\" source.\n>> \n>> I cannot read the word without a hyphen, i.e.e.g., \"in-memory\".\n>\n> Fair. I think I'll keep it as `odb_source_inmemory` in the sources,\n> which I find easier ot parse than `odb_source_in_memory`, but will adapt\n> to \"in-memory\" in prose. I already did this for most of the part, but\n> not in the cover letter indeed.\n\nFair.\n\nFWIW, we do the same for \"in core\" or \"in-core\" in prose, and\n\"incore\" in identifier names, so the above is understandable\nposition to take.\n\nBut stepping back a bit, does this new \"in memory\" refer to a\nconcept that is different from what the rest of the system uses \"in\ncore\" to represent?\n"},{"id":"541211","messageId":"adc3mAItBiKMUFNJ@pks.im","threadId":"65423","inReplyTo":"xmqq5x61xgvv.fsf@gitster.g","subject":"Re: [PATCH 00/16] odb: introduce \"inmemory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T05:22:32Z","receivedAt":"2026-04-09T05:22:46Z","isPatch":true,"body":"On Wed, Apr 08, 2026 at 02:48:52PM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > On Fri, Apr 03, 2026 at 08:41:16AM -0700, Junio C Hamano wrote:\n> >> Patrick Steinhardt <ps@pks.im> writes:\n> >> \n> >> > this patch series introduces the second object database source type,\n> >> > which is the \"inmemory\" source.\n> >> \n> >> I cannot read the word without a hyphen, i.e.e.g., \"in-memory\".\n> >\n> > Fair. I think I'll keep it as `odb_source_inmemory` in the sources,\n> > which I find easier ot parse than `odb_source_in_memory`, but will adapt\n> > to \"in-memory\" in prose. I already did this for most of the part, but\n> > not in the cover letter indeed.\n> \n> Fair.\n> \n> FWIW, we do the same for \"in core\" or \"in-core\" in prose, and\n> \"incore\" in identifier names, so the above is understandable\n> position to take.\n> \n> But stepping back a bit, does this new \"in memory\" refer to a\n> concept that is different from what the rest of the system uses \"in\n> core\" to represent?\n\nNo, in principle it's not any different. One of the reasons I decided to\ngo with \"in memory\" though is that this backend may eventually be\n(power-)user-facing via the planned \"objectStorage\" extension.\n\nThis extension will work similar to how the \"refStorage\" extension\nworks, where every backend has a schema followed by an optional payload.\nSo for the files backend it would be \"files://<path>\", and if one wants\nto configure a temporary ODB source that doesn't store objects it would\nbe \"inmemory://\". And overall, I think that \"inmemory\" is a lot easier\nto understand intuitively compared to \"incore\".\n\nThe counter argument may be that this really only is for power users\nanyway, as it's a rather risky thing to do (e.g. you must not update any\nrefs), and such power users may understand the concept of \"in-core\". But\neven there I feel like it makes sense to rather say \"in-memory\".\n\nPatrick\n"},{"id":"541212","messageId":"adc3pDxks6rCrZo6@pks.im","threadId":"65423","inReplyTo":"ada_W-IWfNKUKnVK@denethor","subject":"Re: [PATCH 01/16] odb: introduce \"inmemory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T05:22:44Z","receivedAt":"2026-04-09T05:22:49Z","isPatch":true,"body":"On Wed, Apr 08, 2026 at 04:00:48PM -0500, Justin Tobler wrote:\n> On 26/04/03 08:01AM, Patrick Steinhardt wrote:\n> > Next to our typical object database sources, each object database also\n> > has an implicit source of \"cached\" objects. These cached objects only\n> > exist in memory and some use cases:\n> > \n> >   - They contain evergreen objects that we expect to always exist, like\n> >     for example the empty tree.\n> > \n> >   - They can be used to store temporary objects that we don't want to\n> >     persist to disk.\n> > \n> > Overall, their use is somewhat restricted though. For example, we don't\n> > provide the ability to use it as a temporary object database source that\n> > allows the user to write objects, but discard them after Git exists. So\n> > while these cached objects behave almost like a source, they aren't used\n> > as one.\n> \n> I find the wording of the second bullet point and paragraph above a\n> little confusing. Are there existing uses where new objects are written\n> to only the cache?\n\nYes, there's a single user with git-blame(1). I'll mention that user\nexplcitly.\n\n> > @@ -1123,9 +1126,11 @@ void odb_free(struct object_database *o)\n> >  \todb_close(o);\n> >  \todb_free_sources(o);\n> >  \n> > -\tfor (size_t i = 0; i < o->cached_object_nr; i++)\n> > -\t\tfree((char *) o->cached_objects[i].value.buf);\n> > -\tfree(o->cached_objects);\n> > +\tfor (size_t i = 0; i < o->inmemory_objects->objects_nr; i++)\n> > +\t\tfree((char *) o->inmemory_objects->objects[i].value.buf);\n> > +\tfree(o->inmemory_objects->objects);\n> > +\tfree(o->inmemory_objects->base.path);\n> > +\tfree(o->inmemory_objects);\n> \n> Should we have some sort of `odb_source_inmemory_release()`?\n\nYup, this is coming in subsequent commits.\n\n> > diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> > new file mode 100644\n> > index 0000000000..c7ac5c24f0\n> > --- /dev/null\n> > +++ b/odb/source-inmemory.c\n> > @@ -0,0 +1,12 @@\n> > +#include \"git-compat-util.h\"\n> > +#include \"odb/source-inmemory.h\"\n> > +\n> > +struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n> > +{\n> > +\tstruct odb_source_inmemory *source;\n> > +\n> > +\tCALLOC_ARRAY(source, 1);\n> > +\todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n> \n> huh, so we set the path for the `struct odb_source` to \"source\". In the\n> context of an inmemory source, a path doesn't make much sense. I suspect\n> though that storing a path is likely only useful the context of the\n> files ODB source. Is there reason for us to still keep this around in\n> the generic ODB source?\n\nThere are two reasons for the \"path\" field to exist:\n\n  - It is used to compare sources with one another to figure out whether\n    two sources are actually the same. This is used when reloading\n    sources. This usage makes sense in principle, but it's wrong that we\n    consider this to be a \"path\" -- it should rather be considered an\n    opaque \"payload\".\n\n  - The path field is used in a bunch of sites to actually figure out\n    paths. This is plain wrong, as we cannot guarantee that the field\n    even is a path for backends that don't store data on the filesystem.\n\nIt's one of the topics that we've got on our plate, to disentangle this.\nThe goal is ultimately to move the path into the files backend, fix up\ncallers to do the right thing (TM) and then convert the current path\nfield that we have into a payload.\n\nPatrick\n"},{"id":"541213","messageId":"adc3qfDPinq1aakq@pks.im","threadId":"65423","inReplyTo":"adbDpPQgvPfctxQS@denethor","subject":"Re: [PATCH 03/16] odb: fix unnecessary call to `find_cached_object()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T05:22:49Z","receivedAt":"2026-04-09T05:22:54Z","isPatch":true,"body":"On Wed, Apr 08, 2026 at 04:13:45PM -0500, Justin Tobler wrote:\n> On 26/04/03 08:01AM, Patrick Steinhardt wrote:\n> > diff --git a/odb.c b/odb.c\n> > index d321242353..21cdedc31c 100644\n> > --- a/odb.c\n> > +++ b/odb.c\n> > @@ -774,8 +774,7 @@ int odb_pretend_object(struct object_database *odb,\n> >  \tchar *co_buf;\n> >  \n> >  \thash_object_file(odb->repo->hash_algo, buf, len, type, oid);\n> > -\tif (odb_has_object(odb, oid, 0) ||\n> > -\t    find_cached_object(odb, oid))\n> > +\tif (odb_has_object(odb, oid, 0))\n> \n> Nice, odb_has_object() does indeed already check the object cache so\n> that makes the explicit find_cached_object() redundant.\n> \n> If a future where temporary objects could be written to the inmemory ODB\n> source, would there ever be a reason for odb_has_object() to\n> differentiate between inmemory and real objects?\n\nWe could in theory just append the in-memory source to the normal list\nof sources, and that would ensure that all the usual operations would\nknow to also consider this source. But there's a couple of points that\nspeak against it, at least for now:\n\n  - Callers that explicitly want to explicitly write temporary objects\n    need to have a handle to the in-memory source. That handle would be\n    hard to obtain if we were to only store the source in the list of\n    sources.\n\n  - It would be a change in behaviour if functions like\n    `odb_for_each_object()` were to also enumerate in-memory objects.\n\nThe former one could be solved by having both the direct pointer and\nkeep the source in the list. The latter can be solved by having a\nseparate flag for `odb_for_each_object()` that tells the ODB that we\nwant to exclude/include in-memory objects.\n\nBut overall it feels like this would only complicate things without much\nof a tangible benefit.\n\nPatrick\n"},{"id":"541214","messageId":"adc3sHcAj3OwhxUy@pks.im","threadId":"65423","inReplyTo":"adbG1gIAALhMINlv@denethor","subject":"Re: [PATCH 05/16] odb/source-inmemory: implement `read_object_stream()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T05:22:56Z","receivedAt":"2026-04-09T05:23:01Z","isPatch":true,"body":"On Wed, Apr 08, 2026 at 04:24:13PM -0500, Justin Tobler wrote:\n> On 26/04/03 08:01AM, Patrick Steinhardt wrote:\n> > Implement the `read_object_stream()` callback function for the inmemory\n> > source.\n> \n> Hmmm, if the whole object is already in memory, outside providing a\n> complete ODB source interface, is there really much reason for streaming\n> the object in practice?\n\nI cannot think of any, but wanted to provide this function anyway so\nthat the backend is complete.\n\nPatrick\n"},{"id":"541216","messageId":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH v2 00/17] odb: introduce \"in-memory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:21Z","receivedAt":"2026-04-09T07:24:34Z","isPatch":true,"body":"Hi,\n\nthis patch series introduces the second object database source type,\nwhich is the \"in-memory\" source.\n\nThis source may seem somewhat odd at first: it always starts out empty,\nand any object written into it will only exist in memory until the\nprocess exits. But the source already serves a purpose in our codebase,\nwhere some commands, for example git-blame(1), write an in-memory\nworktree commit.\n\nFurthermore, I think that going forward it can serve more purposes as we\nnow have an easy way to write and read objects that will not get\npersisted. I could see that this may be useful when for example\nre-merging diffs. But eventually, once we have the object storage format\nextension wired up, callers might even want to manually set up an\nin-memory database as the primary ODB for write operations so that no\ndata will be persisted in an arbitrary write.\n\nLast but not least, this patch series also serves the purpose of\neventually getting rid of the `struct object_info::whence` member.\nInstead, we'll simply yield the ODB source a specific object has been\nread from, together with some backend-specific data, which gives\nstrictly more information compared to the status quo.\n\nThe series is based onb15384c06f (A bit more post -rc1, 2026-04-08)\nwith jt/odb-transaction-write at ddf6aee9c6 (odb/transaction: make\n`write_object_stream()` pluggable, 2026-04-02) merged into it.\n\nChanges in v2:\n  - Fix handling of object IDs when writing objects.\n  - I've changed the base of this series to include Justin's\n    refactorings for the ODB write streams. I've updated the above\n    paragraph detailing the merge base accordingly. @Junio: I'm fine to\n    defer this patch series a bit until Justin's patch series has been\n    merged to `next` in case this causes inconvenience.\n  - Use \"in-memory\" instead of \"inmemory\" in commit messages.\n  - Link to v1: https://patch.msgid.link/20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (17):\n      odb: introduce \"in-memory\" source\n      odb/source-inmemory: implement `free()` callback\n      odb: fix unnecessary call to `find_cached_object()`\n      odb/source-inmemory: implement `read_object_info()` callback\n      odb/source-inmemory: implement `read_object_stream()` callback\n      odb/source-inmemory: implement `write_object()` callback\n      odb/source-inmemory: implement `write_object()` callback\n      odb/source-inmemory: implement `write_object_stream()` callback\n      cbtree: allow using arbitrary wrapper structures for nodes\n      oidtree: add ability to store data\n      odb/source-inmemory: convert to use oidtree\n      odb/source-inmemory: implement `for_each_object()` callback\n      odb/source-inmemory: implement `find_abbrev_len()` callback\n      odb/source-inmemory: implement `count_objects()` callback\n      odb/source-inmemory: implement `freshen_object()` callback\n      odb/source-inmemory: stub out remaining functions\n      odb: generic in-memory source\n\n Makefile                 |   1 +\n cbtree.c                 |  25 +++-\n cbtree.h                 |  11 +-\n loose.c                  |   2 +-\n meson.build              |   1 +\n object-file.c            |   3 +-\n odb.c                    |  82 ++--------\n odb.h                    |   4 +-\n odb/source-inmemory.c    | 378 +++++++++++++++++++++++++++++++++++++++++++++++\n odb/source-inmemory.h    |  33 +++++\n odb/source.h             |   3 +\n oidtree.c                |  66 ++++++---\n oidtree.h                |  12 +-\n t/unit-tests/u-oidtree.c |  26 +++-\n 14 files changed, 532 insertions(+), 115 deletions(-)\n\nRange-diff versus v1:\n\n 1:  b7cd1ae8d1 !  1:  df8567d908 odb: introduce \"inmemory\" source\n    @@ Metadata\n     Author: Patrick Steinhardt <ps@pks.im>\n     \n      ## Commit message ##\n    -    odb: introduce \"inmemory\" source\n    +    odb: introduce \"in-memory\" source\n     \n         Next to our typical object database sources, each object database also\n         has an implicit source of \"cached\" objects. These cached objects only\n    @@ Commit message\n             for example the empty tree.\n     \n           - They can be used to store temporary objects that we don't want to\n    -        persist to disk.\n    +        persist to disk, which is used by git-blame(1) to create a fake\n    +        worktree commit.\n     \n         Overall, their use is somewhat restricted though. For example, we don't\n         provide the ability to use it as a temporary object database source that\n    @@ Commit message\n         as one.\n     \n         This is about to change over the following commits, where we will turn\n    -    cached objects into a new \"inmemory\" source. This will allow us to use\n    +    cached objects into a new \"in-memory\" source. This will allow us to use\n         it exactly the same as any other source by providing the same common\n         interface as the \"files\" source.\n     \n    -    For now, the inmemory source only hosts the cached objects and doesn't\n    +    For now, the in-memory source only hosts the cached objects and doesn't\n         provide any logic yet. This will change with subsequent commits, where\n         we move respective functionality into the source.\n     \n    @@ Makefile: LIB_OBJS += object.o\n      LIB_OBJS += odb/source-files.o\n     +LIB_OBJS += odb/source-inmemory.o\n      LIB_OBJS += odb/streaming.o\n    + LIB_OBJS += odb/transaction.o\n      LIB_OBJS += oid-array.o\n    - LIB_OBJS += oidmap.o\n     \n      ## meson.build ##\n     @@ meson.build: libgit_sources = [\n    @@ meson.build: libgit_sources = [\n        'odb/source-files.c',\n     +  'odb/source-inmemory.c',\n        'odb/streaming.c',\n    +   'odb/transaction.c',\n        'oid-array.c',\n    -   'oidmap.c',\n     \n      ## odb.c ##\n     @@\n 2:  298758b4d5 !  2:  e1ffe26ca9 odb/source-inmemory: implement `free()` callback\n    @@ Metadata\n      ## Commit message ##\n         odb/source-inmemory: implement `free()` callback\n     \n    -    Implement the `free()` callback function for the \"inmemory\" source.\n    +    Implement the `free()` callback function for the \"in-memory\" source.\n     \n         Note that this requires us to define `struct cached_object_entry` in\n         \"odb/source-inmemory.h\", as it is accessed in both \"odb.c\" and\n 3:  b57997d027 =  3:  f58424bb80 odb: fix unnecessary call to `find_cached_object()`\n 4:  9ae26b9aa1 !  4:  786a240391 odb/source-inmemory: implement `read_object_info()` callback\n    @@ Metadata\n      ## Commit message ##\n         odb/source-inmemory: implement `read_object_info()` callback\n     \n    -    Implement the `read_object_info()` callback function for the inmemory\n    +    Implement the `read_object_info()` callback function for the in-memory\n         source.\n     \n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n 5:  5d9781009e !  5:  22d3e7134b odb/source-inmemory: implement `read_object_stream()` callback\n    @@ Metadata\n      ## Commit message ##\n         odb/source-inmemory: implement `read_object_stream()` callback\n     \n    -    Implement the `read_object_stream()` callback function for the inmemory\n    +    Implement the `read_object_stream()` callback function for the in-memory\n         source.\n     \n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n 6:  bc9620c608 !  6:  139e7f2beb odb/source-inmemory: implement `write_object()` callback\n    @@ Metadata\n      ## Commit message ##\n         odb/source-inmemory: implement `write_object()` callback\n     \n    -    Implement the `write_object()` callback function for the inmemory\n    +    Implement the `write_object()` callback function for the in-memory\n         source.\n     \n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n -:  ---------- >  7:  7f5ab16d1c odb/source-inmemory: implement `write_object()` callback\n 7:  6d9f8634e1 !  8:  6006f5e782 odb/source-inmemory: implement `write_object_stream()` callback\n    @@ Metadata\n      ## Commit message ##\n         odb/source-inmemory: implement `write_object_stream()` callback\n     \n    -    Implement the `write_object_stream()` callback function for the inmemory\n    +    Implement the `write_object_stream()` callback function for the in-memory\n         source.\n     \n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n    @@ odb/source-inmemory.c: static int odb_source_inmemory_write_object(struct odb_so\n     +\t\t\t\t\t\t   size_t len,\n     +\t\t\t\t\t\t   struct object_id *oid)\n     +{\n    ++\tchar buf[16384];\n     +\tsize_t total_read = 0;\n     +\tchar *data;\n     +\tint ret;\n     +\n     +\tCALLOC_ARRAY(data, len);\n     +\twhile (!stream->is_finished) {\n    -+\t\tunsigned long bytes_read;\n    -+\t\tconst void *in;\n    ++\t\tssize_t bytes_read;\n     +\n    -+\t\tin = stream->read(stream, &bytes_read);\n    ++\t\tbytes_read = odb_write_stream_read(stream, buf, sizeof(buf));\n     +\t\tif (total_read + bytes_read > len) {\n     +\t\t\tret = error(\"object stream yielded more bytes than expected\");\n     +\t\t\tgoto out;\n     +\t\t}\n     +\n    -+\t\tmemcpy(data, in, bytes_read);\n    ++\t\tmemcpy(data, buf, bytes_read);\n     +\t\ttotal_read += bytes_read;\n     +\t}\n     +\n 8:  45f9c761ce =  9:  392d9bf6ed cbtree: allow using arbitrary wrapper structures for nodes\n 9:  5eb7742886 = 10:  9fd88ffd16 oidtree: add ability to store data\n10:  4f95cd0a51 ! 11:  6d4a77b47c odb/source-inmemory: convert to use oidtree\n    @@ Metadata\n      ## Commit message ##\n         odb/source-inmemory: convert to use oidtree\n     \n    -    The inmemory source stores its objects in a simple array that we grow as\n    +    The in-memory source stores its objects in a simple array that we grow as\n         needed. This has a couple of downsides:\n     \n           - The object lookup is O(n). This doesn't matter in practice because\n    @@ odb/source-inmemory.c: static int odb_source_inmemory_write_object(struct odb_so\n     -\tstruct cached_object_entry *object;\n     +\tstruct inmemory_object *object;\n      \n    + \thash_object_file(source->odb->repo->hash_algo, buf, len, type, oid);\n    + \n     -\tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n     -\t\t   inmemory->objects_alloc);\n     -\tobject = &inmemory->objects[inmemory->objects_nr++];\n11:  fc231e22dc ! 12:  5f345d76ef odb/source-inmemory: implement `for_each_object()` callback\n    @@ Metadata\n      ## Commit message ##\n         odb/source-inmemory: implement `for_each_object()` callback\n     \n    -    Implement the `for_each_object()` callback function for the inmemory\n    +    Implement the `for_each_object()` callback function for the in-memory\n         source.\n     \n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n12:  c2437b2ba5 ! 13:  b428a1760b odb/source-inmemory: implement `find_abbrev_len()` callback\n    @@ Metadata\n      ## Commit message ##\n         odb/source-inmemory: implement `find_abbrev_len()` callback\n     \n    -    Implement the `find_abbrev_len()` callback function for the inmemory\n    +    Implement the `find_abbrev_len()` callback function for the in-memory\n         source.\n     \n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n13:  fee0586da7 ! 14:  564cc60392 odb/source-inmemory: implement `count_objects()` callback\n    @@ Metadata\n      ## Commit message ##\n         odb/source-inmemory: implement `count_objects()` callback\n     \n    -    Implement the `count_objects()` callback function for the inmemory\n    +    Implement the `count_objects()` callback function for the in-memory\n         source.\n     \n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n14:  634392eaf9 ! 15:  9ddfb6f67b odb/source-inmemory: implement `freshen_object()` callback\n    @@ Metadata\n      ## Commit message ##\n         odb/source-inmemory: implement `freshen_object()` callback\n     \n    -    Implement the `freshen_object()` callback function for the inmemory\n    +    Implement the `freshen_object()` callback function for the in-memory\n         source.\n     \n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n15:  3d1f08a849 = 16:  d76329a424 odb/source-inmemory: stub out remaining functions\n16:  29deff493d ! 17:  41cd562975 odb: generic inmemory source\n    @@ Metadata\n     Author: Patrick Steinhardt <ps@pks.im>\n     \n      ## Commit message ##\n    -    odb: generic inmemory source\n    +    odb: generic in-memory source\n     \n         Make the in-memory source generic.\n     \n\n---\nbase-commit: a3ebc5a08e67ccac4c915622049a968a31e48662\nchange-id: 20260401-b4-pks-odb-source-inmemory-7b17c83d9e43\n\n"},{"id":"541217","messageId":"20260409-b4-pks-odb-source-inmemory-v2-1-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 01/17] odb: introduce \"in-memory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:22Z","receivedAt":"2026-04-09T07:24:35Z","isPatch":true,"body":"Next to our typical object database sources, each object database also\nhas an implicit source of \"cached\" objects. These cached objects only\nexist in memory and some use cases:\n\n  - They contain evergreen objects that we expect to always exist, like\n    for example the empty tree.\n\n  - They can be used to store temporary objects that we don't want to\n    persist to disk, which is used by git-blame(1) to create a fake\n    worktree commit.\n\nOverall, their use is somewhat restricted though. For example, we don't\nprovide the ability to use it as a temporary object database source that\nallows the user to write objects, but discard them after Git exists. So\nwhile these cached objects behave almost like a source, they aren't used\nas one.\n\nThis is about to change over the following commits, where we will turn\ncached objects into a new \"in-memory\" source. This will allow us to use\nit exactly the same as any other source by providing the same common\ninterface as the \"files\" source.\n\nFor now, the in-memory source only hosts the cached objects and doesn't\nprovide any logic yet. This will change with subsequent commits, where\nwe move respective functionality into the source.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n Makefile              |  1 +\n meson.build           |  1 +\n odb.c                 | 21 +++++++++++++--------\n odb.h                 |  4 ++--\n odb/source-inmemory.c | 12 ++++++++++++\n odb/source-inmemory.h | 35 +++++++++++++++++++++++++++++++++++\n odb/source.h          |  3 +++\n 7 files changed, 67 insertions(+), 10 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 22a8993482..3cda12c455 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1218,6 +1218,7 @@ LIB_OBJS += object.o\n LIB_OBJS += odb.o\n LIB_OBJS += odb/source.o\n LIB_OBJS += odb/source-files.o\n+LIB_OBJS += odb/source-inmemory.o\n LIB_OBJS += odb/streaming.o\n LIB_OBJS += odb/transaction.o\n LIB_OBJS += oid-array.o\ndiff --git a/meson.build b/meson.build\nindex 6dc23b3af2..ffa73ce7ce 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -404,6 +404,7 @@ libgit_sources = [\n   'odb.c',\n   'odb/source.c',\n   'odb/source-files.c',\n+  'odb/source-inmemory.c',\n   'odb/streaming.c',\n   'odb/transaction.c',\n   'oid-array.c',\ndiff --git a/odb.c b/odb.c\nindex 40a5e9c4e0..60e1eead25 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -14,6 +14,7 @@\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/source-inmemory.h\"\n #include \"packfile.h\"\n #include \"path.h\"\n #include \"promisor-remote.h\"\n@@ -53,9 +54,9 @@ static const struct cached_object *find_cached_object(struct object_database *ob\n \t\t.type = OBJ_TREE,\n \t\t.buf = \"\",\n \t};\n-\tconst struct cached_object_entry *co = object_store->cached_objects;\n+\tconst struct cached_object_entry *co = object_store->inmemory_objects->objects;\n \n-\tfor (size_t i = 0; i < object_store->cached_object_nr; i++, co++)\n+\tfor (size_t i = 0; i < object_store->inmemory_objects->objects_nr; i++, co++)\n \t\tif (oideq(&co->oid, oid))\n \t\t\treturn &co->value;\n \n@@ -792,9 +793,10 @@ int odb_pretend_object(struct object_database *odb,\n \t    find_cached_object(odb, oid))\n \t\treturn 0;\n \n-\tALLOC_GROW(odb->cached_objects,\n-\t\t   odb->cached_object_nr + 1, odb->cached_object_alloc);\n-\tco = &odb->cached_objects[odb->cached_object_nr++];\n+\tALLOC_GROW(odb->inmemory_objects->objects,\n+\t\t   odb->inmemory_objects->objects_nr + 1,\n+\t\t   odb->inmemory_objects->objects_alloc);\n+\tco = &odb->inmemory_objects->objects[odb->inmemory_objects->objects_nr++];\n \tco->value.size = len;\n \tco->value.type = type;\n \tco_buf = xmalloc(len);\n@@ -1083,6 +1085,7 @@ struct object_database *odb_new(struct repository *repo,\n \to->sources = odb_source_new(o, primary_source, true);\n \to->sources_tail = &o->sources->next;\n \to->alternate_db = xstrdup_or_null(secondary_sources);\n+\to->inmemory_objects = odb_source_inmemory_new(o);\n \n \tfree(to_free);\n \n@@ -1123,9 +1126,11 @@ void odb_free(struct object_database *o)\n \todb_close(o);\n \todb_free_sources(o);\n \n-\tfor (size_t i = 0; i < o->cached_object_nr; i++)\n-\t\tfree((char *) o->cached_objects[i].value.buf);\n-\tfree(o->cached_objects);\n+\tfor (size_t i = 0; i < o->inmemory_objects->objects_nr; i++)\n+\t\tfree((char *) o->inmemory_objects->objects[i].value.buf);\n+\tfree(o->inmemory_objects->objects);\n+\tfree(o->inmemory_objects->base.path);\n+\tfree(o->inmemory_objects);\n \n \tstring_list_clear(&o->submodule_source_paths, 0);\n \ndiff --git a/odb.h b/odb.h\nindex 9eb8355aca..c3a7edf9c8 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -8,6 +8,7 @@\n #include \"thread-utils.h\"\n \n struct cached_object_entry;\n+struct odb_source_inmemory;\n struct packed_git;\n struct repository;\n struct strbuf;\n@@ -80,8 +81,7 @@ struct object_database {\n \t * to write them into the object store (e.g. a browse-only\n \t * application).\n \t */\n-\tstruct cached_object_entry *cached_objects;\n-\tsize_t cached_object_nr, cached_object_alloc;\n+\tstruct odb_source_inmemory *inmemory_objects;\n \n \t/*\n \t * A fast, rough count of the number of objects in the repository.\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nnew file mode 100644\nindex 0000000000..c7ac5c24f0\n--- /dev/null\n+++ b/odb/source-inmemory.c\n@@ -0,0 +1,12 @@\n+#include \"git-compat-util.h\"\n+#include \"odb/source-inmemory.h\"\n+\n+struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n+{\n+\tstruct odb_source_inmemory *source;\n+\n+\tCALLOC_ARRAY(source, 1);\n+\todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n+\n+\treturn source;\n+}\ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nnew file mode 100644\nindex 0000000000..95477bf36d\n--- /dev/null\n+++ b/odb/source-inmemory.h\n@@ -0,0 +1,35 @@\n+#ifndef ODB_SOURCE_INMEMORY_H\n+#define ODB_SOURCE_INMEMORY_H\n+\n+#include \"odb/source.h\"\n+\n+struct cached_object_entry;\n+\n+/*\n+ * An inmemory source that you can write objects to that shall be made\n+ * available for reading, but that shouldn't ever be persisted to disk. Note\n+ * that any objects written to this source will be stored in memory, so the\n+ * number of objects you can store is limited by available system memory.\n+ */\n+struct odb_source_inmemory {\n+\tstruct odb_source base;\n+\n+\tstruct cached_object_entry *objects;\n+\tsize_t objects_nr, objects_alloc;\n+};\n+\n+/* Create a new in-memory object database source. */\n+struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb);\n+\n+/*\n+ * Cast the given object database source to the inmemory backend. This will\n+ * cause a BUG in case the source doesn't use this backend.\n+ */\n+static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n+{\n+\tif (source->type != ODB_SOURCE_INMEMORY)\n+\t\tBUG(\"trying to downcast source of type '%d' to inmemory\", source->type);\n+\treturn container_of(source, struct odb_source_inmemory, base);\n+}\n+\n+#endif\ndiff --git a/odb/source.h b/odb/source.h\nindex f706e0608a..cd14f9e046 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -13,6 +13,9 @@ enum odb_source_type {\n \n \t/* The \"files\" backend that uses loose objects and packfiles. */\n \tODB_SOURCE_FILES,\n+\n+\t/* The \"inmemory\" backend that stores objects in memory. */\n+\tODB_SOURCE_INMEMORY,\n };\n \n struct object_id;\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541218","messageId":"20260409-b4-pks-odb-source-inmemory-v2-2-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 02/17] odb/source-inmemory: implement `free()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:23Z","receivedAt":"2026-04-09T07:24:38Z","isPatch":true,"body":"Implement the `free()` callback function for the \"in-memory\" source.\n\nNote that this requires us to define `struct cached_object_entry` in\n\"odb/source-inmemory.h\", as it is accessed in both \"odb.c\" and\n\"odb/source-inmemory.c\" now. This will be fixed in subsequent commits\nthough.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                 | 25 ++++---------------------\n odb/source-inmemory.c | 12 ++++++++++++\n odb/source-inmemory.h |  9 ++++++++-\n 3 files changed, 24 insertions(+), 22 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 60e1eead25..1d65825ed3 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -32,21 +32,6 @@\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct odb_source *, 1, fspathhash, fspatheq)\n \n-/*\n- * This is meant to hold a *small* number of objects that you would\n- * want odb_read_object() to be able to return, but yet you do not want\n- * to write them into the object store (e.g. a browse-only\n- * application).\n- */\n-struct cached_object_entry {\n-\tstruct object_id oid;\n-\tstruct cached_object {\n-\t\tenum object_type type;\n-\t\tconst void *buf;\n-\t\tunsigned long size;\n-\t} value;\n-};\n-\n static const struct cached_object *find_cached_object(struct object_database *object_store,\n \t\t\t\t\t\t      const struct object_id *oid)\n {\n@@ -1109,6 +1094,10 @@ static void odb_free_sources(struct object_database *o)\n \t\todb_source_free(o->sources);\n \t\to->sources = next;\n \t}\n+\n+\todb_source_free(&o->inmemory_objects->base);\n+\to->inmemory_objects = NULL;\n+\n \tkh_destroy_odb_path_map(o->source_by_path);\n \to->source_by_path = NULL;\n }\n@@ -1126,12 +1115,6 @@ void odb_free(struct object_database *o)\n \todb_close(o);\n \todb_free_sources(o);\n \n-\tfor (size_t i = 0; i < o->inmemory_objects->objects_nr; i++)\n-\t\tfree((char *) o->inmemory_objects->objects[i].value.buf);\n-\tfree(o->inmemory_objects->objects);\n-\tfree(o->inmemory_objects->base.path);\n-\tfree(o->inmemory_objects);\n-\n \tstring_list_clear(&o->submodule_source_paths, 0);\n \n \tfree(o);\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex c7ac5c24f0..ccbb622eae 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,6 +1,16 @@\n #include \"git-compat-util.h\"\n #include \"odb/source-inmemory.h\"\n \n+static void odb_source_inmemory_free(struct odb_source *source)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tfor (size_t i = 0; i < inmemory->objects_nr; i++)\n+\t\tfree((char *) inmemory->objects[i].value.buf);\n+\tfree(inmemory->objects);\n+\tfree(inmemory->base.path);\n+\tfree(inmemory);\n+}\n+\n struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n {\n \tstruct odb_source_inmemory *source;\n@@ -8,5 +18,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tCALLOC_ARRAY(source, 1);\n \todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n \n+\tsource->base.free = odb_source_inmemory_free;\n+\n \treturn source;\n }\ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex 95477bf36d..14dc06f7c3 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -3,7 +3,14 @@\n \n #include \"odb/source.h\"\n \n-struct cached_object_entry;\n+struct cached_object_entry {\n+\tstruct object_id oid;\n+\tstruct cached_object {\n+\t\tenum object_type type;\n+\t\tconst void *buf;\n+\t\tunsigned long size;\n+\t} value;\n+};\n \n /*\n  * An inmemory source that you can write objects to that shall be made\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541219","messageId":"20260409-b4-pks-odb-source-inmemory-v2-3-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 03/17] odb: fix unnecessary call to `find_cached_object()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:24Z","receivedAt":"2026-04-09T07:24:40Z","isPatch":true,"body":"The function `odb_pretend_object()` writes an object into the in-memory\nobject database source. The effect of this is that the object will now\nbecome readable, but it won't ever be persisted to disk.\n\nBefore storing the object, we first verify whether the object already\nexists. This is done by calling `odb_has_object()` to check all sources,\nfollowed by `find_cached_object()` to check whether we have already\nstored the object in our in-memory source.\n\nThis is unnecessary though, as `odb_has_object()` already checks the\nin-memory source transitively via:\n\n  - `odb_has_object()`\n  - `odb_read_object_info_extended()`\n  - `do_oid_object_info_extended()`\n  - `find_cached_object()`\n\nDrop the explicit call to `find_cached_object()`.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c | 3 +--\n 1 file changed, 1 insertion(+), 2 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 1d65825ed3..ea3fcf5e11 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -774,8 +774,7 @@ int odb_pretend_object(struct object_database *odb,\n \tchar *co_buf;\n \n \thash_object_file(odb->repo->hash_algo, buf, len, type, oid);\n-\tif (odb_has_object(odb, oid, 0) ||\n-\t    find_cached_object(odb, oid))\n+\tif (odb_has_object(odb, oid, 0))\n \t\treturn 0;\n \n \tALLOC_GROW(odb->inmemory_objects->objects,\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541220","messageId":"20260409-b4-pks-odb-source-inmemory-v2-4-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 04/17] odb/source-inmemory: implement `read_object_info()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:25Z","receivedAt":"2026-04-09T07:24:43Z","isPatch":true,"body":"Implement the `read_object_info()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                 | 39 +------------------------------------\n odb/source-inmemory.c | 53 +++++++++++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 54 insertions(+), 38 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex ea3fcf5e11..6a3912adac 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -32,25 +32,6 @@\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct odb_source *, 1, fspathhash, fspatheq)\n \n-static const struct cached_object *find_cached_object(struct object_database *object_store,\n-\t\t\t\t\t\t      const struct object_id *oid)\n-{\n-\tstatic const struct cached_object empty_tree = {\n-\t\t.type = OBJ_TREE,\n-\t\t.buf = \"\",\n-\t};\n-\tconst struct cached_object_entry *co = object_store->inmemory_objects->objects;\n-\n-\tfor (size_t i = 0; i < object_store->inmemory_objects->objects_nr; i++, co++)\n-\t\tif (oideq(&co->oid, oid))\n-\t\t\treturn &co->value;\n-\n-\tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n-\t\treturn &empty_tree;\n-\n-\treturn NULL;\n-}\n-\n int odb_mkstemp(struct object_database *odb,\n \t\tstruct strbuf *temp_filename, const char *pattern)\n {\n@@ -570,7 +551,6 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \t\t\t\t       const struct object_id *oid,\n \t\t\t\t       struct object_info *oi, unsigned flags)\n {\n-\tconst struct cached_object *co;\n \tconst struct object_id *real = oid;\n \tint already_retried = 0;\n \n@@ -580,25 +560,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \tif (is_null_oid(real))\n \t\treturn -1;\n \n-\tco = find_cached_object(odb, real);\n-\tif (co) {\n-\t\tif (oi) {\n-\t\t\tif (oi->typep)\n-\t\t\t\t*(oi->typep) = co->type;\n-\t\t\tif (oi->sizep)\n-\t\t\t\t*(oi->sizep) = co->size;\n-\t\t\tif (oi->disk_sizep)\n-\t\t\t\t*(oi->disk_sizep) = 0;\n-\t\t\tif (oi->delta_base_oid)\n-\t\t\t\toidclr(oi->delta_base_oid, odb->repo->hash_algo);\n-\t\t\tif (oi->contentp)\n-\t\t\t\t*oi->contentp = xmemdupz(co->buf, co->size);\n-\t\t\tif (oi->mtimep)\n-\t\t\t\t*oi->mtimep = 0;\n-\t\t\toi->whence = OI_CACHED;\n-\t\t}\n+\tif (!odb_source_read_object_info(&odb->inmemory_objects->base, oid, oi, flags))\n \t\treturn 0;\n-\t}\n \n \todb_prepare_alternates(odb);\n \ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex ccbb622eae..12c80f9b34 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,5 +1,57 @@\n #include \"git-compat-util.h\"\n+#include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n+#include \"repository.h\"\n+\n+static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n+\t\t\t\t\t\t      const struct object_id *oid)\n+{\n+\tstatic const struct cached_object empty_tree = {\n+\t\t.type = OBJ_TREE,\n+\t\t.buf = \"\",\n+\t};\n+\tconst struct cached_object_entry *co = source->objects;\n+\n+\tfor (size_t i = 0; i < source->objects_nr; i++, co++)\n+\t\tif (oideq(&co->oid, oid))\n+\t\t\treturn &co->value;\n+\n+\tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n+\t\treturn &empty_tree;\n+\n+\treturn NULL;\n+}\n+\n+static int odb_source_inmemory_read_object_info(struct odb_source *source,\n+\t\t\t\t\t\tconst struct object_id *oid,\n+\t\t\t\t\t\tstruct object_info *oi,\n+\t\t\t\t\t\tenum object_info_flags flags UNUSED)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tconst struct cached_object *object;\n+\n+\tobject = find_cached_object(inmemory, oid);\n+\tif (!object)\n+\t\treturn -1;\n+\n+\tif (oi) {\n+\t\tif (oi->typep)\n+\t\t\t*(oi->typep) = object->type;\n+\t\tif (oi->sizep)\n+\t\t\t*(oi->sizep) = object->size;\n+\t\tif (oi->disk_sizep)\n+\t\t\t*(oi->disk_sizep) = 0;\n+\t\tif (oi->delta_base_oid)\n+\t\t\toidclr(oi->delta_base_oid, source->odb->repo->hash_algo);\n+\t\tif (oi->contentp)\n+\t\t\t*oi->contentp = xmemdupz(object->buf, object->size);\n+\t\tif (oi->mtimep)\n+\t\t\t*oi->mtimep = 0;\n+\t\toi->whence = OI_CACHED;\n+\t}\n+\n+\treturn 0;\n+}\n \n static void odb_source_inmemory_free(struct odb_source *source)\n {\n@@ -19,6 +71,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n \n \tsource->base.free = odb_source_inmemory_free;\n+\tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541221","messageId":"20260409-b4-pks-odb-source-inmemory-v2-5-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 05/17] odb/source-inmemory: implement `read_object_stream()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:26Z","receivedAt":"2026-04-09T07:24:46Z","isPatch":true,"body":"Implement the `read_object_stream()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 50 ++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 50 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 12c80f9b34..4a68169430 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,6 +1,7 @@\n #include \"git-compat-util.h\"\n #include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n+#include \"odb/streaming.h\"\n #include \"repository.h\"\n \n static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n@@ -53,6 +54,54 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \treturn 0;\n }\n \n+struct odb_read_stream_inmemory {\n+\tstruct odb_read_stream base;\n+\tconst void *buf;\n+\tsize_t offset;\n+};\n+\n+static ssize_t odb_read_stream_inmemory_read(struct odb_read_stream *stream,\n+\t\t\t\t\t     char *buf, size_t buf_len)\n+{\n+\tstruct odb_read_stream_inmemory *inmemory =\n+\t\tcontainer_of(stream, struct odb_read_stream_inmemory, base);\n+\tsize_t bytes = buf_len;\n+\n+\tif (buf_len > inmemory->base.size - inmemory->offset)\n+\t\tbytes = inmemory->base.size - inmemory->offset;\n+\tmemcpy(buf, inmemory->buf, bytes);\n+\n+\treturn bytes;\n+}\n+\n+static int odb_read_stream_inmemory_close(struct odb_read_stream *stream UNUSED)\n+{\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t\t\t  struct odb_source *source,\n+\t\t\t\t\t\t  const struct object_id *oid)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tstruct odb_read_stream_inmemory *stream;\n+\tconst struct cached_object *object;\n+\n+\tobject = find_cached_object(inmemory, oid);\n+\tif (!object)\n+\t\treturn -1;\n+\n+\tCALLOC_ARRAY(stream, 1);\n+\tstream->base.read = odb_read_stream_inmemory_read;\n+\tstream->base.close = odb_read_stream_inmemory_close;\n+\tstream->base.size = object->size;\n+\tstream->base.type = object->type;\n+\tstream->buf = object->buf;\n+\n+\t*out = &stream->base;\n+\treturn 0;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n@@ -72,6 +121,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \n \tsource->base.free = odb_source_inmemory_free;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n+\tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541223","messageId":"20260409-b4-pks-odb-source-inmemory-v2-6-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 06/17] odb/source-inmemory: implement `write_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:27Z","receivedAt":"2026-04-09T07:24:49Z","isPatch":true,"body":"Implement the `write_object()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                 | 16 ++--------------\n odb/source-inmemory.c | 22 ++++++++++++++++++++++\n 2 files changed, 24 insertions(+), 14 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 6a3912adac..24e929f03c 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -733,24 +733,12 @@ int odb_pretend_object(struct object_database *odb,\n \t\t       void *buf, unsigned long len, enum object_type type,\n \t\t       struct object_id *oid)\n {\n-\tstruct cached_object_entry *co;\n-\tchar *co_buf;\n-\n \thash_object_file(odb->repo->hash_algo, buf, len, type, oid);\n \tif (odb_has_object(odb, oid, 0))\n \t\treturn 0;\n \n-\tALLOC_GROW(odb->inmemory_objects->objects,\n-\t\t   odb->inmemory_objects->objects_nr + 1,\n-\t\t   odb->inmemory_objects->objects_alloc);\n-\tco = &odb->inmemory_objects->objects[odb->inmemory_objects->objects_nr++];\n-\tco->value.size = len;\n-\tco->value.type = type;\n-\tco_buf = xmalloc(len);\n-\tmemcpy(co_buf, buf, len);\n-\tco->value.buf = co_buf;\n-\toidcpy(&co->oid, oid);\n-\treturn 0;\n+\treturn odb_source_write_object(&odb->inmemory_objects->base,\n+\t\t\t\t       buf, len, type, oid, NULL, 0);\n }\n \n void *odb_read_object(struct object_database *odb,\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 4a68169430..d2fc4c4054 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -102,6 +102,27 @@ static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n \treturn 0;\n }\n \n+static int odb_source_inmemory_write_object(struct odb_source *source,\n+\t\t\t\t\t    const void *buf, unsigned long len,\n+\t\t\t\t\t    enum object_type type,\n+\t\t\t\t\t    struct object_id *oid,\n+\t\t\t\t\t    struct object_id *compat_oid UNUSED,\n+\t\t\t\t\t    enum odb_write_object_flags flags UNUSED)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tstruct cached_object_entry *object;\n+\n+\tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n+\t\t   inmemory->objects_alloc);\n+\tobject = &inmemory->objects[inmemory->objects_nr++];\n+\tobject->value.size = len;\n+\tobject->value.type = type;\n+\tobject->value.buf = xmemdupz(buf, len);\n+\toidcpy(&object->oid, oid);\n+\n+\treturn 0;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n@@ -122,6 +143,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.free = odb_source_inmemory_free;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n+\tsource->base.write_object = odb_source_inmemory_write_object;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541224","messageId":"20260409-b4-pks-odb-source-inmemory-v2-7-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 07/17] odb/source-inmemory: implement `write_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:28Z","receivedAt":"2026-04-09T07:24:50Z","isPatch":true,"body":"Implement the `write_object()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex d2fc4c4054..96e8efd327 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,4 +1,5 @@\n #include \"git-compat-util.h\"\n+#include \"object-file.h\"\n #include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n #include \"odb/streaming.h\"\n@@ -112,6 +113,8 @@ static int odb_source_inmemory_write_object(struct odb_source *source,\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n \tstruct cached_object_entry *object;\n \n+\thash_object_file(source->odb->repo->hash_algo, buf, len, type, oid);\n+\n \tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n \t\t   inmemory->objects_alloc);\n \tobject = &inmemory->objects[inmemory->objects_nr++];\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541225","messageId":"20260409-b4-pks-odb-source-inmemory-v2-8-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 08/17] odb/source-inmemory: implement `write_object_stream()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:29Z","receivedAt":"2026-04-09T07:24:53Z","isPatch":true,"body":"Implement the `write_object_stream()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 40 ++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 40 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 96e8efd327..578ceea550 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -126,6 +126,45 @@ static int odb_source_inmemory_write_object(struct odb_source *source,\n \treturn 0;\n }\n \n+static int odb_source_inmemory_write_object_stream(struct odb_source *source,\n+\t\t\t\t\t\t   struct odb_write_stream *stream,\n+\t\t\t\t\t\t   size_t len,\n+\t\t\t\t\t\t   struct object_id *oid)\n+{\n+\tchar buf[16384];\n+\tsize_t total_read = 0;\n+\tchar *data;\n+\tint ret;\n+\n+\tCALLOC_ARRAY(data, len);\n+\twhile (!stream->is_finished) {\n+\t\tssize_t bytes_read;\n+\n+\t\tbytes_read = odb_write_stream_read(stream, buf, sizeof(buf));\n+\t\tif (total_read + bytes_read > len) {\n+\t\t\tret = error(\"object stream yielded more bytes than expected\");\n+\t\t\tgoto out;\n+\t\t}\n+\n+\t\tmemcpy(data, buf, bytes_read);\n+\t\ttotal_read += bytes_read;\n+\t}\n+\n+\tif (total_read != len) {\n+\t\tret = error(\"object stream yielded less bytes than expected\");\n+\t\tgoto out;\n+\t}\n+\n+\tret = odb_source_inmemory_write_object(source, data, len, OBJ_BLOB, oid,\n+\t\t\t\t\t       NULL, 0);\n+\tif (ret < 0)\n+\t\tgoto out;\n+\n+out:\n+\tfree(data);\n+\treturn ret;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n@@ -147,6 +186,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n+\tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541226","messageId":"20260409-b4-pks-odb-source-inmemory-v2-9-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 09/17] cbtree: allow using arbitrary wrapper structures for nodes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:30Z","receivedAt":"2026-04-09T07:24:56Z","isPatch":true,"body":"The cbtree subsystem allows the user to store arbitrary data in a\nprefix-free set of strings. This is used by us to store object IDs in a\nway that we can easily iterate through them in lexicograph order, and so\nthat we can easily perform lookups with shortened object IDs.\n\nIn its current form, it is not easily possible to store arbitrary data\nwith the tree nodes. There are a couple of approaches such a caller\ncould try to use, but none of them really work:\n\n  - One may embed the `struct cb_node` in a custom structure. This does\n    not work though as `struct cb_node` contains a flex array, and\n    embedding such a struct in another struct is forbidden.\n\n  - One may use a `union` over `struct cb_node` and ones own data type,\n    which _is_ allowed even if the struct contains a flex array. This\n    does not work though, as the compiler may align members of the\n    struct so that the node key would not immediately start where the\n    flex array starts.\n\n  - One may allocate `struct cb_node` such that it has room for both its\n    key and the custom data. This has the downside though that if the\n    custom data is itself a pointer to allocated memory, then the leak\n    checker will not consider the pointer to be alive anymore.\n\nRefactor the cbtree to drop the flex array and instead take in an\nexplicit offset for where to find the key, which allows the caller to\nembed `struct cb_node` is a wrapper struct.\n\nNote that this change has the downside that we now have a bit of padding\nin our structure, which grows the size from 60 to 64 bytes on a 64 bit\nsystem. On the other hand though, it allows us to get rid of the memory\ncopies that we previously had to do to ensure proper alignment. This\nseems like a reasonable tradeoff.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n cbtree.c  | 25 ++++++++++++++++++-------\n cbtree.h  | 11 ++++++-----\n oidtree.c | 33 ++++++++++++++-------------------\n 3 files changed, 38 insertions(+), 31 deletions(-)\n\ndiff --git a/cbtree.c b/cbtree.c\nindex 4ab794bddc..8f5edbb80a 100644\n--- a/cbtree.c\n+++ b/cbtree.c\n@@ -7,6 +7,11 @@\n #include \"git-compat-util.h\"\n #include \"cbtree.h\"\n \n+static inline uint8_t *cb_node_key(struct cb_tree *t, struct cb_node *node)\n+{\n+\treturn (uint8_t *) node + t->key_offset;\n+}\n+\n static struct cb_node *cb_node_of(const void *p)\n {\n \treturn (struct cb_node *)((uintptr_t)p - 1);\n@@ -33,6 +38,7 @@ struct cb_node *cb_insert(struct cb_tree *t, struct cb_node *node, size_t klen)\n \tuint8_t c;\n \tint newdirection;\n \tstruct cb_node **wherep, *p;\n+\tuint8_t *node_key, *p_key;\n \n \tassert(!((uintptr_t)node & 1)); /* allocations must be aligned */\n \n@@ -41,23 +47,26 @@ struct cb_node *cb_insert(struct cb_tree *t, struct cb_node *node, size_t klen)\n \t\treturn NULL;\t/* success */\n \t}\n \n+\tnode_key = cb_node_key(t, node);\n+\n \t/* see if a node already exists */\n-\tp = cb_internal_best_match(t->root, node->k, klen);\n+\tp = cb_internal_best_match(t->root, node_key, klen);\n+\tp_key = cb_node_key(t, p);\n \n \t/* find first differing byte */\n \tfor (newbyte = 0; newbyte < klen; newbyte++) {\n-\t\tif (p->k[newbyte] != node->k[newbyte])\n+\t\tif (p_key[newbyte] != node_key[newbyte])\n \t\t\tgoto different_byte_found;\n \t}\n \treturn p;\t/* element exists, let user deal with it */\n \n different_byte_found:\n-\tnewotherbits = p->k[newbyte] ^ node->k[newbyte];\n+\tnewotherbits = p_key[newbyte] ^ node_key[newbyte];\n \tnewotherbits |= newotherbits >> 1;\n \tnewotherbits |= newotherbits >> 2;\n \tnewotherbits |= newotherbits >> 4;\n \tnewotherbits = (newotherbits & ~(newotherbits >> 1)) ^ 255;\n-\tc = p->k[newbyte];\n+\tc = p_key[newbyte];\n \tnewdirection = (1 + (newotherbits | c)) >> 8;\n \n \tnode->byte = newbyte;\n@@ -78,7 +87,7 @@ struct cb_node *cb_insert(struct cb_tree *t, struct cb_node *node, size_t klen)\n \t\t\tbreak;\n \t\tif (q->byte == newbyte && q->otherbits > newotherbits)\n \t\t\tbreak;\n-\t\tc = q->byte < klen ? node->k[q->byte] : 0;\n+\t\tc = q->byte < klen ? node_key[q->byte] : 0;\n \t\tdirection = (1 + (q->otherbits | c)) >> 8;\n \t\twherep = q->child + direction;\n \t}\n@@ -93,7 +102,7 @@ struct cb_node *cb_lookup(struct cb_tree *t, const uint8_t *k, size_t klen)\n {\n \tstruct cb_node *p = cb_internal_best_match(t->root, k, klen);\n \n-\treturn p && !memcmp(p->k, k, klen) ? p : NULL;\n+\treturn p && !memcmp(cb_node_key(t, p), k, klen) ? p : NULL;\n }\n \n static int cb_descend(struct cb_node *p, cb_iter fn, void *arg)\n@@ -115,6 +124,7 @@ int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n \tstruct cb_node *p = t->root;\n \tstruct cb_node *top = p;\n \tsize_t i = 0;\n+\tuint8_t *p_key;\n \n \tif (!p)\n \t\treturn 0; /* empty tree */\n@@ -130,8 +140,9 @@ int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n \t\t\ttop = p;\n \t}\n \n+\tp_key = cb_node_key(t, p);\n \tfor (i = 0; i < klen; i++) {\n-\t\tif (p->k[i] != kpfx[i])\n+\t\tif (p_key[i] != kpfx[i])\n \t\t\treturn 0; /* \"best\" match failed */\n \t}\n \ndiff --git a/cbtree.h b/cbtree.h\nindex c374b1b3db..3ce0d6b287 100644\n--- a/cbtree.h\n+++ b/cbtree.h\n@@ -23,18 +23,19 @@ struct cb_node {\n \t */\n \tuint32_t byte;\n \tuint8_t otherbits;\n-\tuint8_t k[FLEX_ARRAY]; /* arbitrary data, unaligned */\n };\n \n struct cb_tree {\n \tstruct cb_node *root;\n+\tptrdiff_t key_offset;\n };\n \n-#define CBTREE_INIT { 0 }\n-\n-static inline void cb_init(struct cb_tree *t)\n+static inline void cb_init(struct cb_tree *t,\n+\t\t\t   ptrdiff_t key_offset)\n {\n-\tstruct cb_tree blank = CBTREE_INIT;\n+\tstruct cb_tree blank = {\n+\t\t.key_offset = key_offset,\n+\t};\n \tmemcpy(t, &blank, sizeof(*t));\n }\n \ndiff --git a/oidtree.c b/oidtree.c\nindex ab9fe7ec7a..117649753f 100644\n--- a/oidtree.c\n+++ b/oidtree.c\n@@ -6,9 +6,14 @@\n #include \"oidtree.h\"\n #include \"hash.h\"\n \n+struct oidtree_node {\n+\tstruct cb_node base;\n+\tstruct object_id key;\n+};\n+\n void oidtree_init(struct oidtree *ot)\n {\n-\tcb_init(&ot->tree);\n+\tcb_init(&ot->tree, offsetof(struct oidtree_node, key));\n \tmem_pool_init(&ot->mem_pool, 0);\n }\n \n@@ -22,20 +27,13 @@ void oidtree_clear(struct oidtree *ot)\n \n void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n {\n-\tstruct cb_node *on;\n-\tstruct object_id k;\n+\tstruct oidtree_node *on;\n \n \tif (!oid->algo)\n \t\tBUG(\"oidtree_insert requires oid->algo\");\n \n-\ton = mem_pool_alloc(&ot->mem_pool, sizeof(*on) + sizeof(*oid));\n-\n-\t/*\n-\t * Clear the padding and copy the result in separate steps to\n-\t * respect the 4-byte alignment needed by struct object_id.\n-\t */\n-\toidcpy(&k, oid);\n-\tmemcpy(on->k, &k, sizeof(k));\n+\ton = mem_pool_alloc(&ot->mem_pool, sizeof(*on));\n+\toidcpy(&on->key, oid);\n \n \t/*\n \t * n.b. Current callers won't get us duplicates, here.  If a\n@@ -43,7 +41,7 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n \t * that won't be freed until oidtree_clear.  Currently it's not\n \t * worth maintaining a free list\n \t */\n-\tcb_insert(&ot->tree, on, sizeof(*oid));\n+\tcb_insert(&ot->tree, &on->base, sizeof(*oid));\n }\n \n bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n@@ -73,21 +71,18 @@ struct oidtree_each_data {\n \n static int iter(struct cb_node *n, void *cb_data)\n {\n+\tstruct oidtree_node *node = container_of(n, struct oidtree_node, base);\n \tstruct oidtree_each_data *data = cb_data;\n-\tstruct object_id k;\n-\n-\t/* Copy to provide 4-byte alignment needed by struct object_id. */\n-\tmemcpy(&k, n->k, sizeof(k));\n \n-\tif (data->algo != GIT_HASH_UNKNOWN && data->algo != k.algo)\n+\tif (data->algo != GIT_HASH_UNKNOWN && data->algo != node->key.algo)\n \t\treturn 0;\n \n \tif (data->last_nibble_at) {\n-\t\tif ((k.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n+\t\tif ((node->key.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n \t\t\treturn 0;\n \t}\n \n-\treturn data->cb(&k, data->cb_data);\n+\treturn data->cb(&node->key, data->cb_data);\n }\n \n int oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541227","messageId":"20260409-b4-pks-odb-source-inmemory-v2-10-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 10/17] oidtree: add ability to store data","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:31Z","receivedAt":"2026-04-09T07:24:58Z","isPatch":true,"body":"The oidtree data structure is currently only used to store object IDs,\nwithout any associated data. So consequently, it can only really be used\nto track which object IDs exist, and we can use the tree structure to\nefficiently operate on OID prefixes.\n\nBut there are valid use cases where we want to both:\n\n  - Store object IDs in a sorted order.\n\n  - Associated arbitrary data with them.\n\nRefactor the oidtree interface so that it allows us to store arbitrary\npayloads within the respective nodes. This will be used in the next\ncommit.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c                  |  2 +-\n object-file.c            |  3 ++-\n oidtree.c                | 37 ++++++++++++++++++++++++++++++++-----\n oidtree.h                | 12 ++++++++++--\n t/unit-tests/u-oidtree.c | 26 +++++++++++++++++++++++---\n 5 files changed, 68 insertions(+), 12 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 07333be696..f7a3dd1a72 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -57,7 +57,7 @@ static int insert_loose_map(struct odb_source *source,\n \tinserted |= insert_oid_pair(map->to_compat, oid, compat_oid);\n \tinserted |= insert_oid_pair(map->to_storage, compat_oid, oid);\n \tif (inserted)\n-\t\toidtree_insert(files->loose->cache, compat_oid);\n+\t\toidtree_insert(files->loose->cache, compat_oid, NULL);\n \n \treturn inserted;\n }\ndiff --git a/object-file.c b/object-file.c\nindex 3e70e5d668..d04ab57253 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1857,6 +1857,7 @@ static int for_each_object_wrapper_cb(const struct object_id *oid,\n }\n \n static int for_each_prefixed_object_wrapper_cb(const struct object_id *oid,\n+\t\t\t\t\t       void *node_data UNUSED,\n \t\t\t\t\t       void *cb_data)\n {\n \tstruct for_each_object_wrapper_data *data = cb_data;\n@@ -2002,7 +2003,7 @@ static int append_loose_object(const struct object_id *oid,\n \t\t\t       const char *path UNUSED,\n \t\t\t       void *data)\n {\n-\toidtree_insert(data, oid);\n+\toidtree_insert(data, oid, NULL);\n \treturn 0;\n }\n \ndiff --git a/oidtree.c b/oidtree.c\nindex 117649753f..e43f18026e 100644\n--- a/oidtree.c\n+++ b/oidtree.c\n@@ -9,6 +9,7 @@\n struct oidtree_node {\n \tstruct cb_node base;\n \tstruct object_id key;\n+\tvoid *data;\n };\n \n void oidtree_init(struct oidtree *ot)\n@@ -25,15 +26,22 @@ void oidtree_clear(struct oidtree *ot)\n \t}\n }\n \n-void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n+struct oidtree_data {\n+\tstruct object_id oid;\n+};\n+\n+void oidtree_insert(struct oidtree *ot, const struct object_id *oid,\n+\t\t    void *data)\n {\n \tstruct oidtree_node *on;\n+\tstruct cb_node *node;\n \n \tif (!oid->algo)\n \t\tBUG(\"oidtree_insert requires oid->algo\");\n \n \ton = mem_pool_alloc(&ot->mem_pool, sizeof(*on));\n \toidcpy(&on->key, oid);\n+\ton->data = data;\n \n \t/*\n \t * n.b. Current callers won't get us duplicates, here.  If a\n@@ -41,13 +49,19 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n \t * that won't be freed until oidtree_clear.  Currently it's not\n \t * worth maintaining a free list\n \t */\n-\tcb_insert(&ot->tree, &on->base, sizeof(*oid));\n+\tnode = cb_insert(&ot->tree, &on->base, sizeof(*oid));\n+\tif (node) {\n+\t\tstruct oidtree_node *preexisting = container_of(node, struct oidtree_node, base);\n+\t\tpreexisting->data = data;\n+\t}\n }\n \n-bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n+static struct oidtree_node *oidtree_lookup(struct oidtree *ot,\n+\t\t\t\t\t   const struct object_id *oid)\n {\n \tstruct object_id k;\n \tsize_t klen = sizeof(k);\n+\tstruct cb_node *node;\n \n \toidcpy(&k, oid);\n \n@@ -58,7 +72,20 @@ bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n \tklen += BUILD_ASSERT_OR_ZERO(offsetof(struct object_id, hash) <\n \t\t\t\toffsetof(struct object_id, algo));\n \n-\treturn !!cb_lookup(&ot->tree, (const uint8_t *)&k, klen);\n+\tnode = cb_lookup(&ot->tree, (const uint8_t *)&k, klen);\n+\treturn node ? container_of(node, struct oidtree_node, base) : NULL;\n+}\n+\n+bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n+{\n+\tstruct oidtree_node *node = oidtree_lookup(ot, oid);\n+\treturn node ? 1 : 0;\n+}\n+\n+void *oidtree_get(struct oidtree *ot, const struct object_id *oid)\n+{\n+\tstruct oidtree_node *node = oidtree_lookup(ot, oid);\n+\treturn node ? node->data : NULL;\n }\n \n struct oidtree_each_data {\n@@ -82,7 +109,7 @@ static int iter(struct cb_node *n, void *cb_data)\n \t\t\treturn 0;\n \t}\n \n-\treturn data->cb(&node->key, data->cb_data);\n+\treturn data->cb(&node->key, node->data, data->cb_data);\n }\n \n int oidtree_each(struct oidtree *ot, const struct object_id *prefix,\ndiff --git a/oidtree.h b/oidtree.h\nindex 2b7bad2e60..baa5a436ea 100644\n--- a/oidtree.h\n+++ b/oidtree.h\n@@ -29,18 +29,26 @@ void oidtree_init(struct oidtree *ot);\n  */\n void oidtree_clear(struct oidtree *ot);\n \n-/* Insert the object ID into the tree. */\n-void oidtree_insert(struct oidtree *ot, const struct object_id *oid);\n+/*\n+ * Insert the object ID into the tree and store the given pointer alongside\n+ * with it. The data pointer of any preexisting entry will be overwritten.\n+ */\n+void oidtree_insert(struct oidtree *ot, const struct object_id *oid,\n+\t\t    void *data);\n \n /* Check whether the tree contains the given object ID. */\n bool oidtree_contains(struct oidtree *ot, const struct object_id *oid);\n \n+/* Get the payload stored with the given object ID. */\n+void *oidtree_get(struct oidtree *ot, const struct object_id *oid);\n+\n /*\n  * Callback function used for `oidtree_each()`. Returning a non-zero exit code\n  * will cause iteration to stop. The exit code will be propagated to the caller\n  * of `oidtree_each()`.\n  */\n typedef int (*oidtree_each_cb)(const struct object_id *oid,\n+\t\t\t       void *node_data,\n \t\t\t       void *cb_data);\n \n /*\ndiff --git a/t/unit-tests/u-oidtree.c b/t/unit-tests/u-oidtree.c\nindex d4d05c7dc3..f0d5ebb733 100644\n--- a/t/unit-tests/u-oidtree.c\n+++ b/t/unit-tests/u-oidtree.c\n@@ -19,7 +19,7 @@ static int fill_tree_loc(struct oidtree *ot, const char *hexes[], size_t n)\n \tfor (size_t i = 0; i < n; i++) {\n \t\tstruct object_id oid;\n \t\tcl_parse_any_oid(hexes[i], &oid);\n-\t\toidtree_insert(ot, &oid);\n+\t\toidtree_insert(ot, &oid, NULL);\n \t}\n \treturn 0;\n }\n@@ -38,9 +38,9 @@ struct expected_hex_iter {\n \tconst char *query;\n };\n \n-static int check_each_cb(const struct object_id *oid, void *data)\n+static int check_each_cb(const struct object_id *oid, void *node_data UNUSED, void *cb_data)\n {\n-\tstruct expected_hex_iter *hex_iter = data;\n+\tstruct expected_hex_iter *hex_iter = cb_data;\n \tstruct object_id expected;\n \n \tcl_assert(hex_iter->i < hex_iter->expected_hexes.nr);\n@@ -105,3 +105,23 @@ void test_oidtree__each(void)\n \tcheck_each(&ot, \"32100\", \"321\", NULL);\n \tcheck_each(&ot, \"32\", \"320\", \"321\", NULL);\n }\n+\n+void test_oidtree__insert_overwrites_data(void)\n+{\n+\tstruct object_id oid;\n+\tstruct oidtree ot;\n+\tint a, b;\n+\n+\tcl_parse_any_oid(\"1\", &oid);\n+\n+\toidtree_init(&ot);\n+\n+\toidtree_insert(&ot, &oid, NULL);\n+\tcl_assert_equal_p(oidtree_get(&ot, &oid), NULL);\n+\toidtree_insert(&ot, &oid, &a);\n+\tcl_assert_equal_p(oidtree_get(&ot, &oid), &a);\n+\toidtree_insert(&ot, &oid, &b);\n+\tcl_assert_equal_p(oidtree_get(&ot, &oid), &b);\n+\n+\toidtree_clear(&ot);\n+}\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541228","messageId":"20260409-b4-pks-odb-source-inmemory-v2-11-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 11/17] odb/source-inmemory: convert to use oidtree","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:32Z","receivedAt":"2026-04-09T07:25:01Z","isPatch":true,"body":"The in-memory source stores its objects in a simple array that we grow as\nneeded. This has a couple of downsides:\n\n  - The object lookup is O(n). This doesn't matter in practice because\n    we only store a small number of objects.\n\n  - We don't have an easy way to iterate over all objects in\n    lexicographic order.\n\n  - We don't have an easy way to compute unique object ID prefixes.\n\nRefactor the code to use an oidtree instead. This is the same data\nstructure used by our loose object source, and thus it means we get a\nbunch of functionality for free.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 72 +++++++++++++++++++++++++++++++++++++--------------\n odb/source-inmemory.h | 13 ++--------\n 2 files changed, 54 insertions(+), 31 deletions(-)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 578ceea550..0420b98d00 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -3,20 +3,29 @@\n #include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n #include \"odb/streaming.h\"\n+#include \"oidtree.h\"\n #include \"repository.h\"\n \n-static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n-\t\t\t\t\t\t      const struct object_id *oid)\n+struct inmemory_object {\n+\tenum object_type type;\n+\tconst void *buf;\n+\tunsigned long size;\n+};\n+\n+static const struct inmemory_object *find_cached_object(struct odb_source_inmemory *source,\n+\t\t\t\t\t\t\tconst struct object_id *oid)\n {\n-\tstatic const struct cached_object empty_tree = {\n+\tstatic const struct inmemory_object empty_tree = {\n \t\t.type = OBJ_TREE,\n \t\t.buf = \"\",\n \t};\n-\tconst struct cached_object_entry *co = source->objects;\n+\tconst struct inmemory_object *object;\n \n-\tfor (size_t i = 0; i < source->objects_nr; i++, co++)\n-\t\tif (oideq(&co->oid, oid))\n-\t\t\treturn &co->value;\n+\tif (source->objects) {\n+\t\tobject = oidtree_get(source->objects, oid);\n+\t\tif (object)\n+\t\t\treturn object;\n+\t}\n \n \tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n \t\treturn &empty_tree;\n@@ -30,7 +39,7 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \t\t\t\t\t\tenum object_info_flags flags UNUSED)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n-\tconst struct cached_object *object;\n+\tconst struct inmemory_object *object;\n \n \tobject = find_cached_object(inmemory, oid);\n \tif (!object)\n@@ -86,7 +95,7 @@ static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n \tstruct odb_read_stream_inmemory *stream;\n-\tconst struct cached_object *object;\n+\tconst struct inmemory_object *object;\n \n \tobject = find_cached_object(inmemory, oid);\n \tif (!object)\n@@ -111,17 +120,23 @@ static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    enum odb_write_object_flags flags UNUSED)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n-\tstruct cached_object_entry *object;\n+\tstruct inmemory_object *object;\n \n \thash_object_file(source->odb->repo->hash_algo, buf, len, type, oid);\n \n-\tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n-\t\t   inmemory->objects_alloc);\n-\tobject = &inmemory->objects[inmemory->objects_nr++];\n-\tobject->value.size = len;\n-\tobject->value.type = type;\n-\tobject->value.buf = xmemdupz(buf, len);\n-\toidcpy(&object->oid, oid);\n+\tif (!inmemory->objects) {\n+\t\tCALLOC_ARRAY(inmemory->objects, 1);\n+\t\toidtree_init(inmemory->objects);\n+\t} else if (oidtree_contains(inmemory->objects, oid)) {\n+\t\treturn 0;\n+\t}\n+\n+\tCALLOC_ARRAY(object, 1);\n+\tobject->size = len;\n+\tobject->type = type;\n+\tobject->buf = xmemdupz(buf, len);\n+\n+\toidtree_insert(inmemory->objects, oid, object);\n \n \treturn 0;\n }\n@@ -165,12 +180,29 @@ static int odb_source_inmemory_write_object_stream(struct odb_source *source,\n \treturn ret;\n }\n \n+static int inmemory_object_free(const struct object_id *oid UNUSED,\n+\t\t\t\tvoid *node_data,\n+\t\t\t\tvoid *cb_data UNUSED)\n+{\n+\tstruct inmemory_object *object = node_data;\n+\tfree((void *) object->buf);\n+\tfree(object);\n+\treturn 0;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n-\tfor (size_t i = 0; i < inmemory->objects_nr; i++)\n-\t\tfree((char *) inmemory->objects[i].value.buf);\n-\tfree(inmemory->objects);\n+\n+\tif (inmemory->objects) {\n+\t\tstruct object_id null_oid = { 0 };\n+\n+\t\toidtree_each(inmemory->objects, &null_oid, 0,\n+\t\t\t     inmemory_object_free, NULL);\n+\t\toidtree_clear(inmemory->objects);\n+\t\tfree(inmemory->objects);\n+\t}\n+\n \tfree(inmemory->base.path);\n \tfree(inmemory);\n }\ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex 14dc06f7c3..02cf586b63 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -3,14 +3,7 @@\n \n #include \"odb/source.h\"\n \n-struct cached_object_entry {\n-\tstruct object_id oid;\n-\tstruct cached_object {\n-\t\tenum object_type type;\n-\t\tconst void *buf;\n-\t\tunsigned long size;\n-\t} value;\n-};\n+struct oidtree;\n \n /*\n  * An inmemory source that you can write objects to that shall be made\n@@ -20,9 +13,7 @@ struct cached_object_entry {\n  */\n struct odb_source_inmemory {\n \tstruct odb_source base;\n-\n-\tstruct cached_object_entry *objects;\n-\tsize_t objects_nr, objects_alloc;\n+\tstruct oidtree *objects;\n };\n \n /* Create a new in-memory object database source. */\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541229","messageId":"20260409-b4-pks-odb-source-inmemory-v2-12-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 12/17] odb/source-inmemory: implement `for_each_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:33Z","receivedAt":"2026-04-09T07:25:04Z","isPatch":true,"body":"Implement the `for_each_object()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 86 +++++++++++++++++++++++++++++++++++++++++----------\n 1 file changed, 70 insertions(+), 16 deletions(-)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 0420b98d00..d1674836cc 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -33,6 +33,28 @@ static const struct inmemory_object *find_cached_object(struct odb_source_inmemo\n \treturn NULL;\n }\n \n+static void populate_object_info(struct odb_source_inmemory *source,\n+\t\t\t\t struct object_info *oi,\n+\t\t\t\t const struct inmemory_object *object)\n+{\n+\tif (!oi)\n+\t\treturn;\n+\n+\tif (oi->typep)\n+\t\t*(oi->typep) = object->type;\n+\tif (oi->sizep)\n+\t\t*(oi->sizep) = object->size;\n+\tif (oi->disk_sizep)\n+\t\t*(oi->disk_sizep) = 0;\n+\tif (oi->delta_base_oid)\n+\t\toidclr(oi->delta_base_oid, source->base.odb->repo->hash_algo);\n+\tif (oi->contentp)\n+\t\t*oi->contentp = xmemdupz(object->buf, object->size);\n+\tif (oi->mtimep)\n+\t\t*oi->mtimep = 0;\n+\toi->whence = OI_CACHED;\n+}\n+\n static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \t\t\t\t\t\tconst struct object_id *oid,\n \t\t\t\t\t\tstruct object_info *oi,\n@@ -45,22 +67,7 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \tif (!object)\n \t\treturn -1;\n \n-\tif (oi) {\n-\t\tif (oi->typep)\n-\t\t\t*(oi->typep) = object->type;\n-\t\tif (oi->sizep)\n-\t\t\t*(oi->sizep) = object->size;\n-\t\tif (oi->disk_sizep)\n-\t\t\t*(oi->disk_sizep) = 0;\n-\t\tif (oi->delta_base_oid)\n-\t\t\toidclr(oi->delta_base_oid, source->odb->repo->hash_algo);\n-\t\tif (oi->contentp)\n-\t\t\t*oi->contentp = xmemdupz(object->buf, object->size);\n-\t\tif (oi->mtimep)\n-\t\t\t*oi->mtimep = 0;\n-\t\toi->whence = OI_CACHED;\n-\t}\n-\n+\tpopulate_object_info(inmemory, oi, object);\n \treturn 0;\n }\n \n@@ -112,6 +119,52 @@ static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n \treturn 0;\n }\n \n+struct odb_source_inmemory_for_each_object_data {\n+\tstruct odb_source_inmemory *inmemory;\n+\tconst struct object_info *request;\n+\todb_for_each_object_cb cb;\n+\tvoid *cb_data;\n+};\n+\n+static int odb_source_inmemory_for_each_object_cb(const struct object_id *oid,\n+\t\t\t\t\t\t  void *node_data, void *cb_data)\n+{\n+\tstruct odb_source_inmemory_for_each_object_data *data = cb_data;\n+\tstruct inmemory_object *object = node_data;\n+\n+\tif (data->request) {\n+\t\tstruct object_info oi = *data->request;\n+\t\tpopulate_object_info(data->inmemory, &oi, object);\n+\t\treturn data->cb(oid, &oi, data->cb_data);\n+\t} else {\n+\t\treturn data->cb(oid, NULL, data->cb_data);\n+\t}\n+}\n+\n+static int odb_source_inmemory_for_each_object(struct odb_source *source,\n+\t\t\t\t\t       const struct object_info *request,\n+\t\t\t\t\t       odb_for_each_object_cb cb,\n+\t\t\t\t\t       void *cb_data,\n+\t\t\t\t\t       const struct odb_for_each_object_options *opts)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tstruct odb_source_inmemory_for_each_object_data payload = {\n+\t\t.inmemory = inmemory,\n+\t\t.request = request,\n+\t\t.cb = cb,\n+\t\t.cb_data = cb_data,\n+\t};\n+\tstruct object_id null_oid = { 0 };\n+\n+\tif ((opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY) ||\n+\t    (opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY && !source->local))\n+\t\treturn 0;\n+\n+\treturn oidtree_each(inmemory->objects,\n+\t\t\t    opts->prefix ? opts->prefix : &null_oid, opts->prefix_hex_len,\n+\t\t\t    odb_source_inmemory_for_each_object_cb, &payload);\n+}\n+\n static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    const void *buf, unsigned long len,\n \t\t\t\t\t    enum object_type type,\n@@ -217,6 +270,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.free = odb_source_inmemory_free;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n+\tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541230","messageId":"20260409-b4-pks-odb-source-inmemory-v2-13-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 13/17] odb/source-inmemory: implement `find_abbrev_len()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:34Z","receivedAt":"2026-04-09T07:25:06Z","isPatch":true,"body":"Implement the `find_abbrev_len()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 39 +++++++++++++++++++++++++++++++++++++++\n 1 file changed, 39 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex d1674836cc..a8eba373ee 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -165,6 +165,44 @@ static int odb_source_inmemory_for_each_object(struct odb_source *source,\n \t\t\t    odb_source_inmemory_for_each_object_cb, &payload);\n }\n \n+struct find_abbrev_len_data {\n+\tconst struct object_id *oid;\n+\tunsigned len;\n+};\n+\n+static int find_abbrev_len_cb(const struct object_id *oid,\n+\t\t\t      struct object_info *oi UNUSED,\n+\t\t\t      void *cb_data)\n+{\n+\tstruct find_abbrev_len_data *data = cb_data;\n+\tunsigned len = oid_common_prefix_hexlen(oid, data->oid);\n+\tif (len != hash_algos[oid->algo].hexsz && len >= data->len)\n+\t\tdata->len = len + 1;\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t\t       const struct object_id *oid,\n+\t\t\t\t\t       unsigned min_len,\n+\t\t\t\t\t       unsigned *out)\n+{\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = oid,\n+\t\t.prefix_hex_len = min_len,\n+\t};\n+\tstruct find_abbrev_len_data data = {\n+\t\t.oid = oid,\n+\t\t.len = min_len,\n+\t};\n+\tint ret;\n+\n+\tret = odb_source_inmemory_for_each_object(source, NULL, find_abbrev_len_cb,\n+\t\t\t\t\t\t  &data, &opts);\n+\t*out = data.len;\n+\n+\treturn ret;\n+}\n+\n static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    const void *buf, unsigned long len,\n \t\t\t\t\t    enum object_type type,\n@@ -271,6 +309,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n+\tsource->base.find_abbrev_len = odb_source_inmemory_find_abbrev_len;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541231","messageId":"20260409-b4-pks-odb-source-inmemory-v2-14-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 14/17] odb/source-inmemory: implement `count_objects()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:35Z","receivedAt":"2026-04-09T07:25:09Z","isPatch":true,"body":"Implement the `count_objects()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 20 ++++++++++++++++++++\n 1 file changed, 20 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex a8eba373ee..f038debaa3 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -203,6 +203,25 @@ static int odb_source_inmemory_find_abbrev_len(struct odb_source *source,\n \treturn ret;\n }\n \n+static int count_objects_cb(const struct object_id *oid UNUSED,\n+\t\t\t    struct object_info *oi UNUSED,\n+\t\t\t    void *cb_data)\n+{\n+\tunsigned long *counter = cb_data;\n+\t(*counter)++;\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_count_objects(struct odb_source *source,\n+\t\t\t\t\t     enum odb_count_objects_flags flags UNUSED,\n+\t\t\t\t\t     unsigned long *out)\n+{\n+\tstruct odb_for_each_object_options opts = { 0 };\n+\t*out = 0;\n+\treturn odb_source_inmemory_for_each_object(source, NULL, count_objects_cb,\n+\t\t\t\t\t\t   out, &opts);\n+}\n+\n static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    const void *buf, unsigned long len,\n \t\t\t\t\t    enum object_type type,\n@@ -310,6 +329,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n \tsource->base.find_abbrev_len = odb_source_inmemory_find_abbrev_len;\n+\tsource->base.count_objects = odb_source_inmemory_count_objects;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541232","messageId":"20260409-b4-pks-odb-source-inmemory-v2-15-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 15/17] odb/source-inmemory: implement `freshen_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:36Z","receivedAt":"2026-04-09T07:25:11Z","isPatch":true,"body":"Implement the `freshen_object()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 10 ++++++++++\n 1 file changed, 10 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex f038debaa3..15a6a5ae64 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -290,6 +290,15 @@ static int odb_source_inmemory_write_object_stream(struct odb_source *source,\n \treturn ret;\n }\n \n+static int odb_source_inmemory_freshen_object(struct odb_source *source,\n+\t\t\t\t\t      const struct object_id *oid)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tif (find_cached_object(inmemory, oid))\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n static int inmemory_object_free(const struct object_id *oid UNUSED,\n \t\t\t\tvoid *node_data,\n \t\t\t\tvoid *cb_data UNUSED)\n@@ -332,6 +341,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.count_objects = odb_source_inmemory_count_objects;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n+\tsource->base.freshen_object = odb_source_inmemory_freshen_object;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541233","messageId":"20260409-b4-pks-odb-source-inmemory-v2-16-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 16/17] odb/source-inmemory: stub out remaining functions","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:37Z","receivedAt":"2026-04-09T07:25:14Z","isPatch":true,"body":"Stub out remaining functions that we either don't need or that are\nbasically no-ops.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 31 +++++++++++++++++++++++++++++++\n 1 file changed, 31 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 15a6a5ae64..1140b1b916 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -299,6 +299,32 @@ static int odb_source_inmemory_freshen_object(struct odb_source *source,\n \treturn 0;\n }\n \n+static int odb_source_inmemory_begin_transaction(struct odb_source *source UNUSED,\n+\t\t\t\t\t\t struct odb_transaction **out UNUSED)\n+{\n+\treturn error(\"inmemory source does not support transactions\");\n+}\n+\n+static int odb_source_inmemory_read_alternates(struct odb_source *source UNUSED,\n+\t\t\t\t\t       struct strvec *out UNUSED)\n+{\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_write_alternate(struct odb_source *source UNUSED,\n+\t\t\t\t\t       const char *alternate UNUSED)\n+{\n+\treturn error(\"inmemory source does not support alternates\");\n+}\n+\n+static void odb_source_inmemory_close(struct odb_source *source UNUSED)\n+{\n+}\n+\n+static void odb_source_inmemory_reprepare(struct odb_source *source UNUSED)\n+{\n+}\n+\n static int inmemory_object_free(const struct object_id *oid UNUSED,\n \t\t\t\tvoid *node_data,\n \t\t\t\tvoid *cb_data UNUSED)\n@@ -334,6 +360,8 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n \n \tsource->base.free = odb_source_inmemory_free;\n+\tsource->base.close = odb_source_inmemory_close;\n+\tsource->base.reprepare = odb_source_inmemory_reprepare;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n@@ -342,6 +370,9 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \tsource->base.freshen_object = odb_source_inmemory_freshen_object;\n+\tsource->base.begin_transaction = odb_source_inmemory_begin_transaction;\n+\tsource->base.read_alternates = odb_source_inmemory_read_alternates;\n+\tsource->base.write_alternate = odb_source_inmemory_write_alternate;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541234","messageId":"20260409-b4-pks-odb-source-inmemory-v2-17-f02b4f1c0f13@pks.im","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"[PATCH v2 17/17] odb: generic in-memory source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T07:24:38Z","receivedAt":"2026-04-09T07:25:16Z","isPatch":true,"body":"Make the in-memory source generic.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c | 8 ++++----\n odb.h | 2 +-\n 2 files changed, 5 insertions(+), 5 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 24e929f03c..965ef68e4e 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -560,7 +560,7 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \tif (is_null_oid(real))\n \t\treturn -1;\n \n-\tif (!odb_source_read_object_info(&odb->inmemory_objects->base, oid, oi, flags))\n+\tif (!odb_source_read_object_info(odb->inmemory_objects, oid, oi, flags))\n \t\treturn 0;\n \n \todb_prepare_alternates(odb);\n@@ -737,7 +737,7 @@ int odb_pretend_object(struct object_database *odb,\n \tif (odb_has_object(odb, oid, 0))\n \t\treturn 0;\n \n-\treturn odb_source_write_object(&odb->inmemory_objects->base,\n+\treturn odb_source_write_object(odb->inmemory_objects,\n \t\t\t\t       buf, len, type, oid, NULL, 0);\n }\n \n@@ -1020,7 +1020,7 @@ struct object_database *odb_new(struct repository *repo,\n \to->sources = odb_source_new(o, primary_source, true);\n \to->sources_tail = &o->sources->next;\n \to->alternate_db = xstrdup_or_null(secondary_sources);\n-\to->inmemory_objects = odb_source_inmemory_new(o);\n+\to->inmemory_objects = &odb_source_inmemory_new(o)->base;\n \n \tfree(to_free);\n \n@@ -1045,7 +1045,7 @@ static void odb_free_sources(struct object_database *o)\n \t\to->sources = next;\n \t}\n \n-\todb_source_free(&o->inmemory_objects->base);\n+\todb_source_free(o->inmemory_objects);\n \to->inmemory_objects = NULL;\n \n \tkh_destroy_odb_path_map(o->source_by_path);\ndiff --git a/odb.h b/odb.h\nindex c3a7edf9c8..73553ed5a7 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -81,7 +81,7 @@ struct object_database {\n \t * to write them into the object store (e.g. a browse-only\n \t * application).\n \t */\n-\tstruct odb_source_inmemory *inmemory_objects;\n+\tstruct odb_source *inmemory_objects;\n \n \t/*\n \t * A fast, rough count of the number of objects in the repository.\n\n-- \n2.54.0.rc0.680.geaeac8ef83.dirty\n\n"},{"id":"541236","messageId":"CAOLa=ZRkctXNkpTqiTSTkvskajPZTid9WTG3fKr0YV641_5qrw@mail.gmail.com","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-1-f02b4f1c0f13@pks.im","subject":"Re: [PATCH v2 01/17] odb: introduce \"in-memory\" source","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-04-09T09:26:50Z","receivedAt":"2026-04-09T09:26:53Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Next to our typical object database sources, each object database also\n> has an implicit source of \"cached\" objects. These cached objects only\n> exist in memory and some use cases:\n>\n>   - They contain evergreen objects that we expect to always exist, like\n>     for example the empty tree.\n>\n>   - They can be used to store temporary objects that we don't want to\n>     persist to disk, which is used by git-blame(1) to create a fake\n>     worktree commit.\n>\n> Overall, their use is somewhat restricted though. For example, we don't\n> provide the ability to use it as a temporary object database source that\n> allows the user to write objects, but discard them after Git exists. So\n> while these cached objects behave almost like a source, they aren't used\n> as one.\n>\n> This is about to change over the following commits, where we will turn\n> cached objects into a new \"in-memory\" source. This will allow us to use\n> it exactly the same as any other source by providing the same common\n> interface as the \"files\" source.\n>\n> For now, the in-memory source only hosts the cached objects and doesn't\n> provide any logic yet. This will change with subsequent commits, where\n> we move respective functionality into the source.\n\n[snip]\n\n> diff --git a/odb.c b/odb.c\n> index 40a5e9c4e0..60e1eead25 100644\n> --- a/odb.c\n> +++ b/odb.c\n> @@ -14,6 +14,7 @@\n>  #include \"object-file.h\"\n>  #include \"object-name.h\"\n>  #include \"odb.h\"\n> +#include \"odb/source-inmemory.h\"\n>  #include \"packfile.h\"\n>  #include \"path.h\"\n>  #include \"promisor-remote.h\"\n> @@ -53,9 +54,9 @@ static const struct cached_object *find_cached_object(struct object_database *ob\n>  \t\t.type = OBJ_TREE,\n>  \t\t.buf = \"\",\n>  \t};\n> -\tconst struct cached_object_entry *co = object_store->cached_objects;\n> +\tconst struct cached_object_entry *co = object_store->inmemory_objects->objects;\n>\n> -\tfor (size_t i = 0; i < object_store->cached_object_nr; i++, co++)\n> +\tfor (size_t i = 0; i < object_store->inmemory_objects->objects_nr; i++, co++)\n>  \t\tif (oideq(&co->oid, oid))\n>  \t\t\treturn &co->value;\n>\n> @@ -792,9 +793,10 @@ int odb_pretend_object(struct object_database *odb,\n>  \t    find_cached_object(odb, oid))\n>  \t\treturn 0;\n>\n> -\tALLOC_GROW(odb->cached_objects,\n> -\t\t   odb->cached_object_nr + 1, odb->cached_object_alloc);\n> -\tco = &odb->cached_objects[odb->cached_object_nr++];\n> +\tALLOC_GROW(odb->inmemory_objects->objects,\n> +\t\t   odb->inmemory_objects->objects_nr + 1,\n> +\t\t   odb->inmemory_objects->objects_alloc);\n> +\tco = &odb->inmemory_objects->objects[odb->inmemory_objects->objects_nr++];\n\nOkay so we introduce the inmemory object storage and directly write\nobjects to it. I guess in the upcoming commits, we'll swap to using the\nAPI as we implement them.\n\nMakes sense for now.\n\n>  \tco->value.size = len;\n>  \tco->value.type = type;\n>  \tco_buf = xmalloc(len);\n> @@ -1083,6 +1085,7 @@ struct object_database *odb_new(struct repository *repo,\n>  \to->sources = odb_source_new(o, primary_source, true);\n>  \to->sources_tail = &o->sources->next;\n>  \to->alternate_db = xstrdup_or_null(secondary_sources);\n> +\to->inmemory_objects = odb_source_inmemory_new(o);\n>\n>  \tfree(to_free);\n>\n> @@ -1123,9 +1126,11 @@ void odb_free(struct object_database *o)\n>  \todb_close(o);\n>  \todb_free_sources(o);\n>\n> -\tfor (size_t i = 0; i < o->cached_object_nr; i++)\n> -\t\tfree((char *) o->cached_objects[i].value.buf);\n> -\tfree(o->cached_objects);\n> +\tfor (size_t i = 0; i < o->inmemory_objects->objects_nr; i++)\n> +\t\tfree((char *) o->inmemory_objects->objects[i].value.buf);\n> +\tfree(o->inmemory_objects->objects);\n> +\tfree(o->inmemory_objects->base.path);\n> +\tfree(o->inmemory_objects);\n>\n>  \tstring_list_clear(&o->submodule_source_paths, 0);\n>\n> diff --git a/odb.h b/odb.h\n> index 9eb8355aca..c3a7edf9c8 100644\n> --- a/odb.h\n> +++ b/odb.h\n> @@ -8,6 +8,7 @@\n>  #include \"thread-utils.h\"\n>\n>  struct cached_object_entry;\n> +struct odb_source_inmemory;\n>  struct packed_git;\n>  struct repository;\n>  struct strbuf;\n> @@ -80,8 +81,7 @@ struct object_database {\n>  \t * to write them into the object store (e.g. a browse-only\n>  \t * application).\n>  \t */\n> -\tstruct cached_object_entry *cached_objects;\n> -\tsize_t cached_object_nr, cached_object_alloc;\n> +\tstruct odb_source_inmemory *inmemory_objects;\n>\n>  \t/*\n>  \t * A fast, rough count of the number of objects in the repository.\n> diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> new file mode 100644\n> index 0000000000..c7ac5c24f0\n> --- /dev/null\n> +++ b/odb/source-inmemory.c\n> @@ -0,0 +1,12 @@\n> +#include \"git-compat-util.h\"\n> +#include \"odb/source-inmemory.h\"\n> +\n> +struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n> +{\n> +\tstruct odb_source_inmemory *source;\n> +\n> +\tCALLOC_ARRAY(source, 1);\n> +\todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n> +\n> +\treturn source;\n> +}\n> diff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\n> new file mode 100644\n> index 0000000000..95477bf36d\n> --- /dev/null\n> +++ b/odb/source-inmemory.h\n> @@ -0,0 +1,35 @@\n> +#ifndef ODB_SOURCE_INMEMORY_H\n> +#define ODB_SOURCE_INMEMORY_H\n> +\n> +#include \"odb/source.h\"\n> +\n> +struct cached_object_entry;\n> +\n> +/*\n> + * An inmemory source that you can write objects to that shall be made\n> + * available for reading, but that shouldn't ever be persisted to disk. Note\n> + * that any objects written to this source will be stored in memory, so the\n> + * number of objects you can store is limited by available system memory.\n> + */\n> +struct odb_source_inmemory {\n> +\tstruct odb_source base;\n> +\n> +\tstruct cached_object_entry *objects;\n> +\tsize_t objects_nr, objects_alloc;\n> +};\n> +\n> +/* Create a new in-memory object database source. */\n> +struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb);\n> +\n> +/*\n> + * Cast the given object database source to the inmemory backend. This will\n> + * cause a BUG in case the source doesn't use this backend.\n> + */\n> +static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n> +{\n> +\tif (source->type != ODB_SOURCE_INMEMORY)\n> +\t\tBUG(\"trying to downcast source of type '%d' to inmemory\", source->type);\n> +\treturn container_of(source, struct odb_source_inmemory, base);\n> +}\n> +\n\nInteresting, in the refs namespace the downcast functions are added to\nthe source file (.c). This works too, is there any reason though?\n\n[snip]\n"},{"id":"541237","messageId":"CAOLa=ZRwv_NYqtNyvhi=5auLhVx+FDbt+RP6Kj_ZqjF=VsefyA@mail.gmail.com","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-4-f02b4f1c0f13@pks.im","subject":"Re: [PATCH v2 04/17] odb/source-inmemory: implement `read_object_info()` callback","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-04-09T09:40:01Z","receivedAt":"2026-04-09T09:40:03Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n[snip]\n\n> diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> index ccbb622eae..12c80f9b34 100644\n> --- a/odb/source-inmemory.c\n> +++ b/odb/source-inmemory.c\n> @@ -1,5 +1,57 @@\n>  #include \"git-compat-util.h\"\n> +#include \"odb.h\"\n>  #include \"odb/source-inmemory.h\"\n> +#include \"repository.h\"\n> +\n> +static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n> +\t\t\t\t\t\t      const struct object_id *oid)\n> +{\n> +\tstatic const struct cached_object empty_tree = {\n> +\t\t.type = OBJ_TREE,\n> +\t\t.buf = \"\",\n> +\t};\n> +\tconst struct cached_object_entry *co = source->objects;\n> +\n> +\tfor (size_t i = 0; i < source->objects_nr; i++, co++)\n> +\t\tif (oideq(&co->oid, oid))\n> +\t\t\treturn &co->value;\n> +\n> +\tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n> +\t\treturn &empty_tree;\n> +\n\nSilly questiong, would it make more sense to check for empty_tree before\niterating over all objects?\n\nThe rest looks good\n\n[snip]\n"},{"id":"541239","messageId":"CAOLa=ZSHAF25zbJ=KHp=u0pFCpAHb-jd45A3dxTSn9pwKHkxFQ@mail.gmail.com","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-5-f02b4f1c0f13@pks.im","subject":"Re: [PATCH v2 05/17] odb/source-inmemory: implement `read_object_stream()` callback","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-04-09T09:49:32Z","receivedAt":"2026-04-09T09:49:33Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Implement the `read_object_stream()` callback function for the in-memory\n> source.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  odb/source-inmemory.c | 50 ++++++++++++++++++++++++++++++++++++++++++++++++++\n>  1 file changed, 50 insertions(+)\n>\n> diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> index 12c80f9b34..4a68169430 100644\n> --- a/odb/source-inmemory.c\n> +++ b/odb/source-inmemory.c\n> @@ -1,6 +1,7 @@\n>  #include \"git-compat-util.h\"\n>  #include \"odb.h\"\n>  #include \"odb/source-inmemory.h\"\n> +#include \"odb/streaming.h\"\n>  #include \"repository.h\"\n>\n>  static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n> @@ -53,6 +54,54 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n>  \treturn 0;\n>  }\n>\n> +struct odb_read_stream_inmemory {\n> +\tstruct odb_read_stream base;\n> +\tconst void *buf;\n> +\tsize_t offset;\n> +};\n> +\n\nTo stream objects, we have a new structure which is used in the callback.\n\n> +static ssize_t odb_read_stream_inmemory_read(struct odb_read_stream *stream,\n> +\t\t\t\t\t     char *buf, size_t buf_len)\n> +{\n> +\tstruct odb_read_stream_inmemory *inmemory =\n> +\t\tcontainer_of(stream, struct odb_read_stream_inmemory, base);\n> +\tsize_t bytes = buf_len;\n\n\n\n> +\tif (buf_len > inmemory->base.size - inmemory->offset)\n> +\t\tbytes = inmemory->base.size - inmemory->offset;\n> +\tmemcpy(buf, inmemory->buf, bytes);\n> +\n\nShouldn't the offset also be set and we only memcpy offset onwards?\n\n> +\treturn bytes;\n> +}\n> +\n> +static int odb_read_stream_inmemory_close(struct odb_read_stream *stream UNUSED)\n> +{\n> +\treturn 0;\n> +}\n> +\n> +static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n> +\t\t\t\t\t\t  struct odb_source *source,\n> +\t\t\t\t\t\t  const struct object_id *oid)\n> +{\n> +\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n> +\tstruct odb_read_stream_inmemory *stream;\n> +\tconst struct cached_object *object;\n> +\n> +\tobject = find_cached_object(inmemory, oid);\n> +\tif (!object)\n> +\t\treturn -1;\n> +\n> +\tCALLOC_ARRAY(stream, 1);\n> +\tstream->base.read = odb_read_stream_inmemory_read;\n> +\tstream->base.close = odb_read_stream_inmemory_close;\n> +\tstream->base.size = object->size;\n> +\tstream->base.type = object->type;\n> +\tstream->buf = object->buf;\n> +\n\nSo the object is simply mapped to the structure which is propagated in\n`read()`. Since we don't copy any new data over, `close()` has nothing\nto do.\n\n[snip]\n"},{"id":"541240","messageId":"CAOLa=ZQHyhDGGLLcGBjFwG9FOtvjpyjgmrnOO_u3rwZyAYoDHQ@mail.gmail.com","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-7-f02b4f1c0f13@pks.im","subject":"Re: [PATCH v2 07/17] odb/source-inmemory: implement `write_object()` callback","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-04-09T10:27:27Z","receivedAt":"2026-04-09T10:27:29Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Implement the `write_object()` callback function for the in-memory\n> source.\n>\n\nrebase error? Seems like the commit message as the last commit.\n\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  odb/source-inmemory.c | 3 +++\n>  1 file changed, 3 insertions(+)\n>\n> diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> index d2fc4c4054..96e8efd327 100644\n> --- a/odb/source-inmemory.c\n> +++ b/odb/source-inmemory.c\n> @@ -1,4 +1,5 @@\n>  #include \"git-compat-util.h\"\n> +#include \"object-file.h\"\n>  #include \"odb.h\"\n>  #include \"odb/source-inmemory.h\"\n>  #include \"odb/streaming.h\"\n> @@ -112,6 +113,8 @@ static int odb_source_inmemory_write_object(struct odb_source *source,\n>  \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n>  \tstruct cached_object_entry *object;\n>\n> +\thash_object_file(source->odb->repo->hash_algo, buf, len, type, oid);\n> +\n>  \tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n>  \t\t   inmemory->objects_alloc);\n>  \tobject = &inmemory->objects[inmemory->objects_nr++];\n>\n> --\n> 2.54.0.rc0.680.geaeac8ef83.dirty\n"},{"id":"541241","messageId":"adeCXRAWvho7Qqaz@pks.im","threadId":"65423","inReplyTo":"CAOLa=ZRkctXNkpTqiTSTkvskajPZTid9WTG3fKr0YV641_5qrw@mail.gmail.com","subject":"Re: [PATCH v2 01/17] odb: introduce \"in-memory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T10:41:33Z","receivedAt":"2026-04-09T10:41:40Z","isPatch":true,"body":"On Thu, Apr 09, 2026 at 05:26:50AM -0400, Karthik Nayak wrote:\n> > diff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\n> > new file mode 100644\n> > index 0000000000..95477bf36d\n> > --- /dev/null\n> > +++ b/odb/source-inmemory.h\n> > @@ -0,0 +1,35 @@\n> > +#ifndef ODB_SOURCE_INMEMORY_H\n> > +#define ODB_SOURCE_INMEMORY_H\n> > +\n> > +#include \"odb/source.h\"\n> > +\n> > +struct cached_object_entry;\n> > +\n> > +/*\n> > + * An inmemory source that you can write objects to that shall be made\n> > + * available for reading, but that shouldn't ever be persisted to disk. Note\n> > + * that any objects written to this source will be stored in memory, so the\n> > + * number of objects you can store is limited by available system memory.\n> > + */\n> > +struct odb_source_inmemory {\n> > +\tstruct odb_source base;\n> > +\n> > +\tstruct cached_object_entry *objects;\n> > +\tsize_t objects_nr, objects_alloc;\n> > +};\n> > +\n> > +/* Create a new in-memory object database source. */\n> > +struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb);\n> > +\n> > +/*\n> > + * Cast the given object database source to the inmemory backend. This will\n> > + * cause a BUG in case the source doesn't use this backend.\n> > + */\n> > +static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n> > +{\n> > +\tif (source->type != ODB_SOURCE_INMEMORY)\n> > +\t\tBUG(\"trying to downcast source of type '%d' to inmemory\", source->type);\n> > +\treturn container_of(source, struct odb_source_inmemory, base);\n> > +}\n> > +\n> \n> Interesting, in the refs namespace the downcast functions are added to\n> the source file (.c). This works too, is there any reason though?\n\nBy having it static inline over here we can basically ensure that the\ncompiler can inline this call everywhere. I doubt that it really matters\nin the end, but I guess it doesn't hurt, either.\n\nPatrick\n"},{"id":"541242","messageId":"adeCY8QyvDnQdJU2@pks.im","threadId":"65423","inReplyTo":"CAOLa=ZRwv_NYqtNyvhi=5auLhVx+FDbt+RP6Kj_ZqjF=VsefyA@mail.gmail.com","subject":"Re: [PATCH v2 04/17] odb/source-inmemory: implement `read_object_info()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T10:41:39Z","receivedAt":"2026-04-09T10:41:44Z","isPatch":true,"body":"On Thu, Apr 09, 2026 at 05:40:01AM -0400, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> > index ccbb622eae..12c80f9b34 100644\n> > --- a/odb/source-inmemory.c\n> > +++ b/odb/source-inmemory.c\n> > @@ -1,5 +1,57 @@\n> >  #include \"git-compat-util.h\"\n> > +#include \"odb.h\"\n> >  #include \"odb/source-inmemory.h\"\n> > +#include \"repository.h\"\n> > +\n> > +static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n> > +\t\t\t\t\t\t      const struct object_id *oid)\n> > +{\n> > +\tstatic const struct cached_object empty_tree = {\n> > +\t\t.type = OBJ_TREE,\n> > +\t\t.buf = \"\",\n> > +\t};\n> > +\tconst struct cached_object_entry *co = source->objects;\n> > +\n> > +\tfor (size_t i = 0; i < source->objects_nr; i++, co++)\n> > +\t\tif (oideq(&co->oid, oid))\n> > +\t\t\treturn &co->value;\n> > +\n> > +\tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n> > +\t\treturn &empty_tree;\n> > +\n> \n> Silly questiong, would it make more sense to check for empty_tree before\n> iterating over all objects?\n> \n> The rest looks good\n\nMaybe? I guess for now reading the empty tree is the most important use\ncase we have for the in-memory backend, as we only write in-memory\nobjects in a single caller. On the other hand, `source->objects_nr`\nwould be zero in all the other cases, and jumping over the loop should\nbe fast enough to not matter in practice.\n\nAn alternative I was thinking about is to store the empty tree the same\nway as we store all the other objects so that we don't have to special\ncase anything. That has the benefit that we can actually modify the tree\nobject, too, which may eventually become relevant with regards to an\nobject's mtime that we may want to update. The downside is that we have\nanother allocation here and need to eagerly initialize the data\nstructure that stores the objects.\n\nPatrick\n"},{"id":"541243","messageId":"adeCaAonGgwJFbQ5@pks.im","threadId":"65423","inReplyTo":"CAOLa=ZSHAF25zbJ=KHp=u0pFCpAHb-jd45A3dxTSn9pwKHkxFQ@mail.gmail.com","subject":"Re: [PATCH v2 05/17] odb/source-inmemory: implement `read_object_stream()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T10:41:44Z","receivedAt":"2026-04-09T10:41:49Z","isPatch":true,"body":"On Thu, Apr 09, 2026 at 05:49:32AM -0400, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > Implement the `read_object_stream()` callback function for the in-memory\n> > source.\n> >\n> > Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> > ---\n> >  odb/source-inmemory.c | 50 ++++++++++++++++++++++++++++++++++++++++++++++++++\n> >  1 file changed, 50 insertions(+)\n> >\n> > diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> > index 12c80f9b34..4a68169430 100644\n> > --- a/odb/source-inmemory.c\n> > +++ b/odb/source-inmemory.c\n> > @@ -53,6 +54,54 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n> >  \treturn 0;\n> >  }\n> >\n> > +struct odb_read_stream_inmemory {\n> > +\tstruct odb_read_stream base;\n> > +\tconst void *buf;\n> > +\tsize_t offset;\n> > +};\n> > +\n> \n> To stream objects, we have a new structure which is used in the callback.\n> \n> > +static ssize_t odb_read_stream_inmemory_read(struct odb_read_stream *stream,\n> > +\t\t\t\t\t     char *buf, size_t buf_len)\n> > +{\n> > +\tstruct odb_read_stream_inmemory *inmemory =\n> > +\t\tcontainer_of(stream, struct odb_read_stream_inmemory, base);\n> > +\tsize_t bytes = buf_len;\n> \n> \n> \n> > +\tif (buf_len > inmemory->base.size - inmemory->offset)\n> > +\t\tbytes = inmemory->base.size - inmemory->offset;\n> > +\tmemcpy(buf, inmemory->buf, bytes);\n> > +\n> \n> Shouldn't the offset also be set and we only memcpy offset onwards?\n\nOh, good catch. We don't have any users of this API yet, which is why it\nwent undetected. Will fix, thanks.\n\nPatrick\n"},{"id":"541244","messageId":"adeCbStzfZS40IYj@pks.im","threadId":"65423","inReplyTo":"CAOLa=ZQHyhDGGLLcGBjFwG9FOtvjpyjgmrnOO_u3rwZyAYoDHQ@mail.gmail.com","subject":"Re: [PATCH v2 07/17] odb/source-inmemory: implement `write_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T10:41:49Z","receivedAt":"2026-04-09T10:41:54Z","isPatch":true,"body":"On Thu, Apr 09, 2026 at 06:27:27AM -0400, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > Implement the `write_object()` callback function for the in-memory\n> > source.\n> >\n> \n> rebase error? Seems like the commit message as the last commit.\n\nI saw the empty new commit in the range diff, but somehow didn't get\nwhat was happening. But yes, this obviously needs to be squashed into\nthe preceding commit, thanks!\n\nPatrick\n"},{"id":"541248","messageId":"CAOLa=ZTcXerM6_zof5q6Kfav4N=MWZSjTJWSZJPLWZbW=+sNHA@mail.gmail.com","threadId":"65423","inReplyTo":"adeCY8QyvDnQdJU2@pks.im","subject":"Re: [PATCH v2 04/17] odb/source-inmemory: implement `read_object_info()` callback","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-04-09T11:22:42Z","receivedAt":"2026-04-09T11:22:44Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> On Thu, Apr 09, 2026 at 05:40:01AM -0400, Karthik Nayak wrote:\n>> Patrick Steinhardt <ps@pks.im> writes:\n>> > diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n>> > index ccbb622eae..12c80f9b34 100644\n>> > --- a/odb/source-inmemory.c\n>> > +++ b/odb/source-inmemory.c\n>> > @@ -1,5 +1,57 @@\n>> >  #include \"git-compat-util.h\"\n>> > +#include \"odb.h\"\n>> >  #include \"odb/source-inmemory.h\"\n>> > +#include \"repository.h\"\n>> > +\n>> > +static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n>> > +\t\t\t\t\t\t      const struct object_id *oid)\n>> > +{\n>> > +\tstatic const struct cached_object empty_tree = {\n>> > +\t\t.type = OBJ_TREE,\n>> > +\t\t.buf = \"\",\n>> > +\t};\n>> > +\tconst struct cached_object_entry *co = source->objects;\n>> > +\n>> > +\tfor (size_t i = 0; i < source->objects_nr; i++, co++)\n>> > +\t\tif (oideq(&co->oid, oid))\n>> > +\t\t\treturn &co->value;\n>> > +\n>> > +\tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n>> > +\t\treturn &empty_tree;\n>> > +\n>>\n>> Silly questiong, would it make more sense to check for empty_tree before\n>> iterating over all objects?\n>>\n>> The rest looks good\n>\n> Maybe? I guess for now reading the empty tree is the most important use\n> case we have for the in-memory backend, as we only write in-memory\n> objects in a single caller. On the other hand, `source->objects_nr`\n> would be zero in all the other cases, and jumping over the loop should\n> be fast enough to not matter in practice.\n>\n\nThat was what I understood, okay so it's fine as is.\n\n> An alternative I was thinking about is to store the empty tree the same\n> way as we store all the other objects so that we don't have to special\n> case anything. That has the benefit that we can actually modify the tree\n> object, too, which may eventually become relevant with regards to an\n> object's mtime that we may want to update. The downside is that we have\n> another allocation here and need to eagerly initialize the data\n> structure that stores the objects.\n>\n> Patrick\n\nThat would be good too, but I also think maybe it is fine to just leave\nit. This is simple enough.\n"},{"id":"541250","messageId":"CAOLa=ZREniG1jkqk4SW6W1s6hLHh42fLQK+8tox59jprn2hPPg@mail.gmail.com","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-9-f02b4f1c0f13@pks.im","subject":"Re: [PATCH v2 09/17] cbtree: allow using arbitrary wrapper structures for nodes","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-04-09T11:36:50Z","receivedAt":"2026-04-09T11:36:55Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n[snip]\n\n> diff --git a/cbtree.h b/cbtree.h\n> index c374b1b3db..3ce0d6b287 100644\n> --- a/cbtree.h\n> +++ b/cbtree.h\n> @@ -23,18 +23,19 @@ struct cb_node {\n>  \t */\n>  \tuint32_t byte;\n>  \tuint8_t otherbits;\n> -\tuint8_t k[FLEX_ARRAY]; /* arbitrary data, unaligned */\n>  };\n>\n\nSeems like we need to update the comments at the top of the header file\nwhich still talks about this field.\n\n[snip]\n"},{"id":"541251","messageId":"CAOLa=ZTOOTqNv7j-DFdC2cMje=5MBdNBkqp9PhgUC7dQFKLz3A@mail.gmail.com","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im","subject":"Re: [PATCH v2 00/17] odb: introduce \"in-memory\" source","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-04-09T11:44:17Z","receivedAt":"2026-04-09T11:44:19Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Hi,\n>\n> this patch series introduces the second object database source type,\n> which is the \"in-memory\" source.\n>\n> This source may seem somewhat odd at first: it always starts out empty,\n> and any object written into it will only exist in memory until the\n> process exits. But the source already serves a purpose in our codebase,\n> where some commands, for example git-blame(1), write an in-memory\n> worktree commit.\n>\n> Furthermore, I think that going forward it can serve more purposes as we\n> now have an easy way to write and read objects that will not get\n> persisted. I could see that this may be useful when for example\n> re-merging diffs. But eventually, once we have the object storage format\n> extension wired up, callers might even want to manually set up an\n> in-memory database as the primary ODB for write operations so that no\n> data will be persisted in an arbitrary write.\n>\n> Last but not least, this patch series also serves the purpose of\n> eventually getting rid of the `struct object_info::whence` member.\n> Instead, we'll simply yield the ODB source a specific object has been\n> read from, together with some backend-specific data, which gives\n> strictly more information compared to the status quo.\n>\n> The series is based onb15384c06f (A bit more post -rc1, 2026-04-08)\n> with jt/odb-transaction-write at ddf6aee9c6 (odb/transaction: make\n> `write_object_stream()` pluggable, 2026-04-02) merged into it.\n>\n\nWas a nice read, only a few comments from me. Should be good with a\nre-roll!\n\n[snip]\n"},{"id":"541252","messageId":"adeRjBmWsRX1rLDi@pks.im","threadId":"65423","inReplyTo":"CAOLa=ZREniG1jkqk4SW6W1s6hLHh42fLQK+8tox59jprn2hPPg@mail.gmail.com","subject":"Re: [PATCH v2 09/17] cbtree: allow using arbitrary wrapper structures for nodes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T11:46:20Z","receivedAt":"2026-04-09T11:46:25Z","isPatch":true,"body":"On Thu, Apr 09, 2026 at 07:36:50AM -0400, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> [snip]\n> \n> > diff --git a/cbtree.h b/cbtree.h\n> > index c374b1b3db..3ce0d6b287 100644\n> > --- a/cbtree.h\n> > +++ b/cbtree.h\n> > @@ -23,18 +23,19 @@ struct cb_node {\n> >  \t */\n> >  \tuint32_t byte;\n> >  \tuint8_t otherbits;\n> > -\tuint8_t k[FLEX_ARRAY]; /* arbitrary data, unaligned */\n> >  };\n> >\n> \n> Seems like we need to update the comments at the top of the header file\n> which still talks about this field.\n\nGood eyes, will adapt.\n\nPatrick\n"},{"id":"541253","messageId":"adeR-VSCkH8BrFGE@pks.im","threadId":"65423","inReplyTo":"CAOLa=ZTOOTqNv7j-DFdC2cMje=5MBdNBkqp9PhgUC7dQFKLz3A@mail.gmail.com","subject":"Re: [PATCH v2 00/17] odb: introduce \"in-memory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-09T11:48:09Z","receivedAt":"2026-04-09T11:48:14Z","isPatch":true,"body":"On Thu, Apr 09, 2026 at 07:44:17AM -0400, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > Hi,\n> >\n> > this patch series introduces the second object database source type,\n> > which is the \"in-memory\" source.\n> >\n> > This source may seem somewhat odd at first: it always starts out empty,\n> > and any object written into it will only exist in memory until the\n> > process exits. But the source already serves a purpose in our codebase,\n> > where some commands, for example git-blame(1), write an in-memory\n> > worktree commit.\n> >\n> > Furthermore, I think that going forward it can serve more purposes as we\n> > now have an easy way to write and read objects that will not get\n> > persisted. I could see that this may be useful when for example\n> > re-merging diffs. But eventually, once we have the object storage format\n> > extension wired up, callers might even want to manually set up an\n> > in-memory database as the primary ODB for write operations so that no\n> > data will be persisted in an arbitrary write.\n> >\n> > Last but not least, this patch series also serves the purpose of\n> > eventually getting rid of the `struct object_info::whence` member.\n> > Instead, we'll simply yield the ODB source a specific object has been\n> > read from, together with some backend-specific data, which gives\n> > strictly more information compared to the status quo.\n> >\n> > The series is based onb15384c06f (A bit more post -rc1, 2026-04-08)\n> > with jt/odb-transaction-write at ddf6aee9c6 (odb/transaction: make\n> > `write_object_stream()` pluggable, 2026-04-02) merged into it.\n> >\n> \n> Was a nice read, only a few comments from me. Should be good with a\n> re-roll!\n\nThanks! Will send the new version tomorrow to wait for some more\nfeedback.\n\nPatrick\n"},{"id":"541270","messageId":"xmqqjyugw8jt.fsf@gitster.g","threadId":"65423","inReplyTo":"adc3mAItBiKMUFNJ@pks.im","subject":"Re: [PATCH 00/16] odb: introduce \"inmemory\" source","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-09T13:46:30Z","receivedAt":"2026-04-09T13:46:36Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>> But stepping back a bit, does this new \"in memory\" refer to a\n>> concept that is different from what the rest of the system uses \"in\n>> core\" to represent?\n>\n> No, in principle it's not any different. One of the reasons I decided to\n> go with \"in memory\" though is that this backend may eventually be\n> (power-)user-facing via the planned \"objectStorage\" extension.\n\nDoesn't \n\n    git grep -E -e 'in[- ]?core' -- ':!Documentation/RelNotes' ':!t'\n\ngive many hits that we want to be in line with in the codebase\nanyway, and even in some user-facing things?  I just noticed an\noption \"--no-kept-objects=in-core\" (which I didn't know about ;-).\n"},{"id":"541291","messageId":"xmqqzf3bsz3f.fsf@gitster.g","threadId":"65423","inReplyTo":"20260409-b4-pks-odb-source-inmemory-v2-16-f02b4f1c0f13@pks.im","subject":"Re: [PATCH v2 16/17] odb/source-inmemory: stub out remaining functions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-09T19:39:00Z","receivedAt":"2026-04-09T19:39:03Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Stub out remaining functions that we either don't need or that are\n> basically no-ops.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  odb/source-inmemory.c | 31 +++++++++++++++++++++++++++++++\n>  1 file changed, 31 insertions(+)\n>\n> diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> index 15a6a5ae64..1140b1b916 100644\n> --- a/odb/source-inmemory.c\n> +++ b/odb/source-inmemory.c\n> @@ -299,6 +299,32 @@ static int odb_source_inmemory_freshen_object(struct odb_source *source,\n>  \treturn 0;\n>  }\n>  \n> +static int odb_source_inmemory_begin_transaction(struct odb_source *source UNUSED,\n> +\t\t\t\t\t\t struct odb_transaction **out UNUSED)\n> +{\n> +\treturn error(\"inmemory source does not support transactions\");\n> +}\n> +\n> +static int odb_source_inmemory_read_alternates(struct odb_source *source UNUSED,\n> +\t\t\t\t\t       struct strvec *out UNUSED)\n> +{\n> +\treturn 0;\n> +}\n> +\n> +static int odb_source_inmemory_write_alternate(struct odb_source *source UNUSED,\n> +\t\t\t\t\t       const char *alternate UNUSED)\n> +{\n> +\treturn error(\"inmemory source does not support alternates\");\n> +}\n\nOK, 00/17 said it only fixed log message, but the messages or\nanything end-user facing should consistently say \"in-memory\".\n\nOr, \"in-core\", if \"incore\" is chosen as part of identifiers to be\nconsistent with the rest of the system.\n"},{"id":"541324","messageId":"adiCR-rx3OwYQP9H@pks.im","threadId":"65423","inReplyTo":"xmqqjyugw8jt.fsf@gitster.g","subject":"Re: [PATCH 00/16] odb: introduce \"inmemory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T04:53:27Z","receivedAt":"2026-04-10T04:53:32Z","isPatch":true,"body":"On Thu, Apr 09, 2026 at 06:46:30AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> >> But stepping back a bit, does this new \"in memory\" refer to a\n> >> concept that is different from what the rest of the system uses \"in\n> >> core\" to represent?\n> >\n> > No, in principle it's not any different. One of the reasons I decided to\n> > go with \"in memory\" though is that this backend may eventually be\n> > (power-)user-facing via the planned \"objectStorage\" extension.\n> \n> Doesn't \n> \n>     git grep -E -e 'in[- ]?core' -- ':!Documentation/RelNotes' ':!t'\n> \n> give many hits that we want to be in line with in the codebase\n> anyway, and even in some user-facing things?  I just noticed an\n> option \"--no-kept-objects=in-core\" (which I didn't know about ;-).\n\nMost of the hits are in our code though, and end users wouldn't\ntypically see those. So what I think is more relevant is documentation\nor options like the one you pointed out. But \"--no-kept-objects=\" is not\neven documented, which basically leaves us with the following hits:\n\n  Documentation/gitformat-pack.adoc:write a cruft pack. Crucially, the set of in-core kept packs is exactly the set\n  Documentation/technical/parallel-checkout.adoc:parallelize the work of uncompressing the blobs, applying in-core\n  Documentation/technical/racy-git.adoc:because in-core timestamps can have finer granularity than\n  Documentation/technical/racy-git.adoc:([PATCH] Sync in core time granularity with filesystems,\n\nI think that these hits are all related to what we're doing here, as\nwe're talking about object data that we handle in-core.\n\nI initially said \"it's not any different\", but thinking a bit more about\nit I think there is a slight difference: in-core could be any object\nthat we have parsed from the object database, even if it's backed by an\nactual on-disk object. In-memory objects may not even have been parsed\nat all, so technically speaking they may not even be in-core.\n\nSo in summary:\n\n  - I think that end users have not really been exposed to the concept\n    of \"in-core\".\n\n  - The concepts of \"in-core\" and the ODB source here are slightly\n    different, as any object is treated as \"in-core\" that has been\n    parsed.\n\n  - The concept of \"in-memory\" is easier for the end user to understand\n    in the context of the ODB, as it's a more general concept compared\n    to the very Git-specific \"in-core\" term\".\n\nHope that makes sense :)\n\nThanks!\n\nPatrick\n"},{"id":"541325","messageId":"adiCTHTDzmnvMOIt@pks.im","threadId":"65423","inReplyTo":"xmqqzf3bsz3f.fsf@gitster.g","subject":"Re: [PATCH v2 16/17] odb/source-inmemory: stub out remaining functions","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T04:53:32Z","receivedAt":"2026-04-10T04:53:36Z","isPatch":true,"body":"On Thu, Apr 09, 2026 at 12:39:00PM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > Stub out remaining functions that we either don't need or that are\n> > basically no-ops.\n> >\n> > Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> > ---\n> >  odb/source-inmemory.c | 31 +++++++++++++++++++++++++++++++\n> >  1 file changed, 31 insertions(+)\n> >\n> > diff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\n> > index 15a6a5ae64..1140b1b916 100644\n> > --- a/odb/source-inmemory.c\n> > +++ b/odb/source-inmemory.c\n> > @@ -299,6 +299,32 @@ static int odb_source_inmemory_freshen_object(struct odb_source *source,\n> >  \treturn 0;\n> >  }\n> >  \n> > +static int odb_source_inmemory_begin_transaction(struct odb_source *source UNUSED,\n> > +\t\t\t\t\t\t struct odb_transaction **out UNUSED)\n> > +{\n> > +\treturn error(\"inmemory source does not support transactions\");\n> > +}\n> > +\n> > +static int odb_source_inmemory_read_alternates(struct odb_source *source UNUSED,\n> > +\t\t\t\t\t       struct strvec *out UNUSED)\n> > +{\n> > +\treturn 0;\n> > +}\n> > +\n> > +static int odb_source_inmemory_write_alternate(struct odb_source *source UNUSED,\n> > +\t\t\t\t\t       const char *alternate UNUSED)\n> > +{\n> > +\treturn error(\"inmemory source does not support alternates\");\n> > +}\n> \n> OK, 00/17 said it only fixed log message, but the messages or\n> anything end-user facing should consistently say \"in-memory\".\n> \n> Or, \"in-core\", if \"incore\" is chosen as part of identifiers to be\n> consistent with the rest of the system.\n\nGood catch, will fix. Thanks!\n\nPatrick\n"},{"id":"541346","messageId":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im","subject":"[PATCH v3 00/17] odb: introduce \"in-memory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:30Z","receivedAt":"2026-04-10T12:12:50Z","isPatch":true,"body":"Hi,\n\nthis patch series introduces the second object database source type,\nwhich is the \"in-memory\" source.\n\nThis source may seem somewhat odd at first: it always starts out empty,\nand any object written into it will only exist in memory until the\nprocess exits. But the source already serves a purpose in our codebase,\nwhere some commands, for example git-blame(1), write an in-memory\nworktree commit.\n\nFurthermore, I think that going forward it can serve more purposes as we\nnow have an easy way to write and read objects that will not get\npersisted. I could see that this may be useful when for example\nre-merging diffs. But eventually, once we have the object storage format\nextension wired up, callers might even want to manually set up an\nin-memory database as the primary ODB for write operations so that no\ndata will be persisted in an arbitrary write.\n\nLast but not least, this patch series also serves the purpose of\neventually getting rid of the `struct object_info::whence` member.\nInstead, we'll simply yield the ODB source a specific object has been\nread from, together with some backend-specific data, which gives\nstrictly more information compared to the status quo.\n\nThe series is based onb15384c06f (A bit more post -rc1, 2026-04-08)\nwith jt/odb-transaction-write at ddf6aee9c6 (odb/transaction: make\n`write_object_stream()` pluggable, 2026-04-02) merged into it.\n\nChanges in v2:\n  - Fix handling of object IDs when writing objects.\n  - I've changed the base of this series to include Justin's\n    refactorings for the ODB write streams. I've updated the above\n    paragraph detailing the merge base accordingly. @Junio: I'm fine to\n    defer this patch series a bit until Justin's patch series has been\n    merged to `next` in case this causes inconvenience.\n  - Use \"in-memory\" instead of \"inmemory\" in commit messages.\n  - Link to v1: https://patch.msgid.link/20260403-b4-pks-odb-source-inmemory-v1-0-8b8d1abaa25e@pks.im\n\nChanges in v3:\n  - Fix a couple more instances where we were saying \"inmemory\" in\n    prose.\n  - Fix streaming interface when reading an object.\n  - Add unit tests to exercise full functionality of the new source.\n    Some of the functionality isn't exercised in our code base yet, so\n    this allows us to verify that things work as expected.\n  - Link to v2: https://patch.msgid.link/20260409-b4-pks-odb-source-inmemory-v2-0-f02b4f1c0f13@pks.im\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (17):\n      odb: introduce \"in-memory\" source\n      odb/source-inmemory: implement `free()` callback\n      odb: fix unnecessary call to `find_cached_object()`\n      odb/source-inmemory: implement `read_object_info()` callback\n      odb/source-inmemory: implement `read_object_stream()` callback\n      odb/source-inmemory: implement `write_object()` callback\n      odb/source-inmemory: implement `write_object_stream()` callback\n      cbtree: allow using arbitrary wrapper structures for nodes\n      oidtree: add ability to store data\n      odb/source-inmemory: convert to use oidtree\n      odb/source-inmemory: implement `for_each_object()` callback\n      odb/source-inmemory: implement `find_abbrev_len()` callback\n      odb/source-inmemory: implement `count_objects()` callback\n      odb/source-inmemory: implement `freshen_object()` callback\n      odb/source-inmemory: stub out remaining functions\n      odb: generic in-memory source\n      t/unit-tests: add tests for the in-memory object source\n\n Makefile                      |   2 +\n cbtree.c                      |  25 ++-\n cbtree.h                      |  17 +-\n loose.c                       |   2 +-\n meson.build                   |   1 +\n object-file.c                 |   3 +-\n odb.c                         |  82 ++-------\n odb.h                         |   4 +-\n odb/source-inmemory.c         | 382 ++++++++++++++++++++++++++++++++++++++++++\n odb/source-inmemory.h         |  33 ++++\n odb/source.h                  |   3 +\n oidtree.c                     |  66 +++++---\n oidtree.h                     |  12 +-\n t/meson.build                 |   1 +\n t/unit-tests/u-odb-inmemory.c | 313 ++++++++++++++++++++++++++++++++++\n t/unit-tests/u-oidtree.c      |  26 ++-\n 16 files changed, 854 insertions(+), 118 deletions(-)\n\nRange-diff versus v2:\n\n 1:  b18e427c69 !  1:  155b2cdf81 odb: introduce \"in-memory\" source\n    @@ odb/source-inmemory.h (new)\n     +struct cached_object_entry;\n     +\n     +/*\n    -+ * An inmemory source that you can write objects to that shall be made\n    ++ * An in-memory source that you can write objects to that shall be made\n     + * available for reading, but that shouldn't ever be persisted to disk. Note\n     + * that any objects written to this source will be stored in memory, so the\n     + * number of objects you can store is limited by available system memory.\n    @@ odb/source-inmemory.h (new)\n     +struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb);\n     +\n     +/*\n    -+ * Cast the given object database source to the inmemory backend. This will\n    ++ * Cast the given object database source to the in-memory backend. This will\n     + * cause a BUG in case the source doesn't use this backend.\n     + */\n     +static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n     +{\n     +\tif (source->type != ODB_SOURCE_INMEMORY)\n    -+\t\tBUG(\"trying to downcast source of type '%d' to inmemory\", source->type);\n    ++\t\tBUG(\"trying to downcast source of type '%d' to in-memory\", source->type);\n     +\treturn container_of(source, struct odb_source_inmemory, base);\n     +}\n     +\n    @@ odb/source.h: enum odb_source_type {\n      \t/* The \"files\" backend that uses loose objects and packfiles. */\n      \tODB_SOURCE_FILES,\n     +\n    -+\t/* The \"inmemory\" backend that stores objects in memory. */\n    ++\t/* The \"in-memory\" backend that stores objects in memory. */\n     +\tODB_SOURCE_INMEMORY,\n      };\n      \n 2:  8fd337da90 !  2:  c66edd10a8 odb/source-inmemory: implement `free()` callback\n    @@ odb/source-inmemory.h\n     +};\n      \n      /*\n    -  * An inmemory source that you can write objects to that shall be made\n    +  * An in-memory source that you can write objects to that shall be made\n 3:  f4ae2a2bde =  3:  a86549f39c odb: fix unnecessary call to `find_cached_object()`\n 4:  8600b88530 =  4:  49ac739dd2 odb/source-inmemory: implement `read_object_info()` callback\n 5:  ab33c0b7ee !  5:  321ef11be3 odb/source-inmemory: implement `read_object_stream()` callback\n    @@ odb/source-inmemory.c: static int odb_source_inmemory_read_object_info(struct od\n      \n     +struct odb_read_stream_inmemory {\n     +\tstruct odb_read_stream base;\n    -+\tconst void *buf;\n    ++\tconst unsigned char *buf;\n     +\tsize_t offset;\n     +};\n     +\n    @@ odb/source-inmemory.c: static int odb_source_inmemory_read_object_info(struct od\n     +\n     +\tif (buf_len > inmemory->base.size - inmemory->offset)\n     +\t\tbytes = inmemory->base.size - inmemory->offset;\n    -+\tmemcpy(buf, inmemory->buf, bytes);\n    ++\n    ++\tmemcpy(buf, inmemory->buf + inmemory->offset, bytes);\n    ++\tinmemory->offset += bytes;\n     +\n     +\treturn bytes;\n     +}\n 6:  983f886eeb !  6:  506df5e488 odb/source-inmemory: implement `write_object()` callback\n    @@ odb.c: int odb_pretend_object(struct object_database *odb,\n      void *odb_read_object(struct object_database *odb,\n     \n      ## odb/source-inmemory.c ##\n    +@@\n    + #include \"git-compat-util.h\"\n    ++#include \"object-file.h\"\n    + #include \"odb.h\"\n    + #include \"odb/source-inmemory.h\"\n    + #include \"odb/streaming.h\"\n     @@ odb/source-inmemory.c: static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n      \treturn 0;\n      }\n    @@ odb/source-inmemory.c: static int odb_source_inmemory_read_object_stream(struct\n     +\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n     +\tstruct cached_object_entry *object;\n     +\n    ++\thash_object_file(source->odb->repo->hash_algo, buf, len, type, oid);\n    ++\n     +\tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n     +\t\t   inmemory->objects_alloc);\n     +\tobject = &inmemory->objects[inmemory->objects_nr++];\n 7:  68edefa269 <  -:  ---------- odb/source-inmemory: implement `write_object()` callback\n 8:  18d451152b !  7:  21eef34c1b odb/source-inmemory: implement `write_object_stream()` callback\n    @@ odb/source-inmemory.c: static int odb_source_inmemory_write_object(struct odb_so\n     +\t\t\tgoto out;\n     +\t\t}\n     +\n    -+\t\tmemcpy(data, buf, bytes_read);\n    ++\t\tmemcpy(data + total_read, buf, bytes_read);\n     +\t\ttotal_read += bytes_read;\n     +\t}\n     +\n 9:  cee53b9853 !  8:  504e34d116 cbtree: allow using arbitrary wrapper structures for nodes\n    @@ cbtree.c: int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n      \n     \n      ## cbtree.h ##\n    +@@\n    +  *\n    +  * This is adapted to store arbitrary data (not just NUL-terminated C strings\n    +  * and allocates no memory internally.  The user needs to allocate\n    +- * \"struct cb_node\" and fill cb_node.k[] with arbitrary match data\n    +- * for memcmp.\n    +- * If \"klen\" is variable, then it should be embedded into \"c_node.k[]\"\n    ++ * \"struct cb_node\" and provide `key_offset` to indicate where the key can be\n    ++ * found relative to the `struct cb_node` for memcmp.\n    ++ * If \"klen\" is variable, then it should be embedded into the key.\n    +  * Recursion is bound by the maximum value of \"klen\" used.\n    +  */\n    + #ifndef CBTREE_H\n     @@ cbtree.h: struct cb_node {\n      \t */\n      \tuint32_t byte;\n10:  8ad5b81b13 =  9:  9bdd475a92 oidtree: add ability to store data\n11:  1ed2d23137 ! 10:  956b989529 odb/source-inmemory: convert to use oidtree\n    @@ odb/source-inmemory.h\n     +struct oidtree;\n      \n      /*\n    -  * An inmemory source that you can write objects to that shall be made\n    +  * An in-memory source that you can write objects to that shall be made\n     @@ odb/source-inmemory.h: struct cached_object_entry {\n       */\n      struct odb_source_inmemory {\n12:  99fbb1cc35 ! 11:  bec1428116 odb/source-inmemory: implement `for_each_object()` callback\n    @@ odb/source-inmemory.c: static int odb_source_inmemory_read_object_stream(struct\n     +\tif ((opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY) ||\n     +\t    (opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY && !source->local))\n     +\t\treturn 0;\n    ++\tif (!inmemory->objects)\n    ++\t\treturn 0;\n     +\n     +\treturn oidtree_each(inmemory->objects,\n     +\t\t\t    opts->prefix ? opts->prefix : &null_oid, opts->prefix_hex_len,\n13:  c87a621f39 = 12:  32dada3c27 odb/source-inmemory: implement `find_abbrev_len()` callback\n14:  9b88f0c07b = 13:  43127840c0 odb/source-inmemory: implement `count_objects()` callback\n15:  3c9493f2bb = 14:  439acbd068 odb/source-inmemory: implement `freshen_object()` callback\n16:  f2b6317104 ! 15:  12c1b6ffd2 odb/source-inmemory: stub out remaining functions\n    @@ odb/source-inmemory.c: static int odb_source_inmemory_freshen_object(struct odb_\n     +static int odb_source_inmemory_begin_transaction(struct odb_source *source UNUSED,\n     +\t\t\t\t\t\t struct odb_transaction **out UNUSED)\n     +{\n    -+\treturn error(\"inmemory source does not support transactions\");\n    ++\treturn error(\"in-memory source does not support transactions\");\n     +}\n     +\n     +static int odb_source_inmemory_read_alternates(struct odb_source *source UNUSED,\n    @@ odb/source-inmemory.c: static int odb_source_inmemory_freshen_object(struct odb_\n     +static int odb_source_inmemory_write_alternate(struct odb_source *source UNUSED,\n     +\t\t\t\t\t       const char *alternate UNUSED)\n     +{\n    -+\treturn error(\"inmemory source does not support alternates\");\n    ++\treturn error(\"in-memory source does not support alternates\");\n     +}\n     +\n     +static void odb_source_inmemory_close(struct odb_source *source UNUSED)\n17:  81da5d5048 = 16:  ef37a61e7f odb: generic in-memory source\n -:  ---------- > 17:  51b51e0382 t/unit-tests: add tests for the in-memory object source\n\n---\nbase-commit: a3ebc5a08e67ccac4c915622049a968a31e48662\nchange-id: 20260401-b4-pks-odb-source-inmemory-7b17c83d9e43\n\n"},{"id":"541347","messageId":"20260410-b4-pks-odb-source-inmemory-v3-1-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 01/17] odb: introduce \"in-memory\" source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:31Z","receivedAt":"2026-04-10T12:12:51Z","isPatch":true,"body":"Next to our typical object database sources, each object database also\nhas an implicit source of \"cached\" objects. These cached objects only\nexist in memory and some use cases:\n\n  - They contain evergreen objects that we expect to always exist, like\n    for example the empty tree.\n\n  - They can be used to store temporary objects that we don't want to\n    persist to disk, which is used by git-blame(1) to create a fake\n    worktree commit.\n\nOverall, their use is somewhat restricted though. For example, we don't\nprovide the ability to use it as a temporary object database source that\nallows the user to write objects, but discard them after Git exists. So\nwhile these cached objects behave almost like a source, they aren't used\nas one.\n\nThis is about to change over the following commits, where we will turn\ncached objects into a new \"in-memory\" source. This will allow us to use\nit exactly the same as any other source by providing the same common\ninterface as the \"files\" source.\n\nFor now, the in-memory source only hosts the cached objects and doesn't\nprovide any logic yet. This will change with subsequent commits, where\nwe move respective functionality into the source.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n Makefile              |  1 +\n meson.build           |  1 +\n odb.c                 | 21 +++++++++++++--------\n odb.h                 |  4 ++--\n odb/source-inmemory.c | 12 ++++++++++++\n odb/source-inmemory.h | 35 +++++++++++++++++++++++++++++++++++\n odb/source.h          |  3 +++\n 7 files changed, 67 insertions(+), 10 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 22a8993482..3cda12c455 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1218,6 +1218,7 @@ LIB_OBJS += object.o\n LIB_OBJS += odb.o\n LIB_OBJS += odb/source.o\n LIB_OBJS += odb/source-files.o\n+LIB_OBJS += odb/source-inmemory.o\n LIB_OBJS += odb/streaming.o\n LIB_OBJS += odb/transaction.o\n LIB_OBJS += oid-array.o\ndiff --git a/meson.build b/meson.build\nindex 6dc23b3af2..ffa73ce7ce 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -404,6 +404,7 @@ libgit_sources = [\n   'odb.c',\n   'odb/source.c',\n   'odb/source-files.c',\n+  'odb/source-inmemory.c',\n   'odb/streaming.c',\n   'odb/transaction.c',\n   'oid-array.c',\ndiff --git a/odb.c b/odb.c\nindex 40a5e9c4e0..60e1eead25 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -14,6 +14,7 @@\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/source-inmemory.h\"\n #include \"packfile.h\"\n #include \"path.h\"\n #include \"promisor-remote.h\"\n@@ -53,9 +54,9 @@ static const struct cached_object *find_cached_object(struct object_database *ob\n \t\t.type = OBJ_TREE,\n \t\t.buf = \"\",\n \t};\n-\tconst struct cached_object_entry *co = object_store->cached_objects;\n+\tconst struct cached_object_entry *co = object_store->inmemory_objects->objects;\n \n-\tfor (size_t i = 0; i < object_store->cached_object_nr; i++, co++)\n+\tfor (size_t i = 0; i < object_store->inmemory_objects->objects_nr; i++, co++)\n \t\tif (oideq(&co->oid, oid))\n \t\t\treturn &co->value;\n \n@@ -792,9 +793,10 @@ int odb_pretend_object(struct object_database *odb,\n \t    find_cached_object(odb, oid))\n \t\treturn 0;\n \n-\tALLOC_GROW(odb->cached_objects,\n-\t\t   odb->cached_object_nr + 1, odb->cached_object_alloc);\n-\tco = &odb->cached_objects[odb->cached_object_nr++];\n+\tALLOC_GROW(odb->inmemory_objects->objects,\n+\t\t   odb->inmemory_objects->objects_nr + 1,\n+\t\t   odb->inmemory_objects->objects_alloc);\n+\tco = &odb->inmemory_objects->objects[odb->inmemory_objects->objects_nr++];\n \tco->value.size = len;\n \tco->value.type = type;\n \tco_buf = xmalloc(len);\n@@ -1083,6 +1085,7 @@ struct object_database *odb_new(struct repository *repo,\n \to->sources = odb_source_new(o, primary_source, true);\n \to->sources_tail = &o->sources->next;\n \to->alternate_db = xstrdup_or_null(secondary_sources);\n+\to->inmemory_objects = odb_source_inmemory_new(o);\n \n \tfree(to_free);\n \n@@ -1123,9 +1126,11 @@ void odb_free(struct object_database *o)\n \todb_close(o);\n \todb_free_sources(o);\n \n-\tfor (size_t i = 0; i < o->cached_object_nr; i++)\n-\t\tfree((char *) o->cached_objects[i].value.buf);\n-\tfree(o->cached_objects);\n+\tfor (size_t i = 0; i < o->inmemory_objects->objects_nr; i++)\n+\t\tfree((char *) o->inmemory_objects->objects[i].value.buf);\n+\tfree(o->inmemory_objects->objects);\n+\tfree(o->inmemory_objects->base.path);\n+\tfree(o->inmemory_objects);\n \n \tstring_list_clear(&o->submodule_source_paths, 0);\n \ndiff --git a/odb.h b/odb.h\nindex 9eb8355aca..c3a7edf9c8 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -8,6 +8,7 @@\n #include \"thread-utils.h\"\n \n struct cached_object_entry;\n+struct odb_source_inmemory;\n struct packed_git;\n struct repository;\n struct strbuf;\n@@ -80,8 +81,7 @@ struct object_database {\n \t * to write them into the object store (e.g. a browse-only\n \t * application).\n \t */\n-\tstruct cached_object_entry *cached_objects;\n-\tsize_t cached_object_nr, cached_object_alloc;\n+\tstruct odb_source_inmemory *inmemory_objects;\n \n \t/*\n \t * A fast, rough count of the number of objects in the repository.\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nnew file mode 100644\nindex 0000000000..c7ac5c24f0\n--- /dev/null\n+++ b/odb/source-inmemory.c\n@@ -0,0 +1,12 @@\n+#include \"git-compat-util.h\"\n+#include \"odb/source-inmemory.h\"\n+\n+struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n+{\n+\tstruct odb_source_inmemory *source;\n+\n+\tCALLOC_ARRAY(source, 1);\n+\todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n+\n+\treturn source;\n+}\ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nnew file mode 100644\nindex 0000000000..15db068ef7\n--- /dev/null\n+++ b/odb/source-inmemory.h\n@@ -0,0 +1,35 @@\n+#ifndef ODB_SOURCE_INMEMORY_H\n+#define ODB_SOURCE_INMEMORY_H\n+\n+#include \"odb/source.h\"\n+\n+struct cached_object_entry;\n+\n+/*\n+ * An in-memory source that you can write objects to that shall be made\n+ * available for reading, but that shouldn't ever be persisted to disk. Note\n+ * that any objects written to this source will be stored in memory, so the\n+ * number of objects you can store is limited by available system memory.\n+ */\n+struct odb_source_inmemory {\n+\tstruct odb_source base;\n+\n+\tstruct cached_object_entry *objects;\n+\tsize_t objects_nr, objects_alloc;\n+};\n+\n+/* Create a new in-memory object database source. */\n+struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb);\n+\n+/*\n+ * Cast the given object database source to the in-memory backend. This will\n+ * cause a BUG in case the source doesn't use this backend.\n+ */\n+static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n+{\n+\tif (source->type != ODB_SOURCE_INMEMORY)\n+\t\tBUG(\"trying to downcast source of type '%d' to in-memory\", source->type);\n+\treturn container_of(source, struct odb_source_inmemory, base);\n+}\n+\n+#endif\ndiff --git a/odb/source.h b/odb/source.h\nindex f706e0608a..0a440884e4 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -13,6 +13,9 @@ enum odb_source_type {\n \n \t/* The \"files\" backend that uses loose objects and packfiles. */\n \tODB_SOURCE_FILES,\n+\n+\t/* The \"in-memory\" backend that stores objects in memory. */\n+\tODB_SOURCE_INMEMORY,\n };\n \n struct object_id;\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541348","messageId":"20260410-b4-pks-odb-source-inmemory-v3-2-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 02/17] odb/source-inmemory: implement `free()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:32Z","receivedAt":"2026-04-10T12:12:54Z","isPatch":true,"body":"Implement the `free()` callback function for the \"in-memory\" source.\n\nNote that this requires us to define `struct cached_object_entry` in\n\"odb/source-inmemory.h\", as it is accessed in both \"odb.c\" and\n\"odb/source-inmemory.c\" now. This will be fixed in subsequent commits\nthough.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                 | 25 ++++---------------------\n odb/source-inmemory.c | 12 ++++++++++++\n odb/source-inmemory.h |  9 ++++++++-\n 3 files changed, 24 insertions(+), 22 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 60e1eead25..1d65825ed3 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -32,21 +32,6 @@\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct odb_source *, 1, fspathhash, fspatheq)\n \n-/*\n- * This is meant to hold a *small* number of objects that you would\n- * want odb_read_object() to be able to return, but yet you do not want\n- * to write them into the object store (e.g. a browse-only\n- * application).\n- */\n-struct cached_object_entry {\n-\tstruct object_id oid;\n-\tstruct cached_object {\n-\t\tenum object_type type;\n-\t\tconst void *buf;\n-\t\tunsigned long size;\n-\t} value;\n-};\n-\n static const struct cached_object *find_cached_object(struct object_database *object_store,\n \t\t\t\t\t\t      const struct object_id *oid)\n {\n@@ -1109,6 +1094,10 @@ static void odb_free_sources(struct object_database *o)\n \t\todb_source_free(o->sources);\n \t\to->sources = next;\n \t}\n+\n+\todb_source_free(&o->inmemory_objects->base);\n+\to->inmemory_objects = NULL;\n+\n \tkh_destroy_odb_path_map(o->source_by_path);\n \to->source_by_path = NULL;\n }\n@@ -1126,12 +1115,6 @@ void odb_free(struct object_database *o)\n \todb_close(o);\n \todb_free_sources(o);\n \n-\tfor (size_t i = 0; i < o->inmemory_objects->objects_nr; i++)\n-\t\tfree((char *) o->inmemory_objects->objects[i].value.buf);\n-\tfree(o->inmemory_objects->objects);\n-\tfree(o->inmemory_objects->base.path);\n-\tfree(o->inmemory_objects);\n-\n \tstring_list_clear(&o->submodule_source_paths, 0);\n \n \tfree(o);\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex c7ac5c24f0..ccbb622eae 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,6 +1,16 @@\n #include \"git-compat-util.h\"\n #include \"odb/source-inmemory.h\"\n \n+static void odb_source_inmemory_free(struct odb_source *source)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tfor (size_t i = 0; i < inmemory->objects_nr; i++)\n+\t\tfree((char *) inmemory->objects[i].value.buf);\n+\tfree(inmemory->objects);\n+\tfree(inmemory->base.path);\n+\tfree(inmemory);\n+}\n+\n struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n {\n \tstruct odb_source_inmemory *source;\n@@ -8,5 +18,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tCALLOC_ARRAY(source, 1);\n \todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n \n+\tsource->base.free = odb_source_inmemory_free;\n+\n \treturn source;\n }\ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex 15db068ef7..d1b05a3996 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -3,7 +3,14 @@\n \n #include \"odb/source.h\"\n \n-struct cached_object_entry;\n+struct cached_object_entry {\n+\tstruct object_id oid;\n+\tstruct cached_object {\n+\t\tenum object_type type;\n+\t\tconst void *buf;\n+\t\tunsigned long size;\n+\t} value;\n+};\n \n /*\n  * An in-memory source that you can write objects to that shall be made\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541349","messageId":"20260410-b4-pks-odb-source-inmemory-v3-3-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 03/17] odb: fix unnecessary call to `find_cached_object()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:33Z","receivedAt":"2026-04-10T12:12:56Z","isPatch":true,"body":"The function `odb_pretend_object()` writes an object into the in-memory\nobject database source. The effect of this is that the object will now\nbecome readable, but it won't ever be persisted to disk.\n\nBefore storing the object, we first verify whether the object already\nexists. This is done by calling `odb_has_object()` to check all sources,\nfollowed by `find_cached_object()` to check whether we have already\nstored the object in our in-memory source.\n\nThis is unnecessary though, as `odb_has_object()` already checks the\nin-memory source transitively via:\n\n  - `odb_has_object()`\n  - `odb_read_object_info_extended()`\n  - `do_oid_object_info_extended()`\n  - `find_cached_object()`\n\nDrop the explicit call to `find_cached_object()`.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c | 3 +--\n 1 file changed, 1 insertion(+), 2 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 1d65825ed3..ea3fcf5e11 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -774,8 +774,7 @@ int odb_pretend_object(struct object_database *odb,\n \tchar *co_buf;\n \n \thash_object_file(odb->repo->hash_algo, buf, len, type, oid);\n-\tif (odb_has_object(odb, oid, 0) ||\n-\t    find_cached_object(odb, oid))\n+\tif (odb_has_object(odb, oid, 0))\n \t\treturn 0;\n \n \tALLOC_GROW(odb->inmemory_objects->objects,\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541350","messageId":"20260410-b4-pks-odb-source-inmemory-v3-4-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 04/17] odb/source-inmemory: implement `read_object_info()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:34Z","receivedAt":"2026-04-10T12:12:59Z","isPatch":true,"body":"Implement the `read_object_info()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                 | 39 +------------------------------------\n odb/source-inmemory.c | 53 +++++++++++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 54 insertions(+), 38 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex ea3fcf5e11..6a3912adac 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -32,25 +32,6 @@\n KHASH_INIT(odb_path_map, const char * /* key: odb_path */,\n \tstruct odb_source *, 1, fspathhash, fspatheq)\n \n-static const struct cached_object *find_cached_object(struct object_database *object_store,\n-\t\t\t\t\t\t      const struct object_id *oid)\n-{\n-\tstatic const struct cached_object empty_tree = {\n-\t\t.type = OBJ_TREE,\n-\t\t.buf = \"\",\n-\t};\n-\tconst struct cached_object_entry *co = object_store->inmemory_objects->objects;\n-\n-\tfor (size_t i = 0; i < object_store->inmemory_objects->objects_nr; i++, co++)\n-\t\tif (oideq(&co->oid, oid))\n-\t\t\treturn &co->value;\n-\n-\tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n-\t\treturn &empty_tree;\n-\n-\treturn NULL;\n-}\n-\n int odb_mkstemp(struct object_database *odb,\n \t\tstruct strbuf *temp_filename, const char *pattern)\n {\n@@ -570,7 +551,6 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \t\t\t\t       const struct object_id *oid,\n \t\t\t\t       struct object_info *oi, unsigned flags)\n {\n-\tconst struct cached_object *co;\n \tconst struct object_id *real = oid;\n \tint already_retried = 0;\n \n@@ -580,25 +560,8 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \tif (is_null_oid(real))\n \t\treturn -1;\n \n-\tco = find_cached_object(odb, real);\n-\tif (co) {\n-\t\tif (oi) {\n-\t\t\tif (oi->typep)\n-\t\t\t\t*(oi->typep) = co->type;\n-\t\t\tif (oi->sizep)\n-\t\t\t\t*(oi->sizep) = co->size;\n-\t\t\tif (oi->disk_sizep)\n-\t\t\t\t*(oi->disk_sizep) = 0;\n-\t\t\tif (oi->delta_base_oid)\n-\t\t\t\toidclr(oi->delta_base_oid, odb->repo->hash_algo);\n-\t\t\tif (oi->contentp)\n-\t\t\t\t*oi->contentp = xmemdupz(co->buf, co->size);\n-\t\t\tif (oi->mtimep)\n-\t\t\t\t*oi->mtimep = 0;\n-\t\t\toi->whence = OI_CACHED;\n-\t\t}\n+\tif (!odb_source_read_object_info(&odb->inmemory_objects->base, oid, oi, flags))\n \t\treturn 0;\n-\t}\n \n \todb_prepare_alternates(odb);\n \ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex ccbb622eae..12c80f9b34 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,5 +1,57 @@\n #include \"git-compat-util.h\"\n+#include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n+#include \"repository.h\"\n+\n+static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n+\t\t\t\t\t\t      const struct object_id *oid)\n+{\n+\tstatic const struct cached_object empty_tree = {\n+\t\t.type = OBJ_TREE,\n+\t\t.buf = \"\",\n+\t};\n+\tconst struct cached_object_entry *co = source->objects;\n+\n+\tfor (size_t i = 0; i < source->objects_nr; i++, co++)\n+\t\tif (oideq(&co->oid, oid))\n+\t\t\treturn &co->value;\n+\n+\tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n+\t\treturn &empty_tree;\n+\n+\treturn NULL;\n+}\n+\n+static int odb_source_inmemory_read_object_info(struct odb_source *source,\n+\t\t\t\t\t\tconst struct object_id *oid,\n+\t\t\t\t\t\tstruct object_info *oi,\n+\t\t\t\t\t\tenum object_info_flags flags UNUSED)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tconst struct cached_object *object;\n+\n+\tobject = find_cached_object(inmemory, oid);\n+\tif (!object)\n+\t\treturn -1;\n+\n+\tif (oi) {\n+\t\tif (oi->typep)\n+\t\t\t*(oi->typep) = object->type;\n+\t\tif (oi->sizep)\n+\t\t\t*(oi->sizep) = object->size;\n+\t\tif (oi->disk_sizep)\n+\t\t\t*(oi->disk_sizep) = 0;\n+\t\tif (oi->delta_base_oid)\n+\t\t\toidclr(oi->delta_base_oid, source->odb->repo->hash_algo);\n+\t\tif (oi->contentp)\n+\t\t\t*oi->contentp = xmemdupz(object->buf, object->size);\n+\t\tif (oi->mtimep)\n+\t\t\t*oi->mtimep = 0;\n+\t\toi->whence = OI_CACHED;\n+\t}\n+\n+\treturn 0;\n+}\n \n static void odb_source_inmemory_free(struct odb_source *source)\n {\n@@ -19,6 +71,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n \n \tsource->base.free = odb_source_inmemory_free;\n+\tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541351","messageId":"20260410-b4-pks-odb-source-inmemory-v3-5-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 05/17] odb/source-inmemory: implement `read_object_stream()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:35Z","receivedAt":"2026-04-10T12:13:01Z","isPatch":true,"body":"Implement the `read_object_stream()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 52 +++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 52 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 12c80f9b34..39f0e799c7 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,6 +1,7 @@\n #include \"git-compat-util.h\"\n #include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n+#include \"odb/streaming.h\"\n #include \"repository.h\"\n \n static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n@@ -53,6 +54,56 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \treturn 0;\n }\n \n+struct odb_read_stream_inmemory {\n+\tstruct odb_read_stream base;\n+\tconst unsigned char *buf;\n+\tsize_t offset;\n+};\n+\n+static ssize_t odb_read_stream_inmemory_read(struct odb_read_stream *stream,\n+\t\t\t\t\t     char *buf, size_t buf_len)\n+{\n+\tstruct odb_read_stream_inmemory *inmemory =\n+\t\tcontainer_of(stream, struct odb_read_stream_inmemory, base);\n+\tsize_t bytes = buf_len;\n+\n+\tif (buf_len > inmemory->base.size - inmemory->offset)\n+\t\tbytes = inmemory->base.size - inmemory->offset;\n+\n+\tmemcpy(buf, inmemory->buf + inmemory->offset, bytes);\n+\tinmemory->offset += bytes;\n+\n+\treturn bytes;\n+}\n+\n+static int odb_read_stream_inmemory_close(struct odb_read_stream *stream UNUSED)\n+{\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n+\t\t\t\t\t\t  struct odb_source *source,\n+\t\t\t\t\t\t  const struct object_id *oid)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tstruct odb_read_stream_inmemory *stream;\n+\tconst struct cached_object *object;\n+\n+\tobject = find_cached_object(inmemory, oid);\n+\tif (!object)\n+\t\treturn -1;\n+\n+\tCALLOC_ARRAY(stream, 1);\n+\tstream->base.read = odb_read_stream_inmemory_read;\n+\tstream->base.close = odb_read_stream_inmemory_close;\n+\tstream->base.size = object->size;\n+\tstream->base.type = object->type;\n+\tstream->buf = object->buf;\n+\n+\t*out = &stream->base;\n+\treturn 0;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n@@ -72,6 +123,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \n \tsource->base.free = odb_source_inmemory_free;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n+\tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541352","messageId":"20260410-b4-pks-odb-source-inmemory-v3-6-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 06/17] odb/source-inmemory: implement `write_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:36Z","receivedAt":"2026-04-10T12:13:04Z","isPatch":true,"body":"Implement the `write_object()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                 | 16 ++--------------\n odb/source-inmemory.c | 25 +++++++++++++++++++++++++\n 2 files changed, 27 insertions(+), 14 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 6a3912adac..24e929f03c 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -733,24 +733,12 @@ int odb_pretend_object(struct object_database *odb,\n \t\t       void *buf, unsigned long len, enum object_type type,\n \t\t       struct object_id *oid)\n {\n-\tstruct cached_object_entry *co;\n-\tchar *co_buf;\n-\n \thash_object_file(odb->repo->hash_algo, buf, len, type, oid);\n \tif (odb_has_object(odb, oid, 0))\n \t\treturn 0;\n \n-\tALLOC_GROW(odb->inmemory_objects->objects,\n-\t\t   odb->inmemory_objects->objects_nr + 1,\n-\t\t   odb->inmemory_objects->objects_alloc);\n-\tco = &odb->inmemory_objects->objects[odb->inmemory_objects->objects_nr++];\n-\tco->value.size = len;\n-\tco->value.type = type;\n-\tco_buf = xmalloc(len);\n-\tmemcpy(co_buf, buf, len);\n-\tco->value.buf = co_buf;\n-\toidcpy(&co->oid, oid);\n-\treturn 0;\n+\treturn odb_source_write_object(&odb->inmemory_objects->base,\n+\t\t\t\t       buf, len, type, oid, NULL, 0);\n }\n \n void *odb_read_object(struct object_database *odb,\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 39f0e799c7..4848011df5 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -1,4 +1,5 @@\n #include \"git-compat-util.h\"\n+#include \"object-file.h\"\n #include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n #include \"odb/streaming.h\"\n@@ -104,6 +105,29 @@ static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n \treturn 0;\n }\n \n+static int odb_source_inmemory_write_object(struct odb_source *source,\n+\t\t\t\t\t    const void *buf, unsigned long len,\n+\t\t\t\t\t    enum object_type type,\n+\t\t\t\t\t    struct object_id *oid,\n+\t\t\t\t\t    struct object_id *compat_oid UNUSED,\n+\t\t\t\t\t    enum odb_write_object_flags flags UNUSED)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tstruct cached_object_entry *object;\n+\n+\thash_object_file(source->odb->repo->hash_algo, buf, len, type, oid);\n+\n+\tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n+\t\t   inmemory->objects_alloc);\n+\tobject = &inmemory->objects[inmemory->objects_nr++];\n+\tobject->value.size = len;\n+\tobject->value.type = type;\n+\tobject->value.buf = xmemdupz(buf, len);\n+\toidcpy(&object->oid, oid);\n+\n+\treturn 0;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n@@ -124,6 +148,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.free = odb_source_inmemory_free;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n+\tsource->base.write_object = odb_source_inmemory_write_object;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541353","messageId":"20260410-b4-pks-odb-source-inmemory-v3-7-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 07/17] odb/source-inmemory: implement `write_object_stream()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:37Z","receivedAt":"2026-04-10T12:13:07Z","isPatch":true,"body":"Implement the `write_object_stream()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 40 ++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 40 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 4848011df5..d05a13df45 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -128,6 +128,45 @@ static int odb_source_inmemory_write_object(struct odb_source *source,\n \treturn 0;\n }\n \n+static int odb_source_inmemory_write_object_stream(struct odb_source *source,\n+\t\t\t\t\t\t   struct odb_write_stream *stream,\n+\t\t\t\t\t\t   size_t len,\n+\t\t\t\t\t\t   struct object_id *oid)\n+{\n+\tchar buf[16384];\n+\tsize_t total_read = 0;\n+\tchar *data;\n+\tint ret;\n+\n+\tCALLOC_ARRAY(data, len);\n+\twhile (!stream->is_finished) {\n+\t\tssize_t bytes_read;\n+\n+\t\tbytes_read = odb_write_stream_read(stream, buf, sizeof(buf));\n+\t\tif (total_read + bytes_read > len) {\n+\t\t\tret = error(\"object stream yielded more bytes than expected\");\n+\t\t\tgoto out;\n+\t\t}\n+\n+\t\tmemcpy(data + total_read, buf, bytes_read);\n+\t\ttotal_read += bytes_read;\n+\t}\n+\n+\tif (total_read != len) {\n+\t\tret = error(\"object stream yielded less bytes than expected\");\n+\t\tgoto out;\n+\t}\n+\n+\tret = odb_source_inmemory_write_object(source, data, len, OBJ_BLOB, oid,\n+\t\t\t\t\t       NULL, 0);\n+\tif (ret < 0)\n+\t\tgoto out;\n+\n+out:\n+\tfree(data);\n+\treturn ret;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n@@ -149,6 +188,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n+\tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541354","messageId":"20260410-b4-pks-odb-source-inmemory-v3-8-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 08/17] cbtree: allow using arbitrary wrapper structures for nodes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:38Z","receivedAt":"2026-04-10T12:13:09Z","isPatch":true,"body":"The cbtree subsystem allows the user to store arbitrary data in a\nprefix-free set of strings. This is used by us to store object IDs in a\nway that we can easily iterate through them in lexicograph order, and so\nthat we can easily perform lookups with shortened object IDs.\n\nIn its current form, it is not easily possible to store arbitrary data\nwith the tree nodes. There are a couple of approaches such a caller\ncould try to use, but none of them really work:\n\n  - One may embed the `struct cb_node` in a custom structure. This does\n    not work though as `struct cb_node` contains a flex array, and\n    embedding such a struct in another struct is forbidden.\n\n  - One may use a `union` over `struct cb_node` and ones own data type,\n    which _is_ allowed even if the struct contains a flex array. This\n    does not work though, as the compiler may align members of the\n    struct so that the node key would not immediately start where the\n    flex array starts.\n\n  - One may allocate `struct cb_node` such that it has room for both its\n    key and the custom data. This has the downside though that if the\n    custom data is itself a pointer to allocated memory, then the leak\n    checker will not consider the pointer to be alive anymore.\n\nRefactor the cbtree to drop the flex array and instead take in an\nexplicit offset for where to find the key, which allows the caller to\nembed `struct cb_node` is a wrapper struct.\n\nNote that this change has the downside that we now have a bit of padding\nin our structure, which grows the size from 60 to 64 bytes on a 64 bit\nsystem. On the other hand though, it allows us to get rid of the memory\ncopies that we previously had to do to ensure proper alignment. This\nseems like a reasonable tradeoff.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n cbtree.c  | 25 ++++++++++++++++++-------\n cbtree.h  | 17 +++++++++--------\n oidtree.c | 33 ++++++++++++++-------------------\n 3 files changed, 41 insertions(+), 34 deletions(-)\n\ndiff --git a/cbtree.c b/cbtree.c\nindex 4ab794bddc..8f5edbb80a 100644\n--- a/cbtree.c\n+++ b/cbtree.c\n@@ -7,6 +7,11 @@\n #include \"git-compat-util.h\"\n #include \"cbtree.h\"\n \n+static inline uint8_t *cb_node_key(struct cb_tree *t, struct cb_node *node)\n+{\n+\treturn (uint8_t *) node + t->key_offset;\n+}\n+\n static struct cb_node *cb_node_of(const void *p)\n {\n \treturn (struct cb_node *)((uintptr_t)p - 1);\n@@ -33,6 +38,7 @@ struct cb_node *cb_insert(struct cb_tree *t, struct cb_node *node, size_t klen)\n \tuint8_t c;\n \tint newdirection;\n \tstruct cb_node **wherep, *p;\n+\tuint8_t *node_key, *p_key;\n \n \tassert(!((uintptr_t)node & 1)); /* allocations must be aligned */\n \n@@ -41,23 +47,26 @@ struct cb_node *cb_insert(struct cb_tree *t, struct cb_node *node, size_t klen)\n \t\treturn NULL;\t/* success */\n \t}\n \n+\tnode_key = cb_node_key(t, node);\n+\n \t/* see if a node already exists */\n-\tp = cb_internal_best_match(t->root, node->k, klen);\n+\tp = cb_internal_best_match(t->root, node_key, klen);\n+\tp_key = cb_node_key(t, p);\n \n \t/* find first differing byte */\n \tfor (newbyte = 0; newbyte < klen; newbyte++) {\n-\t\tif (p->k[newbyte] != node->k[newbyte])\n+\t\tif (p_key[newbyte] != node_key[newbyte])\n \t\t\tgoto different_byte_found;\n \t}\n \treturn p;\t/* element exists, let user deal with it */\n \n different_byte_found:\n-\tnewotherbits = p->k[newbyte] ^ node->k[newbyte];\n+\tnewotherbits = p_key[newbyte] ^ node_key[newbyte];\n \tnewotherbits |= newotherbits >> 1;\n \tnewotherbits |= newotherbits >> 2;\n \tnewotherbits |= newotherbits >> 4;\n \tnewotherbits = (newotherbits & ~(newotherbits >> 1)) ^ 255;\n-\tc = p->k[newbyte];\n+\tc = p_key[newbyte];\n \tnewdirection = (1 + (newotherbits | c)) >> 8;\n \n \tnode->byte = newbyte;\n@@ -78,7 +87,7 @@ struct cb_node *cb_insert(struct cb_tree *t, struct cb_node *node, size_t klen)\n \t\t\tbreak;\n \t\tif (q->byte == newbyte && q->otherbits > newotherbits)\n \t\t\tbreak;\n-\t\tc = q->byte < klen ? node->k[q->byte] : 0;\n+\t\tc = q->byte < klen ? node_key[q->byte] : 0;\n \t\tdirection = (1 + (q->otherbits | c)) >> 8;\n \t\twherep = q->child + direction;\n \t}\n@@ -93,7 +102,7 @@ struct cb_node *cb_lookup(struct cb_tree *t, const uint8_t *k, size_t klen)\n {\n \tstruct cb_node *p = cb_internal_best_match(t->root, k, klen);\n \n-\treturn p && !memcmp(p->k, k, klen) ? p : NULL;\n+\treturn p && !memcmp(cb_node_key(t, p), k, klen) ? p : NULL;\n }\n \n static int cb_descend(struct cb_node *p, cb_iter fn, void *arg)\n@@ -115,6 +124,7 @@ int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n \tstruct cb_node *p = t->root;\n \tstruct cb_node *top = p;\n \tsize_t i = 0;\n+\tuint8_t *p_key;\n \n \tif (!p)\n \t\treturn 0; /* empty tree */\n@@ -130,8 +140,9 @@ int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n \t\t\ttop = p;\n \t}\n \n+\tp_key = cb_node_key(t, p);\n \tfor (i = 0; i < klen; i++) {\n-\t\tif (p->k[i] != kpfx[i])\n+\t\tif (p_key[i] != kpfx[i])\n \t\t\treturn 0; /* \"best\" match failed */\n \t}\n \ndiff --git a/cbtree.h b/cbtree.h\nindex c374b1b3db..4647d4a32f 100644\n--- a/cbtree.h\n+++ b/cbtree.h\n@@ -6,9 +6,9 @@\n  *\n  * This is adapted to store arbitrary data (not just NUL-terminated C strings\n  * and allocates no memory internally.  The user needs to allocate\n- * \"struct cb_node\" and fill cb_node.k[] with arbitrary match data\n- * for memcmp.\n- * If \"klen\" is variable, then it should be embedded into \"c_node.k[]\"\n+ * \"struct cb_node\" and provide `key_offset` to indicate where the key can be\n+ * found relative to the `struct cb_node` for memcmp.\n+ * If \"klen\" is variable, then it should be embedded into the key.\n  * Recursion is bound by the maximum value of \"klen\" used.\n  */\n #ifndef CBTREE_H\n@@ -23,18 +23,19 @@ struct cb_node {\n \t */\n \tuint32_t byte;\n \tuint8_t otherbits;\n-\tuint8_t k[FLEX_ARRAY]; /* arbitrary data, unaligned */\n };\n \n struct cb_tree {\n \tstruct cb_node *root;\n+\tptrdiff_t key_offset;\n };\n \n-#define CBTREE_INIT { 0 }\n-\n-static inline void cb_init(struct cb_tree *t)\n+static inline void cb_init(struct cb_tree *t,\n+\t\t\t   ptrdiff_t key_offset)\n {\n-\tstruct cb_tree blank = CBTREE_INIT;\n+\tstruct cb_tree blank = {\n+\t\t.key_offset = key_offset,\n+\t};\n \tmemcpy(t, &blank, sizeof(*t));\n }\n \ndiff --git a/oidtree.c b/oidtree.c\nindex ab9fe7ec7a..117649753f 100644\n--- a/oidtree.c\n+++ b/oidtree.c\n@@ -6,9 +6,14 @@\n #include \"oidtree.h\"\n #include \"hash.h\"\n \n+struct oidtree_node {\n+\tstruct cb_node base;\n+\tstruct object_id key;\n+};\n+\n void oidtree_init(struct oidtree *ot)\n {\n-\tcb_init(&ot->tree);\n+\tcb_init(&ot->tree, offsetof(struct oidtree_node, key));\n \tmem_pool_init(&ot->mem_pool, 0);\n }\n \n@@ -22,20 +27,13 @@ void oidtree_clear(struct oidtree *ot)\n \n void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n {\n-\tstruct cb_node *on;\n-\tstruct object_id k;\n+\tstruct oidtree_node *on;\n \n \tif (!oid->algo)\n \t\tBUG(\"oidtree_insert requires oid->algo\");\n \n-\ton = mem_pool_alloc(&ot->mem_pool, sizeof(*on) + sizeof(*oid));\n-\n-\t/*\n-\t * Clear the padding and copy the result in separate steps to\n-\t * respect the 4-byte alignment needed by struct object_id.\n-\t */\n-\toidcpy(&k, oid);\n-\tmemcpy(on->k, &k, sizeof(k));\n+\ton = mem_pool_alloc(&ot->mem_pool, sizeof(*on));\n+\toidcpy(&on->key, oid);\n \n \t/*\n \t * n.b. Current callers won't get us duplicates, here.  If a\n@@ -43,7 +41,7 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n \t * that won't be freed until oidtree_clear.  Currently it's not\n \t * worth maintaining a free list\n \t */\n-\tcb_insert(&ot->tree, on, sizeof(*oid));\n+\tcb_insert(&ot->tree, &on->base, sizeof(*oid));\n }\n \n bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n@@ -73,21 +71,18 @@ struct oidtree_each_data {\n \n static int iter(struct cb_node *n, void *cb_data)\n {\n+\tstruct oidtree_node *node = container_of(n, struct oidtree_node, base);\n \tstruct oidtree_each_data *data = cb_data;\n-\tstruct object_id k;\n-\n-\t/* Copy to provide 4-byte alignment needed by struct object_id. */\n-\tmemcpy(&k, n->k, sizeof(k));\n \n-\tif (data->algo != GIT_HASH_UNKNOWN && data->algo != k.algo)\n+\tif (data->algo != GIT_HASH_UNKNOWN && data->algo != node->key.algo)\n \t\treturn 0;\n \n \tif (data->last_nibble_at) {\n-\t\tif ((k.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n+\t\tif ((node->key.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n \t\t\treturn 0;\n \t}\n \n-\treturn data->cb(&k, data->cb_data);\n+\treturn data->cb(&node->key, data->cb_data);\n }\n \n int oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541355","messageId":"20260410-b4-pks-odb-source-inmemory-v3-9-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 09/17] oidtree: add ability to store data","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:39Z","receivedAt":"2026-04-10T12:13:12Z","isPatch":true,"body":"The oidtree data structure is currently only used to store object IDs,\nwithout any associated data. So consequently, it can only really be used\nto track which object IDs exist, and we can use the tree structure to\nefficiently operate on OID prefixes.\n\nBut there are valid use cases where we want to both:\n\n  - Store object IDs in a sorted order.\n\n  - Associated arbitrary data with them.\n\nRefactor the oidtree interface so that it allows us to store arbitrary\npayloads within the respective nodes. This will be used in the next\ncommit.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c                  |  2 +-\n object-file.c            |  3 ++-\n oidtree.c                | 37 ++++++++++++++++++++++++++++++++-----\n oidtree.h                | 12 ++++++++++--\n t/unit-tests/u-oidtree.c | 26 +++++++++++++++++++++++---\n 5 files changed, 68 insertions(+), 12 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 07333be696..f7a3dd1a72 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -57,7 +57,7 @@ static int insert_loose_map(struct odb_source *source,\n \tinserted |= insert_oid_pair(map->to_compat, oid, compat_oid);\n \tinserted |= insert_oid_pair(map->to_storage, compat_oid, oid);\n \tif (inserted)\n-\t\toidtree_insert(files->loose->cache, compat_oid);\n+\t\toidtree_insert(files->loose->cache, compat_oid, NULL);\n \n \treturn inserted;\n }\ndiff --git a/object-file.c b/object-file.c\nindex 3e70e5d668..d04ab57253 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1857,6 +1857,7 @@ static int for_each_object_wrapper_cb(const struct object_id *oid,\n }\n \n static int for_each_prefixed_object_wrapper_cb(const struct object_id *oid,\n+\t\t\t\t\t       void *node_data UNUSED,\n \t\t\t\t\t       void *cb_data)\n {\n \tstruct for_each_object_wrapper_data *data = cb_data;\n@@ -2002,7 +2003,7 @@ static int append_loose_object(const struct object_id *oid,\n \t\t\t       const char *path UNUSED,\n \t\t\t       void *data)\n {\n-\toidtree_insert(data, oid);\n+\toidtree_insert(data, oid, NULL);\n \treturn 0;\n }\n \ndiff --git a/oidtree.c b/oidtree.c\nindex 117649753f..e43f18026e 100644\n--- a/oidtree.c\n+++ b/oidtree.c\n@@ -9,6 +9,7 @@\n struct oidtree_node {\n \tstruct cb_node base;\n \tstruct object_id key;\n+\tvoid *data;\n };\n \n void oidtree_init(struct oidtree *ot)\n@@ -25,15 +26,22 @@ void oidtree_clear(struct oidtree *ot)\n \t}\n }\n \n-void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n+struct oidtree_data {\n+\tstruct object_id oid;\n+};\n+\n+void oidtree_insert(struct oidtree *ot, const struct object_id *oid,\n+\t\t    void *data)\n {\n \tstruct oidtree_node *on;\n+\tstruct cb_node *node;\n \n \tif (!oid->algo)\n \t\tBUG(\"oidtree_insert requires oid->algo\");\n \n \ton = mem_pool_alloc(&ot->mem_pool, sizeof(*on));\n \toidcpy(&on->key, oid);\n+\ton->data = data;\n \n \t/*\n \t * n.b. Current callers won't get us duplicates, here.  If a\n@@ -41,13 +49,19 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n \t * that won't be freed until oidtree_clear.  Currently it's not\n \t * worth maintaining a free list\n \t */\n-\tcb_insert(&ot->tree, &on->base, sizeof(*oid));\n+\tnode = cb_insert(&ot->tree, &on->base, sizeof(*oid));\n+\tif (node) {\n+\t\tstruct oidtree_node *preexisting = container_of(node, struct oidtree_node, base);\n+\t\tpreexisting->data = data;\n+\t}\n }\n \n-bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n+static struct oidtree_node *oidtree_lookup(struct oidtree *ot,\n+\t\t\t\t\t   const struct object_id *oid)\n {\n \tstruct object_id k;\n \tsize_t klen = sizeof(k);\n+\tstruct cb_node *node;\n \n \toidcpy(&k, oid);\n \n@@ -58,7 +72,20 @@ bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n \tklen += BUILD_ASSERT_OR_ZERO(offsetof(struct object_id, hash) <\n \t\t\t\toffsetof(struct object_id, algo));\n \n-\treturn !!cb_lookup(&ot->tree, (const uint8_t *)&k, klen);\n+\tnode = cb_lookup(&ot->tree, (const uint8_t *)&k, klen);\n+\treturn node ? container_of(node, struct oidtree_node, base) : NULL;\n+}\n+\n+bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n+{\n+\tstruct oidtree_node *node = oidtree_lookup(ot, oid);\n+\treturn node ? 1 : 0;\n+}\n+\n+void *oidtree_get(struct oidtree *ot, const struct object_id *oid)\n+{\n+\tstruct oidtree_node *node = oidtree_lookup(ot, oid);\n+\treturn node ? node->data : NULL;\n }\n \n struct oidtree_each_data {\n@@ -82,7 +109,7 @@ static int iter(struct cb_node *n, void *cb_data)\n \t\t\treturn 0;\n \t}\n \n-\treturn data->cb(&node->key, data->cb_data);\n+\treturn data->cb(&node->key, node->data, data->cb_data);\n }\n \n int oidtree_each(struct oidtree *ot, const struct object_id *prefix,\ndiff --git a/oidtree.h b/oidtree.h\nindex 2b7bad2e60..baa5a436ea 100644\n--- a/oidtree.h\n+++ b/oidtree.h\n@@ -29,18 +29,26 @@ void oidtree_init(struct oidtree *ot);\n  */\n void oidtree_clear(struct oidtree *ot);\n \n-/* Insert the object ID into the tree. */\n-void oidtree_insert(struct oidtree *ot, const struct object_id *oid);\n+/*\n+ * Insert the object ID into the tree and store the given pointer alongside\n+ * with it. The data pointer of any preexisting entry will be overwritten.\n+ */\n+void oidtree_insert(struct oidtree *ot, const struct object_id *oid,\n+\t\t    void *data);\n \n /* Check whether the tree contains the given object ID. */\n bool oidtree_contains(struct oidtree *ot, const struct object_id *oid);\n \n+/* Get the payload stored with the given object ID. */\n+void *oidtree_get(struct oidtree *ot, const struct object_id *oid);\n+\n /*\n  * Callback function used for `oidtree_each()`. Returning a non-zero exit code\n  * will cause iteration to stop. The exit code will be propagated to the caller\n  * of `oidtree_each()`.\n  */\n typedef int (*oidtree_each_cb)(const struct object_id *oid,\n+\t\t\t       void *node_data,\n \t\t\t       void *cb_data);\n \n /*\ndiff --git a/t/unit-tests/u-oidtree.c b/t/unit-tests/u-oidtree.c\nindex d4d05c7dc3..f0d5ebb733 100644\n--- a/t/unit-tests/u-oidtree.c\n+++ b/t/unit-tests/u-oidtree.c\n@@ -19,7 +19,7 @@ static int fill_tree_loc(struct oidtree *ot, const char *hexes[], size_t n)\n \tfor (size_t i = 0; i < n; i++) {\n \t\tstruct object_id oid;\n \t\tcl_parse_any_oid(hexes[i], &oid);\n-\t\toidtree_insert(ot, &oid);\n+\t\toidtree_insert(ot, &oid, NULL);\n \t}\n \treturn 0;\n }\n@@ -38,9 +38,9 @@ struct expected_hex_iter {\n \tconst char *query;\n };\n \n-static int check_each_cb(const struct object_id *oid, void *data)\n+static int check_each_cb(const struct object_id *oid, void *node_data UNUSED, void *cb_data)\n {\n-\tstruct expected_hex_iter *hex_iter = data;\n+\tstruct expected_hex_iter *hex_iter = cb_data;\n \tstruct object_id expected;\n \n \tcl_assert(hex_iter->i < hex_iter->expected_hexes.nr);\n@@ -105,3 +105,23 @@ void test_oidtree__each(void)\n \tcheck_each(&ot, \"32100\", \"321\", NULL);\n \tcheck_each(&ot, \"32\", \"320\", \"321\", NULL);\n }\n+\n+void test_oidtree__insert_overwrites_data(void)\n+{\n+\tstruct object_id oid;\n+\tstruct oidtree ot;\n+\tint a, b;\n+\n+\tcl_parse_any_oid(\"1\", &oid);\n+\n+\toidtree_init(&ot);\n+\n+\toidtree_insert(&ot, &oid, NULL);\n+\tcl_assert_equal_p(oidtree_get(&ot, &oid), NULL);\n+\toidtree_insert(&ot, &oid, &a);\n+\tcl_assert_equal_p(oidtree_get(&ot, &oid), &a);\n+\toidtree_insert(&ot, &oid, &b);\n+\tcl_assert_equal_p(oidtree_get(&ot, &oid), &b);\n+\n+\toidtree_clear(&ot);\n+}\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541356","messageId":"20260410-b4-pks-odb-source-inmemory-v3-10-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 10/17] odb/source-inmemory: convert to use oidtree","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:40Z","receivedAt":"2026-04-10T12:13:14Z","isPatch":true,"body":"The in-memory source stores its objects in a simple array that we grow as\nneeded. This has a couple of downsides:\n\n  - The object lookup is O(n). This doesn't matter in practice because\n    we only store a small number of objects.\n\n  - We don't have an easy way to iterate over all objects in\n    lexicographic order.\n\n  - We don't have an easy way to compute unique object ID prefixes.\n\nRefactor the code to use an oidtree instead. This is the same data\nstructure used by our loose object source, and thus it means we get a\nbunch of functionality for free.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 72 +++++++++++++++++++++++++++++++++++++--------------\n odb/source-inmemory.h | 13 ++--------\n 2 files changed, 54 insertions(+), 31 deletions(-)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex d05a13df45..3b51cc7fef 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -3,20 +3,29 @@\n #include \"odb.h\"\n #include \"odb/source-inmemory.h\"\n #include \"odb/streaming.h\"\n+#include \"oidtree.h\"\n #include \"repository.h\"\n \n-static const struct cached_object *find_cached_object(struct odb_source_inmemory *source,\n-\t\t\t\t\t\t      const struct object_id *oid)\n+struct inmemory_object {\n+\tenum object_type type;\n+\tconst void *buf;\n+\tunsigned long size;\n+};\n+\n+static const struct inmemory_object *find_cached_object(struct odb_source_inmemory *source,\n+\t\t\t\t\t\t\tconst struct object_id *oid)\n {\n-\tstatic const struct cached_object empty_tree = {\n+\tstatic const struct inmemory_object empty_tree = {\n \t\t.type = OBJ_TREE,\n \t\t.buf = \"\",\n \t};\n-\tconst struct cached_object_entry *co = source->objects;\n+\tconst struct inmemory_object *object;\n \n-\tfor (size_t i = 0; i < source->objects_nr; i++, co++)\n-\t\tif (oideq(&co->oid, oid))\n-\t\t\treturn &co->value;\n+\tif (source->objects) {\n+\t\tobject = oidtree_get(source->objects, oid);\n+\t\tif (object)\n+\t\t\treturn object;\n+\t}\n \n \tif (oid->algo && oideq(oid, hash_algos[oid->algo].empty_tree))\n \t\treturn &empty_tree;\n@@ -30,7 +39,7 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \t\t\t\t\t\tenum object_info_flags flags UNUSED)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n-\tconst struct cached_object *object;\n+\tconst struct inmemory_object *object;\n \n \tobject = find_cached_object(inmemory, oid);\n \tif (!object)\n@@ -88,7 +97,7 @@ static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n \tstruct odb_read_stream_inmemory *stream;\n-\tconst struct cached_object *object;\n+\tconst struct inmemory_object *object;\n \n \tobject = find_cached_object(inmemory, oid);\n \tif (!object)\n@@ -113,17 +122,23 @@ static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    enum odb_write_object_flags flags UNUSED)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n-\tstruct cached_object_entry *object;\n+\tstruct inmemory_object *object;\n \n \thash_object_file(source->odb->repo->hash_algo, buf, len, type, oid);\n \n-\tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n-\t\t   inmemory->objects_alloc);\n-\tobject = &inmemory->objects[inmemory->objects_nr++];\n-\tobject->value.size = len;\n-\tobject->value.type = type;\n-\tobject->value.buf = xmemdupz(buf, len);\n-\toidcpy(&object->oid, oid);\n+\tif (!inmemory->objects) {\n+\t\tCALLOC_ARRAY(inmemory->objects, 1);\n+\t\toidtree_init(inmemory->objects);\n+\t} else if (oidtree_contains(inmemory->objects, oid)) {\n+\t\treturn 0;\n+\t}\n+\n+\tCALLOC_ARRAY(object, 1);\n+\tobject->size = len;\n+\tobject->type = type;\n+\tobject->buf = xmemdupz(buf, len);\n+\n+\toidtree_insert(inmemory->objects, oid, object);\n \n \treturn 0;\n }\n@@ -167,12 +182,29 @@ static int odb_source_inmemory_write_object_stream(struct odb_source *source,\n \treturn ret;\n }\n \n+static int inmemory_object_free(const struct object_id *oid UNUSED,\n+\t\t\t\tvoid *node_data,\n+\t\t\t\tvoid *cb_data UNUSED)\n+{\n+\tstruct inmemory_object *object = node_data;\n+\tfree((void *) object->buf);\n+\tfree(object);\n+\treturn 0;\n+}\n+\n static void odb_source_inmemory_free(struct odb_source *source)\n {\n \tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n-\tfor (size_t i = 0; i < inmemory->objects_nr; i++)\n-\t\tfree((char *) inmemory->objects[i].value.buf);\n-\tfree(inmemory->objects);\n+\n+\tif (inmemory->objects) {\n+\t\tstruct object_id null_oid = { 0 };\n+\n+\t\toidtree_each(inmemory->objects, &null_oid, 0,\n+\t\t\t     inmemory_object_free, NULL);\n+\t\toidtree_clear(inmemory->objects);\n+\t\tfree(inmemory->objects);\n+\t}\n+\n \tfree(inmemory->base.path);\n \tfree(inmemory);\n }\ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex d1b05a3996..a88fc2e320 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -3,14 +3,7 @@\n \n #include \"odb/source.h\"\n \n-struct cached_object_entry {\n-\tstruct object_id oid;\n-\tstruct cached_object {\n-\t\tenum object_type type;\n-\t\tconst void *buf;\n-\t\tunsigned long size;\n-\t} value;\n-};\n+struct oidtree;\n \n /*\n  * An in-memory source that you can write objects to that shall be made\n@@ -20,9 +13,7 @@ struct cached_object_entry {\n  */\n struct odb_source_inmemory {\n \tstruct odb_source base;\n-\n-\tstruct cached_object_entry *objects;\n-\tsize_t objects_nr, objects_alloc;\n+\tstruct oidtree *objects;\n };\n \n /* Create a new in-memory object database source. */\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541357","messageId":"20260410-b4-pks-odb-source-inmemory-v3-11-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 11/17] odb/source-inmemory: implement `for_each_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:41Z","receivedAt":"2026-04-10T12:13:16Z","isPatch":true,"body":"Implement the `for_each_object()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 88 +++++++++++++++++++++++++++++++++++++++++----------\n 1 file changed, 72 insertions(+), 16 deletions(-)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 3b51cc7fef..f60eecbdbb 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -33,6 +33,28 @@ static const struct inmemory_object *find_cached_object(struct odb_source_inmemo\n \treturn NULL;\n }\n \n+static void populate_object_info(struct odb_source_inmemory *source,\n+\t\t\t\t struct object_info *oi,\n+\t\t\t\t const struct inmemory_object *object)\n+{\n+\tif (!oi)\n+\t\treturn;\n+\n+\tif (oi->typep)\n+\t\t*(oi->typep) = object->type;\n+\tif (oi->sizep)\n+\t\t*(oi->sizep) = object->size;\n+\tif (oi->disk_sizep)\n+\t\t*(oi->disk_sizep) = 0;\n+\tif (oi->delta_base_oid)\n+\t\toidclr(oi->delta_base_oid, source->base.odb->repo->hash_algo);\n+\tif (oi->contentp)\n+\t\t*oi->contentp = xmemdupz(object->buf, object->size);\n+\tif (oi->mtimep)\n+\t\t*oi->mtimep = 0;\n+\toi->whence = OI_CACHED;\n+}\n+\n static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \t\t\t\t\t\tconst struct object_id *oid,\n \t\t\t\t\t\tstruct object_info *oi,\n@@ -45,22 +67,7 @@ static int odb_source_inmemory_read_object_info(struct odb_source *source,\n \tif (!object)\n \t\treturn -1;\n \n-\tif (oi) {\n-\t\tif (oi->typep)\n-\t\t\t*(oi->typep) = object->type;\n-\t\tif (oi->sizep)\n-\t\t\t*(oi->sizep) = object->size;\n-\t\tif (oi->disk_sizep)\n-\t\t\t*(oi->disk_sizep) = 0;\n-\t\tif (oi->delta_base_oid)\n-\t\t\toidclr(oi->delta_base_oid, source->odb->repo->hash_algo);\n-\t\tif (oi->contentp)\n-\t\t\t*oi->contentp = xmemdupz(object->buf, object->size);\n-\t\tif (oi->mtimep)\n-\t\t\t*oi->mtimep = 0;\n-\t\toi->whence = OI_CACHED;\n-\t}\n-\n+\tpopulate_object_info(inmemory, oi, object);\n \treturn 0;\n }\n \n@@ -114,6 +121,54 @@ static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n \treturn 0;\n }\n \n+struct odb_source_inmemory_for_each_object_data {\n+\tstruct odb_source_inmemory *inmemory;\n+\tconst struct object_info *request;\n+\todb_for_each_object_cb cb;\n+\tvoid *cb_data;\n+};\n+\n+static int odb_source_inmemory_for_each_object_cb(const struct object_id *oid,\n+\t\t\t\t\t\t  void *node_data, void *cb_data)\n+{\n+\tstruct odb_source_inmemory_for_each_object_data *data = cb_data;\n+\tstruct inmemory_object *object = node_data;\n+\n+\tif (data->request) {\n+\t\tstruct object_info oi = *data->request;\n+\t\tpopulate_object_info(data->inmemory, &oi, object);\n+\t\treturn data->cb(oid, &oi, data->cb_data);\n+\t} else {\n+\t\treturn data->cb(oid, NULL, data->cb_data);\n+\t}\n+}\n+\n+static int odb_source_inmemory_for_each_object(struct odb_source *source,\n+\t\t\t\t\t       const struct object_info *request,\n+\t\t\t\t\t       odb_for_each_object_cb cb,\n+\t\t\t\t\t       void *cb_data,\n+\t\t\t\t\t       const struct odb_for_each_object_options *opts)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tstruct odb_source_inmemory_for_each_object_data payload = {\n+\t\t.inmemory = inmemory,\n+\t\t.request = request,\n+\t\t.cb = cb,\n+\t\t.cb_data = cb_data,\n+\t};\n+\tstruct object_id null_oid = { 0 };\n+\n+\tif ((opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY) ||\n+\t    (opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY && !source->local))\n+\t\treturn 0;\n+\tif (!inmemory->objects)\n+\t\treturn 0;\n+\n+\treturn oidtree_each(inmemory->objects,\n+\t\t\t    opts->prefix ? opts->prefix : &null_oid, opts->prefix_hex_len,\n+\t\t\t    odb_source_inmemory_for_each_object_cb, &payload);\n+}\n+\n static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    const void *buf, unsigned long len,\n \t\t\t\t\t    enum object_type type,\n@@ -219,6 +274,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.free = odb_source_inmemory_free;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n+\tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541358","messageId":"20260410-b4-pks-odb-source-inmemory-v3-12-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 12/17] odb/source-inmemory: implement `find_abbrev_len()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:42Z","receivedAt":"2026-04-10T12:13:20Z","isPatch":true,"body":"Implement the `find_abbrev_len()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 39 +++++++++++++++++++++++++++++++++++++++\n 1 file changed, 39 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex f60eecbdbb..44d9bbedec 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -169,6 +169,44 @@ static int odb_source_inmemory_for_each_object(struct odb_source *source,\n \t\t\t    odb_source_inmemory_for_each_object_cb, &payload);\n }\n \n+struct find_abbrev_len_data {\n+\tconst struct object_id *oid;\n+\tunsigned len;\n+};\n+\n+static int find_abbrev_len_cb(const struct object_id *oid,\n+\t\t\t      struct object_info *oi UNUSED,\n+\t\t\t      void *cb_data)\n+{\n+\tstruct find_abbrev_len_data *data = cb_data;\n+\tunsigned len = oid_common_prefix_hexlen(oid, data->oid);\n+\tif (len != hash_algos[oid->algo].hexsz && len >= data->len)\n+\t\tdata->len = len + 1;\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t\t       const struct object_id *oid,\n+\t\t\t\t\t       unsigned min_len,\n+\t\t\t\t\t       unsigned *out)\n+{\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = oid,\n+\t\t.prefix_hex_len = min_len,\n+\t};\n+\tstruct find_abbrev_len_data data = {\n+\t\t.oid = oid,\n+\t\t.len = min_len,\n+\t};\n+\tint ret;\n+\n+\tret = odb_source_inmemory_for_each_object(source, NULL, find_abbrev_len_cb,\n+\t\t\t\t\t\t  &data, &opts);\n+\t*out = data.len;\n+\n+\treturn ret;\n+}\n+\n static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    const void *buf, unsigned long len,\n \t\t\t\t\t    enum object_type type,\n@@ -275,6 +313,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n+\tsource->base.find_abbrev_len = odb_source_inmemory_find_abbrev_len;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541359","messageId":"20260410-b4-pks-odb-source-inmemory-v3-13-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 13/17] odb/source-inmemory: implement `count_objects()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:43Z","receivedAt":"2026-04-10T12:13:22Z","isPatch":true,"body":"Implement the `count_objects()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 20 ++++++++++++++++++++\n 1 file changed, 20 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 44d9bbedec..674dbcad30 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -207,6 +207,25 @@ static int odb_source_inmemory_find_abbrev_len(struct odb_source *source,\n \treturn ret;\n }\n \n+static int count_objects_cb(const struct object_id *oid UNUSED,\n+\t\t\t    struct object_info *oi UNUSED,\n+\t\t\t    void *cb_data)\n+{\n+\tunsigned long *counter = cb_data;\n+\t(*counter)++;\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_count_objects(struct odb_source *source,\n+\t\t\t\t\t     enum odb_count_objects_flags flags UNUSED,\n+\t\t\t\t\t     unsigned long *out)\n+{\n+\tstruct odb_for_each_object_options opts = { 0 };\n+\t*out = 0;\n+\treturn odb_source_inmemory_for_each_object(source, NULL, count_objects_cb,\n+\t\t\t\t\t\t   out, &opts);\n+}\n+\n static int odb_source_inmemory_write_object(struct odb_source *source,\n \t\t\t\t\t    const void *buf, unsigned long len,\n \t\t\t\t\t    enum object_type type,\n@@ -314,6 +333,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n \tsource->base.find_abbrev_len = odb_source_inmemory_find_abbrev_len;\n+\tsource->base.count_objects = odb_source_inmemory_count_objects;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541360","messageId":"20260410-b4-pks-odb-source-inmemory-v3-14-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 14/17] odb/source-inmemory: implement `freshen_object()` callback","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:44Z","receivedAt":"2026-04-10T12:13:24Z","isPatch":true,"body":"Implement the `freshen_object()` callback function for the in-memory\nsource.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 10 ++++++++++\n 1 file changed, 10 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 674dbcad30..8934e0f547 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -294,6 +294,15 @@ static int odb_source_inmemory_write_object_stream(struct odb_source *source,\n \treturn ret;\n }\n \n+static int odb_source_inmemory_freshen_object(struct odb_source *source,\n+\t\t\t\t\t      const struct object_id *oid)\n+{\n+\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n+\tif (find_cached_object(inmemory, oid))\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n static int inmemory_object_free(const struct object_id *oid UNUSED,\n \t\t\t\tvoid *node_data,\n \t\t\t\tvoid *cb_data UNUSED)\n@@ -336,6 +345,7 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.count_objects = odb_source_inmemory_count_objects;\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n+\tsource->base.freshen_object = odb_source_inmemory_freshen_object;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541361","messageId":"20260410-b4-pks-odb-source-inmemory-v3-15-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 15/17] odb/source-inmemory: stub out remaining functions","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:45Z","receivedAt":"2026-04-10T12:13:26Z","isPatch":true,"body":"Stub out remaining functions that we either don't need or that are\nbasically no-ops.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-inmemory.c | 31 +++++++++++++++++++++++++++++++\n 1 file changed, 31 insertions(+)\n\ndiff --git a/odb/source-inmemory.c b/odb/source-inmemory.c\nindex 8934e0f547..e004566d76 100644\n--- a/odb/source-inmemory.c\n+++ b/odb/source-inmemory.c\n@@ -303,6 +303,32 @@ static int odb_source_inmemory_freshen_object(struct odb_source *source,\n \treturn 0;\n }\n \n+static int odb_source_inmemory_begin_transaction(struct odb_source *source UNUSED,\n+\t\t\t\t\t\t struct odb_transaction **out UNUSED)\n+{\n+\treturn error(\"in-memory source does not support transactions\");\n+}\n+\n+static int odb_source_inmemory_read_alternates(struct odb_source *source UNUSED,\n+\t\t\t\t\t       struct strvec *out UNUSED)\n+{\n+\treturn 0;\n+}\n+\n+static int odb_source_inmemory_write_alternate(struct odb_source *source UNUSED,\n+\t\t\t\t\t       const char *alternate UNUSED)\n+{\n+\treturn error(\"in-memory source does not support alternates\");\n+}\n+\n+static void odb_source_inmemory_close(struct odb_source *source UNUSED)\n+{\n+}\n+\n+static void odb_source_inmemory_reprepare(struct odb_source *source UNUSED)\n+{\n+}\n+\n static int inmemory_object_free(const struct object_id *oid UNUSED,\n \t\t\t\tvoid *node_data,\n \t\t\t\tvoid *cb_data UNUSED)\n@@ -338,6 +364,8 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \todb_source_init(&source->base, odb, ODB_SOURCE_INMEMORY, \"source\", false);\n \n \tsource->base.free = odb_source_inmemory_free;\n+\tsource->base.close = odb_source_inmemory_close;\n+\tsource->base.reprepare = odb_source_inmemory_reprepare;\n \tsource->base.read_object_info = odb_source_inmemory_read_object_info;\n \tsource->base.read_object_stream = odb_source_inmemory_read_object_stream;\n \tsource->base.for_each_object = odb_source_inmemory_for_each_object;\n@@ -346,6 +374,9 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n \tsource->base.write_object = odb_source_inmemory_write_object;\n \tsource->base.write_object_stream = odb_source_inmemory_write_object_stream;\n \tsource->base.freshen_object = odb_source_inmemory_freshen_object;\n+\tsource->base.begin_transaction = odb_source_inmemory_begin_transaction;\n+\tsource->base.read_alternates = odb_source_inmemory_read_alternates;\n+\tsource->base.write_alternate = odb_source_inmemory_write_alternate;\n \n \treturn source;\n }\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541362","messageId":"20260410-b4-pks-odb-source-inmemory-v3-16-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 16/17] odb: generic in-memory source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:46Z","receivedAt":"2026-04-10T12:13:30Z","isPatch":true,"body":"Make the in-memory source generic.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c | 8 ++++----\n odb.h | 2 +-\n 2 files changed, 5 insertions(+), 5 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex 24e929f03c..965ef68e4e 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -560,7 +560,7 @@ static int do_oid_object_info_extended(struct object_database *odb,\n \tif (is_null_oid(real))\n \t\treturn -1;\n \n-\tif (!odb_source_read_object_info(&odb->inmemory_objects->base, oid, oi, flags))\n+\tif (!odb_source_read_object_info(odb->inmemory_objects, oid, oi, flags))\n \t\treturn 0;\n \n \todb_prepare_alternates(odb);\n@@ -737,7 +737,7 @@ int odb_pretend_object(struct object_database *odb,\n \tif (odb_has_object(odb, oid, 0))\n \t\treturn 0;\n \n-\treturn odb_source_write_object(&odb->inmemory_objects->base,\n+\treturn odb_source_write_object(odb->inmemory_objects,\n \t\t\t\t       buf, len, type, oid, NULL, 0);\n }\n \n@@ -1020,7 +1020,7 @@ struct object_database *odb_new(struct repository *repo,\n \to->sources = odb_source_new(o, primary_source, true);\n \to->sources_tail = &o->sources->next;\n \to->alternate_db = xstrdup_or_null(secondary_sources);\n-\to->inmemory_objects = odb_source_inmemory_new(o);\n+\to->inmemory_objects = &odb_source_inmemory_new(o)->base;\n \n \tfree(to_free);\n \n@@ -1045,7 +1045,7 @@ static void odb_free_sources(struct object_database *o)\n \t\to->sources = next;\n \t}\n \n-\todb_source_free(&o->inmemory_objects->base);\n+\todb_source_free(o->inmemory_objects);\n \to->inmemory_objects = NULL;\n \n \tkh_destroy_odb_path_map(o->source_by_path);\ndiff --git a/odb.h b/odb.h\nindex c3a7edf9c8..73553ed5a7 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -81,7 +81,7 @@ struct object_database {\n \t * to write them into the object store (e.g. a browse-only\n \t * application).\n \t */\n-\tstruct odb_source_inmemory *inmemory_objects;\n+\tstruct odb_source *inmemory_objects;\n \n \t/*\n \t * A fast, rough count of the number of objects in the repository.\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541363","messageId":"20260410-b4-pks-odb-source-inmemory-v3-17-22fd0fad58fe@pks.im","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"[PATCH v3 17/17] t/unit-tests: add tests for the in-memory object source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-04-10T12:12:47Z","receivedAt":"2026-04-10T12:13:32Z","isPatch":true,"body":"While the in-memory object source is a full-fledged source, our code\nbase only exercises parts of its functionality because we only use it in\ngit-blame(1). Implement unit tests to verify that the yet-unused\nfunctionality of the backend works as expected.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n Makefile                      |   1 +\n t/meson.build                 |   1 +\n t/unit-tests/u-odb-inmemory.c | 313 ++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 315 insertions(+)\n\ndiff --git a/Makefile b/Makefile\nindex 3cda12c455..68b4daa1ad 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1529,6 +1529,7 @@ CLAR_TEST_SUITES += u-hash\n CLAR_TEST_SUITES += u-hashmap\n CLAR_TEST_SUITES += u-list-objects-filter-options\n CLAR_TEST_SUITES += u-mem-pool\n+CLAR_TEST_SUITES += u-odb-inmemory\n CLAR_TEST_SUITES += u-oid-array\n CLAR_TEST_SUITES += u-oidmap\n CLAR_TEST_SUITES += u-oidtree\ndiff --git a/t/meson.build b/t/meson.build\nindex 7528e5cda5..db5e01c49b 100644\n--- a/t/meson.build\n+++ b/t/meson.build\n@@ -6,6 +6,7 @@ clar_test_suites = [\n   'unit-tests/u-hashmap.c',\n   'unit-tests/u-list-objects-filter-options.c',\n   'unit-tests/u-mem-pool.c',\n+  'unit-tests/u-odb-inmemory.c',\n   'unit-tests/u-oid-array.c',\n   'unit-tests/u-oidmap.c',\n   'unit-tests/u-oidtree.c',\ndiff --git a/t/unit-tests/u-odb-inmemory.c b/t/unit-tests/u-odb-inmemory.c\nnew file mode 100644\nindex 0000000000..482502ef4b\n--- /dev/null\n+++ b/t/unit-tests/u-odb-inmemory.c\n@@ -0,0 +1,313 @@\n+#include \"unit-test.h\"\n+#include \"hex.h\"\n+#include \"odb/source-inmemory.h\"\n+#include \"odb/streaming.h\"\n+#include \"oidset.h\"\n+#include \"repository.h\"\n+#include \"strbuf.h\"\n+\n+#define RANDOM_OID \"da39a3ee5e6b4b0d3255bfef95601890afd80709\"\n+#define FOOBAR_OID \"f6ea0495187600e7b2288c8ac19c5886383a4632\"\n+\n+static struct repository repo = {\n+\t.hash_algo = &hash_algos[GIT_HASH_SHA1],\n+};\n+static struct object_database *odb;\n+\n+static void cl_assert_object_info(struct odb_source_inmemory *source,\n+\t\t\t\t  const struct object_id *oid,\n+\t\t\t\t  enum object_type expected_type,\n+\t\t\t\t  const char *expected_content)\n+{\n+\tenum object_type actual_type;\n+\tunsigned long actual_size;\n+\tvoid *actual_content;\n+\tstruct object_info oi = {\n+\t\t.typep = &actual_type,\n+\t\t.sizep = &actual_size,\n+\t\t.contentp = &actual_content,\n+\t};\n+\n+\tcl_must_pass(odb_source_read_object_info(&source->base, oid, &oi, 0));\n+\tcl_assert_equal_u(actual_size, strlen(expected_content));\n+\tcl_assert_equal_u(actual_type, expected_type);\n+\tcl_assert_equal_s((char *) actual_content, expected_content);\n+\n+\tfree(actual_content);\n+}\n+\n+void test_odb_inmemory__initialize(void)\n+{\n+\todb = odb_new(&repo, \"\", \"\");\n+}\n+\n+void test_odb_inmemory__cleanup(void)\n+{\n+\todb_free(odb);\n+}\n+\n+void test_odb_inmemory__new(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tcl_assert_equal_i(source->base.type, ODB_SOURCE_INMEMORY);\n+\todb_source_free(&source->base);\n+}\n+\n+void test_odb_inmemory__read_missing_object(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tstruct object_id oid;\n+\tconst char *end;\n+\n+\tcl_must_pass(parse_oid_hex_algop(RANDOM_OID, &oid, &end, repo.hash_algo));\n+\tcl_must_fail(odb_source_read_object_info(&source->base, &oid, NULL, 0));\n+\n+\todb_source_free(&source->base);\n+}\n+\n+void test_odb_inmemory__read_empty_tree(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tcl_assert_object_info(source, repo.hash_algo->empty_tree, OBJ_TREE, \"\");\n+\todb_source_free(&source->base);\n+}\n+\n+void test_odb_inmemory__read_written_object(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tconst char data[] = \"foobar\";\n+\tstruct object_id written_oid;\n+\n+\tcl_must_pass(odb_source_write_object(&source->base, data, strlen(data),\n+\t\t\t\t\t     OBJ_BLOB, &written_oid, NULL, 0));\n+\tcl_assert_equal_s(oid_to_hex(&written_oid), FOOBAR_OID);\n+\tcl_assert_object_info(source, &written_oid, OBJ_BLOB, \"foobar\");\n+\n+\todb_source_free(&source->base);\n+}\n+\n+void test_odb_inmemory__read_stream_object(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tstruct odb_read_stream *stream;\n+\tstruct object_id written_oid;\n+\tconst char data[] = \"foobar\";\n+\tchar buf[3] = { 0 };\n+\n+\tcl_must_pass(odb_source_write_object(&source->base, data, strlen(data),\n+\t\t\t\t\t     OBJ_BLOB, &written_oid, NULL, 0));\n+\n+\tcl_must_pass(odb_source_read_object_stream(&stream, &source->base,\n+\t\t\t\t\t\t   &written_oid));\n+\tcl_assert_equal_i(stream->type, OBJ_BLOB);\n+\tcl_assert_equal_u(stream->size, 6);\n+\n+\tcl_assert_equal_i(odb_read_stream_read(stream, buf, 2), 2);\n+\tcl_assert_equal_s(buf, \"fo\");\n+\tcl_assert_equal_i(odb_read_stream_read(stream, buf, 2), 2);\n+\tcl_assert_equal_s(buf, \"ob\");\n+\tcl_assert_equal_i(odb_read_stream_read(stream, buf, 2), 2);\n+\tcl_assert_equal_s(buf, \"ar\");\n+\tcl_assert_equal_i(odb_read_stream_read(stream, buf, 2), 0);\n+\n+\todb_read_stream_close(stream);\n+\todb_source_free(&source->base);\n+}\n+\n+static int add_one_object(const struct object_id *oid,\n+\t\t\t  struct object_info *oi UNUSED,\n+\t\t\t  void *payload)\n+{\n+\tstruct oidset *actual_oids = payload;\n+\tcl_must_pass(oidset_insert(actual_oids, oid));\n+\treturn 0;\n+}\n+\n+void test_odb_inmemory__for_each_object(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tstruct odb_for_each_object_options opts = { 0 };\n+\tstruct oidset expected_oids = OIDSET_INIT;\n+\tstruct oidset actual_oids = OIDSET_INIT;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\n+\tcl_must_pass(odb_source_for_each_object(&source->base, NULL,\n+\t\t\t\t\t\tadd_one_object, &actual_oids, &opts));\n+\tcl_assert_equal_u(oidset_size(&actual_oids), 0);\n+\n+\tfor (int i = 0; i < 10; i++) {\n+\t\tstruct object_id written_oid;\n+\n+\t\tstrbuf_reset(&buf);\n+\t\tstrbuf_addf(&buf, \"%d\", i);\n+\n+\t\tcl_must_pass(odb_source_write_object(&source->base, buf.buf, buf.len,\n+\t\t\t\t\t\t     OBJ_BLOB, &written_oid, NULL, 0));\n+\t\tcl_must_pass(oidset_insert(&expected_oids, &written_oid));\n+\t}\n+\n+\tcl_must_pass(odb_source_for_each_object(&source->base, NULL,\n+\t\t\t\t\t\tadd_one_object, &actual_oids, &opts));\n+\tcl_assert_equal_b(oidset_equal(&expected_oids, &actual_oids), true);\n+\n+\todb_source_free(&source->base);\n+\toidset_clear(&expected_oids);\n+\toidset_clear(&actual_oids);\n+\tstrbuf_release(&buf);\n+}\n+\n+static int abort_after_two_objects(const struct object_id *oid UNUSED,\n+\t\t\t\t   struct object_info *oi UNUSED,\n+\t\t\t\t   void *payload)\n+{\n+\tunsigned *counter = payload;\n+\t(*counter)++;\n+\tif (*counter == 2)\n+\t\treturn 123;\n+\treturn 0;\n+}\n+\n+void test_odb_inmemory__for_each_object_can_abort_iteration(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tstruct odb_for_each_object_options opts = { 0 };\n+\tstruct object_id written_oid;\n+\tunsigned counter = 0;\n+\n+\tcl_must_pass(odb_source_write_object(&source->base, \"1\", 1,\n+\t\t\t\t\t     OBJ_BLOB, &written_oid, NULL, 0));\n+\tcl_must_pass(odb_source_write_object(&source->base, \"2\", 1,\n+\t\t\t\t\t     OBJ_BLOB, &written_oid, NULL, 0));\n+\tcl_must_pass(odb_source_write_object(&source->base, \"3\", 1,\n+\t\t\t\t\t     OBJ_BLOB, &written_oid, NULL, 0));\n+\n+\tcl_assert_equal_i(odb_source_for_each_object(&source->base, NULL,\n+\t\t\t\t\t\t     abort_after_two_objects,\n+\t\t\t\t\t\t     &counter, &opts),\n+\t\t\t  123);\n+\tcl_assert_equal_u(counter, 2);\n+\n+\todb_source_free(&source->base);\n+}\n+\n+void test_odb_inmemory__count_objects(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tstruct object_id written_oid;\n+\tunsigned long count;\n+\n+\tcl_must_pass(odb_source_count_objects(&source->base, 0, &count));\n+\tcl_assert_equal_u(count, 0);\n+\n+\tcl_must_pass(odb_source_write_object(&source->base, \"1\", 1,\n+\t\t\t\t\t     OBJ_BLOB, &written_oid, NULL, 0));\n+\tcl_must_pass(odb_source_write_object(&source->base, \"2\", 1,\n+\t\t\t\t\t     OBJ_BLOB, &written_oid, NULL, 0));\n+\tcl_must_pass(odb_source_write_object(&source->base, \"3\", 1,\n+\t\t\t\t\t     OBJ_BLOB, &written_oid, NULL, 0));\n+\n+\tcl_must_pass(odb_source_count_objects(&source->base, 0, &count));\n+\tcl_assert_equal_u(count, 3);\n+\n+\todb_source_free(&source->base);\n+}\n+\n+void test_odb_inmemory__find_abbrev_len(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tstruct object_id oid1, oid2;\n+\tunsigned abbrev_len;\n+\n+\t/*\n+\t * The two blobs we're about to write share the first 10 hex characters\n+\t * of their object IDs (\"a09f43dc45\"), so at least 11 characters are\n+\t * needed to tell them apart:\n+\t *\n+\t *   \"368317\" -> a09f43dc4562d45115583f5094640ae237df55f7\n+\t *   \"514796\" -> a09f43dc45fef837235eb7e6b1a6ca5e169a3981\n+\t *\n+\t * With only one blob written we expect a length of 4.\n+\t */\n+\tcl_must_pass(odb_source_write_object(&source->base, \"368317\", strlen(\"368317\"),\n+\t\t\t\t\t     OBJ_BLOB, &oid1, NULL, 0));\n+\tcl_must_pass(odb_source_find_abbrev_len(&source->base, &oid1, 4,\n+\t\t\t\t\t\t&abbrev_len));\n+\tcl_assert_equal_u(abbrev_len, 4);\n+\n+\t/*\n+\t * With both objects present, the shared 10-character prefix means we\n+\t * need at least 11 characters to uniquely identify either object.\n+\t */\n+\tcl_must_pass(odb_source_write_object(&source->base, \"514796\", strlen(\"514796\"),\n+\t\t\t\t\t     OBJ_BLOB, &oid2, NULL, 0));\n+\tcl_must_pass(odb_source_find_abbrev_len(&source->base, &oid1, 4,\n+\t\t\t\t\t\t&abbrev_len));\n+\tcl_assert_equal_u(abbrev_len, 11);\n+\n+\todb_source_free(&source->base);\n+}\n+\n+void test_odb_inmemory__freshen_object(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tstruct object_id written_oid;\n+\tstruct object_id oid;\n+\tconst char *end;\n+\n+\tcl_must_pass(parse_oid_hex_algop(RANDOM_OID, &oid, &end, repo.hash_algo));\n+\tcl_assert_equal_i(odb_source_freshen_object(&source->base, &oid), 0);\n+\n+\tcl_must_pass(odb_source_write_object(&source->base, \"foobar\",\n+\t\t\t\t\t     strlen(\"foobar\"), OBJ_BLOB,\n+\t\t\t\t\t     &written_oid, NULL, 0));\n+\tcl_assert_equal_i(odb_source_freshen_object(&source->base,\n+\t\t\t\t\t\t    &written_oid), 1);\n+\n+\todb_source_free(&source->base);\n+}\n+\n+struct membuf_write_stream {\n+\tstruct odb_write_stream base;\n+\tconst char *buf;\n+\tsize_t offset;\n+\tsize_t size;\n+};\n+\n+static ssize_t membuf_write_stream_read(struct odb_write_stream *stream,\n+\t\t\t\t\tunsigned char *buf, size_t len)\n+{\n+\tstruct membuf_write_stream *s = container_of(stream, struct membuf_write_stream, base);\n+\tsize_t chunk_size = 2;\n+\n+\tif (chunk_size > len)\n+\t\tchunk_size = len;\n+\tif (chunk_size > s->size - s->offset)\n+\t\tchunk_size = s->size - s->offset;\n+\n+\tmemcpy(buf, s->buf + s->offset, chunk_size);\n+\n+\ts->offset += chunk_size;\n+\tif (s->offset == s->size)\n+\t\ts->base.is_finished = 1;\n+\n+\treturn chunk_size;\n+}\n+\n+void test_odb_inmemory__write_object_stream(void)\n+{\n+\tstruct odb_source_inmemory *source = odb_source_inmemory_new(odb);\n+\tconst char data[] = \"foobar\";\n+\tstruct membuf_write_stream stream = {\n+\t\t.base.read = membuf_write_stream_read,\n+\t\t.buf = data,\n+\t\t.size = strlen(data),\n+\t};\n+\tstruct object_id written_oid;\n+\n+\tcl_must_pass(odb_source_write_object_stream(&source->base, &stream.base,\n+\t\t\t\t\t\t    strlen(data), &written_oid));\n+\tcl_assert_equal_s(oid_to_hex(&written_oid), FOOBAR_OID);\n+\tcl_assert_object_info(source, &written_oid, OBJ_BLOB, \"foobar\");\n+\n+\todb_source_free(&source->base);\n+}\n\n-- \n2.54.0.rc0.707.g0fbf48f4d6.dirty\n\n"},{"id":"541531","messageId":"CAOLa=ZSrThty13-C_WVa5dvakZAtidwOXWnUrOA4LGX93DvmGQ@mail.gmail.com","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-0-22fd0fad58fe@pks.im","subject":"Re: [PATCH v3 00/17] odb: introduce \"in-memory\" source","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-04-14T08:27:26Z","receivedAt":"2026-04-14T08:27:29Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n[snip]\n\n> Range-diff versus v2:\n>\n>  1:  b18e427c69 !  1:  155b2cdf81 odb: introduce \"in-memory\" source\n>     @@ odb/source-inmemory.h (new)\n>      +struct cached_object_entry;\n>      +\n>      +/*\n>     -+ * An inmemory source that you can write objects to that shall be made\n>     ++ * An in-memory source that you can write objects to that shall be made\n>      + * available for reading, but that shouldn't ever be persisted to disk. Note\n>      + * that any objects written to this source will be stored in memory, so the\n>      + * number of objects you can store is limited by available system memory.\n>     @@ odb/source-inmemory.h (new)\n>      +struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb);\n>      +\n>      +/*\n>     -+ * Cast the given object database source to the inmemory backend. This will\n>     ++ * Cast the given object database source to the in-memory backend. This will\n>      + * cause a BUG in case the source doesn't use this backend.\n>      + */\n>      +static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n>      +{\n>      +\tif (source->type != ODB_SOURCE_INMEMORY)\n>     -+\t\tBUG(\"trying to downcast source of type '%d' to inmemory\", source->type);\n>     ++\t\tBUG(\"trying to downcast source of type '%d' to in-memory\", source->type);\n>      +\treturn container_of(source, struct odb_source_inmemory, base);\n>      +}\n>      +\n>     @@ odb/source.h: enum odb_source_type {\n>       \t/* The \"files\" backend that uses loose objects and packfiles. */\n>       \tODB_SOURCE_FILES,\n>      +\n>     -+\t/* The \"inmemory\" backend that stores objects in memory. */\n>     ++\t/* The \"in-memory\" backend that stores objects in memory. */\n>      +\tODB_SOURCE_INMEMORY,\n>       };\n>\n>  2:  8fd337da90 !  2:  c66edd10a8 odb/source-inmemory: implement `free()` callback\n>     @@ odb/source-inmemory.h\n>      +};\n>\n>       /*\n>     -  * An inmemory source that you can write objects to that shall be made\n>     +  * An in-memory source that you can write objects to that shall be made\n>  3:  f4ae2a2bde =  3:  a86549f39c odb: fix unnecessary call to `find_cached_object()`\n>  4:  8600b88530 =  4:  49ac739dd2 odb/source-inmemory: implement `read_object_info()` callback\n>  5:  ab33c0b7ee !  5:  321ef11be3 odb/source-inmemory: implement `read_object_stream()` callback\n>     @@ odb/source-inmemory.c: static int odb_source_inmemory_read_object_info(struct od\n>\n>      +struct odb_read_stream_inmemory {\n>      +\tstruct odb_read_stream base;\n>     -+\tconst void *buf;\n>     ++\tconst unsigned char *buf;\n\nOkay this does make more sense.\n\n>      +\tsize_t offset;\n>      +};\n>      +\n>     @@ odb/source-inmemory.c: static int odb_source_inmemory_read_object_info(struct od\n>      +\n>      +\tif (buf_len > inmemory->base.size - inmemory->offset)\n>      +\t\tbytes = inmemory->base.size - inmemory->offset;\n>     -+\tmemcpy(buf, inmemory->buf, bytes);\n>     ++\n>     ++\tmemcpy(buf, inmemory->buf + inmemory->offset, bytes);\n>     ++\tinmemory->offset += bytes;\n\nNow, we also use the offset correctly.\n\n>      +\n>      +\treturn bytes;\n>      +}\n>  6:  983f886eeb !  6:  506df5e488 odb/source-inmemory: implement `write_object()` callback\n>     @@ odb.c: int odb_pretend_object(struct object_database *odb,\n>       void *odb_read_object(struct object_database *odb,\n>\n>       ## odb/source-inmemory.c ##\n>     +@@\n>     + #include \"git-compat-util.h\"\n>     ++#include \"object-file.h\"\n>     + #include \"odb.h\"\n>     + #include \"odb/source-inmemory.h\"\n>     + #include \"odb/streaming.h\"\n>      @@ odb/source-inmemory.c: static int odb_source_inmemory_read_object_stream(struct odb_read_stream **out,\n>       \treturn 0;\n>       }\n>     @@ odb/source-inmemory.c: static int odb_source_inmemory_read_object_stream(struct\n>      +\tstruct odb_source_inmemory *inmemory = odb_source_inmemory_downcast(source);\n>      +\tstruct cached_object_entry *object;\n>      +\n>     ++\thash_object_file(source->odb->repo->hash_algo, buf, len, type, oid);\n>     ++\n>      +\tALLOC_GROW(inmemory->objects, inmemory->objects_nr + 1,\n>      +\t\t   inmemory->objects_alloc);\n>      +\tobject = &inmemory->objects[inmemory->objects_nr++];\n>  7:  68edefa269 <  -:  ---------- odb/source-inmemory: implement `write_object()` callback\n>  8:  18d451152b !  7:  21eef34c1b odb/source-inmemory: implement `write_object_stream()` callback\n>     @@ odb/source-inmemory.c: static int odb_source_inmemory_write_object(struct odb_so\n>      +\t\t\tgoto out;\n>      +\t\t}\n>      +\n>     -+\t\tmemcpy(data, buf, bytes_read);\n>     ++\t\tmemcpy(data + total_read, buf, bytes_read);\n>      +\t\ttotal_read += bytes_read;\n>      +\t}\n>      +\n>  9:  cee53b9853 !  8:  504e34d116 cbtree: allow using arbitrary wrapper structures for nodes\n>     @@ cbtree.c: int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n>\n>\n>       ## cbtree.h ##\n>     +@@\n>     +  *\n>     +  * This is adapted to store arbitrary data (not just NUL-terminated C strings\n>     +  * and allocates no memory internally.  The user needs to allocate\n>     +- * \"struct cb_node\" and fill cb_node.k[] with arbitrary match data\n>     +- * for memcmp.\n>     +- * If \"klen\" is variable, then it should be embedded into \"c_node.k[]\"\n>     ++ * \"struct cb_node\" and provide `key_offset` to indicate where the key can be\n>     ++ * found relative to the `struct cb_node` for memcmp.\n>     ++ * If \"klen\" is variable, then it should be embedded into the key.\n>     +  * Recursion is bound by the maximum value of \"klen\" used.\n>     +  */\n\nWe fix up the comments here also.\n\n>     + #ifndef CBTREE_H\n>      @@ cbtree.h: struct cb_node {\n>       \t */\n>       \tuint32_t byte;\n> 10:  8ad5b81b13 =  9:  9bdd475a92 oidtree: add ability to store data\n> 11:  1ed2d23137 ! 10:  956b989529 odb/source-inmemory: convert to use oidtree\n>     @@ odb/source-inmemory.h\n>      +struct oidtree;\n>\n>       /*\n>     -  * An inmemory source that you can write objects to that shall be made\n>     +  * An in-memory source that you can write objects to that shall be made\n>      @@ odb/source-inmemory.h: struct cached_object_entry {\n>        */\n>       struct odb_source_inmemory {\n> 12:  99fbb1cc35 ! 11:  bec1428116 odb/source-inmemory: implement `for_each_object()` callback\n>     @@ odb/source-inmemory.c: static int odb_source_inmemory_read_object_stream(struct\n>      +\tif ((opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY) ||\n>      +\t    (opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY && !source->local))\n>      +\t\treturn 0;\n>     ++\tif (!inmemory->objects)\n>     ++\t\treturn 0;\n>      +\n>      +\treturn oidtree_each(inmemory->objects,\n>      +\t\t\t    opts->prefix ? opts->prefix : &null_oid, opts->prefix_hex_len,\n> 13:  c87a621f39 = 12:  32dada3c27 odb/source-inmemory: implement `find_abbrev_len()` callback\n> 14:  9b88f0c07b = 13:  43127840c0 odb/source-inmemory: implement `count_objects()` callback\n> 15:  3c9493f2bb = 14:  439acbd068 odb/source-inmemory: implement `freshen_object()` callback\n> 16:  f2b6317104 ! 15:  12c1b6ffd2 odb/source-inmemory: stub out remaining functions\n>     @@ odb/source-inmemory.c: static int odb_source_inmemory_freshen_object(struct odb_\n>      +static int odb_source_inmemory_begin_transaction(struct odb_source *source UNUSED,\n>      +\t\t\t\t\t\t struct odb_transaction **out UNUSED)\n>      +{\n>     -+\treturn error(\"inmemory source does not support transactions\");\n>     ++\treturn error(\"in-memory source does not support transactions\");\n>      +}\n>      +\n>      +static int odb_source_inmemory_read_alternates(struct odb_source *source UNUSED,\n>     @@ odb/source-inmemory.c: static int odb_source_inmemory_freshen_object(struct odb_\n>      +static int odb_source_inmemory_write_alternate(struct odb_source *source UNUSED,\n>      +\t\t\t\t\t       const char *alternate UNUSED)\n>      +{\n>     -+\treturn error(\"inmemory source does not support alternates\");\n>     ++\treturn error(\"in-memory source does not support alternates\");\n>      +}\n>      +\n>      +static void odb_source_inmemory_close(struct odb_source *source UNUSED)\n> 17:  81da5d5048 = 16:  ef37a61e7f odb: generic in-memory source\n>  -:  ---------- > 17:  51b51e0382 t/unit-tests: add tests for the in-memory object source\n\nThe range diff looks good. I'll have a look at the unit test patch\nindependently. Thanks\n"},{"id":"541532","messageId":"CAOLa=ZQnrtU5MP-J2-8rffbBacSUbm=m503k_v-TYSR4Qy781A@mail.gmail.com","threadId":"65423","inReplyTo":"20260410-b4-pks-odb-source-inmemory-v3-17-22fd0fad58fe@pks.im","subject":"Re: [PATCH v3 17/17] t/unit-tests: add tests for the in-memory object source","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-04-14T08:45:16Z","receivedAt":"2026-04-14T08:45:18Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> While the in-memory object source is a full-fledged source, our code\n> base only exercises parts of its functionality because we only use it in\n> git-blame(1). Implement unit tests to verify that the yet-unused\n> functionality of the backend works as expected.\n>\n\nThis patch seems extensive and good!\n\nOverall I'm happy with this version.\n"}]}