{"thread":{"id":"65199","subject":"[PATCH 0/6] odb: introduce generic object counting","startedAt":"2026-03-10T15:18:28Z","lastAt":"2026-03-13T11:52:47Z","messageCount":27,"participants":["Patrick Steinhardt","Junio C Hamano","Toon Claes"],"isPatch":true,"patchVersion":1,"patchTotal":6},"messages":[{"id":"538460","messageId":"20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im","threadId":"65199","inReplyTo":null,"subject":"[PATCH 0/6] odb: introduce generic object counting","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-10T15:18:20Z","receivedAt":"2026-03-10T15:18:28Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Hi,\n\nthis small patch series introduces generic object counting for pluggable\nobject databases. The series is built on top of d181b9354c (The 13th\nbatch, 2026-03-09) with ps/odb-sources at d6fc6fe6f8 (odb/source: make\n`begin_transaction()` function pluggable, 2026-03-05) merged into it.\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (6):\n      odb: stop including \"odb/source.h\"\n      packfile: extract logic to count number of objects\n      object-file: extract logic to approximate object count\n      object-file: generalize counting objects\n      odb/source: introduce generic object counting\n      odb: introduce generic object counting\n\n builtin/gc.c                | 44 +++++++++----------------\n builtin/multi-pack-index.c  |  1 +\n builtin/submodule--helper.c |  1 +\n commit-graph.c              |  3 +-\n object-file.c               | 57 ++++++++++++++++++++++++++++++++\n object-file.h               | 14 ++++++++\n object-name.c               |  6 +++-\n odb.c                       | 37 ++++++++++++++++++++-\n odb.h                       | 76 +++++++++++++++++++++++++++++++++++++++++--\n odb/source-files.c          | 30 +++++++++++++++++\n odb/source.h                | 79 ++++++++++++++++-----------------------------\n odb/streaming.c             |  1 +\n packfile.c                  | 48 +++++++++++++--------------\n packfile.h                  | 16 +++++----\n repository.c                |  1 +\n submodule-config.c          |  1 +\n tmp-objdir.c                |  1 +\n 17 files changed, 299 insertions(+), 117 deletions(-)\n\n\n---\nbase-commit: 2247f478a898a7f8f8322cc51bdeb1cc773d8f4a\nchange-id: 20260224-b4-pks-odb-source-count-objects-479fe682cf6f\n\n"},{"id":"538461","messageId":"20260310-b4-pks-odb-source-count-objects-v1-1-109e07d425f4@pks.im","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im","subject":"[PATCH 1/6] odb: stop including \"odb/source.h\"","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-10T15:18:21Z","receivedAt":"2026-03-10T15:18:30Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"The \"odb.h\" header currently includes the \"odb/source.h\" file. This is\nsomewhat roundabout though: most callers shouldn't have to care about\nthe `struct odb_source`, but should rather use the ODB-level functions.\nFurthermore, it means that a couple of definitions have to live on the\nsource level even though they should be part of the generic interface.\n\nReverse the relation between \"odb/source.h\" and \"odb.h\" and move the\nenums and typedefs that relate to the generic interfaces back into\n\"odb.h\". Add the necessary includes to all files that rely on the\ntransitive include.\n\nSuggested-by: Justin Tobler <jltobler@gmail.com>\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/multi-pack-index.c  |  1 +\n builtin/submodule--helper.c |  1 +\n odb.h                       | 50 ++++++++++++++++++++++++++++++++++++++++++-\n odb/source.h                | 52 +--------------------------------------------\n odb/streaming.c             |  1 +\n repository.c                |  1 +\n submodule-config.c          |  1 +\n tmp-objdir.c                |  1 +\n 8 files changed, 56 insertions(+), 52 deletions(-)\n\ndiff --git a/builtin/multi-pack-index.c b/builtin/multi-pack-index.c\nindex 5f364aa816..3fcb207f1a 100644\n--- a/builtin/multi-pack-index.c\n+++ b/builtin/multi-pack-index.c\n@@ -9,6 +9,7 @@\n #include \"strbuf.h\"\n #include \"trace2.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"replace-object.h\"\n #include \"repository.h\"\n \ndiff --git a/builtin/submodule--helper.c b/builtin/submodule--helper.c\nindex 143f7cb3cc..4957487536 100644\n--- a/builtin/submodule--helper.c\n+++ b/builtin/submodule--helper.c\n@@ -29,6 +29,7 @@\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"advice.h\"\n #include \"branch.h\"\n #include \"list-objects-filter-options.h\"\ndiff --git a/odb.h b/odb.h\nindex 86e0365c24..7a583e3873 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -3,7 +3,6 @@\n \n #include \"hashmap.h\"\n #include \"object.h\"\n-#include \"odb/source.h\"\n #include \"oidset.h\"\n #include \"oidmap.h\"\n #include \"string-list.h\"\n@@ -12,6 +11,7 @@\n struct oidmap;\n struct oidtree;\n struct strbuf;\n+struct strvec;\n struct repository;\n struct multi_pack_index;\n \n@@ -339,6 +339,42 @@ struct object_info {\n  */\n #define OBJECT_INFO_INIT { 0 }\n \n+/* Flags that can be passed to `odb_read_object_info_extended()`. */\n+enum object_info_flags {\n+\t/* Invoke lookup_replace_object() on the given hash. */\n+\tOBJECT_INFO_LOOKUP_REPLACE = (1 << 0),\n+\n+\t/* Do not reprepare object sources when the first lookup has failed. */\n+\tOBJECT_INFO_QUICK = (1 << 1),\n+\n+\t/*\n+\t * Do not attempt to fetch the object if missing (even if fetch_is_missing is\n+\t * nonzero).\n+\t */\n+\tOBJECT_INFO_SKIP_FETCH_OBJECT = (1 << 2),\n+\n+\t/* Die if object corruption (not just an object being missing) was detected. */\n+\tOBJECT_INFO_DIE_IF_CORRUPT = (1 << 3),\n+\n+\t/*\n+\t * We have already tried reading the object, but it couldn't be found\n+\t * via any of the attached sources, and are now doing a second read.\n+\t * This second read asks the individual sources to also evaluate\n+\t * whether any on-disk state may have changed that may have caused the\n+\t * object to appear.\n+\t *\n+\t * This flag is for internal use, only. The second read only occurs\n+\t * when `OBJECT_INFO_QUICK` was not passed.\n+\t */\n+\tOBJECT_INFO_SECOND_READ = (1 << 4),\n+\n+\t/*\n+\t * This is meant for bulk prefetching of missing blobs in a partial\n+\t * clone. Implies OBJECT_INFO_SKIP_FETCH_OBJECT and OBJECT_INFO_QUICK.\n+\t */\n+\tOBJECT_INFO_FOR_PREFETCH = (OBJECT_INFO_SKIP_FETCH_OBJECT | OBJECT_INFO_QUICK),\n+};\n+\n /*\n  * Read object info from the object database and populate the `object_info`\n  * structure. Returns 0 on success, a negative error code otherwise.\n@@ -432,6 +468,18 @@ enum odb_for_each_object_flags {\n \tODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS = (1<<4),\n };\n \n+/*\n+ * A callback function that can be used to iterate through objects. If given,\n+ * the optional `oi` parameter will be populated the same as if you would call\n+ * `odb_read_object_info()`.\n+ *\n+ * Returning a non-zero error code will cause iteration to abort. The error\n+ * code will be propagated.\n+ */\n+typedef int (*odb_for_each_object_cb)(const struct object_id *oid,\n+\t\t\t\t      struct object_info *oi,\n+\t\t\t\t      void *cb_data);\n+\n /*\n  * Iterate through all objects contained in the object database. Note that\n  * objects may be iterated over multiple times in case they are either stored\ndiff --git a/odb/source.h b/odb/source.h\nindex caac558149..a1fd9dd920 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -2,6 +2,7 @@\n #define ODB_SOURCE_H\n \n #include \"object.h\"\n+#include \"odb.h\"\n \n enum odb_source_type {\n \t/*\n@@ -14,61 +15,10 @@ enum odb_source_type {\n \tODB_SOURCE_FILES,\n };\n \n-/* Flags that can be passed to `odb_read_object_info_extended()`. */\n-enum object_info_flags {\n-\t/* Invoke lookup_replace_object() on the given hash. */\n-\tOBJECT_INFO_LOOKUP_REPLACE = (1 << 0),\n-\n-\t/* Do not reprepare object sources when the first lookup has failed. */\n-\tOBJECT_INFO_QUICK = (1 << 1),\n-\n-\t/*\n-\t * Do not attempt to fetch the object if missing (even if fetch_is_missing is\n-\t * nonzero).\n-\t */\n-\tOBJECT_INFO_SKIP_FETCH_OBJECT = (1 << 2),\n-\n-\t/* Die if object corruption (not just an object being missing) was detected. */\n-\tOBJECT_INFO_DIE_IF_CORRUPT = (1 << 3),\n-\n-\t/*\n-\t * We have already tried reading the object, but it couldn't be found\n-\t * via any of the attached sources, and are now doing a second read.\n-\t * This second read asks the individual sources to also evaluate\n-\t * whether any on-disk state may have changed that may have caused the\n-\t * object to appear.\n-\t *\n-\t * This flag is for internal use, only. The second read only occurs\n-\t * when `OBJECT_INFO_QUICK` was not passed.\n-\t */\n-\tOBJECT_INFO_SECOND_READ = (1 << 4),\n-\n-\t/*\n-\t * This is meant for bulk prefetching of missing blobs in a partial\n-\t * clone. Implies OBJECT_INFO_SKIP_FETCH_OBJECT and OBJECT_INFO_QUICK.\n-\t */\n-\tOBJECT_INFO_FOR_PREFETCH = (OBJECT_INFO_SKIP_FETCH_OBJECT | OBJECT_INFO_QUICK),\n-};\n-\n struct object_id;\n-struct object_info;\n struct odb_read_stream;\n-struct odb_transaction;\n-struct odb_write_stream;\n struct strvec;\n \n-/*\n- * A callback function that can be used to iterate through objects. If given,\n- * the optional `oi` parameter will be populated the same as if you would call\n- * `odb_read_object_info()`.\n- *\n- * Returning a non-zero error code will cause iteration to abort. The error\n- * code will be propagated.\n- */\n-typedef int (*odb_for_each_object_cb)(const struct object_id *oid,\n-\t\t\t\t      struct object_info *oi,\n-\t\t\t\t      void *cb_data);\n-\n /*\n  * The source is the part of the object database that stores the actual\n  * objects. It thus encapsulates the logic to read and write the specific\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex a4355cd245..5927a12954 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -7,6 +7,7 @@\n #include \"environment.h\"\n #include \"repository.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"odb/streaming.h\"\n #include \"replace-object.h\"\n \ndiff --git a/repository.c b/repository.c\nindex e7fa42c14f..05c26bdbc3 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -2,6 +2,7 @@\n #include \"abspath.h\"\n #include \"repository.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"config.h\"\n #include \"object.h\"\n #include \"lockfile.h\"\ndiff --git a/submodule-config.c b/submodule-config.c\nindex 1f19fe2077..72a46b7a54 100644\n--- a/submodule-config.c\n+++ b/submodule-config.c\n@@ -14,6 +14,7 @@\n #include \"strbuf.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"parse-options.h\"\n #include \"thread-utils.h\"\n #include \"tree-walk.h\"\ndiff --git a/tmp-objdir.c b/tmp-objdir.c\nindex e436eed07e..d199d39e7c 100644\n--- a/tmp-objdir.c\n+++ b/tmp-objdir.c\n@@ -11,6 +11,7 @@\n #include \"strvec.h\"\n #include \"quote.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"repository.h\"\n \n struct tmp_objdir {\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538462","messageId":"20260310-b4-pks-odb-source-count-objects-v1-2-109e07d425f4@pks.im","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im","subject":"[PATCH 2/6] packfile: extract logic to count number of objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-10T15:18:22Z","receivedAt":"2026-03-10T15:18:31Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"In a subsequent commit we're about to introduce a new\n`odb_source_count_objects()` function so that we can make the logic\npluggable. Prepare for this change by extracting the logic that we have\nto count packed objects into a standalone function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n packfile.c | 45 +++++++++++++++++++++++++++++++++++----------\n packfile.h |  9 +++++++++\n 2 files changed, 44 insertions(+), 10 deletions(-)\n\ndiff --git a/packfile.c b/packfile.c\nindex 215a23e42b..1ee5dd3da3 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1101,6 +1101,36 @@ struct packfile_list_entry *packfile_store_get_packs(struct packfile_store *stor\n \treturn store->packs.head;\n }\n \n+int packfile_store_count_objects(struct packfile_store *store,\n+\t\t\t\t unsigned long *out)\n+{\n+\tstruct packfile_list_entry *e;\n+\tstruct multi_pack_index *m;\n+\tunsigned long count = 0;\n+\tint ret;\n+\n+\tm = get_multi_pack_index(store->source);\n+\tif (m)\n+\t\tcount += m->num_objects + m->num_objects_in_base;\n+\n+\tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n+\t\tif (e->pack->multi_pack_index)\n+\t\t\tcontinue;\n+\t\tif (open_pack_index(e->pack)) {\n+\t\t\tret = -1;\n+\t\t\tgoto out;\n+\t\t}\n+\n+\t\tcount += e->pack->num_objects;\n+\t}\n+\n+\t*out = count;\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n /*\n  * Give a fast, rough count of the number of objects in the repository. This\n  * ignores loose objects completely. If you have a lot of them, then either\n@@ -1113,21 +1143,16 @@ unsigned long repo_approximate_object_count(struct repository *r)\n \tif (!r->objects->approximate_object_count_valid) {\n \t\tstruct odb_source *source;\n \t\tunsigned long count = 0;\n-\t\tstruct packed_git *p;\n \n \t\todb_prepare_alternates(r->objects);\n-\n \t\tfor (source = r->objects->sources; source; source = source->next) {\n-\t\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n-\t\t\tif (m)\n-\t\t\t\tcount += m->num_objects + m->num_objects_in_base;\n-\t\t}\n+\t\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\t\t\tunsigned long c;\n \n-\t\trepo_for_each_pack(r, p) {\n-\t\t\tif (p->multi_pack_index || open_pack_index(p))\n-\t\t\t\tcontinue;\n-\t\t\tcount += p->num_objects;\n+\t\t\tif (!packfile_store_count_objects(files->packed, &c))\n+\t\t\t\tcount += c;\n \t\t}\n+\n \t\tr->objects->approximate_object_count = count;\n \t\tr->objects->approximate_object_count_valid = 1;\n \t}\ndiff --git a/packfile.h b/packfile.h\nindex 8b04a258a7..1da8c729cb 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -268,6 +268,15 @@ enum kept_pack_type {\n \tKEPT_PACK_IN_CORE = (1 << 1),\n };\n \n+/*\n+ * Count the number objects contained in the given packfile store. If\n+ * successful, the number of objects will be written to the `out` pointer.\n+ *\n+ * Return 0 on success, a negative error code otherwise.\n+ */\n+int packfile_store_count_objects(struct packfile_store *store,\n+\t\t\t\t unsigned long *out);\n+\n /*\n  * Retrieve the cache of kept packs from the given packfile store. Accepts a\n  * combination of `kept_pack_type` flags. The cache is computed on demand and\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538463","messageId":"20260310-b4-pks-odb-source-count-objects-v1-3-109e07d425f4@pks.im","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im","subject":"[PATCH 3/6] object-file: extract logic to approximate object count","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-10T15:18:23Z","receivedAt":"2026-03-10T15:18:34Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"In \"builtin/gc.c\" we have some logic that checks whether we need to\nrepack objects. This is done by counting the number of objects that we\nhave and checking whether it exceeds a certain threshold. We don't\nreally need an accurate object count though, which is why we only\nopen a single object diretcroy shard and then extrapolate from there.\n\nExtract this logic into a new function that is owned by the loose object\ndatabase source. This is done to prepare for a subsequent change, where\nwe'll introduce object counting on the object database source level.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/gc.c  | 37 +++++++++----------------------------\n object-file.c | 41 +++++++++++++++++++++++++++++++++++++++++\n object-file.h | 13 +++++++++++++\n 3 files changed, 63 insertions(+), 28 deletions(-)\n\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex fb329c2cff..a08c7554cb 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -467,37 +467,18 @@ static int rerere_gc_condition(struct gc_config *cfg UNUSED)\n static int too_many_loose_objects(int limit)\n {\n \t/*\n-\t * Quickly check if a \"gc\" is needed, by estimating how\n-\t * many loose objects there are.  Because SHA-1 is evenly\n-\t * distributed, we can check only one and get a reasonable\n-\t * estimate.\n+\t * This is weird, but stems from legacy behaviour: the GC auto\n+\t * threshold was always essentially interpreted as if it was rounded up\n+\t * to the next multiple 256 of, so we retain this behaviour for now.\n \t */\n-\tDIR *dir;\n-\tstruct dirent *ent;\n-\tint auto_threshold;\n-\tint num_loose = 0;\n-\tint needed = 0;\n-\tconst unsigned hexsz_loose = the_hash_algo->hexsz - 2;\n-\tchar *path;\n-\n-\tpath = repo_git_path(the_repository, \"objects/17\");\n-\tdir = opendir(path);\n-\tfree(path);\n-\tif (!dir)\n+\tint auto_threshold = DIV_ROUND_UP(limit, 256) * 256;\n+\tunsigned long loose_count;\n+\n+\tif (odb_source_loose_approximate_object_count(the_repository->objects->sources,\n+\t\t\t\t\t\t      &loose_count) < 0)\n \t\treturn 0;\n \n-\tauto_threshold = DIV_ROUND_UP(limit, 256);\n-\twhile ((ent = readdir(dir)) != NULL) {\n-\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz_loose ||\n-\t\t    ent->d_name[hexsz_loose] != '\\0')\n-\t\t\tcontinue;\n-\t\tif (++num_loose > auto_threshold) {\n-\t\t\tneeded = 1;\n-\t\t\tbreak;\n-\t\t}\n-\t}\n-\tclosedir(dir);\n-\treturn needed;\n+\treturn loose_count > auto_threshold;\n }\n \n static struct packed_git *find_base_packs(struct string_list *packs,\ndiff --git a/object-file.c b/object-file.c\nindex a3ff7f586c..da67e3c9ff 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1868,6 +1868,47 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t\t     NULL, NULL, &data);\n }\n \n+int odb_source_loose_approximate_object_count(struct odb_source *source,\n+\t\t\t\t\t      unsigned long *out)\n+{\n+\tconst unsigned hexsz = source->odb->repo->hash_algo->hexsz - 2;\n+\tunsigned long count = 0;\n+\tstruct dirent *ent;\n+\tchar *path = NULL;\n+\tDIR *dir = NULL;\n+\tint ret;\n+\n+\tpath = xstrfmt(\"%s/17\", source->path);\n+\n+\tdir = opendir(path);\n+\tif (!dir) {\n+\t\tif (errno == ENOENT) {\n+\t\t\t*out = 0;\n+\t\t\tret = 0;\n+\t\t\tgoto out;\n+\t\t}\n+\n+\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n+\t\tgoto out;\n+\t}\n+\n+\twhile ((ent = readdir(dir)) != NULL) {\n+\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n+\t\t    ent->d_name[hexsz] != '\\0')\n+\t\t\tcontinue;\n+\t\tcount++;\n+\t}\n+\n+\t*out = count * 256;\n+\tret = 0;\n+\n+out:\n+\tif (dir)\n+\t\tclosedir(dir);\n+\tfree(path);\n+\treturn ret;\n+}\n+\n static int append_loose_object(const struct object_id *oid,\n \t\t\t       const char *path UNUSED,\n \t\t\t       void *data)\ndiff --git a/object-file.h b/object-file.h\nindex ff6da65296..b870ea9fa8 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -139,6 +139,19 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t     void *cb_data,\n \t\t\t\t     unsigned flags);\n \n+/*\n+ * Count the number of loose objects in this source.\n+ *\n+ * The object count is approximated by opening a single sharding directory for\n+ * loose objects and scanning its contents. The result is then extrapolated by\n+ * 256. This should generally work as a reasonable estimate given that the\n+ * object hash is supposed to be indistinguishable from random.\n+ *\n+ * Returns 0 on success, a negative error code otherwise.\n+ */\n+int odb_source_loose_approximate_object_count(struct odb_source *source,\n+\t\t\t\t\t      unsigned long *out);\n+\n /**\n  * format_object_header() is a thin wrapper around s xsnprintf() that\n  * writes the initial \"<type> <obj-len>\" part of the loose object\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538464","messageId":"20260310-b4-pks-odb-source-count-objects-v1-4-109e07d425f4@pks.im","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im","subject":"[PATCH 4/6] object-file: generalize counting objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-10T15:18:24Z","receivedAt":"2026-03-10T15:18:37Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Generalize the function introduced in the preceding commit to not only\nbe able to approximate the number of loose objects, but to also provide\nan accurate count. The behaviour can be toggled via a new flag.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/gc.c  |  5 +++--\n object-file.c | 58 +++++++++++++++++++++++++++++++++++++---------------------\n object-file.h |  5 +++--\n odb.h         |  9 +++++++++\n 4 files changed, 52 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex a08c7554cb..3a64d28da8 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -474,8 +474,9 @@ static int too_many_loose_objects(int limit)\n \tint auto_threshold = DIV_ROUND_UP(limit, 256) * 256;\n \tunsigned long loose_count;\n \n-\tif (odb_source_loose_approximate_object_count(the_repository->objects->sources,\n-\t\t\t\t\t\t      &loose_count) < 0)\n+\tif (odb_source_loose_count_objects(the_repository->objects->sources,\n+\t\t\t\t\t   ODB_COUNT_OBJECTS_APPROXIMATE,\n+\t\t\t\t\t   &loose_count) < 0)\n \t\treturn 0;\n \n \treturn loose_count > auto_threshold;\ndiff --git a/object-file.c b/object-file.c\nindex da67e3c9ff..d35cec201f 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1868,40 +1868,56 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t\t     NULL, NULL, &data);\n }\n \n-int odb_source_loose_approximate_object_count(struct odb_source *source,\n-\t\t\t\t\t      unsigned long *out)\n+static int count_loose_object(const struct object_id *oid UNUSED,\n+\t\t\t      struct object_info *oi UNUSED,\n+\t\t\t      void *payload)\n+{\n+\tunsigned long *count = payload;\n+\t(*count)++;\n+\treturn 0;\n+}\n+\n+int odb_source_loose_count_objects(struct odb_source *source,\n+\t\t\t\t   enum odb_count_objects_flags flags,\n+\t\t\t\t   unsigned long *out)\n {\n \tconst unsigned hexsz = source->odb->repo->hash_algo->hexsz - 2;\n-\tunsigned long count = 0;\n-\tstruct dirent *ent;\n \tchar *path = NULL;\n \tDIR *dir = NULL;\n \tint ret;\n \n-\tpath = xstrfmt(\"%s/17\", source->path);\n+\tif (flags & ODB_COUNT_OBJECTS_APPROXIMATE) {\n+\t\tunsigned long count = 0;\n+\t\tstruct dirent *ent;\n \n-\tdir = opendir(path);\n-\tif (!dir) {\n-\t\tif (errno == ENOENT) {\n-\t\t\t*out = 0;\n-\t\t\tret = 0;\n+\t\tpath = xstrfmt(\"%s/17\", source->path);\n+\n+\t\tdir = opendir(path);\n+\t\tif (!dir) {\n+\t\t\tif (errno == ENOENT) {\n+\t\t\t\t*out = 0;\n+\t\t\t\tret = 0;\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\n+\t\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n \t\t\tgoto out;\n \t\t}\n \n-\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n-\t\tgoto out;\n-\t}\n+\t\twhile ((ent = readdir(dir)) != NULL) {\n+\t\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n+\t\t\t    ent->d_name[hexsz] != '\\0')\n+\t\t\t\tcontinue;\n+\t\t\tcount++;\n+\t\t}\n \n-\twhile ((ent = readdir(dir)) != NULL) {\n-\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n-\t\t    ent->d_name[hexsz] != '\\0')\n-\t\t\tcontinue;\n-\t\tcount++;\n+\t\t*out = count * 256;\n+\t\tret = 0;\n+\t} else {\n+\t\tret = odb_source_loose_for_each_object(source, NULL, count_loose_object,\n+\t\t\t\t\t\t       out, 0);\n \t}\n \n-\t*out = count * 256;\n-\tret = 0;\n-\n out:\n \tif (dir)\n \t\tclosedir(dir);\ndiff --git a/object-file.h b/object-file.h\nindex b870ea9fa8..f8d8805a18 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -149,8 +149,9 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n  *\n  * Returns 0 on success, a negative error code otherwise.\n  */\n-int odb_source_loose_approximate_object_count(struct odb_source *source,\n-\t\t\t\t\t      unsigned long *out);\n+int odb_source_loose_count_objects(struct odb_source *source,\n+\t\t\t\t   enum odb_count_objects_flags flags,\n+\t\t\t\t   unsigned long *out);\n \n /**\n  * format_object_header() is a thin wrapper around s xsnprintf() that\ndiff --git a/odb.h b/odb.h\nindex 7a583e3873..e6057477f6 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -500,6 +500,15 @@ int odb_for_each_object(struct object_database *odb,\n \t\t\tvoid *cb_data,\n \t\t\tunsigned flags);\n \n+enum odb_count_objects_flags {\n+\t/*\n+\t * Instead of providing an accurate count, allow the number of objects\n+\t * to be approximated. Details of how this approximation works are\n+\t * subject to the specific source's implementation.\n+\t */\n+\tODB_COUNT_OBJECTS_APPROXIMATE = (1 << 0),\n+};\n+\n enum {\n \t/*\n \t * By default, `odb_write_object()` does not actually write anything\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538465","messageId":"20260310-b4-pks-odb-source-count-objects-v1-5-109e07d425f4@pks.im","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im","subject":"[PATCH 5/6] odb/source: introduce generic object counting","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-10T15:18:25Z","receivedAt":"2026-03-10T15:18:39Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Introduce generic object counting on the object database source level\nwith a new backend-specific callback function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.c | 30 ++++++++++++++++++++++++++++++\n odb/source.h       | 27 +++++++++++++++++++++++++++\n packfile.c         |  4 ++--\n packfile.h         |  1 +\n 4 files changed, 60 insertions(+), 2 deletions(-)\n\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex 14cb9adeca..c08d8993e3 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -93,6 +93,35 @@ static int odb_source_files_for_each_object(struct odb_source *source,\n \treturn 0;\n }\n \n+static int odb_source_files_count_objects(struct odb_source *source,\n+\t\t\t\t\t  enum odb_count_objects_flags flags,\n+\t\t\t\t\t  unsigned long *out)\n+{\n+\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\tunsigned long count;\n+\tint ret;\n+\n+\tret = packfile_store_count_objects(files->packed, flags, &count);\n+\tif (ret < 0)\n+\t\tgoto out;\n+\n+\tif (!(flags & ODB_COUNT_OBJECTS_APPROXIMATE)) {\n+\t\tunsigned long loose_count;\n+\n+\t\tret = odb_source_loose_count_objects(source, flags, &loose_count);\n+\t\tif (ret < 0)\n+\t\t\tgoto out;\n+\n+\t\tcount += loose_count;\n+\t}\n+\n+\t*out = count;\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n static int odb_source_files_freshen_object(struct odb_source *source,\n \t\t\t\t\t   const struct object_id *oid)\n {\n@@ -220,6 +249,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \tfiles->base.read_object_info = odb_source_files_read_object_info;\n \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\n \tfiles->base.for_each_object = odb_source_files_for_each_object;\n+\tfiles->base.count_objects = odb_source_files_count_objects;\n \tfiles->base.freshen_object = odb_source_files_freshen_object;\n \tfiles->base.write_object = odb_source_files_write_object;\n \tfiles->base.write_object_stream = odb_source_files_write_object_stream;\ndiff --git a/odb/source.h b/odb/source.h\nindex a1fd9dd920..96c906e7a1 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -142,6 +142,21 @@ struct odb_source {\n \t\t\t       void *cb_data,\n \t\t\t       unsigned flags);\n \n+\t/*\n+\t * This callback is expected to count objects in the given object\n+\t * database source. The callback function does not have to guarantee\n+\t * that only unique objects are counted. The result shall be assigned\n+\t * to the `out` pointer.\n+\t *\n+\t * Accepts `enum odb_count_objects_flag` flags to alter the behaviour.\n+\t *\n+\t * The callback is expected to return 0 on success, or a negative error\n+\t * code otherwise.\n+\t */\n+\tint (*count_objects)(struct odb_source *source,\n+\t\t\t     enum odb_count_objects_flags flags,\n+\t\t\t     unsigned long *out);\n+\n \t/*\n \t * This callback is expected to freshen the given object so that its\n \t * last access time is set to the current time. This is used to ensure\n@@ -333,6 +348,18 @@ static inline int odb_source_for_each_object(struct odb_source *source,\n \treturn source->for_each_object(source, request, cb, cb_data, flags);\n }\n \n+/*\n+ * Count the number of objects in the given object database source.\n+ *\n+ * Returns 0 on success, a negative error code otherwise.\n+ */\n+static inline int odb_source_count_objects(struct odb_source *source,\n+\t\t\t\t\t   enum odb_count_objects_flags flags,\n+\t\t\t\t\t   unsigned long *out)\n+{\n+\treturn source->count_objects(source, flags, out);\n+}\n+\n /*\n  * Freshen an object in the object database by updating its timestamp.\n  * Returns 1 in case the object has been freshened, 0 in case the object does\ndiff --git a/packfile.c b/packfile.c\nindex 1ee5dd3da3..8ee462303a 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1102,6 +1102,7 @@ struct packfile_list_entry *packfile_store_get_packs(struct packfile_store *stor\n }\n \n int packfile_store_count_objects(struct packfile_store *store,\n+\t\t\t\t enum odb_count_objects_flags flags UNUSED,\n \t\t\t\t unsigned long *out)\n {\n \tstruct packfile_list_entry *e;\n@@ -1146,10 +1147,9 @@ unsigned long repo_approximate_object_count(struct repository *r)\n \n \t\todb_prepare_alternates(r->objects);\n \t\tfor (source = r->objects->sources; source; source = source->next) {\n-\t\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n \t\t\tunsigned long c;\n \n-\t\t\tif (!packfile_store_count_objects(files->packed, &c))\n+\t\t\tif (!odb_source_count_objects(source, ODB_COUNT_OBJECTS_APPROXIMATE, &c))\n \t\t\t\tcount += c;\n \t\t}\n \ndiff --git a/packfile.h b/packfile.h\nindex 1da8c729cb..74b6bc58c5 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -275,6 +275,7 @@ enum kept_pack_type {\n  * Return 0 on success, a negative error code otherwise.\n  */\n int packfile_store_count_objects(struct packfile_store *store,\n+\t\t\t\t enum odb_count_objects_flags flags,\n \t\t\t\t unsigned long *out);\n \n /*\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538466","messageId":"20260310-b4-pks-odb-source-count-objects-v1-6-109e07d425f4@pks.im","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im","subject":"[PATCH 6/6] odb: introduce generic object counting","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-10T15:18:26Z","receivedAt":"2026-03-10T15:18:41Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Similar to the preceding commit, introduce counting of objects on the\nobject database level, replacing the logic that we have in\n`repo_approximate_object_count()`.\n\nNote that the function knows to cache the object count. It's unclear\nwhether this cache is really required as we shouldn't have that many\ncases where we count objects repeatedly. But to be on the safe side the\ncaching mechanism is retained, with the only excepting being that we\nalso have to use the passed flags as caching key.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/gc.c   |  6 +++++-\n commit-graph.c |  3 ++-\n object-name.c  |  6 +++++-\n odb.c          | 37 ++++++++++++++++++++++++++++++++++++-\n odb.h          | 17 +++++++++++++++--\n packfile.c     | 27 ---------------------------\n packfile.h     |  6 ------\n 7 files changed, 63 insertions(+), 39 deletions(-)\n\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex 3a64d28da8..cb9ca89a97 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -574,9 +574,13 @@ static uint64_t total_ram(void)\n static uint64_t estimate_repack_memory(struct gc_config *cfg,\n \t\t\t\t       struct packed_git *pack)\n {\n-\tunsigned long nr_objects = repo_approximate_object_count(the_repository);\n+\tunsigned long nr_objects;\n \tsize_t os_cache, heap;\n \n+\tif (odb_count_objects(the_repository->objects,\n+\t\t\t      ODB_COUNT_OBJECTS_APPROXIMATE, &nr_objects) < 0)\n+\t\treturn 0;\n+\n \tif (!pack || !nr_objects)\n \t\treturn 0;\n \ndiff --git a/commit-graph.c b/commit-graph.c\nindex f8e24145a5..c030003330 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -2607,7 +2607,8 @@ int write_commit_graph(struct odb_source *source,\n \t\t\treplace = ctx.opts->split_flags & COMMIT_GRAPH_SPLIT_REPLACE;\n \t}\n \n-\tctx.approx_nr_objects = repo_approximate_object_count(r);\n+\tif (odb_count_objects(r->objects, ODB_COUNT_OBJECTS_APPROXIMATE, &ctx.approx_nr_objects) < 0)\n+\t\tctx.approx_nr_objects = 0;\n \n \tif (ctx.append && g) {\n \t\tfor (i = 0; i < g->num_commits; i++) {\ndiff --git a/object-name.c b/object-name.c\nindex 7b14c3bf9b..e5adec4c9d 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -837,7 +837,11 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \tconst unsigned hexsz = algo->hexsz;\n \n \tif (len < 0) {\n-\t\tunsigned long count = repo_approximate_object_count(r);\n+\t\tunsigned long count;\n+\n+\t\tif (odb_count_objects(r->objects, ODB_COUNT_OBJECTS_APPROXIMATE, &count) < 0)\n+\t\t\tcount = 0;\n+\n \t\t/*\n \t\t * Add one because the MSB only tells us the highest bit set,\n \t\t * not including the value of all the _other_ bits (so \"15\"\ndiff --git a/odb.c b/odb.c\nindex 84a31084d3..350e23f3c0 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -917,6 +917,41 @@ int odb_for_each_object(struct object_database *odb,\n \treturn 0;\n }\n \n+int odb_count_objects(struct object_database *odb,\n+\t\t      enum odb_count_objects_flags flags,\n+\t\t      unsigned long *out)\n+{\n+\tstruct odb_source *source;\n+\tunsigned long count = 0;\n+\tint ret;\n+\n+\tif (odb->object_count_valid && odb->object_count_flags == flags) {\n+\t\t*out = odb->object_count;\n+\t\treturn 0;\n+\t}\n+\n+\todb_prepare_alternates(odb);\n+\tfor (source = odb->sources; source; source = source->next) {\n+\t\tunsigned long c;\n+\n+\t\tret = odb_source_count_objects(source, flags, &c);\n+\t\tif (ret < 0)\n+\t\t\tgoto out;\n+\n+\t\tcount += c;\n+\t}\n+\n+\todb->object_count = count;\n+\todb->object_count_valid = 1;\n+\todb->object_count_flags = flags;\n+\n+\t*out = count;\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n void odb_assert_oid_type(struct object_database *odb,\n \t\t\t const struct object_id *oid, enum object_type expect)\n {\n@@ -1030,7 +1065,7 @@ void odb_reprepare(struct object_database *o)\n \tfor (source = o->sources; source; source = source->next)\n \t\todb_source_reprepare(source);\n \n-\to->approximate_object_count_valid = 0;\n+\to->object_count_valid = 0;\n \n \tobj_read_unlock();\n }\ndiff --git a/odb.h b/odb.h\nindex e6057477f6..7b004f1cf4 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -112,8 +112,9 @@ struct object_database {\n \t * These two fields are not meant for direct access. Use\n \t * repo_approximate_object_count() instead.\n \t */\n-\tunsigned long approximate_object_count;\n-\tunsigned approximate_object_count_valid : 1;\n+\tunsigned long object_count;\n+\tunsigned object_count_flags;\n+\tunsigned object_count_valid : 1;\n \n \t/*\n \t * Submodule source paths that will be added as additional sources to\n@@ -509,6 +510,18 @@ enum odb_count_objects_flags {\n \tODB_COUNT_OBJECTS_APPROXIMATE = (1 << 0),\n };\n \n+/*\n+ * Count the number of objects in the given object database. This object count\n+ * may double-count objects that are stored in multiple backends, or which are\n+ * stored multiple times in a single backend.\n+ *\n+ * Returns 0 on success, a negative error code otherwise. The number of objects\n+ * will be assigned to the `out` pointer on success.\n+ */\n+int odb_count_objects(struct object_database *odb,\n+\t\t      enum odb_count_objects_flags flags,\n+\t\t      unsigned long *out);\n+\n enum {\n \t/*\n \t * By default, `odb_write_object()` does not actually write anything\ndiff --git a/packfile.c b/packfile.c\nindex 8ee462303a..d4de9f3ffe 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1132,33 +1132,6 @@ int packfile_store_count_objects(struct packfile_store *store,\n \treturn ret;\n }\n \n-/*\n- * Give a fast, rough count of the number of objects in the repository. This\n- * ignores loose objects completely. If you have a lot of them, then either\n- * you should repack because your performance will be awful, or they are\n- * all unreachable objects about to be pruned, in which case they're not really\n- * interesting as a measure of repo size in the first place.\n- */\n-unsigned long repo_approximate_object_count(struct repository *r)\n-{\n-\tif (!r->objects->approximate_object_count_valid) {\n-\t\tstruct odb_source *source;\n-\t\tunsigned long count = 0;\n-\n-\t\todb_prepare_alternates(r->objects);\n-\t\tfor (source = r->objects->sources; source; source = source->next) {\n-\t\t\tunsigned long c;\n-\n-\t\t\tif (!odb_source_count_objects(source, ODB_COUNT_OBJECTS_APPROXIMATE, &c))\n-\t\t\t\tcount += c;\n-\t\t}\n-\n-\t\tr->objects->approximate_object_count = count;\n-\t\tr->objects->approximate_object_count_valid = 1;\n-\t}\n-\treturn r->objects->approximate_object_count;\n-}\n-\n unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \t\tunsigned long len, enum object_type *type, unsigned long *sizep)\n {\ndiff --git a/packfile.h b/packfile.h\nindex 74b6bc58c5..a16ec3950d 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -375,12 +375,6 @@ int packfile_store_for_each_object(struct packfile_store *store,\n #define PACKDIR_FILE_GARBAGE 4\n extern void (*report_garbage)(unsigned seen_bits, const char *path);\n \n-/*\n- * Give a rough count of objects in the repository. This sacrifices accuracy\n- * for speed.\n- */\n-unsigned long repo_approximate_object_count(struct repository *r);\n-\n void pack_report(struct repository *repo);\n \n /*\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538483","messageId":"xmqqjyvjvau1.fsf@gitster.g","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-3-109e07d425f4@pks.im","subject":"Re: [PATCH 3/6] object-file: extract logic to approximate object count","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-10T17:44:06Z","receivedAt":"2026-03-10T17:44:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>  static int too_many_loose_objects(int limit)\n>  {\n> ...\n> +\tint auto_threshold = DIV_ROUND_UP(limit, 256) * 256;\n> +\tunsigned long loose_count;\n> +\n> +\tif (odb_source_loose_approximate_object_count(the_repository->objects->sources,\n> +\t\t\t\t\t\t      &loose_count) < 0)\n>  \t\treturn 0;\n>  \n> -\tauto_threshold = DIV_ROUND_UP(limit, 256);\n> -\twhile ((ent = readdir(dir)) != NULL) {\n> -\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz_loose ||\n> -\t\t    ent->d_name[hexsz_loose] != '\\0')\n> -\t\t\tcontinue;\n> -\t\tif (++num_loose > auto_threshold) {\n> -\t\t\tneeded = 1;\n> -\t\t\tbreak;\n> -\t\t}\n> -\t}\n> -\tclosedir(dir);\n> -\treturn needed;\n> +\treturn loose_count > auto_threshold;\n>  }\n\nWe used to sample one shared directory and stopped when we know we\nhave more than auto_threshold, which is roughly 1/256 of the given\nlimit.  Now, we ask \"approximate\" function to count and then compare\nthe result with the same auto_threshold (i.e., 1/256 of the given\nlimit), which means we expect approximate function to count only\n1/256 of the total loose objects somehow?  Let's keep reading.\n\n>  static struct packed_git *find_base_packs(struct string_list *packs,\n> diff --git a/object-file.c b/object-file.c\n> index a3ff7f586c..da67e3c9ff 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1868,6 +1868,47 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n>  \t\t\t\t\t     NULL, NULL, &data);\n>  }\n>  \n> +int odb_source_loose_approximate_object_count(struct odb_source *source,\n> +\t\t\t\t\t      unsigned long *out)\n> +{\n> +\tconst unsigned hexsz = source->odb->repo->hash_algo->hexsz - 2;\n> +\tunsigned long count = 0;\n> +\tstruct dirent *ent;\n> +\tchar *path = NULL;\n> +\tDIR *dir = NULL;\n> +\tint ret;\n> +\n> +\tpath = xstrfmt(\"%s/17\", source->path);\n> +\n> +\tdir = opendir(path);\n> +\tif (!dir) {\n> +\t\tif (errno == ENOENT) {\n> +\t\t\t*out = 0;\n> +\t\t\tret = 0;\n> +\t\t\tgoto out;\n> +\t\t}\n> +\n> +\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n> +\t\tgoto out;\n> +\t}\n> +\n> +\twhile ((ent = readdir(dir)) != NULL) {\n> +\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n> +\t\t    ent->d_name[hexsz] != '\\0')\n> +\t\t\tcontinue;\n> +\t\tcount++;\n> +\t}\n\nThis counts one shared (\"17\" that is randomly picked) fully and then ...\n\n> +\t*out = count * 256;\n\n... estimate that the entire world would probably have 256 times as\nmany as the objects in that one shared.\n\nAh, my earlier read of the caller was confused.  auto_threshold used\nto be 1/256 of the limit, but now the number used is computed in a\nstrange arithmetic, \"DIV_ROUND_UP(limit,256) * 256\".  Not directly\nusing \"limit\" fooled me into thinking that it somehow kept using the\nsame 1/256 of the limit.\n\nSo we are answering \"do we have too many?\" question using roughly\nthe same criteria as before, not 1/256 off as I suspected earlier.\n\nThe old implementation exited early as soon as the threshold was\nhit.  While scanning a single shard directory is likely fast enough\nthat this may not matter in practice, it is a slight change in\nbehaviour. If a repository has an extremely large number of loose\nobjects (e.g. tens of thousands in shard 17), this will now count\nall of them instead of stopping at ~30 (if the limit set to around\n7000 objects).\n\nGiven that this is an \"auto\" GC check, the performance difference is\nprobably negligible, but I thought it worth pointing out.\n"},{"id":"538488","messageId":"xmqqfr67vahm.fsf@gitster.g","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-5-109e07d425f4@pks.im","subject":"Re: [PATCH 5/6] odb/source: introduce generic object counting","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-10T17:51:33Z","receivedAt":"2026-03-10T17:51:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> +static int odb_source_files_count_objects(struct odb_source *source,\n> +\t\t\t\t\t  enum odb_count_objects_flags flags,\n> +\t\t\t\t\t  unsigned long *out)\n> +{\n> +\tstruct odb_source_files *files = odb_source_files_downcast(source);\n> +\tunsigned long count;\n> +\tint ret;\n> +\n> +\tret = packfile_store_count_objects(files->packed, flags, &count);\n> +\tif (ret < 0)\n> +\t\tgoto out;\n> +\n> +\tif (!(flags & ODB_COUNT_OBJECTS_APPROXIMATE)) {\n> +\t\tunsigned long loose_count;\n> +\n> +\t\tret = odb_source_loose_count_objects(source, flags, &loose_count);\n> +\t\tif (ret < 0)\n> +\t\t\tgoto out;\n> +\n> +\t\tcount += loose_count;\n> +\t}\n> +\n> +\t*out = count;\n> +\tret = 0;\n> +\n> +out:\n> +\treturn ret;\n> +}\n\nThe design to assume that the majority of objects should be in the\npackfiles and the number of loose objects can be ignored when we are\ngetting approximation is inherited from the world before this\nseries, I think, which is a valid choice for this series to make.\n\nAs your \"get an approximate count of loose objects\" counts a single\nshared fully, instead of punting as soon as the limit is hit, we\ncould ask that function and add it in when the APPROXIMATE flag is\npassed, and get a bit more accurate number cheaply even when we are\napproximating.  I am not sure what the pros and cons of doing so\nmyself, but you may already have thought about it and rejected it,\nperhaps?\n\nThanks for a pleasant read.\nQueued.\n\n"},{"id":"538557","messageId":"abEPSZdahq50L3aF@pks.im","threadId":"65199","inReplyTo":"xmqqfr67vahm.fsf@gitster.g","subject":"Re: [PATCH 5/6] odb/source: introduce generic object counting","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-11T06:44:25Z","receivedAt":"2026-03-11T06:44:30Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Tue, Mar 10, 2026 at 10:51:33AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > +static int odb_source_files_count_objects(struct odb_source *source,\n> > +\t\t\t\t\t  enum odb_count_objects_flags flags,\n> > +\t\t\t\t\t  unsigned long *out)\n> > +{\n> > +\tstruct odb_source_files *files = odb_source_files_downcast(source);\n> > +\tunsigned long count;\n> > +\tint ret;\n> > +\n> > +\tret = packfile_store_count_objects(files->packed, flags, &count);\n> > +\tif (ret < 0)\n> > +\t\tgoto out;\n> > +\n> > +\tif (!(flags & ODB_COUNT_OBJECTS_APPROXIMATE)) {\n> > +\t\tunsigned long loose_count;\n> > +\n> > +\t\tret = odb_source_loose_count_objects(source, flags, &loose_count);\n> > +\t\tif (ret < 0)\n> > +\t\t\tgoto out;\n> > +\n> > +\t\tcount += loose_count;\n> > +\t}\n> > +\n> > +\t*out = count;\n> > +\tret = 0;\n> > +\n> > +out:\n> > +\treturn ret;\n> > +}\n> \n> The design to assume that the majority of objects should be in the\n> packfiles and the number of loose objects can be ignored when we are\n> getting approximation is inherited from the world before this\n> series, I think, which is a valid choice for this series to make.\n> \n> As your \"get an approximate count of loose objects\" counts a single\n> shared fully, instead of punting as soon as the limit is hit, we\n> could ask that function and add it in when the APPROXIMATE flag is\n> passed, and get a bit more accurate number cheaply even when we are\n> approximating.  I am not sure what the pros and cons of doing so\n> myself, but you may already have thought about it and rejected it,\n> perhaps?\n\nYeah, I was very torn on this myself, and I switched back and forth\nmultiple times. I don't really think there is a large downside if we\nstarted to also count loose objects here. The performance overhead\nshould be negligible, and it may arrive at a result that is closer to\nthe real world.\n\nBut the loose object counting approximation is somewhat vague overall\nwith the way we extrapolate the object count, and as you mentioned it\nmatches the old semantics to not include loose objects. So that's why I\ndecided against including it, so that we retain semantics.\n\nThat being said, with the current set of users it doesn't really matter\ntoo much which of both approaches we pick.\n\nPatrick\n"},{"id":"538594","messageId":"871phqmtcu.fsf@iotcl.com","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-2-109e07d425f4@pks.im","subject":"Re: [PATCH 2/6] packfile: extract logic to count number of objects","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-03-11T12:41:05Z","receivedAt":"2026-03-11T12:41:24Z","isPatch":true,"sender":{"key":"toon@iotcl.com","avatar":"https://avatars.githubusercontent.com/u/121621?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> In a subsequent commit we're about to introduce a new\n> `odb_source_count_objects()` function so that we can make the logic\n> pluggable. Prepare for this change by extracting the logic that we have\n> to count packed objects into a standalone function.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  packfile.c | 45 +++++++++++++++++++++++++++++++++++----------\n>  packfile.h |  9 +++++++++\n>  2 files changed, 44 insertions(+), 10 deletions(-)\n>\n> diff --git a/packfile.c b/packfile.c\n> index 215a23e42b..1ee5dd3da3 100644\n> --- a/packfile.c\n> +++ b/packfile.c\n> @@ -1101,6 +1101,36 @@ struct packfile_list_entry *packfile_store_get_packs(struct packfile_store *stor\n>  \treturn store->packs.head;\n>  }\n>  \n> +int packfile_store_count_objects(struct packfile_store *store,\n> +\t\t\t\t unsigned long *out)\n> +{\n> +\tstruct packfile_list_entry *e;\n> +\tstruct multi_pack_index *m;\n> +\tunsigned long count = 0;\n> +\tint ret;\n> +\n> +\tm = get_multi_pack_index(store->source);\n> +\tif (m)\n> +\t\tcount += m->num_objects + m->num_objects_in_base;\n\nTo make sure I understand correctly:\n\n`m->num_objects` indicates how many objects are in the current pack, and\n`m->num_objects_in_base` how many are in it's base (and that accumulates\nwhat's in the bases of the base?).\n\n> +\tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n> +\t\tif (e->pack->multi_pack_index)\n> +\t\t\tcontinue;\n\nBecause we added the count through the midx already, we skip any\npackfile that's included in the midx.\n\nBut some packfiles are not in the midx so we fall through for those.\n\nMakes sense.\n\n\n-- \nCheers,\nToon\n"},{"id":"538595","messageId":"87v7f2lei6.fsf@iotcl.com","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-3-109e07d425f4@pks.im","subject":"Re: [PATCH 3/6] object-file: extract logic to approximate object count","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-03-11T12:47:13Z","receivedAt":"2026-03-11T12:47:31Z","isPatch":true,"sender":{"key":"toon@iotcl.com","avatar":"https://avatars.githubusercontent.com/u/121621?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> In \"builtin/gc.c\" we have some logic that checks whether we need to\n> repack objects. This is done by counting the number of objects that we\n> have and checking whether it exceeds a certain threshold. We don't\n> really need an accurate object count though, which is why we only\n> open a single object diretcroy shard and then extrapolate from there.\n\ns/diretcroy/directory/\n\n>\n> Extract this logic into a new function that is owned by the loose object\n> database source. This is done to prepare for a subsequent change, where\n> we'll introduce object counting on the object database source level.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  builtin/gc.c  | 37 +++++++++----------------------------\n>  object-file.c | 41 +++++++++++++++++++++++++++++++++++++++++\n>  object-file.h | 13 +++++++++++++\n>  3 files changed, 63 insertions(+), 28 deletions(-)\n>\n> diff --git a/builtin/gc.c b/builtin/gc.c\n> index fb329c2cff..a08c7554cb 100644\n> --- a/builtin/gc.c\n> +++ b/builtin/gc.c\n> @@ -467,37 +467,18 @@ static int rerere_gc_condition(struct gc_config *cfg UNUSED)\n>  static int too_many_loose_objects(int limit)\n>  {\n>  \t/*\n> -\t * Quickly check if a \"gc\" is needed, by estimating how\n> -\t * many loose objects there are.  Because SHA-1 is evenly\n> -\t * distributed, we can check only one and get a reasonable\n> -\t * estimate.\n> +\t * This is weird, but stems from legacy behaviour: the GC auto\n> +\t * threshold was always essentially interpreted as if it was rounded up\n> +\t * to the next multiple 256 of, so we retain this behaviour for now.\n>  \t */\n> -\tDIR *dir;\n> -\tstruct dirent *ent;\n> -\tint auto_threshold;\n> -\tint num_loose = 0;\n> -\tint needed = 0;\n> -\tconst unsigned hexsz_loose = the_hash_algo->hexsz - 2;\n> -\tchar *path;\n> -\n> -\tpath = repo_git_path(the_repository, \"objects/17\");\n> -\tdir = opendir(path);\n> -\tfree(path);\n> -\tif (!dir)\n> +\tint auto_threshold = DIV_ROUND_UP(limit, 256) * 256;\n> +\tunsigned long loose_count;\n> +\n> +\tif (odb_source_loose_approximate_object_count(the_repository->objects->sources,\n> +\t\t\t\t\t\t      &loose_count) < 0)\n>  \t\treturn 0;\n>  \n> -\tauto_threshold = DIV_ROUND_UP(limit, 256);\n> -\twhile ((ent = readdir(dir)) != NULL) {\n> -\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz_loose ||\n> -\t\t    ent->d_name[hexsz_loose] != '\\0')\n> -\t\t\tcontinue;\n> -\t\tif (++num_loose > auto_threshold) {\n> -\t\t\tneeded = 1;\n> -\t\t\tbreak;\n> -\t\t}\n> -\t}\n> -\tclosedir(dir);\n> -\treturn needed;\n> +\treturn loose_count > auto_threshold;\n>  }\n>  \n>  static struct packed_git *find_base_packs(struct string_list *packs,\n> diff --git a/object-file.c b/object-file.c\n> index a3ff7f586c..da67e3c9ff 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1868,6 +1868,47 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n>  \t\t\t\t\t     NULL, NULL, &data);\n>  }\n>  \n> +int odb_source_loose_approximate_object_count(struct odb_source *source,\n> +\t\t\t\t\t      unsigned long *out)\n> +{\n> +\tconst unsigned hexsz = source->odb->repo->hash_algo->hexsz - 2;\n> +\tunsigned long count = 0;\n> +\tstruct dirent *ent;\n> +\tchar *path = NULL;\n> +\tDIR *dir = NULL;\n> +\tint ret;\n> +\n> +\tpath = xstrfmt(\"%s/17\", source->path);\n> +\n> +\tdir = opendir(path);\n> +\tif (!dir) {\n> +\t\tif (errno == ENOENT) {\n> +\t\t\t*out = 0;\n> +\t\t\tret = 0;\n> +\t\t\tgoto out;\n> +\t\t}\n> +\n> +\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n> +\t\tgoto out;\n> +\t}\n> +\n> +\twhile ((ent = readdir(dir)) != NULL) {\n> +\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n> +\t\t    ent->d_name[hexsz] != '\\0')\n> +\t\t\tcontinue;\n> +\t\tcount++;\n> +\t}\n> +\n> +\t*out = count * 256;\n\nThis makes the number way larger, but I don't think we need to worry\ngetting anywhere near ULONG_MAX, because I would expect to have Git\ncoming to a grind way before that happens (not to mention filesystems\nwould get unhappy about it too).\n\n-- \nCheers,\nToon\n"},{"id":"538601","messageId":"87pl5albfz.fsf@iotcl.com","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-4-109e07d425f4@pks.im","subject":"Re: [PATCH 4/6] object-file: generalize counting objects","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-03-11T13:53:20Z","receivedAt":"2026-03-11T13:53:32Z","isPatch":true,"sender":{"key":"toon@iotcl.com","avatar":"https://avatars.githubusercontent.com/u/121621?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Generalize the function introduced in the preceding commit to not only\n> be able to approximate the number of loose objects, but to also provide\n> an accurate count. The behaviour can be toggled via a new flag.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  builtin/gc.c  |  5 +++--\n>  object-file.c | 58 +++++++++++++++++++++++++++++++++++++---------------------\n>  object-file.h |  5 +++--\n>  odb.h         |  9 +++++++++\n>  4 files changed, 52 insertions(+), 25 deletions(-)\n>\n> diff --git a/builtin/gc.c b/builtin/gc.c\n> index a08c7554cb..3a64d28da8 100644\n> --- a/builtin/gc.c\n> +++ b/builtin/gc.c\n> @@ -474,8 +474,9 @@ static int too_many_loose_objects(int limit)\n>  \tint auto_threshold = DIV_ROUND_UP(limit, 256) * 256;\n>  \tunsigned long loose_count;\n>  \n> -\tif (odb_source_loose_approximate_object_count(the_repository->objects->sources,\n> -\t\t\t\t\t\t      &loose_count) < 0)\n> +\tif (odb_source_loose_count_objects(the_repository->objects->sources,\n> +\t\t\t\t\t   ODB_COUNT_OBJECTS_APPROXIMATE,\n> +\t\t\t\t\t   &loose_count) < 0)\n>  \t\treturn 0;\n>  \n>  \treturn loose_count > auto_threshold;\n> diff --git a/object-file.c b/object-file.c\n> index da67e3c9ff..d35cec201f 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -1868,40 +1868,56 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n>  \t\t\t\t\t     NULL, NULL, &data);\n>  }\n>  \n> -int odb_source_loose_approximate_object_count(struct odb_source *source,\n> -\t\t\t\t\t      unsigned long *out)\n> +static int count_loose_object(const struct object_id *oid UNUSED,\n> +\t\t\t      struct object_info *oi UNUSED,\n> +\t\t\t      void *payload)\n> +{\n> +\tunsigned long *count = payload;\n> +\t(*count)++;\n> +\treturn 0;\n> +}\n> +\n> +int odb_source_loose_count_objects(struct odb_source *source,\n> +\t\t\t\t   enum odb_count_objects_flags flags,\n> +\t\t\t\t   unsigned long *out)\n>  {\n>  \tconst unsigned hexsz = source->odb->repo->hash_algo->hexsz - 2;\n> -\tunsigned long count = 0;\n> -\tstruct dirent *ent;\n>  \tchar *path = NULL;\n>  \tDIR *dir = NULL;\n>  \tint ret;\n>  \n> -\tpath = xstrfmt(\"%s/17\", source->path);\n> +\tif (flags & ODB_COUNT_OBJECTS_APPROXIMATE) {\n> +\t\tunsigned long count = 0;\n> +\t\tstruct dirent *ent;\n>  \n> -\tdir = opendir(path);\n> -\tif (!dir) {\n> -\t\tif (errno == ENOENT) {\n> -\t\t\t*out = 0;\n> -\t\t\tret = 0;\n> +\t\tpath = xstrfmt(\"%s/17\", source->path);\n> +\n> +\t\tdir = opendir(path);\n> +\t\tif (!dir) {\n> +\t\t\tif (errno == ENOENT) {\n> +\t\t\t\t*out = 0;\n> +\t\t\t\tret = 0;\n> +\t\t\t\tgoto out;\n> +\t\t\t}\n> +\n> +\t\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n>  \t\t\tgoto out;\n>  \t\t}\n>  \n> -\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n> -\t\tgoto out;\n> -\t}\n> +\t\twhile ((ent = readdir(dir)) != NULL) {\n> +\t\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n> +\t\t\t    ent->d_name[hexsz] != '\\0')\n> +\t\t\t\tcontinue;\n> +\t\t\tcount++;\n> +\t\t}\n>  \n> -\twhile ((ent = readdir(dir)) != NULL) {\n> -\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n> -\t\t    ent->d_name[hexsz] != '\\0')\n> -\t\t\tcontinue;\n> -\t\tcount++;\n> +\t\t*out = count * 256;\n> +\t\tret = 0;\n> +\t} else {\n> +\t\tret = odb_source_loose_for_each_object(source, NULL, count_loose_object,\n> +\t\t\t\t\t\t       out, 0);\n\nIsn't `*out` uninitialized here? Should we add `*out = 0;` before this\nline?\n\n>  \t}\n>  \n> -\t*out = count * 256;\n> -\tret = 0;\n> -\n>  out:\n>  \tif (dir)\n>  \t\tclosedir(dir);\n> diff --git a/object-file.h b/object-file.h\n> index b870ea9fa8..f8d8805a18 100644\n> --- a/object-file.h\n> +++ b/object-file.h\n> @@ -149,8 +149,9 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n>   *\n>   * Returns 0 on success, a negative error code otherwise.\n>   */\n> -int odb_source_loose_approximate_object_count(struct odb_source *source,\n> -\t\t\t\t\t      unsigned long *out);\n> +int odb_source_loose_count_objects(struct odb_source *source,\n> +\t\t\t\t   enum odb_count_objects_flags flags,\n> +\t\t\t\t   unsigned long *out);\n>  \n>  /**\n>   * format_object_header() is a thin wrapper around s xsnprintf() that\n> diff --git a/odb.h b/odb.h\n> index 7a583e3873..e6057477f6 100644\n> --- a/odb.h\n> +++ b/odb.h\n> @@ -500,6 +500,15 @@ int odb_for_each_object(struct object_database *odb,\n>  \t\t\tvoid *cb_data,\n>  \t\t\tunsigned flags);\n>  \n> +enum odb_count_objects_flags {\n> +\t/*\n> +\t * Instead of providing an accurate count, allow the number of objects\n> +\t * to be approximated. Details of how this approximation works are\n> +\t * subject to the specific source's implementation.\n> +\t */\n> +\tODB_COUNT_OBJECTS_APPROXIMATE = (1 << 0),\n> +};\n> +\n>  enum {\n>  \t/*\n>  \t * By default, `odb_write_object()` does not actually write anything\n>\n> -- \n> 2.53.0.880.g73c4285caa.dirty\n>\n>\n\n-- \nCheers,\nToon\n"},{"id":"538603","messageId":"abF0ROcaUpYxdAQq@pks.im","threadId":"65199","inReplyTo":"871phqmtcu.fsf@iotcl.com","subject":"Re: [PATCH 2/6] packfile: extract logic to count number of objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-11T13:55:16Z","receivedAt":"2026-03-11T13:55:21Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Mar 11, 2026 at 01:41:05PM +0100, Toon Claes wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > In a subsequent commit we're about to introduce a new\n> > `odb_source_count_objects()` function so that we can make the logic\n> > pluggable. Prepare for this change by extracting the logic that we have\n> > to count packed objects into a standalone function.\n> >\n> > Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> > ---\n> >  packfile.c | 45 +++++++++++++++++++++++++++++++++++----------\n> >  packfile.h |  9 +++++++++\n> >  2 files changed, 44 insertions(+), 10 deletions(-)\n> >\n> > diff --git a/packfile.c b/packfile.c\n> > index 215a23e42b..1ee5dd3da3 100644\n> > --- a/packfile.c\n> > +++ b/packfile.c\n> > @@ -1101,6 +1101,36 @@ struct packfile_list_entry *packfile_store_get_packs(struct packfile_store *stor\n> >  \treturn store->packs.head;\n> >  }\n> >  \n> > +int packfile_store_count_objects(struct packfile_store *store,\n> > +\t\t\t\t unsigned long *out)\n> > +{\n> > +\tstruct packfile_list_entry *e;\n> > +\tstruct multi_pack_index *m;\n> > +\tunsigned long count = 0;\n> > +\tint ret;\n> > +\n> > +\tm = get_multi_pack_index(store->source);\n> > +\tif (m)\n> > +\t\tcount += m->num_objects + m->num_objects_in_base;\n> \n> To make sure I understand correctly:\n> \n> `m->num_objects` indicates how many objects are in the current pack, and\n> `m->num_objects_in_base` how many are in it's base (and that accumulates\n> what's in the bases of the base?).\n\nYup, that's exactly right, `num_objects_in_base` is basically recursive\nacross the bitmap layers.\n\n> > +\tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n> > +\t\tif (e->pack->multi_pack_index)\n> > +\t\t\tcontinue;\n> \n> Because we added the count through the midx already, we skip any\n> packfile that's included in the midx.\n> \n> But some packfiles are not in the midx so we fall through for those.\n> \n> Makes sense.\n\nCorrect.\n\nPatrick\n"},{"id":"538604","messageId":"abF084pB38G5Nyv6@pks.im","threadId":"65199","inReplyTo":"87v7f2lei6.fsf@iotcl.com","subject":"Re: [PATCH 3/6] object-file: extract logic to approximate object count","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-11T13:58:11Z","receivedAt":"2026-03-11T13:58:16Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Mar 11, 2026 at 01:47:13PM +0100, Toon Claes wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > In \"builtin/gc.c\" we have some logic that checks whether we need to\n> > repack objects. This is done by counting the number of objects that we\n> > have and checking whether it exceeds a certain threshold. We don't\n> > really need an accurate object count though, which is why we only\n> > open a single object diretcroy shard and then extrapolate from there.\n> \n> s/diretcroy/directory/\n\nThanks, fixed locally.\n\n> > diff --git a/object-file.c b/object-file.c\n> > index a3ff7f586c..da67e3c9ff 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1868,6 +1868,47 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n> >  \t\t\t\t\t     NULL, NULL, &data);\n> >  }\n> >  \n> > +int odb_source_loose_approximate_object_count(struct odb_source *source,\n> > +\t\t\t\t\t      unsigned long *out)\n> > +{\n> > +\tconst unsigned hexsz = source->odb->repo->hash_algo->hexsz - 2;\n> > +\tunsigned long count = 0;\n> > +\tstruct dirent *ent;\n> > +\tchar *path = NULL;\n> > +\tDIR *dir = NULL;\n> > +\tint ret;\n> > +\n> > +\tpath = xstrfmt(\"%s/17\", source->path);\n> > +\n> > +\tdir = opendir(path);\n> > +\tif (!dir) {\n> > +\t\tif (errno == ENOENT) {\n> > +\t\t\t*out = 0;\n> > +\t\t\tret = 0;\n> > +\t\t\tgoto out;\n> > +\t\t}\n> > +\n> > +\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n> > +\t\tgoto out;\n> > +\t}\n> > +\n> > +\twhile ((ent = readdir(dir)) != NULL) {\n> > +\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n> > +\t\t    ent->d_name[hexsz] != '\\0')\n> > +\t\t\tcontinue;\n> > +\t\tcount++;\n> > +\t}\n> > +\n> > +\t*out = count * 256;\n> \n> This makes the number way larger, but I don't think we need to worry\n> getting anywhere near ULONG_MAX, because I would expect to have Git\n> coming to a grind way before that happens (not to mention filesystems\n> would get unhappy about it too).\n\nYup. Even if `unsigned long` was 32 bits that would be >128 million\nloose objects in a single directory. I agree that this is probably going\nto make some things in Git unhappy. So we could have overflow checks\nhere, but I'm not sure it's worth it.\n\nThanks!\n\nPatrick\n"},{"id":"538605","messageId":"abF1opEays8LQYbr@pks.im","threadId":"65199","inReplyTo":"87pl5albfz.fsf@iotcl.com","subject":"Re: [PATCH 4/6] object-file: generalize counting objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-11T14:01:06Z","receivedAt":"2026-03-11T14:01:11Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Mar 11, 2026 at 02:53:20PM +0100, Toon Claes wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > diff --git a/object-file.c b/object-file.c\n> > index da67e3c9ff..d35cec201f 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -1868,40 +1868,56 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n> >  \t\t\t\t\t     NULL, NULL, &data);\n> >  }\n> >  \n> > -int odb_source_loose_approximate_object_count(struct odb_source *source,\n> > -\t\t\t\t\t      unsigned long *out)\n> > +static int count_loose_object(const struct object_id *oid UNUSED,\n> > +\t\t\t      struct object_info *oi UNUSED,\n> > +\t\t\t      void *payload)\n> > +{\n> > +\tunsigned long *count = payload;\n> > +\t(*count)++;\n> > +\treturn 0;\n> > +}\n> > +\n> > +int odb_source_loose_count_objects(struct odb_source *source,\n> > +\t\t\t\t   enum odb_count_objects_flags flags,\n> > +\t\t\t\t   unsigned long *out)\n> >  {\n> >  \tconst unsigned hexsz = source->odb->repo->hash_algo->hexsz - 2;\n> > -\tunsigned long count = 0;\n> > -\tstruct dirent *ent;\n> >  \tchar *path = NULL;\n> >  \tDIR *dir = NULL;\n> >  \tint ret;\n> >  \n> > -\tpath = xstrfmt(\"%s/17\", source->path);\n> > +\tif (flags & ODB_COUNT_OBJECTS_APPROXIMATE) {\n> > +\t\tunsigned long count = 0;\n> > +\t\tstruct dirent *ent;\n> >  \n> > -\tdir = opendir(path);\n> > -\tif (!dir) {\n> > -\t\tif (errno == ENOENT) {\n> > -\t\t\t*out = 0;\n> > -\t\t\tret = 0;\n> > +\t\tpath = xstrfmt(\"%s/17\", source->path);\n> > +\n> > +\t\tdir = opendir(path);\n> > +\t\tif (!dir) {\n> > +\t\t\tif (errno == ENOENT) {\n> > +\t\t\t\t*out = 0;\n> > +\t\t\t\tret = 0;\n> > +\t\t\t\tgoto out;\n> > +\t\t\t}\n> > +\n> > +\t\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n> >  \t\t\tgoto out;\n> >  \t\t}\n> >  \n> > -\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n> > -\t\tgoto out;\n> > -\t}\n> > +\t\twhile ((ent = readdir(dir)) != NULL) {\n> > +\t\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n> > +\t\t\t    ent->d_name[hexsz] != '\\0')\n> > +\t\t\t\tcontinue;\n> > +\t\t\tcount++;\n> > +\t\t}\n> >  \n> > -\twhile ((ent = readdir(dir)) != NULL) {\n> > -\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n> > -\t\t    ent->d_name[hexsz] != '\\0')\n> > -\t\t\tcontinue;\n> > -\t\tcount++;\n> > +\t\t*out = count * 256;\n> > +\t\tret = 0;\n> > +\t} else {\n> > +\t\tret = odb_source_loose_for_each_object(source, NULL, count_loose_object,\n> > +\t\t\t\t\t\t       out, 0);\n> \n> Isn't `*out` uninitialized here? Should we add `*out = 0;` before this\n> line?\n\nOh, indeed. Will fix, thanks!\n\nPatrick\n"},{"id":"538613","messageId":"87ldfyl86v.fsf@iotcl.com","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-5-109e07d425f4@pks.im","subject":"Re: [PATCH 5/6] odb/source: introduce generic object counting","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-03-11T15:03:36Z","receivedAt":"2026-03-11T15:03:48Z","isPatch":true,"sender":{"key":"toon@iotcl.com","avatar":"https://avatars.githubusercontent.com/u/121621?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Introduce generic object counting on the object database source level\n> with a new backend-specific callback function.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  odb/source-files.c | 30 ++++++++++++++++++++++++++++++\n>  odb/source.h       | 27 +++++++++++++++++++++++++++\n>  packfile.c         |  4 ++--\n>  packfile.h         |  1 +\n>  4 files changed, 60 insertions(+), 2 deletions(-)\n>\n> diff --git a/odb/source-files.c b/odb/source-files.c\n> index 14cb9adeca..c08d8993e3 100644\n> --- a/odb/source-files.c\n> +++ b/odb/source-files.c\n> @@ -93,6 +93,35 @@ static int odb_source_files_for_each_object(struct odb_source *source,\n>  \treturn 0;\n>  }\n>  \n> +static int odb_source_files_count_objects(struct odb_source *source,\n> +\t\t\t\t\t  enum odb_count_objects_flags flags,\n> +\t\t\t\t\t  unsigned long *out)\n> +{\n> +\tstruct odb_source_files *files = odb_source_files_downcast(source);\n\nI didn't read the other series this depends on, but good to see\nodb_source_files_downcast() BUGs when the source isn't a 'files'\nsource.\n\n> +\tunsigned long count;\n> +\tint ret;\n> +\n> +\tret = packfile_store_count_objects(files->packed, flags, &count);\n> +\tif (ret < 0)\n> +\t\tgoto out;\n> +\n> +\tif (!(flags & ODB_COUNT_OBJECTS_APPROXIMATE)) {\n> +\t\tunsigned long loose_count;\n> +\n> +\t\tret = odb_source_loose_count_objects(source, flags, &loose_count);\n> +\t\tif (ret < 0)\n> +\t\t\tgoto out;\n> +\n> +\t\tcount += loose_count;\n> +\t}\n> +\n> +\t*out = count;\n> +\tret = 0;\n> +\n> +out:\n> +\treturn ret;\n> +}\n> +\n>  static int odb_source_files_freshen_object(struct odb_source *source,\n>  \t\t\t\t\t   const struct object_id *oid)\n>  {\n> @@ -220,6 +249,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n>  \tfiles->base.read_object_info = odb_source_files_read_object_info;\n>  \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\n>  \tfiles->base.for_each_object = odb_source_files_for_each_object;\n> +\tfiles->base.count_objects = odb_source_files_count_objects;\n>  \tfiles->base.freshen_object = odb_source_files_freshen_object;\n>  \tfiles->base.write_object = odb_source_files_write_object;\n>  \tfiles->base.write_object_stream = odb_source_files_write_object_stream;\n> diff --git a/odb/source.h b/odb/source.h\n> index a1fd9dd920..96c906e7a1 100644\n> --- a/odb/source.h\n> +++ b/odb/source.h\n> @@ -142,6 +142,21 @@ struct odb_source {\n>  \t\t\t       void *cb_data,\n>  \t\t\t       unsigned flags);\n>  \n> +\t/*\n> +\t * This callback is expected to count objects in the given object\n> +\t * database source. The callback function does not have to guarantee\n> +\t * that only unique objects are counted. The result shall be assigned\n> +\t * to the `out` pointer.\n> +\t *\n> +\t * Accepts `enum odb_count_objects_flag` flags to alter the behaviour.\n> +\t *\n> +\t * The callback is expected to return 0 on success, or a negative error\n> +\t * code otherwise.\n> +\t */\n> +\tint (*count_objects)(struct odb_source *source,\n> +\t\t\t     enum odb_count_objects_flags flags,\n> +\t\t\t     unsigned long *out);\n> +\n>  \t/*\n>  \t * This callback is expected to freshen the given object so that its\n>  \t * last access time is set to the current time. This is used to ensure\n> @@ -333,6 +348,18 @@ static inline int odb_source_for_each_object(struct odb_source *source,\n>  \treturn source->for_each_object(source, request, cb, cb_data, flags);\n>  }\n>  \n> +/*\n> + * Count the number of objects in the given object database source.\n> + *\n> + * Returns 0 on success, a negative error code otherwise.\n> + */\n> +static inline int odb_source_count_objects(struct odb_source *source,\n> +\t\t\t\t\t   enum odb_count_objects_flags flags,\n> +\t\t\t\t\t   unsigned long *out)\n> +{\n> +\treturn source->count_objects(source, flags, out);\n> +}\n> +\n>  /*\n>   * Freshen an object in the object database by updating its timestamp.\n>   * Returns 1 in case the object has been freshened, 0 in case the object does\n> diff --git a/packfile.c b/packfile.c\n> index 1ee5dd3da3..8ee462303a 100644\n> --- a/packfile.c\n> +++ b/packfile.c\n> @@ -1102,6 +1102,7 @@ struct packfile_list_entry *packfile_store_get_packs(struct packfile_store *stor\n>  }\n>  \n>  int packfile_store_count_objects(struct packfile_store *store,\n> +\t\t\t\t enum odb_count_objects_flags flags UNUSED,\n>  \t\t\t\t unsigned long *out)\n>  {\n>  \tstruct packfile_list_entry *e;\n> @@ -1146,10 +1147,9 @@ unsigned long repo_approximate_object_count(struct repository *r)\n>  \n>  \t\todb_prepare_alternates(r->objects);\n>  \t\tfor (source = r->objects->sources; source; source = source->next) {\n> -\t\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n>  \t\t\tunsigned long c;\n>  \n> -\t\t\tif (!packfile_store_count_objects(files->packed, &c))\n> +\t\t\tif (!odb_source_count_objects(source, ODB_COUNT_OBJECTS_APPROXIMATE, &c))\n>  \t\t\t\tcount += c;\n>  \t\t}\n>  \n> diff --git a/packfile.h b/packfile.h\n> index 1da8c729cb..74b6bc58c5 100644\n> --- a/packfile.h\n> +++ b/packfile.h\n> @@ -275,6 +275,7 @@ enum kept_pack_type {\n>   * Return 0 on success, a negative error code otherwise.\n>   */\n>  int packfile_store_count_objects(struct packfile_store *store,\n> +\t\t\t\t enum odb_count_objects_flags flags,\n>  \t\t\t\t unsigned long *out);\n>  \n>  /*\n>\n> -- \n> 2.53.0.880.g73c4285caa.dirty\n>\n\nOkay.\n\n-- \nCheers,\nToon\n"},{"id":"538619","messageId":"87fr66l6xh.fsf@iotcl.com","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-6-109e07d425f4@pks.im","subject":"Re: [PATCH 6/6] odb: introduce generic object counting","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-03-11T15:30:50Z","receivedAt":"2026-03-11T15:31:04Z","isPatch":true,"sender":{"key":"toon@iotcl.com","avatar":"https://avatars.githubusercontent.com/u/121621?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Similar to the preceding commit, introduce counting of objects on the\n> object database level, replacing the logic that we have in\n> `repo_approximate_object_count()`.\n>\n> Note that the function knows to cache the object count. It's unclear\n> whether this cache is really required as we shouldn't have that many\n> cases where we count objects repeatedly. But to be on the safe side the\n> caching mechanism is retained, with the only excepting being that we\n> also have to use the passed flags as caching key.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  builtin/gc.c   |  6 +++++-\n>  commit-graph.c |  3 ++-\n>  object-name.c  |  6 +++++-\n>  odb.c          | 37 ++++++++++++++++++++++++++++++++++++-\n>  odb.h          | 17 +++++++++++++++--\n>  packfile.c     | 27 ---------------------------\n>  packfile.h     |  6 ------\n>  7 files changed, 63 insertions(+), 39 deletions(-)\n>\n> diff --git a/builtin/gc.c b/builtin/gc.c\n> index 3a64d28da8..cb9ca89a97 100644\n> --- a/builtin/gc.c\n> +++ b/builtin/gc.c\n> @@ -574,9 +574,13 @@ static uint64_t total_ram(void)\n>  static uint64_t estimate_repack_memory(struct gc_config *cfg,\n>  \t\t\t\t       struct packed_git *pack)\n>  {\n> -\tunsigned long nr_objects = repo_approximate_object_count(the_repository);\n> +\tunsigned long nr_objects;\n>  \tsize_t os_cache, heap;\n>  \n> +\tif (odb_count_objects(the_repository->objects,\n> +\t\t\t      ODB_COUNT_OBJECTS_APPROXIMATE, &nr_objects) < 0)\n> +\t\treturn 0;\n> +\n>  \tif (!pack || !nr_objects)\n>  \t\treturn 0;\n>  \n> diff --git a/commit-graph.c b/commit-graph.c\n> index f8e24145a5..c030003330 100644\n> --- a/commit-graph.c\n> +++ b/commit-graph.c\n> @@ -2607,7 +2607,8 @@ int write_commit_graph(struct odb_source *source,\n>  \t\t\treplace = ctx.opts->split_flags & COMMIT_GRAPH_SPLIT_REPLACE;\n>  \t}\n>  \n> -\tctx.approx_nr_objects = repo_approximate_object_count(r);\n> +\tif (odb_count_objects(r->objects, ODB_COUNT_OBJECTS_APPROXIMATE, &ctx.approx_nr_objects) < 0)\n> +\t\tctx.approx_nr_objects = 0;\n>  \n>  \tif (ctx.append && g) {\n>  \t\tfor (i = 0; i < g->num_commits; i++) {\n> diff --git a/object-name.c b/object-name.c\n> index 7b14c3bf9b..e5adec4c9d 100644\n> --- a/object-name.c\n> +++ b/object-name.c\n> @@ -837,7 +837,11 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n>  \tconst unsigned hexsz = algo->hexsz;\n>  \n>  \tif (len < 0) {\n> -\t\tunsigned long count = repo_approximate_object_count(r);\n> +\t\tunsigned long count;\n> +\n> +\t\tif (odb_count_objects(r->objects, ODB_COUNT_OBJECTS_APPROXIMATE, &count) < 0)\n> +\t\t\tcount = 0;\n> +\n>  \t\t/*\n>  \t\t * Add one because the MSB only tells us the highest bit set,\n>  \t\t * not including the value of all the _other_ bits (so \"15\"\n> diff --git a/odb.c b/odb.c\n> index 84a31084d3..350e23f3c0 100644\n> --- a/odb.c\n> +++ b/odb.c\n> @@ -917,6 +917,41 @@ int odb_for_each_object(struct object_database *odb,\n>  \treturn 0;\n>  }\n>  \n> +int odb_count_objects(struct object_database *odb,\n> +\t\t      enum odb_count_objects_flags flags,\n> +\t\t      unsigned long *out)\n> +{\n> +\tstruct odb_source *source;\n> +\tunsigned long count = 0;\n> +\tint ret;\n> +\n> +\tif (odb->object_count_valid && odb->object_count_flags == flags) {\n> +\t\t*out = odb->object_count;\n> +\t\treturn 0;\n> +\t}\n> +\n> +\todb_prepare_alternates(odb);\n> +\tfor (source = odb->sources; source; source = source->next) {\n> +\t\tunsigned long c;\n> +\n> +\t\tret = odb_source_count_objects(source, flags, &c);\n> +\t\tif (ret < 0)\n> +\t\t\tgoto out;\n> +\n> +\t\tcount += c;\n> +\t}\n> +\n> +\todb->object_count = count;\n> +\todb->object_count_valid = 1;\n> +\todb->object_count_flags = flags;\n> +\n> +\t*out = count;\n> +\tret = 0;\n> +\n> +out:\n> +\treturn ret;\n> +}\n> +\n>  void odb_assert_oid_type(struct object_database *odb,\n>  \t\t\t const struct object_id *oid, enum object_type expect)\n>  {\n> @@ -1030,7 +1065,7 @@ void odb_reprepare(struct object_database *o)\n>  \tfor (source = o->sources; source; source = source->next)\n>  \t\todb_source_reprepare(source);\n>  \n> -\to->approximate_object_count_valid = 0;\n> +\to->object_count_valid = 0;\n>  \n>  \tobj_read_unlock();\n>  }\n> diff --git a/odb.h b/odb.h\n> index e6057477f6..7b004f1cf4 100644\n> --- a/odb.h\n> +++ b/odb.h\n> @@ -112,8 +112,9 @@ struct object_database {\n>  \t * These two fields are not meant for direct access. Use\n>  \t * repo_approximate_object_count() instead.\n\nThis comment needs updating now.\n\nOtherwise no comments here.\n\n-- \nCheers,\nToon\n"},{"id":"538723","messageId":"abJj1THZrxtued7v@pks.im","threadId":"65199","inReplyTo":"87fr66l6xh.fsf@iotcl.com","subject":"Re: [PATCH 6/6] odb: introduce generic object counting","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-12T06:57:25Z","receivedAt":"2026-03-12T06:57:31Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Mar 11, 2026 at 04:30:50PM +0100, Toon Claes wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > diff --git a/odb.h b/odb.h\n> > index e6057477f6..7b004f1cf4 100644\n> > --- a/odb.h\n> > +++ b/odb.h\n> > @@ -112,8 +112,9 @@ struct object_database {\n> >  \t * These two fields are not meant for direct access. Use\n> >  \t * repo_approximate_object_count() instead.\n> \n> This comment needs updating now.\n> \n> Otherwise no comments here.\n\nGood catch, will fix. Thanks!\n\nPatrick\n"},{"id":"538728","messageId":"20260312-b4-pks-odb-source-count-objects-v2-0-5914f69256bf@pks.im","threadId":"65199","inReplyTo":"20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im","subject":"[PATCH v2 0/6] odb: introduce generic object counting","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-12T08:42:55Z","receivedAt":"2026-03-12T08:43:04Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Hi,\n\nthis small patch series introduces generic object counting for pluggable\nobject databases. The series is built on top of d181b9354c (The 13th\nbatch, 2026-03-09) with ps/odb-sources at d6fc6fe6f8 (odb/source: make\n`begin_transaction()` function pluggable, 2026-03-05) merged into it.\n\nChanges in v2:\n  - Properly initialize `out` pointer when counting loose objects.\n  - Fix a stale comment.\n  - Fix a commit message type.\n  - Link to v1: https://lore.kernel.org/r/20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (6):\n      odb: stop including \"odb/source.h\"\n      packfile: extract logic to count number of objects\n      object-file: extract logic to approximate object count\n      object-file: generalize counting objects\n      odb/source: introduce generic object counting\n      odb: introduce generic object counting\n\n builtin/gc.c                | 44 +++++++++----------------\n builtin/multi-pack-index.c  |  1 +\n builtin/submodule--helper.c |  1 +\n commit-graph.c              |  3 +-\n object-file.c               | 58 +++++++++++++++++++++++++++++++++\n object-file.h               | 14 ++++++++\n object-name.c               |  6 +++-\n odb.c                       | 37 ++++++++++++++++++++-\n odb.h                       | 78 +++++++++++++++++++++++++++++++++++++++++---\n odb/source-files.c          | 30 +++++++++++++++++\n odb/source.h                | 79 ++++++++++++++++-----------------------------\n odb/streaming.c             |  1 +\n packfile.c                  | 48 +++++++++++++--------------\n packfile.h                  | 16 +++++----\n repository.c                |  1 +\n submodule-config.c          |  1 +\n tmp-objdir.c                |  1 +\n 17 files changed, 301 insertions(+), 118 deletions(-)\n\nRange-diff versus v1:\n\n1:  3d5f8733d5 = 1:  1639bb1725 odb: stop including \"odb/source.h\"\n2:  2fe42618f4 = 2:  056b6f3ae3 packfile: extract logic to count number of objects\n3:  4cbf727523 ! 3:  069c908771 object-file: extract logic to approximate object count\n    @@ Commit message\n         repack objects. This is done by counting the number of objects that we\n         have and checking whether it exceeds a certain threshold. We don't\n         really need an accurate object count though, which is why we only\n    -    open a single object diretcroy shard and then extrapolate from there.\n    +    open a single object directory shard and then extrapolate from there.\n     \n         Extract this logic into a new function that is owned by the loose object\n         database source. This is done to prepare for a subsequent change, where\n4:  c1a527877a ! 4:  6f293f1352 object-file: generalize counting objects\n    @@ object-file.c: int odb_source_loose_for_each_object(struct odb_source *source,\n     +\t\t*out = count * 256;\n     +\t\tret = 0;\n     +\t} else {\n    ++\t\t*out = 0;\n     +\t\tret = odb_source_loose_for_each_object(source, NULL, count_loose_object,\n     +\t\t\t\t\t\t       out, 0);\n      \t}\n5:  c12d4ec401 = 5:  0bda4a3d01 odb/source: introduce generic object counting\n6:  91307f205d ! 6:  89164dad76 odb: introduce generic object counting\n    @@ odb.c: void odb_reprepare(struct object_database *o)\n     \n      ## odb.h ##\n     @@ odb.h: struct object_database {\n    + \t/*\n    + \t * A fast, rough count of the number of objects in the repository.\n      \t * These two fields are not meant for direct access. Use\n    - \t * repo_approximate_object_count() instead.\n    +-\t * repo_approximate_object_count() instead.\n    ++\t * odb_count_objects() instead.\n      \t */\n     -\tunsigned long approximate_object_count;\n     -\tunsigned approximate_object_count_valid : 1;\n\n---\nbase-commit: 2247f478a898a7f8f8322cc51bdeb1cc773d8f4a\nchange-id: 20260224-b4-pks-odb-source-count-objects-479fe682cf6f\n\n"},{"id":"538729","messageId":"20260312-b4-pks-odb-source-count-objects-v2-1-5914f69256bf@pks.im","threadId":"65199","inReplyTo":"20260312-b4-pks-odb-source-count-objects-v2-0-5914f69256bf@pks.im","subject":"[PATCH v2 1/6] odb: stop including \"odb/source.h\"","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-12T08:42:56Z","receivedAt":"2026-03-12T08:43:07Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"The \"odb.h\" header currently includes the \"odb/source.h\" file. This is\nsomewhat roundabout though: most callers shouldn't have to care about\nthe `struct odb_source`, but should rather use the ODB-level functions.\nFurthermore, it means that a couple of definitions have to live on the\nsource level even though they should be part of the generic interface.\n\nReverse the relation between \"odb/source.h\" and \"odb.h\" and move the\nenums and typedefs that relate to the generic interfaces back into\n\"odb.h\". Add the necessary includes to all files that rely on the\ntransitive include.\n\nSuggested-by: Justin Tobler <jltobler@gmail.com>\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/multi-pack-index.c  |  1 +\n builtin/submodule--helper.c |  1 +\n odb.h                       | 50 ++++++++++++++++++++++++++++++++++++++++++-\n odb/source.h                | 52 +--------------------------------------------\n odb/streaming.c             |  1 +\n repository.c                |  1 +\n submodule-config.c          |  1 +\n tmp-objdir.c                |  1 +\n 8 files changed, 56 insertions(+), 52 deletions(-)\n\ndiff --git a/builtin/multi-pack-index.c b/builtin/multi-pack-index.c\nindex 5f364aa816..3fcb207f1a 100644\n--- a/builtin/multi-pack-index.c\n+++ b/builtin/multi-pack-index.c\n@@ -9,6 +9,7 @@\n #include \"strbuf.h\"\n #include \"trace2.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"replace-object.h\"\n #include \"repository.h\"\n \ndiff --git a/builtin/submodule--helper.c b/builtin/submodule--helper.c\nindex 143f7cb3cc..4957487536 100644\n--- a/builtin/submodule--helper.c\n+++ b/builtin/submodule--helper.c\n@@ -29,6 +29,7 @@\n #include \"object-file.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"advice.h\"\n #include \"branch.h\"\n #include \"list-objects-filter-options.h\"\ndiff --git a/odb.h b/odb.h\nindex 86e0365c24..7a583e3873 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -3,7 +3,6 @@\n \n #include \"hashmap.h\"\n #include \"object.h\"\n-#include \"odb/source.h\"\n #include \"oidset.h\"\n #include \"oidmap.h\"\n #include \"string-list.h\"\n@@ -12,6 +11,7 @@\n struct oidmap;\n struct oidtree;\n struct strbuf;\n+struct strvec;\n struct repository;\n struct multi_pack_index;\n \n@@ -339,6 +339,42 @@ struct object_info {\n  */\n #define OBJECT_INFO_INIT { 0 }\n \n+/* Flags that can be passed to `odb_read_object_info_extended()`. */\n+enum object_info_flags {\n+\t/* Invoke lookup_replace_object() on the given hash. */\n+\tOBJECT_INFO_LOOKUP_REPLACE = (1 << 0),\n+\n+\t/* Do not reprepare object sources when the first lookup has failed. */\n+\tOBJECT_INFO_QUICK = (1 << 1),\n+\n+\t/*\n+\t * Do not attempt to fetch the object if missing (even if fetch_is_missing is\n+\t * nonzero).\n+\t */\n+\tOBJECT_INFO_SKIP_FETCH_OBJECT = (1 << 2),\n+\n+\t/* Die if object corruption (not just an object being missing) was detected. */\n+\tOBJECT_INFO_DIE_IF_CORRUPT = (1 << 3),\n+\n+\t/*\n+\t * We have already tried reading the object, but it couldn't be found\n+\t * via any of the attached sources, and are now doing a second read.\n+\t * This second read asks the individual sources to also evaluate\n+\t * whether any on-disk state may have changed that may have caused the\n+\t * object to appear.\n+\t *\n+\t * This flag is for internal use, only. The second read only occurs\n+\t * when `OBJECT_INFO_QUICK` was not passed.\n+\t */\n+\tOBJECT_INFO_SECOND_READ = (1 << 4),\n+\n+\t/*\n+\t * This is meant for bulk prefetching of missing blobs in a partial\n+\t * clone. Implies OBJECT_INFO_SKIP_FETCH_OBJECT and OBJECT_INFO_QUICK.\n+\t */\n+\tOBJECT_INFO_FOR_PREFETCH = (OBJECT_INFO_SKIP_FETCH_OBJECT | OBJECT_INFO_QUICK),\n+};\n+\n /*\n  * Read object info from the object database and populate the `object_info`\n  * structure. Returns 0 on success, a negative error code otherwise.\n@@ -432,6 +468,18 @@ enum odb_for_each_object_flags {\n \tODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS = (1<<4),\n };\n \n+/*\n+ * A callback function that can be used to iterate through objects. If given,\n+ * the optional `oi` parameter will be populated the same as if you would call\n+ * `odb_read_object_info()`.\n+ *\n+ * Returning a non-zero error code will cause iteration to abort. The error\n+ * code will be propagated.\n+ */\n+typedef int (*odb_for_each_object_cb)(const struct object_id *oid,\n+\t\t\t\t      struct object_info *oi,\n+\t\t\t\t      void *cb_data);\n+\n /*\n  * Iterate through all objects contained in the object database. Note that\n  * objects may be iterated over multiple times in case they are either stored\ndiff --git a/odb/source.h b/odb/source.h\nindex caac558149..a1fd9dd920 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -2,6 +2,7 @@\n #define ODB_SOURCE_H\n \n #include \"object.h\"\n+#include \"odb.h\"\n \n enum odb_source_type {\n \t/*\n@@ -14,61 +15,10 @@ enum odb_source_type {\n \tODB_SOURCE_FILES,\n };\n \n-/* Flags that can be passed to `odb_read_object_info_extended()`. */\n-enum object_info_flags {\n-\t/* Invoke lookup_replace_object() on the given hash. */\n-\tOBJECT_INFO_LOOKUP_REPLACE = (1 << 0),\n-\n-\t/* Do not reprepare object sources when the first lookup has failed. */\n-\tOBJECT_INFO_QUICK = (1 << 1),\n-\n-\t/*\n-\t * Do not attempt to fetch the object if missing (even if fetch_is_missing is\n-\t * nonzero).\n-\t */\n-\tOBJECT_INFO_SKIP_FETCH_OBJECT = (1 << 2),\n-\n-\t/* Die if object corruption (not just an object being missing) was detected. */\n-\tOBJECT_INFO_DIE_IF_CORRUPT = (1 << 3),\n-\n-\t/*\n-\t * We have already tried reading the object, but it couldn't be found\n-\t * via any of the attached sources, and are now doing a second read.\n-\t * This second read asks the individual sources to also evaluate\n-\t * whether any on-disk state may have changed that may have caused the\n-\t * object to appear.\n-\t *\n-\t * This flag is for internal use, only. The second read only occurs\n-\t * when `OBJECT_INFO_QUICK` was not passed.\n-\t */\n-\tOBJECT_INFO_SECOND_READ = (1 << 4),\n-\n-\t/*\n-\t * This is meant for bulk prefetching of missing blobs in a partial\n-\t * clone. Implies OBJECT_INFO_SKIP_FETCH_OBJECT and OBJECT_INFO_QUICK.\n-\t */\n-\tOBJECT_INFO_FOR_PREFETCH = (OBJECT_INFO_SKIP_FETCH_OBJECT | OBJECT_INFO_QUICK),\n-};\n-\n struct object_id;\n-struct object_info;\n struct odb_read_stream;\n-struct odb_transaction;\n-struct odb_write_stream;\n struct strvec;\n \n-/*\n- * A callback function that can be used to iterate through objects. If given,\n- * the optional `oi` parameter will be populated the same as if you would call\n- * `odb_read_object_info()`.\n- *\n- * Returning a non-zero error code will cause iteration to abort. The error\n- * code will be propagated.\n- */\n-typedef int (*odb_for_each_object_cb)(const struct object_id *oid,\n-\t\t\t\t      struct object_info *oi,\n-\t\t\t\t      void *cb_data);\n-\n /*\n  * The source is the part of the object database that stores the actual\n  * objects. It thus encapsulates the logic to read and write the specific\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex a4355cd245..5927a12954 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -7,6 +7,7 @@\n #include \"environment.h\"\n #include \"repository.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"odb/streaming.h\"\n #include \"replace-object.h\"\n \ndiff --git a/repository.c b/repository.c\nindex e7fa42c14f..05c26bdbc3 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -2,6 +2,7 @@\n #include \"abspath.h\"\n #include \"repository.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"config.h\"\n #include \"object.h\"\n #include \"lockfile.h\"\ndiff --git a/submodule-config.c b/submodule-config.c\nindex 1f19fe2077..72a46b7a54 100644\n--- a/submodule-config.c\n+++ b/submodule-config.c\n@@ -14,6 +14,7 @@\n #include \"strbuf.h\"\n #include \"object-name.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"parse-options.h\"\n #include \"thread-utils.h\"\n #include \"tree-walk.h\"\ndiff --git a/tmp-objdir.c b/tmp-objdir.c\nindex e436eed07e..d199d39e7c 100644\n--- a/tmp-objdir.c\n+++ b/tmp-objdir.c\n@@ -11,6 +11,7 @@\n #include \"strvec.h\"\n #include \"quote.h\"\n #include \"odb.h\"\n+#include \"odb/source.h\"\n #include \"repository.h\"\n \n struct tmp_objdir {\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538730","messageId":"20260312-b4-pks-odb-source-count-objects-v2-2-5914f69256bf@pks.im","threadId":"65199","inReplyTo":"20260312-b4-pks-odb-source-count-objects-v2-0-5914f69256bf@pks.im","subject":"[PATCH v2 2/6] packfile: extract logic to count number of objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-12T08:42:57Z","receivedAt":"2026-03-12T08:43:09Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"In a subsequent commit we're about to introduce a new\n`odb_source_count_objects()` function so that we can make the logic\npluggable. Prepare for this change by extracting the logic that we have\nto count packed objects into a standalone function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n packfile.c | 45 +++++++++++++++++++++++++++++++++++----------\n packfile.h |  9 +++++++++\n 2 files changed, 44 insertions(+), 10 deletions(-)\n\ndiff --git a/packfile.c b/packfile.c\nindex 215a23e42b..1ee5dd3da3 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1101,6 +1101,36 @@ struct packfile_list_entry *packfile_store_get_packs(struct packfile_store *stor\n \treturn store->packs.head;\n }\n \n+int packfile_store_count_objects(struct packfile_store *store,\n+\t\t\t\t unsigned long *out)\n+{\n+\tstruct packfile_list_entry *e;\n+\tstruct multi_pack_index *m;\n+\tunsigned long count = 0;\n+\tint ret;\n+\n+\tm = get_multi_pack_index(store->source);\n+\tif (m)\n+\t\tcount += m->num_objects + m->num_objects_in_base;\n+\n+\tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n+\t\tif (e->pack->multi_pack_index)\n+\t\t\tcontinue;\n+\t\tif (open_pack_index(e->pack)) {\n+\t\t\tret = -1;\n+\t\t\tgoto out;\n+\t\t}\n+\n+\t\tcount += e->pack->num_objects;\n+\t}\n+\n+\t*out = count;\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n /*\n  * Give a fast, rough count of the number of objects in the repository. This\n  * ignores loose objects completely. If you have a lot of them, then either\n@@ -1113,21 +1143,16 @@ unsigned long repo_approximate_object_count(struct repository *r)\n \tif (!r->objects->approximate_object_count_valid) {\n \t\tstruct odb_source *source;\n \t\tunsigned long count = 0;\n-\t\tstruct packed_git *p;\n \n \t\todb_prepare_alternates(r->objects);\n-\n \t\tfor (source = r->objects->sources; source; source = source->next) {\n-\t\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n-\t\t\tif (m)\n-\t\t\t\tcount += m->num_objects + m->num_objects_in_base;\n-\t\t}\n+\t\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\t\t\tunsigned long c;\n \n-\t\trepo_for_each_pack(r, p) {\n-\t\t\tif (p->multi_pack_index || open_pack_index(p))\n-\t\t\t\tcontinue;\n-\t\t\tcount += p->num_objects;\n+\t\t\tif (!packfile_store_count_objects(files->packed, &c))\n+\t\t\t\tcount += c;\n \t\t}\n+\n \t\tr->objects->approximate_object_count = count;\n \t\tr->objects->approximate_object_count_valid = 1;\n \t}\ndiff --git a/packfile.h b/packfile.h\nindex 8b04a258a7..1da8c729cb 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -268,6 +268,15 @@ enum kept_pack_type {\n \tKEPT_PACK_IN_CORE = (1 << 1),\n };\n \n+/*\n+ * Count the number objects contained in the given packfile store. If\n+ * successful, the number of objects will be written to the `out` pointer.\n+ *\n+ * Return 0 on success, a negative error code otherwise.\n+ */\n+int packfile_store_count_objects(struct packfile_store *store,\n+\t\t\t\t unsigned long *out);\n+\n /*\n  * Retrieve the cache of kept packs from the given packfile store. Accepts a\n  * combination of `kept_pack_type` flags. The cache is computed on demand and\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538731","messageId":"20260312-b4-pks-odb-source-count-objects-v2-3-5914f69256bf@pks.im","threadId":"65199","inReplyTo":"20260312-b4-pks-odb-source-count-objects-v2-0-5914f69256bf@pks.im","subject":"[PATCH v2 3/6] object-file: extract logic to approximate object count","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-12T08:42:58Z","receivedAt":"2026-03-12T08:43:12Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"In \"builtin/gc.c\" we have some logic that checks whether we need to\nrepack objects. This is done by counting the number of objects that we\nhave and checking whether it exceeds a certain threshold. We don't\nreally need an accurate object count though, which is why we only\nopen a single object directory shard and then extrapolate from there.\n\nExtract this logic into a new function that is owned by the loose object\ndatabase source. This is done to prepare for a subsequent change, where\nwe'll introduce object counting on the object database source level.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/gc.c  | 37 +++++++++----------------------------\n object-file.c | 41 +++++++++++++++++++++++++++++++++++++++++\n object-file.h | 13 +++++++++++++\n 3 files changed, 63 insertions(+), 28 deletions(-)\n\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex fb329c2cff..a08c7554cb 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -467,37 +467,18 @@ static int rerere_gc_condition(struct gc_config *cfg UNUSED)\n static int too_many_loose_objects(int limit)\n {\n \t/*\n-\t * Quickly check if a \"gc\" is needed, by estimating how\n-\t * many loose objects there are.  Because SHA-1 is evenly\n-\t * distributed, we can check only one and get a reasonable\n-\t * estimate.\n+\t * This is weird, but stems from legacy behaviour: the GC auto\n+\t * threshold was always essentially interpreted as if it was rounded up\n+\t * to the next multiple 256 of, so we retain this behaviour for now.\n \t */\n-\tDIR *dir;\n-\tstruct dirent *ent;\n-\tint auto_threshold;\n-\tint num_loose = 0;\n-\tint needed = 0;\n-\tconst unsigned hexsz_loose = the_hash_algo->hexsz - 2;\n-\tchar *path;\n-\n-\tpath = repo_git_path(the_repository, \"objects/17\");\n-\tdir = opendir(path);\n-\tfree(path);\n-\tif (!dir)\n+\tint auto_threshold = DIV_ROUND_UP(limit, 256) * 256;\n+\tunsigned long loose_count;\n+\n+\tif (odb_source_loose_approximate_object_count(the_repository->objects->sources,\n+\t\t\t\t\t\t      &loose_count) < 0)\n \t\treturn 0;\n \n-\tauto_threshold = DIV_ROUND_UP(limit, 256);\n-\twhile ((ent = readdir(dir)) != NULL) {\n-\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz_loose ||\n-\t\t    ent->d_name[hexsz_loose] != '\\0')\n-\t\t\tcontinue;\n-\t\tif (++num_loose > auto_threshold) {\n-\t\t\tneeded = 1;\n-\t\t\tbreak;\n-\t\t}\n-\t}\n-\tclosedir(dir);\n-\treturn needed;\n+\treturn loose_count > auto_threshold;\n }\n \n static struct packed_git *find_base_packs(struct string_list *packs,\ndiff --git a/object-file.c b/object-file.c\nindex a3ff7f586c..da67e3c9ff 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1868,6 +1868,47 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t\t     NULL, NULL, &data);\n }\n \n+int odb_source_loose_approximate_object_count(struct odb_source *source,\n+\t\t\t\t\t      unsigned long *out)\n+{\n+\tconst unsigned hexsz = source->odb->repo->hash_algo->hexsz - 2;\n+\tunsigned long count = 0;\n+\tstruct dirent *ent;\n+\tchar *path = NULL;\n+\tDIR *dir = NULL;\n+\tint ret;\n+\n+\tpath = xstrfmt(\"%s/17\", source->path);\n+\n+\tdir = opendir(path);\n+\tif (!dir) {\n+\t\tif (errno == ENOENT) {\n+\t\t\t*out = 0;\n+\t\t\tret = 0;\n+\t\t\tgoto out;\n+\t\t}\n+\n+\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n+\t\tgoto out;\n+\t}\n+\n+\twhile ((ent = readdir(dir)) != NULL) {\n+\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n+\t\t    ent->d_name[hexsz] != '\\0')\n+\t\t\tcontinue;\n+\t\tcount++;\n+\t}\n+\n+\t*out = count * 256;\n+\tret = 0;\n+\n+out:\n+\tif (dir)\n+\t\tclosedir(dir);\n+\tfree(path);\n+\treturn ret;\n+}\n+\n static int append_loose_object(const struct object_id *oid,\n \t\t\t       const char *path UNUSED,\n \t\t\t       void *data)\ndiff --git a/object-file.h b/object-file.h\nindex ff6da65296..b870ea9fa8 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -139,6 +139,19 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t     void *cb_data,\n \t\t\t\t     unsigned flags);\n \n+/*\n+ * Count the number of loose objects in this source.\n+ *\n+ * The object count is approximated by opening a single sharding directory for\n+ * loose objects and scanning its contents. The result is then extrapolated by\n+ * 256. This should generally work as a reasonable estimate given that the\n+ * object hash is supposed to be indistinguishable from random.\n+ *\n+ * Returns 0 on success, a negative error code otherwise.\n+ */\n+int odb_source_loose_approximate_object_count(struct odb_source *source,\n+\t\t\t\t\t      unsigned long *out);\n+\n /**\n  * format_object_header() is a thin wrapper around s xsnprintf() that\n  * writes the initial \"<type> <obj-len>\" part of the loose object\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538732","messageId":"20260312-b4-pks-odb-source-count-objects-v2-4-5914f69256bf@pks.im","threadId":"65199","inReplyTo":"20260312-b4-pks-odb-source-count-objects-v2-0-5914f69256bf@pks.im","subject":"[PATCH v2 4/6] object-file: generalize counting objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-12T08:42:59Z","receivedAt":"2026-03-12T08:43:15Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Generalize the function introduced in the preceding commit to not only\nbe able to approximate the number of loose objects, but to also provide\nan accurate count. The behaviour can be toggled via a new flag.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/gc.c  |  5 +++--\n object-file.c | 59 ++++++++++++++++++++++++++++++++++++++---------------------\n object-file.h |  5 +++--\n odb.h         |  9 +++++++++\n 4 files changed, 53 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex a08c7554cb..3a64d28da8 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -474,8 +474,9 @@ static int too_many_loose_objects(int limit)\n \tint auto_threshold = DIV_ROUND_UP(limit, 256) * 256;\n \tunsigned long loose_count;\n \n-\tif (odb_source_loose_approximate_object_count(the_repository->objects->sources,\n-\t\t\t\t\t\t      &loose_count) < 0)\n+\tif (odb_source_loose_count_objects(the_repository->objects->sources,\n+\t\t\t\t\t   ODB_COUNT_OBJECTS_APPROXIMATE,\n+\t\t\t\t\t   &loose_count) < 0)\n \t\treturn 0;\n \n \treturn loose_count > auto_threshold;\ndiff --git a/object-file.c b/object-file.c\nindex da67e3c9ff..569ce6eaed 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1868,40 +1868,57 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t\t     NULL, NULL, &data);\n }\n \n-int odb_source_loose_approximate_object_count(struct odb_source *source,\n-\t\t\t\t\t      unsigned long *out)\n+static int count_loose_object(const struct object_id *oid UNUSED,\n+\t\t\t      struct object_info *oi UNUSED,\n+\t\t\t      void *payload)\n+{\n+\tunsigned long *count = payload;\n+\t(*count)++;\n+\treturn 0;\n+}\n+\n+int odb_source_loose_count_objects(struct odb_source *source,\n+\t\t\t\t   enum odb_count_objects_flags flags,\n+\t\t\t\t   unsigned long *out)\n {\n \tconst unsigned hexsz = source->odb->repo->hash_algo->hexsz - 2;\n-\tunsigned long count = 0;\n-\tstruct dirent *ent;\n \tchar *path = NULL;\n \tDIR *dir = NULL;\n \tint ret;\n \n-\tpath = xstrfmt(\"%s/17\", source->path);\n+\tif (flags & ODB_COUNT_OBJECTS_APPROXIMATE) {\n+\t\tunsigned long count = 0;\n+\t\tstruct dirent *ent;\n \n-\tdir = opendir(path);\n-\tif (!dir) {\n-\t\tif (errno == ENOENT) {\n-\t\t\t*out = 0;\n-\t\t\tret = 0;\n+\t\tpath = xstrfmt(\"%s/17\", source->path);\n+\n+\t\tdir = opendir(path);\n+\t\tif (!dir) {\n+\t\t\tif (errno == ENOENT) {\n+\t\t\t\t*out = 0;\n+\t\t\t\tret = 0;\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\n+\t\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n \t\t\tgoto out;\n \t\t}\n \n-\t\tret = error_errno(\"cannot open object shard '%s'\", path);\n-\t\tgoto out;\n-\t}\n+\t\twhile ((ent = readdir(dir)) != NULL) {\n+\t\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n+\t\t\t    ent->d_name[hexsz] != '\\0')\n+\t\t\t\tcontinue;\n+\t\t\tcount++;\n+\t\t}\n \n-\twhile ((ent = readdir(dir)) != NULL) {\n-\t\tif (strspn(ent->d_name, \"0123456789abcdef\") != hexsz ||\n-\t\t    ent->d_name[hexsz] != '\\0')\n-\t\t\tcontinue;\n-\t\tcount++;\n+\t\t*out = count * 256;\n+\t\tret = 0;\n+\t} else {\n+\t\t*out = 0;\n+\t\tret = odb_source_loose_for_each_object(source, NULL, count_loose_object,\n+\t\t\t\t\t\t       out, 0);\n \t}\n \n-\t*out = count * 256;\n-\tret = 0;\n-\n out:\n \tif (dir)\n \t\tclosedir(dir);\ndiff --git a/object-file.h b/object-file.h\nindex b870ea9fa8..f8d8805a18 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -149,8 +149,9 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n  *\n  * Returns 0 on success, a negative error code otherwise.\n  */\n-int odb_source_loose_approximate_object_count(struct odb_source *source,\n-\t\t\t\t\t      unsigned long *out);\n+int odb_source_loose_count_objects(struct odb_source *source,\n+\t\t\t\t   enum odb_count_objects_flags flags,\n+\t\t\t\t   unsigned long *out);\n \n /**\n  * format_object_header() is a thin wrapper around s xsnprintf() that\ndiff --git a/odb.h b/odb.h\nindex 7a583e3873..e6057477f6 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -500,6 +500,15 @@ int odb_for_each_object(struct object_database *odb,\n \t\t\tvoid *cb_data,\n \t\t\tunsigned flags);\n \n+enum odb_count_objects_flags {\n+\t/*\n+\t * Instead of providing an accurate count, allow the number of objects\n+\t * to be approximated. Details of how this approximation works are\n+\t * subject to the specific source's implementation.\n+\t */\n+\tODB_COUNT_OBJECTS_APPROXIMATE = (1 << 0),\n+};\n+\n enum {\n \t/*\n \t * By default, `odb_write_object()` does not actually write anything\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538733","messageId":"20260312-b4-pks-odb-source-count-objects-v2-5-5914f69256bf@pks.im","threadId":"65199","inReplyTo":"20260312-b4-pks-odb-source-count-objects-v2-0-5914f69256bf@pks.im","subject":"[PATCH v2 5/6] odb/source: introduce generic object counting","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-12T08:43:00Z","receivedAt":"2026-03-12T08:43:17Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Introduce generic object counting on the object database source level\nwith a new backend-specific callback function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.c | 30 ++++++++++++++++++++++++++++++\n odb/source.h       | 27 +++++++++++++++++++++++++++\n packfile.c         |  4 ++--\n packfile.h         |  1 +\n 4 files changed, 60 insertions(+), 2 deletions(-)\n\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex 14cb9adeca..c08d8993e3 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -93,6 +93,35 @@ static int odb_source_files_for_each_object(struct odb_source *source,\n \treturn 0;\n }\n \n+static int odb_source_files_count_objects(struct odb_source *source,\n+\t\t\t\t\t  enum odb_count_objects_flags flags,\n+\t\t\t\t\t  unsigned long *out)\n+{\n+\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\tunsigned long count;\n+\tint ret;\n+\n+\tret = packfile_store_count_objects(files->packed, flags, &count);\n+\tif (ret < 0)\n+\t\tgoto out;\n+\n+\tif (!(flags & ODB_COUNT_OBJECTS_APPROXIMATE)) {\n+\t\tunsigned long loose_count;\n+\n+\t\tret = odb_source_loose_count_objects(source, flags, &loose_count);\n+\t\tif (ret < 0)\n+\t\t\tgoto out;\n+\n+\t\tcount += loose_count;\n+\t}\n+\n+\t*out = count;\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n static int odb_source_files_freshen_object(struct odb_source *source,\n \t\t\t\t\t   const struct object_id *oid)\n {\n@@ -220,6 +249,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \tfiles->base.read_object_info = odb_source_files_read_object_info;\n \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\n \tfiles->base.for_each_object = odb_source_files_for_each_object;\n+\tfiles->base.count_objects = odb_source_files_count_objects;\n \tfiles->base.freshen_object = odb_source_files_freshen_object;\n \tfiles->base.write_object = odb_source_files_write_object;\n \tfiles->base.write_object_stream = odb_source_files_write_object_stream;\ndiff --git a/odb/source.h b/odb/source.h\nindex a1fd9dd920..96c906e7a1 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -142,6 +142,21 @@ struct odb_source {\n \t\t\t       void *cb_data,\n \t\t\t       unsigned flags);\n \n+\t/*\n+\t * This callback is expected to count objects in the given object\n+\t * database source. The callback function does not have to guarantee\n+\t * that only unique objects are counted. The result shall be assigned\n+\t * to the `out` pointer.\n+\t *\n+\t * Accepts `enum odb_count_objects_flag` flags to alter the behaviour.\n+\t *\n+\t * The callback is expected to return 0 on success, or a negative error\n+\t * code otherwise.\n+\t */\n+\tint (*count_objects)(struct odb_source *source,\n+\t\t\t     enum odb_count_objects_flags flags,\n+\t\t\t     unsigned long *out);\n+\n \t/*\n \t * This callback is expected to freshen the given object so that its\n \t * last access time is set to the current time. This is used to ensure\n@@ -333,6 +348,18 @@ static inline int odb_source_for_each_object(struct odb_source *source,\n \treturn source->for_each_object(source, request, cb, cb_data, flags);\n }\n \n+/*\n+ * Count the number of objects in the given object database source.\n+ *\n+ * Returns 0 on success, a negative error code otherwise.\n+ */\n+static inline int odb_source_count_objects(struct odb_source *source,\n+\t\t\t\t\t   enum odb_count_objects_flags flags,\n+\t\t\t\t\t   unsigned long *out)\n+{\n+\treturn source->count_objects(source, flags, out);\n+}\n+\n /*\n  * Freshen an object in the object database by updating its timestamp.\n  * Returns 1 in case the object has been freshened, 0 in case the object does\ndiff --git a/packfile.c b/packfile.c\nindex 1ee5dd3da3..8ee462303a 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1102,6 +1102,7 @@ struct packfile_list_entry *packfile_store_get_packs(struct packfile_store *stor\n }\n \n int packfile_store_count_objects(struct packfile_store *store,\n+\t\t\t\t enum odb_count_objects_flags flags UNUSED,\n \t\t\t\t unsigned long *out)\n {\n \tstruct packfile_list_entry *e;\n@@ -1146,10 +1147,9 @@ unsigned long repo_approximate_object_count(struct repository *r)\n \n \t\todb_prepare_alternates(r->objects);\n \t\tfor (source = r->objects->sources; source; source = source->next) {\n-\t\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n \t\t\tunsigned long c;\n \n-\t\t\tif (!packfile_store_count_objects(files->packed, &c))\n+\t\t\tif (!odb_source_count_objects(source, ODB_COUNT_OBJECTS_APPROXIMATE, &c))\n \t\t\t\tcount += c;\n \t\t}\n \ndiff --git a/packfile.h b/packfile.h\nindex 1da8c729cb..74b6bc58c5 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -275,6 +275,7 @@ enum kept_pack_type {\n  * Return 0 on success, a negative error code otherwise.\n  */\n int packfile_store_count_objects(struct packfile_store *store,\n+\t\t\t\t enum odb_count_objects_flags flags,\n \t\t\t\t unsigned long *out);\n \n /*\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538734","messageId":"20260312-b4-pks-odb-source-count-objects-v2-6-5914f69256bf@pks.im","threadId":"65199","inReplyTo":"20260312-b4-pks-odb-source-count-objects-v2-0-5914f69256bf@pks.im","subject":"[PATCH v2 6/6] odb: introduce generic object counting","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-12T08:43:01Z","receivedAt":"2026-03-12T08:43:20Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Similar to the preceding commit, introduce counting of objects on the\nobject database level, replacing the logic that we have in\n`repo_approximate_object_count()`.\n\nNote that the function knows to cache the object count. It's unclear\nwhether this cache is really required as we shouldn't have that many\ncases where we count objects repeatedly. But to be on the safe side the\ncaching mechanism is retained, with the only excepting being that we\nalso have to use the passed flags as caching key.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/gc.c   |  6 +++++-\n commit-graph.c |  3 ++-\n object-name.c  |  6 +++++-\n odb.c          | 37 ++++++++++++++++++++++++++++++++++++-\n odb.h          | 19 ++++++++++++++++---\n packfile.c     | 27 ---------------------------\n packfile.h     |  6 ------\n 7 files changed, 64 insertions(+), 40 deletions(-)\n\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex 3a64d28da8..cb9ca89a97 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -574,9 +574,13 @@ static uint64_t total_ram(void)\n static uint64_t estimate_repack_memory(struct gc_config *cfg,\n \t\t\t\t       struct packed_git *pack)\n {\n-\tunsigned long nr_objects = repo_approximate_object_count(the_repository);\n+\tunsigned long nr_objects;\n \tsize_t os_cache, heap;\n \n+\tif (odb_count_objects(the_repository->objects,\n+\t\t\t      ODB_COUNT_OBJECTS_APPROXIMATE, &nr_objects) < 0)\n+\t\treturn 0;\n+\n \tif (!pack || !nr_objects)\n \t\treturn 0;\n \ndiff --git a/commit-graph.c b/commit-graph.c\nindex f8e24145a5..c030003330 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -2607,7 +2607,8 @@ int write_commit_graph(struct odb_source *source,\n \t\t\treplace = ctx.opts->split_flags & COMMIT_GRAPH_SPLIT_REPLACE;\n \t}\n \n-\tctx.approx_nr_objects = repo_approximate_object_count(r);\n+\tif (odb_count_objects(r->objects, ODB_COUNT_OBJECTS_APPROXIMATE, &ctx.approx_nr_objects) < 0)\n+\t\tctx.approx_nr_objects = 0;\n \n \tif (ctx.append && g) {\n \t\tfor (i = 0; i < g->num_commits; i++) {\ndiff --git a/object-name.c b/object-name.c\nindex 7b14c3bf9b..e5adec4c9d 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -837,7 +837,11 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \tconst unsigned hexsz = algo->hexsz;\n \n \tif (len < 0) {\n-\t\tunsigned long count = repo_approximate_object_count(r);\n+\t\tunsigned long count;\n+\n+\t\tif (odb_count_objects(r->objects, ODB_COUNT_OBJECTS_APPROXIMATE, &count) < 0)\n+\t\t\tcount = 0;\n+\n \t\t/*\n \t\t * Add one because the MSB only tells us the highest bit set,\n \t\t * not including the value of all the _other_ bits (so \"15\"\ndiff --git a/odb.c b/odb.c\nindex 84a31084d3..350e23f3c0 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -917,6 +917,41 @@ int odb_for_each_object(struct object_database *odb,\n \treturn 0;\n }\n \n+int odb_count_objects(struct object_database *odb,\n+\t\t      enum odb_count_objects_flags flags,\n+\t\t      unsigned long *out)\n+{\n+\tstruct odb_source *source;\n+\tunsigned long count = 0;\n+\tint ret;\n+\n+\tif (odb->object_count_valid && odb->object_count_flags == flags) {\n+\t\t*out = odb->object_count;\n+\t\treturn 0;\n+\t}\n+\n+\todb_prepare_alternates(odb);\n+\tfor (source = odb->sources; source; source = source->next) {\n+\t\tunsigned long c;\n+\n+\t\tret = odb_source_count_objects(source, flags, &c);\n+\t\tif (ret < 0)\n+\t\t\tgoto out;\n+\n+\t\tcount += c;\n+\t}\n+\n+\todb->object_count = count;\n+\todb->object_count_valid = 1;\n+\todb->object_count_flags = flags;\n+\n+\t*out = count;\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n void odb_assert_oid_type(struct object_database *odb,\n \t\t\t const struct object_id *oid, enum object_type expect)\n {\n@@ -1030,7 +1065,7 @@ void odb_reprepare(struct object_database *o)\n \tfor (source = o->sources; source; source = source->next)\n \t\todb_source_reprepare(source);\n \n-\to->approximate_object_count_valid = 0;\n+\to->object_count_valid = 0;\n \n \tobj_read_unlock();\n }\ndiff --git a/odb.h b/odb.h\nindex e6057477f6..9aee260105 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -110,10 +110,11 @@ struct object_database {\n \t/*\n \t * A fast, rough count of the number of objects in the repository.\n \t * These two fields are not meant for direct access. Use\n-\t * repo_approximate_object_count() instead.\n+\t * odb_count_objects() instead.\n \t */\n-\tunsigned long approximate_object_count;\n-\tunsigned approximate_object_count_valid : 1;\n+\tunsigned long object_count;\n+\tunsigned object_count_flags;\n+\tunsigned object_count_valid : 1;\n \n \t/*\n \t * Submodule source paths that will be added as additional sources to\n@@ -509,6 +510,18 @@ enum odb_count_objects_flags {\n \tODB_COUNT_OBJECTS_APPROXIMATE = (1 << 0),\n };\n \n+/*\n+ * Count the number of objects in the given object database. This object count\n+ * may double-count objects that are stored in multiple backends, or which are\n+ * stored multiple times in a single backend.\n+ *\n+ * Returns 0 on success, a negative error code otherwise. The number of objects\n+ * will be assigned to the `out` pointer on success.\n+ */\n+int odb_count_objects(struct object_database *odb,\n+\t\t      enum odb_count_objects_flags flags,\n+\t\t      unsigned long *out);\n+\n enum {\n \t/*\n \t * By default, `odb_write_object()` does not actually write anything\ndiff --git a/packfile.c b/packfile.c\nindex 8ee462303a..d4de9f3ffe 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1132,33 +1132,6 @@ int packfile_store_count_objects(struct packfile_store *store,\n \treturn ret;\n }\n \n-/*\n- * Give a fast, rough count of the number of objects in the repository. This\n- * ignores loose objects completely. If you have a lot of them, then either\n- * you should repack because your performance will be awful, or they are\n- * all unreachable objects about to be pruned, in which case they're not really\n- * interesting as a measure of repo size in the first place.\n- */\n-unsigned long repo_approximate_object_count(struct repository *r)\n-{\n-\tif (!r->objects->approximate_object_count_valid) {\n-\t\tstruct odb_source *source;\n-\t\tunsigned long count = 0;\n-\n-\t\todb_prepare_alternates(r->objects);\n-\t\tfor (source = r->objects->sources; source; source = source->next) {\n-\t\t\tunsigned long c;\n-\n-\t\t\tif (!odb_source_count_objects(source, ODB_COUNT_OBJECTS_APPROXIMATE, &c))\n-\t\t\t\tcount += c;\n-\t\t}\n-\n-\t\tr->objects->approximate_object_count = count;\n-\t\tr->objects->approximate_object_count_valid = 1;\n-\t}\n-\treturn r->objects->approximate_object_count;\n-}\n-\n unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \t\tunsigned long len, enum object_type *type, unsigned long *sizep)\n {\ndiff --git a/packfile.h b/packfile.h\nindex 74b6bc58c5..a16ec3950d 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -375,12 +375,6 @@ int packfile_store_for_each_object(struct packfile_store *store,\n #define PACKDIR_FILE_GARBAGE 4\n extern void (*report_garbage)(unsigned seen_bits, const char *path);\n \n-/*\n- * Give a rough count of objects in the repository. This sacrifices accuracy\n- * for speed.\n- */\n-unsigned long repo_approximate_object_count(struct repository *r);\n-\n void pack_report(struct repository *repo);\n \n /*\n\n-- \n2.53.0.880.g73c4285caa.dirty\n\n"},{"id":"538878","messageId":"874imkt08r.fsf@iotcl.com","threadId":"65199","inReplyTo":"20260312-b4-pks-odb-source-count-objects-v2-0-5914f69256bf@pks.im","subject":"Re: [PATCH v2 0/6] odb: introduce generic object counting","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-03-13T11:52:36Z","receivedAt":"2026-03-13T11:52:47Z","isPatch":true,"sender":{"key":"toon@iotcl.com","avatar":"https://avatars.githubusercontent.com/u/121621?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Hi,\n>\n> this small patch series introduces generic object counting for pluggable\n> object databases. The series is built on top of d181b9354c (The 13th\n> batch, 2026-03-09) with ps/odb-sources at d6fc6fe6f8 (odb/source: make\n> `begin_transaction()` function pluggable, 2026-03-05) merged into it.\n>\n> Changes in v2:\n>   - Properly initialize `out` pointer when counting loose objects.\n>   - Fix a stale comment.\n>   - Fix a commit message type.\n>   - Link to v1: https://lore.kernel.org/r/20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im\n\nLooking at the range-diff, all my concerns are addressed. Thanks!\n\n-- \nCheers,\nToon\n"}]}