From: Patrick Steinhardt Date: Thu, 12 Mar 2026 08:42:55 GMT Subject: [PATCH v2 0/6] odb: introduce generic object counting Message-ID: <20260312-b4-pks-odb-source-count-objects-v2-0-5914f69256bf@pks.im> In-Reply-To: <20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im> Hi, this small patch series introduces generic object counting for pluggable object databases. The series is built on top of d181b9354c (The 13th batch, 2026-03-09) with ps/odb-sources at d6fc6fe6f8 (odb/source: make `begin_transaction()` function pluggable, 2026-03-05) merged into it. Changes in v2: - Properly initialize `out` pointer when counting loose objects. - Fix a stale comment. - Fix a commit message type. - Link to v1: https://lore.kernel.org/r/20260310-b4-pks-odb-source-count-objects-v1-0-109e07d425f4@pks.im Thanks! Patrick --- Patrick Steinhardt (6): odb: stop including "odb/source.h" packfile: extract logic to count number of objects object-file: extract logic to approximate object count object-file: generalize counting objects odb/source: introduce generic object counting odb: introduce generic object counting builtin/gc.c | 44 +++++++++---------------- builtin/multi-pack-index.c | 1 + builtin/submodule--helper.c | 1 + commit-graph.c | 3 +- object-file.c | 58 +++++++++++++++++++++++++++++++++ object-file.h | 14 ++++++++ object-name.c | 6 +++- odb.c | 37 ++++++++++++++++++++- odb.h | 78 +++++++++++++++++++++++++++++++++++++++++--- odb/source-files.c | 30 +++++++++++++++++ odb/source.h | 79 ++++++++++++++++----------------------------- odb/streaming.c | 1 + packfile.c | 48 +++++++++++++-------------- packfile.h | 16 +++++---- repository.c | 1 + submodule-config.c | 1 + tmp-objdir.c | 1 + 17 files changed, 301 insertions(+), 118 deletions(-) Range-diff versus v1: 1: 3d5f8733d5 = 1: 1639bb1725 odb: stop including "odb/source.h" 2: 2fe42618f4 = 2: 056b6f3ae3 packfile: extract logic to count number of objects 3: 4cbf727523 ! 3: 069c908771 object-file: extract logic to approximate object count @@ Commit message repack objects. This is done by counting the number of objects that we have and checking whether it exceeds a certain threshold. We don't really need an accurate object count though, which is why we only - open a single object diretcroy shard and then extrapolate from there. + open a single object directory shard and then extrapolate from there. Extract this logic into a new function that is owned by the loose object database source. This is done to prepare for a subsequent change, where 4: c1a527877a ! 4: 6f293f1352 object-file: generalize counting objects @@ object-file.c: int odb_source_loose_for_each_object(struct odb_source *source, + *out = count * 256; + ret = 0; + } else { ++ *out = 0; + ret = odb_source_loose_for_each_object(source, NULL, count_loose_object, + out, 0); } 5: c12d4ec401 = 5: 0bda4a3d01 odb/source: introduce generic object counting 6: 91307f205d ! 6: 89164dad76 odb: introduce generic object counting @@ odb.c: void odb_reprepare(struct object_database *o) ## odb.h ## @@ odb.h: struct object_database { + /* + * A fast, rough count of the number of objects in the repository. * These two fields are not meant for direct access. Use - * repo_approximate_object_count() instead. +- * repo_approximate_object_count() instead. ++ * odb_count_objects() instead. */ - unsigned long approximate_object_count; - unsigned approximate_object_count_valid : 1; --- base-commit: 2247f478a898a7f8f8322cc51bdeb1cc773d8f4a change-id: 20260224-b4-pks-odb-source-count-objects-479fe682cf6f