{"thread":{"id":"65299","subject":"[PATCH 00/14] odb: generic object name handling","startedAt":"2026-03-19T06:53:07Z","lastAt":"2026-03-31T06:04:48Z","messageCount":52,"participants":["Patrick Steinhardt","Junio C Hamano","Karthik Nayak","brian m. carlson","Toon Claes"],"isPatch":true,"patchVersion":1,"patchTotal":14},"messages":[{"id":"539352","messageId":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":null,"subject":"[PATCH 00/14] odb: generic object name handling","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:52:58Z","receivedAt":"2026-03-19T06:53:07Z","isPatch":true,"body":"Hi,\n\nthis patch series refactors handling of object names to become pluggable\nand thus generic. This includes:\n\n  - Disambiguation of object names with a common prefix. This is\n    required to list candidate objects in case the user has passed a\n    non-unique prefix.\n\n  - Abbreviating an object ID to the shortest prefix required while\n    staying unique.\n\nThe logic to compute these operations is specific to the backend, but\nnot generic. This patch series fixes that by moving the functionality\ninto the respective backends.\n\nThis patch series may feel somewhat unexiting, but it's not. Especially\nabbreviating object IDs is done in lots of places, so this functionality\nis overall quite critical. So starting with this series, it is now\npossible to do all kinds of local work with an alternative backend:\ngit-commit(1), git-log(1), git-rev-parse(1), git-merge(1) and many other\ncommands now work as expected. My MongoDB proof of concept [1] only\nrequires two commits (the object format extension) on top. And no, I\ndon't endorse MongoDB or propose it as a future potential backend. It\nsimply had a good C API that was easy to use.\n\nOf course, other functionality, especially everything that involves\npackfiles, doesn't yet work.\n\nThis patch series is built on top of ca1db8a0f7 (The 17th batch,\n2026-03-16) with ps/object-counting at 6801ffd37d (odb: introduce\ngeneric object counting, 2026-03-12) merged into it.\n\nThanks!\n\nPatrick\n\n[1]: https://gitlab.com/gitlab-org/git/-/merge_requests/454\n\n---\nPatrick Steinhardt (14):\n      oidtree: modernize the code a bit\n      oidtree: extend iteration to allow for arbitrary return codes\n      odb: introduce `struct odb_for_each_object_options`\n      object-name: move logic to iterate through loose prefixed objects\n      object-name: move logic to iterate through packed prefixed objects\n      object-name: extract function to parse object ID prefixes\n      object-name: backend-generic `repo_collect_ambiguous()`\n      object-name: backend-generic `get_short_oid()`\n      object-name: merge `update_candidates()` and `match_prefix()`\n      object-name: abbreviate loose object names without `disambiguate_state`\n      object-name: simplify computing common prefixes\n      object-name: move logic to compute loose abbreviation length\n      object-file: move logic to compute packed abbreviation length\n      odb: introduce generic `odb_find_abbrev_len()`\n\n builtin/cat-file.c       |   7 +-\n builtin/pack-objects.c   |  12 +-\n cbtree.c                 |  21 ++-\n cbtree.h                 |  11 +-\n commit-graph.c           |   5 +-\n hash.c                   |  18 ++\n hash.h                   |   3 +\n object-file.c            |  73 +++++++-\n object-file.h            |  21 ++-\n object-name.c            | 437 ++++++++---------------------------------------\n odb.c                    |  99 ++++++++++-\n odb.h                    |  39 +++++\n odb/source-files.c       |  33 +++-\n odb/source.h             |  30 +++-\n oidtree.c                |  65 +++----\n oidtree.h                |  48 +++++-\n packfile.c               | 297 +++++++++++++++++++++++++++++++-\n packfile.h               |   7 +-\n t/unit-tests/u-oidtree.c |  18 +-\n 19 files changed, 774 insertions(+), 470 deletions(-)\n\n\n---\nbase-commit: b052aca69d64d2d8e28e7ce97dcb1beb3d94515a\nchange-id: 20260313-b4-pks-odb-source-abbrev-a84c51222bca\n\n"},{"id":"539353","messageId":"20260319-b4-pks-odb-source-abbrev-v1-1-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 01/14] oidtree: modernize the code a bit","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:52:59Z","receivedAt":"2026-03-19T06:53:08Z","isPatch":true,"body":"The \"oidtree.c\" subsystem is rather small and self-contained and tends\nto just work. It thus doesn't typically receive a lot of attention,\nwhich has as a consequence that it's coding style is somewhat dated\nnowadays.\n\nModernize the style of this subsystem a bit:\n\n  - Rename the `oidtree_iter()` function to `oidtree_each_cb()`.\n\n  - Rename `struct oidtree_iter_data` to `struct oidtree_each_data` to\n    match the renamed callback function type.\n\n  - Rename parameters and variables to clarify their intent.\n\n  - Add comments that explain what some of the functions do.\n\n  - Adapt the return value of `oidtree_contains()` to be a boolean.\n\nThis prepares for some changes to the subsystem that'll happen in the\nnext commit.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n oidtree.c                | 61 ++++++++++++++++++++++++------------------------\n oidtree.h                | 42 +++++++++++++++++++++++++++------\n t/unit-tests/u-oidtree.c | 14 +++++------\n 3 files changed, 73 insertions(+), 44 deletions(-)\n\ndiff --git a/oidtree.c b/oidtree.c\nindex 324de94934..a4d10cd429 100644\n--- a/oidtree.c\n+++ b/oidtree.c\n@@ -6,14 +6,6 @@\n #include \"oidtree.h\"\n #include \"hash.h\"\n \n-struct oidtree_iter_data {\n-\toidtree_iter fn;\n-\tvoid *arg;\n-\tsize_t *last_nibble_at;\n-\tuint32_t algo;\n-\tuint8_t last_byte;\n-};\n-\n void oidtree_init(struct oidtree *ot)\n {\n \tcb_init(&ot->tree);\n@@ -54,8 +46,7 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n \tcb_insert(&ot->tree, on, sizeof(*oid));\n }\n \n-\n-int oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n+bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n {\n \tstruct object_id k;\n \tsize_t klen = sizeof(k);\n@@ -69,41 +60,51 @@ int oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n \tklen += BUILD_ASSERT_OR_ZERO(offsetof(struct object_id, hash) <\n \t\t\t\toffsetof(struct object_id, algo));\n \n-\treturn cb_lookup(&ot->tree, (const uint8_t *)&k, klen) ? 1 : 0;\n+\treturn !!cb_lookup(&ot->tree, (const uint8_t *)&k, klen);\n }\n \n-static enum cb_next iter(struct cb_node *n, void *arg)\n+struct oidtree_each_data {\n+\toidtree_each_cb cb;\n+\tvoid *cb_data;\n+\tsize_t *last_nibble_at;\n+\tuint32_t algo;\n+\tuint8_t last_byte;\n+};\n+\n+static enum cb_next iter(struct cb_node *n, void *cb_data)\n {\n-\tstruct oidtree_iter_data *x = arg;\n+\tstruct oidtree_each_data *data = cb_data;\n \tstruct object_id k;\n \n \t/* Copy to provide 4-byte alignment needed by struct object_id. */\n \tmemcpy(&k, n->k, sizeof(k));\n \n-\tif (x->algo != GIT_HASH_UNKNOWN && x->algo != k.algo)\n+\tif (data->algo != GIT_HASH_UNKNOWN && data->algo != k.algo)\n \t\treturn CB_CONTINUE;\n \n-\tif (x->last_nibble_at) {\n-\t\tif ((k.hash[*x->last_nibble_at] ^ x->last_byte) & 0xf0)\n+\tif (data->last_nibble_at) {\n+\t\tif ((k.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n \t\t\treturn CB_CONTINUE;\n \t}\n \n-\treturn x->fn(&k, x->arg);\n+\treturn data->cb(&k, data->cb_data);\n }\n \n-void oidtree_each(struct oidtree *ot, const struct object_id *oid,\n-\t\t\tsize_t oidhexsz, oidtree_iter fn, void *arg)\n+void oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n+\t\t  size_t prefix_hex_len, oidtree_each_cb cb, void *cb_data)\n {\n-\tsize_t klen = oidhexsz / 2;\n-\tstruct oidtree_iter_data x = { 0 };\n-\tassert(oidhexsz <= GIT_MAX_HEXSZ);\n-\n-\tx.fn = fn;\n-\tx.arg = arg;\n-\tx.algo = oid->algo;\n-\tif (oidhexsz & 1) {\n-\t\tx.last_byte = oid->hash[klen];\n-\t\tx.last_nibble_at = &klen;\n+\tstruct oidtree_each_data data = {\n+\t\t.cb = cb,\n+\t\t.cb_data = cb_data,\n+\t\t.algo = prefix->algo,\n+\t};\n+\tsize_t klen = prefix_hex_len / 2;\n+\tassert(prefix_hex_len <= GIT_MAX_HEXSZ);\n+\n+\tif (prefix_hex_len & 1) {\n+\t\tdata.last_byte = prefix->hash[klen];\n+\t\tdata.last_nibble_at = &klen;\n \t}\n-\tcb_each(&ot->tree, (const uint8_t *)oid, klen, iter, &x);\n+\n+\tcb_each(&ot->tree, prefix->hash, klen, iter, &data);\n }\ndiff --git a/oidtree.h b/oidtree.h\nindex 77898f510a..0651401017 100644\n--- a/oidtree.h\n+++ b/oidtree.h\n@@ -5,18 +5,46 @@\n #include \"hash.h\"\n #include \"mem-pool.h\"\n \n+/*\n+ * OID trees are an efficient storage for object IDs that use a critbit tree\n+ * internally. Common prefixes are duplicated and object IDs are stored in a\n+ * way that allow easy iteration over the objects in lexicographic order. As a\n+ * consequence, operations that want to enumerate all object IDs that match a\n+ * given prefix can be answered efficiently.\n+ *\n+ * Note that it is not (yet) possible to store data other than the object IDs\n+ * themselves in this tree.\n+ */\n struct oidtree {\n \tstruct cb_tree tree;\n \tstruct mem_pool mem_pool;\n };\n \n-void oidtree_init(struct oidtree *);\n-void oidtree_clear(struct oidtree *);\n-void oidtree_insert(struct oidtree *, const struct object_id *);\n-int oidtree_contains(struct oidtree *, const struct object_id *);\n+/* Initialize the oidtree so that it is ready for use. */\n+void oidtree_init(struct oidtree *ot);\n \n-typedef enum cb_next (*oidtree_iter)(const struct object_id *, void *data);\n-void oidtree_each(struct oidtree *, const struct object_id *,\n-\t\t\tsize_t oidhexsz, oidtree_iter, void *data);\n+/*\n+ * Release all memory associated with the oidtree and reinitialize it for\n+ * subsequent use.\n+ */\n+void oidtree_clear(struct oidtree *ot);\n+\n+/* Insert the object ID into the tree. */\n+void oidtree_insert(struct oidtree *ot, const struct object_id *oid);\n+\n+/* Check whether the tree contains the given object ID. */\n+bool oidtree_contains(struct oidtree *ot, const struct object_id *oid);\n+\n+/* Callback function used for `oidtree_each()`. */\n+typedef enum cb_next (*oidtree_each_cb)(const struct object_id *oid,\n+\t\t\t\t\tvoid *cb_data);\n+\n+/*\n+ * Iterate through all object IDs in the tree whose prefix matches the given\n+ * object ID prefix and invoke the callback function on each of them.\n+ */\n+void oidtree_each(struct oidtree *ot,\n+\t\t  const struct object_id *prefix, size_t prefix_hex_len,\n+\t\t  oidtree_each_cb cb, void *cb_data);\n \n #endif /* OIDTREE_H */\ndiff --git a/t/unit-tests/u-oidtree.c b/t/unit-tests/u-oidtree.c\nindex e6eede2740..def47c6795 100644\n--- a/t/unit-tests/u-oidtree.c\n+++ b/t/unit-tests/u-oidtree.c\n@@ -24,7 +24,7 @@ static int fill_tree_loc(struct oidtree *ot, const char *hexes[], size_t n)\n \treturn 0;\n }\n \n-static void check_contains(struct oidtree *ot, const char *hex, int expected)\n+static void check_contains(struct oidtree *ot, const char *hex, bool expected)\n {\n \tstruct object_id oid;\n \n@@ -88,12 +88,12 @@ void test_oidtree__cleanup(void)\n void test_oidtree__contains(void)\n {\n \tFILL_TREE(&ot, \"444\", \"1\", \"2\", \"3\", \"4\", \"5\", \"a\", \"b\", \"c\", \"d\", \"e\");\n-\tcheck_contains(&ot, \"44\", 0);\n-\tcheck_contains(&ot, \"441\", 0);\n-\tcheck_contains(&ot, \"440\", 0);\n-\tcheck_contains(&ot, \"444\", 1);\n-\tcheck_contains(&ot, \"4440\", 1);\n-\tcheck_contains(&ot, \"4444\", 0);\n+\tcheck_contains(&ot, \"44\", false);\n+\tcheck_contains(&ot, \"441\", false);\n+\tcheck_contains(&ot, \"440\", false);\n+\tcheck_contains(&ot, \"444\", true);\n+\tcheck_contains(&ot, \"4440\", true);\n+\tcheck_contains(&ot, \"4444\", false);\n }\n \n void test_oidtree__each(void)\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539354","messageId":"20260319-b4-pks-odb-source-abbrev-v1-2-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 02/14] oidtree: extend iteration to allow for arbitrary return codes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:00Z","receivedAt":"2026-03-19T06:53:11Z","isPatch":true,"body":"The interface `cb_each()` iterates through a crit-bit tree and calls a\nspecific callback function for each of the contained items. The callback\nfunction is expected to return either:\n\n  - `CB_CONTINUE` in case iteration shall continue.\n\n  - `CB_BREAK` to abort iteration.\n\nThis is needlessly restrictive though, as callers may want to return\narbitrary values and have them be bubbled up to the `cb_each()` call\nsite. In fact, this is a rather common pattern we have: whenever such a\ncallback function returns a non-zero error code, we abort iteration and\nbubble up the code as-is.\n\nRefactor both the crit-bit tree and oidtree subsystems to behave\naccordingly.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n cbtree.c                 | 21 ++++++++++++---------\n cbtree.h                 | 11 +++--------\n object-name.c            |  4 ++--\n oidtree.c                | 12 ++++++------\n oidtree.h                | 18 ++++++++++++------\n t/unit-tests/u-oidtree.c |  4 ++--\n 6 files changed, 37 insertions(+), 33 deletions(-)\n\ndiff --git a/cbtree.c b/cbtree.c\nindex cf8cf75b89..4ab794bddc 100644\n--- a/cbtree.c\n+++ b/cbtree.c\n@@ -96,26 +96,28 @@ struct cb_node *cb_lookup(struct cb_tree *t, const uint8_t *k, size_t klen)\n \treturn p && !memcmp(p->k, k, klen) ? p : NULL;\n }\n \n-static enum cb_next cb_descend(struct cb_node *p, cb_iter fn, void *arg)\n+static int cb_descend(struct cb_node *p, cb_iter fn, void *arg)\n {\n \tif (1 & (uintptr_t)p) {\n \t\tstruct cb_node *q = cb_node_of(p);\n-\t\tenum cb_next n = cb_descend(q->child[0], fn, arg);\n-\n-\t\treturn n == CB_BREAK ? n : cb_descend(q->child[1], fn, arg);\n+\t\tint ret = cb_descend(q->child[0], fn, arg);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t\treturn cb_descend(q->child[1], fn, arg);\n \t} else {\n \t\treturn fn(p, arg);\n \t}\n }\n \n-void cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n-\t\t\tcb_iter fn, void *arg)\n+int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n+\t    cb_iter fn, void *arg)\n {\n \tstruct cb_node *p = t->root;\n \tstruct cb_node *top = p;\n \tsize_t i = 0;\n \n-\tif (!p) return; /* empty tree */\n+\tif (!p)\n+\t\treturn 0; /* empty tree */\n \n \t/* Walk tree, maintaining top pointer */\n \twhile (1 & (uintptr_t)p) {\n@@ -130,7 +132,8 @@ void cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n \n \tfor (i = 0; i < klen; i++) {\n \t\tif (p->k[i] != kpfx[i])\n-\t\t\treturn; /* \"best\" match failed */\n+\t\t\treturn 0; /* \"best\" match failed */\n \t}\n-\tcb_descend(top, fn, arg);\n+\n+\treturn cb_descend(top, fn, arg);\n }\ndiff --git a/cbtree.h b/cbtree.h\nindex 43193abdda..4f644d6e45 100644\n--- a/cbtree.h\n+++ b/cbtree.h\n@@ -30,11 +30,6 @@ struct cb_tree {\n \tstruct cb_node *root;\n };\n \n-enum cb_next {\n-\tCB_CONTINUE = 0,\n-\tCB_BREAK = 1\n-};\n-\n #define CBTREE_INIT { 0 }\n \n static inline void cb_init(struct cb_tree *t)\n@@ -46,9 +41,9 @@ static inline void cb_init(struct cb_tree *t)\n struct cb_node *cb_lookup(struct cb_tree *, const uint8_t *k, size_t klen);\n struct cb_node *cb_insert(struct cb_tree *, struct cb_node *, size_t klen);\n \n-typedef enum cb_next (*cb_iter)(struct cb_node *, void *arg);\n+typedef int (*cb_iter)(struct cb_node *, void *arg);\n \n-void cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n-\t\tcb_iter, void *arg);\n+int cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n+\t    cb_iter, void *arg);\n \n #endif /* CBTREE_H */\ndiff --git a/object-name.c b/object-name.c\nindex e5adec4c9d..a24a1b48e1 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -103,12 +103,12 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n \n static int match_hash(unsigned, const unsigned char *, const unsigned char *);\n \n-static enum cb_next match_prefix(const struct object_id *oid, void *arg)\n+static int match_prefix(const struct object_id *oid, void *arg)\n {\n \tstruct disambiguate_state *ds = arg;\n \t/* no need to call match_hash, oidtree_each did prefix match */\n \tupdate_candidates(ds, oid);\n-\treturn ds->ambiguous ? CB_BREAK : CB_CONTINUE;\n+\treturn ds->ambiguous;\n }\n \n static void find_short_object_filename(struct disambiguate_state *ds)\ndiff --git a/oidtree.c b/oidtree.c\nindex a4d10cd429..ab9fe7ec7a 100644\n--- a/oidtree.c\n+++ b/oidtree.c\n@@ -71,7 +71,7 @@ struct oidtree_each_data {\n \tuint8_t last_byte;\n };\n \n-static enum cb_next iter(struct cb_node *n, void *cb_data)\n+static int iter(struct cb_node *n, void *cb_data)\n {\n \tstruct oidtree_each_data *data = cb_data;\n \tstruct object_id k;\n@@ -80,18 +80,18 @@ static enum cb_next iter(struct cb_node *n, void *cb_data)\n \tmemcpy(&k, n->k, sizeof(k));\n \n \tif (data->algo != GIT_HASH_UNKNOWN && data->algo != k.algo)\n-\t\treturn CB_CONTINUE;\n+\t\treturn 0;\n \n \tif (data->last_nibble_at) {\n \t\tif ((k.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n-\t\t\treturn CB_CONTINUE;\n+\t\t\treturn 0;\n \t}\n \n \treturn data->cb(&k, data->cb_data);\n }\n \n-void oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n-\t\t  size_t prefix_hex_len, oidtree_each_cb cb, void *cb_data)\n+int oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n+\t\t size_t prefix_hex_len, oidtree_each_cb cb, void *cb_data)\n {\n \tstruct oidtree_each_data data = {\n \t\t.cb = cb,\n@@ -106,5 +106,5 @@ void oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n \t\tdata.last_nibble_at = &klen;\n \t}\n \n-\tcb_each(&ot->tree, prefix->hash, klen, iter, &data);\n+\treturn cb_each(&ot->tree, prefix->hash, klen, iter, &data);\n }\ndiff --git a/oidtree.h b/oidtree.h\nindex 0651401017..2b7bad2e60 100644\n--- a/oidtree.h\n+++ b/oidtree.h\n@@ -35,16 +35,22 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid);\n /* Check whether the tree contains the given object ID. */\n bool oidtree_contains(struct oidtree *ot, const struct object_id *oid);\n \n-/* Callback function used for `oidtree_each()`. */\n-typedef enum cb_next (*oidtree_each_cb)(const struct object_id *oid,\n-\t\t\t\t\tvoid *cb_data);\n+/*\n+ * Callback function used for `oidtree_each()`. Returning a non-zero exit code\n+ * will cause iteration to stop. The exit code will be propagated to the caller\n+ * of `oidtree_each()`.\n+ */\n+typedef int (*oidtree_each_cb)(const struct object_id *oid,\n+\t\t\t       void *cb_data);\n \n /*\n  * Iterate through all object IDs in the tree whose prefix matches the given\n  * object ID prefix and invoke the callback function on each of them.\n+ *\n+ * Returns any non-zero exit code from the provided callback function.\n  */\n-void oidtree_each(struct oidtree *ot,\n-\t\t  const struct object_id *prefix, size_t prefix_hex_len,\n-\t\t  oidtree_each_cb cb, void *cb_data);\n+int oidtree_each(struct oidtree *ot,\n+\t\t const struct object_id *prefix, size_t prefix_hex_len,\n+\t\t oidtree_each_cb cb, void *cb_data);\n \n #endif /* OIDTREE_H */\ndiff --git a/t/unit-tests/u-oidtree.c b/t/unit-tests/u-oidtree.c\nindex def47c6795..d4d05c7dc3 100644\n--- a/t/unit-tests/u-oidtree.c\n+++ b/t/unit-tests/u-oidtree.c\n@@ -38,7 +38,7 @@ struct expected_hex_iter {\n \tconst char *query;\n };\n \n-static enum cb_next check_each_cb(const struct object_id *oid, void *data)\n+static int check_each_cb(const struct object_id *oid, void *data)\n {\n \tstruct expected_hex_iter *hex_iter = data;\n \tstruct object_id expected;\n@@ -49,7 +49,7 @@ static enum cb_next check_each_cb(const struct object_id *oid, void *data)\n \t\t\t &expected);\n \tcl_assert_equal_s(oid_to_hex(oid), oid_to_hex(&expected));\n \thex_iter->i += 1;\n-\treturn CB_CONTINUE;\n+\treturn 0;\n }\n \n LAST_ARG_MUST_BE_NULL\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539355","messageId":"20260319-b4-pks-odb-source-abbrev-v1-3-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 03/14] odb: introduce `struct odb_for_each_object_options`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:01Z","receivedAt":"2026-03-19T06:53:13Z","isPatch":true,"body":"The `odb_for_each_object()` function only accepts a bitset of flags. In\na subsequent commit we'll want to change object iteration to also\nsupport iterating over only those objects that have a specific prefix.\nWhile we could of course add the prefix to the function signature, or\nalternative introduce a new function, both of these options don't really\nseem to be that sensible.\n\nInstead, introduce a new `struct odb_for_each_object_options` that can\nbe passed to a new `odb_for_each_object_ext()` function. Splice through\nthe options structure into the respective object database sources.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/cat-file.c     |  7 +++++--\n builtin/pack-objects.c | 12 +++++++-----\n commit-graph.c         |  5 ++++-\n object-file.c          |  6 +++---\n object-file.h          |  2 +-\n odb.c                  | 26 +++++++++++++++++++-------\n odb.h                  | 16 ++++++++++++++++\n odb/source-files.c     |  8 ++++----\n odb/source.h           |  6 +++---\n packfile.c             | 12 ++++++------\n packfile.h             |  2 +-\n 11 files changed, 69 insertions(+), 33 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex b6f12f41d6..cd13a3a89f 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -848,6 +848,9 @@ static void batch_each_object(struct batch_options *opt,\n \t\t.callback = callback,\n \t\t.payload = _payload,\n \t};\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.flags = flags,\n+\t};\n \tstruct bitmap_index *bitmap = NULL;\n \tstruct odb_source *source;\n \n@@ -860,7 +863,7 @@ static void batch_each_object(struct batch_options *opt,\n \todb_prepare_alternates(the_repository->objects);\n \tfor (source = the_repository->objects->sources; source; source = source->next) {\n \t\tint ret = odb_source_loose_for_each_object(source, NULL, batch_one_object_oi,\n-\t\t\t\t\t\t\t   &payload, flags);\n+\t\t\t\t\t\t\t   &payload, &opts);\n \t\tif (ret)\n \t\t\tbreak;\n \t}\n@@ -884,7 +887,7 @@ static void batch_each_object(struct batch_options *opt,\n \t\tfor (source = the_repository->objects->sources; source; source = source->next) {\n \t\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n \t\t\tint ret = packfile_store_for_each_object(files->packed, &oi,\n-\t\t\t\t\t\t\t\t batch_one_object_oi, &payload, flags);\n+\t\t\t\t\t\t\t\t batch_one_object_oi, &payload, &opts);\n \t\t\tif (ret)\n \t\t\t\tbreak;\n \t\t}\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex cd013c0b68..3bb57ff183 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -4344,6 +4344,12 @@ static void add_objects_in_unpacked_packs(void)\n {\n \tstruct odb_source *source;\n \ttime_t mtime;\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.flags = ODB_FOR_EACH_OBJECT_PACK_ORDER |\n+\t\t\t ODB_FOR_EACH_OBJECT_LOCAL_ONLY |\n+\t\t\t ODB_FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS |\n+\t\t\t ODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS,\n+\t};\n \tstruct object_info oi = {\n \t\t.mtimep = &mtime,\n \t};\n@@ -4356,11 +4362,7 @@ static void add_objects_in_unpacked_packs(void)\n \t\t\tcontinue;\n \n \t\tif (packfile_store_for_each_object(files->packed, &oi,\n-\t\t\t\t\t\t   add_object_in_unpacked_pack, NULL,\n-\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_PACK_ORDER |\n-\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_LOCAL_ONLY |\n-\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS |\n-\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS))\n+\t\t\t\t\t\t   add_object_in_unpacked_pack, NULL, &opts))\n \t\t\tdie(_(\"cannot open pack index\"));\n \t}\n }\ndiff --git a/commit-graph.c b/commit-graph.c\nindex c030003330..df4b4a125e 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1969,6 +1969,9 @@ static void fill_oids_from_all_packs(struct write_commit_graph_context *ctx)\n {\n \tstruct odb_source *source;\n \tenum object_type type;\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.flags = ODB_FOR_EACH_OBJECT_PACK_ORDER,\n+\t};\n \tstruct object_info oi = {\n \t\t.typep = &type,\n \t};\n@@ -1983,7 +1986,7 @@ static void fill_oids_from_all_packs(struct write_commit_graph_context *ctx)\n \tfor (source = ctx->r->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n \t\tpackfile_store_for_each_object(files->packed, &oi, add_packed_commits_oi,\n-\t\t\t\t\t       ctx, ODB_FOR_EACH_OBJECT_PACK_ORDER);\n+\t\t\t\t\t       ctx, &opts);\n \t}\n \n \tif (ctx->progress_done < ctx->approx_nr_objects)\ndiff --git a/object-file.c b/object-file.c\nindex f0b029ff0b..ddcc8e81b4 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1849,7 +1849,7 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t     const struct object_info *request,\n \t\t\t\t     odb_for_each_object_cb cb,\n \t\t\t\t     void *cb_data,\n-\t\t\t\t     unsigned flags)\n+\t\t\t\t     const struct odb_for_each_object_options *opts)\n {\n \tstruct for_each_object_wrapper_data data = {\n \t\t.source = source,\n@@ -1859,9 +1859,9 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t};\n \n \t/* There are no loose promisor objects, so we can return immediately. */\n-\tif ((flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY))\n+\tif ((opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY))\n \t\treturn 0;\n-\tif ((flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY) && !source->local)\n+\tif ((opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY) && !source->local)\n \t\treturn 0;\n \n \treturn for_each_loose_file_in_source(source, for_each_object_wrapper_cb,\ndiff --git a/object-file.h b/object-file.h\nindex f8d8805a18..46dfa7b632 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -137,7 +137,7 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t     const struct object_info *request,\n \t\t\t\t     odb_for_each_object_cb cb,\n \t\t\t\t     void *cb_data,\n-\t\t\t\t     unsigned flags);\n+\t\t\t\t     const struct odb_for_each_object_options *opts);\n \n /*\n  * Count the number of loose objects in this source.\ndiff --git a/odb.c b/odb.c\nindex 350e23f3c0..3019957b87 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -896,20 +896,20 @@ int odb_freshen_object(struct object_database *odb,\n \treturn 0;\n }\n \n-int odb_for_each_object(struct object_database *odb,\n-\t\t\tconst struct object_info *request,\n-\t\t\todb_for_each_object_cb cb,\n-\t\t\tvoid *cb_data,\n-\t\t\tunsigned flags)\n+int odb_for_each_object_ext(struct object_database *odb,\n+\t\t\t    const struct object_info *request,\n+\t\t\t    odb_for_each_object_cb cb,\n+\t\t\t    void *cb_data,\n+\t\t\t    const struct odb_for_each_object_options *opts)\n {\n \tint ret;\n \n \todb_prepare_alternates(odb);\n \tfor (struct odb_source *source = odb->sources; source; source = source->next) {\n-\t\tif (flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY && !source->local)\n+\t\tif (opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY && !source->local)\n \t\t\tcontinue;\n \n-\t\tret = odb_source_for_each_object(source, request, cb, cb_data, flags);\n+\t\tret = odb_source_for_each_object(source, request, cb, cb_data, opts);\n \t\tif (ret)\n \t\t\treturn ret;\n \t}\n@@ -917,6 +917,18 @@ int odb_for_each_object(struct object_database *odb,\n \treturn 0;\n }\n \n+int odb_for_each_object(struct object_database *odb,\n+\t\t\tconst struct object_info *request,\n+\t\t\todb_for_each_object_cb cb,\n+\t\t\tvoid *cb_data,\n+\t\t\tunsigned flags)\n+{\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.flags = flags,\n+\t};\n+\treturn odb_for_each_object_ext(odb, request, cb, cb_data, &opts);\n+}\n+\n int odb_count_objects(struct object_database *odb,\n \t\t      enum odb_count_objects_flags flags,\n \t\t      unsigned long *out)\ndiff --git a/odb.h b/odb.h\nindex 9aee260105..a19a8bb50d 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -481,6 +481,15 @@ typedef int (*odb_for_each_object_cb)(const struct object_id *oid,\n \t\t\t\t      struct object_info *oi,\n \t\t\t\t      void *cb_data);\n \n+/*\n+ * Options that can be passed to `odb_for_each_object()` and its\n+ * backend-specific implementations.\n+ */\n+struct odb_for_each_object_options {\n+\t/* A bitfield of `odb_for_each_object_flags`. */\n+\tenum odb_for_each_object_flags flags;\n+};\n+\n /*\n  * Iterate through all objects contained in the object database. Note that\n  * objects may be iterated over multiple times in case they are either stored\n@@ -495,6 +504,13 @@ typedef int (*odb_for_each_object_cb)(const struct object_id *oid,\n  * Returns 0 on success, a negative error code in case a failure occurred, or\n  * an arbitrary non-zero error code returned by the callback itself.\n  */\n+int odb_for_each_object_ext(struct object_database *odb,\n+\t\t\t    const struct object_info *request,\n+\t\t\t    odb_for_each_object_cb cb,\n+\t\t\t    void *cb_data,\n+\t\t\t    const struct odb_for_each_object_options *opts);\n+\n+/* Same as `odb_for_each_object_ext()` with `opts.flags` set to the given flags. */\n int odb_for_each_object(struct object_database *odb,\n \t\t\tconst struct object_info *request,\n \t\t\todb_for_each_object_cb cb,\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex c08d8993e3..e90bb689bb 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -75,18 +75,18 @@ static int odb_source_files_for_each_object(struct odb_source *source,\n \t\t\t\t\t    const struct object_info *request,\n \t\t\t\t\t    odb_for_each_object_cb cb,\n \t\t\t\t\t    void *cb_data,\n-\t\t\t\t\t    unsigned flags)\n+\t\t\t\t\t    const struct odb_for_each_object_options *opts)\n {\n \tstruct odb_source_files *files = odb_source_files_downcast(source);\n \tint ret;\n \n-\tif (!(flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY)) {\n-\t\tret = odb_source_loose_for_each_object(source, request, cb, cb_data, flags);\n+\tif (!(opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY)) {\n+\t\tret = odb_source_loose_for_each_object(source, request, cb, cb_data, opts);\n \t\tif (ret)\n \t\t\treturn ret;\n \t}\n \n-\tret = packfile_store_for_each_object(files->packed, request, cb, cb_data, flags);\n+\tret = packfile_store_for_each_object(files->packed, request, cb, cb_data, opts);\n \tif (ret)\n \t\treturn ret;\n \ndiff --git a/odb/source.h b/odb/source.h\nindex 96c906e7a1..ee5d6ed530 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -140,7 +140,7 @@ struct odb_source {\n \t\t\t       const struct object_info *request,\n \t\t\t       odb_for_each_object_cb cb,\n \t\t\t       void *cb_data,\n-\t\t\t       unsigned flags);\n+\t\t\t       const struct odb_for_each_object_options *opts);\n \n \t/*\n \t * This callback is expected to count objects in the given object\n@@ -343,9 +343,9 @@ static inline int odb_source_for_each_object(struct odb_source *source,\n \t\t\t\t\t     const struct object_info *request,\n \t\t\t\t\t     odb_for_each_object_cb cb,\n \t\t\t\t\t     void *cb_data,\n-\t\t\t\t\t     unsigned flags)\n+\t\t\t\t\t     const struct odb_for_each_object_options *opts)\n {\n-\treturn source->for_each_object(source, request, cb, cb_data, flags);\n+\treturn source->for_each_object(source, request, cb, cb_data, opts);\n }\n \n /*\ndiff --git a/packfile.c b/packfile.c\nindex d4de9f3ffe..a6f3d2035d 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2375,7 +2375,7 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \t\t\t\t   const struct object_info *request,\n \t\t\t\t   odb_for_each_object_cb cb,\n \t\t\t\t   void *cb_data,\n-\t\t\t\t   unsigned flags)\n+\t\t\t\t   const struct odb_for_each_object_options *opts)\n {\n \tstruct packfile_store_for_each_object_wrapper_data data = {\n \t\t.store = store,\n@@ -2391,15 +2391,15 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n \t\tstruct packed_git *p = e->pack;\n \n-\t\tif ((flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY) && !p->pack_local)\n+\t\tif ((opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY) && !p->pack_local)\n \t\t\tcontinue;\n-\t\tif ((flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY) &&\n+\t\tif ((opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY) &&\n \t\t    !p->pack_promisor)\n \t\t\tcontinue;\n-\t\tif ((flags & ODB_FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS) &&\n+\t\tif ((opts->flags & ODB_FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS) &&\n \t\t    p->pack_keep_in_core)\n \t\t\tcontinue;\n-\t\tif ((flags & ODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS) &&\n+\t\tif ((opts->flags & ODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS) &&\n \t\t    p->pack_keep)\n \t\t\tcontinue;\n \t\tif (open_pack_index(p)) {\n@@ -2408,7 +2408,7 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \t\t}\n \n \t\tret = for_each_object_in_pack(p, packfile_store_for_each_object_wrapper,\n-\t\t\t\t\t      &data, flags);\n+\t\t\t\t\t      &data, opts->flags);\n \t\tif (ret)\n \t\t\tgoto out;\n \t}\ndiff --git a/packfile.h b/packfile.h\nindex a16ec3950d..fa41dfda38 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -367,7 +367,7 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \t\t\t\t   const struct object_info *request,\n \t\t\t\t   odb_for_each_object_cb cb,\n \t\t\t\t   void *cb_data,\n-\t\t\t\t   unsigned flags);\n+\t\t\t\t   const struct odb_for_each_object_options *opts);\n \n /* A hook to report invalid files in pack directory */\n #define PACKDIR_FILE_PACK 1\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539356","messageId":"20260319-b4-pks-odb-source-abbrev-v1-4-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 04/14] object-name: move logic to iterate through loose prefixed objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:02Z","receivedAt":"2026-03-19T06:53:16Z","isPatch":true,"body":"The logic to iterate through loose objects that have a certain prefix is\ncurrently hosted in \"object-name.c\". This logic reaches into specifics\nof the loose object source, so it breaks once a different backend is\nused for the object storage.\n\nMove the logic to iterate through loose objects with a prefix into\n\"object-file.c\". This is done by extending the for-each-object options\nto support an optional prefix that is then honored by the loose source.\nNaturally, we'll also have this support in the packfile store. This is\ndone in the next commit.\n\nFurthermore, there are no users of the loose cache outside of\n\"object-file.c\" anymore. As such, convert `odb_source_loose_cache()` to\nhave file scope.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-file.c | 29 +++++++++++++++++++++++++++--\n object-file.h |  7 -------\n object-name.c | 10 ++++++----\n odb.h         |  7 +++++++\n 4 files changed, 40 insertions(+), 13 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex ddcc8e81b4..8a9e68a768 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -33,6 +33,9 @@\n /* The maximum size for an object header. */\n #define MAX_HEADER_LEN 32\n \n+static struct oidtree *odb_source_loose_cache(struct odb_source *source,\n+\t\t\t\t\t      const struct object_id *oid);\n+\n static int get_conv_flags(unsigned flags)\n {\n \tif (flags & INDEX_RENORMALIZE)\n@@ -1845,6 +1848,23 @@ static int for_each_object_wrapper_cb(const struct object_id *oid,\n \t}\n }\n \n+static int for_each_prefixed_object_wrapper_cb(const struct object_id *oid,\n+\t\t\t\t\t       void *cb_data)\n+{\n+\tstruct for_each_object_wrapper_data *data = cb_data;\n+\tif (data->request) {\n+\t\tstruct object_info oi = *data->request;\n+\n+\t\tif (odb_source_loose_read_object_info(data->source,\n+\t\t\t\t\t\t      oid, &oi, 0) < 0)\n+\t\t\treturn -1;\n+\n+\t\treturn data->cb(oid, &oi, data->cb_data);\n+\t} else {\n+\t\treturn data->cb(oid, NULL, data->cb_data);\n+\t}\n+}\n+\n int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t     const struct object_info *request,\n \t\t\t\t     odb_for_each_object_cb cb,\n@@ -1864,6 +1884,11 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \tif ((opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY) && !source->local)\n \t\treturn 0;\n \n+\tif (opts->prefix)\n+\t\treturn oidtree_each(odb_source_loose_cache(source, opts->prefix),\n+\t\t\t\t    opts->prefix, opts->prefix_hex_len,\n+\t\t\t\t    for_each_prefixed_object_wrapper_cb, &data);\n+\n \treturn for_each_loose_file_in_source(source, for_each_object_wrapper_cb,\n \t\t\t\t\t     NULL, NULL, &data);\n }\n@@ -1934,8 +1959,8 @@ static int append_loose_object(const struct object_id *oid,\n \treturn 0;\n }\n \n-struct oidtree *odb_source_loose_cache(struct odb_source *source,\n-\t\t\t\t       const struct object_id *oid)\n+static struct oidtree *odb_source_loose_cache(struct odb_source *source,\n+\t\t\t\t\t      const struct object_id *oid)\n {\n \tstruct odb_source_files *files = odb_source_files_downcast(source);\n \tint subdir_nr = oid->hash[0];\ndiff --git a/object-file.h b/object-file.h\nindex 46dfa7b632..f11ad58f6c 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -74,13 +74,6 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \t\t\t\t  struct odb_write_stream *stream, size_t len,\n \t\t\t\t  struct object_id *oid);\n \n-/*\n- * Populate and return the loose object cache array corresponding to the\n- * given object ID.\n- */\n-struct oidtree *odb_source_loose_cache(struct odb_source *source,\n-\t\t\t\t       const struct object_id *oid);\n-\n /*\n  * Put in `buf` the name of the file in the local object database that\n  * would be used to store a loose object with the specified oid.\ndiff --git a/object-name.c b/object-name.c\nindex a24a1b48e1..929a68dbd0 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -16,7 +16,6 @@\n #include \"remote.h\"\n #include \"dir.h\"\n #include \"oid-array.h\"\n-#include \"oidtree.h\"\n #include \"packfile.h\"\n #include \"pretty.h\"\n #include \"object-file.h\"\n@@ -103,7 +102,7 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n \n static int match_hash(unsigned, const unsigned char *, const unsigned char *);\n \n-static int match_prefix(const struct object_id *oid, void *arg)\n+static int match_prefix(const struct object_id *oid, struct object_info *oi UNUSED, void *arg)\n {\n \tstruct disambiguate_state *ds = arg;\n \t/* no need to call match_hash, oidtree_each did prefix match */\n@@ -113,11 +112,14 @@ static int match_prefix(const struct object_id *oid, void *arg)\n \n static void find_short_object_filename(struct disambiguate_state *ds)\n {\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = &ds->bin_pfx,\n+\t\t.prefix_hex_len = ds->len,\n+\t};\n \tstruct odb_source *source;\n \n \tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next)\n-\t\toidtree_each(odb_source_loose_cache(source, &ds->bin_pfx),\n-\t\t\t\t&ds->bin_pfx, ds->len, match_prefix, ds);\n+\t\todb_source_loose_for_each_object(source, NULL, match_prefix, ds, &opts);\n }\n \n static int match_hash(unsigned len, const unsigned char *a, const unsigned char *b)\ndiff --git a/odb.h b/odb.h\nindex a19a8bb50d..e80fd8f7ab 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -488,6 +488,13 @@ typedef int (*odb_for_each_object_cb)(const struct object_id *oid,\n struct odb_for_each_object_options {\n \t/* A bitfield of `odb_for_each_object_flags`. */\n \tenum odb_for_each_object_flags flags;\n+\n+\t/*\n+\t * If set, only iterate through objects whose first `prefix_hex_len`\n+\t * hex characters matches the given prefix.\n+\t */\n+\tconst struct object_id *prefix;\n+\tsize_t prefix_hex_len;\n };\n \n /*\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539357","messageId":"20260319-b4-pks-odb-source-abbrev-v1-5-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 05/14] object-name: move logic to iterate through packed prefixed objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:03Z","receivedAt":"2026-03-19T06:53:18Z","isPatch":true,"body":"Similar to the preceding commit, move the logic to iterate through\nobjects that have a given prefix into \"packfile.c\".\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c |  94 +++----------------------------\n packfile.c    | 174 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 181 insertions(+), 87 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex 929a68dbd0..ff0de06ff9 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -100,8 +100,6 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n \t/* otherwise, current can be discarded and candidate is still good */\n }\n \n-static int match_hash(unsigned, const unsigned char *, const unsigned char *);\n-\n static int match_prefix(const struct object_id *oid, struct object_info *oi UNUSED, void *arg)\n {\n \tstruct disambiguate_state *ds = arg;\n@@ -122,103 +120,25 @@ static void find_short_object_filename(struct disambiguate_state *ds)\n \t\todb_source_loose_for_each_object(source, NULL, match_prefix, ds, &opts);\n }\n \n-static int match_hash(unsigned len, const unsigned char *a, const unsigned char *b)\n-{\n-\tdo {\n-\t\tif (*a != *b)\n-\t\t\treturn 0;\n-\t\ta++;\n-\t\tb++;\n-\t\tlen -= 2;\n-\t} while (len > 1);\n-\tif (len)\n-\t\tif ((*a ^ *b) & 0xf0)\n-\t\t\treturn 0;\n-\treturn 1;\n-}\n-\n-static void unique_in_midx(struct multi_pack_index *m,\n-\t\t\t   struct disambiguate_state *ds)\n-{\n-\tfor (; m; m = m->base_midx) {\n-\t\tuint32_t num, i, first = 0;\n-\t\tconst struct object_id *current = NULL;\n-\t\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n-\t\t\tds->repo->hash_algo->hexsz : ds->len;\n-\n-\t\tif (!m->num_objects)\n-\t\t\tcontinue;\n-\n-\t\tnum = m->num_objects + m->num_objects_in_base;\n-\n-\t\tbsearch_one_midx(&ds->bin_pfx, m, &first);\n-\n-\t\t/*\n-\t\t * At this point, \"first\" is the location of the lowest\n-\t\t * object with an object name that could match\n-\t\t * \"bin_pfx\".  See if we have 0, 1 or more objects that\n-\t\t * actually match(es).\n-\t\t */\n-\t\tfor (i = first; i < num && !ds->ambiguous; i++) {\n-\t\t\tstruct object_id oid;\n-\t\t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n-\t\t\tif (!match_hash(len, ds->bin_pfx.hash, current->hash))\n-\t\t\t\tbreak;\n-\t\t\tupdate_candidates(ds, current);\n-\t\t}\n-\t}\n-}\n-\n-static void unique_in_pack(struct packed_git *p,\n-\t\t\t   struct disambiguate_state *ds)\n-{\n-\tuint32_t num, i, first = 0;\n-\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n-\t\tds->repo->hash_algo->hexsz : ds->len;\n-\n-\tif (p->multi_pack_index)\n-\t\treturn;\n-\n-\tif (open_pack_index(p) || !p->num_objects)\n-\t\treturn;\n-\n-\tnum = p->num_objects;\n-\tbsearch_pack(&ds->bin_pfx, p, &first);\n-\n-\t/*\n-\t * At this point, \"first\" is the location of the lowest object\n-\t * with an object name that could match \"bin_pfx\".  See if we have\n-\t * 0, 1 or more objects that actually match(es).\n-\t */\n-\tfor (i = first; i < num && !ds->ambiguous; i++) {\n-\t\tstruct object_id oid;\n-\t\tnth_packed_object_id(&oid, p, i);\n-\t\tif (!match_hash(len, ds->bin_pfx.hash, oid.hash))\n-\t\t\tbreak;\n-\t\tupdate_candidates(ds, &oid);\n-\t}\n-}\n-\n static void find_short_packed_object(struct disambiguate_state *ds)\n {\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = &ds->bin_pfx,\n+\t\t.prefix_hex_len = ds->len,\n+\t};\n \tstruct odb_source *source;\n-\tstruct packed_git *p;\n \n \t/* Skip, unless oids from the storage hash algorithm are wanted */\n \tif (ds->bin_pfx.algo && (&hash_algos[ds->bin_pfx.algo] != ds->repo->hash_algo))\n \t\treturn;\n \n \todb_prepare_alternates(ds->repo->objects);\n-\tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next) {\n-\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n-\t\tif (m)\n-\t\t\tunique_in_midx(m, ds);\n-\t}\n+\tfor (source = ds->repo->objects->sources; source; source = source->next) {\n+\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n \n-\trepo_for_each_pack(ds->repo, p) {\n+\t\tpackfile_store_for_each_object(files->packed, NULL, match_prefix, ds, &opts);\n \t\tif (ds->ambiguous)\n \t\t\tbreak;\n-\t\tunique_in_pack(p, ds);\n \t}\n }\n \ndiff --git a/packfile.c b/packfile.c\nindex a6f3d2035d..2539a371c1 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2371,6 +2371,177 @@ static int packfile_store_for_each_object_wrapper(const struct object_id *oid,\n \t}\n }\n \n+static int match_hash(unsigned len, const unsigned char *a, const unsigned char *b)\n+{\n+\tdo {\n+\t\tif (*a != *b)\n+\t\t\treturn 0;\n+\t\ta++;\n+\t\tb++;\n+\t\tlen -= 2;\n+\t} while (len > 1);\n+\tif (len)\n+\t\tif ((*a ^ *b) & 0xf0)\n+\t\t\treturn 0;\n+\treturn 1;\n+}\n+\n+static int for_each_prefixed_object_in_midx(\n+\tstruct packfile_store *store,\n+\tstruct multi_pack_index *m,\n+\tconst struct odb_for_each_object_options *opts,\n+\tstruct packfile_store_for_each_object_wrapper_data *data)\n+{\n+\tint ret;\n+\n+\tfor (; m; m = m->base_midx) {\n+\t\tuint32_t num, i, first = 0;\n+\t\tint len = opts->prefix_hex_len > m->source->odb->repo->hash_algo->hexsz ?\n+\t\t\tm->source->odb->repo->hash_algo->hexsz : opts->prefix_hex_len;\n+\n+\t\tif (!m->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tnum = m->num_objects + m->num_objects_in_base;\n+\n+\t\tbsearch_one_midx(opts->prefix, m, &first);\n+\n+\t\t/*\n+\t\t * At this point, \"first\" is the location of the lowest\n+\t\t * object with an object name that could match \"opts->prefix\".\n+\t\t * See if we have 0, 1 or more objects that actually match(es).\n+\t\t */\n+\t\tfor (i = first; i < num; i++) {\n+\t\t\tconst struct object_id *current = NULL;\n+\t\t\tstruct object_id oid;\n+\n+\t\t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n+\n+\t\t\tif (!match_hash(len, opts->prefix->hash, current->hash))\n+\t\t\t\tbreak;\n+\n+\t\t\tif (data->request) {\n+\t\t\t\tstruct object_info oi = *data->request;\n+\n+\t\t\t\tret = packfile_store_read_object_info(store, current,\n+\t\t\t\t\t\t\t\t      &oi, 0);\n+\t\t\t\tif (ret)\n+\t\t\t\t\tgoto out;\n+\n+\t\t\t\tret = data->cb(&oid, &oi, data->cb_data);\n+\t\t\t\tif (ret)\n+\t\t\t\t\tgoto out;\n+\t\t\t} else {\n+\t\t\t\tret = data->cb(&oid, NULL, data->cb_data);\n+\t\t\t\tif (ret)\n+\t\t\t\t\tgoto out;\n+\t\t\t}\n+\t\t}\n+\t}\n+\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n+static int for_each_prefixed_object_in_pack(\n+\tstruct packfile_store *store,\n+\tstruct packed_git *p,\n+\tconst struct odb_for_each_object_options *opts,\n+\tstruct packfile_store_for_each_object_wrapper_data *data)\n+{\n+\tuint32_t num, i, first = 0;\n+\tint len = opts->prefix_hex_len > p->repo->hash_algo->hexsz ?\n+\t\tp->repo->hash_algo->hexsz : opts->prefix_hex_len;\n+\tint ret;\n+\n+\tnum = p->num_objects;\n+\tbsearch_pack(opts->prefix, p, &first);\n+\n+\t/*\n+\t * At this point, \"first\" is the location of the lowest object\n+\t * with an object name that could match \"bin_pfx\".  See if we have\n+\t * 0, 1 or more objects that actually match(es).\n+\t */\n+\tfor (i = first; i < num; i++) {\n+\t\tstruct object_id oid;\n+\n+\t\tnth_packed_object_id(&oid, p, i);\n+\t\tif (!match_hash(len, opts->prefix->hash, oid.hash))\n+\t\t\tbreak;\n+\n+\t\tif (data->request) {\n+\t\t\tstruct object_info oi = *data->request;\n+\n+\t\t\tret = packfile_store_read_object_info(store, &oid, &oi, 0);\n+\t\t\tif (ret)\n+\t\t\t\tgoto out;\n+\n+\t\t\tret = data->cb(&oid, &oi, data->cb_data);\n+\t\t\tif (ret)\n+\t\t\t\tgoto out;\n+\t\t} else {\n+\t\t\tret = data->cb(&oid, NULL, data->cb_data);\n+\t\t\tif (ret)\n+\t\t\t\tgoto out;\n+\t\t}\n+\t}\n+\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n+static int packfile_store_for_each_prefixed_object(\n+\tstruct packfile_store *store,\n+\tconst struct odb_for_each_object_options *opts,\n+\tstruct packfile_store_for_each_object_wrapper_data *data)\n+{\n+\tstruct packfile_list_entry *e;\n+\tstruct multi_pack_index *m;\n+\tbool pack_errors = false;\n+\tint ret;\n+\n+\tif (opts->flags)\n+\t\tBUG(\"flags unsupported\");\n+\n+\tstore->skip_mru_updates = true;\n+\n+\tm = get_multi_pack_index(store->source);\n+\tif (m) {\n+\t\tret = for_each_prefixed_object_in_midx(store, m, opts, data);\n+\t\tif (ret)\n+\t\t\tgoto out;\n+\t}\n+\n+\tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n+\t\tif (e->pack->multi_pack_index)\n+\t\t\tcontinue;\n+\n+\t\tif (open_pack_index(e->pack)) {\n+\t\t\tpack_errors = true;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tif (!e->pack->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tret = for_each_prefixed_object_in_pack(store, e->pack, opts, data);\n+\t\tif (ret)\n+\t\t\tgoto out;\n+\t}\n+\n+\tret = 0;\n+\n+out:\n+\tstore->skip_mru_updates = false;\n+\tif (!ret && pack_errors)\n+\t\tret = -1;\n+\treturn ret;\n+}\n+\n int packfile_store_for_each_object(struct packfile_store *store,\n \t\t\t\t   const struct object_info *request,\n \t\t\t\t   odb_for_each_object_cb cb,\n@@ -2386,6 +2557,9 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \tstruct packfile_list_entry *e;\n \tint pack_errors = 0, ret;\n \n+\tif (opts->prefix)\n+\t\treturn packfile_store_for_each_prefixed_object(store, opts, &data);\n+\n \tstore->skip_mru_updates = true;\n \n \tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539358","messageId":"20260319-b4-pks-odb-source-abbrev-v1-6-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 06/14] object-name: extract function to parse object ID prefixes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:04Z","receivedAt":"2026-03-19T06:53:21Z","isPatch":true,"body":"Extract the logic that parses an object ID prefix into a new function.\nThis function will be used by a second callsite in a subsequent commit.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 60 +++++++++++++++++++++++++++++++++++++----------------------\n 1 file changed, 38 insertions(+), 22 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex ff0de06ff9..fd1b010ab3 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -270,41 +270,57 @@ int set_disambiguate_hint_config(const char *var, const char *value)\n \treturn error(\"unknown hint type for '%s': %s\", var, value);\n }\n \n+static int parse_oid_prefix(const char *name, int len,\n+\t\t\t    const struct git_hash_algo *algo,\n+\t\t\t    char *hex_out,\n+\t\t\t    struct object_id *oid_out)\n+{\n+\tfor (int i = 0; i < len; i++) {\n+\t\tunsigned char c = name[i];\n+\t\tunsigned char val;\n+\t\tif (c >= '0' && c <= '9') {\n+\t\t\tval = c - '0';\n+\t\t} else if (c >= 'a' && c <= 'f') {\n+\t\t\tval = c - 'a' + 10;\n+\t\t} else if (c >= 'A' && c <='F') {\n+\t\t\tval = c - 'A' + 10;\n+\t\t\tc -= 'A' - 'a';\n+\t\t} else {\n+\t\t\treturn -1;\n+\t\t}\n+\n+\t\tif (hex_out)\n+\t\t\thex_out[i] = c;\n+\t\tif (oid_out) {\n+\t\t\tif (!(i & 1))\n+\t\t\t\tval <<= 4;\n+\t\t\toid_out->hash[i >> 1] |= val;\n+\t\t}\n+\t}\n+\n+\tif (hex_out)\n+\t\thex_out[len] = '\\0';\n+\tif (oid_out)\n+\t\toid_out->algo = algo ? hash_algo_by_ptr(algo) : GIT_HASH_UNKNOWN;\n+\n+\treturn 0;\n+}\n+\n static int init_object_disambiguation(struct repository *r,\n \t\t\t\t      const char *name, int len,\n \t\t\t\t      const struct git_hash_algo *algo,\n \t\t\t\t      struct disambiguate_state *ds)\n {\n-\tint i;\n-\n \tif (len < MINIMUM_ABBREV || len > GIT_MAX_HEXSZ)\n \t\treturn -1;\n \n \tmemset(ds, 0, sizeof(*ds));\n \n-\tfor (i = 0; i < len ;i++) {\n-\t\tunsigned char c = name[i];\n-\t\tunsigned char val;\n-\t\tif (c >= '0' && c <= '9')\n-\t\t\tval = c - '0';\n-\t\telse if (c >= 'a' && c <= 'f')\n-\t\t\tval = c - 'a' + 10;\n-\t\telse if (c >= 'A' && c <='F') {\n-\t\t\tval = c - 'A' + 10;\n-\t\t\tc -= 'A' - 'a';\n-\t\t}\n-\t\telse\n-\t\t\treturn -1;\n-\t\tds->hex_pfx[i] = c;\n-\t\tif (!(i & 1))\n-\t\t\tval <<= 4;\n-\t\tds->bin_pfx.hash[i >> 1] |= val;\n-\t}\n+\tif (parse_oid_prefix(name, len, algo, ds->hex_pfx, &ds->bin_pfx) < 0)\n+\t\treturn -1;\n \n \tds->len = len;\n-\tds->hex_pfx[len] = '\\0';\n \tds->repo = r;\n-\tds->bin_pfx.algo = algo ? hash_algo_by_ptr(algo) : GIT_HASH_UNKNOWN;\n \todb_prepare_alternates(r->objects);\n \treturn 0;\n }\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539359","messageId":"20260319-b4-pks-odb-source-abbrev-v1-7-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 07/14] object-name: backend-generic `repo_collect_ambiguous()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:05Z","receivedAt":"2026-03-19T06:53:23Z","isPatch":true,"body":"The function `repo_collect_ambiguous()` is responsible for collecting\nobjects whose IDs match a specific prefix. The information is then\nused to inform the user about which objects they could have meant in\ncase a short object ID is ambiguous.\n\nThe logic to do this uses the object disambiguation infrastructure and\ncalls into backend-specific functions to iterate through loose and\npacked objects. This isn't really required anymore though: all we want\nto do is to enumerate objects that have such a prefix and then append\nthose objects to a `struct oid_array`. This can be trivially achieved\nin a generic way now that `odb_for_each_object()` has learned to yield\nonly objects that much such a prefix.\n\nRefactor the code to use the backend-generic infrastructure instead.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 19 ++++++++++---------\n 1 file changed, 10 insertions(+), 9 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex fd1b010ab3..4c3ace150e 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -448,8 +448,8 @@ static int collect_ambiguous(const struct object_id *oid, void *data)\n \treturn 0;\n }\n \n-static int repo_collect_ambiguous(struct repository *r UNUSED,\n-\t\t\t\t  const struct object_id *oid,\n+static int repo_collect_ambiguous(const struct object_id *oid,\n+\t\t\t\t  struct object_info *oi UNUSED,\n \t\t\t\t  void *data)\n {\n \treturn collect_ambiguous(oid, data);\n@@ -586,18 +586,19 @@ int repo_for_each_abbrev(struct repository *r, const char *prefix,\n \t\t\t const struct git_hash_algo *algo,\n \t\t\t each_abbrev_fn fn, void *cb_data)\n {\n+\tstruct object_id prefix_oid = { 0 };\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = &prefix_oid,\n+\t\t.prefix_hex_len = strlen(prefix),\n+\t};\n \tstruct oid_array collect = OID_ARRAY_INIT;\n-\tstruct disambiguate_state ds;\n \tint ret;\n \n-\tif (init_object_disambiguation(r, prefix, strlen(prefix), algo, &ds) < 0)\n+\tif (parse_oid_prefix(prefix, opts.prefix_hex_len, algo, NULL, &prefix_oid) < 0)\n \t\treturn -1;\n \n-\tds.always_call_fn = 1;\n-\tds.fn = repo_collect_ambiguous;\n-\tds.cb_data = &collect;\n-\tfind_short_object_filename(&ds);\n-\tfind_short_packed_object(&ds);\n+\tif (odb_for_each_object_ext(r->objects, NULL, repo_collect_ambiguous, &collect, &opts) < 0)\n+\t\treturn -1;\n \n \tret = oid_array_for_each_unique(&collect, fn, cb_data);\n \toid_array_clear(&collect);\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539360","messageId":"20260319-b4-pks-odb-source-abbrev-v1-8-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 08/14] object-name: backend-generic `get_short_oid()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:06Z","receivedAt":"2026-03-19T06:53:25Z","isPatch":true,"body":"The function `get_short_oid()` takes as input an abbreviated object ID\nand tries to turn that object ID into the full object ID. This is done\nby iterating through all objects that have the user-provided prefix. If\nthat yields exactly one object we know that the abbreviated object ID is\nunambiguous, otherwise it is ambiguous and we print the list of objects\nthat match the prefix.\n\nWe iterate through all objects with the given prefix by calling both\n`find_short_packed_object()` and `find_short_object_filename()`, which\nis of course specific to the \"files\" backend. But we now have a generic\nway to iterate through objects with a specific prefix.\n\nRefactor the code to use `odb_for_each_object()` instead so that it\nworks with object backends different than the \"files\" backend.\n\nRemove the now-unused `find_short_packed_object()` function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 32 ++++++--------------------------\n 1 file changed, 6 insertions(+), 26 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex 4c3ace150e..7a224ab4af 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -120,28 +120,6 @@ static void find_short_object_filename(struct disambiguate_state *ds)\n \t\todb_source_loose_for_each_object(source, NULL, match_prefix, ds, &opts);\n }\n \n-static void find_short_packed_object(struct disambiguate_state *ds)\n-{\n-\tstruct odb_for_each_object_options opts = {\n-\t\t.prefix = &ds->bin_pfx,\n-\t\t.prefix_hex_len = ds->len,\n-\t};\n-\tstruct odb_source *source;\n-\n-\t/* Skip, unless oids from the storage hash algorithm are wanted */\n-\tif (ds->bin_pfx.algo && (&hash_algos[ds->bin_pfx.algo] != ds->repo->hash_algo))\n-\t\treturn;\n-\n-\todb_prepare_alternates(ds->repo->objects);\n-\tfor (source = ds->repo->objects->sources; source; source = source->next) {\n-\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\n-\t\tpackfile_store_for_each_object(files->packed, NULL, match_prefix, ds, &opts);\n-\t\tif (ds->ambiguous)\n-\t\t\tbreak;\n-\t}\n-}\n-\n static int finish_object_disambiguation(struct disambiguate_state *ds,\n \t\t\t\t\tstruct object_id *oid)\n {\n@@ -499,6 +477,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \t\t\t\t\t struct object_id *oid,\n \t\t\t\t\t unsigned flags)\n {\n+\tstruct odb_for_each_object_options opts = { 0 };\n \tint status;\n \tstruct disambiguate_state ds;\n \tint quietly = !!(flags & GET_OID_QUIETLY);\n@@ -526,8 +505,10 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \telse\n \t\tds.fn = default_disambiguate_hint;\n \n-\tfind_short_object_filename(&ds);\n-\tfind_short_packed_object(&ds);\n+\topts.prefix = &ds.bin_pfx;\n+\topts.prefix_hex_len = ds.len;\n+\n+\todb_for_each_object_ext(r->objects, NULL, match_prefix, &ds, &opts);\n \tstatus = finish_object_disambiguation(&ds, oid);\n \n \t/*\n@@ -537,8 +518,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \t */\n \tif (status == MISSING_OBJECT) {\n \t\todb_reprepare(r->objects);\n-\t\tfind_short_object_filename(&ds);\n-\t\tfind_short_packed_object(&ds);\n+\t\todb_for_each_object_ext(r->objects, NULL, match_prefix, &ds, &opts);\n \t\tstatus = finish_object_disambiguation(&ds, oid);\n \t}\n \n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539361","messageId":"20260319-b4-pks-odb-source-abbrev-v1-9-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 09/14] object-name: merge `update_candidates()` and `match_prefix()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:07Z","receivedAt":"2026-03-19T06:53:28Z","isPatch":true,"body":"There's only a single callsite for `match_prefix()`, and that function\nis a rather trivial wrapper of `update_candidates()`. Merge these two\nfunctions into a single `update_disambiguate_state()` function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 34 ++++++++++++++++++----------------\n 1 file changed, 18 insertions(+), 16 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex 7a224ab4af..f55a332032 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -51,27 +51,31 @@ struct disambiguate_state {\n \tunsigned always_call_fn:1;\n };\n \n-static void update_candidates(struct disambiguate_state *ds, const struct object_id *current)\n+static int update_disambiguate_state(const struct object_id *current,\n+\t\t\t\t     struct object_info *oi UNUSED,\n+\t\t\t\t     void *cb_data)\n {\n+\tstruct disambiguate_state *ds = cb_data;\n+\n \t/* The hash algorithm of current has already been filtered */\n \tif (ds->always_call_fn) {\n \t\tds->ambiguous = ds->fn(ds->repo, current, ds->cb_data) ? 1 : 0;\n-\t\treturn;\n+\t\treturn ds->ambiguous;\n \t}\n \tif (!ds->candidate_exists) {\n \t\t/* this is the first candidate */\n \t\toidcpy(&ds->candidate, current);\n \t\tds->candidate_exists = 1;\n-\t\treturn;\n+\t\treturn 0;\n \t} else if (oideq(&ds->candidate, current)) {\n \t\t/* the same as what we already have seen */\n-\t\treturn;\n+\t\treturn 0;\n \t}\n \n \tif (!ds->fn) {\n \t\t/* cannot disambiguate between ds->candidate and current */\n \t\tds->ambiguous = 1;\n-\t\treturn;\n+\t\treturn ds->ambiguous;\n \t}\n \n \tif (!ds->candidate_checked) {\n@@ -84,7 +88,7 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n \t\t/* discard the candidate; we know it does not satisfy fn */\n \t\toidcpy(&ds->candidate, current);\n \t\tds->candidate_checked = 0;\n-\t\treturn;\n+\t\treturn 0;\n \t}\n \n \t/* if we reach this point, we know ds->candidate satisfies fn */\n@@ -95,17 +99,12 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n \t\t */\n \t\tds->candidate_ok = 0;\n \t\tds->ambiguous = 1;\n+\t\treturn ds->ambiguous;\n \t}\n \n \t/* otherwise, current can be discarded and candidate is still good */\n-}\n \n-static int match_prefix(const struct object_id *oid, struct object_info *oi UNUSED, void *arg)\n-{\n-\tstruct disambiguate_state *ds = arg;\n-\t/* no need to call match_hash, oidtree_each did prefix match */\n-\tupdate_candidates(ds, oid);\n-\treturn ds->ambiguous;\n+\treturn 0;\n }\n \n static void find_short_object_filename(struct disambiguate_state *ds)\n@@ -117,7 +116,8 @@ static void find_short_object_filename(struct disambiguate_state *ds)\n \tstruct odb_source *source;\n \n \tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next)\n-\t\todb_source_loose_for_each_object(source, NULL, match_prefix, ds, &opts);\n+\t\todb_source_loose_for_each_object(source, NULL, update_disambiguate_state,\n+\t\t\t\t\t\t ds, &opts);\n }\n \n static int finish_object_disambiguation(struct disambiguate_state *ds,\n@@ -508,7 +508,8 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \topts.prefix = &ds.bin_pfx;\n \topts.prefix_hex_len = ds.len;\n \n-\todb_for_each_object_ext(r->objects, NULL, match_prefix, &ds, &opts);\n+\todb_for_each_object_ext(r->objects, NULL, update_disambiguate_state,\n+\t\t\t\t&ds, &opts);\n \tstatus = finish_object_disambiguation(&ds, oid);\n \n \t/*\n@@ -518,7 +519,8 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \t */\n \tif (status == MISSING_OBJECT) {\n \t\todb_reprepare(r->objects);\n-\t\todb_for_each_object_ext(r->objects, NULL, match_prefix, &ds, &opts);\n+\t\todb_for_each_object_ext(r->objects, NULL, update_disambiguate_state,\n+\t\t\t\t\t&ds, &opts);\n \t\tstatus = finish_object_disambiguation(&ds, oid);\n \t}\n \n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539362","messageId":"20260319-b4-pks-odb-source-abbrev-v1-10-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 10/14] object-name: abbreviate loose object names without `disambiguate_state`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:08Z","receivedAt":"2026-03-19T06:53:31Z","isPatch":true,"body":"The function `find_short_object_filename()` takes an object ID and\ncomputes the minimum required object name length to make it unique. This\nis done by reusing the object disambiguation infrastructure, where we\niterate through every loose object and then update the disambiguate\nstate one by one.\n\nUltimately, we don't care about the disambiguate state though. It is\nused because this infrastructure knows how to enumerate only those\nobjects that match a given prefix. But now that we have extended the\n`odb_for_each_object()` function to do this for us we have an easier way\nto do this. Consequently, we really only use the disambiguate state now\nto propagate `struct min_abbrev_data`.\n\nRefactor the code and drop this indirection so that we use `struct\nmin_abbrev_data` directly. This also allows us to drop some now-unused\nlogic from the disambiguate infrastructure.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 54 ++++++++++++++++++++----------------------------------\n 1 file changed, 20 insertions(+), 34 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex f55a332032..d82fb49f39 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -48,7 +48,6 @@ struct disambiguate_state {\n \tunsigned candidate_ok:1;\n \tunsigned disambiguate_fn_used:1;\n \tunsigned ambiguous:1;\n-\tunsigned always_call_fn:1;\n };\n \n static int update_disambiguate_state(const struct object_id *current,\n@@ -58,10 +57,6 @@ static int update_disambiguate_state(const struct object_id *current,\n \tstruct disambiguate_state *ds = cb_data;\n \n \t/* The hash algorithm of current has already been filtered */\n-\tif (ds->always_call_fn) {\n-\t\tds->ambiguous = ds->fn(ds->repo, current, ds->cb_data) ? 1 : 0;\n-\t\treturn ds->ambiguous;\n-\t}\n \tif (!ds->candidate_exists) {\n \t\t/* this is the first candidate */\n \t\toidcpy(&ds->candidate, current);\n@@ -107,19 +102,6 @@ static int update_disambiguate_state(const struct object_id *current,\n \treturn 0;\n }\n \n-static void find_short_object_filename(struct disambiguate_state *ds)\n-{\n-\tstruct odb_for_each_object_options opts = {\n-\t\t.prefix = &ds->bin_pfx,\n-\t\t.prefix_hex_len = ds->len,\n-\t};\n-\tstruct odb_source *source;\n-\n-\tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next)\n-\t\todb_source_loose_for_each_object(source, NULL, update_disambiguate_state,\n-\t\t\t\t\t\t ds, &opts);\n-}\n-\n static int finish_object_disambiguation(struct disambiguate_state *ds,\n \t\t\t\t\tstruct object_id *oid)\n {\n@@ -632,11 +614,26 @@ static int extend_abbrev_len(const struct object_id *oid,\n \treturn 0;\n }\n \n-static int repo_extend_abbrev_len(struct repository *r UNUSED,\n-\t\t\t\t  const struct object_id *oid,\n-\t\t\t\t  void *cb_data)\n+static int extend_abbrev_len_loose(const struct object_id *oid,\n+\t\t\t\t   struct object_info *oi UNUSED,\n+\t\t\t\t   void *cb_data)\n {\n-\treturn extend_abbrev_len(oid, cb_data);\n+\tstruct min_abbrev_data *data = cb_data;\n+\textend_abbrev_len(oid, data);\n+\treturn 0;\n+}\n+\n+static void find_abbrev_len_loose(struct min_abbrev_data *mad)\n+{\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = mad->oid,\n+\t\t.prefix_hex_len = mad->cur_len,\n+\t};\n+\tstruct odb_source *source;\n+\n+\tfor (source = mad->repo->objects->sources; source; source = source->next)\n+\t\todb_source_loose_for_each_object(source, NULL, extend_abbrev_len_loose,\n+\t\t\t\t\t\t mad, &opts);\n }\n \n static void find_abbrev_len_for_midx(struct multi_pack_index *m,\n@@ -752,9 +749,7 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n {\n \tconst struct git_hash_algo *algo =\n \t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n-\tstruct disambiguate_state ds;\n \tstruct min_abbrev_data mad;\n-\tstruct object_id oid_ret;\n \tconst unsigned hexsz = algo->hexsz;\n \n \tif (len < 0) {\n@@ -794,16 +789,7 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \tmad.oid = oid;\n \n \tfind_abbrev_len_packed(&mad);\n-\n-\tif (init_object_disambiguation(r, hex, mad.cur_len, algo, &ds) < 0)\n-\t\treturn -1;\n-\n-\tds.fn = repo_extend_abbrev_len;\n-\tds.always_call_fn = 1;\n-\tds.cb_data = (void *)&mad;\n-\n-\tfind_short_object_filename(&ds);\n-\t(void)finish_object_disambiguation(&ds, &oid_ret);\n+\tfind_abbrev_len_loose(&mad);\n \n \thex[mad.cur_len] = 0;\n \treturn mad.cur_len;\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539363","messageId":"20260319-b4-pks-odb-source-abbrev-v1-11-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 11/14] object-name: simplify computing common prefixes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:09Z","receivedAt":"2026-03-19T06:53:33Z","isPatch":true,"body":"The function `extend_abbrev_len()` computes the length of common hex\ncharacters between two object IDs. This is done by:\n\n  - Making the caller provide the `hex` string for the needle object ID.\n\n  - Comparing every hex position of the haystack object ID with\n    `get_hex_char_from_oid()`.\n\nTurning the binary representation into hex first is roundabout though:\nwe can simply compare the binary representation and give some special\nattention to the final nibble.\n\nIntroduce a new function `oid_common_prefix_hexlen()` that does exactly\nthis and refactor the code to use the new function. This allows us to\ndrop the `struct min_abbrev_data::hex` field. Furthermore, this function\nwill be used in by some other callsites in subsequent commits.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n hash.c        | 18 ++++++++++++++++++\n hash.h        |  3 +++\n object-name.c | 23 +++--------------------\n 3 files changed, 24 insertions(+), 20 deletions(-)\n\ndiff --git a/hash.c b/hash.c\nindex 553f2008ea..e925b9754e 100644\n--- a/hash.c\n+++ b/hash.c\n@@ -317,3 +317,21 @@ const struct git_hash_algo *unsafe_hash_algo(const struct git_hash_algo *algop)\n \t/* Otherwise use the default one. */\n \treturn algop;\n }\n+\n+unsigned oid_common_prefix_hexlen(const struct object_id *a,\n+\t\t\t\t  const struct object_id *b)\n+{\n+\tunsigned rawsz = hash_algos[a->algo].rawsz;\n+\n+\tfor (unsigned i = 0; i < rawsz; i++) {\n+\t\tif (a->hash[i] == b->hash[i])\n+\t\t\tcontinue;\n+\n+\t\tif ((a->hash[i] ^ b->hash[i]) & 0xf0)\n+\t\t\treturn i * 2;\n+\t\telse\n+\t\t\treturn i * 2 + 1;\n+\t}\n+\n+\treturn rawsz * 2;\n+}\ndiff --git a/hash.h b/hash.h\nindex d51efce1d3..c082a53c9a 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -396,6 +396,9 @@ static inline int oideq(const struct object_id *oid1, const struct object_id *oi\n \treturn !memcmp(oid1->hash, oid2->hash, GIT_MAX_RAWSZ);\n }\n \n+unsigned oid_common_prefix_hexlen(const struct object_id *a,\n+\t\t\t\t  const struct object_id *b);\n+\n static inline void oidcpy(struct object_id *dst, const struct object_id *src)\n {\n \tmemcpy(dst->hash, src->hash, GIT_MAX_RAWSZ);\ndiff --git a/object-name.c b/object-name.c\nindex d82fb49f39..32e9c23e40 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -585,32 +585,16 @@ static unsigned msb(unsigned long val)\n struct min_abbrev_data {\n \tunsigned int init_len;\n \tunsigned int cur_len;\n-\tchar *hex;\n \tstruct repository *repo;\n \tconst struct object_id *oid;\n };\n \n-static inline char get_hex_char_from_oid(const struct object_id *oid,\n-\t\t\t\t\t unsigned int pos)\n-{\n-\tstatic const char hex[] = \"0123456789abcdef\";\n-\n-\tif ((pos & 1) == 0)\n-\t\treturn hex[oid->hash[pos >> 1] >> 4];\n-\telse\n-\t\treturn hex[oid->hash[pos >> 1] & 0xf];\n-}\n-\n static int extend_abbrev_len(const struct object_id *oid,\n \t\t\t     struct min_abbrev_data *mad)\n {\n-\tunsigned int i = mad->init_len;\n-\twhile (mad->hex[i] && mad->hex[i] == get_hex_char_from_oid(oid, i))\n-\t\ti++;\n-\n-\tif (mad->hex[i] && i >= mad->cur_len)\n-\t\tmad->cur_len = i + 1;\n-\n+\tunsigned len = oid_common_prefix_hexlen(oid, mad->oid);\n+\tif (len != hash_algos[oid->algo].hexsz && len >= mad->cur_len)\n+\t\tmad->cur_len = len + 1;\n \treturn 0;\n }\n \n@@ -785,7 +769,6 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \tmad.repo = r;\n \tmad.init_len = len;\n \tmad.cur_len = len;\n-\tmad.hex = hex;\n \tmad.oid = oid;\n \n \tfind_abbrev_len_packed(&mad);\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539364","messageId":"20260319-b4-pks-odb-source-abbrev-v1-12-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 12/14] object-name: move logic to compute loose abbreviation length","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:10Z","receivedAt":"2026-03-19T06:53:35Z","isPatch":true,"body":"The function `repo_find_unique_abbrev_r()` takes as input an object ID\nas well as a minimum object ID length and returns the minimum required\nprefix to make the object ID unique.\n\nThe logic that computes the abbreviation length for loose objects is\ndeeply tied to the loose object storage format. As such, it would fail\nin case a different object storage format was used.\n\nPrepare for making this logic generic to the backend by moving the logic\ninto a new `odb_source_loose_find_abbrev_len()` function that is part of\n\"object-file.c\".\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-file.c | 38 ++++++++++++++++++++++++++++++++++++++\n object-file.h | 12 ++++++++++++\n object-name.c | 27 ++++-----------------------\n 3 files changed, 54 insertions(+), 23 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 8a9e68a768..35be7e58cb 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1951,6 +1951,44 @@ int odb_source_loose_count_objects(struct odb_source *source,\n \treturn ret;\n }\n \n+struct find_abbrev_len_data {\n+\tconst struct object_id *oid;\n+\tunsigned len;\n+};\n+\n+static int find_abbrev_len_cb(const struct object_id *oid,\n+\t\t\t      struct object_info *oi UNUSED,\n+\t\t\t      void *cb_data)\n+{\n+\tstruct find_abbrev_len_data *data = cb_data;\n+\tunsigned len = oid_common_prefix_hexlen(oid, data->oid);\n+\tif (len != hash_algos[oid->algo].hexsz && len >= data->len)\n+\t\tdata->len = len + 1;\n+\treturn 0;\n+}\n+\n+int odb_source_loose_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     unsigned min_len,\n+\t\t\t\t     unsigned *out)\n+{\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = oid,\n+\t\t.prefix_hex_len = min_len,\n+\t};\n+\tstruct find_abbrev_len_data data = {\n+\t\t.oid = oid,\n+\t\t.len = min_len,\n+\t};\n+\tint ret;\n+\n+\tret = odb_source_loose_for_each_object(source, NULL, find_abbrev_len_cb,\n+\t\t\t\t\t       &data, &opts);\n+\t*out = data.len;\n+\n+\treturn ret;\n+}\n+\n static int append_loose_object(const struct object_id *oid,\n \t\t\t       const char *path UNUSED,\n \t\t\t       void *data)\ndiff --git a/object-file.h b/object-file.h\nindex f11ad58f6c..3686f182e4 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -146,6 +146,18 @@ int odb_source_loose_count_objects(struct odb_source *source,\n \t\t\t\t   enum odb_count_objects_flags flags,\n \t\t\t\t   unsigned long *out);\n \n+/*\n+ * Find the shortest unique prefix for the given object ID, where `min_len` is\n+ * the minimum length that the prefix should have.\n+ *\n+ * Returns 0 on success, in which case the computed length will be written to\n+ * `out`. Otherwise, a negative error code is returned.\n+ */\n+int odb_source_loose_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     unsigned min_len,\n+\t\t\t\t     unsigned *out);\n+\n /**\n  * format_object_header() is a thin wrapper around s xsnprintf() that\n  * writes the initial \"<type> <obj-len>\" part of the loose object\ndiff --git a/object-name.c b/object-name.c\nindex 32e9c23e40..4e21dbfa97 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -598,28 +598,6 @@ static int extend_abbrev_len(const struct object_id *oid,\n \treturn 0;\n }\n \n-static int extend_abbrev_len_loose(const struct object_id *oid,\n-\t\t\t\t   struct object_info *oi UNUSED,\n-\t\t\t\t   void *cb_data)\n-{\n-\tstruct min_abbrev_data *data = cb_data;\n-\textend_abbrev_len(oid, data);\n-\treturn 0;\n-}\n-\n-static void find_abbrev_len_loose(struct min_abbrev_data *mad)\n-{\n-\tstruct odb_for_each_object_options opts = {\n-\t\t.prefix = mad->oid,\n-\t\t.prefix_hex_len = mad->cur_len,\n-\t};\n-\tstruct odb_source *source;\n-\n-\tfor (source = mad->repo->objects->sources; source; source = source->next)\n-\t\todb_source_loose_for_each_object(source, NULL, extend_abbrev_len_loose,\n-\t\t\t\t\t\t mad, &opts);\n-}\n-\n static void find_abbrev_len_for_midx(struct multi_pack_index *m,\n \t\t\t\t     struct min_abbrev_data *mad)\n {\n@@ -772,7 +750,10 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \tmad.oid = oid;\n \n \tfind_abbrev_len_packed(&mad);\n-\tfind_abbrev_len_loose(&mad);\n+\n+\todb_prepare_alternates(r->objects);\n+\tfor (struct odb_source *s = r->objects->sources; s; s = s->next)\n+\t\todb_source_loose_find_abbrev_len(s, mad.oid, mad.cur_len, &mad.cur_len);\n \n \thex[mad.cur_len] = 0;\n \treturn mad.cur_len;\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539365","messageId":"20260319-b4-pks-odb-source-abbrev-v1-13-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 13/14] object-file: move logic to compute packed abbreviation length","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:11Z","receivedAt":"2026-03-19T06:53:38Z","isPatch":true,"body":"Same as the preceding commit, move the logic that computes the minimum\nrequired prefix length to make a given object ID unique for the packfile\nstore into a new function `packfile_store_find_abbrev_len()` that is\npart of \"packfile.c\". This prepares for making the logic fully generic\nvia pluggable object databases.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 135 ++++++----------------------------------------------------\n packfile.c    | 111 +++++++++++++++++++++++++++++++++++++++++++++++\n packfile.h    |   5 +++\n 3 files changed, 128 insertions(+), 123 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex 4e21dbfa97..bb2294a193 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -582,115 +582,6 @@ static unsigned msb(unsigned long val)\n \treturn r;\n }\n \n-struct min_abbrev_data {\n-\tunsigned int init_len;\n-\tunsigned int cur_len;\n-\tstruct repository *repo;\n-\tconst struct object_id *oid;\n-};\n-\n-static int extend_abbrev_len(const struct object_id *oid,\n-\t\t\t     struct min_abbrev_data *mad)\n-{\n-\tunsigned len = oid_common_prefix_hexlen(oid, mad->oid);\n-\tif (len != hash_algos[oid->algo].hexsz && len >= mad->cur_len)\n-\t\tmad->cur_len = len + 1;\n-\treturn 0;\n-}\n-\n-static void find_abbrev_len_for_midx(struct multi_pack_index *m,\n-\t\t\t\t     struct min_abbrev_data *mad)\n-{\n-\tfor (; m; m = m->base_midx) {\n-\t\tint match = 0;\n-\t\tuint32_t num, first = 0;\n-\t\tstruct object_id oid;\n-\t\tconst struct object_id *mad_oid;\n-\n-\t\tif (!m->num_objects)\n-\t\t\tcontinue;\n-\n-\t\tnum = m->num_objects + m->num_objects_in_base;\n-\t\tmad_oid = mad->oid;\n-\t\tmatch = bsearch_one_midx(mad_oid, m, &first);\n-\n-\t\t/*\n-\t\t * first is now the position in the packfile where we\n-\t\t * would insert mad->hash if it does not exist (or the\n-\t\t * position of mad->hash if it does exist). Hence, we\n-\t\t * consider a maximum of two objects nearby for the\n-\t\t * abbreviation length.\n-\t\t */\n-\t\tmad->init_len = 0;\n-\t\tif (!match) {\n-\t\t\tif (nth_midxed_object_oid(&oid, m, first))\n-\t\t\t\textend_abbrev_len(&oid, mad);\n-\t\t} else if (first < num - 1) {\n-\t\t\tif (nth_midxed_object_oid(&oid, m, first + 1))\n-\t\t\t\textend_abbrev_len(&oid, mad);\n-\t\t}\n-\t\tif (first > 0) {\n-\t\t\tif (nth_midxed_object_oid(&oid, m, first - 1))\n-\t\t\t\textend_abbrev_len(&oid, mad);\n-\t\t}\n-\t\tmad->init_len = mad->cur_len;\n-\t}\n-}\n-\n-static void find_abbrev_len_for_pack(struct packed_git *p,\n-\t\t\t\t     struct min_abbrev_data *mad)\n-{\n-\tint match = 0;\n-\tuint32_t num, first = 0;\n-\tstruct object_id oid;\n-\tconst struct object_id *mad_oid;\n-\n-\tif (p->multi_pack_index)\n-\t\treturn;\n-\n-\tif (open_pack_index(p) || !p->num_objects)\n-\t\treturn;\n-\n-\tnum = p->num_objects;\n-\tmad_oid = mad->oid;\n-\tmatch = bsearch_pack(mad_oid, p, &first);\n-\n-\t/*\n-\t * first is now the position in the packfile where we would insert\n-\t * mad->hash if it does not exist (or the position of mad->hash if\n-\t * it does exist). Hence, we consider a maximum of two objects\n-\t * nearby for the abbreviation length.\n-\t */\n-\tmad->init_len = 0;\n-\tif (!match) {\n-\t\tif (!nth_packed_object_id(&oid, p, first))\n-\t\t\textend_abbrev_len(&oid, mad);\n-\t} else if (first < num - 1) {\n-\t\tif (!nth_packed_object_id(&oid, p, first + 1))\n-\t\t\textend_abbrev_len(&oid, mad);\n-\t}\n-\tif (first > 0) {\n-\t\tif (!nth_packed_object_id(&oid, p, first - 1))\n-\t\t\textend_abbrev_len(&oid, mad);\n-\t}\n-\tmad->init_len = mad->cur_len;\n-}\n-\n-static void find_abbrev_len_packed(struct min_abbrev_data *mad)\n-{\n-\tstruct packed_git *p;\n-\n-\todb_prepare_alternates(mad->repo->objects);\n-\tfor (struct odb_source *source = mad->repo->objects->sources; source; source = source->next) {\n-\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n-\t\tif (m)\n-\t\t\tfind_abbrev_len_for_midx(m, mad);\n-\t}\n-\n-\trepo_for_each_pack(mad->repo, p)\n-\t\tfind_abbrev_len_for_pack(p, mad);\n-}\n-\n void strbuf_repo_add_unique_abbrev(struct strbuf *sb, struct repository *repo,\n \t\t\t\t   const struct object_id *oid, int abbrev_len)\n {\n@@ -707,14 +598,14 @@ void strbuf_add_unique_abbrev(struct strbuf *sb, const struct object_id *oid,\n }\n \n int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n-\t\t\t      const struct object_id *oid, int len)\n+\t\t\t      const struct object_id *oid, int min_len)\n {\n \tconst struct git_hash_algo *algo =\n \t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n-\tstruct min_abbrev_data mad;\n \tconst unsigned hexsz = algo->hexsz;\n+\tunsigned len;\n \n-\tif (len < 0) {\n+\tif (min_len < 0) {\n \t\tunsigned long count;\n \n \t\tif (odb_count_objects(r->objects, ODB_COUNT_OBJECTS_APPROXIMATE, &count) < 0)\n@@ -738,25 +629,23 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \t\t */\n \t\tif (len < FALLBACK_DEFAULT_ABBREV)\n \t\t\tlen = FALLBACK_DEFAULT_ABBREV;\n+\t} else {\n+\t\tlen = min_len;\n \t}\n \n \toid_to_hex_r(hex, oid);\n \tif (len >= hexsz || !len)\n \t\treturn hexsz;\n \n-\tmad.repo = r;\n-\tmad.init_len = len;\n-\tmad.cur_len = len;\n-\tmad.oid = oid;\n-\n-\tfind_abbrev_len_packed(&mad);\n-\n \todb_prepare_alternates(r->objects);\n-\tfor (struct odb_source *s = r->objects->sources; s; s = s->next)\n-\t\todb_source_loose_find_abbrev_len(s, mad.oid, mad.cur_len, &mad.cur_len);\n+\tfor (struct odb_source *s = r->objects->sources; s; s = s->next) {\n+\t\tstruct odb_source_files *files = odb_source_files_downcast(s);\n+\t\tpackfile_store_find_abbrev_len(files->packed, oid, len, &len);\n+\t\todb_source_loose_find_abbrev_len(s, oid, len, &len);\n+\t}\n \n-\thex[mad.cur_len] = 0;\n-\treturn mad.cur_len;\n+\thex[len] = 0;\n+\treturn len;\n }\n \n const char *repo_find_unique_abbrev(struct repository *r,\ndiff --git a/packfile.c b/packfile.c\nindex 2539a371c1..ee9c7ea1d1 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2597,6 +2597,117 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \treturn ret;\n }\n \n+static int extend_abbrev_len(const struct object_id *a,\n+\t\t\t     const struct object_id *b,\n+\t\t\t     unsigned *out)\n+{\n+\tunsigned len = oid_common_prefix_hexlen(a, b);\n+\tif (len != hash_algos[a->algo].hexsz && len >= *out)\n+\t\t*out = len + 1;\n+\treturn 0;\n+}\n+\n+static void find_abbrev_len_for_midx(struct multi_pack_index *m,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     unsigned min_len,\n+\t\t\t\t     unsigned *out)\n+{\n+\tunsigned len = min_len;\n+\n+\tfor (; m; m = m->base_midx) {\n+\t\tint match = 0;\n+\t\tuint32_t num, first = 0;\n+\t\tstruct object_id found_oid;\n+\n+\t\tif (!m->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tnum = m->num_objects + m->num_objects_in_base;\n+\t\tmatch = bsearch_one_midx(oid, m, &first);\n+\n+\t\t/*\n+\t\t * first is now the position in the packfile where we\n+\t\t * would insert the object ID if it does not exist (or the\n+\t\t * position of the object ID if it does exist). Hence, we\n+\t\t * consider a maximum of two objects nearby for the\n+\t\t * abbreviation length.\n+\t\t */\n+\n+\t\tif (!match) {\n+\t\t\tif (nth_midxed_object_oid(&found_oid, m, first))\n+\t\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t\t} else if (first < num - 1) {\n+\t\t\tif (nth_midxed_object_oid(&found_oid, m, first + 1))\n+\t\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t\t}\n+\t\tif (first > 0) {\n+\t\t\tif (nth_midxed_object_oid(&found_oid, m, first - 1))\n+\t\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t\t}\n+\t}\n+\n+\t*out = len;\n+}\n+\n+static void find_abbrev_len_for_pack(struct packed_git *p,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     unsigned min_len,\n+\t\t\t\t     unsigned *out)\n+{\n+\tint match;\n+\tuint32_t num, first = 0;\n+\tstruct object_id found_oid;\n+\tunsigned len = min_len;\n+\n+\tnum = p->num_objects;\n+\tmatch = bsearch_pack(oid, p, &first);\n+\n+\t/*\n+\t * first is now the position in the packfile where we would insert\n+\t * the object ID if it does not exist (or the position of mad->hash if\n+\t * it does exist). Hence, we consider a maximum of two objects\n+\t * nearby for the abbreviation length.\n+\t */\n+\tif (!match) {\n+\t\tif (!nth_packed_object_id(&found_oid, p, first))\n+\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t} else if (first < num - 1) {\n+\t\tif (!nth_packed_object_id(&found_oid, p, first + 1))\n+\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t}\n+\tif (first > 0) {\n+\t\tif (!nth_packed_object_id(&found_oid, p, first - 1))\n+\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t}\n+\n+\t*out = len;\n+}\n+\n+int packfile_store_find_abbrev_len(struct packfile_store *store,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   unsigned min_len,\n+\t\t\t\t   unsigned *out)\n+{\n+\tstruct packfile_list_entry *e;\n+\tstruct multi_pack_index *m;\n+\n+\tm = get_multi_pack_index(store->source);\n+\tif (m)\n+\t\tfind_abbrev_len_for_midx(m, oid, min_len, &min_len);\n+\n+\tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n+\t\tif (e->pack->multi_pack_index)\n+\t\t\tcontinue;\n+\t\tif (open_pack_index(e->pack) || !e->pack->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tfind_abbrev_len_for_pack(e->pack, oid, min_len, &min_len);\n+\t}\n+\n+\t*out = min_len;\n+\treturn 0;\n+}\n+\n struct add_promisor_object_data {\n \tstruct repository *repo;\n \tstruct oidset *set;\ndiff --git a/packfile.h b/packfile.h\nindex fa41dfda38..45b35973f0 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -369,6 +369,11 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \t\t\t\t   void *cb_data,\n \t\t\t\t   const struct odb_for_each_object_options *opts);\n \n+int packfile_store_find_abbrev_len(struct packfile_store *store,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   unsigned min_len,\n+\t\t\t\t   unsigned *out);\n+\n /* A hook to report invalid files in pack directory */\n #define PACKDIR_FILE_PACK 1\n #define PACKDIR_FILE_IDX 2\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539366","messageId":"20260319-b4-pks-odb-source-abbrev-v1-14-5ddebad292b0@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH 14/14] odb: introduce generic `odb_find_abbrev_len()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T06:53:12Z","receivedAt":"2026-03-19T06:53:41Z","isPatch":true,"body":"Introduce a new generic `odb_find_abbrev_len()` function as well as\nsource-specific callback functions. This makes the logic to compute the\nrequired prefix length to make a given object unique fully pluggable.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c      | 57 +++---------------------------------------\n odb.c              | 73 ++++++++++++++++++++++++++++++++++++++++++++++++++++++\n odb.h              | 16 ++++++++++++\n odb/source-files.c | 25 +++++++++++++++++++\n odb/source.h       | 24 ++++++++++++++++++\n 5 files changed, 142 insertions(+), 53 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex bb2294a193..f6e1f29e1f 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -15,10 +15,9 @@\n #include \"refs.h\"\n #include \"remote.h\"\n #include \"dir.h\"\n+#include \"odb.h\"\n #include \"oid-array.h\"\n-#include \"packfile.h\"\n #include \"pretty.h\"\n-#include \"object-file.h\"\n #include \"read-cache-ll.h\"\n #include \"repo-settings.h\"\n #include \"repository.h\"\n@@ -569,19 +568,6 @@ int repo_for_each_abbrev(struct repository *r, const char *prefix,\n \treturn ret;\n }\n \n-/*\n- * Return the slot of the most-significant bit set in \"val\". There are various\n- * ways to do this quickly with fls() or __builtin_clzl(), but speed is\n- * probably not a big deal here.\n- */\n-static unsigned msb(unsigned long val)\n-{\n-\tunsigned r = 0;\n-\twhile (val >>= 1)\n-\t\tr++;\n-\treturn r;\n-}\n-\n void strbuf_repo_add_unique_abbrev(struct strbuf *sb, struct repository *repo,\n \t\t\t\t   const struct object_id *oid, int abbrev_len)\n {\n@@ -602,49 +588,14 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n {\n \tconst struct git_hash_algo *algo =\n \t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n-\tconst unsigned hexsz = algo->hexsz;\n \tunsigned len;\n \n-\tif (min_len < 0) {\n-\t\tunsigned long count;\n-\n-\t\tif (odb_count_objects(r->objects, ODB_COUNT_OBJECTS_APPROXIMATE, &count) < 0)\n-\t\t\tcount = 0;\n-\n-\t\t/*\n-\t\t * Add one because the MSB only tells us the highest bit set,\n-\t\t * not including the value of all the _other_ bits (so \"15\"\n-\t\t * is only one off of 2^4, but the MSB is the 3rd bit.\n-\t\t */\n-\t\tlen = msb(count) + 1;\n-\t\t/*\n-\t\t * We now know we have on the order of 2^len objects, which\n-\t\t * expects a collision at 2^(len/2). But we also care about hex\n-\t\t * chars, not bits, and there are 4 bits per hex. So all\n-\t\t * together we need to divide by 2 and round up.\n-\t\t */\n-\t\tlen = DIV_ROUND_UP(len, 2);\n-\t\t/*\n-\t\t * For very small repos, we stick with our regular fallback.\n-\t\t */\n-\t\tif (len < FALLBACK_DEFAULT_ABBREV)\n-\t\t\tlen = FALLBACK_DEFAULT_ABBREV;\n-\t} else {\n-\t\tlen = min_len;\n-\t}\n+\tif (odb_find_abbrev_len(r->objects, oid, min_len, &len) < 0)\n+\t\tlen = algo->hexsz;\n \n \toid_to_hex_r(hex, oid);\n-\tif (len >= hexsz || !len)\n-\t\treturn hexsz;\n-\n-\todb_prepare_alternates(r->objects);\n-\tfor (struct odb_source *s = r->objects->sources; s; s = s->next) {\n-\t\tstruct odb_source_files *files = odb_source_files_downcast(s);\n-\t\tpackfile_store_find_abbrev_len(files->packed, oid, len, &len);\n-\t\todb_source_loose_find_abbrev_len(s, oid, len, &len);\n-\t}\n-\n \thex[len] = 0;\n+\n \treturn len;\n }\n \ndiff --git a/odb.c b/odb.c\nindex 3019957b87..3f94a53df1 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -12,6 +12,7 @@\n #include \"midx.h\"\n #include \"object-file-convert.h\"\n #include \"object-file.h\"\n+#include \"object-name.h\"\n #include \"odb.h\"\n #include \"packfile.h\"\n #include \"path.h\"\n@@ -964,6 +965,78 @@ int odb_count_objects(struct object_database *odb,\n \treturn ret;\n }\n \n+/*\n+ * Return the slot of the most-significant bit set in \"val\". There are various\n+ * ways to do this quickly with fls() or __builtin_clzl(), but speed is\n+ * probably not a big deal here.\n+ */\n+static unsigned msb(unsigned long val)\n+{\n+\tunsigned r = 0;\n+\twhile (val >>= 1)\n+\t\tr++;\n+\treturn r;\n+}\n+\n+int odb_find_abbrev_len(struct object_database *odb,\n+\t\t\tconst struct object_id *oid,\n+\t\t\tint min_length,\n+\t\t\tunsigned *out)\n+{\n+\tconst struct git_hash_algo *algo =\n+\t\toid->algo ? &hash_algos[oid->algo] : odb->repo->hash_algo;\n+\tconst unsigned hexsz = algo->hexsz;\n+\tunsigned len;\n+\tint ret;\n+\n+\tif (min_length < 0) {\n+\t\tunsigned long count;\n+\n+\t\tif (odb_count_objects(odb, ODB_COUNT_OBJECTS_APPROXIMATE, &count) < 0)\n+\t\t\tcount = 0;\n+\n+\t\t/*\n+\t\t * Add one because the MSB only tells us the highest bit set,\n+\t\t * not including the value of all the _other_ bits (so \"15\"\n+\t\t * is only one off of 2^4, but the MSB is the 3rd bit.\n+\t\t */\n+\t\tlen = msb(count) + 1;\n+\t\t/*\n+\t\t * We now know we have on the order of 2^len objects, which\n+\t\t * expects a collision at 2^(len/2). But we also care about hex\n+\t\t * chars, not bits, and there are 4 bits per hex. So all\n+\t\t * together we need to divide by 2 and round up.\n+\t\t */\n+\t\tlen = DIV_ROUND_UP(len, 2);\n+\t\t/*\n+\t\t * For very small repos, we stick with our regular fallback.\n+\t\t */\n+\t\tif (len < FALLBACK_DEFAULT_ABBREV)\n+\t\t\tlen = FALLBACK_DEFAULT_ABBREV;\n+\t} else {\n+\t\tlen = min_length;\n+\t}\n+\n+\tif (len >= hexsz || !len) {\n+\t\t*out = hexsz;\n+\t\tret = 0;\n+\t\tgoto out;\n+\t}\n+\n+\todb_prepare_alternates(odb);\n+\tfor (struct odb_source *source = odb->sources; source; source = source->next) {\n+\t\tret = odb_source_find_abbrev_len(source, oid, len, &len);\n+\t\tif (ret)\n+\t\t\tgoto out;\n+\t}\n+\n+\tret = 0;\n+\t*out = len;\n+\n+out:\n+\treturn ret;\n+}\n+\n void odb_assert_oid_type(struct object_database *odb,\n \t\t\t const struct object_id *oid, enum object_type expect)\n {\ndiff --git a/odb.h b/odb.h\nindex e80fd8f7ab..984bafca9d 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -545,6 +545,22 @@ int odb_count_objects(struct object_database *odb,\n \t\t      enum odb_count_objects_flags flags,\n \t\t      unsigned long *out);\n \n+/*\n+ * Given an object ID, find the minimum required length required to make the\n+ * object ID unique across the whole object database.\n+ *\n+ * The `min_len` determines the minimum abbreviated length that'll be returned\n+ * by this function. If `min_len < 0`, then the function will set a sensible\n+ * default minimum abbreviation length.\n+ *\n+ * Returns 0 on success, a negative error code otherwise. The computed length\n+ * will be assigned to `*out`.\n+ */\n+int odb_find_abbrev_len(struct object_database *odb,\n+\t\t\tconst struct object_id *oid,\n+\t\t\tint min_len,\n+\t\t\tunsigned *out);\n+\n enum {\n \t/*\n \t * By default, `odb_write_object()` does not actually write anything\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex e90bb689bb..76797569de 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -122,6 +122,30 @@ static int odb_source_files_count_objects(struct odb_source *source,\n \treturn ret;\n }\n \n+static int odb_source_files_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t\t    const struct object_id *oid,\n+\t\t\t\t\t    unsigned min_len,\n+\t\t\t\t\t    unsigned *out)\n+{\n+\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\tunsigned len = min_len;\n+\tint ret;\n+\n+\tret = packfile_store_find_abbrev_len(files->packed, oid, len, &len);\n+\tif (ret < 0)\n+\t\tgoto out;\n+\n+\tret = odb_source_loose_find_abbrev_len(source, oid, len, &len);\n+\tif (ret < 0)\n+\t\tgoto out;\n+\n+\t*out = len;\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n static int odb_source_files_freshen_object(struct odb_source *source,\n \t\t\t\t\t   const struct object_id *oid)\n {\n@@ -250,6 +274,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\n \tfiles->base.for_each_object = odb_source_files_for_each_object;\n \tfiles->base.count_objects = odb_source_files_count_objects;\n+\tfiles->base.find_abbrev_len = odb_source_files_find_abbrev_len;\n \tfiles->base.freshen_object = odb_source_files_freshen_object;\n \tfiles->base.write_object = odb_source_files_write_object;\n \tfiles->base.write_object_stream = odb_source_files_write_object_stream;\ndiff --git a/odb/source.h b/odb/source.h\nindex ee5d6ed530..a9d7d0b96f 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -157,6 +157,18 @@ struct odb_source {\n \t\t\t     enum odb_count_objects_flags flags,\n \t\t\t     unsigned long *out);\n \n+\t/*\n+\t * This callback is expected to find the minimum required length to\n+\t * make the given object ID unique.\n+\t *\n+\t * The callback is expected to return a negative error code in case it\n+\t * failed, 0 otherwise.\n+\t */\n+\tint (*find_abbrev_len)(struct odb_source *source,\n+\t\t\t       const struct object_id *oid,\n+\t\t\t       unsigned min_length,\n+\t\t\t       unsigned *out);\n+\n \t/*\n \t * This callback is expected to freshen the given object so that its\n \t * last access time is set to the current time. This is used to ensure\n@@ -360,6 +372,18 @@ static inline int odb_source_count_objects(struct odb_source *source,\n \treturn source->count_objects(source, flags, out);\n }\n \n+/*\n+ * Determine the minimum required length to make the given object ID unique in\n+ * the given source. Returns 0 on success, a negative error code otherwise.\n+ */\n+static inline int odb_source_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t\t     unsigned min_len,\n+\t\t\t\t\t     unsigned *out)\n+{\n+\treturn source->find_abbrev_len(source, oid, min_len, out);\n+}\n+\n /*\n  * Freshen an object in the object database by updating its timestamp.\n  * Returns 1 in case the object has been freshened, 0 in case the object does\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539384","messageId":"xmqqse9vnbgb.fsf@gitster.g","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-3-5ddebad292b0@pks.im","subject":"Re: [PATCH 03/14] odb: introduce `struct odb_for_each_object_options`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-19T14:25:08Z","receivedAt":"2026-03-19T14:25:12Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> While we could of course add the prefix to the function signature, or\n> alternative introduce a new function, both of these options don't really\n> seem to be that sensible.\n\n\"alternative\" -> \"alternatigvely\"?\n\n> Instead, introduce a new `struct odb_for_each_object_options` that can\n> be passed to a new `odb_for_each_object_ext()` function. Splice through\n> the options structure into the respective object database sources.\n\nA lot of churn, but we only need to suffer once and reap a lot of\nbenefit later, I guess ;-).\n"},{"id":"539385","messageId":"xmqqo6kjnbdz.fsf@gitster.g","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-7-5ddebad292b0@pks.im","subject":"Re: [PATCH 07/14] object-name: backend-generic `repo_collect_ambiguous()`","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-19T14:26:32Z","receivedAt":"2026-03-19T14:26:35Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> those objects to a `struct oid_array`. This can be trivially achieved\n> in a generic way now that `odb_for_each_object()` has learned to yield\n> only objects that much such a prefix.\n\n\"much\" -> \"match\"?\n\n"},{"id":"539386","messageId":"abwPR1NgOShKXh8P@pks.im","threadId":"65299","inReplyTo":"xmqqse9vnbgb.fsf@gitster.g","subject":"Re: [PATCH 03/14] odb: introduce `struct odb_for_each_object_options`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T14:59:19Z","receivedAt":"2026-03-19T14:59:25Z","isPatch":true,"body":"On Thu, Mar 19, 2026 at 07:25:08AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > While we could of course add the prefix to the function signature, or\n> > alternative introduce a new function, both of these options don't really\n> > seem to be that sensible.\n> \n> \"alternative\" -> \"alternatigvely\"?\n\nI prefer \"alternatively\" :) Will fix.\n\n> > Instead, introduce a new `struct odb_for_each_object_options` that can\n> > be passed to a new `odb_for_each_object_ext()` function. Splice through\n> > the options structure into the respective object database sources.\n> \n> A lot of churn, but we only need to suffer once and reap a lot of\n> benefit later, I guess ;-).\n\nRight, that's the idea. I also got the intent to eventually support\nobject filters in `odb_for_each_object_ext()`, which will be required\nfor example by git-cat-file(1). This would require splicing through\nanother parameter, but with this change here it will only require us to\nadd another new field to the options structure.\n\nPatrick\n"},{"id":"539387","messageId":"abwPTMKSmxb61Od0@pks.im","threadId":"65299","inReplyTo":"xmqqo6kjnbdz.fsf@gitster.g","subject":"Re: [PATCH 07/14] object-name: backend-generic `repo_collect_ambiguous()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-19T14:59:24Z","receivedAt":"2026-03-19T14:59:29Z","isPatch":true,"body":"On Thu, Mar 19, 2026 at 07:26:32AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > those objects to a `struct oid_array`. This can be trivially achieved\n> > in a generic way now that `odb_for_each_object()` has learned to yield\n> > only objects that much such a prefix.\n> \n> \"much\" -> \"match\"?\n\nIndeed. Thanks!\n\nPatrick\n"},{"id":"539397","messageId":"xmqqeclfn6nn.fsf@gitster.g","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-1-5ddebad292b0@pks.im","subject":"Re: [PATCH 01/14] oidtree: modernize the code a bit","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-19T16:08:44Z","receivedAt":"2026-03-19T16:08:47Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> +void oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n> +\t\t  size_t prefix_hex_len, oidtree_each_cb cb, void *cb_data)\n>  {\n> -\tsize_t klen = oidhexsz / 2;\n> -\tstruct oidtree_iter_data x = { 0 };\n> -\tassert(oidhexsz <= GIT_MAX_HEXSZ);\n> -\n> -\tx.fn = fn;\n> -\tx.arg = arg;\n> -\tx.algo = oid->algo;\n> -\tif (oidhexsz & 1) {\n> -\t\tx.last_byte = oid->hash[klen];\n> -\t\tx.last_nibble_at = &klen;\n> +\tstruct oidtree_each_data data = {\n> +\t\t.cb = cb,\n> +\t\t.cb_data = cb_data,\n> +\t\t.algo = prefix->algo,\n> +\t};\n> +\tsize_t klen = prefix_hex_len / 2;\n> +\tassert(prefix_hex_len <= GIT_MAX_HEXSZ);\n\nI know the original also used GIT_MAX_HEXSZ to clamp the length for\nsanity, but because we know what algorithm is in use, I wonder if we\nwant to use the limit more specific to it.\n\n"},{"id":"539401","messageId":"xmqqa4w3n5se.fsf@gitster.g","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-2-5ddebad292b0@pks.im","subject":"Re: [PATCH 02/14] oidtree: extend iteration to allow for arbitrary return codes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-19T16:27:29Z","receivedAt":"2026-03-19T16:27:32Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> diff --git a/cbtree.h b/cbtree.h\n> index 43193abdda..4f644d6e45 100644\n> --- a/cbtree.h\n> +++ b/cbtree.h\n> @@ -30,11 +30,6 @@ struct cb_tree {\n>  \tstruct cb_node *root;\n>  };\n>  \n> -enum cb_next {\n> -\tCB_CONTINUE = 0,\n> -\tCB_BREAK = 1\n> -};\n> -\n>  #define CBTREE_INIT { 0 }\n>  \n>  static inline void cb_init(struct cb_tree *t)\n> @@ -46,9 +41,9 @@ static inline void cb_init(struct cb_tree *t)\n>  struct cb_node *cb_lookup(struct cb_tree *, const uint8_t *k, size_t klen);\n>  struct cb_node *cb_insert(struct cb_tree *, struct cb_node *, size_t klen);\n>  \n\nThis change does make sense,...\n\n> -typedef enum cb_next (*cb_iter)(struct cb_node *, void *arg);\n> +typedef int (*cb_iter)(struct cb_node *, void *arg);\n\n... but it probably now warrants a bit of comment as readers cannot\nguess from the values in cb_next enum that is gone.\n\n   cb_each() keeps iterating while this function returns 0; a non-zero\n   value returned from the iterator callback is relayed back to the\n   caller of cb_each() after immediately stopping iteration.\n\n> -void cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n> -\t\tcb_iter, void *arg);\n> +int cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n> +\t    cb_iter, void *arg);\n\nor something like that, perhaps.\n\n"},{"id":"539402","messageId":"CAOLa=ZQFUH4k5xDqv0rozSXcbsGFsUzVa-fnzfcb=+957zwHRg@mail.gmail.com","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-2-5ddebad292b0@pks.im","subject":"Re: [PATCH 02/14] oidtree: extend iteration to allow for arbitrary return codes","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-03-19T16:27:51Z","receivedAt":"2026-03-19T16:27:54Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> The interface `cb_each()` iterates through a crit-bit tree and calls a\n> specific callback function for each of the contained items. The callback\n> function is expected to return either:\n>\n>   - `CB_CONTINUE` in case iteration shall continue.\n>\n>   - `CB_BREAK` to abort iteration.\n>\n> This is needlessly restrictive though, as callers may want to return\n> arbitrary values and have them be bubbled up to the `cb_each()` call\n> site. In fact, this is a rather common pattern we have: whenever such a\n> callback function returns a non-zero error code, we abort iteration and\n> bubble up the code as-is.\n>\n> Refactor both the crit-bit tree and oidtree subsystems to behave\n> accordingly.\n>\n\nOkay so this patch simply irradiates the need for a specific enum to\nreplace it the standard where non-zero error code stops the iteration.\nOkay makes sense.\n\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  cbtree.c                 | 21 ++++++++++++---------\n>  cbtree.h                 | 11 +++--------\n>  object-name.c            |  4 ++--\n>  oidtree.c                | 12 ++++++------\n>  oidtree.h                | 18 ++++++++++++------\n>  t/unit-tests/u-oidtree.c |  4 ++--\n>  6 files changed, 37 insertions(+), 33 deletions(-)\n>\n> diff --git a/cbtree.c b/cbtree.c\n> index cf8cf75b89..4ab794bddc 100644\n> --- a/cbtree.c\n> +++ b/cbtree.c\n> @@ -96,26 +96,28 @@ struct cb_node *cb_lookup(struct cb_tree *t, const uint8_t *k, size_t klen)\n>  \treturn p && !memcmp(p->k, k, klen) ? p : NULL;\n>  }\n>\n> -static enum cb_next cb_descend(struct cb_node *p, cb_iter fn, void *arg)\n> +static int cb_descend(struct cb_node *p, cb_iter fn, void *arg)\n>  {\n>  \tif (1 & (uintptr_t)p) {\n>  \t\tstruct cb_node *q = cb_node_of(p);\n> -\t\tenum cb_next n = cb_descend(q->child[0], fn, arg);\n> -\n> -\t\treturn n == CB_BREAK ? n : cb_descend(q->child[1], fn, arg);\n> +\t\tint ret = cb_descend(q->child[0], fn, arg);\n> +\t\tif (ret)\n> +\t\t\treturn ret;\n> +\t\treturn cb_descend(q->child[1], fn, arg);\n>  \t} else {\n>  \t\treturn fn(p, arg);\n>  \t}\n>  }\n>\n> -void cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n> -\t\t\tcb_iter fn, void *arg)\n> +int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n> +\t    cb_iter fn, void *arg)\n>  {\n>  \tstruct cb_node *p = t->root;\n>  \tstruct cb_node *top = p;\n>  \tsize_t i = 0;\n>\n> -\tif (!p) return; /* empty tree */\n> +\tif (!p)\n> +\t\treturn 0; /* empty tree */\n>\n>  \t/* Walk tree, maintaining top pointer */\n>  \twhile (1 & (uintptr_t)p) {\n> @@ -130,7 +132,8 @@ void cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n>\n>  \tfor (i = 0; i < klen; i++) {\n>  \t\tif (p->k[i] != kpfx[i])\n> -\t\t\treturn; /* \"best\" match failed */\n> +\t\t\treturn 0; /* \"best\" match failed */\n>  \t}\n> -\tcb_descend(top, fn, arg);\n> +\n> +\treturn cb_descend(top, fn, arg);\n>  }\n> diff --git a/cbtree.h b/cbtree.h\n> index 43193abdda..4f644d6e45 100644\n> --- a/cbtree.h\n> +++ b/cbtree.h\n> @@ -30,11 +30,6 @@ struct cb_tree {\n>  \tstruct cb_node *root;\n>  };\n>\n> -enum cb_next {\n> -\tCB_CONTINUE = 0,\n> -\tCB_BREAK = 1\n> -};\n> -\n>  #define CBTREE_INIT { 0 }\n>\n>  static inline void cb_init(struct cb_tree *t)\n> @@ -46,9 +41,9 @@ static inline void cb_init(struct cb_tree *t)\n>  struct cb_node *cb_lookup(struct cb_tree *, const uint8_t *k, size_t klen);\n>  struct cb_node *cb_insert(struct cb_tree *, struct cb_node *, size_t klen);\n>\n> -typedef enum cb_next (*cb_iter)(struct cb_node *, void *arg);\n> +typedef int (*cb_iter)(struct cb_node *, void *arg);\n>\n> -void cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n> -\t\tcb_iter, void *arg);\n> +int cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n> +\t    cb_iter, void *arg);\n>\n>  #endif /* CBTREE_H */\n> diff --git a/object-name.c b/object-name.c\n> index e5adec4c9d..a24a1b48e1 100644\n> --- a/object-name.c\n> +++ b/object-name.c\n> @@ -103,12 +103,12 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n>\n>  static int match_hash(unsigned, const unsigned char *, const unsigned char *);\n>\n> -static enum cb_next match_prefix(const struct object_id *oid, void *arg)\n> +static int match_prefix(const struct object_id *oid, void *arg)\n>  {\n>  \tstruct disambiguate_state *ds = arg;\n>  \t/* no need to call match_hash, oidtree_each did prefix match */\n>  \tupdate_candidates(ds, oid);\n> -\treturn ds->ambiguous ? CB_BREAK : CB_CONTINUE;\n> +\treturn ds->ambiguous;\n>  }\n>\n>  static void find_short_object_filename(struct disambiguate_state *ds)\n> diff --git a/oidtree.c b/oidtree.c\n> index a4d10cd429..ab9fe7ec7a 100644\n> --- a/oidtree.c\n> +++ b/oidtree.c\n> @@ -71,7 +71,7 @@ struct oidtree_each_data {\n>  \tuint8_t last_byte;\n>  };\n>\n> -static enum cb_next iter(struct cb_node *n, void *cb_data)\n> +static int iter(struct cb_node *n, void *cb_data)\n>  {\n>  \tstruct oidtree_each_data *data = cb_data;\n>  \tstruct object_id k;\n> @@ -80,18 +80,18 @@ static enum cb_next iter(struct cb_node *n, void *cb_data)\n>  \tmemcpy(&k, n->k, sizeof(k));\n>\n>  \tif (data->algo != GIT_HASH_UNKNOWN && data->algo != k.algo)\n> -\t\treturn CB_CONTINUE;\n> +\t\treturn 0;\n>\n>  \tif (data->last_nibble_at) {\n>  \t\tif ((k.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n> -\t\t\treturn CB_CONTINUE;\n> +\t\t\treturn 0;\n>  \t}\n>\n>  \treturn data->cb(&k, data->cb_data);\n>  }\n>\n> -void oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n> -\t\t  size_t prefix_hex_len, oidtree_each_cb cb, void *cb_data)\n> +int oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n> +\t\t size_t prefix_hex_len, oidtree_each_cb cb, void *cb_data)\n>  {\n>  \tstruct oidtree_each_data data = {\n>  \t\t.cb = cb,\n> @@ -106,5 +106,5 @@ void oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n>  \t\tdata.last_nibble_at = &klen;\n>  \t}\n>\n> -\tcb_each(&ot->tree, prefix->hash, klen, iter, &data);\n> +\treturn cb_each(&ot->tree, prefix->hash, klen, iter, &data);\n>  }\n> diff --git a/oidtree.h b/oidtree.h\n> index 0651401017..2b7bad2e60 100644\n> --- a/oidtree.h\n> +++ b/oidtree.h\n> @@ -35,16 +35,22 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid);\n>  /* Check whether the tree contains the given object ID. */\n>  bool oidtree_contains(struct oidtree *ot, const struct object_id *oid);\n>\n> -/* Callback function used for `oidtree_each()`. */\n> -typedef enum cb_next (*oidtree_each_cb)(const struct object_id *oid,\n> -\t\t\t\t\tvoid *cb_data);\n> +/*\n> + * Callback function used for `oidtree_each()`. Returning a non-zero exit code\n> + * will cause iteration to stop. The exit code will be propagated to the caller\n> + * of `oidtree_each()`.\n> + */\n> +typedef int (*oidtree_each_cb)(const struct object_id *oid,\n> +\t\t\t       void *cb_data);\n>\n>  /*\n>   * Iterate through all object IDs in the tree whose prefix matches the given\n>   * object ID prefix and invoke the callback function on each of them.\n> + *\n> + * Returns any non-zero exit code from the provided callback function.\n>   */\n> -void oidtree_each(struct oidtree *ot,\n> -\t\t  const struct object_id *prefix, size_t prefix_hex_len,\n> -\t\t  oidtree_each_cb cb, void *cb_data);\n> +int oidtree_each(struct oidtree *ot,\n> +\t\t const struct object_id *prefix, size_t prefix_hex_len,\n> +\t\t oidtree_each_cb cb, void *cb_data);\n>\n>  #endif /* OIDTREE_H */\n> diff --git a/t/unit-tests/u-oidtree.c b/t/unit-tests/u-oidtree.c\n> index def47c6795..d4d05c7dc3 100644\n> --- a/t/unit-tests/u-oidtree.c\n> +++ b/t/unit-tests/u-oidtree.c\n> @@ -38,7 +38,7 @@ struct expected_hex_iter {\n>  \tconst char *query;\n>  };\n>\n> -static enum cb_next check_each_cb(const struct object_id *oid, void *data)\n> +static int check_each_cb(const struct object_id *oid, void *data)\n>  {\n>  \tstruct expected_hex_iter *hex_iter = data;\n>  \tstruct object_id expected;\n> @@ -49,7 +49,7 @@ static enum cb_next check_each_cb(const struct object_id *oid, void *data)\n>  \t\t\t &expected);\n>  \tcl_assert_equal_s(oid_to_hex(oid), oid_to_hex(&expected));\n>  \thex_iter->i += 1;\n> -\treturn CB_CONTINUE;\n> +\treturn 0;\n>  }\n>\n>  LAST_ARG_MUST_BE_NULL\n>\n> --\n> 2.53.0.1055.ga2ffed1127.dirty\n\nThe patch looks good.\n"},{"id":"539478","messageId":"abzryk0qjlbvy8OL@pks.im","threadId":"65299","inReplyTo":"xmqqeclfn6nn.fsf@gitster.g","subject":"Re: [PATCH 01/14] oidtree: modernize the code a bit","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T06:40:10Z","receivedAt":"2026-03-20T06:40:16Z","isPatch":true,"body":"On Thu, Mar 19, 2026 at 09:08:44AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > +void oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n> > +\t\t  size_t prefix_hex_len, oidtree_each_cb cb, void *cb_data)\n> >  {\n> > -\tsize_t klen = oidhexsz / 2;\n> > -\tstruct oidtree_iter_data x = { 0 };\n> > -\tassert(oidhexsz <= GIT_MAX_HEXSZ);\n> > -\n> > -\tx.fn = fn;\n> > -\tx.arg = arg;\n> > -\tx.algo = oid->algo;\n> > -\tif (oidhexsz & 1) {\n> > -\t\tx.last_byte = oid->hash[klen];\n> > -\t\tx.last_nibble_at = &klen;\n> > +\tstruct oidtree_each_data data = {\n> > +\t\t.cb = cb,\n> > +\t\t.cb_data = cb_data,\n> > +\t\t.algo = prefix->algo,\n> > +\t};\n> > +\tsize_t klen = prefix_hex_len / 2;\n> > +\tassert(prefix_hex_len <= GIT_MAX_HEXSZ);\n> \n> I know the original also used GIT_MAX_HEXSZ to clamp the length for\n> sanity, but because we know what algorithm is in use, I wonder if we\n> want to use the limit more specific to it.\n\nThat assumes that the passed prefix OID actually has an algorithm\nattached to it, and that may not be the case. We could initialize the\noverall oidtree with a hash algorithm in `oidtree_init()`, and if so we\ncan then become a bit more thorough with our asserts.\n\nBut I feel like that would go beyond the smallish cleanups that I'm\ndoing in this patch.\n\nPatrick\n"},{"id":"539479","messageId":"abzr0EaDfn2Dpo03@pks.im","threadId":"65299","inReplyTo":"xmqqa4w3n5se.fsf@gitster.g","subject":"Re: [PATCH 02/14] oidtree: extend iteration to allow for arbitrary return codes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T06:40:16Z","receivedAt":"2026-03-20T06:40:20Z","isPatch":true,"body":"On Thu, Mar 19, 2026 at 09:27:29AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > diff --git a/cbtree.h b/cbtree.h\n> > index 43193abdda..4f644d6e45 100644\n> > --- a/cbtree.h\n> > +++ b/cbtree.h\n> > @@ -30,11 +30,6 @@ struct cb_tree {\n> >  \tstruct cb_node *root;\n> >  };\n> >  \n> > -enum cb_next {\n> > -\tCB_CONTINUE = 0,\n> > -\tCB_BREAK = 1\n> > -};\n> > -\n> >  #define CBTREE_INIT { 0 }\n> >  \n> >  static inline void cb_init(struct cb_tree *t)\n> > @@ -46,9 +41,9 @@ static inline void cb_init(struct cb_tree *t)\n> >  struct cb_node *cb_lookup(struct cb_tree *, const uint8_t *k, size_t klen);\n> >  struct cb_node *cb_insert(struct cb_tree *, struct cb_node *, size_t klen);\n> >  \n> \n> This change does make sense,...\n> \n> > -typedef enum cb_next (*cb_iter)(struct cb_node *, void *arg);\n> > +typedef int (*cb_iter)(struct cb_node *, void *arg);\n> \n> ... but it probably now warrants a bit of comment as readers cannot\n> guess from the values in cb_next enum that is gone.\n> \n>    cb_each() keeps iterating while this function returns 0; a non-zero\n>    value returned from the iterator callback is relayed back to the\n>    caller of cb_each() after immediately stopping iteration.\n> \n> > -void cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n> > -\t\tcb_iter, void *arg);\n> > +int cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n> > +\t    cb_iter, void *arg);\n> \n> or something like that, perhaps.\n\nMakes sense, will do.\n\nPatrick\n"},{"id":"539480","messageId":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im","subject":"[PATCH v2 00/14] odb: generic object name handling","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:26Z","receivedAt":"2026-03-20T07:07:42Z","isPatch":true,"body":"Hi,\n\nthis patch series refactors handling of object names to become pluggable\nand thus generic. This includes:\n\n  - Disambiguation of object names with a common prefix. This is\n    required to list candidate objects in case the user has passed a\n    non-unique prefix.\n\n  - Abbreviating an object ID to the shortest prefix required while\n    staying unique.\n\nThe logic to compute these operations is specific to the backend, but\nnot generic. This patch series fixes that by moving the functionality\ninto the respective backends.\n\nThis patch series may feel somewhat unexiting, but it's not. Especially\nabbreviating object IDs is done in lots of places, so this functionality\nis overall quite critical. So starting with this series, it is now\npossible to do all kinds of local work with an alternative backend:\ngit-commit(1), git-log(1), git-rev-parse(1), git-merge(1) and many other\ncommands now work as expected. My MongoDB proof of concept [1] only\nrequires two commits (the object format extension) on top. And no, I\ndon't endorse MongoDB or propose it as a future potential backend. It\nsimply had a good C API that was easy to use.\n\nOf course, other functionality, especially everything that involves\npackfiles, doesn't yet work.\n\nThis patch series is built on top of ca1db8a0f7 (The 17th batch,\n2026-03-16) with ps/object-counting at 6801ffd37d (odb: introduce\ngeneric object counting, 2026-03-12) merged into it.\n\nChanges in v2:\n  - Document `cb_iter` callback.\n  - Fix left-over conversion of `odb_source_loose_for_each_object()`.\n  - commit message typo fixes.\n  - Link to v1: https://lore.kernel.org/r/20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im\n\nThanks!\n\nPatrick\n\n[1]: https://gitlab.com/gitlab-org/git/-/merge_requests/454\n\n---\nPatrick Steinhardt (14):\n      oidtree: modernize the code a bit\n      oidtree: extend iteration to allow for arbitrary return codes\n      odb: introduce `struct odb_for_each_object_options`\n      object-name: move logic to iterate through loose prefixed objects\n      object-name: move logic to iterate through packed prefixed objects\n      object-name: extract function to parse object ID prefixes\n      object-name: backend-generic `repo_collect_ambiguous()`\n      object-name: backend-generic `get_short_oid()`\n      object-name: merge `update_candidates()` and `match_prefix()`\n      object-name: abbreviate loose object names without `disambiguate_state`\n      object-name: simplify computing common prefixes\n      object-name: move logic to compute loose abbreviation length\n      object-file: move logic to compute packed abbreviation length\n      odb: introduce generic `odb_find_abbrev_len()`\n\n builtin/cat-file.c       |   7 +-\n builtin/pack-objects.c   |  12 +-\n cbtree.c                 |  21 ++-\n cbtree.h                 |  17 +-\n commit-graph.c           |   5 +-\n hash.c                   |  18 ++\n hash.h                   |   3 +\n object-file.c            |  76 ++++++++-\n object-file.h            |  21 ++-\n object-name.c            | 437 ++++++++---------------------------------------\n odb.c                    |  99 ++++++++++-\n odb.h                    |  39 +++++\n odb/source-files.c       |  33 +++-\n odb/source.h             |  30 +++-\n oidtree.c                |  65 +++----\n oidtree.h                |  48 +++++-\n packfile.c               | 297 +++++++++++++++++++++++++++++++-\n packfile.h               |   7 +-\n t/unit-tests/u-oidtree.c |  18 +-\n 19 files changed, 782 insertions(+), 471 deletions(-)\n\nRange-diff versus v1:\n\n 1:  5ce8cced1e =  1:  755acf126c oidtree: modernize the code a bit\n 2:  26c4377ff1 !  2:  6ab9e5d41c oidtree: extend iteration to allow for arbitrary return codes\n    @@ cbtree.h: static inline void cb_init(struct cb_tree *t)\n      struct cb_node *cb_insert(struct cb_tree *, struct cb_node *, size_t klen);\n      \n     -typedef enum cb_next (*cb_iter)(struct cb_node *, void *arg);\n    ++/*\n    ++ * Callback invoked by `cb_each()` for each node in the critbit tree. A return\n    ++ * value of 0 will cause the iteration to continue, a non-zero return code will\n    ++ * cause iteration to abort. The error code will be relayed back from\n    ++ * `cb_each()` in that case.\n    ++ */\n     +typedef int (*cb_iter)(struct cb_node *, void *arg);\n      \n     -void cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n 3:  da7b74b572 !  3:  9caf0288e4 odb: introduce `struct odb_for_each_object_options`\n    @@ Commit message\n         a subsequent commit we'll want to change object iteration to also\n         support iterating over only those objects that have a specific prefix.\n         While we could of course add the prefix to the function signature, or\n    -    alternative introduce a new function, both of these options don't really\n    -    seem to be that sensible.\n    +    alternatively introduce a new function, both of these options don't\n    +    really seem to be that sensible.\n     \n         Instead, introduce a new `struct odb_for_each_object_options` that can\n         be passed to a new `odb_for_each_object_ext()` function. Splice through\n    @@ object-file.c: int odb_source_loose_for_each_object(struct odb_source *source,\n      \t\treturn 0;\n      \n      \treturn for_each_loose_file_in_source(source, for_each_object_wrapper_cb,\n    +@@ object-file.c: int odb_source_loose_count_objects(struct odb_source *source,\n    + \t\t*out = count * 256;\n    + \t\tret = 0;\n    + \t} else {\n    ++\t\tstruct odb_for_each_object_options opts = { 0 };\n    + \t\t*out = 0;\n    + \t\tret = odb_source_loose_for_each_object(source, NULL, count_loose_object,\n    +-\t\t\t\t\t\t       out, 0);\n    ++\t\t\t\t\t\t       out, &opts);\n    + \t}\n    + \n    + out:\n     \n      ## object-file.h ##\n     @@ object-file.h: int odb_source_loose_for_each_object(struct odb_source *source,\n 4:  e4dedd4686 =  4:  232bcf662e object-name: move logic to iterate through loose prefixed objects\n 5:  7f9bfca9fd =  5:  be9b546c67 object-name: move logic to iterate through packed prefixed objects\n 6:  3eb88fa774 =  6:  60672e6cd1 object-name: extract function to parse object ID prefixes\n 7:  8fa16eb02a !  7:  38032fee38 object-name: backend-generic `repo_collect_ambiguous()`\n    @@ Commit message\n         to do is to enumerate objects that have such a prefix and then append\n         those objects to a `struct oid_array`. This can be trivially achieved\n         in a generic way now that `odb_for_each_object()` has learned to yield\n    -    only objects that much such a prefix.\n    +    only objects that match such a prefix.\n     \n         Refactor the code to use the backend-generic infrastructure instead.\n     \n 8:  0b0ce71bdd =  8:  98af219ca1 object-name: backend-generic `get_short_oid()`\n 9:  74733d34de =  9:  12866c582b object-name: merge `update_candidates()` and `match_prefix()`\n10:  2ca53a01d4 = 10:  6b4eda123a object-name: abbreviate loose object names without `disambiguate_state`\n11:  a391911854 = 11:  7aa921504b object-name: simplify computing common prefixes\n12:  9ad66df1ca = 12:  d1bf2706f8 object-name: move logic to compute loose abbreviation length\n13:  cf8a33ab67 = 13:  e59e7dffb4 object-file: move logic to compute packed abbreviation length\n14:  acd07686db = 14:  726be6de40 odb: introduce generic `odb_find_abbrev_len()`\n\n---\nbase-commit: b052aca69d64d2d8e28e7ce97dcb1beb3d94515a\nchange-id: 20260313-b4-pks-odb-source-abbrev-a84c51222bca\n\n"},{"id":"539481","messageId":"20260320-b4-pks-odb-source-abbrev-v2-1-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 01/14] oidtree: modernize the code a bit","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:27Z","receivedAt":"2026-03-20T07:07:45Z","isPatch":true,"body":"The \"oidtree.c\" subsystem is rather small and self-contained and tends\nto just work. It thus doesn't typically receive a lot of attention,\nwhich has as a consequence that it's coding style is somewhat dated\nnowadays.\n\nModernize the style of this subsystem a bit:\n\n  - Rename the `oidtree_iter()` function to `oidtree_each_cb()`.\n\n  - Rename `struct oidtree_iter_data` to `struct oidtree_each_data` to\n    match the renamed callback function type.\n\n  - Rename parameters and variables to clarify their intent.\n\n  - Add comments that explain what some of the functions do.\n\n  - Adapt the return value of `oidtree_contains()` to be a boolean.\n\nThis prepares for some changes to the subsystem that'll happen in the\nnext commit.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n oidtree.c                | 61 ++++++++++++++++++++++++------------------------\n oidtree.h                | 42 +++++++++++++++++++++++++++------\n t/unit-tests/u-oidtree.c | 14 +++++------\n 3 files changed, 73 insertions(+), 44 deletions(-)\n\ndiff --git a/oidtree.c b/oidtree.c\nindex 324de94934..a4d10cd429 100644\n--- a/oidtree.c\n+++ b/oidtree.c\n@@ -6,14 +6,6 @@\n #include \"oidtree.h\"\n #include \"hash.h\"\n \n-struct oidtree_iter_data {\n-\toidtree_iter fn;\n-\tvoid *arg;\n-\tsize_t *last_nibble_at;\n-\tuint32_t algo;\n-\tuint8_t last_byte;\n-};\n-\n void oidtree_init(struct oidtree *ot)\n {\n \tcb_init(&ot->tree);\n@@ -54,8 +46,7 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid)\n \tcb_insert(&ot->tree, on, sizeof(*oid));\n }\n \n-\n-int oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n+bool oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n {\n \tstruct object_id k;\n \tsize_t klen = sizeof(k);\n@@ -69,41 +60,51 @@ int oidtree_contains(struct oidtree *ot, const struct object_id *oid)\n \tklen += BUILD_ASSERT_OR_ZERO(offsetof(struct object_id, hash) <\n \t\t\t\toffsetof(struct object_id, algo));\n \n-\treturn cb_lookup(&ot->tree, (const uint8_t *)&k, klen) ? 1 : 0;\n+\treturn !!cb_lookup(&ot->tree, (const uint8_t *)&k, klen);\n }\n \n-static enum cb_next iter(struct cb_node *n, void *arg)\n+struct oidtree_each_data {\n+\toidtree_each_cb cb;\n+\tvoid *cb_data;\n+\tsize_t *last_nibble_at;\n+\tuint32_t algo;\n+\tuint8_t last_byte;\n+};\n+\n+static enum cb_next iter(struct cb_node *n, void *cb_data)\n {\n-\tstruct oidtree_iter_data *x = arg;\n+\tstruct oidtree_each_data *data = cb_data;\n \tstruct object_id k;\n \n \t/* Copy to provide 4-byte alignment needed by struct object_id. */\n \tmemcpy(&k, n->k, sizeof(k));\n \n-\tif (x->algo != GIT_HASH_UNKNOWN && x->algo != k.algo)\n+\tif (data->algo != GIT_HASH_UNKNOWN && data->algo != k.algo)\n \t\treturn CB_CONTINUE;\n \n-\tif (x->last_nibble_at) {\n-\t\tif ((k.hash[*x->last_nibble_at] ^ x->last_byte) & 0xf0)\n+\tif (data->last_nibble_at) {\n+\t\tif ((k.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n \t\t\treturn CB_CONTINUE;\n \t}\n \n-\treturn x->fn(&k, x->arg);\n+\treturn data->cb(&k, data->cb_data);\n }\n \n-void oidtree_each(struct oidtree *ot, const struct object_id *oid,\n-\t\t\tsize_t oidhexsz, oidtree_iter fn, void *arg)\n+void oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n+\t\t  size_t prefix_hex_len, oidtree_each_cb cb, void *cb_data)\n {\n-\tsize_t klen = oidhexsz / 2;\n-\tstruct oidtree_iter_data x = { 0 };\n-\tassert(oidhexsz <= GIT_MAX_HEXSZ);\n-\n-\tx.fn = fn;\n-\tx.arg = arg;\n-\tx.algo = oid->algo;\n-\tif (oidhexsz & 1) {\n-\t\tx.last_byte = oid->hash[klen];\n-\t\tx.last_nibble_at = &klen;\n+\tstruct oidtree_each_data data = {\n+\t\t.cb = cb,\n+\t\t.cb_data = cb_data,\n+\t\t.algo = prefix->algo,\n+\t};\n+\tsize_t klen = prefix_hex_len / 2;\n+\tassert(prefix_hex_len <= GIT_MAX_HEXSZ);\n+\n+\tif (prefix_hex_len & 1) {\n+\t\tdata.last_byte = prefix->hash[klen];\n+\t\tdata.last_nibble_at = &klen;\n \t}\n-\tcb_each(&ot->tree, (const uint8_t *)oid, klen, iter, &x);\n+\n+\tcb_each(&ot->tree, prefix->hash, klen, iter, &data);\n }\ndiff --git a/oidtree.h b/oidtree.h\nindex 77898f510a..0651401017 100644\n--- a/oidtree.h\n+++ b/oidtree.h\n@@ -5,18 +5,46 @@\n #include \"hash.h\"\n #include \"mem-pool.h\"\n \n+/*\n+ * OID trees are an efficient storage for object IDs that use a critbit tree\n+ * internally. Common prefixes are duplicated and object IDs are stored in a\n+ * way that allow easy iteration over the objects in lexicographic order. As a\n+ * consequence, operations that want to enumerate all object IDs that match a\n+ * given prefix can be answered efficiently.\n+ *\n+ * Note that it is not (yet) possible to store data other than the object IDs\n+ * themselves in this tree.\n+ */\n struct oidtree {\n \tstruct cb_tree tree;\n \tstruct mem_pool mem_pool;\n };\n \n-void oidtree_init(struct oidtree *);\n-void oidtree_clear(struct oidtree *);\n-void oidtree_insert(struct oidtree *, const struct object_id *);\n-int oidtree_contains(struct oidtree *, const struct object_id *);\n+/* Initialize the oidtree so that it is ready for use. */\n+void oidtree_init(struct oidtree *ot);\n \n-typedef enum cb_next (*oidtree_iter)(const struct object_id *, void *data);\n-void oidtree_each(struct oidtree *, const struct object_id *,\n-\t\t\tsize_t oidhexsz, oidtree_iter, void *data);\n+/*\n+ * Release all memory associated with the oidtree and reinitialize it for\n+ * subsequent use.\n+ */\n+void oidtree_clear(struct oidtree *ot);\n+\n+/* Insert the object ID into the tree. */\n+void oidtree_insert(struct oidtree *ot, const struct object_id *oid);\n+\n+/* Check whether the tree contains the given object ID. */\n+bool oidtree_contains(struct oidtree *ot, const struct object_id *oid);\n+\n+/* Callback function used for `oidtree_each()`. */\n+typedef enum cb_next (*oidtree_each_cb)(const struct object_id *oid,\n+\t\t\t\t\tvoid *cb_data);\n+\n+/*\n+ * Iterate through all object IDs in the tree whose prefix matches the given\n+ * object ID prefix and invoke the callback function on each of them.\n+ */\n+void oidtree_each(struct oidtree *ot,\n+\t\t  const struct object_id *prefix, size_t prefix_hex_len,\n+\t\t  oidtree_each_cb cb, void *cb_data);\n \n #endif /* OIDTREE_H */\ndiff --git a/t/unit-tests/u-oidtree.c b/t/unit-tests/u-oidtree.c\nindex e6eede2740..def47c6795 100644\n--- a/t/unit-tests/u-oidtree.c\n+++ b/t/unit-tests/u-oidtree.c\n@@ -24,7 +24,7 @@ static int fill_tree_loc(struct oidtree *ot, const char *hexes[], size_t n)\n \treturn 0;\n }\n \n-static void check_contains(struct oidtree *ot, const char *hex, int expected)\n+static void check_contains(struct oidtree *ot, const char *hex, bool expected)\n {\n \tstruct object_id oid;\n \n@@ -88,12 +88,12 @@ void test_oidtree__cleanup(void)\n void test_oidtree__contains(void)\n {\n \tFILL_TREE(&ot, \"444\", \"1\", \"2\", \"3\", \"4\", \"5\", \"a\", \"b\", \"c\", \"d\", \"e\");\n-\tcheck_contains(&ot, \"44\", 0);\n-\tcheck_contains(&ot, \"441\", 0);\n-\tcheck_contains(&ot, \"440\", 0);\n-\tcheck_contains(&ot, \"444\", 1);\n-\tcheck_contains(&ot, \"4440\", 1);\n-\tcheck_contains(&ot, \"4444\", 0);\n+\tcheck_contains(&ot, \"44\", false);\n+\tcheck_contains(&ot, \"441\", false);\n+\tcheck_contains(&ot, \"440\", false);\n+\tcheck_contains(&ot, \"444\", true);\n+\tcheck_contains(&ot, \"4440\", true);\n+\tcheck_contains(&ot, \"4444\", false);\n }\n \n void test_oidtree__each(void)\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539482","messageId":"20260320-b4-pks-odb-source-abbrev-v2-2-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 02/14] oidtree: extend iteration to allow for arbitrary return codes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:28Z","receivedAt":"2026-03-20T07:07:48Z","isPatch":true,"body":"The interface `cb_each()` iterates through a crit-bit tree and calls a\nspecific callback function for each of the contained items. The callback\nfunction is expected to return either:\n\n  - `CB_CONTINUE` in case iteration shall continue.\n\n  - `CB_BREAK` to abort iteration.\n\nThis is needlessly restrictive though, as callers may want to return\narbitrary values and have them be bubbled up to the `cb_each()` call\nsite. In fact, this is a rather common pattern we have: whenever such a\ncallback function returns a non-zero error code, we abort iteration and\nbubble up the code as-is.\n\nRefactor both the crit-bit tree and oidtree subsystems to behave\naccordingly.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n cbtree.c                 | 21 ++++++++++++---------\n cbtree.h                 | 17 +++++++++--------\n object-name.c            |  4 ++--\n oidtree.c                | 12 ++++++------\n oidtree.h                | 18 ++++++++++++------\n t/unit-tests/u-oidtree.c |  4 ++--\n 6 files changed, 43 insertions(+), 33 deletions(-)\n\ndiff --git a/cbtree.c b/cbtree.c\nindex cf8cf75b89..4ab794bddc 100644\n--- a/cbtree.c\n+++ b/cbtree.c\n@@ -96,26 +96,28 @@ struct cb_node *cb_lookup(struct cb_tree *t, const uint8_t *k, size_t klen)\n \treturn p && !memcmp(p->k, k, klen) ? p : NULL;\n }\n \n-static enum cb_next cb_descend(struct cb_node *p, cb_iter fn, void *arg)\n+static int cb_descend(struct cb_node *p, cb_iter fn, void *arg)\n {\n \tif (1 & (uintptr_t)p) {\n \t\tstruct cb_node *q = cb_node_of(p);\n-\t\tenum cb_next n = cb_descend(q->child[0], fn, arg);\n-\n-\t\treturn n == CB_BREAK ? n : cb_descend(q->child[1], fn, arg);\n+\t\tint ret = cb_descend(q->child[0], fn, arg);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t\treturn cb_descend(q->child[1], fn, arg);\n \t} else {\n \t\treturn fn(p, arg);\n \t}\n }\n \n-void cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n-\t\t\tcb_iter fn, void *arg)\n+int cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n+\t    cb_iter fn, void *arg)\n {\n \tstruct cb_node *p = t->root;\n \tstruct cb_node *top = p;\n \tsize_t i = 0;\n \n-\tif (!p) return; /* empty tree */\n+\tif (!p)\n+\t\treturn 0; /* empty tree */\n \n \t/* Walk tree, maintaining top pointer */\n \twhile (1 & (uintptr_t)p) {\n@@ -130,7 +132,8 @@ void cb_each(struct cb_tree *t, const uint8_t *kpfx, size_t klen,\n \n \tfor (i = 0; i < klen; i++) {\n \t\tif (p->k[i] != kpfx[i])\n-\t\t\treturn; /* \"best\" match failed */\n+\t\t\treturn 0; /* \"best\" match failed */\n \t}\n-\tcb_descend(top, fn, arg);\n+\n+\treturn cb_descend(top, fn, arg);\n }\ndiff --git a/cbtree.h b/cbtree.h\nindex 43193abdda..c374b1b3db 100644\n--- a/cbtree.h\n+++ b/cbtree.h\n@@ -30,11 +30,6 @@ struct cb_tree {\n \tstruct cb_node *root;\n };\n \n-enum cb_next {\n-\tCB_CONTINUE = 0,\n-\tCB_BREAK = 1\n-};\n-\n #define CBTREE_INIT { 0 }\n \n static inline void cb_init(struct cb_tree *t)\n@@ -46,9 +41,15 @@ static inline void cb_init(struct cb_tree *t)\n struct cb_node *cb_lookup(struct cb_tree *, const uint8_t *k, size_t klen);\n struct cb_node *cb_insert(struct cb_tree *, struct cb_node *, size_t klen);\n \n-typedef enum cb_next (*cb_iter)(struct cb_node *, void *arg);\n+/*\n+ * Callback invoked by `cb_each()` for each node in the critbit tree. A return\n+ * value of 0 will cause the iteration to continue, a non-zero return code will\n+ * cause iteration to abort. The error code will be relayed back from\n+ * `cb_each()` in that case.\n+ */\n+typedef int (*cb_iter)(struct cb_node *, void *arg);\n \n-void cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n-\t\tcb_iter, void *arg);\n+int cb_each(struct cb_tree *, const uint8_t *kpfx, size_t klen,\n+\t    cb_iter, void *arg);\n \n #endif /* CBTREE_H */\ndiff --git a/object-name.c b/object-name.c\nindex e5adec4c9d..a24a1b48e1 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -103,12 +103,12 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n \n static int match_hash(unsigned, const unsigned char *, const unsigned char *);\n \n-static enum cb_next match_prefix(const struct object_id *oid, void *arg)\n+static int match_prefix(const struct object_id *oid, void *arg)\n {\n \tstruct disambiguate_state *ds = arg;\n \t/* no need to call match_hash, oidtree_each did prefix match */\n \tupdate_candidates(ds, oid);\n-\treturn ds->ambiguous ? CB_BREAK : CB_CONTINUE;\n+\treturn ds->ambiguous;\n }\n \n static void find_short_object_filename(struct disambiguate_state *ds)\ndiff --git a/oidtree.c b/oidtree.c\nindex a4d10cd429..ab9fe7ec7a 100644\n--- a/oidtree.c\n+++ b/oidtree.c\n@@ -71,7 +71,7 @@ struct oidtree_each_data {\n \tuint8_t last_byte;\n };\n \n-static enum cb_next iter(struct cb_node *n, void *cb_data)\n+static int iter(struct cb_node *n, void *cb_data)\n {\n \tstruct oidtree_each_data *data = cb_data;\n \tstruct object_id k;\n@@ -80,18 +80,18 @@ static enum cb_next iter(struct cb_node *n, void *cb_data)\n \tmemcpy(&k, n->k, sizeof(k));\n \n \tif (data->algo != GIT_HASH_UNKNOWN && data->algo != k.algo)\n-\t\treturn CB_CONTINUE;\n+\t\treturn 0;\n \n \tif (data->last_nibble_at) {\n \t\tif ((k.hash[*data->last_nibble_at] ^ data->last_byte) & 0xf0)\n-\t\t\treturn CB_CONTINUE;\n+\t\t\treturn 0;\n \t}\n \n \treturn data->cb(&k, data->cb_data);\n }\n \n-void oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n-\t\t  size_t prefix_hex_len, oidtree_each_cb cb, void *cb_data)\n+int oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n+\t\t size_t prefix_hex_len, oidtree_each_cb cb, void *cb_data)\n {\n \tstruct oidtree_each_data data = {\n \t\t.cb = cb,\n@@ -106,5 +106,5 @@ void oidtree_each(struct oidtree *ot, const struct object_id *prefix,\n \t\tdata.last_nibble_at = &klen;\n \t}\n \n-\tcb_each(&ot->tree, prefix->hash, klen, iter, &data);\n+\treturn cb_each(&ot->tree, prefix->hash, klen, iter, &data);\n }\ndiff --git a/oidtree.h b/oidtree.h\nindex 0651401017..2b7bad2e60 100644\n--- a/oidtree.h\n+++ b/oidtree.h\n@@ -35,16 +35,22 @@ void oidtree_insert(struct oidtree *ot, const struct object_id *oid);\n /* Check whether the tree contains the given object ID. */\n bool oidtree_contains(struct oidtree *ot, const struct object_id *oid);\n \n-/* Callback function used for `oidtree_each()`. */\n-typedef enum cb_next (*oidtree_each_cb)(const struct object_id *oid,\n-\t\t\t\t\tvoid *cb_data);\n+/*\n+ * Callback function used for `oidtree_each()`. Returning a non-zero exit code\n+ * will cause iteration to stop. The exit code will be propagated to the caller\n+ * of `oidtree_each()`.\n+ */\n+typedef int (*oidtree_each_cb)(const struct object_id *oid,\n+\t\t\t       void *cb_data);\n \n /*\n  * Iterate through all object IDs in the tree whose prefix matches the given\n  * object ID prefix and invoke the callback function on each of them.\n+ *\n+ * Returns any non-zero exit code from the provided callback function.\n  */\n-void oidtree_each(struct oidtree *ot,\n-\t\t  const struct object_id *prefix, size_t prefix_hex_len,\n-\t\t  oidtree_each_cb cb, void *cb_data);\n+int oidtree_each(struct oidtree *ot,\n+\t\t const struct object_id *prefix, size_t prefix_hex_len,\n+\t\t oidtree_each_cb cb, void *cb_data);\n \n #endif /* OIDTREE_H */\ndiff --git a/t/unit-tests/u-oidtree.c b/t/unit-tests/u-oidtree.c\nindex def47c6795..d4d05c7dc3 100644\n--- a/t/unit-tests/u-oidtree.c\n+++ b/t/unit-tests/u-oidtree.c\n@@ -38,7 +38,7 @@ struct expected_hex_iter {\n \tconst char *query;\n };\n \n-static enum cb_next check_each_cb(const struct object_id *oid, void *data)\n+static int check_each_cb(const struct object_id *oid, void *data)\n {\n \tstruct expected_hex_iter *hex_iter = data;\n \tstruct object_id expected;\n@@ -49,7 +49,7 @@ static enum cb_next check_each_cb(const struct object_id *oid, void *data)\n \t\t\t &expected);\n \tcl_assert_equal_s(oid_to_hex(oid), oid_to_hex(&expected));\n \thex_iter->i += 1;\n-\treturn CB_CONTINUE;\n+\treturn 0;\n }\n \n LAST_ARG_MUST_BE_NULL\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539483","messageId":"20260320-b4-pks-odb-source-abbrev-v2-3-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 03/14] odb: introduce `struct odb_for_each_object_options`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:29Z","receivedAt":"2026-03-20T07:07:50Z","isPatch":true,"body":"The `odb_for_each_object()` function only accepts a bitset of flags. In\na subsequent commit we'll want to change object iteration to also\nsupport iterating over only those objects that have a specific prefix.\nWhile we could of course add the prefix to the function signature, or\nalternatively introduce a new function, both of these options don't\nreally seem to be that sensible.\n\nInstead, introduce a new `struct odb_for_each_object_options` that can\nbe passed to a new `odb_for_each_object_ext()` function. Splice through\nthe options structure into the respective object database sources.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n builtin/cat-file.c     |  7 +++++--\n builtin/pack-objects.c | 12 +++++++-----\n commit-graph.c         |  5 ++++-\n object-file.c          |  9 +++++----\n object-file.h          |  2 +-\n odb.c                  | 26 +++++++++++++++++++-------\n odb.h                  | 16 ++++++++++++++++\n odb/source-files.c     |  8 ++++----\n odb/source.h           |  6 +++---\n packfile.c             | 12 ++++++------\n packfile.h             |  2 +-\n 11 files changed, 71 insertions(+), 34 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex b6f12f41d6..cd13a3a89f 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -848,6 +848,9 @@ static void batch_each_object(struct batch_options *opt,\n \t\t.callback = callback,\n \t\t.payload = _payload,\n \t};\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.flags = flags,\n+\t};\n \tstruct bitmap_index *bitmap = NULL;\n \tstruct odb_source *source;\n \n@@ -860,7 +863,7 @@ static void batch_each_object(struct batch_options *opt,\n \todb_prepare_alternates(the_repository->objects);\n \tfor (source = the_repository->objects->sources; source; source = source->next) {\n \t\tint ret = odb_source_loose_for_each_object(source, NULL, batch_one_object_oi,\n-\t\t\t\t\t\t\t   &payload, flags);\n+\t\t\t\t\t\t\t   &payload, &opts);\n \t\tif (ret)\n \t\t\tbreak;\n \t}\n@@ -884,7 +887,7 @@ static void batch_each_object(struct batch_options *opt,\n \t\tfor (source = the_repository->objects->sources; source; source = source->next) {\n \t\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n \t\t\tint ret = packfile_store_for_each_object(files->packed, &oi,\n-\t\t\t\t\t\t\t\t batch_one_object_oi, &payload, flags);\n+\t\t\t\t\t\t\t\t batch_one_object_oi, &payload, &opts);\n \t\t\tif (ret)\n \t\t\t\tbreak;\n \t\t}\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex cd013c0b68..3bb57ff183 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -4344,6 +4344,12 @@ static void add_objects_in_unpacked_packs(void)\n {\n \tstruct odb_source *source;\n \ttime_t mtime;\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.flags = ODB_FOR_EACH_OBJECT_PACK_ORDER |\n+\t\t\t ODB_FOR_EACH_OBJECT_LOCAL_ONLY |\n+\t\t\t ODB_FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS |\n+\t\t\t ODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS,\n+\t};\n \tstruct object_info oi = {\n \t\t.mtimep = &mtime,\n \t};\n@@ -4356,11 +4362,7 @@ static void add_objects_in_unpacked_packs(void)\n \t\t\tcontinue;\n \n \t\tif (packfile_store_for_each_object(files->packed, &oi,\n-\t\t\t\t\t\t   add_object_in_unpacked_pack, NULL,\n-\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_PACK_ORDER |\n-\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_LOCAL_ONLY |\n-\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS |\n-\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS))\n+\t\t\t\t\t\t   add_object_in_unpacked_pack, NULL, &opts))\n \t\t\tdie(_(\"cannot open pack index\"));\n \t}\n }\ndiff --git a/commit-graph.c b/commit-graph.c\nindex c030003330..df4b4a125e 100644\n--- a/commit-graph.c\n+++ b/commit-graph.c\n@@ -1969,6 +1969,9 @@ static void fill_oids_from_all_packs(struct write_commit_graph_context *ctx)\n {\n \tstruct odb_source *source;\n \tenum object_type type;\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.flags = ODB_FOR_EACH_OBJECT_PACK_ORDER,\n+\t};\n \tstruct object_info oi = {\n \t\t.typep = &type,\n \t};\n@@ -1983,7 +1986,7 @@ static void fill_oids_from_all_packs(struct write_commit_graph_context *ctx)\n \tfor (source = ctx->r->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n \t\tpackfile_store_for_each_object(files->packed, &oi, add_packed_commits_oi,\n-\t\t\t\t\t       ctx, ODB_FOR_EACH_OBJECT_PACK_ORDER);\n+\t\t\t\t\t       ctx, &opts);\n \t}\n \n \tif (ctx->progress_done < ctx->approx_nr_objects)\ndiff --git a/object-file.c b/object-file.c\nindex f0b029ff0b..56cbb27ab9 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1849,7 +1849,7 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t     const struct object_info *request,\n \t\t\t\t     odb_for_each_object_cb cb,\n \t\t\t\t     void *cb_data,\n-\t\t\t\t     unsigned flags)\n+\t\t\t\t     const struct odb_for_each_object_options *opts)\n {\n \tstruct for_each_object_wrapper_data data = {\n \t\t.source = source,\n@@ -1859,9 +1859,9 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t};\n \n \t/* There are no loose promisor objects, so we can return immediately. */\n-\tif ((flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY))\n+\tif ((opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY))\n \t\treturn 0;\n-\tif ((flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY) && !source->local)\n+\tif ((opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY) && !source->local)\n \t\treturn 0;\n \n \treturn for_each_loose_file_in_source(source, for_each_object_wrapper_cb,\n@@ -1914,9 +1914,10 @@ int odb_source_loose_count_objects(struct odb_source *source,\n \t\t*out = count * 256;\n \t\tret = 0;\n \t} else {\n+\t\tstruct odb_for_each_object_options opts = { 0 };\n \t\t*out = 0;\n \t\tret = odb_source_loose_for_each_object(source, NULL, count_loose_object,\n-\t\t\t\t\t\t       out, 0);\n+\t\t\t\t\t\t       out, &opts);\n \t}\n \n out:\ndiff --git a/object-file.h b/object-file.h\nindex f8d8805a18..46dfa7b632 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -137,7 +137,7 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t     const struct object_info *request,\n \t\t\t\t     odb_for_each_object_cb cb,\n \t\t\t\t     void *cb_data,\n-\t\t\t\t     unsigned flags);\n+\t\t\t\t     const struct odb_for_each_object_options *opts);\n \n /*\n  * Count the number of loose objects in this source.\ndiff --git a/odb.c b/odb.c\nindex 350e23f3c0..3019957b87 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -896,20 +896,20 @@ int odb_freshen_object(struct object_database *odb,\n \treturn 0;\n }\n \n-int odb_for_each_object(struct object_database *odb,\n-\t\t\tconst struct object_info *request,\n-\t\t\todb_for_each_object_cb cb,\n-\t\t\tvoid *cb_data,\n-\t\t\tunsigned flags)\n+int odb_for_each_object_ext(struct object_database *odb,\n+\t\t\t    const struct object_info *request,\n+\t\t\t    odb_for_each_object_cb cb,\n+\t\t\t    void *cb_data,\n+\t\t\t    const struct odb_for_each_object_options *opts)\n {\n \tint ret;\n \n \todb_prepare_alternates(odb);\n \tfor (struct odb_source *source = odb->sources; source; source = source->next) {\n-\t\tif (flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY && !source->local)\n+\t\tif (opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY && !source->local)\n \t\t\tcontinue;\n \n-\t\tret = odb_source_for_each_object(source, request, cb, cb_data, flags);\n+\t\tret = odb_source_for_each_object(source, request, cb, cb_data, opts);\n \t\tif (ret)\n \t\t\treturn ret;\n \t}\n@@ -917,6 +917,18 @@ int odb_for_each_object(struct object_database *odb,\n \treturn 0;\n }\n \n+int odb_for_each_object(struct object_database *odb,\n+\t\t\tconst struct object_info *request,\n+\t\t\todb_for_each_object_cb cb,\n+\t\t\tvoid *cb_data,\n+\t\t\tunsigned flags)\n+{\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.flags = flags,\n+\t};\n+\treturn odb_for_each_object_ext(odb, request, cb, cb_data, &opts);\n+}\n+\n int odb_count_objects(struct object_database *odb,\n \t\t      enum odb_count_objects_flags flags,\n \t\t      unsigned long *out)\ndiff --git a/odb.h b/odb.h\nindex 9aee260105..a19a8bb50d 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -481,6 +481,15 @@ typedef int (*odb_for_each_object_cb)(const struct object_id *oid,\n \t\t\t\t      struct object_info *oi,\n \t\t\t\t      void *cb_data);\n \n+/*\n+ * Options that can be passed to `odb_for_each_object()` and its\n+ * backend-specific implementations.\n+ */\n+struct odb_for_each_object_options {\n+\t/* A bitfield of `odb_for_each_object_flags`. */\n+\tenum odb_for_each_object_flags flags;\n+};\n+\n /*\n  * Iterate through all objects contained in the object database. Note that\n  * objects may be iterated over multiple times in case they are either stored\n@@ -495,6 +504,13 @@ typedef int (*odb_for_each_object_cb)(const struct object_id *oid,\n  * Returns 0 on success, a negative error code in case a failure occurred, or\n  * an arbitrary non-zero error code returned by the callback itself.\n  */\n+int odb_for_each_object_ext(struct object_database *odb,\n+\t\t\t    const struct object_info *request,\n+\t\t\t    odb_for_each_object_cb cb,\n+\t\t\t    void *cb_data,\n+\t\t\t    const struct odb_for_each_object_options *opts);\n+\n+/* Same as `odb_for_each_object_ext()` with `opts.flags` set to the given flags. */\n int odb_for_each_object(struct object_database *odb,\n \t\t\tconst struct object_info *request,\n \t\t\todb_for_each_object_cb cb,\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex c08d8993e3..e90bb689bb 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -75,18 +75,18 @@ static int odb_source_files_for_each_object(struct odb_source *source,\n \t\t\t\t\t    const struct object_info *request,\n \t\t\t\t\t    odb_for_each_object_cb cb,\n \t\t\t\t\t    void *cb_data,\n-\t\t\t\t\t    unsigned flags)\n+\t\t\t\t\t    const struct odb_for_each_object_options *opts)\n {\n \tstruct odb_source_files *files = odb_source_files_downcast(source);\n \tint ret;\n \n-\tif (!(flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY)) {\n-\t\tret = odb_source_loose_for_each_object(source, request, cb, cb_data, flags);\n+\tif (!(opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY)) {\n+\t\tret = odb_source_loose_for_each_object(source, request, cb, cb_data, opts);\n \t\tif (ret)\n \t\t\treturn ret;\n \t}\n \n-\tret = packfile_store_for_each_object(files->packed, request, cb, cb_data, flags);\n+\tret = packfile_store_for_each_object(files->packed, request, cb, cb_data, opts);\n \tif (ret)\n \t\treturn ret;\n \ndiff --git a/odb/source.h b/odb/source.h\nindex 96c906e7a1..ee5d6ed530 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -140,7 +140,7 @@ struct odb_source {\n \t\t\t       const struct object_info *request,\n \t\t\t       odb_for_each_object_cb cb,\n \t\t\t       void *cb_data,\n-\t\t\t       unsigned flags);\n+\t\t\t       const struct odb_for_each_object_options *opts);\n \n \t/*\n \t * This callback is expected to count objects in the given object\n@@ -343,9 +343,9 @@ static inline int odb_source_for_each_object(struct odb_source *source,\n \t\t\t\t\t     const struct object_info *request,\n \t\t\t\t\t     odb_for_each_object_cb cb,\n \t\t\t\t\t     void *cb_data,\n-\t\t\t\t\t     unsigned flags)\n+\t\t\t\t\t     const struct odb_for_each_object_options *opts)\n {\n-\treturn source->for_each_object(source, request, cb, cb_data, flags);\n+\treturn source->for_each_object(source, request, cb, cb_data, opts);\n }\n \n /*\ndiff --git a/packfile.c b/packfile.c\nindex d4de9f3ffe..a6f3d2035d 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2375,7 +2375,7 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \t\t\t\t   const struct object_info *request,\n \t\t\t\t   odb_for_each_object_cb cb,\n \t\t\t\t   void *cb_data,\n-\t\t\t\t   unsigned flags)\n+\t\t\t\t   const struct odb_for_each_object_options *opts)\n {\n \tstruct packfile_store_for_each_object_wrapper_data data = {\n \t\t.store = store,\n@@ -2391,15 +2391,15 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n \t\tstruct packed_git *p = e->pack;\n \n-\t\tif ((flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY) && !p->pack_local)\n+\t\tif ((opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY) && !p->pack_local)\n \t\t\tcontinue;\n-\t\tif ((flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY) &&\n+\t\tif ((opts->flags & ODB_FOR_EACH_OBJECT_PROMISOR_ONLY) &&\n \t\t    !p->pack_promisor)\n \t\t\tcontinue;\n-\t\tif ((flags & ODB_FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS) &&\n+\t\tif ((opts->flags & ODB_FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS) &&\n \t\t    p->pack_keep_in_core)\n \t\t\tcontinue;\n-\t\tif ((flags & ODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS) &&\n+\t\tif ((opts->flags & ODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS) &&\n \t\t    p->pack_keep)\n \t\t\tcontinue;\n \t\tif (open_pack_index(p)) {\n@@ -2408,7 +2408,7 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \t\t}\n \n \t\tret = for_each_object_in_pack(p, packfile_store_for_each_object_wrapper,\n-\t\t\t\t\t      &data, flags);\n+\t\t\t\t\t      &data, opts->flags);\n \t\tif (ret)\n \t\t\tgoto out;\n \t}\ndiff --git a/packfile.h b/packfile.h\nindex a16ec3950d..fa41dfda38 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -367,7 +367,7 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \t\t\t\t   const struct object_info *request,\n \t\t\t\t   odb_for_each_object_cb cb,\n \t\t\t\t   void *cb_data,\n-\t\t\t\t   unsigned flags);\n+\t\t\t\t   const struct odb_for_each_object_options *opts);\n \n /* A hook to report invalid files in pack directory */\n #define PACKDIR_FILE_PACK 1\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539484","messageId":"20260320-b4-pks-odb-source-abbrev-v2-4-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 04/14] object-name: move logic to iterate through loose prefixed objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:30Z","receivedAt":"2026-03-20T07:07:53Z","isPatch":true,"body":"The logic to iterate through loose objects that have a certain prefix is\ncurrently hosted in \"object-name.c\". This logic reaches into specifics\nof the loose object source, so it breaks once a different backend is\nused for the object storage.\n\nMove the logic to iterate through loose objects with a prefix into\n\"object-file.c\". This is done by extending the for-each-object options\nto support an optional prefix that is then honored by the loose source.\nNaturally, we'll also have this support in the packfile store. This is\ndone in the next commit.\n\nFurthermore, there are no users of the loose cache outside of\n\"object-file.c\" anymore. As such, convert `odb_source_loose_cache()` to\nhave file scope.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-file.c | 29 +++++++++++++++++++++++++++--\n object-file.h |  7 -------\n object-name.c | 10 ++++++----\n odb.h         |  7 +++++++\n 4 files changed, 40 insertions(+), 13 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 56cbb27ab9..13732f324f 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -33,6 +33,9 @@\n /* The maximum size for an object header. */\n #define MAX_HEADER_LEN 32\n \n+static struct oidtree *odb_source_loose_cache(struct odb_source *source,\n+\t\t\t\t\t      const struct object_id *oid);\n+\n static int get_conv_flags(unsigned flags)\n {\n \tif (flags & INDEX_RENORMALIZE)\n@@ -1845,6 +1848,23 @@ static int for_each_object_wrapper_cb(const struct object_id *oid,\n \t}\n }\n \n+static int for_each_prefixed_object_wrapper_cb(const struct object_id *oid,\n+\t\t\t\t\t       void *cb_data)\n+{\n+\tstruct for_each_object_wrapper_data *data = cb_data;\n+\tif (data->request) {\n+\t\tstruct object_info oi = *data->request;\n+\n+\t\tif (odb_source_loose_read_object_info(data->source,\n+\t\t\t\t\t\t      oid, &oi, 0) < 0)\n+\t\t\treturn -1;\n+\n+\t\treturn data->cb(oid, &oi, data->cb_data);\n+\t} else {\n+\t\treturn data->cb(oid, NULL, data->cb_data);\n+\t}\n+}\n+\n int odb_source_loose_for_each_object(struct odb_source *source,\n \t\t\t\t     const struct object_info *request,\n \t\t\t\t     odb_for_each_object_cb cb,\n@@ -1864,6 +1884,11 @@ int odb_source_loose_for_each_object(struct odb_source *source,\n \tif ((opts->flags & ODB_FOR_EACH_OBJECT_LOCAL_ONLY) && !source->local)\n \t\treturn 0;\n \n+\tif (opts->prefix)\n+\t\treturn oidtree_each(odb_source_loose_cache(source, opts->prefix),\n+\t\t\t\t    opts->prefix, opts->prefix_hex_len,\n+\t\t\t\t    for_each_prefixed_object_wrapper_cb, &data);\n+\n \treturn for_each_loose_file_in_source(source, for_each_object_wrapper_cb,\n \t\t\t\t\t     NULL, NULL, &data);\n }\n@@ -1935,8 +1960,8 @@ static int append_loose_object(const struct object_id *oid,\n \treturn 0;\n }\n \n-struct oidtree *odb_source_loose_cache(struct odb_source *source,\n-\t\t\t\t       const struct object_id *oid)\n+static struct oidtree *odb_source_loose_cache(struct odb_source *source,\n+\t\t\t\t\t      const struct object_id *oid)\n {\n \tstruct odb_source_files *files = odb_source_files_downcast(source);\n \tint subdir_nr = oid->hash[0];\ndiff --git a/object-file.h b/object-file.h\nindex 46dfa7b632..f11ad58f6c 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -74,13 +74,6 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \t\t\t\t  struct odb_write_stream *stream, size_t len,\n \t\t\t\t  struct object_id *oid);\n \n-/*\n- * Populate and return the loose object cache array corresponding to the\n- * given object ID.\n- */\n-struct oidtree *odb_source_loose_cache(struct odb_source *source,\n-\t\t\t\t       const struct object_id *oid);\n-\n /*\n  * Put in `buf` the name of the file in the local object database that\n  * would be used to store a loose object with the specified oid.\ndiff --git a/object-name.c b/object-name.c\nindex a24a1b48e1..929a68dbd0 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -16,7 +16,6 @@\n #include \"remote.h\"\n #include \"dir.h\"\n #include \"oid-array.h\"\n-#include \"oidtree.h\"\n #include \"packfile.h\"\n #include \"pretty.h\"\n #include \"object-file.h\"\n@@ -103,7 +102,7 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n \n static int match_hash(unsigned, const unsigned char *, const unsigned char *);\n \n-static int match_prefix(const struct object_id *oid, void *arg)\n+static int match_prefix(const struct object_id *oid, struct object_info *oi UNUSED, void *arg)\n {\n \tstruct disambiguate_state *ds = arg;\n \t/* no need to call match_hash, oidtree_each did prefix match */\n@@ -113,11 +112,14 @@ static int match_prefix(const struct object_id *oid, void *arg)\n \n static void find_short_object_filename(struct disambiguate_state *ds)\n {\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = &ds->bin_pfx,\n+\t\t.prefix_hex_len = ds->len,\n+\t};\n \tstruct odb_source *source;\n \n \tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next)\n-\t\toidtree_each(odb_source_loose_cache(source, &ds->bin_pfx),\n-\t\t\t\t&ds->bin_pfx, ds->len, match_prefix, ds);\n+\t\todb_source_loose_for_each_object(source, NULL, match_prefix, ds, &opts);\n }\n \n static int match_hash(unsigned len, const unsigned char *a, const unsigned char *b)\ndiff --git a/odb.h b/odb.h\nindex a19a8bb50d..e80fd8f7ab 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -488,6 +488,13 @@ typedef int (*odb_for_each_object_cb)(const struct object_id *oid,\n struct odb_for_each_object_options {\n \t/* A bitfield of `odb_for_each_object_flags`. */\n \tenum odb_for_each_object_flags flags;\n+\n+\t/*\n+\t * If set, only iterate through objects whose first `prefix_hex_len`\n+\t * hex characters matches the given prefix.\n+\t */\n+\tconst struct object_id *prefix;\n+\tsize_t prefix_hex_len;\n };\n \n /*\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539485","messageId":"20260320-b4-pks-odb-source-abbrev-v2-5-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 05/14] object-name: move logic to iterate through packed prefixed objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:31Z","receivedAt":"2026-03-20T07:07:56Z","isPatch":true,"body":"Similar to the preceding commit, move the logic to iterate through\nobjects that have a given prefix into \"packfile.c\".\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c |  94 +++----------------------------\n packfile.c    | 174 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 2 files changed, 181 insertions(+), 87 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex 929a68dbd0..ff0de06ff9 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -100,8 +100,6 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n \t/* otherwise, current can be discarded and candidate is still good */\n }\n \n-static int match_hash(unsigned, const unsigned char *, const unsigned char *);\n-\n static int match_prefix(const struct object_id *oid, struct object_info *oi UNUSED, void *arg)\n {\n \tstruct disambiguate_state *ds = arg;\n@@ -122,103 +120,25 @@ static void find_short_object_filename(struct disambiguate_state *ds)\n \t\todb_source_loose_for_each_object(source, NULL, match_prefix, ds, &opts);\n }\n \n-static int match_hash(unsigned len, const unsigned char *a, const unsigned char *b)\n-{\n-\tdo {\n-\t\tif (*a != *b)\n-\t\t\treturn 0;\n-\t\ta++;\n-\t\tb++;\n-\t\tlen -= 2;\n-\t} while (len > 1);\n-\tif (len)\n-\t\tif ((*a ^ *b) & 0xf0)\n-\t\t\treturn 0;\n-\treturn 1;\n-}\n-\n-static void unique_in_midx(struct multi_pack_index *m,\n-\t\t\t   struct disambiguate_state *ds)\n-{\n-\tfor (; m; m = m->base_midx) {\n-\t\tuint32_t num, i, first = 0;\n-\t\tconst struct object_id *current = NULL;\n-\t\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n-\t\t\tds->repo->hash_algo->hexsz : ds->len;\n-\n-\t\tif (!m->num_objects)\n-\t\t\tcontinue;\n-\n-\t\tnum = m->num_objects + m->num_objects_in_base;\n-\n-\t\tbsearch_one_midx(&ds->bin_pfx, m, &first);\n-\n-\t\t/*\n-\t\t * At this point, \"first\" is the location of the lowest\n-\t\t * object with an object name that could match\n-\t\t * \"bin_pfx\".  See if we have 0, 1 or more objects that\n-\t\t * actually match(es).\n-\t\t */\n-\t\tfor (i = first; i < num && !ds->ambiguous; i++) {\n-\t\t\tstruct object_id oid;\n-\t\t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n-\t\t\tif (!match_hash(len, ds->bin_pfx.hash, current->hash))\n-\t\t\t\tbreak;\n-\t\t\tupdate_candidates(ds, current);\n-\t\t}\n-\t}\n-}\n-\n-static void unique_in_pack(struct packed_git *p,\n-\t\t\t   struct disambiguate_state *ds)\n-{\n-\tuint32_t num, i, first = 0;\n-\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n-\t\tds->repo->hash_algo->hexsz : ds->len;\n-\n-\tif (p->multi_pack_index)\n-\t\treturn;\n-\n-\tif (open_pack_index(p) || !p->num_objects)\n-\t\treturn;\n-\n-\tnum = p->num_objects;\n-\tbsearch_pack(&ds->bin_pfx, p, &first);\n-\n-\t/*\n-\t * At this point, \"first\" is the location of the lowest object\n-\t * with an object name that could match \"bin_pfx\".  See if we have\n-\t * 0, 1 or more objects that actually match(es).\n-\t */\n-\tfor (i = first; i < num && !ds->ambiguous; i++) {\n-\t\tstruct object_id oid;\n-\t\tnth_packed_object_id(&oid, p, i);\n-\t\tif (!match_hash(len, ds->bin_pfx.hash, oid.hash))\n-\t\t\tbreak;\n-\t\tupdate_candidates(ds, &oid);\n-\t}\n-}\n-\n static void find_short_packed_object(struct disambiguate_state *ds)\n {\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = &ds->bin_pfx,\n+\t\t.prefix_hex_len = ds->len,\n+\t};\n \tstruct odb_source *source;\n-\tstruct packed_git *p;\n \n \t/* Skip, unless oids from the storage hash algorithm are wanted */\n \tif (ds->bin_pfx.algo && (&hash_algos[ds->bin_pfx.algo] != ds->repo->hash_algo))\n \t\treturn;\n \n \todb_prepare_alternates(ds->repo->objects);\n-\tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next) {\n-\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n-\t\tif (m)\n-\t\t\tunique_in_midx(m, ds);\n-\t}\n+\tfor (source = ds->repo->objects->sources; source; source = source->next) {\n+\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n \n-\trepo_for_each_pack(ds->repo, p) {\n+\t\tpackfile_store_for_each_object(files->packed, NULL, match_prefix, ds, &opts);\n \t\tif (ds->ambiguous)\n \t\t\tbreak;\n-\t\tunique_in_pack(p, ds);\n \t}\n }\n \ndiff --git a/packfile.c b/packfile.c\nindex a6f3d2035d..2539a371c1 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2371,6 +2371,177 @@ static int packfile_store_for_each_object_wrapper(const struct object_id *oid,\n \t}\n }\n \n+static int match_hash(unsigned len, const unsigned char *a, const unsigned char *b)\n+{\n+\tdo {\n+\t\tif (*a != *b)\n+\t\t\treturn 0;\n+\t\ta++;\n+\t\tb++;\n+\t\tlen -= 2;\n+\t} while (len > 1);\n+\tif (len)\n+\t\tif ((*a ^ *b) & 0xf0)\n+\t\t\treturn 0;\n+\treturn 1;\n+}\n+\n+static int for_each_prefixed_object_in_midx(\n+\tstruct packfile_store *store,\n+\tstruct multi_pack_index *m,\n+\tconst struct odb_for_each_object_options *opts,\n+\tstruct packfile_store_for_each_object_wrapper_data *data)\n+{\n+\tint ret;\n+\n+\tfor (; m; m = m->base_midx) {\n+\t\tuint32_t num, i, first = 0;\n+\t\tint len = opts->prefix_hex_len > m->source->odb->repo->hash_algo->hexsz ?\n+\t\t\tm->source->odb->repo->hash_algo->hexsz : opts->prefix_hex_len;\n+\n+\t\tif (!m->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tnum = m->num_objects + m->num_objects_in_base;\n+\n+\t\tbsearch_one_midx(opts->prefix, m, &first);\n+\n+\t\t/*\n+\t\t * At this point, \"first\" is the location of the lowest\n+\t\t * object with an object name that could match \"opts->prefix\".\n+\t\t * See if we have 0, 1 or more objects that actually match(es).\n+\t\t */\n+\t\tfor (i = first; i < num; i++) {\n+\t\t\tconst struct object_id *current = NULL;\n+\t\t\tstruct object_id oid;\n+\n+\t\t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n+\n+\t\t\tif (!match_hash(len, opts->prefix->hash, current->hash))\n+\t\t\t\tbreak;\n+\n+\t\t\tif (data->request) {\n+\t\t\t\tstruct object_info oi = *data->request;\n+\n+\t\t\t\tret = packfile_store_read_object_info(store, current,\n+\t\t\t\t\t\t\t\t      &oi, 0);\n+\t\t\t\tif (ret)\n+\t\t\t\t\tgoto out;\n+\n+\t\t\t\tret = data->cb(&oid, &oi, data->cb_data);\n+\t\t\t\tif (ret)\n+\t\t\t\t\tgoto out;\n+\t\t\t} else {\n+\t\t\t\tret = data->cb(&oid, NULL, data->cb_data);\n+\t\t\t\tif (ret)\n+\t\t\t\t\tgoto out;\n+\t\t\t}\n+\t\t}\n+\t}\n+\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n+static int for_each_prefixed_object_in_pack(\n+\tstruct packfile_store *store,\n+\tstruct packed_git *p,\n+\tconst struct odb_for_each_object_options *opts,\n+\tstruct packfile_store_for_each_object_wrapper_data *data)\n+{\n+\tuint32_t num, i, first = 0;\n+\tint len = opts->prefix_hex_len > p->repo->hash_algo->hexsz ?\n+\t\tp->repo->hash_algo->hexsz : opts->prefix_hex_len;\n+\tint ret;\n+\n+\tnum = p->num_objects;\n+\tbsearch_pack(opts->prefix, p, &first);\n+\n+\t/*\n+\t * At this point, \"first\" is the location of the lowest object\n+\t * with an object name that could match \"bin_pfx\".  See if we have\n+\t * 0, 1 or more objects that actually match(es).\n+\t */\n+\tfor (i = first; i < num; i++) {\n+\t\tstruct object_id oid;\n+\n+\t\tnth_packed_object_id(&oid, p, i);\n+\t\tif (!match_hash(len, opts->prefix->hash, oid.hash))\n+\t\t\tbreak;\n+\n+\t\tif (data->request) {\n+\t\t\tstruct object_info oi = *data->request;\n+\n+\t\t\tret = packfile_store_read_object_info(store, &oid, &oi, 0);\n+\t\t\tif (ret)\n+\t\t\t\tgoto out;\n+\n+\t\t\tret = data->cb(&oid, &oi, data->cb_data);\n+\t\t\tif (ret)\n+\t\t\t\tgoto out;\n+\t\t} else {\n+\t\t\tret = data->cb(&oid, NULL, data->cb_data);\n+\t\t\tif (ret)\n+\t\t\t\tgoto out;\n+\t\t}\n+\t}\n+\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n+static int packfile_store_for_each_prefixed_object(\n+\tstruct packfile_store *store,\n+\tconst struct odb_for_each_object_options *opts,\n+\tstruct packfile_store_for_each_object_wrapper_data *data)\n+{\n+\tstruct packfile_list_entry *e;\n+\tstruct multi_pack_index *m;\n+\tbool pack_errors = false;\n+\tint ret;\n+\n+\tif (opts->flags)\n+\t\tBUG(\"flags unsupported\");\n+\n+\tstore->skip_mru_updates = true;\n+\n+\tm = get_multi_pack_index(store->source);\n+\tif (m) {\n+\t\tret = for_each_prefixed_object_in_midx(store, m, opts, data);\n+\t\tif (ret)\n+\t\t\tgoto out;\n+\t}\n+\n+\tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n+\t\tif (e->pack->multi_pack_index)\n+\t\t\tcontinue;\n+\n+\t\tif (open_pack_index(e->pack)) {\n+\t\t\tpack_errors = true;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tif (!e->pack->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tret = for_each_prefixed_object_in_pack(store, e->pack, opts, data);\n+\t\tif (ret)\n+\t\t\tgoto out;\n+\t}\n+\n+\tret = 0;\n+\n+out:\n+\tstore->skip_mru_updates = false;\n+\tif (!ret && pack_errors)\n+\t\tret = -1;\n+\treturn ret;\n+}\n+\n int packfile_store_for_each_object(struct packfile_store *store,\n \t\t\t\t   const struct object_info *request,\n \t\t\t\t   odb_for_each_object_cb cb,\n@@ -2386,6 +2557,9 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \tstruct packfile_list_entry *e;\n \tint pack_errors = 0, ret;\n \n+\tif (opts->prefix)\n+\t\treturn packfile_store_for_each_prefixed_object(store, opts, &data);\n+\n \tstore->skip_mru_updates = true;\n \n \tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539486","messageId":"20260320-b4-pks-odb-source-abbrev-v2-6-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 06/14] object-name: extract function to parse object ID prefixes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:32Z","receivedAt":"2026-03-20T07:07:59Z","isPatch":true,"body":"Extract the logic that parses an object ID prefix into a new function.\nThis function will be used by a second callsite in a subsequent commit.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 60 +++++++++++++++++++++++++++++++++++++----------------------\n 1 file changed, 38 insertions(+), 22 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex ff0de06ff9..fd1b010ab3 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -270,41 +270,57 @@ int set_disambiguate_hint_config(const char *var, const char *value)\n \treturn error(\"unknown hint type for '%s': %s\", var, value);\n }\n \n+static int parse_oid_prefix(const char *name, int len,\n+\t\t\t    const struct git_hash_algo *algo,\n+\t\t\t    char *hex_out,\n+\t\t\t    struct object_id *oid_out)\n+{\n+\tfor (int i = 0; i < len; i++) {\n+\t\tunsigned char c = name[i];\n+\t\tunsigned char val;\n+\t\tif (c >= '0' && c <= '9') {\n+\t\t\tval = c - '0';\n+\t\t} else if (c >= 'a' && c <= 'f') {\n+\t\t\tval = c - 'a' + 10;\n+\t\t} else if (c >= 'A' && c <='F') {\n+\t\t\tval = c - 'A' + 10;\n+\t\t\tc -= 'A' - 'a';\n+\t\t} else {\n+\t\t\treturn -1;\n+\t\t}\n+\n+\t\tif (hex_out)\n+\t\t\thex_out[i] = c;\n+\t\tif (oid_out) {\n+\t\t\tif (!(i & 1))\n+\t\t\t\tval <<= 4;\n+\t\t\toid_out->hash[i >> 1] |= val;\n+\t\t}\n+\t}\n+\n+\tif (hex_out)\n+\t\thex_out[len] = '\\0';\n+\tif (oid_out)\n+\t\toid_out->algo = algo ? hash_algo_by_ptr(algo) : GIT_HASH_UNKNOWN;\n+\n+\treturn 0;\n+}\n+\n static int init_object_disambiguation(struct repository *r,\n \t\t\t\t      const char *name, int len,\n \t\t\t\t      const struct git_hash_algo *algo,\n \t\t\t\t      struct disambiguate_state *ds)\n {\n-\tint i;\n-\n \tif (len < MINIMUM_ABBREV || len > GIT_MAX_HEXSZ)\n \t\treturn -1;\n \n \tmemset(ds, 0, sizeof(*ds));\n \n-\tfor (i = 0; i < len ;i++) {\n-\t\tunsigned char c = name[i];\n-\t\tunsigned char val;\n-\t\tif (c >= '0' && c <= '9')\n-\t\t\tval = c - '0';\n-\t\telse if (c >= 'a' && c <= 'f')\n-\t\t\tval = c - 'a' + 10;\n-\t\telse if (c >= 'A' && c <='F') {\n-\t\t\tval = c - 'A' + 10;\n-\t\t\tc -= 'A' - 'a';\n-\t\t}\n-\t\telse\n-\t\t\treturn -1;\n-\t\tds->hex_pfx[i] = c;\n-\t\tif (!(i & 1))\n-\t\t\tval <<= 4;\n-\t\tds->bin_pfx.hash[i >> 1] |= val;\n-\t}\n+\tif (parse_oid_prefix(name, len, algo, ds->hex_pfx, &ds->bin_pfx) < 0)\n+\t\treturn -1;\n \n \tds->len = len;\n-\tds->hex_pfx[len] = '\\0';\n \tds->repo = r;\n-\tds->bin_pfx.algo = algo ? hash_algo_by_ptr(algo) : GIT_HASH_UNKNOWN;\n \todb_prepare_alternates(r->objects);\n \treturn 0;\n }\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539487","messageId":"20260320-b4-pks-odb-source-abbrev-v2-7-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 07/14] object-name: backend-generic `repo_collect_ambiguous()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:33Z","receivedAt":"2026-03-20T07:08:01Z","isPatch":true,"body":"The function `repo_collect_ambiguous()` is responsible for collecting\nobjects whose IDs match a specific prefix. The information is then\nused to inform the user about which objects they could have meant in\ncase a short object ID is ambiguous.\n\nThe logic to do this uses the object disambiguation infrastructure and\ncalls into backend-specific functions to iterate through loose and\npacked objects. This isn't really required anymore though: all we want\nto do is to enumerate objects that have such a prefix and then append\nthose objects to a `struct oid_array`. This can be trivially achieved\nin a generic way now that `odb_for_each_object()` has learned to yield\nonly objects that match such a prefix.\n\nRefactor the code to use the backend-generic infrastructure instead.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 19 ++++++++++---------\n 1 file changed, 10 insertions(+), 9 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex fd1b010ab3..4c3ace150e 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -448,8 +448,8 @@ static int collect_ambiguous(const struct object_id *oid, void *data)\n \treturn 0;\n }\n \n-static int repo_collect_ambiguous(struct repository *r UNUSED,\n-\t\t\t\t  const struct object_id *oid,\n+static int repo_collect_ambiguous(const struct object_id *oid,\n+\t\t\t\t  struct object_info *oi UNUSED,\n \t\t\t\t  void *data)\n {\n \treturn collect_ambiguous(oid, data);\n@@ -586,18 +586,19 @@ int repo_for_each_abbrev(struct repository *r, const char *prefix,\n \t\t\t const struct git_hash_algo *algo,\n \t\t\t each_abbrev_fn fn, void *cb_data)\n {\n+\tstruct object_id prefix_oid = { 0 };\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = &prefix_oid,\n+\t\t.prefix_hex_len = strlen(prefix),\n+\t};\n \tstruct oid_array collect = OID_ARRAY_INIT;\n-\tstruct disambiguate_state ds;\n \tint ret;\n \n-\tif (init_object_disambiguation(r, prefix, strlen(prefix), algo, &ds) < 0)\n+\tif (parse_oid_prefix(prefix, opts.prefix_hex_len, algo, NULL, &prefix_oid) < 0)\n \t\treturn -1;\n \n-\tds.always_call_fn = 1;\n-\tds.fn = repo_collect_ambiguous;\n-\tds.cb_data = &collect;\n-\tfind_short_object_filename(&ds);\n-\tfind_short_packed_object(&ds);\n+\tif (odb_for_each_object_ext(r->objects, NULL, repo_collect_ambiguous, &collect, &opts) < 0)\n+\t\treturn -1;\n \n \tret = oid_array_for_each_unique(&collect, fn, cb_data);\n \toid_array_clear(&collect);\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539488","messageId":"20260320-b4-pks-odb-source-abbrev-v2-8-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 08/14] object-name: backend-generic `get_short_oid()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:34Z","receivedAt":"2026-03-20T07:08:04Z","isPatch":true,"body":"The function `get_short_oid()` takes as input an abbreviated object ID\nand tries to turn that object ID into the full object ID. This is done\nby iterating through all objects that have the user-provided prefix. If\nthat yields exactly one object we know that the abbreviated object ID is\nunambiguous, otherwise it is ambiguous and we print the list of objects\nthat match the prefix.\n\nWe iterate through all objects with the given prefix by calling both\n`find_short_packed_object()` and `find_short_object_filename()`, which\nis of course specific to the \"files\" backend. But we now have a generic\nway to iterate through objects with a specific prefix.\n\nRefactor the code to use `odb_for_each_object()` instead so that it\nworks with object backends different than the \"files\" backend.\n\nRemove the now-unused `find_short_packed_object()` function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 32 ++++++--------------------------\n 1 file changed, 6 insertions(+), 26 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex 4c3ace150e..7a224ab4af 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -120,28 +120,6 @@ static void find_short_object_filename(struct disambiguate_state *ds)\n \t\todb_source_loose_for_each_object(source, NULL, match_prefix, ds, &opts);\n }\n \n-static void find_short_packed_object(struct disambiguate_state *ds)\n-{\n-\tstruct odb_for_each_object_options opts = {\n-\t\t.prefix = &ds->bin_pfx,\n-\t\t.prefix_hex_len = ds->len,\n-\t};\n-\tstruct odb_source *source;\n-\n-\t/* Skip, unless oids from the storage hash algorithm are wanted */\n-\tif (ds->bin_pfx.algo && (&hash_algos[ds->bin_pfx.algo] != ds->repo->hash_algo))\n-\t\treturn;\n-\n-\todb_prepare_alternates(ds->repo->objects);\n-\tfor (source = ds->repo->objects->sources; source; source = source->next) {\n-\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\n-\t\tpackfile_store_for_each_object(files->packed, NULL, match_prefix, ds, &opts);\n-\t\tif (ds->ambiguous)\n-\t\t\tbreak;\n-\t}\n-}\n-\n static int finish_object_disambiguation(struct disambiguate_state *ds,\n \t\t\t\t\tstruct object_id *oid)\n {\n@@ -499,6 +477,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \t\t\t\t\t struct object_id *oid,\n \t\t\t\t\t unsigned flags)\n {\n+\tstruct odb_for_each_object_options opts = { 0 };\n \tint status;\n \tstruct disambiguate_state ds;\n \tint quietly = !!(flags & GET_OID_QUIETLY);\n@@ -526,8 +505,10 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \telse\n \t\tds.fn = default_disambiguate_hint;\n \n-\tfind_short_object_filename(&ds);\n-\tfind_short_packed_object(&ds);\n+\topts.prefix = &ds.bin_pfx;\n+\topts.prefix_hex_len = ds.len;\n+\n+\todb_for_each_object_ext(r->objects, NULL, match_prefix, &ds, &opts);\n \tstatus = finish_object_disambiguation(&ds, oid);\n \n \t/*\n@@ -537,8 +518,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \t */\n \tif (status == MISSING_OBJECT) {\n \t\todb_reprepare(r->objects);\n-\t\tfind_short_object_filename(&ds);\n-\t\tfind_short_packed_object(&ds);\n+\t\todb_for_each_object_ext(r->objects, NULL, match_prefix, &ds, &opts);\n \t\tstatus = finish_object_disambiguation(&ds, oid);\n \t}\n \n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539489","messageId":"20260320-b4-pks-odb-source-abbrev-v2-9-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 09/14] object-name: merge `update_candidates()` and `match_prefix()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:35Z","receivedAt":"2026-03-20T07:08:07Z","isPatch":true,"body":"There's only a single callsite for `match_prefix()`, and that function\nis a rather trivial wrapper of `update_candidates()`. Merge these two\nfunctions into a single `update_disambiguate_state()` function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 34 ++++++++++++++++++----------------\n 1 file changed, 18 insertions(+), 16 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex 7a224ab4af..f55a332032 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -51,27 +51,31 @@ struct disambiguate_state {\n \tunsigned always_call_fn:1;\n };\n \n-static void update_candidates(struct disambiguate_state *ds, const struct object_id *current)\n+static int update_disambiguate_state(const struct object_id *current,\n+\t\t\t\t     struct object_info *oi UNUSED,\n+\t\t\t\t     void *cb_data)\n {\n+\tstruct disambiguate_state *ds = cb_data;\n+\n \t/* The hash algorithm of current has already been filtered */\n \tif (ds->always_call_fn) {\n \t\tds->ambiguous = ds->fn(ds->repo, current, ds->cb_data) ? 1 : 0;\n-\t\treturn;\n+\t\treturn ds->ambiguous;\n \t}\n \tif (!ds->candidate_exists) {\n \t\t/* this is the first candidate */\n \t\toidcpy(&ds->candidate, current);\n \t\tds->candidate_exists = 1;\n-\t\treturn;\n+\t\treturn 0;\n \t} else if (oideq(&ds->candidate, current)) {\n \t\t/* the same as what we already have seen */\n-\t\treturn;\n+\t\treturn 0;\n \t}\n \n \tif (!ds->fn) {\n \t\t/* cannot disambiguate between ds->candidate and current */\n \t\tds->ambiguous = 1;\n-\t\treturn;\n+\t\treturn ds->ambiguous;\n \t}\n \n \tif (!ds->candidate_checked) {\n@@ -84,7 +88,7 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n \t\t/* discard the candidate; we know it does not satisfy fn */\n \t\toidcpy(&ds->candidate, current);\n \t\tds->candidate_checked = 0;\n-\t\treturn;\n+\t\treturn 0;\n \t}\n \n \t/* if we reach this point, we know ds->candidate satisfies fn */\n@@ -95,17 +99,12 @@ static void update_candidates(struct disambiguate_state *ds, const struct object\n \t\t */\n \t\tds->candidate_ok = 0;\n \t\tds->ambiguous = 1;\n+\t\treturn ds->ambiguous;\n \t}\n \n \t/* otherwise, current can be discarded and candidate is still good */\n-}\n \n-static int match_prefix(const struct object_id *oid, struct object_info *oi UNUSED, void *arg)\n-{\n-\tstruct disambiguate_state *ds = arg;\n-\t/* no need to call match_hash, oidtree_each did prefix match */\n-\tupdate_candidates(ds, oid);\n-\treturn ds->ambiguous;\n+\treturn 0;\n }\n \n static void find_short_object_filename(struct disambiguate_state *ds)\n@@ -117,7 +116,8 @@ static void find_short_object_filename(struct disambiguate_state *ds)\n \tstruct odb_source *source;\n \n \tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next)\n-\t\todb_source_loose_for_each_object(source, NULL, match_prefix, ds, &opts);\n+\t\todb_source_loose_for_each_object(source, NULL, update_disambiguate_state,\n+\t\t\t\t\t\t ds, &opts);\n }\n \n static int finish_object_disambiguation(struct disambiguate_state *ds,\n@@ -508,7 +508,8 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \topts.prefix = &ds.bin_pfx;\n \topts.prefix_hex_len = ds.len;\n \n-\todb_for_each_object_ext(r->objects, NULL, match_prefix, &ds, &opts);\n+\todb_for_each_object_ext(r->objects, NULL, update_disambiguate_state,\n+\t\t\t\t&ds, &opts);\n \tstatus = finish_object_disambiguation(&ds, oid);\n \n \t/*\n@@ -518,7 +519,8 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \t */\n \tif (status == MISSING_OBJECT) {\n \t\todb_reprepare(r->objects);\n-\t\todb_for_each_object_ext(r->objects, NULL, match_prefix, &ds, &opts);\n+\t\todb_for_each_object_ext(r->objects, NULL, update_disambiguate_state,\n+\t\t\t\t\t&ds, &opts);\n \t\tstatus = finish_object_disambiguation(&ds, oid);\n \t}\n \n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539490","messageId":"20260320-b4-pks-odb-source-abbrev-v2-10-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 10/14] object-name: abbreviate loose object names without `disambiguate_state`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:36Z","receivedAt":"2026-03-20T07:08:10Z","isPatch":true,"body":"The function `find_short_object_filename()` takes an object ID and\ncomputes the minimum required object name length to make it unique. This\nis done by reusing the object disambiguation infrastructure, where we\niterate through every loose object and then update the disambiguate\nstate one by one.\n\nUltimately, we don't care about the disambiguate state though. It is\nused because this infrastructure knows how to enumerate only those\nobjects that match a given prefix. But now that we have extended the\n`odb_for_each_object()` function to do this for us we have an easier way\nto do this. Consequently, we really only use the disambiguate state now\nto propagate `struct min_abbrev_data`.\n\nRefactor the code and drop this indirection so that we use `struct\nmin_abbrev_data` directly. This also allows us to drop some now-unused\nlogic from the disambiguate infrastructure.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 54 ++++++++++++++++++++----------------------------------\n 1 file changed, 20 insertions(+), 34 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex f55a332032..d82fb49f39 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -48,7 +48,6 @@ struct disambiguate_state {\n \tunsigned candidate_ok:1;\n \tunsigned disambiguate_fn_used:1;\n \tunsigned ambiguous:1;\n-\tunsigned always_call_fn:1;\n };\n \n static int update_disambiguate_state(const struct object_id *current,\n@@ -58,10 +57,6 @@ static int update_disambiguate_state(const struct object_id *current,\n \tstruct disambiguate_state *ds = cb_data;\n \n \t/* The hash algorithm of current has already been filtered */\n-\tif (ds->always_call_fn) {\n-\t\tds->ambiguous = ds->fn(ds->repo, current, ds->cb_data) ? 1 : 0;\n-\t\treturn ds->ambiguous;\n-\t}\n \tif (!ds->candidate_exists) {\n \t\t/* this is the first candidate */\n \t\toidcpy(&ds->candidate, current);\n@@ -107,19 +102,6 @@ static int update_disambiguate_state(const struct object_id *current,\n \treturn 0;\n }\n \n-static void find_short_object_filename(struct disambiguate_state *ds)\n-{\n-\tstruct odb_for_each_object_options opts = {\n-\t\t.prefix = &ds->bin_pfx,\n-\t\t.prefix_hex_len = ds->len,\n-\t};\n-\tstruct odb_source *source;\n-\n-\tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next)\n-\t\todb_source_loose_for_each_object(source, NULL, update_disambiguate_state,\n-\t\t\t\t\t\t ds, &opts);\n-}\n-\n static int finish_object_disambiguation(struct disambiguate_state *ds,\n \t\t\t\t\tstruct object_id *oid)\n {\n@@ -632,11 +614,26 @@ static int extend_abbrev_len(const struct object_id *oid,\n \treturn 0;\n }\n \n-static int repo_extend_abbrev_len(struct repository *r UNUSED,\n-\t\t\t\t  const struct object_id *oid,\n-\t\t\t\t  void *cb_data)\n+static int extend_abbrev_len_loose(const struct object_id *oid,\n+\t\t\t\t   struct object_info *oi UNUSED,\n+\t\t\t\t   void *cb_data)\n {\n-\treturn extend_abbrev_len(oid, cb_data);\n+\tstruct min_abbrev_data *data = cb_data;\n+\textend_abbrev_len(oid, data);\n+\treturn 0;\n+}\n+\n+static void find_abbrev_len_loose(struct min_abbrev_data *mad)\n+{\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = mad->oid,\n+\t\t.prefix_hex_len = mad->cur_len,\n+\t};\n+\tstruct odb_source *source;\n+\n+\tfor (source = mad->repo->objects->sources; source; source = source->next)\n+\t\todb_source_loose_for_each_object(source, NULL, extend_abbrev_len_loose,\n+\t\t\t\t\t\t mad, &opts);\n }\n \n static void find_abbrev_len_for_midx(struct multi_pack_index *m,\n@@ -752,9 +749,7 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n {\n \tconst struct git_hash_algo *algo =\n \t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n-\tstruct disambiguate_state ds;\n \tstruct min_abbrev_data mad;\n-\tstruct object_id oid_ret;\n \tconst unsigned hexsz = algo->hexsz;\n \n \tif (len < 0) {\n@@ -794,16 +789,7 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \tmad.oid = oid;\n \n \tfind_abbrev_len_packed(&mad);\n-\n-\tif (init_object_disambiguation(r, hex, mad.cur_len, algo, &ds) < 0)\n-\t\treturn -1;\n-\n-\tds.fn = repo_extend_abbrev_len;\n-\tds.always_call_fn = 1;\n-\tds.cb_data = (void *)&mad;\n-\n-\tfind_short_object_filename(&ds);\n-\t(void)finish_object_disambiguation(&ds, &oid_ret);\n+\tfind_abbrev_len_loose(&mad);\n \n \thex[mad.cur_len] = 0;\n \treturn mad.cur_len;\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539491","messageId":"20260320-b4-pks-odb-source-abbrev-v2-11-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 11/14] object-name: simplify computing common prefixes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:37Z","receivedAt":"2026-03-20T07:08:13Z","isPatch":true,"body":"The function `extend_abbrev_len()` computes the length of common hex\ncharacters between two object IDs. This is done by:\n\n  - Making the caller provide the `hex` string for the needle object ID.\n\n  - Comparing every hex position of the haystack object ID with\n    `get_hex_char_from_oid()`.\n\nTurning the binary representation into hex first is roundabout though:\nwe can simply compare the binary representation and give some special\nattention to the final nibble.\n\nIntroduce a new function `oid_common_prefix_hexlen()` that does exactly\nthis and refactor the code to use the new function. This allows us to\ndrop the `struct min_abbrev_data::hex` field. Furthermore, this function\nwill be used in by some other callsites in subsequent commits.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n hash.c        | 18 ++++++++++++++++++\n hash.h        |  3 +++\n object-name.c | 23 +++--------------------\n 3 files changed, 24 insertions(+), 20 deletions(-)\n\ndiff --git a/hash.c b/hash.c\nindex 553f2008ea..e925b9754e 100644\n--- a/hash.c\n+++ b/hash.c\n@@ -317,3 +317,21 @@ const struct git_hash_algo *unsafe_hash_algo(const struct git_hash_algo *algop)\n \t/* Otherwise use the default one. */\n \treturn algop;\n }\n+\n+unsigned oid_common_prefix_hexlen(const struct object_id *a,\n+\t\t\t\t  const struct object_id *b)\n+{\n+\tunsigned rawsz = hash_algos[a->algo].rawsz;\n+\n+\tfor (unsigned i = 0; i < rawsz; i++) {\n+\t\tif (a->hash[i] == b->hash[i])\n+\t\t\tcontinue;\n+\n+\t\tif ((a->hash[i] ^ b->hash[i]) & 0xf0)\n+\t\t\treturn i * 2;\n+\t\telse\n+\t\t\treturn i * 2 + 1;\n+\t}\n+\n+\treturn rawsz * 2;\n+}\ndiff --git a/hash.h b/hash.h\nindex d51efce1d3..c082a53c9a 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -396,6 +396,9 @@ static inline int oideq(const struct object_id *oid1, const struct object_id *oi\n \treturn !memcmp(oid1->hash, oid2->hash, GIT_MAX_RAWSZ);\n }\n \n+unsigned oid_common_prefix_hexlen(const struct object_id *a,\n+\t\t\t\t  const struct object_id *b);\n+\n static inline void oidcpy(struct object_id *dst, const struct object_id *src)\n {\n \tmemcpy(dst->hash, src->hash, GIT_MAX_RAWSZ);\ndiff --git a/object-name.c b/object-name.c\nindex d82fb49f39..32e9c23e40 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -585,32 +585,16 @@ static unsigned msb(unsigned long val)\n struct min_abbrev_data {\n \tunsigned int init_len;\n \tunsigned int cur_len;\n-\tchar *hex;\n \tstruct repository *repo;\n \tconst struct object_id *oid;\n };\n \n-static inline char get_hex_char_from_oid(const struct object_id *oid,\n-\t\t\t\t\t unsigned int pos)\n-{\n-\tstatic const char hex[] = \"0123456789abcdef\";\n-\n-\tif ((pos & 1) == 0)\n-\t\treturn hex[oid->hash[pos >> 1] >> 4];\n-\telse\n-\t\treturn hex[oid->hash[pos >> 1] & 0xf];\n-}\n-\n static int extend_abbrev_len(const struct object_id *oid,\n \t\t\t     struct min_abbrev_data *mad)\n {\n-\tunsigned int i = mad->init_len;\n-\twhile (mad->hex[i] && mad->hex[i] == get_hex_char_from_oid(oid, i))\n-\t\ti++;\n-\n-\tif (mad->hex[i] && i >= mad->cur_len)\n-\t\tmad->cur_len = i + 1;\n-\n+\tunsigned len = oid_common_prefix_hexlen(oid, mad->oid);\n+\tif (len != hash_algos[oid->algo].hexsz && len >= mad->cur_len)\n+\t\tmad->cur_len = len + 1;\n \treturn 0;\n }\n \n@@ -785,7 +769,6 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \tmad.repo = r;\n \tmad.init_len = len;\n \tmad.cur_len = len;\n-\tmad.hex = hex;\n \tmad.oid = oid;\n \n \tfind_abbrev_len_packed(&mad);\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539492","messageId":"20260320-b4-pks-odb-source-abbrev-v2-12-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 12/14] object-name: move logic to compute loose abbreviation length","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:38Z","receivedAt":"2026-03-20T07:08:16Z","isPatch":true,"body":"The function `repo_find_unique_abbrev_r()` takes as input an object ID\nas well as a minimum object ID length and returns the minimum required\nprefix to make the object ID unique.\n\nThe logic that computes the abbreviation length for loose objects is\ndeeply tied to the loose object storage format. As such, it would fail\nin case a different object storage format was used.\n\nPrepare for making this logic generic to the backend by moving the logic\ninto a new `odb_source_loose_find_abbrev_len()` function that is part of\n\"object-file.c\".\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-file.c | 38 ++++++++++++++++++++++++++++++++++++++\n object-file.h | 12 ++++++++++++\n object-name.c | 27 ++++-----------------------\n 3 files changed, 54 insertions(+), 23 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 13732f324f..4f77ce0982 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1952,6 +1952,44 @@ int odb_source_loose_count_objects(struct odb_source *source,\n \treturn ret;\n }\n \n+struct find_abbrev_len_data {\n+\tconst struct object_id *oid;\n+\tunsigned len;\n+};\n+\n+static int find_abbrev_len_cb(const struct object_id *oid,\n+\t\t\t      struct object_info *oi UNUSED,\n+\t\t\t      void *cb_data)\n+{\n+\tstruct find_abbrev_len_data *data = cb_data;\n+\tunsigned len = oid_common_prefix_hexlen(oid, data->oid);\n+\tif (len != hash_algos[oid->algo].hexsz && len >= data->len)\n+\t\tdata->len = len + 1;\n+\treturn 0;\n+}\n+\n+int odb_source_loose_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     unsigned min_len,\n+\t\t\t\t     unsigned *out)\n+{\n+\tstruct odb_for_each_object_options opts = {\n+\t\t.prefix = oid,\n+\t\t.prefix_hex_len = min_len,\n+\t};\n+\tstruct find_abbrev_len_data data = {\n+\t\t.oid = oid,\n+\t\t.len = min_len,\n+\t};\n+\tint ret;\n+\n+\tret = odb_source_loose_for_each_object(source, NULL, find_abbrev_len_cb,\n+\t\t\t\t\t       &data, &opts);\n+\t*out = data.len;\n+\n+\treturn ret;\n+}\n+\n static int append_loose_object(const struct object_id *oid,\n \t\t\t       const char *path UNUSED,\n \t\t\t       void *data)\ndiff --git a/object-file.h b/object-file.h\nindex f11ad58f6c..3686f182e4 100644\n--- a/object-file.h\n+++ b/object-file.h\n@@ -146,6 +146,18 @@ int odb_source_loose_count_objects(struct odb_source *source,\n \t\t\t\t   enum odb_count_objects_flags flags,\n \t\t\t\t   unsigned long *out);\n \n+/*\n+ * Find the shortest unique prefix for the given object ID, where `min_len` is\n+ * the minimum length that the prefix should have.\n+ *\n+ * Returns 0 on success, in which case the computed length will be written to\n+ * `out`. Otherwise, a negative error code is returned.\n+ */\n+int odb_source_loose_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     unsigned min_len,\n+\t\t\t\t     unsigned *out);\n+\n /**\n  * format_object_header() is a thin wrapper around s xsnprintf() that\n  * writes the initial \"<type> <obj-len>\" part of the loose object\ndiff --git a/object-name.c b/object-name.c\nindex 32e9c23e40..4e21dbfa97 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -598,28 +598,6 @@ static int extend_abbrev_len(const struct object_id *oid,\n \treturn 0;\n }\n \n-static int extend_abbrev_len_loose(const struct object_id *oid,\n-\t\t\t\t   struct object_info *oi UNUSED,\n-\t\t\t\t   void *cb_data)\n-{\n-\tstruct min_abbrev_data *data = cb_data;\n-\textend_abbrev_len(oid, data);\n-\treturn 0;\n-}\n-\n-static void find_abbrev_len_loose(struct min_abbrev_data *mad)\n-{\n-\tstruct odb_for_each_object_options opts = {\n-\t\t.prefix = mad->oid,\n-\t\t.prefix_hex_len = mad->cur_len,\n-\t};\n-\tstruct odb_source *source;\n-\n-\tfor (source = mad->repo->objects->sources; source; source = source->next)\n-\t\todb_source_loose_for_each_object(source, NULL, extend_abbrev_len_loose,\n-\t\t\t\t\t\t mad, &opts);\n-}\n-\n static void find_abbrev_len_for_midx(struct multi_pack_index *m,\n \t\t\t\t     struct min_abbrev_data *mad)\n {\n@@ -772,7 +750,10 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \tmad.oid = oid;\n \n \tfind_abbrev_len_packed(&mad);\n-\tfind_abbrev_len_loose(&mad);\n+\n+\todb_prepare_alternates(r->objects);\n+\tfor (struct odb_source *s = r->objects->sources; s; s = s->next)\n+\t\todb_source_loose_find_abbrev_len(s, mad.oid, mad.cur_len, &mad.cur_len);\n \n \thex[mad.cur_len] = 0;\n \treturn mad.cur_len;\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539493","messageId":"20260320-b4-pks-odb-source-abbrev-v2-13-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 13/14] object-file: move logic to compute packed abbreviation length","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:39Z","receivedAt":"2026-03-20T07:08:18Z","isPatch":true,"body":"Same as the preceding commit, move the logic that computes the minimum\nrequired prefix length to make a given object ID unique for the packfile\nstore into a new function `packfile_store_find_abbrev_len()` that is\npart of \"packfile.c\". This prepares for making the logic fully generic\nvia pluggable object databases.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c | 135 ++++++----------------------------------------------------\n packfile.c    | 111 +++++++++++++++++++++++++++++++++++++++++++++++\n packfile.h    |   5 +++\n 3 files changed, 128 insertions(+), 123 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex 4e21dbfa97..bb2294a193 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -582,115 +582,6 @@ static unsigned msb(unsigned long val)\n \treturn r;\n }\n \n-struct min_abbrev_data {\n-\tunsigned int init_len;\n-\tunsigned int cur_len;\n-\tstruct repository *repo;\n-\tconst struct object_id *oid;\n-};\n-\n-static int extend_abbrev_len(const struct object_id *oid,\n-\t\t\t     struct min_abbrev_data *mad)\n-{\n-\tunsigned len = oid_common_prefix_hexlen(oid, mad->oid);\n-\tif (len != hash_algos[oid->algo].hexsz && len >= mad->cur_len)\n-\t\tmad->cur_len = len + 1;\n-\treturn 0;\n-}\n-\n-static void find_abbrev_len_for_midx(struct multi_pack_index *m,\n-\t\t\t\t     struct min_abbrev_data *mad)\n-{\n-\tfor (; m; m = m->base_midx) {\n-\t\tint match = 0;\n-\t\tuint32_t num, first = 0;\n-\t\tstruct object_id oid;\n-\t\tconst struct object_id *mad_oid;\n-\n-\t\tif (!m->num_objects)\n-\t\t\tcontinue;\n-\n-\t\tnum = m->num_objects + m->num_objects_in_base;\n-\t\tmad_oid = mad->oid;\n-\t\tmatch = bsearch_one_midx(mad_oid, m, &first);\n-\n-\t\t/*\n-\t\t * first is now the position in the packfile where we\n-\t\t * would insert mad->hash if it does not exist (or the\n-\t\t * position of mad->hash if it does exist). Hence, we\n-\t\t * consider a maximum of two objects nearby for the\n-\t\t * abbreviation length.\n-\t\t */\n-\t\tmad->init_len = 0;\n-\t\tif (!match) {\n-\t\t\tif (nth_midxed_object_oid(&oid, m, first))\n-\t\t\t\textend_abbrev_len(&oid, mad);\n-\t\t} else if (first < num - 1) {\n-\t\t\tif (nth_midxed_object_oid(&oid, m, first + 1))\n-\t\t\t\textend_abbrev_len(&oid, mad);\n-\t\t}\n-\t\tif (first > 0) {\n-\t\t\tif (nth_midxed_object_oid(&oid, m, first - 1))\n-\t\t\t\textend_abbrev_len(&oid, mad);\n-\t\t}\n-\t\tmad->init_len = mad->cur_len;\n-\t}\n-}\n-\n-static void find_abbrev_len_for_pack(struct packed_git *p,\n-\t\t\t\t     struct min_abbrev_data *mad)\n-{\n-\tint match = 0;\n-\tuint32_t num, first = 0;\n-\tstruct object_id oid;\n-\tconst struct object_id *mad_oid;\n-\n-\tif (p->multi_pack_index)\n-\t\treturn;\n-\n-\tif (open_pack_index(p) || !p->num_objects)\n-\t\treturn;\n-\n-\tnum = p->num_objects;\n-\tmad_oid = mad->oid;\n-\tmatch = bsearch_pack(mad_oid, p, &first);\n-\n-\t/*\n-\t * first is now the position in the packfile where we would insert\n-\t * mad->hash if it does not exist (or the position of mad->hash if\n-\t * it does exist). Hence, we consider a maximum of two objects\n-\t * nearby for the abbreviation length.\n-\t */\n-\tmad->init_len = 0;\n-\tif (!match) {\n-\t\tif (!nth_packed_object_id(&oid, p, first))\n-\t\t\textend_abbrev_len(&oid, mad);\n-\t} else if (first < num - 1) {\n-\t\tif (!nth_packed_object_id(&oid, p, first + 1))\n-\t\t\textend_abbrev_len(&oid, mad);\n-\t}\n-\tif (first > 0) {\n-\t\tif (!nth_packed_object_id(&oid, p, first - 1))\n-\t\t\textend_abbrev_len(&oid, mad);\n-\t}\n-\tmad->init_len = mad->cur_len;\n-}\n-\n-static void find_abbrev_len_packed(struct min_abbrev_data *mad)\n-{\n-\tstruct packed_git *p;\n-\n-\todb_prepare_alternates(mad->repo->objects);\n-\tfor (struct odb_source *source = mad->repo->objects->sources; source; source = source->next) {\n-\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n-\t\tif (m)\n-\t\t\tfind_abbrev_len_for_midx(m, mad);\n-\t}\n-\n-\trepo_for_each_pack(mad->repo, p)\n-\t\tfind_abbrev_len_for_pack(p, mad);\n-}\n-\n void strbuf_repo_add_unique_abbrev(struct strbuf *sb, struct repository *repo,\n \t\t\t\t   const struct object_id *oid, int abbrev_len)\n {\n@@ -707,14 +598,14 @@ void strbuf_add_unique_abbrev(struct strbuf *sb, const struct object_id *oid,\n }\n \n int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n-\t\t\t      const struct object_id *oid, int len)\n+\t\t\t      const struct object_id *oid, int min_len)\n {\n \tconst struct git_hash_algo *algo =\n \t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n-\tstruct min_abbrev_data mad;\n \tconst unsigned hexsz = algo->hexsz;\n+\tunsigned len;\n \n-\tif (len < 0) {\n+\tif (min_len < 0) {\n \t\tunsigned long count;\n \n \t\tif (odb_count_objects(r->objects, ODB_COUNT_OBJECTS_APPROXIMATE, &count) < 0)\n@@ -738,25 +629,23 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \t\t */\n \t\tif (len < FALLBACK_DEFAULT_ABBREV)\n \t\t\tlen = FALLBACK_DEFAULT_ABBREV;\n+\t} else {\n+\t\tlen = min_len;\n \t}\n \n \toid_to_hex_r(hex, oid);\n \tif (len >= hexsz || !len)\n \t\treturn hexsz;\n \n-\tmad.repo = r;\n-\tmad.init_len = len;\n-\tmad.cur_len = len;\n-\tmad.oid = oid;\n-\n-\tfind_abbrev_len_packed(&mad);\n-\n \todb_prepare_alternates(r->objects);\n-\tfor (struct odb_source *s = r->objects->sources; s; s = s->next)\n-\t\todb_source_loose_find_abbrev_len(s, mad.oid, mad.cur_len, &mad.cur_len);\n+\tfor (struct odb_source *s = r->objects->sources; s; s = s->next) {\n+\t\tstruct odb_source_files *files = odb_source_files_downcast(s);\n+\t\tpackfile_store_find_abbrev_len(files->packed, oid, len, &len);\n+\t\todb_source_loose_find_abbrev_len(s, oid, len, &len);\n+\t}\n \n-\thex[mad.cur_len] = 0;\n-\treturn mad.cur_len;\n+\thex[len] = 0;\n+\treturn len;\n }\n \n const char *repo_find_unique_abbrev(struct repository *r,\ndiff --git a/packfile.c b/packfile.c\nindex 2539a371c1..ee9c7ea1d1 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2597,6 +2597,117 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \treturn ret;\n }\n \n+static int extend_abbrev_len(const struct object_id *a,\n+\t\t\t     const struct object_id *b,\n+\t\t\t     unsigned *out)\n+{\n+\tunsigned len = oid_common_prefix_hexlen(a, b);\n+\tif (len != hash_algos[a->algo].hexsz && len >= *out)\n+\t\t*out = len + 1;\n+\treturn 0;\n+}\n+\n+static void find_abbrev_len_for_midx(struct multi_pack_index *m,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     unsigned min_len,\n+\t\t\t\t     unsigned *out)\n+{\n+\tunsigned len = min_len;\n+\n+\tfor (; m; m = m->base_midx) {\n+\t\tint match = 0;\n+\t\tuint32_t num, first = 0;\n+\t\tstruct object_id found_oid;\n+\n+\t\tif (!m->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tnum = m->num_objects + m->num_objects_in_base;\n+\t\tmatch = bsearch_one_midx(oid, m, &first);\n+\n+\t\t/*\n+\t\t * first is now the position in the packfile where we\n+\t\t * would insert the object ID if it does not exist (or the\n+\t\t * position of the object ID if it does exist). Hence, we\n+\t\t * consider a maximum of two objects nearby for the\n+\t\t * abbreviation length.\n+\t\t */\n+\n+\t\tif (!match) {\n+\t\t\tif (nth_midxed_object_oid(&found_oid, m, first))\n+\t\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t\t} else if (first < num - 1) {\n+\t\t\tif (nth_midxed_object_oid(&found_oid, m, first + 1))\n+\t\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t\t}\n+\t\tif (first > 0) {\n+\t\t\tif (nth_midxed_object_oid(&found_oid, m, first - 1))\n+\t\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t\t}\n+\t}\n+\n+\t*out = len;\n+}\n+\n+static void find_abbrev_len_for_pack(struct packed_git *p,\n+\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t     unsigned min_len,\n+\t\t\t\t     unsigned *out)\n+{\n+\tint match;\n+\tuint32_t num, first = 0;\n+\tstruct object_id found_oid;\n+\tunsigned len = min_len;\n+\n+\tnum = p->num_objects;\n+\tmatch = bsearch_pack(oid, p, &first);\n+\n+\t/*\n+\t * first is now the position in the packfile where we would insert\n+\t * the object ID if it does not exist (or the position of mad->hash if\n+\t * it does exist). Hence, we consider a maximum of two objects\n+\t * nearby for the abbreviation length.\n+\t */\n+\tif (!match) {\n+\t\tif (!nth_packed_object_id(&found_oid, p, first))\n+\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t} else if (first < num - 1) {\n+\t\tif (!nth_packed_object_id(&found_oid, p, first + 1))\n+\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t}\n+\tif (first > 0) {\n+\t\tif (!nth_packed_object_id(&found_oid, p, first - 1))\n+\t\t\textend_abbrev_len(&found_oid, oid, &len);\n+\t}\n+\n+\t*out = len;\n+}\n+\n+int packfile_store_find_abbrev_len(struct packfile_store *store,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   unsigned min_len,\n+\t\t\t\t   unsigned *out)\n+{\n+\tstruct packfile_list_entry *e;\n+\tstruct multi_pack_index *m;\n+\n+\tm = get_multi_pack_index(store->source);\n+\tif (m)\n+\t\tfind_abbrev_len_for_midx(m, oid, min_len, &min_len);\n+\n+\tfor (e = packfile_store_get_packs(store); e; e = e->next) {\n+\t\tif (e->pack->multi_pack_index)\n+\t\t\tcontinue;\n+\t\tif (open_pack_index(e->pack) || !e->pack->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tfind_abbrev_len_for_pack(e->pack, oid, min_len, &min_len);\n+\t}\n+\n+\t*out = min_len;\n+\treturn 0;\n+}\n+\n struct add_promisor_object_data {\n \tstruct repository *repo;\n \tstruct oidset *set;\ndiff --git a/packfile.h b/packfile.h\nindex fa41dfda38..45b35973f0 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -369,6 +369,11 @@ int packfile_store_for_each_object(struct packfile_store *store,\n \t\t\t\t   void *cb_data,\n \t\t\t\t   const struct odb_for_each_object_options *opts);\n \n+int packfile_store_find_abbrev_len(struct packfile_store *store,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   unsigned min_len,\n+\t\t\t\t   unsigned *out);\n+\n /* A hook to report invalid files in pack directory */\n #define PACKDIR_FILE_PACK 1\n #define PACKDIR_FILE_IDX 2\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539494","messageId":"20260320-b4-pks-odb-source-abbrev-v2-14-fe65dcd8c735@pks.im","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"[PATCH v2 14/14] odb: introduce generic `odb_find_abbrev_len()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T07:07:40Z","receivedAt":"2026-03-20T07:08:21Z","isPatch":true,"body":"Introduce a new generic `odb_find_abbrev_len()` function as well as\nsource-specific callback functions. This makes the logic to compute the\nrequired prefix length to make a given object unique fully pluggable.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n object-name.c      | 57 +++---------------------------------------\n odb.c              | 73 ++++++++++++++++++++++++++++++++++++++++++++++++++++++\n odb.h              | 16 ++++++++++++\n odb/source-files.c | 25 +++++++++++++++++++\n odb/source.h       | 24 ++++++++++++++++++\n 5 files changed, 142 insertions(+), 53 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex bb2294a193..f6e1f29e1f 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -15,10 +15,9 @@\n #include \"refs.h\"\n #include \"remote.h\"\n #include \"dir.h\"\n+#include \"odb.h\"\n #include \"oid-array.h\"\n-#include \"packfile.h\"\n #include \"pretty.h\"\n-#include \"object-file.h\"\n #include \"read-cache-ll.h\"\n #include \"repo-settings.h\"\n #include \"repository.h\"\n@@ -569,19 +568,6 @@ int repo_for_each_abbrev(struct repository *r, const char *prefix,\n \treturn ret;\n }\n \n-/*\n- * Return the slot of the most-significant bit set in \"val\". There are various\n- * ways to do this quickly with fls() or __builtin_clzl(), but speed is\n- * probably not a big deal here.\n- */\n-static unsigned msb(unsigned long val)\n-{\n-\tunsigned r = 0;\n-\twhile (val >>= 1)\n-\t\tr++;\n-\treturn r;\n-}\n-\n void strbuf_repo_add_unique_abbrev(struct strbuf *sb, struct repository *repo,\n \t\t\t\t   const struct object_id *oid, int abbrev_len)\n {\n@@ -602,49 +588,14 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n {\n \tconst struct git_hash_algo *algo =\n \t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n-\tconst unsigned hexsz = algo->hexsz;\n \tunsigned len;\n \n-\tif (min_len < 0) {\n-\t\tunsigned long count;\n-\n-\t\tif (odb_count_objects(r->objects, ODB_COUNT_OBJECTS_APPROXIMATE, &count) < 0)\n-\t\t\tcount = 0;\n-\n-\t\t/*\n-\t\t * Add one because the MSB only tells us the highest bit set,\n-\t\t * not including the value of all the _other_ bits (so \"15\"\n-\t\t * is only one off of 2^4, but the MSB is the 3rd bit.\n-\t\t */\n-\t\tlen = msb(count) + 1;\n-\t\t/*\n-\t\t * We now know we have on the order of 2^len objects, which\n-\t\t * expects a collision at 2^(len/2). But we also care about hex\n-\t\t * chars, not bits, and there are 4 bits per hex. So all\n-\t\t * together we need to divide by 2 and round up.\n-\t\t */\n-\t\tlen = DIV_ROUND_UP(len, 2);\n-\t\t/*\n-\t\t * For very small repos, we stick with our regular fallback.\n-\t\t */\n-\t\tif (len < FALLBACK_DEFAULT_ABBREV)\n-\t\t\tlen = FALLBACK_DEFAULT_ABBREV;\n-\t} else {\n-\t\tlen = min_len;\n-\t}\n+\tif (odb_find_abbrev_len(r->objects, oid, min_len, &len) < 0)\n+\t\tlen = algo->hexsz;\n \n \toid_to_hex_r(hex, oid);\n-\tif (len >= hexsz || !len)\n-\t\treturn hexsz;\n-\n-\todb_prepare_alternates(r->objects);\n-\tfor (struct odb_source *s = r->objects->sources; s; s = s->next) {\n-\t\tstruct odb_source_files *files = odb_source_files_downcast(s);\n-\t\tpackfile_store_find_abbrev_len(files->packed, oid, len, &len);\n-\t\todb_source_loose_find_abbrev_len(s, oid, len, &len);\n-\t}\n-\n \thex[len] = 0;\n+\n \treturn len;\n }\n \ndiff --git a/odb.c b/odb.c\nindex 3019957b87..3f94a53df1 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -12,6 +12,7 @@\n #include \"midx.h\"\n #include \"object-file-convert.h\"\n #include \"object-file.h\"\n+#include \"object-name.h\"\n #include \"odb.h\"\n #include \"packfile.h\"\n #include \"path.h\"\n@@ -964,6 +965,78 @@ int odb_count_objects(struct object_database *odb,\n \treturn ret;\n }\n \n+/*\n+ * Return the slot of the most-significant bit set in \"val\". There are various\n+ * ways to do this quickly with fls() or __builtin_clzl(), but speed is\n+ * probably not a big deal here.\n+ */\n+static unsigned msb(unsigned long val)\n+{\n+\tunsigned r = 0;\n+\twhile (val >>= 1)\n+\t\tr++;\n+\treturn r;\n+}\n+\n+int odb_find_abbrev_len(struct object_database *odb,\n+\t\t\tconst struct object_id *oid,\n+\t\t\tint min_length,\n+\t\t\tunsigned *out)\n+{\n+\tconst struct git_hash_algo *algo =\n+\t\toid->algo ? &hash_algos[oid->algo] : odb->repo->hash_algo;\n+\tconst unsigned hexsz = algo->hexsz;\n+\tunsigned len;\n+\tint ret;\n+\n+\tif (min_length < 0) {\n+\t\tunsigned long count;\n+\n+\t\tif (odb_count_objects(odb, ODB_COUNT_OBJECTS_APPROXIMATE, &count) < 0)\n+\t\t\tcount = 0;\n+\n+\t\t/*\n+\t\t * Add one because the MSB only tells us the highest bit set,\n+\t\t * not including the value of all the _other_ bits (so \"15\"\n+\t\t * is only one off of 2^4, but the MSB is the 3rd bit.\n+\t\t */\n+\t\tlen = msb(count) + 1;\n+\t\t/*\n+\t\t * We now know we have on the order of 2^len objects, which\n+\t\t * expects a collision at 2^(len/2). But we also care about hex\n+\t\t * chars, not bits, and there are 4 bits per hex. So all\n+\t\t * together we need to divide by 2 and round up.\n+\t\t */\n+\t\tlen = DIV_ROUND_UP(len, 2);\n+\t\t/*\n+\t\t * For very small repos, we stick with our regular fallback.\n+\t\t */\n+\t\tif (len < FALLBACK_DEFAULT_ABBREV)\n+\t\t\tlen = FALLBACK_DEFAULT_ABBREV;\n+\t} else {\n+\t\tlen = min_length;\n+\t}\n+\n+\tif (len >= hexsz || !len) {\n+\t\t*out = hexsz;\n+\t\tret = 0;\n+\t\tgoto out;\n+\t}\n+\n+\todb_prepare_alternates(odb);\n+\tfor (struct odb_source *source = odb->sources; source; source = source->next) {\n+\t\tret = odb_source_find_abbrev_len(source, oid, len, &len);\n+\t\tif (ret)\n+\t\t\tgoto out;\n+\t}\n+\n+\tret = 0;\n+\t*out = len;\n+\n+out:\n+\treturn ret;\n+}\n+\n void odb_assert_oid_type(struct object_database *odb,\n \t\t\t const struct object_id *oid, enum object_type expect)\n {\ndiff --git a/odb.h b/odb.h\nindex e80fd8f7ab..984bafca9d 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -545,6 +545,22 @@ int odb_count_objects(struct object_database *odb,\n \t\t      enum odb_count_objects_flags flags,\n \t\t      unsigned long *out);\n \n+/*\n+ * Given an object ID, find the minimum required length required to make the\n+ * object ID unique across the whole object database.\n+ *\n+ * The `min_len` determines the minimum abbreviated length that'll be returned\n+ * by this function. If `min_len < 0`, then the function will set a sensible\n+ * default minimum abbreviation length.\n+ *\n+ * Returns 0 on success, a negative error code otherwise. The computed length\n+ * will be assigned to `*out`.\n+ */\n+int odb_find_abbrev_len(struct object_database *odb,\n+\t\t\tconst struct object_id *oid,\n+\t\t\tint min_len,\n+\t\t\tunsigned *out);\n+\n enum {\n \t/*\n \t * By default, `odb_write_object()` does not actually write anything\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex e90bb689bb..76797569de 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -122,6 +122,30 @@ static int odb_source_files_count_objects(struct odb_source *source,\n \treturn ret;\n }\n \n+static int odb_source_files_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t\t    const struct object_id *oid,\n+\t\t\t\t\t    unsigned min_len,\n+\t\t\t\t\t    unsigned *out)\n+{\n+\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\tunsigned len = min_len;\n+\tint ret;\n+\n+\tret = packfile_store_find_abbrev_len(files->packed, oid, len, &len);\n+\tif (ret < 0)\n+\t\tgoto out;\n+\n+\tret = odb_source_loose_find_abbrev_len(source, oid, len, &len);\n+\tif (ret < 0)\n+\t\tgoto out;\n+\n+\t*out = len;\n+\tret = 0;\n+\n+out:\n+\treturn ret;\n+}\n+\n static int odb_source_files_freshen_object(struct odb_source *source,\n \t\t\t\t\t   const struct object_id *oid)\n {\n@@ -250,6 +274,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\n \tfiles->base.for_each_object = odb_source_files_for_each_object;\n \tfiles->base.count_objects = odb_source_files_count_objects;\n+\tfiles->base.find_abbrev_len = odb_source_files_find_abbrev_len;\n \tfiles->base.freshen_object = odb_source_files_freshen_object;\n \tfiles->base.write_object = odb_source_files_write_object;\n \tfiles->base.write_object_stream = odb_source_files_write_object_stream;\ndiff --git a/odb/source.h b/odb/source.h\nindex ee5d6ed530..a9d7d0b96f 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -157,6 +157,18 @@ struct odb_source {\n \t\t\t     enum odb_count_objects_flags flags,\n \t\t\t     unsigned long *out);\n \n+\t/*\n+\t * This callback is expected to find the minimum required length to\n+\t * make the given object ID unique.\n+\t *\n+\t * The callback is expected to return a negative error code in case it\n+\t * failed, 0 otherwise.\n+\t */\n+\tint (*find_abbrev_len)(struct odb_source *source,\n+\t\t\t       const struct object_id *oid,\n+\t\t\t       unsigned min_length,\n+\t\t\t       unsigned *out);\n+\n \t/*\n \t * This callback is expected to freshen the given object so that its\n \t * last access time is set to the current time. This is used to ensure\n@@ -360,6 +372,18 @@ static inline int odb_source_count_objects(struct odb_source *source,\n \treturn source->count_objects(source, flags, out);\n }\n \n+/*\n+ * Determine the minimum required length to make the given object ID unique in\n+ * the given source. Returns 0 on success, a negative error code otherwise.\n+ */\n+static inline int odb_source_find_abbrev_len(struct odb_source *source,\n+\t\t\t\t\t     const struct object_id *oid,\n+\t\t\t\t\t     unsigned min_len,\n+\t\t\t\t\t     unsigned *out)\n+{\n+\treturn source->find_abbrev_len(source, oid, min_len, out);\n+}\n+\n /*\n  * Freshen an object in the object database by updating its timestamp.\n  * Returns 1 in case the object has been freshened, 0 in case the object does\n\n-- \n2.53.0.1055.ga2ffed1127.dirty\n\n"},{"id":"539501","messageId":"CAOLa=ZRpCunxE_F1AG-aFqHiyVZ=c+T_wFwpfxC7vFm7dKxqAw@mail.gmail.com","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-3-5ddebad292b0@pks.im","subject":"Re: [PATCH 03/14] odb: introduce `struct odb_for_each_object_options`","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-03-20T09:01:42Z","receivedAt":"2026-03-20T09:01:44Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> The `odb_for_each_object()` function only accepts a bitset of flags. In\n> a subsequent commit we'll want to change object iteration to also\n> support iterating over only those objects that have a specific prefix.\n> While we could of course add the prefix to the function signature, or\n> alternative introduce a new function, both of these options don't really\n> seem to be that sensible.\n>\n> Instead, introduce a new `struct odb_for_each_object_options` that can\n> be passed to a new `odb_for_each_object_ext()` function. Splice through\n> the options structure into the respective object database sources.\n>\n\nYeah I like this pattern, really cleans up the arguments sent into a\nfunction. Making future additions to the struct produce a localized diff\nrather than modifying the function params each time. Nice.\n\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  builtin/cat-file.c     |  7 +++++--\n>  builtin/pack-objects.c | 12 +++++++-----\n>  commit-graph.c         |  5 ++++-\n>  object-file.c          |  6 +++---\n>  object-file.h          |  2 +-\n>  odb.c                  | 26 +++++++++++++++++++-------\n>  odb.h                  | 16 ++++++++++++++++\n>  odb/source-files.c     |  8 ++++----\n>  odb/source.h           |  6 +++---\n>  packfile.c             | 12 ++++++------\n>  packfile.h             |  2 +-\n>  11 files changed, 69 insertions(+), 33 deletions(-)\n>\n> diff --git a/builtin/cat-file.c b/builtin/cat-file.c\n> index b6f12f41d6..cd13a3a89f 100644\n> --- a/builtin/cat-file.c\n> +++ b/builtin/cat-file.c\n> @@ -848,6 +848,9 @@ static void batch_each_object(struct batch_options *opt,\n>  \t\t.callback = callback,\n>  \t\t.payload = _payload,\n>  \t};\n> +\tstruct odb_for_each_object_options opts = {\n> +\t\t.flags = flags,\n> +\t};\n>  \tstruct bitmap_index *bitmap = NULL;\n>  \tstruct odb_source *source;\n>\n> @@ -860,7 +863,7 @@ static void batch_each_object(struct batch_options *opt,\n>  \todb_prepare_alternates(the_repository->objects);\n>  \tfor (source = the_repository->objects->sources; source; source = source->next) {\n>  \t\tint ret = odb_source_loose_for_each_object(source, NULL, batch_one_object_oi,\n> -\t\t\t\t\t\t\t   &payload, flags);\n> +\t\t\t\t\t\t\t   &payload, &opts);\n>  \t\tif (ret)\n>  \t\t\tbreak;\n>  \t}\n> @@ -884,7 +887,7 @@ static void batch_each_object(struct batch_options *opt,\n>  \t\tfor (source = the_repository->objects->sources; source; source = source->next) {\n>  \t\t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n>  \t\t\tint ret = packfile_store_for_each_object(files->packed, &oi,\n> -\t\t\t\t\t\t\t\t batch_one_object_oi, &payload, flags);\n> +\t\t\t\t\t\t\t\t batch_one_object_oi, &payload, &opts);\n>  \t\t\tif (ret)\n>  \t\t\t\tbreak;\n>  \t\t}\n> diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> index cd013c0b68..3bb57ff183 100644\n> --- a/builtin/pack-objects.c\n> +++ b/builtin/pack-objects.c\n> @@ -4344,6 +4344,12 @@ static void add_objects_in_unpacked_packs(void)\n>  {\n>  \tstruct odb_source *source;\n>  \ttime_t mtime;\n> +\tstruct odb_for_each_object_options opts = {\n> +\t\t.flags = ODB_FOR_EACH_OBJECT_PACK_ORDER |\n> +\t\t\t ODB_FOR_EACH_OBJECT_LOCAL_ONLY |\n> +\t\t\t ODB_FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS |\n> +\t\t\t ODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS,\n> +\t};\n>  \tstruct object_info oi = {\n>  \t\t.mtimep = &mtime,\n>  \t};\n> @@ -4356,11 +4362,7 @@ static void add_objects_in_unpacked_packs(void)\n>  \t\t\tcontinue;\n>\n>  \t\tif (packfile_store_for_each_object(files->packed, &oi,\n> -\t\t\t\t\t\t   add_object_in_unpacked_pack, NULL,\n> -\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_PACK_ORDER |\n> -\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_LOCAL_ONLY |\n> -\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_SKIP_IN_CORE_KEPT_PACKS |\n> -\t\t\t\t\t\t   ODB_FOR_EACH_OBJECT_SKIP_ON_DISK_KEPT_PACKS))\n> +\t\t\t\t\t\t   add_object_in_unpacked_pack, NULL, &opts))\n\nPlus this is so much easier to read now.\n\n[snip]\n\nThe rest looked good.\n"},{"id":"539502","messageId":"CAOLa=ZSC=uTPg516dGSEnRcWgr4J9iDBQ+8D=o0F+jT1jkK-sA@mail.gmail.com","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-7-5ddebad292b0@pks.im","subject":"Re: [PATCH 07/14] object-name: backend-generic `repo_collect_ambiguous()`","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-03-20T09:23:13Z","receivedAt":"2026-03-20T09:23:15Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> The function `repo_collect_ambiguous()` is responsible for collecting\n> objects whose IDs match a specific prefix. The information is then\n> used to inform the user about which objects they could have meant in\n> case a short object ID is ambiguous.\n>\n> The logic to do this uses the object disambiguation infrastructure and\n> calls into backend-specific functions to iterate through loose and\n> packed objects. This isn't really required anymore though: all we want\n> to do is to enumerate objects that have such a prefix and then append\n> those objects to a `struct oid_array`. This can be trivially achieved\n> in a generic way now that `odb_for_each_object()` has learned to yield\n> only objects that much such a prefix.\n>\n> Refactor the code to use the backend-generic infrastructure instead.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  object-name.c | 19 ++++++++++---------\n>  1 file changed, 10 insertions(+), 9 deletions(-)\n>\n> diff --git a/object-name.c b/object-name.c\n> index fd1b010ab3..4c3ace150e 100644\n> --- a/object-name.c\n> +++ b/object-name.c\n> @@ -448,8 +448,8 @@ static int collect_ambiguous(const struct object_id *oid, void *data)\n>  \treturn 0;\n>  }\n>\n> -static int repo_collect_ambiguous(struct repository *r UNUSED,\n> -\t\t\t\t  const struct object_id *oid,\n> +static int repo_collect_ambiguous(const struct object_id *oid,\n> +\t\t\t\t  struct object_info *oi UNUSED,\n>  \t\t\t\t  void *data)\n\nSo these modifications are so that it matches the callback function\ntype.\n\n>  {\n>  \treturn collect_ambiguous(oid, data);\n> @@ -586,18 +586,19 @@ int repo_for_each_abbrev(struct repository *r, const char *prefix,\n>  \t\t\t const struct git_hash_algo *algo,\n>  \t\t\t each_abbrev_fn fn, void *cb_data)\n>  {\n> +\tstruct object_id prefix_oid = { 0 };\n> +\tstruct odb_for_each_object_options opts = {\n> +\t\t.prefix = &prefix_oid,\n> +\t\t.prefix_hex_len = strlen(prefix),\n> +\t};\n>  \tstruct oid_array collect = OID_ARRAY_INIT;\n> -\tstruct disambiguate_state ds;\n>  \tint ret;\n>\n> -\tif (init_object_disambiguation(r, prefix, strlen(prefix), algo, &ds) < 0)\n> +\tif (parse_oid_prefix(prefix, opts.prefix_hex_len, algo, NULL, &prefix_oid) < 0)\n>  \t\treturn -1;\n>\n> -\tds.always_call_fn = 1;\n> -\tds.fn = repo_collect_ambiguous;\n> -\tds.cb_data = &collect;\n> -\tfind_short_object_filename(&ds);\n> -\tfind_short_packed_object(&ds);\n> +\tif (odb_for_each_object_ext(r->objects, NULL, repo_collect_ambiguous, &collect, &opts) < 0)\n> +\t\treturn -1;\n>\n\nAnd finally we simply call the generic `odb_for_each_object_ext()` with\nthe prefix options set.\n\n>  \tret = oid_array_for_each_unique(&collect, fn, cb_data);\n>  \toid_array_clear(&collect);\n>\n> --\n> 2.53.0.1055.ga2ffed1127.dirty\n"},{"id":"539503","messageId":"CAOLa=ZRU3=FqDo8SiJ=+qTsU79NEfoyAVp1uZYBX57SNPTZomw@mail.gmail.com","threadId":"65299","inReplyTo":"20260319-b4-pks-odb-source-abbrev-v1-11-5ddebad292b0@pks.im","subject":"Re: [PATCH 11/14] object-name: simplify computing common prefixes","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-03-20T10:01:48Z","receivedAt":"2026-03-20T10:01:51Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> The function `extend_abbrev_len()` computes the length of common hex\n> characters between two object IDs. This is done by:\n>\n>   - Making the caller provide the `hex` string for the needle object ID.\n>\n>   - Comparing every hex position of the haystack object ID with\n>     `get_hex_char_from_oid()`.\n>\n> Turning the binary representation into hex first is roundabout though:\n> we can simply compare the binary representation and give some special\n> attention to the final nibble.\n>\n> Introduce a new function `oid_common_prefix_hexlen()` that does exactly\n> this and refactor the code to use the new function. This allows us to\n> drop the `struct min_abbrev_data::hex` field. Furthermore, this function\n> will be used in by some other callsites in subsequent commits.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  hash.c        | 18 ++++++++++++++++++\n>  hash.h        |  3 +++\n>  object-name.c | 23 +++--------------------\n>  3 files changed, 24 insertions(+), 20 deletions(-)\n>\n> diff --git a/hash.c b/hash.c\n> index 553f2008ea..e925b9754e 100644\n> --- a/hash.c\n> +++ b/hash.c\n> @@ -317,3 +317,21 @@ const struct git_hash_algo *unsafe_hash_algo(const struct git_hash_algo *algop)\n>  \t/* Otherwise use the default one. */\n>  \treturn algop;\n>  }\n> +\n> +unsigned oid_common_prefix_hexlen(const struct object_id *a,\n> +\t\t\t\t  const struct object_id *b)\n> +{\n> +\tunsigned rawsz = hash_algos[a->algo].rawsz;\n> +\n> +\tfor (unsigned i = 0; i < rawsz; i++) {\n> +\t\tif (a->hash[i] == b->hash[i])\n> +\t\t\tcontinue;\n> +\n\nInstead of transforming the bytes into 2 hex components we now compare\nthe bytes themselves and perhaps then compare parts of it?\n\n> +\t\tif ((a->hash[i] ^ b->hash[i]) & 0xf0)\n\nOkay so if the 4 MSB are the same then we end up here and return i * 2.\nMakes sense.\n\n> +\t\t\treturn i * 2;\n> +\t\telse\n> +\t\t\treturn i * 2 + 1;\n\nIf not, its the 4 LSB.\n\n> +\t}\n> +\n> +\treturn rawsz * 2;\n> +}\n> diff --git a/hash.h b/hash.h\n> index d51efce1d3..c082a53c9a 100644\n> --- a/hash.h\n> +++ b/hash.h\n> @@ -396,6 +396,9 @@ static inline int oideq(const struct object_id *oid1, const struct object_id *oi\n>  \treturn !memcmp(oid1->hash, oid2->hash, GIT_MAX_RAWSZ);\n>  }\n>\n> +unsigned oid_common_prefix_hexlen(const struct object_id *a,\n> +\t\t\t\t  const struct object_id *b);\n> +\n>  static inline void oidcpy(struct object_id *dst, const struct object_id *src)\n>  {\n>  \tmemcpy(dst->hash, src->hash, GIT_MAX_RAWSZ);\n> diff --git a/object-name.c b/object-name.c\n> index d82fb49f39..32e9c23e40 100644\n> --- a/object-name.c\n> +++ b/object-name.c\n> @@ -585,32 +585,16 @@ static unsigned msb(unsigned long val)\n>  struct min_abbrev_data {\n>  \tunsigned int init_len;\n>  \tunsigned int cur_len;\n> -\tchar *hex;\n>  \tstruct repository *repo;\n>  \tconst struct object_id *oid;\n>  };\n>\n> -static inline char get_hex_char_from_oid(const struct object_id *oid,\n> -\t\t\t\t\t unsigned int pos)\n> -{\n> -\tstatic const char hex[] = \"0123456789abcdef\";\n> -\n> -\tif ((pos & 1) == 0)\n\nSo this basically alternates between odd/even positions.\n\nSo walking with an example:\n\nif we have '10101011 11111010 10101010'\npos 1 should get the hex for '1010'\npos 2 should get the hex for '1011'\npos 3 should get the hex for '1111'\n...\n\n> -\t\treturn hex[oid->hash[pos >> 1] >> 4];\n\nSo for pos 1 '1010', to obtain the byte we first do 'pos >> 1'. Then we\nonly care about the 4 MSB so we do `oid->hash[pos >> 1] >> 4`.\n\nFinally we map it to the hex[] char array.\n\n> -\telse\n> -\t\treturn hex[oid->hash[pos >> 1] & 0xf];\n> -}\n> -\n>  static int extend_abbrev_len(const struct object_id *oid,\n>  \t\t\t     struct min_abbrev_data *mad)\n>  {\n> -\tunsigned int i = mad->init_len;\n> -\twhile (mad->hex[i] && mad->hex[i] == get_hex_char_from_oid(oid, i))\n> -\t\ti++;\n> -\n\nSo earlier, we were iterating through the mad->hex which already had the\ncomputed hex. So to compare it with oid, we needed to get the hex at\neach position i of the oid.\n\nThat's where `get_hex_char_from_oid()` came in place.\n\n> -\tif (mad->hex[i] && i >= mad->cur_len)\n> -\t\tmad->cur_len = i + 1;\n> -\n> +\tunsigned len = oid_common_prefix_hexlen(oid, mad->oid);\n> +\tif (len != hash_algos[oid->algo].hexsz && len >= mad->cur_len)\n> +\t\tmad->cur_len = len + 1;\n>  \treturn 0;\n>  }\n>\n> @@ -785,7 +769,6 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n>  \tmad.repo = r;\n>  \tmad.init_len = len;\n>  \tmad.cur_len = len;\n> -\tmad.hex = hex;\n>  \tmad.oid = oid;\n>\n>  \tfind_abbrev_len_packed(&mad);\n>\n> --\n> 2.53.0.1055.ga2ffed1127.dirty\n\nThe patch looks good.\n"},{"id":"539504","messageId":"CAOLa=ZSeMS2iKzgMUWix_Sx+e24863PsOazRLrqHtS5hYSUk3A@mail.gmail.com","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-0-fe65dcd8c735@pks.im","subject":"Re: [PATCH v2 00/14] odb: generic object name handling","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-03-20T10:04:02Z","receivedAt":"2026-03-20T10:04:04Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Hi,\n>\n> this patch series refactors handling of object names to become pluggable\n> and thus generic. This includes:\n>\n>   - Disambiguation of object names with a common prefix. This is\n>     required to list candidate objects in case the user has passed a\n>     non-unique prefix.\n>\n>   - Abbreviating an object ID to the shortest prefix required while\n>     staying unique.\n>\n> The logic to compute these operations is specific to the backend, but\n> not generic. This patch series fixes that by moving the functionality\n> into the respective backends.\n>\n> This patch series may feel somewhat unexiting, but it's not. Especially\n> abbreviating object IDs is done in lots of places, so this functionality\n> is overall quite critical. So starting with this series, it is now\n> possible to do all kinds of local work with an alternative backend:\n> git-commit(1), git-log(1), git-rev-parse(1), git-merge(1) and many other\n> commands now work as expected. My MongoDB proof of concept [1] only\n> requires two commits (the object format extension) on top. And no, I\n> don't endorse MongoDB or propose it as a future potential backend. It\n> simply had a good C API that was easy to use.\n>\n> Of course, other functionality, especially everything that involves\n> packfiles, doesn't yet work.\n>\n> This patch series is built on top of ca1db8a0f7 (The 17th batch,\n> 2026-03-16) with ps/object-counting at 6801ffd37d (odb: introduce\n> generic object counting, 2026-03-12) merged into it.\n>\n> Changes in v2:\n>   - Document `cb_iter` callback.\n>   - Fix left-over conversion of `odb_source_loose_for_each_object()`.\n>   - commit message typo fixes.\n>   - Link to v1: https://lore.kernel.org/r/20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im\n\nI only got around to reviewing v1 now, but the range-diff here looks\ngood.\n\n- Karthik\n\n[snip]\n"},{"id":"539508","messageId":"ab0htJ0fafNdsXeR@pks.im","threadId":"65299","inReplyTo":"CAOLa=ZRU3=FqDo8SiJ=+qTsU79NEfoyAVp1uZYBX57SNPTZomw@mail.gmail.com","subject":"Re: [PATCH 11/14] object-name: simplify computing common prefixes","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T10:30:12Z","receivedAt":"2026-03-20T10:30:18Z","isPatch":true,"body":"On Fri, Mar 20, 2026 at 03:01:48AM -0700, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > diff --git a/hash.c b/hash.c\n> > index 553f2008ea..e925b9754e 100644\n> > --- a/hash.c\n> > +++ b/hash.c\n> > @@ -317,3 +317,21 @@ const struct git_hash_algo *unsafe_hash_algo(const struct git_hash_algo *algop)\n> >  \t/* Otherwise use the default one. */\n> >  \treturn algop;\n> >  }\n> > +\n> > +unsigned oid_common_prefix_hexlen(const struct object_id *a,\n> > +\t\t\t\t  const struct object_id *b)\n> > +{\n> > +\tunsigned rawsz = hash_algos[a->algo].rawsz;\n> > +\n> > +\tfor (unsigned i = 0; i < rawsz; i++) {\n> > +\t\tif (a->hash[i] == b->hash[i])\n> > +\t\t\tcontinue;\n> > +\n> \n> Instead of transforming the bytes into 2 hex components we now compare\n> the bytes themselves and perhaps then compare parts of it?\n\nYes, exactly. It should be more performant overall compared to first\nconverting to their respective hex presentations, even though I doubt it\nreally matters in practice.\n\n> > +\t\tif ((a->hash[i] ^ b->hash[i]) & 0xf0)\n> \n> Okay so if the 4 MSB are the same then we end up here and return i * 2.\n> Makes sense.\n> \n> > +\t\t\treturn i * 2;\n> > +\t\telse\n> > +\t\t\treturn i * 2 + 1;\n> \n> If not, its the 4 LSB.\n\nYup.\n\nThanks!\n\nPatrick\n"},{"id":"539509","messageId":"ab0hy6AitZFMf3RO@pks.im","threadId":"65299","inReplyTo":"CAOLa=ZSeMS2iKzgMUWix_Sx+e24863PsOazRLrqHtS5hYSUk3A@mail.gmail.com","subject":"Re: [PATCH v2 00/14] odb: generic object name handling","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-20T10:30:35Z","receivedAt":"2026-03-20T10:30:41Z","isPatch":true,"body":"On Fri, Mar 20, 2026 at 03:04:02AM -0700, Karthik Nayak wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > Hi,\n> >\n> > this patch series refactors handling of object names to become pluggable\n> > and thus generic. This includes:\n> >\n> >   - Disambiguation of object names with a common prefix. This is\n> >     required to list candidate objects in case the user has passed a\n> >     non-unique prefix.\n> >\n> >   - Abbreviating an object ID to the shortest prefix required while\n> >     staying unique.\n> >\n> > The logic to compute these operations is specific to the backend, but\n> > not generic. This patch series fixes that by moving the functionality\n> > into the respective backends.\n> >\n> > This patch series may feel somewhat unexiting, but it's not. Especially\n> > abbreviating object IDs is done in lots of places, so this functionality\n> > is overall quite critical. So starting with this series, it is now\n> > possible to do all kinds of local work with an alternative backend:\n> > git-commit(1), git-log(1), git-rev-parse(1), git-merge(1) and many other\n> > commands now work as expected. My MongoDB proof of concept [1] only\n> > requires two commits (the object format extension) on top. And no, I\n> > don't endorse MongoDB or propose it as a future potential backend. It\n> > simply had a good C API that was easy to use.\n> >\n> > Of course, other functionality, especially everything that involves\n> > packfiles, doesn't yet work.\n> >\n> > This patch series is built on top of ca1db8a0f7 (The 17th batch,\n> > 2026-03-16) with ps/object-counting at 6801ffd37d (odb: introduce\n> > generic object counting, 2026-03-12) merged into it.\n> >\n> > Changes in v2:\n> >   - Document `cb_iter` callback.\n> >   - Fix left-over conversion of `odb_source_loose_for_each_object()`.\n> >   - commit message typo fixes.\n> >   - Link to v1: https://lore.kernel.org/r/20260319-b4-pks-odb-source-abbrev-v1-0-5ddebad292b0@pks.im\n> \n> I only got around to reviewing v1 now, but the range-diff here looks\n> good.\n\nThanks for your review!\n\nPatrick\n"},{"id":"539579","messageId":"ab3KjF_1WW7hQBaA@fruit.crustytoothpaste.net","threadId":"65299","inReplyTo":"abzryk0qjlbvy8OL@pks.im","subject":"Re: [PATCH 01/14] oidtree: modernize the code a bit","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-03-20T22:30:36Z","receivedAt":"2026-03-20T22:30:45Z","isPatch":true,"body":"On 2026-03-20 at 06:40:10, Patrick Steinhardt wrote:\n> On Thu, Mar 19, 2026 at 09:08:44AM -0700, Junio C Hamano wrote:\n> > I know the original also used GIT_MAX_HEXSZ to clamp the length for\n> > sanity, but because we know what algorithm is in use, I wonder if we\n> > want to use the limit more specific to it.\n> \n> That assumes that the passed prefix OID actually has an algorithm\n> attached to it, and that may not be the case. We could initialize the\n> overall oidtree with a hash algorithm in `oidtree_init()`, and if so we\n> can then become a bit more thorough with our asserts.\n>\n> But I feel like that would go beyond the smallish cleanups that I'm\n> doing in this patch.\n\nWe should stop assuming that a zero `algo` field in `struct object_id`\nmeans `the_hash_algo` because that makes libification hard and our Rust\ncode doesn't support it (because accessing mutable globals without a\nlock is unsafe)[0].  So in general, I would be fine with forcing callers\nto set an algorithm per OID, both here and elsewhere in our code.\n\nHowever, I am also fine with doing that in a different series for the\nsake of minimalism in this one.  I will probably get to that at some\npoint if nobody else does.\n\n[0] Rust also typically initializes all fields explicitly (and\nzero-initialization is also unsafe), so there's no urge to be lazy and\ndo `memset(p, 0, sizeof(*p))`, which is the usual source of the zero\n`algo` fields in our codebase.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"539700","messageId":"acDcNkIojSNiryFV@pks.im","threadId":"65299","inReplyTo":"ab3KjF_1WW7hQBaA@fruit.crustytoothpaste.net","subject":"Re: [PATCH 01/14] oidtree: modernize the code a bit","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-23T06:22:46Z","receivedAt":"2026-03-23T06:22:52Z","isPatch":true,"body":"On Fri, Mar 20, 2026 at 10:30:36PM +0000, brian m. carlson wrote:\n> On 2026-03-20 at 06:40:10, Patrick Steinhardt wrote:\n> > On Thu, Mar 19, 2026 at 09:08:44AM -0700, Junio C Hamano wrote:\n> > > I know the original also used GIT_MAX_HEXSZ to clamp the length for\n> > > sanity, but because we know what algorithm is in use, I wonder if we\n> > > want to use the limit more specific to it.\n> > \n> > That assumes that the passed prefix OID actually has an algorithm\n> > attached to it, and that may not be the case. We could initialize the\n> > overall oidtree with a hash algorithm in `oidtree_init()`, and if so we\n> > can then become a bit more thorough with our asserts.\n> >\n> > But I feel like that would go beyond the smallish cleanups that I'm\n> > doing in this patch.\n> \n> We should stop assuming that a zero `algo` field in `struct object_id`\n> means `the_hash_algo` because that makes libification hard and our Rust\n> code doesn't support it (because accessing mutable globals without a\n> lock is unsafe)[0].  So in general, I would be fine with forcing callers\n> to set an algorithm per OID, both here and elsewhere in our code.\n> \n> However, I am also fine with doing that in a different series for the\n> sake of minimalism in this one.  I will probably get to that at some\n> point if nobody else does.\n\nYeah, I fully agree that we should get rid of this assumption. Thanks!\n\nPatrick\n"},{"id":"540392","messageId":"87cy0lmjwd.fsf@toon--20250203-5JQV3.mail-host-address-is-not-set","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-8-fe65dcd8c735@pks.im","subject":"Re: [PATCH v2 08/14] object-name: backend-generic `get_short_oid()`","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-03-30T15:12:02Z","receivedAt":"2026-03-30T15:12:14Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> The function `get_short_oid()` takes as input an abbreviated object ID\n> and tries to turn that object ID into the full object ID. This is done\n> by iterating through all objects that have the user-provided prefix. If\n> that yields exactly one object we know that the abbreviated object ID is\n> unambiguous, otherwise it is ambiguous and we print the list of objects\n> that match the prefix.\n>\n> We iterate through all objects with the given prefix by calling both\n> `find_short_packed_object()` and `find_short_object_filename()`, which\n> is of course specific to the \"files\" backend. But we now have a generic\n> way to iterate through objects with a specific prefix.\n>\n> Refactor the code to use `odb_for_each_object()` instead so that it\n> works with object backends different than the \"files\" backend.\n>\n> Remove the now-unused `find_short_packed_object()` function.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  object-name.c | 32 ++++++--------------------------\n>  1 file changed, 6 insertions(+), 26 deletions(-)\n>\n> diff --git a/object-name.c b/object-name.c\n> index 4c3ace150e..7a224ab4af 100644\n> --- a/object-name.c\n> +++ b/object-name.c\n>\n> [snip]\n>\n> @@ -499,6 +477,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n>  \t\t\t\t\t struct object_id *oid,\n>  \t\t\t\t\t unsigned flags)\n>  {\n> +\tstruct odb_for_each_object_options opts = { 0 };\n>  \tint status;\n>  \tstruct disambiguate_state ds;\n>  \tint quietly = !!(flags & GET_OID_QUIETLY);\n> @@ -526,8 +505,10 @@ static enum get_oid_result get_short_oid(struct repository *r,\n>  \telse\n>  \t\tds.fn = default_disambiguate_hint;\n>  \n> -\tfind_short_object_filename(&ds);\n> -\tfind_short_packed_object(&ds);\n> +\topts.prefix = &ds.bin_pfx;\n\nThis `ds` is initialized by init_object_disambiguation(), which calls\nparse_oid_prefix() already. That's nice to see!\n\n-- \nCheers,\nToon\n"},{"id":"540414","messageId":"878qb9mcbv.fsf@toon--20250203-5JQV3.mail-host-address-is-not-set","threadId":"65299","inReplyTo":"20260320-b4-pks-odb-source-abbrev-v2-5-fe65dcd8c735@pks.im","subject":"Re: [PATCH v2 05/14] object-name: move logic to iterate through packed prefixed objects","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-03-30T17:55:32Z","receivedAt":"2026-03-30T17:55:46Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> +static int for_each_prefixed_object_in_pack(\n> +\tstruct packfile_store *store,\n> +\tstruct packed_git *p,\n> +\tconst struct odb_for_each_object_options *opts,\n> +\tstruct packfile_store_for_each_object_wrapper_data *data)\n> +{\n> +\tuint32_t num, i, first = 0;\n> +\tint len = opts->prefix_hex_len > p->repo->hash_algo->hexsz ?\n> +\t\tp->repo->hash_algo->hexsz : opts->prefix_hex_len;\n> +\tint ret;\n> +\n> +\tnum = p->num_objects;\n> +\tbsearch_pack(opts->prefix, p, &first);\n> +\n> +\t/*\n> +\t * At this point, \"first\" is the location of the lowest object\n> +\t * with an object name that could match \"bin_pfx\".  See if we have\n\n\"bin_pfx\" isn't used no more, should be \"opts->prefix\".\n\n-- \nCheers,\nToon\n"},{"id":"540426","messageId":"874ilxm4wp.fsf@toon--20250203-5JQV3.mail-host-address-is-not-set","threadId":"65299","inReplyTo":"ab0hy6AitZFMf3RO@pks.im","subject":"Re: [PATCH v2 00/14] odb: generic object name handling","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-03-30T20:35:50Z","receivedAt":"2026-03-30T20:36:00Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> On Fri, Mar 20, 2026 at 03:04:02AM -0700, Karthik Nayak wrote:\n>> \n>> I only got around to reviewing v1 now, but the range-diff here looks\n>> good.\n>\n> Thanks for your review!\n\nI did a full review of v2 and only got one comment about a stale code\ncomment. Not worth a reroll if you ask me. Looks good to me.\n\n-- \nCheers,\nToon\n"},{"id":"540427","messageId":"xmqqh5pxdobq.fsf@gitster.g","threadId":"65299","inReplyTo":"874ilxm4wp.fsf@toon--20250203-5JQV3.mail-host-address-is-not-set","subject":"Re: [PATCH v2 00/14] odb: generic object name handling","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-03-30T21:01:13Z","receivedAt":"2026-03-30T21:01:15Z","isPatch":true,"body":"Toon Claes <toon@iotcl.com> writes:\n\n> Patrick Steinhardt <ps@pks.im> writes:\n>\n>> On Fri, Mar 20, 2026 at 03:04:02AM -0700, Karthik Nayak wrote:\n>>> \n>>> I only got around to reviewing v1 now, but the range-diff here looks\n>>> good.\n>>\n>> Thanks for your review!\n>\n> I did a full review of v2 and only got one comment about a stale code\n> comment. Not worth a reroll if you ask me. Looks good to me.\n\nThanks, all.  Let me take a final look over the patches and then\nmark the topic for 'next', then.\n"},{"id":"540467","messageId":"actj-nfq_PR863ag@pks.im","threadId":"65299","inReplyTo":"874ilxm4wp.fsf@toon--20250203-5JQV3.mail-host-address-is-not-set","subject":"Re: [PATCH v2 00/14] odb: generic object name handling","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-31T06:04:42Z","receivedAt":"2026-03-31T06:04:48Z","isPatch":true,"body":"On Mon, Mar 30, 2026 at 10:35:50PM +0200, Toon Claes wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > On Fri, Mar 20, 2026 at 03:04:02AM -0700, Karthik Nayak wrote:\n> >> \n> >> I only got around to reviewing v1 now, but the range-diff here looks\n> >> good.\n> >\n> > Thanks for your review!\n> \n> I did a full review of v2 and only got one comment about a stale code\n> comment. Not worth a reroll if you ask me. Looks good to me.\n\nThanks for your review!\n\nPatrick\n"}]}