{"thread":{"id":"65358","subject":"[PATCH] odb: add write_packfile, for_each_unique_abbrev, convert_object_id","startedAt":"2026-03-26T13:07:21Z","lastAt":"2026-03-27T07:04:11Z","messageCount":5,"participants":["Aaron Paterson via GitGitGadget","Patrick Steinhardt","apaterson@pm.me"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"540072","messageId":"pull.2074.git.1774530437562.gitgitgadget@gmail.com","threadId":"65358","inReplyTo":null,"subject":"[PATCH] odb: add write_packfile, for_each_unique_abbrev, convert_object_id","fromName":"Aaron Paterson via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-26T13:07:17Z","receivedAt":"2026-03-26T13:07:21Z","isPatch":true,"body":"From: Aaron Paterson <apaterson@pm.me>\n\nAdd three vtable methods to odb_source that were not part of the\nrecent ps/odb-sources and ps/object-counting series:\n\n - write_packfile: ingest a pack from a file descriptor. The files\n   backend chooses between index-pack (large packs) and\n   unpack-objects (small packs below fetch.unpackLimit). Options\n   cover thin-pack fixing, promisor marking, fsck, lockfile\n   capture, and shallow file passing.\n\n - for_each_unique_abbrev: iterate objects matching a hex prefix\n   for disambiguation. Searches loose objects via oidtree, then\n   multi-pack indices, then non-MIDX packs.\n\n - convert_object_id: translate between hash algorithms using the\n   loose object map. Used during SHA-1 to SHA-256 migration.\n\nAlso add ODB_SOURCE_HELPER to the source type enum, preparing for\nthe helper backend in the next commit.\n\nThe write_packfile vtable method replaces the pattern where callers\nspawn index-pack/unpack-objects directly. fast-import already uses\nodb_write_packfile() and this allows non-files backends to handle\npack ingestion through their own mechanism.\n\nSigned-off-by: Aaron Paterson <apaterson@pm.me>\n---\n    odb: add write_packfile, for_each_unique_abbrev, convert_object_id\n    \n    This adds three ODB source vtable methods that were not part of the\n    recent ps/odb-sources and ps/object-counting series, plus caller routing\n    for object-name.c.\n    \n    New vtable methods:\n    \n     * write_packfile: Ingest a pack from a file descriptor. The files\n       backend chooses between index-pack (large packs) and unpack-objects\n       (small packs below fetch.unpackLimit). Options cover thin-pack\n       fixing, promisor marking, fsck, lockfile capture, and shallow file\n       passing. Non-files backends can handle pack ingestion through their\n       own mechanism.\n    \n     * for_each_unique_abbrev: Iterate objects matching a hex prefix for\n       disambiguation. The files backend searches loose objects via oidtree,\n       multi-pack indices, then non-MIDX packs.\n    \n     * convert_object_id: Translate between hash algorithms using the loose\n       object map. Used during SHA-1 to SHA-256 migration.\n    \n    Caller routing in object-name.c:\n    \n    The abbreviation and disambiguation paths in object-name.c\n    (find_short_object_filename, find_abbrev_len_packed, and\n    find_short_packed_object) directly access files-backend internals (loose\n    cache, pack store, MIDX). These are converted to dispatch through the\n    for_each_unique_abbrev vtable method, so that non-files backends\n    participate in abbreviation and disambiguation through proper\n    abstraction rather than being skipped.\n    \n    This addresses Patrick's feedback on the previous submission [1]: the\n    correct fix for downcast sites is proper vtable abstraction, not\n    skipping non-files backends.\n    \n    Additional:\n    \n     * ODB_SOURCE_HELPER added to the source type enum\n     * odb/source-type.h extracted to avoid circular includes with\n       repository.h\n     * OBJECT_INFO_KEPT_ONLY flag for backends that track kept status\n     * self_contained_out output field on odb_write_packfile_options\n    \n    Motivation: These methods are needed by the local helper backend series\n    [2], which delegates object and reference storage to external git-local-\n    helper processes. sqlite-git [3] is a working proof of concept that\n    stores objects, refs, and reflogs in a single SQLite database with full\n    worktree support.\n    \n    CC: Junio C Hamano gitster@pobox.com, Patrick Steinhardt ps@pks.im\n    \n    [1] https://github.com/gitgitgadget/git/pull/2068.patch [2]\n    https://github.com/gitgitgadget/git/compare/master...MayCXC:git:ps/series-2-helpers-v3.patch\n    [3] https://github.com/MayCXC/sqlite-git\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2074%2FMayCXC%2Fps%2Fseries-1-vtable-v3-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2074/MayCXC/ps/series-1-vtable-v3-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/2074\n\n object-name.c      |  79 ++++++++++----\n odb.c              |  26 +++++\n odb.h              |  26 +++++\n odb/source-files.c | 259 +++++++++++++++++++++++++++++++++++++++++++++\n odb/source.h       | 108 +++++++++++++++++++\n 5 files changed, 480 insertions(+), 18 deletions(-)\n\ndiff --git a/object-name.c b/object-name.c\nindex e5adec4c9d..8f503b985f 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -20,6 +20,7 @@\n #include \"packfile.h\"\n #include \"pretty.h\"\n #include \"object-file.h\"\n+#include \"odb/source.h\"\n #include \"read-cache-ll.h\"\n #include \"repo-settings.h\"\n #include \"repository.h\"\n@@ -111,13 +112,28 @@ static enum cb_next match_prefix(const struct object_id *oid, void *arg)\n \treturn ds->ambiguous ? CB_BREAK : CB_CONTINUE;\n }\n \n+static int disambiguate_cb(const struct object_id *oid,\n+\t\t\t   struct object_info *oi UNUSED, void *data)\n+{\n+\tstruct disambiguate_state *ds = data;\n+\tupdate_candidates(ds, oid);\n+\treturn ds->ambiguous ? 1 : 0;\n+}\n+\n static void find_short_object_filename(struct disambiguate_state *ds)\n {\n \tstruct odb_source *source;\n \n-\tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next)\n-\t\toidtree_each(odb_source_loose_cache(source, &ds->bin_pfx),\n-\t\t\t\t&ds->bin_pfx, ds->len, match_prefix, ds);\n+\tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next) {\n+\t\tif (source->for_each_unique_abbrev) {\n+\t\t\todb_source_for_each_unique_abbrev(\n+\t\t\t\tsource, &ds->bin_pfx, ds->len,\n+\t\t\t\tdisambiguate_cb, ds);\n+\t\t} else {\n+\t\t\toidtree_each(odb_source_loose_cache(source, &ds->bin_pfx),\n+\t\t\t\t\t&ds->bin_pfx, ds->len, match_prefix, ds);\n+\t\t}\n+\t}\n }\n \n static int match_hash(unsigned len, const unsigned char *a, const unsigned char *b)\n@@ -208,15 +224,23 @@ static void find_short_packed_object(struct disambiguate_state *ds)\n \n \todb_prepare_alternates(ds->repo->objects);\n \tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next) {\n-\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n-\t\tif (m)\n-\t\t\tunique_in_midx(m, ds);\n+\t\tif (source->for_each_unique_abbrev) {\n+\t\t\todb_source_for_each_unique_abbrev(\n+\t\t\t\tsource, &ds->bin_pfx, ds->len,\n+\t\t\t\tdisambiguate_cb, ds);\n+\t\t} else {\n+\t\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n+\t\t\tif (m)\n+\t\t\t\tunique_in_midx(m, ds);\n+\t\t}\n \t}\n \n-\trepo_for_each_pack(ds->repo, p) {\n-\t\tif (ds->ambiguous)\n-\t\t\tbreak;\n-\t\tunique_in_pack(p, ds);\n+\tif (!ds->repo->objects->sources->for_each_unique_abbrev) {\n+\t\trepo_for_each_pack(ds->repo, p) {\n+\t\t\tif (ds->ambiguous)\n+\t\t\t\tbreak;\n+\t\t\tunique_in_pack(p, ds);\n+\t\t}\n \t}\n }\n \n@@ -796,19 +820,38 @@ static void find_abbrev_len_for_pack(struct packed_git *p,\n \tmad->init_len = mad->cur_len;\n }\n \n-static void find_abbrev_len_packed(struct min_abbrev_data *mad)\n+static int abbrev_len_cb(const struct object_id *oid,\n+\t\t\t struct object_info *oi UNUSED, void *data)\n {\n-\tstruct packed_git *p;\n+\tstruct min_abbrev_data *mad = data;\n+\textend_abbrev_len(oid, mad);\n+\treturn 0;\n+}\n \n+static void find_abbrev_len_packed(struct min_abbrev_data *mad)\n+{\n \todb_prepare_alternates(mad->repo->objects);\n-\tfor (struct odb_source *source = mad->repo->objects->sources; source; source = source->next) {\n-\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n-\t\tif (m)\n-\t\t\tfind_abbrev_len_for_midx(m, mad);\n+\n+\tfor (struct odb_source *source = mad->repo->objects->sources;\n+\t     source; source = source->next) {\n+\t\tif (source->for_each_unique_abbrev) {\n+\t\t\tmad->init_len = 0;\n+\t\t\todb_source_for_each_unique_abbrev(\n+\t\t\t\tsource, mad->oid, mad->cur_len,\n+\t\t\t\tabbrev_len_cb, mad);\n+\t\t\tmad->init_len = mad->cur_len;\n+\t\t} else {\n+\t\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n+\t\t\tif (m)\n+\t\t\t\tfind_abbrev_len_for_midx(m, mad);\n+\t\t}\n \t}\n \n-\trepo_for_each_pack(mad->repo, p)\n-\t\tfind_abbrev_len_for_pack(p, mad);\n+\tif (!mad->repo->objects->sources->for_each_unique_abbrev) {\n+\t\tstruct packed_git *p;\n+\t\trepo_for_each_pack(mad->repo, p)\n+\t\t\tfind_abbrev_len_for_pack(p, mad);\n+\t}\n }\n \n void strbuf_repo_add_unique_abbrev(struct strbuf *sb, struct repository *repo,\ndiff --git a/odb.c b/odb.c\nindex 350e23f3c0..3032d5492c 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -981,6 +981,32 @@ int odb_write_object_stream(struct object_database *odb,\n \treturn odb_source_write_object_stream(odb->sources, stream, len, oid);\n }\n \n+int odb_write_packfile(struct object_database *odb,\n+\t\t       int pack_fd,\n+\t\t       struct odb_write_packfile_options *opts)\n+{\n+\treturn odb_source_write_packfile(odb->sources, pack_fd, opts);\n+}\n+\n+int odb_for_each_unique_abbrev(struct object_database *odb,\n+\t\t\t       const struct object_id *oid_prefix,\n+\t\t\t       unsigned int prefix_len,\n+\t\t\t       odb_for_each_object_cb cb,\n+\t\t\t       void *cb_data)\n+{\n+\tint ret;\n+\n+\todb_prepare_alternates(odb);\n+\tfor (struct odb_source *source = odb->sources; source; source = source->next) {\n+\t\tret = odb_source_for_each_unique_abbrev(source, oid_prefix,\n+\t\t\t\t\t\t\tprefix_len, cb, cb_data);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t}\n+\n+\treturn 0;\n+}\n+\n struct object_database *odb_new(struct repository *repo,\n \t\t\t\tconst char *primary_source,\n \t\t\t\tconst char *secondary_sources)\ndiff --git a/odb.h b/odb.h\nindex 9aee260105..99d6674706 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -374,6 +374,13 @@ enum object_info_flags {\n \t * clone. Implies OBJECT_INFO_SKIP_FETCH_OBJECT and OBJECT_INFO_QUICK.\n \t */\n \tOBJECT_INFO_FOR_PREFETCH = (OBJECT_INFO_SKIP_FETCH_OBJECT | OBJECT_INFO_QUICK),\n+\n+\t/*\n+\t * Only consider objects marked as \"kept\" (surviving GC). Used by\n+\t * helper backends that track kept status per object. Backends that\n+\t * do not support kept tracking should return -1 (not found).\n+\t */\n+\tOBJECT_INFO_KEPT_ONLY = (1 << 5),\n };\n \n /*\n@@ -570,6 +577,25 @@ int odb_write_object_stream(struct object_database *odb,\n \t\t\t    struct odb_write_stream *stream, size_t len,\n \t\t\t    struct object_id *oid);\n \n+/*\n+ * Ingest a pack from a file descriptor into the primary source.\n+ * Returns 0 on success, a negative error code otherwise.\n+ */\n+struct odb_write_packfile_options;\n+int odb_write_packfile(struct object_database *odb,\n+\t\t       int pack_fd,\n+\t\t       struct odb_write_packfile_options *opts);\n+\n+/*\n+ * Iterate over all objects across all sources whose ID starts with\n+ * the given prefix. Used for object name disambiguation.\n+ */\n+int odb_for_each_unique_abbrev(struct object_database *odb,\n+\t\t\t       const struct object_id *oid_prefix,\n+\t\t\t       unsigned int prefix_len,\n+\t\t\t       odb_for_each_object_cb cb,\n+\t\t\t       void *cb_data);\n+\n void parse_alternates(const char *string,\n \t\t      int sep,\n \t\t      const char *relative_base,\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex c08d8993e3..e450c87f91 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -1,14 +1,21 @@\n #include \"git-compat-util.h\"\n #include \"abspath.h\"\n #include \"chdir-notify.h\"\n+#include \"config.h\"\n #include \"gettext.h\"\n #include \"lockfile.h\"\n+#include \"loose.h\"\n+#include \"midx.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n #include \"odb/source.h\"\n #include \"odb/source-files.h\"\n+#include \"pack-objects.h\"\n #include \"packfile.h\"\n+#include \"run-command.h\"\n #include \"strbuf.h\"\n+#include \"strvec.h\"\n+#include \"oidtree.h\"\n #include \"write-or-die.h\"\n \n static void odb_source_files_reparent(const char *name UNUSED,\n@@ -232,6 +239,255 @@ out:\n \treturn ret;\n }\n \n+static int odb_source_files_write_packfile(struct odb_source *source,\n+\t\t\t\t\t   int pack_fd,\n+\t\t\t\t\t   struct odb_write_packfile_options *opts)\n+{\n+\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\tstruct child_process cmd = CHILD_PROCESS_INIT;\n+\tint fsck_objects = 0;\n+\tint use_index_pack = 1;\n+\tint ret;\n+\n+\tif (opts && opts->nr_objects) {\n+\t\tint transfer_unpack_limit = -1;\n+\t\tint fetch_unpack_limit = -1;\n+\t\tint unpack_limit = 100;\n+\n+\t\trepo_config_get_int(source->odb->repo, \"fetch.unpacklimit\",\n+\t\t\t\t    &fetch_unpack_limit);\n+\t\trepo_config_get_int(source->odb->repo, \"transfer.unpacklimit\",\n+\t\t\t\t    &transfer_unpack_limit);\n+\t\tif (0 <= fetch_unpack_limit)\n+\t\t\tunpack_limit = fetch_unpack_limit;\n+\t\telse if (0 <= transfer_unpack_limit)\n+\t\t\tunpack_limit = transfer_unpack_limit;\n+\n+\t\tif (opts->nr_objects < (unsigned int)unpack_limit &&\n+\t\t    !opts->from_promisor && !opts->lockfile_out)\n+\t\t\tuse_index_pack = 0;\n+\t}\n+\n+\tcmd.in = pack_fd;\n+\tcmd.git_cmd = 1;\n+\n+\tif (!use_index_pack) {\n+\t\tstrvec_push(&cmd.args, \"unpack-objects\");\n+\t\tif (opts && opts->quiet)\n+\t\t\tstrvec_push(&cmd.args, \"-q\");\n+\t\tif (opts && opts->pack_header_version)\n+\t\t\tstrvec_pushf(&cmd.args, \"--pack_header=%\"PRIu32\",%\"PRIu32,\n+\t\t\t\t     opts->pack_header_version,\n+\t\t\t\t     opts->pack_header_entries);\n+\t\trepo_config_get_bool(source->odb->repo, \"transfer.fsckobjects\",\n+\t\t\t\t     &fsck_objects);\n+\t\trepo_config_get_bool(source->odb->repo, \"receive.fsckobjects\",\n+\t\t\t\t     &fsck_objects);\n+\t\tif (fsck_objects)\n+\t\t\tstrvec_push(&cmd.args, \"--strict\");\n+\t\tif (opts && opts->max_input_size)\n+\t\t\tstrvec_pushf(&cmd.args, \"--max-input-size=%lu\",\n+\t\t\t\t     opts->max_input_size);\n+\t\tret = run_command(&cmd);\n+\t\tif (ret)\n+\t\t\treturn error(_(\"unpack-objects failed\"));\n+\t\treturn 0;\n+\t}\n+\n+\tstrvec_push(&cmd.args, \"index-pack\");\n+\tstrvec_push(&cmd.args, \"--stdin\");\n+\tstrvec_push(&cmd.args, \"--keep=write_packfile\");\n+\n+\tif (opts && opts->pack_header_version)\n+\t\tstrvec_pushf(&cmd.args, \"--pack_header=%\"PRIu32\",%\"PRIu32,\n+\t\t\t     opts->pack_header_version,\n+\t\t\t     opts->pack_header_entries);\n+\n+\tif (opts) {\n+\t\tif (opts->use_thin_pack)\n+\t\t\tstrvec_push(&cmd.args, \"--fix-thin\");\n+\t\tif (opts->from_promisor)\n+\t\t\tstrvec_push(&cmd.args, \"--promisor\");\n+\t\tif (opts->check_self_contained)\n+\t\t\tstrvec_push(&cmd.args, \"--check-self-contained-and-connected\");\n+\t\tif (opts->max_input_size)\n+\t\t\tstrvec_pushf(&cmd.args, \"--max-input-size=%lu\",\n+\t\t\t\t     opts->max_input_size);\n+\t\tif (opts->shallow_file)\n+\t\t\tstrvec_pushf(&cmd.env, \"GIT_SHALLOW_FILE=%s\",\n+\t\t\t\t     opts->shallow_file);\n+\t\tif (opts->report_end_of_input)\n+\t\t\tstrvec_push(&cmd.args, \"--report-end-of-input\");\n+\t\tif (opts->fsck_objects)\n+\t\t\tfsck_objects = 1;\n+\t}\n+\n+\tif (!fsck_objects) {\n+\t\trepo_config_get_bool(source->odb->repo, \"transfer.fsckobjects\",\n+\t\t\t\t     &fsck_objects);\n+\t\trepo_config_get_bool(source->odb->repo, \"fetch.fsckobjects\",\n+\t\t\t\t     &fsck_objects);\n+\t}\n+\tif (fsck_objects)\n+\t\tstrvec_push(&cmd.args, \"--strict\");\n+\n+\tif (opts && opts->lockfile_out) {\n+\t\tcmd.out = -1;\n+\t\tret = start_command(&cmd);\n+\t\tif (ret)\n+\t\t\treturn error(_(\"index-pack failed to start\"));\n+\t\t*opts->lockfile_out = index_pack_lockfile(source->odb->repo,\n+\t\t\t\t\t\t\t  cmd.out, NULL);\n+\t\tclose(cmd.out);\n+\t\tret = finish_command(&cmd);\n+\t} else {\n+\t\tret = run_command(&cmd);\n+\t}\n+\n+\tif (ret)\n+\t\treturn error(_(\"index-pack failed\"));\n+\n+\tif (opts && opts->check_self_contained)\n+\t\topts->self_contained_out = 1;\n+\n+\tpackfile_store_reprepare(files->packed);\n+\treturn 0;\n+}\n+\n+static int match_hash_prefix(unsigned len, const unsigned char *a,\n+\t\t\t     const unsigned char *b)\n+{\n+\twhile (len > 1) {\n+\t\tif (*a != *b)\n+\t\t\treturn 0;\n+\t\ta++; b++; len -= 2;\n+\t}\n+\tif (len)\n+\t\tif ((*a ^ *b) & 0xf0)\n+\t\t\treturn 0;\n+\treturn 1;\n+}\n+\n+struct abbrev_cb_data {\n+\todb_for_each_object_cb cb;\n+\tvoid *cb_data;\n+\tint ret;\n+};\n+\n+static enum cb_next abbrev_loose_cb(const struct object_id *oid, void *data)\n+{\n+\tstruct abbrev_cb_data *d = data;\n+\td->ret = d->cb(oid, NULL, d->cb_data);\n+\treturn d->ret ? CB_BREAK : CB_CONTINUE;\n+}\n+\n+static int odb_source_files_for_each_unique_abbrev(struct odb_source *source,\n+\t\t\t\t\t\t   const struct object_id *oid_prefix,\n+\t\t\t\t\t\t   unsigned int prefix_len,\n+\t\t\t\t\t\t   odb_for_each_object_cb cb,\n+\t\t\t\t\t\t   void *cb_data)\n+{\n+\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\tstruct multi_pack_index *m;\n+\tstruct packfile_list_entry *entry;\n+\tunsigned int hexsz = source->odb->repo->hash_algo->hexsz;\n+\tunsigned int len = prefix_len > hexsz ? hexsz : prefix_len;\n+\n+\t/* Search loose objects */\n+\t{\n+\t\tstruct oidtree *tree = odb_source_loose_cache(source, oid_prefix);\n+\t\tif (tree) {\n+\t\t\tstruct abbrev_cb_data d = { cb, cb_data, 0 };\n+\t\t\toidtree_each(tree, oid_prefix, prefix_len, abbrev_loose_cb, &d);\n+\t\t\tif (d.ret)\n+\t\t\t\treturn d.ret;\n+\t\t}\n+\t}\n+\n+\t/* Search multi-pack indices */\n+\tm = get_multi_pack_index(source);\n+\tfor (; m; m = m->base_midx) {\n+\t\tuint32_t num, i, first = 0;\n+\n+\t\tif (!m->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tnum = m->num_objects + m->num_objects_in_base;\n+\t\tbsearch_one_midx(oid_prefix, m, &first);\n+\n+\t\tfor (i = first; i < num; i++) {\n+\t\t\tstruct object_id oid;\n+\t\t\tconst struct object_id *current;\n+\t\t\tint ret;\n+\n+\t\t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n+\t\t\tif (!match_hash_prefix(len, oid_prefix->hash, current->hash))\n+\t\t\t\tbreak;\n+\t\t\tret = cb(current, NULL, cb_data);\n+\t\t\tif (ret)\n+\t\t\t\treturn ret;\n+\t\t}\n+\t}\n+\n+\t/* Search packs not covered by MIDX */\n+\tfor (entry = packfile_store_get_packs(files->packed); entry; entry = entry->next) {\n+\t\tstruct packed_git *p = entry->pack;\n+\t\tuint32_t num, i, first = 0;\n+\n+\t\tif (p->multi_pack_index)\n+\t\t\tcontinue;\n+\t\tif (open_pack_index(p) || !p->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tnum = p->num_objects;\n+\t\tbsearch_pack(oid_prefix, p, &first);\n+\n+\t\tfor (i = first; i < num; i++) {\n+\t\t\tstruct object_id oid;\n+\t\t\tint ret;\n+\n+\t\t\tnth_packed_object_id(&oid, p, i);\n+\t\t\tif (!match_hash_prefix(len, oid_prefix->hash, oid.hash))\n+\t\t\t\tbreak;\n+\t\t\tret = cb(&oid, NULL, cb_data);\n+\t\t\tif (ret)\n+\t\t\t\treturn ret;\n+\t\t}\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int odb_source_files_convert_object_id(struct odb_source *source,\n+\t\t\t\t\t      const struct object_id *src,\n+\t\t\t\t\t      const struct git_hash_algo *to,\n+\t\t\t\t\t      struct object_id *dest)\n+{\n+\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\tstruct loose_object_map *map;\n+\tkh_oid_map_t *hash_map;\n+\tkhiter_t pos;\n+\n+\tif (!files->loose || !files->loose->map)\n+\t\treturn -1;\n+\n+\tmap = files->loose->map;\n+\n+\tif (to == source->odb->repo->compat_hash_algo)\n+\t\thash_map = map->to_compat;\n+\telse if (to == source->odb->repo->hash_algo)\n+\t\thash_map = map->to_storage;\n+\telse\n+\t\treturn -1;\n+\n+\tpos = kh_get_oid_map(hash_map, *src);\n+\tif (pos == kh_end(hash_map))\n+\t\treturn -1;\n+\n+\toidcpy(dest, kh_value(hash_map, pos));\n+\treturn 0;\n+}\n+\n struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \t\t\t\t\t      const char *path,\n \t\t\t\t\t      bool local)\n@@ -256,6 +512,9 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \tfiles->base.begin_transaction = odb_source_files_begin_transaction;\n \tfiles->base.read_alternates = odb_source_files_read_alternates;\n \tfiles->base.write_alternate = odb_source_files_write_alternate;\n+\tfiles->base.write_packfile = odb_source_files_write_packfile;\n+\tfiles->base.for_each_unique_abbrev = odb_source_files_for_each_unique_abbrev;\n+\tfiles->base.convert_object_id = odb_source_files_convert_object_id;\n \n \t/*\n \t * Ideally, we would only ever store absolute paths in the source. This\ndiff --git a/odb/source.h b/odb/source.h\nindex 96c906e7a1..8b898f80ed 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -13,12 +13,42 @@ enum odb_source_type {\n \n \t/* The \"files\" backend that uses loose objects and packfiles. */\n \tODB_SOURCE_FILES,\n+\n+\t/* An external helper process (git-local-<name>). */\n+\tODB_SOURCE_HELPER,\n };\n \n struct object_id;\n struct odb_read_stream;\n struct strvec;\n \n+/*\n+ * Options for write_packfile. When NULL is passed, the backend\n+ * uses sensible defaults.\n+ */\n+struct odb_write_packfile_options {\n+\tunsigned int nr_objects;\n+\tuint32_t pack_header_version;\n+\tuint32_t pack_header_entries;\n+\tint use_thin_pack;\n+\tint from_promisor;\n+\tint fsck_objects;\n+\tint check_self_contained;\n+\tunsigned long max_input_size;\n+\tint quiet;\n+\tint show_progress;\n+\tint report_end_of_input;\n+\tconst char *shallow_file;\n+\tchar **lockfile_out;\n+\n+\t/*\n+\t * Output: set to 1 by the backend if the ingested pack was\n+\t * verified as self-contained (all referenced objects present).\n+\t * Used by the transport layer to skip connectivity checks.\n+\t */\n+\tint self_contained_out;\n+};\n+\n /*\n  * The source is the part of the object database that stores the actual\n  * objects. It thus encapsulates the logic to read and write the specific\n@@ -237,6 +267,45 @@ struct odb_source {\n \t */\n \tint (*write_alternate)(struct odb_source *source,\n \t\t\t       const char *alternate);\n+\n+\t/*\n+\t * Ingest a pack from a file descriptor. Each backend chooses\n+\t * its own ingestion strategy:\n+\t *\n+\t *   - The files backend spawns index-pack (large packs) or\n+\t *     unpack-objects (small packs), then registers the result.\n+\t *\n+\t *   - Non-files backends may parse the pack and write each\n+\t *     object individually through write_object.\n+\t *\n+\t * Returns 0 on success, a negative error code otherwise.\n+\t */\n+\tint (*write_packfile)(struct odb_source *source,\n+\t\t\t      int pack_fd,\n+\t\t\t      struct odb_write_packfile_options *opts);\n+\n+\t/*\n+\t * Iterate over all objects whose object ID starts with the\n+\t * given prefix. Used for object name disambiguation.\n+\t *\n+\t * Returns 0 on success, a negative error code in case\n+\t * iteration has failed, or a non-zero value from the callback.\n+\t */\n+\tint (*for_each_unique_abbrev)(struct odb_source *source,\n+\t\t\t\t      const struct object_id *oid_prefix,\n+\t\t\t\t      unsigned int prefix_len,\n+\t\t\t\t      odb_for_each_object_cb cb,\n+\t\t\t\t      void *cb_data);\n+\n+\t/*\n+\t * Translate an object ID from one hash algorithm to another\n+\t * using the source's internal mapping (for SHA-1/SHA-256\n+\t * migration). Returns 0 on success, -1 if no mapping exists.\n+\t */\n+\tint (*convert_object_id)(struct odb_source *source,\n+\t\t\t\t const struct object_id *src,\n+\t\t\t\t const struct git_hash_algo *to,\n+\t\t\t\t struct object_id *dest);\n };\n \n /*\n@@ -442,4 +511,43 @@ static inline int odb_source_begin_transaction(struct odb_source *source,\n \treturn source->begin_transaction(source, out);\n }\n \n+/*\n+ * Ingest a pack from a file descriptor into the given source. Returns 0 on\n+ * success, a negative error code otherwise.\n+ */\n+static inline int odb_source_write_packfile(struct odb_source *source,\n+\t\t\t\t\t    int pack_fd,\n+\t\t\t\t\t    struct odb_write_packfile_options *opts)\n+{\n+\treturn source->write_packfile(source, pack_fd, opts);\n+}\n+\n+/*\n+ * Iterate over all objects in the source whose ID starts with the given\n+ * prefix. Used for object name disambiguation.\n+ */\n+static inline int odb_source_for_each_unique_abbrev(struct odb_source *source,\n+\t\t\t\t\t\t    const struct object_id *oid_prefix,\n+\t\t\t\t\t\t    unsigned int prefix_len,\n+\t\t\t\t\t\t    odb_for_each_object_cb cb,\n+\t\t\t\t\t\t    void *cb_data)\n+{\n+\treturn source->for_each_unique_abbrev(source, oid_prefix, prefix_len,\n+\t\t\t\t\t      cb, cb_data);\n+}\n+\n+/*\n+ * Translate an object ID between hash algorithms using the source's mapping.\n+ * Returns 0 on success, -1 if no mapping exists.\n+ */\n+static inline int odb_source_convert_object_id(struct odb_source *source,\n+\t\t\t\t\t       const struct object_id *src,\n+\t\t\t\t\t       const struct git_hash_algo *to,\n+\t\t\t\t\t       struct object_id *dest)\n+{\n+\tif (!source->convert_object_id)\n+\t\treturn -1;\n+\treturn source->convert_object_id(source, src, to, dest);\n+}\n+\n #endif\n\nbase-commit: 41688c1a2312f62f44435e1a6d03b4b904b5b0ec\n-- \ngitgitgadget\n"},{"id":"540073","messageId":"pull.2074.v2.git.1774532383055.gitgitgadget@gmail.com","threadId":"65358","inReplyTo":"pull.2074.git.1774530437562.gitgitgadget@gmail.com","subject":"[PATCH v2] odb: add write_packfile, for_each_unique_abbrev, convert_object_id","fromName":"Aaron Paterson via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-03-26T13:39:43Z","receivedAt":"2026-03-26T13:39:46Z","isPatch":true,"body":"From: Aaron Paterson <apaterson@pm.me>\n\nAdd three vtable methods to odb_source that were not part of the\nrecent ps/odb-sources and ps/object-counting series:\n\n - write_packfile: ingest a pack from a file descriptor. The files\n   backend chooses between index-pack (large packs) and\n   unpack-objects (small packs below fetch.unpackLimit). Options\n   cover thin-pack fixing, promisor marking, fsck, lockfile\n   capture, and shallow file passing.\n\n - for_each_unique_abbrev: iterate objects matching a hex prefix\n   for disambiguation. Searches loose objects via oidtree, then\n   multi-pack indices, then non-MIDX packs.\n\n - convert_object_id: translate between hash algorithms using the\n   loose object map. Used during SHA-1 to SHA-256 migration.\n\nAlso add ODB_SOURCE_HELPER to the source type enum, preparing for\nthe helper backend in the next commit.\n\nThe write_packfile vtable method replaces the pattern where callers\nspawn index-pack/unpack-objects directly. fast-import already uses\nodb_write_packfile() and this allows non-files backends to handle\npack ingestion through their own mechanism.\n\nSigned-off-by: Aaron Paterson <apaterson@pm.me>\n---\n    odb: add write_packfile, for_each_unique_abbrev, convert_object_id\n    \n    This adds three ODB source vtable methods that were not part of the\n    recent ps/odb-sources and ps/object-counting series, plus caller routing\n    for object-name.c and fast-import.c.\n    \n    New vtable methods:\n    \n     * write_packfile: Ingest a pack from a file descriptor. The files\n       backend chooses between index-pack (large packs) and unpack-objects\n       (small packs below fetch.unpackLimit). Options cover thin-pack\n       fixing, promisor marking, fsck, lockfile capture, and shallow file\n       passing. Non-files backends can handle pack ingestion through their\n       own mechanism.\n    \n     * for_each_unique_abbrev: Iterate objects matching a hex prefix for\n       disambiguation. The files backend searches loose objects via oidtree,\n       multi-pack indices, then non-MIDX packs.\n    \n     * convert_object_id: Translate between hash algorithms using the loose\n       object map. Used during SHA-1 to SHA-256 migration.\n    \n    Caller routing:\n    \n     * object-name.c: The abbreviation and disambiguation paths\n       (find_short_object_filename, find_abbrev_len_packed, and\n       find_short_packed_object) directly access files-backend internals\n       (loose cache, pack store, MIDX). These are converted to dispatch\n       through the for_each_unique_abbrev vtable method, so that non-files\n       backends participate through proper abstraction rather than being\n       skipped.\n    \n     * fast-import.c: end_packfile() replaced direct pack indexing,\n       registration, and odb_source_files_downcast() with a call to\n       odb_write_packfile(). gfi_unpack_entry() falls back to\n       odb_read_object() when the pack slot is NULL (non-files backends\n       ingest packs without registering them on disk).\n    \n    This addresses Patrick's feedback on the previous submission [1]: the\n    correct fix for downcast sites is proper vtable abstraction, not\n    skipping non-files backends.\n    \n    Additional:\n    \n     * ODB_SOURCE_HELPER added to the source type enum\n     * odb/source-type.h extracted to avoid circular includes with\n       repository.h\n     * OBJECT_INFO_KEPT_ONLY flag for backends that track kept status\n     * self_contained_out output field on odb_write_packfile_options\n    \n    Motivation: These methods are needed by the local helper backend series\n    (Series 2) [2], which delegates object and reference storage to external\n    git-local- helper processes. sqlite-git [3] is a working proof of\n    concept that stores objects, refs, and reflogs in a single SQLite\n    database with full worktree support.\n    \n    CC: Junio C Hamano gitster@pobox.com, Patrick Steinhardt ps@pks.im\n    \n    [1] https://github.com/gitgitgadget/git/pull/2068.patch [2]\n    https://github.com/gitgitgadget/git/compare/master...MayCXC:git:ps/series-2-helpers-v3.patch\n    [3] https://github.com/MayCXC/sqlite-git\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2074%2FMayCXC%2Fps%2Fseries-1-vtable-v3-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2074/MayCXC/ps/series-1-vtable-v3-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/2074\n\nRange-diff vs v1:\n\n 1:  146c7ed0b2 ! 1:  5b3e9a8298 odb: add write_packfile, for_each_unique_abbrev, convert_object_id\n     @@ Commit message\n      \n          Signed-off-by: Aaron Paterson <apaterson@pm.me>\n      \n     + ## builtin/fast-import.c ##\n     +@@ builtin/fast-import.c: static void end_packfile(void)\n     + \trunning = 1;\n     + \tclear_delta_base_cache();\n     + \tif (object_count) {\n     +-\t\tstruct odb_source_files *files = odb_source_files_downcast(pack_data->repo->objects->sources);\n     +-\t\tstruct packed_git *new_p;\n     + \t\tstruct object_id cur_pack_oid;\n     +-\t\tchar *idx_name;\n     + \t\tint i;\n     + \t\tstruct branch *b;\n     + \t\tstruct tag *t;\n     +@@ builtin/fast-import.c: static void end_packfile(void)\n     + \t\t\t\t\t object_count, cur_pack_oid.hash,\n     + \t\t\t\t\t pack_size);\n     + \n     +-\t\tif (object_count <= unpack_limit) {\n     +-\t\t\tif (!loosen_small_pack(pack_data)) {\n     +-\t\t\t\tinvalidate_pack_id(pack_id);\n     +-\t\t\t\tgoto discard_pack;\n     +-\t\t\t}\n     +-\t\t}\n     ++\t\tif (lseek(pack_data->pack_fd, 0, SEEK_SET) < 0)\n     ++\t\t\tdie_errno(_(\"failed seeking to start of '%s'\"),\n     ++\t\t\t\t  pack_data->pack_name);\n     + \n     +-\t\tclose(pack_data->pack_fd);\n     +-\t\tidx_name = keep_pack(create_index());\n     ++\t\tif (odb_write_packfile(the_repository->objects,\n     ++\t\t\t\t       pack_data->pack_fd, NULL))\n     ++\t\t\tdie(_(\"failed to ingest pack\"));\n     + \n     +-\t\t/* Register the packfile with core git's machinery. */\n     +-\t\tnew_p = packfile_store_load_pack(files->packed, idx_name, 1);\n     +-\t\tif (!new_p)\n     +-\t\t\tdie(_(\"core Git rejected index %s\"), idx_name);\n     +-\t\tall_packs[pack_id] = new_p;\n     +-\t\tfree(idx_name);\n     ++\t\t/*\n     ++\t\t * Non-files backends do not register a pack on disk,\n     ++\t\t * so NULL out the slot to prevent use-after-free in\n     ++\t\t * gfi_unpack_entry.\n     ++\t\t */\n     ++\t\tall_packs[pack_id] = NULL;\n     + \n     + \t\t/* Print the boundary */\n     + \t\tif (pack_edges) {\n     +-\t\t\tfprintf(pack_edges, \"%s:\", new_p->pack_name);\n     ++\t\t\tfprintf(pack_edges, \"pack-%s:\",\n     ++\t\t\t\thash_to_hex(pack_data->hash));\n     + \t\t\tfor (i = 0; i < branch_table_sz; i++) {\n     + \t\t\t\tfor (b = branch_table[i]; b; b = b->table_next_branch) {\n     + \t\t\t\t\tif (b->pack_id == pack_id)\n     +@@ builtin/fast-import.c: static void *gfi_unpack_entry(\n     + {\n     + \tenum object_type type;\n     + \tstruct packed_git *p = all_packs[oe->pack_id];\n     ++\tif (!p) {\n     ++\t\t/*\n     ++\t\t * Pack was ingested by a non-files backend via\n     ++\t\t * odb_write_packfile() and is no longer on disk.\n     ++\t\t * Read the object back through the ODB instead.\n     ++\t\t */\n     ++\t\tenum object_type type;\n     ++\t\tenum object_type odb_type;\n     ++\t\treturn odb_read_object(the_repository->objects,\n     ++\t\t\t\t       &oe->idx.oid, &odb_type, sizep);\n     ++\t}\n     + \tif (p == pack_data && p->pack_size < (pack_size + the_hash_algo->rawsz)) {\n     + \t\t/* The object is stored in the packfile we are writing to\n     + \t\t * and we have modified it since the last time we scanned\n     +\n       ## object-name.c ##\n      @@\n       #include \"packfile.h\"\n     @@ odb.c: int odb_write_object_stream(struct object_database *odb,\n       \t\t\t\tconst char *secondary_sources)\n      \n       ## odb.h ##\n     -@@ odb.h: enum object_info_flags {\n     - \t * clone. Implies OBJECT_INFO_SKIP_FETCH_OBJECT and OBJECT_INFO_QUICK.\n     - \t */\n     - \tOBJECT_INFO_FOR_PREFETCH = (OBJECT_INFO_SKIP_FETCH_OBJECT | OBJECT_INFO_QUICK),\n     -+\n     -+\t/*\n     -+\t * Only consider objects marked as \"kept\" (surviving GC). Used by\n     -+\t * helper backends that track kept status per object. Backends that\n     -+\t * do not support kept tracking should return -1 (not found).\n     -+\t */\n     -+\tOBJECT_INFO_KEPT_ONLY = (1 << 5),\n     - };\n     - \n     - /*\n      @@ odb.h: int odb_write_object_stream(struct object_database *odb,\n       \t\t\t    struct odb_write_stream *stream, size_t len,\n       \t\t\t    struct object_id *oid);\n\n\n builtin/fast-import.c |  43 ++++---\n object-name.c         |  79 ++++++++++---\n odb.c                 |  26 +++++\n odb.h                 |  19 ++++\n odb/source-files.c    | 259 ++++++++++++++++++++++++++++++++++++++++++\n odb/source.h          | 108 ++++++++++++++++++\n 6 files changed, 498 insertions(+), 36 deletions(-)\n\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex 9fc6c35b74..160495d9b1 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -876,10 +876,7 @@ static void end_packfile(void)\n \trunning = 1;\n \tclear_delta_base_cache();\n \tif (object_count) {\n-\t\tstruct odb_source_files *files = odb_source_files_downcast(pack_data->repo->objects->sources);\n-\t\tstruct packed_git *new_p;\n \t\tstruct object_id cur_pack_oid;\n-\t\tchar *idx_name;\n \t\tint i;\n \t\tstruct branch *b;\n \t\tstruct tag *t;\n@@ -891,26 +888,25 @@ static void end_packfile(void)\n \t\t\t\t\t object_count, cur_pack_oid.hash,\n \t\t\t\t\t pack_size);\n \n-\t\tif (object_count <= unpack_limit) {\n-\t\t\tif (!loosen_small_pack(pack_data)) {\n-\t\t\t\tinvalidate_pack_id(pack_id);\n-\t\t\t\tgoto discard_pack;\n-\t\t\t}\n-\t\t}\n+\t\tif (lseek(pack_data->pack_fd, 0, SEEK_SET) < 0)\n+\t\t\tdie_errno(_(\"failed seeking to start of '%s'\"),\n+\t\t\t\t  pack_data->pack_name);\n \n-\t\tclose(pack_data->pack_fd);\n-\t\tidx_name = keep_pack(create_index());\n+\t\tif (odb_write_packfile(the_repository->objects,\n+\t\t\t\t       pack_data->pack_fd, NULL))\n+\t\t\tdie(_(\"failed to ingest pack\"));\n \n-\t\t/* Register the packfile with core git's machinery. */\n-\t\tnew_p = packfile_store_load_pack(files->packed, idx_name, 1);\n-\t\tif (!new_p)\n-\t\t\tdie(_(\"core Git rejected index %s\"), idx_name);\n-\t\tall_packs[pack_id] = new_p;\n-\t\tfree(idx_name);\n+\t\t/*\n+\t\t * Non-files backends do not register a pack on disk,\n+\t\t * so NULL out the slot to prevent use-after-free in\n+\t\t * gfi_unpack_entry.\n+\t\t */\n+\t\tall_packs[pack_id] = NULL;\n \n \t\t/* Print the boundary */\n \t\tif (pack_edges) {\n-\t\t\tfprintf(pack_edges, \"%s:\", new_p->pack_name);\n+\t\t\tfprintf(pack_edges, \"pack-%s:\",\n+\t\t\t\thash_to_hex(pack_data->hash));\n \t\t\tfor (i = 0; i < branch_table_sz; i++) {\n \t\t\t\tfor (b = branch_table[i]; b; b = b->table_next_branch) {\n \t\t\t\t\tif (b->pack_id == pack_id)\n@@ -1239,6 +1235,17 @@ static void *gfi_unpack_entry(\n {\n \tenum object_type type;\n \tstruct packed_git *p = all_packs[oe->pack_id];\n+\tif (!p) {\n+\t\t/*\n+\t\t * Pack was ingested by a non-files backend via\n+\t\t * odb_write_packfile() and is no longer on disk.\n+\t\t * Read the object back through the ODB instead.\n+\t\t */\n+\t\tenum object_type type;\n+\t\tenum object_type odb_type;\n+\t\treturn odb_read_object(the_repository->objects,\n+\t\t\t\t       &oe->idx.oid, &odb_type, sizep);\n+\t}\n \tif (p == pack_data && p->pack_size < (pack_size + the_hash_algo->rawsz)) {\n \t\t/* The object is stored in the packfile we are writing to\n \t\t * and we have modified it since the last time we scanned\ndiff --git a/object-name.c b/object-name.c\nindex e5adec4c9d..8f503b985f 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -20,6 +20,7 @@\n #include \"packfile.h\"\n #include \"pretty.h\"\n #include \"object-file.h\"\n+#include \"odb/source.h\"\n #include \"read-cache-ll.h\"\n #include \"repo-settings.h\"\n #include \"repository.h\"\n@@ -111,13 +112,28 @@ static enum cb_next match_prefix(const struct object_id *oid, void *arg)\n \treturn ds->ambiguous ? CB_BREAK : CB_CONTINUE;\n }\n \n+static int disambiguate_cb(const struct object_id *oid,\n+\t\t\t   struct object_info *oi UNUSED, void *data)\n+{\n+\tstruct disambiguate_state *ds = data;\n+\tupdate_candidates(ds, oid);\n+\treturn ds->ambiguous ? 1 : 0;\n+}\n+\n static void find_short_object_filename(struct disambiguate_state *ds)\n {\n \tstruct odb_source *source;\n \n-\tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next)\n-\t\toidtree_each(odb_source_loose_cache(source, &ds->bin_pfx),\n-\t\t\t\t&ds->bin_pfx, ds->len, match_prefix, ds);\n+\tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next) {\n+\t\tif (source->for_each_unique_abbrev) {\n+\t\t\todb_source_for_each_unique_abbrev(\n+\t\t\t\tsource, &ds->bin_pfx, ds->len,\n+\t\t\t\tdisambiguate_cb, ds);\n+\t\t} else {\n+\t\t\toidtree_each(odb_source_loose_cache(source, &ds->bin_pfx),\n+\t\t\t\t\t&ds->bin_pfx, ds->len, match_prefix, ds);\n+\t\t}\n+\t}\n }\n \n static int match_hash(unsigned len, const unsigned char *a, const unsigned char *b)\n@@ -208,15 +224,23 @@ static void find_short_packed_object(struct disambiguate_state *ds)\n \n \todb_prepare_alternates(ds->repo->objects);\n \tfor (source = ds->repo->objects->sources; source && !ds->ambiguous; source = source->next) {\n-\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n-\t\tif (m)\n-\t\t\tunique_in_midx(m, ds);\n+\t\tif (source->for_each_unique_abbrev) {\n+\t\t\todb_source_for_each_unique_abbrev(\n+\t\t\t\tsource, &ds->bin_pfx, ds->len,\n+\t\t\t\tdisambiguate_cb, ds);\n+\t\t} else {\n+\t\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n+\t\t\tif (m)\n+\t\t\t\tunique_in_midx(m, ds);\n+\t\t}\n \t}\n \n-\trepo_for_each_pack(ds->repo, p) {\n-\t\tif (ds->ambiguous)\n-\t\t\tbreak;\n-\t\tunique_in_pack(p, ds);\n+\tif (!ds->repo->objects->sources->for_each_unique_abbrev) {\n+\t\trepo_for_each_pack(ds->repo, p) {\n+\t\t\tif (ds->ambiguous)\n+\t\t\t\tbreak;\n+\t\t\tunique_in_pack(p, ds);\n+\t\t}\n \t}\n }\n \n@@ -796,19 +820,38 @@ static void find_abbrev_len_for_pack(struct packed_git *p,\n \tmad->init_len = mad->cur_len;\n }\n \n-static void find_abbrev_len_packed(struct min_abbrev_data *mad)\n+static int abbrev_len_cb(const struct object_id *oid,\n+\t\t\t struct object_info *oi UNUSED, void *data)\n {\n-\tstruct packed_git *p;\n+\tstruct min_abbrev_data *mad = data;\n+\textend_abbrev_len(oid, mad);\n+\treturn 0;\n+}\n \n+static void find_abbrev_len_packed(struct min_abbrev_data *mad)\n+{\n \todb_prepare_alternates(mad->repo->objects);\n-\tfor (struct odb_source *source = mad->repo->objects->sources; source; source = source->next) {\n-\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n-\t\tif (m)\n-\t\t\tfind_abbrev_len_for_midx(m, mad);\n+\n+\tfor (struct odb_source *source = mad->repo->objects->sources;\n+\t     source; source = source->next) {\n+\t\tif (source->for_each_unique_abbrev) {\n+\t\t\tmad->init_len = 0;\n+\t\t\todb_source_for_each_unique_abbrev(\n+\t\t\t\tsource, mad->oid, mad->cur_len,\n+\t\t\t\tabbrev_len_cb, mad);\n+\t\t\tmad->init_len = mad->cur_len;\n+\t\t} else {\n+\t\t\tstruct multi_pack_index *m = get_multi_pack_index(source);\n+\t\t\tif (m)\n+\t\t\t\tfind_abbrev_len_for_midx(m, mad);\n+\t\t}\n \t}\n \n-\trepo_for_each_pack(mad->repo, p)\n-\t\tfind_abbrev_len_for_pack(p, mad);\n+\tif (!mad->repo->objects->sources->for_each_unique_abbrev) {\n+\t\tstruct packed_git *p;\n+\t\trepo_for_each_pack(mad->repo, p)\n+\t\t\tfind_abbrev_len_for_pack(p, mad);\n+\t}\n }\n \n void strbuf_repo_add_unique_abbrev(struct strbuf *sb, struct repository *repo,\ndiff --git a/odb.c b/odb.c\nindex 350e23f3c0..3032d5492c 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -981,6 +981,32 @@ int odb_write_object_stream(struct object_database *odb,\n \treturn odb_source_write_object_stream(odb->sources, stream, len, oid);\n }\n \n+int odb_write_packfile(struct object_database *odb,\n+\t\t       int pack_fd,\n+\t\t       struct odb_write_packfile_options *opts)\n+{\n+\treturn odb_source_write_packfile(odb->sources, pack_fd, opts);\n+}\n+\n+int odb_for_each_unique_abbrev(struct object_database *odb,\n+\t\t\t       const struct object_id *oid_prefix,\n+\t\t\t       unsigned int prefix_len,\n+\t\t\t       odb_for_each_object_cb cb,\n+\t\t\t       void *cb_data)\n+{\n+\tint ret;\n+\n+\todb_prepare_alternates(odb);\n+\tfor (struct odb_source *source = odb->sources; source; source = source->next) {\n+\t\tret = odb_source_for_each_unique_abbrev(source, oid_prefix,\n+\t\t\t\t\t\t\tprefix_len, cb, cb_data);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t}\n+\n+\treturn 0;\n+}\n+\n struct object_database *odb_new(struct repository *repo,\n \t\t\t\tconst char *primary_source,\n \t\t\t\tconst char *secondary_sources)\ndiff --git a/odb.h b/odb.h\nindex 9aee260105..b7f1a24006 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -570,6 +570,25 @@ int odb_write_object_stream(struct object_database *odb,\n \t\t\t    struct odb_write_stream *stream, size_t len,\n \t\t\t    struct object_id *oid);\n \n+/*\n+ * Ingest a pack from a file descriptor into the primary source.\n+ * Returns 0 on success, a negative error code otherwise.\n+ */\n+struct odb_write_packfile_options;\n+int odb_write_packfile(struct object_database *odb,\n+\t\t       int pack_fd,\n+\t\t       struct odb_write_packfile_options *opts);\n+\n+/*\n+ * Iterate over all objects across all sources whose ID starts with\n+ * the given prefix. Used for object name disambiguation.\n+ */\n+int odb_for_each_unique_abbrev(struct object_database *odb,\n+\t\t\t       const struct object_id *oid_prefix,\n+\t\t\t       unsigned int prefix_len,\n+\t\t\t       odb_for_each_object_cb cb,\n+\t\t\t       void *cb_data);\n+\n void parse_alternates(const char *string,\n \t\t      int sep,\n \t\t      const char *relative_base,\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex c08d8993e3..e450c87f91 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -1,14 +1,21 @@\n #include \"git-compat-util.h\"\n #include \"abspath.h\"\n #include \"chdir-notify.h\"\n+#include \"config.h\"\n #include \"gettext.h\"\n #include \"lockfile.h\"\n+#include \"loose.h\"\n+#include \"midx.h\"\n #include \"object-file.h\"\n #include \"odb.h\"\n #include \"odb/source.h\"\n #include \"odb/source-files.h\"\n+#include \"pack-objects.h\"\n #include \"packfile.h\"\n+#include \"run-command.h\"\n #include \"strbuf.h\"\n+#include \"strvec.h\"\n+#include \"oidtree.h\"\n #include \"write-or-die.h\"\n \n static void odb_source_files_reparent(const char *name UNUSED,\n@@ -232,6 +239,255 @@ out:\n \treturn ret;\n }\n \n+static int odb_source_files_write_packfile(struct odb_source *source,\n+\t\t\t\t\t   int pack_fd,\n+\t\t\t\t\t   struct odb_write_packfile_options *opts)\n+{\n+\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\tstruct child_process cmd = CHILD_PROCESS_INIT;\n+\tint fsck_objects = 0;\n+\tint use_index_pack = 1;\n+\tint ret;\n+\n+\tif (opts && opts->nr_objects) {\n+\t\tint transfer_unpack_limit = -1;\n+\t\tint fetch_unpack_limit = -1;\n+\t\tint unpack_limit = 100;\n+\n+\t\trepo_config_get_int(source->odb->repo, \"fetch.unpacklimit\",\n+\t\t\t\t    &fetch_unpack_limit);\n+\t\trepo_config_get_int(source->odb->repo, \"transfer.unpacklimit\",\n+\t\t\t\t    &transfer_unpack_limit);\n+\t\tif (0 <= fetch_unpack_limit)\n+\t\t\tunpack_limit = fetch_unpack_limit;\n+\t\telse if (0 <= transfer_unpack_limit)\n+\t\t\tunpack_limit = transfer_unpack_limit;\n+\n+\t\tif (opts->nr_objects < (unsigned int)unpack_limit &&\n+\t\t    !opts->from_promisor && !opts->lockfile_out)\n+\t\t\tuse_index_pack = 0;\n+\t}\n+\n+\tcmd.in = pack_fd;\n+\tcmd.git_cmd = 1;\n+\n+\tif (!use_index_pack) {\n+\t\tstrvec_push(&cmd.args, \"unpack-objects\");\n+\t\tif (opts && opts->quiet)\n+\t\t\tstrvec_push(&cmd.args, \"-q\");\n+\t\tif (opts && opts->pack_header_version)\n+\t\t\tstrvec_pushf(&cmd.args, \"--pack_header=%\"PRIu32\",%\"PRIu32,\n+\t\t\t\t     opts->pack_header_version,\n+\t\t\t\t     opts->pack_header_entries);\n+\t\trepo_config_get_bool(source->odb->repo, \"transfer.fsckobjects\",\n+\t\t\t\t     &fsck_objects);\n+\t\trepo_config_get_bool(source->odb->repo, \"receive.fsckobjects\",\n+\t\t\t\t     &fsck_objects);\n+\t\tif (fsck_objects)\n+\t\t\tstrvec_push(&cmd.args, \"--strict\");\n+\t\tif (opts && opts->max_input_size)\n+\t\t\tstrvec_pushf(&cmd.args, \"--max-input-size=%lu\",\n+\t\t\t\t     opts->max_input_size);\n+\t\tret = run_command(&cmd);\n+\t\tif (ret)\n+\t\t\treturn error(_(\"unpack-objects failed\"));\n+\t\treturn 0;\n+\t}\n+\n+\tstrvec_push(&cmd.args, \"index-pack\");\n+\tstrvec_push(&cmd.args, \"--stdin\");\n+\tstrvec_push(&cmd.args, \"--keep=write_packfile\");\n+\n+\tif (opts && opts->pack_header_version)\n+\t\tstrvec_pushf(&cmd.args, \"--pack_header=%\"PRIu32\",%\"PRIu32,\n+\t\t\t     opts->pack_header_version,\n+\t\t\t     opts->pack_header_entries);\n+\n+\tif (opts) {\n+\t\tif (opts->use_thin_pack)\n+\t\t\tstrvec_push(&cmd.args, \"--fix-thin\");\n+\t\tif (opts->from_promisor)\n+\t\t\tstrvec_push(&cmd.args, \"--promisor\");\n+\t\tif (opts->check_self_contained)\n+\t\t\tstrvec_push(&cmd.args, \"--check-self-contained-and-connected\");\n+\t\tif (opts->max_input_size)\n+\t\t\tstrvec_pushf(&cmd.args, \"--max-input-size=%lu\",\n+\t\t\t\t     opts->max_input_size);\n+\t\tif (opts->shallow_file)\n+\t\t\tstrvec_pushf(&cmd.env, \"GIT_SHALLOW_FILE=%s\",\n+\t\t\t\t     opts->shallow_file);\n+\t\tif (opts->report_end_of_input)\n+\t\t\tstrvec_push(&cmd.args, \"--report-end-of-input\");\n+\t\tif (opts->fsck_objects)\n+\t\t\tfsck_objects = 1;\n+\t}\n+\n+\tif (!fsck_objects) {\n+\t\trepo_config_get_bool(source->odb->repo, \"transfer.fsckobjects\",\n+\t\t\t\t     &fsck_objects);\n+\t\trepo_config_get_bool(source->odb->repo, \"fetch.fsckobjects\",\n+\t\t\t\t     &fsck_objects);\n+\t}\n+\tif (fsck_objects)\n+\t\tstrvec_push(&cmd.args, \"--strict\");\n+\n+\tif (opts && opts->lockfile_out) {\n+\t\tcmd.out = -1;\n+\t\tret = start_command(&cmd);\n+\t\tif (ret)\n+\t\t\treturn error(_(\"index-pack failed to start\"));\n+\t\t*opts->lockfile_out = index_pack_lockfile(source->odb->repo,\n+\t\t\t\t\t\t\t  cmd.out, NULL);\n+\t\tclose(cmd.out);\n+\t\tret = finish_command(&cmd);\n+\t} else {\n+\t\tret = run_command(&cmd);\n+\t}\n+\n+\tif (ret)\n+\t\treturn error(_(\"index-pack failed\"));\n+\n+\tif (opts && opts->check_self_contained)\n+\t\topts->self_contained_out = 1;\n+\n+\tpackfile_store_reprepare(files->packed);\n+\treturn 0;\n+}\n+\n+static int match_hash_prefix(unsigned len, const unsigned char *a,\n+\t\t\t     const unsigned char *b)\n+{\n+\twhile (len > 1) {\n+\t\tif (*a != *b)\n+\t\t\treturn 0;\n+\t\ta++; b++; len -= 2;\n+\t}\n+\tif (len)\n+\t\tif ((*a ^ *b) & 0xf0)\n+\t\t\treturn 0;\n+\treturn 1;\n+}\n+\n+struct abbrev_cb_data {\n+\todb_for_each_object_cb cb;\n+\tvoid *cb_data;\n+\tint ret;\n+};\n+\n+static enum cb_next abbrev_loose_cb(const struct object_id *oid, void *data)\n+{\n+\tstruct abbrev_cb_data *d = data;\n+\td->ret = d->cb(oid, NULL, d->cb_data);\n+\treturn d->ret ? CB_BREAK : CB_CONTINUE;\n+}\n+\n+static int odb_source_files_for_each_unique_abbrev(struct odb_source *source,\n+\t\t\t\t\t\t   const struct object_id *oid_prefix,\n+\t\t\t\t\t\t   unsigned int prefix_len,\n+\t\t\t\t\t\t   odb_for_each_object_cb cb,\n+\t\t\t\t\t\t   void *cb_data)\n+{\n+\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\tstruct multi_pack_index *m;\n+\tstruct packfile_list_entry *entry;\n+\tunsigned int hexsz = source->odb->repo->hash_algo->hexsz;\n+\tunsigned int len = prefix_len > hexsz ? hexsz : prefix_len;\n+\n+\t/* Search loose objects */\n+\t{\n+\t\tstruct oidtree *tree = odb_source_loose_cache(source, oid_prefix);\n+\t\tif (tree) {\n+\t\t\tstruct abbrev_cb_data d = { cb, cb_data, 0 };\n+\t\t\toidtree_each(tree, oid_prefix, prefix_len, abbrev_loose_cb, &d);\n+\t\t\tif (d.ret)\n+\t\t\t\treturn d.ret;\n+\t\t}\n+\t}\n+\n+\t/* Search multi-pack indices */\n+\tm = get_multi_pack_index(source);\n+\tfor (; m; m = m->base_midx) {\n+\t\tuint32_t num, i, first = 0;\n+\n+\t\tif (!m->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tnum = m->num_objects + m->num_objects_in_base;\n+\t\tbsearch_one_midx(oid_prefix, m, &first);\n+\n+\t\tfor (i = first; i < num; i++) {\n+\t\t\tstruct object_id oid;\n+\t\t\tconst struct object_id *current;\n+\t\t\tint ret;\n+\n+\t\t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n+\t\t\tif (!match_hash_prefix(len, oid_prefix->hash, current->hash))\n+\t\t\t\tbreak;\n+\t\t\tret = cb(current, NULL, cb_data);\n+\t\t\tif (ret)\n+\t\t\t\treturn ret;\n+\t\t}\n+\t}\n+\n+\t/* Search packs not covered by MIDX */\n+\tfor (entry = packfile_store_get_packs(files->packed); entry; entry = entry->next) {\n+\t\tstruct packed_git *p = entry->pack;\n+\t\tuint32_t num, i, first = 0;\n+\n+\t\tif (p->multi_pack_index)\n+\t\t\tcontinue;\n+\t\tif (open_pack_index(p) || !p->num_objects)\n+\t\t\tcontinue;\n+\n+\t\tnum = p->num_objects;\n+\t\tbsearch_pack(oid_prefix, p, &first);\n+\n+\t\tfor (i = first; i < num; i++) {\n+\t\t\tstruct object_id oid;\n+\t\t\tint ret;\n+\n+\t\t\tnth_packed_object_id(&oid, p, i);\n+\t\t\tif (!match_hash_prefix(len, oid_prefix->hash, oid.hash))\n+\t\t\t\tbreak;\n+\t\t\tret = cb(&oid, NULL, cb_data);\n+\t\t\tif (ret)\n+\t\t\t\treturn ret;\n+\t\t}\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static int odb_source_files_convert_object_id(struct odb_source *source,\n+\t\t\t\t\t      const struct object_id *src,\n+\t\t\t\t\t      const struct git_hash_algo *to,\n+\t\t\t\t\t      struct object_id *dest)\n+{\n+\tstruct odb_source_files *files = odb_source_files_downcast(source);\n+\tstruct loose_object_map *map;\n+\tkh_oid_map_t *hash_map;\n+\tkhiter_t pos;\n+\n+\tif (!files->loose || !files->loose->map)\n+\t\treturn -1;\n+\n+\tmap = files->loose->map;\n+\n+\tif (to == source->odb->repo->compat_hash_algo)\n+\t\thash_map = map->to_compat;\n+\telse if (to == source->odb->repo->hash_algo)\n+\t\thash_map = map->to_storage;\n+\telse\n+\t\treturn -1;\n+\n+\tpos = kh_get_oid_map(hash_map, *src);\n+\tif (pos == kh_end(hash_map))\n+\t\treturn -1;\n+\n+\toidcpy(dest, kh_value(hash_map, pos));\n+\treturn 0;\n+}\n+\n struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \t\t\t\t\t      const char *path,\n \t\t\t\t\t      bool local)\n@@ -256,6 +512,9 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \tfiles->base.begin_transaction = odb_source_files_begin_transaction;\n \tfiles->base.read_alternates = odb_source_files_read_alternates;\n \tfiles->base.write_alternate = odb_source_files_write_alternate;\n+\tfiles->base.write_packfile = odb_source_files_write_packfile;\n+\tfiles->base.for_each_unique_abbrev = odb_source_files_for_each_unique_abbrev;\n+\tfiles->base.convert_object_id = odb_source_files_convert_object_id;\n \n \t/*\n \t * Ideally, we would only ever store absolute paths in the source. This\ndiff --git a/odb/source.h b/odb/source.h\nindex 96c906e7a1..8b898f80ed 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -13,12 +13,42 @@ enum odb_source_type {\n \n \t/* The \"files\" backend that uses loose objects and packfiles. */\n \tODB_SOURCE_FILES,\n+\n+\t/* An external helper process (git-local-<name>). */\n+\tODB_SOURCE_HELPER,\n };\n \n struct object_id;\n struct odb_read_stream;\n struct strvec;\n \n+/*\n+ * Options for write_packfile. When NULL is passed, the backend\n+ * uses sensible defaults.\n+ */\n+struct odb_write_packfile_options {\n+\tunsigned int nr_objects;\n+\tuint32_t pack_header_version;\n+\tuint32_t pack_header_entries;\n+\tint use_thin_pack;\n+\tint from_promisor;\n+\tint fsck_objects;\n+\tint check_self_contained;\n+\tunsigned long max_input_size;\n+\tint quiet;\n+\tint show_progress;\n+\tint report_end_of_input;\n+\tconst char *shallow_file;\n+\tchar **lockfile_out;\n+\n+\t/*\n+\t * Output: set to 1 by the backend if the ingested pack was\n+\t * verified as self-contained (all referenced objects present).\n+\t * Used by the transport layer to skip connectivity checks.\n+\t */\n+\tint self_contained_out;\n+};\n+\n /*\n  * The source is the part of the object database that stores the actual\n  * objects. It thus encapsulates the logic to read and write the specific\n@@ -237,6 +267,45 @@ struct odb_source {\n \t */\n \tint (*write_alternate)(struct odb_source *source,\n \t\t\t       const char *alternate);\n+\n+\t/*\n+\t * Ingest a pack from a file descriptor. Each backend chooses\n+\t * its own ingestion strategy:\n+\t *\n+\t *   - The files backend spawns index-pack (large packs) or\n+\t *     unpack-objects (small packs), then registers the result.\n+\t *\n+\t *   - Non-files backends may parse the pack and write each\n+\t *     object individually through write_object.\n+\t *\n+\t * Returns 0 on success, a negative error code otherwise.\n+\t */\n+\tint (*write_packfile)(struct odb_source *source,\n+\t\t\t      int pack_fd,\n+\t\t\t      struct odb_write_packfile_options *opts);\n+\n+\t/*\n+\t * Iterate over all objects whose object ID starts with the\n+\t * given prefix. Used for object name disambiguation.\n+\t *\n+\t * Returns 0 on success, a negative error code in case\n+\t * iteration has failed, or a non-zero value from the callback.\n+\t */\n+\tint (*for_each_unique_abbrev)(struct odb_source *source,\n+\t\t\t\t      const struct object_id *oid_prefix,\n+\t\t\t\t      unsigned int prefix_len,\n+\t\t\t\t      odb_for_each_object_cb cb,\n+\t\t\t\t      void *cb_data);\n+\n+\t/*\n+\t * Translate an object ID from one hash algorithm to another\n+\t * using the source's internal mapping (for SHA-1/SHA-256\n+\t * migration). Returns 0 on success, -1 if no mapping exists.\n+\t */\n+\tint (*convert_object_id)(struct odb_source *source,\n+\t\t\t\t const struct object_id *src,\n+\t\t\t\t const struct git_hash_algo *to,\n+\t\t\t\t struct object_id *dest);\n };\n \n /*\n@@ -442,4 +511,43 @@ static inline int odb_source_begin_transaction(struct odb_source *source,\n \treturn source->begin_transaction(source, out);\n }\n \n+/*\n+ * Ingest a pack from a file descriptor into the given source. Returns 0 on\n+ * success, a negative error code otherwise.\n+ */\n+static inline int odb_source_write_packfile(struct odb_source *source,\n+\t\t\t\t\t    int pack_fd,\n+\t\t\t\t\t    struct odb_write_packfile_options *opts)\n+{\n+\treturn source->write_packfile(source, pack_fd, opts);\n+}\n+\n+/*\n+ * Iterate over all objects in the source whose ID starts with the given\n+ * prefix. Used for object name disambiguation.\n+ */\n+static inline int odb_source_for_each_unique_abbrev(struct odb_source *source,\n+\t\t\t\t\t\t    const struct object_id *oid_prefix,\n+\t\t\t\t\t\t    unsigned int prefix_len,\n+\t\t\t\t\t\t    odb_for_each_object_cb cb,\n+\t\t\t\t\t\t    void *cb_data)\n+{\n+\treturn source->for_each_unique_abbrev(source, oid_prefix, prefix_len,\n+\t\t\t\t\t      cb, cb_data);\n+}\n+\n+/*\n+ * Translate an object ID between hash algorithms using the source's mapping.\n+ * Returns 0 on success, -1 if no mapping exists.\n+ */\n+static inline int odb_source_convert_object_id(struct odb_source *source,\n+\t\t\t\t\t       const struct object_id *src,\n+\t\t\t\t\t       const struct git_hash_algo *to,\n+\t\t\t\t\t       struct object_id *dest)\n+{\n+\tif (!source->convert_object_id)\n+\t\treturn -1;\n+\treturn source->convert_object_id(source, src, to, dest);\n+}\n+\n #endif\n\nbase-commit: 41688c1a2312f62f44435e1a6d03b4b904b5b0ec\n-- \ngitgitgadget\n"},{"id":"540074","messageId":"acU7eJ0MpUVhCs6-@pks.im","threadId":"65358","inReplyTo":"pull.2074.v2.git.1774532383055.gitgitgadget@gmail.com","subject":"Re: [PATCH v2] odb: add write_packfile, for_each_unique_abbrev, convert_object_id","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-26T13:58:16Z","receivedAt":"2026-03-26T13:58:22Z","isPatch":true,"body":"On Thu, Mar 26, 2026 at 01:39:43PM +0000, Aaron Paterson via GitGitGadget wrote:\n> From: Aaron Paterson <apaterson@pm.me>\n> \n> Add three vtable methods to odb_source that were not part of the\n> recent ps/odb-sources and ps/object-counting series:\n> \n>  - write_packfile: ingest a pack from a file descriptor. The files\n>    backend chooses between index-pack (large packs) and\n>    unpack-objects (small packs below fetch.unpackLimit). Options\n>    cover thin-pack fixing, promisor marking, fsck, lockfile\n>    capture, and shallow file passing.\n> \n>  - for_each_unique_abbrev: iterate objects matching a hex prefix\n>    for disambiguation. Searches loose objects via oidtree, then\n>    multi-pack indices, then non-MIDX packs.\n> \n>  - convert_object_id: translate between hash algorithms using the\n>    loose object map. Used during SHA-1 to SHA-256 migration.\n\nThis will conflict with ps/odb-generic-object-name-handling, which\nalready introduces generic callbacks for `for_each_unique_abbrev()`.\nThere's also ongoing work by Justin to handle writing packfiles via the\nODB transaction interface.\n\n> Also add ODB_SOURCE_HELPER to the source type enum, preparing for\n> the helper backend in the next commit.\n\nHuh.\n\n> The write_packfile vtable method replaces the pattern where callers\n> spawn index-pack/unpack-objects directly. fast-import already uses\n> odb_write_packfile() and this allows non-files backends to handle\n> pack ingestion through their own mechanism.\n\nI'm again a bit puzzled, same as with your previous patch series. It\nwould be nice to collaborate on this topic, but that will require a bit\nmore coordination than just sending in a patch series as things are\nquite in flux here.\n\nPatrick\n"},{"id":"540082","messageId":"DpRTZzuEPU7m8kvCckzHYEK380REXLfunHXO4hE3qgAZsKPSNtyqkBT2oRzusMxvIDgLVkt4FOis0HKBTEJPIRQgcbYM6QZdpzN6-Y-F7WE=@pm.me","threadId":"65358","inReplyTo":"acU7eJ0MpUVhCs6-@pks.im","subject":"Re: [PATCH v2] odb: add write_packfile, for_each_unique_abbrev, convert_object_id","fromName":"","fromEmail":"apaterson@pm.me","sentAt":"2026-03-26T14:21:07Z","receivedAt":"2026-03-26T14:21:18Z","isPatch":true,"body":"Of course, and my apologies, gitgadget is not formatting these messages as clearly as I would like them to be.\n\nBoth this series and the last were adapted from my fork that supports [1] with a feature similar to gitremote-helpers. My hope is that the fork can converge with master so that sqlite-git can become redistributable. The local backends vtable was already a step in this direction, so the question is if letting users bring their own local backends, the way they currently can with helpers for remote backends, is in scope for git core.\n\nEither way, it sounds like series 1 will be covered by upstream, so next I would like to contribute support for git-local-* helpers. This allows users to create .git repositories with storage formats other than packs and builtin alternatives like reftables, which seems appropriate as direct sqlite support would probably be out of scope for core. Local helpers are already implemented in [2] but if it makes sense to hold off and rebuild it after e.g. ps/odb-generic-object-name-handling is merged, I am not in such a rush.\n\n[1] https://github.com/mayCXC/sqlite-git\n[2] https://github.com/gitgitgadget/git/compare/master...MayCXC:git:ps/series-2-helpers-v3.patch\n\n- Aaron\n\nOn Thursday, March 26th, 2026 at 7:58 AM, Patrick Steinhardt <ps@pks.im> wrote:\n\n> On Thu, Mar 26, 2026 at 01:39:43PM +0000, Aaron Paterson via GitGitGadget wrote:\n> > From: Aaron Paterson <apaterson@pm.me>\n> >\n> > Add three vtable methods to odb_source that were not part of the\n> > recent ps/odb-sources and ps/object-counting series:\n> >\n> >  - write_packfile: ingest a pack from a file descriptor. The files\n> >    backend chooses between index-pack (large packs) and\n> >    unpack-objects (small packs below fetch.unpackLimit). Options\n> >    cover thin-pack fixing, promisor marking, fsck, lockfile\n> >    capture, and shallow file passing.\n> >\n> >  - for_each_unique_abbrev: iterate objects matching a hex prefix\n> >    for disambiguation. Searches loose objects via oidtree, then\n> >    multi-pack indices, then non-MIDX packs.\n> >\n> >  - convert_object_id: translate between hash algorithms using the\n> >    loose object map. Used during SHA-1 to SHA-256 migration.\n> \n> This will conflict with ps/odb-generic-object-name-handling, which\n> already introduces generic callbacks for `for_each_unique_abbrev()`.\n> There's also ongoing work by Justin to handle writing packfiles via the\n> ODB transaction interface.\n> \n> > Also add ODB_SOURCE_HELPER to the source type enum, preparing for\n> > the helper backend in the next commit.\n> \n> Huh.\n> \n> > The write_packfile vtable method replaces the pattern where callers\n> > spawn index-pack/unpack-objects directly. fast-import already uses\n> > odb_write_packfile() and this allows non-files backends to handle\n> > pack ingestion through their own mechanism.\n> \n> I'm again a bit puzzled, same as with your previous patch series. It\n> would be nice to collaborate on this topic, but that will require a bit\n> more coordination than just sending in a patch series as things are\n> quite in flux here.\n> \n> Patrick\n> \n"},{"id":"540162","messageId":"acYr5DZTPuOlyxvi@pks.im","threadId":"65358","inReplyTo":"DpRTZzuEPU7m8kvCckzHYEK380REXLfunHXO4hE3qgAZsKPSNtyqkBT2oRzusMxvIDgLVkt4FOis0HKBTEJPIRQgcbYM6QZdpzN6-Y-F7WE=@pm.me","subject":"Re: [PATCH v2] odb: add write_packfile, for_each_unique_abbrev, convert_object_id","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-03-27T07:04:04Z","receivedAt":"2026-03-27T07:04:11Z","isPatch":true,"body":"On Thu, Mar 26, 2026 at 02:21:07PM +0000, apaterson@pm.me wrote:\n> Of course, and my apologies, gitgadget is not formatting these\n> messages as clearly as I would like them to be.\n> \n> Both this series and the last were adapted from my fork that supports\n> [1] with a feature similar to gitremote-helpers. My hope is that the\n> fork can converge with master so that sqlite-git can become\n> redistributable. The local backends vtable was already a step in this\n> direction, so the question is if letting users bring their own local\n> backends, the way they currently can with helpers for remote backends,\n> is in scope for git core.\n\nThanks for the context! This also matches with our eventual goal, even\nthough we rather envision that it makes more sense to maybe use a plugin\nin the form of a shared object instead of using a helper executable.\n\n> Either way, it sounds like series 1 will be covered by upstream, so\n> next I would like to contribute support for git-local-* helpers. This\n> allows users to create .git repositories with storage formats other\n> than packs and builtin alternatives like reftables, which seems\n> appropriate as direct sqlite support would probably be out of scope\n> for core. Local helpers are already implemented in [2] but if it makes\n> sense to hold off and rebuild it after e.g.\n> ps/odb-generic-object-name-handling is merged, I am not in such a\n> rush.\n\nI've currently got around 10 more patch series pending that are mostly\nready to be sent out, but that all build on one another. As said, there\nis a ton of stuff changing in the area of pluggable object databases,\nand I expect it'll probably take two more Git releases until we have\nfully carved out the foundation. Once that's done I think it should\nbecome quieter, and at that point it'll become easier to also do\ndrive-by contributions without requiring too much coordination.\n\nYou can have a look at [1], which is our (non-official and\nGitLab-specific) epic for the work that we have planned over the next\nfew months. Maybe it helps you a bit to figure out where we're going.\n\nMore concretely, next steps will be:\n\n  - I plan to turn in-memory, loose and packed backends into proper ODB\n    sources.\n\n  - I plan to introduce backend-specific consistency checks.\n\n  - I plan to introduce backend-specific logic for optimizations.\n\n  - I plan to introduce backend-specific logic of generating packfiles.\n\n  - Justin is revamping how writes work and plans to refactor existing\n    callers that do ad-hoc transactions. This will also eventually cover\n    writing packfiles into the ODB.\n\nIf you'd like to get involved earlier I'd propose that we sync off-list\nto figure out how to collaborate without stepping on each others toes\nall the time :)\n\nThanks!\n\nPatrick\n"}]}