{"thread":{"id":"66173","subject":"[PATCH] rev-parse: compute object names in another hash algorithm","startedAt":"2026-08-14T13:48:07Z","lastAt":"2026-08-14T13:48:07Z","messageCount":1,"participants":["Dimitri John Ledkov"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"550614","messageId":"20260814134804.680154-1-dimitri.ledkov@chainguard.dev","threadId":"66173","inReplyTo":null,"subject":"[PATCH] rev-parse: compute object names in another hash algorithm","fromName":"Dimitri John Ledkov","fromEmail":"dimitri.ledkov@chainguard.dev","sentAt":"2026-08-14T13:48:04Z","receivedAt":"2026-08-14T13:48:07Z","isPatch":true,"body":"\"git rev-parse --output-object-format=sha256 HEAD^{tree}\" could not\nanswer in a SHA-1 repository.  And vice-versa.  Without\nextensions.compatObjectFormat the option was rejected outright, and with\nit the answer was wrong: the extension records a mapping only for objects\nwritten while it is enabled, so for the objects a repository already\ncontained the lookup missed, repo_oid_to_algop() left the object ID\nuntouched, and rev-parse printed the SHA-1 name with a zero exit code.\n\nThe machinery to convert object contents between algorithms is already\nhere.  convert_tree_object() and friends rewrite the object IDs a tree,\ncommit or tag refers to, calling repo_oid_to_algop() for each one, so the\nrecursion is written; what is missing is a base case.  Give\nrepo_oid_to_algop() one: on a lookup miss, read the object, convert its\ncontents, and hash the result.  A blob needs no conversion at all, as its\ncontents are the same either way, only a rehash.\n\nNames computed this way are remembered during the process execution,\nso that a tree that reaches the same subtree by several paths pays for\nit once.  These are not stored permamently.  In the future a side-car\ncache could be added for these results, or eventual dual-format v3\ncould be used to lazy compute/convert and store these.\n\nCommits are refused rather than computed.  A commit names its parents, so\nconverting one converts the entire history behind it; that is a repository\nconversion, not an answer about a single object, and recursing over it\nhere would be bounded only by the length of the history.  A tree holding a\nsubmodule is refused for the same reason, and now says which entry is at\nfault instead of reporting an unreadable object.\n\nSince names can now be computed, accept any algorithm git knows rather\nthan only the configured compatibility one, and check what\nrepo_oid_to_algop() returns so that a name we cannot produce is an error\ninstead of the storage name.\n\nOn git.git's tree at 9936c1b52a (3089 files, 26MB) this names the tree\nin 0.26s, against 0.25s for reading and hashing that content by hand,\nso close to the floor for the work involved. On linux.git tree of\n1.6GiB it takes 8s to compute.\n\nInitially, I have implemented contrib git-sha256-tree.sh script using\na temporary git repo and perform fast-export/fast-import of a single\ncommit and its tree.  It is a lot slower due to needless work of\nfast-import that is discarded.  However, a contrib script carries less\nmaintainance and is more portable to existing installations.  If there\nis interest, I can publish that implementation as well.  The goal is\nto help with interop, and have the ability to record tree names in\neither format today; to check again later if and when a given project\nswitches to sha256 format.\n\nAssisted-by: Claude Opus 5 xHigh\nSigned-off-by: Dimitri John Ledkov <dimitri.ledkov@chainguard.dev>\n---\n Documentation/git-rev-parse.adoc |  13 +++\n builtin/rev-parse.c              |  23 ++++-\n object-file-convert.c            | 145 +++++++++++++++++++++++++++++--\n object-file-convert.h            |   6 ++\n repository.c                     |   3 +\n repository.h                     |   8 ++\n t/meson.build                    |   1 +\n t/t1018-output-object-format.sh  | 118 +++++++++++++++++++++++++\n 8 files changed, 308 insertions(+), 9 deletions(-)\n create mode 100755 t/t1018-output-object-format.sh\n\ndiff --git a/Documentation/git-rev-parse.adoc b/Documentation/git-rev-parse.adoc\nindex 5398691f3f..c18ae9ca81 100644\n--- a/Documentation/git-rev-parse.adoc\n+++ b/Documentation/git-rev-parse.adoc\n@@ -181,6 +181,19 @@ Specifying \"sha256\" translates if necessary and returns a sha256 oid.\n +\n Specifying \"storage\" translates if necessary and returns an oid in\n encoded in the storage hash algorithm.\n++\n+A translation that `extensions.compatObjectFormat` did not record, which is\n+every object a repository already contained when that extension was turned\n+on, is computed on demand.  Since an object's name in another algorithm is\n+defined over the names of the objects it refers to, this reads everything\n+reachable from the object, and so costs time proportional to the total size\n+of that content.\n++\n+Commits cannot be named this way.  A commit refers to its parents, so naming\n+one would mean converting the whole history behind it, which is a repository\n+conversion rather than a question about a single object.  For the same\n+reason a tree containing a submodule cannot be named, as its gitlink refers\n+to a commit in another repository.\n \n Options for Objects\n ~~~~~~~~~~~~~~~~~~~\ndiff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\nindex 43693454d5..7dd06f324e 100644\n--- a/builtin/rev-parse.c\n+++ b/builtin/rev-parse.c\n@@ -876,8 +876,21 @@ int cmd_rev_parse(int argc,\n \t\t\t\t\tflags |= GET_OID_HASH_ANY;\n \t\t\t\t\toutput_algo = compat;\n \t\t\t\t\tcontinue;\n+\t\t\t\t} else {\n+\t\t\t\t\t/*\n+\t\t\t\t\t * Names in another algorithm can be\n+\t\t\t\t\t * computed on demand, so accept any\n+\t\t\t\t\t * algorithm we know rather than only\n+\t\t\t\t\t * the configured compatibility one.\n+\t\t\t\t\t */\n+\t\t\t\t\tuint32_t algo = hash_algo_by_name(arg);\n+\n+\t\t\t\t\tif (algo == GIT_HASH_UNKNOWN)\n+\t\t\t\t\t\tdie(_(\"unsupported object format: %s\"), arg);\n+\t\t\t\t\tflags |= GET_OID_HASH_ANY;\n+\t\t\t\t\toutput_algo = &hash_algos[algo];\n+\t\t\t\t\tcontinue;\n \t\t\t\t}\n-\t\t\t\telse die(_(\"unsupported object format: %s\"), arg);\n \t\t\t}\n \t\t\tif (opt_with_value(arg, \"--short\", &arg)) {\n \t\t\t\tfilter &= ~(DO_FLAGS|DO_NOREV);\n@@ -1162,9 +1175,11 @@ int cmd_rev_parse(int argc,\n \t\t}\n \t\tif (!repo_get_oid_with_flags(the_repository, name, &oid,\n \t\t\t\t\t     flags)) {\n-\t\t\tif (output_algo)\n-\t\t\t\trepo_oid_to_algop(the_repository, &oid,\n-\t\t\t\t\t\t  output_algo, &oid);\n+\t\t\tif (output_algo &&\n+\t\t\t    repo_oid_to_algop(the_repository, &oid,\n+\t\t\t\t\t      output_algo, &oid))\n+\t\t\t\tdie(_(\"cannot express %s as a %s object name\"),\n+\t\t\t\t    name, output_algo->name);\n \t\t\tif (verify)\n \t\t\t\trevs_count++;\n \t\t\telse\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 63ee18630b..1f0475b5a0 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -10,7 +10,120 @@\n #include \"loose.h\"\n #include \"commit.h\"\n #include \"gpg-interface.h\"\n+#include \"object-file.h\"\n #include \"object-file-convert.h\"\n+#include \"odb.h\"\n+#include \"oidmap.h\"\n+\n+struct compat_oid_cache_entry {\n+\tstruct oidmap_entry entry;\n+\tstruct object_id compat_oid;\n+};\n+\n+void repo_clear_compat_oid_cache(struct repository *repo)\n+{\n+\tif (!repo->compat_oid_cache)\n+\t\treturn;\n+\toidmap_clear(repo->compat_oid_cache, 1);\n+\tFREE_AND_NULL(repo->compat_oid_cache);\n+}\n+\n+static int lookup_computed_oid(struct repository *repo,\n+\t\t\t       const struct object_id *src,\n+\t\t\t       struct object_id *dest)\n+{\n+\tstruct compat_oid_cache_entry *found;\n+\n+\tif (!repo->compat_oid_cache)\n+\t\treturn -1;\n+\tfound = oidmap_get(repo->compat_oid_cache, src);\n+\tif (!found)\n+\t\treturn -1;\n+\toidcpy(dest, &found->compat_oid);\n+\treturn 0;\n+}\n+\n+static void remember_computed_oid(struct repository *repo,\n+\t\t\t\t  const struct object_id *src,\n+\t\t\t\t  const struct object_id *dest)\n+{\n+\tstruct compat_oid_cache_entry *added;\n+\n+\tif (!repo->compat_oid_cache) {\n+\t\tCALLOC_ARRAY(repo->compat_oid_cache, 1);\n+\t\toidmap_init(repo->compat_oid_cache, 0);\n+\t}\n+\tCALLOC_ARRAY(added, 1);\n+\toidcpy(&added->entry.oid, src);\n+\toidcpy(&added->compat_oid, dest);\n+\toidmap_put(repo->compat_oid_cache, added);\n+}\n+\n+/*\n+ * Compute the name an object has under another hash algorithm, by converting\n+ * its contents and hashing the result.  Objects the contents refer to are\n+ * resolved by recursing through repo_oid_to_algop(), so a tree costs a walk\n+ * of everything reachable from it.\n+ *\n+ * Commits are deliberately not handled.  A commit names its parents, so\n+ * converting one converts the whole history behind it; that is a repository\n+ * conversion rather than something a caller asking for a single object name\n+ * should trigger, and recursing over it here would also be unbounded.\n+ */\n+static int compute_oid_to_algop(struct repository *repo,\n+\t\t\t\tconst struct object_id *src,\n+\t\t\t\tconst struct git_hash_algo *from,\n+\t\t\t\tconst struct git_hash_algo *to,\n+\t\t\t\tstruct object_id *dest)\n+{\n+\tstruct strbuf converted = STRBUF_INIT;\n+\tenum object_type type;\n+\tsize_t size;\n+\tvoid *buf;\n+\tint ret = -1;\n+\n+\t/* We can only read objects that are stored the way the repo stores them. */\n+\tif (from != repo->hash_algo)\n+\t\treturn -1;\n+\n+\tbuf = odb_read_object(repo->objects, src, &type, &size);\n+\tif (!buf)\n+\t\treturn error(_(\"unable to read %s\"), oid_to_hex(src));\n+\n+\tswitch (type) {\n+\tcase OBJ_BLOB:\n+\t\t/*\n+\t\t * A blob's contents are the same under either algorithm, so\n+\t\t * there is nothing to convert, only to rehash.\n+\t\t */\n+\t\thash_object_file(to, buf, size, OBJ_BLOB, dest);\n+\t\tret = 0;\n+\t\tbreak;\n+\tcase OBJ_TREE:\n+\tcase OBJ_TAG:\n+\t\tif (!convert_object_file(repo, &converted, from, to, buf, size,\n+\t\t\t\t\t type, 1)) {\n+\t\t\thash_object_file(to, converted.buf, converted.len, type,\n+\t\t\t\t\t dest);\n+\t\t\tret = 0;\n+\t\t}\n+\t\tbreak;\n+\tcase OBJ_COMMIT:\n+\t\terror(_(\"cannot compute the %s name of commit %s\"),\n+\t\t      to->name, oid_to_hex(src));\n+\t\tbreak;\n+\tdefault:\n+\t\terror(_(\"unknown type for object %s\"), oid_to_hex(src));\n+\t\tbreak;\n+\t}\n+\n+\tfree(buf);\n+\tstrbuf_release(&converted);\n+\n+\tif (!ret)\n+\t\tremember_computed_oid(repo, src, dest);\n+\treturn ret;\n+}\n \n int repo_oid_to_algop(struct repository *repo, const struct object_id *srcoid,\n \t\t      const struct git_hash_algo *to, struct object_id *dest)\n@@ -43,14 +156,25 @@ int repo_oid_to_algop(struct repository *repo, const struct object_id *srcoid,\n \t\t * let's reload the map to see if the object has appeared.\n \t\t */\n \t\trepo_read_loose_object_map(repo);\n-\t\tif (repo_loose_object_map_oid(repo, src, to, dest))\n-\t\t\treturn -1;\n+\t\tif (repo_loose_object_map_oid(repo, src, to, dest)) {\n+\t\t\t/*\n+\t\t\t * The map only covers objects written while\n+\t\t\t * extensions.compatObjectFormat was in effect, so it\n+\t\t\t * cannot answer for objects a repository already had.\n+\t\t\t * Compute the name instead, remembering it so that\n+\t\t\t * trees sharing a subtree only pay for it once.\n+\t\t\t */\n+\t\t\tif (!lookup_computed_oid(repo, src, dest))\n+\t\t\t\treturn 0;\n+\t\t\treturn compute_oid_to_algop(repo, src, from, to, dest);\n+\t\t}\n \t}\n \treturn 0;\n }\n \n static int decode_tree_entry_raw(struct object_id *oid, const char **path,\n-\t\t\t\t size_t *len, const struct git_hash_algo *algo,\n+\t\t\t\t size_t *len, uint16_t *modep,\n+\t\t\t\t const struct git_hash_algo *algo,\n \t\t\t\t const char *buf, unsigned long size)\n {\n \tuint16_t mode;\n@@ -64,6 +188,7 @@ static int decode_tree_entry_raw(struct object_id *oid, const char **path,\n \tif (!*path || !**path)\n \t\treturn -1;\n \t*len = strlen(*path) + 1;\n+\t*modep = mode;\n \n \toidread(oid, (const unsigned char *)*path + *len, algo);\n \treturn 0;\n@@ -81,10 +206,20 @@ static int convert_tree_object(struct repository *repo,\n \t\tstruct object_id entry_oid, mapped_oid;\n \t\tconst char *path = NULL;\n \t\tsize_t pathlen;\n+\t\tuint16_t mode;\n \n-\t\tif (decode_tree_entry_raw(&entry_oid, &path, &pathlen, from, p,\n-\t\t\t\t\t  end - p))\n+\t\tif (decode_tree_entry_raw(&entry_oid, &path, &pathlen, &mode,\n+\t\t\t\t\t  from, p, end - p))\n \t\t\treturn error(_(\"failed to decode tree entry\"));\n+\t\t/*\n+\t\t * A gitlink names a commit in the submodule's repository,\n+\t\t * which we cannot read, so say so rather than complaining\n+\t\t * about a missing object.\n+\t\t */\n+\t\tif (S_ISGITLINK(mode))\n+\t\t\treturn error(_(\"cannot map submodule entry '%s'; convert \"\n+\t\t\t\t       \"commit %s in the submodule repository first\"),\n+\t\t\t\t     path, oid_to_hex(&entry_oid));\n \t\tif (repo_oid_to_algop(repo, &entry_oid, to, &mapped_oid))\n \t\t\treturn error(_(\"failed to map tree entry for %s\"), oid_to_hex(&entry_oid));\n \t\tstrbuf_add(out, p, path - p);\ndiff --git a/object-file-convert.h b/object-file-convert.h\nindex 9b3cc5e533..af384c362b 100644\n--- a/object-file-convert.h\n+++ b/object-file-convert.h\n@@ -10,6 +10,12 @@ struct strbuf;\n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \t\t      const struct git_hash_algo *to, struct object_id *dest);\n \n+/*\n+ * Release the object names repo_oid_to_algop() computed on demand.  Called\n+ * from repo_clear().\n+ */\n+void repo_clear_compat_oid_cache(struct repository *repo);\n+\n /*\n  * Convert an object file from one hash algorithm to another algorithm.\n  * Return -1 on failure, 0 on success.\ndiff --git a/repository.c b/repository.c\nindex 651b0f6933..16201ec5ea 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -7,6 +7,7 @@\n #include \"config.h\"\n #include \"gettext.h\"\n #include \"object.h\"\n+#include \"object-file-convert.h\"\n #include \"lockfile.h\"\n #include \"path.h\"\n #include \"read-cache-ll.h\"\n@@ -386,6 +387,8 @@ void repo_clear(struct repository *repo)\n \todb_free(repo->objects);\n \trepo->objects = NULL;\n \n+\trepo_clear_compat_oid_cache(repo);\n+\n \tparsed_object_pool_clear(repo->parsed_objects);\n \tFREE_AND_NULL(repo->parsed_objects);\n \ndiff --git a/repository.h b/repository.h\nindex 3b467a2513..1a89559cd5 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -12,6 +12,7 @@ struct index_state;\n struct lock_file;\n struct pathspec;\n struct object_database;\n+struct oidmap;\n struct submodule_cache;\n struct promisor_remote_config;\n struct remote_state;\n@@ -167,6 +168,13 @@ struct repository {\n \t/* Repository's compatibility hash algorithm. */\n \tconst struct git_hash_algo *compat_hash_algo;\n \n+\t/*\n+\t * Object names computed in another hash algorithm on demand, keyed by\n+\t * the name the object has in this repository.  Owned and populated by\n+\t * repo_oid_to_algop().\n+\t */\n+\tstruct oidmap *compat_oid_cache;\n+\n \t/* Repository-specific configuration values. */\n \tstruct repo_config_values config_values_private_;\n \ndiff --git a/t/meson.build b/t/meson.build\nindex a25f37d2f5..bc3e644709 100644\n--- a/t/meson.build\n+++ b/t/meson.build\n@@ -172,6 +172,7 @@ integration_tests = [\n   't1015-read-index-unmerged.sh',\n   't1016-compatObjectFormat.sh',\n   't1017-cat-file-remote-object-info.sh',\n+  't1018-output-object-format.sh',\n   't1020-subdirectory.sh',\n   't1022-read-tree-partial-clone.sh',\n   't1050-large.sh',\ndiff --git a/t/t1018-output-object-format.sh b/t/t1018-output-object-format.sh\nnew file mode 100755\nindex 0000000000..45ddc510dd\n--- /dev/null\n+++ b/t/t1018-output-object-format.sh\n@@ -0,0 +1,118 @@\n+#!/bin/sh\n+\n+test_description='rev-parse --output-object-format names objects in another algorithm\n+\n+The repositories here do not set extensions.compatObjectFormat, so nothing\n+is recorded in the loose object map and every name has to be computed.\n+'\n+\n+TEST_PASSES_SANITIZE_LEAK=true\n+. ./test-lib.sh\n+\n+# The object formats are pinned rather than left to the default so that\n+# these tests behave the same under GIT_TEST_DEFAULT_HASH=sha256.\n+test_expect_success 'setup' '\n+\tfor fmt in sha1 sha256\n+\tdo\n+\t\tgit init --object-format=$fmt ${fmt}repo &&\n+\t\tmkdir -p ${fmt}repo/sub &&\n+\t\ttest_write_lines hello >${fmt}repo/a.txt &&\n+\t\ttest_write_lines world >${fmt}repo/sub/b.txt &&\n+\t\ttest_write_lines exec >${fmt}repo/sub/run.sh &&\n+\t\tchmod +x ${fmt}repo/sub/run.sh &&\n+\t\ttest_write_lines dup >${fmt}repo/sub/dup1 &&\n+\t\ttest_write_lines dup >${fmt}repo/dup2 &&\n+\t\tgit -C ${fmt}repo add -A &&\n+\t\tgit -C ${fmt}repo commit -m initial || return 1\n+\tdone\n+'\n+\n+test_expect_success 'name a tree in the other algorithm' '\n+\tgit -C sha256repo rev-parse HEAD^{tree} >expect &&\n+\tgit -C sha1repo rev-parse --output-object-format=sha256 HEAD^{tree} >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'name a subtree given as <rev>:<path>' '\n+\tgit -C sha256repo rev-parse HEAD:sub >expect &&\n+\tgit -C sha1repo rev-parse --output-object-format=sha256 HEAD:sub >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'name a blob in the other algorithm' '\n+\tgit -C sha256repo rev-parse HEAD:a.txt >expect &&\n+\tgit -C sha1repo rev-parse --output-object-format=sha256 HEAD:a.txt >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'name the empty tree' '\n+\tempty1=$(git -C sha1repo hash-object -t tree /dev/null) &&\n+\tgit -C sha256repo hash-object -t tree /dev/null >expect &&\n+\tgit -C sha1repo rev-parse --output-object-format=sha256 \"$empty1\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'the conversion runs in both directions' '\n+\tgit -C sha1repo rev-parse HEAD^{tree} >expect &&\n+\tgit -C sha256repo rev-parse --output-object-format=sha1 HEAD^{tree} >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success 'asking for the storage algorithm is a no-op' '\n+\tgit -C sha1repo rev-parse HEAD^{tree} >expect &&\n+\tgit -C sha1repo rev-parse --output-object-format=storage HEAD^{tree} >actual &&\n+\ttest_cmp expect actual &&\n+\tgit -C sha1repo rev-parse --output-object-format=sha1 HEAD^{tree} >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+# Before the names could be computed, a lookup miss left the object ID\n+# untouched and rev-parse printed the storage name with a zero exit code.\n+test_expect_success 'a name that cannot be computed is an error, not a wrong answer' '\n+\ttest_must_fail git -C sha1repo rev-parse --output-object-format=sha256 \\\n+\t\tHEAD >actual 2>err &&\n+\ttest_grep \"cannot compute the sha256 name of commit\" err &&\n+\ttest_grep \"cannot express HEAD as a sha256 object name\" err &&\n+\ttest_must_be_empty actual\n+'\n+\n+test_expect_success 'a tree containing a submodule reports the submodule' '\n+\tgit init --object-format=sha1 withsub &&\n+\tgitlink=$(git -C sha1repo rev-parse HEAD) &&\n+\t(\n+\t\tcd withsub &&\n+\t\ttest_write_lines x >f &&\n+\t\tgit add f &&\n+\t\tgit update-index --add --cacheinfo 160000,$gitlink,modpath &&\n+\t\tgit commit -m withsub\n+\t) &&\n+\ttest_must_fail git -C withsub rev-parse --output-object-format=sha256 \\\n+\t\tHEAD^{tree} 2>err &&\n+\ttest_grep \"cannot map submodule entry .modpath.\" err\n+'\n+\n+test_expect_success 'an unknown algorithm is rejected' '\n+\ttest_must_fail git -C sha1repo rev-parse --output-object-format=md5 \\\n+\t\tHEAD^{tree} 2>err &&\n+\ttest_grep \"unsupported object format: md5\" err\n+'\n+\n+test_expect_success 'a missing algorithm is rejected' '\n+\ttest_must_fail git -C sha1repo rev-parse --output-object-format \\\n+\t\tHEAD^{tree} 2>err &&\n+\ttest_grep \"no object format specified\" err &&\n+\ttest_must_fail git -C sha1repo rev-parse --output-object-format= \\\n+\t\tHEAD^{tree} 2>err &&\n+\ttest_grep \"unsupported object format:\" err\n+'\n+\n+test_expect_success 'names computed here match a real conversion' '\n+\tgit -C sha1repo fast-export --full-tree HEAD >stream &&\n+\tgit init --bare --object-format=sha256 converted.git &&\n+\tgit -C converted.git fast-import --quiet <stream &&\n+\tgit -C converted.git rev-parse HEAD^{tree} >expect &&\n+\tgit -C sha1repo rev-parse --output-object-format=sha256 HEAD^{tree} >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+test_done\n-- \n2.53.0\n\n"}]}